EDBT 2026 Demo / reviewers in the wild / expert
Quanshi Zhang
dblp:46/7793
· DBLP profile ↗
80ranked-venue papers
25as first author
47since 2021 · last 2026
0000-0002-6108-2738ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 22 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 12 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorSystems, architecture and hardware · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Attribution Explanations for Deep Neural Networks: A Theoretical PerspectiveabstractAttribution explanation is a typical approach for interpreting deep neural networks (DNNs), aiming to quantify the contribution score of individual input variables to model predictions. Despite extensive methodological development, a fundamental faithfulness problem remains unresolved: whether existing attribution methods faithfully reflect the true decision-making logic of DNNs, which significantly limits their reliability and practical adoption. These concerns largely stem from three core challenges: the lack of a unified theoretical framework, clear theoretical rationales, and principled faithfulness evaluation in the absence of ground truth. Recently, a growing body of theoretical studies has begun to address these issues, marking an important shift toward principled understanding of attribution methods. In this survey, we provide a comprehensive review of these advances, with a particular emphasis on three interconnected directions: (i) Theoretical unification, which uncovers key commonalities and differences among attribution methods; (ii) Theoretical rationale, which clarifies the mathematical and conceptual justifications underlying existing methods; (iii) Theoretical evaluation, which rigorously proves whether attribution methods satisfy established faithfulness principles. Beyond a comprehensive review, we provide practical recommendations and a case study illustrating how theoretical findings can be translated into operational decision rules for method design, selection, and usage. We conclude with a discussion of promising open problems for further work. Huiqi Deng, Hongbin Pei, Quanshi Zhang, Mengnan Du |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Monitoring Primitive Interactions During the Training of DNNsabstractThis paper focuses on the newly emerged research topic, i.e., whether the complex decision-making logic of a DNN can be mathematically summarized into a few simple logics. Beyond the explanation of a static DNN, in this paper, we hope to show that the seemingly complex learning dynamics of a DNN can be faithfully represented as the change of a few primitive interaction patterns encoded by the DNN. Therefore, we redefine the interaction of principal feature components in intermediate-layer features, which enables us to concisely summarize the highly complex dynamics of interactions throughout the learning of the DNN. The mathematical faithfulness of the new interaction is experimentally verified. From the perspective of learning efficiency, we find that the interactions naturally belong to five groups (reliable, withdrawn, forgotten, betraying, and fluctuating interactions), each representing a distinct type of dynamics of an interaction being learned and/or being forgotten. This provides deep insights into the learning process of a DNN. Jie Ren 0018, Xinhao Zheng, Jiyu Liu, Andrew Lizarraga, Ying Nian Wu, Liang Lin 0004, Quanshi Zhang |
AAAI | 7 |
| 2025 | Towards Attributions of Input Variables in a CoalitionabstractThis paper focuses on the fundamental challenge of partitioning input variables in attribution methods for Explainable AI, particularly in Shapley value-based approaches. Previous methods always compute attributions given a predefined partition but lack theoretical guidance on how to form meaningful variable partitions. We identify that attribution conflicts arise when the attribution of a coalition differs from the sum of its individual variables' attributions. To address this, we analyze the numerical effects of AND-OR interactions in AI models and extend the Shapley value to a new attribution metric for variable coalitions. Our theoretical findings reveal that specific interactions cause attribution conflicts, and we propose three metrics to evaluate coalition faithfulness. Experiments on synthetic data, NLP, image classification, and the game of Go validate our approach, demonstrating consistency with human intuition and practical applicability. Xinhao Zheng, Huiqi Deng, Quanshi Zhang |
ICML | 3 |
| 2025 | Towards the Resistance of Neural Network Fingerprinting to Fine-tuningabstractThis paper proves a new fingerprinting method to embed the ownership information into a deep neural network (DNN) with theoretically guaranteed robustness to fine-tuning. Specifically, we prove that when the input feature of a convolutional layer only contains low-frequency components, specific frequency components of the convolutional filter will not be changed by gradient descent during the fine-tuning process, where we propose a revised Fourier transform to extract frequency components from the convolutional filter. Additionally, we also prove that these frequency components are equivariant to weight scaling and weight permutations. In this way, we design a fingerprint module to embed the fingerprint information into specific frequency components of convolutional filters. Preliminary experiments demonstrate the effectiveness of our method. Ling Tang 0002, Yuefeng Chen, Hui Xue 0001, Quanshi Zhang |
NeurIPS | 4 |
| 2025 | Towards the first principles of explaining DNNs: interactions explain the learning dynamicsabstractMost explanation methods are designed in an empirical manner, so exploring whether there exists a first-principles explanation of a deep neural network (DNN) becomes the next core scientific problem in explainable artificial intelligence (XAI). Although it is still an open problem, in this paper, we discuss whether the interaction-based explanation can serve as the first-principles explanation of a DNN. The strong explanatory power of interaction theory comes from the following aspects: (1) it establishes a new axiomatic system to quantify the decision-making logic of a DNN into a set of symbolic interaction concepts; (2) it simultaneously explains various deep learning phenomena, such as generalization power, adversarial sensitivity, representation bottleneck, and learning dynamics; (3) it provides mathematical tools that uniformly explain the mechanisms of various empirical attribution methods and empirical adversarial-transferability-boosting methods; (4) it explains the extremely complex learning dynamics of a DNN by analyzing the two-phase dynamics of interaction complexity, which further reveals the internal mechanism of why and how the generalization power/adversarial sensitivity of a DNN changes during the learning process. Huilin Zhou, Qihan Ren, Quanshi Zhang |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2025 | Interpretable Rotation-Equivariant Multiary-Valued Network for Attribute ObfuscationabstractThis paper focuses on the problem of preventing information leakage in neural networks, i.e., assuming that attackers have obtained intermediate-layer features of a neural network, and preventing attackers from inverting these features to the input with private information. We propose a generic method to slightly revise each arbitrary traditional neural network into a multiary-valued rotation-equivariant neural network (RENN) for preventing information leakage. Specifically, we convert real-valued features in the network into multi-ary features, and each element in the feature vector is a multi-ary number. We hide the input information into a certain phase of the multi-ary feature, and rotate the multi-ary feature for attribute obfuscation in the encryption process. The rotation axis and angle can be considered as the private key. In this way, even when attackers have obtained network parameters and intermediate-layer features, they still cannot extract input information without knowing the rotation information. More crucially, the encryption operation does not damage the spatial correlations between features, so that the encrypted features can be easily processed by convolution operations in the neural network without difficulties. In order to implement successful encryption and decryption, the RENN is designed to satisfy the rotation equivariance property. To this end, we propose a set of rules to revise classic operations in the neural network to ensure the rotation equivariance property. Besides, we prove that the $d$d-ary RENN is downward compatible with the $d^{\prime }$d'-ary RENN when $d^{\prime }< d$d' Quanshi Zhang, Hao Zhang 0063, Yiting Chen 0003, Qihan Ren, Jie Ren 0018, Xu Cheng 0005, Liyao Xiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Clarifying the Behavior and the Difficulty of Adversarial TrainingabstractAdversarial training is usually difficult to optimize. This paper provides conceptual and analytic insights into the difficulty of adversarial training via a simple theoretical study, where we derive an approximate dynamics of a recursive multi-step attack in a simple setting. Despite the simplicity of our theory, it still reveals verifiable predictions about various phenomena in adversarial training under real-world settings. First, compared to vanilla training, adversarial training is more likely to boost the influence of input samples with large gradient norms in an exponential manner. Besides, adversarial training also strengthens the influence of the Hessian matrix of the loss w.r.t. network parameters, which is more likely to make network parameters oscillate and boosts the difficulty of adversarial training. Xu Cheng 0005, Hao Zhang 0063, Wen Shen 0002, Quanshi Zhang |
AAAI | 5 |
| 2024 | Batch Normalization Is Blind to the First and Second Derivatives of the LossabstractWe prove that when we do the Taylor series expansion of the loss function, the BN operation will block the influence of the first-order term and most influence of the second-order term of the loss. We also find that such a problem is caused by the standardization phase of the BN operation. We believe that proving the blocking of certain loss terms provides an analytic perspective for potential detects of a deep model with BN operations, although the blocking problem is not fully equivalent to significant damages in all tasks on benchmark datasets. Experiments show that the BN operation significantly affects feature representations in specific tasks. Zhanpeng Zhou, Wen Shen 0002, Huixin Chen, Ling Tang 0002, Yuefeng Chen, Quanshi Zhang |
AAAI | 6 |
| 2024 | Explaining Generalization Power of a DNN Using Interactive ConceptsabstractThis paper explains the generalization power of a deep neural network (DNN) from the perspective of interactions. Although there is no universally accepted definition of the concepts encoded by a DNN, the sparsity of interactions in a DNN has been proved, i.e., the output score of a DNN can be well explained by a small number of interactions between input variables. In this way, to some extent, we can consider such interactions as interactive concepts encoded by the DNN. Therefore, in this paper, we derive an analytic explanation of inconsistency of concepts of different complexities. This may shed new lights on using the generalization power of concepts to explain the generalization power of the entire DNN. Besides, we discover that the DNN with stronger generalization power usually learns simple concepts more quickly and encodes fewer complex concepts. We also discover the detouring dynamics of learning complex concepts, which explains both the high learning difficulty and the low generalization power of complex concepts. The code will be released when the paper is accepted. Huilin Zhou, Hao Zhang 0063, Huiqi Deng, Dongrui Liu, Wen Shen 0002, Shih-Han Chan, Quanshi Zhang |
AAAI | 7 |
| 2024 | Defining and extracting generalizable interaction primitives from DNNsabstractFaithfully summarizing the knowledge encoded by a deep neural network (DNN) into a few symbolic primitive patterns without losing much information represents a core challenge in explainable AI. To this end, Ren et al. (2024) have derived a series of theorems to prove that the inference score of a DNN can be explained as a small set of interactions between input variables. However, the lack of generalization power makes it still hard to consider such interactions as faithful primitive patterns encoded by the DNN. Therefore, given different DNNs trained for the same task, we develop a new method to extract interactions that are shared by these DNNs. Experiments show that the extracted interactions can better reflect common knowledge shared by different DNNs. Siyu Lou, Benhao Huang, Quanshi Zhang |
ICLR | 4 |
| 2024 | Where We Have Arrived in Proving the Emergence of Sparse Interaction Primitives in DNNsabstractThis study aims to prove the emergence of symbolic concepts (or more precisely, sparse primitive inference patterns) in well-trained deep neural networks (DNNs). Specifically, we prove the following three conditions for the emergence. (i) The high-order derivatives of the network output with respect to the input variables are all zero. (ii) The DNN can be used on occluded samples, and when the input sample is less occluded, the DNN will yield higher confidence. (iii) The confidence of the DNN does not significantly degrade on occluded samples. These conditions are quite common, and we prove that under these conditions, the DNN will only encode a relatively small number of sparse interactions between input variables. Moreover, we can consider such interactions as symbolic primitive inference patterns encoded by a DNN, because we show that inference scores of the DNN on an exponentially large number of randomly masked samples can always be well mimicked by numerical effects of just a few interactions. Qihan Ren, Jiayang Gao, Wen Shen 0002, Quanshi Zhang |
ICLR | 4 |
| 2024 | Layerwise Change of Knowledge in Neural NetworksabstractThis paper aims to explain how a deep neural network (DNN) gradually extracts new knowledge and forgets noisy features through layers in forward propagation. Up to now, although how to define knowledge encoded by the DNN has not reached a consensus so far, previous studies have derived a series of mathematical evidences to take interactions as symbolic primitive inference patterns encoded by a DNN. We extend the definition of interactions and, for the first time, extract interactions encoded by intermediate layers. We quantify and track the newly emerged interactions and the forgotten interactions in each layer during the forward propagation, which shed new light on the learning behavior of DNNs. The layer-wise change of interactions also reveals the change of the generalization capacity and instability of feature representations of a DNN. Xu Cheng 0005, Lei Cheng 0006, Zhaoran Peng, Tian Han 0001, Quanshi Zhang |
ICML | 6 |
| 2024 | Towards the Dynamics of a DNN Learning Symbolic InteractionsabstractThis study proves the two-phase dynamics of a deep neural network (DNN) learning interactions. Despite the long disappointing view of the faithfulness of post-hoc explanation of a DNN, a series of theorems have been proven [27] in recent years to show that for a given input sample, a small set of interactions between input variables can be considered as primitive inference patterns that faithfully represent a DNN's detailed inference logic on that sample. Particularly, Zhang et al. [41] have observed that various DNNs all learn interactions of different complexities in two distinct phases, and this two-phase dynamics well explains how a DNN changes from under-fitting to over-fitting. Therefore, in this study, we mathematically prove the two-phase dynamics of interactions, providing a theoretical mechanism for how the generalization power of a DNN changes during the training process. Experiments show that our theory well predicts the real dynamics of interactions on different DNNs trained for various tasks. Qihan Ren, Dongrui Liu, Quanshi Zhang |
NeurIPS | 6 |
| 2024 | Unifying Fourteen Post-Hoc Attribution Methods With Taylor InteractionsabstractVarious attribution methods have been developed to explain deep neural networks (DNNs) by inferring the attribution/importance/contribution score of each input variable to the final output. However, existing attribution methods are often built upon different heuristics. There remains a lack of a unified theoretical understanding of why these methods are effective and how they are related. Furthermore, there is still no universally accepted criterion to compare whether one attribution method is preferable over another. In this paper, we resort to Taylor interactions and for the first time, we discover that fourteen existing attribution methods, which define attributions based on fully different heuristics, actually share the same core mechanism. Specifically, we prove that attribution scores of input variables estimated by the fourteen attribution methods can all be mathematically reformulated as a weighted allocation of two typical types of effects, i.e., independent effects of each input variable and interaction effects between input variables. The essential difference among these attribution methods lies in the weights of allocating different effects. Inspired by these insights, we propose three principles for fairly allocating the effects, which serve as new criteria to evaluate the faithfulness of attribution methods. In summary, this study can be considered as a new unified perspective to revisit fourteen attribution methods, which theoretically clarifies essential similarities and differences among these methods. Besides, the proposed new principles enable people to make a direct and fair comparison among different methods under the unified perspective. Huiqi Deng, Na Zou 0001, Mengnan Du, Weifu Chen, Guo-Can Feng, Zheyang Li, Quanshi Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Interpretable Rotation-Equivariant Quaternion Neural Networks for 3D Point Cloud ProcessingabstractThis study proposes a set of generic rules to revise existing neural networks for 3D point cloud processing to rotation-equivariant quaternion neural networks (REQNNs), in order to make feature representations of neural networks to be rotation-equivariant and permutation-invariant. Rotation equivariance of features means that the feature computed on a rotated input point cloud is the same as applying the same rotation transformation to the feature computed on the original input point cloud. We find that the rotation-equivariance of features is naturally satisfied, if a neural network uses quaternion features. Interestingly, we prove that such a network revision also makes gradients of features in the REQNN to be rotation-equivariant w.r.t. inputs, and the training of the REQNN to be rotation-invariant w.r.t. inputs. Besides, permutation-invariance examines whether the intermediate-layer features are invariant, when we reorder input points. We also evaluate the stability of knowledge representations of REQNNs, and the robustness of REQNNs to adversarial rotation attacks. Experiments have shown that REQNNs outperform traditional neural networks in both terms of classification accuracy and robustness on rotated testing samples. Wen Shen 0002, Zhihua Wei 0001, Qihan Ren, Shikun Huang, Quanshi Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | An Introspective Data Augmentation Method for Training Math Word Problem SolversabstractThough quite challenging, training a deep neural network for automatically solving Math Word Problems (MWPs) has increasingly attracted attention due to its significance in investigating how a machine can understand and reason complex problems like a human. However, the data volume of existing high-quality MWP datasets is far from sufficient to train a robust solver since collecting these datasets would cost a very high price, i.e., they require professional knowledge that accords with the educational standard and massive accessible data. This data bottleneck inspires us to consider using cost-effective data augmentation methods to improve the utilization of the existing data and enhance the performance of an MWP solver. Nevertheless, the traditional input-based data augmentation methods for training natural image or language models are incompetent for training MWP solvers due to the following two reasons. First, MWPs are concise yet comprehensive, so these data augmentation methods are prone to make them more ambiguous. Second, the mathematical dependencies grounded in the problems must be maintained when a batch of augmented examples is generated during the data augmentation process. To address these issues, we propose a simple yet effective data augmentation method called the Introspective Data Augmentation Method (IDAM) that allows the MWP examples to be augmented in latent space during the training of the neural network, instead of making perturbations over the input data. In particular, our IDAM is capable of applying different data augmentation operations on the latent feature representations of MWPs to produce new examples. Moreover, a new training objective is developed to constrain the mathematical dependency consistency between the original MWP and the produced ones. Extensive experiments conducted on standard benchmarks demonstrate the effectiveness of IDAM in generally improving the performance of existing MWP solvers without any elaborated model crafting. Jinghui Qin, Zhongzhan Huang, Quanshi Zhang, Liang Lin 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Learning to Prevent Input Leakages in the Mobile Cloud InferenceabstractPowered by machine learning services in the cloud, numerous learning-driven mobile applications are gaining popularity in the market. As deep learning tasks are mostly computation-intensive, it has become a trend to process raw data on devices and send the deep neural network (DNN) features to the cloud, where the features are further processed to return final results. However, there is always an unexpected leakage with the release of features, by which an adversary could infer much information on the original data. We propose a privacy-preserving framework on top of the mobile cloud infrastructure from the perspective of DNN structures. Our framework aims to learn a policy to modify the base DNNs to prevent information leakage while maintaining high inference accuracy. The policy can also be readily transferred to large-size DNNs and large-scale datasets to speed up learning. Extensive evaluations on a variety of DNNs have shown that our framework successfully finds privacy-preserving DNN structures to defend privacy attacks. Liyao Xiang, Shuang Zhang 0007, Quanshi Zhang |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Defining and Quantifying the Emergence of Sparse Concepts in DNNsabstractThis paper aims to illustrate the concept-emerging phenomenon in a trained DNN. Specifically, we find that the inference score of a DNN can be disentangled into the effects of a few interactive concepts. These concepts can be understood as causal patterns in a sparse, symbolic causal graph, which explains the DNN. The faithfulness of using such a causal graph to explain the DNN is theoretically guaranteed, because we prove that the causal graph can well mimic the DNN's outputs on an exponential number of different masked samples. Besides, such a causal graph can be further simplified and re-written as an And-Or graph (AOG), without losing much explanation accuracy. The code is released at https://github.com/sjtu-xai-lab/aog. Jie Ren 0018, Qirui Chen, Huiqi Deng, Quanshi Zhang |
CVPR | 5 |
| 2023 | Can We Faithfully Represent Absence States to Compute Shapley Values on a DNN?
Jie Ren 0018, Zhanpeng Zhou, Qirui Chen, Quanshi Zhang |
ICLR | 4 |
| 2023 | HarsanyiNet: Computing Accurate Shapley Values in a Single Forward PropagationabstractThe Shapley value is widely regarded as a trustworthy attribution metric. However, when people use Shapley values to explain the attribution of input variables of a deep neural network (DNN), it usually requires a very high computational cost to approximate relatively accurate Shapley values in real-world applications. Therefore, we propose a novel network architecture, the HarsanyiNet, which makes inferences on the input sample and simultaneously computes the exact Shapley values of the input variables in a single forward propagation. The HarsanyiNet is designed on the theoretical foundation that the Shapley value can be reformulated as the redistribution of Harsanyi interactions encoded by the network. Siyu Lou, Keyan Zhang, Quanshi Zhang |
ICML | 5 |
| 2023 | Does a Neural Network Really Encode Symbolic Concepts?abstractRecently, a series of studies have tried to extract interactions between input variables modeled by a DNN and define such interactions as concepts encoded by the DNN. However, strictly speaking, there still lacks a solid guarantee whether such interactions indeed represent meaningful concepts. Therefore, in this paper, we examine the trustworthiness of interaction concepts from four perspectives. Extensive empirical studies have verified that a well-trained DNN usually encodes sparse, transferable, and discriminative concepts, which is partially aligned with human intuition. The code is released at https://github.com/sjtu-xai-lab/interaction-concept. Quanshi Zhang |
ICML | 2 |
| 2023 | Bayesian Neural Networks Avoid Encoding Complex and Perturbation-Sensitive ConceptsabstractIn this paper, we focus on mean-field variational Bayesian Neural Networks (BNNs) and explore the representation capacity of such BNNs by investigating which types of concepts are less likely to be encoded by the BNN. It has been observed and studied that a relatively small set of interactive concepts usually emerge in the knowledge representation of a sufficiently-trained neural network, and such concepts can faithfully explain the network output. Based on this, our study proves that compared to standard deep neural networks (DNNs), it is less likely for BNNs to encode complex concepts. Experiments verify our theoretical proofs. Note that the tendency to encode less complex concepts does not necessarily imply weak representation power, considering that complex concepts exhibit low generalization power and high adversarial vulnerability. The code is available at https://github.com/sjtu-xai-lab/BNN-concepts. Qihan Ren, Huiqi Deng, Yunuo Chen 0002, Siyu Lou, Quanshi Zhang |
ICML | 5 |
| 2023 | Defects of Convolutional Decoder Networks in Frequency RepresentationabstractIn this paper, we prove the representation defects of a cascaded convolutional decoder network, considering the capacity of representing different frequency components of an input sample. We conduct the discrete Fourier transform on each channel of the feature map in an intermediate layer of the decoder network. Then, we extend the 2D circular convolution theorem to represent the forward and backward propagations through convolutional layers in the frequency domain. Based on this, we prove three defects in representing feature spectrums. First, we prove that the convolution operation, the zero-padding operation, and a set of other settings all make a convolutional decoder network more likely to weaken high-frequency components. Second, we prove that the upsampling operation generates a feature spectrum, in which strong signals repetitively appear at certain frequencies. Third, we prove that if the frequency components in the input sample and frequency components in the target output for regression have a small shift, then the decoder usually cannot be effectively learned. Ling Tang 0002, Wen Shen 0002, Zhanpeng Zhou, Yuefeng Chen, Quanshi Zhang |
ICML | 5 |
| 2023 | Towards the Difficulty for a Deep Neural Network to Learn Concepts of Different ComplexitiesabstractThis paper theoretically explains the intuition that simple concepts are more likely to be learned by deep neural networks (DNNs) than complex concepts. In fact, recent studies have observed [24, 15] and proved [26] the emergence of interactive concepts in a DNN, i.e., it is proven that a DNN usually only encodes a small number of interactive concepts, and can be considered to use their interaction effects to compute inference scores. Each interactive concept is encoded by the DNN to represent the collaboration between a set of input variables. Therefore, in this study, we aim to theoretically explain that interactive concepts involving more input variables (i.e., more complex concepts) are more difficult to learn. Our finding clarifies the exact conceptual complexity that boosts the learning difficulty. Dongrui Liu, Huiqi Deng, Xu Cheng 0005, Qihan Ren, Kangrui Wang, Quanshi Zhang |
NeurIPS | 6 |
| 2023 | Network Transplanting for the Functionally Modular Architecture
Quanshi Zhang, Xu Cheng 0005, Xin Wang 0108, Yu Yang 0007, Ying Nian Wu |
PRCV (3) | 1 |
| 2023 | Quantifying the Knowledge in a DNN to Explain Knowledge Distillation for ClassificationabstractCompared to traditional learning from scratch, knowledge distillation sometimes makes the DNN achieve superior performance. In this paper, we provide a new perspective to explain the success of knowledge distillation based on the information theory, i.e., quantifying knowledge points encoded in intermediate layers of a DNN for classification. To this end, we consider the signal processing in a DNN as a layer-wise process of discarding information. A knowledge point is referred to as an input unit, the information of which is discarded much less than that of other input units. Thus, we propose three hypotheses for knowledge distillation based on the quantification of knowledge points. 1. The DNN learning from knowledge distillation encodes more knowledge points than the DNN learning from scratch. 2. Knowledge distillation makes the DNN more likely to learn different knowledge points simultaneously. In comparison, the DNN learning from scratch tends to encode various knowledge points sequentially. 3. The DNN learning from knowledge distillation is often more stably optimized than the DNN learning from scratch. To verify the above hypotheses, we design three types of metrics with annotations of foreground objects to analyze feature representations of the DNN, i.e., the quantity and the quality of knowledge points, the learning speed of different knowledge points, and the stability of optimization directions. In experiments, we diagnosed various DNNs on different classification tasks, including image classification, 3D point cloud classification, binary sentiment classification, and question answering, which verified the above hypotheses. Quanshi Zhang, Xu Cheng 0005, Yilan Chen 0002, Zhefan Rao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Interpretable Generative Adversarial NetworksabstractLearning a disentangled representation is still a challenge in the field of the interpretability of generative adversarial networks (GANs). This paper proposes a generic method to modify a traditional GAN into an interpretable GAN, which ensures that filters in an intermediate layer of the generator encode disentangled localized visual concepts. Each filter in the layer is supposed to consistently generate image regions corresponding to the same visual concept when generating different images. The interpretable GAN learns to automatically discover meaningful visual concepts without any annotations of visual concepts. The interpretable GAN enables people to modify a specific visual concept on generated images by manipulating feature maps of the corresponding filters in the layer. Our method can be broadly applied to different types of GANs. Experiments have demonstrated the effectiveness of our method. Chao Li 0028, Kelu Yao, Jin Wang 0039, Boyu Diao, Yongjun Xu 0001, Quanshi Zhang |
AAAI | 6 |
| 2022 | Exploring Image Regions Not Well Encoded by an INNabstractThis paper proposes a method to clarify image regions that are not well encoded by an invertible neural network (INN), i.e., image regions that significantly decrease the likelihood of the input image. The proposed method can diagnose the limitation of the representation capacity of an INN. Given an input image, our method extracts image regions, which are not well encoded, by maximizing the likelihood of the image. We explicitly model the distribution of not-well-encoded regions. A metric is proposed to evaluate the extraction of the not-well-encoded regions. Finally, we use the proposed method to analyze several state-of-the-art INNs trained on various benchmark datasets. Zenan Ling, Quanshi Zhang |
AISTATS | 4 |
| 2022 | RASAT: Integrating Relational Structures into Pretrained Seq2Seq Model for Text-to-SQLabstractJiexing Qi, Jingyao Tang, Ziwei He, Xiangpeng Wan, Yu Cheng, Chenghu Zhou, Xinbing Wang, Quanshi Zhang, Zhouhan Lin. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jiexing Qi, Ziwei He, Xiangpeng Wan, Yu Cheng 0003, Chenghu Zhou, Xinbing Wang, Quanshi Zhang, Zhouhan Lin |
EMNLP | 8 |
| 2022 | Discovering and Explaining the Representation Bottleneck of DNNS
Huiqi Deng, Qihan Ren, Hao Zhang 0063, Quanshi Zhang |
ICLR | 4 |
| 2022 | Towards Theoretical Analysis of Transformation Complexity of ReLU DNNsabstractThis paper aims to theoretically analyze the complexity of feature transformations encoded in piecewise linear DNNs with ReLU layers. We propose metrics to measure three types of complexities of transformations based on the information theory. We further discover and prove the strong correlation between the complexity and the disentanglement of transformations. Based on the proposed metrics, we analyze two typical phenomena of the change of the transformation complexity during the training process, and explore the ceiling of a DNN’s complexity. The proposed metrics can also be used as a loss to learn a DNN with the minimum complexity, which also controls the over-fitting level of the DNN and influences adversarial robustness, adversarial transferability, and knowledge consistency. Comprehensive comparative studies have provided new perspectives to understand the DNN. The code is released at https://github.com/sjtu-XAI-lab/transformation-complexity. Jie Ren 0018, Shih-Han Chan, Quanshi Zhang |
ICML | 5 |
| 2022 | Quantification and Analysis of Layer-wise and Pixel-wise Information DiscardingabstractThis paper presents a method to explain how the information of each input variable is gradually discarded during the forward propagation in a deep neural network (DNN), which provides new perspectives to explain DNNs. We define two types of entropy-based metrics, i.e. (1) the discarding of pixel-wise information used in the forward propagation, and (2) the uncertainty of the input reconstruction, to measure input information contained by a specific layer from two perspectives. Unlike previous attribution metrics, the proposed metrics ensure the fairness of comparisons between different layers of different DNNs. We can use these metrics to analyze the efficiency of information processing in DNNs, which exhibits strong connections to the performance of DNNs. We analyze information discarding in a pixel-wise manner, which is different from the information bottleneck theory measuring feature information w.r.t. the sample distribution. Experiments have shown the effectiveness of our metrics in analyzing classic DNNs and explaining existing deep-learning techniques. The code is available at https://github.com/haotianSustc/deepinfo. Hao Zhang 0063, Yinqing Zhang, Quanshi Zhang |
ICML | 5 |
| 2022 | Rapid detection and recognition of whole brain activity in a freely behaving Caenorhabditis elegansabstractAdvanced volumetric imaging methods and genetically encoded activity indicators have permitted a comprehensive characterization of whole brain activity at single neuron resolution in Caenorhabditis elegans. The constant motion and deformation of the nematode nervous system, however, impose a great challenge for consistent identification of densely packed neurons in a behaving animal. Here, we propose a cascade solution for long-term and rapid recognition of head ganglion neurons in a freely moving C. elegans. First, potential neuronal regions from a stack of fluorescence images are detected by a deep learning algorithm. Second, 2-dimensional neuronal regions are fused into 3-dimensional neuron entities. Third, by exploiting the neuronal density distribution surrounding a neuron and relative positional information between neurons, a multi-class artificial neural network transforms engineered neuronal feature vectors into digital neuronal identities. With a small number of training samples, our bottom-up approach is able to process each volume-1024 × 1024 × 18 in voxels-in less than 1 second and achieves an accuracy of 91% in neuronal detection and above 80% in neuronal tracking over a long video recording. Our work represents a step towards rapid and fully automated algorithms for decoding whole brain activity underlying naturalistic behaviors. Yuxiang Wu, Xin Wang 0108, Chengtian Lang, Quanshi Zhang |
PLoS Comput. Biol. | 5 |
| 2021 | Interpreting Multivariate Shapley Interactions in DNNsabstractThis paper aims to explain deep neural networks (DNNs) from the perspective of multivariate interactions. In this paper, we define and quantify the significance of interactions among multiple input variables of the DNN. Input variables with strong interactions usually form a coalition and reflect prototype features, which are memorized and used by the DNN for inference. We define the significance of interactions based on the Shapley value, which is designed to assign the attribution value of each input variable to the inference. We have conducted experiments with various DNNs. Experimental results have demonstrated the effectiveness of the proposed method. Hao Zhang 0063, Yichen Xie 0002, Longjie Zheng, Die Zhang, Quanshi Zhang |
AAAI | 5 |
| 2021 | Building Interpretable Interaction Trees for Deep NLP ModelsabstractThis paper proposes a method to disentangle and quantify interactions among words that are encoded inside a DNN for natural language processing. We construct a tree to encode salient interactions extracted by the DNN. Six metrics are proposed to analyze properties of interactions between constituents in a sentence. The interaction is defined based on Shapley values of words, which are considered as an unbiased estimation of word contributions to the network prediction. Our method is used to quantify word interactions encoded inside the BERT, ELMo, LSTM, CNN, and Transformer networks. Experimental results have provided a new perspective to understand these DNNs, and have demonstrated the effectiveness of our method. Die Zhang, Hao Zhang 0063, Huilin Zhou, Xiaoyi Bao, Da Huo 0002, Ruizhao Chen, Xu Cheng 0005, Mengyue Wu, Quanshi Zhang |
AAAI | 9 |
| 2021 | Verifiability and Predictability: Interpreting Utilities of Network Architectures for Point Cloud ProcessingabstractIn this paper, we diagnose deep neural networks for 3D point cloud processing to explore utilities of different intermediate-layer network architectures. We propose a number of hypotheses on the effects of specific intermediate-layer network architectures on the representation capacity of DNNs. In order to prove the hypotheses, we design five metrics to diagnose various types of DNNs from the following perspectives, information discarding, information concentration, rotation robustness, adversarial robustness, and neighborhood inconsistency. We conduct comparative studies based on such metrics to verify the hypotheses. We further use the verified hypotheses to revise intermediate-layer architectures of existing DNNs and improve their utilities. Experiments demonstrate the effectiveness of our method. The code will be released when this paper is accepted. Wen Shen 0002, Zhihua Wei 0001, Shikun Huang, Panyue Chen, Quanshi Zhang |
CVPR | 7 |
| 2021 | Interpreting Attributions and Interactions of Adversarial AttacksabstractThis paper aims to explain adversarial attacks in terms of how adversarial perturbations contribute to the attacking task. We estimate attributions of different image regions to the decrease of the attacking cost based on the Shapley value. We define and quantify interactions among adversarial perturbation pixels, and decompose the entire perturbation map into relatively independent perturbation components. The decomposition of the perturbation map shows that adversarially-trained DNNs have more perturbation components in the foreground than normally-trained DNNs. Moreover, compared to the normally-trained DNN, the adversarially-trained DNN have more components which mainly decrease the score of the true category. Above analyses provide new insights into the understanding of adversarial attacks. Xin Wang 0108, Shuyun Lin, Hao Zhang 0063, Quanshi Zhang |
ICCV | 5 |
| 2021 | A Unified Approach to Interpreting and Boosting Adversarial Transferability
Xin Wang 0108, Jie Ren 0018, Shuyun Lin, Xiangming Zhu 0002, Yisen Wang 0001, Quanshi Zhang |
ICLR | 6 |
| 2021 | Interpreting and Boosting Dropout from a Game-Theoretic View
Hao Zhang 0063, Yinchao Ma, Yichen Xie 0002, Quanshi Zhang |
ICLR | 6 |
| 2021 | Interpreting and Disentangling Feature Components of Various Complexity from DNNsabstractThis paper aims to define, visualize, and analyze the feature complexity that is learned by a DNN. We propose a generic definition for the feature complexity. Given the feature of a certain layer in the DNN, our method decomposes and visualizes feature components of different complexity orders from the feature. The feature decomposition enables us to evaluate the reliability, the effectiveness, and the significance of over-fitting of these feature components. Furthermore, such analysis helps to improve the performance of DNNs. As a generic method, the feature complexity also provides new insights into existing deep-learning techniques, such as network compression and knowledge distillation. Jie Ren 0018, Zexu Liu, Quanshi Zhang |
ICML | 4 |
| 2021 | Interpretable Compositional Convolutional Neural NetworksabstractThis paper proposes a method to modify a traditional convolutional neural network (CNN) into an interpretable compositional CNN, in order to learn filters that encode meaningful visual patterns in intermediate convolutional layers. In a compositional CNN, each filter is supposed to consistently represent a specific compositional object part or image region with a clear meaning. The compositional CNN learns from image labels for classification without any annotations of parts or regions for supervision. Our method can be broadly applied to different types of CNNs. Experiments have demonstrated the effectiveness of our method. The code will be released when the paper is accepted. Wen Shen 0002, Zhihua Wei 0001, Shikun Huang, Quanshi Zhang |
IJCAI | 7 |
| 2021 | Visualizing the Emergence of Intermediate Visual Patterns in DNNsabstractThis paper proposes a method to visualize the discrimination power of intermediate-layer visual patterns encoded by a DNN. Specifically, we visualize (1) how the DNN gradually learns regional visual patterns in each intermediate layer during the training process, and (2) the effects of the DNN using non-discriminative patterns in low layers to construct disciminative patterns in middle/high layers through the forward propagation. Based on our visualization method, we can quantify knowledge points (i.e. the number of discriminative visual patterns) learned by the DNN to evaluate the representation capacity of the DNN. Furthermore, this method also provides new insights into signal-processing behaviors of existing deep-learning techniques, such as adversarial attacks and knowledge distillation. Shaobo Wang 0001, Quanshi Zhang |
NeurIPS | 3 |
| 2021 | Towards a Unified Game-Theoretic View of Adversarial Perturbations and RobustnessabstractThis paper provides a unified view to explain different adversarial attacks and defense methods, i.e. the view of multi-order interactions between input variables of DNNs. Based on the multi-order interaction, we discover that adversarial attacks mainly affect high-order interactions to fool the DNN. Furthermore, we find that the robustness of adversarially trained DNNs comes from category-specific low-order interactions. Our findings provide a potential method to unify adversarial perturbations and robustness, which can explain the existing robustness-boosting methods in a principle way. Besides, our findings also make a revision of previous inaccurate understanding of the shape bias of adversarially learned features. Our code is available online at https://github.com/Jie-Ren/A-Unified-Game-Theoretic-Interpretation-of-Adversarial-Robustness. Jie Ren 0018, Die Zhang, Yisen Wang 0001, Zhanpeng Zhou, Yiting Chen 0003, Xu Cheng 0005, Xin Wang 0108, Quanshi Zhang |
NeurIPS | 11 |
| 2021 | Interpreting Representation Quality of DNNs for 3D Point Cloud ProcessingabstractIn this paper, we evaluate the quality of knowledge representations encoded in deep neural networks (DNNs) for 3D point cloud processing. We propose a method to disentangle the overall model vulnerability into the sensitivity to the rotation, the translation, the scale, and local 3D structures. Besides, we also propose metrics to evaluate the spatial smoothness of encoding 3D structures, and the representation complexity of the DNN. Based on such analysis, experiments expose representation problems with classic DNNs, and explain the utility of the adversarial training. The code will be released when this paper is accepted. Wen Shen 0002, Qihan Ren, Dongrui Liu, Quanshi Zhang |
NeurIPS | 4 |
| 2021 | Mining Interpretable AOG Representations From Convolutional Networks via Active Question AnsweringabstractIn this paper, we present a method to mine object-part patterns from conv-layers of a pre-trained convolutional neural network (CNN). The mined object-part patterns are organized by an And-Or graph (AOG). This interpretable AOG representation consists of a four-layer semantic hierarchy, i.e., semantic parts, part templates, latent patterns, and neural units. The AOG associates each object part with certain neural units in feature maps of conv-layers. The AOG is constructed with very few annotations (e.g., 3-20) of object parts. We develop a question-answering (QA) method that uses active human-computer communications to mine patterns from a pre-trained CNN, in order to explain features in conv-layers incrementally. During the learning process, our QA method uses the current AOG for part localization. The QA method actively identifies objects, whose feature maps cannot be explained by the AOG. Then, our method asks people to annotate parts on the unexplained objects, and uses answers to discover CNN patterns corresponding to newly labeled parts. In this way, our method gradually grows new branches and refines existing branches on the AOG to semanticize CNN representations. In experiments, our method exhibited a high learning efficiency. Our method used about 1/6- 1/3 of the part annotations for training, but achieved similar or better part-localization performance than fast-RCNN methods. Quanshi Zhang, Jie Ren 0018, Ge Huang, Ruiming Cao, Ying Nian Wu, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Extraction of an Explanatory Graph to Interpret a CNNabstractin a conv-layer usually represents a mixture of object parts. We develop a simple yet effective method to learn an explanatory graph, which automatically disentangles object parts from each filter without any part annotations. Specifically, given the feature map of a filter, we mine neural activations from the feature map, which correspond to different object parts. The explanatory graph is constructed to organize each mined part as a graph node. Each edge connects two nodes, whose corresponding object parts usually co-activate and keep a stable spatial relationship. Experiments show that each graph node consistently represented the same object part through different images, which boosted the transferability of CNN features. The explanatory graph transferred features of object parts to the task of part localization, and our method significantly outperformed other approaches. Quanshi Zhang, Xin Wang 0108, Ruiming Cao, Ying Nian Wu, Feng Shi 0006, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Interpretable CNNs for Object ClassificationabstractThis paper proposes a generic method to learn interpretable convolutional filters in a deep convolutional neural network (CNN) for object classification, where each interpretable filter encodes features of a specific object part. Our method does not require additional annotations of object parts or textures for supervision. Instead, we use the same training data as traditional CNNs. Our method automatically assigns each interpretable filter in a high conv-layer with an object part of a certain category during the learning process. Such explicit knowledge representations in conv-layers of the CNN help people clarify the logic encoded in the CNN, i.e., answering what patterns the CNN extracts from an input image and uses for prediction. We have tested our method using different benchmark CNNs with various architectures to demonstrate the broad applicability of our method. Experiments have shown that our interpretable filters are much more semantically meaningful than traditional filters. Quanshi Zhang, Xin Wang 0108, Ying Nian Wu, Huilin Zhou, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Explaining Knowledge Distillation by Quantifying the KnowledgeabstractThis paper presents a method to interpret the success of knowledge distillation by quantifying and analyzing task-relevant and task-irrelevant visual concepts that are encoded in intermediate layers of a deep neural network (DNN). More specifically, three hypotheses are proposed as follows. 1. Knowledge distillation makes the DNN learn more visual concepts than learning from raw data. 2. Knowledge distillation ensures that the DNN is prone to learning various visual concepts simultaneously. Whereas, in the scenario of learning from raw data, the DNN learns visual concepts sequentially. 3. Knowledge distillation yields more stable optimization directions than learning from raw data. Accordingly, we design three types of mathematical metrics to evaluate feature representations of the DNN. In experiments, we diagnosed various DNNs, and above hypotheses were verified. Xu Cheng 0005, Zhefan Rao, Yilan Chen 0002, Quanshi Zhang |
CVPR | 4 |
| 2020 | 3D-Rotation-Equivariant Quaternion Neural Networks
Wen Shen 0002, Shikun Huang, Zhihua Wei 0001, Quanshi Zhang |
ECCV (20) | 5 |
| 2020 | Knowledge Consistency between Neural Networks and Beyond
Ruofan Liang, Tianlin Li, Quanshi Zhang |
ICLR | 5 |
| 2020 | Interpretable Complex-Valued Neural Networks for Privacy Protection
Liyao Xiang, Hao Zhang 0063, Jie Ren 0018, Quanshi Zhang |
ICLR | 6 |
| 2019 | Interpreting CNNs via Decision TreesabstractThis paper aims to quantitatively explain the rationales of each prediction that is made by a pre-trained convolutional neural network (CNN). We propose to learn a decision tree, which clarifies the specific reason for each prediction made by the CNN at the semantic level. I.e., the decision tree decomposes feature representations in high conv-layers of the CNN into elementary concepts of object parts. In this way, the decision tree tells people which object parts activate which filters for the prediction and how much each object part contributes to the prediction score. Such semantic and quantitative explanations for CNN predictions have specific values beyond the traditional pixel-level analysis of CNNs. More specifically, our method mines all potential decision modes of the CNN, where each mode represents a typical case of how the CNN uses object parts for prediction. The decision tree organizes all potential decision modes in a coarse-to-fine manner to explain CNN predictions at different fine-grained levels. Experiments have demonstrated the effectiveness of the proposed method. Quanshi Zhang, Yu Yang 0007, Ying Nian Wu |
CVPR | 1 |
| 2019 | Explaining Neural Networks Semantically and QuantitativelyabstractThis paper presents a method to pursue a semantic and quantitative explanation for the knowledge encoded in a convolutional neural network (CNN). The estimation of the specific rationale of each prediction made by the CNN presents a key issue of understanding neural networks, and it is of significant values in real applications. In this study, we propose to distill knowledge from the CNN into an explainable additive model, which explains the CNN prediction quantitatively. We discuss the problem of the biased interpretation of CNN predictions. To overcome the biased interpretation, we develop prior losses to guide the learning of the explainable additive model. Experimental results have demonstrated the effectiveness of our method. Runjin Chen, Hao Chen 0099, Ge Huang, Jie Ren 0018, Quanshi Zhang |
ICCV | 5 |
| 2019 | Towards a Deep and Unified Understanding of Deep Neural Models in NLPabstractWe define a unified information-based measure to provide quantitative explanations on how intermediate layers of deep Natural Language Processing (NLP) models leverage information of input words. Our method advances existing explanation methods by addressing issues in coherency and generality. Explanations generated by using our method are consistent and faithful across different timestamps, layers, and models. We show how our method can be applied to four widely used models in NLP and explain their performances on three real-world benchmark datasets. Chaoyu Guan, Xiting Wang, Quanshi Zhang, Runjin Chen, Di He 0001, Xing Xie 0001 |
ICML | 3 |
| 2019 | Visual graph mining for graph matching
Quanshi Zhang, Xuan Song 0001, Yu Yang 0007, Ryosuke Shibasaki |
Comput. Vis. Image Underst. | 1 |
| 2018 | Interpreting CNN Knowledge via an Explanatory GraphabstractThis paper learns a graphical model, namely an explanatory graph, which reveals the knowledge hierarchy hidden inside a pre-trained CNN. Considering that each filter in a conv-layer of a pre-trained CNN usually represents a mixture of object parts, we propose a simple yet efficient method to automatically disentangles different part patterns from each filter, and construct an explanatory graph. In the explanatory graph, each node represents a part pattern, and each edge encodes co-activation relationships and spatial relationships between patterns. More importantly, we learn the explanatory graph for a pre-trained CNN in an unsupervised manner, i.e., without a need of annotating object parts. Experiments show that each graph node consistently represents the same object part through different images. We transfer part patterns in the explanatory graph to the task of part localization, and our method significantly outperforms other approaches. Quanshi Zhang, Ruiming Cao, Feng Shi 0006, Ying Nian Wu, Song-Chun Zhu |
AAAI | 1 |
| 2018 | Examining CNN Representations With Respect to Dataset BiasabstractGiven a pre-trained CNN without any testing samples, this paper proposes a simple yet effective method to diagnose feature representations of the CNN. We aim to discover representation flaws caused by potential dataset bias. More specifically, when the CNN is trained to estimate image attributes, we mine latent relationships between representations of different attributes inside the CNN. Then, we compare the mined attribute relationships with ground-truth attribute relationships to discover the CNN's blind spots and failure modes due to dataset bias. In fact, representation flaws caused by dataset bias cannot be examined by conventional evaluation strategies based on testing images, because testing images may also have a similar bias. Experiments have demonstrated the effectiveness of our method. Quanshi Zhang, Wenguan Wang, Song-Chun Zhu |
AAAI | 1 |
| 2018 | Interpretable Convolutional Neural NetworksabstractThis paper proposes a method to modify a traditional convolutional neural network (CNN) into an interpretable CNN, in order to clarify knowledge representations in high conv-layers of the CNN. In an interpretable CNN, each filter in a high conv-layer represents a specific object part. Our interpretable CNNs use the same training data as ordinary CNNs without a need for any annotations of object parts or textures for supervision. The interpretable CNN automatically assigns each filter in a high conv-layer with an object part during the learning process. We can apply our method to different types of CNNs with various structures. The explicit knowledge representation in an interpretable CNN can help people understand the logic inside a CNN, i.e. what patterns are memorized by the CNN for prediction. Experiments have shown that filters in an interpretable CNN are more semantically meaningful than those in a traditional CNN. The code is available at https://github.com/zqs1022/interpretableCNN. Quanshi Zhang, Ying Nian Wu, Song-Chun Zhu |
CVPR | 1 |
| 2018 | Mining deep And-Or object structures via cost-sensitive question-answer-based active annotations
Quanshi Zhang, Ying Nian Wu, Hao Zhang 0063, Song-Chun Zhu |
Comput. Vis. Image Underst. | 1 |
| 2018 | Visual interpretability for deep learning: a surveyabstractThis paper reviews recent studies in understanding neural-network representations and learning neural networks with interpretable/disentangled middle-layer representations. Although deep neural networks have exhibited superior performance in various tasks, interpretability is always Achilles’ heel of deep neural networks. At present, deep neural networks obtain high discrimination power at the cost of a low interpretability of their black-box representations. We believe that high model interpretability may help people break several bottlenecks of deep learning, e.g., learning from a few annotations, learning via human–computer communications at the semantic level, and semantically debugging network representations. We focus on convolutional neural networks (CNNs), and revisit the visualization of CNN representations, methods of diagnosing representations of pre-trained CNNs, approaches for disentangling pre-trained CNN representations, learning of CNNs with disentangled representations, and middle-to-end learning based on model interpretability. Finally, we discuss prospective trends in explainable artificial intelligence. Quanshi Zhang, Song-Chun Zhu |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2017 | Growing Interpretable Part Graphs on ConvNets via Multi-Shot LearningabstractThis paper proposes a learning strategy that embeds object-part concepts into a pre-trained convolutional neural network (CNN), in an attempt to 1) explore explicit semantics hidden in CNN units and 2) gradually transform the pre-trained CNN into a semantically interpretable graphical model for hierarchical object understanding. Given part annotations on very few (e.g., 3-12) objects, our method mines certain latent patterns from the pre-trained CNN and associates them with different semantic parts. We use a four-layer And-Or graph to organize the CNN units, so as to clarify their internal semantic hierarchy. Our method is guided by a small number of part annotations, and it achieves superior part-localization performance (about 13%-107% improvement in part center prediction on the PASCAL VOC and ImageNet datasets) Quanshi Zhang, Ruiming Cao, Ying Nian Wu, Song-Chun Zhu |
AAAI | 1 |
| 2017 | Mining Object Parts from CNNs via Active Question-AnsweringabstractGiven a convolutional neural network (CNN) that is pre-trained for object classification, this paper proposes to use active question-answering to semanticize neural patterns in conv-layers of the CNN and mine part concepts. For each part concept, we mine neural patterns in the pre-trained CNN, which are related to the target part, and use these patterns to construct an And-Or graph (AOG) to represent a four-layer semantic hierarchy of the part. As an interpretable model, the AOG associates different CNN units with different explicit object parts. We use an active human-computer communication to incrementally grow such an AOG on the pre-trained CNN as follows. We allow the computer to actively identify objects, whose neural patterns cannot be explained by the current AOG. Then, the computer asks human about the unexplained objects, and uses the answers to automatically discover certain CNN patterns corresponding to the missing knowledge. We incrementally grow the AOG to encode new knowledge discovered during the active-learning process. In experiments, our method exhibits high learning efficiency. Our method uses about 1/6-1/3 of the part annotations for training, but achieves similar or better part-localization performance than fast-RCNN methods. Quanshi Zhang, Ruiming Cao, Ying Nian Wu, Song-Chun Zhu |
CVPR | 1 |
| 2017 | Prediction and Simulation of Human Mobility Following Natural DisastersabstractIn recent decades, the frequency and intensity of natural disasters has increased significantly, and this trend is expected to continue. Therefore, understanding and predicting human behavior and mobility during a disaster will play a vital role in planning effective humanitarian relief, disaster management, and long-term societal reconstruction. However, such research is very difficult to perform owing to the uniqueness of various disasters and the unavailability of reliable and large-scale human mobility data. In this study, we collect big and heterogeneous data (e.g., GPS records of 1.6 million users 1 over 3 years, data on earthquakes that have occurred in Japan over 4 years, news report data, and transportation network data) to study human mobility following natural disasters. An empirical analysis is conducted to explore the basic laws governing human mobility following disasters, and an effective human mobility model is developed to predict and simulate population movements. The experimental results demonstrate the efficiency of our model, and they suggest that human mobility following disasters can be significantly more predictable and be more easily simulated than previously thought. Xuan Song 0001, Quanshi Zhang, Yoshihide Sekimoto, Ryosuke Shibasaki, Nicholas Jing Yuan, Xing Xie 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2016 | Object Discovery: Soft Attributed Graph MiningabstractWe categorize this research in terms of its contribution to both graph theory and computer vision. From the theoretical perspective, this study can be considered as the first attempt to formulate the idea of mining maximal frequent subgraphs in the challenging domain of messy visual data, and as a conceptual extension to the unsupervised learning of graph matching. We define a soft attributed pattern (SAP) to represent the common subgraph pattern among a set of attributed relational graphs (ARGs), considering both their structure and attributes. Regarding the differences between ARGs with fuzzy attributes and conventional labeled graphs, we propose a new mining strategy that directly extracts the SAP with the maximal graph size without applying node enumeration. Given an initial graph template and a number of ARGs, we develop an unsupervised method to modify the graph template into the maximal-size SAP. From a practical perspective, this research develops a general platform for learning the category model (i.e., the SAP) from cluttered visual data (i.e., the ARGs) without labeling "what is where," thereby opening the possibility for a series of applications in the era of big visual data. Experiments demonstrate the superior performance of the proposed method on RGB/RGB-D images and videos. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | A Simulator of Human Emergency Mobility Following Disasters: Knowledge Transfer from Big Disaster DataabstractThe frequency and intensity of natural disasters has significantly increased over the past decades and this trend is predicted to continue. Facing these possible and unexpected disasters, understanding and simulating of human emergency mobility following disasters will becomethe critical issue for planning effective humanitarian relief, disaster management, and long-term societal reconstruction. However, due to the uniquenessof various disasters and the unavailability of reliable and large scale human mobility data, such kind of research is very difficult to be performed. Hence, in this paper,we collect big and heterogeneous data (e.g. 1.6 million users' GPS records in three years, 17520 times of Japan earthquake data in four years, news reporting data, transportation network data and etc.) to capture and analyze human emergency mobility following different disasters. By mining these big data, we aim to understand what basic laws govern human mobility following disasters, and develop a general model of human emergency mobility for generating and simulating large amount of human emergency movements. The experimental results and validations demonstrate the efficiency of our simulation model, and suggest that human mobility following disasters may be significantly morepredictable and can be easier simulated than previously thought. Xuan Song 0001, Quanshi Zhang, Yoshihide Sekimoto, Ryosuke Shibasaki, Nicholas Jing Yuan, Xing Xie 0001 |
AAAI | 2 |
| 2015 | Mining And-Or Graphs for Graph Matching and Object DiscoveryabstractThis paper reformulates the theory of graph mining on the technical basis of graph matching, and extends its scope of applications to computer vision. Given a set of attributed relational graphs (ARGs), we propose to use a hierarchical And-Or Graph (AoG) to model the pattern of maximal-size common subgraphs embedded in the ARGs, and we develop a general method to mine the AoG model from the unlabeled ARGs. This method provides a general solution to the problem of mining hierarchical models from unannotated visual data without exhaustive search of objects. We apply our method to RGB/RGB-D images and videos to demonstrate its generality and the wide range of applicability. The code will be available at https://sites.google.com/site/quanshizhang/mining-and-or-graphs. Quanshi Zhang, Ying Nian Wu, Song-Chun Zhu |
ICCV | 1 |
| 2015 | From RGB-D Images to RGB Images: Single Labeling for Mining Visual ModelsabstractMining object-level knowledge, that is, building a comprehensive category model base, from a large set of cluttered scenes presents a considerable challenge to the field of artificial intelligence. How to initiate model learning with the least human supervision (i.e., manual labeling) and how to encode the structural knowledge are two elements of this challenge, as they largely determine the scalability and applicability of any solution. In this article, we propose a model-learning method that starts from a single-labeled object for each category, and mines further model knowledge from a number of informally captured, cluttered scenes. However, in these scenes, target objects are relatively small and have large variations in texture, scale, and rotation. Thus, to reduce the model bias normally associated with less supervised learning methods, we use the robust 3D shape in RGB-D images to guide our model learning, then apply the properly trained category models to both object detection and recognition in more conventional RGB images. In addition to model training for their own categories, the knowledge extracted from the RGB-D images can also be transferred to guide model learning for a new category, in which only RGB images without depth information in the new category are provided for training. Preliminary testing shows that the proposed method performs as well as fully supervised learning methods. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2014 | Intelligent System for Urban Emergency Management during Large-Scale DisasterabstractThe frequency and intensity of natural disasters has significantly increased over the past decades and this trend is predicted to continue. Facing these possible and unexpected disasters, urban emergency management has become the especially important issue for the whole governments around the world.In this paper, we present a novel intelligent system for urban emergency management during the large-scale disasters. The proposed systemstores and manages the global positioning system (GPS) records from mobile devices used by approximately 1.6 million people throughout Japan over one year. By mining and analyzing population movements after the Great East Japan Earthquake, our system can automatically learn a probabilistic model to better understand and simulate human mobility during the emergency situations. Based on the learning model, population mobility in various urban areas impacted by the earthquake throughout Japan can be automatically simulated or predicted. On the basis of such kind of system, it is easy for us to find some new features or population mobility patterns after the recent and unprecedented composite disasters, which are likely to provide valuable experience and play a vital role for future disaster management worldwide. Xuan Song 0001, Quanshi Zhang, Yoshihide Sekimoto, Ryosuke Shibasaki |
AAAI | 2 |
| 2014 | When 3D Reconstruction Meets Ubiquitous RGB-D Imagesabstract3D reconstruction from a single image is a classical problem in computer vision. However, it still poses great challenges for the reconstruction of daily-use objects with irregular shapes. In this paper, we propose to learn 3D reconstruction knowledge from informally captured RGB-D images, which will probably be ubiquitously used in daily life. The learning of 3D reconstruction is defined as a category modeling problem, in which a model for each category is trained to encode category-specific knowledge for 3D reconstruction. The category model estimates the pixel-level 3D structure of an object from its 2D appearance, by taking into account considerable variations in rotation, 3D structure, and texture. Learning 3D reconstruction from ubiquitous RGB-D images creates a new set of challenges. Experimental results have demonstrated the effectiveness of the proposed approach. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
CVPR | 1 |
| 2014 | Attributed Graph Mining and Matching: An Attempt to Define and Extract Soft Attributed PatternsabstractGraph matching and graph mining are two typical areas in artificial intelligence. In this paper, we define the soft attributed pattern (SAP) to describe the common subgraph pattern among a set of attributed relational graphs (ARGs), considering both the graphical structure and graph attributes. We propose a direct solution to extract the SAP with the maximal graph size without node enumeration. Given an initial graph template and a number of ARGs, we modify the graph template into the maximal SAP among the ARGs in an unsupervised fashion. The maximal SAP extraction is equivalent to learning a graphical model (i.e. an object model) from large ARGs (i.e. cluttered RGB/RGB-D images) for graph matching, which extends the concept of "unsupervised learning for graph matching." Furthermore, this study can be also regarded as the first known approach to formulating "maximal graph mining" in the graph domain of ARGs. Our method exhibits superior performance on RGB and RGB-D images. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
CVPR | 1 |
| 2014 | Start from minimum labeling: Learning of 3D object models and point labeling from a large and complex environmentabstractA large category model base can provide object-level knowledge for various perception tasks of the intelligent vehicle system. The automatic and efficient construction of such a model base is highly desirable but challenging. This paper presents a novel semi-supervised approach to discover possible prototype models of 3D object structures from the point cloud of a large and complex environment, given a limited number of seeds in an object category. Our method incrementally trains the models while simultaneously collecting object samples. Considering the bias problem of model learning caused by bias accumulation in a sample collection, we propose to gradually differentiate the standard category model into several sub-category models to represent different intra-category structural styles. Thus, new sub-categories are discovered and modeled, old models are improved, and redundant models for similar structures are deleted iteratively during the learning process. This multiple-model strategy provides several interactive options for the category boundary to deal with the bias problem. Experimental results demonstrate the effectiveness and high efficiency of our approach to model mining from “big point cloud data”. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
ICRA | 1 |
| 2014 | Prediction of human emergency behavior and their mobility following large-scale disasterabstractThe frequency and intensity of natural disasters has significantly increased over the past decades and this trend is predicted to continue. Facing these possible and unexpected disasters, accurately predicting human emergency behavior and their mobility will become the critical issue for planning effective humanitarian relief, disaster management, and long-term societal reconstruction. In this paper, we build up a large human mobility database (GPS records of 1.6 million users over one year) and several different datasets to capture and analyze human emergency behavior and their mobility following the Great East Japan Earthquake and Fukushima nuclear accident. Based on our empirical analysis through these data, we find that human behavior and their mobility following large-scale disaster sometimes correlate with their mobility patterns during normal times, and are also highly impacted by their social relationship, intensity of disaster, damage level, government appointed shelters, news reporting, large population flow and etc. On the basis of these findings, we develop a model of human behavior that takes into account these factors for accurately predicting human emergency behavior and their mobility following large-scale disaster. The experimental results and validations demonstrate the efficiency of our behavior model, and suggest that human behavior and their movements during disasters may be significantly more predictable than previously thought. Xuan Song 0001, Quanshi Zhang, Yoshihide Sekimoto, Ryosuke Shibasaki |
KDD | 2 |
| 2013 | Category Modeling from Just a Single Labeling: Use Depth Information to Guide the Learning of 2D ModelsabstractAn object model base that covers a large number of object categories is of great value for many computer vision tasks. As artifacts are usually designed to have various textures, their structure is the primary distinguishing feature between different categories. Thus, how to encode this structural information and how to start the model learning with a minimum of human labeling become two key challenges for the construction of the model base. We design a graphical model that uses object edges to represent object structures, and this paper aims to incrementally learn this category model from one labeled object and a number of casually captured scenes. However, the incremental model learning may be biased due to the limited human labeling. Therefore, we propose a new strategy that uses the depth information in RGBD images to guide the model learning for object detection in ordinary RGB images. In experiments, the proposed method achieves superior performance as good as the supervised methods that require the labeling of all target objects. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Ryosuke Shibasaki, Huijing Zhao |
CVPR | 1 |
| 2013 | Learning Graph Matching: Oriented to Category Modeling from Cluttered ScenesabstractAlthough graph matching is a fundamental problem in pattern recognition, and has drawn broad interest from many fields, the problem of learning graph matching has not received much attention. In this paper, we redefine the learning of graph matching as a model learning problem. In addition to conventional training of matching parameters, our approach modifies the graph structure and attributes to generate a graphical model. In this way, the model learning is oriented toward both matching and recognition performance, and can proceed in an unsupervised fashion. Experiments demonstrate that our approach outperforms conventional methods for learning graph matching. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
ICCV | 1 |
| 2013 | Unsupervised 3D category discovery and point labeling from a large urban environmentabstractThe building of an object-level knowledge base is the foundation of a new methodology for many perception tasks in artificial intelligence, and is an area that has received increasing attention in recent years. In this paper, we propose, for the first time, to mine category shape patterns directly from a large urban environment, thus constructing a category structure base. Conventionally, category patterns are learned from a large collection of object samples, but automatic object collection requires prior knowledge of category structures. To solve this chicken-and-egg problem, we learn shape patterns from raw segmentations, and then refine these segmentations based on the pattern knowledge. In the process, we solve two challenging problems of knowledge mining. First, as some categories have large intra-category structure variations, we design an entropy-based method to determine the structure variation for each category, in order to establish the correct range of sample collection. Second, because incorrect segmentation is unavoidable without prior knowledge, we propose a novel unsupervised method that uses a pattern competition strategy to identify and subtract shape patterns formed by incorrectly segmented objects. This ensures that shape patterns are meaningful at the object level. Experimental results demonstrated the effectiveness of the proposed method for category structure mining in a large urban environment. Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Ryosuke Shibasaki |
ICRA | 1 |
| 2013 | Modeling and probabilistic reasoning of population evacuation during large-scale disasterabstractThe Great East Japan Earthquake and the Fukushima nuclear accident cause large human population movements and evacuations. Understanding and predicting these movements is critical for planning effective humanitarian relief, disaster management, and long-term societal reconstruction. In this paper, we construct a large human mobility database that stores and manages GPS records from mobile devices used by approximately 1.6 million people throughout Japan from 1 August 2010 to 31 July 2011. By mining this enormous set of Auto-GPS mobile sensor data, the short-term and long-term evacuation behaviors for individuals throughout Japan during this disaster are able to be automatically discovered. To better understand and simulate human mobility during the disasters, we develop a probabilistic model that is able to be effectively trained by the discovered evacuations via machine learning technique. Based on our training model, population mobility in various cities impacted by the disasters throughout the country is able to be automatically simulated or predicted. On the basis of the whole database, developed model, and experimental results, it is easy for us to find some new features or population mobility patterns after the recent severe earthquake, tsunami and release of radioactivity in Japan, which are likely to play a vital role in future disaster relief and management worldwide. Xuan Song 0001, Quanshi Zhang, Yoshihide Sekimoto, Teerayut Horanont, Satoshi Ueyama, Ryosuke Shibasaki |
KDD | 2 |
| 2013 | Unsupervised skeleton extraction and motion capture from 3D deformable matching
Quanshi Zhang, Xuan Song 0001, Xiaowei Shao, Ryosuke Shibasaki, Huijing Zhao |
Neurocomputing | 1 |
| 2013 | A fully online and unsupervised system for large and high-density area surveillance: Tracking, semantic scene learning and abnormality detectionabstractFor reasons of public security, an intelligent surveillance system that can cover a large, crowded public area has become an urgent need. In this article, we propose a novel laser-based system that can simultaneously perform tracking, semantic scene learning, and abnormality detection in a fully online and unsupervised way. Furthermore, these three tasks cooperate with each other in one framework to improve their respective performances. The proposed system has the following key advantages over previous ones: (1) It can cover quite a large area (more than 60×35m), and simultaneously perform robust tracking, semantic scene learning, and abnormality detection in a high-density situation. (2) The overall system can vary with time, incrementally learn the structure of the scene, and perform fully online abnormal activity detection and tracking. This feature makes our system suitable for real-time applications. (3) The surveillance tasks are carried out in a fully unsupervised manner, so that there is no need for manual labeling and the construction of huge training datasets. We successfully apply the proposed system to the JR subway station in Tokyo, and demonstrate that it can cover an area of 60×35m, robustly track more than 150 targets at the same time, and simultaneously perform online semantic scene learning and abnormality detection with no human intervention. Xuan Song 0001, Xiaowei Shao, Quanshi Zhang, Ryosuke Shibasaki, Huijing Zhao, Jinshi Cui, Hongbin Zha |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2012 | Laser-based intelligent surveillance and abnormality detection in extremely crowded scenariosabstractAbnormal activity detection plays a crucial role in surveillance applications, and a surveillance system that can perform robustly in the extremely crowded area has become an urgent need for public security. In this paper, we propose a novel laser-based system which can simultaneously perform the tracking, semantic scene learning and abnormality detection in the large and crowded environment. In our system, a novel abnormality detection model is proposed, and it considers and combines various factors that will influence human activity. Moreover, this model intensively investigate the relationship between pedestrians' social behaviors and their walking scenarios. We successfully applied the proposed system to the JR subway station of Tokyo, which can cover a 60×35m area, robustly track more than 180 targets at the same time and simultaneously perform the online semantic scene learning and abnormality detection with no human intervention. Xuan Song 0001, Xiaowei Shao, Quanshi Zhang, Ryosuke Shibasaki, Huijing Zhao, Hongbin Zha |
ICRA | 3 |
| 2009 | Moving object classification using horizontal laser scan dataabstractMotivated by two potential applications, i.e. enhancing driving safety and traffic data collection, a system has been developed using a single-layer horizontal laser scanner as the major sensor for both localization and perception of the surroundings in a large dynamic urban environment. This research focuses on a classification method, that given a stream of laser measurements, classify the moving object into either a person, a group of people, a bicycle or a car. In this research, a number of features are defined after examining the property of data appearance. A classification method is proposed after examining the likelihood measures between each pair of feature and class. Experimental results are presented, demonstrating that the algorithm has efficiency with respect to both driving safety and traffic data collection in highly dynamic environment. Huijing Zhao, Quanshi Zhang, Masaki Chiba, Ryosuke Shibasaki, Jinshi Cui, Hongbin Zha |
ICRA | 2 |