Bo Zhang 0010

dblp:36/2259-10 · DBLP profile ↗
← Back
197ranked-venue papers
13as first author
21since 2021 · last 2024
0000-0002-1566-3275ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 120 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 79 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 28 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 2Human-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2024 Language-Driven Anchors for Zero-Shot Adversarial Robustness
abstract
Deep Neural Networks (DNNs) are known to be susceptible to adversarial attacks. Previous researches mainly fo-cus on improving adversarial robustness in the fully super-vised setting, leaving the challenging domain of zero-shot adversarial robustness an open question. In this work, we investigate this domain by leveraging the recent advances in large vision-language models, such as CLIP, to introduce zero-shot adversarial robustness to DNNs. We pro-pose LAAT, a Language-driven, Anchor-based Adversarial Training strategy. LAAT utilizes the features of a text en-coder for each category as fixed anchors (normalized feature embeddings) for each category, which are then employed for adversarial training. By leveraging the semantic consistency of the text encoders, LAAT aims to enhance the adversarial robustness of the image model on novel cate-gories. However, naively using text encoders leads to poor results. Through analysis, we identified the issue to be the high cosine similarity between text encoders. We then design an expansion algorithm and an alignment cross-entropy loss to alleviate the problem. Our experimental results demonstrated that LAAT significantly improves zero-shot adversarial robustness over state-of-the-art methods. LAAT has the potential to enhance adversarial robustness by large-scale multimodal models, especially when labeled data is unavailable during training. Code is available at https://github.com/LixiaoTHU/LAAT.
Xiao Li 0028, Wei Zhang 0370, Zhanhao Hu, Bo Zhang 0010, Xiaolin Hu 0001
CVPR5
2024 GATS: Generative Audience Targeting System for Online Advertising
Zhongde Chen, Bo Zhang 0010, Yankun Ren, Xin Dong 0012, Lei Cheng 0005, Xinxing Yang, Jun Zhou 0011, Linjian Mo
SIGIR3
2024 Enhancing Sequential Recommenders with Augmented Knowledge from Aligned Large Language Models
abstract
Recommender systems are widely used in various online platforms. In the context of sequential recommendation, it is essential to accurately capture the chronological patterns in user activities to generate relevant recommendations. Conventional ID-based sequential recommenders have shown promise but lack comprehensive real-world knowledge about items, limiting their effectiveness. Recent advancements in Large Language Models (LLMs) offer the potential to bridge this gap by leveraging the extensive real-world knowledge encapsulated in LLMs. However, integrating LLMs into sequential recommender systems comes with its own challenges, including inadequate representation of sequential behavior patterns and long inference latency. In this paper, we propose SeRALM (Enhancing Sequential Recommenders with Augmented Knowledge from Aligned Large Language Models) to address these challenges. SeRALM integrates LLMs with conventional ID-based sequential recommenders for sequential recommendation tasks. We combine text-format knowledge generated by LLMs with item IDs and feed this enriched data into ID-based recommenders, benefitting from the strengths of both paradigms. Moreover, we develop a theoretically underpinned alignment training method to refine LLMs' generation using feedback from ID-based recommenders for better knowledge augmentation. We also present an asynchronous technique to expedite the alignment training process. Experimental results on public benchmarks demonstrate that SeRALM significantly improves the performances of ID-based sequential recommenders. Further, a series of ablation studies and analyses corroborate SeRALM's proficiency in steering LLMs to generate more pertinent and advantageous knowledge across diverse scenarios.
Yankun Ren, Zhongde Chen, Xinxing Yang, Lei Cheng 0005, Bo Zhang 0010, Linjian Mo, Jun Zhou 0011
SIGIR7
2024 Understanding adversarial attacks on observations in deep reinforcement learning
You Qiaoben, Chengyang Ying, Xinning Zhou, Hang Su 0006, Jun Zhu 0001, Bo Zhang 0010
Sci. China Inf. Sci.6
2024 To Boost Zero-Shot Generalization for Embodied Reasoning With Vision-Language Pre-Training
abstract
Recently, there exists an increased research interest in embodied artificial intelligence (EAI), which involves an agent learning to perform a specific task when dynamically interacting with the surrounding 3D environment. There into, a new challenge is that many unseen objects may appear due to the increased number of object categories in 3D scenes. It makes developing models with strong zero-shot generalization ability to new objects necessary. Existing work tries to achieve this goal by providing embodied agents with massive high-quality human annotations closely related to the task to be learned, while it is too costly in practice. Inspired by recent advances in pre-trained models in 2D visual tasks, we attempt to boost zero-shot generalization for embodied reasoning with vision-language pre-training that can encode common sense as general prior knowledge. To further improve its performance on a specific task, we rectify the pre-trained representation through masked scene graph modeling (MSGM) in a self-supervised manner, where the task-specific knowledge is learned from iterative message passing. Our method can improve a variety of representative embodied reasoning tasks by a large margin (e.g., over 5.0% w.r.t. answer accuracy on MP3D-EQA dataset that consists of many real-world scenes with a large number of new objects during testing), and achieve the new state-of-the-art performance.
Xingxing Zhang 0001, Siyang Zhang, Jun Zhu 0001, Bo Zhang 0010
IEEE Trans. Image Process.5
2024 Probabilistic Neural-Symbolic Models With Inductive Posterior Constraints
abstract
Neural-symbolic models provide a powerful tool to tackle complex visual reasoning tasks by combining symbolic program execution for reasoning and deep representation learning for visual recognition. A probabilistic formulation of such models with stochastic latent variables can obtain an interpretable and legible reasoning system with less supervision. However, it is still nontrivial to generate reasonable symbolic structures without the guidance of domain knowledge, since it generally involves an optimization problem with both continuous and discrete variables. Despite the challenges, the interpretability of such symbolic structures provides an interface to regularize their generation by domain knowledge. In this article, we propose to incorporate the available domain knowledge into the learning process of probabilistic neural-symbolic (PNS) models via posterior constraints that directly regularize the structure posterior. In this way, our model is able to identify a middle point where the structure generation process mainly learns from data but also selectively borrows information from domain knowledge. We further present inductive reasoning where the posterior constraints can be automatically reweighted to handle noisy annotations. The experimental results show that our method achieves state-of-the-art performance on major abstract reasoning datasets and enjoys good generalization capability and data efficiency.
Hang Su 0006, Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
IEEE Trans. Neural Networks Learn. Syst.5
2023 Physically Realizable Natural-Looking Clothing Textures Evade Person Detectors via 3D Modeling
abstract
Recent works have proposed to craft adversarial clothes for evading person detectors, while they are either only effective at limited viewing angles or very conspicuous to humans. We aim to craft adversarial texture for clothes based on 3D modeling, an idea that has been used to craft rigid adversarial objects such as a 3D-printed turtle. Unlike rigid objects, humans and clothes are non-rigid, leading to difficulties in physical realization. In order to craft natural-looking adversarial clothes that can evade person detectors at multiple viewing angles, we propose adversarial camou-flage textures (AdvCaT) that resemble one kind of the typical textures of daily clothes, camouflage textures. We leverage the Voronoi diagram and Gumbel-softmax trick to parameterize the camouflage textures and optimize the parameters via 3D modeling. Moreover, we propose an efficient augmentation pipeline on 3D meshes combining topologically plausible projection (TopoProj) and Thin Plate Spline (TPS) to narrow the gap between digital and real-world objects. We printed the developed 3D texture pieces on fabric materials and tailored them into T-shirts and trousers. Experiments show high attack success rates of these clothes against multiple detectors.
Zhanhao Hu, Wenda Chu, Xiaopei Zhu, Bo Zhang 0010, Xiaolin Hu 0001
CVPR5
2023 A Collaborative Transfer Learning Framework for Cross-domain Recommendation
abstract
In the recommendation systems, there are multiple business domains to meet the diverse interests and needs of users, and the click-through rate(CTR) of each domain can be quite different, which leads to the demand for CTR prediction modeling for different business domains. The industry solution is to use domain-specific models or transfer learning techniques for each domain. The disadvantage of the former is that the data from other domains is not utilized by a single domain model, while the latter leverage all the data from different domains, but the fine-tuned model of transfer learning may trap the model in a local optimum of the source domain, making it difficult to fit the target domain. Meanwhile, significant differences in data quantity and feature schemas between different domains, known as domain shift, may lead to negative transfer in the process of transferring. To overcome these challenges, we propose the Collaborative Cross-Domain Transfer Learning Framework (CCTL). CCTL evaluates the information gain of the source domain on the target domain using a symmetric companion network and adjusts the information transfer weight of each source domain sample using the information flow network. This approach enables full utilization of other domain data while avoiding negative migration. Additionally, a representation enhancement network is used as an auxiliary task to preserve domain-specific features. Comprehensive experiments on both public and real-world industrial datasets, CCTL achieved SOTA score on offline metrics. At the same time, the CCTL algorithm has been deployed in Meituan, bringing 4.37% CTR and 5.43% GMV lift, which is significant to the business.
Wei Zhang 0370, Pengye Zhang, Bo Zhang 0010, Dong Wang 0022
KDD3
2023 Toward the third generation artificial intelligence
Bo Zhang 0010, Jun Zhu 0001, Hang Su 0006
Sci. China Inf. Sci.1
2023 Recognizing Object by Components With Human Prior Knowledge Enhances Adversarial Robustness of Deep Neural Networks
abstract
Adversarial attacks can easily fool object recognition systems based on deep neural networks (DNNs). Although many defense methods have been proposed in recent years, most of them can still be adaptively evaded. One reason for the weak adversarial robustness may be that DNNs are only supervised by category labels and do not have part-based inductive bias like the recognition process of humans. Inspired by a well-known theory in cognitive psychology - recognition-by-components, we propose a novel object recognition model ROCK (Recognizing Object by Components with human prior Knowledge). It first segments parts of objects from images, then scores part segmentation results with predefined human prior knowledge, and finally outputs prediction based on the scores. The first stage of ROCK corresponds to the process of decomposing objects into parts in human vision. The second stage corresponds to the decision process of the human brain. ROCK shows better robustness than classical recognition models across various attack settings. These results encourage researchers to rethink the rationality of currently widely-used DNN-based object recognition models and explore the potential of part-based models, once important but recently ignored, for improving robustness.
Xiao Li 0028, Ziqi Wang 0003, Bo Zhang 0010, Fuchun Sun 0001, Xiaolin Hu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Graph Based Long-Term And Short-Term Interest Model for Click-Through Rate Prediction
abstract
Click-through rate (CTR) prediction aims to predict the probability that the user will click an item, which has been one of the key tasks in online recommender and advertising systems. In such systems, rich user behavior (viz. long- and short-term) has been proved to be of great value in capturing user interests. Both industry and academy have paid much attention to this topic and propose different approaches to modeling with long-term and short-term user behavior data. But there are still some unresolved issues. More specially, (1) rule and truncation based methods to extract information from long-term behavior are easy to cause information loss, and (2) single feedback behavior regardless of scenario to extract information from short-term behavior lead to information confusion and noise. To fill this gap, we propose a Graph based Long-term and Short-term interest Model, termed GLSM. It consists of a multi-interest graph structure for capturing long-term user behavior, a multi-scenario heterogeneous sequence model for modeling short-term information, then an adaptive fusion mechanism to fused information from long-term and short-term behaviors. Comprehensive experiments on real-world datasets, GLSM achieved SOTA score on offline metrics. At the same time, the GLSM algorithm has been deployed in our industrial application, bringing 4.9% CTR and 4.3% GMV lift, which is significant to the business
Huinan Sun, Guangliang Yu, Pengye Zhang, Bo Zhang 0010, Dong Wang 0022
CIKM4
2022 Adversarial Texture for Fooling Person Detectors in the Physical World
abstract
Nowadays, cameras equipped with AI systems can capture and analyze images to detect people automatically. However, the AI system can make mistakes when receiving deliberately designed patterns in the real world, i.e., physical adversarial examples. Prior works have shown that it is possible to print adversarial patches on clothes to evade DNN-based person detectors. However, these adversarial examples could have catastrophic drops in the attack success rate when the viewing angle (i.e., the camera's angle towards the object) changes. To perform a multi-angle attack, we propose Adversarial Texture (AdvTexture). AdvTexture can cover clothes with arbitrary shapes so that people wearing such clothes can hide from person detectors from different viewing angles. We propose a generative method, named Toroidal-Cropping-based Expandable Generative Attack (TC-EGA), to craft AdvTexture with repetitive structures. We printed several pieces of cloth with AdvTexure and then made T-shirts, skirts, and dresses in the physical world. Experiments showed that these clothes could fool person detectors in the physical world.
Zhanhao Hu, Xiaopei Zhu, Fuchun Sun 0001, Bo Zhang 0010, Xiaolin Hu 0001
CVPR5
2022 Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models
Fan Bao, Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
ICLR4
2022 Estimating the Optimal Covariance with Imperfect Mean in Diffusion Probabilistic Models
abstract
Diffusion probabilistic models (DPMs) are a class of powerful deep generative models (DGMs). Despite their success, the iterative generation process over the full timesteps is much less efficient than other DGMs such as GANs. Thus, the generation performance on a subset of timesteps is crucial, which is greatly influenced by the covariance design in DPMs. In this work, we consider diagonal and full covariances to improve the expressive power of DPMs. We derive the optimal result for such covariances, and then correct it when the mean of DPMs is imperfect. Both the optimal and the corrected ones can be decomposed into terms of conditional expectations over functions of noise. Building upon it, we propose to estimate the optimal covariance and its correction given imperfect mean by learning these conditional expectations. Our method can be applied to DPMs with both discrete and continuous timesteps. We consider the diagonal covariance in our implementation for computational efficiency. For an efficient practical implementation, we adopt a parameter sharing scheme and a two-stage training process. Empirically, our method outperforms a wide variety of covariance design on likelihood results, and improves the sample quality especially on a small number of timesteps.
Fan Bao, Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
ICML5
2022 Fast Lossless Neural Compression with Integer-Only Discrete Flows
abstract
By applying entropy codecs with learned data distributions, neural compressors have significantly outperformed traditional codecs in terms of compression ratio. However, the high inference latency of neural networks hinders the deployment of neural compressors in practical applications. In this work, we propose Integer-only Discrete Flows (IODF) an efficient neural compressor with integer-only arithmetic. Our work is built upon integer discrete flows, which consists of invertible transformations between discrete random variables. We propose efficient invertible transformations with integer-only arithmetic based on 8-bit quantization. Our invertible transformation is equipped with learnable binary gates to remove redundant filters during inference. We deploy IODF with TensorRT on GPUs, achieving $10\times$ inference speedup compared to the fastest existing neural compressors, while retaining the high compression rates on ImageNet32 and ImageNet64.
Jianfei Chen 0001, Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
ICML5
2022 Amplification trojan network: Attack deep neural networks by amplifying their inherent weakness
Zhanhao Hu, Jun Zhu 0001, Bo Zhang 0010, Xiaolin Hu 0001
Neurocomputing3
2022 Triple Generative Adversarial Networks
abstract
We propose a unified game-theoretical framework to perform classification and conditional image generation given limited supervision. It is formulated as a three-player minimax game consisting of a generator, a classifier and a discriminator, and therefore is referred to as Triple Generative Adversarial Network (Triple-GAN). The generator and the classifier characterize the conditional distributions between images and labels to perform conditional generation and classification, respectively. The discriminator solely focuses on identifying fake image-label pairs. Theoretically, the three-player formulation guarantees consistency. Namely, under a nonparametric assumption, the unique equilibrium of the game is that the distributions characterized by the generator and the classifier converge to the data distribution. As a byproduct of the three-player formulation, Triple-GAN is flexible to incorporate different semi-supervised classifiers and GAN architectures. We evaluate Triple-GAN in two challenging settings, namely, semi-supervised learning and the extremely low data regime. In both settings, Triple-GAN can achieve excellent classification results and generate meaningful samples in a specific class simultaneously. In particular, using a commonly adopted 13-layer CNN classifier, Triple-GAN outperforms extensive semi-supervised learning methods substantially on several benchmarks no matter data augmentation is applied or not.
Chongxuan Li, Kun Xu 0004, Jun Zhu 0001, Bo Zhang 0010
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 User-Level Privacy-Preserving Federated Learning: Analysis and Performance Optimization
abstract
Federated learning (FL), as a type of collaborative machine learning framework, is capable of preserving private data from mobile terminals (MTs) while training the data into useful models. Nevertheless, from a viewpoint of information theory, it is still possible for a curious server to infer private information from the shared models uploaded by MTs. To address this problem, we first make use of the concept of local differential privacy (LDP), and propose a user-level differential privacy (UDP) algorithm by adding artificial noise to the shared models before uploading them to servers. According to our analysis, the UDP framework can realize$(\epsilon _{i}, \delta _{i})$-LDP for the$i$th MT with adjustable privacy protection levels by varying the variances of the artificial noise processes. We then derive a theoretical convergence upper-bound for the UDP algorithm. It reveals that there exists an optimal number of communication rounds to achieve the best learning performance. More importantly, we propose a communication rounds discounting (CRD) method. Compared with the heuristic search method, the proposed CRD method can achieve a much better trade-off between the computational complexity of searching and the convergence performance. Extensive experiments indicate that our UDP algorithm using the proposed CRD method can effectively improve both the training efficiency and model quality for the given privacy protection levels.
Kang Wei 0004, Jun Li 0004, Ming Ding 0001, Chuan Ma 0001, Hang Su 0006, Bo Zhang 0010, H. Vincent Poor
IEEE Trans. Mob. Comput.6
2021 MagDR: Mask-Guided Detection and Reconstruction for Defending Deepfakes
abstract
1Deepfakes raised serious concerns on the authenticity of visual contents. Prior works revealed the possibility to disrupt deepfakes by adding adversarial perturbations to the source data, but we argue that the threat has not been eliminated yet. This paper presents MagDR, a mask-guided detection and reconstruction pipeline for defending deepfakes from adversarial attacks. MagDR starts with a detection module that defines a few criteria to judge the abnormality of the output of deepfakes, and then uses it to guide a learnable reconstruction procedure. Adaptive masks are extracted to capture the change in local facial regions. In experiments, MagDR defends three main tasks of deepfakes, and the learned reconstruction pipeline transfers across input data, showing promising performance in defending both black-box and white-box attacks.
Lingxi Xie, Shanmin Pang, Bo Zhang 0010
CVPR5
2021 Variational (Gradient) Estimate of the Score Function in Energy-based Latent Variable Models
abstract
This paper presents new estimates of the score function and its gradient with respect to the model parameters in a general energy-based latent variable model (EBLVM). The score function and its gradient can be expressed as combinations of expectation and covariance terms over the (generally intractable) posterior of the latent variables. New estimates are obtained by introducing a variational posterior to approximate the true posterior in these terms. The variational posterior is trained to minimize a certain divergence (e.g., the KL divergence) between itself and the true posterior. Theoretically, the divergence characterizes upper bounds of the bias of the estimates. In principle, our estimates can be applied to a wide range of objectives, including kernelized Stein discrepancy (KSD), score matching (SM)-based methods and exact Fisher divergence with a minimal model assumption. In particular, these estimates applied to SM-based methods outperform existing methods in learning EBLVMs on several image datasets.
Fan Bao, Kun Xu 0004, Chongxuan Li, Lanqing Hong, Jun Zhu 0001, Bo Zhang 0010
ICML6
2021 Stability and Generalization of Bilevel Programming in Hyperparameter Optimization
abstract
The (gradient-based) bilevel programming framework is widely used in hyperparameter optimization and has achieved excellent performance empirically. Previous theoretical work mainly focuses on its optimization properties, while leaving the analysis on generalization largely open. This paper attempts to address the issue by presenting an expectation bound w.r.t. the validation set based on uniform stability. Our results can explain some mysterious behaviours of the bilevel programming in practice, for instance, overfitting to the validation set. We also present an expectation bound for the classical cross-validation algorithm. Our results suggest that gradient-based algorithms can be better than cross-validation under certain conditions in a theoretical perspective. Furthermore, we prove that regularization terms in both the outer and inner levels can relieve the overfitting problem in gradient-based algorithms. In experiments on feature learning and data reweighting for noisy labels, we corroborate our theoretical findings.
Fan Bao, Guoqiang Wu, Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
NeurIPS5
2020 Dynamic Network Pruning with Interpretable Layerwise Channel Selection
abstract
Dynamic network pruning achieves runtime acceleration by dynamically determining the inference paths based on different inputs. However, previous methods directly generate continuous decision values for each weight channel, which cannot reflect a clear and interpretable pruning process. In this paper, we propose to explicitly model the discrete weight channel selections, which encourages more diverse weights utilization, and achieves more sparse runtime inference paths. Meanwhile, with the help of interpretable layerwise channel selections in the dynamic network, we can visualize the network decision paths explicitly for model interpretability. We observe that there are clear differences in the layerwise decisions between normal and adversarial examples. Therefore, we propose a novel adversarial example detection algorithm by discriminating the runtime decision features. Experiments show that our dynamic network achieves higher prediction accuracy under the similar computing budgets on CIFAR10 and ImageNet datasets compared to traditional static pruning methods and other dynamic pruning approaches. The proposed adversarial detection algorithm can significantly improve the state-of-the-art detection rate across multiple attacks, which provides an opportunity to build an interpretable and robust model.
Xiaolin Hu 0001, Bo Zhang 0010, Hang Su 0006
AAAI4
2020 Pruning from Scratch
abstract
Network pruning is an important research field aiming at reducing computational costs of neural networks. Conventional approaches follow a fixed paradigm which first trains a large and redundant network, and then determines which units (e.g., channels) are less important and thus can be removed. In this work, we find that pre-training an over-parameterized model is not necessary for obtaining the target pruned structure. In fact, a fully-trained over-parameterized model will reduce the search space for the pruned structure. We empirically show that more diverse pruned structures can be directly pruned from randomly initialized weights, including potential models with better performance. Therefore, we propose a novel network pruning pipeline which allows pruning from scratch with little training overhead. In the experiments for compressing classification models on CIFAR10 and ImageNet datasets, our approach not only greatly reduces the pre-training burden of traditional pruning methods, but also achieves similar or even higher accuracy under the same computation budgets. Our results facilitate the community to rethink the effectiveness of existing techniques used for network pruning.
Lingxi Xie, Jun Zhou 0011, Hang Su 0006, Bo Zhang 0010, Xiaolin Hu 0001
AAAI6
2020 A Wasserstein Minimum Velocity Approach to Learning Unnormalized Models
abstract
Score matching provides an effective approach to learning flexible unnormalized models, but its scalability is limited by the need to evaluate a second-order derivative. In this paper, we present a scalable approximation to a general family of learning objectives including score matching, by observing a new connection between these objectives and Wasserstein gradient flows. We present applications with promise in learning neural density estimators on manifolds, and training implicit variational and Wasserstein auto-encoders with a manifold-valued prior.
Ziyu Wang 0006, Shuyu Cheng, Yueru Li, Jun Zhu 0001, Bo Zhang 0010
AISTATS5
2020 Training Interpretable Convolutional Neural Networks by Differentiating Class-Specific Filters
Zhihao Ouyang, Yuyuan Zeng, Hang Su 0006, Shutao Xia, Jun Zhu 0001, Bo Zhang 0010
ECCV (2)8
2020 To Relieve Your Headache of Training an MRF, Take AdVIL
Chongxuan Li, Kun Xu 0004, Max Welling, Jun Zhu 0001, Bo Zhang 0010
ICLR6
2020 Understanding and Stabilizing GANs' Training Dynamics Using Control Theory
abstract
Generative adversarial networks (GANs) are effective in generating realistic images but the training is often unstable. There are existing efforts that model the training dynamics of GANs in the parameter space but the analysis cannot directly motivate practically effective stabilizing methods. To this end, we present a conceptually novel perspective from control theory to directly model the dynamics of GANs in the frequency domain and provide simple yet effective methods to stabilize GAN’s training. We first analyze the training dynamic of a prototypical Dirac GAN and adopt the widely-used closed-loop control (CLC) to improve its stability. We then extend CLC to stabilize the training dynamic of normal GANs, which can be implemented as an L2 regularizer on the output of the discriminator. Empirical results show that our method can effectively stabilize the training and obtain state-of-the-art performance on data generation tasks.
Kun Xu 0004, Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
ICML4
2020 Discrete Memory Addressing Variational Autoencoder for Visual Concept Learning
abstract
A substantial aspect of general intelligence is the ability to summarize basic building blocks from various high-level concepts. Artificial vision systems with such hierarchical property can not only perform accurate reasoning for complex observations, but also learn useful low-level knowledge shared across scenes. To achieve this goal, we propose a discrete memory addressing VAE model (DM-VAE) for explicitly memorizing and reasoning about shared primitives in images. A time-persistence memory module is used to store the learned abstract knowledge and to interact with the generative model. The model decides what to pay attention to at each step, and constructs the primitive library automatically as the learning progresses in a fully unsupervised setting. While performing inference, the model attempts to interpret a new observation as a combination of previously learned elements. We further derive a proper variational lower bound which can be optimized efficiently. We conduct visual comprehension experiments on images and demonstrate that our model is able to search, identify, and memorize semantically meaningful primitive concepts.
Yanze Min, Hang Su 0006, Jun Zhu 0001, Bo Zhang 0010
IJCNN4
2020 Bi-level Score Matching for Learning Energy-based Latent Variable Models
abstract
Score matching (SM) provides a compelling approach to learn energy-based models (EBMs) by avoiding the calculation of partition function. However, it remains largely open to learn energy-based latent variable models (EBLVMs), except some special cases. This paper presents a bi-level score matching (BiSM) method to learn EBLVMs with general structures by reformulating SM as a bi-level optimization problem. The higher level introduces a variational posterior of the latent variables and optimizes a modified SM objective, and the lower level optimizes the variational posterior to fit the true posterior. To solve BiSM efficiently, we develop a stochastic optimization algorithm with gradient unrolling. Theoretically, we analyze the consistency of BiSM and the convergence of the stochastic algorithm. Empirically, we show the promise of BiSM in Gaussian restricted Boltzmann machines and highly nonstructural EBLVMs parameterized by deep convolutional neural networks. BiSM is comparable to the widely adopted contrastive divergence and SM methods when they are applicable; and can learn complex EBLVMs with intractable posteriors to generate natural images.
Fan Bao, Chongxuan Li, Taufik Xu, Hang Su 0006, Jun Zhu 0001, Bo Zhang 0010
NeurIPS6
2020 Learning Implicit Generative Models by Teaching Density Estimators
Kun Xu 0004, Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
ECML/PKDD (2)5
2020 MaHRL: Multi-goals Abstraction Based Deep Hierarchical Reinforcement Learning for Recommendations
abstract
As huge commercial value of the recommender system, there has been growing interest to improve its performance in recent years. The majority of existing methods have achieved great improvement on the metric of click, but perform poorly on the metric of conversion possibly due to its extremely sparse feedback signal. To track this challenge, we design a novel deep hierarchical reinforcement learning based recommendation framework to model consumers' hierarchical purchase interest. Specifically, the high-level agent catches long-term sparse conversion interest, and automatically sets abstract goals for low-level agent, while the low-level agent follows the abstract goals and catches short-term click interest via interacting with real-time environment. To solve the inherent problem in hierarchical reinforcement learning, we propose a novel multi-goals abstraction based deep hierarchical reinforcement learning algorithm (MaHRL). Our proposed algorithm contains three contributions: 1) the high-level agent generates multiple goals to guide the low-level agent in different sub-periods, which reduces the difficulty of approaching high-level goals; 2) different goals share the same state encoder structure and its parameters, which increases the update frequency of the high-level agent and thus accelerates the convergence of our proposed algorithm; 3) an appreciated reward assignment mechanism is designed to allocate rewards in each goal so as to coordinate different goals in a consistent direction. We evaluate our proposed algorithm based on a real-world e-commerce dataset and validate its effectiveness.
Dongyang Zhao, Liang Zhang 0042, Bo Zhang 0010, Lizhou Zheng, Yongjun Bao, Weipeng Yan
SIGIR3
2020 Interpret Neural Networks by Extracting Critical Subnetworks
abstract
In recent years, deep neural networks have achieved excellent performance in many fields of artificial intelligence. The requirements for the interpretability and robustness of neural networks are also increasing. In this paper, we propose to understand the functional mechanism of neural networks by extracting critical subnetworks. Specifically, we denote the critical subnetworks as a group of important channels across layers such that if they were suppressed to zeros, the final test performance would deteriorate severely. This novel perspective can not only reveal the layerwise semantic behavior within the model but also present more accurate visual explanations appearing in the data through attribution methods. Moreover, we propose two adversarial example detection methods based on the properties of sample-specific and class-specific subnetworks, which provides the possibility for increasing the model robustness.
Hang Su 0006, Bo Zhang 0010, Xiaolin Hu 0001
IEEE Trans. Image Process.3
2020 Learning Reliable Visual Saliency For Model Explanations
abstract
By highlighting important features that contribute to model prediction, visual saliency is used as a natural form to interpret the working mechanism of deep neural networks. Numerous methods have been proposed to achieve better saliency results. However, we find that previous visual saliency methods are not reliable enough to provide meaningful interpretation through a simple sanity check: saliency methods are required to explain the output of non-maximum prediction classes, which are usually not ground-truth classes. For example, let the methods interpret an image of “dog” given a wrong class label “fish” as the query. This procedure can test whether these methods reliably interpret model's predictions based on existing features that appear in the data. Our experiments show that previous methods failed to pass the test by generating similar saliency maps or scattered patterns. This false saliency response can be dangerous in certain scenarios, such as medical diagnosis. We find that these failure cases are mainly due to the attribution vanishing and adversarial noise within these methods. In order to learn reliable visual saliency, we propose a simple method that requires the output of the model to be close to the original output while learning an explanatory saliency mask. To enhance the smoothness of the optimized saliency masks, we then propose a simple Hierarchical Attribution Fusion (HAF) technique. In order to fully evaluate the reliability of visual saliency methods, we propose a new task Disturbed Weakly Supervised Object Localization (D-WSOL) to measure whether these methods can correctly attribute the model's output to existing features. Experiments show that previous methods fail to meet this standard, and our approach helps to improve the reliability by suppressing false saliency responses. After observing a significant layout difference in saliency masks between real and adversarial samples. we propose to train a simple CNN on these learned hierarchical attribution masks to distinguish adversarial samples. Experiments show that our method can improve detection performance over other approaches significantly.
Hang Su 0006, Bo Zhang 0010, Xiaolin Hu 0001
IEEE Trans. Multim.3
2019 Regularized Adversarial Sampling and Deep Time-aware Attention for Click-Through Rate Prediction
abstract
Improving the performance of click-through rate (CTR) prediction remains one of the core tasks in online advertising systems. With the rise of deep learning, CTR prediction models with deep networks remarkably enhance model capacities. In deep CTR models, exploiting users' historical data is essential for learning users' behaviors and interests. As existing CTR prediction works neglect the importance of the temporal signals when embed users' historical clicking records, we propose a time-aware attention model which explicitly uses absolute temporal signals for expressing the users' periodic behaviors and relative temporal signals for expressing the temporal relation between items. Besides, we propose a regularized adversarial sampling strategy for negative sampling which eases the classification imbalance of CTR data and can make use of the strong guidance provided by the observed negative CTR samples. The adversarial sampling strategy significantly improves the training efficiency, and can be co-trained with the time-aware attention model seamlessly. Experiments are conducted on real-world CTR datasets from both in-station and out-station advertising places.
Yikai Wang 0001, Liang Zhang 0042, Quanyu Dai, Fuchun Sun 0001, Bo Zhang 0010, Weipeng Yan, Yongjun Bao
CIKM5
2019 Function Space Particle Optimization for Bayesian Neural Networks
Ziyu Wang 0006, Tongzheng Ren, Jun Zhu 0001, Bo Zhang 0010
ICLR (Poster)4
2019 Joint Cluster Unary Loss for Efficient Cross-Modal Hashing
abstract
Recently, cross-modal deep hashing has received broad attention for solving cross-modal retrieval problems efficiently. Most cross-modal hashing methods generate $O(n^2)$ data pairs and $O(n^3)$ data triplets for training, but the training procedure is less efficient because the complexity is high for large-scale dataset. In this paper, we propose a novel and efficient cross-modal hashing algorithm named Joint Cluster Cross-Modal Hashing (JCCH). First, We introduce the Cross-Modal Unary Loss (CMUL) with $O(n)$ complexity to bridge the traditional triplet loss and classification-based unary loss, and the JCCH algorithm is introduced with CMUL. Second, a more accurate bound of the triplet loss for structured multilabel data is introduced in CMUL. The resultant hashcodes form several clusters in which the hashcodes in the same cluster share similar semantic information, and the heterogeneity gap on different modalities is diminished by sharing the clusters. Experiments on large-scale datasets show that the proposed method is superior over or comparable with state-of-the-art cross-modal hashing methods, and training with the proposed method is more efficient than others.
Jianmin Li 0001, Bo Zhang 0010
ICMR3
2019 Multi-objects Generation with Amortized Structural Regularization
abstract
Deep generative models (DGMs) have shown promise in image generation. However, most of the existing methods learn a model by simply optimizing a divergence between the marginal distributions of the model and the data, and often fail to capture rich structures, such as attributes of objects and their relationships, in an image. Human knowledge is a crucial element to the success of DGMs to infer these structures, especially in unsupervised learning. In this paper, we propose amortized structural regularization (ASR), which adopts posterior regularization (PR) to embed human knowledge into DGMs via a set of structural constraints. We derive a lower bound of the regularized log-likelihood in PR and adopt the amortized inference technique to jointly optimize the generative model and an auxiliary recognition model for inference efficiently. Empirical results show that ASR outperforms the DGM baselines in terms of inference performance and sample quality.
Taufik Xu, Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
NeurIPS4
2019 A hierarchical sparse coding model predicts acoustic feature encoding in both auditory midbrain and cortex
abstract
The auditory pathway consists of multiple stages, from the cochlear nucleus to the auditory cortex. Neurons acting at different stages have different functions and exhibit different response properties. It is unclear whether these stages share a common encoding mechanism. We trained an unsupervised deep learning model consisting of alternating sparse coding and max pooling layers on cochleogram-filtered human speech. Evaluation of the response properties revealed that computing units in lower layers exhibited spectro-temporal receptive fields (STRFs) similar to those of inferior colliculus neurons measured in physiological experiments, including properties such as sound onset and termination, checkerboard pattern, and spectral motion. Units in upper layers tended to be tuned to phonetic features such as plosivity and nasality, resembling the results of field recording in human auditory cortex. Variation of the sparseness level of the units in each higher layer revealed a positive correlation between the sparseness level and the strength of phonetic feature encoding. The activities of the units in the top layer, but not other layers, correlated with the dynamics of the first two formants (F1, F2) of all phonemes, indicating the encoding of phoneme dynamics in these units. These results suggest that the principles of sparse coding and max pooling may be universal in the human auditory pathway.
Qingtian Zhang, Xiaolin Hu 0001, Bo Zhang 0010
PLoS Comput. Biol.4
2019 Semantic Cluster Unary Loss for Efficient Deep Hashing
abstract
Hashing method maps similar data to binary hashcodes with smaller hamming distance, which has received broad attention due to its low storage cost and fast retrieval speed. With the rapid development of deep learning, deep hashing methods have achieved promising results in efficient information retrieval. Most existing deep hashing methods adopt pairwise or triplet losses to deal with similarities underlying the data, but their training are difficult and less efficient because O(n2) data pairs and O(n3) triplets are involved. To address these issues, we propose a novel deep hashing algorithm with unary loss which can be trained very efficiently. First of all, we introduce a Unary Upper Bound of the traditional triplet loss, thus reducing the complexity to O(n) and bridging the classificationbased unary loss and the triplet loss. Second, we propose a novel Semantic Cluster Deep Hashing (SCDH) algorithm by introducing a modified Unary Upper Bound loss, named Semantic Cluster Unary Loss (SCUL). The resultant hashcodes form several compact clusters, which means hashcodes in the same cluster have similar semantic information. We also demonstrate that the proposed SCDH is easy to be extended to semi-supervised settings by incorporating the state-of-the-art semi-supervised learning algorithms. Experiments on large-scale datasets show that the proposed method is superior to state-of-the-art hashing algorithms.
Jianmin Li 0001, Bo Zhang 0010
IEEE Trans. Image Process.3
2018 Collaborative Filtering With User-Item Co-Autoregressive Models
abstract
Deep neural networks have shown promise in collaborative filtering (CF). However, existing neural approaches are either user-based or item-based, which cannot leverage all the underlying information explicitly. We propose CF-UIcA, a neural co-autoregressive model for CF tasks, which exploits the structural correlation in the domains of both users and items. The co-autoregression allows extra desired properties to be incorporated for different tasks. Furthermore, we develop an efficient stochastic learning algorithm to handle large scale datasets. We evaluate CF-UIcA on two popular benchmarks: MovieLens 1M and Netflix, and achieve state-of-the-art performance in both rating prediction and top-N recommendation tasks, which demonstrates the effectiveness of CF-UIcA.
Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
AAAI5
2018 Textbook Question Answering Under Instructor Guidance With Memory Networks
abstract
Textbook Question Answering (TQA) is a task to choose the most proper answers by reading a multi-modal context of abundant essays and images. TQA serves as a favorable test bed for visual and textual reasoning. However, most of the current methods are incapable of reasoning over the long contexts and images. To address this issue, we propose a novel approach of Instructor Guidance with Memory Networks (IGMN) which conducts the TQA task by finding contradictions between the candidate answers and their corresponding context. We build the Contradiction Entity-Relationship Graph (CERG) to extend the passage-level multi-modal contradictions to an essay level. The machine thus performs as an instructor to extract the essay-level contradictions as the Guidance. Afterwards, we exploit the memory networks to capture the information in the Guidance, and use the attention mechanisms to jointly reason over the global features of the multi-modal input. Extensive experiments demonstrate that our method outperforms the state-of-the-arts on the TQA dataset. The source code is available at https://github.com/freerailway/igmn.
Juzheng Li, Hang Su 0006, Jun Zhu 0001, Bo Zhang 0010
CVPR5
2018 Smooth Neighbors on Teacher Graphs for Semi-Supervised Learning
abstract
The recently proposed self-ensembling methods have achieved promising results in deep semi-supervised learning, which penalize inconsistent predictions of unlabeled data under different perturbations. However, they only consider adding perturbations to each single data point, while ignoring the connections between data samples. In this paper, we propose a novel method, called Smooth Neighbors on Teacher Graphs (SNTG). In SNTG, a graph is constructed based on the predictions of the teacher model, i.e., the implicit self-ensemble of models. Then the graph serves as a similarity measure with respect to which the representations of "similar" neighboring points are learned to be smooth on the low-dimensional manifold. We achieve state-of-the-art results on semi-supervised learning benchmarks. The error rates are 9.89%, 3.99% for CIFAR-10 with 4000 labels, SVHN with 500 labels, respectively. In particular, the improvements are significant when the labels are fewer. For the non-augmented MNIST with only 20 labels, the error rate is reduced from previous 4.81% to 1.36%. Our method also shows robustness to noisy labels.
Yucen Luo, Jun Zhu 0001, Mengxi Li, Bo Zhang 0010
CVPR5
2018 Interpret Neural Networks by Identifying Critical Data Routing Paths
abstract
Interpretability of a deep neural network aims to explain the rationale behind its decisions and enable the users to understand the intelligent agents, which has become an important issue due to its importance in practical applications. To address this issue, we develop a Distillation Guided Routing method, which is a flexible framework to interpret a deep neural network by identifying critical data routing paths and analyzing the functional processing behavior of the corresponding layers. Specifically, we propose to discover the critical nodes on the data routing paths during network inferring prediction for individual input samples by learning associated control gates for each layer's output channel. The routing paths can, therefore, be represented based on the responses of concatenated control gates from all the layers, which reflect the network's semantic selectivity regarding to the input patterns and more detailed functional process across different layer levels. Based on the discoveries, we propose an adversarial sample detection algorithm by learning a classifier to discriminate whether the critical data routing paths are from real or adversarial samples. Experiments demonstrate that our algorithm can effectively achieve high defense rate with minor training overhead.
Hang Su 0006, Bo Zhang 0010, Xiaolin Hu 0001
CVPR3
2018 Essay-Anchor Attentive Multi-Modal Bilinear Pooling for Textbook Question Answering
abstract
Textbook Question Answering (TQA) [1] is a newly proposed task to answer arbitrary questions in middle school curricula, which has particular challenges to understand the long essays in additional to the images. Bilinear models [2], [3] are effective at learning high-level associations between questions and images, but are inefficient to handle the long essays. In this paper, we propose an Essay-anchor Attentive Multi-modal Bilinear pooling (EAMB), a novel method to encode the long essays into the joint space of the questions and images. The essay-anchors, embedded from the keywords, represent the essay information in a latent space. We propose a novel network architecture to pay special attention on the keywords in the questions, consequently encoding the essay information into the question features, and thus the joint space with the images. We then use the bilinear models to extract the multi-modal interactions to obtain the answers. EAMB successfully utilizes the redundancy of the pre-trained word embedding space to represent the essay-anchors. This avoids the extra learning difficulties from exploiting large network structures. Quantitative and qualitative experiments show the outperforming effects of EAMB on the TQA dataset.
Juzheng Li, Hang Su 0006, Jun Zhu 0001, Bo Zhang 0010
ICME4
2018 Message Passing Stein Variational Gradient Descent
abstract
Stein variational gradient descent (SVGD) is a recently proposed particle-based Bayesian inference method, which has attracted a lot of interest due to its remarkable approximation ability and particle efficiency compared to traditional variational inference and Markov Chain Monte Carlo methods. However, we observed that particles of SVGD tend to collapse to modes of the target distribution, and this particle degeneracy phenomenon becomes more severe with higher dimensions. Our theoretical analysis finds out that there exists a negative correlation between the dimensionality and the repulsive force of SVGD which should be blamed for this phenomenon. We propose Message Passing SVGD (MP-SVGD) to solve this problem. By leveraging the conditional independence structure of probabilistic graphical models (PGMs), MP-SVGD converts the original high-dimensional global inference problem into a set of local ones over the Markov blanket with lower dimensions. Experimental results show its advantages of preventing vanishing repulsive force in high-dimensional space over SVGD, and its particle efficiency and approximation flexibility over other inference methods on graphical models.
Jingwei Zhuo, Chang Liu 0030, Jiaxin Shi, Jun Zhu 0001, Ning Chen 0002, Bo Zhang 0010
ICML6
2018 Graphical Generative Adversarial Networks
abstract
We propose Graphical Generative Adversarial Networks (Graphical-GAN) to model structured data. Graphical-GAN conjoins the power of Bayesian networks on compactly representing the dependency structures among random variables and that of generative adversarial networks on learning expressive dependency functions. We introduce a structured recognition model to infer the posterior distribution of latent variables given observations. We generalize the Expectation Propagation (EP) algorithm to learn the generative model and recognition model jointly. Finally, we present two important instances of Graphical-GAN, i.e. Gaussian Mixture GAN (GMGAN) and State Space GAN (SSGAN), which can successfully learn the discrete and temporal structures on visual datasets, respectively.
Chongxuan Li, Max Welling, Jun Zhu 0001, Bo Zhang 0010
NeurIPS4
2018 Semi-crowdsourced Clustering with Deep Generative Models
abstract
We consider the semi-supervised clustering problem where crowdsourcing provides noisy information about the pairwise comparisons on a small subset of data, i.e., whether a sample pair is in the same cluster. We propose a new approach that includes a deep generative model (DGM) to characterize low-level features of the data, and a statistical relational model for noisy pairwise annotations on its subset. The two parts share the latent variables. To make the model automatically trade-off between its complexity and fitting data, we also develop its fully Bayesian variant. The challenge of inference is addressed by fast (natural-gradient) stochastic variational inference algorithms, where we effectively combine variational message passing for the relational part and amortized learning of the DGM under a unified framework. Empirical results on synthetic and real-world datasets show that our model outperforms previous crowdsourced clustering methods.
Yucen Luo, Tian Tian 0001, Jiaxin Shi, Jun Zhu 0001, Bo Zhang 0010
NeurIPS5
2018 Visual instance mining from the graph perspective
Wei Li 0152, Jianmin Li 0001, Changhu Wang, Lei Zhang 0001, Bo Zhang 0010
Multim. Syst.5
2018 Max-Margin Deep Generative Models for (Semi-)Supervised Learning
abstract
Deep generative models (DGMs) can effectively capture the underlying distributions of complex data by learning multilayered representations and performing inference. However, it is relatively insufficient to boost the discriminative ability of DGMs. This paper presents max-margin deep generative models (mmDGMs) and a class-conditional variant (mmDCGMs), which explore the strongly discriminative principle of max-margin learning to improve the predictive performance of DGMs in both supervised and semi-supervised learning, while retaining the generative capability. In semi-supervised learning, we use the predictions of a max-margin classifier as the missing labels instead of performing full posterior inference for efficiency; we also introduce additional max-margin and label-balance regularization terms of unlabeled data for effectiveness. We develop an efficient doubly stochastic subgradient algorithm for the piecewise linear objectives in different settings. Empirical results on various datasets demonstrate that: (1) max-margin learning can significantly improve the prediction performance of DGMs and meanwhile retain the generative ability; (2) in supervised learning, mmDGMs are competitive to the best fully discriminative networks when employing convolutional neural networks as the generative and recognition models; and (3) in semi-supervised learning, mmDCGMs can perform efficient inference and achieve state-of-the-art classification results on several benchmarks.
Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Scalable Discrete Supervised Multimedia Hash Learning With Clustering
abstract
The hashing method maps similar data of various types to binary hashcodes with smaller hamming distance, and it has received broad attention due to its low-storage cost and fast retrieval speed. However, the existing limitations make the present algorithms difficult to deal with for large-scale data sets: 1) discrete constraints are involved in the learning of the hash function and 2) pairwise or triplet similarity is adopted to generate efficient hashcodes, resulting in both time and space complexity greater than O(n2). To address these issues, we propose a novel discrete supervised hash learning framework that can be scalable to large-scale data sets of various types. First, the discrete learning procedure is decomposed into a binary classifier learning scheme and binary codes learning scheme, which makes the learning procedure more efficient. Second, by adopting the asymmetric low-rank matrix factorization, we propose the fast clustering-based batch coordinate descent method, such that the time and space complexity are reduced to O(n). The proposed framework also provides a flexible paradigm to incorporate with arbitrary hash function, including deep neural networks and kernel methods, as well as any types of data to hash, including images and videos. Experiments on large-scale data sets demonstrate that the proposed method is superior or comparable with the state-of-the-art hashing algorithms.
Jianmin Li 0001, Mengqing Jiang, Peijiang Yuan, Bo Zhang 0010
IEEE Trans. Circuits Syst. Video Technol.5
2018 Learning Deep Generative Models With Doubly Stochastic Gradient MCMC
abstract
Deep generative models (DGMs), which are often organized in a hierarchical manner, provide a principled framework of capturing the underlying causal factors of data. Recent work on DGMs focussed on the development of efficient and scalable variational inference methods that learn a single model under some mean-field or parameterization assumptions. However, little work has been done on extending Markov chain Monte Carlo (MCMC) methods to Bayesian DGMs, which enjoy many advantages compared with variational methods. We present doubly stochastic gradient MCMC, a simple and generic method for (approximate) Bayesian inference of DGMs in a collapsed continuous parameter space. At each MCMC sampling step, the algorithm randomly draws a mini-batch of data samples to estimate the gradient of log-posterior and further estimates the intractable expectation over hidden variables via a neural adaptive importance sampler, where the proposal distribution is parameterized by a deep neural network and learnt jointly along with the sampling process. We demonstrate the effectiveness of learning various DGMs on a wide range of tasks, including density estimation, data generation, and missing data imputation. Our method outperforms many state-of-the-art competitors.
Jun Zhu 0001, Bo Zhang 0010
IEEE Trans. Neural Networks Learn. Syst.3
2017 Improving Interpretability of Deep Neural Networks with Semantic Information
abstract
Interpretability of deep neural networks (DNNs) is essential since it enables users to understand the overall strengths and weaknesses of the models, conveys an understanding of how the models will behave in the future, and how to diagnose and correct potential problems. However, it is challenging to reason about what a DNN actually does due to its opaque or black-box nature. To address this issue, we propose a novel technique to improve the interpretability of DNNs by leveraging the rich semantic information embedded in human descriptions. By concentrating on the video captioning task, we first extract a set of semantically meaningful topics from the human descriptions that cover a wide range of visual concepts, and integrate them into the model with an interpretive loss. We then propose a prediction difference maximization algorithm to interpret the learned features of each neuron. Experimental results demonstrate its effectiveness in video captioning using the interpretable features, which can also be transferred to video action recognition. By clearly understanding the learned features, users can easily revise false predictions via a human-in-the-loop procedure.
Yinpeng Dong, Hang Su 0006, Jun Zhu 0001, Bo Zhang 0010
CVPR4
2017 Semi-supervised Max-margin Topic Model with Manifold Posterior Regularization
abstract
Supervised topic models leverage label information to learn discriminative latent topic representations. As collecting a fully labeled dataset is often time-consuming, semi-supervised learning is of high interest. In this paper, we present an effective semi-supervised max-margin topic model by naturally introducing manifold posterior regularization to a regularized Bayesian topic model, named LapMedLDA. The model jointly learns latent topics and a related classifier with only a small fraction of labeled documents. To perform the approximate inference, we derive an efficient stochastic gradient MCMC method. Unlike the previous semi-supervised topic models, our model adopts a tight coupling between the generative topic model and the discriminative classifier. Extensive experiments demonstrate that such tight coupling brings significant benefits in quantitative and qualitative performance.
Wenbo Hu 0001, Jun Zhu 0001, Hang Su 0006, Jingwei Zhuo, Bo Zhang 0010
IJCAI5
2017 Forecast the Plausible Paths in Crowd Scenes
abstract
Forecasting the future plausible paths of pedestrians in crowd scenes is of wide applications, but it still remains as a challenging task due to the complexities and uncertainties of crowd motions. To address these issues, we propose to explore the inherent crowd dynamics via a social-aware recurrent Gaussian process model, which facilitates the path prediction by taking advantages of the interplay between the rich prior knowledge and motion uncertainties. Specifically, we derive a social-aware LSTM to explore the crowd dynamic, resulting in a hidden feature embedding the rich prior in massive data. Afterwards, we integrate the descriptor into deep Gaussian processes with motion uncertainties appropriately harnessed. Crowd motion forecasting is implemented by regressing relative motion against the current positions, yielding the predicted paths based on a functional object associated with a distribution. Extensive experiments on public datasets demonstrate that our method obtains the state-of-the-art performance in both structured and unstructured scenes by exploring the complex and uncertain motion patterns, even if the occlusion is serious or the observed trajectories are noisy.
Hang Su 0006, Jun Zhu 0001, Yinpeng Dong, Bo Zhang 0010
IJCAI4
2017 Triple Generative Adversarial Nets
abstract
Generative Adversarial Nets (GANs) have shown promise in image generation and semi-supervised learning (SSL). However, existing GANs in SSL have two problems: (1) the generator and the discriminator (i.e. the classifier) may not be optimal at the same time; and (2) the generator cannot control the semantics of the generated samples. The problems essentially arise from the two-player formulation, where a single discriminator shares incompatible roles of identifying fake samples and predicting labels and it only estimates the data without considering the labels. To address the problems, we present triple generative adversarial net (Triple-GAN), which consists of three players---a generator, a discriminator and a classifier. The generator and the classifier characterize the conditional distributions between images and labels, and the discriminator solely focuses on identifying fake image-label pairs. We design compatible utilities to ensure that the distributions characterized by the classifier and the generator both converge to the data distribution. Our results on various datasets demonstrate that Triple-GAN as a unified model can simultaneously (1) achieve the state-of-the-art classification results among deep generative models, and (2) disentangle the classes and styles of the input and transfer smoothly in the data space via interpolation in the latent space class-conditionally.
Chongxuan Li, Taufik Xu, Jun Zhu 0001, Bo Zhang 0010
NIPS4
2017 Fast sampling methods for Bayesian max-margin models
Wenbo Hu 0001, Jun Zhu 0001, Bo Zhang 0010
Expert Syst. Appl.3
2017 Towards Reversal-Invariant Image Representation
Lingxi Xie, Jingdong Wang 0001, Weiyao Lin, Bo Zhang 0010, Qi Tian 0001
Int. J. Comput. Vis.4
2016 Discriminative Nonparametric Latent Feature Relational Models with Data Augmentation
abstract
We present a discriminative nonparametric latent feature relational model (LFRM) for link prediction to automatically infer the dimensionality of latent features. Under the generic RegBayes (regularized Bayesian inference) framework, we handily incorporate the prediction loss with probabilistic inference of a Bayesian model; set distinct regularization parameters for different types of links to handle the imbalance issue in real networks; and unify the analysis of both the smooth logistic log-loss and the piecewise linear hinge loss. For the nonconjugate posterior inference, we present a simple Gibbs sampler via data augmentation, without making restricting assumptions as done in variational methods. We further develop an approximate sampler using stochastic gradient Langevin dynamics to handle large networks with hundreds of thousands of entities and millions of links, orders of magnitude larger than what existing LFRM models can process. Extensive studies on various real networks show promising performance.
Ning Chen 0002, Jun Zhu 0001, Jiaming Song, Bo Zhang 0010
AAAI5
2016 Jointly Modeling Topics and Intents with Global Order Structure
abstract
Modeling document structure is of great importance for discourse analysis and related applications. The goal of this research is to capture the document intent structure by modeling documents as a mixture of topic words and rhetorical words. While the topics are relatively unchanged through one document, the rhetorical functions of sentences usually change following certain orders in discourse. We propose GMM-LDA, a topic modeling based Bayesian unsupervised model, to analyze the document intent structure cooperated with order information. Our model is flexible that has the ability to combine the annotations and do supervised learning. Additionally, entropic regularization can be introduced to model the significant divergence between topics and intents. We perform experiments in both unsupervised and supervised settings, results show the superiority of our model over several state-of-the-art baselines.
Jun Zhu 0001, Nan Yang 0002, Tian Tian 0001, Ming Zhou 0001, Bo Zhang 0010
AAAI6
2016 Discriminative Deep Random Walk for Network Classification
abstract
Deep Random Walk (DeepWalk) can learn a latent space representation for describing the topological structure of a network.However, for relational network classification, DeepWalk can be suboptimal as it lacks a mechanism to optimize the objective of the target task.In this paper, we present Discriminative Deep Random Walk (DDRW), a novel method for relational network classification.By solving a joint optimization problem, DDRW can learn the latent space representations that well capture the topological structure and meanwhile are discriminative for the network classification task.Our experimental results on several real social networks demonstrate that DDRW significantly outperforms DeepWalk on multilabel network classification tasks, while retaining the topological structure in the latent space.DDRW is stable and consistently outperforms the baseline methods by various percentages of labeled data.DDRW is also an online method that is scalable and can be naturally parallelized.
Juzheng Li, Jun Zhu 0001, Bo Zhang 0010
ACL (1)3
2016 Segment-Level Sequence Modeling using Gated Recursive Semi-Markov Conditional Random Fields
abstract
Most of the sequence tagging tasks in natural language processing require to recognize segments with certain syntactic role or semantic meaning in a sentence.They are usually tackled with Conditional Random Fields (CRFs), which do indirect word-level modeling over word-level features and thus cannot make full use of segment-level information.Semi-Markov Conditional Random Fields (Semi-CRFs) model segments directly but extracting segment-level features for Semi-CRFs is still a very challenging problem.This paper presents Gated Recursive Semi-CRFs (grSemi-CRFs), which model segments directly and automatically learn segmentlevel features through a gated recursive convolutional neural network.Our experiments on text chunking and named entity recognition (NER) demonstrate that grSemi-CRFs generally outperform other neural models.
Jingwei Zhuo, Jun Zhu 0001, Bo Zhang 0010, Zaiqing Nie
ACL (1)4
2016 Efficient and Robust Semi-supervised Learning Over a Sparse-Regularized Graph
Hang Su 0006, Jun Zhu 0001, Zhaozheng Yin, Yinpeng Dong, Bo Zhang 0010
ECCV (8)5
2016 Scalable Discrete Supervised Hash Learning with Asymmetric Matrix Factorization
abstract
Hashing methods map similar data to binary hashcodes with smaller hamming distance, and it has received a broad attention due to its low storage cost and fast retrieval speed. However, the existing limitations make the present algorithms difficult to deal with large-scale datasets: (1) discrete constraints are involved in the learning of the hash function, (2) pairwise or triplet similarity is adopted to generate efficient hashcodes, resulting both time and space complexity are greater than O(n2). To address these issues, we propose a novel discrete supervised hash learning framework which can be scalable to large-scale datasets. First, the learning procedure is decomposed into a binary classifier learning scheme and hashcodes learning scheme. Second, we adopt the Asymmetric Low-rank Matrix Factorization and propose the Fast Clustering-based Batch Coordinate Descent method, such that the time and space complexity is reduced to O(n). The proposed framework also provides a flexible paradigm to incorporate with arbitrary hash function, including deep neural networks. Experiments on large-scale datasets demonstrate that the proposed method is superior or comparable with state-of-the-art hashing algorithms.
Jianmin Li 0001, Jinma Guo, Bo Zhang 0010
ICDM4
2016 Learning to Generate with Memory
abstract
Memory units have been widely used to enrich the capabilities of deep networks on capturing long-term dependencies in reasoning and prediction tasks, but little investigation exists on deep generative models (DGMs) which are good at inferring high-level invariant representations from unlabeled data. This paper presents a deep generative model with a possibly large external memory and an attention mechanism to capture the local detail information that is often lost in the bottom-up abstraction process in representation learning. By adopting a smooth attention model, the whole network is trained end-to-end by optimizing a variational bound of data likelihood via auto-encoding variational Bayesian methods, where an asymmetric recognition network is learnt jointly to infer high-level invariant representations. The asymmetric architecture can reduce the competition between bottom-up invariant feature extraction and top-down generation of instance details. Our experiments on several datasets demonstrate that memory can significantly boost the performance of DGMs on various tasks, including density estimation, image generation, and missing value imputation, and DGMs with memory can achieve state-of-the-art quantitative results.
Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010
ICML3
2016 Crowd Scene Understanding with Coherent Recurrent Neural Networks
Hang Su 0006, Yinpeng Dong, Jun Zhu 0001, Haibin Ling, Bo Zhang 0010
IJCAI5
2016 Incorporating visual adjectives for image classification
Lingxi Xie, Jingdong Wang 0001, Bo Zhang 0010, Qi Tian 0001
Neurocomputing3
2016 BitHash: An efficient bitwise Locality Sensitive Hashing method with applications
Wenhao Zhang 0003, Jianqiu Ji, Jun Zhu 0001, Jianmin Li 0001, Hua Xu 0003, Bo Zhang 0010
Knowl. Based Syst.6
2016 Simple Techniques Make Sense: Feature Pooling and Normalization for Image Classification
abstract
Image classification is a fundamental task in computer vision, implying a wide range of challenging problems, such as object recognition, scene understanding, and image tagging. One of the most popular approaches to image classification, the bag-of-features (BoF) model, represents an image with a long feature vector and adopts machine learning algorithms for training and testing. Owing to its simplicity and scalability, the BoF model is widely used in both academic research studies and industrial applications. This paper discusses the feature summarization stage, including pooling and normalization, in the BoF model. We show that these two modules, although devalued sometimes, have important impacts on image classification performance. We present two algorithms, i.e., generalized regular spatial pooling for constructing a better group of spatial bins and hierarchical feature normalization for assigning proper weights for regional feature normalization. Both algorithms are independent of the descriptor extraction and feature encoding stages, and therefore, they could be freely transplanted onto many other classification frameworks based on local feature statistics. We further provide insightful discussions for the nature of designing efficient image classification models. Experiments verify that the proposed algorithm achieves state-of-the-art results on a wide range of image classification data sets.
Lingxi Xie, Qi Tian 0001, Bo Zhang 0010
IEEE Trans. Circuits Syst. Video Technol.3
2016 Balancing Convergence and Diversity in Decomposition-Based Many-Objective Optimizers
abstract
The decomposition-based multiobjective evolutionary algorithms (MOEAs) generally make use of aggregation functions to decompose a multiobjective optimization problem into multiple single-objective optimization problems. However, due to the nature of contour lines for the adopted aggregation functions, they usually fail to preserve the diversity in high-dimensional objective space even by using diverse weight vectors. To address this problem, we propose to maintain the desired diversity of solutions in their evolutionary process explicitly by exploiting the perpendicular distance from the solution to the weight vector in the objective space, which achieves better balance between convergence and diversity in many-objective optimization. The idea is implemented to enhance two well-performing decomposition-based algorithms, i.e., MOEA, based on decomposition and ensemble fitness ranking. The two enhanced algorithms are compared to several state-of-the-art algorithms and a series of comparative experiments are conducted on a number of test problems from two well-known test suites. The experimental results show that the two proposed algorithms are generally more effective than their predecessors in balancing convergence and diversity, and they are also very competitive against other existing algorithms for solving many-objective optimization problems.
Yuan Yuan 0004, Hua Xu 0003, Bo Wang 0051, Bo Zhang 0010, Xin Yao 0001
IEEE Trans. Evol. Comput.4
2015 Scalable Visual Instance Mining with Instance Graph
abstract
In this paper we address the problem of visual instance mining, which is to automatically discover frequently appearing visual instances from a large collection of images. We propose a scalable mining method by leveraging the graph structure with images as vertices. Different from most existing work that focused on either instance-level similarities or image-level context properties, our graph captures both information. The instance-level information is integrated during the construction of a weighted and undirected instance graph based on the similarity between augmented local features, while the image-level context is explored with a greedy breadth-first search algorithm to discover clusters of visual instances from the graph. This method is capable of mining challenging small visual instances with diverse variations. We evaluated our method on two fully annotated datasets and outperformed the state of the arts on both datasets with higher recalls. We also applied our method on a one-million Flickr dataset and proved its scalability.
Wei Li 0152, Changhu Wang, Lei Zhang 0001, Yong Rui, Bo Zhang 0010
BMVC5
2015 RIDE: Reversal Invariant Descriptor Enhancement
abstract
In many fine-grained object recognition datasets, image orientation (left/right) might vary from sample to sample. Since handcrafted descriptors such as SIFT are not reversal invariant, the stability of image representation based on them is consequently limited. A popular solution is to augment the datasets by adding a left-right reversed copy for each original image. This strategy improves recognition accuracy to some extent, but also brings the price of almost doubled time and memory consumptions. In this paper, we present RIDE (Reversal Invariant Descriptor Enhancement) for fine-grained object recognition. RIDE is a generalized algorithm which cancels out the impact of image reversal by estimating the orientation of local descriptors, and guarantees to produce the identical representation for an image and its left-right reversed copy. Experimental results reveal the consistent accuracy gain of RIDE with various types of descriptors. We also provide insightful discussions on the working mechanism of RIDE and its generalization to other applications.
Lingxi Xie, Jingdong Wang 0001, Weiyao Lin, Bo Zhang 0010, Qi Tian 0001
ICCV4
2015 Adaptive Dropout Rates for Learning with Corrupted Features
Jingwei Zhuo, Jun Zhu 0001, Bo Zhang 0010
IJCAI3
2015 Regularizing neural networks with adaptive local drop
abstract
Neural network (NN) models have shown good performance on many image recognition benchmarks. Given large image datasets, these models typically have millions or billions of parameters that can easily lead to over-fitting without regularization. Dropout and DropConnect show their effectiveness of regularizing large fully connected layers within neural networks. In Dropout, each neural activation within the network is randomly set to zero with a probability during training. In DropConnect, a generalization of Dropout, each connection weight within the network is randomly set to zero with a probability instead. Both of the probabilities in Dropout and DropConnect are universal predefined constants. We propose Adaptive Local Drop (ALDrop), a novel regularization method that sets each connection weight within the network with a learned probability adaptive to the input image dataset using a locality-based measure. Experiments on several image recognition benchmarks show that our model outperforms Dropout and DropConnect.
Binbin Cao, Jianmin Li 0001, Bo Zhang 0010
IJCNN3
2015 Interlinked Convolutional Neural Networks for Face Parsing
abstract
Face parsing is a basic task in face image analysis. It amounts to labeling each pixel with appropriate facial parts such as eyes and nose. In the paper, we present a interlinked convolutional neural network (iCNN) for solving this problem in an end-to-end fashion. It consists of multiple convolutional neural networks (CNNs) taking input in different scales. A special interlinking layer is designed to allow the CNNs to exchange information, enabling them to integrate local and contextual information efficiently. The hallmark of iCNN is the extensive use of downsampling and upsampling in the interlinking layers, while traditional CNNs usually uses downsampling only. A two-stage pipeline is proposed for face parsing and both stages use iCNN. The first stage localizes facial parts in the size-reduced image and the second stage labels the pixels in the identified facial parts in the original image. On a benchmark dataset we have obtained better results than the state-of-the-art methods.
Yisu Zhou, Xiaolin Hu 0001, Bo Zhang 0010
ISNN3
2015 Image Classification and Retrieval are ONE
abstract
In this paper, we demonstrate that the essentials of image classification and retrieval are the same, since both tasks could be tackled by measuring the similarity between images. To this end, we propose ONE (Online Nearest-neighbor Estimation), a unified algorithm for both image classification and retrieval. ONE is surprisingly simple, which only involves manual object definition, regional description and nearest-neighbor search. We take advantage of PCA and PQ approximation and GPU parallelization to scale our algorithm up to large-scale image search. Experimental results verify that ONE achieves state-of-the-art accuracy in a wide range of image classification and retrieval benchmarks.
Lingxi Xie, Richang Hong, Bo Zhang 0010, Qi Tian 0001
ICMR3
2015 Max-Margin Deep Generative Models
abstract
Deep generative models (DGMs) are effective on learning multilayered representations of complex data and performing inference of input data by exploring the generative ability. However, little work has been done on examining or empowering the discriminative ability of DGMs on making accurate predictions. This paper presents max-margin deep generative models (mmDGMs), which explore the strongly discriminative principle of max-margin learning to improve the discriminative power of DGMs, while retaining the generative capability. We develop an efficient doubly stochastic subgradient algorithm for the piecewise linear objective. Empirical results on MNIST and SVHN datasets demonstrate that (1) max-margin learning can significantly improve the prediction performance of DGMs and meanwhile retain the generative ability; and (2) mmDGMs are competitive to the state-of-the-art fully discriminative networks by employing deep convolutional neural networks (CNNs) as both recognition and generative models.
Chongxuan Li, Jun Zhu 0001, Tianlin Shi, Bo Zhang 0010
NIPS4
2015 Convolutional Neural Networks with Intra-Layer Recurrent Connections for Scene Labeling
abstract
Scene labeling is a challenging computer vision task. It requires the use of both local discriminative features and global context information. We adopt a deep recurrent convolutional neural network (RCNN) for this task, which is originally proposed for object recognition. Different from traditional convolutional neural networks (CNN), this model has intra-layer recurrent connections in the convolutional layers. Therefore each convolutional layer becomes a two-dimensional recurrent neural network. The units receive constant feed-forward inputs from the previous layer and recurrent inputs from their neighborhoods. While recurrent iterations proceed, the region of context captured by each unit expands. In this way, feature extraction and context modulation are seamlessly integrated, which is different from typical methods that entail separate modules for the two steps. To further utilize the context, a multi-scale RCNN is proposed. Over two benchmark datasets, Standford Background and Sift Flow, the model outperforms many state-of-the-art models in accuracy and efficiency.
Xiaolin Hu 0001, Bo Zhang 0010
NIPS3
2015 Discriminative Relational Topic Models
abstract
Relational topic models (RTMs) provide a probabilistic generative process to describe both the link structure and document contents for document networks, and they have shown promise on predicting network structures and discovering latent topic representations. However, existing RTMs have limitations in both the restricted model expressiveness and incapability of dealing with imbalanced network data. To expand the scope and improve the inference accuracy of RTMs, this paper presents three extensions: 1) unlike the common link likelihood with a diagonal weight matrix that allows the-same-topic interactions only, we generalize it to use a full weight matrix that captures all pairwise topic interactions and is applicable to asymmetric networks; 2) instead of doing standard Bayesian inference, we perform regularized Bayesian inference (RegBayes) with a regularization parameter to deal with the imbalanced link structure issue in real networks and improve the discriminative ability of learned latent representations; and 3) instead of doing variational approximation with strict mean-field assumptions, we present collapsed Gibbs sampling algorithms for the generalized relational topic models by exploring data augmentation without making restricting assumptions. Under the generic RegBayes framework, we carefully investigate two popular discriminative loss functions, namely, the logistic log-loss and the max-margin hinge loss. Experimental results on several real network datasets demonstrate the significance of these extensions on improving prediction performance.
Ning Chen 0002, Jun Zhu 0001, Fei Xia 0005, Bo Zhang 0010
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 Angular-Similarity-Preserving Binary Signatures for Linear Subspaces
abstract
We propose a similarity-preserving binary signature method for linear subspaces. In computer vision and pattern recognition, linear subspace is a very important representation for many kinds of data, such as face images, action and gesture videos, and so on. When there is a large amount of subspace data and the ambient dimension is high, the cost of computing the pairwise similarity between the subspaces would be high and it requires a large storage space for storing the subspaces. In this paper, we first define the angular similarity and angular distance between the subspaces. Then, based on this similarity definition, we develop a similarity-preserving binary signature method for linear subspaces, which transforms a linear subspace into a compact binary signature, and the Hamming distance between two signatures provides an unbiased estimate of the angular similarity between the two subspaces. We also provide a lower bound of the signature length sufficient to guarantee uniform distance-preservation between every pair of subspaces in a set. Experiments on face recognition, gesture recognition, and action recognition verify the effectiveness of the proposed method.
Jianqiu Ji, Jianmin Li 0001, Qi Tian 0001, Shuicheng Yan, Bo Zhang 0010
IEEE Trans. Image Process.5
2015 Heterogeneous Graph Propagation for Large-Scale Web Image Search
abstract
State-of-the-art web image search frameworks are often based on the bag-of-visual-words (BoVWs) model and the inverted index structure. Despite the simplicity, efficiency, and scalability, they often suffer from low precision and/or recall, due to the limited stability of local features and the considerable information loss on the quantization stage. To refine the quality of retrieved images, various postprocessing methods have been adopted after the initial search process. In this paper, we investigate the online querying process from a graph-based perspective. We introduce a heterogeneous graph model containing both image and feature nodes explicitly, and propose an efficient reranking approach consisting of two successive modules, i.e., incremental query expansion and image-feature voting, to improve the recall and precision, respectively. Compared with the conventional reranking algorithms, our method does not require using geometric information of visual words, therefore enjoys low consumptions of both time and memory. Moreover, our method is independent of the initial search process, and could cooperate with many BoVW-based image search pipelines, or adopted after other postprocessing algorithms. We evaluate our approach on large-scale image search tasks and verify its competitive search performance.
Lingxi Xie, Qi Tian 0001, Wengang Zhou 0001, Bo Zhang 0010
IEEE Trans. Image Process.4
2015 Disease Inference from Health-Related Questions via Sparse Deep Learning
abstract
Automatic disease inference is of importance to bridge the gap between what online health seekers with unusual symptoms need and what busy human doctors with biased expertise can offer. However, accurately and efficiently inferring diseases is non-trivial, especially for community-based health services due to the vocabulary gap, incomplete information, correlated medical concepts, and limited high quality training samples. In this paper, we first report a user study on the information needs of health seekers in terms of questions and then select those that ask for possible diseases of their manifested symptoms for further analytic. We next propose a novel deep learning scheme to infer the possible diseases given the questions of health seekers. The proposed scheme is comprised of two key components. The first globally mines the discriminant medical signatures from raw features. The second deems the raw features and their signatures as input nodes in one layer and hidden nodes in the subsequent layer, respectively. Meanwhile, it learns the inter-relations between these two layers via pre-training with pseudo-labeled data. Following that, the hidden nodes serve as raw features for the more abstract signature mining. With incremental and alternative repeating of these two components, our scheme builds a sparsely connected deep architecture with three hidden layers. Overall, it well fits specific tasks with fine-tuning. Extensive experiments on a real-world dataset labeled by online doctors show the significant performance gains of our scheme.
Liqiang Nie, Meng Wang 0001, Shuicheng Yan, Bo Zhang 0010, Tat-Seng Chua
IEEE Trans. Knowl. Data Eng.5
2015 Partial-Duplicate Clustering and Visual Pattern Discovery on Web Scale Image Database
abstract
In this paper, we study the problem of discovering visual patterns and partial-duplicate images, which is fundamental to visual concept representation and image parsing, but very challenging when the database is extremely large, such as billions of images indexed by a commercial search engine. Although extensive research with sophisticated algorithms has been conducted for either partial-duplicate clustering or visual pattern discovery, most of them can not be easily extended to this scale, since both are clustering problems in nature and require pairwise comparisons. To tackle this computational challenge, we introduce a novel and highly parallelizable framework to discover partial-duplicate images and visual patterns in a unified way in distributed computing systems. We emphasize the nested property of local features, and propose the generalized nested feature (GNF) as a mid-level representation for regions and local patterns. Initial coarse clusters are then discovered by GNFs, upon which$n$-gram GNF is defined to represent co-occurrent visual patterns. After that, efficient merging and refining algorithms are used to get the partial-duplicate clusters, and logical combinations of probabilistic GNF models are leveraged to represent the visual patterns of partially duplicate images. Extensive experiments show the parallelizable property and effectiveness of the algorithms on both partial-duplicate clustering and visual pattern discovery. With 2000 machines, it costs about eight and 400 minutes to process one million and 40 million images respectively, which is quite efficient compared to previous methods.
Wei Li 0152, Changhu Wang, Lei Zhang 0001, Yong Rui, Bo Zhang 0010
IEEE Trans. Multim.5
2015 Fine-Grained Image Search
abstract
Large-scale image search has been attracting lots of attention from both academic and commercial fields. The conventional bag-of-visual-words (BoVW) model with inverted index is verified efficient at retrieving near-duplicate images, but it is less capable of discovering fine-grained concepts in the query and returning semantically matched search results. In this paper, we suggest that instance search should return not only near-duplicate images, but also fine-grained results, which is usually the actual intention of a user. We propose a new and interesting problem named fine-grained image search, which means that we prefer those images containing the same fine-grained concept with the query. We formulate the problem by constructing a hierarchical database and defining an evaluation method. We thereafter introduce a baseline system using fine-grained classification scores to represent and co-index images so that the semantic attributes are better incorporated in the online querying stage. Large-scale experiments reveal that promising search results are achieved with reasonable time and memory consumption. We hope this paper will be the foundation for future work on image search. We also expect more follow-up efforts along this research topic and look forward to commercial fine-grained image search engines.
Lingxi Xie, Jingdong Wang 0001, Bo Zhang 0010, Qi Tian 0001
IEEE Trans. Multim.3
2015 Comparison of ℓ1-Norm SVR and Sparse Coding Algorithms for Linear Regression
abstract
Support vector regression (SVR) is a popular function estimation technique based on Vapnik's concept of support vector machine. Among many variants, the l1-norm SVR is known to be good at selecting useful features when the features are redundant. Sparse coding (SC) is a technique widely used in many areas and a number of efficient algorithms are available. Both l1-norm SVR and SC can be used for linear regression. In this brief, the close connection between the l1-norm SVR and SC is revealed and some typical algorithms are compared for linear regression. The results show that the SC algorithms outperform the Newton linear programming algorithm, an efficient l1-norm SVR algorithm, in efficiency. The algorithms are then used to design the radial basis function (RBF) neural networks. Experiments on some benchmark data sets demonstrate the high efficiency of the SC algorithms. In particular, one of the SC algorithms, the orthogonal matching pursuit is two orders of magnitude faster than a well-known RBF network designing algorithm, the orthogonal least squares algorithm.
Qingtian Zhang, Xiaolin Hu 0001, Bo Zhang 0010
IEEE Trans. Neural Networks Learn. Syst.3
2014 Dropout Training for Support Vector Machines
abstract
Dropout and other feature noising schemes have shown promising results in controlling over-fitting by artificially corrupting the training data. Though extensive theoretical and empirical studies have been performed for generalized linear models, little work has been done for support vector machines (SVMs), one of the most successful approaches for supervised learning. This paper presents dropout training for linear SVMs. To deal with the intractable expectation of the non-smooth hinge loss under corrupting distributions, we develop an iteratively re-weighted least square (IRLS) algorithm by exploring data augmentation techniques. Our algorithm iteratively minimizes the expectation of a re-weighted least square problem, where the re-weights have closed-form solutions. The similar ideas are applied to develop a new IRLS algorithm for the expected logistic loss under corrupting distributions. Our algorithms offer insights on the connection and difference between the hinge loss and logistic loss in dropout training. Empirical results on several real datasets demonstrate the effectiveness of dropout training on significantly boosting the classification accuracy of linear SVMs.
Ning Chen 0002, Jun Zhu 0001, Jianfei Chen 0001, Bo Zhang 0010
AAAI4
2014 Similarity-Preserving Binary Signature for Linear Subspaces
abstract
Linear subspace is an important representation for many kinds of real-world data in computer vision and pattern recognition, e.g. faces, motion videos, speeches. In this paper, first we define pairwise angular similarity and angular distance for linear subspaces. The angular distance satisfies non-negativity, identity of indiscernibles, symmetry and triangle inequality, and thus it is a metric. Then we propose a method to compress linear subspaces into compact similarity-preserving binary signatures, between which the normalized Hamming distance is an unbiased estimator of the angular distance. We provide a lower bound on the length of the binary signatures which suffices to guarantee uniform distance-preservation within a set of subspaces. Experiments on face recognition demonstrate the effectiveness of the binary signature in terms of recognition accuracy, speed and storage requirement. The results show that, compared with the exact method, the approximation with the binary signatures achieves an order of magnitude speed-up, while requiring significantly smaller amount of storage space, yet it still accurately preserves the similarity, and achieves high recognition accuracy comparable to the exact method in face recognition.
Jianqiu Ji, Jianmin Li 0001, Shuicheng Yan, Qi Tian 0001, Bo Zhang 0010
AAAI5
2014 Orientational Pyramid Matching for Recognizing Indoor Scenes
abstract
Scene recognition is a basic task towards image understanding. Spatial Pyramid Matching (SPM) has been shown to be an efficient solution for spatial context modeling. In this paper, we introduce an alternative approach, Orientational Pyramid Matching (OPM), for orientational context modeling. Our approach is motivated by the observation that the 3D orientations of objects are a crucial factor to discriminate indoor scenes. The novelty lies in that OPM uses the 3D orientations to form the pyramid and produce the pooling regions, which is unlike SPM that uses the spatial positions to form the pyramid. Experimental results on challenging scene classification tasks show that OPM achieves the performance comparable with SPM and that OPM and SPM make complementary contributions so that their combination gives the state-of-the-art performance.
Lingxi Xie, Jingdong Wang 0001, Baining Guo, Bo Zhang 0010, Qi Tian 0001
CVPR4
2014 Max-SIFT: Flipping invariant descriptors for Web logo search
abstract
Logo search is widely required in many real-world applications. As a special case of near-duplicate images, logo pictures have some particular properties, for instance, suffering from flipping operations, e.g., geometry-inverted and brightness-inverted operations. Such operations completely change the spatial structure of local descriptors, such as SIFT, so that image search algorithms based on Bag-of-Visual-Words (BoVW) often fail to retrieve the flipped logos. We propose a novel descriptor named Max-SIFT, which finds the maximal SIFT value sequence for detecting flipping operations. Compared with previous algorithms, our algorithm is extremely easy to implement yet very efficient to carry out. We evaluate the improved descriptor on a large-scale Web logo search dataset, and demonstrate that our method enjoys good performance and low computational costs.
Lingxi Xie, Qi Tian 0001, Bo Zhang 0010
ICIP3
2014 Max-Margin Infinite Hidden Markov Models
abstract
Infinite hidden Markov models (iHMMs) are nonparametric Bayesian extensions of hidden Markov models (HMMs) with an infinite number of states. Though flexible in describing sequential data, the generative formulation of iHMMs could limit their discriminative ability in sequential prediction tasks. Our paper introduces max-margin infinite HMMs (M2iHMMs), new infinite HMMs that explore the max-margin principle for discriminative learning. By using the theory of Gibbs classifiers and data augmentation, we develop efficient beam sampling algorithms without making restricting mean-field assumptions or truncated approximation. For single variate classification, M2iHMMs reduce to a new formulation of DP mixtures of max-margin machines. Empirical results on synthetic and real data sets show that our methods obtain superior performance than other competitors in both single variate classification and sequential prediction tasks.
Aonan Zhang, Jun Zhu 0001, Bo Zhang 0010
ICML3
2014 Distributed Bayesian Posterior Sampling via Moment Sharing
Minjie Xu, Balaji Lakshminarayanan, Yee Whye Teh, Jun Zhu 0001, Bo Zhang 0010
NIPS5
2014 Fast and accurate near-duplicate image search with affinity propagation on the ImageWeb
Lingxi Xie, Qi Tian 0001, Wengang Zhou 0001, Bo Zhang 0010
Comput. Vis. Image Underst.4
2014 Modeling response properties of V2 neurons using a hierarchical K-means model
Xiaolin Hu 0001, Jianwei Zhang 0001, Peng Qi 0004, Bo Zhang 0010
Neurocomputing4
2014 Gibbs max-margin topic models with data augmentation
Jun Zhu 0001, Ning Chen 0002, Hugh Perkins, Bo Zhang 0010
J. Mach. Learn. Res.4
2014 Batch-Orthogonal Locality-Sensitive Hashingfor Angular Similarity
abstract
Sign-random-projection locality-sensitive hashing (SRP-LSH) is a widely used hashing method, which provides an unbiased estimate of pairwise angular similarity, yet may suffer from its large estimation variance. We propose in this work batch-orthogonal locality-sensitive hashing (BOLSH), as a significant improvement of SRP-LSH. Instead of independent random projections, BOLSH makes use of batch-orthogonalized random projections, i.e, we divide random projection vectors into several batches and orthogonalize the vectors in each batch respectively. These batch-orthogonalized random projections partition the data space into regular regions, and thus provide a more accurate estimator. We prove theoretically that BOLSH still provides an unbiased estimate of pairwise angular similarity, with a smaller variance for any angle in (0, π), compared with SRP-LSH. Furthermore, we give a lower bound on the reduction of variance. The extensive experiments on real data well validate that with the same length of binary code, BOLSH may achieve significant mean squared error reduction in estimating pairwise angular similarity. Moreover, BOLSH shows the superiority in extensive approximate nearest neighbor (ANN) retrieval experiments.
Jianqiu Ji, Shuicheng Yan, Jianmin Li 0001, Guangyu Gao, Qi Tian 0001, Bo Zhang 0010
IEEE Trans. Pattern Anal. Mach. Intell.6
2014 Spatial Pooling of Heterogeneous Features for Image Classification
abstract
In image classification tasks, one of the most successful algorithms is the bag-of-features (BoFs) model. Although the BoF model has many advantages, such as simplicity, generality, and scalability, it still suffers from several drawbacks, including the limited semantic description of local descriptors, lack of robust structures upon single visual words, and missing of efficient spatial weighting. To overcome these shortcomings, various techniques have been proposed, such as extracting multiple descriptors, spatial context modeling, and interest region detection. Though they have been proven to improve the BoF model to some extent, there still lacks a coherent scheme to integrate each individual module together. To address the problems above, we propose a novel framework with spatial pooling of complementary features. Our model expands the traditional BoF model on three aspects. First, we propose a new scheme for combining texture and edge-based local features together at the descriptor extraction level. Next, we build geometric visual phrases to model spatial context upon complementary features for midlevel image representation. Finally, based on a smoothed edgemap, a simple and effective spatial weighting scheme is performed to capture the image saliency. We test the proposed framework on several benchmark data sets for image classification. The extensive results show the superior performance of our algorithm over the state-of-the-art methods.
Lingxi Xie, Qi Tian 0001, Meng Wang 0001, Bo Zhang 0010
IEEE Trans. Image Process.4
2014 Learning Harmonium Models With Infinite Latent Features
abstract
Undirected latent variable models represent an important class of graphical models that have been successfully developed to deal with various tasks. One common challenge in learning such models is to determine the number of hidden units that are unknown a priori. Although Bayesian nonparametrics have provided promising results in bypassing the model selection problem in learning directed Bayesian Networks, very little effort has been made toward applying Bayesian nonparametrics to learn undirected latent variable models. In this paper, we present the infinite exponential family Harmonium (iEFH), a bipartite undirected latent variable model that automatically determines the number of latent units from an unbounded pool. We also present two important extensions of iEFH to 1) multiview iEFH for dealing with heterogeneous data, and 2) infinite maximum-margin Harmonium (iMMH) for incorporating supervising side information to learn predictive latent features. We develop variational inference algorithms to learn model parameters. Our methods are computationally competitive because of the avoidance of selecting the number of latent units. Our extensive experiments on real image datasets and text datasets appear to demonstrate the benefits of iEFH and iMMH inherited from Bayesian nonparametrics and max-margin learning. Such results were not available until now and contribute to expanding the scope of Bayesian nonparametrics to learn the structures of undirected latent variable models.
Ning Chen 0002, Jun Zhu 0001, Fuchun Sun 0001, Bo Zhang 0010
IEEE Trans. Neural Networks Learn. Syst.4
2013 Improved Bayesian Logistic Supervised Topic Models with Data Augmentation
Jun Zhu 0001, Xun Zheng, Bo Zhang 0010
ACL (1)3
2013 Hierarchical Part Matching for Fine-Grained Visual Categorization
abstract
As a special topic in computer vision, fine-grained visual categorization (FGVC) has been attracting growing attention these years. Different with traditional image classification tasks in which objects have large inter-class variation, the visual concepts in the fine-grained datasets, such as hundreds of bird species, often have very similar semantics. Due to the large inter-class similarity, it is very difficult to classify the objects without locating really discriminative features, therefore it becomes more important for the algorithm to make full use of the part information in order to train a robust model. In this paper, we propose a powerful flowchart named Hierarchical Part Matching (HPM) to cope with fine-grained classification tasks. We extend the Bag-of-Features (BoF) model by introducing several novel modules to integrate into image representation, including foreground inference and segmentation, Hierarchical Structure Learning (HSL), and Geometric Phrase Pooling (GPP). We verify in experiments that our algorithm achieves the state-of-the-art classification accuracy in the Caltech-UCSD-Birds-200-2011 dataset by making full use of the ground-truth part annotations.
Lingxi Xie, Qi Tian 0001, Richang Hong, Shuicheng Yan, Bo Zhang 0010
ICCV5
2013 Min-Max Hash for Jaccard Similarity
abstract
Min-wise hash is a widely-used hashing method for scalable similarity search in terms of Jaccard similarity, while in practice it is necessary to compute many such hash functions for certain precision, leading to expensive computational cost. In this paper, we introduce an effective method, i.e. the min-max hash method, which significantly reduces the hashing time by half, yet it has a provably slightly smaller variance in estimating pair wise Jaccard similarity. In addition, the estimator of min-max hash only contains pair wise equality checking, thus it is especially suitable for approximate nearest neighbor search. Since min-max hash is equally simple as min-wise hash, many extensions based on min-wise hash can be easily adapted to min-max hash, and we show how to combine it with b-bit minwise hash. Experiments show that with the same length of hash code, min-max hash reduces the hashing time to half as much as that of min-wise hash, while achieving smaller mean squared error (MSE) in estimating pair wise Jaccard similarity, and better best approximate ratio (BAR) in approximate nearest neighbor search.
Jianqiu Ji, Jianmin Li 0001, Shuicheng Yan, Qi Tian 0001, Bo Zhang 0010
ICDM5
2013 Feature normalization for part-based image classification
abstract
Part-based Bag-of-Features (BoF) models such as Spatial Pyramid Matching (SPM) play an important role in image classification. Before sending the feature vectors into classifiers for training and testing, it is required to normalize them in order to approximately equalize ranges of the attributes and make them have comparable effects in distance computation. Although some works have been focused on general feature normalization, we do not see any discussion on specialized normalization algorithms for part-based BoF models. In this paper, we fill in the blank with extensive experiments and discussions. Based on solid normalization parameters (power and coefficient), we further study two straightforward part-based properties, i.e., the independent assumption and the hierarchical-contribution assumption, to scale the feature super-vectors separately. Finally, we test our algorithm on challenging image sets, i.e., Caltech 101 and CUB-200-2011, for general and fine-grained classification, and show its efficiency, scalability and adaptability in both scenarios.
Lingxi Xie, Qi Tian 0001, Bo Zhang 0010
ICIP3
2013 Fast Max-Margin Matrix Factorization with Data Augmentation
abstract
Existing max-margin matrix factorization (M3F) methods either are computationally inefficient or need a model selection procedure to determine the number of latent factors. In this paper we present a probabilistic M3F model that admits a highly efficient Gibbs sampling algorithm through data augmentation. We further extend our approach to incorporate Bayesian nonparametrics and build accordingly a truncation-free nonparametric M3F model where the number of latent factors is literally unbounded and inferred from data. Empirical studies on two large real-world data sets verify the efficacy of our proposed methods.
Minjie Xu, Jun Zhu 0001, Bo Zhang 0010
ICML (3)3
2013 Gibbs Max-Margin Topic Models with Fast Sampling Algorithms
abstract
Existing max-margin supervised topic models rely on an iterative procedure to solve multiple latent SVM subproblems with additional mean-field assumptions on the desired posterior distributions. This paper presents Gibbs max-margin supervised topic models by minimizing an expected margin loss, an upper bound of the existing margin loss derived from an expected prediction rule. By introducing augmented variables, we develop simple and fast Gibbs sampling algorithms with no restricting assumptions and no need to solve SVM subproblems for both classification and regression. Empirical results demonstrate significant improvements on time efficiency. The classification performance is also significantly improved over competitors.
Jun Zhu 0001, Ning Chen 0002, Hugh Perkins, Bo Zhang 0010
ICML (1)4
2013 Restricted Boltzmann Machine with Adaptive Local Hidden Units
Binbin Cao, Jianmin Li 0001, Jun Wu 0022, Bo Zhang 0010
ICONIP (2)4
2013 Generalized Relational Topic Models with Data Augmentation
Ning Chen 0002, Jun Zhu 0001, Fei Xia 0005, Bo Zhang 0010
IJCAI4
2013 Scalable inference in max-margin topic models
abstract
Topic models have played a pivotal role in analyzing large collections of complex data. Besides discovering latent semantics, supervised topic models (STMs) can make predictions on unseen test data. By marrying with advanced learning techniques, the predictive strengths of STMs have been dramatically enhanced, such as max-margin supervised topic models, state-of-the-art methods that integrate max-margin learning with topic models. Though powerful, max-margin STMs have a hard non-smooth learning problem. Existing algorithms rely on solving multiple latent SVM subproblems in an EM-type procedure, which can be too slow to be applicable to large-scale categorization tasks.
Jun Zhu 0001, Xun Zheng, Li Zhou 0006, Bo Zhang 0010
KDD4
2013 Scalable Inference for Logistic-Normal Topic Models
abstract
Logistic-normal topic models can effectively discover correlation structures among latent topics. However, their inference remains a challenge because of the non-conjugacy between the logistic-normal prior and multinomial topic mixing proportions. Existing algorithms either make restricting mean-field assumptions or are not scalable to large-scale applications. This paper presents a partially collapsed Gibbs sampling algorithm that approaches the provably correct distribution by exploring the ideas of data augmentation. To improve time efficiency, we further present a parallel implementation that can deal with large-scale applications and learn the correlation structures of thousands of topics from millions of documents. Extensive empirical results demonstrate the promise.
Jianfei Chen 0001, Jun Zhu 0001, Xun Zheng, Bo Zhang 0010
NIPS5
2013 Sparse Relational Topic Models for Document Networks
Aonan Zhang, Jun Zhu 0001, Bo Zhang 0010
ECML/PKDD (1)3
2013 Sparse online topic models
abstract
Topic models have shown great promise in discovering latent semantic structures from complex data corpora, ranging from text documents and web news articles to images, videos, and even biological data. In order to deal with massive data collections and dynamic text streams, probabilistic online topic models such as online latent Dirichlet allocation (OLDA) have recently been developed. However, due to normalization constraints, OLDA can be ineffective in controlling the sparsity of discovered representations, a desirable property for learning interpretable semantic patterns, especially when the total number of topics is large. In contrast, sparse topical coding (STC) has been successfully introduced as a non-probabilistic topic model for effectively discovering sparse latent patterns by using sparsity-inducing regularization. But, unfortunately STC cannot scale to very large datasets or deal with online text streams, partly due to its batch learning procedure. In this paper, we present a sparse online topic model, which directly controls the sparsity of latent semantic patterns by imposing sparsity-inducing regularization and learns the topical dictionary by an online algorithm. The online algorithm is efficient and guaranteed to converge. Extensive empirical results of the sparse online topic model as well as its collapsed and supervised extensions on a large-scale Wikipedia dataset and the medium-sized 20Newsgroups dataset demonstrate appealing performance.
Aonan Zhang, Jun Zhu 0001, Bo Zhang 0010
WWW3
2013 Learning a Contextual Multi-Thread Model for Movie/TV Scene Segmentation
abstract
Compared with general videos, movies and TV shows attract a significantly larger portion of people across time and contain very rich and interesting narrative patterns of shots and scenes. In this paper, we aim to recover the inherent structure of scenes and shots in such video narratives. The obtained structure could be useful for subsequent video analysis tasks such as tracking objects across cuts, action retrieval, as well as enriching user browsing and video editing interfaces. Recent research on this problem has mainly focused on combining multiple cues such as scripts, subtitles, sound, or human faces. However, considering that visual information is sufficient for human to identify scene boundaries and some cues are not always available, we are motivated to design a purely visual approach. Observing that dialog patterns occur frequently in a movie/TV show to form a scene, we propose a probabilistic framework to imitate the authoring process. The multi-thread shot model and contextual visual dynamics are embedded into a unified framework to capture the video hierarchy. We devise an efficient algorithm to jointly learn the parameters of the unified model. Experiments on two large datasets containing six movies and 24 episodes of Lost, a popular TV show with complex plot structures, are conducted. Comparative results show that, leveraging only visual cues, our method could successfully recover complicated shot threads and outperform several approaches. Moreover, our method is fast and advantageous for large-scale computation.
Cailiang Liu, Dong Wang 0022, Jun Zhu 0001, Bo Zhang 0010
IEEE Trans. Multim.4
2012 Hierarchical K-Means Algorithm for Modeling Visual Area V2 Neurons
Xiaolin Hu 0001, Peng Qi 0004, Bo Zhang 0010
ICONIP (3)3
2012 A new challenge of information processing under the 21st century
abstract
In web era, we are confronted with a huge amount of raw data and a tremendous change of man-machine interaction modes. We have to deal with the content (semantics) of data rather than their form alone. Traditional information processing approaches face a new challenge since they cannot deal with the semantic meaning or content of information. But humans can handle such a problem easily. So it's needed a new information processing strategy that correlated with the content of information by learning some mechanisms from human beings. Therefore, we need (1) a set of robust detectors for detecting semantically meaningful features such as boundaries, shapes, etc. in images, words, sentences, etc. in text, and (2) a set of methods that can effectively analyze and exploit the information structures that encode the content of information. During the past 40 years the probability theory has made a great progress. It has provided a set of mathematical tools for representing and analyzing information structures. In the talk we will discuss what difficulty we face, what we can do, and how we should do in the content-based information processing.
Bo Zhang 0010
KDD1
2012 Spatial pooling of heterogeneous features for image applications
abstract
The Bag-of-Features (BoF) model has played an important role for image representation in many multimedia applications. It has been extensively applied to many tasks including image classification, image retrieval, scene understanding, and so on. Despite the advantages of this model such as simplicity, efficiency and generality, there are also notable drawbacks for this model, including poor power of semantic expression of local descriptors, and lack of robust structures upon single visual words. To overcome these problems, various techniques have been proposed, such as multiple descriptors, spatial context modeling and interest region detection. Though they have been proven to improve the BoF model to some extent, there still lacks a coherent scheme to integrate each individual module.
Lingxi Xie, Qi Tian 0001, Bo Zhang 0010
ACM Multimedia3
2012 Super-Bit Locality-Sensitive Hashing
abstract
Sign-random-projection locality-sensitive hashing (SRP-LSH) is a probabilistic dimension reduction method which provides an unbiased estimate of angular similarity, yet suffers from the large variance of its estimation. In this work, we propose the Super-Bit locality-sensitive hashing (SBLSH). It is easy to implement, which orthogonalizes the random projection vectors in batches, and it is theoretically guaranteed that SBLSH also provides an unbiased estimate of angular similarity, yet with a smaller variance when the angle to estimate is within $(0,\pi/2]$. The extensive experiments on real data well validate that given the same length of binary code, SBLSH may achieve significant mean squared error reduction in estimating pairwise angular similarity. Moreover, SBLSH shows the superiority over SRP-LSH in approximate nearest neighbor (ANN) retrieval experiments.
Jianqiu Ji, Jianmin Li 0001, Shuicheng Yan, Bo Zhang 0010, Qi Tian 0001
NIPS4
2012 Nonparametric Max-Margin Matrix Factorization for Collaborative Prediction
abstract
We present a probabilistic formulation of max-margin matrix factorization and build accordingly a nonparametric Bayesian model which automatically resolves the unknown number of latent factors. Our work demonstrates a successful example that integrates Bayesian nonparametrics and max-margin learning, which are conventionally two separate paradigms and enjoy complementary advantages. We develop an efcient variational algorithm for posterior inference, and our extensive empirical studies on large-scale MovieLens and EachMovie data sets appear to justify the aforementioned dual advantages.
Minjie Xu, Jun Zhu 0001, Bo Zhang 0010
NIPS3
2011 Learning complex image patterns with Scale and Shift Invariant Sparse Coding
abstract
The image patches learned by recent works are usually only bar-like or Gabor-like patterns. However those simple patterns are not meaningful enough to capture higher level information. In this study, we try to learn more complex image patterns from unaligned images. We propose Scale And Shift Invariant Sparse Coding (SASISC), which aligns basis patches at proper locations and scales to reconstruct the whole image. The experiment results on unaligned images show that SASISC can explain the images much better than the original sparse coding, and can extract more complex image patterns.
Bo Zhang 0010
ICIP2
2011 Learning Hierarchical Dictionary for Shape Patterns
Bo Zhang 0010
ISNN (2)2
2011 The structural analysis of fuzzy measures
Ling Zhang 0001, Bo Zhang 0010, Yanping Zhang 0001
Sci. China Inf. Sci.2
2010 Robust semantic sketch based specific image retrieval
abstract
Specific images refer to images one has a certain episodic memory about, e.g. a picture one has ever seen before. Specific image retrieval is a frequent daily information need and the episodic memory is the key to find a specific image. In this paper, we propose a novel semantic sketch-based interface to incorporate the episodic memory for specific image retrieval. The interface allows a user to specify the semantic category and rough area/color of the objects in his memory. To bridge the semantic gap between the query sketch and database images, in the back end, a sampling method selects exemplars from a reference dataset which contains many object instances with user-provided tags and bounding boxes. After that, an exemplar matching algorithm ranks images to retrieve the target image to match the user's memory. In practice, we have observed that query sketches are usually error prone. That is, the position or the color of an object may not be accurate. Meanwhile, the annotations in the reference dataset are also noisy. Thus, the search algorithm has to handle two kinds of errors: 1) reference dataset label noise; 2) user sketch error such as position or scale. For the former, we propose a robust sampling method. For the latter, we derive an efficient spatial reranking algorithm to tolerate inaccurate user sketches. Detailed experimental results on the LabelMe dataset show that the proposed approach is robust to both kinds of errors.
Cailiang Liu, Dong Wang 0022, Changhu Wang, Lei Zhang 0001, Bo Zhang 0010
ICME6
2010 Learning Vocabulary-Based Hashing with AdaBoost
Yingyu Liang, Jianmin Li 0001, Bo Zhang 0010
MMM3
2010 Fuzzy tolerance quotient spaces and fuzzy subsets
Ling Zhang 0001, Bo Zhang 0010
Sci. China Inf. Sci.2
2010 A Gaussian Attractor Network for Memory and Recognition with Experience-Dependent Learning
abstract
Attractor networks are widely believed to underlie the memory systems of animals across different species. Existing models have succeeded in qualitatively modeling properties of attractor dynamics, but their computational abilities often suffer from poor representations for realistic complex patterns, spurious attractors, low storage capacity, and difficulty in identifying attractive fields of attractors. We propose a simple two-layer architecture, gaussian attractor network, which has no spurious attractors if patterns to be stored are uncorrelated and can store as many patterns as the number of neurons in the output layer. Meanwhile the attractive fields can be precisely quantified and manipulated. Equipped with experience-dependent unsupervised learning strategies, the network can exhibit both discrete and continuous attractor dynamics. A testable prediction based on numerical simulations is that there exist neurons in the brain that can discriminate two similar stimuli at first but cannot after extensive exposure to physically intermediate stimuli. Inspired by this network, we found that adding some local feedbacks to a well-known hierarchical visual recognition model, HMAX, can enable the model to reproduce some recent experimental results related to high-level visual perception.
Xiaolin Hu 0001, Bo Zhang 0010
Neural Comput.2
2010 Design of recurrent neural networks for solving constrained least absolute deviation problems
abstract
Recurrent neural networks for solving constrained least absolute deviation (LAD) problems or L(1)-norm optimization problems have attracted much interest in recent years. But so far most neural networks can only deal with some special linear constraints efficiently. In this paper, two neural networks are proposed for solving LAD problems with various linear constraints including equality, two-sided inequality and bound constraints. When tailored to solve some special cases of LAD problems in which not all types of constraints are present, the two networks can yield simpler architectures than most existing ones in the literature. In particular, for solving problems with both equality and one-sided inequality constraints, another network is invented. All of the networks proposed in this paper are rigorously shown to be capable of solving the corresponding problems. The different networks designed for solving the same types of problems possess the same structural complexity, which is due to the fact these architectures share the same computing blocks and only differ in connections between some blocks. By this means, some flexibility for circuits realization is provided. Numerical simulations are carried out to illustrate the theoretical results and compare the convergence rates of the networks.
Xiaolin Hu 0001, Changyin Sun 0001, Bo Zhang 0010
IEEE Trans. Neural Networks3
2009 Another Simple Recurrent Neural Network for Quadratic and Linear Programming
Xiaolin Hu 0001, Bo Zhang 0010
ISNN (3)2
2009 Primal sparse Max-margin Markov networks
abstract
Max-margin Markov networks (M3N) have shown great promise in structured prediction and relational learning. Due to the KKT conditions, the M3N enjoys dual sparsity. However, the existing M3N formulation does not enjoy primal sparsity, which is a desirable property for selecting significant features and reducing the risk of over-fitting. In this paper, we present an l1-norm regularized max-margin Markov network (l1-M3N), which enjoys dual and primal sparsity simultaneously. To learn an l1-M3N, we present three methods including projected sub-gradient, cutting-plane, and a novel EM-style algorithm, which is based on an equivalence between l1-M3N and an adaptive M3N. We perform extensive empirical studies on both synthetic and real data sets. Our experimental results show that: (1) l1-M3N can effectively select significant features; (2) l1-M3N can perform as well as the pseudo-primal sparse Laplace M3N in prediction accuracy, while consistently outperforms other competing methods that enjoy either primal or dual sparsity; and (3) the EM-algorithm is more robust than the other two in pre-diction accuracy and time efficiency.
Jun Zhu 0001, Eric P. Xing, Bo Zhang 0010
KDD3
2009 Vocabulary-based hashing for image search
abstract
This paper proposes a hash function family based on feature vocabularies and investigates the application in building indexes for image search. Each hash function is associated with a set of feature points, i.e. a vocabulary, and maps an input point to the ID of the nearest one in the vocabulary. The function family can be employed to build a high-dimensional index for approximate nearest neighbor search. Then we concentrate on its application in image search. Guiding rules for the construction of the vocabularies are derived, which improve the effectiveness of the approach in this context by taking advantage of the data distribution. The rules are applied to design an algorithm for vocabulary construction in practice. Experiments show promising performance of the approach and the effectiveness of the guiding rules. Comparison with the popular Euclidean locality-sensitive hashing also shows the advantage of our approach in image search.
Yingyu Liang, Jianmin Li 0001, Bo Zhang 0010
ACM Multimedia3
2009 Motion Planning with Obstacle Avoidance for Kinematically Redundant Manipulators Based on Two Recurrent Neural Networks
abstract
Inverse kinematic motion planning of redundant manipulators by using recurrent neural networks in the presence of obstacles and uncertainties is a real-time nonlinear optimization problem. To tackle this problem, two subproblems should be resolved in real time. One is the determination of critical points on a given manipulator closest to obstacles, and the other is the computation of joint velocities of the manipulator which can direct the manipulator following a desired trajectory and away from obstacles if it is getting close to them. Different from our previous approaches where the critical points on the manipulator were assumed to be known, these points are to be computed by using a recurrent neural network in the paper. A time-varying quadratic programming problem is formulated for avoiding polyhedral obstacles. In view that the problem is not strictly convex, an existing recurrent neural network, general projection neural network, is applied for solving it. By introducing a velocity smoothing technique into our previous quadratic programming formulation of the joint velocity assignment problem, a recently developed recurrent neural network, improved dual neural network, is proposed to solve it, which features lower structural complexity compared with existing neural networks. Moreover, The effectiveness of the proposed neural networks is demonstrated by simulations on the Mitsubishi PA10-7C manipulator.
Xiaolin Hu 0001, Jun Wang 0002, Bo Zhang 0010
SMC3
2009 StatSnowball: a statistical approach to extracting entity relationships
abstract
Traditional relation extraction methods require pre-specified relations and relation-specific human-tagged examples. Bootstrapping systems significantly reduce the number of training examples, but they usually apply heuristic-based methods to combine a set of strict hard rules, which limit the ability to generalize and thus generate a low recall. Furthermore, existing bootstrapping methods do not perform open information extraction (Open IE), which can identify various types of relations without requiring pre-specifications. In this paper, we propose a statistical extraction framework called Statistical Snowball (StatSnowball), which is a bootstrapping system and can perform both traditional relation extraction and Open IE.
Jun Zhu 0001, Zaiqing Nie, Xiaojiang Liu, Bo Zhang 0010, Ji-Rong Wen
WWW4
2009 Query representation by structured concept threads with application to interactive video retrieval
Dong Wang 0022, Jianmin Li 0001, Bo Zhang 0010, Xirong Li 0001
J. Vis. Commun. Image Represent.4
2009 A New Recurrent Neural Network for Solving Convex Quadratic Programming Problems With an Application to the k -Winners-Take-All Problem
abstract
In this paper, a new recurrent neural network is proposed for solving convex quadratic programming (QP) problems. Compared with existing neural networks, the proposed one features global convergence property under weak conditions, low structural complexity, and no calculation of matrix inverse. It serves as a competitive alternative in the neural network family for solving linear or quadratic programming problems. In addition, it is found that by some variable substitution, the proposed network turns out to be an existing model for solving minimax problems. In this sense, it can be also viewed as a special case of the minimax neural network. Based on this scheme, a k-winners-take-all ( k-WTA) network with O(n) complexity is designed, which is characterized by simple structure, global convergence, and capability to deal with some ill cases. Numerical simulations are provided to validate the theoretical results obtained. More importantly, the network design method proposed in this paper has great potential to inspire other competitive inventions along the same line.
Xiaolin Hu 0001, Bo Zhang 0010
IEEE Trans. Neural Networks2
2009 An Alternative Recurrent Neural Network for Solving Variational Inequalities and Related Optimization Problems
abstract
There exist many recurrent neural networks for solving optimization-related problems. In this paper, we present a method for deriving such networks from existing ones by changing connections between computing blocks. Although the dynamic systems may become much different, some distinguished properties may be retained. One example is discussed to solve variational inequalities and related optimization problems with mixed linear and nonlinear constraints. A new network is obtained from two classical models by this means, and its performance is comparable to its predecessors. Thus, an alternative choice for circuits implementation is offered to accomplish such computing tasks.
Xiaolin Hu 0001, Bo Zhang 0010
IEEE Trans. Syst. Man Cybern. Part B2
2008 Scene understanding with discriminative structured prediction
abstract
Spatial priors play crucial roles in many high-level vision tasks, e.g. scene understanding. Usually, learning spatial priors relies on training a structured output model. In this paper, two special cases of discriminative structured output model, i.e. conditional random fields (CRFs) and max-margin Markov networks (M3N), are demonstrated to perform image scene understanding. The two models are empirically compared in a fair manner, i.e. using the common feature representation and the same optimization algorithm. Particularly, we adopt online exponentiated gradient (EG) algorithm to solve the convex duals of both models. We describe the general procedure of EG algorithm and present a two-stage training procedure to overcome the degeneration of EG when exact inference is intractable. Experiments on a large scale image region annotation task are carried out. The results show that both models yield encouraging results but CRFs slightly outperforms M3N.
Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010
CVPR3
2008 Laplace maximum margin Markov networks
abstract
We propose Laplace max-margin Markov networks (LapM3N), and a general class of Bayesian M3N (BM3N) of which the LapM3N is a special case with sparse structural bias, for robust structured prediction. BM3N generalizes extant structured prediction rules based on point estimator to a Bayes-predictor using a learnt distribution of rules. We present a novel Structured Maximum Entropy Discrimination (SMED) formalism for combining Bayesian and max-margin learning of Markov networks for structured prediction, and our approach subsumes the conventional M3N as a special case. An efficient learning algorithm based on variational inference and standard convex-optimization solvers for M3N, and a generalization bound are offered. Our method outperforms competing ones on both synthetic and real OCR data.
Jun Zhu 0001, Eric P. Xing, Bo Zhang 0010
ICML3
2008 Three Global Exponential Convergence Results of the GPNN for Solving Generalized Linear Variational Inequalities
Xiaolin Hu 0001, Zhigang Zeng, Bo Zhang 0010
ISNN (1)3
2008 Partially Observed Maximum Entropy Discrimination Markov Networks
abstract
Learning graphical models with hidden variables can offer semantic insights to complex data and lead to salient structured predictors without relying on expensive, sometime unattainable fully annotated training data. While likelihood-based methods have been extensively explored, to our knowledge, learning structured prediction models with latent variables based on the max-margin principle remains largely an open problem. In this paper, we present a partially observed Maximum Entropy Discrimination Markov Network (PoMEN) model that attempts to combine the advantages of Bayesian and margin based paradigms for learning Markov networks from partially labeled data. PoMEN leads to an averaging prediction rule that resembles a Bayes predictor that is more robust to overfitting, but is also built on the desirable discriminative laws resemble those of the M$^3$N. We develop an EM-style algorithm utilizing existing convex optimization algorithms for M$^3$N as a subroutine. We demonstrate competent performance of PoMEN over existing methods on a real-world web data extraction task.
Jun Zhu 0001, Eric P. Xing, Bo Zhang 0010
NIPS3
2008 Dynamic Hierarchical Markov Random Fields for Integrated Web Data Extraction
Jun Zhu 0001, Zaiqing Nie, Bo Zhang 0010, Ji-Rong Wen
J. Mach. Learn. Res.3
2007 AHP: A New Strategy for the Semantic Concept Detection in Video
abstract
The analytic hierarchy process (AHP) is a method to help people make a complex decision by analyzing and synthesizing multiple criteria for the decision objective in a hierarchy. We adopt the AHP as a new strategy for the semantic concept detection (SCD) in video so that multiple factors involved in SCD, including multiple modalities and relating concepts, can be hierarchically analyzed and synthesized. In this paper, we first explain, by an example, why and how the SCD problem can be analyzed by the AHP. Then, following the idea of the AHP, we develop a new rank aggregation (RA) method, called AHP-RA. Experimental results of RA for SCD in video show the effectiveness of this method.
Dayong Ding, Bo Zhang 0010, Jinglan Wu
ICME2
2007 Dynamic hierarchical Markov random fields and their application to web data extraction
abstract
Hierarchical models have been extensively studied in various domains. However, existing models assume fixed model structures or incorporate structural uncertainty generatively. In this paper, we propose Dynamic Hierarchical Markov Random Fields (DHMRFs) to incorporate structural uncertainty in a discriminative manner. DHMRFs consist of two parts -- structure model and class label model. Both are defined as exponential family distributions. Conditioned on observations, DHMRFs relax the independence assumption as made in directed models. As exact inference is intractable, a variational method is developed to learn parameters and to find the MAP model structure and label assignment. We apply the model to a real-world web data extraction task, which automatically extracts product items for sale on the Web. The results show promise.
Jun Zhu 0001, Zaiqing Nie, Bo Zhang 0010, Ji-Rong Wen
ICML3
2007 Webpage understanding: an integrated approach
abstract
Recent work has shown the effectiveness of leveraging layout and tag-tree structure for segmenting webpages and labeling HTML elements. However, how to effectively segment and label the text contents inside HTML elements is still an open problem. Since many text contents on a webpage are often text fragments and not strictly grammatical, traditional natural language processing techniques, that typically expect grammatical sentences, are no longer directly applicable. In this paper, we examine how to use layout and tag-tree structure in a principled way to help understand text contents on webpages. We propose to segment and label the page structure and the text content of a webpage in a joint discriminative probabilistic model. In this model, semantic labels of page structure can be leveraged to help text content understanding, and semantic labels ofthe text phrases can be used in page structure understanding tasks such as data record detection. Thus, integration of both page structure and text content understanding leads to an integrated solution of webpage understanding. Experimental results on research homepage extraction show the feasibility and promise of our approach.
Jun Zhu 0001, Bo Zhang 0010, Zaiqing Nie, Ji-Rong Wen, Hsiao-Wuen Hon
KDD2
2007 Quotient Space Based Multi-granular Analysis
Ling Zhang 0001, Bo Zhang 0010
KSEM2
2007 The importance of query-concept-mapping for automatic video retrieval
abstract
A new video retrieval paradigm of query-by-concept emerges recently. However, it remains unclear how to exploit the detected concepts in retrieval given a multimedia query. In this paper, we point out that it is important to map the query to a few relevant concepts instead of search with all concepts. In addition, we show that solving this problem through both text and image inputs are effective for search, and it is possible to determine the number of related concepts by a language modeling approach. Experimental evidence is obtained on the automatic search task of TRECVID 2006 using a large lexicon of 311 learned semantic concept detectors.
Dong Wang 0022, Xirong Li 0001, Jianmin Li 0001, Bo Zhang 0010
ACM Multimedia4
2007 Gradual transition detection with conditional random fields
abstract
In this paper, we view gradual transition detection as a sequence labeling problem and propose to use Conditional Random Fields (CRFs) for this purpose. CRFs is a state-of-the-art sequence labeling approach. It provides a unified way to integrate various useful clues to form a decision system. Moreover, it has principled way for parameter estimation and inference. Compared to rule-based approaches, gradual transition detection with CRFs requires fewer human interactions while designing the system. The experiments on TRECVID platform show that CRFs can achieve comparable performance to that of the state-of-the-art approaches.
Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010
ACM Multimedia3
2007 Exploiting spatial context constraints for automatic image region annotation
abstract
In this paper we conduct a relatively complete study on how to exploit spatial context constraints for automated image region annotation. We present a straight forward method to regularize the segmented regions into 2D lattice layout, so that simple grid-structure graphical models can be employed to characterize the spatial dependencies. We show how to represent the spatial context constraints in various graphical models and also present the related learning and inference algorithms. Different from most of the existing work, we specifically investigate how to combine the classification performance of discriminative learning and the representation capability of graphical models. To reliably evaluate the proposed approaches, we create a moderate scale image set with region-level ground truth. The experimental results show that (i) spatial context constraints indeed help for accurate region annotation, (ii) the approaches combining the merits of discriminative learning and context constraints perform best, (iii) image retrieval can benefit from accurate region-level annotation.
Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010
ACM Multimedia3
2007 Multi-modal and Multi-granular Learning
Bo Zhang 0010, Ling Zhang 0001
PAKDD1
2007 A Formal Study of Shot Boundary Detection
abstract
This paper conducts a formal study of the shot boundary detection problem. First, a general formal framework of shot boundary detection techniques is proposed. Three critical techniques, i.e., the representation of visual content, the construction of continuity signal and the classification of continuity values, are identified and formulated in the perspective of pattern recognition. Meanwhile, the major challenges to the framework are identified. Second, a comprehensive review of the existing approaches is conducted. The representative approaches are categorized and compared according to their roles in the formal framework. Based on the comparison of the existing approaches, optimal criteria for each module of the framework are discussed, which will provide practical guide for developing novel methods. Third, with all the above issues considered, we present a unified shot boundary detection system based on graph partition model. Extensive experiments are carried out on the platform of TRECVID. The experiments not only verify the optimal criteria discussed above, but also show that the proposed approach is among the best in the evaluation of TRECVID 2005. Finally, we conclude the paper and present some further discussions on what shot boundary detection can learn from other related fields
Jinhui Yuan, Wujie Zheng, Jianmin Li 0001, Fuzong Lin, Bo Zhang 0010
IEEE Trans. Circuits Syst. Video Technol.7
2007 Statistically Robust Detection of Multiplicative Spread-Spectrum Watermarks
abstract
The uncertainties in host signal modeling due to inherent model errors and various attack distortions have prompted the introduction of robust statistics theory in the context of watermark detection. Specifically, the$\epsilon$-contamination model was applied to describe the host signals, and statistically robust (SR) watermark detectors assuming known embedding strengths were derived. In this work, we investigate the robust detection structure for multiplicative watermarking. A detection-simulation (DS)-based approach to determine the contamination factor is also presented. Moreover, considering that the strengths of the watermark signals may be adapted to host signals and will very likely change after being distorted by attacks, we go further to propose the asymptotically robust detector for multiplicative watermarks, which can be viewed as the SR counterpart of the locally most powerful watermark detector in the same sense that the SR detector with full knowledge of the watermark strengths is the corresponding parallel for the optimum detector. Experiments on real images demonstrate the superiority of the new schemes over the conventional ones.
Xingliang Huang, Bo Zhang 0010
IEEE Trans. Inf. Forensics Secur.2
2006 Multiple-Instance Learning Via Random Walk
Dong Wang 0022, Jianmin Li 0001, Bo Zhang 0010
ECML3
2006 Visual Object Recognition in Diverse Scenes with Multiple Instance Learning
abstract
Visual object recognition is important to the robot industry and is a prerequisite for other robot functionalities, such as grasping and manipulation. Object representation and a learning technique are two indispensable parts for this demanding task while arbitrary object appearance and diverse scenes with cluttered background are two great challenges. However, compared with object representation, the learning technique is less developed to deal with these challenges. This paper extends the multiple instance learning (MIL) technique to the multi-class classification scenario and introduces this multi-class MIL framework to the object recognition domain for the first time. This framework is independent of object representation and is useful for object/background discrimination in unseen scenes. Preliminary experiments show that it compares favorably with the supervised learning approach which takes whole images as the classifier training input
Dong Wang 0022, Bo Zhang 0010, Jianwei Zhang 0001
IROS2
2006 A Constructive Learning Algorithm for Text Categorization
Weijun Chen 0001, Bo Zhang 0010
ISNN (2)2
2006 Simultaneous record detection and attribute labeling in web data extraction
abstract
Recent work has shown the feasibility and promise of templateindependent Web data extraction. However, existing approaches use decoupled strategies – attempting to do data record detection and attribute labeling in two separate phases. In this paper, we show that separately extracting data records and attributes is highly ineffective and propose a probabilistic model to perform these two tasks simultaneously. In our approach, record detection can benefit from the availability of semantics required in attribute labeling and, at the same time, the accuracy of attribute labeling can be improved when data records are labeled in a collective manner. The proposed model is called Hierarchical Conditional Random Fields. It can efficiently integrate all useful features by learning their importance, and it can also incorporate hierarchical interactions which are very important for Web data extraction. We empirically compare the proposed model with existing decoupled approaches for product information extraction, and the results show significant improvements in both record detection and attribute labeling.
Jun Zhu 0001, Zaiqing Nie, Ji-Rong Wen, Bo Zhang 0010, Wei-Ying Ma
KDD4
2006 Learning concepts from large scale imbalanced data sets using support cluster machines
abstract
This paper considers the problem of using Support Vector Machines (SVMs) to learn concepts from large scale imbalanced data sets. The objective of this paper is twofold. Firstly, we investigate the effects of large scale and imbalance on SVMs. We highlight the role of linear non-separability in this problem. Secondly, we develop a both practical and theoretical guaranteed meta-algorithm to handle the trouble of scale and imbalance. The approach is named Support Cluster Machines (SCMs). It incorporates the informative and the representative under-sampling mechanisms to speedup the training procedure. The SCMs differs from the previous similar ideas in two ways, (a) the theoretical foundation has been provided, and (b) the clustering is performed in the feature space rather than in the input space. The theoretical analysis not only provides justification, but also guides the technical choices of the proposed approach. Finally, experiments on both the synthetic and the TRECVID data are carried out. The results support the previous analysis and show that the SCMs are efficient and effective while dealing with large scale imbalanced data sets.
Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010
ACM Multimedia3
2006 A semi-supervised incremental learning framework for sports video view classification
abstract
Sports videos have special characteristics such as well-defined video structure, specialized sports syntax, and typically having some canonical view types. In this paper, we propose a semi-supervised incremental learning framework for sports video view classification. Baseball is selected as an example to explain the main ideas. In order to obtain an optimal model based on a small number of pre-labeled training samples, the semi-supervised incremental learning framework explores the local distributed properties of the video sequences and sufficiently utilizes the information of a positive model pool and a negative model pool. After each round of online optimization process for the under-investigating video, a locally-optimized positive model and a set of negative models are added into the positive model pool and the negative model pool according to some heuristic criteria, respectively. Experiments results on real sports video data show that the proposed system is effective and promising
Jun Wu 0022, Bo Zhang 0010, Xian-Sheng Hua 0001, Jianwei Zhang 0001
MMM2
2005 AP-Based Borda Voting Method for Feature Extraction in TRECVID-2004
Dayong Ding, Dong Wang 0022, Fuzong Lin, Bo Zhang 0010
ECIR5
2005 Temporal Shot Clustering Analysis for Video Concept Detection
Dayong Ding, Bo Zhang 0010
ECIR3
2005 Perceptual Watermarking Using a Wavelet Visible Difference Predictor
abstract
This paper proposes a new pixel-wise perceptual mask based on a wavelet visible difference predictor (WVDP) for watermarking. The mask is very effective in that the embedding energy can be sufficiently exerted, and that the annoying global parameter controlling the watermark strength as in usual schemes can be dropped. The watermark sequence is drawn from a uniform distribution and added to all the detail bands after being masked. In the detection phase, a correlation detector is used, and the original image is not required. Experimental results show that our scheme provides very good performance both in terms of watermark unobtrusiveness and robustness.
Xingliang Huang, Bo Zhang 0010
ICASSP (2)2
2005 Two kinds of timing cues and their usage in concept detection in news video
abstract
Two open problems remain unsolved in the content based video retrieval area. Firstly, how to find useful information and express it by different features. Secondly, how to fuse the heterogeneous information together to boost the retrieval performance beyond any single component. The paper presents two kinds of timing information and their use in concept detection in news video, and a novel non-linear information fusion method to combine timing cues with other information from different sources. Experiments on the TRECVID 2004 dataset show that timing cues can boost performance when combined with other information.
Dong Wang 0022, Dayong Ding, Fuzong Lin, Bo Zhang 0010
ICASSP (2)6
2005 2D Conditional Random Fields for Web information extraction
abstract
The Web contains an abundance of useful semistructured information about real world objects, and our empirical study shows that strong sequence characteristics exist for Web information about objects of the same type across different Web sites. Conditional Random Fields (CRFs) are the state of the art approaches taking the sequence characteristics to do better labeling. However, as the information on a Web page is two-dimensionally laid out, previous linear-chain CRFs have their limitations for Web information extraction. To better incorporate the two-dimensional neighborhood interactions, this paper presents a two-dimensional CRF model to automatically extract object information from the Web. We empirically compare the proposed model with existing linear-chain CRF models for product information extraction, and the results show the effectiveness of our model.
Jun Zhu 0001, Zaiqing Nie, Ji-Rong Wen, Bo Zhang 0010, Wei-Ying Ma
ICML4
2005 Robust Detection of Transform Domain Additive Watermarks
Xingliang Huang, Bo Zhang 0010
IWDW2
2005 A unified shot boundary detection framework based on graph partition model
abstract
In this paper, we propose a unified shot boundary detection framework by extending the previous work of graph partition model with temporal constraints. To detect both the abrupt transitions (CUTs) and gradual transitions (GTs, excluding fade out/in) in a unified way, we incorporate temporal multi-resolution analysis into the model. Furthermore, instead of ad-hoc thresholding scheme, we construct a novel kind of feature to characterize shot transitions and employ support vector machine (SVM) with active leaning strategy to classify boundaries and non-boundaries. Extensive experiments have been carried out on the platform of TRECVID benchmark. The experimental results show that the proposed framework outperforms some others and achieves satisfactory results.
Jinhui Yuan, Jianmin Li 0001, Fuzong Lin, Bo Zhang 0010
ACM Multimedia4
2005 Graph Partition Model for Robust Temporal Data Segmentation
Jinhui Yuan, Bo Zhang 0010, Fuzong Lin
PAKDD2
2005 The structure analysis of fuzzy sets
Ling Zhang 0001, Bo Zhang 0010
Int. J. Approx. Reason.2
2005 Fuzzy reasoning model under quotient space structure
Ling Zhang 0001, Bo Zhang 0010
Inf. Sci.2
2005 A Quotient Space Approximation Model of Multiresolution Signal Analysis
Ling Zhang 0001, Bo Zhang 0010
J. Comput. Sci. Technol.2
2005 A unified framework for image retrieval using keyword and visual features
abstract
In this paper, a unified image retrieval framework based on both keyword annotations and visual features is proposed. In this framework, a set of statistical models are built based on visual features of a small set of manually labeled images to represent semantic concepts and used to propagate keywords to other unlabeled images. These models are updated periodically when more images implicitly labeled by users become available through relevance feedback. In this sense, the keyword models serve the function of accumulation and memorization of knowledge learned from user-provided relevance feedback. Furthermore, two sets of effective and efficient similarity measures and relevance feedback schemes are proposed for query by keyword scenario and query by image example scenario, respectively. Keyword models are combined with visual features in these schemes. In particular, a new, entropy-based active learning strategy is introduced to improve the efficiency of relevance feedback for query by keyword. Furthermore, a new algorithm is proposed to estimate the keyword features of the search concept for query by image example. It is shown to be more appropriate than two existing relevance feedback algorithms. Experimental results demonstrate the effectiveness of the proposed framework.
Mingjing Li, HongJiang Zhang, Bo Zhang 0010
IEEE Trans. Image Process.4
2004 Entropy-based active learning with support vector machines for content-based image retrieval
abstract
An entropy-based active learning scheme with support vector machines (SVMs) is proposed for relevance feedback in content-based image retrieval. The main issue in active learning for image retrieval is how to choose images for the user to label in the next interaction. According to information theory, we proposed an entropy-based criterion for good request selection. To apply the criterion with SVMs, probabilistic outputs are required. Since standard SVMs do not provide such outputs, two techniques are used to produce probabilities. One is to train the parameters of an additional sigmoid function. The other is to use the notion of version space. Experimental results on a database of 10,000 general-purpose images demonstrate the effectiveness of the proposed active learning scheme.
Mingjing Li, HongJiang Zhang, Bo Zhang 0010
ICME4
2004 Improvements to Bennett?s Nearest Point Algorithm for Support Vector Machines
Jianmin Li 0001, Jianwei Zhang 0001, Bo Zhang 0010
ISNN (1)3
2004 An online-optimized incremental learning framework for video semantic classification
abstract
This paper considers the problems of feature variation and concept uncertainty in typical learning-based video semantic classification schemes. We proposed a new online semantic classification framework, termed OOIL (for Online-Optimized Incremental Learning), in which two sets of optimized classification models, local and global, are online trained by sufficiently exploiting both local and global statistic characteristics of videos. The global models are pre-trained on a relatively small set of pre-labeled samples. And the local models are optimized for the under-test video or video segment by checking a small portion of unlabeled samples in this video, while they are also applied to incrementally update the global models. Experiments have illustrated promising results on simulated data as well as real sports videos.
Jun Wu 0022, Xian-Sheng Hua 0001, HongJiang Zhang, Bo Zhang 0010
ACM Multimedia4
2004 The Quotient Space Theory of Problem Solving
Ling Zhang 0001, Bo Zhang 0010
Fundam. Informaticae2
2004 Relevance feedback in region-based image retrieval
abstract
Relevance feedback and region-based representations are two effective ways to improve the accuracy of content-based image retrieval systems. Although these two techniques have been successfully investigated and developed in the last few years, little attention has been paid to combining them together. We argue that integrating these two approaches and allowing them to benefit from each other will yield better performance than using either of them alone. To do that, on the one hand, two relevance feedback algorithms are proposed based on region representations. One is inspired from the query point movement method. By assembling all of the segmented regions of positive examples together and reweighting the regions to emphasize the latest ones, a pseudo image is formed as the new query. An incremental clustering technique is also considered to improve the retrieval efficiency. The other is the introduction of existing support vector machine-based algorithms. A new kernel is proposed so as to enable the algorithms to be applicable to region-based representations. On the other hand, a rational region weighting scheme based on users' feedback information is proposed. The region weights that somewhat coincide with human perception not only can be used in a query session, but can also be memorized and accumulated for future queries. Experimental results on a database of 10 000 general-purpose images demonstrate the effectiveness of the proposed framework.
Mingjing Li, HongJiang Zhang, Bo Zhang 0010
IEEE Trans. Circuits Syst. Video Technol.4
2004 An efficient and effective region-based image retrieval framework
abstract
An image retrieval framework that integrates efficient region-based representation in terms of storage and complexity and effective on-line learning capability is proposed. The framework consists of methods for region-based image representation and comparison, indexing using modified inverted files, relevance feedback, and learning region weighting. By exploiting a vector quantization method, both compact and sparse (vector) region-based image representations are achieved. Using the compact representation, an indexing scheme similar to the inverted file technology and an image similarity measure based on Earth Mover's Distance are presented. Moreover, the vector representation facilitates a weighted query point movement algorithm and the compact representation enables a classification-based algorithm for relevance feedback. Based on users' feedback information, a region weighting strategy is also introduced to optimally weight the regions and enable the system to self-improve. Experimental results on a database of 10,000 general-purposed images demonstrate the efficiency and effectiveness of the proposed framework.
Mingjing Li, HongJiang Zhang, Bo Zhang 0010
IEEE Trans. Image Process.4
2003 Support vector machines for region-based image retrieval
abstract
In this paper, the application of support vector machines (SVM) in relevance feedback for region-based image retrieval is investigated. Both the one class SVM as a class distribution estimator and two classes SVM as a classifier are taken into account. For the latter, two representative display strategies are studied. Since the common kernels often rely on inner product or L/sub p/ norm in the input space, they are infeasible in the region-based image retrieval systems that use variable-length representations. To resolve the issue, a new kind of kernel that is a generalization of Gaussian kernel is proposed. Experimental results on a database of 10,000 general-purpose images demonstrate the effectiveness and robustness of the proposed approach.
Mingjing Li, HongJiang Zhang, Bo Zhang 0010
ICME4
2003 Nonlinear Speech Model Based on Support Vector Machine and Wavelet Transform
abstract
To improve the naturalness of reconstructed speech, nonlinear speech models are paid more and more attention in recent years. A nonlinear speech model for speech synthesis based on support vector machine (SVM) is presented firstly. After speech signal is embedded into phase space, nonlinear map in the model is obtained with support vector regression. It is shown in the experiments that for some pieces of speech, not only can speech be perfectly reconstructed by the system, but also jitter and shimmer in the original signal is preserved. However, the output of the system is quite different from the original one for other pieces. The reason is that the sub-bands with different frequency in the original signal can not be perfectly described by a SVM-based autoregressive model trained with one set of training parameters. Consequently, a multi-band model is then proposed. After the original speech is decomposed into several bands through wavelet packet decomposition, a nonlinear dynamical model based on SVM is constructed for each sub-band signal. It is shown in the experiments that the stability of such system is improved.
Jianmin Li 0001, Bo Zhang 0010, Fuzong Lin
ICTAI2
2003 Alternating Feature Spaces in Relevance Feedback
Fang Qian, Mingjing Li, HongJiang Zhang, Wei-Ying Ma, Bo Zhang 0010
Multim. Tools Appl.5
2002 Generation of Chinese prosodic phrasing rules by an extension matrix algorithm
abstract
This paper presents a new rule induction algorithm based on the extension matrix theory, and uses it to learn prosodic phrasing rules automatically for Chinese text-to-speech. Firstly, the basic idea of our algorithm is introduced. Secondly, we collected 937 sentences from news programs and built a corpus for modeling Chinese prosody, a group of feature variables are also proposed. Lastly, the data is divided into two parts: training set and test set, and the experimental results show that our method achieves higher comprehensibility, better accuracy and fewer rules than other algorithms. And the generated rules are quite similar to hand-crafted ones.
Weijun Chen 0001, Fuzong Lin, Jianmin Li 0001, Bo Zhang 0010
ICASSP4
2002 Learning region weighting from relevance feedback in image retrieval
abstract
The region-based approach to image retrieval has emerged as one of the most active research directions in the past few years. There are two crucial problems in the region-based systems: the weighting of regions and the use of relevance feedback. The former plays an important role in computing the region-based similarity of two images, while the latter can improve the efficiency and effectiveness of any CBIR system, if properly employed. In this paper, we propose Key-Region, a novel region weighting scheme that is based on the user's relevance feedback information. The region weight that coincides with human perception can not only be used in a query session, but also be memorized and accumulated for future queries. Experimental results on a database of about 10,000 general-purposed images show the effectiveness of our weighting scheme.
Mingjing Li, HongJiang Zhang, Bo Zhang 0010
ICASSP4
2002 Gaussian mixture model for relevance feedback in image retrieval
abstract
Relevance feedback (RF) has become a powerful technique in content-based image retrieval. Most RF methods assume that positive images follow the single Gaussian distribution, which is not sufficient to model the actual distribution of images due to the gap between the semantic concept and low-level features. In this paper, the Gaussian mixture model (GMM) is applied to represent the distribution of positive images in relevance feedback, and a novel method is proposed to estimate the parameters of the GMM. Both positive and negative examples are used to estimate the number of Gaussian components. Furthermore, due to the lack of training samples, unlabeled data are also incorporated to estimate the covariance matrices. Experimental results show that our GMM-based RF method outperforms that based on a single Gaussian model.
Fang Qian, Mingjing Li, Lei Zhang 0001, HongJiang Zhang, Bo Zhang 0010
ICME (1)5
2002 An effective region-based image retrieval framework
abstract
We present a region-based image retrieval framework that integrates efficient region-based representation in terms of storage and retrieval and effective on-line learning capability. The framework consists of methods for image segmentation and grouping, indexing using modified inverted file, relevance feedback, and continuous learning. By exploiting a vector quantization method, a compact region-based image representation is achieved. Based on this representation, an indexing scheme similar to the inverted file technology is proposed. In addition, it supports relevance feedback based on the vector model with a weighting scheme. A continuous learning strategy is also proposed to enable the system to self improve. Experimental results on a database of 10,000 general-purposed images demonstrate the efficiency and effectiveness of the proposed framework.
Mingjing Li, HongJiang Zhang, Bo Zhang 0010
ACM Multimedia4
2002 Relationship Between Support Vector Set and Kernel Functions in SVM
Ling Zhang 0001, Bo Zhang 0010
J. Comput. Sci. Technol.2
2001 Support vector machine learning for image retrieval
abstract
A novel method of relevance feedback is presented based on support vector machine learning in the content-based image retrieval system. A SVM classifier can be learned from training data of relevance images and irrelevance images marked by users. Using the classifier, the system can retrieve more images relevant to the query in the database efficiently. Experiments were carried out on a large-size database of 9918 images. It shows that the interactive learning and retrieval process can find correct images increasingly. It also shows the generalization ability of SVM under the condition of limited training samples.
Lei Zhang 0001, Fuzong Lin, Bo Zhang 0010
ICIP (2)3
2001 Training prosodic phrasing rules for Chinese TTS systems
abstract
This paper describes several experiments designed to train prosodic phrasing models for Chinese TTS systems and to investigate the underlying rules that control Chinese prosody. First, we collected 559 sentences from news programs and built a large corpus for modeling Chinese prosody. Second, we selected 20 features and used classification and regression trees (CART) and transformational rule-based learning (TRBL) techniques to generate phrasing rules automatically. Lastly, we propose a computer aided error-driven method of designing rule templates, and integrate it into the TRBL algorithm. The experimental results show that we achieve a high success rate of 94.5%, and we also get a set of well comprehensible rule templates which may give us insights into the relationship between Chinese syntax and p rosody.
Weijun Chen 0001, Fuzong Lin, Jianmin Li 0001, Bo Zhang 0010
INTERSPEECH4
2000 A Neural Network Based Classifier for Handwritten Chinese Character Recognition
abstract
In this paper, a geometrical approach for building neural networks is proposed. With the proposed approach, it is very easy to construct an efficient neural classifier to solve the handwritten Chinese character recognition problem, as well as other pattern recognition problems of large scale. Experiments are conducted to evaluate the performance of the proposed approach and results obtained are promising.
Mingrui Wu, Bo Zhang 0010, Ling Zhang 0001
ICPR2
1999 Neural Network Based Classifers for a Vast Amount of Data
Ling Zhang 0001, Bo Zhang 0010
PAKDD2
1999 A geometrical representation of McCulloch-Pitts neural model and its applications
abstract
In this paper, a geometrical representation of McCulloch-Pitts neural model is presented. From the representation, a clear visual picture and interpretation of the model can be seen. Two interesting applications based on the interpretation are discussed. They are 1) a new design principle of feedforward neural networks and 2) a new proof of mapping abilities of three-layer feedforward neural networks.
Ling Zhang 0001, Bo Zhang 0010
IEEE Trans. Neural Networks2
1996 Generating and coding of fractal graphs by neural network and mathematical morphology methods
abstract
We present an algorithm for generating a class of self-similar (fractal) graphs using simple probabilistic logic neuron networks and show that the graphs can be represented by a set of compressed encoding. An algorithm for quickly finding the coding, i.e., recognizing the corresponding graphs, is given and the coding are shown to be optimal (i.e., of minimal length). The same graphs can also be generated by a mathematical morphology method. These results may possibly have applications in image compression and pattern recognition.
Ling Zhang 0001, Bo Zhang 0010
IEEE Trans. Neural Networks2
1995 The generation of a sort of fractal graphs
Bo Zhang 0010, Ling Zhang 0001
J. Comput. Sci. Technol.1
1995 The complexity of learning in PLN networks
Bo Zhang 0010, Ling Zhang 0001, Huai Zhang
Neural Networks1
1995 Programming based learning algorithms of neural networks with self-feedback connections
abstract
Discusses the learning problem of neural networks with self-feedback connections and shows that when the neural network is used as associative memory, the learning problem can be transformed into some sort of programming (optimization) problem. Thus, the rather mature optimization technique in programming mathematics can be used for solving the learning problem of neural networks with self-feedback connections. Two learning algorithms based on programming technique are presented. Their complexity is just polynomial. Then, the optimization of the radius of attraction of the training samples is discussed using quadratic programming techniques and the corresponding algorithm is given. Finally, the comparison is made between the given learning algorithm and some other known algorithms.
Bo Zhang 0010, Ling Zhang 0001, Fachao Wu
IEEE Trans. Neural Networks1
1994 Real-time collision-free path planning for robots in configuration space
Wei Li 0006, Bo Zhang 0010, Hilmar Jaschek
J. Comput. Sci. Technol.2
1993 On memory capacity of the Probabilistic Logic Neuron network
Bo Zhang 0010, Ling Zhang 0001
J. Comput. Sci. Technol.1
1993 The complexity of recognition in the single-layered PLN network with feedback connections
Bo Zhang 0010, Ling Zhang 0001
J. Comput. Sci. Technol.1
1992 A quantitative analysis of the behaviors of the PLN network
Bo Zhang 0010, Ling Zhang 0001, Huai Zhang
Neural Networks1
1990 Hierarchy and statistical heuristic search
Bo Zhang 0010, Ling Zhang 0001
Future Gener. Comput. Syst.1
1989 Motion Planning of Multi-Joint Robotic Arm with Topological Dimension Reduction Method
Bo Zhang 0010, Ling Zhang 0001
IJCAI1
1985 A Weighted Technique in Heuristic Search
Bo Zhang 0010, Ling Zhang 0001
IJCAI1
1985 A New Heuristic Search Technique-Algorithm SA
abstract
In this paper, we present a new heuristic searching algorithm by introducing the statistical inference method on the basis of algorithm A (or A*). It is called algorithm SA. In a simplified search space, a uniform m-ary tree, we obtain the following result. Using algorithm SA, a goal node can be found with probability one, and its mean complexity is O(N·ln N) where N is the depth at which the goal is located.
Bo Zhang 0010, Ling Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
1984 The Successive SA* Search and its Computational Complexity
Ling Zhang 0001, Bo Zhang 0010
ECAI2
1984 Planning Collision-Free Paths for Robotic Arm Among Obstacles
abstract
A theory for planning collision-free paths of a moving object among obstacles is described. Using the concepts of state space and rotation mapping, the relationship between the positions and the corresponding collision-free orientations of a moving object among obstacles is represented as some set of a state space. This set is called the rotation mapping graph (RMG) of that object. The problem of finding collision-free paths for an object translating and rotating among obstacles is thus transformed to that of considering the connectivity of the RMG. Since the connectivity of the graph can be solved by topological methods, the problem of planning collision-free paths is easily solved in theory. Using this theory, a topological method for planning collision-free paths of a rod-object translating and rotating among obstacles is presented. If a nonrigid robotic arm is viewed as a composite rod with some degrees of freedom, the planning of collision-free paths of a robotic arm can be solved in a similar way to a rod.
Robert T. Chien, Ling Zhang 0001, Bo Zhang 0010
IEEE Trans. Pattern Anal. Mach. Intell.3
1983 The Statistical Inference Method in Heuristic Search Techniques
Ling Zhang 0001, Bo Zhang 0010
IJCAI2