EDBT 2026 Demo / reviewers in the wild / expert
Tian Han 0001
dblp:65/4065-1
· DBLP profile ↗
31ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0001-8080-3791ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LDSwap: A Semantic-Related Latent Code Disentangling Method in StyleSpace Towards High-Resolution Face SwappingabstractRecent StyleGAN-based face swapping methods have been able to generate very realistic high-resolution face swapping results, but they are often plagued by the challenge of maintaining various attributes (such as expression, pose, and illumination). One reason is that these methods usually focus on the latent codes of facial semantic features corresponding to the W/W + space, and latent codes in these spaces are often highly entangled. To address this issue, we propose a new method, LDSwap. for disentangling and re-fusing latent codes in the StyleSpace. The semantic-related latent code disentangling module (SLDM) we propose can successfully achieve facial semantic feature exchange and reorganization by disentangling latent codes. In addition, we propose a channel-split adaptive feature fusion module (CAFF) that adaptively learns and refuses spatial information in the target image. This module can learn spatial features from the target image without interference from the features of the target face region. Through qualitative and quantitative evaluation, we demonstrate that LDSwap shows significant improvements over three state-of-the-art methods in maintaining the appearance of semantic features. Chaolin Yang, Lifa Sun, Jinghua Zhong, Tian Han 0001 |
Comput. Vis. Media | 7 |
| 2024 | Learning Multimodal Latent Generative Models with Energy-Based Prior
Shiyu Yuan, Jiali Cui, Hanao Li, Tian Han 0001 |
ECCV (89) | 4 |
| 2024 | Layerwise Change of Knowledge in Neural NetworksabstractThis paper aims to explain how a deep neural network (DNN) gradually extracts new knowledge and forgets noisy features through layers in forward propagation. Up to now, although how to define knowledge encoded by the DNN has not reached a consensus so far, previous studies have derived a series of mathematical evidences to take interactions as symbolic primitive inference patterns encoded by a DNN. We extend the definition of interactions and, for the first time, extract interactions encoded by intermediate layers. We quantify and track the newly emerged interactions and the forgotten interactions in each layer during the forward propagation, which shed new light on the learning behavior of DNNs. The layer-wise change of interactions also reveals the change of the generalization capacity and instability of feature representations of a DNN. Xu Cheng 0005, Lei Cheng 0006, Zhaoran Peng, Tian Han 0001, Quanshi Zhang |
ICML | 5 |
| 2024 | Learning Latent Space Hierarchical EBM Diffusion ModelsabstractThis work studies the learning problem of the energy-based prior model and the multi-layer generator model. The multi-layer generator model, which contains multiple layers of latent variables organized in a top-down hierarchical structure, typically assumes the Gaussian prior model. Such a prior model can be limited in modelling expressivity, which results in a gap between the generator posterior and the prior model, known as the prior hole problem. Recent works have explored learning the energy-based (EBM) prior model as a second-stage, complementary model to bridge the gap. However, the EBM defined on a multi-layer latent space can be highly multi-modal, which makes sampling from such marginal EBM prior challenging in practice, resulting in ineffectively learned EBM. To tackle the challenge, we propose to leverage the diffusion probabilistic scheme to mitigate the burden of EBM sampling and thus facilitate EBM learning. Our extensive experiments demonstrate a superior performance of our diffusion-learned EBM prior on various challenging tasks. Jiali Cui, Tian Han 0001 |
ICML | 2 |
| 2024 | Improving Adversarial Energy-Based Model via Diffusion ProcessabstractGenerative models have shown strong generation ability while efficient likelihood estimation is less explored. Energy-based models (EBMs) define a flexible energy function to parameterize unnormalized densities efficiently but are notorious for being difficult to train. Adversarial EBMs introduce a generator to form a minimax training game to avoid expensive MCMC sampling used in traditional EBMs, but a noticeable gap between adversarial EBMs and other strong generative models still exists. Inspired by diffusion-based models, we embedded EBMs into each denoising step to split a long-generated process into several smaller steps. Besides, we employ a symmetric Jeffrey divergence and introduce a variational posterior distribution for the generator's training to address the main challenges that exist in adversarial EBMs. Our experiments show significant improvement in generation compared to existing adversarial EBMs, while also providing a useful energy function for efficient density estimation. Cong Geng, Tian Han 0001, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Søren Hauberg, Bo Li 0115 |
ICML | 2 |
| 2023 | Learning Joint Latent Space EBM Prior Model for Multi-layer GeneratorabstractThis paper studies the fundamental problem of learning multi-layer generator models. The multi-layer generator model builds multiple layers of latent variables as a prior model on top of the generator, which benefits learning complex data distribution and hierarchical representations. However, such a prior model usually focuses on modeling inter-layer relations between latent variables by assuming non-informative (conditional) Gaussian distributions, which can be limited in model expressivity. To tackle this issue and learn more expressive prior models, we propose an energy-based model (EBM) on the joint latent space over all layers of latent variables with the multi-layer generator as its backbone. Such joint latent space EBM prior model captures the intra-layer contextual relations at each layer through layer-wise energy terms, and latent variables across different layers are jointly corrected. We develop a joint training scheme via maximum likelihood estimation (MLE), which involves Markov Chain Monte Carlo (MCMC) sampling for both prior and posterior distributions of the latent variables from different layers. To ensure efficient inference and learning, we further propose a variational training scheme where an inference model is used to amortize the costly posterior MCMC sampling. Our experiments demonstrate that the learned model can be expressive in generating high-quality images and capturing hierarchical features for better outlier detection. Jiali Cui, Ying Nian Wu, Tian Han 0001 |
CVPR | 3 |
| 2023 | Learning Hierarchical Features with Joint Latent Space Energy-Based PriorabstractThis paper studies the fundamental problem of multilayer generator models in learning hierarchical representations. The multi-layer generator model that consists of multiple layers of latent variables organized in a top-down architecture tends to learn multiple levels of data abstraction. However, such multi-layer latent variables are typically parameterized to be Gaussian, which can be less informative in capturing complex abstractions, resulting in limited success in hierarchical representation learning. On the other hand, the energy-based (EBM) prior is known to be expressive in capturing the data regularities, but it often lacks the hierarchical structure to capture different levels of hierarchical representations. In this paper, we propose a joint latent space EBM prior model with multi-layer latent variables for effective hierarchical representation learning. We develop a variational joint learning scheme that seamlessly integrates an inference model for efficient inference. Our experiments demonstrate that the proposed joint EBM prior is effective and expressive in capturing hierarchical representations and modelling data distribution. Jiali Cui, Ying Nian Wu, Tian Han 0001 |
ICCV | 3 |
| 2023 | Learning Energy-based Model via Dual-MCMC TeachingabstractThis paper studies the fundamental learning problem of the energy-based model (EBM). Learning the EBM can be achieved using the maximum likelihood estimation (MLE), which typically involves the Markov Chain Monte Carlo (MCMC) sampling, such as the Langevin dynamics. However, the noise-initialized Langevin dynamics can be challenging in practice and hard to mix. This motivates the exploration of joint training with the generator model where the generator model serves as a complementary model to bypass MCMC sampling. However, such a method can be less accurate than the MCMC and result in biased EBM learning. While the generator can also serve as an initializer model for better MCMC sampling, its learning can be biased since it only matches the EBM and has no access to empirical training examples. Such biased generator learning may limit the potential of learning the EBM. To address this issue, we present a joint learning framework that interweaves the maximum likelihood learning algorithm for both the EBM and the complementary generator model. In particular, the generator model is learned by MLE to match both the EBM and the empirical data distribution, making it a more informative initializer for MCMC sampling of EBM. Learning generator with observed examples typically requires inference of the generator posterior. To ensure accurate and efficient inference, we adopt the MCMC posterior sampling and introduce a complementary inference model to initialize such latent MCMC sampling. We show that three separate models can be seamlessly integrated into our joint framework through two (dual-) MCMC teaching, enabling effective and efficient EBM learning. Jiali Cui, Tian Han 0001 |
NeurIPS | 2 |
| 2023 | Molecule Design by Latent Space Energy-Based Modeling and Gradual Distribution ShiftingabstractGeneration of molecules with desired chemical and biological properties such as high drug-likeness, high binding affinity to target proteins, is critical for drug discovery. In this paper, we propose a probabilistic generative model to capture the joint distribution of molecules and their properties. Our model assumes an energy-based model (EBM) in the latent space. Conditional on the latent vector, the molecule and its properties are modeled by a molecule generation model and a property regression model respectively. To search for molecules with desired properties, we propose a sampling with gradual distribution shifting (SGDS) algorithm, so that after learning the model initially on the training data of existing molecules and their properties, the proposed algorithm gradually shifts the model distribution towards the region supported by molecules with desired values of properties. Our experiments show that our method achieves very strong performances on various molecule design tasks. Deqian Kong, Bo Pang 0004, Tian Han 0001, Ying Nian Wu |
UAI | 3 |
| 2022 | Context-Aware Health Event Prediction via Transition Functions on Dynamic Disease GraphsabstractWith the wide application of electronic health records (EHR) in healthcare facilities, health event prediction with deep learning has gained more and more attention. A common feature of EHR data used for deep-learning-based predictions is historical diagnoses. Existing work mainly regards a diagnosis as an independent disease and does not consider clinical relations among diseases in a visit. Many machine learning approaches assume disease representations are static in different visits of a patient. However, in real practice, multiple diseases that are frequently diagnosed at the same time reflect hidden patterns that are conducive to prognosis. Moreover, the development of a disease is not static since some diseases can emerge or disappear and show various symptoms in different visits of a patient. To effectively utilize this combinational disease information and explore the dynamics of diseases, we propose a novel context-aware learning framework using transition functions on dynamic disease graphs. Specifically, we construct a global disease co-occurrence graph with multiple node properties for disease combinations. We design dynamic subgraphs for each patient's visit to leverage global and local contexts. We further define three diagnosis roles in each visit based on the variation of node properties to model disease transition processes. Experimental results on two real-world EHR datasets show that the proposed model outperforms state of the art in predicting health events. Chang Lu 0004, Tian Han 0001, Yue Ning 0001 |
AAAI | 2 |
| 2022 | Learning from the Tangram to Solve Mini Visual TasksabstractCurrent pre-training methods in computer vision focus on natural images in the daily-life context. However, abstract diagrams such as icons and symbols are common and important in the real world. We are inspired by Tangram, a game that requires replicating an abstract pattern from seven dissected shapes. By recording human experience in solving tangram puzzles, we present the Tangram dataset and show that a pre-trained neural model on the Tangram helps solve some mini visual tasks based on low-resolution vision. Extensive experiments demonstrate that our proposed method generates intelligent solutions for aesthetic tasks such as folding clothes and evaluating room layouts. The pre-trained feature extractor can facilitate the convergence of few-shot learning tasks on human handwriting and improve the accuracy in identifying icons by their contours. The Tangram dataset is available at https://github.com/yizhouzhao/Tangram. Liang Qiu 0001, Pan Lu, Feng Shi 0006, Tian Han 0001, Song-Chun Zhu |
AAAI | 5 |
| 2022 | Adaptive Multi-stage Density Ratio Estimation for Learning Latent Space Energy-based ModelabstractThis paper studies the fundamental problem of learning energy-based model (EBM) in the latent space of the generator model. Learning such prior model typically requires running costly Markov Chain Monte Carlo (MCMC). Instead, we propose to use noise contrastive estimation (NCE) to discriminatively learn the EBM through density ratio estimation between the latent prior density and latent posterior density. However, the NCE typically fails to accurately estimate such density ratio given large gap between two densities. To effectively tackle this issue and further learn more expressive prior model, we develop the adaptive multi-stage density ratio estimation which breaks the estimation into multiple stages and learn different stages of density ratio sequentially and adaptively. The latent prior model can be gradually learned using ratio estimated in previous stage so that the final latent space EBM prior can be naturally formed by product of ratios in different stages. The proposed method enables informative and much sharper prior than existing baselines, and can be trained efficiently. Our experiments demonstrate strong performances in terms of image generation and reconstruction as well as anomaly detection. Zhisheng Xiao, Tian Han 0001 |
NeurIPS | 2 |
| 2022 | Deformable Generator Networks: Unsupervised Disentanglement of Appearance and GeometryabstractWe present a deformable generator model to disentangle the appearance and geometric information for both image and video data in a purely unsupervised manner. The appearance generator network models the information related to appearance, including color, illumination, identity or category, while the geometric generator performs geometric warping, such as rotation and stretching, through generating deformation field which is used to warp the generated appearance to obtain the final image or video sequences. Two generators take independent latent vectors as input to disentangle the appearance and geometric information from image or video sequences. For video data, a nonlinear transition model is introduced to both the appearance and geometric generators to capture the dynamics over time. The proposed scheme is general and can be easily integrated into different generative models. An extensive set of qualitative and quantitative experiments shows that the appearance and geometric information can be well disentangled, and the learned geometric generator can be conveniently transferred to other image datasets that share similar structure regularity to facilitate knowledge transfer tasks. Xianglei Xing, Ruiqi Gao, Tian Han 0001, Song-Chun Zhu, Ying Nian Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Generative Text Modeling through Short Run InferenceabstractLatent variable models for text, when trained successfully, accurately model the data distribution and capture global semantic and syntactic features of sentences.The prominent approach to train such models is variational autoencoders (VAE).It is nevertheless challenging to train and often results in a trivial local optimum where the latent variable is ignored and its posterior collapses into the prior, an issue known as posterior collapse.Various techniques have been proposed to mitigate this issue.Most of them focus on improving the inference model to yield latent codes of higher quality.The present work proposes a short run dynamics for inference.It is initialized from the prior distribution of the latent variable and then runs a small number (e.g., 20) of Langevin dynamics steps guided by its posterior distribution.The major advantage of our method is that it does not require a separate inference model or assume simple geometry of the posterior distribution, thus rendering an automatic, natural and flexible inference engine.We show that the models trained with short run dynamics more accurately model the data, compared to strong language model and VAE baselines, and exhibit no sign of posterior collapse.Analyses of the latent space show that interpolation in the latent space is able to generate coherent sentences with smooth transition and demonstrate improved classification over strong baselines with latent features from unsupervised pretraining.These results together expose a well-structured latent space of our generative model. Bo Pang 0004, Erik Nijkamp, Tian Han 0001, Ying Nian Wu |
EACL | 3 |
| 2021 | Lavs: A Lightweight Audio-Visual Saliency Prediction ModelabstractAudio information is essential for guiding human attention and visual perception, which has been verified by many comprehensive psychological studies. However, the audio modality has been rather neglected in modeling visual attention, most of the current visual attention models heavily depend on visual information. Additionally, current existing high-performing visual attention models rely on deeper convolution neural networks (CNNs), benefiting from their extraordinary feature learning ability but incurring high computational cost. To this end, we propose a novel lightweight audio-visual saliency (LAVS) model to efficiently address the problem of fixation prediction in videos. To the best of our knowledge, our proposed model constitutes the first attempt to exploit a lightweight network and combines the visual and audio cues to perform saliency estimation in videos. Specifically, our proposed model consists of four modules, which are spatial-temporal visual saliency estimation module, audio features extraction module, source sound localization module, and audio-visual saliency fusion module. Extensive experiments across datasets validate the effectiveness and real-time performance of the proposed LAVS model, which outperforms the other state-of-the-art methods. Dandan Zhu 0001, Defang Zhao, Xiongkuo Min, Tian Han 0001, Qiangqiang Zhou, Shaobo Yu, Yongqing Chen, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 4 |
| 2021 | Towards multi-scale deep features learning with correlation metric for person re-identification
Dandan Zhu 0001, Qiangqiang Zhou, Tian Han 0001, Yongqing Chen, Defang Zhao, Xiaokang Yang 0001 |
Knowl. Based Syst. | 3 |
| 2020 | On the Anatomy of MCMC-Based Maximum Likelihood Learning of Energy-Based ModelsabstractThis study investigates the effects of Markov chain Monte Carlo (MCMC) sampling in unsupervised Maximum Likelihood (ML) learning. Our attention is restricted to the family of unnormalized probability densities for which the negative log density (or energy function) is a ConvNet. We find that many of the techniques used to stabilize training in previous studies are not necessary. ML learning with a ConvNet potential requires only a few hyper-parameters and no regularization. Using this minimal framework, we identify a variety of ML learning outcomes that depend solely on the implementation of MCMC sampling.On one hand, we show that it is easy to train an energy-based model which can sample realistic images with short-run Langevin. ML can be effective and stable even when MCMC samples have much higher energy than true steady-state samples throughout training. Based on this insight, we introduce an ML method with purely noise-initialized MCMC, high-quality short-run synthesis, and the same budget as ML with informative MCMC initialization such as CD or PCD. Unlike previous models, our energy model can obtain realistic high-diversity samples from a noise signal after training.On the other hand, ConvNet potentials learned with non-convergent MCMC do not have a valid steady-state and cannot be considered approximate unnormalized densities of the training data because long-run MCMC samples differ greatly from observed images. We show that it is much harder to train a ConvNet potential to learn a steady-state over realistic images. To our knowledge, long-run MCMC samples of all previous models lose the realism of short-run samples. With correct tuning of Langevin noise, we train the first ConvNet potentials for which long-run and steady-state MCMC samples are realistic images. Erik Nijkamp, Mitch Hill, Tian Han 0001, Song-Chun Zhu, Ying Nian Wu |
AAAI | 3 |
| 2020 | Joint Training of Variational Auto-Encoder and Latent Energy-Based ModelabstractThis paper proposes a joint training method to learn both the variational auto-encoder (VAE) and the latent energy-based model (EBM). The joint training of VAE and latent EBM are based on an objective function that consists of three Kullback-Leibler divergences between three joint distributions on the latent vector and the image, and the objective function is of an elegant symmetric and anti-symmetric form of divergence triangle that seamlessly integrates variational and adversarial learning. In this joint training scheme, the latent EBM serves as a critic of the generator model, while the generator model and the inference model in VAE serve as the approximate synthesis sampler and inference sampler of the latent EBM. Our experiments show that the joint training greatly improves the synthesis quality of the VAE. It also enables learning of an energy function that is capable of detecting out of sample examples for anomaly detection. Tian Han 0001, Erik Nijkamp, Linqi Zhou, Bo Pang 0004, Song-Chun Zhu, Ying Nian Wu |
CVPR | 1 |
| 2020 | Learning Multi-layer Latent Variable Model via Variational Optimization of Short Run MCMC for Approximate Inference
Erik Nijkamp, Bo Pang 0004, Tian Han 0001, Linqi Zhou, Song-Chun Zhu, Ying Nian Wu |
ECCV (6) | 3 |
| 2020 | Ransp: Ranking Attention Network For Saliency Prediction On Omnidirectional ImagesabstractVarious convolutional neural network (CNN)-based methods have shown the ability to boost the performance of saliency prediction on omnidirectional images (ODIs). However, these methods are limited by sub-optimal accuracy, because not all the features extracted by the CNN model are not useful for the final fine-grained saliency prediction. Features are redundant and have negative impact on the final fine-grained saliency prediction. To tackle this problem, we propose a novel Ranking Attention Network for saliency prediction (RANSP) of head fixations on ODIs. Specifically, the part-guided attention (PA) module and channel-wise feature (CF) extraction module are integrated in a unified framework and are trained in an end-to-end manner for fine-grained saliency prediction. To better utilize the channel-wise feature map, we further propose a new Ranking Attention Module (RAM), which automatically ranks and selects these maps based on scores for fine-grained saliency prediction. Extensive experiments are conducted to show the effectiveness of our method for saliency prediction of ODIs. Dandan Zhu 0001, Yongqing Chen, Tian Han 0001, Defang Zhao, Yucheng Zhu, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 3 |
| 2020 | Saliency Prediction on Omnidirectional Images with Brain-Like Shallow Neural NetworkabstractDeep feedforward convolutional neural networks (CNNs) perform well in the saliency prediction of omnidirectional images (ODIs), and have become the leading class of candidate models of the visual processing mechanism in the primate ventral stream. These CNNs have evolved from shallow network architecture to extremely deep and branching architecture to achieve superb performance in various vision tasks, yet it is unclear how brain-like they are. In particular, these deep feedforward CNNs are difficult to mapping to ventral stream structure of the brain visual system due to their vast number of layers and missing biologically-important connections, such as recurrence. To tackle this issue, some brain-like shallow neural networks are introduced. In this paper, we propose a novel brain-like network model for saliency prediction of head fixations on ODIs. Specifically, our proposed model consists of three modules: a CORnet-S module, a template feature extraction module and a ranking attention module (RAM). The CORnet-S module is a lightweight artificial neural network (ANN) with four anatomically mapped areas (V1, V2, V4 and IT) and it can simulate the visual processing mechanism of ventral visual stream in the human brain. The template features extraction module is introduced to extract attention maps of ODIs and provide guidance for the feature ranking in the following RAM module. The RAM module is used to rank and select features that are important for fine-grained saliency prediction. Extensive experiments have validated the effectiveness of the proposed model in predicting saliency maps of ODIs, and the proposed model outperforms other state-of-the-art methods with similar scale. Dandan Zhu 0001, Yongqing Chen, Xiongkuo Min, Defang Zhao, Yucheng Zhu, Qiangqiang Zhou, Xiaokang Yang 0001, Tian Han 0001 |
ICPR | 8 |
| 2020 | Learning Latent Space Energy-Based Prior ModelabstractWe propose an energy-based model (EBM) in the latent space of a generator model, so that the EBM serves as a prior model that stands on the top-down network of the generator model. Both the latent space EBM and the top-down network can be learned jointly by maximum likelihood, which involves short-run MCMC sampling from both the prior and posterior distributions of the latent vector. Due to the low dimensionality of the latent space and the expressiveness of the top-down network, a simple EBM in latent space can capture regularities in the data effectively, and MCMC sampling in latent space is efficient and mixes well. We show that the learned model exhibits strong performances in terms of image and text generation and anomaly detection. The one-page code can be found in supplementary materials. Bo Pang 0004, Tian Han 0001, Erik Nijkamp, Song-Chun Zhu, Ying Nian Wu |
NeurIPS | 2 |
| 2019 | Divergence Triangle for Joint Training of Generator Model, Energy-Based Model, and Inferential ModelabstractThis paper proposes the divergence triangle as a framework for joint training of a generator model, energy-based model and inference model. The divergence triangle is a compact and symmetric (anti-symmetric) objective function that seamlessly integrates variational learning, adversarial learning, wake-sleep algorithm, and contrastive divergence in a unified probabilistic formulation. This unification makes the processes of sampling, inference, and energy evaluation readily available without the need for costly Markov chain Monte Carlo methods. Our experiments demonstrate that the divergence triangle is capable of learning (1) an energy-based model with well-formed energy landscape, (2) direct sampling in the form of a generator network, and (3) feed-forward inference that faithfully reconstructs observed as well as synthesized data. Tian Han 0001, Erik Nijkamp, Xiaolin Fang 0002, Mitch Hill, Song-Chun Zhu, Ying Nian Wu |
CVPR | 1 |
| 2019 | Unsupervised Disentangling of Appearance and Geometry by Deformable Generator NetworkabstractWe present a deformable generator model to disentangle the appearance and geometric information in purely unsupervised manner. The appearance generator models the appearance related information, including color, illumination, identity or category, of an image, while the geometric generator performs geometric related warping, such as rotation and stretching, through generating displacement of the coordinates of each pixel to obtain the final image. Two generators act upon independent latent factors to extract disentangled appearance and geometric information from image. The proposed scheme is general and can be easily integrated into different generative models. An extensive set of qualitative and quantitative experiments show that the appearance and geometric information can be well disentangled, and the learned geometric generator can be conveniently transferred to the other image datasets to facilitate knowledge transfer tasks. Xianglei Xing, Tian Han 0001, Ruiqi Gao, Song-Chun Zhu, Ying Nian Wu |
CVPR | 2 |
| 2019 | Learning Generator Networks for Dynamic PatternsabstractWe address the problem of learning dynamic patterns from unlabeled video sequences, either in the form of generating new video sequences, or recovering incomplete video sequences. This problem is challenging because the appearances and motions in the video sequences can be very complex. We propose to use the alternating back-propagation algorithm to learn the generator network with the spatial-temporal convolutional architecture. The proposed method is efficient and flexible. It can not only generate realistic video sequences, but can also recover the incomplete video sequences in the testing stage or even in the learning stage. The proposed algorithm can be further improved by using learned initialization which is useful for the recovery tasks. Further, the proposed algorithm can naturally help to learn the shared representation between different modalities. Our experiments show that our method is competitive with the existing state of the art methods both qualitatively and quantitatively. Tian Han 0001, Yang Lu 0006, Xianglei Xing, Ying Nian Wu |
WACV | 1 |
| 2019 | Replicating Neuroscience Observations on ML/MF and AM Face Patches by Deep Generative Modelabstractpaper (Chang & Tsao, 2017) reports an interesting discovery. For the face stimuli generated by a pretrained active appearance model (AAM), the responses of neurons in the areas of the primate brain that are responsible for face recognition exhibit a strong linear relationship with the shape variables and appearance variables of the AAM that generates the face stimuli. In this letter, we show that this behavior can be replicated by a deep generative model, the generator network, that assumes that the observed signals are generated by latent random variables via a top-down convolutional neural network. Specifically, we learn the generator network from the face images generated by a pretrained AAM model using a variational autoencoder, and we show that the inferred latent variables of the learned generator network have a strong linear relationship with the shape and appearance variables of the AAM model that generates the face images. Unlike the AAM model, which has an explicit shape model where the shape variables generate the control points or landmarks, the generator network has no such shape model and shape variables. Yet it can learn the shape knowledge in the sense that some of the latent variables of the learned generator network capture the shape variations in the face images generated by AAM. Tian Han 0001, Xianglei Xing, Ying Nian Wu |
Neural Comput. | 1 |
| 2018 | Learning Multi-view Generator Network for Shared RepresentationabstractMulti-view representation learning is challenging because different views contain both the common structure and the complex view specific information. The traditional generative models may not be effective in such situation, since view-specific and common information cannot be well separated, which may cause problems for downstream vision tasks. In this paper, we introduce a multi-view generator model to solve the problem of multi-view generation and recognition in a unified framework. We propose a multi-view alternating back-propagation algorithm to learn multi-view generator networks by allowing them to share common latent factors. Our experiments show that the proposed method is effective for both image generation and recognition. Specifically, we first qualitatively demonstrate that our model can rotate and complete faces accurately. Then we show that our model can achieve state-of-art or competitive recognition performances through quantitative comparisons. Tian Han 0001, Xianglei Xing, Ying Nian Wu |
ICPR | 1 |
| 2018 | Replicating Active Appearance Model by Generator NetworkabstractA recent Cell paper [Chang and Tsao, 2017] reports an interesting discovery. For the face stimuli generated by a pre-trained active appearance model (AAM), the responses of neurons in the areas of the primate brain that are responsible for face recognition exhibit strong linear relationship with the shape variables and appearance variables of the AAM that generates the face stimuli. In this paper, we show that this behavior can be replicated by a deep generative model called the generator network, which assumes that the observed signals are generated by latent random variables via a top-down convolutional neural network. Specifically, we learn the generator network from the face images generated by a pre-trained AAM model using variational auto-encoder, and we show that the inferred latent variables of the learned generator network have strong linear relationship with the shape and appearance variables of the AAM model that generates the face images. Unlike the AAM model that has an explicit shape model where the shape variables generate the control points or landmarks, the generator network has no such shape model and shape variables. Yet the generator network can learn the shape knowledge in the sense that some of the latent variables of the learned generator network capture the shape variations in the face images generated by AAM. Tian Han 0001, Ying Nian Wu |
IJCAI | 1 |
| 2017 | Alternating Back-Propagation for Generator NetworkabstractThis paper proposes an alternating back-propagation algorithm for learning the generator network model. The model is a non-linear generalization of factor analysis. In this model, the mapping from the continuous latent factors to the observed signal is parametrized by a convolutional neural network. The alternating back-propagation algorithm iterates the following two steps: (1) Inferential back-propagation, which infers the latent factors by Langevin dynamics or gradient descent. (2) Learning back-propagation, which updates the parameters given the inferred latent factors by gradient descent. The gradient computations in both steps are powered by back-propagation, and they share most of their code in common. We show that the alternating back-propagation algorithm can learn realistic generator models of natural images, video sequences, and sounds. Moreover, it can also be used to learn from incomplete or indirect training data. Tian Han 0001, Yang Lu 0006, Song-Chun Zhu, Ying Nian Wu |
AAAI | 1 |
| 2012 | Quasi-regular Facade Structure Extraction
Tian Han 0001, Chiew-Lan Tai, Long Quan |
ACCV (4) | 1 |
| 2012 | Parsing façade with rank-one approximationabstractThe binary split grammar is powerful to parse façade in a broad range of types, whose structure is characterized by repetitive patterns with different layouts. We notice that, as far as two labels are concerned, BSG parsing is equivalent to approximating a façade by a matrix with multiple rank-one patterns. Then, we propose an efficient algorithm to decompose an arbitrary matrix into a rank-one matrix and a residual matrix, whose magnitude is small in the sense of l0-norm. Next, we develop a block-wise partition method to parse a more general façade. Our method leverages on the recent breakthroughs in convex optimization that can effectively decompose a matrix into a low-rank and sparse matrix pair. The rank-one block-wise parsing not only leads to the detection of repetitive patterns, but also gives an accurate façade segmentation. Experiments on intensive façade data sets have demonstrated that our method outperforms the state-of-the-art techniques and benchmarks both in robustness and efficiency. Tian Han 0001, Long Quan, Chiew-Lan Tai |
CVPR | 2 |