EDBT 2026 Demo / reviewers in the wild / expert
Han Zhang 0010
dblp:26/4189-10
· DBLP profile ↗
55ranked-venue papers
8as first author
38since 2021 · last 2026
0000-0001-7072-2189ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 7 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 5 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DICE: Discrete Inversion Enabling Controllable Editing for Masked Generative ModelsabstractRecent advances in discrete diffusion models have demonstrated strong performance in image generation and masked language modeling, yet they remain limited in their capacity for controlled content editing. We propose DICE (Discrete Inversion for Controllable Editing), a novel framework that pioneers precise inversion capabilities for discrete diffusion models, including both masked generative and multinomial diffusion variants. Our key innovation lies in capturing noise sequences and masking patterns during reverse diffusion process, enabling both accurate reconstruction and flexible editing without relying on predefined masks or attention-based manipulations. Through comprehensive experiments across image and text modalities using models such as Paella, VQ-Diffusion, RoBERTa and LLaDA, we demonstrate that DICE successfully maintains high fidelity to the original data while significantly expanding editing capabilities. These results establish new possibilities for fine-grained content manipulation in discrete spaces. Xiaoxiao He, Quan Dao, Ligong Han, Song Wen 0001, Minhao Bai, Di Liu 0003, Han Zhang 0010, Felix Juefei-Xu, Chaowei Tan, Bo Liu 0005, Martin Renqiang Min, Kang Li 0004, Faez Ahmed, Akash Srivastava, Hongdong Li, Junzhou Huang, Dimitris N. Metaxas |
WACV | 7 |
| 2026 | Source-Free Active Domain Adaptation via Influential-Points-Guided Progressive Teacher for Medical Image SegmentationabstractDomain adaptation in medical image segmentation enables pre-trained models to generalize to new target domains. Given limited annotated data and privacy constraints, Source-Free Active Domain Adaptation (SFADA) methods provide promising solutions by selecting a few target samples for labeling without accessing source samples. However, in a fully source-free setting, existing works have not fully explored how to select these target samples in a class-balanced manner and how to conduct robust model adaptation using both labeled and unlabeled samples. In this study, we discover that boundary samples with source-like semantics but sharp predictive discrepancies are beneficial for SFADA. We define these samples as the most influential points and propose a slice-wise framework using influential points learning to explore them. Specifically, we detect source-like samples to retain source-specific knowledge. For each target sample, an adaptive K-nearest neighbor algorithm based on local density is introduced to construct neighborhoods of source-like samples for knowledge transfer. We then propose a class-balanced Kullback-Leibler divergence for these neighborhoods, calculating it to obtain an influential score ranking. A diverse subset of the highest-ranked target samples (considered influential points) is manually annotated. Furthermore, we design a progressive teacher model to facilitate SFADA for medical image segmentation. With the guidance of influential points, this model independently generates and utilizes pseudo-labels to mitigate error accumulation. To further suppress noise, curriculum learning is incorporated into the model to progressively leverage reliable supervision signals from pseudo-labels. Experiments on multiple benchmarks demonstrate that our method outperforms state-of-the-art methods even with only 2.5% of the labeling budget. Yong Chen 0024, Xiangde Luo, Renyi Chen, Yiyue Li, Han Zhang 0010, He Lyu, Huan Song, Kang Li 0004 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | BACON: Improving Clarity of Image Captions via Bag-of-Concept GraphsabstractAdvancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions often carry lengthy, intertwined contexts that are difficult to parse and frequently overlook essential cues, posing a great barrier for models like GroundingDINO and SDXL, which lack the strong text encoding and syntax analysis needed to fully leverage dense captions. To address this, we propose BACON, a prompting method that breaks down VLM-generated captions into disentangled, structured elements such as objects, relationships, styles, and themes. This approach not only minimizes confusion from handling complex contexts but also allows for efficient transfer into a JSON dictionary, enabling models without linguistic processing capabilities to easily access key information. We annotated 100,000 image-caption pairs using BACON with GPT-4V and trained an LLaVA captioner on this dataset, enabling it to produce BACON-style captions without relying on costly GPT-4V. Evaluations of overall quality, precision, and recall—as well as user studies—demonstrate that the resulting caption model consistently outperforms other SOTA VLM models in generating high-quality captions. Besides, we show that BACON-style captions exhibit better clarity when applied to various models, enabling them to accomplish previously unattainable tasks or surpass existing SOTA solutions without training. For example, BACON-style captions help GroundingDINO achieve 1.51× higher recall scores on open-vocabulary object detection tasks compared to leading methods. Zhantao Yang, Ruili Feng, Huangji Wang, Zhicai Wang, Shangwen Zhu, Han Zhang 0010, Jie Xiao 0002, Pingyu Wu, Kai Zhu 0004, Jixuan Chen, Chen-Wei Xie, Hongyang Zhang 0001, Yu Liu 0063, Fan Cheng 0002 |
CVPR | 7 |
| 2025 | Accelerating Diffusion Sampling via Exploiting Local Transition CoherenceabstractText-based diffusion models have made significant breakthroughs in generating high-quality images and videos from textual descriptions. However, the lengthy sampling time of the denoising process remains a significant bottleneck in practical applications. Previous methods either ignore the statistical relationships between adjacent steps or rely on attention or feature similarity between them, which often only works with specific network structures. To address this issue, we discover a new statistical relationship in the transition operator between adjacent steps, focusing on the relationship of the outputs from the network. This relationship does not impose any requirements on the network structure. Based on this observation, we propose a novel training-free acceleration method called LTC-Accel, which uses the identified relationship to estimate the current transition operator based on adjacent steps. Due to no specific assumptions regarding the network structure, LTC-Accel is applicable to almost all diffusion-based methods and orthogonal to almost all existing acceleration techniques, making it easy to combine with them. Experimental results demonstrate that LTC-Accel significantly speeds up sampling in text-to-image and text-to-video synthesis while maintaining competitive sample quality. Specifically, LTC-Accel achieves a speedup of 1.67-fold in Stable Diffusion v2 and a speedup of 1.55-fold in video generation models. When combined with distillation models, LTC-Accel achieves a remarkable 10-fold speedup in video generation, allowing real-time generation of more than 16FPS. Shangwen Zhu, Han Zhang 0010, Zhantao Yang, Qianyu Peng, Zhao Pu, Huangji Wang, Fan Cheng 0002 |
ICCV | 2 |
| 2025 | Epsilon-VAE: Denoising as Visual DecodingabstractIn generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data, it reduces redundancy and emphasizes key features for high-quality generation. Current visual tokenization methods rely on a traditional autoencoder framework, where the encoder compresses data into latent representations, and the decoder reconstructs the original input. In this work, we offer a new perspective by proposing denoising as decoding, shifting from single-step reconstruction to iterative refinement. Specifically, we replace the decoder with a diffusion process that iteratively refines noise to recover the original image, guided by the latents provided by the encoder. We evaluate our approach by assessing both reconstruction (rFID) and generation quality (FID), comparing it to state-of-the-art autoencoding approaches. By adopting iterative reconstruction through diffusion, our autoencoder, namely Epsilon-VAE, achieves high reconstruction quality, which in turn enhances downstream generation quality by 22% at the same compression rates or provides 2.3x inference speedup through increasing compression rates. We hope this work offers new insights into integrating iterative generation and autoencoding for improved compression and generation. Long Zhao 0003, Sanghyun Woo, Ziyu Wan, Yandong Li, Han Zhang 0010, Boqing Gong, Hartwig Adam, Xuhui Jia, Ting Liu 0005 |
ICML | 5 |
| 2025 | DiffOSeg: Omni Medical Image Segmentation via Multi-Expert Collaboration Diffusion Model
Han Zhang 0010, Xiangde Luo, Yong Chen 0024, Kang Li 0004 |
MICCAI (13) | 1 |
| 2025 | The Matrix: Infinite-Horizon World Generation with Real-Time Moving ControlabstractWe present The Matrix, a foundational realistic world simulator capable of generating infinitely long 720p high-fidelity real-scene video streams with real-time, responsive control in both first- and third-person perspectives. Trained on limited supervised data from video games like Forza Horizon 5 and Cyberpunk 2077, complemented by large-scale unsupervised footage from real-world settings like Tokyo streets, The Matrix allows users to traverse diverse terrains—deserts, grasslands, water bodies, and urban landscapes—in continuous, uncut hour-long sequences. With speeds of up to 16 FPS, the system supports real-time interactivity and demonstrates zero-shot generalization, translating virtual game environments to real-world contexts where collecting continuous movement data is often infeasible. For example, The Matrix can simulate a BMW X3 driving through an office setting—an environment present in neither gaming data nor real-world sources. This approach showcases the potential of game data to advance robust world models, bridging the gap between simulations and real-world applications in scenarios with limited data. Ruili Feng, Han Zhang 0010, Zhilei Shu, Zhantao Yang, Longxiang Tang, Zhicai Wang, Andy Zheng, Jie Xiao 0002, Ruihang Chu, Yu Liu 0063, Hongyang Zhang 0001 |
NeurIPS | 2 |
| 2024 | Parrot: Pareto-Optimal Multi-reward Reinforcement Learning Framework for Text-to-Image Generation
Yinxiao Li, Junjie Ke, Innfarn Yoo, Han Zhang 0010, Qifei Wang, Fei Deng 0001, Glenn Entis, Junfeng He, Gang Li 0021, Sangpil Kim, Irfan A. Essa, Feng Yang 0008 |
ECCV (38) | 5 |
| 2024 | DreamClean: Restoring Clean Image Using Deep Diffusion PriorabstractImage restoration poses a garners substantial interest due to the exponential surge in demands for recovering high-quality images from diverse mobile camera devices, adverse lighting conditions, suboptimal shooting environments, and frequent image compression for efficient transmission purposes. Yet this problem gathers significant challenges as people are blind to the type of restoration the images suffer, which, is usually the case in real-day scenarios and is most urgent to solve for this field. Current research, however, heavily relies on prior knowledge of the restoration type, either explicitly through rules or implicitly through the availability of degraded-clean image pairs to define the restoration process, and consumes considerable effort to collect image pairs of vast degradation types. This paper introduces DreamClean, a training-free method that needs no degradation prior knowledge but yields high-fidelity and generality towards various types of image degradation. DreamClean embeds the degraded image back to the latent of pre-trained diffusion models and re-sample it through a carefully designed diffusion process that mimics those generating clean images. Thanks to the rich image prior in diffusion models and our novel Variance Preservation Sampling (VPS) technique, DreamClean manages to handle various different degradation types at one time and reaches far more satisfied final quality than previous competitors. DreamClean relies on elegant theoretical supports to assure its convergence to clean image when VPS has appropriate parameters, and also enjoys superior experimental performance over various challenging tasks that could be overwhelming for previous methods when degradation prior is unavailable. Jie Xiao 0002, Ruili Feng, Han Zhang 0010, Zhantao Yang, Yurui Zhu, Xueyang Fu, Kai Zhu 0004, Yu Liu 0063, Zhengjun Zha |
ICLR | 3 |
| 2024 | Leveraging Unpaired Data for Vision-Language Generative Models via Cycle ConsistencyabstractCurrent vision-language generative models rely on expansive corpora of $\textit{paired}$ image-text data to attain optimal performance and generalization capabilities. However, automatically collecting such data (e.g. via large-scale web scraping) leads to low quality and poor image-text correlation, while human annotation is more accurate but requires significant manual effort and expense. We introduce $\textbf{ITIT}$ ($\textbf{I}$n$\textbf{T}$egrating $\textbf{I}$mage $\textbf{T}$ext): an innovative training paradigm grounded in the concept of cycle consistency which allows vision-language training on $\textit{unpaired}$ image and text data. ITIT is comprised of a joint image-text encoder with disjoint image and text decoders that enable bidirectional image-to-text and text-to-image generation in a single framework. During training, ITIT leverages a small set of paired image-text data to ensure its output matches the input reasonably well in both directions. Simultaneously, the model is also trained on much larger datasets containing only images or texts. This is achieved by enforcing cycle consistency between the original unpaired samples and the cycle-generated counterparts. For instance, it generates a caption for a given input image and then uses the caption to create an output image, and enforces similarity between the input and output images. Our experiments show that ITIT with unpaired datasets exhibits similar scaling behavior as using high-quality paired data. We demonstrate image generation and captioning performance on par with state-of-the-art text-to-image and image-to-text models with orders of magnitude fewer (only 3M) paired image-text data. Code will be released at https://github.com/LTH14/itit. Tianhong Li, Sangnie Bhardwaj, Yonglong Tian, Han Zhang 0010, Jarred Barber, Dina Katabi, Guillaume Lajoie, Huiwen Chang, Dilip Krishnan |
ICLR | 4 |
| 2024 | Lipschitz Singularities in Diffusion ModelsabstractDiffusion models, which employ stochastic differential equations to sample images through integrals, have emerged as a dominant class of generative models. However, the rationality of the diffusion process itself receives limited attention, leaving the question of whether the problem is well-posed and well-conditioned. In this paper, we uncover a vexing propensity of diffusion models: they frequently exhibit the infinite Lipschitz near the zero point of timesteps. We provide theoretical proofs to illustrate the presence of infinite Lipschitz constants and empirical results to confirm it. The Lipschitz singularities pose a threat to the stability and accuracy during both the training and inference processes of diffusion models. Therefore, the mitigation of Lipschitz singularities holds great potential for enhancing the performance of diffusion models. To address this challenge, we propose a novel approach, dubbed E-TSDM, which alleviates the Lipschitz singularities of the diffusion model near the zero point. Remarkably, our technique yields a substantial improvement in performance. Moreover, as a byproduct of our method, we achieve a dramatic reduction in the Fréchet Inception Distance of acceleration methods relying on network Lipschitz, including DDIM and DPM-Solver, by over 33\%. Extensive experiments on diverse datasets validate our theory and method. Our work may advance the understanding of the general diffusion process, and also provide insights for the design of diffusion models. Zhantao Yang, Ruili Feng, Han Zhang 0010, Yujun Shen, Kai Zhu 0004, Lianghua Huang, Yu Liu 0063, Deli Zhao, Jingren Zhou 0001, Fan Cheng 0002 |
ICLR | 3 |
| 2024 | CCM: Real-Time Controllable Visual Content Creation Using Text-to-Image Consistency ModelsabstractConsistency Models (CMs) have showed a promise in creating high-quality images with few steps. However, the way to add new conditional controls to the pre-trained CMs has not been explored. In this paper, we explore the pivotal subject of leveraging the generative capacity and efficiency of consistency models to facilitate controllable visual content creation via ControlNet. First, it is observed that ControlNet trained for diffusion models (DMs) can be directly applied to CMs for high-level semantic controls but sacrifice image low-level details and realism. To tackle with this issue, we develop a CMs-tailored training strategy for ControlNet using the consistency training. It is substantiated that ControlNet can be successfully established through the consistency training technique. Besides, a unified adapter can be trained utilizing the consistency training, which enhances the adaptation of DM’s ControlNet. We quantitatively and qualitatively evaluate all strategies across various conditional controls, including sketch, hed, canny, depth, human pose, low-resolution image and masked image, with the pre-trained text-to-image latent consistency models. Jie Xiao 0002, Kai Zhu 0004, Han Zhang 0010, Yujun Shen, Zhantao Yang, Ruili Feng, Yu Liu 0063, Xueyang Fu, Zhengjun Zha |
ICML | 3 |
| 2024 | Steering Prototypes with Prompt-tuning for Rehearsal-free Continual LearningabstractIn the context of continual learning, prototypes—as representative class embeddings—offer advantages in memory conservation and the mitigation of catastrophic forgetting. However, challenges related to semantic drift and prototype interference persist. In this study, we introduce the Contrastive Prototypical Prompt (CPP) approach. Through task-specific prompt-tuning, underpinned by a contrastive learning objective, we effectively address both aforementioned challenges. Our evaluations on four challenging class-incremental benchmarks reveal that CPP achieves a significant 4% to 6% improvement over state-of-the-art methods. Importantly, CPP operates without a rehearsal buffer and narrows the performance divergence between continual and offline joint-learning, suggesting an innovative scheme for Transformer-based continual learning systems1. Zhuowei Li 0002, Long Zhao 0003, Han Zhang 0010, Di Liu 0003, Ting Liu 0005, Dimitris N. Metaxas |
WACV | 4 |
| 2023 | MAGE: MAsked Generative Encoder to Unify Representation Learning and Image SynthesisabstractGenerative modeling and representation learning are two key tasks in computer vision. However, these models are typically trained independently, which ignores the potential for each task to help the other, and leads to training and model maintenance overheads. In this work, we propose MAsked Generative Encoder (MAGE), the first framework to unify SOTA image generation and self-supervised representation learning. Our key insight is that using variable masking ratios in masked image modeling pre-training can allow generative training (very high masking ratio) and representation learning (lower masking ratio) under the same training framework. Inspired by previous generative models, MAGE uses semantic tokens learned by a vector-quantized GAN at inputs and outputs, combining this with masking. We can further improve the representation by adding a contrastive loss to the encoder output. We extensively evaluate the generation and representation learning capabilities of MAGE. On ImageNet-1K, a single MAGE ViT-L model obtains 9.10 FID in the task of class-unconditional image generation and 78.9% top-1 accuracy for linear probing, achieving state-of-the-art performance in both image generation and representation learning. Code is available at https://github.com/LTHl4/mage. Tianhong Li, Huiwen Chang, Shlok Kumar Mishra, Han Zhang 0010, Dina Katabi, Dilip Krishnan |
CVPR | 4 |
| 2023 | Visual Prompt Tuning for Generative Transfer LearningabstractLearning generative image models from various domains efficiently needs transferring knowledge from an image synthesis model trained on a large dataset. We present a recipe for learning vision transformers by generative knowledge transfer. We base our framework on generative vision transformers representing an image as a sequence of visual tokens with the autoregressive or non-autoregressive transformers. To adapt to a new domain, we employ prompt tuning, which prepends learnable tokens called prompts to the image token sequence and introduces a new prompt design for our task. We study on a variety of visual domains with varying amounts of training images. We show the effectiveness of knowledge transfer and a significantly better image generation quality.11https://github.com/google-research/generative_transfer Kihyuk Sohn, Huiwen Chang, José Lezama, Luisa Polania, Han Zhang 0010, Yuan Hao, Irfan A. Essa, Lu Jiang 0004 |
CVPR | 5 |
| 2023 | MAGVIT: Masked Generative Video TransformerabstractWe introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model. We introduce a 3D tokenizer to quantize a video into spatial-temporal visual tokens and propose an embedding method for masked video token modeling to facilitate multi-task learning. We conduct extensive experiments to demonstrate the quality, efficiency, and flexibility of MAGVIT. Our experiments show that (i) MAGVIT performs favorably against state-of-the-art approaches and establishes the best-published FVD on three video generation benchmarks, including the challenging Kinetics-600. (ii) MAGVIT outperforms existing methods in inference time by two orders of magnitude against diffusion models and by 60x against autoregressive models. (iii) A single MAGVIT model supports ten diverse generation tasks and generalizes across videos from different visual domains. The source code and trained models will be released to the public at https://magvit.cs.cmu.edu. Lijun Yu, Yong Cheng 0003, Kihyuk Sohn, José Lezama, Han Zhang 0010, Huiwen Chang, Alex Hauptmann 0001, Ming-Hsuan Yang 0001, Yuan Hao, Irfan A. Essa, Lu Jiang 0004 |
CVPR | 5 |
| 2023 | Dimensionality-Varying Diffusion ProcessabstractDiffusion models, which learn to reverse a signal destruction process to generate new data, typically require the signal at each step to have the same dimension. We argue that, considering the spatial redundancy in image signals, there is no need to maintain a high dimensionality in the evolution process, especially in the early generation phase. To this end, we make a theoretical generalization of the forward diffusion process via signal decomposition. Concretely, we manage to decompose an image into multiple orthogonal components and control the attenuation of each component when perturbing the image. That way, along with the noise strength increasing, we are able to diminish those inconsequential components and thus use a lower-dimensional signal to represent the source, barely losing information. Such a reformulation allows to vary dimensions in both training and inference of diffusion models. Extensive experiments on a range of datasets suggest that our approach substantially reduces the computational cost and achieves on-par or even better synthesis performance compared to baseline methods. We also show that our strategy facilitates high-resolution image synthesis and improves FID of diffusion model trained on FFHQ at$1024\times 1024$resolution from 52.40 to 10.46. Code is available at https://github.com/damo-vilab/dvdp. Han Zhang 0010, Ruili Feng, Zhantao Yang, Lianghua Huang, Yu Liu 0063, Yujun Shen, Deli Zhao, Jingren Zhou 0001, Fan Cheng 0002 |
CVPR | 1 |
| 2023 | SVDiff: Compact Parameter Space for Diffusion Fine-TuningabstractDiffusion models have achieved remarkable success in text-to-image generation, enabling the creation of high-quality images from text prompts or other modalities. However, existing methods for customizing these models are limited by handling multiple personalized subjects and the risk of overfitting. Moreover, their large number of parameters is inefficient for model storage. In this paper, we propose a novel approach to address these limitations in existing text-to-image diffusion models for personalization. Our method involves fine-tuning the singular values of the weight matrices, leading to a compact and efficient parameter space that reduces the risk of overfitting and language-drifting. We also propose a Cut-Mix-Unmix data-augmentation technique to enhance the quality of multi-subject image generation and a simple text-based image editing framework. Our proposed SVDiff method has a significantly smaller model size compared to existing methods (≈2,200 times fewer parameters compared with vanilla DreamBooth), making it more practical for real-world applications. Ligong Han, Yinxiao Li, Han Zhang 0010, Peyman Milanfar, Dimitris N. Metaxas, Feng Yang 0008 |
ICCV | 3 |
| 2023 | VQ3D: Learning a 3D-Aware Generative Model on ImageNetabstractRecent work has shown the possibility of training generative models of 3D content from 2D image collections on small datasets corresponding to a single object class, such as human faces, animal faces, or cars. However, these models struggle on larger, more complex datasets. To model diverse and unconstrained image collections such as ImageNet, we present VQ3D, which introduces a NeRF-based decoder into a two-stage vector-quantized autoencoder. Our Stage 1 allows for the reconstruction of an input image and the ability to change the camera position around the image, and our Stage 2 allows for the generation of new 3D scenes. VQ3D is capable of generating and reconstructing 3D-aware images from the 1000-class ImageNet dataset of 1.2 million training images, and achieves a competitive ImageNet generation FID score of 16.8. Our project webpage is at this url. Kyle Sargent, Jing Yu Koh, Han Zhang 0010, Huiwen Chang, Charles Herrmann, Pratul P. Srinivasan, Jiajun Wu 0001, Deqing Sun |
ICCV | 3 |
| 2023 | Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang 0010, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, Dumitru Erhan |
ICLR | 5 |
| 2023 | Muse: Text-To-Image Generation via Masked Generative TransformersabstractWe present Muse, a text-to-image Transformermodel that achieves state-of-the-art image genera-tion performance while being significantly moreefficient than diffusion or autoregressive models.Muse is trained on a masked modeling task indiscrete token space: given the text embeddingextracted from a pre-trained large language model(LLM), Muse learns to predict randomly maskedimage tokens. Compared to pixel-space diffusionmodels, such as Imagen and DALL-E 2, Muse issignificantly more efficient due to the use of dis-crete tokens and requires fewer sampling itera-tions; compared to autoregressive models such asParti, Muse is more efficient due to the use of par-allel decoding. The use of a pre-trained LLM en-ables fine-grained language understanding, whichtranslates to high-fidelity image generation andthe understanding of visual concepts such as ob-jects, their spatial relationships, pose, cardinalityetc. Our 900M parameter model achieves a newSOTA on CC3M, with an FID score of 6.06. TheMuse 3B parameter model achieves an FID of7.88 on zero-shot COCO evaluation, along with aCLIP score of 0.32. Muse also directly enables anumber of image editing applications without theneed to fine-tune or invert the model: inpainting,outpainting, and mask-free editing. More resultsand videos demonstrating editing are available at https://muse-icml.github.io/ Huiwen Chang, Han Zhang 0010, Jarred Barber, Aaron Maschinot, José Lezama, Lu Jiang 0004, Ming-Hsuan Yang 0001, Kevin Murphy 0002, William T. Freeman, Michael Rubinstein, Yuanzhen Li, Dilip Krishnan |
ICML | 2 |
| 2023 | StoryBench: A Multifaceted Benchmark for Continuous Story VisualizationabstractGenerating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark for video generation requires data annotated over time, which contrasts with the single caption used often in video datasets. To fill this gap, we collect comprehensive human annotations on three existing datasets, and introduce StoryBench: a new, challenging multi-task benchmark to reliably evaluate forthcoming text-to-video models. Our benchmark includes three video generation tasks of increasing difficulty: action execution, where the next action must be generated starting from a conditioning video; story continuation, where a sequence of actions must be executed starting from a conditioning video; and story generation, where a video must be generated from only text prompts. We evaluate small yet strong text-to-video baselines, and show the benefits of training on story-like data algorithmically generated from existing video captions. Finally, we establish guidelines for human evaluation of video stories, and reaffirm the need of better automatic metrics for video generation. StoryBench aims at encouraging future research efforts in this exciting new area. Emanuele Bugliarello, Hernan Moraldo, Ruben Villegas, Mohammad Babaeizadeh, Mohammad Taghi Saffar, Han Zhang 0010, Dumitru Erhan, Vittorio Ferrari, Pieter-Jan Kindermans, Paul Voigtlaender |
NeurIPS | 6 |
| 2022 | Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual UnderstandingabstractHierarchical structures are popular in recent vision transformers, however, they require sophisticated designs and massive datasets to work well. In this paper, we explore the idea of nesting basic local transformers on non-overlapping image blocks and aggregating them in a hierarchical way. We find that the block aggregation function plays a critical role in enabling cross-block non-local information communication. This observation leads us to design a simplified architecture that requires minor code changes upon the original vision transformer. The benefits of the proposed judiciously-selected design are threefold: (1) NesT converges faster and requires much less training data to achieve good generalization on both ImageNet and small datasets like CIFAR; (2) when extending our key ideas to image generation, NesT leads to a strong decoder that is 8 times faster than previous transformer-based generators; and (3) we show that decoupling the feature learning and abstraction processes via this nested hierarchy in our design enables constructing a novel method (named GradCAT) for visually interpreting the learned model. Source code is available https://github.com/google-research/nested-transformer. Han Zhang 0010, Long Zhao 0003, Ting Chen 0001, Sercan Ö. Arik, Tomas Pfister |
AAAI | 2 |
| 2022 | Learning to Prompt for Continual LearningabstractThe mainstream paradigm behind continual learning has been to adapt the model parameters to non-stationary data distributions, where catastrophic forgetting is the central challenge. Typical methods rely on a rehearsal buffer or known task identity at test time to retrieve learned knowl-edge and address forgetting, while this work presents a new paradigm for continual learning that aims to train a more succinct memory system without accessing task identity at test time. Our method learns to dynamically prompt (L2P) a pre-trained model to learn tasks sequen-tially under different task transitions. In our proposed framework, prompts are small learnable parameters, which are maintained in a memory space. The objective is to optimize prompts to instruct the model prediction and ex-plicitly manage task-invariant and task-specific knowledge while maintaining model plasticity. We conduct comprehen-sive experiments under popular image classification bench-marks with different challenging continual learning set-tings, where L2P consistently outperforms prior state-of-the-art methods. Surprisingly, L2P achieves competitive results against rehearsal-based methods even without a re-hearsal buffer and is directly applicable to challenging task-agnostic continual learning. Source code is available at https://github.com/google-research/12p. Zifeng Wang 0002, Chen-Yu Lee, Han Zhang 0010, Ruoxi Sun 0002, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer G. Dy, Tomas Pfister |
CVPR | 4 |
| 2022 | MaskGIT: Masked Generative Image TransformerabstractGenerative transformers have experienced rapid popularity growth in the computer vision community in synthesizing high-fidelity and high-resolution images. The best generative transformer models so far, however, still treat an image naively as a sequence of tokens, and decode an image sequentially following the raster scan ordering (i.e. line-by-line). We find this strategy neither optimal nor efficient. This paper proposes a novel image synthesis paradigm using a bidirectional transformer decoder, which we term MaskGIT. During training, MaskGIT learns to predict randomly masked tokens by attending to tokens in all directions. At inference time, the model begins with generating all tokens of an image simultaneously, and then refines the image iteratively conditioned on the previous generation. Our experiments demonstrate that MaskGIT significantly outperforms the state-of-the-art transformer model on the ImageNet dataset, and accelerates autoregressive decoding by up to 48x. Besides, we illustrate that MaskGIT can be easily extended to various image editing tasks, such as inpainting, extrapolation, and image manipulation. Project page: masked-generative-image-transformer.github.io. Huiwen Chang, Han Zhang 0010, Lu Jiang 0004, Ce Liu 0001, William T. Freeman |
CVPR | 2 |
| 2022 | MAXIM: Multi-Axis MLP for Image ProcessingabstractRecent progress on Transformers and multilayer perceptron (MLP) models provide new network architectural designs for computer vision tasks. Although these models proved to be effective in many vision tasks such as image recognition, there remain challenges in adapting them for lowlevel vision. The inflexibility to support high-resolution images and limitations of local attention are perhaps the main bottlenecks. In this work, we present a multi-axis MLP based architecture called MAXIM, that can serve as an efficient and flexible general-purpose vision backbone for image processing tasks. MAXIM uses a UNet-shaped hierarchical structure and supports long-range interactions enabled by spatially-gated MLPs. Specifically, MAXIM contains two MLP-based building blocks: a multi-axis gated MLP that allows for efficient and scalable spatial mixing of local and global visual cues, and a cross-gating block, an alternative to cross-attention, which accounts for cross-feature conditioning. Both these modules are exclusively based on MLPs, but also benefit from being both global and ‘fully-convolutional’, two properties that are desirable for image processing. Our extensive experimental results show that the proposed MAXIM model achieves state-of-the-art performance on more than ten benchmarks across a range of image processing tasks, including denoising, deblurring, de raining, dehazing, and enhancement while requiring fewer or comparable numbers of parameters and FLOPs than competitive models. The source code and trained models will be available at https://github.com/google-research/maxim. Zhengzhong Tu, Hossein Talebi, Han Zhang 0010, Feng Yang 0008, Peyman Milanfar, Alan C. Bovik, Yinxiao Li |
CVPR | 3 |
| 2022 | DualPrompt: Complementary Prompting for Rehearsal-Free Continual Learning
Zifeng Wang 0002, Sayna Ebrahimi, Ruoxi Sun 0002, Han Zhang 0010, Chen-Yu Lee, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer G. Dy, Tomas Pfister |
ECCV (26) | 5 |
| 2022 | BLT: Bidirectional Layout Transformer for Controllable Layout Generation
Xiang Kong, Lu Jiang 0004, Huiwen Chang, Han Zhang 0010, Yuan Hao, Haifeng Gong, Irfan A. Essa |
ECCV (17) | 4 |
| 2022 | MaxViT: Multi-axis Vision Transformer
Zhengzhong Tu, Hossein Talebi, Han Zhang 0010, Feng Yang 0008, Peyman Milanfar, Alan C. Bovik, Yinxiao Li |
ECCV (24) | 3 |
| 2022 | Learning Instance-Specific Adaptation for Cross-Domain Segmentation
Yuliang Zou, Chun-Liang Li, Han Zhang 0010, Tomas Pfister, Jia-Bin Huang 0001 |
ECCV (33) | 4 |
| 2022 | ViTGAN: Training GANs with Vision Transformers
Kwonjoon Lee, Huiwen Chang, Lu Jiang 0004, Han Zhang 0010, Zhuowen Tu, Ce Liu 0001 |
ICLR | 4 |
| 2022 | Vector-quantized Image Modeling with Improved VQGAN
Jing Yu Koh, Han Zhang 0010, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge |
ICLR | 4 |
| 2022 | Deep image synthesis from intuitive user input: A review and perspectivesabstractIn many applications of computer graphics, art, and design, it is desirable for a user to provide intuitive non-image input, such as text, sketch, stroke, graph, or layout, and have a computer system automatically generate photo-realistic images according to that input. While classically, works that allow such automatic image content generation have followed a framework of image retrieval and composition, recent advances in deep generative models such as generative adversarial networks (GANs), variational autoencoders (VAEs), and flow-based methods have enabled more powerful and versatile image generation approaches. This paper reviews recent works for image synthesis given intuitive user input, covering advances in input versatility, image generation methodology, benchmark datasets, and evaluation metrics. This motivates new perspectives on input representation and interactivity, cross fertilization between major image generation paradigms, and evaluation and comparison of generation methods. Yuan Xue 0002, Han Zhang 0010, Tao Xu 0029, Song-Hai Zhang, Sharon X. Huang |
Comput. Vis. Media | 3 |
| 2021 | Improved Consistency Regularization for GANsabstractRecent work has increased the performance of Generative Adversarial Networks (GANs) by enforcing a consistency cost on the discriminator. We improve on this technique in several ways. We first show that consistency regularization can introduce artifacts into the GAN samples and explain how to fix this issue. We then propose several modifications to the consistency regularization procedure designed to improve its performance. We carry out extensive experiments quantifying the benefit of our improvements. For unconditional image synthesis on CIFAR-10 and CelebA, our modifications yield the best known FID scores on various GAN architectures. For conditional image synthesis on CIFAR-10, we improve the state-of-the-art FID score from 11.48 to 9.21. Finally, on ImageNet-2012, we apply our technique to the original BigGAN model and improve the FID from 6.66 to 5.38, which is the best score at that model size. Zhengli Zhao, Sameer Singh 0001, Honglak Lee, Augustus Odena, Han Zhang 0010 |
AAAI | 6 |
| 2021 | Cross-Modal Contrastive Learning for Text-to-Image GenerationabstractThe output of text-to-image synthesis systems should be coherent, clear, photo-realistic scenes with high semantic fidelity to their conditioned text descriptions. Our Cross-Modal Contrastive Generative Adversarial Network (XMC-GAN) addresses this challenge by maximizing the mutual information between image and text. It does this via multiple contrastive losses which capture inter-modality and intra-modality correspondences. XMC-GAN uses an attentional self-modulation generator, which enforces strong text-image correspondence, and a contrastive discriminator, which acts as a critic as well as a feature encoder for contrastive learning. The quality of XMC-GAN’s output is a major step up from previous models, as we show on three challenging datasets. On MS-COCO, not only does XMC-GAN improve state-of-the-art FID from 24.70 to 9.33, but– more importantly–people prefer XMC-GAN by 77.3% for image quality and 74.1% for image-text alignment, compared to three other recent models. XMC-GAN also generalizes to the challenging Localized Narratives dataset (which has longer, more detailed descriptions), improving state-of-the-art FID from 48.70 to 14.12. Lastly, we train and evaluate XMC-GAN on the challenging Open Images data, establishing a strong benchmark FID score of 26.91. Han Zhang 0010, Jing Yu Koh, Jason Baldridge, Honglak Lee, Yinfei Yang |
CVPR | 1 |
| 2021 | Revisiting Hierarchical Approach for Persistent Long-Term Video Prediction
Wonkwang Lee, Whie Jung, Han Zhang 0010, Ting Chen 0001, Jing Yu Koh, Thomas E. Huang, Hyungsuk Yoon, Honglak Lee, Seunghoon Hong |
ICLR | 3 |
| 2021 | PseudoSeg: Designing Pseudo Labels for Semantic Segmentation
Yuliang Zou, Han Zhang 0010, Chun-Liang Li, Xiao Bian, Jia-Bin Huang 0001, Tomas Pfister |
ICLR | 3 |
| 2021 | Improved Transformer for High-Resolution GANsabstractAttention-based models, exemplified by the Transformer, can effectively model long range dependency, but suffer from the quadratic complexity of self-attention operation, making them difficult to be adopted for high-resolution image generation based on Generative Adversarial Networks (GANs). In this paper, we introduce two key ingredients to Transformer to address this challenge. First, in low-resolution stages of the generative process, standard global self-attention is replaced with the proposed multi-axis blocked self-attention which allows efficient mixing of local and global attention. Second, in high-resolution stages, we drop self-attention while only keeping multi-layer perceptrons reminiscent of the implicit neural function. To further improve the performance, we introduce an additional self-modulation component based on cross-attention. The resulting model, denoted as HiT, has a nearly linear computational complexity with respect to the image size and thus directly scales to synthesizing high definition images. We show in the experiments that the proposed HiT achieves state-of-the-art FID scores of 30.83 and 2.95 on unconditional ImageNet $128 \times 128$ and FFHQ $256 \times 256$, respectively, with a reasonable throughput. We believe the proposed HiT is an important milestone for generators in GANs which are completely free of convolutions. Our code is made publicly available at https://github.com/google-research/hit-gan. Long Zhao 0003, Ting Chen 0001, Dimitris N. Metaxas, Han Zhang 0010 |
NeurIPS | 5 |
| 2020 | Your Local GAN: Designing Two Dimensional Local Attention Mechanisms for Generative ModelsabstractWe introduce a new local sparse attention layer that preserves two-dimensional geometry and locality. We show that by just replacing the dense attention layer of SAGAN with our construction, we obtain very significant FID, Inception score and pure visual improvements. FID score is improved from 18.65 to 15.94 on ImageNet, keeping all other parameters the same. The sparse attention patterns that we propose for our new layer are designed using a novel information theoretic criterion that uses information flow graphs. We also present a novel way to invert Generative Adversarial Networks with attention. Our method uses the attention layer of the discriminator to create an innovative loss function. This allows us to visualize the newly introduced attention heads and show that they indeed capture interesting aspects of two-dimensional geometry of real images. Giannis Daras, Augustus Odena, Han Zhang 0010, Alexandros G. Dimakis |
CVPR | 3 |
| 2020 | Distilling Effective Supervision From Severe Label NoiseabstractCollecting large-scale data with clean labels for supervised training of neural networks is practically challenging. Although noisy labels are usually cheap to acquire, existing methods suffer a lot from label noise. This paper targets at the challenge of robust training at high label noise regimes. The key insight to achieve this goal is to wisely leverage a small trusted set to estimate exemplar weights and pseudo labels for noisy data in order to reuse them for supervised training. We present a holistic framework to train deep neural networks in a way that is highly invulnerable to label noise. Our method sets the new state of the art on various types of label noise and achieves excellent performance on large-scale datasets with real-world label noise. For instance, on CIFAR100 with a 40% uniform noise ratio and only 10 trusted labeled data per class, our method achieves 80.2% classification accuracy, where the error rate is only 1.4% higher than a neural network trained without label noise. Moreover, increasing the noise ratio to 80%, our method still maintains a high accuracy of 75.5%, compared to the previous best accuracy 48.2%. Han Zhang 0010, Sercan Ö. Arik, Honglak Lee, Tomas Pfister |
CVPR | 2 |
| 2020 | ReMixMatch: Semi-Supervised Learning with Distribution Matching and Augmentation Anchoring
David Berthelot, Nicholas Carlini, Ekin Dogus Cubuk, Alexey Kurakin, Kihyuk Sohn, Han Zhang 0010, Colin Raffel |
ICLR | 6 |
| 2020 | Consistency Regularization for Generative Adversarial Networks
Han Zhang 0010, Augustus Odena, Honglak Lee |
ICLR | 1 |
| 2020 | Small-GAN: Speeding up GAN Training using Core-SetsabstractRecent work suggests that Generative Adversarial Networks (GANs) benefit disproportionately from large mini-batch sizes. This finding is interesting but also discouraging – large batch sizes are slow and expensive to emulate on conventional hardware. Thus, it would be nice if there were some trick by which we could generate batches that were effectively big though small in practice. In this work, we propose such a trick, inspired by the use of Coreset-selection in active learning. When training a GAN, we draw a large batch of samples from the prior and then compress that batch using Coreset-selection. To create effectively large batches of real images, we create a cached dataset of Inception activations of each training image, randomly project them down to a smaller dimension, and then use Coreset-selection on those projected embeddings at training time. We conduct experiments showing that this technique substantially reduces training time and memory usage for modern GAN variants, that it reduces the fraction of dropped modes in a synthetic dataset, and that it helps us use GANs to reach a new state of the art in anomaly detection. Samarth Sinha, Han Zhang 0010, Anirudh Goyal, Yoshua Bengio, Hugo Larochelle, Augustus Odena |
ICML | 2 |
| 2020 | FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceabstractSemi-supervised learning (SSL) provides an effective means of leveraging unlabeled data to improve a model’s performance. This domain has seen fast progress recently, at the cost of requiring more complex methods. In this paper we propose FixMatch, an algorithm that is a significant simplification of existing SSL methods. FixMatch first generates pseudo-labels using the model’s predictions on weakly-augmented unlabeled images. For a given image, the pseudo-label is only retained if the model produces a high-confidence prediction. The model is then trained to predict the pseudo-label when fed a strongly-augmented version of the same image. Despite its simplicity, we show that FixMatch achieves state-of-the-art performance across a variety of standard semi-supervised learning benchmarks, including 94.93% accuracy on CIFAR-10 with 250 labels and 88.61% accuracy with 40 – just 4 labels per class. We carry out an extensive ablation study to tease apart the experimental factors that are most important to FixMatch’s success. The code is available at https://github.com/google-research/fixmatch. Kihyuk Sohn, David Berthelot, Nicholas Carlini, Han Zhang 0010, Colin Raffel, Ekin Dogus Cubuk, Alexey Kurakin, Chun-Liang Li |
NeurIPS | 5 |
| 2019 | Co-Occurrent Features in Semantic SegmentationabstractRecent work has achieved great success in utilizing global contextual information for semantic segmentation, including increasing the receptive field and aggregating pyramid feature representations. In this paper, we go beyond global context and explore the fine-grained representation using co-occurrent features by introducing Co-occurrent Feature Model, which predicts the distribution of co-occurrent features for a given target. To leverage the semantic context in the co-occurrent features, we build an Aggregated Co-occurrent Feature (ACF) Module by aggregating the probability of the co-occurrent feature with the co-occurrent context. ACF Module learns a fine-grained spatial invariant representation to capture co-occurrent context information across the scene. Our approach significantly improves the segmentation results using FCN and achieves superior performance 54.0% mIoU on Pascal Context, 87.2% mIoU on Pascal VOC 2012 and 44.89% mIoU on ADE20K datasets. The source code and complete system will be publicly available upon publication. Hang Zhang 0005, Han Zhang 0010, Chenguang Wang 0001, Junyuan Xie |
CVPR | 2 |
| 2019 | Self-Attention Generative Adversarial NetworksabstractIn this paper, we propose the Self-Attention Generative Adversarial Network (SAGAN) which allows attention-driven, long-range dependency modeling for image generation tasks. Traditional convolutional GANs generate high-resolution details as a function of only spatially local points in lower-resolution feature maps. In SAGAN, details can be generated using cues from all feature locations. Moreover, the discriminator can check that highly detailed features in distant portions of the image are consistent with each other. Furthermore, recent work has shown that generator conditioning affects GAN performance. Leveraging this insight, we apply spectral normalization to the GAN generator and find that this improves training dynamics. The proposed SAGAN performs better than prior work, boosting the best published Inception score from 36.8 to 52.52 and reducing Fréchet Inception distance from 27.62 to 18.65 on the challenging ImageNet dataset. Visualization of the attention layers shows that the generator leverages neighborhoods that correspond to object shapes rather than local regions of fixed shape. Han Zhang 0010, Ian J. Goodfellow, Dimitris N. Metaxas, Augustus Odena |
ICML | 1 |
| 2019 | StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial NetworksabstractAlthough Generative Adversarial Networks (GANs) have shown remarkable success in various tasks, they still face challenges in generating high quality images. In this paper, we propose Stacked Generative Adversarial Networks (StackGANs) aimed at generating high-resolution photo-realistic images. First, we propose a two-stage generative adversarial network architecture, StackGAN-v1, for text-to-image synthesis. The Stage-I GAN sketches the primitive shape and colors of a scene based on a given text description, yielding low-resolution images. The Stage-II GAN takes Stage-I results and the text description as inputs, and generates high-resolution images with photo-realistic details. Second, an advanced multi-stage generative adversarial network architecture, StackGAN-v2, is proposed for both conditional and unconditional generative tasks. Our StackGAN-v2 consists of multiple generators and multiple discriminators arranged in a tree-like structure; images at multiple scales corresponding to the same scene are generated from different branches of the tree. StackGAN-v2 shows more stable training behavior than StackGAN-v1 by jointly approximating multiple distributions. Extensive experiments demonstrate that the proposed stacked generative adversarial networks significantly outperform other state-of-the-art methods in generating photo-realistic images. Han Zhang 0010, Tao Xu 0029, Hongsheng Li 0001, Shaoting Zhang 0001, Xiaogang Wang 0001, Sharon X. Huang, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | AttnGAN: Fine-Grained Text to Image Generation With Attentional Generative Adversarial NetworksabstractIn this paper, we propose an Attentional Generative Adversarial Network (AttnGAN) that allows attention-driven, multi-stage refinement for fine-grained text-to-image generation. With a novel attentional generative network, the AttnGAN can synthesize fine-grained details at different sub-regions of the image by paying attentions to the relevant words in the natural language description. In addition, a deep attentional multimodal similarity model is proposed to compute a fine-grained image-text matching loss for training the generator. The proposed AttnGAN significantly outperforms the previous state of the art, boosting the best reported inception score by 14.14% on the CUB dataset and 170.25% on the more challenging COCO dataset. A detailed analysis is also performed by visualizing the attention layers of the AttnGAN. It for the first time shows that the layered attentional GAN is able to automatically select the condition at the word level for generating different parts of the image. Tao Xu 0029, Pengchuan Zhang, Qiuyuan Huang, Han Zhang 0010, Zhe Gan, Sharon X. Huang, Xiaodong He 0001 |
CVPR | 4 |
| 2018 | Improving GANs Using Optimal Transport
Tim Salimans, Han Zhang 0010, Alec Radford, Dimitris N. Metaxas |
ICLR (Poster) | 2 |
| 2017 | Link the Head to the "Beak": Zero Shot Learning from Noisy Text Description at Part PrecisionabstractIn this paper, we study learning visual classifiers from unstructured text descriptions at part precision with no training images. We propose a learning framework that is able to connect text terms to its relevant parts and suppress connections to non-visual text terms without any part-text annotations. For instance, this learning process enables terms like beak to be sparsely linked to the visual representation of parts like head, while reduces the effect of non-visual terms like migrate on classifier prediction. Images are encoded by a part-based CNN that detect bird parts and learn part-specific representation. Part-based visual classifiers are predicted from text descriptions of unseen visual classifiers to facilitate classification without training images (also known as zero-shot recognition). We performed our experiments on CUBirds 2011 dataset and improves the state-of-the-art textbased zero-shot recognition results from 34.7% to 43.6%. We also created large scale benchmarks on North American Bird Images augmented with text descriptions, where we also show that our approach outperforms existing methods. Our code, data, and models are publically available link [1]. Yizhe Zhu, Han Zhang 0010, Ahmed M. Elgammal |
CVPR | 3 |
| 2017 | StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial NetworksabstractSynthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications. Samples generated by existing textto- image approaches can roughly reflect the meaning of the given descriptions, but they fail to contain necessary details and vivid object parts. In this paper, we propose Stacked Generative Adversarial Networks (StackGAN) to generate 256.256 photo-realistic images conditioned on text descriptions. We decompose the hard problem into more manageable sub-problems through a sketch-refinement process. The Stage-I GAN sketches the primitive shape and colors of the object based on the given text description, yielding Stage-I low-resolution images. The Stage-II GAN takes Stage-I results and text descriptions as inputs, and generates high-resolution images with photo-realistic details. It is able to rectify defects in Stage-I results and add compelling details with the refinement process. To improve the diversity of the synthesized images and stabilize the training of the conditional-GAN, we introduce a novel Conditioning Augmentation technique that encourages smoothness in the latent conditioning manifold. Extensive experiments and comparisons with state-of-the-arts on benchmark datasets demonstrate that the proposed method achieves significant improvements on generating photo-realistic images conditioned on text descriptions. Han Zhang 0010, Tao Xu 0029, Hongsheng Li 0001 |
ICCV | 1 |
| 2017 | Multi-feature based benchmark for cervical dysplasia classification evaluation
Tao Xu 0029, Han Zhang 0010, Cheng Xin, L. Rodney Long, Zhiyun Xue, Sameer K. Antani, Sharon X. Huang |
Pattern Recognit. | 2 |
| 2016 | SPDA-CNN: Unifying Semantic Part Detection and Abstraction for Fine-Grained RecognitionabstractMost convolutional neural networks (CNNs) lack midlevel layers that model semantic parts of objects. This limits CNN-based methods from reaching their full potential in detecting and utilizing small semantic parts in recognition. Introducing such mid-level layers can facilitate the extraction of part-specific features which can be utilized for better recognition performance. This is particularly important in the domain of fine-grained recognition. In this paper, we propose a new CNN architecture that integrates semantic part detection and abstraction (SPDACNN) for fine-grained classification. The proposed network has two sub-networks: one for detection and one for recognition. The detection sub-network has a novel top-down proposal method to generate small semantic part candidates for detection. The classification sub-network introduces novel part layers that extract features from parts detected by the detection sub-network, and combine them for recognition. As a result, the proposed architecture provides an end-to-end network that performs detection, localization of multiple semantic parts, and whole object recognition within one framework that shares the computation of convolutional filters. Our method outperforms state-of-theart methods with a large margin for small parts detection (e.g. our precision of 93.40% vs the best previous precision of 74.00% for detecting the head on CUB-2011). It also compares favorably to the existing state-of-the-art on finegrained classification, e.g. it achieves 85.14% accuracy on CUB-2011. Han Zhang 0010, Tao Xu 0029, Sharon X. Huang, Shaoting Zhang 0001, Ahmed M. Elgammal, Dimitris N. Metaxas |
CVPR | 1 |
| 2016 | Multimodal Deep Learning for Cervical Dysplasia Diagnosis
Tao Xu 0029, Han Zhang 0010, Sharon X. Huang, Shaoting Zhang 0001, Dimitris N. Metaxas |
MICCAI (2) | 2 |
| 2012 | Video Browser Showdown by NUS
Jin Yuan 0002, Huan-Bo Luan, Dejun Hou, Han Zhang 0010, Yantao Zheng, Zhengjun Zha, Tat-Seng Chua |
MMM | 4 |