VLDB 2026 Research / reviewers in the wild / expert
Weili Nie
dblp:147/4786
· DBLP profile ↗
34ranked-venue papers
12as first author
24since 2021 · last 2025
0000-0002-0030-3189ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 9 first-author · 22 since 2021Computer networks · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video RepresentationsabstractExisting video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability. In this work, we propose to decompose videos into visual primitives – blob video representation, a general representation for controllable video generation. Based on blob conditions, we develop a blob-grounded video diffusion model named BlobGEN-Vid that allows users to control object motions and fine-grained object appearance. In particular, we introduce a masked 3D attention module that effectively improves regional consistency across frames. In addition, we introduce a learnable module to interpolate text embeddings so that users can control semantics in specific frames and obtain smooth object transitions. We show that our framework is model-agnostic and can build BlobGEN-Vid on both U-Net and DiT-based video diffusion models. Extensive experimental results show that BlobGEN-Vid achieves superior zero-shot video generation ability and state-of-the-art layout controllability on multiple benchmarks. When combined with an LLM for layout planning, our framework even outperforms proprietary text-to-video generators regarding compositional accuracy. Our project page: blobgen-vid.github.io Weixi Feng, Chao Liu 0064, Sifei Liu, William Yang Wang, Arash Vahdat, Weili Nie |
CVPR | 6 |
| 2025 | Truncated Consistency ModelsabstractConsistency models have recently been introduced to accelerate the generation speed of diffusion models by directly predicting the solution (data) of the probability flow ODE (PF ODE) from initial noise.
However, the training of consistency models requires learning to map all intermediate points along PF ODE trajectories to their corresponding endpoints. This task is much more challenging than the ultimate objective of one-step generation, which only concerns the PF ODE's noise-to-data mapping.
We empirically find that this training paradigm limits the one-step generation performance of consistency models.
To address this issue, we generalize consistency training to the truncated time range, which allows the model to ignore denoising tasks at earlier time steps and focus its capacity on generation.
We propose a new parameterization of the consistency function and a two-stage training procedure that prevent the truncated-time training from collapsing to a trivial solution.
Experiments on CIFAR-10 and ImageNet $64\times64$ datasets show that our method achieves better one-step and two-step FIDs than the state-of-the-art consistency models such as iCT-deep,
using more than 2$\times$ smaller networks. Tomas Geffner, Giulia Fanti, Karsten Kreis, Arash Vahdat, Weili Nie |
ICLR | 7 |
| 2025 | T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory StitchingabstractSampling from diffusion probabilistic models (DPMs) is often expensive for high-quality image generation and typically requires many steps with a large model. In this paper, we introduce sampling Trajectory Stitching (T-Stitch), a simple yet efficient technique to improve the sampling efficiency with little or no generation degradation. Instead of solely using a large DPM for the entire sampling trajectory, T-Stitch first leverages a smaller DPM in the initial steps as a cheap drop-in replacement of the larger DPM and switches to the larger DPM at a later stage. Our key insight is that different diffusion models learn similar encodings under the same training data distribution and smaller models are capable of generating good global structures in the early steps. Extensive experiments demonstrate that T-Stitch is training-free, generally applicable for different architectures, and complements most existing fast sampling techniques with flexible speed and quality trade-offs. On DiT-XL, for example, 40% of the early timesteps can be safely replaced with a 10x faster DiT-S without performance drop on class-conditional ImageNet generation. We further show that our method can also be used as a drop-in technique to not only accelerate the popular pretrained stable diffusion (SD) models but also improve the prompt alignment of stylized SD models from the public model zoo. Finally, the explicit model allocation strategy of T-Stitch significantly reduces the need of training or searching, delivering high deployment efficiency. Zizheng Pan, Bohan Zhuang, De-An Huang, Weili Nie, Zhiding Yu, Chaowei Xiao, Jianfei Cai 0001, Anima Anandkumar |
ICLR | 4 |
| 2025 | Energy-Based Diffusion Language Models for Text GenerationabstractDespite remarkable progress in autoregressive language models, alternative generative paradigms beyond left-to-right generation are still being actively explored. Discrete diffusion models, with the capacity for parallel generation, have recently emerged as a promising alternative. Unfortunately, these models still underperform the autoregressive counterparts, with the performance gap increasing when reducing the number of sampling steps. Our analysis reveals that this degradation is a consequence of an imperfect approximation used by diffusion models. In this work, we propose Energy-based Diffusion Language Model (EDLM), an energy-based model operating at the full sequence level for each diffusion step, introduced to improve the underlying approximation used by diffusion models. More specifically, we introduce an EBM in a residual form, and show that its parameters can be obtained by leveraging a pretrained autoregressive model or by finetuning a bidirectional transformer via noise contrastive estimation. We also propose an efficient generation algorithm via parallel important sampling. Comprehensive experiments on language modeling benchmarks show that our model can consistently outperform state-of-the-art diffusion models by a significant margin, and approaches autoregressive models' perplexity. We further show that, without any generation performance drop, our framework offers a 1.3x sampling speedup over existing diffusion models. Reproduced code is available at https://github.com/MinkaiXu/Energy-Diffusion-LLM. Minkai Xu, Tomas Geffner, Karsten Kreis, Weili Nie, Jure Leskovec, Stefano Ermon, Arash Vahdat |
ICLR | 4 |
| 2025 | GenMol: A Drug Discovery Generalist with Discrete DiffusionabstractDrug discovery is a complex process that involves multiple stages and tasks. However, existing molecular generative models can only tackle some of these tasks. We present Generalist Molecular generative model (GenMol), a versatile framework that uses only a single discrete diffusion model to handle diverse drug discovery scenarios. GenMol generates Sequential Attachment-based Fragment Embedding (SAFE) sequences through non-autoregressive bidirectional parallel decoding, thereby allowing the utilization of a molecular context that does not rely on the specific token ordering while having better sampling efficiency. GenMol uses fragments as basic building blocks for molecules and introduces fragment remasking, a strategy that optimizes molecules by regenerating masked fragments, enabling effective exploration of chemical space. We further propose molecular context guidance (MCG), a guidance method tailored for masked discrete diffusion of GenMol. GenMol significantly outperforms the previous GPT-based model in de novo generation and fragment-constrained generation, and achieves state-of-the-art performance in goal-directed hit generation and lead optimization. These results demonstrate that GenMol can tackle a wide range of drug discovery tasks, providing a unified and versatile approach for molecular design. Seul Lee, Karsten Kreis, Srimukh Prasad Veccham, Meng Liu 0015, Danny Reidenbach, Yuxing Peng 0005, Saee Gopal Paliwal, Weili Nie, Arash Vahdat |
ICML | 8 |
| 2025 | Intelligent Recognition of Gram-Stained Microscopic Images Based on DIBAS Dataset
Weili Nie |
ISNN | 1 |
| 2024 | Efficient Video Diffusion Models via Content-Frame Motion-Latent DecompositionabstractVideo diffusion models have recently made great progress in generation quality, but are still limited by the high memory and computational requirements. This is because current video diffusion models often attempt to process high-dimensional videos directly. To tackle this issue, we propose content-motion latent diffusion model (CMD), a novel efficient extension of pretrained image diffusion models for video generation. Specifically, we propose an autoencoder that succinctly encodes a video as a combination of a content frame (like an image) and a low-dimensional motion latent representation. The former represents the common content, and the latter represents the underlying motion in the video, respectively. We generate the content frame by fine-tuning a pretrained image diffusion model, and we generate the motion latent representation by training a new lightweight diffusion model. A key innovation here is the design of a compact latent space that can directly utilizes a pretrained image diffusion model, which has not been done in previous latent video diffusion models. This leads to considerably better quality generation and reduced computational costs. For instance, CMD can sample a video 7.7$\times$ faster than prior approaches by generating a video of 512$\times$1024 resolution and length 16 in 3.1 seconds. Moreover, CMD achieves an FVD score of 238.3 on WebVid-10M, 18.5% better than the previous state-of-the-art of 292.4. Sihyun Yu, Weili Nie, De-An Huang, Jinwoo Shin, Anima Anandkumar |
ICLR | 2 |
| 2024 | Compositional Text-to-Image Generation with Dense Blob RepresentationsabstractExisting text-to-image models struggle to follow complex text prompts, raising the need for extra grounding inputs for better controllability. In this work, we propose to decompose a scene into visual primitives - denoted as dense blob representations - that contain fine-grained details of the scene while being modular, human-interpretable, and easy-to-construct. Based on blob representations, we develop a blob-grounded text-to-image diffusion model, termed BlobGEN, for compositional generation. Particularly, we introduce a new masked cross-attention module to disentangle the fusion between blob representations and visual features. To leverage the compositionality of large language models (LLMs), we introduce a new in-context learning approach to generate blob representations from text prompts. Our extensive experiments show that BlobGEN achieves superior zero-shot generation quality and better layout-guided controllability on MS-COCO. When augmented by LLMs, our method exhibits superior numerical and spatial correctness on compositional image generation benchmarks. Weili Nie, Sifei Liu, Morteza Mardani, Chao Liu 0064, Benjamin Eckart, Arash Vahdat |
ICML | 1 |
| 2024 | Warped Diffusion: Solving Video Inverse Problems with Image Diffusion ModelsabstractUsing image models naively for solving inverse video problems often suffers from flickering, texture-sticking, and temporal inconsistency in generated videos. To tackle these problems, in this paper, we view frames as continuous functions in the 2D space, and videos as a sequence of continuous warping transformations between different frames. This perspective allows us to train function space diffusion models only on **images** and utilize them to solve temporally correlated inverse problems. The function space diffusion models need to be equivariant with respect to the underlying spatial transformations. To ensure temporal consistency, we introduce a simple post-hoc test-time guidance towards (self)-equivariant solutions. Our method allows us to deploy state-of-the-art latent diffusion models such as Stable Diffusion XL to solve video inverse problems. We demonstrate the effectiveness of our method for video inpainting and $8\times$ video super-resolution, outperforming existing techniques based on noise transformations. We provide generated video results in the following URL: https://giannisdaras.github.io/warped_diffusion.github.io/. Giannis Daras, Weili Nie, Karsten Kreis, Alexandros G. Dimakis, Morteza Mardani, Nikola B. Kovachki, Arash Vahdat |
NeurIPS | 2 |
| 2024 | Aligning Target-Aware Molecule Diffusion Models with Exact Energy OptimizationabstractGenerating ligand molecules for specific protein targets, known as structure-based drug design, is a fundamental problem in therapeutics development and biological discovery. Recently, target-aware generative models, especially diffusion models, have shown great promise in modeling protein-ligand interactions and generating candidate drugs. However, existing models primarily focus on learning the chemical distribution of all drug candidates, which lacks effective steerability on the chemical quality of model generations. In this paper, we propose a novel and general alignment framework to align pretrained target diffusion models with preferred functional properties, named AliDiff. AliDiff shifts the target-conditioned chemical distribution towards regions with higher binding affinity and structural rationality, specified by user-defined reward functions, via the preference optimization approach. To avoid the overfitting problem in common preference optimization objectives, we further develop an improved Exact Energy Preference Optimization method to yield an exact and efficient alignment of the diffusion models, and provide the closed-form expression for the converged distribution. Empirical studies on the CrossDocked2020 benchmark show that AliDiff can generate molecules with state-of-the-art binding energies with up to -7.07 Avg. Vina Score, while maintaining strong molecular properties. Code is available at https://github.com/MinkaiXu/AliDiff. Siyi Gu, Minkai Xu, Alexander S. Powers, Weili Nie, Tomas Geffner, Karsten Kreis, Jure Leskovec, Arash Vahdat, Stefano Ermon |
NeurIPS | 4 |
| 2024 | Molecule Generation with Fragment Retrieval AugmentationabstractFragment-based drug discovery, in which molecular fragments are assembled into new molecules with desirable biochemical properties, has achieved great success. However, many fragment-based molecule generation methods show limited exploration beyond the existing fragments in the database as they only reassemble or slightly modify the given ones. To tackle this problem, we propose a new fragment-based molecule generation framework with retrieval augmentation, namely *Fragment Retrieval-Augmented Generation* (*f*-RAG). *f*-RAG is based on a pre-trained molecular generative model that proposes additional fragments from input fragments to complete and generate a new molecule. Given a fragment vocabulary, *f*-RAG retrieves two types of fragments: (1) *hard fragments*, which serve as building blocks that will be explicitly included in the newly generated molecule, and (2) *soft fragments*, which serve as reference to guide the generation of new fragments through a trainable *fragment injection module*. To extrapolate beyond the existing fragments, *f*-RAG updates the fragment vocabulary with generated fragments via an iterative refinement process which is further enhanced with post-hoc genetic fragment modification. *f*-RAG can achieve an improved exploration-exploitation trade-off by maintaining a pool of fragments and expanding it with novel and high-quality fragments through a strong generative prior. Seul Lee, Karsten Kreis, Srimukh Prasad Veccham, Meng Liu 0015, Danny Reidenbach, Saee Gopal Paliwal, Arash Vahdat, Weili Nie |
NeurIPS | 8 |
| 2024 | BlobGEN-3D: Compositional 3D-Consistent Freeview Image Generation with 3D Blobs
Chao Liu 0064, Weili Nie, Sifei Liu, Abhishek Badki, Hang Su 0005, Morteza Mardani, Benjamin Eckart, Arash Vahdat |
SIGGRAPH Asia | 2 |
| 2024 | DiffUHaul: A Training-Free Method for Object Dragging in Images
Omri Avrahami, Rinon Gal, Gal Chechik, Ohad Fried, Dani Lischinski, Arash Vahdat, Weili Nie |
SIGGRAPH Asia | 7 |
| 2023 | Retrieval-based Controllable Molecule Generation
Zichao Wang 0001, Weili Nie, Jarren Zhuoran Qiao, Chaowei Xiao, Richard G. Baraniuk, Anima Anandkumar |
ICLR | 2 |
| 2023 | Defending against Adversarial Audio via Diffusion Model
Shutong Wu, Jiongxiao Wang, Wei Ping, Weili Nie, Chaowei Xiao |
ICLR | 4 |
| 2023 | DensePure: Understanding Diffusion Models for Adversarial Robustness
Chaowei Xiao, Zhongzhu Chen, Jiongxiao Wang, Weili Nie, Mingyan Liu, Anima Anandkumar, Bo Li 0026, Dawn Song |
ICLR | 5 |
| 2023 | I2SB: Image-to-Image Schrödinger BridgeabstractWe propose Image-to-Image Schrödinger Bridge (I$^2$SB), a new class of conditional diffusion models that directly learn the nonlinear diffusion processes between two given distributions. These diffusion bridges are particularly useful for image restoration, as the degraded images are structurally informative priors for reconstructing the clean images. I$^2$SB belongs to a tractable class of Schrödinger bridge, the nonlinear extension to score-based models, whose marginal distributions can be computed analytically given boundary pairs. This results in a simulation-free framework for nonlinear diffusions, where the I$^2$SB training becomes scalable by adopting practical techniques used in standard diffusion models. We validate I$^2$SB in solving various image restoration tasks, including inpainting, super-resolution, deblurring, and JPEG restoration on ImageNet 256$\times$256 and show that I$^2$SB surpasses standard conditional diffusion models with more interpretable generative processes. Moreover, I$^2$SB matches the performance of inverse methods that additionally require the knowledge of the corruption operators. Our work opens up new algorithmic opportunities for developing efficient nonlinear diffusion models on a large scale. Project page and codes: https://i2sb.github.io/ Guan-Horng Liu, Arash Vahdat, De-An Huang, Evangelos A. Theodorou, Weili Nie, Anima Anandkumar |
ICML | 5 |
| 2023 | A Critical Revisit of Adversarial Robustness in 3D Point Cloud Recognition with Diffusion-Driven Purificationabstract3D point clouds serve as a crucial data representation in numerous real-world applications such as autonomous driving, robotics, and medical imaging. While the advancements in deep learning have spurred the utilization of 3D point clouds, deep models are notoriously vulnerable to adversarial attacks. Various defense solutions have been proposed to build robust models against adversarial attacks. In this work, we pinpoint a major limitation of the leading empirical defense, adversarial training, when applied to 3D point cloud models: gradient obfuscation, which significantly hampers robustness against potent attacks. To bridge the gap, we propose PointDP, a purification strategy that leverages diffusion models to defend against 3D adversarial attacks. Since PointDP does not rely on predefined adversarial examples for training, it can defend against a variety of threats. We conduct a comprehensive evaluation of PointDP across six representative 3D point cloud architectures, employing sixteen strong and adaptive attacks to manifest its foundational robustness. Our evaluation shows that PointDP achieves significantly better (i.e., 12.6%-40.3%) adversarial robustness than state-of-the-art methods under strong attacks bounded by different $\ell_p$ norms. Jiongxiao Wang, Weili Nie, Zhiding Yu, Z. Morley Mao, Chaowei Xiao |
ICML | 3 |
| 2023 | Fast Sampling of Diffusion Models via Operator LearningabstractDiffusion models have found widespread adoption in various areas. However, their sampling process is slow because it requires hundreds to thousands of network evaluations to emulate a continuous process defined by differential equations. In this work, we use neural operators, an efficient method to solve the probability flow differential equations, to accelerate the sampling process of diffusion models. Compared to other fast sampling methods that have a sequential nature, we are the first to propose a parallel decoding method that generates images with only one model forward pass. We propose diffusion model sampling with neural operator (DSNO) that maps the initial condition, i.e., Gaussian distribution, to the continuous-time solution trajectory of the reverse diffusion process. To model the temporal correlations along the trajectory, we introduce temporal convolution layers that are parameterized in the Fourier space into the given diffusion model backbone. We show our method achieves state-of-the-art FID of 3.78 for CIFAR-10 and 7.83 for ImageNet-64 in the one-model-evaluation setting. Hongkai Zheng, Weili Nie, Arash Vahdat, Kamyar Azizzadenesheli, Anima Anandkumar |
ICML | 2 |
| 2022 | Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object InteractionsabstractA significant gap remains between today's visual pattern recognition models and humanlevel visual cognition especially when it comes to fewshot learning and compositional reasoning of novel concepts. We introduce Bongard-HOI, a new visual reasoning benchmark that focuses on compositional learning of humanobject interactions (HOIs) from natural images. It is inspired by two desirable characteristics from the classical Bongard problems (BPs): 1) fewshot concept learning, and 2) contextdependent reasoning. We carefully curate the fewshot instances with hard negatives, where positive and negative images only disagree on action labels, making mere recognition of object categories insufficient to complete our benchmarks. We also design multiple test sets to systematically study the generalization of visual learning models, where we vary the overlap of the HOI concepts between the training and test sets of fewshot instances, from partial to no overlaps. Bongard-HOI presents a substantial challenge to today's visual recognition models. The state-of-the-art HOI detection model achieves only 62% accuracy on fewshot binary prediction while even amateur human testers on MTurk have 91% accuracy. With the Bongard-HOI benchmark, we hope to further advance research efforts in visual reasoning, especially in holistic perception-reasoning systems and better representation learning. Huaizu Jiang, Xiaojian Ma 0001, Weili Nie, Zhiding Yu, Yuke Zhu, Anima Anandkumar |
CVPR | 3 |
| 2022 | RelViT: Concept-guided Vision Transformer for Visual Relational Reasoning
Xiaojian Ma 0001, Weili Nie, Zhiding Yu, Huaizu Jiang, Chaowei Xiao, Yuke Zhu, Song-Chun Zhu, Anima Anandkumar |
ICLR | 2 |
| 2022 | Diffusion Models for Adversarial PurificationabstractAdversarial purification refers to a class of defense methods that remove adversarial perturbations using a generative model. These methods do not make assumptions on the form of attack and the classification model, and thus can defend pre-existing classifiers against unseen threats. However, their performance currently falls behind adversarial training methods. In this work, we propose DiffPure that uses diffusion models for adversarial purification: Given an adversarial example, we first diffuse it with a small amount of noise following a forward diffusion process, and then recover the clean image through a reverse generative process. To evaluate our method against strong adaptive attacks in an efficient and scalable way, we propose to use the adjoint method to compute full gradients of the reverse generative process. Extensive experiments on three image datasets including CIFAR-10, ImageNet and CelebA-HQ with three classifier architectures including ResNet, WideResNet and ViT demonstrate that our method achieves the state-of-the-art results, outperforming current adversarial training and adversarial purification methods, often by a large margin. Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, Anima Anandkumar |
ICML | 1 |
| 2022 | Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsabstractPre-trained vision-language models (e.g., CLIP) have shown promising zero-shot generalization in many downstream tasks with properly designed text prompts. Instead of relying on hand-engineered prompts, recent works learn prompts using the training data from downstream tasks. While effective, training on domain-specific data reduces a model's generalization capability to unseen new domains. In this work, we propose test-time prompt tuning (TPT), a method that can learn adaptive prompts on the fly with a single test sample. TPT optimizes the prompt by minimizing the entropy with confidence selection so that the model has consistent predictions across different augmented views of each test sample. In evaluating generalization to natural distribution shifts, TPT improves the zero-shot top-1 accuracy of CLIP by 3.6\% on average, surpassing previous prompt tuning approaches that require additional task-specific training data. In evaluating cross-dataset generalization with unseen categories, TPTperforms on par with the state-of-the-art approaches that use additional training data. Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, Chaowei Xiao |
NeurIPS | 2 |
| 2021 | Controllable and Compositional Generation with Latent-Space Energy-Based ModelsabstractControllable generation is one of the key requirements for successful adoption of deep generative models in real-world applications, but it still remains as a great challenge. In particular, the compositional ability to generate novel concept combinations is out of reach for most current models. In this work, we use energy-based models (EBMs) to handle compositional generation over a set of attributes. To make them scalable to high-resolution image generation, we introduce an EBM in the latent space of a pre-trained generative model such as StyleGAN. We propose a novel EBM formulation representing the joint distribution of data and attributes together, and we show how sampling from it is formulated as solving an ordinary differential equation (ODE). Given a pre-trained generator, all we need for controllable generation is to train an attribute classifier. Sampling with ODEs is done efficiently in the latent space and is robust to hyperparameters. Thus, our method is simple, fast to train, and efficient to sample. Experimental results show that our method outperforms the state-of-the-art in both conditional sampling and sequential editing. In compositional generation, our method excels at zero-shot generation of unseen attribute combinations. Also, by composing energy functions with logical operators, this work is the first to achieve such compositionality in generating photo-realistic images of resolution 1024x1024. Weili Nie, Arash Vahdat, Anima Anandkumar |
NeurIPS | 1 |
| 2020 | Semi-Supervised StyleGAN for Disentanglement LearningabstractDisentanglement learning is crucial for obtaining disentangled representations and controllable generation. Current disentanglement methods face several inherent limitations: difficulty with high-resolution images, primarily focusing on learning disentangled representations, and non-identifiability due to the unsupervised setting. To alleviate these limitations, we design new architectures and loss functions based on StyleGAN (Karras et al., 2019), for semi-supervised high-resolution disentanglement learning. We create two complex high-resolution synthetic datasets for systematic testing. We investigate the impact of limited supervision and find that using only 0.25% 2.5% of labeled data is sufficient for good disentanglement on both synthetic and real datasets. We propose new metrics to quantify generator controllability, and observe there may exist a crucial trade-off between disentangled representation learning and controllable generation. We also consider semantic fine-grained image editing to achieve better generalization to unseen images. Weili Nie, Tero Karras, Animesh Garg, Shoubhik Debnath, Anjul Patney, Ankit B. Patel, Anima Anandkumar |
ICML | 1 |
| 2020 | Bongard-LOGO: A New Benchmark for Human-Level Concept Learning and ReasoningabstractHumans have an inherent ability to learn novel concepts from only a few samples and generalize these concepts to different situations. Even though today's machine learning models excel with a plethora of training data on standard recognition tasks, a considerable gap exists between machine-level pattern recognition and human-level concept learning. To narrow this gap, the Bongard Problems (BPs) were introduced as an inspirational challenge for visual cognition in intelligent systems. Albeit new advances in representation learning and learning to learn, BPs remain a daunting challenge for modern AI. Inspired by the original one hundred BPs, we propose a new benchmark Bongard-LOGO for human-level concept learning and reasoning. We develop a program-guided generation technique to produce a large set of human-interpretable visual cognition problems in action-oriented LOGO language. Our benchmark captures three core properties of human cognition: 1) context-dependent perception, in which the same object may have disparate interpretations given different contexts; 2) analogy-making perception, in which some meaningful concepts are traded off for other meaningful concepts; and 3) perception with a few samples but infinite vocabulary. In experiments, we show that the state-of-the-art deep learning methods perform substantially worse than human subjects, implying that they fail to capture core human cognition properties. Finally, we discuss research directions towards a general architecture for visual reasoning to tackle this benchmark. Weili Nie, Zhiding Yu, Ankit B. Patel, Yuke Zhu, Anima Anandkumar |
NeurIPS | 1 |
| 2019 | RelGAN: Relational Generative Adversarial Networks for Text Generation
Weili Nie, Nina Narodytska |
ICLR (Poster) | 1 |
| 2019 | Towards a Better Understanding and Regularization of GAN Training Dynamics
Weili Nie |
UAI | 1 |
| 2018 | A Theoretical Explanation for Perplexing Behaviors of Backpropagation-based VisualizationsabstractBackpropagation-based visualizations have been proposed to interpret convolutional neural networks (CNNs), however a theory is missing to justify their behaviors: Guided backpropagation (GBP) and deconvolutional network (DeconvNet) generate more human-interpretable but less class-sensitive visualizations than saliency map. Motivated by this, we develop a theoretical explanation revealing that GBP and DeconvNet are essentially doing (partial) image recovery which is unrelated to the network decisions. Specifically, our analysis shows that the backward ReLU introduced by GBP and DeconvNet, and the local connections in CNNs are the two main causes of compelling visualizations. Extensive experiments are provided that support the theoretical analysis. Weili Nie |
ICML | 1 |
| 2018 | QG-net: a data-driven question generation model for educational contentabstractThe ever growing amount of educational content renders it increasingly difficult to manually generate sufficient practice or quiz questions to accompany it. This paper introduces QG-Net, a recurrent neural network-based model specifically designed for automatically generating quiz questions from educational content such as textbooks. QG-Net, when trained on a publicly available, general-purpose question/answer dataset and without further fine-tuning, is capable of generating high quality questions from textbooks, where the content is significantly different from the training data. Indeed, QG-Net outperforms state-of-the-art neural network-based and rules-based systems for question generation, both when evaluated using standard benchmark datasets and when using human evaluators. QG-Net also scales favorably to applications with large amounts of educational content, since its performance improves with the amount of training data. Zichao Wang 0001, Andrew S. Lan, Weili Nie, Andrew E. Waters, Phillip Grimaldi, Richard G. Baraniuk |
L@S | 3 |
| 2016 | User-Centric Cross-Tier Base Station Clustering and Cooperation in Heterogeneous Networks: Rate Improvement and Energy SavingabstractHeterogeneous cellular networks (HetNets) are to be deployed for future wireless communication to meet the ever-increasing mobile traffic demand. However, the dense and random deployment of small cells and their uncoordinated operation raise important concerns about various costs issues, among which notably is energy efficiency. Base station (BS) cooperation is set to play a key role in managing interference in HetNets. In this paper, we consider BS cooperation in the downlink HetNets where BSs from different tiers within the respective cooperative clusters jointly transmit the same data to a typical user, and in particular focus on the optimization of the energy efficiency performance. First, based on a proposed clustering model, we derive the spectral efficiency using tools from stochastic geometry. Furthermore, we formulate a power minimization problem with a minimum spectral efficiency constraint and derive the optimal received signal strength (RSS) thresholds under certain approximation. Building upon these results, we could address the problem of how to design appropriate RSS thresholds, taking into account the tradeoff between spectral efficiency and energy efficiency. Simulations show that the proposed clustering model is more energy-saving than the geometric clustering model, and deploying a multitier HetNet is significantly more energy-saving compared to a macro-only network. Weili Nie, Fu-Chun Zheng, Xiaoming Wang 0011, Wenyi Zhang 0001, Shi Jin 0002 |
IEEE J. Sel. Areas Commun. | 1 |
| 2015 | A Comparison Study of Coupled and Decoupled Uplink-Downlink Access in Heterogeneous Cellular NetworksabstractThe rapid evolution of cellular networks has brought great changes to mobile network architecture. One trend is the dense deployment of base stations (BSs) in heterogeneous cellular network (HetNets) architecture. On the other hand, the booming mobile Internet applications introduce increasingly significant imbalance in regard to Signal to Interference and Noise Ratio (SINR) statistics and traffic load between uplink (UL) and downlink (DL) in HetNets. These evolutions inspire us to exploit decoupling of UL and DL in HetNets for improving system performance. In this paper, we conduct a comparison study for the system performance of the decoupled UL/DL access (DUDA) mode and traditional coupled UL/DL access (CUDA) mode based on stochastic geometry theory. Compared to existing related work, we establish an analytical model for CUDA mode as a comparison reference and consider a more realistic system model, where we employ dynamic transmit power control in UL transmission by applying fractional power control (FPC) to model a location-dependent per-mobile power state. Numerical results reveal that DUDA mode significantly outperforms CUDA mode in terms of system rate, spectral efficiency (SE) and energy efficiency (EE) in HetNets. In addition, results also show that DUDA mode can improve load balance and fairness. Simulation results further validate the accuracy of our analytical model. Lan Zhang 0005, Gang Feng 0004, Weili Nie, Shuang Qin |
GLOBECOM | 3 |
| 2015 | Local delay and energy efficiency analysis in HetNets with random DTX schemeabstractHeterogeneous cellular networks (HetNets) are to be deployed for future wireless communication to meet the ever-increasing mobile traffic demand. However, the dense and random deployment of small cells and their uncoordinated operation raise important concerns about energy efficiency. On the other hand, discontinuous transmission (DTX) mode at the base station (BS) serves as an effective technology to improve the energy efficiency of overall system. In this paper, we investigate the energy efficiency under the finite local delay constraint in the downlink HetNets with random DTX scheme. Using a stochastic geometry based model, we derive the local delay and energy efficiency in the general case and obtain closed-form expressions in some special cases. These results give some useful insights on the system performance, taking the tradeoff between local delay and energy efficiency into account. Furthermore, we provide the low-rate and high-rate asymptotic behavior of the maximum energy efficiency. It is analytically shown that it is less energy-efficient to apply random DTX scheme in the low-rate regime. However, in the high-rate regime, random DTX scheme is essential to achieve the finite local delay and higher energy efficiency. Weili Nie, Yi Zhong 0001, Fu-Chun Zheng, Wenyi Zhang 0001 |
ICC | 1 |
| 2014 | Energy-efficient base station cooperation in downlink heterogeneous cellular networksabstractHeterogeneous cellular networks (HetNets) are to be deployed for future wireless communication to meet the ever-increasing mobile traffic demand. However, the dense and random deployment of small cells and their uncoordinated operation raise important concerns about energy efficiency. In this paper, we consider the base station (BS) cooperation solution for improving energy efficiency of the HetNets where BSs from each tier within the cooperative cluster jointly transmit the same data to a typical user. Firstly, based on the proposed clustering model, we precisely derive the ergodic rate expression using tools from stochastic geometry. Furthermore, we formulate a power minimization problem with minimum ergodic rate constraint and derive a closed-form approximated result of the optimal cooperative radii. Building upon these results, we could effectively address the problem how to design appropriate cooperative radii, taking into account the trade-off of ergodic rate and energy efficiency. Simulation results also indicate that under the proposed clustering model, deploying a two-tier HetNet is more energy-saving compared to a macro-only network. Weili Nie, Xiaoming Wang 0011, Fu-Chun Zheng, Wenyi Zhang 0001 |
GLOBECOM | 1 |