EDBT 2026 Demo / reviewers in the wild / expert
Kuan-Chieh Wang
dblp:13/7562
· DBLP profile ↗
29ranked-venue papers
4as first author
19since 2021 · last 2025
0000-0002-6785-8146ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NeuHMR: Neural Rendering-Guided Human Motion ReconstructionabstractReconstructing 3D human movements from video sequences is an important task in the fields of computer vision, graphics, and biomechanics. Although much progress has been made to infer 3D human mesh based on visual contexts provided in video sequences, generalization to in-the-wild videos still remains challenging for existing human mesh recovery (HMR) methods. To overcome inaccurate prediction, they can perform a second step optimization that refines the inaccurate estimations continuously at test time. Most optimization methods seek fitting of the body joints in the image space with respect to pseudo ground truth predicted by an off-the-shelf key point detector. However, state-of-theart detectors still introduce errors, especially for challenging poses. In this work, we rethink the dependency on the 2D key point fitting paradigm and present NeuHMR, an optimization-based mesh recovery framework based on recent advances in neural rendering. Our method builds on Human Neural Radiance Fields that allow the refinement of human meshes through animatable$2 D$renderings. We evaluated our method on two common benchmarks and validated its effectiveness. Tiange Xiang, Kuan-Chieh Wang, Jaewoo Heo, Ehsan Adeli-Mosabbeb, Serena Yeung-Levy, Scott L. Delp, Li Fei-Fei 0001 |
3DV | 2 |
| 2025 | Omni-ID: Holistic Identity Representation Designed for Generative TasksabstractWe introduce Omni-ID, a novel facial representation designed specifically for generative tasks. Omni-ID encodes holistic information about an individual’s appearance across diverse expressions and poses within a fixed-size representation. It consolidates information from a varied number of unstructured input images into a structured representation, where each entry represents certain global or local identity features. Our approach uses a few-to-many identity reconstruction training paradigm, where a limited set of input images is used to reconstruct multiple target images of the same individual in various poses and expressions. A multi-decoder framework is introduced to leverage the complementary strengths of diverse decoders during training. Unlike conventional representations, such as ArcFace and CLIP, which are typically learned through discriminative or contrastive objectives, Omni-ID is optimized with a generative objective, resulting in a more comprehensive and nuanced identity capture for generative tasks. Trained on our MFHQ dataset – a multi-view facial image collection, Omni-ID demonstrates substantial improvements over conventional representations across various generative tasks. Guocheng Qian, Kuan-Chieh Wang, Or Patashnik, Negin Heravi, Daniil Ostashev, Sergey Tulyakov, Daniel Cohen-Or, Kfir Aberman |
CVPR | 2 |
| 2025 | I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion ModelsabstractThis paper presents ThinkDiff, a novel alignment paradigm that empowers text-to-image diffusion models with multimodal in-context understanding and reasoning capabilities by integrating the strengths of vision-language models (VLMs). Existing multimodal diffusion finetuning methods largely focus on pixel-level reconstruction rather than in-context reasoning, and are constrained by the complexity and limited availability of reasoning-based datasets. ThinkDiff addresses these challenges by leveraging vision-language training as a proxy task, aligning VLMs with the decoder of an encoder-decoder large language model (LLM) instead of a diffusion decoder. This proxy task builds on the observation that the LLM decoder shares the same input feature space with diffusion decoders that use the corresponding LLM encoder for prompt embedding. As a result, aligning VLMs with diffusion decoders can be simplified through alignment with the LLM decoder. Without complex training and datasets, ThinkDiff effectively unleashes understanding, reasoning, and composing capabilities in diffusion models. Experiments demonstrate that ThinkDiff significantly improves accuracy from 19.2% to 46.3% on the challenging CoBSAT benchmark for multimodal in-context reasoning generation, with only 5 hours of training on 4 A100 GPUs. Additionally, ThinkDiff demonstrates exceptional performance in composing multiple images and texts into logically coherent images. Project page: https://mizhenxing.github.io/ThinkDiff. Zhenxing Mi, Kuan-Chieh Wang, Guocheng Qian, Hanrong Ye, Runtao Liu, Sergey Tulyakov, Kfir Aberman, Dan Xu 0002 |
ICML | 2 |
| 2025 | Preventing Shortcuts in Adapter Training via Providing the ShortcutsabstractAdapter-based training has emerged as a key mechanism for extending the capabilities of powerful foundation image generators, enabling personalized and stylized text-to-image synthesis. These adapters are typically trained to capture a specific target attribute, such as subject identity, using single-image reconstruction objectives. However, because the input image inevitably contains a mixture of visual factors, adapters are prone to entangle the target attribute with incidental ones, such as pose, expression, and lighting. This spurious correlation problem limits generalization and obstructs the model's ability to adhere to the input text prompt. In this work, we uncover a simple yet effective solution: provide the very shortcuts we wish to eliminate during adapter training. In Shortcut-Rerouted Adapter Training, confounding factors are routed through auxiliary modules, such as ControlNet or LoRA, eliminating the incentive for the adapter to internalize them. The auxiliary modules are then removed during inference. When applied to tasks like facial and full-body identity injection, our approach improves generation quality, diversity, and prompt adherence. These results point to a general design principle in the era of large models: when seeking disentangled representations, the most effective path may be to establish shortcuts for what should NOT be learned. Anujraaj Goyal, Guocheng Qian, Huseyin Coskun, Aarush Gupta, Himmy Tam, Daniil Ostashev, Ju Hu, Dhritiman Sagar, Sergey Tulyakov, Kfir Aberman, Kuan-Chieh Wang |
NeurIPS | 11 |
| 2025 | Object-level Visual Prompts for Compositional Image GenerationabstractWe introduce a method for composing object-level visual prompts within a text-to-image diffusion model. Our approach addresses the task of generating semantically coherent compositions across diverse scenes and styles, similar to the versatility and expressiveness offered by text prompts. A key challenge in this task is to preserve the identity of the objects depicted in the input visual prompts, while also generating diverse compositions across different images. To address this challenge, we introduce a new KV-mixed cross-attention mechanism, in which keys and values are learned from distinct visual representations. The keys are derived from an encoder with a small bottleneck for layout control, whereas the values come from a larger bottleneck encoder that captures fine-grained appearance details. By mixing keys and values from these complementary sources, our model preserves the identity of the visual prompts while supporting flexible variations in object arrangement, pose, and composition. During inference, we further propose object-level compositional guidance to improve the method’s identity preservation and layout correctness. Results show that our technique produces diverse scene compositions that preserve the unique characteristics of each visual prompt, expanding the creative potential of text-to-image generation. Gaurav Parmar, Or Patashnik, Kuan-Chieh Wang, Daniil Ostashev, Srinivasa G. Narasimhan, Jun-Yan Zhu, Daniel Cohen-Or, Kfir Aberman |
SIGGRAPH Asia | 3 |
| 2025 | Enhancing Virtualization Security Through System Call-Based Anomaly Detection in ContainersabstractIn the current era of micro-services, containerized applications face unprecedented security challenges due to shared kernels and limited isolation. This research proposes a container security framework based on monitoring system call sequences to detect anomalies in micro-service containers. We introduce a custom dataset named XXXX, which capture container captures system call sequences behavior in micro-services containers and simulated attacks. The framework includes real-time system call monitors, parsers, dashboards, and an unsupervised anomaly detection model using unsupervised learning with autoencoders to enhance the detection capability of unknown vulnerabilities. It leverages containerization benefits-simplicity, scalability, and automation. Our evaluation emphasizes false alarm rate and average detection time. Results show that the attack detection performance of most containers meets expectations, though the detection time of one subset had slightly longer detection time due to the intrinsic complexity of vulnerabilities. This work offers valuable insights for improving container security in microservice systems. Kuan-Chieh Wang, Po-Kai Hsu, Jhen-Jie Hsieh, Po-Shen Chen, Tze-Rong Jian, Kun-Hsiang Huang, Min-Te Sun, Chun-Ying Huang |
TENCON | 2 |
| 2024 | Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
James Burgess, Kuan-Chieh Wang, Serena Yeung-Levy |
ECCV (64) | 2 |
| 2024 | VIMI: Grounding Video Generation through Multi-modal InstructionabstractYuwei Fang, Willi Menapace, Aliaksandr Siarohin, Tsai-Shien Chen, Kuan-Chieh Wang, Ivan Skorokhodov, Graham Neubig, Sergey Tulyakov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yuwei Fang, Willi Menapace, Aliaksandr Siarohin, Tsai-Shien Chen, Kuan-Chieh Wang, Ivan Skorokhodov, Graham Neubig, Sergey Tulyakov |
EMNLP | 5 |
| 2024 | Interpreting the Weight Space of Customized Diffusion ModelsabstractWe investigate the space of weights spanned by a large collection of customized diffusion models. We populate this space by creating a dataset of over 60,000 models, each of which is a base model fine-tuned to insert a different person's visual identity. We model the underlying manifold of these weights as a subspace, which we term $\textit{weights2weights}$. We demonstrate three immediate applications of this space that result in new diffusion models -- sampling, editing, and inversion. First, sampling a set of weights from this space results in a new model encoding a novel identity. Next, we find linear directions in this space corresponding to semantic edits of the identity (e.g., adding a beard), resulting in a new model with the original identity edited. Finally, we show that inverting a single image into this space encodes a realistic identity into a model, even if the input image is out of distribution (e.g., a painting). We further find that these linear properties of the diffusion model weight space extend to other visual concepts. Our results indicate that the weight space of fine-tuned diffusion models can behave as an interpretable $\textit{meta}$-latent space producing new models. Amil Dravid, Yossi Gandelsman, Kuan-Chieh Wang, Rameen Abdal, Gordon Wetzstein, Alexei A. Efros, Kfir Aberman |
NeurIPS | 3 |
| 2024 | MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation
Kuan-Chieh Wang, Daniil Ostashev, Yuwei Fang, Sergey Tulyakov, Kfir Aberman |
SIGGRAPH Asia | 1 |
| 2023 | NeMo: 3D Neural Motion Fields from Multiple Video Instances of the Same ActionabstractThe task of reconstructing 3D human motion has wide-ranging applications. The gold standard Motion capture (MoCap) systems are accurate but inaccessible to the general public due to their cost, hardware, and space constraints. In contrast, monocular human mesh recovery (HMR) methods are much more accessible than MoCap as they take single-view videos as inputs. Replacing the multi-view MoCap systems with a monocular HMR method would break the current barriers to collecting accurate 3D motion thus making exciting applications like motion analysis and motion-driven animation accessible to the general public. However, the performance of existing HMR methods degrades when the video contains challenging and dynamic motion that is not in existing MoCap datasets used for training. This reduces its appeal as dynamic motion is frequently the target in 3D motion recovery in the aforementioned applications. Our study aims to bridge the gap between monocular HMR and multi-view MoCap systems by leveraging information shared across multiple video instances of the same action. We introduce the Neural Motion (NeMo) field. It is optimized to represent the underlying 3D motions across a set of videos of the same action. Empirically, we show that NeMo can recover 3D motion in sports using videos from the Penn Action dataset, where NeMo outperforms existing HMR methods in terms of 2D keypoint detection. To further validate NeMo using 3D metrics, we collected a small MoCap dataset mimicking actions in Penn Action, and show that NeMo achieves better 3D reconstruction compared to various baselines. Kuan-Chieh Wang, Zhenzhen Weng, Maria Xenochristou, João Pedro Araújo 0001, Jeffrey Gu, C. Karen Liu, Serena Yeung-Levy |
CVPR | 1 |
| 2023 | PROB: Probabilistic Objectness for Open World Object DetectionabstractOpen World Object Detection (OWOD) is a new and challenging computer vision task that bridges the gap between classic object detection (OD) benchmarks and object detection in the real world. In addition to detecting and classifying seen/labeled objects, OWOD algorithms are expected to detect novel/unknown objects - which can be classified and incrementally learned. In standard OD, object proposals not overlapping with a labeled object are automatically classified as background. Therefore, simply applying OD methods to OWOD fails as unknown objects would be predicted as background. The challenge of detecting unknown objects stems from the lack of supervision in distinguishing unknown objects and background object proposals. Previous OWOD methods have attempted to overcome this issue by generating supervision using pseudo-labeling - however, unknown object detection has remained low. Probabilistic/generative models may provide a solution for this challenge. Herein, we introduce a novel probabilistic framework for objectness estimation, where we alternate between probability distribution estimation and objectness likelihood maximization of known objects in the embedded feature space - ultimately allowing us to estimate the objectness probability of different proposals. The resulting Probabilistic Objectness transformer-based open-world detector, PROB, integrates our framework into traditional object detection models, adapting them for the open-world setting. Comprehensive experiments on OWOD benchmarks show that PROB outperforms all existing OWOD methods in both unknown object detection (~ 2 × unknown recall) and known object detection (~ 10% mAP). Our code is available at https://github.com/orrzohar/PROB. Orr Zohar, Kuan-Chieh Wang, Serena Yeung-Levy |
CVPR | 2 |
| 2023 | Generalizable Neural Fields as Partially Observed Neural ProcessesabstractNeural fields, which represent signals as a function parameterized by a neural network, are a promising alternative to traditional discrete vector or grid-based representations. Compared to discrete representations, neural representations both scale well with increasing resolution, are continuous, and can be many-times differentiable. However, given a dataset of signals that we would like to represent, having to optimize a separate neural field for each signal is inefficient, and cannot capitalize on shared information or structures among signals. Existing generalization methods view this as a meta-learning problem and employ gradient-based meta-learning to learn an initialization which is then fine-tuned with test-time optimization, or learn hypernetworks to produce the weights of a neural field. We instead propose a new paradigm that views the large-scale training of neural representations as a part of a partially-observed neural process framework, and leverage neural process algorithms to solve this task. We demonstrate that this approach outperforms both state-of-the-art gradient-based meta-learning approaches and hypernetwork approaches. Jeffrey Gu, Kuan-Chieh Wang, Serena Yeung-Levy |
ICCV | 2 |
| 2023 | Diagnosing and Rectifying Vision Models using Language
Jeff Z. HaoChen, Shih-Cheng Huang, Kuan-Chieh Wang, James Zou 0001, Serena Yeung-Levy |
ICLR | 4 |
| 2023 | LOVM: Language-Only Vision Model SelectionabstractPre-trained multi-modal vision-language models (VLMs) are becoming increasingly popular due to their exceptional performance on downstream vision applications, particularly in the few- and zero-shot settings. However, selecting the best-performing VLM for some downstream applications is non-trivial, as it is dataset and task-dependent. Meanwhile, the exhaustive evaluation of all available VLMs on a novel application is not only time and computationally demanding but also necessitates the collection of a labeled dataset for evaluation. As the number of open-source VLM variants increases, there is a need for an efficient model selection strategy that does not require access to a curated evaluation dataset. This paper proposes a novel task and benchmark for efficiently evaluating VLMs' zero-shot performance on downstream applications without access to the downstream task dataset. Specifically, we introduce a new task LOVM: Language-Only Vision Model Selection , where methods are expected to perform both model selection and performance prediction based solely on a text description of the desired downstream application. We then introduced an extensive LOVM benchmark consisting of ground-truth evaluations of 35 pre-trained VLMs and 23 datasets, where methods are expected to rank the pre-trained VLMs and predict their zero-shot performance. Orr Zohar, Shih-Cheng Huang, Kuan-Chieh Wang, Serena Yeung-Levy |
NeurIPS | 3 |
| 2022 | Domain Adaptive 3D Pose Augmentation for In-the-Wild Human Mesh RecoveryabstractThe ability to perceive 3D human bodies from a single image has a multitude of applications ranging from entertainment and robotics to neuroscience and healthcare. A fundamental challenge in human mesh recovery is in collecting the ground truth 3D mesh targets required for training, which requires burdensome motion capturing systems and is often limited to indoor laboratories. As a result, while progress is made on benchmark datasets collected in these restrictive settings, models fail to generalize to real-world "in-the-wild" scenarios due to distribution shifts. We propose Domain Adaptive 3D Pose Augmentation (DAPA), a data augmentation method that enhances the model's generalization ability in in-the-wild scenarios. DAPA combines the strength of methods based on synthetic datasets by getting direct supervision from the synthesized meshes, and domain adaptation methods by using ground truth 2D keypoints from the target dataset. We show quantitatively that finetuning with DAPA effectively improves results on benchmarks 3DPW [38]and AGORA[32]. We further demonstrate the utility of DAPA on a challenging dataset curated from videos of real-world parent-child interaction. Zhenzhen Weng, Kuan-Chieh Wang, Angjoo Kanazawa, Serena Yeung-Levy |
3DV | 2 |
| 2021 | Understanding and Mitigating Exploding Inverses in Invertible Neural NetworksabstractInvertible neural networks (INNs) have been used to design generative models, implement memory-saving gradient computation, and solve inverse problems. In this work, we show that commonly-used INN architectures suffer from exploding inverses and are thus prone to becoming numerically non-invertible. Across a wide range of INN use-cases, we reveal failures including the non-applicability of the change-of-variables formula on in- and out-of-distribution (OOD) data, incorrect gradients for memory-saving backprop, and the inability to sample from normalizing flow models. We further derive bi-Lipschitz properties of atomic building blocks of common architectures. These insights into the stability of INNs then provide ways forward to remedy these failures. For tasks where local invertibility is sufficient, like memory-saving backprop, we propose a flexible and efficient regularizer. For problems where global invertibility is necessary, such as applying normalizing flows on OOD data, we show the importance of designing stable INN building blocks. Jens Behrmann, Paul Vicol, Kuan-Chieh Wang, Roger B. Grosse, Jörn-Henrik Jacobsen |
AISTATS | 3 |
| 2021 | Variational Model Inversion AttacksabstractGiven the ubiquity of deep neural networks, it is important that these models do not reveal information about sensitive data that they have been trained on. In model inversion attacks, a malicious user attempts to recover the private dataset used to train a supervised neural network. A successful model inversion attack should generate realistic and diverse samples that accurately describe each of the classes in the private dataset. In this work, we provide a probabilistic interpretation of model inversion attacks, and formulate a variational objective that accounts for both diversity and accuracy. In order to optimize this variational objective, we choose a variational family defined in the code space of a deep generative model, trained on a public auxiliary dataset that shares some structural similarity with the target dataset. Empirically, our method substantially improves performance in terms of target attack accuracy, sample realism, and diversity on datasets of faces and chest X-ray images. Kuan-Chieh Wang, Ke Li 0011, Ashish Khisti, Richard S. Zemel, Alireza Makhzani |
NeurIPS | 1 |
| 2021 | Grad2Task: Improved Few-shot Text Classification Using Gradients for Task RepresentationabstractLarge pretrained language models (LMs) like BERT have improved performance in many disparate natural language processing (NLP) tasks. However, fine tuning such models requires a large number of training examples for each target task. Simultaneously, many realistic NLP problems are "few shot", without a sufficiently large training set. In this work, we propose a novel conditional neural process-based approach for few-shot text classification that learns to transfer from other diverse tasks with rich annotation. Our key idea is to represent each task using gradient information from a base model and to train an adaptation network that modulates a text classifier conditioned on the task representation. While previous task-aware few-shot learners represent tasks by input encoding, our novel task representation is more powerful, as the gradient captures input-output relationships of a task. Experimental results show that our approach outperforms traditional fine-tuning, sequential transfer learning, and state-of-the-art meta learning approaches on a collection of diverse few-shot tasks. We further conducted analysis and ablations to justify our design choices. Jixuan Wang, Kuan-Chieh Wang, Frank Rudzicz, Michael Brudno |
NeurIPS | 2 |
| 2020 | Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi 0002, Kevin Swersky |
ICLR | 2 |
| 2020 | Learning the Stein Discrepancy for Training and Evaluating Energy-Based Models without SamplingabstractWe present a new method for evaluating and training unnormalized density models. Our approach only requires access to the gradient of the unnormalized model’s log-density. We estimate the Stein discrepancy between the data density p(x) and the model density q(x) based on a vector function of the data. We parameterize this function with a neural network and fit its parameters to maximize this discrepancy. This yields a novel goodness-of-fit test which outperforms existing methods on high dimensional data. Furthermore, optimizing q(x) to minimize this discrepancy produces a novel method for training unnormalized models. This training method can fit large unnormalized models faster than existing approaches. The ability to both learn and compare models is a unique feature of the proposed method. Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Richard S. Zemel |
ICML | 2 |
| 2020 | A crosswalk pedestrian recognition system by using deep learning and zebra-crossing recognition techniquesabstractSummary Pedestrian detection is essential for improving pedestrian safety in an intelligent traffic system. The efficiency of the system is affected by real‐time processing and the error rate of detection. These concerns have not been completely addressed in previous studies. Therefore, this study proposes a real‐time pedestrian recognition system that ensures high accuracy by using a deep learning classifier and zebra‐crossing recognition techniques. The proposed system was designed to improve pedestrian safety and reduce accidents at intersections. Environmental feature vectors were first used to detect zebra crossings and to determine crossing areas. An adaptive mapping technique was then used to map the pedestrian waiting area based on the crossing area. A dual camera mechanism was used to maintain detection accuracy and improve system fault tolerance. Finally, the you‐only‐look‐once model was used to recognize pedestrians at intersections. A system prototype was implemented to verify the feasibility of the proposed system. The results revealed that the proposed scheme outperforms the conventional histogram of oriented gradients and Haarcascade schemes. Chyi-Ren Dow, Ngo Huu Huy, Liang-Hsuan Lee, Po-Yu Lai, Kuan-Chieh Wang, Van-Tung Bui |
Softw. Pract. Exp. | 5 |
| 2019 | Centroid-based Deep Metric Learning for Speaker RecognitionabstractSpeaker embedding models that utilize neural networks to map utterances to a space where distances reflect similarity between speakers have driven recent progress in the speaker recognition task. However, there is still a significant performance gap between recognizing speakers in the training set and unseen speakers. The latter case corresponds to the few-shot learning task, where a trained model is evaluated on unseen classes. Here, we optimize a speaker embedding model with prototypical network loss (PNL), a state-of-the-art approach for the few-shot image classification task. The resulting embedding model outperforms the state-of-the-art triplet loss based models in both speaker verification and identification tasks, for both seen and unseen speakers. Jixuan Wang, Kuan-Chieh Wang, Marc T. Law, Frank Rudzicz, Michael Brudno |
ICASSP | 2 |
| 2019 | Ultra-Low Complexity Block-Based Lane Detection and Departure Warning SystemabstractThis paper proposes an ultra-low complexity block-based lane detection and departure warning system. Based on the distribution of the lane markings in the region close to the vehicle, a parameterized region of interest (ROI) is determined. The lane markings in the ROI are enhanced by increasing the pixel intensity for detection in various environmental conditions. To reduce the computational burden, the ROI is partitioned into non-overlapping blocks and two simplified masks are proposed to obtain the block gradients and block angles. The driving conditions are classified into four classes to simplify the lane detection process and the proposed lane departure warning system is based on the lane detection results. The experimental results reveal that the average lane detection rate and the departure warning rate are 96.12% and 98.60%, respectively. With a 1920 × 1080 resolution, the average processing time is 4.28 ms per frame. Chung-Bin Wu, Li-Hung Wang, Kuan-Chieh Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Neural Relational Inference for Interacting SystemsabstractInteracting systems are prevalent in nature, from dynamical systems in physics to complex societal dynamics. The interplay of components can give rise to complex behavior, which can often be explained using a simple model of the system’s constituent parts. In this work, we introduce the neural relational inference (NRI) model: an unsupervised model that learns to infer interactions while simultaneously learning the dynamics purely from observational data. Our model takes the form of a variational auto-encoder, in which the latent code represents the underlying interaction graph and the reconstruction is based on graph neural networks. In experiments on simulated physical systems, we show that our NRI model can accurately recover ground-truth interactions in an unsupervised manner. We further demonstrate that we can find an interpretable structure and predict complex dynamics in real motion capture and sports tracking data. Thomas Kipf, Ethan Fetaya, Kuan-Chieh Wang, Max Welling, Richard S. Zemel |
ICML | 3 |
| 2018 | Adversarial Distillation of Bayesian Neural Network PosteriorsabstractBayesian neural networks (BNNs) allow us to reason about uncertainty in a principled way. Stochastic Gradient Langevin Dynamics (SGLD) enables efficient BNN learning by drawing samples from the BNN posterior using mini-batches. However, SGLD and its extensions require storage of many copies of the model parameters, a potentially prohibitive cost, especially for large neural networks. We propose a framework, Adversarial Posterior Distillation, to distill the SGLD samples using a Generative Adversarial Network (GAN). At test-time, samples are generated by the GAN. We show that this distillation framework incurs no loss in performance on recent BNN applications including anomaly detection, active learning, and defense against adversarial attacks. By construction, our framework distills not only the Bayesian predictive distribution, but the posterior itself. This allows one to compute quantities such as the approximate model variance, which is useful in downstream tasks. To our knowledge, these are the first results applying MCMC-based BNNs to the aforementioned applications. Kuan-Chieh Wang, Paul Vicol, James Lucas, Li Gu, Roger B. Grosse, Richard S. Zemel |
ICML | 1 |
| 2017 | Dualing GANsabstractGenerative adversarial nets (GANs) are a promising technique for modeling a distribution from samples. It is however well known that GAN training suffers from instability due to the nature of its saddle point formulation. In this paper, we explore ways to tackle the instability problem by dualizing the discriminator. We start from linear discriminators in which case conjugate duality provides a mechanism to reformulate the saddle point objective into a maximization problem, such that both the generator and the discriminator of this ‘dualing GAN’ act in concert. We then demonstrate how to extend this intuition to non-linear formulations. For GANs with linear discriminators our approach is able to remove the instability in training, while for GANs with nonlinear discriminators our approach provides an alternative to the commonly used GAN training algorithm. Yujia Li 0001, Alexander G. Schwing, Kuan-Chieh Wang, Richard S. Zemel |
NIPS | 3 |
| 2011 | Green Power Management with Dynamic Resource Allocation for Cloud Virtual MachinesabstractWith the development of electronics in governments and business, the implementation of these services are increasing demand for servers. Continued expansion of servers represents our need for more space, power, air conditioning, network, human resources and other infrastructure. Regardless of how powerful servers now become, we do not make good use of all resources and strive for the waste. In this paper, the Green Power Management (GPM) is proposed for load balancing for virtual machine management on cloud. It includes three main phrases: (1) supporting green power mechanism, (2) implementing virtual machine resource monitor onto Open Nebula with web-based interface, and (3) integrating a Dynamic Resource Allocation (DRA) and Open Nebula functions as bases instead of traditionally booting physical machines with command mode. Chao-Tung Yang, Kuan-Chieh Wang, Hsiang-Yao Cheng, Cheng-Ta Kuo, William C. Chu |
HPCC | 2 |
| 2011 | Implementation of a Green Power Management Algorithm for Virtual Machines on Cloud Computing
Chao-Tung Yang, Kuan-Chieh Wang, Hsiang-Yao Cheng, Cheng-Ta Kuo, Ching-Hsien Hsu |
UIC | 2 |