EDBT 2026 Demo / reviewers in the wild / expert
Ruixiang Zhang
dblp:20/9860
· DBLP profile ↗
30ranked-venue papers
11as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 6 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Oriented Tiny Object Detection: A Dataset, Benchmark, and Dynamic Unbiased LearningabstractDetecting oriented tiny objects, which are limited in appearance information yet prevalent in real-world applications, remains an intricate and under-explored problem. To address this, we systematically introduce a new dataset, a benchmark, and a dynamic coarse-to-fine learning scheme in this study. Our proposed dataset, AI-TOD-R, features the smallest object sizes among all oriented object detection datasets. Based on AI-TOD-R, we present a benchmark spanning a broad range of detection paradigms, including both fully-supervised and label-efficient approaches. Through investigation, we identify a learning bias presents across various learning pipelines: confident objects become increasingly confident, while vulnerable oriented tiny objects are further marginalized, hindering their detection performance. To mitigate this issue, we propose a Dynamic Coarse-to-Fine Learning (DCFL) scheme towards unbiased learning. DCFL dynamically updates prior positions to better align with the limited areas of oriented tiny objects, and it assigns samples in a way that balances both quantity and quality across different object shapes, thus mitigating biases in prior settings and sample selection. Extensive experiments across 10 challenging object detection datasets demonstrate that DCFL achieves state-of-the-art accuracy, high efficiency, and remarkable versatility. Chang Xu 0027, Ruixiang Zhang, Wen Yang 0001, Jian Ding 0001, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | DN-TOD: Robust tiny object detection amidst label noise
Chang Xu 0027, Wen Yang 0001, Ruixiang Zhang, Yan Zhang 0115, Gui-Song Xia |
Pattern Recognit. | 4 |
| 2025 | Composition and Control with Distilled Energy Diffusion Models and Sequential Monte CarloabstractDiffusion models may be formulated as a time-indexed sequence of energy-based models, where the score corresponds to the negative gradient of an energy function. As opposed to learning the score directly, an energy parameterization is attractive as the energy itself can be used to control generation via Monte Carlo samplers. Architectural constraints and training instability in energy parameterized models have so far yielded inferior performance compared to directly approximating the score or denoiser. We address these deficiencies by introducing a novel training regime for the energy function through distillation of pre-trained diffusion models, resembling a Helmholtz decomposition of the score vector field. We further showcase the synergies between energy and score by casting the diffusion sampling procedure as a Feynman Kac Model where sampling is controlled using potentials from the learnt energy functions. The Feynman Kac model formalism enables composition and low temperature sampling through sequential Monte Carlo. James Thornton, Louis Béthune, Ruixiang Zhang, Arwen Bradley, Preetum Nakkiran, Shuangfei Zhai |
AISTATS | 3 |
| 2025 | ChipChat: Low-Latency Cascaded Conversational Agent in MLXabstractThe emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages, cascaded systems (CSs) continue to outperform them in language understanding tasks, despite being constrained by sequential processing latency. In this work, we introduce ChipChat, a novel low-latency CS that overcomes traditional bottlenecks through architectural innovations and streaming optimizations. Our system integrates streaming (a) conversational speech recognition with mixture-of-experts, (b) state-action augmented LLM, (c) text-to-speech synthesis, (d) neural vocoder, and (e) speaker modeling. Implemented using MLX, ChipChat achieves subsecond response latency on a Mac Studio without dedicated GPUs, while preserving user privacy through complete on-device processing. Our work shows that strategically redesigned CSs can overcome their historical latency limitations, offering a promising path forward for practical voice-based AI agents. Tatiana Likhomanenko, Luke Carlson, He Bai 0002, Zijin Gu, Han Tran, Zakaria Aldeneh, Yizhe Zhang 0002, Ruixiang Zhang, Huangjie Zheng, Navdeep Jaitly |
ASRU | 8 |
| 2025 | Normalizing Flows are Capable Generative ModelsabstractNormalizing Flows (NFs) are likelihood-based models for continuous inputs. They have demonstrated promising results on both density estimation and generative modeling tasks, but have received relatively little attention in recent years. In this work, we demonstrate that NFs are more powerful than previously believed. We present TarFlow: a simple and scalable architecture that enables highly performant NF models. TarFlow can be thought of as a Transformer-based variant of Masked Autoregressive Flows (MAFs): it consists of a stack of autoregressive Transformer blocks on image patches, alternating the autoregression direction between layers. TarFlow is straightforward to train end-to-end, and capable of directly modeling and generating pixels. We also propose three key techniques to improve sample quality: Gaussian noise augmentation during training, a post training denoising procedure, and an effective guidance method for both class-conditional and unconditional settings. Putting these together, TarFlow sets new state-of-the-art results on likelihood estimation for images, beating the previous best methods by a large margin, and generates samples with quality and diversity comparable to diffusion models, for the first time with a stand-alone NF model. We make our code available at https://github.com/apple/ml-tarflow. Shuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot, Jiatao Gu, Huangjie Zheng, Tianrong Chen, Miguel Ángel Bautista 0001, Navdeep Jaitly, Joshua M. Susskind |
ICML | 2 |
| 2025 | Target Concrete Score Matching: A Holistic Framework for Discrete DiffusionabstractDiscrete diffusion is a promising framework for modeling and generating discrete data. In this work, we present Target Concrete Score Matching (TCSM), a novel and versatile objective for training and fine-tuning discrete diffusion models. TCSM provides a general framework with broad applicability. It supports pre-training discrete diffusion models directly from data samples, and many existing discrete diffusion approaches naturally emerge as special cases of our more general TCSM framework. Furthermore, the same TCSM objective extends to post-training of discrete diffusion models, including fine-tuning using reward functions or preference data, and distillation of knowledge from pre-trained autoregressive models. These new capabilities stem from the core idea of TCSM, estimating the concrete score of the target distribution, which resides in the original (clean) data space. This allows seamless integration with reward functions and pre-trained models, which inherently only operate in the clean data space rather than the noisy intermediate spaces of diffusion processes. Our experiments on language modeling tasks demonstrate that TCSM matches or surpasses current methods. Additionally, TCSM is versatile, applicable to both pre-training and post-training scenarios, offering greater flexibility and sample efficiency. Ruixiang Zhang, Shuangfei Zhai, Yizhe Zhang 0002, James Thornton, Zijing Ou, Joshua M. Susskind, Navdeep Jaitly |
ICML | 1 |
| 2025 | STARFlow: Scaling Latent Normalizing Flows for High-resolution Image SynthesisabstractWe present STARFlow, a scalable generative model based on normalizing flows that achieves strong performance on high-resolution image synthesis.
STARFlow's main building block is Transformer Autoregressive Flow (TARFlow), which combines normalizing flows with Autoregressive Transformer architectures and has recently achieved impressive results in image modeling. In this work, we first establish the theoretical universality of TARFlow for modeling continuous distributions. Building on this foundation, we introduce a set of architectural and algorithmic innovations that significantly enhance the scalability: (1) a deep-shallow design where a deep Transformer block captures most of the model’s capacity, followed by a few shallow Transformer blocks that are computationally cheap yet contribute non-negligibly, (2) learning in the latent space of pretrained autoencoders, which proves far more effective than modeling pixels directly, and (3) a novel guidance algorithm that substantially improves sample quality. Crucially, our model remains a single, end-to-end normalizing flow, allowing exact maximum likelihood training in continuous space without discretization. STARFlow achieves competitive results in both class- and text-conditional image generation, with sample quality approaching that of state-of-the-art diffusion models. To our knowledge, this is the **first** successful demonstration of normalizing flows at this scale and resolution. Code and weights available at https://github.com/apple/ml-starflow. Jiatao Gu, Tianrong Chen, David Berthelot, Huangjie Zheng, Ruixiang Zhang, Laurent Dinh, Miguel Ángel Bautista 0001, Joshua M. Susskind, Shuangfei Zhai |
NeurIPS | 6 |
| 2025 | Discrete Neural Flow Samplers with Locally Equivariant TransformerabstractSampling from unnormalised discrete distributions is a fundamental problem across various domains.
While Markov chain Monte Carlo offers a principled approach, it often suffers from slow mixing and poor convergence.
In this paper, we propose Discrete Neural Flow Samplers (DNFS), a trainable and efficient framework for discrete sampling. DNFS learns the rate matrix of a continuous-time Markov chain such that the resulting dynamics satisfy the Kolmogorov equation.
As this objective involves the intractable partition function, we then employ control variates to reduce the variance of its Monte Carlo estimation, leading to a coordinate descent learning algorithm.
To further facilitate computational efficiency, we propose locally equivaraint Transformer, a novel parameterisation of the rate matrix that significantly improves training efficiency while preserving powerful network expressiveness.
Empirically, we demonstrate the efficacy of DNFS in a wide range of applications, including sampling from unnormalised distributions, training discrete energy-based models, and solving combinatorial optimisation problems. Zijing Ou, Ruixiang Zhang, Yingzhen Li |
NeurIPS | 2 |
| 2025 | Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive FlowsabstractAutoregressive models have driven remarkable progress in language modeling. Their foundational reliance on discrete tokens, unidirectional context, and single-pass decoding, while central to their success, also inspires the exploration of a design space that could offer new axes of modeling flexibility.
In this work, we explore an alternative paradigm, shifting language modeling from a discrete token space to a continuous latent space.
We propose a novel framework that employs transformer-based autoregressive normalizing flows to model these continuous representations.
This approach unlocks substantial flexibility, enabling the construction of models that can capture global bi-directional context through stacked, alternating-direction autoregressive transformations, support block-wise generation with flexible token patch sizes, and facilitate a hierarchical multi-pass generation process.
We further propose new mixture-based coupling transformations designed to capture complex dependencies within the latent space shaped by discrete data, and demonstrate theoretical connections to conventional discrete autoregressive models.
Extensive experiments on language modeling benchmarks demonstrate strong likelihood performance and highlight the flexible modeling capabilities inherent in our framework. Ruixiang Zhang, Shuangfei Zhai, Jiatao Gu, Yizhe Zhang 0002, Huangjie Zheng, Tianrong Chen, Miguel Ángel Bautista 0001, Joshua M. Susskind, Navdeep Jaitly |
NeurIPS | 1 |
| 2025 | A Dual Two-Stage Attention-based Model for interpretable hard landing prediction from flight data
Jiaxing Shang, Xiaoquan Li, Ruixiang Zhang, Linjiang Zheng, Xu Li 0014, Riquan Zhang, Xinbin Zhao, Fan Li 0020 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Minimizing Sample Redundancy for Label-Efficient Object Detection in Aerial ImagesabstractObjects in aerial images tend to be densely scattered and appear in arbitrary orientations, making the annotation process quite costly. To reduce the annotation cost, existing methods propose randomly annotating a proportion of images or objects for aerial object detection with fewer label usage. These approaches, however, can lead to redundancy in labels and inherit the biases associated with the imbalance in datasets. To minimize sample redundancy and alleviate data imbalance, we propose a novel labeling pattern that acquires heterogeneous object labels in a class-orthogonal manner, preserving a broader diversity of samples for each category with less annotation effort. To improve data utility, we design a Dynamic Multi-View Learning (DML) strategy to overcome the sample quantity-quality dilemma in current pseudo-labeling methods—a high pseudo-label threshold reduces sample quantity, while low thresholds compromise sample quality. First, DML separates model predictions into multiple hierarchies for finer screening, mitigating the suppression of unlabeled objects in binary pseudo-label strategies. With this separation, DML learns to construct a new view by injecting high-quality samples and masking low-quality regions in this view, simultaneously expanding sample quantity while ensuring sample quality. Unlike previous methods that mine pseudo labels solely from unlabelled regions, DML releases this constraint by learning to expand high-quality samples with a dynamic view. Extensive experiments on five benchmark datasets validate our method’s state-of-the-art accuracy and label efficiency. Notably, with approximately 5% DOTA-v2.0 annotations, DML achieves nearly 90% of the fully supervised performance. The codes will be available at https://github.com/ZhangRuixiang-WHU/ALOD_DML/. Ruixiang Zhang, Chang Xu 0027, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | AODet: Anti-Occlusion for Enhanced Small Object Detection in Drone-Based RGBT ImageryabstractDrone-based RGBT person detection promotes various applications such as search and rescue due to its maneuverability. While existing research predominantly concentrates on refining fusion strategies and bolstering learning mechanisms for small objects, the pervasive yet unique occlusion challenge in drone-based RGBT settings remains inadequately addressed. In this work, we address the unique challenge of occlusion in the context of RGBT small object detection, particularly emphasizing its vulnerability and the distinct characteristics it exhibits across different modalities. We propose AODet, a novel Anti-Occlusion Detector meticulously crafted to tackle the challenges posed by occlusion in drone-based RGBT object detection. Our proposed approach significantly improves the detection performance of RGBT small objects, surpassing strong baselines on two large-scale datasets, VTUAV-det and RGBTDronePerson, by 1.30 points and 2.24 points in mAPsand ${\text{mAP}}_{50}^{{\text{tiny}}}$, respectively. Ziming Gui, Yan Zhang 0115, Xu Lei 0002, Ruixiang Zhang, Wen Yang 0001 |
IGARSS | 4 |
| 2024 | Cross-Resolution Distillation for Building Extraction in Medium-Resolution Satellite ImageryabstractExtracting buildings in Earth observation is a vital task for various applications. Medium-resolution remote sensing images can offer notable advantages like reduced storage needs and faster acquisition, facilitating more rapid revisits of a certain area. However, they often lack detailed building features, posing challenges for accurate segmentation. To address this, we present CDNet (Cross-resolution Distillation Network), a novel framework designed to boost building extraction accuracy on medium-resolution images, resulting in computational efficiency and low time cost. CDNet operates on a teacher-student network paradigm, employing cross-resolution knowledge distillation. The teacher network processes super-resolved images, endowing the student network with crucial priors to adeptly handle lower-resolution originals. Our method achieves a cutting-edge performance on the Multi-Temporal Urban Development SpaceNet (MUDS) Dataset, showcasing exceptional accuracy with a mean Intersection over Union (mIoU) of 62.97 and Boundary over Union (BIoU) of 29.35, respectively. Shuailin Chen, Ruixiang Zhang, Pingping Huang |
IGARSS | 3 |
| 2024 | Decoupled Multi-Teacher: Cross-Modal Learning Enhanced Object Detection in SAR ImageryabstractObject Detection in Synthetic Aperture Radar (SAR) images holds significant potential for diverse remote sensing applications. Nevertheless, the substantial cost associated with object-level annotation poses a formidable challenge in SAR object detection. While cross-modal learning offers a promising solution to address this annotation hurdle by using annotated optical images, existing methods falter in achieving satisfactory results due to the significant modal gap between optical and SAR. In this study, we introduce the Decoupled Multi-Teacher (DMT) framework, specifically tailored for object detection in SAR images. By strategically decoupling cross-modal learning from detector training, DMT adeptly mitigates challenges associated with co-training disparate modalities. Our experiments showcase the substantial improvements achieved by the proposed method in SAR image object detection. Ruixiang Zhang, Pingping Huang |
IGARSS | 1 |
| 2024 | Beyond Dehazing: Learning Intrinsic Hazy Robustness for Aerial Object DetectionabstractAccurate object detection in aerial imagery is crucial across numerous applications. However, haze can significantly degrade the performance of normal detectors, presenting a substantial obstacle in real-world scenarios. Previous solutions often resort to image dehazing as a pre-processing step to enhance image quality for subsequent detection. Despite being logically intuitive, their performance is limited due to the inherent objective mismatch between low-level image restoration tasks and high-level object detection tasks. In this article, we present haze-robust aerial object detection (HRAOD) to directly enhance detection robustness under hazy conditions. HRAOD constructs a clean-to-hazy distillation framework, enabling the detector to “see through haze,” without relying on the explicit image dehazing process. To address the challenge of extracting informative hazy features from blurry and low-contrast hazy images, we introduce a gradient-guided feature imitation method to emphasize the desired objects. Moreover, recognizing that different regions suffer from varying degradation degrees and pose distinct detection difficulties, we further propose a degradation-weighted response distillation method to mimic the normal predictions according to the degradation pattern adaptively. Due to the scarcity of hazy aerial data, we curate two remote sensing hazy aerial datasets, namely DOTA-Haze and SODA-A-Haze, and one drone hazy aerial dataset, DroneVehicle-Haze, for simulation. Extensive experimental results demonstrate the superiority of our method. Specifically, our HRAOD outperforms the state-of-the-art “dehaze + detect” method by 13.1 points in mAP on the DOTA-Haze dataset without incurring additional inference costs. HRAOD also performs favorably against other methods on real-world hazy scenes. Yan Zhang 0115, Ruixiang Zhang, Wen Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning
Ting Chen 0007, Ruixiang Zhang, Geoffrey E. Hinton |
ICLR | 2 |
| 2023 | Robust and Controllable Object-Centric Learning through Energy-based Models
Ruixiang Zhang, Tong Che, Boris Ivanovic, Renhao Wang, Marco Pavone 0001, Yoshua Bengio, Liam Paull |
ICLR | 1 |
| 2022 | Learning Representation from Neural Fisher Kernel with Low-rank Approximation
Ruixiang Zhang, Shuangfei Zhai, Etai Littwin, Joshua M. Susskind |
ICLR | 1 |
| 2022 | Optical-Enhanced Oil Tank Detection in High-Resolution SAR ImagesabstractIn recent years, object detection in high-resolution SAR images has made significant progress, especially after the introduction of deep learning. However, objects like dense oil tanks, which are compactly arranged in SAR images, are still challenging to recognize due to the unique imaging mechanism of SAR. Inspired by human learning from comparison, we propose a multi-stage framework for oil tank detection in SAR images using optical image enhancement. Specifically, in the training stage, we build a teacher-student network to align the semantic information between the two modalities, where the optical features are used to guide the corresponding SAR feature learning. While in the inference stage, the learned network detects oil tanks using only SAR images as input. Besides, a pre-training stage before training is applied to further improve the network’s ability for SAR feature extraction, which is realized by the proposed paired optical-SAR self-supervised learning. To verify the effectiveness of the proposed method, we perform experiments on our newly built SpaceNet6-OTD dataset. Extensive experiments demonstrate that the proposed method can effectively improve the accuracy of detecting oil tanks in SAR images. Datasets, codes, and more results will be released at: https://EIS-VIPG.github.io/SpaceNet6-OTD/. Ruixiang Zhang, Haowen Guo, Wen Yang 0001, Huai Yu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Hierarchical Prediction and Adversarial Learning For Conditional Response GenerationabstractThere are a variety of underlying factors influencing what and how people communicate in their daily life. The ability to capture and utilize these factors enables the conversational systems to generate favorable responses and set up amicable connections with users. In this work, we investigate two major factors in response generation, i.e., emotion and intention. To explore the dependency between them, we develop a hierarchical variational model that predicts in sequence the emotion and intention to be conveyed in a response. The response can then be generated word-by-word based on the predictions. We also apply a novel adversarial-augmented inference network to facilitate model training. The experimental results demonstrate the effectiveness of the proposed model as well as the novel adversarial objective. The hypothesis that emotion shapes human communication behavior is also validated. Yanran Li, Ruixiang Zhang, Wenjie Li 0002, Ziqiang Cao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative ModelsabstractAI Safety is a major concern in many deep learning applications such as autonomous driving. Given a trained deep learning model, an important natural problem is how to reliably verify the model's prediction. In this paper, we propose a novel framework --- deep verifier networks (DVN) to detect unreliable inputs or predictions of deep discriminative models, using separately trained deep generative models. Our proposed model is based on conditional variational auto-encoders with disentanglement constraints to separate the label information from the latent representation. We give both intuitive and theoretical justifications for the model. Our verifier network is trained independently with the prediction model, which eliminates the need of retraining the verifier network for a new model. We test the verifier network on both out-of-distribution detection and adversarial example detection problems, as well as anomaly detection problems in structured prediction tasks such as image caption generation. We achieve state-of-the-art results in all of these problems. Tong Che, Xiaofeng Liu 0001, Site Li, Yubin Ge, Ruixiang Zhang, Caiming Xiong, Yoshua Bengio |
AAAI | 5 |
| 2021 | RPVNet: A Deep and Efficient Range-Point-Voxel Fusion Network for LiDAR Point Cloud SegmentationabstractPoint clouds can be represented in many forms (views), typically, point-based sets, voxel-based cells or range-based images(i.e., panoramic view). The point-based view is geometrically accurate, but it is disordered, which makes it difficult to find local neighbors efficiently. The voxel-based view is regular, but sparse, and computation grows cubicly when voxel resolution increases. The range-based view is regular and generally dense, however spherical projection makes physical dimensions distorted. Both voxel-and range-based views suffer from quantization loss, especially for voxels when facing large-scale scenes. In order to utilize different view’s advantages and alleviate their own shortcomings in fine-grained segmentation task, we propose a novel range-point-voxel fusion network, namely RPVNet. In this network, we devise a deep fusion framework with multiple and mutual information interactions among these three views, and propose a gated fusion module (termed as GFM), which can adaptively merge the three features based on concurrent inputs. Moreover, the proposed RPV interaction mechanism is highly efficient, and we summarize it to a more general formulation. By leveraging this efficient interaction and relatively lower voxel resolution, our method is also proved to be more efficient. Finally, we evaluated the proposed model on two large-scale datasets, i.e., SemanticKITTI and nuScenes, and it shows state-of-the-art performance on both of them. Note that, our method currently ranks 1st on SemanticKITTI leaderboard without any extra tricks. Jianyun Xu, Ruixiang Zhang, Jian Dou, Yushi Zhu, Shiliang Pu |
ICCV | 2 |
| 2020 | Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic PerspectiveabstractLearning controllable and generalizable representation of multivariate data with desired structural properties remains a fundamental problem in machine learning. In this paper, we present a novel framework for learning generative models with various underlying structures in the latent space. Learning controllable and generalizable representation of multivariate data with desired structural properties remains a fundamental problem in machine learning. In this paper, we present a novel framework for learning generative models with various underlying structures in the latent space. We represent the inductive bias in the form of mask variables to model the dependency structure in the graphical model and extend the theory of multivariate information bottleneck (Friedman et al., 2001) to enforce it. Our model provides a principled approach to learn a set of semantically meaningful latent factors that reflect various types of desired structures like capturing correlation or encoding invariance, while also offering the flexibility to automatically estimate the dependency structure from data. We show that our framework unifies many existing generative models and can be applied to a variety of tasks, including multimodal data modeling, algorithmic fairness, and out-of-distribution generalization. Ruixiang Zhang, Masanori Koyama, Katsuhiko Ishiguro |
ICML | 1 |
| 2020 | Perceptual Generative AutoencodersabstractModern generative models are usually designed to match target distributions directly in the data space, where the intrinsic dimension of data can be much lower than the ambient dimension. We argue that this discrepancy may contribute to the difficulties in training generative models. We therefore propose to map both the generated and target distributions to the latent space using the encoder of a standard autoencoder, and train the generator (or decoder) to match the target distribution in the latent space. Specifically, we enforce the consistency in both the data space and the latent space with theoretically justified data and latent reconstruction losses. The resulting generative model, which we call a perceptual generative autoencoder (PGA), is then trained with a maximum likelihood or variational autoencoder (VAE) objective. With maximum likelihood, PGAs generalize the idea of reversible generative models to unrestricted neural network architectures and arbitrary number of latent dimensions. When combined with VAEs, PGAs substantially improve over the baseline VAEs in terms of sample quality. Compared to other autoencoder-based generative models using simple priors, PGAs achieve state-of-the-art FID scores on CIFAR-10 and CelebA. Ruixiang Zhang, Zongpeng Li, Yoshua Bengio, Liam Paull |
ICML | 2 |
| 2020 | Tiny Object Detection in Aerial ImagesabstractObject detection in Earth Vision has achieved great progress in recent years. However, tiny object detection in aerial images remains a very challenging problem since the tiny objects contain a small number of pixels and are easily confused with the background. To advance tiny object detection research in aerial images, we present a new dataset for Tiny Object Detection in Aerial Images (AI-TOD). Specifically, AI-TOD comes with 700,621 object instances for eight categories across 28,036 aerial images. Compared to existing object detection datasets in aerial images, the mean size of objects in AI-TOD is about 12.8 pixels, which is much smaller than others. To build a benchmark for tiny object detection in aerial images, we evaluate the state-of-the-art object detectors on our AI-TOD dataset. Experimental results show that direct application of these approaches on AI-TOD produces suboptimal object detection results, thus new specialized detectors for tiny object detection need to be designed. Therefore, we propose a multiple center points based learning network (M-CenterNet) to improve the localization performance of tiny object detection, and experimental results show the significant performance gain over the competitors. Jinwang Wang, Wen Yang 0001, Haowen Guo, Ruixiang Zhang, Gui-Song Xia |
ICPR | 4 |
| 2020 | Edge-Driven Object Matching for UAV Images and Satellite SAR ImagesabstractThe task of matching between images acquired by terminal equipment and satellites is important and challenging due to the dramatic viewpoint changes and unknown orientations, especially with different imaging sensors. In this paper, we firstly present the task to match the optical/infrared images acquired by UAVs with satellite SAR images. Many previous works mainly focused on matching with the images of the same modality, and may not perform well to our task. To overcome the difficulties caused by the diversity among the three modalities of data, we mine the common features of them and propose a novel edge-driven matching framework to find the correspondence between the UAV images and SAR images. Experimental results demonstrate the effectiveness and superiority of our method. Ruixiang Zhang, Huai Yu, Wen Yang 0001, Heng-Chao Li 0001 |
IGARSS | 1 |
| 2020 | Your GAN is Secretly an Energy-based Model and You Should Use Discriminator Driven Latent SamplingabstractWe show that the sum of the implicit generator log-density $\log p_g$ of a GAN with the logit score of the discriminator defines an energy function which yields the true data density when the generator is imperfect but the discriminator is optimal, thus making it possible to improve on the typical generator (with implicit density $p_g$). To make that practical, we show that sampling from this modified density can be achieved by sampling in latent space according to an energy-based model induced by the sum of the latent prior log-density and the discriminator output score. This can be achieved by running a Langevin MCMC in latent space and then applying the generator function, which we call Discriminator Driven Latent Sampling~(DDLS). We show that DDLS is highly efficient compared to previous methods which work in the high-dimensional pixel space and can be applied to improve on previously trained GANs of many types. We evaluate DDLS on both synthetic and real-world datasets qualitatively and quantitatively. On CIFAR-10, DDLS substantially improves the Inception Score of an off-the-shelf pre-trained SN-GAN~\citep{sngan} from $8.22$ to $9.09$ which is even comparable to the class-conditional BigGAN~\citep{biggan} model. This achieves a new state-of-the-art in unconditional image synthesis setting without introducing extra parameters or additional training. Tong Che, Ruixiang Zhang, Jascha Sohl-Dickstein, Hugo Larochelle, Liam Paull, Yoshua Bengio |
NeurIPS | 2 |
| 2019 | Mental Retrieval of Large-Scale Satellite Images Via Learned Sketch-Image Deep FeaturesabstractSearching targets of interest in large-scale satellite images is an imperative task, which becomes a challenging issue when the targets reside only in the mind of the user as a set of subjective visual patterns. In this paper, we take the advantage of hand-drawn sketches' strong intuition of describing mental target to address the problem of no available exemplar query. We introduce a multi-level-of-detail model to learn a cross-domain representation for bridging the gap between sketches and satellite images. To train the model, we propose a novel method of generating satellite images with corresponding level of details based on generative adversarial network. Experiments on both large-scale satellite images and commonly used RS datasets demonstrate the effectiveness and superiority of our method. Ruixiang Zhang, Wen Yang 0001, Gui-Song Xia |
IGARSS | 2 |
| 2018 | MetaGAN: An Adversarial Approach to Few-Shot LearningabstractIn this paper, we propose a conceptually simple and general framework called MetaGAN for few-shot learning problems. Most state-of-the-art few-shot classification models can be integrated with MetaGAN in a principled and straightforward way. By introducing an adversarial generator conditioned on tasks, we augment vanilla few-shot classification models with the ability to discriminate between real and fake data. We argue that this GAN-based approach can help few-shot classifiers to learn sharper decision boundary, which could generalize better. We show that with our MetaGAN framework, we can extend supervised few-shot learning models to naturally cope with unsupervised data. Different from previous work in semi-supervised few-shot learning, our algorithms can deal with semi-supervision at both sample-level and task-level. We give theoretical justifications of the strength of MetaGAN, and validate the effectiveness of MetaGAN on challenging few-shot image classification benchmarks. Ruixiang Zhang, Tong Che, Zoubin Ghahramani, Yoshua Bengio, Yangqiu Song |
NeurIPS | 1 |
| 2011 | On the Number of Ordinary Circles Determined by n Points
Ruixiang Zhang |
Discret. Comput. Geom. | 1 |