Sicheng Zhu

dblp:219/8361 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 7 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Domain-Auxiliary Infrared Moving Small Target Detection by Learning to Overlook Domain Discrepancy
abstract
Currently, almost all traditional infrared small target detection methods work on the assumption that training and test sets always belong to the same domain, and training samples are sufficient. However, in real applications, a new detection task could often have no sufficient training samples from a special domain. In this situation, adopting the auxiliary data from big-sample domains is usually believed to be one of the most potential solutions. However, exceeding expectations, it is found that simply adding auxiliary samples cannot often be always effective, even causing performance decline, due to existing infrared domain shift. To overcome this unexpected problem, we propose the first infrared moving small target detection framework with domain-auxiliary supports by Learning to Overlook Domain Discrepancy (Loddis). This framework consists of three primary processing stages: correlation weakening, domain confusing, and target consistency contrastive learning. Breaking through traditional learning paradigm, through auxiliary data, it enables the model to focus more on targets themselves, and less on image backgrounds, minimizing the sensitivity to domain discrepancy. The extensive experiments on 6 different-domain datasets show the effectiveness and superiority of the proposed Loddis framework for infrared small target detection.
Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001
AAAI4
2026 SeViL: Semi-supervised Vision-Language Learning with Text Prompt Guiding for Moving Infrared Small Target Detection
abstract
Unlike traditional object detection, moving infrared small target detection is highly challenging due to tiny target size and limited labeled samples. Currently, most existing methods mainly focus on the pure-vision features usually by fully-supervised learning, heavily relying on extensive high-cost manual annotations. Moreover, they almost have not concerned the potentials of multi-modal (e.g., vision and text) learning yet. To address these issues, inspired by prevalent vision-language models, we propose the first semi-supervised vision-language (SeViL) framework with adaptive text prompt guiding. Breaking through traditional pure-vision modality, it takes text prompts as prior knowledge to adaptively enhance target regions and then filter the low-quality pseudo-labels generated on unlabeled data. In the meanwhile, we employ an adaptive cross-modal masking strategy to align text and vision features, promoting cross-modal deep interactions. Remarkably, our extensive experiments on three public datasets (DAUB, ITSDT-15K and IRDST) verify that our new scheme could outperform other semi-supervised ones, and even achieve comparable performance to fully-supervised state-of-the-art (SOTA) methods, with only 10% labeled training samples.
Luping Ji, Jianghong Huang, Sicheng Zhu
AAAI4
2026 Cross-domain Joint Learning with Prototype-guided Mixture-of-Experts for Infrared Moving Small Target Detection
abstract
Infrared small target detection often faces significant domain gaps across datasets due to varying sensors and scene distributions. Currently, most existing methods are typically based on single-domain learning (i.e., training and test are on the same dataset), requiring training separate detectors when considering different datasets. However, they overlook the valuable public knowledge across domains and limit the applicability in multiple infrared scenarios. To break through single-domain learning, implementing only one universal detector simultaneously on multiple datasets, as the first exploration, we propose a cross-domain joint learning task framework with prototype-guided Mixture-of-Experts (CoMoE). Specifically, it designs a hyperspherical prototype learning to adaptively maintain both domain-specific prototypes and global prototypes, enhancing cross-domain feature representation. Meanwhile, a domain-aware Mixture-of-Experts with Top-K routing strategy is proposed to select the optimal domain experts. Moreover, to enhance cross-domain feature alignment, we design an adaptive cross-domain feature modulation with noise-guided contrastive learning. The extensive experiments on a newly constructed benchmark comprising three datasets verify the superiority of our CoMoE, even under limited data settings. It could often surpass general joint learning methods, and state-of-the-art (SOTA) single-domain ones.
Luping Ji, Jianghong Huang, Sicheng Zhu, Mao Ye 0001
AAAI4
2025 Can Watermarking Large Language Models Prevent Copyrighted Text Generation and Hide Training Data?
abstract
Large Language Models (LLMs) have demonstrated impressive capabilities in generating diverse and contextually rich text. However, concerns regarding copyright infringement arise as LLMs may inadvertently produce copyrighted material. In this paper, we first investigate the effectiveness of watermarking LLMs as a deterrent against the generation of copyrighted texts. Through theoretical analysis and empirical evaluation, we demonstrate that incorporating watermarks into LLMs significantly reduces the likelihood of generating copyrighted content, thereby addressing a critical concern in the deployment of LLMs. However, we also find that watermarking can have unintended consequences on Membership Inference Attacks (MIAs), which aim to discern whether a sample was part of the pretraining dataset and may be used to detect copyright violations. Surprisingly, we find that watermarking adversely affects the success rate of MIAs, complicating the task of detecting copyrighted text in the pretraining dataset. These results reveal the complex interplay between different regulatory measures, which may impact each other in unforeseen ways. Finally, we propose an adaptive technique to improve the success rate of a recent MIA under watermarking. Our findings underscore the importance of developing adaptive methods to study critical problems in LLMs with potential legal implications.
Michael-Andrei Panaitescu-Liess, Zora Che, Bang An 0001, Yuancheng Xu, Pankayaraj Pathmanathan, Souradip Chakraborty, Sicheng Zhu, Tom Goldstein, Furong Huang
AAAI7
2025 GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-Time Alignment
abstract
Large Language Models (LLMs) exhibit impressive capabilities but require careful alignment with human preferences. Traditional training-time methods finetune LLMs using human preference datasets but incur significant training costs and require repeated training to handle diverse user preferences. Test-time alignment methods address this by using reward models (RMs) to guide frozen LLMs without retraining. However, existing test-time approaches rely on trajectory-level RMs which are designed to evaluate complete responses, making them unsuitable for autoregressive text generation that requires computing next-token rewards from partial responses. To address this, we introduce GenARM, a test-time alignment approach that leverages the Autoregressive Reward Model—a novel reward parametrization designed to predict next-token rewards for efficient and effective autoregressive generation. Theoretically, we demonstrate that this parametrization can provably guide frozen LLMs toward any distribution achievable by traditional RMs within the KL-regularized reinforcement learning framework. Experimental results show that GenARM significantly outperforms prior test-time alignment baselines and matches the performance of training-time methods. Additionally, GenARM enables efficient weak-to-strong guidance, aligning larger LLMs with smaller RMs without the high costs of training larger models. Furthermore, GenARM supports multi-objective alignment, allowing real-time trade-offs between preference dimensions and catering to diverse user preferences without retraining. Our project page is available at: https://genarm.github.io.
Yuancheng Xu, T. W. U. Madhushani, Alec Koppel, Sicheng Zhu, Bang An 0001, Furong Huang, Sumitra Ganesh
ICLR4
2025 PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
abstract
Michael-Andrei Panaitescu-Liess, Pankayaraj Pathmanathan, Yigitcan Kaya, Zora Che, Bang An, Sicheng Zhu, Aakriti Agrawal, Furong Huang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Michael-Andrei Panaitescu-Liess, Pankayaraj Pathmanathan, Yigitcan Kaya, Zora Che, Bang An 0001, Sicheng Zhu, Aakriti Agrawal, Furong Huang
NAACL (Long Papers)6
2025 AdvPrefix: An Objective for Nuanced LLM Jailbreaks
abstract
Many jailbreak attacks on large language models (LLMs) rely on a common objective: making the model respond with the prefix ``Sure, here is (harmful request)''. While straightforward, this objective has two limitations: limited control over model behaviors, yielding incomplete or unrealistic jailbroken responses, and a rigid format that hinders optimization. We introduce AdvPrefix, a plug-and-play prefix-forcing objective that selects one or more model-dependent prefixes by combining two criteria: high prefilling attack success rates and low negative log-likelihood. AdvPrefix integrates seamlessly into existing jailbreak attacks to mitigate the previous limitations for free. For example, replacing GCG's default prefixes on Llama-3 improves nuanced attack success rates from 14\% to 80\%, revealing that current safety alignment fails to generalize to new prefixes. Code and selected prefixes are released.
Sicheng Zhu, Brandon Amos, Yuandong Tian, Chuan Guo 0001, Ivan Evtimov
NeurIPS1
2025 Moving infrared dim and small target detection by mixed spatio-temporal encoding
Luping Ji, Shengjia Chen, Sicheng Zhu
Eng. Appl. Artif. Intell.5
2025 Spatial-temporal-channel collaborative feature learning with transformers for infrared small target detection
Sicheng Zhu, Luping Ji, Shengjia Chen
Image Vis. Comput.1
2025 Language-Driven Motion Prior Knowledge Learning for Moving Infrared Small Target Detection
abstract
Different from traditional object detection, pure vision is often not enough to infrared small target detection (ISTD), due to the small target size and weak background contrast. For promoting performance, more target representations are often needed. Currently, motion representations have been proved to be one of the most potential feature patterns for infrared small targets. Besides vision features, existing methods have an obvious weakness that they could only capture coarse motion representations from the temporal domain. By vision features, fine motion representations could often be more effective to enhance detection performance. To overcome this weakness, and inspired by prevalent vision-language models (VLMs), the first vision-language framework with motion prior knowledge learning (MoPKL) was proposed in our previous work. To further extend it, we repropose an improved version, i.e., iMoPKL. Breaking through traditional pure-vision modality, it utilizes the homogeneous language descriptions, specially formatted for moving targets, to directionally guide vision channels to learn the motion prior knowledge of targets. In detail, it learns the distribution of target motion reconstruction corresponding to the language description as a type of prior knowledge. With the facilitation of language-driven motion alignment, the motion of infrared small targets could be further refined by motion-relation learning, to generate more fine motion representations. The extensive experiments on ITSDT-15K, DAUB-R, and IRDST-H show that our improvement version is effective. It could often obviously outperform the other methods, including our original MoPKL. Our source codes are available athttps://github.com/UESTC-nnLab/MoPKL
Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001, Yongsheng Sang
IEEE Trans. Geosci. Remote. Sens.4
2025 Weakly Supervised Contrastive Learning With Quantity Prompts for Moving Infrared Small Target Detection
Luping Ji, Shengjia Chen, Sicheng Zhu, Jianghong Huang, Mao Ye 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Semi-Supervised Multiview Prototype Learning With Motion Reconstruction for Moving Infrared Small Target Detection
abstract
Moving infrared small target detection is critical for various applications, e.g., remote sensing and military. Due to tiny target size and limited labeled data, accurately detecting targets is highly challenging. Currently, existing methods primarily focus on fully-supervised learning, which relies heavily on numerous annotated frames for training. However, annotating a large number of frames for each video is often expensive, time-consuming, and redundant, especially for low-quality infrared images. To break through traditional fully-supervised framework, we propose a new semi-supervised multi-view prototype (S2MVP) learning scheme that incorporates motion reconstruction. In our scheme, we design a bi-temporal motion perceptor based on bidirectional ConvGRU cells to effectively model the motion paradigms of targets by perceiving both forward and backward. Additionally, to explore the potential of unlabeled data, it generates the multi-view feature prototypes of targets as soft labels to guide feature learning by calculating cosine similarity. Imitating human visual system, it retains only the feature prototypes of recent frames. Moreover, it eliminates noisy pseudo-labels to enhance the quality of pseudo-labels through anomaly-driven pseudo-label filtering. Furthermore, we develop a target-aware motion reconstruction loss to provide additional supervision and prevent the loss of target details. To our best knowledge, the proposed S2MVP is the first work to utilize large-scale unlabeled video frames to detect moving infrared small targets. Although 10% labeled training samples are used, the experiments on three public benchmarks (DAUB, ITSDT-15K and IRDST) verify the superiority of our scheme compared to other methods. Source codes are available at https://github.com/UESTC-nnLab/S2MVP.
Luping Ji, Jianghong Huang, Shengjia Chen, Sicheng Zhu, Mao Ye 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 MICPL: Motion-Inspired Cross-Pattern Learning for Small-Object Detection in Satellite Videos
abstract
For small-object detection, vision patterns can only provide limited support to feature learning. Most prior schemes mainly depend on a single vision pattern to learn object features, seldom considering more latent motion patterns. In the real world, humans often efficiently perceive small objects through multipattern signals. Inspired by this observation, this article attempts to address small-object detection from a new prospective of latent pattern learning. To fulfill this purpose, it regards a real-world moving object as the spatiotemporal sequences of a static object to capture latent motion patterns. In view of this, we propose a motion-inspired cross-pattern learning (MICPL) scheme to capture the motion patterns for moving small-object scenarios. This scheme mainly consists of two crucial parts: motion pattern mining (MPM) and motion-vision adaption. The former is designed to effectively mine the motion pattern from time-dependent representation space. The latter is devised to correlate between motion patterns and vision semantics. In the meanwhile, we explore their cross-pattern interactions to guide MICPL to capture motion patterns effectively. Comparison experiments verify that, cooperated by motion pattern, even a simple detector could often refresh state-of-the-art (SOTA) results on moving small-object detection. Moreover, the experiments on two small-object-related tasks further prove the adaptivity and advantages of our cross-pattern feature learning scheme. Our source codes are available at https://github.com/UESTC-nnLab/MICPL.
Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
abstract
Vision-language models like CLIP are widely used in zero-shot image classification due to their ability to understand various visual concepts and natural language descriptions. However, how to fully leverage CLIP's unprecedented human-like understanding capabilities to achieve better performance is still an open question. This paper draws inspiration from the human visual perception process: when classifying an object, humans first infer contextual attributes (e.g., background and orientation) which help separate the foreground object from the background, and then classify the object based on this information. Inspired by it, we observe that providing CLIP with contextual attributes improves zero-shot image classification and mitigates reliance on spurious features. We also observe that CLIP itself can reasonably infer the attributes from an image. With these observations, we propose a training-free, two-step zero-shot classification method PerceptionCLIP. Given an image, it first infers contextual attributes (e.g., background) and then performs object classification conditioning on them. Our experiments show that PerceptionCLIP achieves better generalization, group robustness, and interpretability.
Bang An 0001, Sicheng Zhu, Michael-Andrei Panaitescu-Liess, Chaithanya Kumar Mummadi, Furong Huang
ICLR2
2024 Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds
abstract
Group robustness has become a major concern in machine learning (ML) as conventional training paradigms were found to produce high error on minority groups. Without explicit group annotations, proposed solutions rely on heuristics that aim to identify and then amplify the minority samples during training. In our work, we first uncover a critical shortcoming of these methods: an inability to distinguish legitimate minority samples from poison samples in the training set. By amplifying poison samples as well, group robustness methods inadvertently boost the success rate of an adversary---e.g., from 0\% without amplification to over 97\% with it. Notably, we supplement our empirical evidence with an impossibility result proving this inability of a standard heuristic under some assumptions. Moreover, scrutinizing recent poisoning defenses both in centralized and federated learning, we observe that they rely on similar heuristics to identify which samples should be eliminated as poisons. In consequence, minority samples are eliminated along with poisons, which damages group robustness---e.g., from 55\% without the removal of the minority samples to 41\% with it. Finally, as they pursue opposing goals using similar heuristics, our attempt to alleviate the trade-off by combining group robustness methods and poisoning defenses falls short. By exposing this tension, we also hope to highlight how benchmark-driven ML scholarship can obscure the trade-offs among different metrics with potentially detrimental consequences.
Michael-Andrei Panaitescu-Liess, Yigitcan Kaya, Sicheng Zhu, Furong Huang, Tudor Dumitras
ICLR3
2024 WAVES: Benchmarking the Robustness of Image Watermarks
abstract
In the burgeoning age of generative AI, watermarks act as identifiers of provenance and artificial content. We present WAVES (Watermark Analysis via Enhanced Stress-testing), a benchmark for assessing image watermark robustness, overcoming the limitations of current evaluation methods. WAVES integrates detection and identification tasks and establishes a standardized evaluation protocol comprised of a diverse range of stress tests. The attacks in WAVES range from traditional image distortions to advanced, novel variations of diffusive, and adversarial attacks. Our evaluation examines two pivotal dimensions: the degree of image quality degradation and the efficacy of watermark detection after attacks. Our novel, comprehensive evaluation reveals previously undetected vulnerabilities of several modern watermarking algorithms. We envision WAVES as a toolkit for the future development of robust watermarks.
Bang An 0001, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal, Yuancheng Xu, Chenghao Deng, Sicheng Zhu, Abdirisak Mohamed, Yuxin Wen, Tom Goldstein, Furong Huang
ICML7
2024 Position: On the Possibilities of AI-Generated Text Detection
abstract
Our study addresses the challenge of distinguishing human-written text from Large Language Model (LLM) outputs. We provide evidence that this differentiation is consistently feasible, except when human and machine text distributions are indistinguishable across their entire support. Employing information theory, we show that while detecting machine-generated text becomes harder as it nears human quality, it remains possible with adequate text data. We introduce guidelines on the required text data quantity, either through sample size or sequence length, for reliable AI text detection, through derivations of sample complexity bounds. This research paves the way for advanced detection methods. Our comprehensive empirical tests, conducted across various datasets (Xsum, Squad, IMDb, and Kaggle FakeNews) and with several state-of-the-art text generators (GPT-2, GPT-3.5-Turbo, Llama, Llama-2-13B-Chat-HF, Llama-2-70B-Chat-HF), assess the viability of enhanced detection methods against detectors like RoBERTa-Large/Base-Detector and GPTZero, with increasing sample sizes and sequence lengths. Our findings align with OpenAI’s empirical data related to sequence length, marking the first theoretical substantiation for these observations.
Souradip Chakraborty, Amrit Singh Bedi, Sicheng Zhu, Bang An 0001, Dinesh Manocha, Furong Huang
ICML3
2024 Spatio-temporal fusion with motion masks for the moving small target detection from remote-sensing videos
Sicheng Zhu, Luping Ji, Jiewen Zhu, Shengjia Chen, Haohao Ren
Eng. Appl. Artif. Intell.1
2024 TMP: Temporal Motion Perception with spatial auxiliary enhancement for moving Infrared dim-small target detection
Sicheng Zhu, Luping Ji, Jiewen Zhu, Shengjia Chen
Expert Syst. Appl.1
2024 Toward Dense Moving Infrared Small Target Detection: New Datasets and Baseline
abstract
As an important research branch of infrared small target detection, dense target detection (e.g., drone swarm detection) has always been a topic worth exploring. Currently, existing datasets cover only one or several (sparse) targets, with almost no dataset available for the research on dense small target detection. To advance this kind of search, for the first time, we synthesize two special dense moving target datasets (DMIST-60 and DMIST-100) on DAUB data. They both contain far more than 50 infrared small targets per frame. In the meantime, for evaluating our new datasets and flourishing detection methodology research, we propose a linking-aware sliced network (LASNet) as the baseline of our datasets. It mainly consists of visual feature extraction, motion feature extraction and motion-affinity fusion. The comprehensive experiments on our synthesized datasets confirm: i) both datasets are practical and effective for dense moving infrared small target detection and ii) proposed LASNet could always obviously outperform other compared methods in both sparse and dense target scenarios. Our new datasets and source codes are currently available athttps://github.com/UESTC-nnLab/DMIST.
Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001, Haohao Ren, Yongsheng Sang
IEEE Trans. Geosci. Remote. Sens.3
2024 Triple-Domain Feature Learning With Frequency-Aware Memory Enhancement for Moving Infrared Small Target Detection
abstract
As a subfield of object detection, moving infrared small target detection (ISTD) presents significant challenges due to tiny target sizes and low contrast against backgrounds. Currently existing methods primarily rely on the features extracted only from spatiotemporal domain. Frequency domain has hardly been concerned yet, although it has been widely applied in image processing. To extend feature source domains and enhance feature representation, we propose a new triple-domain strategy (Tridos) with the frequency-aware memory enhancement on spatiotemporal domain for ISTD. In this scheme, it effectively detaches and enhances frequency features by a local-global frequency-aware module (LGFM) with Fourier transform (FT). Inspired by human visual system (HVS), our memory enhancement is designed to capture the spatial relationships of infrared targets among video frames. Furthermore, it encodes temporal dynamics motion features via differential learning and residual enhancing. In addition, we further design a residual compensation to reconcile possible cross-domain feature mismatches. To our best knowledge, proposed Tridos is the first work to explore infrared target feature learning comprehensively in spatiotemporal-frequency domains. The extensive experiments on three datasets (i.e., DAUB, ITSDT-15K, and IRDST) validate that our triple-domain infrared feature learning scheme could often be obviously superior to state-of-the-art (SOTA) ones. Source codes are available athttps://github.com/UESTC-nnLab/Tridos.
Luping Ji, Shengjia Chen, Sicheng Zhu, Mao Ye 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Learning Unforeseen Robustness from Out-of-distribution Data Using Equivariant Domain Translator
abstract
Current approaches for training robust models are typically tailored to scenarios where data variations are accessible in the training set. While shown effective in achieving robustness to these foreseen variations, these approaches are ineffective in learning unforeseen robustness, i.e., robustness to data variations without known characterization or training examples reflecting them. In this work, we learn unforeseen robustness by harnessing the variations in the abundant out-of-distribution data. To overcome the main challenge of using such data, the domain gap, we use a domain translator to bridge it and bound the unforeseen robustness on the target distribution. As implied by our analysis, we propose a two-step algorithm that first trains an equivariant domain translator to map out-of-distribution data to the target distribution while preserving the considered variation, and then regularizes a model’s output consistency on the domain-translated data to improve its robustness. We empirically show the effectiveness of our approach in improving unforeseen and foreseen robustness compared to existing approaches. Additionally, we show that training the equivariant domain translator serves as an effective criterion for source data selection.
Sicheng Zhu, Bang An 0001, Furong Huang, Sanghyun Hong 0001
ICML1
2021 Understanding the Generalization Benefit of Model Invariance from a Data Perspective
abstract
Machine learning models that are developed to be invariant under certain types of data transformations have shown improved generalization in practice. However, a principled understanding of why invariance benefits generalization is limited. Given a dataset, there is often no principled way to select "suitable" data transformations under which model invariance guarantees better generalization. This paper studies the generalization benefit of model invariance by introducing the sample cover induced by transformations, i.e., a representative subset of a dataset that can approximately recover the whole dataset using transformations. For any data transformations, we provide refined generalization bounds for invariant models based on the sample cover. We also characterize the "suitability" of a set of data transformations by the sample covering number induced by transformations, i.e., the smallest size of its induced sample covers. We show that we may tighten the generalization bounds for "suitable" transformations that have a small sample covering number. In addition, our proposed sample covering number can be empirically evaluated and thus provides a guidance for selecting transformations to develop model invariance for better generalization. In experiments on multiple datasets, we evaluate sample covering numbers for some commonly used transformations and show that the smaller sample covering number for a set of transformations (e.g., the 3D-view transformation) indicates a smaller gap between the test and training error for invariant models, which verifies our propositions.
Sicheng Zhu, Bang An 0001, Furong Huang
NeurIPS1
2020 Learning Adversarially Robust Representations via Worst-Case Mutual Information Maximization
abstract
Training machine learning models that are robust against adversarial inputs poses seemingly insurmountable challenges. To better understand adversarial robustness, we consider the underlying problem of learning robust representations. We develop a notion of representation vulnerability that captures the maximum change of mutual information between the input and output distributions, under the worst-case input perturbation. Then, we prove a theorem that establishes a lower bound on the minimum adversarial risk that can be achieved for any downstream classifier based on its representation vulnerability. We propose an unsupervised learning method for obtaining intrinsically robust representations by maximizing the worst-case mutual information between the input and output distributions. Experiments on downstream classification tasks support the robustness of the representations found using unsupervised learning with our training principle.
Sicheng Zhu, Xiao Zhang 0016, David Evans 0001
ICML1