EDBT 2026 Demo / reviewers in the wild / expert
Chen Huang 0001
dblp:05/8125-1
· DBLP profile ↗
33ranked-venue papers
17as first author
14since 2021 · last 2025
0000-0002-4978-493XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 15 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 8 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without ForgettingabstractVision foundation models pre-trained on massive data encode rich representations of real-world concepts, which can be adapted to downstream tasks by fine-tuning. However, fine-tuning foundation models on one task often leads to the issue of *concept forgetting* on other tasks. Recent methods of robust fine-tuning aim to mitigate forgetting of prior knowledge without affecting the fine-tuning performance. Knowledge is often preserved by matching the original and fine-tuned model weights or feature pairs. However, such point-wise matching can be too strong, without explicit awareness of the feature neighborhood structures that encode rich knowledge as well. We propose a novel regularization method **Proxy-FDA** that explicitly preserves the structural knowledge in feature space. Proxy-FDA performs Feature Distribution Alignment (using nearest neighbor graphs) between the pre-trained and fine-tuned feature spaces, and the alignment is further improved by informative proxies that are generated dynamically to increase data diversity. Experiments show that Proxy-FDA significantly reduces concept forgetting during fine-tuning, and we find a strong correlation between forgetting and a distributional distance metric (in comparison to L2 distance). We further demonstrate Proxy-FDA's benefits in various fine-tuning settings (end-to-end, few-shot and continual tuning) and across different tasks like image classification, captioning and VQA. Chen Huang 0001, Skyler Seto, Hadi Pouransari, Mehrdad Farajtabar, Raviteja Vemulapalli, Fartash Faghri, Oncel Tuzel, Barry-John Theobald, Joshua M. Susskind |
ICML | 1 |
| 2024 | LiDAR: Sensing Linear Probing Performance in Joint Embedding SSL ArchitecturesabstractJoint embedding (JE) architectures have emerged as a promising avenue for ac-
quiring transferable data representations. A key obstacle to using JE methods,
however, is the inherent challenge of evaluating learned representations without
access to a downstream task, and an annotated dataset. Without efficient and re-
liable evaluation, it is difficult to iterate on architectural and training choices for
JE methods. In this paper, we introduce LiDAR (Linear Discriminant Analysis
Rank), a metric designed to measure the quality of representations within JE archi-
tectures. Our metric addresses several shortcomings of recent approaches based
on feature covariance rank by discriminating between informative and uninforma-
tive features. In essence, LiDAR quantifies the rank of the Linear Discriminant
Analysis (LDA) matrix associated with the surrogate SSL task—a measure that
intuitively captures the information content as it pertains to solving the SSL task.
We empirically demonstrate that LiDAR significantly surpasses naive rank based
approaches in its predictive power of optimal hyperparameters. Our proposed cri-
terion presents a more robust and intuitive means of assessing the quality of rep-
resentations within JE architectures, which we hope facilitates broader adoption
of these powerful techniques in various domains. Vimal Thilak, Chen Huang 0001, Omid Saremi, Laurent Dinh, Hanlin Goh, Preetum Nakkiran, Joshua M. Susskind, Etai Littwin |
ICLR | 2 |
| 2024 | Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD GeneralizationabstractExisting vision-language models exhibit strong generalization on a variety of visual domains and tasks. However, such models mainly perform zero-shot recognition in a closed-set manner, and thus struggle to handle open-domain visual concepts by design. There are recent finetuning methods, such as prompt learning, that not only study the discrimination between in-distribution (ID) and out-of-distribution (OOD) samples, but also show some improvements in both ID and OOD accuracies. In this paper, we first demonstrate that vision-language models, after long enough finetuning but without proper regularization, tend to overfit the known classes in the given dataset, with degraded performance on unknown classes. Then we propose a novel approach OGEN to address this pitfall, with the main focus on improving the OOD GENeralization of finetuned models. Specifically, a class-conditional feature generator is introduced to synthesize OOD features using just the class name of any unknown class. Such synthesized features will provide useful knowledge about unknowns and help regularize the decision boundary between ID and OOD data when optimized jointly. Equally important is our adaptive self-distillation mechanism to regularize our feature generation model during joint optimization, i.e., adaptively transferring knowledge between model states to further prevent overfitting. Experiments validate that our method yields convincing gains in OOD generalization performance in different settings. Code: https://github.com/apple/ml-ogen. Yuhang Zang, Hanlin Goh, Joshua M. Susskind, Chen Huang 0001 |
ICLR | 4 |
| 2024 | Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIPabstractLarge pretrained vision-language models like CLIP have shown promising generalization capability, but may struggle in specialized domains (e.g., satellite imagery) or fine-grained classification (e.g., car models) where the visual concepts are unseen or under-represented during pretraining. Prompt learning offers a parameter-efficient finetuning framework that can adapt CLIP to downstream tasks even when limited annotation data are available. In this paper, we improve prompt learning by distilling the textual knowledge from natural language prompts (either human- or LLM-generated) to provide rich priors for those under-represented concepts. We first obtain a prompt ``summary'' aligned to each input image via a learned prompt aggregator. Then we jointly train a prompt generator, optimized to produce a prompt embedding that stays close to the aggregated summary while minimizing task loss at the same time. We dub such prompt embedding as Aggregate-and-Adapted Prompt Embedding (AAPE). AAPE is shown to be able to generalize to different downstream data distributions and tasks, including vision-language understanding tasks (e.g., few-shot classification, VQA) and generation tasks (image captioning) where AAPE achieves competitive performance. We also show AAPE is particularly helpful to handle non-canonical and OOD examples. Furthermore, AAPE learning eliminates LLM-based inference cost as required by baselines, and scales better with data and LLM model size. Chen Huang 0001, Skyler Seto, Samira Abnar, David Grangier, Navdeep Jaitly, Joshua M. Susskind |
NeurIPS | 1 |
| 2024 | How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation NetworksabstractTwo competing paradigms exist for self-supervised learning of data representations.
Joint Embedding Predictive Architectures (JEPAs) is a class of architectures in which semantically similar inputs are encoded into representations that are predictive of each other. A recent successful approach that falls under the JEPA framework is self-distillation, where an online encoder is trained to predict the output of the target encoder, sometimes with a lightweight predictor network. This is contrasted with the Masked Auto Encoder (MAE) paradigm, where an encoder and decoder are trained to reconstruct missing parts of the input in ambient space rather than its latent representation. A common motivation for using the JEPA approach over MAE is that the JEPA objective prioritizes abstract features over fine-grained pixel information (which can be unpredictable and uninformative).
In this work, we seek to understand the mechanism behind this empirical observation by analyzing deep linear models. We uncover a surprising mechanism: in a simplified linear setting where both approaches learn similar representations, JEPAs are biased to learn high influence features, or features characterized by having high regression coefficients. Our results point to a distinct implicit bias of predicting in latent space that may shed light on its success in practice. Etai Littwin, Omid Saremi, Madhu Advani, Vimal Thilak, Preetum Nakkiran, Chen Huang 0001, Joshua M. Susskind |
NeurIPS | 6 |
| 2023 | MAST: Masked Augmentation Subspace Training for Generalizable Self-Supervised Priors
Chen Huang 0001, Hanlin Goh, Jiatao Gu, Joshua M. Susskind |
ICLR | 1 |
| 2023 | DUET: 2D Structured and Approximately Equivariant RepresentationsabstractMultiview Self-Supervised Learning (MSSL) is based on learning invariances with respect to a set of input transformations. However, invariance partially or totally removes transformation-related information from the representations, which might harm performance for specific downstream tasks that require such information. We propose 2D strUctured and EquivarianT representations (coined DUET), which are 2d representations organized in a matrix structure, and equivariant with respect to transformations acting on the input data. DUET representations maintain information about an input transformation, while remaining semantically expressive. Compared to SimCLR (Chen et al., 2020) (unstructured and invariant) and ESSL (Dangovski et al., 2022) (unstructured and equivariant), the structured and equivariant nature of DUET representations enables controlled generation with lower reconstruction error, while controllability is not possible with SimCLR or ESSL. DUET also achieves higher accuracy for several discriminative tasks, and improves transfer learning. Xavier Suau, Federico Danieli, Arno Blaas, Chen Huang 0001, Jason Ramapuram, Dan Busbridge, Luca Zappella |
ICML | 5 |
| 2023 | Semi-Supervised and Long-Tailed Object Detection with CascadeMatch
Yuhang Zang, Kaiyang Zhou, Chen Huang 0001, Chen Change Loy |
Int. J. Comput. Vis. | 3 |
| 2022 | Open-Vocabulary DETR with Conditional Matching
Yuhang Zang, Wei Li 0319, Kaiyang Zhou, Chen Huang 0001, Chen Change Loy |
ECCV (9) | 4 |
| 2022 | Efficient Representation Learning via Adaptive Context PoolingabstractSelf-attention mechanisms model long-range context by using pairwise attention between all input tokens. In doing so, they assume a fixed attention granularity defined by the individual tokens (e.g., text characters or image pixels), which may not be optimal for modeling complex dependencies at higher levels. In this paper, we propose ContextPool to address this problem by adapting the attention granularity for each token. Inspired by the success of ConvNets that are combined with pooling to capture long-range dependencies, we learn to pool neighboring features for each token before computing attention in a given attention layer. The pooling weights and support size are adaptively determined, allowing the pooled features to encode meaningful context with varying scale. We show that ContextPool makes attention models more expressive, achieving strong performance often with fewer layers and thus significantly reduced cost. Experiments validate that our ContextPool module, when plugged into transformer models, matches or surpasses state-of-the-art performance using less compute on several language and image benchmarks, outperforms recent works with learned context sizes or sparse attention patterns, and is also applicable to ConvNets for efficient feature learning. Chen Huang 0001, Walter Talbott, Navdeep Jaitly, Joshua M. Susskind |
ICML | 1 |
| 2022 | Position Prediction as an Effective Pretraining StrategyabstractTransformers \cite{transformer} have gained increasing popularity in a wide range of applications, including Natural Language Processing (NLP), Computer Vision and Speech Recognition, because of their powerful representational capacity. However, harnessing this representational capacity effectively requires a large amount of data, strong regularization, or both, to mitigate overfitting. Recently, the power of the Transformer has been unlocked by self-supervised pretraining strategies based on masked autoencoders which rely on reconstructing masked inputs, directly, or contrastively from unmasked content. This pretraining strategy which has been used in BERT models in NLP \cite{bert}, Wav2Vec models in Speech \cite{wv2v2} and, recently, in MAE models in Vision \cite{beit, mae}, forces the model to learn about relationships between the content in different parts of the input using autoencoding related objectives. In this paper, we propose a novel, but surprisingly simple alternative to content reconstruction – that of predicting locations from content, without providing positional information for it. Doing so requires the Transformer to understand the positional relationships between different parts of the input, from their content alone. This amounts to an efficient implementation where the pretext task is a classification problem among all possible positions for each input token. We experiment on both Vision and Speech benchmarks, where our approach brings improvements over strong supervised training baselines and is comparable to modern unsupervised/self-supervised pretraining methods. Our method also enables Transformers trained without position embeddings to outperform ones trained with full position information. Shuangfei Zhai, Navdeep Jaitly, Jason Ramapuram, Dan Busbridge, Tatiana Likhomanenko, Joseph Y. Cheng, Walter Talbott, Chen Huang 0001, Hanlin Goh, Joshua M. Susskind |
ICML | 8 |
| 2021 | MetricOpt: Learning To Optimize Black-Box Evaluation MetricsabstractWe study the problem of directly optimizing arbitrary non-differentiable task evaluation metrics such as misclassification rate and recall. Our method, named MetricOpt, operates in a black-box setting where the computational details of the target metric are unknown. We achieve this by learning a differentiable value function, which maps compact task-specific model parameters to metric observations. The learned value function is easily pluggable into existing optimizers like SGD and Adam, and is effective for rapidly finetuning a pre-trained model. This leads to consistent improvements since the value function provides effective metric supervision during finetuning, and helps to correct the potential bias of loss-only supervision. MetricOpt achieves state-of-the-art performance on a variety of metrics for (image) classification, image retrieval and object detection. Solid benefits are found over competing methods, which often involve complex loss design or adaptation. MetricOpt also generalizes well to new tasks and model architectures. Chen Huang 0001, Shuangfei Zhai, Pengsheng Guo, Joshua M. Susskind |
CVPR | 1 |
| 2021 | FASA: Feature Augmentation and Sampling Adaptation for Long-Tailed Instance SegmentationabstractRecent methods for long-tailed instance segmentation still struggle on rare object classes with few training data. We propose a simple yet effective method, Feature Augmentation and Sampling Adaptation (FASA), that addresses the data scarcity issue by augmenting the feature space especially for rare classes. Both the Feature Augmentation (FA) and feature sampling components are adaptive to the actual training status — FA is informed by the feature mean and variance of observed real samples from past iterations, and we sample the generated virtual features in a loss-adapted manner to avoid over-fitting. FASA does not require any elaborate loss design, and removes the need for inter-class transfer learning that often involves large cost and manually-defined head/tail class groups. We show FASA is a fast, generic method that can be easily plugged into standard or long-tailed segmentation frameworks, with consistent performance gains and little added cost. FASA is also applicable to other tasks like long-tailed classification with state-of-the-art performance.12 Yuhang Zang, Chen Huang 0001, Chen Change Loy |
ICCV | 2 |
| 2021 | Mask-aware photorealistic facial attribute manipulationabstractThe technique of facial attribute manipulation has found increasing application, but it remains challenging to restrict editing of attributes so that a face’s unique details are preserved. In this paper, we introduce our method, which we call a mask-adversarial autoencoder (M-AAE). It combines a variational autoencoder (VAE) and a generative adversarial network (GAN) for photorealistic image generation. We use partial dilated layers to modify a few pixels in the feature maps of an encoder, changing the attribute strength continuously without hindering global information. Our training objectives for the VAE and GAN are reinforced by supervision of face recognition loss and cycle consistency loss, to faithfully preserve facial details. Moreover, we generate facial masks to enforce background consistency, which allows our training to focus on the foreground face rather than the background. Experimental results demonstrate that our method can generate high-quality images with varying attributes, and outperforms existing methods in detail preservation. Ruoqi Sun, Chen Huang 0001, Hengliang Zhu, Lizhuang Ma |
Comput. Vis. Media | 2 |
| 2020 | Deep Imbalanced Learning for Face Recognition and Attribute PredictionabstractData for face analysis often exhibit highly-skewed class distribution, i.e., most data belong to a few majority classes, while the minority classes only contain a scarce amount of instances. To mitigate this issue, contemporary deep learning methods typically follow classic strategies such as class re-sampling or cost-sensitive training. In this paper, we conduct extensive and systematic experiments to validate the effectiveness of these classic schemes for representation learning on class-imbalanced data. We further demonstrate that more discriminative deep representation can be learned by enforcing a deep network to maintain inter-cluster margins both within and between classes. This tight constraint effectively reduces the class imbalance inherent in the local data neighborhood, thus carving much more balanced class boundaries locally. We show that it is easy to deploy angular margins between the cluster distributions on a hypersphere manifold. Such learned Cluster-based Large Margin Local Embedding (CLMLE), when combined with a simple k-nearest cluster algorithm, shows significant improvements in accuracy over existing methods on both face recognition and face attribute prediction tasks that exhibit imbalanced class distribution. Chen Huang 0001, Chen Change Loy, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Dense Intrinsic Appearance Flow for Human Pose TransferabstractWe present a novel approach for the task of human pose transfer, which aims at synthesizing a new image of a person from an input image of that person and a target pose. We address the issues of limited correspondences identified between keypoints only and invisible pixels due to self-occlusion. Unlike existing methods, we propose to estimate dense and intrinsic 3D appearance flow to better guide the transfer of pixels between poses. In particular, we wish to generate the 3D flow from just the reference and target poses. Training a network for this purpose is non-trivial, especially when the annotations for 3D appearance flow are scarce by nature. We address this problem through a flow synthesis stage. This is achieved by fitting a 3D model to the given pose pair and project them back to the 2D plane to compute the dense appearance flow for training. The synthesized ground-truths are then used to train a feedforward network for efficient mapping from the input and target skeleton poses to the 3D appearance flow. With the appearance flow, we perform feature warping on the input image and generate a photorealistic image of the target pose. Extensive results on DeepFashion and Market-1501 datasets demonstrate the effectiveness of our approach over existing methods. Our code is available at http://mmlab.ie.cuhk.edu.hk/projects/pose-transfer. Chen Huang 0001, Chen Change Loy |
CVPR | 2 |
| 2019 | Not All Areas Are Equal: Transfer Learning for Semantic Segmentation via Hierarchical Region SelectionabstractThe success of deep neural networks for semantic segmentation heavily relies on large-scale and well-labeled datasets, which are hard to collect in practice. Synthetic data offers an alternative to obtain ground-truth labels for free. However, models directly trained on synthetic data often struggle to generalize to real images. In this paper, we consider transfer learning for semantic segmentation that aims to mitigate the gap between abundant synthetic data (source domain) and limited real data (target domain). Unlike previous approaches that either learn mappings to target domain or finetune on target images, our proposed method jointly learn from real images and selectively from realistic pixels in synthetic images to adapt to the target domain. Our key idea is to have weighting networks to score how similar the synthetic pixels are to real ones, and learn such weighting at pixel-, region- and image-levels. We jointly learn these hierarchical weighting networks and segmentation network in an end-to-end manner. Extensive experiments demonstrate that our proposed approach significantly outperforms other existing baselines, and is applicable to scenarios with extremely limited real images. Ruoqi Sun, Xinge Zhu, Chongruo Wu, Chen Huang 0001, Jianping Shi, Lizhuang Ma |
CVPR | 4 |
| 2019 | Addressing the Loss-Metric Mismatch with Adaptive Loss AlignmentabstractIn most machine learning training paradigms a fixed, often handcrafted, loss function is assumed to be a good proxy for an underlying evaluation metric. In this work we assess this assumption by meta-learning an adaptive loss function to directly optimize the evaluation metric. We propose a sample efficient reinforcement learning approach for adapting the loss dynamically during training. We empirically show how this formulation improves performance by simultaneously optimizing the evaluation metric and smoothing the loss landscape. We verify our method in metric learning and classification scenarios, showing considerable improvements over the state-of-the-art on a diverse set of tasks. Importantly, our method is applicable to a wide range of loss functions and evaluation metrics. Furthermore, the learned policies are transferable across tasks and data, demonstrating the versatility of the method. Chen Huang 0001, Shuangfei Zhai, Walter Talbott, Miguel Ángel Bautista 0001, Shih-Yu Sun, Carlos Guestrin, Joshua M. Susskind |
ICML | 1 |
| 2018 | Pose Guided Human Video Generation
Ceyuan Yang, Zhe Wang 0006, Xinge Zhu, Chen Huang 0001, Jianping Shi, Dahua Lin |
ECCV (10) | 4 |
| 2018 | Ensemble Knowledge Transfer for Semantic SegmentationabstractSemantic segmentation networks are usually learned in a strictly supervised manner, i.e., they are trained and tested on similar data distributions. Performance drops drastically in the presence of domain shifts. In this paper, we explore methods for learning across train and test distributions that dramatically differ in scene structure, viewpoints, and objects statistics. Motivated by the proliferation of aerial drone robotics, we consider the target task of semantic segmentation from aerial viewpoints. Inspired by the impact of Cityscapes [11], we introduce AeroScapes, a new dataset of 3269 images of aerial scenes (captured with a fleet of drones) annotated with dense semantic segmentations. Our dataset differs from existing segmentation datasets (that focus on ground-view or indoorscene domains) in terms of viewpoint, scene composition, and object scales. We propose a simple but effective approach for transferring knowledge from such diverse domains (for which considerable annotated training data exists) to our target task. To do so, we train multiple models for aerial segmentation via progressive fine-tuning through each source domain. We then treat these collections of models as an ensemble that can be aggregated to significantly improve performance. We demonstrate large absolute improvements (8.12%) over widely-used standard baselines. Ishan Nigam, Chen Huang 0001, Deva Ramanan |
WACV | 2 |
| 2018 | Discriminative Sparse Neighbor Approximation for Imbalanced LearningabstractData imbalance is common in many vision tasks where one or more classes are rare. Without addressing this issue, conventional methods tend to be biased toward the majority class with poor predictive accuracy for the minority class. These methods further deteriorate on small, imbalanced data that have a large degree of class overlap. In this paper, we propose a novel discriminative sparse neighbor approximation (DSNA) method to ameliorate the effect of class-imbalance during prediction. Specifically, given a test sample, we first traverse it through a cost-sensitive decision forest to collect a good subset of training examples in its local neighborhood. Then, we generate from this subset several class-discriminating but overlapping clusters and model each as an affine subspace. From these subspaces, the proposed DSNA iteratively seeks an optimal approximation of the test sample and outputs an unbiased prediction. We show that our method not only effectively mitigates the imbalance issue, but also allows the prediction to extrapolate to unseen data. The latter capability is crucial for achieving accurate prediction on small data set with limited samples. The proposed imbalanced learning method can be applied to both classification and regression tasks at a wide range of imbalance levels. It significantly outperforms the state-of-the-art methods that do not possess an imbalance handling mechanism, and is found to perform comparably or even better than recent deep learning methods by using hand-crafted features only. Chen Huang 0001, Chen Change Loy, Xiaoou Tang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Need for Speed: A Benchmark for Higher Frame Rate Object TrackingabstractIn this paper, we propose the first higher frame rate video dataset (called Need for Speed - NfS) and benchmark for visual object tracking. The dataset consists of 100 videos (380K frames) captured with now commonly available higher frame rate (240 FPS) cameras from real world scenarios. All frames are annotated with axis aligned bounding boxes and all sequences are manually labelled with nine visual attributes - such as occlusion, fast motion, background clutter, etc. Our benchmark provides an extensive evaluation of many recent and state-of-the-art trackers on higher frame rate sequences. We ranked each of these trackers according to their tracking accuracy and real-time performance. One of our surprising conclusions is that at higher frame rates, simple trackers such as correlation filters outperform complex methods based on deep networks. This suggests that for practical applications (such as in robotics or embedded vision), one needs to carefully tradeoff bandwidth constraints associated with higher frame rate acquisition, computational costs of real-time analysis, and the required application accuracy. Our dataset and benchmark allows for the first time (to our knowledge) systematic exploration of such issues, and will be made available to allow for further research in this space. Hamed Kiani Galoogahi, Ashton Fagg, Chen Huang 0001, Deva Ramanan, Simon Lucey |
ICCV | 3 |
| 2017 | Learning Policies for Adaptive Tracking with Deep Feature CascadesabstractVisual object tracking is a fundamental and time-critical vision task. Recent years have seen many shallow tracking methods based on real-time pixel-based correlation filters, as well as deep methods that have top performance but need a high-end GPU. In this paper, we learn to improve the speed of deep trackers without losing accuracy. Our fundamental insight is to take an adaptive approach, where easy frames are processed with cheap features (such as pixel values), while challenging frames are processed with invariant but expensive deep features. We formulate the adaptive tracking problem as a decision-making process, and learn an agent to decide whether to locate objects with high confidence on an early layer, or continue processing subsequent layers of a network. This significantly reduces the feedforward cost for easy frames with distinct or slow-moving objects. We train the agent offline in a reinforcement learning fashion, and further demonstrate that learning all deep layers (so as to provide good features for adaptive tracking) can lead to near real-time average tracking speed of 23 fps on a single CPU while achieving state-of-the-art performance. Perhaps most tellingly, our approach provides a 100X speedup for almost 50% of the time, indicating the power of an adaptive approach. Chen Huang 0001, Simon Lucey, Deva Ramanan |
ICCV | 1 |
| 2017 | Learning to Disambiguate by Asking Discriminative QuestionsabstractThe ability to ask questions is a powerful tool to gather information in order to learn about the world and resolve ambiguities. In this paper, we explore a novel problem of generating discriminative questions to help disambiguate visual instances. Our work can be seen as a complement and new extension to the rich research studies on image captioning and question answering. We introduce the first large-scale dataset with over 10,000 carefully annotated images-question tuples to facilitate benchmarking. In particular, each tuple consists of a pair of images and 4.6 discriminative questions (as positive samples) and 5.9 non-discriminative questions (as negative samples) on average. In addition, we present an effective method for visual discriminative question generation. The method can be trained in a weakly supervised manner without discriminative images-question tuples but just existing visual question answering datasets. Promising results are shown against representative baselines through quantitative evaluations and user studies. Chen Huang 0001, Xiaoou Tang, Chen Change Loy |
ICCV | 2 |
| 2016 | Learning Deep Representation for Imbalanced ClassificationabstractData in vision domain often exhibit highly-skewed class distribution, i.e., most data belong to a few majority classes, while the minority classes only contain a scarce amount of instances. To mitigate this issue, contemporary classification methods based on deep convolutional neural network (CNN) typically follow classic strategies such as class re-sampling or cost-sensitive training. In this paper, we conduct extensive and systematic experiments to validate the effectiveness of these classic schemes for representation learning on class-imbalanced data. We further demonstrate that more discriminative deep representation can be learned by enforcing a deep network to maintain both intercluster and inter-class margins. This tighter constraint effectively reduces the class imbalance inherent in the local data neighborhood. We show that the margins can be easily deployed in standard deep learning framework through quintuplet instance sampling and the associated triple-header hinge loss. The representation learned by our approach, when combined with a simple k-nearest neighbor (kNN) algorithm, shows significant improvements over existing methods on both high-and low-level vision classification tasks that exhibit imbalanced class distribution. Chen Huang 0001, Chen Change Loy, Xiaoou Tang |
CVPR | 1 |
| 2016 | Unsupervised Learning of Discriminative Attributes and Visual RepresentationsabstractAttributes offer useful mid-level features to interpret visual data. While most attribute learning methods are supervised by costly human-generated labels, we introduce a simple yet powerful unsupervised approach to learn and predict visual attributes directly from data. Given a large unlabeled image collection as input, we train deep Convolutional Neural Networks (CNNs) to output a set of discriminative, binary attributes often with semantic meanings. Specifically, we first train a CNN coupled with unsupervised discriminative clustering, and then use the cluster membership as a soft supervision to discover shared attributes from the clusters while maximizing their separability. The learned attributes are shown to be capable of encoding rich imagery properties from both natural images and contour patches. The visual representations learned in this way are also transferrable to other tasks such as object detection. We show other convincing results on the related tasks of image retrieval and classification, and contour detection. Chen Huang 0001, Chen Change Loy, Xiaoou Tang |
CVPR | 1 |
| 2016 | Human Attribute Recognition by Deep Hierarchical Contexts
Chen Huang 0001, Chen Change Loy, Xiaoou Tang |
ECCV (6) | 2 |
| 2016 | Local Similarity-Aware Deep Feature EmbeddingabstractExisting deep embedding methods in vision tasks are capable of learning a compact Euclidean space from images, where Euclidean distances correspond to a similarity metric. To make learning more effective and efficient, hard sample mining is usually employed, with samples identified through computing the Euclidean feature distance. However, the global Euclidean distance cannot faithfully characterize the true feature similarity in a complex visual feature space, where the intraclass distance in a high-density region may be larger than the interclass distance in low-density regions. In this paper, we introduce a Position-Dependent Deep Metric (PDDM) unit, which is capable of learning a similarity metric adaptive to local feature structure. The metric can be used to select genuinely hard samples in a local neighborhood to guide the deep embedding learning in an online and robust manner. The new layer is appealing in that it is pluggable to any convolutional networks and is trained end-to-end. Our local similarity-aware feature embedding not only demonstrates faster convergence and boosted performance on two complex image retrieval datasets, its large margin nature also leads to superior generalization results under the large and open set scenarios of transfer learning and zero-shot learning on ImageNet 2010 and ImageNet-10K datasets. Chen Huang 0001, Chen Change Loy, Xiaoou Tang |
NIPS | 1 |
| 2014 | Generalized joint kernel regression and adaptive dictionary learning for single-image super-resolution
Chen Huang 0001, Yicong Liang, Xiaoqing Ding, Chi Fang |
Signal Process. | 1 |
| 2014 | Robust Image Restoration via Adaptive Low-Rank Approximation and Joint Kernel RegressionabstractIn recent years, image priors based on nonlocal self-similarity and low-rank approximation have been proven as powerful tools for image restoration. Many restoration methods group similar patches as a matrix and recover the underlying low-rank structure from the corrupted matrix via rank minimization. However, both the nonlocally redundant and low-rank properties are highly content dependent, and whether they can faithfully characterize a wide range of natural images still remains unclear. In this paper, we analyze these two properties and provide quantifications of them in a data-driven and parametric way, respectively, obtaining the new measures of regional redundancy and nonlocal patch rank. Leveraging these prior leads to an adaptive image restoration method with content-awareness. In particular, our method iteratively removes outliers and recovers latent fine details. To handle outliers, we propose an adaptive low-rank and sparse matrix approximation algorithm to encourage the estimated nonlocal rank in the patch matrix. The guidance of regional redundancy further gives rise to the “denoise” quality. In the detail recovery step, we propose an adaptive joint kernel regression algorithm using the redundancy measure to determine the confidence of each regression group. It also bridges the gap between our online and offline dictionary learning schemes. Experiments on synthetic and real-world images show the efficacy of our method in image deblurring and super-resolution tasks, especially when subject to practical outliers such as rain drops. Chen Huang 0001, Xiaoqing Ding, Chi Fang, Di Wen 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Single-Image Super-Resolution via Adaptive Joint Kernel RegressionabstractSingle image super-resolution (SR) methods can be broadly categorized into three classes: interpolation-based methods, reconstruction-based methods [7], and example-based methods [2, 3, 6]. The reconstruction-based methods often incorporate prior knowledge to regularize the ill-posed problem. For example, Zhang et al. [7] assembled the Steering Kernel Regression [5] (SKR)-based local prior and Nonlocal Means [1] (NLM)based nonlocal prior. The example-based methods strongly rely on the chosen dictionary for satisfactory results. This paper focuses on learning good image priors and robust dictionaries for SR reconstruction. Among the extensively studied natural image priors, we choose to exploit the local structural regularity prior and nonlocal self-similarity prior in a coherent framework. We propose in this paper an Adaptive Joint Kernel Regression (AJKR)based prior to simultaneously exploit both image statistics. Our approach differs from others in several ways: 1) we combine a set of NLM-generalized local kernel regressors, which are more consistent with our nonlocal collaborative framework; 2) the proposed regional redundancy measure introduces higher-order statistics at the region level for each regression group, making the overall framework more adaptive (see Fig. 1(a)); and 3) an adaptive PCA-based dictionary learning scheme is adopted to bridge the gap of dictionaries learned online and offline by mixing them, and more importantly such a scheme together with its induced sparsity prior can adapt to the AJKR process in response to the regional redundancy measure (see the block diagram in Fig. 1(b)). The imaging model for SR (assume Y ∈ Rm is low resolution (LR) image, X ∈ Rn is high resolution (HR) image) is usually expressed as Chen Huang 0001, Xiaoqing Ding, Chi Fang |
BMVC | 1 |
| 2012 | Pose robust face tracking by combining view-based AAMs and temporal filters
Chen Huang 0001, Xiaoqing Ding, Chi Fang |
Comput. Vis. Image Underst. | 1 |
| 2010 | Head Pose Estimation Based on Random Forests for Multiclass ClassificationabstractHead pose estimation remains a unique challenge for computer vision system due to identity variation, illumination changes, noise, etc. Previous statistical approaches like PCA, linear discriminative analysis (LDA) and machine learning methods, including SVM and Adaboost, cannot achieve both accuracy and robustness that well. In this paper, we propose to use Gabor feature based random forests as the classification technique since they naturally handle such multi-class classification problem and are accurate and fast. The two sources of randomness, random inputs and random features, make random forests robust and able to deal with large feature spaces. Besides, we implement LDA as the node test to improve the discriminative power of individual trees in the forest, with each node generating both constant and variant number of children nodes. Experiments are carried out on two public databases to show the proposed algorithm outperforms other approaches in both accuracy and computational efficiency. Chen Huang 0001, Xiaoqing Ding, Chi Fang |
ICPR | 1 |