EDBT 2026 Demo / reviewers in the wild / expert
Chao Xu 0006
dblp:79/1442-6
· DBLP profile ↗
141ranked-venue papers
3as first author
37since 2021 · last 2026
0000-0002-6522-0982ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 94 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 86 · 33 since 2021Databases, data management, data science and information retrieval · 9Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward a Unified Complementary Fusion Framework for Robust Polarimetric ImagingabstractPolarization, as an intrinsic property of light alongside amplitude and phase, has demonstrated great potential in a variety of downstream applications by providing valuable physical cues encoded in the degree of polarization (DoP) and the angle of polarization (AoP). Polarimetric imaging aims to acquire these polarimetric parameters by capturing polarized snapshots. However, compared to conventional imaging, it faces greater difficulties due to the presence of polarizers, which attenuate light intensity in a spatially variant manner. Such attenuation complicates exposure control: a short exposure leads to low signal-to-noise ratio and color distortion, whereas a relatively long exposure increases the risk of motion blur and saturation. To address these challenges, this work proposes PolFusion+, a unified framework that robustly produces clean and sharp polarized snapshots by complementarily fusing a degraded pair of short-exposed noisy and long-exposed blurry inputs. Building upon a polarization-aware three-phase fusion scheme, PolFusion+ introduces two key advancements. First, to handle saturation in the blurry snapshot, the irradiance restoration phase extracts and rectifies color information from both inputs, effectively mitigating saturation-induced degradation. Second, to ensure physically faithful polarization reconstruction, the framework explicitly models the individual characteristics and interdependencies of the DoP and AoP, enabling their joint restoration. These improvements are supported by a degradation-oriented neural network tailored to the fusion scheme. Experimental results demonstrate that PolFusion+ achieves state-of-the-art performance, effectively benefiting downstream applications. Chu Zhou, Minggui Teng, Chao Xu 0006, Boxin Shi, Imari Sato |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | DiC: Rethinking Conv3x3 Designs in Diffusion ModelsabstractDiffusion models have shown exceptional performance in visual generation tasks. Recently, these models have shifted from traditional U-Shaped CNN-Attention hybrid structures to fully transformer-based isotropic architectures. While these transformers exhibit strong scalability and performance, their reliance on complicated self-attention operation results in slow inference speeds. Contrary to these works, we rethink one of the simplest yet fastest module in deep learning, 3x3 Convolution, to construct a scaled-up purely convolutional diffusion model. We first discover that an Encoder-Decoder Hourglass design outperforms scalable isotropic architectures for Conv3x3, but still under-performing our expectation. Further improving the architecture, we introduce sparse skip connections to reduce redundancy and improve scalability. Based on the architecture, we introduce conditioning improvements including stage-specific embeddings, mid-block condition injection, and conditional gating. These improvements lead to our proposed Diffusion CNN (DiC), which serves as a swift yet competitive diffusion architecture baseline. Experiments on various scales and settings show that DiC surpasses existing diffusion transformers by considerable margins in terms of performance while keeping a good speed advantage. Project page: https://github.com/YuchuanTian/DiC Yuchuan Tian, Chao Xu 0006, Hanting Chen |
CVPR | 5 |
| 2025 | Event-Based Visual Vibrometry
Peiqi Duan 0002, Yeliduosi Xiaokaiti, Chao Xu 0006, Boxin Shi |
ICCV | 4 |
| 2025 | U-REPA: Aligning Diffusion U-Nets to ViTsabstractRepresentation Alignment (REPA) that aligns Diffusion Transformer (DiT) hidden-states with ViT visual encoders has proven highly effective in DiT training, demonstrating superior convergence properties, but it has not been validated on the canonical diffusion U-Net architecture that shows faster convergence compared to DiTs. However, adapting REPA to U-Net architectures presents unique challenges: (1) different block functionalities necessitate revised alignment strategies; (2) spatial-dimension inconsistencies emerge from U-Net's spatial downsampling operations; (3) space gaps between U-Net and ViT hinder the effectiveness of tokenwise alignment. To encounter these challenges, we propose U-REPA, a representation alignment paradigm that bridges U-Net hidden states and ViT features as follows: Firstly, we propose via observation that due to skip connection, the middle stage of U-Net is the best alignment option. Secondly, we propose upsampling of U-Net features after passing them through MLPs. Thirdly, we observe difficulty when performing tokenwise similarity alignment, and further introduces a manifold loss that regularizes the relative similarity between samples. Experiments indicate that the resulting U-REPA could achieve excellent generation quality and greatly accelerates the convergence speed. With CFG guidance interval, U-REPA could reach FID<1.5 in 200 epochs or 1M iterations on ImageNet 256 $\times$ 256, and needs only half the total epochs to perform better than REPA under \textit{sd-vae-ft-ema}. Yuchuan Tian, Hanting Chen, Mengyu Zheng, Chao Xu 0006, Yunhe Wang 0001 |
NeurIPS | 5 |
| 2025 | Spk2ImgMamba: Spiking Camera Image Reconstruction with Multi-Scale State Space ModelsabstractAs a bio-inspired vision sensor, the spiking camera has showcased remarkable capability in high-speed imaging with a sampling rate of 40,000 Hz. Reconstructing clear images from continuous spike streams, which is obtained by each photosensor continuously detecting photons and firing them asynchronously, has garnered significant attention. Despite promising results, existing spike-to-image reconstruction methods face challenges in balancing global receptive fields and efficient computation due to the inherent limitations of their backbones. Recently, due to powerful long-range modeling and linear complexity, the state space model (SSM) has emerged as a competitive alternative to CNNs and Transformers. In this paper, we propose a lightweight spike-to-image reconstruction network that harnesses Mamba as the backbone. Our approach sequentially executes three core modules: temporal information integration, spatial feature enhancement, and progressive image reconstruction. The former accumulates cues across diverse temporal windows to explore both long-term and short-term contexts. Subsequently, to model global dependencies while heightening local detail perception, we develop a multi-scale SSM block characterized by multi-scale multi-direction scanning, which effectively boosts spatial feature representations. Finally, intensity images are decoded progressively from the enhanced light-intensity features. Extensive experiments on both synthetic and real-captured data demonstrate that our approach achieves state-of-the-art performance, with only 10% of the network parameters and nearly two orders of magnitude less computational effort. The code will be available at https://github.com/interstellarH/Spk2ImgMamba. Jiaoyang Yin, Bin Fan 0002, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi |
WACV | 3 |
| 2025 | Learning to Deblur Polarized Images
Chu Zhou, Minggui Teng, Chao Xu 0006, Imari Sato, Boxin Shi |
Int. J. Comput. Vis. | 4 |
| 2025 | Self-Supervised Learning for Rolling Shutter Temporal Super-ResolutionabstractMost cameras on portable devices adopt a rolling shutter (RS) mechanism, encoding sufficient temporal dynamic information through sequential readouts. This advantage can be exploited to recover a temporal sequence of latent global shutter (GS) images. Existing methods rely on fully supervised learning, necessitating specialized optical devices to collect paired RS-GS images as ground-truth, which is too costly to scale. In this paper, we propose a self-supervised learning framework for the first time to produce a high frame rate GS video from two consecutive RS images, unleashing the potential of RS cameras. Specifically, we first develop the unified warping model of RS2GS and GS2RS, enabling the complement conversions of RS2GS and GS2RS to be incorporated into a uniform network model. Then, based on the cycle consistency constraint, given a triplet of consecutive RS frames, we minimize the discrepancy between the input middle RS frame and its cycle reconstruction, generated by interpolating back from the predicted two intermediate GS frames. Experiments on various benchmarks show that our approach achieves comparable or better performance than state-of-the-art supervised methods while enjoying stronger generalization capabilities. Moreover, our approach makes it possible to recover smooth and distortion-free videos from two adjacent RS frames in the real-world BS-RSC dataset, surpassing prior limitations. Bin Fan 0002, Ying Guo 0014, Yuchao Dai, Chao Xu 0006, Boxin Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Image Processing GNN: Breaking Rigidity in Super-ResolutionabstractSuper-Resolution (SR) reconstructs high-resolution images from low-resolution ones. CNNs and window-attention methods are two major categories of canonical SR models. However, these measures are rigid: in both operations, each pixel gathers the same number of neighboring pixels, hindering their effectiveness in SR tasks. Alternatively, we leverage the flexibility of graphs and propose the Image Processing GNN (IPG) model to break the rigidity that dominates previous SR methods. Firstly, SR is unbalanced in that most reconstruction efforts are concentrated to a small proportion of detail-rich image parts. Hence, we leverage degree flexibility by assigning higher node degrees to detail-rich image nodes. Then in order to construct graphs for SR-effective aggregation, we treat images as pixel node sets rather than patch nodes. Lastly, we hold that both local and global information are crucial for SR performance. In the hope of gathering pixel information from both local and global scales efficiently via flexible graphs, we search node connections within nearby regions to construct local graphs; and find connections within a strided sampling space of the whole image for global graphs. The flexibility of graphs boosts the SR performance of the IPG model. Experiment results on various datasets demonstrates that the proposed IPG outperforms State-of-the-Art baselines. Codes are available at this link. Yuchuan Tian, Hanting Chen, Chao Xu 0006, Yunhe Wang 0001 |
CVPR | 3 |
| 2024 | EvDiG: Event-guided Direct and Global Components SeparationabstractSeparating the direct and global components of a scene aids in shape recovery and basic material understanding. Conventional methods capture multiple frames under high frequency illumination patterns or shadows, requiring the scene to keep stationary during the image acquisition process. Single-frame methods simplify the capture procedure but yield lower-quality separation results. In this paper, we leverage the event camera to facilitate the separation of direct and global components, enabling video-rate separation of high quality. In detail, we adopt an event camera to record rapid illumination changes caused by the shadow of a line occluder sweeping over the scene, and reconstruct the coarse separation results through event accumulation. We then design a network to resolve the noise in the coarse sep-aration results and restore color information. A real-world dataset is collected using a hybrid camera system for network training and evaluation. Experimental results show superior performance over state-of-the-art methods. Peiqi Duan 0002, Chu Zhou, Chao Xu 0006, Boxin Shi |
CVPR | 5 |
| 2024 | Multiscale Positive-Unlabeled Detection of AI-Generated TextsabstractRecent releases of Large Language Models (LLMs), e.g. ChatGPT, are astonishing at generating human-like texts, but they may impact the authenticity of texts. Previous works proposed methods to detect these AI-generated texts, including simple ML classifiers, pretrained-model-based zero-shot methods, and finetuned language classification models. However, mainstream detectors always fail on short texts, like SMSes, Tweets, and reviews. In this paper, a Multiscale Positive-Unlabeled (MPU) training framework is proposed to address the difficulty of short-text detection without sacrificing long-texts. Firstly, we acknowledge the human-resemblance property of short machine texts, and rephrase AI text detection as a partial Positive-Unlabeled (PU) problem by regarding these short machine texts as partially "unlabeled". Then in this PU context, we propose the length-sensitive Multiscale PU Loss, where a recurrent model in abstraction is used to estimate positive priors of scale-variant corpora. Additionally, we introduce a Text Multiscaling module to enrich training corpora. Experiments show that our MPU method augments detection performance on long AI-generated texts, and significantly improves short-text detection of language model detectors. Language Models trained with MPU could outcompete existing detectors on various short-text and long-text detection benchmarks. The codes are available at https://github.com/mindspore-lab/mindone/tree/master/examples/detect_chatgpt and https://github.com/YuchuanTian/AIGC_text_detector. Yuchuan Tian, Hanting Chen, Xutao Wang, Zheyuan Bai, Chao Xu 0006, Yunhe Wang 0001 |
ICLR | 7 |
| 2024 | Spatio-Temporal Interactive Learning for Efficient Image Reconstruction of Spiking CamerasabstractThe spiking camera is an emerging neuromorphic vision sensor that records high-speed motion scenes by asynchronously firing continuous binary spike streams. Prevailing image reconstruction methods, generating intermediate frames from these spike streams, often rely on complex step-by-step network architectures that overlook the intrinsic collaboration of spatio-temporal complementary information. In this paper, we propose an efficient spatio-temporal interactive reconstruction network to jointly perform inter-frame feature alignment and intra-frame feature filtering in a coarse-to-fine manner. Specifically, it starts by extracting hierarchical features from a concise hybrid spike representation, then refines the motion fields and target frames scale-by-scale, ultimately obtaining a full-resolution output. Meanwhile, we introduce a symmetric interactive attention block and a multi-motion field estimation block to further enhance the interaction capability of the overall network. Experiments on synthetic and real-captured data show that our approach exhibits excellent performance while maintaining low model complexity. Bin Fan 0002, Jiaoyang Yin, Yuchao Dai, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi |
NeurIPS | 4 |
| 2024 | MemoryFormer : Minimize Transformer Computation by Removing Fully-Connected LayersabstractIn order to reduce the computational complexity of large language models, great efforts have been made to to improve the efficiency of transformer models such as linear attention and flash-attention. However, the model size and corresponding computational complexity are constantly scaled up in pursuit of higher performance. In this work, we present MemoryFormer, a novel transformer architecture which significantly reduces the computational complexity (FLOPs) from a new perspective. We eliminate nearly all the computations of the transformer model except for the necessary computation required by the multi-head attention operation. This is made possible by utilizing an alternative method for feature transformation to replace the linear projection of fully-connected layers. Specifically, we first construct a group of in-memory lookup tables that store a large amount of discrete vectors to replace the weight matrix used in linear projection. We then use a hash algorithm to retrieve a correlated subset of vectors dynamically based on the input embedding. The retrieved vectors combined together will form the output embedding, which provides an estimation of the result of matrix multiplication operation in a fully-connected layer. Compared to conducting matrix multiplication, retrieving data blocks from memory is a much cheaper operation which requires little computations. We train MemoryFormer from scratch and conduct extensive experiments on various benchmarks to demonstrate the effectiveness of the proposed model. Yehui Tang 0001, Haochen Qin, Zhenli Zhou, Chao Xu 0006, Kai Han 0002, Heng Liao, Yunhe Wang 0001 |
NeurIPS | 5 |
| 2024 | U-DiTs: Downsample Tokens in U-Shaped Diffusion TransformersabstractDiffusion Transformers (DiTs) introduce the transformer architecture to diffusion tasks for latent-space image generation. With an isotropic architecture that chains a series of transformer blocks, DiTs demonstrate competitive performance and good scalability; but meanwhile, the abandonment of U-Net by DiTs and their following improvements is worth rethinking. To this end, we conduct a simple toy experiment by comparing a U-Net architectured DiT with an isotropic one. It turns out that the U-Net architecture only gain a slight advantage amid the U-Net inductive bias, indicating potential redundancies within the U-Net-style DiT. Inspired by the discovery that U-Net backbone features are low-frequency-dominated, we perform token downsampling on the query-key-value tuple for self-attention and bring further improvements despite a considerable amount of reduction in computation. Based on self-attention with downsampled tokens, we propose a series of U-shaped DiTs (U-DiTs) in the paper and conduct extensive experiments to demonstrate the extraordinary performance of U-DiT models. The proposed U-DiT could outperform DiT-XL with only 1/6 of its computation cost. Codes are available at https://github.com/YuchuanTian/U-DiT. Yuchuan Tian, Zhijun Tu, Hanting Chen, Jie Hu 0021, Chao Xu 0006, Yunhe Wang 0001 |
NeurIPS | 5 |
| 2024 | Unified Video Reconstruction for Rolling Shutter and Global Shutter CamerasabstractCurrently, the general domain of video reconstruction (VR) is fragmented into different shutters spanning global shutter and rolling shutter cameras. Despite rapid progress in the state-of-the-art, existing methods overwhelmingly follow shutter-specific paradigms and cannot conceptually generalize to other shutter types, hindering the uniformity of VR models. In this paper, we propose UniVR, a versatile framework to handle various shutters through unified modeling and shared parameters. Specifically, UniVR encodes diverse shutter types into a unified space via a tractable shutter adapter, which is parameter-free and thus can be seamlessly delivered to current well-established VR architectures for cross-shutter transfer. To demonstrate its effectiveness, we conceptualize UniVR as three shutter-generic VR methods, namely Uni-SoftSplat, Uni-SuperSloMo, and Uni-RIFE. Extensive experimental results demonstrate that the pre-trained model without any fine-tuning can achieve reasonable performance even on novel shutters. After fine-tuning, new state-of-the-art performances are established that go beyond shutter-specific methods and enjoy strong generalization. The code is available at https://github.com/GitCVfb/UniVR. Bin Fan 0002, Zhexiong Wan, Boxin Shi, Chao Xu 0006, Yuchao Dai |
IEEE Trans. Image Process. | 4 |
| 2023 | Polarization-Aware Low-Light Image EnhancementabstractPolarization-based vision algorithms have found uses in various applications since polarization provides additional physical constraints. However, in low-light conditions, their performance would be severely degenerated since the captured polarized images could be noisy, leading to noticeable degradation in the degree of polarization (DoP) and the angle of polarization (AoP). Existing low-light image enhancement methods cannot handle the polarized images well since they operate in the intensity domain, without effectively exploiting the information provided by polarization. In this paper, we propose a Stokes-domain enhancement pipeline along with a dual-branch neural network to handle the problem in a polarization-aware manner. Two application scenarios (reflection removal and shape from polarization) are presented to show how our enhancement can improve their results. Chu Zhou, Minggui Teng, Youwei Lyu, Si Li 0001, Chao Xu 0006, Boxin Shi |
AAAI | 5 |
| 2023 | Network Expansion For Practical Training AccelerationabstractRecently, the sizes of deep neural networks and training datasets both increase drastically to pursue better performance in a practical sense. With the prevalence of transformer-based models in vision tasks, even more pressure is laid on the GPU platforms to train these heavy models, which consumes a large amount of time and computing resources as well. Therefore, it's crucial to accelerate the training process of deep neural networks. In this paper, we propose a general network expansion method to reduce the practical time cost of the model training process. Specifically, we utilize both width- and depth-level sparsity of dense models to accelerate the training of deep neural networks. Firstly, we pick a sparse sub-network from the original dense model by reducing the number of parameters as the starting point of training. Then the sparse architecture will gradually expand during the training procedure and finally grow into a dense one. We design different expanding strategies to grow CNNs and ViTs respectively, due to the great heterogeneity in between the two architectures. Our method can be easily integrated into popular deep learning frameworks, which saves considerable training time and hardware resources. Extensive experiments show that our acceleration method can significantly speed up the training process of modern vision models on general GPU devices with negligible performance drop (e.g. 1.42× faster for ResNet-101 and 1.34× faster for DeiT-base on ImageNet-1k). The code is available at https://github.com/huawei-noah/Efficient-Computing/tree/master/TrainingAcceleration/NetworkExpansion and https://gitee.com/mindspore/hub/blob/master/mshub_res/assets/noah-cvlab/gpu/1.8/networkexpansion_v1.0_imagenet2012.md Yehui Tang 0001, Kai Han 0002, Chao Xu 0006, Yunhe Wang 0001 |
CVPR | 4 |
| 2023 | Towards Higher Ranks via Adversarial Weight PruningabstractConvolutional Neural Networks (CNNs) are hard to deploy on edge devices due to its high computation and storage complexities. As a common practice for model compression, network pruning consists of two major categories: unstructured and structured pruning, where unstructured pruning constantly performs better. However, unstructured pruning presents a structured pattern at high pruning rates, which limits its performance. To this end, we propose a Rank-based PruninG (RPG) method to maintain the ranks of sparse weights in an adversarial manner. In each step, we minimize the low-rank approximation error for the weight matrices using singular value decomposition, and maximize their distance by pushing the weight matrices away from its low rank approximation. This rank-based optimization objective guides sparse weights towards a high-rank topology. The proposed method is conducted in a gradual pruning fashion to stabilize the change of rank during training. Experimental results on various datasets and different tasks demonstrate the effectiveness of our algorithm in high sparsity. The proposed RPG outperforms the state-of-the-art performance by 1.13\% top-1 accuracy on ImageNet in ResNet-50 with 98\% sparsity. The codes are available at https://github.com/huawei-noah/Efficient-Computing/tree/master/Pruning/RPG and https://gitee.com/mindspore/models/tree/master/research/cv/RPG. Yuchuan Tian, Hanting Chen, Tianyu Guo 0001, Chao Xu 0006, Yunhe Wang 0001 |
NeurIPS | 4 |
| 2023 | Deblurring Low-Light Images with Events
Chu Zhou, Minggui Teng, Jin Han 0001, Jinxiu Liang, Chao Xu 0006, Boxin Shi |
Int. J. Comput. Vis. | 5 |
| 2023 | Hybrid High Dynamic Range Imaging fusing Neuromorphic and Conventional ImagesabstractReconstruction of high dynamic range image from a single low dynamic range image captured by a conventional RGB camera, which suffers from over- or under-exposure, is an ill-posed problem. In contrast, recent neuromorphic cameras like event camera and spike camera can record high dynamic range scenes in the form of intensity maps, but with much lower spatial resolution and no color information. In this article, we propose a hybrid imaging system (denoted as NeurImg) that captures and fuses the visual information from a neuromorphic camera and ordinary images from an RGB camera to reconstruct high-quality high dynamic range images and videos. The proposed NeurImg-HDR+ network consists of specially designed modules, which bridges the domain gaps on resolution, dynamic range, and color representation between two types of sensors and images to reconstruct high-resolution, high dynamic range images and videos. We capture a test dataset of hybrid signals on various HDR scenes using the hybrid camera, and analyze the advantages of the proposed fusing strategy by comparing it to state-of-the-art inverse tone mapping methods and merging two low dynamic range images approaches. Quantitative and qualitative experiments on both synthetic data and real-world scenarios demonstrate the effectiveness of the proposed hybrid high dynamic range imaging system. Code and dataset can be found at: https://github.com/hjynwa/NeurImg-HDR. Jin Han 0001, Yixin Yang 0008, Peiqi Duan 0002, Chu Zhou, Lei Ma 0008, Chao Xu 0006, Tiejun Huang 0001, Imari Sato, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Polarization Guided HDR Reconstruction via Pixel-Wise DepolarizationabstractTaking photos with digital cameras often accompanies saturated pixels due to their limited dynamic range, and it is far too ill-posed to restore them. Capturing multiple low dynamic range images with bracketed exposures can make the problem less ill-posed, however, it is prone to ghosting artifacts caused by spatial misalignment among images. A polarization camera can capture four spatially-aligned and temporally-synchronized polarized images with different polarizer angles in a single shot, which can be used for ghost-free high dynamic range (HDR) reconstruction. However, real-world scenarios are still challenging since existing polarization-based HDR reconstruction methods treat all pixels in the same manner and only utilize the spatially-variant exposures of the polarized images (without fully exploiting the degree of polarization (DoP) and the angle of polarization (AoP) of the incoming light to the sensor, which encode abundant structural and contextual information of the scene) to handle the problem still in an ill-posed manner. In this paper, we propose a pixel-wise depolarization strategy to solve the polarization guided HDR reconstruction problem, by classifying the pixels based on their levels of ill-posedness in HDR reconstruction procedure and applying different solutions to different classes. To utilize the strategy with better generalization ability and higher robustness, we propose a network-physics-hybrid polarization-based HDR reconstruction pipeline along with a neural network tailored to it, fully exploiting the DoP and AoP. Experimental results show that our approach achieves state-of-the-art performance on both synthetic and real-world images. Chu Zhou, Yufei Han 0002, Minggui Teng, Jin Han 0001, Si Li 0001, Chao Xu 0006, Boxin Shi |
IEEE Trans. Image Process. | 6 |
| 2022 | Source-Free Domain Adaptation via Distribution EstimationabstractDomain Adaptation aims to transfer the knowledge learned from a labeled source domain to an unlabeled target domain whose data distributions are different. However, the training data in source domain required by most of the existing methods is usually unavailable in real-world applications due to privacy preserving policies. Recently, Source-Free Domain Adaptation (SFDA) has drawn much attention, which tries to tackle domain adaptation problem without using source data. In this work, we propose a novel framework called SFDA-DE to address SFDA task via source Distribution Estimation. Firstly, we produce robust pseudo-labels for target data with spherical k-means clustering, whose initial class centers are the weight vectors (anchors) learned by the classifier of pretrained model. Furthermore, we propose to estimate the class-conditioned feature distribution of source domain by exploiting target data and corresponding anchors. Finally, we sample surrogate features from the estimated distribution, which are then utilized to align two domains by minimizing a contrastive adaptation loss function. Extensive experiments show that the proposed method achieves state-of-the-art performance on multiple DA benchmarks, and even outperforms traditional DA methods which require plenty of source data. Yixing Xu, Yehui Tang 0001, Chao Xu 0006, Yunhe Wang 0001, Dacheng Tao |
CVPR | 4 |
| 2022 | Hire-MLP: Vision MLP via Hierarchical RearrangementabstractPrevious vision MLPs such as MLP-Mixer and ResMLP accept linearly flattened image patches as input, making them inflexible for different input sizes and hard to capture spatial information. Such approach withholds MLPs from getting comparable performance with their transformer-based counterparts and prevents them from becoming a general backbone for computer vision. This paper presents Hire-MLP, a simple yet competitive vision MLP architecture via Hierarchical rearrangement, which contains two levels of rearrangements. Specifically, the inner-region rearrangement is proposed to capture local information inside a spatial region, and the cross-region rearrangement is proposed to enable information communication between different regions and capture global context by circularly shifting all tokens along spatial directions. Extensive experiments demonstrate the effectiveness of Hire-MLP as a versatile backbone for various vision tasks. In particular, Hire-MLP achieves competitive results on image classification, object detection and semantic segmentation tasks, e.g., 83.8% top-1 accuracy on ImageNet, 51.7% box AP and 44.8% mask AP on COCO val2017, and 49.9% mIoU on ADE20K, surpassing previous transformer-based and MLP-based models with better trade-off for accuracy and throughput. Jianyuan Guo, Yehui Tang 0001, Kai Han 0002, Xinghao Chen 0001, Han Wu 0009, Chao Xu 0006, Chang Xu 0002, Yunhe Wang 0001 |
CVPR | 6 |
| 2022 | Patch Slimming for Efficient Vision TransformersabstractThis paper studies the efficiency problem for visual transformers by excavating redundant calculation in given networks. The recent transformer architecture has demonstrated its effectiveness for achieving excellent performance on a series of computer vision tasks. However, similar to that of convolutional neural networks, the huge computational cost of vision transformers is still a severe issue. Considering that the attention mechanism aggregates different patches layer-by-layer, we present a novel patch slimming approach that discards useless patches in a topdown paradigm. We first identify the effective patches in the last layer and then use them to guide the patch selection process of previous layers. For each layer, the impact of a patch on the final output feature is approximated and patches with less impacts will be removed. Experimental results on benchmark datasets demonstrate that the proposed method can significantly reduce the computational costs of vision transformers without affecting their performances. For example, over 45% FLOPs of the ViT-Ti model can be reduced with only 0.2% top-1 accuracy drop on the ImageNet dataset. Yehui Tang 0001, Kai Han 0002, Yunhe Wang 0001, Chang Xu 0002, Jianyuan Guo, Chao Xu 0006, Dacheng Tao |
CVPR | 6 |
| 2022 | An Image Patch is a Wave: Phase-Aware Vision MLPabstractIn the field of computer vision, recent works show that a pure MLP architecture mainly stacked by fully-connected layers can achieve competing performance with CNN and transformer. An input image of vision MLP is usually split into multiple tokens (patches), while the existing MLP models directly aggregate them with fixed weights, neglecting the varying semantic information of tokens from different images. To dynamically aggregate tokens, we propose to represent each token as a wave function with two parts, amplitude and phase. Amplitude is the original feature and the phase term is a complex value changing according to the semantic contents of input images. Introducing the phase term can dynamically modulate the relationship between tokens and fixed weights in MLP. Based on the wave-like token representation, we establish a novel Wave-MLP architecture for vision tasks. Extensive experiments demonstrate that the proposed Wave-MLP is superior to the state-of-the-art MLP architectures on various vision tasks such as image classification, object detection and semantic segmentation. The source code is available at https://github.com/huawei-noah/CV-Backbones/tree/master/wavemlp_pytorch and https://gitee.com/mindspore/models/tree/master/research/cv/wave_mlp. Yehui Tang 0001, Kai Han 0002, Jianyuan Guo, Chang Xu 0002, Yanxi Li 0001, Chao Xu 0006, Yunhe Wang 0001 |
CVPR | 6 |
| 2022 | Federated Learning with Positive and Unlabeled DataabstractWe study the problem of learning from positive and unlabeled (PU) data in the federated setting, where each client only labels a little part of their dataset due to the limitation of resources and time. Different from the settings in traditional PU learning where the negative class consists of a single class, the negative samples which cannot be identified by a client in the federated setting may come from multiple classes which are unknown to the client. Therefore, existing PU learning methods can be hardly applied in this situation. To address this problem, we propose a novel framework, namely Federated learning with Positive and Unlabeled data (FedPU), to minimize the expected risk of multiple negative classes by leveraging the labeled data in other clients. We theoretically analyze the generalization bound of the proposed FedPU. Empirical experiments show that the FedPU can achieve much better performance than conventional supervised and semi-supervised federated learning methods. Xinyang Lin, Hanting Chen, Yixing Xu, Chao Xu 0006, Xiaolin Gui, Yiping Deng, Yunhe Wang 0001 |
ICML | 4 |
| 2022 | GhostNetV2: Enhance Cheap Operation with Long-Range AttentionabstractLight-weight convolutional neural networks (CNNs) are specially designed for applications on mobile devices with faster inference speed. The convolutional operation can only capture local information in a window region, which prevents performance from being further improved. Introducing self-attention into convolution can capture global information well, but it will largely encumber the actual speed. In this paper, we propose a hardware-friendly attention mechanism (dubbed DFC attention) and then present a new GhostNetV2 architecture for mobile applications. The proposed DFC attention is constructed based on fully-connected layers, which can not only execute fast on common hardware but also capture the dependence between long-range pixels. We further revisit the expressiveness bottleneck in previous GhostNet and propose to enhance expanded features produced by cheap operations with DFC attention, so that a GhostNetV2 block can aggregate local and long-range information simultaneously. Extensive experiments demonstrate the superiority of GhostNetV2 over existing architectures. For example, it achieves 75.3% top-1 accuracy on ImageNet with 167M FLOPs, significantly suppressing GhostNetV1 (74.5%) with a similar computational cost. The source code will be available at https://github.com/huawei-noah/Efficient-AI-Backbones/tree/master/ghostnetv2_pytorch and https://gitee.com/mindspore/models/tree/master/research/cv/ghostnetv2. Yehui Tang 0001, Kai Han 0002, Jianyuan Guo, Chang Xu 0002, Chao Xu 0006, Yunhe Wang 0001 |
NeurIPS | 5 |
| 2022 | Optimizing Latent Distributions for Non-Adversarial Generative NetworksabstractThe generator in generative adversarial networks (GANs) is driven by a discriminator to produce high-quality images through an adversarial game. At the same time, the difficulty of reaching a stable generator has been increased. This paper focuses on non-adversarial generative networks that are trained in a plain manner without adversarial loss. The given limited number of real images could be insufficient to fully represent the real data distribution. We therefore investigate a set of distributions in a Wasserstein ball centred on the distribution induced by the training data and propose to optimize the generator over this Wasserstein ball. We theoretically discuss the solvability of the newly defined objective function and develop a tractable reformulation to learn the generator. The connections and differences between the proposed non-adversarial generative networks and GANs are analyzed. Experimental results on real-world datasets demonstrate that the proposed algorithm can effectively learn image generators in a non-adversarial approach, and the generated images are of comparable quality with those from GANs. Tianyu Guo 0001, Chang Xu 0002, Boxin Shi, Chao Xu 0006, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Training Neural Networks by Lifted Proximal Operator MachinesabstractWe present the lifted proximal operator machine (LPOM) to train fully-connected feed-forward neural networks. LPOM represents the activation function as an equivalent proximal operator and adds the proximal operators to the objective function of a network as penalties. LPOM is block multi-convex in all layer-wise weights and activations. This allows us to develop a new block coordinate descent (BCD) method with convergence guarantee to solve it. Due to the novel formulation and solving method, LPOM only uses the activation function itself and does not require any gradient steps. Thus it avoids the gradient vanishing or exploding issues, which are often blamed in gradient-based methods. Also, it can handle various non-decreasing Lipschitz continuous activation functions. Additionally, LPOM is almost as memory-efficient as stochastic gradient descent and its parameter tuning is relatively easy. We further implement and analyze the parallel solution of LPOM. We first propose a general asynchronous-parallel BCD method with convergence guarantee. Then we use it to solve LPOM, resulting in asynchronous-parallel LPOM. For faster speed, we develop the synchronous-parallel LPOM. We validate the advantages of LPOM on various network architectures and datasets. We also apply synchronous-parallel LPOM to autoencoder training and demonstrate its fast convergence and superior performance. Jia Li 0002, Mingqing Xiao 0002, Cong Fang 0001, Yue Dai 0003, Chao Xu 0006, Zhouchen Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Pre-Trained Image Processing TransformerabstractAs the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectiveness over conventional methods. The big progress is mainly contributed to the representation ability of transformer and its variant architectures. In this paper, we study the low-level computer vision task (e.g., denoising, super-resolution and deraining) and develop a new pre-trained model, namely, image processing transformer (IPT). To maximally excavate the capability of transformer, we present to utilize the well-known ImageNet benchmark for generating a large amount of corrupted image pairs. The IPT model is trained on these images with multi-heads and multi-tails. In addition, the contrastive learning is introduced for well adapting to different image processing tasks. The pre-trained model can therefore efficiently employed on desired task after fine-tuning. With only one pre-trained model, IPT outperforms the current state-of-the-art methods on various low-level benchmarks. Code is available at https://github.com/huawei-noah/Pretrained-IPT and https://gitee.com/mindspore/mindspore/tree/master/model_zoo/research/cv/IPT Hanting Chen, Yunhe Wang 0001, Tianyu Guo 0001, Chang Xu 0002, Yiping Deng, Zhenhua Liu 0003, Siwei Ma 0001, Chunjing Xu, Chao Xu 0006, Wen Gao 0001 |
CVPR | 9 |
| 2021 | Learning Student Networks in the WildabstractData-free learning for student networks is a new paradigm for solving users’ anxiety caused by the privacy problem of using original training data. Since the architectures of modern convolutional neural networks (CNNs) are compact and sophisticated, the alternative images or meta-data generated from the teacher network are often broken. Thus, the student network cannot achieve the comparable performance to that of the pre-trained teacher network especially on the large-scale image dataset. Different to previous works, we present to maximally utilize the massive available unlabeled data in the wild. Specifically, we first thoroughly analyze the output differences between teacher and student network on the original data and develop a data collection method. Then, a noisy knowledge distillation algorithm is proposed for achieving the performance of the student network. In practice, an adaptation matrix is learned with the student network for correcting the label noise produced by the teacher network on the collected unlabeled images. The effectiveness of our DFND (Data-Free Noisy Distillation) method is then verified on several benchmarks to demonstrate its superiority over state-of-the-art data-free distillation methods. Experiments on various datasets demonstrate that the student networks learned by the proposed method can achieve comparable performance with those using the original dataset. Code is available at https://github.com/huawei-noah/Data-Efficient-Model-Compression Hanting Chen, Tianyu Guo 0001, Chang Xu 0002, Chunjing Xu, Chao Xu 0006, Yunhe Wang 0001 |
CVPR | 6 |
| 2021 | Manifold Regularized Dynamic Network PruningabstractNeural network pruning is an essential approach for reducing the computational complexity of deep models so that they can be well deployed on resource-limited devices. Compared with conventional methods, the recently developed dynamic pruning methods determine redundant filters variant to each input instance which achieves higher acceleration. Most of the existing methods discover effective subnetworks for each instance independently and do not utilize the relationship between different inputs. To maximally excavate redundancy in the given network architecture, this paper proposes a new paradigm that dynamically removes redundant filters by embedding the manifold information of all instances into the space of pruned networks (dubbed as ManiDP). We first investigate the recognition complexity and feature similarity between images in the training set. Then, the manifold relationship between instances and the pruned sub-networks will be aligned in the training procedure. The effectiveness of the proposed method is verified on several benchmarks, which shows better performance in terms of both accuracy and computational cost compared to the state-of-the-art methods. For example, our method can reduce 55.3% FLOPs of ResNet-34 with only 0.57% top-1 accuracy degradation on ImageNet. The code will be available at https://github.com/huawei-noah/Pruning/tree/master/ManiDP. Yehui Tang 0001, Yunhe Wang 0001, Yixing Xu, Yiping Deng, Chao Xu 0006, Dacheng Tao, Chang Xu 0002 |
CVPR | 5 |
| 2021 | HourNAS: Extremely Fast Neural Architecture Search Through an Hourglass LensabstractNeural Architecture Search (NAS) aims to automatically discover optimal architectures. In this paper, we propose an hourglass-inspired approach (HourNAS) for extremely fast NAS. It is motivated by the fact that the effects of the architecture often proceed from the vital few blocks. Acting like the narrow neck of an hourglass, vital blocks in the guaranteed path from the input to the output of a deep neural network restrict the information flow and influence the network accuracy. The other blocks occupy the major volume of the network and determine the overall network complexity, corresponding to the bulbs of an hourglass. To achieve an extremely fast NAS while preserving the high accuracy, we propose to identify the vital blocks and make them the priority in the architecture search. The search space of those non-vital blocks is further shrunk to only cover the candidates that are affordable under the computational resource constraints. Experimental results on ImageNet show that only using 3 hours (0.1 days) with one GPU, our HourNAS can search an architecture that achieves a 77.0% Top-1 accuracy, which outperforms the state-of-the-art methods. Zhaohui Yang 0003, Yunhe Wang 0001, Xinghao Chen 0001, Jianyuan Guo, Wei Zhang 0196, Chao Xu 0006, Chunjing Xu, Dacheng Tao, Chang Xu 0002 |
CVPR | 6 |
| 2021 | EvIntSR-Net: Event Guided Multiple Latent Frames Reconstruction and Super-resolutionabstractAn event camera detects the scene radiance changes and sends a sequence of asynchronous event streams with high dynamic range, high temporal resolution, and low latency. However, the spatial resolution of event cameras is limited as a trade-off for these outstanding properties. To reconstruct high-resolution intensity images from event data, we propose EvIntSR-Net that converts Event data to multiple latent Intensity frames to achieve Super-Resolution on intensity images in this paper. EvIntSR-Net bridges the domain gap between event streams and intensity frames and learns to merge a sequence of latent intensity frames in a recurrent updating manner. Experimental results show that EvIntSR-Net can reconstruct SR intensity images with higher dynamic range and fewer blurry artifacts by fusing events with intensity frames for both simulated and real-world data. Furthermore, the proposed EvIntSR-Net is able to generate high-frame-rate videos with super-resolved frames. Jin Han 0001, Yixin Yang 0008, Chu Zhou, Chao Xu 0006, Boxin Shi |
ICCV | 4 |
| 2021 | Augmented Shortcuts for Vision TransformersabstractTransformer models have achieved great progress on computer vision tasks recently. The rapid development of vision transformers is mainly contributed by their high representation ability for extracting informative features from input images. However, the mainstream transformer models are designed with deep architectures, and the feature diversity will be continuously reduced as the depth increases, \ie, feature collapse. In this paper, we theoretically analyze the feature collapse phenomenon and study the relationship between shortcuts and feature diversity in these transformer models. Then, we present an augmented shortcut scheme, which inserts additional paths with learnable parameters in parallel on the original shortcuts. To save the computational costs, we further explore an efficient approach that uses the block-circulant projection to implement augmented shortcuts. Extensive experiments conducted on benchmark datasets demonstrate the effectiveness of the proposed method, which brings about 1% accuracy increase of the state-of-the-art visual transformers without obviously increasing their parameters and FLOPs. Yehui Tang 0001, Kai Han 0002, Chang Xu 0002, An Xiao, Yiping Deng, Chao Xu 0006, Yunhe Wang 0001 |
NeurIPS | 6 |
| 2021 | Learning to dehaze with polarizationabstractHaze, a common kind of bad weather caused by atmospheric scattering, decreases the visibility of scenes and degenerates the performance of computer vision algorithms. Single-image dehazing methods have shown their effectiveness in a large variety of scenes, however, they are based on handcrafted priors or learned features, which do not generalize well to real-world images. Polarization information can be used to relieve its ill-posedness, however, real-world images are still challenging since existing polarization-based methods usually assume that the transmitted light is not significantly polarized, and they require specific clues to estimate necessary physical parameters. In this paper, we propose a generalized physical formation model of hazy images and a robust polarization-based dehazing pipeline without the above assumption or requirement, along with a neural network tailored to the pipeline. Experimental results show that our approach achieves state-of-the-art performance on both synthetic data and real-world hazy images. Chu Zhou, Minggui Teng, Yufei Han 0002, Chao Xu 0006, Boxin Shi |
NeurIPS | 4 |
| 2021 | Training object detectors from few weakly-labeled and many unlabeled images
Zhaohui Yang 0003, Miaojing Shi, Chao Xu 0006, Vittorio Ferrari, Yannis Avrithis |
Pattern Recognit. | 3 |
| 2021 | Learning Student Networks via Feature EmbeddingabstractDeep convolutional neural networks have been widely used in numerous applications, but their demanding storage and computational resource requirements prevent their applications on mobile devices. Knowledge distillation aims to optimize a portable student network by taking the knowledge from a well-trained heavy teacher network. Traditional teacher-student-based methods used to rely on additional fully connected layers to bridge intermediate layers of teacher and student networks, which brings in a large number of auxiliary parameters. In contrast, this article aims to propagate information from teacher to student without introducing new variables that need to be optimized. We regard the teacher-student paradigm from a new perspective of feature embedding. By introducing the locality preserving loss, the student network is encouraged to generate the low-dimensional features that could inherit intrinsic properties of their corresponding high-dimensional features from the teacher network. The resulting portable network, thus, can naturally maintain the performance as that of the teacher network. Theoretical analysis is provided to justify the lower computation complexity of the proposed method. Experiments on benchmark data sets and well-trained networks suggest that the proposed algorithm is superior to state-of-the-art teacher-student learning methods in terms of computational and storage complexity. Hanting Chen, Yunhe Wang 0001, Chang Xu 0002, Chao Xu 0006, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Distilling Portable Generative Adversarial Networks for Image TranslationabstractDespite Generative Adversarial Networks (GANs) have been widely used in various image-to-image translation tasks, they can be hardly applied on mobile devices due to their heavy computation and storage cost. Traditional network compression methods focus on visually recognition tasks, but never deal with generation tasks. Inspired by knowledge distillation, a student generator of fewer parameters is trained by inheriting the low-level and high-level information from the original heavy teacher generator. To promote the capability of student generator, we include a student discriminator to measure the distances between real images, and images generated by student and teacher generators. An adversarial learning process is therefore established to optimize student generator and student discriminator. Qualitative and quantitative analysis by conducting experiments on benchmark datasets demonstrate that the proposed method can learn portable generative models with strong performance. Hanting Chen, Yunhe Wang 0001, Han Shu, Changyuan Wen, Chunjing Xu, Boxin Shi, Chao Xu 0006, Chang Xu 0002 |
AAAI | 7 |
| 2020 | Beyond Dropout: Feature Map Distortion to Regularize Deep Neural NetworksabstractDeep neural networks often consist of a great number of trainable parameters for extracting powerful features from given datasets. One one hand, massive trainable parameters significantly enhance the performance of these deep networks. One the other hand, they bring the problem of over-fitting. To this end, dropout based methods disable some elements in the output feature maps during the training phase for reducing the co-adaptation of neurons. Although the generalization ability of the resulting models can be enhanced by these approaches, the conventional binary dropout is not the optimal solution. Therefore, we investigate the empirical Rademacher complexity related to intermediate layers of deep neural networks and propose a feature distortion method for addressing the aforementioned problem. In the training period, randomly selected elements in the feature maps will be replaced with specific values by exploiting the generalization error bound. The superiority of the proposed feature map distortion for producing deep neural network with higher testing performance is analyzed and demonstrated on several benchmark image datasets. Yehui Tang 0001, Yunhe Wang 0001, Yixing Xu, Boxin Shi, Chao Xu 0006, Chunjing Xu, Chang Xu 0002 |
AAAI | 5 |
| 2020 | Reborn Filters: Pruning Convolutional Neural Networks with Limited DataabstractChannel pruning is effective in compressing the pretrained CNNs for their deployment on low-end edge devices. Most existing methods independently prune some of the original channels and need the complete original dataset to fix the performance drop after pruning. However, due to commercial protection or data privacy, users may only have access to a tiny portion of training examples, which could be insufficient for the performance recovery. In this paper, for pruning with limited data, we propose to use all original filters to directly develop new compact filters, named reborn filters, so that all useful structure priors in the original filters can be well preserved into the pruned networks, alleviating the performance drop accordingly. During training, reborn filters can be easily implemented via 1×1 convolutional layers and then be fused in the inference stage for acceleration. Based on reborn filters, the proposed channel pruning algorithm shows its effectiveness and superiority on extensive experiments. Yehui Tang 0001, Shan You, Chang Xu 0002, Jin Han 0001, Chen Qian 0006, Boxin Shi, Chao Xu 0006, Changshui Zhang |
AAAI | 7 |
| 2020 | Frequency Domain Compact 3D Convolutional Neural NetworksabstractThis paper studies the compression and acceleration of 3-dimensional convolutional neural networks (3D CNNs). To reduce the memory cost and computational complexity of deep neural networks, a number of algorithms have been explored by discovering redundant parameters in pre-trained networks. However, most of existing methods are designed for processing neural networks consisting of 2-dimensional convolution filters (i.e. image classification and detection) and cannot be straightforwardly applied for 3-dimensional filters (i.e. time series data). In this paper, we develop a novel approach for eliminating redundancy in the time dimensionality of 3D convolution filters by converting them into the frequency domain through a series of learned optimal transforms with extremely fewer parameters. Moreover, these transforms are forced to be orthogonal, and the calculation of feature maps can be accomplished in the frequency domain to achieve considerable speed-up rates. Experimental results on benchmark 3D CNN models and datasets demonstrate that the proposed Frequency Domain Compact 3D CNNs (FDC3D) can achieve the state-of-the-art performance, \eg a 2x speed-up ratio on the 3D-ResNet-18 without obviously affecting its accuracy. Hanting Chen, Yunhe Wang 0001, Han Shu, Yehui Tang 0001, Chunjing Xu, Boxin Shi, Chao Xu 0006, Qi Tian 0001, Chang Xu 0002 |
CVPR | 7 |
| 2020 | AdderNet: Do We Really Need Multiplications in Deep Learning?abstractCompared with cheap addition operation, multiplication operation is of much higher computation complexity. The widely-used convolutions in deep neural networks are exactly cross-correlation to measure the similarity between input feature and convolution filters, which involves massive multiplications between float values. In this paper, we present adder networks (AdderNets) to trade these massive multiplications in deep neural networks, especially convolutional neural networks (CNNs), for much cheaper additions to reduce computation costs. In AdderNets, we take the L1-norm distance between filters and input feature as the output response. The influence of this new similarity measure on the optimization of neural network have been thoroughly analyzed. To achieve a better performance, we develop a special back-propagation approach for AdderNets by investigating the full-precision gradient. We then propose an adaptive learning rate strategy to enhance the training procedure of AdderNets according to the magnitude of each neuron's gradient. As a result, the proposed AdderNets can achieve 74.9% Top-1 accuracy 91.7% Top-5 accuracy using ResNet-50 on the ImageNet dataset without any multiplication in convolutional layer. The codes are publicly available at: (https://github.com/huaweinoah/AdderNet). Hanting Chen, Yunhe Wang 0001, Chunjing Xu, Boxin Shi, Chao Xu 0006, Qi Tian 0001, Chang Xu 0002 |
CVPR | 5 |
| 2020 | On Positive-Unlabeled Classification in GANabstractThis paper defines a positive and unlabeled classification problem for standard GANs, which then leads to a novel technique to stabilize the training of the discriminator in GANs. Traditionally, real data are taken as positive while generated data are negative. This positive-negative classification criterion was kept fixed all through the learning process of the discriminator without considering the gradually improved quality of generated data, even if they could be more realistic than real data at times. In contrast, it is more reasonable to treat the generated data as unlabeled, which could be positive or negative according to their quality. The discriminator is thus a classifier for this positive and unlabeled classification problem, and we derive a new Positive-Unlabeled GAN (PUGAN). We theoretically discuss the global optimality the proposed model will achieve and the equivalent optimization goal. Empirically, we find that PUGAN can achieve comparable or even better performance than those sophisticated discriminator stabilization methods. Tianyu Guo 0001, Chang Xu 0002, Yunhe Wang 0001, Boxin Shi, Chao Xu 0006, Dacheng Tao |
CVPR | 6 |
| 2020 | Neuromorphic Camera Guided High Dynamic Range ImagingabstractReconstruction of high dynamic range image from a single low dynamic range image captured by a frame-based conventional camera, which suffers from over- or under-exposure, is an ill-posed problem. In contrast, recent neuromorphic cameras are able to record high dynamic range scenes in the form of an intensity map, with much lower spatial resolution, and without color. In this paper, we propose a neuromorphic camera guided high dynamic range imaging pipeline, and a network consisting of specially designed modules according to each step in the pipeline, which bridges the domain gaps on resolution, dynamic range, and color representation between two types of sensors and images. A hybrid camera system has been built to validate that the proposed method is able to reconstruct quantitatively and qualitatively high-quality high dynamic range images by successfully fusing the images and intensity maps for various real-world scenarios. Jin Han 0001, Chu Zhou, Peiqi Duan 0002, Yehui Tang 0001, Chang Xu 0002, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi |
CVPR | 6 |
| 2020 | A Semi-Supervised Assessor of Neural ArchitecturesabstractNeural architecture search (NAS) aims to automatically design deep neural networks of satisfactory performance. Wherein, architecture performance predictor is critical to efficiently value an intermediate neural architecture. But for the training of this predictor, a number of neural architectures and their corresponding real performance often have to be collected. In contrast with classical performance predictor optimized in a fully supervised way, this paper suggests a semi-supervised assessor of neural architectures. We employ an auto-encoder to discover meaningful representations of neural architectures. Taking each neural architecture as an individual instance in the search space, we construct a graph to capture their intrinsic similarities, where both labeled and unlabeled architectures are involved. A graph convolutional neural network is introduced to predict the performance of architectures based on the learned representations and their relation modeled by the graph. Extensive experimental results on the NAS-Benchmark-101 dataset demonstrated that our method is able to make a significant reduction on the required fully trained architectures for finding efficient architectures. Yehui Tang 0001, Yunhe Wang 0001, Yixing Xu, Hanting Chen, Boxin Shi, Chao Xu 0006, Chunjing Xu, Qi Tian 0001, Chang Xu 0002 |
CVPR | 6 |
| 2020 | CARS: Continuous Evolution for Efficient Neural Architecture SearchabstractSearching techniques in most of existing neural architecture search (NAS) algorithms are mainly dominated by differentiable methods for the efficiency reason. In contrast, we develop an efficient continuous evolutionary approach for searching neural networks. Architectures in the population that share parameters within one SuperNet in the latest generation will be tuned over the training dataset with a few epochs. The searching in the next evolution generation will directly inherit both the SuperNet and the population, which accelerates the optimal network generation. The non-dominated sorting strategy is further applied to preserve only results on the Pareto front for accurately updating the SuperNet. Several neural networks with different model sizes and performances will be produced after the continuous search with only 0.4 GPU days. As a result, our framework provides a series of networks with the number of parameters ranging from 3.7M to 5.1M under mobile settings. These networks surpass those produced by the state-of-the-art methods on the benchmark ImageNet dataset. Zhaohui Yang 0003, Yunhe Wang 0001, Xinghao Chen 0001, Boxin Shi, Chao Xu 0006, Chunjing Xu, Qi Tian 0001, Chang Xu 0002 |
CVPR | 5 |
| 2020 | Discernible Image CompressionabstractImage compression, as one of the fundamental low-level image processing tasks, is very essential for computer vision. Tremendous computing and storage resources can be preserved with a trivial amount of visual information. Conventional image compression methods tend to obtain compressed images by minimizing their appearance discrepancy with the corresponding original images, but pay little attention to their efficacy in downstream perception tasks, e.g., image recognition and object detection. Thus, some of compressed images could be recognized with bias. In contrast, this paper aims to produce compressed images by pursuing both appearance and perceptual consistency. Based on the encoder-decoder framework, we propose using a pre-trained CNN to extract features of the original and compressed images, and making them similar. Thus the compressed images are discernible to subsequent tasks, and we name our method as Discernible Image Compression (DIC). In addition, the maximum mean discrepancy (MMD) is employed to minimize the difference between feature distributions. The resulting compression network can generate images with high image quality and preserve the consistent perception in the feature domain, so that these images can be well recognized by pre-trained machine learning models. Experiments on benchmarks demonstrate that images compressed by using the proposed method can also be well recognized by subsequent visual recognition and detection models. For instance, the mAP value of compressed images by DIC is about 0.6% higher than that of using compressed images by conventional methods. Zhaohui Yang 0003, Yunhe Wang 0001, Chang Xu 0002, Chao Xu 0006, Chunjing Xu, Qi Tian 0001 |
ACM Multimedia | 5 |
| 2020 | SCOP: Scientific Control for Reliable Neural Network PruningabstractThis paper proposes a reliable neural network pruning algorithm by setting up a scientific control. Existing pruning methods have developed various hypotheses to approximate the importance of filters to the network and then execute filter pruning accordingly. To increase the reliability of the results, we prefer to have a more rigorous research design by including a scientific control group as an essential part to minimize the effect of all factors except the association between the filter and expected network output. Acting as a control group, knockoff feature is generated to mimic the feature map produced by the network filter, but they are conditionally independent of the example label given the real feature map. We theoretically suggest that the knockoff condition can be approximately preserved given the information propagation of network layers. Besides the real feature map on an intermediate layer, the corresponding knockoff feature is brought in as another auxiliary input signal for the subsequent layers. Redundant filters can be discovered in the adversarial process of different features. Through experiments, we demonstrate the superiority of the proposed algorithm over state-of-the-art methods. For example, our method can reduce 57.8% parameters and 60.2% FLOPs of ResNet-101 with only 0.01% top-1 accuracy loss on ImageNet. Yehui Tang 0001, Yunhe Wang 0001, Yixing Xu, Dacheng Tao, Chunjing Xu, Chao Xu 0006, Chang Xu 0002 |
NeurIPS | 6 |
| 2020 | Searching for Low-Bit Weights in Quantized Neural NetworksabstractQuantized neural networks with low-bit weights and activations are attractive for developing AI accelerators. However, the quantization functions used in most conventional quantization methods are non-differentiable, which increases the optimization difficulty of quantized networks. Compared with full-precision parameters (\emph{i.e.}, 32-bit floating numbers), low-bit values are selected from a much smaller set. For example, there are only 16 possibilities in 4-bit space. Thus, we present to regard the discrete weights in an arbitrary quantized neural network as searchable variables, and utilize a differential method to search them accurately. In particular, each weight is represented as a probability distribution over the discrete value set. The probabilities are optimized during training and the values with the highest probability are selected to establish the desired quantized network. Experimental results on benchmarks demonstrate that the proposed method is able to produce quantized neural networks with higher performance over the state-of-the-arts on both image classification and super-resolution tasks. Zhaohui Yang 0003, Yunhe Wang 0001, Kai Han 0002, Chunjing Xu, Chao Xu 0006, Dacheng Tao, Chang Xu 0002 |
NeurIPS | 5 |
| 2020 | UnModNet: Learning to Unwrap a Modulo Image for High Dynamic Range ImagingabstractA conventional camera often suffers from over- or under-exposure when recording a real-world scene with a very high dynamic range (HDR). In contrast, a modulo camera with a Markov random field (MRF) based unwrapping algorithm can theoretically accomplish unbounded dynamic range but shows degenerate performances when there are modulus-intensity ambiguity, strong local contrast, and color misalignment. In this paper, we reformulate the modulo image unwrapping problem into a series of binary labeling problems and propose a modulo edge-aware model, named as UnModNet, to iteratively estimate the binary rollover masks of the modulo image for unwrapping. Experimental results show that our approach can generate 12-bit HDR images from 8-bit modulo images reliably, and runs much faster than the previous MRF-based algorithm thanks to the GPU acceleration. Chu Zhou, Hang Zhao 0021, Jin Han 0001, Chang Xu 0002, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi |
NeurIPS | 5 |
| 2020 | Robust Student Network LearningabstractDeep neural networks bring in impressive accuracy in various applications, but the success often relies on heavy network architectures. Taking well-trained heavy networks as teachers, classical teacher-student learning paradigm aims to learn a student network that is lightweight yet accurate. In this way, a portable student network with significantly fewer parameters can achieve considerable accuracy, which is comparable to that of a teacher network. However, beyond accuracy, the robustness of the learned student network against perturbation is also essential for practical uses. Existing teacher-student learning frameworks mainly focus on accuracy and compression ratios, but ignore the robustness. In this paper, we make the student network produce more confident predictions with the help of the teacher network, and analyze the lower bound of the perturbation that will destroy the confidence of the student network. Two important objectives regarding prediction scores and gradients of examples are developed to maximize this lower bound, to enhance the robustness of the student network without sacrificing the performance. Experiments on benchmark data sets demonstrate the efficiency of the proposed approach to learning robust student networks that have satisfying accuracy and compact sizes. Tianyu Guo 0001, Chang Xu 0002, Shiyi He, Boxin Shi, Chao Xu 0006, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Smooth Deep Image Generator from NoisesabstractGenerative Adversarial Networks (GANs) have demonstrated a strong ability to fit complex distributions since they were presented, especially in the field of generating natural images. Linear interpolation in the noise space produces a continuously changing in the image space, which is an impressive property of GANs. However, there is no special consideration on this property in the objective function of GANs or its derived models. This paper analyzes the perturbation on the input of the generator and its influence on the generated images. A smooth generator is then developed by investigating the tolerable input perturbation. We further integrate this smooth generator with a gradient penalized discriminator, and design smooth GAN that generates stable and high-quality images. Experiments on real-world image datasets demonstrate the necessity of studying smooth generator and the effectiveness of the proposed algorithm. Tianyu Guo 0001, Chang Xu 0002, Boxin Shi, Chao Xu 0006, Dacheng Tao |
AAAI | 4 |
| 2019 | Revisiting Perspective Information for Efficient Crowd CountingabstractCrowd counting is the task of estimating people numbers in crowd images. Modern crowd counting methods employ deep neural networks to estimate crowd counts via crowd density regressions. A major challenge of this task lies in the perspective distortion, which results in drastic person scale change in an image. Density regression on the small person area is in general very hard. In this work, we propose a perspective-aware convolutional neural network (PACNN) for efficient crowd counting, which integrates the perspective information into density regression to provide additional knowledge of the person scale change in an image. Ground truth perspective maps are firstly generated for training; PACNN is then specifically designed to predict multi-scale perspective maps and encode them as perspective-aware weighting layers in the network to adaptively combine the outputs of multi-scale density maps. The weights are learned at every pixel of the maps such that the final density combination is robust to the perspective distortion. We conduct extensive experiments on the ShanghaiTech, WorldExpo'10, UCF_CC_50, and UCSD datasets, and demonstrate the effectiveness and efficiency of PACNN over the state-of-the-art. Miaojing Shi, Zhaohui Yang 0003, Chao Xu 0006 |
CVPR | 3 |
| 2019 | Data-Free Learning of Student NetworksabstractLearning portable neural networks is very essential for computer vision for the purpose that pre-trained heavy deep models can be well applied on edge devices such as mobile phones and micro sensors. Most existing deep neural network compression and speed-up methods are very effective for training compact deep models, when we can directly access the training dataset. However, training data for the given deep network are often unavailable due to some practice problems (\eg privacy, legal issue, and transmission), and the architecture of the given network are also unknown except some interfaces. To this end, we propose a novel framework for training efficient deep neural networks by exploiting generative adversarial networks (GANs). To be specific, the pre-trained teacher networks are regarded as a fixed discriminator and the generator is utilized for derivating training samples which can obtain the maximum response on the discriminator. Then, an efficient network with smaller model size and computational complexity is trained using the generated data and the teacher network, simultaneously. Efficient student networks learned using the proposed Data-Free Learning (DFL) method achieve 92.22% and 74.47% accuracies without any training data on the CIFAR-10 and CIFAR-100 datasets, respectively. Meanwhile, our student network obtains an 80.56% accuracy on the CelebA benchmark. Hanting Chen, Yunhe Wang 0001, Chang Xu 0002, Zhaohui Yang 0003, Chuanjian Liu, Boxin Shi, Chunjing Xu, Chao Xu 0006, Qi Tian 0001 |
ICCV | 8 |
| 2019 | LegoNet: Efficient Convolutional Neural Networks with Lego FiltersabstractThis paper aims to build efficient convolutional neural networks using a set of Lego filters. Many successful building blocks, e.g., inception and residual modules, have been designed to refresh state-of-the-art records of CNNs on visual recognition tasks. Beyond these high-level modules, we suggest that an ordinary filter in the neural network can be upgraded to a sophisticated module as well. Filter modules are established by assembling a shared set of Lego filters that are often of much lower dimensions. Weights in Lego filters and binary masks to stack Lego filters for these filter modules can be simultaneously optimized in an end-to-end manner as usual. Inspired by network engineering, we develop a split-transform-merge strategy for an efficient convolution by exploiting intermediate Lego feature maps. The compression and acceleration achieved by Lego Networks using the proposed Lego filters have been theoretically discussed. Experimental results on benchmark datasets and deep models demonstrate the advantages of the proposed Lego filters and their potential real-world applications on mobile devices. Zhaohui Yang 0003, Yunhe Wang 0001, Chuanjian Liu, Hanting Chen, Chunjing Xu, Boxin Shi, Chao Xu 0006, Chang Xu 0002 |
ICML | 7 |
| 2019 | Learning from Bad Data via GenerationabstractBad training data would challenge the learning model from understanding the underlying data-generating scheme, which then increases the difficulty in achieving satisfactory performance on unseen test data. We suppose the real data distribution lies in a distribution set supported by the empirical distribution of bad data. A worst-case formulation can be developed over this distribution set, and then be interpreted as a generation task in an adversarial manner. The connections and differences between GANs and our framework have been thoroughly discussed. We further theoretically show the influence of this generation task on learning from bad data and reveal its connection with a data-dependent regularization. Given different distance measures (\eg, Wasserstein distance or JS divergence) of distributions, we can derive different objective functions for the problem. Experimental results on different kinds of bad training data demonstrate the necessity and effectiveness of the proposed method. Tianyu Guo 0001, Chang Xu 0002, Boxin Shi, Chao Xu 0006, Dacheng Tao |
NeurIPS | 4 |
| 2019 | Packing Convolutional Neural Networks in the Frequency DomainabstractDeep convolutional neural networks (CNNs) are successfully used in a number of applications. However, their storage and computational requirements have largely prevented their widespread use on mobile devices. Here we present a series of approaches for compressing and speeding up CNNs in the frequency domain, which focuses not only on smaller weights but on all the weights and their underlying connections. By treating convolution filters as images, we decompose their representations in the frequency domain as common parts (i.e., cluster centers) shared by other similar filters and their individual private parts (i.e., individual residuals). A large number of low-energy frequency coefficients in both parts can be discarded to produce high compression without significantly compression romising accuracy. Furthermore, we explore a data-driven method for removing redundancies in both spatial and frequency domains, which allows us to discard more useless weights by keeping similar accuracies. After obtaining the optimal sparse CNN in the frequency domain, we relax the computational burden of convolution operations in CNNs by linearly combining the convolution responses of discrete cosine transform (DCT) bases. The compression and speed-up ratios of the proposed algorithm are thoroughly analyzed and evaluated on benchmark image datasets to demonstrate its superiority over state-of-the-art methods. Yunhe Wang 0001, Chang Xu 0002, Chao Xu 0006, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Reinforced Multi-Label Image Classification by Exploring CurriculumabstractHumans and animals learn much better when the examples are not randomly presented but organized in a meaningful order which illustrates gradually more concepts, and gradually more complex ones. Inspired by this curriculum learning mechanism, we propose a reinforced multi-label image classification approach imitating human behavior to label image from easy to complex. This approach allows a reinforcement learning agent to sequentially predict labels by fully exploiting image feature and previously predicted labels. The agent discovers the optimal policies through maximizing the long-term reward which reflects prediction accuracies. Experimental results on PASCAL VOC2007 and 2012 demonstrate the necessity of reinforcement multi-label learning and the algorithm’s effectiveness in real-world multi-label image classification tasks. Shiyi He, Chang Xu 0002, Tianyu Guo 0001, Chao Xu 0006, Dacheng Tao |
AAAI | 4 |
| 2018 | Adversarial Learning of Portable Student NetworksabstractEffective methods for learning deep neural networks with fewer parameters are urgently required, since storage and computations of heavy neural networks have largely prevented their widespread use on mobile devices. Compared with algorithms which directly remove weights or filters for obtaining considerable compression and speed-up ratios, training thin deep networks exploiting the student-teacher learning paradigm is more flexible. However, it is very hard to determine which formulation is optimal to measure the information inherited from teacher networks. To overcome this challenge, we utilize the generative adversarial network (GAN) to learn the student network. In practice, the generator is exactly the student network with extremely less parameters and the discriminator is used as a teaching assistant for distinguishing features extracted from student and teacher networks. By simultaneously optimizing the generator and the discriminator, the resulting student network can produce features of input data with the similar distribution as that of features of the teacher network. Extensive experimental results on benchmark datasets demonstrate that the proposed method is capable of learning well-performed portable networks, which is superior to the state-of-the-art methods. Yunhe Wang 0001, Chang Xu 0002, Chao Xu 0006, Dacheng Tao |
AAAI | 3 |
| 2018 | Learning With Single-Teacher Multi-StudentabstractIn this paper we study a new learning problem defined as "Single-Teacher Multi-Student" (STMS) problem, which investigates how to learn a series of student (simple and specific) models from a single teacher (complex and universal) model. Taking the multiclass and binary classification for example, we focus on learning multiple binary classifiers from a single multiclass classifier, where each of binary classifier is responsible for a certain class. This actually derives from some realistic problems, such as identifying the suspect based on a comprehensive face recognition system. By treating the already-trained multiclass classifier as the teacher, and multiple binary classifiers as the students, we propose a gated support vector machine (gSVM) as a solution. A series of gSVMs are learned with the help of single teacher multiclass classifier. The teacher's help is two-fold; first, the teacher's score provides the gated values for students' decision; second, the teacher can guide the students to accommodate training examples with different difficulty degrees. Extensive experiments on real datasets validate its effectiveness. Shan You, Chang Xu 0002, Chao Xu 0006, Dacheng Tao |
AAAI | 3 |
| 2018 | Autoencoder Inspired Unsupervised Feature SelectionabstractHigh-dimensional data in many areas such as computer vision and machine learning tasks brings in computational and analytical difficulty. Feature selection which selects a subset from observed features is a widely used approach for improving performance and effectiveness of machine learning models with high-dimensional data. In this paper, we propose a novel AutoEncoder Feature Selector (AEFS) for unsupervised feature selection which combines autoencoder regression and group lasso tasks. Compared to traditional feature selection methods, AEFS can select the most important features by excavating both linear and nonlinear information among features, which is more flexible than the conventional self-representation method for unsupervised feature selection with only linear assumptions. Experimental results on benchmark dataset show that the proposed method is superior to the state-of-the-art method. Kai Han 0002, Yunhe Wang 0001, Chao Zhang 0001, Chao Xu 0006 |
ICASSP | 5 |
| 2018 | Online Dictionary Learning with ConfidenceabstractOnline dictionary learning has received intensive attention in signal processing field with streaming or dynamic data. Different from classical online dictionary learning methods that treat all atoms equally, in this paper, we present a novel online dictionary learning with a confidence parameter introduced on each of atoms. The confidence indicates the scale of the update of atoms during online learning; frequently estimated atoms are thus supposed to be updated less aggressively than those of low usage frequency. This updating mechanism is beneficial for learning with dirty examples. As a result, the outliers would be prevented from influencing much on the frequently used atoms that have been well estimated through the clean examples. In detail, we employ variance of the atoms to measure the confidence on the quality of the learned atoms. And frequently-used atoms are supposed to have smaller variances (i.e. more confidence) than those of low usage frequency. Our algorithm consists of two main parts, namely, confidence-weighted sparse coding and dictionary update with confidence fine-tuning, each of which can be solved efficiently by our designed optimization methods. In addition, due to the introduced confidence, our algorithm does not have to depend on the widely-used ℓ1norm in order to realize the robustness against the outliers, which is usually computationally expensive in the training cost. Experimental results on synthetic and benchmark datasets demonstrate the imposed confidence on atoms in the dictionary can improve the performance of the learned dictionary. The quality of the obtained dictionary is comparable to that of state-of-the-art robust online dictionary learning methods, however, our method is much faster than those ℓ1norm based ones. Shan You, Chang Xu 0002, Chao Xu 0006 |
ICDM | 3 |
| 2018 | Towards Evolutionary CompressionabstractCompressing convolutional neural networks (CNNs) is essential for transferring the success of CNNs to a wide variety of applications to mobile devices. In contrast to directly recognizing subtle weights or filters as redundant in a given CNN, this paper presents an evolutionary method to automatically eliminate redundant convolution filters. We represent each compressed network as a binary individual of specific fitness. Then, the population is upgraded at each evolutionary iteration using genetic operations. As a result, an extremely compact CNN is generated using the fittest individual, which has the original network structure and can be directly deployed in any off-the-shelf deep learning libraries. In this approach, either large or small convolution filters can be redundant, and filters in the compressed network are more distinct. In addition, since the number of filters in each convolutional layer is reduced, the number of filter channels and the size of feature maps are also decreased, naturally improving both the compression and speed-up ratios. Experiments on benchmark deep CNN models suggest the superiority of the proposed algorithm over the state-of-the-art compression methods, e.g. combined with the parameter refining approach, we can reduce the storage requirement and the floating-point multiplications of ResNet-50 by a factor of 14.64x and 5.19x, respectively, without affecting its accuracy. Yunhe Wang 0001, Chang Xu 0002, Jiayan Qiu, Chao Xu 0006, Dacheng Tao |
KDD | 4 |
| 2018 | Learning Versatile Filters for Efficient Convolutional Neural NetworksabstractThis paper introduces versatile filters to construct efficient convolutional neural network. Considering the demands of efficient deep learning techniques running on cost-effective hardware, a number of methods have been developed to learn compact neural networks. Most of these works aim to slim down filters in different ways, e.g., investigating small, sparse or binarized filters. In contrast, we treat filters from an additive perspective. A series of secondary filters can be derived from a primary filter. These secondary filters all inherit in the primary filter without occupying more storage, but once been unfolded in computation they could significantly enhance the capability of the filter by integrating information extracted from different receptive fields. Besides spatial versatile filters, we additionally investigate versatile filters from the channel perspective. The new techniques are general to upgrade filters in existing CNNs. Experimental results on benchmark datasets and neural networks demonstrate that CNNs constructed with our versatile filters are able to achieve comparable accuracy as that of original filters, but require less memory and FLOPs. Yunhe Wang 0001, Chang Xu 0002, Chunjing Xu, Chao Xu 0006, Dacheng Tao |
NeurIPS | 4 |
| 2018 | Cost-Sensitive Feature Selection by Optimizing F-MeasuresabstractFeature selection is beneficial for improving the performance of general machine learning tasks by extracting an informative subset from the high-dimensional features. Conventional feature selection methods usually ignore the class imbalance problem, thus the selected features will be biased towards the majority class. Considering that F-measure is a more reasonable performance measure than accuracy for imbalanced data, this paper presents an effective feature selection algorithm that explores the class imbalance issue by optimizing F-measures. Since F-measure optimization can be decomposed into a series of cost-sensitive classification problems, we investigate the cost-sensitive feature selection by generating and assigning different costs to each class with rigorous theory guidance. After solving a series of cost-sensitive feature selection problems, features corresponding to the best F-measure will be selected. In this way, the selected features will fully represent the properties of all classes. Experimental results on popular benchmarks and challenging real-world data sets demonstrate the significance of cost-sensitive feature selection for the imbalanced data setting and validate the effectiveness of the proposed method. Meng Liu 0003, Chang Xu 0002, Yong Luo 0002, Chao Xu 0006, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2017 | Cost-Sensitive Feature Selection via F-Measure Optimization ReductionabstractFeature selection aims to select a small subset from the high-dimensional features which can lead to better learning performance, lower computational complexity, and better model readability. The class imbalance problem has been neglected by traditional feature selection methods, therefore the selected features will be biased towards the majority classes. Because of the superiority of F-measure to accuracy for imbalanced data, we propose to use F-measure as the performance measure for feature selection algorithms. As a pseudo-linear function, the optimization of F-measure can be achieved by minimizing the total costs. In this paper, we present a novel cost-sensitive feature selection (CSFS) method which optimizes F-measure instead of accuracy to take class imbalance issue into account. The features will be selected according to optimal F-measure classifier after solving a series of cost-sensitive feature selection sub-problems. The features selected by our method will fully represent the characteristics of not only majority classes, but also minority classes. Extensive experimental results conducted on synthetic, multi-class and multi-label datasets validate the efficiency and significance of our feature selection method. Meng Liu 0003, Chang Xu 0002, Yong Luo 0002, Chao Xu 0006, Yonggang Wen 0001, Dacheng Tao |
AAAI | 4 |
| 2017 | Beyond RPCA: Flattening Complex Noise in the Frequency DomainabstractDiscovering robust low-rank data representations is important in many real-world problems. Traditional robust principal component analysis (RPCA) assumes that the observed data are corrupted by some sparse noise (e.g., Laplacian noise) and utilizes the l1-norm to separate out the noisy compo- nent. Nevertheless, as well as simple Gaussian or Laplacian noise, noise in real-world data is often more complex, and thus the l1 and l2-norms are insufficient for noise charac- terization. This paper presents a more flexible approach to modeling complex noise by investigating their properties in the frequency domain. Although elements of a noise matrix are chaotic in the spatial domain, the absolute values of its alternative coefficients in the frequency domain are constant w.r.t. their variance. Based on this observation, a new robust PCA algorithm is formulated by simultaneously discovering the low-rank and noisy components. Extensive experiments on synthetic data and video background subtraction demon- strate that FRPCA is effective for handles complex noise. Yunhe Wang 0001, Chang Xu 0002, Chao Xu 0006, Dacheng Tao |
AAAI | 3 |
| 2017 | Fast Compressive Phase Retrieval under Bounded NoiseabstractWe study the problem of recovering a t-sparse real vector from m quadratic equations yi=(ai*x)^2 with noisy measurements yi's. This is known as the problem of compressive phase retrieval, and has been widely applied to X-ray diffraction imaging, microscopy, quantum mechanics, etc. The challenge is to design a a) fast and b) noise-tolerant algorithms with c) near-optimal sample complexity. Prior work in this direction typically achieved one or two of these goals, but none of them enjoyed the three performance guarantees simultaneously. In this work, with a particular set of sensing vectors ai's, we give a provable algorithm that is robust to any bounded yet unstructured deterministic noise. Our algorithm requires roughly O(t) measurements and runs in O(tn*log (1/epsilon)) time, where epsilon is the error. This result advances the state-of-the-art work, and guarantees the applicability of our method to large datasets. Experiments on synthetic and real data verify our theory. Hongyang Zhang 0001, Shan You, Zhouchen Lin, Chao Xu 0006 |
AAAI | 4 |
| 2017 | Beyond Filters: Compact Feature Map for Portable Deep ModelabstractConvolutional neural networks (CNNs) have shown extraordinary performance in a number of applications, but they are usually of heavy design for the accuracy reason. Beyond compressing the filters in CNNs, this paper focuses on the redundancy in the feature maps derived from the large number of filters in a layer. We propose to extract intrinsic representation of the feature maps and preserve the discriminability of the features. Circulant matrix is employed to formulate the feature map transformation, which only requires O(dlog d) computation complexity to embed a d-dimensional feature map. The filter is then re-configured to establish the mapping from original input to the new compact feature map, and the resulting network can preserve intrinsic information of the original network with significantly fewer parameters, which not only decreases the online memory for launching CNN but also accelerates the computation speed. Experiments on benchmark image datasets demonstrate the superiority of the proposed algorithm over state-of-the-art methods. Yunhe Wang 0001, Chang Xu 0002, Chao Xu 0006, Dacheng Tao |
ICML | 3 |
| 2017 | Fast SVM Trained by Divide-and-Conquer AnchorsabstractSupporting vector machine (SVM) is the most frequently used classifier for machine learning tasks. However, its training time could become cumbersome when the size of training data is very large. Thus, many kinds of representative subsets are chosen from the original dataset to reduce the training complexity. In this paper, we propose to choose the representative points which are noted as anchors obtained from non-negative matrix factorization (NMF) in a divide-and-conquer framework, and then use the anchors to train an approximate SVM. Our theoretical analysis shows that the solving the DCA-SVM can yield an approximate solution close to the primal SVM. Experimental results on multiple datasets demonstrate that our DCA-SVM is faster than the state-of-the-art algorithms without notably decreasing the accuracy of classification results. Meng Liu 0003, Chang Xu 0002, Chao Xu 0006, Dacheng Tao |
IJCAI | 3 |
| 2017 | Multi-Positive and Unlabeled LearningabstractThe positive and unlabeled (PU) learning problem focuses on learning a classifier from positive and unlabeled data. Some methods have been developed to solve the PU learning problem. However, they are often limited in practical applications, since only binary classes are involved and cannot easily be adapted to multi-class data. Here we propose a one-step method that directly enables multi-class model to be trained using the given input multi-class data and that predicts the label based on the model decision. Specifically, we construct different convex loss functions for labeled and unlabeled data to learn a discriminant function F. The theoretical analysis on the generalization error bound shows that it is no worse than k√k times of the fully supervised multi-class classification methods when the size of the data in k classes is of the same order. Finally, our experimental results demonstrate the significance and effectiveness of the proposed algorithm in synthetic and real-world datasets. Yixing Xu, Chang Xu 0002, Chao Xu 0006, Dacheng Tao |
IJCAI | 3 |
| 2017 | Privileged Multi-label LearningabstractThis paper presents privileged multi-label learning (PrML) to explore and exploit the relationship between labels in multi-label learning problems. We suggest that for each individual label, it cannot only be implicitly connected with other labels via the low-rank constraint over label predictors, but also its performance on examples can receive the explicit comments from other labels together acting as an Oracle teacher. We generate privileged label feature for each example and its individual label, and then integrate it into the framework of low-rank based multi-label learning. The proposed algorithm can therefore comprehensively explore and exploit label relationships by inheriting all the merits of privileged information and low-rank constraints. We show that PrML can be efficiently solved by dual coordinate descent algorithm using iterative optimization strategy with cheap updates. Experiments on benchmark datasets show that through privileged label features, the performance can be significantly improved and PrML is superior to several competing methods in most cases. Shan You, Chang Xu 0002, Yunhe Wang 0001, Chao Xu 0006, Dacheng Tao |
IJCAI | 4 |
| 2017 | Learning from Multiple Teacher NetworksabstractTraining thin deep networks following the student-teacher learning paradigm has received intensive attention because of its excellent performance. However, to the best of our knowledge, most existing work mainly considers one single teacher network. In practice, a student may access multiple teachers, and multiple teacher networks together provide comprehensive guidance that is beneficial for training the student network. In this paper, we present a method to train a thin deep network by incorporating multiple teacher networks not only in output layer by averaging the softened outputs (dark knowledge) from different networks, but also in the intermediate layers by imposing a constraint about the dissimilarity among examples. We suggest that the relative dissimilarity between intermediate representations of different examples serves as a more flexible and appropriate guidance from teacher networks. Then triplets are utilized to encourage the consistence of these relative dissimilarity relationships between the student network and teacher networks. Moreover, we leverage a voting strategy to unify multiple relative dissimilarity information provided by multiple teacher networks, which realizes their incorporation in the intermediate layers. Extensive experimental results demonstrated that our method is capable of generating a well-performed student network, with the classification accuracy comparable or even superior to all teacher networks, yet having much fewer parameters and being much faster in running. Shan You, Chang Xu 0002, Chao Xu 0006, Dacheng Tao |
KDD | 3 |
| 2017 | How do you smile? Towards a comprehensive smile analysis system
Hong Liu 0008, Chao Xu 0006, Yuan Gao 0008, Xuewu Zhang 0003 |
Neurocomputing | 3 |
| 2017 | DCT Regularized Extreme Visual RecoveryabstractHere we study the extreme visual recovery problem, in which over 90% of pixel values in a given image are missing. Existing low rank-based algorithms are only effective for recovering data with at most 90% missing values. Thus, we exploit visual data's smoothness property to help solve this challenging extreme visual recovery problem. Based on the discrete cosine transform (DCT), we propose a novel DCT regularizer that involves all pixels and produces smooth estimations in any view. Our theoretical analysis shows that the total variation regularizer, which only achieves local smoothness, is a special case of the proposed DCT regularizer. We also develop a new visual recovery algorithm by minimizing the DCT regularizer and nuclear norm to achieve a more visually pleasing estimation. Experimental results on a benchmark image data set demonstrate that the proposed approach is superior to the state-of-the-art methods in terms of peak signal-to-noise ratio and structural similarity. Yunhe Wang 0001, Chang Xu 0002, Shan You, Chao Xu 0006, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2016 | Tensor canonical correlation analysis for multi-view dimension reductionabstractCanonical correlation analysis (CCA) has proven an effective tool for two-view dimension reduction due to its profound theoretical foundation and success in practical applications. In respect of multi-view learning, however, it is limited by its capability of only handling data represented by two-view features, while in many real-world applications, the number of views is frequently many more. Although the ad hoc way of simultaneously exploring all possible pairs of features can numerically deal with multi-view data, it ignores the high order statistics (correlation information) which can only be discovered by simultaneously exploring all features. Therefore, in this work, we develop tensor CCA (TCCA) which straightforwardly yet naturally generalizes CCA to handle the data of an arbitrary number of views by analyzing the covariance tensor of the different views. TCCA aims to directly maximize the canonical correlation of multiple (more than two) views. Crucially, we prove that the main problem of multiview canonical correlation maximization is equivalent to finding the best rank-1 approximation of the data covariance tensor, which can be solved efficiently using the well-known alternating least squares (ALS) algorithm. As a consequence, the high order correlation information contained in the different views is explored and thus a more reliable common subspace shared by all features can be obtained. Yong Luo 0002, Dacheng Tao, Kotagiri Ramamohanarao, Chao Xu 0006, Yonggang Wen 0001 |
ICDE | 4 |
| 2016 | Robust Extreme Multi-label LearningabstractTail labels in the multi-label learning problem undermine the low-rank assumption. Nevertheless, this problem has rarely been investigated. In addition to using the low-rank structure to depict label correlations, this paper explores and exploits an additional sparse component to handle tail labels behaving as outliers, in order to make the classical low-rank principle in multi-label learning valid. The divide-and-conquer optimization technique is employed to increase the scalability of the proposed algorithm while theoretically guaranteeing its performance. A theoretical analysis of the generalizability of the proposed algorithm suggests that it can be improved by the low-rank and sparse decomposition given tail labels. Experimental results on real-world data demonstrate the significance of investigating tail labels and the effectiveness of the proposed algorithm. Chang Xu 0002, Dacheng Tao, Chao Xu 0006 |
KDD | 3 |
| 2016 | CNNpack: Packing Convolutional Neural Networks in the Frequency DomainabstractDeep convolutional neural networks (CNNs) are successfully used in a number of applications. However, their storage and computational requirements have largely prevented their widespread use on mobile devices. Here we present an effective CNN compression approach in the frequency domain, which focuses not only on smaller weights but on all the weights and their underlying connections. By treating convolutional filters as images, we decompose their representations in the frequency domain as common parts (i.e., cluster centers) shared by other similar filters and their individual private parts (i.e., individual residuals). A large number of low-energy frequency coefficients in both parts can be discarded to produce high compression without significantly compromising accuracy. We relax the computational burden of convolution operations in CNNs by linearly combining the convolution responses of discrete cosine transform (DCT) bases. The compression and speed-up ratios of the proposed algorithm are thoroughly analyzed and evaluated on benchmark image datasets to demonstrate its superiority over state-of-the-art methods. Yunhe Wang 0001, Chang Xu 0002, Shan You, Dacheng Tao, Chao Xu 0006 |
NIPS | 5 |
| 2016 | Large Margin Multi-Modal Multi-Task Feature Extraction for Image ClassificationabstractThe features used in many image analysis-based applications are frequently of very high dimension. Feature extraction offers several advantages in high-dimensional cases, and many recent studies have used multi-task feature extraction approaches, which often outperform single-task feature extraction approaches. However, most of these methods are limited in that they only consider data represented by a single type of feature, even though features usually represent images from multiple modalities. We, therefore, propose a novel large margin multi-modal multi-task feature extraction (LM3FE) framework for handling multi-modal features for image classification. In particular, LM3FE simultaneously learns the feature extraction matrix for each modality and the modality combination coefficients. In this way, LM3FE not only handles correlated and noisy features, but also utilizes the complementarity of different modalities to further help reduce feature redundancy in each modality. The large margin principle employed also helps to extract strongly predictive features, so that they are more suitable for prediction (e.g., classification). An alternating algorithm is developed for problem optimization, and each subproblem can be efficiently solved. Experiments on two challenging real-world image data sets demonstrate the effectiveness and superiority of the proposed method. Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao, Jie Gui, Chao Xu 0006 |
IEEE Trans. Image Process. | 5 |
| 2016 | DCT Inspired Feature Transform for Image Retrieval and ReconstructionabstractScale invariant feature transform (SIFT) is effective for representing images in computer vision tasks, as one of the most resistant feature descriptions to common image deformations. However, two issues should be addressed: first, feature description based on gradient accumulation is not compact and contains redundancies; second, multiple orientations are often extracted from one local region and therefore produce multiple descriptions, which is not good for memory efficiency. To resolve these two issues, this paper introduces a novel method to determine the dominant orientation for multiple-orientation cases, named discrete cosine transform (DCT) intrinsic orientation, and a new DCT inspired feature transform (DIFT). In each local region, it first computes a unique DCT intrinsic orientation via DCT matrix and rotates the region accordingly, and then describes the rotated region with partial DCT matrix coefficients to produce an optimized low-dimensional descriptor. We test the accuracy and robustness of DIFT on real image matching. Afterward, extensive applications performed on public benchmarks for visual retrieval show that using DCT intrinsic orientation achieves performance on a par with SIFT, but with only 60% of its features; replacing the SIFT description with DIFT reduces dimensions from 128 to 32 and improves precision. Image reconstruction resulting from DIFT is presented to show another of its advantages over SIFT. Yunhe Wang 0001, Miaojing Shi, Shan You, Chao Xu 0006 |
IEEE Trans. Image Process. | 4 |
| 2016 | Local Rademacher Complexity for Multi-Label LearningabstractWe analyze the local Rademacher complexity of empirical risk minimization-based multi-label learning algorithms, and in doing so propose a new algorithm for multi-label learning. Rather than using the trace norm to regularize the multi-label predictor, we instead minimize the tail sum of the singular values of the predictor in multi-label learning. Benefiting from the use of the local Rademacher complexity, our algorithm, therefore, has a sharper generalization error bound. Compared with methods that minimize over all singular values, concentrating on the tail singular values results in better recovery of the low-rank structure of the multi-label predictor, which plays an important role in exploiting label correlations. We propose a new conditional singular value thresholding algorithm to solve the resulting objective function. Moreover, a variance control strategy is employed to reduce the variance of variables in optimization. Empirical studies on real-world data sets validate our theoretical results and demonstrate the effectiveness of the proposed algorithm for multi-label learning. Chang Xu 0002, Tongliang Liu, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Image Process. | 4 |
| 2015 | Low-Rank Multi-View Learning in Matrix Completion for Multi-Label Image ClassificationabstractMulti-label image classification is of significant interest due to its major role in real-world web image analysis applications such as large-scale image retrieval and browsing. Recently, matrix completion (MC) has been developed to deal with multi-label classification tasks. MC has distinct advantages, such as robustness to missing entries in the feature and label spaces and a natural ability to handle multi-label problems. However, current MC-based multi-label image classification methods only consider data represented by a single-view feature, therefore, do not precisely characterize images that contain several semantic concepts. An intuitive way to utilize multiple features taken from different views is to concatenate the different features into a long vector; however, this concatenation is prone to over-fitting and leads to high time complexity in MC-based image classification. Therefore, we present a novel multi-view learning model for MC-based image classification, called low-rank multi-view matrix completion (lrMMC), which first seeks a low-dimensional common representation of all views by utilizing the proposed low-rank multi-view learning (lrMVL) algorithm. In lrMVL, the common subspace is constrained to be low rank so that it is suitable for MC. In addition, combination weights are learned to explore complementarity between different views. An efficient solver based on fixed-point continuation (FPC) is developed for optimization, and the learned low-rank representation is then incorporated into MC-based image classification. Extensive experimentation on the challenging PASCAL VOC' 07 dataset demonstrates the superiority of lrMMC compared to other multi-label image classification approaches. Meng Liu 0003, Yong Luo 0002, Dacheng Tao, Chao Xu 0006, Yonggang Wen 0001 |
AAAI | 4 |
| 2015 | Large-Margin Multi-Label Causal Feature LearningabstractIn multi-label learning, an example is represented by a descriptive feature associated with several labels. Simply considering labels as independent or correlated is crude; it would be beneficial to define and exploit the causality between multiple labels. For example, an image label 'lake' implies the label 'water', but not vice versa. Since the original features are a disorderly mixture of the properties originating from different labels, it is intuitive to factorize these raw features to clearly represent each individual label and its causality relationship.Following the large-margin principle, we propose an effective approach to discover the causal features of multiple labels, thus revealing the causality between labels from the perspective of feature. We show theoretically that the proposed approach is a tight approximation of the empirical multi-label classification error, and the causality revealed strengthens the consistency of the algorithm. Extensive experimentations using synthetic and real-world data demonstrate that the proposed algorithm effectively discovers label causality, generates causal features, and improves multi-label learning. Chang Xu 0002, Dacheng Tao, Chao Xu 0006 |
AAAI | 3 |
| 2015 | Manifold Regularized Transfer Distance Metric LearningabstractThe performance of many computer vision and machine learning algorithms are heavily depend on the distance metric between samples. It is necessary to exploit abundant of side information like pairwise constraints to learn a robust and reliable distance metric[2, 3]. Let D = {(xl i ,xj,yi j)} l i, j=1 denotes the labeled training set for the target task, wherein xi, x j ∈ Rd and yi j = ±1 indicates xl i and xl i are similar/dissimilar to each other. Then, a metric is usually learned to minimize the distance between the data from the same class and maximize their distance otherwise. This leads to the following loss function for learning the metric A: Haibo Shi, Yong Luo 0002, Chao Xu 0006, Yonggang Wen 0001 |
BMVC | 3 |
| 2015 | Multi-view Self-Paced Learning for Clustering
Chang Xu 0002, Dacheng Tao, Chao Xu 0006 |
IJCAI | 3 |
| 2015 | Large-margin multi-view Gaussian process
Chang Xu 0002, Dacheng Tao, Yangxi Li, Chao Xu 0006 |
Multim. Syst. | 4 |
| 2015 | Multi-View Intact Space LearningabstractIt is practical to assume that an individual view is unlikely to be sufficient for effective multi-view learning. Therefore, integration of multi-view information is both valuable and necessary. In this paper, we propose the Multi-view Intact Space Learning (MISL) algorithm, which integrates the encoded complementary information in multiple views to discover a latent intact representation of the data. Even though each view on its own is insufficient, we show theoretically that by combing multiple views we can obtain abundant information for latent intact space learning. Employing the Cauchy loss (a technique used in statistical learning) as the error measurement strengthens robustness to outliers. We propose a new definition of multi-view stability and then derive the generalization error bound based on multi-view stability and Rademacher complexity, and show that the complementarity between multiple views is beneficial for the stability and generalization. MISL is efficiently optimized using a novel Iteratively Reweight Residuals (IRR) technique, whose convergence is theoretically analyzed. Experiments on synthetic data and real-world datasets demonstrate that MISL is an effective and promising algorithm for practical applications. Chang Xu 0002, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Multiview Matrix Completion for Multilabel Image ClassificationabstractThere is growing interest in multilabel image classification due to its critical role in web-based image analytics-based applications, such as large-scale image retrieval and browsing. Matrix completion (MC) has recently been introduced as a method for transductive (semisupervised) multilabel classification, and has several distinct advantages, including robustness to missing data and background noise in both feature and label space. However, it is limited by only considering data represented by a single-view feature, which cannot precisely characterize images containing several semantic concepts. To utilize multiple features taken from different views, we have to concatenate the different features as a long vector. However, this concatenation is prone to over-fitting and often leads to very high time complexity in MC-based image classification. Therefore, we propose to weightedly combine the MC outputs of different views, and present the multiview MC (MVMC) framework for transductive multilabel image classification. To learn the view combination weights effectively, we apply a cross-validation strategy on the labeled set. In particular, MVMC splits the labeled set into two parts, and predicts the labels of one part using the known labels of the other part. The predicted labels are then used to learn the view combination coefficients. In the learning process, we adopt the average precision (AP) loss, which is particular suitable for multilabel image classification, since the ranking-based criteria are critical for evaluating a multilabel classification system. A least squares loss formulation is also presented for the sake of efficiency, and the robustness of the algorithm based on the AP loss compared with the other losses is investigated. Experimental evaluation on two real-world data sets (PASCAL VOC' 07 and MIR Flickr) demonstrate the effectiveness of MVMC for transductive (semisupervised) multilabel image classification, and show that MVMC can exploit complementary properties of different features and output-consistent labels for improved multilabel image classification. Yong Luo 0002, Tongliang Liu, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Image Process. | 4 |
| 2015 | Multi-View Learning With Incomplete ViewsabstractOne underlying assumption of the conventional multi-view learning algorithms is that all examples can be successfully observed on all the views. However, due to various failures or faults in collecting and pre-processing the data on different views, we are more likely to be faced with an incomplete-view setting, where an example could be missing its representation on one view (i.e., missing view) or could be only partially observed on that view (i.e., missing variables). Low-rank assumption used to be effective for recovering the random missing variables of features, but it is disabled by concentrated missing variables and has no effect on missing views. This paper suggests that the key to handling the incomplete-view problem is to exploit the connections between multiple views, enabling the incomplete views to be restored with the help of the complete views. We propose an effective algorithm to accomplish multi-view learning with incomplete views by assuming that different views are generated from a shared subspace. To handle the large-scale problem and obtain fast convergence, we investigate a successive over-relaxation method to solve the objective function. Convergence of the optimization technique is theoretically analyzed. The experimental results on toy data and real-world data sets suggest that studying the incomplete-view problem in multi-view learning is significant and that the proposed algorithm can effectively handle the incomplete views in different applications. Chang Xu 0002, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Image Process. | 3 |
| 2015 | Exploring Spatial Correlation for Visual Object RetrievalabstractBag-of-visual-words (BOVW)-based image representation has received intense attention in recent years and has improved content-based image retrieval (CBIR) significantly. BOVW does not consider the spatial correlation between visual words in natural images and thus biases the generated visual words toward noise when the corresponding visual features are not stable. This article outlines the construction of a visual word co-occurrence matrix by exploring visual word co-occurrence extracted from small affine-invariant regions in a large collection of natural images. Based on this co-occurrence matrix, we first present a novel high-order predictor to accelerate the generation of spatially correlated visual words and a penalty tree (PTree) to continue generating the words after the prediction. Subsequently, we propose two methods of co-occurrence weighting similarity measure for image ranking: Co-Cosine and Co-TFIDF. These two new schemes down-weight the contributions of the words that are less discriminative because of frequent co-occurrences with other words. We conduct experiments on Oxford and Paris Building datasets, in which the ImageNet dataset is used to implement a large-scale evaluation. Cross-dataset evaluations between the Oxford and Paris datasets and Oxford and Holidays datasets are also provided. Thorough experimental results suggest that our method outperforms the state of the art without adding much additional cost to the BOVW model. Miaojing Shi, Xinghai Sun, Dacheng Tao, Chao Xu 0006, George Baciu, Hong Liu 0008 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2015 | Tensor Canonical Correlation Analysis for Multi-View Dimension ReductionabstractCanonical correlation analysis (CCA) has proven an effective tool for two-view dimension reduction due to its profound theoretical foundation and success in practical applications. In respect of multi-view learning, however, it is limited by its capability of only handling data represented by two-view features, while in many real-world applications, the number of views is frequently many more. Although the ad hoc way of simultaneously exploring all possible pairs of features can numerically deal with multi-view data, it ignores the high order statistics (correlation information) which can only be discovered by simultaneously exploring all features. Therefore, in this work, we develop tensor CCA (TCCA) which straightforwardly yet naturally generalizes CCA to handle the data of an arbitrary number of views by analyzing the covariance tensor of the different views. TCCA aims to directly maximize the canonical correlation of multiple (more than two) views. Crucially, we prove that the main problem of multi-view canonical correlation maximization is equivalent to finding the best rank-1 approximation of the data covariance tensor, which can be solved efficiently using the well-known alternating least squares (ALS) algorithm. As a consequence, the high order correlation information contained in the different views is explored and thus a more reliable common subspace shared by all features can be obtained. In addition, a non-linear extension of TCCA is presented. Experiments on various challenge tasks, including large scale biometric structure prediction, internet advertisement classification, and web image annotation, demonstrate the effectiveness of the proposed method. Yong Luo 0002, Dacheng Tao, Kotagiri Ramamohanarao, Chao Xu 0006, Yonggang Wen 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | Database Saliency for Fast Image RetrievalabstractThe bag-of-visual-words (BoW) model is effective for representing images and videos in many computer vision problems, and achieves promising performance in image retrieval. Nevertheless, the level of retrieval efficiency in a large-scale database is not acceptable for practical usage. Considering that the relevant images in the database of a given query are more likely to be distinctive than ambiguous, this paper defines “database saliency” as the distinctiveness score calculated for every image to measure its overall “saliency” in the database. By taking advantage of database saliency, we propose a saliency- inspired fast image retrieval scheme, S-sim, which significantly improves efficiency while retains state-of-the-art accuracy in image retrieval . There are two stages in S-sim: the bottom-up saliency mechanism computes the database saliency value of each image by hierarchically decomposing a posterior probability into local patches and visual words, the concurrent information of visual words is then bottom-up propagated to estimate the distinctiveness, and the top-down saliency mechanism discriminatively expands the query via a very low-dimensional linear SVM trained on the top-ranked images after initial search, ranking images are then sorted on their distances to the decision boundary as well as the database saliency values. We comprehensively evaluate S-sim on common retrieval benchmarks, e.g., Oxford and Paris datasets. Thorough experiments suggest that, because of the offline database saliency computation and online low-dimensional SVM, our approach significantly speeds up online retrieval and outperforms the state-of-the-art BoW-based image retrieval schemes. Yuan Gao 0008, Miaojing Shi, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Multim. | 4 |
| 2014 | Pre-Trained Multi-View Word Embedding Using Two-Side Neural NetworkabstractWord embedding aims to learn a continuous representation for each word. It attracts increasing attention due to its effectiveness in various tasks such as named entity recognition and language modeling. Most existing word embedding results are generally trained on one individual data source such as news pages or Wikipedia articles. However, when we apply them to other tasks such as web search, the performance suffers. To obtain a robust word embedding for different applications, multiple data sources could be leveraged. In this paper, we proposed a two-side multimodal neural network to learn a robust word embedding from multiple data sources including free text, user search queries and search click-through data. This framework takes the word embeddings learned from different data sources as pre-train, and then uses a two-side neural network to unify these embeddings. The pre-trained embeddings are obtained by adapting the recently proposed CBOW algorithm. Since the proposed neural network does not need to re-train word embeddings for a new task, it is highly scalable in real world problem solving. Besides, the network allows weighting different sources differently when applied to different application tasks. Experiments on two real-world applications including web search ranking and word similarity measuring show that our neural network with multiple sources outperforms state-of-the-art word embedding algorithm with each individual source. It also outperforms other competitive baselines using multiple sources. Yong Luo 0002, Jian Tang 0005, Jun Yan 0001, Chao Xu 0006, Zheng Chen 0001 |
AAAI | 4 |
| 2014 | Large-margin Weakly Supervised Dimensionality ReductionabstractThis paper studies dimensionality reduction in a weakly supervised setting, in which the preference relationship between examples is indicated by weak cues. A novel framework is proposed that integrates two aspects of the large margin principle (angle and distance), which simultaneously encourage angle consistency between preference pairs and maximize the distance between examples in preference pairs. Two specific algorithms are developed: an alternating direction method to learn a linear transformation matrix and a gradient boosting technique to optimize a non-linear transformation directly in the function space. Theoretical analysis demonstrates that the proposed large margin optimization criteria can strengthen and improve the robustness and generalization performance of preference learning algorithms on the obtained low-dimensional subspace. Experimental results on real-world datasets demonstrate the significance of studying dimensionality reduction in the weakly supervised setting and the effectiveness of the proposed framework. Chang Xu 0002, Dacheng Tao, Chao Xu 0006, Yong Rui |
ICML | 3 |
| 2014 | Multi-view Multi-task Feature Extraction for Web Image ClassificationabstractThe features used in many multimedia analysis-based applications are frequently of very high dimension. Feature extraction offers several advantages in highly dimensional cases, and many recent studies have used multi-task feature extraction approaches, which often outperform single-task feature extraction approaches. However, most of these methods are limited in that they only consider data represented by a single type of feature, even though features usually represent images from multiple views. We therefore propose a novel multi-view multi-task feature extraction (MVMTFE) framework for handling multi-view features for image classification. In particular, MVMTFE simultaneously learns the feature extraction matrix for each view and the view combination coefficients. In this way, MVMTFE not only handles correlated and noisy features, but also utilizes the complementarity of different views to further help reduce feature redundancy in each view. An alternating algorithm is developed for problem optimization and each sub-problem can be efficiently solved. Experiments on an real-world web image dataset demonstrate the effectiveness and superiority of the proposed method. Zhiqiang Zuo 0001, Yong Luo 0002, Dacheng Tao, Chao Xu 0006 |
ACM Multimedia | 4 |
| 2014 | Large-Margin Multi-ViewInformation BottleneckabstractIn this paper, we extend the theory of the information bottleneck (IB) to learning from examples represented by multi-view features. We formulate the problem as one of encoding a communication system with multiple senders, each of which represents one view of the data. Based on the precise components filtered out from multiple information sources through a "bottleneck", a margin maximization approach is then used to strengthen the discrimination of the encoder by improving the code distance within the frame of coding theory. The resulting algorithm therefore inherits all the merits of the IB principle and coding theory. It has two distinct advantages over existing algorithms, namely, that our method finds a tradeoff between the accuracy and complexity of the multi-view model, and that the encoded multi-view data retains sufficient discrimination for classification. We also derive the robustness and generalization error bound of the proposed algorithm, and reveal the specific properties of multi-view learning. First, the complementarity of multi-view features guarantees the robustness of the algorithm. Second, the consensus of multi-view features reduces the empirical Rademacher complexity of the objective function, enhances the accuracy of the solution, and improves the generalization error bound of the algorithm. The resulting objective function is solved efficiently using the alternating direction method. Experimental results on annotation, classification and recognition tasks demonstrate that the proposed algorithm is promising for practical applications. Chang Xu 0002, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Decomposition-Based Transfer Distance Metric Learning for Image ClassificationabstractDistance metric learning (DML) is a critical factor for image analysis and pattern recognition. To learn a robust distance metric for a target task, we need abundant side information (i.e., the similarity/dissimilarity pairwise constraints over the labeled data), which is usually unavailable in practice due to the high labeling cost. This paper considers the transfer learning setting by exploiting the large quantity of side information from certain related, but different source tasks to help with target metric learning (with only a little side information). The state-of-the-art metric learning algorithms usually fail in this setting because the data distributions of the source task and target task are often quite different. We address this problem by assuming that the target distance metric lies in the space spanned by the eigenvectors of the source metrics (or other randomly generated bases). The target metric is represented as a combination of the base metrics, which are computed using the decomposed components of the source metrics (or simply a set of random bases); we call the proposed method, decomposition-based transfer DML (DTDML). In particular, DTDML learns a sparse combination of the base metrics to construct the target metric by forcing the target metric to be close to an integration of the source metrics. The main advantage of the proposed method compared with existing transfer metric learning approaches is that we directly learn the base metric coefficients instead of the target metric. To this end, far fewer variables need to be learned. We therefore obtain more reliable solutions given the limited side information and the optimization tends to be faster. Experiments on the popular handwritten image (digit, letter) classification and challenge natural image annotation tasks demonstrate the effectiveness of the proposed method. Yong Luo 0002, Tongliang Liu, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Image Process. | 4 |
| 2013 | Vector-Valued Multi-View Semi-Supervsed Learning for Multi-Label Image ClassificationabstractImages are usually associated with multiple labels and comprised of multiple views, due to each image containing several objects (e.g. a pedestrian, bicycle and tree) and multiple visual features (e.g. color, texture and shape). Currently available tools tend to use either labels or features for classification, but both are necessary to describe the image properly. There have been recent successes in using vector-valued functions, which construct matrix-valued kernels, to explore the multi-label structure in the output space. This has motivated us to develop multi-view vector-valued manifold regularization (MV$^3$MR) in order to integrate multiple features. MV$^3$MR exploits the complementary properties of different features, and discovers the intrinsic local geometry of the compact support shared by different features, under the theme of manifold regularization. We validate the effectiveness of the proposed MV$^3$MR methodology for image classification by conducting extensive experiments on two challenge datasets, PASCAL VOC' 07 and MIR Flickr. Yong Luo 0002, Dacheng Tao, Chang Xu 0002, Chao Xu 0006 |
AAAI | 5 |
| 2013 | FIM: A Real-Time Content Based Sample Image Matching SystemabstractSample Image Matching is to decide if a queried image is belongs to the database or not. In this paper, we focus on real-time image matching, which is critical in many real world applications. Although traditional image retrieval methods can be directly utilized for image matching, they usually suffer the high computational cost problem and thus is not applicable here. To resolve this problem, we first introduce ORB, a recently proposed and well-established image feature, for image matching. Then we compare several variants of the descriptors, different size of the codebook, and two approaches to compute the matching scores, based on which we propose a strategy for final matching decision. According to the comparison results, we finally present a real-time image matching system, fast image matching (FIM), which can process about 33 images per second, with a satisfactory accuracy. Yong Luo 0002, Yangxi Li, Jinhui Tu, Chao Xu 0006 |
ICIG | 4 |
| 2013 | MagicBrush: image search by color sketchabstractIn this paper, we showcase the MagicBrush system, a novel painting-based image search engine. This system enables users to draw a color sketch as a query to find images. Different from existing works on sketch-based image retrieval, most of which focus on matching the shape structure without carefully considering other important visual modalities, MagicBrush takes into account the indispensable value of "color" related to "shape", and explores to make use of both the shape and color expectations that users usually have when they're imaging or searching for an image. To achieve this, we 1) develop a user-friendly interface to allow users to easily "paint out" their colorful visual expectations; 2) design a compact feature "color-edge word" to encode both shape and color information in a organic way; and 3) develop a novel matching and index structure to support a real-time response in 6.4 million images. By taking into account both shape and color information, the MagicBrush system helps users to vividly present what they are imagining, and retrieve images in a more natural way. Xinghai Sun, Changhu Wang, Avneesh Sud, Chao Xu 0006, Lei Zhang 0001 |
ACM Multimedia | 4 |
| 2013 | Indexing billions of images for sketch-based retrievalabstractBecause of the popularity of touch-screen devices, it has become a highly desirable feature to retrieve images from a huge repository by matching with a hand-drawn sketch. Although searching images via keywords or an example image has been successfully launched in some commercial search engines of billions of images, it is still very challenging for both academia and industry to develop a sketch-based image retrieval system on a billion-level database. In this work, we systematically study this problem and try to build a system to support query-by-sketch for two billion images. The raw edge pixel and Chamfer matching are selected as the basic representation and matching in this system, owning to the superior performance compared with other methods in extensive experiments. To get a more compact feature and a faster matching, a vector-like Chamfer feature pair is introduced, based on which the complex matching is reformulated as the crossover dot-product of feature pairs. Based on this new formulation, a compact shape code is developed to represent each image/sketch by projecting the Chamfer features to a linear subspace followed by a non-linear source coding. Finally, the multi-probe Kmedoids-LSH is leveraged to index database images, and the compact shape codes are further used for fast reranking. Extensive experiments show the effectiveness of the proposed features and algorithms in building such a sketch-based image search system. Xinghai Sun, Changhu Wang, Chao Xu 0006, Lei Zhang 0001 |
ACM Multimedia | 3 |
| 2013 | A comprehensive study on learning to rank for content-based image retrieval
Yangxi Li, Bo Geng, Chao Xu 0006, Hong Liu 0008 |
Signal Process. | 4 |
| 2013 | Manifold Regularized Multitask Learning for Semi-Supervised Multilabel Image ClassificationabstractIt is a significant challenge to classify images with multiple labels by using only a small number of labeled samples. One option is to learn a binary classifier for each label and use manifold regularization to improve the classification performance by exploring the underlying geometric structure of the data distribution. However, such an approach does not perform well in practice when images from multiple concepts are represented by high-dimensional visual features. Thus, manifold regularization is insufficient to control the model complexity. In this paper, we propose a manifold regularized multitask learning (MRMTL) algorithm. MRMTL learns a discriminative subspace shared by multiple classification tasks by exploiting the common structure of these tasks. It effectively controls the model complexity because different tasks limit one another's search volume, and the manifold regularization ensures that the functions in the shared hypothesis space are smooth along the data manifold. We conduct extensive experiments, on the PASCAL VOC'07 dataset with 20 classes and the MIR dataset with 38 classes, by comparing MRMTL with popular image classification algorithms. The results suggest that MRMTL is effective for image classification. Yong Luo 0002, Dacheng Tao, Bo Geng, Chao Xu 0006, Stephen J. Maybank |
IEEE Trans. Image Process. | 4 |
| 2013 | W-Tree Indexing for Fast Visual Word GenerationabstractThe bag-of-visual-words representation has been widely used in image retrieval and visual recognition. The most time-consuming step in obtaining this representation is the visual word generation, i.e., assigning visual words to the corresponding local features in a high-dimensional space. Recently, structures based on multibranch trees and forests have been adopted to reduce the time cost. However, these approaches cannot perform well without a large number of backtrackings. In this paper, by considering the spatial correlation of local features, we can significantly speed up the time consuming visual word generation process while maintaining accuracy. In particular, visual words associated with certain structures frequently co-occur; hence, we can build a co-occurrence table for each visual word for a large-scale data set. By associating each visual word with a probability according to the corresponding co-occurrence table, we can assign a probabilistic weight to each node of a certain index structure (e.g., a KD-tree and a K-means tree), in order to re-direct the searching path to be close to its global optimum within a small number of backtrackings. We carefully study the proposed scheme by comparing it with the fast library for approximate nearest neighbors and the random KD-trees on the Oxford data set. Thorough experimental results suggest the efficiency and effectiveness of the new scheme. Miaojing Shi, Ruixin Xu, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Image Process. | 4 |
| 2013 | Multiview Vector-Valued Manifold Regularization for Multilabel Image ClassificationabstractIn computer vision, image datasets used for classification are naturally associated with multiple labels and comprised of multiple views, because each image may contain several objects (e.g., pedestrian, bicycle, and tree) and is properly characterized by multiple visual features (e.g., color, texture, and shape). Currently, available tools ignore either the label relationship or the view complementarily. Motivated by the success of the vector-valued function that constructs matrix-valued kernels to explore the multilabel structure in the output space, we introduce multiview vector-valued manifold regularization (MV(3)MR) to integrate multiple features. MV(3)MR exploits the complementary property of different features and discovers the intrinsic local geometry of the compact support shared by different features under the theme of manifold regularization. We conduct extensive experiments on two challenging, but popular, datasets, PASCAL VOC' 07 and MIR Flickr, and validate the effectiveness of the proposed MV(3)MR for image classification. Yong Luo 0002, Dacheng Tao, Chang Xu 0002, Chao Xu 0006, Hong Liu 0008, Yonggang Wen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2012 | Learning to rerank images with enhanced spatial verificationabstractReranking is one of the commonly used schemes to improve the initial ranking performance for content based image retrieval (CBIR). The state-of-the-art reranking methods for CBIR are mainly based on spatial verification and global feature. To mine the complementary properties of different reranking strategies, we combine features representing images from different perspectives with RankSVM to obtain a reranking model to refine the initial ranking list. Besides, compared with traditional spatial verification based methods which measure image similarity only with single inlier's statistical properties, we bind close inlier visual words together to mine more geometric information from images. Through organizing inliers into sequence and computing the relative positions among inliers, we define an efficient similarity measurement with the order consistency between inlier sequences. Experimental results on both Oxford and imageNet datasets demonstrate that our proposed reranking method is effective and promising. Chang Xu 0002, Yangxi Li, Chao Xu 0006 |
ICIP | 4 |
| 2012 | Exploiting visual word co-occurrence for image retrievalabstractBag-of-visual-words (BOVW) based image representation has received intense attention in recent years and has improved content based image retrieval (CBIR) significantly. BOVW does not consider the spatial correlation between visual words in natural images and thus biases the generated visual words towards noise when the corresponding visual features are not stable. In this paper, we construct a visual word co-occurrence table by exploring visual word co-occurrence extracted from small affine-invariant regions in a large collection of natural images. Based on this visual word co-occurrence table, we first present a novel high-order predictor to accelerate the generation of neighboring visual words. A co-occurrence matrix is introduced to refine the similarity measure for image ranking. Like the inverse document frequency (idf), it down-weights the contribution of the words that are less discriminative because of frequent co-occurrence. We conduct experiments on Oxford and Paris Building datasets, in which the ImageNet dataset is used to implement a large scale evaluation. Thorough experimental results suggest that our method outperforms the state-of-the-art, especially when the vocabulary size is comparatively small. In addition, our method is not much more costly than the BOVW model. Miaojing Shi, Xinghai Sun, Dacheng Tao, Chao Xu 0006 |
ACM Multimedia | 4 |
| 2012 | Query difficulty estimation for image retrieval
Yangxi Li, Bo Geng, Linjun Yang, Chao Xu 0006 |
Neurocomputing | 4 |
| 2012 | Ensemble Manifold RegularizationabstractWe propose an automatic approximation of the intrinsic manifold for general semi-supervised learning (SSL) problems. Unfortunately, it is not trivial to define an optimization function to obtain optimal hyperparameters. Usually, cross validation is applied, but it does not necessarily scale up. Other problems derive from the suboptimality incurred by discrete grid search and the overfitting. Therefore, we develop an ensemble manifold regularization (EMR) framework to approximate the intrinsic manifold by combining several initial guesses. Algorithmically, we designed EMR carefully so it 1) learns both the composite manifold and the semi-supervised learner jointly, 2) is fully automatic for learning the intrinsic manifold hyperparameters implicitly, 3) is conditionally optimal for intrinsic manifold approximation under a mild and reasonable assumption, and 4) is scalable for a large number of candidate manifold hyperparameters, from both time and space perspectives. Furthermore, we prove the convergence property of EMR to the deterministic matrix at rate root-n. Extensive experiments over both synthetic and real data sets demonstrate the effectiveness of the proposed framework. Bo Geng, Dacheng Tao, Chao Xu 0006, Linjun Yang, Xian-Sheng Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Ranking Model Adaptation for Domain-Specific SearchabstractWith the explosive emergence of vertical search domains, applying the broad-based ranking model directly to different domains is no longer desirable due to domain differences, while building a unique ranking model for each domain is both laborious for labeling data and time consuming for training models. In this paper, we address these difficulties by proposing a regularization-based algorithm called ranking adaptation SVM (RA-SVM), through which we can adapt an existing ranking model to a new domain, so that the amount of labeled data and the training cost is reduced while the performance is still guaranteed. Our algorithm only requires the prediction from the existing ranking models, rather than their internal representations or the data from auxiliary domains. In addition, we assume that documents similar in the domain-specific feature space should have consistent rankings, and add some constraints to control the margin and slack variables of RA-SVM adaptively. Finally, ranking adaptability measurement is proposed to quantitatively estimate if an existing ranking model can be adapted to a new domain. Experiments performed over Letor and two large scale data sets crawled from a commercial search engine demonstrate the applicabilities of the proposed ranking adaptation algorithms and the ranking adaptability measurement. Bo Geng, Linjun Yang, Chao Xu 0006, Xian-Sheng Hua 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | Parallel Lasso for Large-Scale Video Concept DetectionabstractExisting video concept detectors are generally built upon the kernel based machine learning techniques, e.g., support vector machines, regularized least squares, and logistic regression, just to name a few. However, in order to build robust detectors, the learning process suffers from the scalability issues including the high-dimensional multi-modality visual features and the large-scale keyframe examples. In this paper, we propose parallel lasso (Plasso) by introducing the parallel distributed computation to significantly improve the scalability of lasso (thel1regularized least squares). We apply the parallel incomplete Cholesky factorization to approximate the covariance statistics in the preprocess step, and the parallel primal-dual interior-point method with the Sherman-Morrison-Woodbury formula to optimize the model parameters. For a dataset withnsamples in ad-dimensional space, compared with lasso, Plasso significantly reduces complexities from the originalO(d3) for computational time andO(d2) for storage space toO(h2d/m) andO(hd/m) , respectively, if the system hasmprocessors and the reduced dimensionhis much smaller than the original dimensiond. Furthermore, we develop the kernel extension of the proposed linear algorithm with the sample reweighting schema, and we can achieve similar time and space complexity improvements [time complexity fromO(n3) toO(h2n/m) and the space complexity fromO(n2) toO(hn/m), for a dataset withntraining examples]. Experimental results on TRECVID video concept detection challenges suggest that the proposed method can obtain significant time and space savings for training effective detectors with limited communication overhead. Bo Geng, Yangxi Li, Dacheng Tao, Meng Wang 0001, Zhengjun Zha, Chao Xu 0006 |
IEEE Trans. Multim. | 6 |
| 2012 | Difficulty Guided Image Retrieval Using Linear Multiple Feature EmbeddingabstractExisting image retrieval systems suffer from a performance variance for different queries. Severe performance variance may greatly degrade the effectiveness of the subsequent query-dependent ranking optimization algorithms, especially those that utilize the information mined from the initial search results. In this paper, we tackle this problem by proposing a query difficulty guided image retrieval system, which can predict the queries' ranking performance in terms of their difficulties and adaptively apply ranking optimization approaches. We estimate the query difficulty by comprehensively exploring the information residing in the query image, the retrieval results, and the target database. To handle the high-dimensional and multi-model image features in the large-scale image retrieval setting, we propose a linear multiple feature embedding algorithm which learns a linear transformation from a small set of data by integrating a joint subspace in which the neighborhood information is preserved. The transformation can be effectively and efficiently used to infer the subspace features of the newly observed data in the online setting. We prove the significance of query difficulty to image retrieval by applying it to guide the conduction of three retrieval refinement applications, i.e., reranking, federated search, and query suggestion. Thorough empirical studies on three datasets suggest the effectiveness and scalability of the proposed image query difficulty estimation algorithm, as well as the promising of the image difficulty guided retrieval system. Yangxi Li, Bo Geng, Dacheng Tao, Zhengjun Zha, Linjun Yang, Chao Xu 0006 |
IEEE Trans. Multim. | 6 |
| 2011 | Real-time human tracking based on switching linear dynamic system combined with adaptive Meanshift trackerabstractReal-time human tracking in complex environments usually presents many challenges, such as partial/complete occlusions caused by irregular motion and similar-color distractors. Switching Linear Dynamic System(SLDS) and Meanshift(MS) are two successful approaches, although both have inherent deficiencies, like accumulated prediction and correction errors in SLDS, uncoded attentional and spatial information in Meanshift, respectively. In this paper, a spatial-representive Meanshift and a joint attentional feature histogram from bottom-up and top-down attention models are used to make the tracker more adaptive. Then a tracking algorithm is proposed as adaptive Meanshift embedded SLDS, where four transition matrixes A(st) handle partial/complete occlusions with irregular motion under the framework, and adaptive Meanshift solves short occlusions of similar-color distractors in local search for higher accuracy. Experiments show that this method can work more robustly for partial/complete occlusions among multiple persons compared with adaptive Meanshift and Meanshift embedded SLDS. Hong Liu 0008, Chao Xu 0006 |
ICIP | 3 |
| 2011 | Fast visual word quantization via spatial neighborhood boostingabstractWith the rapid development of bag-of-visual-word model and its wide-spread applications in various computer vision problems such as visual recognition, image retrieval tasks, etc., fast visual word assignment becomes increasingly important, especially for some on-line services and large scale settings. The conventional approximate nearest neighbor mapping techniques purely consider the distribution of image local descriptors in the visual feature space and perform the mapping process independently for each descriptor. In this paper, we propose to involve the spatial correlation information to boost the efficiency of feature quantization. The visual words that frequently co-occur in the same local region of a large number of images are considered as spatial neighborhoods, which can be leveraged to boost the approximate mapping of neighbored local descriptors. Experimental results on a well-known image retrieval dataset demonstrate that, the proposed method is capable of improving the efficiency and precision of visual word assignment. Ruixin Xu, Miaojing Shi, Bo Geng, Chao Xu 0006 |
ICME | 4 |
| 2011 | The role of attractiveness in web image searchabstractExisting web image search engines are mainly designed to optimize topical relevance. However, according to our user study, attractiveness is becoming a more and more important factor for web image search engines to satisfy users' search intentions. Important as it can be, web image attractiveness from the search users' perspective has not been sufficiently recognized in both the industry and the academia. In this paper, we present a definition of web image attractiveness with three levels according to the end users' feedback, including perceptual quality, aesthetic sensitivity and affective tune. Corresponding to each level of the definition, various visual features are investigated on their applicability to attractiveness estimation of web images. To further deal with the unreliability of visual features induced by the large variations of web images, we propose a contextual approach to integrate the visual features with contextual cues mined from image EXIF information and the associated web pages. We explore the role of attractiveness by applying it to various stages of a web image search engine, including the online ranking and the interactive reranking, as well as the offline index selection. Experimental results on three large-scale web image search datasets demonstrate that the incorporation of attractiveness can bring more satisfaction to 80% of the users for ranking/reranking search results and 30.5% index coverage improvement for index selection, compared to the conventional relevance based approaches. Bo Geng, Linjun Yang, Chao Xu 0006, Xian-Sheng Hua 0001, Shipeng Li 0001 |
ACM Multimedia | 3 |
| 2011 | Query expansion by spatial co-occurrence for image retrievalabstractThe well-known bag-of-features (BoF) model is widely utilized for large scale image retrieval. However, BoF model lacks the spatial information of visual words, which is informative for local features to build up meaningful visual patches. To compensate for the spatial information loss, in this paper, we propose a novel query expansion method called Spatial Co-occurrence Query Expansion (SCQE), by utilizing the spatial co-occurrence information of visual words mined from the database images to boost the retrieval performance. In offline phase, for each visual word in the vocabulary, we treat the visual words that are frequently co-occurred with it in the database images as neighbors, base on which a spatial co-occurrence graph is built. In online phase, a query image can be expanded with some spatial co-occurred but unseen visual words according to the spatial co-occurrence graph, and the retrieval performance can be improved by expanding these visual words appropriately. Experimental results demonstrate that, SCQE achieves promising improvements over the typical BoF baseline on two datasets comprising 5K and 505K images respectively. Yingfei Li, Bo Geng, Zhengjun Zha, Yangxi Li, Dacheng Tao, Chao Xu 0006 |
ACM Multimedia | 6 |
| 2011 | Difficulty guided image retrieval using linear multiview embeddingabstractExisting image retrieval systems suffer from a radical performance variance for different queries. The bad initial search results for "difficult" queries may greatly degrade the performance of their subsequent refinements, especially the refinement that utilizes the information mined from the search results, e.g., pseudo relevance feedback based reranking. In this paper, we tackle this problem by proposing a query difficulty guided image retrieval system, which selectively performs reranking according to the estimated query difficulty. To improve the performance of both reranking and difficulty estimation, we apply multiview embedding (ME) to images represented by multiple different features for integrating a joint subspace by preserving the neighborhood information in each feature space. However, existing ME approaches suffer from both "out of sample" and huge computational cost problems, and cannot be applied to online reranking or offline large-scale data processing for practical image retrieval systems. Therefore, we propose a linear multiview embedding algorithm which learns a linear transformation from a small set of data and can effectively infer the subspace features of new data. Empirical evaluations on both Oxford and 500K ImageNet datasets suggest the effectiveness of the proposed difficulty guided retrieval system with LME. Yangxi Li, Bo Geng, Zhengjun Zha, Dacheng Tao, Linjun Yang, Chao Xu 0006 |
ACM Multimedia | 6 |
| 2011 | Shared feature extraction for semi-supervised image classificationabstractMulti-task learning (MTL) plays an important role in image analysis applications, e.g. image classification, face recognition and image annotation. That is because MTL can estimate the latent shared subspace to represent the common features given a set of images from different tasks. However, the geometry of the data probability distribution is always supported on an intrinsic image sub-manifold that is embedded in a high dimensional Euclidean space. Therefore, it is improper to directly apply MTL to multiclass image classification. In this paper, we propose a manifold regularized MTL (MRMTL) algorithm to discover the latent shared subspace by treating the high-dimensional image space as a sub-manifold embedded in an ambient space. We conduct experiments on the PASCAL VOC'07 dataset with 20 classes and the MIR dataset with 38 classes by comparing MRMTL with conventional MTL and several representative image classification algorithms. The results suggest that MRMTL can properly extract the common features for image representation and thus improve the generalization performance of the image classification models. Yong Luo 0002, Dacheng Tao, Bo Geng, Chao Xu 0006, Stephen J. Maybank |
ACM Multimedia | 4 |
| 2011 | Query Difficulty Guided Image Retrieval System
Yangxi Li, Yong Luo 0002, Dacheng Tao, Chao Xu 0006 |
MMM (2) | 4 |
| 2011 | DAML: Domain Adaptation Metric LearningabstractThe state-of-the-art metric-learning algorithms cannot perform well for domain adaptation settings, such as cross-domain face recognition, image annotation, etc., because labeled data in the source domain and unlabeled ones in the target domain are drawn from different, but related distributions. In this paper, we propose the domain adaptation metric learning (DAML), by introducing a data-dependent regularization to the conventional metric learning in the reproducing kernel Hilbert space (RKHS). This data-dependent regularization resolves the distribution difference by minimizing the empirical maximum mean discrepancy between source and target domain data in RKHS. Theoretically, by using the empirical Rademacher complexity, we prove risk bounds for the nearest neighbor classifier that uses the metric learned by DAML. Practically, learning the metric in RKHS does not scale up well. Fortunately, we can prove that learning DAML in RKHS is equivalent to learning DAML in the space spanned by principal components of the kernel principle component analysis (KPCA). Thus, we can apply KPCA to select most important principal components to significantly reduce the time cost of DAML. We perform extensive experiments over four well-known face recognition datasets and a large-scale Web image annotation dataset for the cross-domain face recognition and image annotation tasks under various settings, and the results demonstrate the effectiveness of DAML. Bo Geng, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Image Process. | 3 |
| 2010 | Content-aware Ranking for visual searchabstractThe ranking models of existing image/video search engines are generally based on associated text while the visual content is actually neglected. Imperfect search results frequently appear due to the mismatch between the textual features and the actual visual content. Visual reranking, in which visual information is applied to refine text based search results, has been proven to be effective. However, the improvement brought by visual reranking is limited, and the main reason is that the errors in the text-based results will propagate to the refinement stage. In this paper, we propose a Content-Aware Ranking model based on “learning to rank” framework, in which textual and visual information are simultaneously leveraged in the ranking learning process. We formulate the Content-Aware Ranking based on large margin structured output learning, by modeling the visual information into a regularization term. The direct optimization of the learning problem is nearly infeasible since the number of constraints is huge. The efficient cutting plane algorithm is adopted to learn the model by iteratively adding the most violated constraints. Extensive experimental results on a large-scale dataset collected from a commercial Web image search engine, as well as the TRECVID 2007 video search dataset, demonstrate that the proposed ranking model significantly outperforms the state-of-the-art ranking and reranking methods. Bo Geng, Linjun Yang, Chao Xu 0006, Xian-Sheng Hua 0001 |
CVPR | 3 |
| 2010 | Self-calibrating photometric stereoabstractWe present a self-calibrating photometric stereo method. From a set of images taken from a fixed viewpoint under different and unknown lighting conditions, our method automatically determines a radiometric response function and resolves the generalized bas-relief ambiguity for estimating accurate surface normals and albedos. We show that color and intensity profiles, which are obtained from registered pixels across images, serve as effective cues for addressing these two calibration problems. As a result, we develop a complete auto-calibration method for photometric stereo. The proposed method is useful in many practical scenarios where calibrations are difficult. Experimental results validate the accuracy of the proposed method using various real-world scenes. Boxin Shi, Yasuyuki Matsushita, Chao Xu 0006, Ping Tan 0002 |
CVPR | 4 |
| 2009 | Color Correction and Compression for Multi-view Video Using H.264 Features
Boxin Shi, Yangxi Li, Chao Xu 0006 |
ACCV (3) | 4 |
| 2009 | Ranking model adaptation for domain-specific searchabstractRecently, various domain-specific search engines emerge, which are restricted to specific topicalities or document formats, and vertical to the broad-based search. Simply applying the ranking model trained for the broad-based search to the verticals cannot achieve a sound performance due to the domain differences, while building different ranking models for each domain is both laborious for labeling sufficient training samples and time-consuming or the training process. In this paper, to address the above difficulties, we investigate two problems: (1) whether we can adapt the ranking model learned for existing Web page search or verticals, to the new domain, so that the amount of labeled data and the training cost is reduced, while the performance requirement is still satisfied; and (2) how to adapt the ranking model from auxiliary domains to a new target domain. We address the second problem from the regularization framework and an algorithm called ranking adaptation SVM is proposed. Our algorithm is flexible enough, which needs only the prediction from the existing ranking model, rather than the internal representation of the model or the data from auxiliary domains. The first problem is addressed by the proposed ranking adaptability measurement, which quantitatively estimates if an existing ranking model can be adapted to the new domain. Extensive experiments are performed over Letor benchmark dataset and two large scale datasets crawled from different domains through a commercial internet search engine, where the ranking model learned for one domain will be adapted to the other. The results demonstrate the applicabilities of the proposed ranking model adaptation algorithm and the ranking adaptability measurement. Bo Geng, Linjun Yang, Chao Xu 0006, Xian-Sheng Hua 0001 |
CIKM | 3 |
| 2009 | Ensemble manifold regularizationabstractWe propose an automatic approximation of the intrinsic manifold for general semi-supervised learning problems. Unfortunately, it is not trivial to define an optimization function to obtain optimal hyperparameters. Usually, pure cross-validation is considered but it does not necessarily scale up. A second problem derives from the suboptimality incurred by discrete grid search and overfitting problems. As a consequence, we developed an ensemble manifold regularization (EMR) framework to approximate the intrinsic manifold by combining several initial guesses. Algorithmically, we designed EMR very carefully so that it (a) learns both the composite manifold and the semi-supervised classifier jointly; (b) is fully automatic for learning the intrinsic manifold hyperparameters implicitly; (c) is conditionally optimal for intrinsic manifold approximation under a mild and reasonable assumption; and (d) is scalable for a large number of candidate manifold hyperparameters, from both time and space perspectives. Extensive experiments over both synthetic and real datasets show the effectiveness of the proposed framework. Bo Geng, Chao Xu 0006, Dacheng Tao, Linjun Yang, Xian-Sheng Hua 0001 |
CVPR | 2 |
| 2009 | Integrating Color Constancy into Multi-view Video CodingabstractColor constancy is the ability to remove the dependency of illuminant and show the intrinsic color of objects. A novel multi-view video coding scheme integrated with color constancy algorithms is introduced in this paper. Color constancy algorithms based on gray-edge hypothesis are selected. The new scheme takes color constancy as a preprocessing to make the sequence independent of the illuminant conditions. Based on the examination of the scene change in the sequence, the key frames are picked out and the parameters of color constancy are determined. The illuminant information is sent into the encoder along with the illuminant independent sequence, and then it is used to solve the color cast problem of multi-view coding at the decoder. Both visual quality and coding performance prove that this novel scheme can lower the bitrates and gain better color quality. Yangxi Li, Boxin Shi, Chao Xu 0006 |
ICIG | 3 |
| 2009 | Intrinsic Image Decomposition Using Color Invariant EdgeabstractThe intrinsic image composed of reflectance and shading images plays important roles in various computer vision applications. This paper focuses on solving the problem of intrinsic image decomposition. Based on the assumption that the image derivatives can be classified into either reflectance-related or shading-related, the reflectance and shading image can be restored from the classified derivatives. We improve the classification result using only color information by introducing the color invariant edge. Considering some color invariant properties in the image, the color invariant edge can provide more useful information in guiding the classification and producing more robust decomposition result as it is shown in the experiment. Boxin Shi, Yangxi Li, Chao Xu 0006 |
ICIG | 3 |
| 2009 | Block-based color correction algorithm for multi-view video codingabstractThe color variations among different viewpoints in multiview video sequences may deteriorate the visual quality and coding efficiency. Various color correction methods have been proposed, however, the color appearance and histogram of corrected target frames are not similar enough to the reference frames in details. Focusing on restoring more similar color, a block-based color correction algorithm is proposed. The blocks in reference frames are matched into target frames through spatial prediction, and the colorization scheme is then adopted to expand color as a coarse correction. Finally the mixture with global color transfer result yields the fine correction. The experiment results show this novel method can provide better visual effect in detail and also provide the corrected frames with histograms more similar to reference histograms. Boxin Shi, Yangxi Li, Chao Xu 0006 |
ICME | 4 |
| 2008 | Unbiased active learning for image retrievalabstractIn transductive active learning, after selecting the samples for labeling using existing sample selection strategy such as close-to-boundary, the constructed labeled set will be under a different distribution from the unlabeled set, which violates the i.i.d assumption of existing classifier. In this paper, by explicitly considering the distribution difference, we propose an algorithm called unbiased active learning. In such algorithm, the distribution difference, so-called sample selection bias, is not only considered into the classifier, but also incorporated into the sample selection process for introducing a better sample selection strategy. We apply the proposed method to image retrieval and the experimental results show that our unbiased active learning algorithm outperforms existing approaches. Bo Geng, Linjun Yang, Zhengjun Zha, Chao Xu 0006, Xian-Sheng Hua 0001 |
ICME | 4 |
| 2008 | Comparison between JPEG2000 and H.264 for digital cinemaabstractJPEG2000 and H.264 are the latest image and video coding standards respectively. Digital cinema is a new kind of application for super high definition video. The DCI (Digital Cinema Initiative) specification published in 2005 has selected JPEG2000 instead of H.264 as the video coding standard for digital cinema. It shows that JPEG2000 has a better performance in the field of super high definition video coding. Until now, only a few basic tests have been done to compare between JPEG2000 and H.264. Moreover, the test resolutions and characteristics of input sequences are very limited. In this paper, based on JPEG2000, H.264 intra and inter-frame coding, we compare the coding efficiency and subjective image quality on multiple series of test sequences from low to super high resolutions. The experiment results demonstrate some regularity for JPEG2000 and H.264 video coding, and reveal JPEG2000 is more suitable for digital cinema. Boxin Shi, Chao Xu 0006 |
ICME | 3 |
| 2008 | A wavelet packet based block-partitioning image coding algorithm with rate-distortion optimization
Yongming Yang, Chao Xu 0006 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2007 | Performance Analysis and Architecture Design for Parallel EBCOT Encoder of JPEG2000abstractThe algorithm of embedded block coding with optimization truncation (EBCOT) is one of the key techniques in JPEG2000 standard. The high-speed/performance hardware designs for EBCOT are critical to various high-resolution applications, such as digital cameras, digital video recorders, and HDTV, etc. This paper presents a detailed performance analysis for EBCOT Tier-1 and its bit plane parallel architecture design. By analyzing the bit plane parallel context modeling, we conclude that the difference between the output rate of context modeling module and that of arithmetic coding module degrades the performance of whole EBCOT parallel encoding. We propose two improved methods, referred to as data-pairs ordering (DPO) and flexible MQ (FMQ) coder. It solves the configuration problem between the parallel context modeling module and the sequent arithmetic coding module, takes full advantage of the bit plane parallel encoding technique, and improves the coding speed and efficiency of the EBCOT encoder significantly. The design of parallel EBCOT encoder is tested on the field-programmable gate array (FPGA) platform of Altera Company. The simulation results show that it can on average encode 54 million samples at 55-MHz working frequency. It is equivalent to encoding a 90006000 image per second or encoding 720 p (1280times720, 4:2:2) HDTV picture sequence nearly 30 frames per second. Compared with the conventional bit plane parallel architecture design, the proposed one can reduce execution time by 24%. Yi-Zhen Zhang 0001, Chao Xu 0006, Liang-Bin Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Matrix Factorization for Fast DCT AlgorithmsabstractTwo principles to produce new possibilities for the radix-2 Discrete Cosine Transform (DCT) have been presented in this paper. One is to employ matrix factorization through revealing the intrinsic relationship among several existing famous algorithms, which is regarded as an effective guide for exploring new algorithms. The other is to make use of the orthogonal property of the DCT matrix. As long as the recursive kernel of an algorithm is orthogonal, there must be a twin fast DCT algorithm of it. Matrix factorization is applied through the research and can be used to show how data flows and compute the computational complexity easily. At the end of this paper, we also present a new fast algorithm for DCT. It enjoys the parallel structure which is simpler for programming and hardware implementation and keeps the same numbers of the additions and multiplications as the fastest algorithms. Wenjia Yuan, Pengwei Hao, Chao Xu 0006 |
ICASSP (3) | 3 |
| 2006 | A Memory-Saving System Including DWT and EC for JPEG2000 ImplementationabstractIn this paper, a new system including discrete wavelet transform (DWT) and entropy coder (EC) is presented for JPEG2000 implementation. The system accomplishes a seamless connection between DWT and EC. To reach the goal, an EC-based DWT scheme is proposed, and a new EC architecture is designed that encodes three code-blocks concurrently. The system reduces the memory requirement significantly. Compared to the QCB-DWT scheme, more than 72 percent of memory is removed. And it is a high speed encoding system, able to encode 720×576 video more than 40 frames per second at 54 MHz clock frequency. Chao Xu 0006, Yanju Han, Su-Fang Hu |
ICIP | 1 |
| 2006 | Analysis and Effective Parallel Technique for Rate-Distortion Optimization in JPEG2000abstractThe rate-distortion optimization (RDO) technique plays an important role in JPEG2000 coding systems. In this paper, detailed analysis and parallel technique for RDO is presented. Differing from the usual or so far improved RDO techniques, the proposed technique can process each bit-plane's code-stream and rate-distortion pairs independently, which provides parallel processing and is easy for hardware implementation. It can speed up the RDO processing and meets the requirement of the parallel JPEG2000 encoder. Experiment results show that the PSNR performance is optimal and the coding speed satisfies the requirement of the high resolution real-time applications. Yi-Zhen Zhang 0001, Chao Xu 0006 |
ICIP | 2 |
| 2006 | Fast wavelet packet basis selection for block-partitioning image codingabstractWavelet packet provides an effective representative tool for adaptive waveform analysis of a given signal. To construct an appropriate basis, the basis-selection criterion must be carefully designed. A rate-distortion based basis selection method was proposed in recent years, which can produce an optimal basis in the rate-distortion sense. However, this approach is extremely computationally intensive. Other criterions, such as the entropy based method, do not always produce an effective basis. In this work, we propose a distortion based basis-selection criterion with moderate computational complexity. Two basis selection procedures, i.e. greedy tree growing algorithm and optimal tree pruning algorithm are both evaluated. The experimental results show that, when adopting the same quantization and entropy coding methods, the coding performance of the proposed scheme is inferior to that of the rate-distortion optimized basis by only 0.11dB on the average, while the basis search procedure is sped up by a factor of 10.1 averagely Yongming Yang, Chao Xu 0006 |
ISCAS | 2 |
| 2005 | An improved bit-plane and pass dual parallel architecture for coefficient bit modeling in JPEG2000abstractEmbedded block coding with optimized truncation (EBCOT) is a critical part in JPEG2000 systems. There are bit-plane and pass dual parallel methods that can speed up the encoding, but the acceleration is always companied with the complication of the circuit structure and the increase of the circuit resources. In this paper, we present an improved bit-plane and pass dual parallel architecture (IBPDP), which not only achieves a high encoding speed but also reduces the logic circuit requirement and the coding delay. Experimental results show that about 45% of the logic circuit is reduced and that the average fall of the delay per code-block is 10% compared with BPDP. Yanju Han, Chao Xu 0006, Yi-Zhen Zhang 0001 |
ASP-DAC | 2 |
| 2005 | A wavelet packet based block-partitioning image coding algorithm with rate-distortion optimizationabstractImage coding methods based on wavelet packet decomposition show more potential than those based on usual wavelet transform since wavelet packet presents an effective tool for adaptive waveform analysis. However, most of current work adopts zerotree quantization methods inside the wavelet packet framework, and these quantization methods are proved to be inappropriate for wavelet packet subbands. In this paper, we present an image coding algorithm based on a rate-distortion optimized wavelet packet decomposition and on a block-partitioning coding scheme which quantizes each subband separately. In this way, our algorithm naturally overcomes the difficulty in defining parent-child relationships for wavelet packet subbands, which exists in the zerotree based methods. The experimental results show that the proposed algorithm significantly outperforms SPIHT and JPEG2000 schemes and also surpasses recent wavelet packet based image coding algorithms, in terms of both PSNR and visual quality. Yongming Yang, Chao Xu 0006 |
ICIP (3) | 2 |
| 2005 | Analysis and high performance parallel architecture design for EBCOT in JPEG2000abstractThis paper presents analysis and design of bitplane parallel architecture for embedded block coding with optimization truncation (EBCOT) tier-1 in JPEG2000. By detailed analysis of bitplane parallel context modeling, we discover the reason that causes the degradation of the parallel entropy coding performance. Two improved methods, referred to as data-pairs ordering (DPO) and flexible MQ (FMQ) coder are proposed. The coding speed can be improved greatly by using the improved bitplane parallel coding methods. Compared with the conventional bitplane parallel architecture, our architecture can decrease about 24% of the computation time. Yi-Zhen Zhang 0001, Chao Xu 0006 |
ICIP (3) | 2 |
| 2004 | Bit-plane and pass dual parallel architecture for coefficient bit modeling in JPEG2000abstractIn this paper, the bit-plane and pass dual parallel architecture for coefficient bit modeling in JPEG2000 is proposed. It is a very high speed and efficient structure that is capable of encoding all bits of the wavelet coefficient in only one scan, and largely decreases the memory requirement. Additionally, in order to decrease the logic circuit requirements we propose a partial primitive-parallel technique to replace the whole channel-parallelism. Experimental results show that the architecture can encode about 15 times more than the pass-parallel encoding for the coefficient with 16 bits. And it only requires 3 K bits memory instead of 16 K bits for the pass-parallel encoding. Chao Xu 0006, Yanju Han, Yi-Zhen Zhang 0001 |
ICASSP (5) | 1 |
| 2002 | A fast efficient architecture for MPEG-4 zerotree encoderabstractThis paper presents a novel architecture of MPEG-4 zerotree encoder. Under the architecture, a fast technique of label coefficients is proposed to reduce the recursive scan. It is achieved by exploiting the feature of the MPEG-4 zerotree symbol alphabet. A new combined structure of ZTR address buffer and the significant flag bit is described. It can simplify the skipping of the significant coefficients and locating the descendent coefficients of ZTR/VZTR. Furthermore, a preprocessor is given for independent encoding of each tree in individual bitplane. Chao Xu 0006, Yi-Zhen Zhang 0001, Qing-Yun Shi |
ICME (1) | 1 |