VLDB 2026 Research / reviewers in the wild / expert
Jianchao Yang
dblp:96/3835
· DBLP profile ↗
112ranked-venue papers
18as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 74 · 14 first-author · 7 since 2021Artificial intelligence and machine learning · 55 · 7 first-author · 2 since 2021Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DyGen: A Constant-Time Kernel Generator for Dynamic-Shape Neural NetworksabstractIn recent years, dynamic-shape neural networks have been widely adopted in intelligent applications, such as Mixture-of-Experts based large language models and computer vision tasks. However, in dynamic scenarios, operator shapes are determined at runtime. This leads to prohibitively expensive compilation times for existing static compilers, as they must search across a vast optimization space to identify the best configuration. To address the need for efficient optimization of dynamic-shape neural networks, we present DyGen (Dynamic-shape Kernel Generator)—a lightweight, two-stage compiler plug-in on GPU platforms. In the offline stage, DyGen employs deliberately crafted pruning rules to construct a compact candidate configuration set for the target hardware, then select the configuration of the high-performance kernel to train a configuration generation model. During the online stage, dynamic operator information is directly fed into the generator, which can quickly produce efficient kernel configurations without the need for costly search. Compared to state-of-the-art tensor compilers, DyGen improves inference performance by an average of 36%, while significantly reducing generation overhead from 9 seconds to 0.3 seconds. Yuhan Kang, Dong Chen 0015, Yang Shi 0008, Jianchao Yang, Zeyu Xue, Mei Wen |
DATE | 5 |
| 2026 | Precision boundary modeling for area-efficient Block Floating Point accumulation
Xin Ju 0005, Yasong Cao, Zhongdi Luo, Jianchao Yang, Jingkui Yang, Dong Chen 0015, Mei Wen |
J. Syst. Archit. | 6 |
| 2026 | GroupSD: Self-Distillation From Intermediate ViT Layers for Generalizable Person Re-Identification
Jieru Jia, Jianchao Yang, Chao Li 0070, Qiuqi Ruan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | SparSynergy: Unlocking Flexible and Efficient DNN Acceleration Through Multi-Level SparsityabstractTo more effectively address the computational and memory requirements of deep neural networks (DNNs), leveraging multi-level sparsity-including value-level and bit-level sparsity-has emerged as a pivotal strategy. While substantial research has been dedicated to exploring value-level and bit-level sparsity individually, the combination of both has largely been overlooked until now. In this paper, we propose SparSynergy, which-to the best of our knowledge-is the first accelerator that synergistically integrates multi-level sparsity into a unified framework, maximizing computational efficiency and minimizing memory usage. However, jointly considering multi-level sparsity is non-trivial, as it presents several challenges: (1) increased hardware overhead due to the complexity of incorporating multiple sparsity levels, (2) bandwidth-intensive data transmission during multiplexing, and (3) decreased throughput and scalability caused by bottlenecks in bit-serial computation. Our proposed SparSynergy addresses these challenges by introducing a unified sparsity format and a cooptimized hardware design. Experimental results demonstrate that SparSynergy achieves a 5.38 x geometric mean improvement in the energy-delay product (EDP) when compared with the tensor core, across workloads with varying degrees of sparsity. Furthermore, SparSynergy significantly improves accuracy retention compared to state-of-the-art accelerators for representative DNNs. Jingkui Yang, Mei Wen, Junzhong Shen, Jianchao Yang, Yasong Cao, Minjin Tang, Zhaoyun Chen, Yang Shi 0008 |
DATE | 4 |
| 2025 | A Quality-Aware Sampling Framework for Efficient 3D Point Cloud TransmissionabstractThe large volume of data from the point cloud brings significant demands on network bandwidth. However, the current transmission framework only considers using lossy compression to control the size of data, while ignoring visually redundant information due to the setting of rendering devices. Based on the fact that point overlapping might occur for the case that a dense point cloud is rendered on a relatively low resolution 2D monitor, we propose a novel quality-aware sampling framework for point cloud transmission. When a target visual quality is determined, an optimal sampling module is designed to remove overlapped points with the help of a simple but effective quality model. By taking into account the impact of multiple factors (i.e., sampling, lossy compression, and client rendering resolution), this quality model can predict the final perceptual quality in the client. Based on a newly constructed dataset which consists of 420 samples, experiment results show that the proposed transmission framework can significantly reduce bandwidth cost (e.g., 6.10% to 84.43%) and processing time (e.g., 8.99% to 92.53%) without introducing noticeable distortion under certain rendering conditions, thus achieving higher bandwidth utilization and better real-time performance. Puyue Hou, Qi Yang 0003, Yue Li 0015, Jianchao Yang, Yiling Xu, Tiejun Huang 0001 |
ICASSP | 5 |
| 2025 | SSDViT: Exploring Siamese and Self Distillation in ViTs for Generalizable Person Re-identificationabstractPerson re-identification (re-ID) models often fail to generalize well when deployed to unseen camera networks with domain shift. Domain generalization (DG) aims to address this dilemma by training a model on source domains that learns domain-invariant, and hence generalizable representations. Most methods in the literature are based on Convolutional Neural Networks (CNNs), while the DG performance of Vision Transformers (ViTs) remains relatively unexplored. In contrast to CNNs, ViTs lack explicit inductive biases, which makes it extremely data-hungry and easily overfit to source domains. In this paper, we investigate the generalization ability of ViTs and propose a novel Siamese and Self Distillation Vision Transformer (SSDViT) framework towards addressing the DG re-ID problem. To be specific, Siamese Distillation exploits the weak-to-strong consistency regularization at the image level to enforce the strongly perturbed image to yield consistent prediction with its weakly perturbed version, which provides a simple yet effective solution to introduce invariance bias. On the other hand, Self Distillation seeks to impose consistency constraints at the feature level by leveraging intermediate knowledge to improve the robustness of learned representations. The proposed unified framework pursues the equivalence of predictions at both the image and embedding levels, which underpins the generalization capabilities of learned representations and alleviates the risk of overfitting to source domains. Without bells and whistles, the proposed approach achieves a new state-of-the-art on various DG re-ID benchmarks. Codes are available at https://github.com/yJCTrans/SSDViT. Jieru Jia, Jianchao Yang |
ICASSP | 2 |
| 2025 | Neural Adaptive Contextual Video StreamingabstractVideo streaming services typically employ traditional codecs, such as H.264, to encode videos into multiple bitrate representations. These codecs are tightly limited by discrete quantization parameters (QPs), resulting in encoded rates that do not align with the target bitrate. Additionally, the subpar video quality produced by conventional codecs does not meet the demands of high-resolution communication. Considering the limitations of traditional codecs, we take a fresh new approach to video streaming by leveraging advanced deep learning-based video codecs. Specifically, we develop a neural adaptive contextual video streaming framework that incorporates: 1) an ensemble deep reinforcement learning based adaptive bitrate algorithm named TSAC that enables continuous bitrate adjustment to varying network conditions 2) a two-stage proportional-integral-derivative-based rate control module that dynamically fine-tunes QPs to ensure the encoded bitrate aligning with the target bitrate. Furthermore, we implement intra-GoP and inter-GoP techniques to accelerate the inference process of the contextual video codec for real-time processing needs. Our experiments demonstrate that the average relative error in bitrate remains below 2%, the quality of experience provided by our TSAC agents surpasses that of existing discrete algorithms by 13%-20%. Our optimization techniques enable real-time decoding at approximately 24 frames per second for quad high definition videos. Jianchao Yang, Mufan Liu, Puyue Hou, Yiling Xu, Jun Sun 0005 |
ICASSP | 1 |
| 2025 | SPSA: Exploring Sparse-Packing Computation on Systolic Arrays From ScratchabstractSparse matrix-matrix multiplication (SpMM) and Generalized SpMM (SpGEMM) are essential computational kernels in domains, such as graph analytics and scientific computation. While systolic arrays have traditionally been employed as specialized architectures for complex computing problems like matrix multiplication, they exhibit inefficiency when dealing with sparse matrices. This inefficiency arises from the unnecessary operations performed by processing elements (PEs) that contain zero-valued entries, which do not contribute to the final result. To address this issue, we propose SPSA, a framework that leverages a sparse-packing algorithm suitable for systolic arrays to accelerate sparse matrix computations. Our approach achieves significant reduction of zero-valued items and improves matrix density by packing the rows or columns of the sparse matrix. Furthermore, we have introduced for the first time a data representation format tailored to systolic arrays, called CSXD, which further enhances storage and computational efficiency. Importantly, our adaptation scheme enables acceleration benefits even with limited resources. Through sparse packing, SPSA achieved a$5.2\times $performance improvement compared to the dense baseline, and further reached a$6.4\times $enhancement via CSXD. Simultaneously, CSXD realized an average storage efficiency improvement of$15.0\times $. Through extensive evaluations, SPSA outperforms previous designs on CPU, GPU, and ASIC platforms. Finally, in end-to-end evaluations, SPSA achieved a performance improvement of 3.9 times across the workloads of BERT, VGG19, and ResNet50. Minjin Tang, Mei Wen, Jianchao Yang, Zeyu Xue, Junzhong Shen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Localization of Ground-Based Periodic Pulse Interferers Using Time Difference of Arrival Estimation in SAR Satellite Systems
Shengqi Zhou, Xingyu Lu 0003, Jianchao Yang, Huizhang Yang, Junpeng Du, Lunhao Duan, Wenchao Yu, Ke Tan 0007, Shaojia Ge, Hong Gu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | SSpMM: Efficiently Scalable SpMM Kernels Across Multiple Generations of Tensor CoresabstractSparse-Dense Matrix-Matrix Multiplication (SpMM) has emerged as a foundational primitive in HPC and AI. Recent advancements have aimed to accelerate SpMM by harnessing the powerful Tensor Cores found in modern GPUs. However, despite these efforts, existing methods frequently encounter performance degradation when ported across different Tensor Core architectures. Recognizing that scalable SpMM across multiple generations of Tensor Cores relies on the effective use of general-purpose instructions, we have meticulously developed a SpMM library named SSpMM. However, a significant conflict exists between granularity and performance in current Tensor Core instructions. To resolve this, we introduce the innovative Transpose Mapping Scheme, which elegantly implements fine-grained kernels using coarse-grained instructions. Additionally, we propose the Register Shuffle Method to further enhance performance. Finally, we introduce Sparse Vector Compression, a technique that ensures our kernels are scalable with both structured and unstructured sparsity. Our experimental results, conducted on four generations of Tensor Core GPUs using over 3,000 sparse matrices from well established matrix collections, demonstrate that SSpMM achieves an average speedup of 2.04×, 2.81×, 2.07×, and 1.87×, respectively, over the state-of-the-art SpMM solution. Furthermore, we have integrated SSpMM into PyTorch, achieving a 1.81× speedup in end-to-end Transformer inference compared to cuDNN. Zeyu Xue, Mei Wen, Jianchao Yang, Minjin Tang, Zhongdi Luo, Yang Shi 0008, Zhaoyun Chen, Junzhong Shen, Johannes Langguth |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | A New Method of Noise Frequency Modulated Interference Suppression for SARabstractSynthetic aperture radar (SAR) is vulnerable to interference, including intentional and unintentional-ones. Noise frequency modulated (FM) interference is a kind of intentional interference, which has the characteristics of broadband and randomness, which makes the noise FM signal become a kind of most commonly used interference signal. Noise FM interference will have a serious impact on the SAR image, but the current algorithms for interference suppression are not sufficiently studied. This paper extends a time-domain cancellation algorithm for suppressing the noise FM interference of SAR. This algorithm can reconstruct the noise FM interference signal from the contaminated SAR echo, and then suppress the interference component in the echo by time-domain cancellation. Finally, this paper validates the superior performance of the algorithm by point target simulation and Radarsat-1 data. The proposed method is valid even when the signal-to-interference ratio is lower than -40dB. Lunhao Duan, Xingyu Lu 0003, Shengqi Zhou, Jianchao Yang, Ke Tan 0007, Zheng Dai, Wenchao Yu, Hong Gu 0002 |
IGARSS | 4 |
| 2024 | A Multi-Frame Super-Resolution Imaging Method for Forward-Looking Scanning RadarabstractSuper resolution technology has played a significant role in enhancing the imaging resolution of forward-looking scanning radar. However, a large number of super-resolution methods still rely on single frame scanning echoes. This paper aims to leverage multi-frame real beam images for super-resolution imaging, utilizing the complementary information present in multiple images to construct a higher resolution image. This paper first establishes the multi-frame super-resolution imaging model. Subsequently, a feasible multi-frame super-resolution method was proposed, and motion parameter estimation was performed using the correlated phase method. Finally, the effectiveness of the proposed method was verified through simulation experiments. Ke Tan 0007, Shengqi Zhou, Xingyu Lu 0003, Jianchao Yang, Hong Gu 0002 |
IGARSS | 4 |
| 2024 | RFI Source Localization for SAR: Method and Experiment based on GaoFen-3abstractThe signal emitted by ground radiation sources often interferes with Synthetic Aperture Radar (SAR) satellites, with the most common interference being periodic pulses emitted by ground radars. This paper proposes a method for locating ground-based periodic pulse signal interference sources using SAR echo data. Firstly, We estimate the Time Difference of Arrival (TDOA) of each pulse emitted by the interference source to SAR from the received SAR signals, and we seek the mapping relationship between the coordinates of the interference source (latitude and longitude) and the variations in TDOA. Using this mapping relationship, we achieve the localization of the interference source through a two-dimensional search method. The proposed method in this paper is highly versatile, applicable to single-station SAR satellites, multi-station SAR, and single-station SAR with multiple passes. It is also applicable regardless of the modulation form of the interference signal. Finally, the proposed TDOA-based localization method is experimentally validated for its accuracy based on GaoFen-3 satellite-borne SAR. The results demonstrate that the positioning error using two measurements from the satellite is only 3.708 km. Shengqi Zhou, Jingqiao Wang, Junpeng Du, Xingyu Lu 0003, Jianchao Yang, Ke Tan 0007, Hong Gu 0002 |
IGARSS | 8 |
| 2024 | HyFiSS: A Hybrid Fidelity Stall-Aware Simulator for GPGPUsabstractThe widespread adoption of GPUs has driven the development of GPU simulators, which, in turn, lead advancements in both GPU architectures and software optimization. Trace-driven cycle-accurate Cycle-accurate simulators, which provide detailed microarchitectural models and clock-level precision, come at the cost of extended simulation times and require high computational resources. Their scalability has become a bottleneck. A growing trend is the adoption of cycle-approximate simulators, which introduce mathematical modeling of partial hardware units and utilize sampling to accelerate simulation. However, this approach faces challenges regarding the accuracy of performance predictions. To address these limitations, we introduce HyFiSS, a hybrid fidelity stall-aware GPU simulator. HyFiSS features fine-grained stall events tracking and attribution by constructing a detailed execution pipeline model for various stall events on Streaming Multiprocessors (SMs). It accurately emulates the thread block scheduler behavior using real-time scheduling logs and utilizes sampling based on thread block sets to minimize the precision loss due to fine-grained sampling points on the microarchitectural state. We achieve a balance between reliability, speed, and the level of simulation detail, especially regarding bottlenecks. By evaluating a diverse set of benchmarks, HyFiSS achieves a mean absolute percentage error in predicting active cycles that is comparable to the state-of-the-art cycle-accurate simulator Accel-Sim. Moreover, HyFiSS achieves a substantial 12.8 × speedup in the simulation efficiency compared to Accel-Sim. HyFiSS also requires at least 3.2 × less disk storage than both Accel-Sim and another state-of-the-art cycle-approximate simulator PPT-GPU due to its efficient SASS (Streaming Assembler) traces compression. With precise, per-cycle stall events statistics, HyFiSS can provide accurate GPU performance metrics and stall cause reporting. This significantly simplifies performance analysis, bottleneck identification, and performance optimization tasks for researchers, making it easier to enhance GPU performance effectively. Jianchao Yang, Mei Wen, Dong Chen 0015, Zhaoyun Chen, Zeyu Xue, Junzhong Shen, Yang Shi 0008 |
MICRO | 1 |
| 2023 | Releasing the Potential of Tensor Core for Unstructured SpMM using Tiled-CSR FormatabstractThe GPU has become a popular platform for AI applications, thanks in part to its Tensor Cores that address performance issues. However, the Sparse Matrix Multiplication (SpMM) kernel has remained a bottleneck despite significant advances in computing power. Due to the hardware mechanism of the Tensor Core, its programming granularity does not match SpMM. In this paper, we analyze the reasons why the unstructured SpMM kernel is not suitable for the Tensor Core, and propose the Tiled Compressed Sparse Row (Tiled-CSR) compression format. To address the issue of low non-zero rates in Tiled-CSR format, we exploit the row shuffle algorithm to improve the utilization of Tensor Cores and enhance computing density. We also utilize adaptive memory access modes and 3D-Grid tiling for the SpMM kernel to reduce memory access latency. The experimental results on NVIDIA A100 GPU with matrices in the Deep Learning Matrix Collection (DLMC) demonstrate that the Tiled-CSR format improves the utilization of Tensor Cores under different sparsity, with a maximum of 3.89× at 50% sparsity and a minimum of 1.82× at 90% sparsity compared to the SR-BCRS format. Additionally, our kernel achieves an average speedup of 1.54×(up to 2.12×) over Magicube. Zeyu Xue, Mei Wen, Zhaoyun Chen, Yang Shi 0008, Minjin Tang, Jianchao Yang, Zhongdi Luo |
ICCD | 6 |
| 2022 | Dressing in the Wild by Watching Dance VideosabstractWhile significant progress has been made in garment transfer, one of the most applicable directions of human-centric image generation, existing works overlook the in-the-wild imagery, presenting severe garment-person mis-alignment as well as noticeable degradation in fine texture details. This paper, therefore, attends to virtual try-on in real-world scenes and brings essential improvements in authenticity and naturalness especially for loose garment (e.g., skirts, formal dresses), challenging poses (e.g., cross arms, bent legs), and cluttered backgrounds. Specifically, we find that the pixel flow excels at handling loose gar-ments whereas the vertex flow is preferred for hard poses, and by combining their advantages we propose a novel generative network called wFlow that can effectively push up garment transfer to in-the-wild context. Moreover, former approaches require paired images for training. Instead, we cut down the laboriousness by working on a newly constructed large-scale video dataset named Dance50k with self-supervised cross-frame training and an online cycle op-timization. The proposed Dance50k can boost real-world virtual dressing by covering a wide variety of garments under dancing poses. Extensive experiments demonstrate the superiority of our w Flow in generating realistic garment transfer results for in-the-wild images without resorting to expensive paired datasets.11Xiaodan Liang is the corresponding author. The project page of wFlow is https://awesome-wflow.github.io. Fuwei Zhao, Zhenyu Xie, Xijin Zhang, Daniel K. Du, Xiang Long, Xiaodan Liang, Jianchao Yang |
CVPR | 9 |
| 2022 | Light: A Component Enhances Faster and More Accurate Traffic Measurement*abstractThe greatest challenge when designing an online sketch method for data flow measurement is to reduce the storage cost of sketches with little loss of accuracy and obtain a higher bandwidth. To address this issue, we proposed Light component. By storing elephant and mice flows separately, the accuracy of the Light-enhanced sketches is substantially improved, and the processing speed is significantly faster due to a significant reduction in the required computational overhead. We implement Light-enhanced sketches alongside several existing sketch methods on CPU and compare their performance. Experiments show that, under the same storage conditions, Light-enhanced sketches can greatly reduce the Average Relative Error by 1.80 to 5.07 times, as well as increase the average processing speed of each packet by 4.96 to 9.05 times compared with their original structure. Our approach also achieves a stable performance in the measurement of traffic in different traffic distributions. More importantly, Light component can be deployed to different sketch methods, which demonstrates its reusability. Jianchao Yang, Mei Wen, Yang Shi 0008 |
ICC | 1 |
| 2022 | BP-Im2col: Implicit Im2col Supporting AI Backpropagation on Systolic ArraysabstractState-of-the-art systolic array-based accelerators adopt the traditional im2col algorithm to accelerate the inference of convolutional layers. However, traditional im2col cannot efficiently support AI backpropagation. Backpropagation in convolutional layers involves performing transposed convolution and dilated convolution, which usually introduces plenty of zero-spaces into the feature map or kernel. The zero-space data reorganization interfere with the continuity of training and incur additional and non-negligible overhead in terms of off- and on-chip storage, access and performance. Since countermeasures for backpropagation are rarely proposed, we propose BP-im2col, a novel im2col algorithm for AI backpropagation, and implement it in RTL on a TPU-like accelerator. Experiments on TPU-like accelerator indicate that BP-im2col reduces the backpropagation runtime by 34.9% on average, and reduces the bandwidth of off-chip memory and on-chip buffers by at least 22.7% and 70.6% respectively, over a baseline accelerator adopting the traditional im2col. It further reduces the additional storage overhead in the backpropagation process by at least 74.78%. Jianchao Yang, Mei Wen, Junzhong Shen, Yasong Cao, Minjin Tang, Renyu Yang, Jiawei Fei, Chunyuan Zhang |
ICCD | 1 |
| 2022 | Mentha: Enabling Sparse-Packing Computation on Systolic ArraysabstractGeneralized Sparse Matrix-Matrix Multiplication (SpGEMM) is a critical kernel in domains like graph analytic and scientific computation. As a kind of classical special-purpose architecture, systolic arrays were first used for complex computing problems, e.g., matrix multiplication. However, classical systolic arrays are not efficient enough when handling sparse matrices due to the fact that the PEs containing zero-valued entries perform unnecessary operations that do not contribute to the result. Accordingly, in this paper, we propose Mentha, a framework that enables systolic arrays to accelerate sparse matrix computation by employing a sparse-packing algorithm suitable for various dataflow of systolic array. Firstly, Mentha supports both online and offline methods. By packing the rows or columns of the sparse matrix, the zero-valued items in the matrix are significantly reduced and the density of the matrix is improved. In addition, acceleration benefits can be obtained by the adaptation scheme even with limited resources. Moreover, we reconfigure PEs in systolic arrays at a low cost (1.28x in area, 1.21x in power) and find that our method outperforms TPU-like systolic arrays by 1.2~3.3x in terms of SpMM and 1.3~4.4x in terms of SpGEMM when dealing with moderately sparse matrices (sparsity < 0.9), while its performance is at least 9.7x better than cuSPARSE. Furthermore, experimental results show a FLOPs reduction of roughly 3.4x in the neural network. Minjin Tang, Mei Wen, Yasong Cao, Junzhong Shen, Jianchao Yang, Jiawei Fei, Yang Guo 0003, Sheng Liu 0001 |
ICPP | 5 |
| 2022 | Accurate SAR Image Recovery From RFI Contaminated Raw Data by Using Image Domain Mixed RegularizationsabstractRadio frequency interference (RFI) suppression is a hot topic in synthetic aperture radar (SAR) imaging. Mathematically, the RFI suppression problem can be considered as an underdetermined signal separation problem to extract the signal of interest (SOI) from the RFI contaminated raw data. The regularization-based method can exploit both the prior knowledge of RFI and SOI and, therefore, has the advantage of solving the underdetermined problem and preserving the information of SOI. Current regularization methods make use of the RFI prior well by exploiting low-rank representation (LRR) or sparse representation (SR), but the prior knowledge of SOI has not been sufficiently studied and used. In some literature, the sparsity of the raw data or range profile was exploited to formulate the regularization term, which we found to be inadequate in describing the SOI property. In this article, we explore the features of SAR images and propose an RFI suppression model with a combination of multiple image domain regularizations to preserve different types of targets. An efficient solution to the optimization problem is proposed based on the alternating direction multiplier method (ADMM). The proposed method can accurately recover both the sparse strong targets, and the nonsparse regions in the illuminated area and its performance is validated by measured data. Xingyu Lu 0003, Jianchao Yang, Tat Soon Yeo, Hong Gu 0002, Wenchao Yu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Online Multi-Granularity Distillation for GAN CompressionabstractGenerative Adversarial Networks (GANs) have witnessed prevailing success in yielding outstanding images, however, they are burdensome to deploy on resource-constrained devices due to ponderous computational costs and hulking memory usage. Although recent efforts on compressing GANs have acquired remarkable results, they still exist potential model redundancies and can be further compressed. To solve this issue, we propose a novel online multi-granularity distillation (OMGD) scheme to obtain lightweight GANs, which contributes to generating highfidelity images with low computational demands. We offer the first attempt to popularize single-stage online distillation for GAN-oriented compression, where the progressively promoted teacher generator helps to refine the discriminator-free based student generator. Complementary teacher generators and network layers provide comprehensive and multi-granularity concepts to enhance visual fidelity from diverse dimensions. Experimental results on four benchmark datasets demonstrate that OMGD successes to compress 40× MACs and 82.5× parameters on Pix2Pix and CycleGAN, without loss of image quality. It reveals that OMGD provides a feasible solution for the deployment of real-time image translation on resource-constrained devices. Our code and models are made public at: https://github.com/bytedance/OMGD Yuxi Ren, Jie Wu 0032, Xuefeng Xiao 0001, Jianchao Yang |
ICCV | 4 |
| 2021 | A Super-Resolution Imaging Method for Real-Aperture Scanning Radar Based on MRF Prior ModelabstractDeconvolution technology can be utilized to improve the angular resolution of real-aperture scanning radar (RASR) with high efficiency and low cost. However, it is an ill-posed problem and the solution is sensitive to noise. Regularization methods are considered to be efficient ways to ease the noise sensitivity by absorbing the prior information into the objective function. In this paper, we propose a new super-resolution imaging method for RASA based on the Markov random field (MRF). Compared with the published angular super-resolution methods for RASA, the proposed method takes advantage of the two-dimensional spatial prior information and can recover the shape of scene much better. Simulations are carried out to demonstrate the effectiveness of the proposed method. Ke Tan 0007, Jianchao Yang, Xingyu Lu 0003, Weiming Su, Hong Gu 0002 |
IGARSS | 2 |
| 2021 | Autofocus Method for Sparse Aperture ISAR Based on L0 Norm and NLTV RegularizationabstractAutofocus is one of the key problems in inverse synthetic aperture radar (ISAR) since the noncooperation of the target motion. For sparse aperture ISAR, classical autofocus algorithms are not suitable due to the discontinuity of the azimuth sampling. In this paper, a novel framework is proposed for ISAR autofocus with sparse aperture. The autofocus problem is transformed into an optimization problem with l0norm and nonlocal total variation (NLTV) regularization constraints. Therefore, both spatial sparsity and structural information of the target can be considered in the process. Dual iterative computation which combines regularization method and conjugate gradient (CG) algorithm is applied to reconstruct the image and correct the phase error. Results of real data experiments show the effectiveness of the proposed method. Jianchao Yang, Xingyu Lu 0003, Zheng Dai, Ke Tan 0007, Wenchao Yu |
IGARSS | 1 |
| 2021 | Enhanced LRR-Based RFI Suppression for SAR Imaging Using the Common Sparsity of Range Profiles for Accurate Signal RecoveryabstractThe performance of synthetic aperture radar is vulnerable to radio frequency interference (RFI). In many situations, the RFI has a low-rank property, since the frequency bands occupied by RFI usually remain stable during a short slow time period. Therefore, low-rank representation (LRR)-based methods can be applied to separate RFI and signal of interest (SOI), by minimizing the rank of RFI components with a regularization constraint to protect SOI. However, traditional methods use the sparsity of the raw data or range profile to formulate the regularization term, which fails to describe the properties of SOI accurately. In addition to the sparse property of range profiles, this article explores the common patterns hidden in the range profiles and proposes two new LRR-based RFI suppression optimization models with a well-designed regularization term to describe such common sparsity to protect the SOI. Four methods are proposed to solve the optimization problems based on the alternating direction multiplier (ADM) method, which provides tradeoff between efficiency and accuracy. Compared with traditional LRR-based RFI suppression methods, the proposed methods make a more precise description of the features of SOI, therefore can better protect the information of SOI during the RFI suppression process and improves the imaging quality. The superior performance of the proposed method is validated by measured data in both sparse and nonsparse scenes. Xingyu Lu 0003, Jianchao Yang, Wenchao Yu, Hong Gu 0002, Tat Soon Yeo |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | EnlightenGAN: Deep Light Enhancement Without Paired SupervisionabstractDeep learning-based methods have achieved remarkable success in image restoration and enhancement, but are they still competitive when there is a lack of paired training data? As one such example, this paper explores the low-light image enhancement problem, where in practice it is extremely challenging to simultaneously take a low-light and a normal-light photo of the same visual scene. We propose a highly effective unsupervised generative adversarial network, dubbed EnlightenGAN, that can be trained without low/normal-light image pairs, yet proves to generalize very well on various real-world test images. Instead of supervising the learning using ground truth data, we propose to regularize the unpaired training using the information extracted from the input itself, and benchmark a series of innovations for the low-light image enhancement problem, including a global-local discriminator structure, a self-regularized perceptual loss fusion, and the attention mechanism. Through extensive experiments, our proposed approach outperforms recent methods under a variety of metrics in terms of visual quality and subjective user study. Thanks to the great flexibility brought by unpaired training, EnlightenGAN is demonstrated to be easily adaptable to enhancing real-world images from various domains. Our codes and pre-trained models are available at: https://github.com/VITA-Group/EnlightenGAN. Yifan Jiang 0001, Xinyu Gong, Ding Liu 0001, Yu Cheng 0001, Xiaohui Shen, Jianchao Yang, Pan Zhou 0001, Zhangyang Wang |
IEEE Trans. Image Process. | 7 |
| 2020 | AtomNAS: Fine-Grained End-to-End Neural Architecture Search
Jieru Mei, Yingwei Li 0002, Xiaochen Lian, Xiaojie Jin 0004, Alan L. Yuille, Jianchao Yang |
ICLR | 7 |
| 2020 | Neural Epitome Search for Architecture-Agnostic Network Compression
Daquan Zhou, Xiaojie Jin 0004, Qibin Hou, Jianchao Yang, Jiashi Feng |
ICLR | 5 |
| 2020 | An Efficient Method for Single-Channel SAR Target Reconstruction Under Severe Deceptive JammingabstractDeceptive jamming can severely degrade synthetic aperture radar (SAR) image quality by introducing high-fidelity false targets. In this letter, a simultaneous deceptive jamming suppression and target reconstruction method for a single-channel SAR system is proposed. The signal model is formulated by constructing a joint dictionary based on different time-frequency distributions of the actual targets and false targets. Then, an efficient algorithm is proposed based on the alternating direction method of multipliers (ADMMs) to simultaneously recover the actual and false targets. Several strategies are also proposed to accelerate the computation. Compared with other existing single-channel SAR deceptive jamming suppression methods, the proposed method has lower reconstruction error and computational load. Simulation results demonstrate the superior performance of the proposed algorithm. Xingyu Lu 0003, Yujiu Zhao, Jianchao Yang, Hong Gu 0002, Tat Soon Yeo |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Weak maneuvering target detection in random pulse repetition interval radar
Wenchao Yu, Jianchao Yang |
Signal Process. | 4 |
| 2019 | EIGEN: Ecologically-Inspired GENetic Approach for Neural Network Structure Searching From ScratchabstractDesigning the structure of neural networks is considered one of the most challenging tasks in deep learning, especially when there is few prior knowledge about the task domain. In this paper, we propose an Ecologically-Inspired GENetic (EIGEN) approach that uses the concept of succession, extinction, mimicry, and gene duplication to search neural network structure from scratch with poorly initialized simple network and few constraints forced during the evolution, as we assume no prior knowledge about the task domain. Specifically, we first use primary succession to rapidly evolve a population of poorly initialized neural network structures into a more diverse population, followed by a secondary succession stage for fine-grained searching based on the networks from the primary succession. Extinction is applied in both stages to reduce computational cost. Mimicry is employed during the entire evolution process to help the inferior networks imitate the behavior of a superior network and gene duplication is utilized to duplicate the learned blocks of novel structures, both of which help to find better network structures. Experimental results show that our proposed approach can achieve similar or better performance compared to the existing genetic approaches with dramatically reduced computation cost. For example, the network discovered by our approach on CIFAR-100 dataset achieves 78.1% test accuracy under 120 GPU hours, compared to 77.0% test accuracy in more than 65, 536 GPU hours in [35]. Jian Ren 0005, Zhe Li 0008, Jianchao Yang, Ning Xu 0001, Tianbao Yang, David J. Foran |
CVPR | 3 |
| 2019 | Slimmable Neural Networks
Ning Xu 0001, Jianchao Yang, Thomas S. Huang |
ICLR (Poster) | 4 |
| 2019 | Action Recognition With Spatio-Temporal Visual Attention on Skeleton Image SequencesabstractAction recognition with 3D skeleton sequences became popular due to its speed and robustness. The recently proposed convolutional neural networks (CNNs)-based methods show a good performance in learning spatio-temporal representations for skeleton sequences. Despite the good recognition accuracy achieved by previous CNN-based methods, there existed two problems that potentially limit the performance. First, previous skeleton representations were generated by chaining joints with a fixed order. The corresponding semantic meaning was unclear and the structural information among the joints was lost. Second, previous models did not have an ability to focus on informative joints. The attention mechanism was important for skeleton-based action recognition because different joints contributed unequally toward the correct recognition. To solve these two problems, we proposed a novel CNN-based method for skeleton-based action recognition. We first redesigned the skeleton representations with a depth-first tree traversal order, which enhanced the semantic meaning of skeleton images and better preserved the associated structural information. We then proposed the general two-branch attention architecture that automatically focused on spatio-temporal key stages and filtered out unreliable joint predictions. Based on the proposed general architecture, we designed a global long-sequence attention network with refined branch structures. Furthermore, in order to adjust the kernel's spatio-temporal aspect ratios and better capture long-term dependencies, we proposed a sub-sequence attention network (SSAN) that took sub-image sequences as inputs. We showed that the two-branch attention architecture could be combined with the SSAN to further improve the performance. Our experiment results on the NTU RGB+D data set and the SBU kinetic interaction data set outperformed the state of the art. The model was further validated on noisy estimated poses from the subsets of the UCF101 data set and the kinetics data set. Zhengyuan Yang, Yuncheng Li, Jianchao Yang, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Efficient Video Object Segmentation via Network ModulationabstractVideo object segmentation targets segmenting a specific object throughout a video sequence when given only an annotated first frame. Recent deep learning based approaches find it effective to fine-tune a general-purpose segmentation model on the annotated frame using hundreds of iterations of gradient descent. Despite the high accuracy that these methods achieve, the fine-tuning process is inefficient and fails to meet the requirements of real world applications. We propose a novel approach that uses a single forward pass to adapt the segmentation model to the appearance of a specific object. Specifically, a second meta neural network named modulator is trained to manipulate the intermediate layers of the segmentation network given limited visual and spatial information of the target object. The experiments show that our approach is 70× faster than fine-tuning approaches and achieves similar accuracy. Our model and code have been released at https://github.com/linjieyangsc/video_seg. Yanran Wang 0002, Xuehan Xiong, Jianchao Yang, Aggelos K. Katsaggelos |
CVPR | 4 |
| 2018 | YouTube-VOS: Sequence-to-Sequence Video Object Segmentation
Ning Xu 0007, Yuchen Fan 0001, Jianchao Yang, Dingcheng Yue, Brian L. Price, Scott Cohen, Thomas S. Huang |
ECCV (5) | 4 |
| 2018 | A Hybrid Neural Network for Chroma Intra PredictionabstractFor chroma intra prediction, previous methods exemplified by the Linear Model method (LM) usually assume a linear correlation between the luma and chroma components in a coding block. This assumption is inaccurate for complex image content or large blocks, and restricts the prediction accuracy. In this paper, we propose a chroma intra prediction method by exploiting both spatial and cross-channel correlations using a hybrid neural network. Specifically, we utilize a convolutional neural network to extract features from the reconstructed luma samples of the current block, as well as utilize a fully connected network to extract features from the neighboring reconstructed luma and chroma samples. The extracted twofold features are then fused to predict the chroma samples-Cb and Cr simultaneously. The proposed chroma intra prediction method is integrated into HEVC. Preliminary results show that, compared with HEVC plus LM, the proposed method achieves on average 0.2%, 3.1% and 2.0% BD-rate reduction on Y, Cb and Cr components, respectively, under All-Intra configuration. Yue Li 0015, Li Li 0040, Zhu Li 0001, Jianchao Yang, Ning Xu 0001, Dong Liu 0002, Houqiang Li |
ICIP | 4 |
| 2018 | WSNet: Compact and Efficient Networks Through Weight SamplingabstractWe present a new approach and a novel architecture, termed WSNet, for learning compact and efficient deep neural networks. Existing approaches conventionally learn full model parameters independently and then compress them via ad hoc processing such as model pruning or filter factorization. Alternatively, WSNet proposes learning model parameters by sampling from a compact set of learnable parameters, which naturally enforces parameter sharing throughout the learning process. We demonstrate that such a novel weight sampling approach (and induced WSNet) promotes both weights and computation sharing favorably. By employing this method, we can more efficiently learn much smaller networks with competitive performance compared to baseline networks with equal numbers of convolution filters. Specifically, we consider learning compact and efficient 1D convolutional neural networks for audio classification. Extensive experiments on multiple audio classification datasets verify the effectiveness of WSNet. Combined with weight quantization, the resulted models are up to 180x smaller and theoretically up to 16x faster than the well-established baselines, without noticeable performance drop. Xiaojie Jin 0004, Yingzhen Yang, Ning Xu 0001, Jianchao Yang, Nebojsa Jojic, Jiashi Feng, Shuicheng Yan |
ICML | 4 |
| 2018 | Action Recognition with Visual Attention on Skeleton ImagesabstractAction recognition with 3D skeleton sequences is becoming popular due to its speed and robustness. The recently proposed Convolutional Neural Networks (CNN) based methods have shown good performance in learning spatio-temporal representations for skeleton sequences. Despite the good recognition accuracy achieved by previous CNN based methods, there exist two problems that potentially limit the performance. First, previous skeleton representations are generated by chaining joints with a fixed order. The corresponding semantic meaning is unclear and the structural information among the joints is lost. Second, previous models do not have an ability to focus on informative joints. The attention mechanism is important for skeleton based action recognition because there exist spatio-temporal key stages and the joint predictions can be inaccurate. To solve the two problems, we propose a novel CNN based method for skeleton based action recognition. We first redesign the skeleton representations with a depth-first tree traversal order, which enhances the semantic meaning of skeleton images and better preserves the structural information. We then propose the idea of a two-branch attention architecture that focuses on spatio-temporal key stages and filters out unreliable joint predictions. A base attention model with the simplest structure is first introduced to illustrate the two-branch attention architecture. By improving the structures in both branches, we further propose a Global Long-sequence Attention Network (GLAN). Experiment results on the NTU RGB+D dataset and the SBU Kinetic Interaction dataset show that our proposed approach outperforms the state-of-the-art, as well as the effectiveness of each component. Zhengyuan Yang, Yuncheng Li, Jianchao Yang, Jiebo Luo 0001 |
ICPR | 3 |
| 2018 | Subspace Learning by ℓ0-Induced Sparsity
Yingzhen Yang, Jiashi Feng, Nebojsa Jojic, Jianchao Yang, Thomas S. Huang |
Int. J. Comput. Vis. | 4 |
| 2018 | Local patch encoding-based method for single image super-resolution
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Jianchao Yang, Wenmin Wang 0001, Wen Gao 0001 |
Inf. Sci. | 4 |
| 2018 | Proposal-Free Network for Instance-Level Object SegmentationabstractInstance-level object segmentation is an important yet under-explored task. Most of state-of-the-art methods rely on region proposal methods to extract candidate segments and then utilize object classification to produce final results. Nonetheless, generating reliable region proposals itself is a quite challenging and unsolved task. In this work, we propose a Proposal-Free Network (PFN) to address the instance-level object segmentation problem, which outputs the numbers of instances of different categories and the pixel-level information on i) the coordinates of the instance bounding box each pixel belongs to, and ii) the confidences of different categories for each pixel, based on pixel-to-pixel deep convolutional neural network. All the outputs together, by using any off-the-shelf clustering method for simple post-processing, can naturally generate the ultimate instance-level object segmentation results. The whole PFN can be easily trained without the requirement of a proposal generation stage. Extensive evaluations on the challenging PASCAL VOC 2012 semantic segmentation benchmark demonstrate the effectiveness of the proposed PFN solution without relying on any proposal generation methods. Xiaodan Liang, Liang Lin 0004, Yunchao Wei, Xiaohui Shen, Jianchao Yang, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | Robust emotion recognition from low quality and low bit rate video: A deep learning approachabstractEmotion recognition from facial expressions is tremendously useful, especially when coupled with smart devices and wireless multimedia applications. However, the inadequate network bandwidth often limits the spatial resolution of the transmitted video, which will heavily degrade the recognition reliability. We develop a novel framework to achieve robust emotion recognition from low bit rate video. While video frames are downsampled at the encoder side, the decoder is embedded with a deep network model for joint super-resolution (SR) and recognition. Notably, we propose a novel max-mix training strategy, leading to a single “One-for-All” model that is remarkably robust to a vast range of downsampling factors. That makes our framework well adapted for the varied bandwidths in real transmission scenarios, without hampering scalability or efficiency. The proposed framework is evaluated on the AVEC 2016 benchmark, and demonstrates significantly improved stand-alone recognition performance, as well as rate-distortion (R-D) performance, than either directly recognizing from LR frames, or separating SR and recognition. Bowen Cheng, Zhangyang Wang, Zhaobin Zhang, Zhu Li 0001, Ding Liu 0001, Jianchao Yang, Shuai Huang 0001, Thomas S. Huang |
ACII | 6 |
| 2017 | Dense Captioning with Joint Inference and Visual ContextabstractDense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions. The goal is to densely detect visual concepts (e.g., objects, object parts, and interactions between them) from images, labeling each with a short descriptive phrase. We identify two key challenges of dense captioning that need to be properly addressed when tackling the problem. First, dense visual concept annotations in each image are associated with highly overlapping target regions, making accurate localization of each visual concept challenging. Second, the large amount of visual concepts makes it hard to recognize each of them by appearance alone. We propose a new model pipeline based on two novel ideas, joint inference and context fusion, to alleviate these two challenges. We design our model architecture in a methodical manner and thoroughly evaluate the variations in architecture. Our final model, compact and efficient, achieves state-of-the-art accuracy on Visual Genome [23] for dense captioning with a relative gain of 73% compared to the previous best algorithm. Qualitative experiments also reveal the semantic capabilities of our model in dense captioning. Kevin D. Tang, Jianchao Yang, Li-Jia Li 0001 |
CVPR | 3 |
| 2017 | Learning from Noisy Labels with DistillationabstractThe ability of learning from noisy labels is very useful in many visual recognition tasks, as a vast amount of data with noisy labels are relatively easy to obtain. Traditionally, label noise has been treated as statistical outliers, and techniques such as importance re-weighting and bootstrapping have been proposed to alleviate the problem. According to our observation, the real-world noisy labels exhibit multimode characteristics as the true labels, rather than behaving like independent random outliers. In this work, we propose a unified distillation framework to use “side” information, including a small clean dataset and label relations in knowledge graph, to “hedge the risk” of learning from noisy labels. Unlike the traditional approaches evaluated based on simulated label noises, we propose a suite of new benchmark datasets, in Sports, Species and Artifacts domains, to evaluate the task of learning from noisy labels in the practical setting. The empirical study demonstrates the effectiveness of our proposed method in all the domains. Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao, Jiebo Luo 0001, Li-Jia Li 0001 |
ICCV | 2 |
| 2017 | Support Regularized Sparse Coding and Its Fast Encoder
Yingzhen Yang, Pushmeet Kohli, Jianchao Yang, Thomas S. Huang |
ICLR (Poster) | 4 |
| 2017 | First International ACM Thematic Workshops 2017abstractAs a new addition, this year the ACM Multimedia conference is introducing Thematic Workshops. Inspired by the NIPS model, Thematic Workshops allow papers that could not be accommodated in the main conference to be presented to the research community. Wanmin Wu, Jianchao Yang, Qi Tian 0001, Roger Zimmermann |
ACM Multimedia | 2 |
| 2017 | Neighborhood Regularized l^1-Graph
Yingzhen Yang, Jiashi Feng, Jianchao Yang, Thomas S. Huang |
UAI | 4 |
| 2017 | Learning to Segment Human by Watching YouTubeabstractAn intuition on human segmentation is that when a human is moving in a video, the video-context (e.g., appearance and motion clues) may potentially infer reasonable mask information for the whole human body. Inspired by this, based on popular deep convolutional neural networks (CNN), we explore a very-weakly supervised learning framework for human segmentation task, where only an imperfect human detector is available along with massive weakly-labeled YouTube videos. In our solution, the video-context guided human mask inference and CNN based segmentation network learning iterate to mutually enhance each other until no further improvement gains. In the first step, each video is decomposed into supervoxels by the unsupervised video segmentation. The superpixels within the supervoxels are then classified as human or non-human by graph optimization with unary energies from the imperfect human detection results and the predicted confidence maps by the CNN trained in the previous iteration. In the second step, the video-context derived human masks are used as direct labels to train CNN. Extensive experiments on the challenging PASCAL VOC 2012 semantic segmentation benchmark demonstrate that the proposed framework has already achieved superior results than all previous weakly-supervised methods with object class or bounding box annotations. In addition, by augmenting with the annotated masks from PASCAL VOC 2012, our method reaches a new state-of-the-art performance on the human segmentation task. Xiaodan Liang, Yunchao Wei, Liang Lin 0004, Yunpeng Chen, Xiaohui Shen, Jianchao Yang, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2017 | Human Parsing with Contextualized Convolutional Neural NetworkabstractIn this work, we address the human parsing task with a novel Contextualized Convolutional Neural Network (Co-CNN) architecture, which well integrates the cross-layer context, global image-level context, semantic edge context, within-super-pixel context and cross-super-pixel neighborhood context into a unified network. Given an input human image, Co-CNN produces the pixelwise categorization in an end-to-end way. First, the cross-layer context is captured by our basic local-to-global-to-local structure, which hierarchically combines the global semantic information and the local fine details across different convolutional layers. Second, the global image-level label prediction is used as an auxiliary objective in the intermediate layer of the Co-CNN, and its outputs are further used for guiding the feature learning in subsequent convolutional layers to leverage the global image-level context. Third, semantic edge context is further incorporated into Co-CNN, where the high-level semantic boundaries are leveraged to guide pixel-wise labeling. Finally, to further utilize the local super-pixel contexts, the within-super-pixel smoothing and cross-super-pixel neighbourhood voting are formulated as natural sub-components of the Co-CNN to achieve the local label consistency in both training and testing process. Comprehensive evaluations on two public datasets well demonstrate the significant superiority of our Co-CNN over other state-of-the-arts for human parsing. In particular, the F-1 score on the large dataset [1] reaches 81.72 percent by Co-CNN, significantly higher than 62.81 percent and 64.38 percent by the state-of-the-art algorithms, M-CNN [2] and ATR [1], respectively. By utilizing our newly collected large dataset for training, our Co-CNN can achieve 85.36 percent in F-1 score. Xiaodan Liang, Chunyan Xu, Xiaohui Shen, Jianchao Yang, Jinhui Tang 0001, Liang Lin 0004, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Deep Edge Guided Recurrent Residual Learning for Image Super-ResolutionabstractIn this paper, we consider the image super-resolution (SR) problem. The main challenge of image SR is to recover high-frequency details of a low-resolution (LR) image that are important for human perception. To address this essentially ill-posed problem, we introduce a Deep Edge Guided REcurrent rEsidual (DEGREE) network to progressively recover the high-frequency details. Different from most of the existing methods that aim at predicting high-resolution (HR) images directly, the DEGREE investigates an alternative route to recover the difference between a pair of LR and HR images by recurrent residual learning. DEGREE further augments the SR process with edge-preserving capability, namely the LR image and its edge map can jointly infer the sharp edge details of the HR image during the recurrent recovery process. To speed up its training convergence rate, by-pass connections across the multiple layers of DEGREE are constructed. In addition, we offer an understanding on DEGREE from the view-point of sub-band frequency decomposition on image signal and experimentally demonstrate how the DEGREE can recover different frequency bands separately. Extensive experiments on three benchmark data sets clearly demonstrate the superiority of DEGREE over the well-established baselines and DEGREE also provides new state-of-the-arts on these data sets. We also present addition experiments for JPEG artifacts reduction to demonstrate the good generality and flexibility of our proposed DEGREE network to handle other image processing tasks. Wenhan Yang, Jiashi Feng, Jianchao Yang, Fang Zhao 0006, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan |
IEEE Trans. Image Process. | 3 |
| 2016 | Building a Large Scale Dataset for Image Emotion Recognition: The Fine Print and The BenchmarkabstractPsychological research results have confirmed that people can have different emotional reactions to different visual stimuli. Several papers have been published on the problem of visual emotion analysis. In particular, attempts have been made to analyze and predict people's emotional reaction towards images. To this end, different kinds of hand-tuned features are proposed. The results reported on several carefully selected and labeled small image data sets have confirmed the promise of such features. While the recent successes of many computer vision related tasks are due to the adoption of Convolutional Neural Networks (CNNs), visual emotion analysis has not achieved the same level of success. This may be primarily due to the unavailability of confidently labeled and relatively large image data sets for visual emotion analysis. In this work, we introduce a new data set, which started from 3+ million weakly labeled images of different emotions and ended up 30 times as large as the current largest publicly available visual emotion data set. We hope that this data set encourages further research on visual emotion analysis. We also perform extensive benchmarking analyses on this large data set using the state of the art methods including CNNs. Quanzeng You, Jiebo Luo 0001, Hailin Jin, Jianchao Yang |
AAAI | 4 |
| 2016 | ℓ ^0 ℓ 0 -Sparse Subspace Clustering
Yingzhen Yang, Jiashi Feng, Nebojsa Jojic, Jianchao Yang, Thomas S. Huang |
ECCV (2) | 4 |
| 2016 | Customized expression recognition for performance-driven cutout character animationabstractPerformance-driven character animation enables users to create expressive results by performing the desired motion of the character with their face and/or body. However, for cutout animations where continuous motion is combined with discrete artwork replacements, supporting a performance-driven workflow has some unique requirements. To trigger the appropriate artwork replacements, the system must reliably detect a wide range of customized facial expressions that are challenging for existing recognition methods, which focus on a few canonical expressions (e.g., angry, disgusted, scared, happy, sad and surprised). Also, real usage scenarios require the system to work in realtime with minimal training. In this paper, we propose a novel customized expression recognition technique that meets all of these requirements. We first use a set of handcrafted features combining geometric features derived from facial landmarks and patch-based appearance features through group sparsity-based facial component learning. To improve discrimination and generalization, these handcrafted features are integrated into a custom-designed Deep Convolutional Neural Network (CNN) structure trained from publicly available facial expression datasets. The combined features are fed to an online ensemble of SVMs designed for the few training sample problem and performs in realtime. To improve temporal coherence, we also apply a Hidden Markov Model (HMM) to smooth the recognition results. Our system achieves state-of-the-art performance on canonical expression datasets and promising results on our collected dataset of customized expressions. Xiang Yu 0002, Jianchao Yang, Linjie Luo, Wilmot Li, Jonathan Brandt, Dimitris N. Metaxas |
WACV | 2 |
| 2016 | Cross-modality Consistent Regression for Joint Visual-Textual Sentiment Analysis of Social MultimediaabstractSentiment analysis of online user generated content is important for many social media analytics tasks. Researchers have largely relied on textual sentiment analysis to develop systems to predict political elections, measure economic indicators, and so on. Recently, social media users are increasingly using additional images and videos to express their opinions and share their experiences. Sentiment analysis of such large-scale textual and visual content can help better extract user sentiments toward events or topics. Motivated by the needs to leverage large-scale social multimedia content for sentiment analysis, we propose a cross-modality consistent regression (CCR) model, which is able to utilize both the state-of-the-art visual and textual sentiment analysis techniques. We first fine-tune a convolutional neural network (CNN) for image sentiment analysis and train a paragraph vector model for textual sentiment analysis. On top of them, we train our multi-modality regression model. We use sentimental queries to obtain half a million training samples from Getty Images. We have conducted extensive experiments on both machine weakly labeled and manually labeled image tweets. The results show that the proposed model can achieve better performance than the state-of-the-art textual and visual sentiment analysis algorithms alone. Quanzeng You, Jiebo Luo 0001, Hailin Jin, Jianchao Yang |
WSDM | 4 |
| 2016 | Parsing Based on Parselets: A Unified Deformable Mixture Model for Human ParsingabstractHuman parsing, namely partitioning the human body into semantic regions, has drawn much attention recently for its wide applications in human-centric analysis. Previous works often consider solving the problem of human pose estimation as the prerequisite of human parsing. We argue that these approaches cannot obtain optimal pixel-level parsing due to the inconsistent targets between the different tasks. In this work, we directly address the problem of human parsing by using the novel Parselet representation as the building blocks of our parsing model. Parselets are a group of parsable segments which can generally be obtained by low-level over-segmentation algorithms and bear strong semantic meaning. We then build a deformable mixture parsing model (DMPM) for human parsing to simultaneously handle the deformation and multi-modalities of Parselets. The proposed model has two unique characteristics: (1) the possible numerous modalities of Parselet ensembles are exhibited as the "And-Or" structure of sub-trees; (2) to further solve the practical problem of Parselet occlusion or absence, we directly model the visibility property at some leaf nodes. The DMPM thus directly solves the problem of human parsing by searching for the best graph configuration from a pool of Parselet hypotheses without intermediate tasks. Fast rejection based on hierarchical filtering is employed to ensure the overall efficiency. Comprehensive evaluations on a new large-scale human parsing dataset, which is crawled from the Internet, with high resolution and thoroughly annotated semantic labels at pixel-level, and also a benchmark dataset demonstrate the encouraging performance of the proposed approach. Jian Dong 0011, Qiang Chen 0007, ZhongYang Huang, Jianchao Yang, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Robust Single Image Super-Resolution via Deep Networks With Sparse PriorabstractSingle image super-resolution (SR) is an ill-posed problem, which tries to recover a high-resolution image from its low-resolution observation. To regularize the solution of the problem, previous methods have focused on designing good priors for natural images, such as sparse representation, or directly learning the priors from a large data set with models, such as deep neural networks. In this paper, we argue that domain expertise from the conventional sparse coding model can be combined with the key ingredients of deep learning to achieve further improved results. We demonstrate that a sparse coding model particularly designed for SR can be incarnated as a neural network with the merit of end-to-end optimization over training data. The network has a cascaded structure, which boosts the SR performance for both fixed and incremental scaling factors. The proposed training and testing schemes can be extended for robust handling of images with additional degradation, such as noise and blurring. A subjective assessment is conducted and analyzed in order to thoroughly evaluate various SR techniques. Our proposed model is tested on a wide range of images, and it significantly outperforms the existing state-of-the-art methods for various scaling factors both quantitatively and perceptually. Ding Liu 0001, Bihan Wen, Jianchao Yang, Wei Han 0002, Thomas S. Huang |
IEEE Trans. Image Process. | 4 |
| 2015 | Robust Image Sentiment Analysis Using Progressively Trained and Domain Transferred Deep NetworksabstractSentiment analysis of online user generated content is important for many social media analytics tasks. Researchers have largely relied on textual sentiment analysis to develop systems to predict political elections, measure economic indicators, and so on. Recently, social media users are increasingly using images and videos to express their opinions and share their experiences. Sentiment analysis of such large scale visual content can help better extract user sentiments toward events or topics, such as those in image tweets, so that prediction of sentiment from visual content is complementary to textual sentiment analysis. Motivated by the needs in leveraging large scale yet noisy training data to solve the extremely challenging problem of image sentiment analysis, we employ Convolutional Neural Networks (CNN). We first design a suitable CNN architecture for image sentiment analysis. We obtain half a million training samples by using a baseline sentiment algorithm to label Flickr images. To make use of such noisy machine labeled data, we employ a progressive strategy to fine-tune the deep network. Furthermore, we improve the performance on Twitter images by inducing domain transfer with a small number of manually labeled Twitter images. We have conducted extensive experiments on manually labeled Twitter images. The results show that the proposed CNN can achieve better performance in image sentiment analysis than competing algorithms. Quanzeng You, Jiebo Luo 0001, Hailin Jin, Jianchao Yang |
AAAI | 4 |
| 2015 | Collaborative feature learning from social mediaabstractImage feature representation plays an essential role in image recognition and related tasks. The current state-of-the-art feature learning paradigm is supervised learning from labeled data. However, this paradigm requires large-scale category labels, which limits its applicability to domains where labels are hard to obtain. In this paper, we propose a new data-driven feature learning paradigm which does not rely on category labels. Instead, we learn from user behavior data collected on social media. Concretely, we use the image relationship discovered in the latent space from the user behavior data to guide the image feature learning. We collect a large-scale image and user behavior dataset from Behance.net. The dataset consists of 1.9 million images and over 300 million view records from 1.9 million users. We validate our feature learning paradigm on this dataset and find that the learned feature significantly outperforms the state-of-the-art image features in learning better image similarities. We also show that the learned feature performs competitively on various recognition benchmarks. Hailin Jin, Jianchao Yang, Zhe Lin 0001 |
CVPR | 3 |
| 2015 | Fine-grained recognition without part annotationsabstractScaling up fine-grained recognition to all domains of fine-grained objects is a challenge the computer vision community will need to face in order to realize its goal of recognizing all object categories. Current state-of-the-art techniques rely heavily upon the use of keypoint or part annotations, but scaling up to hundreds or thousands of domains renders this annotation cost-prohibitive for all but the most important categories. In this work we propose a method for fine-grained recognition that uses no part annotations. Our method is based on generating parts using co-segmentation and alignment, which we combine in a discriminative mixture. Experimental results show its efficacy, demonstrating state-of-the-art results even when compared to methods that use part annotations during training. Jonathan Krause, Hailin Jin, Jianchao Yang, Li Fei-Fei 0001 |
CVPR | 3 |
| 2015 | Matching-CNN meets KNN: Quasi-parametric human parsingabstractBoth parametric and non-parametric approaches have demonstrated encouraging performances in the human parsing task, namely segmenting a human image into several semantic regions (e.g., hat, bag, left arm, face). In this work, we aim to develop a new solution with the advantages of both methodologies, namely supervision from annotated data and the flexibility to use newly annotated (possibly uncommon) images, and present a quasi-parametric human parsing model. Under the classic K Nearest Neighbor (KNN)-based nonparametric framework, the parametric Matching Convolutional Neural Network (M-CNN) is proposed to predict the matching confidence and displacements of the best matched region in the testing image for a particular semantic region in one KNN image. Given a testing image, we first retrieve its KNN images from the annotated/manually-parsed human image corpus. Then each semantic region in each KNN image is matched with confidence to the testing image using M-CNN, and the matched regions from all KNN images are further fused, followed by a superpixel smoothing procedure to obtain the ultimate human parsing result. The M-CNN differs from the classic CNN [12] in that the tailored cross image matching filters are introduced to characterize the matching between the testing image and the semantic region of a KNN image. The cross image matching filters are defined at different convolutional layers, each aiming to capture a particular range of displacements. Comprehensive evaluations over a large dataset with 7,700 annotated human images well demonstrate the significant performance gain from the quasi-parametric model over the state-of-the-arts [29, 30], for the human parsing task. Si Liu 0001, Xiaodan Liang, Luoqi Liu, Xiaohui Shen, Jianchao Yang, Changsheng Xu, Liang Lin 0004, Xiaochun Cao, Shuicheng Yan |
CVPR | 5 |
| 2015 | Intra-frame deblurring by leveraging inter-frame camera motionabstractCamera motion introduces motion blur, degrading the quality of video. A video deblurring method is proposed based on two observations: (i) camera motion within capture of each individual frame leads to motion blur; (ii) camera motion between frames yields inter-frame mis-alignment that can be exploited for blur removal. The proposed method effectively leverages the information distributed across multiple video frames due to camera motion, jointly estimating the motion between consecutive frames and blur within each frame. This joint analysis is crucial for achieving effective restoration by leveraging temporal information. Extensive experiments are carried out on synthetic data as well as real-world blurry videos. Comparisons with several state-of-the-art methods verify the effectiveness of the proposed method. Haichao Zhang 0001, Jianchao Yang |
CVPR | 2 |
| 2015 | Human Parsing with Contextualized Convolutional Neural NetworkabstractIn this work, we address the human parsing task with a novel Contextualized Convolutional Neural Network (Co-CNN) architecture, which well integrates the cross-layer context, global image-level context, within-super-pixel context and cross-super-pixel neighborhood context into a unified network. Given an input human image, Co-CNN produces the pixel-wise categorization in an end-to-end way. First, the cross-layer context is captured by our basic local-to-global-to-local structure, which hierarchically combines the global semantic structure and the local fine details within the cross-layers. Second, the global image-level label prediction is used as an auxiliary objective in the intermediate layer of the Co-CNN, and its outputs are further used for guiding the feature learning in subsequent convolutional layers to leverage the global image-level context. Finally, to further utilize the local super-pixel contexts, the within-super-pixel smoothing and cross-super-pixel neighbourhood voting are formulated as natural sub-components of the Co-CNN to achieve the local label consistency in both training and testing process. Comprehensive evaluations on two public datasets well demonstrate the significant superiority of our Co-CNN architecture over other state-of-the-arts for human parsing. In particular, the F-1 score on the large dataset [15] reaches 76.95% by Co-CNN, significantly higher than 62.81% and 64.38% by the state-of-the-art algorithms, M-CNN [21] and ATR [15], respectively. Xiaodan Liang, Chunyan Xu, Xiaohui Shen, Jianchao Yang, Si Liu 0001, Jinhui Tang 0001, Liang Lin 0004, Shuicheng Yan |
ICCV | 4 |
| 2015 | Deep Networks for Image Super-Resolution with Sparse PriorabstractDeep learning techniques have been successfully applied in many areas of computer vision, including low-level image restoration problems. For image super-resolution, several models based on deep neural networks have been recently proposed and attained superior performance that overshadows all previous handcrafted models. The question then arises whether large-capacity and data-driven models have become the dominant solution to the ill-posed super-resolution problem. In this paper, we argue that domain expertise represented by the conventional sparse coding model is still valuable, and it can be combined with the key ingredients of deep learning to achieve further improved results. We show that a sparse coding model particularly designed for super-resolution can be incarnated as a neural network, and trained in a cascaded structure from end to end. The interpretation of the network based on sparse coding leads to much more efficient and effective training, as well as a reduced model size. Our model is evaluated on a wide range of images, and shows clear advantage over existing state-of-the-art methods in terms of both restoration accuracy and human subjective quality. Ding Liu 0001, Jianchao Yang, Wei Han 0002, Thomas S. Huang |
ICCV | 3 |
| 2015 | DeepFont: A System for Font Recognition and SimilarityabstractWe develop the DeepFont system, a large-scale learning-based solution for automatic font identification, organization and selection. In this proposed technical demonstration, we will give our audience a tour to the DeepFont system, with the focus on its impacts on real consumer products, including but not limited to: 1) a cloud-based iOS App for font recognition; 2) a web-based tool for font similarity evaluation and discovery. Zhangyang Wang, Jianchao Yang, Hailin Jin, Jonathan Brandt, Eli Shechtman, Aseem Agarwala, Yuyan Song, Joseph Hsieh, Sarah Kong, Thomas S. Huang |
ACM Multimedia | 2 |
| 2015 | DeepFont: Identify Your Font from An ImageabstractAs font is one of the core design concepts, automatic font identification and similar font suggestion from an image or photo has been on the wish list of many designers. We study the Visual Font Recognition (VFR) problem [4] LFE, and advance the state-of-the-art remarkably by developing the DeepFont system. First of all, we build up the first available large-scale VFR dataset, named AdobeVFR, consisting of both labeled synthetic data and partially labeled real-world data. Next, to combat the domain mismatch between available training and testing data, we introduce a Convolutional Neural Network (CNN) decomposition approach, using a domain adaptation technique based on a Stacked Convolutional Auto-Encoder (SCAE) that exploits a large corpus of unlabeled real-world text images combined with synthetic data preprocessed in a specific way. Moreover, we study a novel learning-based model compression approach, in order to reduce the DeepFont model size without sacrificing its performance. The DeepFont system achieves an accuracy of higher than 80% (top-5) on our collected dataset, and also produces a good font similarity measure for font selection and suggestion. We also achieve around 6 times compression of the model without any visible loss of recognition accuracy. Zhangyang Wang, Jianchao Yang, Hailin Jin, Eli Shechtman, Aseem Agarwala, Jonathan Brandt, Thomas S. Huang |
ACM Multimedia | 2 |
| 2015 | Joint Visual-Textual Sentiment Analysis with Deep Neural NetworksabstractSentiment analysis of online user generated content is important for many social media analytics tasks. Researchers have largely relied on textual sentiment analysis to develop systems to predict political elections, measure economic indicators, and so on. Recently, social media users are increasingly using additional images and videos to express their opinions and share their experiences. Sentiment analysis of such large-scale textual and visual content can help better extract user sentiments toward events or topics. Motivated by the needs to leverage large-scale social multimedia content for sentiment analysis, we utilize both the state-of-the-art visual and textual sentiment analysis techniques for joint visual-textual sentiment analysis. We first fine-tune a convolutional neural network (CNN) for image sentiment analysis and train a paragraph vector model for textual sentiment analysis. We have conducted extensive experiments on both machine weakly labeled and manually labeled image tweets. The results show that joint visual-textual features can achieve the state-of-the-art performance than textual and visual sentiment analysis algorithms alone. Quanzeng You, Jiebo Luo 0001, Hailin Jin, Jianchao Yang |
ACM Multimedia | 4 |
| 2015 | Designing a composite dictionary adaptively from joint examplesabstractWe study the complementary behaviors of external and internal examples in image restoration, and are motivated to formulate a composite dictionary design framework. The composite dictionary consists of the global part learned from external examples, and the sample-specific part learned from internal examples. The dictionary atoms in both parts are further adaptively weighted to emphasize their model statistics. Experiments demonstrate that the joint utilization of external and internal examples leads to substantial improvements, with successful applications in image denoising and super resolution. Zhangyang Wang, Yingzhen Yang, Jianchao Yang, Thomas S. Huang |
VCIP | 3 |
| 2015 | Selective Pooling Vector for Fine-Grained RecognitionabstractWe propose a new framework for image recognition by selectively pooling local visual descriptors, and show its superior discriminative power on fine-grained image classification tasks. The representation is based on selecting the most confident local descriptors for nonlinear function learning using a linear approximation in an embedded higher dimensional space. The advantage of our Selective Pooling Vector over the previous state-of-the-art Super Vector and Fisher Vector representations, is that it ensures a more accurate learning function, which proves to be important for classifying details in fine-grained image recognition. Our experimental results corroborate this claim: with a simple linear SVM as the classifier, the selective pooling vector achieves significant performance gains on standard benchmark datasets for various fine-grained tasks such as the CMU Multi-PIE dataset for face recognition, the Caltech-UCSD Bird dataset and the Stanford Dogs dataset for fine-grained object categorization. On all datasets we outperform the state of the arts and boost the recognition rates to 96.4%, 48.9%, 52.0% respectively. Jianchao Yang, Hailin Jin, Eli Shechtman, Jonathan Brandt, Tony X. Han |
WACV | 2 |
| 2015 | Scalable Similarity Learning Using Large Margin Neighborhood EmbeddingabstractClassifying large-scale image data into object categories is an important problem that has received increasing research attention. Given the huge amount of data, non-parametric approaches such as nearest neighbor classifiers have shown promising results, especially when they are underpinned by a learned distance or similarity measurement. Although metric learning has been well studied in the past decades, most existing algorithms are impractical to handle large-scale data sets. In this paper, we present an image similarity learning method that can scale well in both the number of images and the dimensionality of image descriptors. To this end, similarity comparison is restricted to each sample's local neighbors and a discriminative similarity measure is induced from large margin neighborhood embedding. We also exploit the ensemble of projections so that high-dimensional features can be processed in a set of lower-dimensional subspaces in parallel. The efficiency and scalability of our proposed model are validated on several data sets with scales varying from tens of thousands to one million images. Jianchao Yang, Zhe Lin 0001, Jonathan Brandt, Shiyu Chang, Thomas S. Huang |
WACV | 2 |
| 2015 | Deep Human Parsing with Active Template RegressionabstractIn this work, the human parsing task, namely decomposing a human image into semantic fashion/body regions, is formulated as an active template regression (ATR) problem, where the normalized mask of each fashion/body item is expressed as the linear combination of the learned mask templates, and then morphed to a more precise mask with the active shape parameters, including position, scale and visibility of each semantic region. The mask template coefficients and the active shape parameters together can generate the human parsing results, and are thus called the structure outputs for human parsing. The deep Convolutional Neural Network (CNN) is utilized to build the end-to-end relation between the input human image and the structure outputs for human parsing. More specifically, the structure outputs are predicted by two separate networks. The first CNN network is with max-pooling, and designed to predict the template coefficients for each label mask, while the second CNN network is without max-pooling to preserve sensitivity to label mask position and accurately predict the active shape parameters. For a new image, the structure outputs of the two networks are fused to generate the probability of each label for each pixel, and super-pixel smoothing is finally used to refine the human parsing result. Comprehensive evaluations on a large dataset well demonstrate the significant superiority of the ATR framework over other state-of-the-arts for human parsing. In particular, the F1-score reaches 64.38 percent by our ATR framework, significantly higher than 44.76 percent based on the state-of-the-art algorithm [28]. Xiaodan Liang, Si Liu 0001, Xiaohui Shen, Jianchao Yang, Luoqi Liu, Jian Dong 0011, Liang Lin 0004, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Subcategory-Aware Object DetectionabstractIn this letter, we introduce a subcategory-aware object detection framework to detect generic object classes with high intra-class variance. Motivated by the observation that the object appearance demonstrates some clustering property, we split the training data into subcategories and train a detector for each subcategory. Since the proposed ensemble of detectors relies heavily on subcategory clustering, we propose an effective subcategories generation method that is tuned for the detection task. More specifically, we first initialize subcategories by constrained spectral clustering based on mid-level image features used in object recognition. Then we jointly learn the ensemble detectors and the latent subcategories in an alternative manner. Our performance on the PASCAL VOC 2007 detection challenges and INRIA Person dataset is comparable with state-of-the-art, even with much less computational cost. Xiaoyuan Yu, Jianchao Yang, Zhe Lin 0001, Jiangping Wang, Tianjiang Wang, Thomas S. Huang |
IEEE Signal Process. Lett. | 2 |
| 2015 | Key Point Detection by Max Pooling for TrackingabstractInspired by the recent image feature learning work, we propose a novel key point detection approach for object tracking. Our approach can select mid-level interest key points by max pooling over the local descriptor responses from a set of filters. Linear filters are first learned from targets in first frames. Then max pooling is performed over data driven spatial supporting field to detect discriminant key points, and thus the detected key points bear higher level semantic meanings, which we apply in tracking by structured key point matching. We show that our tracking system is robust to occlusions and cluttered background. Testing on several challenging tracking sequences, we demonstrate that our proposed tracking system can achieve competitive or better performances than the state-of-the-art trackers. Xiaoyuan Yu, Jianchao Yang, Tianjiang Wang, Thomas S. Huang |
IEEE Trans. Cybern. | 2 |
| 2015 | Image-Specific Prior Adaptation for DenoisingabstractImage priors are essential to many image restoration applications, including denoising, deblurring, and inpainting. Existing methods use either priors from the given image (internal) or priors from a separate collection of images (external). We find through statistical analysis that unifying the internal and external patch priors may yield a better patch prior. We propose a novel prior learning algorithm that combines the strength of both internal and external priors. In particular, we first learn a generic Gaussian mixture model from a collection of training images and then adapt the model to the given image by simultaneously adding additional components and refining the component parameters. We apply this image-specific prior to image denoising. The experimental results show that our approach yields better or competitive denoising results in terms of both the peak signal-to-noise ratio and structural similarity. Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2015 | Learning Super-Resolution Jointly From External and Internal ExamplesabstractSingle image super-resolution (SR) aims to estimate a high-resolution (HR) image from a low-resolution (LR) input. Image priors are commonly learned to regularize the, otherwise, seriously ill-posed SR problem, either using external LR-HR pairs or internal similar patterns. We propose joint SR to adaptively combine the advantages of both external and internal SR methods. We define two loss functions using sparse coding-based external examples, and epitomic matching based on internal examples, as well as a corresponding adaptive weight to automatically balance their contributions according to their reconstruction errors. Extensive SR results demonstrate the effectiveness of the proposed method over the existing state-of-the-art methods, and is also verified by our subjective evaluation study. Zhangyang Wang, Yingzhen Yang, Shiyu Chang, Jianchao Yang, Thomas S. Huang |
IEEE Trans. Image Process. | 5 |
| 2015 | Rating Image Aesthetics Using Deep LearningabstractThis paper investigates unified feature learning and classifier training approaches for image aesthetics assessment . Existing methods built upon handcrafted or generic image features and developed machine learning and statistical modeling techniques utilizing training examples. We adopt a novel deep neural network approach to allow unified feature learning and classifier training to estimate image aesthetics. In particular, we develop a double-column deep convolutional neural network to support heterogeneous inputs, i.e., global and local views, in order to capture both global and local characteristics of images . In addition, we employ the style and semantic attributes of images to further boost the aesthetics categorization performance . Experimental results show that our approach produces significantly better results than the earlier reported results on the AVA dataset for both the generic image aesthetics and content -based image aesthetics. Moreover, we introduce a 1.5-million image dataset (IAD) for image aesthetics assessment and we further boost the performance on the AVA test set by training the proposed deep neural networks on the IAD dataset. Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2014 | Data Clustering by Laplacian Regularized L1-GraphabstractL1-Graph has been proven to be effective in data clustering, which partitions the data space by using the sparse representation of the data as the similarity measure. However, the sparse representation is performed for each datum separately without taking into account the geometric structure of the data. Motivated by L1-Graph and manifold leaning, we propose Laplacian Regularized L1-Graph (LRℓ1-Graph) for data clustering. The sparse representations of LRℓ1-Graph are regularized by the geometric information of the data so that they vary smoothly along the geodesics of the data manifold by the graph Laplacian according to the manifold assumption. Moreover, we propose an iterative regularization scheme, where the sparse representation obtained from the previous iteration is used to build the graph Laplacian for the current iteration of regularization. The experimental results on real data sets demonstrate the superiority of our algorithm compared to L1-Graph and other competing clustering methods. Yingzhen Yang, Zhangyang Wang, Jianchao Yang, Jiangping Wang, Shiyu Chang, Thomas S. Huang |
AAAI | 3 |
| 2014 | Regularized l1-Graph for Data Clustering
Yingzhen Yang, Zhangyang Wang, Jianchao Yang, Jiawei Han 0001, Thomas S. Huang |
BMVC | 3 |
| 2014 | Large-Scale Visual Font RecognitionabstractThis paper addresses the large-scale visual font recognition (VFR) problem, which aims at automatic identification of the typeface, weight, and slope of the text in an image or photo without any knowledge of content. Although visual font recognition has many practical applications, it has largely been neglected by the vision community. To address the VFR problem, we construct a large-scale dataset containing 2,420 font classes, which easily exceeds the scale of most image categorization datasets in computer vision. As font recognition is inherently dynamic and open-ended, i.e., new classes and data for existing categories are constantly added to the database over time, we propose a scalable solution based on the nearest class mean classifier (NCM). The core algorithm is built on local feature embedding, local feature metric learning and max-margin template selection, which is naturally amenable to NCM and thus to such open-ended classification problems. The new algorithm can generalize to new classes and new data at little added cost. Extensive experiments demonstrate that our approach is very effective on our synthetic test images, and achieves promising results on real world test images. Jianchao Yang, Hailin Jin, Jonathan Brandt, Eli Shechtman, Aseem Agarwala, Tony X. Han |
CVPR | 2 |
| 2014 | Towards Unified Human Parsing and Pose EstimationabstractWe study the problem of human body configuration analysis, more specifically, human parsing and human pose estimation. These two tasks, ie identifying the semantic regions and body joints respectively over the human body image, are intrinsically highly correlated. However, previous works generally solve these two problems separately or iteratively. In this work, we propose a unified framework for simultaneous human parsing and pose estimation based on semantic parts. By utilizing Parselets and Mixture of Joint-Group Templates as the representations for these semantic parts, we seamlessly formulate the human parsing and pose estimation problem jointly within a unified framework via a tailored and-or graph. A novel Grid Layout Feature is then designed to effectively capture the spatial co-occurrence/occlusion information between/within the Parselets and MJGTs. Thus the mutually complementary nature of these two tasks can be harnessed to boost the performance of each other. The resultant unified model can be solved using the structure learning framework in a principled way. Comprehensive evaluations on two benchmark datasets for both tasks demonstrate the effectiveness of the proposed framework when compared with the state-of-the-art methods. Jian Dong 0011, Qiang Chen 0007, Xiaohui Shen, Jianchao Yang, Shuicheng Yan |
CVPR | 4 |
| 2014 | Investigating Haze-Relevant Features in a Learning Framework for Image DehazingabstractHaze is one of the major factors that degrade outdoor images. Removing haze from a single image is known to be severely ill-posed, and assumptions made in previous methods do not hold in many situations. In this paper, we systematically investigate different haze-relevant features in a learning framework to identify the best feature combination for image dehazing. We show that the dark-channel feature is the most informative one for this task, which confirms the observation of He et al. [8] from a learning perspective, while other haze-relevant features also contribute significantly in a complementary way. We also find that surprisingly, the synthetic hazy image patches we use for feature investigation serve well as training data for realworld images, which allows us to train specific models for specific applications. Experiment results demonstrate that the proposed algorithm outperforms state-of-the-art methods on both synthetic and real-world datasets. Ketan Tang, Jianchao Yang, Jue Wang 0001 |
CVPR | 2 |
| 2014 | Epitomic image colorizationabstractImage colorization adds color to grayscale images. It not only increases the visual appeal of grayscale images, but also enriches the information conveyed by scientific images that lack color information. We develop a new image colorization method, epitomic image colorization, which automatically transfers color from the reference color image to the target grayscale image by a robust feature matching scheme using a new feature representation, namely the heterogeneous feature epitome. As a generative model, heterogeneous feature epitome is a condensed representation of image appearance which is employed for measuring the dissimilarity between reference patches and target patches in a way robust to noise in the reference image. We build a Markov Random Field (MRF) model with the learned heterogeneous feature epitome from the reference image, and inference in the MRF model achieves robust feature matching for transferring color. Our method renders better colorization results than the current state-of-the-art automatic colorization methods in our experiments. Yingzhen Yang, Xinqi Chu, Tian-Tsong Ng, Alex Yong Sang Chia, Jianchao Yang, Hailin Jin, Thomas S. Huang |
ICASSP | 5 |
| 2014 | Non-local compressive sampling recoveryabstractCompressive sampling (CS) aims at acquiring a signal at a sampling rate below the Nyquist rate by exploiting prior knowledge that a signal is sparse or correlated in some domain. Despite the remarkable progress in the theory of CS, the sampling rate on a single image required by CS is still very high in practice. In this paper, a non-local compressive sampling (NLCS) recovery method is proposed to further reduce the sampling rate by exploiting non-local patch correlation and local piecewise smoothness present in natural images. Two non-local sparsity measures, i.e., non-local wavelet sparsity and non-local joint sparsity, are proposed to exploit the patch correlation in NLCS. An efficient iterative algorithm is developed to solve the NLCS recovery problem, which is shown to have stable convergence behavior in experiments. The experimental results show that our NLCS significantly improves the state-of-the-art of image compressive sampling. Xianbiao Shu, Jianchao Yang, Narendra Ahuja |
ICCP | 2 |
| 2014 | RAPID: Rating Pictorial Aesthetics using Deep LearningabstractEffective visual features are essential for computational aesthetic quality rating systems. Existing methods used machine learning and statistical modeling techniques on handcrafted features or generic image descriptors. A recently-published large-scale dataset, the AVA dataset, has further empowered machine learning based approaches. We present the RAPID (RAting PIctorial aesthetics using Deep learning) system, which adopts a novel deep neural network approach to enable automatic feature learning. The central idea is to incorporate heterogeneous inputs generated from the image, which include a global view and a local view, and to unify the feature learning and classifier training using a double-column deep convolutional neural network. In addition, we utilize the style attributes of images to help improve the aesthetic quality categorization accuracy. Experimental results show that our approach significantly outperforms the state of the art on the AVA dataset. Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001 |
ACM Multimedia | 4 |
| 2014 | Scale Adaptive Blind Deblurring
Haichao Zhang 0001, Jianchao Yang |
NIPS | 2 |
| 2014 | A joint perspective towards image super-resolution: Unifying external- and self-examplesabstractExisting example-based super resolution (SR) methods are built upon either external-examples or self-examples. Although effective in certain cases, both methods suffer from their inherent limitation. This paper goes beyond these two classes of most common example-based SR approaches, and proposes a novel joint SR perspective. The joint SR exploits and maximizes the complementary advantages of external- and self-example based methods. We elaborate on exploitable priors for image components of different nature, and formulate their corresponding loss functions mathematically. Equipped with that, we construct a unified SR formulation, and propose an iterative joint super resolution (IJSR) algorithm to solve the optimization. Such a joint perspective approach leads to an impressive improvement of SR results both quantitatively and qualitatively. Zhangyang Wang, Shiyu Chang, Jianchao Yang, Thomas S. Huang |
WACV | 4 |
| 2014 | Supervised super-vector encoding for facial expression recognition
Usman Tariq, Jianchao Yang, Thomas S. Huang |
Pattern Recognit. Lett. | 2 |
| 2013 | Probabilistic Elastic Matching for Pose Variant Face VerificationabstractPose variation remains to be a major challenge for real-world face recognition. We approach this problem through a probabilistic elastic matching method. We take a part based representation by extracting local features (e.g., LBP or SIFT) from densely sampled multi-scale image patches. By augmenting each feature with its location, a Gaussian mixture model (GMM) is trained to capture the spatial-appearance distribution of all face images in the training corpus. Each mixture component of the GMM is confined to be a spherical Gaussian to balance the influence of the appearance and the location terms. Each Gaussian component builds correspondence of a pair of features to be matched between two faces/face tracks. For face verification, we train an SVM on the vector concatenating the difference vectors of all the feature pairs to decide if a pair of faces/face tracks is matched or not. We further propose a joint Bayesian adaptation algorithm to adapt the universally trained GMM to better model the pose variations between the target pair of faces/face tracks, which consistently improves face verification accuracy. Our experiments show that our method outperforms the state-of-the-art in the most restricted protocol on Labeled Face in the Wild (LFW) and the YouTube video face database by a significant margin. Gang Hua 0001, Zhe Lin 0001, Jonathan Brandt, Jianchao Yang |
CVPR | 5 |
| 2013 | Exemplar-Based Face ParsingabstractIn this work, we propose an exemplar-based face image segmentation algorithm. We take inspiration from previous works on image parsing for general scenes. Our approach assumes a database of exemplar face images, each of which is associated with a hand-labeled segmentation map. Given a test image, our algorithm first selects a subset of exemplar images from the database, Our algorithm then computes a nonrigid warp for each exemplar image to align it with the test image. Finally, we propagate labels from the exemplar images to the test image in a pixel-wise manner, using trained weights to modulate and combine label maps from different exemplars. We evaluate our method on two challenging datasets and compare with two face parsing algorithms and a general scene parsing algorithm. We also compare our segmentation results with contour-based face alignment results, that is, we first run the alignment algorithms to extract contour points and then derive segments from the contours. Our algorithm compares favorably with all previous works on all datasets evaluated. Brandon M. Smith 0001, Li Zhang 0003, Jonathan Brandt, Zhe Lin 0001, Jianchao Yang |
CVPR | 5 |
| 2013 | Fast Image Super-Resolution Based on In-Place Example RegressionabstractWe propose a fast regression model for practical single image super-resolution based on in-place examples, by leveraging two fundamental super-resolution approaches- learning from an external database and learning from self-examples. Our in-place self-similarity refines the recently proposed local self-similarity by proving that a patch in the upper scale image have good matches around its origin location in the lower scale image. Based on the in-place examples, a first-order approximation of the nonlinear mapping function from low-to high-resolution image patches is learned. Extensive experiments on benchmark and real-world images demonstrate that our algorithm can produce natural-looking results with sharp edges and preserved fine details, while the current state-of-the-art algorithms are prone to visual artifacts. Furthermore, our model can easily extend to deal with noise by combining the regression results on multiple in-place examples for robust estimation. The algorithm runs fast and is particularly useful for practical applications, where the input images typically contain diverse textures and they are potentially contaminated by noise or compression artifacts. Jianchao Yang, Zhe Lin 0001, Scott Cohen |
CVPR | 1 |
| 2013 | Text Localization in Natural Images Using Stroke Feature Transform and Text Covariance DescriptorsabstractIn this paper, we present a new approach for text localization in natural images, by discriminating text and non-text regions at three levels: pixel, component and text line levels. Firstly, a powerful low-level filter called the Stroke Feature Transform (SFT) is proposed, which extends the widely-used Stroke Width Transform (SWT) by incorporating color cues of text pixels, leading to significantly enhanced performance on inter-component separation and intra-component connection. Secondly, based on the output of SFT, we apply two classifiers, a text component classifier and a text-line classifier, sequentially to extract text regions, eliminating the heuristic procedures that are commonly used in previous approaches. The two classifiers are built upon two novel Text Covariance Descriptors (TCDs) that encode both the heuristic properties and the statistical characteristics of text stokes. Finally, text regions are located by simply thresholding the text-line confident map. Our method was evaluated on two benchmark datasets: ICDAR 2005 and ICDAR 2011, and the corresponding F-measure values are 0.72 and 0.73, respectively, surpassing previous methods in accuracy by a large margin. Zhe Lin 0001, Jianchao Yang, Jue Wang 0001 |
ICCV | 3 |
| 2013 | Probabilistic Elastic Part Model for Unsupervised Face Detector AdaptationabstractWe propose an unsupervised detector adaptation algorithm to adapt any offline trained face detector to a specific collection of images, and hence achieve better accuracy. The core of our detector adaptation algorithm is a probabilistic elastic part (PEP) model, which is offline trained with a set of face examples. It produces a statistically aligned part based face representation, namely the PEP representation. To adapt a general face detector to a collection of images, we compute the PEP representations of the candidate detections from the general face detector, and then train a discriminative classifier with the top positives and negatives. Then we re-rank all the candidate detections with this classifier. This way, a face detector tailored to the statistics of the specific image collection is adapted from the original detector. We present extensive results on three datasets with two state-of-the-art face detectors. The significant improvement of detection accuracy over these state of-the-art face detectors strongly demonstrates the efficacy of the proposed face detector adaptation algorithm. Gang Hua 0001, Zhe Lin 0001, Jonathan Brandt, Jianchao Yang |
ICCV | 5 |
| 2013 | A Max-Margin Perspective on Sparse Representation-Based ClassificationabstractSparse Representation-based Classification (SRC) is a powerful tool in distinguishing signal categories which lie on different subspaces. Despite its wide application to visual recognition tasks, current understanding of SRC is solely based on a reconstructive perspective, which neither offers any guarantee on its classification performance nor provides any insight on how to design a discriminative dictionary for SRC. In this paper, we present a novel perspective towards SRC and interpret it as a margin classifier. The decision boundary and margin of SRC are analyzed in local regions where the support of sparse code is stable. Based on the derived margin, we propose a hinge loss function as the gauge for the classification performance of SRC. A stochastic gradient descent algorithm is implemented to maximize the margin of SRC and obtain more discriminative dictionaries. Experiments validate the effectiveness of the proposed approach in predicting classification performance and improving dictionary quality over reconstructive ones. Classification results competitive with other state-of-the-art sparse coding methods are reported on several data sets. Jianchao Yang, Nasser M. Nasrabadi, Thomas S. Huang |
ICCV | 2 |
| 2013 | Opportunistic sensing for object recognition - A unified formulation for dynamic sensor selection and feature extractionabstractA novel problem of object recognition with dynamically allocated sensing resources is considered in this paper. We call this problem opportunistic sensing since prior knowledge about the correlation between class label and signal distribution is exploited as early as in data acquisition. Two forms of sensing parameters — discrete sensor index and continuous linear measurement vector — are optimized within the same maximum negative entropy framework. The computationally intractable expected entropy is approximated using unscented transform for Gaussian models, and we solve the problem using a gradient-based method. Our formulation is theoretically shown to be closely related to the maximum mutual information criterion for sensor selection and linear feature extraction techniques such as PCA, LDA, and CCA. The proposed approach is validated on multi-view vehicle classification and face recognition datasets, and remarkable improvement over baseline methods is demonstrated in the experiments. Jianchao Yang, Nasser M. Nasrabadi, Jiangping Wang, Thomas S. Huang |
ICME | 2 |
| 2013 | Image and Video Restorations via Nonlocal Kernel RegressionabstractA nonlocal kernel regression (NL-KR) model is presented in this paper for various image and video restoration tasks. The proposed method exploits both the nonlocal self-similarity and local structural regularity properties in natural images. The nonlocal self-similarity is based on the observation that image patches tend to repeat themselves in natural images and videos, and the local structural regularity observes that image patches have regular structures where accurate estimation of pixel values via regression is possible. By unifying both properties explicitly, the proposed NL-KR framework is more robust in image estimation, and the algorithm is applicable to various image and video restoration tasks. In this paper, we apply the proposed model to image and video denoising, deblurring, and superresolution reconstruction. Extensive experimental results on both single images and realistic video sequences demonstrate that the proposed framework performs favorably with previous works both qualitatively and quantitatively. Haichao Zhang 0001, Jianchao Yang, Yanning Zhang 0001, Thomas S. Huang |
IEEE Trans. Cybern. | 2 |
| 2012 | Bilevel sparse coding for coupled feature spacesabstractIn this paper, we propose a bilevel sparse coding model for coupled feature spaces, where we aim to learn dictionaries for sparse modeling in both spaces while enforcing some desired relationships between the two signal spaces. We first present our new general sparse coding model that relates signals from the two spaces by their sparse representations and the corresponding dictionaries. The learning algorithm is formulated as a generic bilevel optimization problem, which is solved by a projected first-order stochastic gradient descent algorithm. This general sparse coding model can be applied to many specific applications involving coupled feature spaces in computer vision and signal processing. In this work, we tailor our general model to learning dictionaries for compressive sensing recovery and single image super-resolution to demonstrate its effectiveness. In both cases, the new sparse coding model remarkably outperforms previous approaches in terms of recovery accuracy. Jianchao Yang, Zhe Lin 0001, Xianbiao Shu, Thomas S. Huang |
CVPR | 1 |
| 2012 | Pooling Robust Shift-Invariant Sparse Representations of Acoustic SignalsabstractIn recent years, designing the coding and pooling structures in layered networks has been shown to be a useful method for learning high-level feature representations for visual data. Yet, such learning structures have not been extensively studied for audio signals. In this paper, we investigate the different pooling strategies based on the sparse coding scheme and propose a temporal pyramid pooling method to extract discriminative and shift-invariant feature representations. We demonstrate the superiority of our new feature representation over traditional features on the acoustic event classification task. Index Terms: sparse coding, pooling, acoustic event classification 1. Po-Sen Huang, Jianchao Yang, Mark Hasegawa-Johnson, Feng Liang 0002, Thomas S. Huang |
INTERSPEECH | 2 |
| 2012 | Coupled Dictionary Training for Image Super-ResolutionabstractIn this paper, we propose a novel coupled dictionary training method for single image super-resolution based on patchwise sparse recovery, where the learned couple dictionaries relate the low- and high-resolution image patch spaces via sparse representation. The learning process enforces that the sparse representation of a low-resolution image patch in terms of the low-resolution dictionary can well reconstruct its underlying high-resolution image patch with the dictionary in the highresolution image patch space. We model the learning problem as a bilevel optimization problem, where the optimization includes an 1-norm minimization problem in its constraints. Implicit differentiation is employed to calculate the desired gradient for stochastic gradient descent. We demonstrate that our coupled dictionary learning method can outperform the existing joint dictionary training method both quantitatively and qualitatively. Furthermore, for real applications, we speed up the algorithm approximately 10 times by learning a neural network model for fast sparse inference and selectively processing only those visually salient regions. Extensive experimental comparisons with stateof- the-art super-resolution algorithms validate the effectiveness of our proposed approach. Jianchao Yang, Zhe Lin 0001, Scott Cohen, Thomas S. Huang |
IEEE Trans. Image Process. | 1 |
| 2012 | Photo Stream Alignment and Summarization for Collaborative Photo Collection and SharingabstractWith the popularity of digital cameras and camera phones, it is common for different people, who may or may not know each other, to attend the same event and take pictures and videos from different spatial or personal perspectives. Within the realm of social media, it is desirable to enable these people to select and share their pictures and videos in order to enrich memories and facilitate social networking. However, it is cumbersome to manually manage these photos from different cameras, of which the clocks settings are often not calibrated. In this paper, we propose automatic algorithms to address the above problems. First, we accurately align different photo streams or sequences from different photographers for the same event in chronological order on a common timeline, while respecting the time constraints within each photo stream. Given the preferred similarity measures (e.g., visual, and spatial similarities), our algorithm performs photo stream alignment via matching on a bipartite kernel sparse representation graph that forces the data connections to be sparse in an explicit fashion. Furthermore, we can produce a summary master stream from the aligned super stream of photos for efficient sharing by removing those redundant photos in the super stream while accounting for the temporal integrity. Based on a similar kernel sparse representation graph, our master stream summarization algorithm performs greedy backward selection to drop redundant photos without affecting the integrity of remaining photos for the entire event. We evaluate our algorithms on real-world personal online albums for 36 events and demonstrate its efficacy in automatically facilitating collaborative photo collection and sharing. Jianchao Yang, Jiebo Luo 0001, Jie Yu 0001, Thomas S. Huang |
IEEE Trans. Multim. | 1 |
| 2011 | Close the loop: Joint blind image restoration and recognition with sparse representation priorabstractMost previous visual recognition systems simply assume ideal inputs without real-world degradations, such as low resolution, motion blur and out-of-focus blur. In presence of such unknown degradations, the conventional approach first resorts to blind image restoration and then feeds the restored image into a classifier. Treating restoration and recognition separately, such a straightforward approach, however, suffers greatly from the defective output of the ill-posed blind image restoration. In this paper, we present a joint blind image restoration and recognition method based on the sparse representation prior to handle the challenging problem of face recognition from low-quality images, where the degradation model is realistic and totally unknown. The sparse representation prior states that the degraded input image, if correctly restored, will have a good sparse representation in terms of the training set, which indicates the identity of the test image. The proposed algorithm achieves simultaneous restoration and recognition by iteratively solving the blind image restoration in pursuit of the sparest representation for recognition. Based on such a sparse representation prior, we demonstrate that the image restoration task and the recognition task can benefit greatly from each other. Extensive experiments on face datasets under various degradations are carried out and the results of our joint model shows significant improvements over conventional methods of treating the two tasks independently. Haichao Zhang 0001, Jianchao Yang, Yanning Zhang 0001, Nasser M. Nasrabadi, Thomas S. Huang |
ICCV | 2 |
| 2011 | Multi-scale Non-Local Kernel Regression for super resolutionabstractIn this paper, we propose an extension of the Non-Local Kernel Regression (NL-KR) method and apply it to super-resolution (SR) tasks. The proposed method extends NL-KR via generalizing the self-similarity from single-scale to multi-scale, and propose an effective SR algorithm using the proposed multi-scale NL-KR model. Experimental results on both synthetic and real images demonstrate the effectiveness of the proposed method. Haichao Zhang 0001, Jianchao Yang, Yanning Zhang 0001, Thomas S. Huang |
ICIP | 2 |
| 2011 | Learning the sparse representation for classificationabstractIn this work, we propose a novel supervised matrix factorization method used directly as a multi-class classifier. The coefficient matrix of the factorization is enforced to be sparse by ℓ1-norm regularization. The basis matrix is composed of atom dictionaries from different classes, which are trained in a jointly supervised manner by penalizing inhomogeneous representations given the labeled data samples. The learned basis matrix models the data of interest as a union of discriminative linear subspaces by sparse projection. The proposed model is based on the observation that many high-dimensional natural signals lie in a much lower dimensional subspaces or union of subspaces. Experiments conducted on several datasets show the effectiveness of such a representation model for classification, which also suggests that a tight reconstructive representation model could be very useful for discriminant analysis. Jianchao Yang, Jiangping Wang, Thomas S. Huang |
ICME | 1 |
| 2011 | Sparse representation based blind image deblurringabstractWe propose a sparse representation based blind image deblurring method. The proposed method exploits the sparsity property of natural images, by assuming that the patches from the natural images can be sparsely represented by an over-complete dictionary. By incorporating this prior into the deblurring process, we can effectively regularize the ill-posed inverse problem and alleviate the undesirable ring effect which is usually suffered by conventional deblurring methods. Experimental results compared with state-of-the-art blind deblurring method demonstrate the effectiveness of the proposed method. Haichao Zhang 0001, Jianchao Yang, Yanning Zhang 0001, Thomas S. Huang |
ICME | 2 |
| 2010 | Locality-constrained Linear Coding for image classificationabstractThe traditional SPM approach based on bag-of-features (BoF) requires nonlinear classifiers to achieve good image classification performance. This paper presents a simple but effective coding scheme called Locality-constrained Linear Coding (LLC) in place of the VQ coding in traditional SPM. LLC utilizes the locality constraints to project each descriptor into its local-coordinate system, and the projected coordinates are integrated by max pooling to generate the final representation. With linear classifier, the proposed approach performs remarkably better than the traditional nonlinear SPM, achieving state-of-the-art performance on several benchmarks. Compared with the sparse coding strategy [22], the objective function used by LLC has an analytical solution. In addition, the paper proposes a fast approximated LLC method by first performing a K-nearest-neighbor search and then solving a constrained least square fitting problem, bearing computational complexity of O(M + K2). Hence even with very large codebooks, our system can still process multiple frames per second. This efficiency significantly adds to the practical values of LLC for real applications. Jinjun Wang, Jianchao Yang, Kai Yu 0001, Fengjun Lv, Thomas S. Huang, Yihong Gong |
CVPR | 2 |
| 2010 | Supervised translation-invariant sparse codingabstractIn this paper, we propose a novel supervised hierarchical sparse coding model based on local image descriptors for classification tasks. The supervised dictionary training is performed via back-projection, by minimizing the training error of classifying the image level features, which are extracted by max pooling over the sparse codes within a spatial pyramid. Such a max pooling procedure across multiple spatial scales offer the model translation invariant properties, similar to the Convolutional Neural Network (CNN). Experiments show that our supervised dictionary improves the performance of the proposed model significantly over the unsupervised dictionary, leading to state-of-the-art performance on diverse image databases. Further more, our supervised model targets learning linear features, implying its great potential in handling large scale datasets in real applications. Jianchao Yang, Kai Yu 0001, Thomas S. Huang |
CVPR | 1 |
| 2010 | Efficient Highly Over-Complete Sparse Coding Using a Mixture Model
Jianchao Yang, Kai Yu 0001, Thomas S. Huang |
ECCV (5) | 1 |
| 2010 | Non-Local Kernel Regression for Image and Video Restoration
Haichao Zhang 0001, Jianchao Yang, Yanning Zhang 0001, Thomas S. Huang |
ECCV (3) | 2 |
| 2010 | Learning With ℓ1-Graph for Image AnalysisabstractThe graph construction procedure essentially determines the potentials of those graph-oriented learning algorithms for image analysis. In this paper, we propose a process to build the so-called directed l1-graph, in which the vertices involve all the samples and the ingoing edge weights to each vertex describe its l1-norm driven reconstruction from the remaining samples and the noise. Then, a series of new algorithms for various machine learning tasks, e.g., data clustering, subspace learning, and semi-supervised learning, are derived upon the l1-graphs. Compared with the conventional k-nearest-neighbor graph and epsilon-ball graph, the l1-graph possesses the advantages: (1) greater robustness to data noise, (2) automatic sparsity, and (3) adaptive neighborhood for individual datum. Extensive experiments on three real-world datasets show the consistent superiority of l1-graph over those classic graphs in data clustering, subspace learning, and semi-supervised learning tasks. Bin Cheng 0001, Jianchao Yang, Shuicheng Yan, Yun Fu 0001, Thomas S. Huang |
IEEE Trans. Image Process. | 2 |
| 2010 | Image Super-Resolution Via Sparse RepresentationabstractThis paper presents a new approach to single-image super-resolution, based on sparse signal representation. Research on image statistics suggests that image patches can be well-represented as a sparse linear combination of elements from an appropriately chosen over-complete dictionary. Inspired by this observation, we seek a sparse representation for each patch of the low-resolution input, and then use the coefficients of this representation to generate the high-resolution output. Theoretical results from compressed sensing suggest that under mild conditions, the sparse representation can be correctly recovered from the downsampled signals. By jointly training two dictionaries for the low- and high-resolution image patches, we can enforce the similarity of sparse representations between the low resolution and high resolution image patch pair with respect to their own dictionaries. Therefore, the sparse representation of a low resolution image patch can be applied with the high resolution image patch dictionary to generate a high resolution image patch. The learned dictionary pair is a more compact representation of the patch pairs, compared to previous approaches, which simply sample a large amount of image patch pairs, reducing the computational cost substantially. The effectiveness of such a sparsity prior is demonstrated for both general image super-resolution and the special case of face hallucination. In both cases, our algorithm generates high-resolution images that are competitive or even superior in quality to images produced by other similar SR methods. In addition, the local sparse modeling of our approach is naturally robust to noise, and therefore the proposed algorithm can handle super-resolution with noisy inputs in a more unified framework. Jianchao Yang, John Wright 0001, Thomas S. Huang, Yi Ma 0001 |
IEEE Trans. Image Process. | 1 |
| 2009 | Linear spatial pyramid matching using sparse coding for image classificationabstractRecently SVMs using spatial pyramid matching (SPM) kernel have been highly successful in image classification. Despite its popularity, these nonlinear SVMs have a complexity O(n2∼ n3) in training and O(n) in testing, where n is the training size, implying that it is nontrivial to scaleup the algorithms to handlemore than thousands of training images. In this paper we develop an extension of the SPM method, by generalizing vector quantization to sparse coding followed by multi-scale spatial max pooling, and propose a linear SPM kernel based on SIFT sparse codes. This new approach remarkably reduces the complexity of SVMs to O(n) in training and a constant in testing. In a number of image categorization experiments, we find that, in terms of classification accuracy, the suggested linear SPM based on sparse coding of SIFT descriptors always significantly outperforms the linear SPM kernel on histograms, and is even better than the nonlinear SPM kernels, leading to state-of-the-art performance on several benchmarks by using a single type of descriptors. Jianchao Yang, Kai Yu 0001, Yihong Gong, Thomas S. Huang |
CVPR | 1 |
| 2009 | Ubiquitously Supervised Subspace LearningabstractIn this paper, our contributions to the subspace learning problem are two-fold. We first justify that most popular subspace learning algorithms, unsupervised or supervised, can be unitedly explained as instances of a ubiquitously supervised prototype. They all essentially minimize the intraclass compactness and at the same time maximize the interclass separability, yet with specialized labeling approaches, such as ground truth, self-labeling, neighborhood propagation, and local subspace approximation. Then, enlightened by this ubiquitously supervised philosophy, we present two categories of novel algorithms for subspace learning, namely, misalignment-robust and semi-supervised subspace learning. The first category is tailored to computer vision applications for improving algorithmic robustness to image misalignments, including image translation, rotation and scaling. The second category naturally integrates the label information from both ground truth and other approaches for unsupervised algorithms. Extensive face recognition experiments on the CMU PIE and FRGC ver1.0 databases demonstrate that the misalignment-robust version algorithms consistently bring encouraging accuracy improvements over the counterparts without considering image misalignments, and also show the advantages of semi-supervised subspace learning over only supervised or unsupervised scheme. Jianchao Yang, Shuicheng Yan, Thomas S. Huang |
IEEE Trans. Image Process. | 1 |
| 2008 | Image super-resolution as sparse representation of raw image patchesabstractThis paper addresses the problem of generating a super-resolution (SR) image from a single low-resolution input image. We approach this problem from the perspective of compressed sensing. The low-resolution image is viewed as downsampled version of a high-resolution image, whose patches are assumed to have a sparse representation with respect to an over-complete dictionary of prototype signal-atoms. The principle of compressed sensing ensures that under mild conditions, the sparse representation can be correctly recovered from the downsampled signal. We will demonstrate the effectiveness of sparsity as a prior for regularizing the otherwise ill-posed super-resolution problem. We further show that a small set of randomly chosen raw patches from training images of similar statistical nature to the input image generally serve as a good dictionary, in the sense that the computed representation is sparse and the recovered high-resolution image is competitive or even superior in quality to images produced by other SR methods. Jianchao Yang, John Wright 0001, Thomas S. Huang, Yi Ma 0001 |
CVPR | 1 |
| 2008 | Non-negative graph embeddingabstractWe introduce a general formulation, called non-negative graph embedding, for non-negative data decomposition by integrating the characteristics of both intrinsic and penalty graphs [17]. In the past, such a decomposition was obtained mostly in an unsupervised manner, such as Non-negative Matrix Factorization (NMF) and its variants, and hence unnecessary to be powerful at classification. In this work, the non-negative data decomposition is studied in a unified way applicable for both unsupervised and supervised/semi-supervised configurations. The ultimate data decomposition is separated into two parts, which separatively preserve the similarities measured by the intrinsic and penalty graphs, and together minimize the data reconstruction error. An iterative procedure is derived for such a purpose, and the algorithmic non-negativity is guaranteed by the non-negative property of the inverse of any M-matrix. Extensive experiments compared with NMF and conventional solutions for graph embedding demonstrate the algorithmic properties in sparsity, classification power, and robustness to image occlusions. Jianchao Yang, Shuicheng Yan, Yun Fu 0001, Xuelong Li 0001, Thomas S. Huang |
CVPR | 1 |
| 2008 | Face hallucination VIA sparse codingabstractIn this paper, we address the problem of hallucinating a high resolution face given a low resolution input face. The problem is approached through sparse coding. To exploit the facial structure, non-negative matrix factorization (NMF) is first employed to learn a localized part-based subspace. This subspace is effective for super-resolving the incoming low resolution face under reconstruction constraints. To further enhance the detailed facial information, we propose a local patch method based on sparse representation with respect to coupled overcomplete patch dictionaries, which can be fast solved through linear programming. Experiments demonstrate that our approach can hallucinate high quality super-resolution faces. Jianchao Yang, Hao Tang 0001, Yi Ma 0001, Thomas S. Huang |
ICIP | 1 |