EDBT 2026 Demo / reviewers in the wild / expert
Xin Lu 0006
dblp:11/1952-6
· DBLP profile ↗
24ranked-venue papers
7as first author
2since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
15 papers |
Image and video processing · 51% Visual content generation and editing · 22% Image and video coding · 11% | |
| Artificial intelligence
15 papers |
Generative modeling · 25% Deep learning architectures and training · 22% Image recognition and object detection · 13% |
Topics — the 30 heaviest of 44, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › image restoration
image inpainting |
1.4 | 4 | 2019 | Free-Form Image Inpainting With Gated Convolution · ICCV 2019 Foreground-Aware Image Inpainting · CVPR 2019 Generative Image Inpainting With Contextual Attention · CVPR 2018 |
Machine learning › Generative modeling
generative adversarial network |
1.0 | 3 | 2019 | Free-Form Image Inpainting With Gated Convolution · ICCV 2019 Generative Image Inpainting With Contextual Attention · CVPR 2018 Diversified Texture Synthesis with Feed-Forward Networks · CVPR 2017 |
Image and video processing
image restoration |
1.0 | 3 | 2023 | On Efficient Transformer-Based Image Pre-training for Low-Level Vision · IJCAI 2023 Image-Specific Prior Adaptation for Denoising · IEEE Trans. Image Process. 2015 Free-Form Image Inpainting With Gated Convolution · ICCV 2019 |
Rendering › photorealistic rendering
bokeh rendering |
0.8 | 1 | 2024 | Dr.Bokeh: DiffeRentiable Occlusion-Aware Bokeh Rendering · CVPR 2024 |
Computational photography and imaging › depth estimation
depth from defocus |
0.8 | 1 | 2024 | Dr.Bokeh: DiffeRentiable Occlusion-Aware Bokeh Rendering · CVPR 2024 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.7 | 2 | 2018 | Generative Image Inpainting With Contextual Attention · CVPR 2018 MAttNet: Modular Attention Network for Referring Expression Comprehension · CVPR 2018 |
Machine learning › Deep learning architectures and training
transformer |
0.7 | 1 | 2023 | On Efficient Transformer-Based Image Pre-training for Low-Level Vision · IJCAI 2023 |
Image and video processing › super-resolution
image super-resolution |
0.7 | 1 | 2023 | On Efficient Transformer-Based Image Pre-training for Low-Level Vision · IJCAI 2023 |
Image and video processing
low-level vision |
0.7 | 1 | 2023 | On Efficient Transformer-Based Image Pre-training for Low-Level Vision · IJCAI 2023 |
Image and video processing › super-resolution › image super-resolution › attention-based super-resolution
transformer-based super-resolution |
0.7 | 1 | 2023 | On Efficient Transformer-Based Image Pre-training for Low-Level Vision · IJCAI 2023 |
Image and video coding › image quality assessment
image aesthetics assessment |
0.7 | 3 | 2016 | Joint Image and Text Representation for Aesthetics Analysis · ACM Multimedia 2016 Rating Image Aesthetics Using Deep Learning · IEEE Trans. Multim. 2015 RAPID: Rating Pictorial Aesthetics using Deep Learning · ACM Multimedia 2014 |
Computer vision › Image recognition and object detection
image aesthetics assessment |
0.4 | 2 | 2015 | Rating Image Aesthetics Using Deep Learning · IEEE Trans. Multim. 2015 Deep Multi-patch Aggregation Network for Image Style, Aesthetics, and Quality Estimation · ICCV 2015 |
Visual content generation and editing › face editing
facial attribute editing |
0.4 | 1 | 2019 | Semantic Component Decomposition for Face Attribute Manipulation · CVPR 2019 |
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
channel pruning |
0.3 | 1 | 2018 | Rethinking the Smaller-Norm-Less-Informative Assumption in Channel Pruning of Convolution Layers · ICLR (Poster) 2018 |
Machine learning › Deep learning architectures and training › attention mechanism
contextual attention |
0.3 | 1 | 2018 | Generative Image Inpainting With Contextual Attention · CVPR 2018 |
Machine learning › Trustworthy machine learning
dataset bias |
0.3 | 1 | 2018 | Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias · ECCV (2) 2018 |
Machine learning › Generative modeling › diffusion model › image restoration
image inpainting |
0.3 | 1 | 2018 | Generative Image Inpainting With Contextual Attention · CVPR 2018 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2018 | Rethinking the Smaller-Norm-Less-Informative Assumption in Channel Pruning of Convolution Layers · ICLR (Poster) 2018 |
Computer vision › 3D vision › motion estimation
optical flow |
0.3 | 1 | 2018 | Flow-Grounded Spatial-Temporal Video Prediction from Still Images · ECCV (9) 2018 |
Computer vision › Vision and language › visual grounding
referring expression comprehension |
0.3 | 1 | 2018 | MAttNet: Modular Attention Network for Referring Expression Comprehension · CVPR 2018 |
Computer vision › Video understanding and tracking
video prediction |
0.3 | 1 | 2018 | Flow-Grounded Spatial-Temporal Video Prediction from Still Images · ECCV (9) 2018 |
Computer vision › Image recognition and object detection
visual emotion recognition |
0.3 | 1 | 2018 | Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias · ECCV (2) 2018 |
Computer vision › Segmentation and scene understanding
referring image segmentation |
0.3 | 1 | 2017 | Recurrent Multimodal Interaction for Referring Image Segmentation · ICCV 2017 |
Computer vision › Segmentation and scene understanding
scene parsing |
0.3 | 1 | 2017 | Scene Parsing with Global Context Embedding · ICCV 2017 |
Visual content generation and editing › style transfer
arbitrary style transfer |
0.3 | 1 | 2017 | Universal Style Transfer via Feature Transforms · NIPS 2017 |
Image and video processing
feature transform |
0.3 | 1 | 2017 | Universal Style Transfer via Feature Transforms · NIPS 2017 |
Image and video processing › image restoration › image inpainting
high-resolution image inpainting |
0.3 | 1 | 2017 | High-Resolution Image Inpainting Using Multi-scale Neural Patch Synthesis · CVPR 2017 |
Visual content generation and editing › image editing
image compositing |
0.3 | 1 | 2017 | Deep Image Harmonization · CVPR 2017 |
Visual content generation and editing › image editing › image compositing
image harmonization |
0.3 | 1 | 2017 | Deep Image Harmonization · CVPR 2017 |
Visual content generation and editing
style transfer |
0.3 | 1 | 2017 | Universal Style Transfer via Feature Transforms · NIPS 2017 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 2.5transformer · 1.3pre-training · 1.3semantic component model · 0.8occlusion-aware rendering · 0.8layered scene representation · 0.8interactive editing · 0.8differentiable rendering · 0.8deep convolutional neural network · 0.4spectral normalization · 0.4generative adversarial network · 0.4gated convolution · 0.4content completion · 0.4visual attention · 0.3language-based attention · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Dr.Bokeh: DiffeRentiable Occlusion-Aware Bokeh RenderingabstractBokeh is widely used in photography to draw attention to the subject while effectively isolating distractions in the background. Computational methods can simulate bokeh effects without relying on a physical camera lens, but the inaccurate lens modeling in existing filtering-based meth-ods leads to artifacts that need post-processing or learning-based methods to fix. We propose Dr.Bokeh, a novel ren-dering method that addresses the issue by directly correcting the defect that violates physics in the current filtering-based bokeh rendering equation. Dr.Bokeh first preprocesses the input RGBD to obtain a layered scene representation. Dr.Bokeh then takes the layered representation and user-defined lens parameters to render photo-realistic lens blur based on the novel occlusion-aware bokeh rendering method. Experiments show that the non-learning based renderer Dr.Bokeh outperforms state-of-the-art bokeh ren-dering algorithms in terms of photo-realism. In addition, extensive quantitative and qualitative evaluations show that the more accurate lens model pushes the limit of depth-from-defocus. Yichen Sheng, Zixun Yu, Lu Ling, Zhiwen Cao, Xuaner Cecilia Zhang, Xin Lu 0006, Ke Xian, Haiting Lin, Bedrich Benes |
CVPR | 6 |
| 2023 | On Efficient Transformer-Based Image Pre-training for Low-Level VisionabstractPre-training has marked numerous state of the arts in high-level computer vision, while few attempts have ever been made to investigate how pre-training acts in image processing systems. In this paper, we tailor transformer-based pre-training regimes that boost various low-level tasks. To comprehensively diagnose the influence of pre-training, we design a whole set of principled evaluation tools that uncover its effects on internal representations. The observations demonstrate that pre-training plays strikingly different roles in low-level tasks. For example, pre-training introduces more local information to intermediate layers in super-resolution (SR), yielding significant performance gains, while pre-training hardly affects internal feature representations in denoising, resulting in limited gains. Further, we explore different methods of pre-training, revealing that multi-related-task pre-training is more effective and data-efficient than other alternatives. Finally, we extend our study to varying data scales and model sizes, as well as comparisons between transformers and CNNs. Based on the study, we successfully develop state-of-the-art models for multiple low-level tasks. Wenbo Li 0002, Xin Lu 0006, Shengju Qian, Jiangbo Lu |
IJCAI | 2 |
| 2019 | Semantic Component Decomposition for Face Attribute ManipulationabstractDeep neural network-based methods were proposed for face attribute manipulation. There still exist, however, two major issues, i.e., insufficient visual quality (or resolution) of the results and lack of user control. They limit the applicability of existing methods since users may have different editing preference on facial attributes. In this paper, we address these issues by proposing a semantic component model. The model decomposes a facial attribute into multiple semantic components, each corresponds to a specific face region. This not only allows for user control of edit strength on different parts based on their preference, but also makes it effective to remove unwanted edit effect. Further, each semantic component is composed of two fundamental elements, which determine the edit effect and region respectively. This property provides fine interactive control. As shown in experiments, our model not only produces high-quality results, but also allows effective user interaction. Ying-Cong Chen, Xiaohui Shen, Zhe Lin 0001, Xin Lu 0006, I-Ming Pao, Jiaya Jia |
CVPR | 4 |
| 2019 | Foreground-Aware Image InpaintingabstractExisting image inpainting methods typically fill holes by borrowing information from surrounding pixels. They often produce unsatisfactory results when the holes overlap with or touch foreground objects due to lack of information about the actual extent of foreground and background regions within the holes. These scenarios, however, are very important in practice, especially for applications such as distracting object removal. To address the problem, we propose a foreground-aware image inpainting system that explicitly disentangles structure inference and content completion. Specifically, our model learns to predict the foreground contour first, and then inpaints the missing region using the predicted contour as guidance. We show that by such disentanglement, the contour completion model predicts reasonable contours of objects, and further substantially improves the performance of image inpainting. Experiments show that our method significantly outperforms existing methods and achieves superior inpainting results on challenging cases with complex compositions. Wei Xiong 0008, Zhe Lin 0001, Jimei Yang, Xin Lu 0006, Connelly Barnes, Jiebo Luo 0001 |
CVPR | 5 |
| 2019 | Free-Form Image Inpainting With Gated ConvolutionabstractWe present a generative image inpainting system to complete images with free-form mask and guidance. The system is based on gated convolutions learned from millions of images without additional labelling efforts. The proposed gated convolution solves the issue of vanilla convolution that treats all input pixels as valid ones, generalizes partial convolution by providing a learnable dynamic feature selection mechanism for each channel at each spatial location across all layers. Moreover, as free-form masks may appear anywhere in images with any shape, global and local GANs designed for a single rectangular mask are not applicable. Thus, we also present a patch-based GAN loss, named SN-PatchGAN, by applying spectral-normalized discriminator on dense image patches. SN-PatchGAN is simple in formulation, fast and stable in training. Results on automatic image inpainting and user-guided extension demonstrate that our system generates higher-quality and more flexible results than previous methods. Our system helps user quickly remove distracting objects, modify image layouts, clear watermarks and edit faces. Code, demo and models are available at: \url{https://github.com/JiahuiYu/generative_inpainting}. Zhe Lin 0001, Jimei Yang, Xiaohui Shen, Xin Lu 0006, Thomas S. Huang |
ICCV | 5 |
| 2018 | MAttNet: Modular Attention Network for Referring Expression ComprehensionabstractIn this paper, we address referring expression comprehension: localizing an image region described by a natural language expression. While most recent work treats expressions as a single unit, we propose to decompose them into three modular components related to subject appearance, location, and relationship to other objects. This allows us to flexibly adapt to expressions containing different types of information in an end-to-end framework. In our model, which we call the Modular Attention Network (MAttNet), two types of attention are utilized: language-based attention that learns the module weights as well as the word/phrase attention that each module should focus on; and visual attention that allows the subject and relationship modules to focus on relevant image components. Module weights combine scores from all three modules dynamically to output an overall score. Experiments show that MAttNet outperforms previous state-of-the-art methods by a large margin on both bounding-box-level and pixel-level comprehension tasks. Demo1 and code2 are provided. Licheng Yu, Zhe Lin 0001, Xiaohui Shen, Jimei Yang, Xin Lu 0006, Mohit Bansal, Tamara L. Berg |
CVPR | 5 |
| 2018 | Generative Image Inpainting With Contextual AttentionabstractRecent deep learning based approaches have shown promising results for the challenging task of inpainting large missing regions in an image. These methods can generate visually plausible image structures and textures, but often create distorted structures or blurry textures inconsistent with surrounding areas. This is mainly due to ineffectiveness of convolutional neural networks in explicitly borrowing or copying information from distant spatial locations. On the other hand, traditional texture and patch synthesis approaches are particularly suitable when it needs to borrow textures from the surrounding regions. Motivated by these observations, we propose a new deep generative model-based approach which can not only synthesize novel image structures but also explicitly utilize surrounding image features as references during network training to make better predictions. The model is a feedforward, fully convolutional neural network which can process images with multiple holes at arbitrary locations and with variable sizes during the test time. Experiments on multiple datasets including faces (CelebA, CelebA-HQ), textures (DTD) and natural images (ImageNet, Places2) demonstrate that our proposed approach generates higher-quality inpainting results than existing ones. Code, demo and models are available at: https://github.com/JiahuiYu/generative_inpainting. Zhe Lin 0001, Jimei Yang, Xiaohui Shen, Xin Lu 0006, Thomas S. Huang |
CVPR | 5 |
| 2018 | Flow-Grounded Spatial-Temporal Video Prediction from Still Images
Yijun Li 0001, Jimei Yang, Xin Lu 0006, Ming-Hsuan Yang 0001 |
ECCV (9) | 5 |
| 2018 | Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias
Rameswar Panda, Jianming Zhang 0001, Joon-Young Lee, Xin Lu 0006, Amit K. Roy-Chowdhury |
ECCV (2) | 5 |
| 2018 | Rethinking the Smaller-Norm-Less-Informative Assumption in Channel Pruning of Convolution Layers
Jianbo Ye, Xin Lu 0006, Zhe Lin 0001, James Z. Wang 0001 |
ICLR (Poster) | 2 |
| 2017 | An investigation into three visual characteristics of complex scenes that evoke human emotionabstractPrior computational studies have examined hundreds of visual characteristics related to color, texture, and composition in an attempt to predict human emotional responses. Beyond those myriad features examined in computer science, roundness, angularity, and visual complexity have also been found to evoke emotions in human perceivers, as demonstrated in psychological studies of facial expressions, dance poses, and even simple synthetic visual patterns. Capturing these characteristics algorithmically to incorporate in computational studies, however, has proven difficult. Here we expand the scope of previous computer vision work by examining these three visual characteristics in computer analysis of complex scenes, and compare the results to the hundreds of visual qualities previously examined. A large collection of ecologically valid stimuli (i.e., photos that humans regularly encounter on the web), named the EmoSet and containing more than 40,000 images crawled from web albums, was generated using crowd-sourcing and subjected to human subject emotion ratings. We developed computational methods to the separate indices of roundness, angularity, and complexity, thereby establishing three new computational constructs. Critically, these three new physically interpretable visual constructs achieve comparable classification accuracy to the hundreds of shape, texture, composition, and facial feature characteristics previously examined. In addition, our experimental results show that color features related most strongly with the positivity of perceived emotions, the texture features related more to calmness or excitement, and roundness, angularity, and simplicity related similarly with both of these emotions dimensions. Xin Lu 0006, Reginald B. Adams Jr., Jia Li 0001, Michelle G. Newman, James Z. Wang 0001 |
ACII | 1 |
| 2017 | Diversified Texture Synthesis with Feed-Forward NetworksabstractRecent progresses on deep discriminative and generative modeling have shown promising results on texture synthesis. However, existing feed-forward based methods trade off generality for efficiency, which suffer from many issues, such as shortage of generality (i.e., build one network per texture), lack of diversity (i.e., always produce visually identical output) and suboptimality (i.e., generate less satisfying visual effects). In this work, we focus on solving these issues for improved texture synthesis. We propose a deep generative feed-forward network which enables efficient synthesis of multiple textures within one single network and meaningful interpolation between them. Meanwhile, a suite of important techniques are introduced to achieve better convergence and diversity. With extensive experiments, we demonstrate the effectiveness of the proposed model and techniques for synthesizing a large number of textures and show its applications with the stylization. Yijun Li 0001, Jimei Yang, Xin Lu 0006, Ming-Hsuan Yang 0001 |
CVPR | 5 |
| 2017 | Deep Image HarmonizationabstractCompositing is one of the most common operations in photo editing. To generate realistic composites, the appearances of foreground and background need to be adjusted to make them compatible. Previous approaches to harmonize composites have focused on learning statistical relationships between hand-crafted appearance features of the foreground and background, which is unreliable especially when the contents in the two layers are vastly different. In this work, we propose an end-to-end deep convolutional neural network for image harmonization, which can capture both the context and semantic information of the composite images during harmonization. We also introduce an efficient way to collect large-scale and high-quality training data that can facilitate the training process. Experiments on the synthesized dataset and real composite images show that the proposed network outperforms previous state-of-the-art methods. Yi-Hsuan Tsai, Xiaohui Shen, Zhe Lin 0001, Kalyan Sunkavalli, Xin Lu 0006, Ming-Hsuan Yang 0001 |
CVPR | 5 |
| 2017 | High-Resolution Image Inpainting Using Multi-scale Neural Patch Synthesis
Chao Yang 0011, Xin Lu 0006, Zhe Lin 0001, Eli Shechtman, Oliver Wang, Hao Li 0015 |
CVPR | 2 |
| 2017 | Scene Parsing with Global Context EmbeddingabstractWe present a scene parsing method that utilizes global context information based on both the parametric and nonparametric models. Compared to previous methods that only exploit the local relationship between objects, we train a context network based on scene similarities to generate feature representations for global contexts. In addition, these learned features are utilized to generate global and spatial priors for explicit classes inference. We then design modules to embed the feature representations and the priors into the segmentation network as additional global context cues. We show that the proposed method can eliminate false positives that are not compatible with the global context representations. Experiments on both the MIT ADE20K and PASCAL Context datasets show that the proposed method performs favorably against existing methods. Wei-Chih Hung, Yi-Hsuan Tsai, Xiaohui Shen, Zhe Lin 0001, Kalyan Sunkavalli, Xin Lu 0006, Ming-Hsuan Yang 0001 |
ICCV | 6 |
| 2017 | Recurrent Multimodal Interaction for Referring Image SegmentationabstractIn this paper we are interested in the problem of image segmentation given natural language descriptions, i.e. referring expressions. Existing works tackle this problem by first modeling images and sentences independently and then segment images by combining these two types of representations. We argue that learning word-to-image interaction is more native in the sense of jointly modeling two modalities for the image segmentation task, and we propose convolutional multimodal LSTM to encode the sequential interactions between individual words, visual information, and spatial information. We show that our proposed model outperforms the baseline model on benchmark datasets. In addition, we analyze the intermediate output of the proposed multimodal LSTM approach and empirically explain how this approach enforces a more effective word-to-image interaction. Chenxi Liu 0001, Zhe Lin 0001, Xiaohui Shen, Jimei Yang, Xin Lu 0006, Alan L. Yuille |
ICCV | 5 |
| 2017 | Universal Style Transfer via Feature TransformsabstractUniversal style transfer aims to transfer arbitrary visual styles to content images. Existing feed-forward based methods, while enjoying the inference efficiency, are mainly limited by inability of generalizing to unseen styles or compromised visual quality. In this paper, we present a simple yet effective method that tackles these limitations without training on any pre-defined styles. The key ingredient of our method is a pair of feature transforms, whitening and coloring, that are embedded to an image reconstruction network. The whitening and coloring transforms reflect direct matching of feature covariance of the content image to a given style image, which shares similar spirits with the optimization of Gram matrix based cost in neural style transfer. We demonstrate the effectiveness of our algorithm by generating high-quality stylized images with comparisons to a number of recent methods. We also analyze our method by visualizing the whitened features and synthesizing textures by simple feature coloring. Yijun Li 0001, Jimei Yang, Xin Lu 0006, Ming-Hsuan Yang 0001 |
NIPS | 5 |
| 2016 | Joint Image and Text Representation for Aesthetics AnalysisabstractImage aesthetics assessment is essential to multimedia applications such as image retrieval, and personalized image search and recommendation. Primarily relying on visual information and manually-supplied ratings, previous studies in this area have not adequately utilized higher-level semantic information. We incorporate additional textual phrases from user comments to jointly represent image aesthetics utilizing multimodal Deep Boltzmann Machine. Given an image, without requiring any associated user comments, the proposed algorithm automatically infers the joint representation and predicts the aesthetics category of the image. We construct the AVA-Comments dataset to systematically evaluate the performance of the proposed algorithm. Experimental results indicate that the proposed joint representation improves the performance of aesthetics assessment on the benchmarking AVA dataset, comparing with only visual features. Xin Lu 0006, Junping Zhang, James Z. Wang 0001 |
ACM Multimedia | 2 |
| 2015 | Deep Multi-patch Aggregation Network for Image Style, Aesthetics, and Quality EstimationabstractThis paper investigates problems of image style, aesthetics, and quality estimation, which require fine-grained details from high-resolution images, utilizing deep neural network training approach. Existing deep convolutional neural networks mostly extracted one patch such as a down-sized crop from each image as a training example. However, one patch may not always well represent the entire image, which may cause ambiguity during training. We propose a deep multi-patch aggregation network training approach, which allows us to train models using multiple patches generated from one image. We achieve this by constructing multiple, shared columns in the neural network and feeding multiple patches to each of the columns. More importantly, we propose two novel network layers (statistics and sorting) to support aggregation of those patches. The proposed deep multi-patch aggregation network integrates shared feature learning and aggregation function learning into a unified framework. We demonstrate the effectiveness of the deep multi-patch aggregation network on the three problems, i.e., image style recognition, aesthetic quality categorization, and image quality estimation. Our models trained using the proposed networks significantly outperformed the state of the art in all three applications. Xin Lu 0006, Zhe Lin 0001, Xiaohui Shen, Radomír Mech, James Z. Wang 0001 |
ICCV | 1 |
| 2015 | Tree-Based Locally Linear Regression for Image DenoisingabstractWe present a new patch-based approach for image denoising that combines similar patches in the same image and from a set of training images. The key idea of our method is that we can partition the training samples according to the clean patches and efficiently learn a denoising operator for each partition. Given a noisy patch, we use self-similarity to compute an initial denoising result which is used to locate the relevant partitions. We apply the corresponding learned denoising operator to the original noisy patch. Our method does not suffer either from the blurring effect that commonly exists in self-similarity based methods or from the training size problem that is associated with training-based methods. We evaluate our method on three benchmark datasets as well as real mobile images. Experimental results show that our approach consistently outperforms BM3D in terms of both peak signal-to-noise ratio and visual quality. Xin Lu 0006, Zhe Lin 0001, Hailin Jin |
WACV | 1 |
| 2015 | Image-Specific Prior Adaptation for DenoisingabstractImage priors are essential to many image restoration applications, including denoising, deblurring, and inpainting. Existing methods use either priors from the given image (internal) or priors from a separate collection of images (external). We find through statistical analysis that unifying the internal and external patch priors may yield a better patch prior. We propose a novel prior learning algorithm that combines the strength of both internal and external priors. In particular, we first learn a generic Gaussian mixture model from a collection of training images and then adapt the model to the given image by simultaneously adding additional components and refining the component parameters. We apply this image-specific prior to image denoising. The experimental results show that our approach yields better or competitive denoising results in terms of both the peak signal-to-noise ratio and structural similarity. Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Rating Image Aesthetics Using Deep LearningabstractThis paper investigates unified feature learning and classifier training approaches for image aesthetics assessment . Existing methods built upon handcrafted or generic image features and developed machine learning and statistical modeling techniques utilizing training examples. We adopt a novel deep neural network approach to allow unified feature learning and classifier training to estimate image aesthetics. In particular, we develop a double-column deep convolutional neural network to support heterogeneous inputs, i.e., global and local views, in order to capture both global and local characteristics of images . In addition, we employ the style and semantic attributes of images to further boost the aesthetics categorization performance . Experimental results show that our approach produces significantly better results than the earlier reported results on the AVA dataset for both the generic image aesthetics and content -based image aesthetics. Moreover, we introduce a 1.5-million image dataset (IAD) for image aesthetics assessment and we further boost the performance on the AVA test set by training the proposed deep neural networks on the IAD dataset. Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2014 | RAPID: Rating Pictorial Aesthetics using Deep LearningabstractEffective visual features are essential for computational aesthetic quality rating systems. Existing methods used machine learning and statistical modeling techniques on handcrafted features or generic image descriptors. A recently-published large-scale dataset, the AVA dataset, has further empowered machine learning based approaches. We present the RAPID (RAting PIctorial aesthetics using Deep learning) system, which adopts a novel deep neural network approach to enable automatic feature learning. The central idea is to incorporate heterogeneous inputs generated from the image, which include a global view and a local view, and to unify the feature learning and classifier training using a double-column deep convolutional neural network. In addition, we utilize the style attributes of images to help improve the aesthetic quality categorization accuracy. Experimental results show that our approach significantly outperforms the state of the art on the AVA dataset. Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001 |
ACM Multimedia | 1 |
| 2012 | On shape and the computability of emotionsabstractWe investigated how shape features in natural images influence emotions aroused in human beings. Shapes and their characteristics such as roundness, angularity, simplicity, and complexity have been postulated to affect the emotional responses of human beings in the field of visual arts and psychology. However, no prior research has modeled the dimensionality of emotions aroused by roundness and angularity. Our contributions include an in-depth statistical analysis to understand the relationship between shapes and emotions. Through experimental results on the International Affective Picture System (IAPS) dataset we provide evidence for the significance of roundness-angularity and simplicity-complexity on predicting emotional content in images. We combine our shape features with other state-of-the-art features to show a gain in prediction and classification accuracy. We model emotions from a dimensional perspective in order to predict valence and arousal ratings which have advantages over modeling the traditional discrete emotional categories. Finally, we distinguish images with strong emotional content from emotionally neutral images with high accuracy. Xin Lu 0006, Poonam Suryanarayan, Reginald B. Adams Jr., Jia Li 0001, Michelle G. Newman, James Z. Wang 0001 |
ACM Multimedia | 1 |