Jingwen Hou

dblp:276/3246 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
19since 2021 · last 2025
0000-0002-6397-0114ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Frequency-Aware Native Resolution Assessment of 8K Omnidirectional Images
abstract
Omnidirectional images (ODIs) serve as fundamental visual medium for presenting virtual reality (VR) contents, supporting fully immersive experiences through 360-degree scene representation. Typically, a high pixel density is essential for visual quality in VR environments, which in turn requires sufficiently high-resolution imagery to achieve. However, capturing native high-resolution ODIs requires expensive omnidirectional cameras with large sensors (e.g., Insta360 TITAN). An alternative approach is to use low-resolution cameras to acquire original images and then enhance their resolution via super-resolution algorithms. In this work, we explore whether super-resolution ODIs can be easily distinguished from native high-resolution ODIs at 8K scale. To this end, we firstly construct the Native Resolution Assessment of 8K Omnidirectional Images (NRA- 8KODI) dataset, whose native 8K ODIs are collected with an Insta360 TITAN camera and 8K super-resolution images are generated from SOTA open-sourced algorithms. Recognizing high-frequency signals are essential for differentiating non-native 8K ODIs, a frequency-aware model is designed to capture high-frequency details. Specially, to maintain high-frequency details kept in high-resolutions while reduce computational costs brought by high-resolutions, we propose a frequency-aware compressor module to suppress feature channels dominated by low-frequency details. Finally, our model achieves 97.2% accuracy in detecting non-native 8K ODIs, implying that super-resolution for ODIs can still be improved for visual experience in VR applications.
Jingwen Hou, Zengliang Li, Jiebin Yan, Weide Liu, Yuming Fang 0001, Wei Zhou 0021
VCIP1
2025 Improving multi-modal brain tumor segmentation via pre-training and knowledge distillation based post-training
Weide Liu, Jingwen Hou, Xiaoyang Zhong, Huijing Zhan, Jun Cheng 0003, Yuming Fang 0001, Guanghui Yue 0001
Neurocomputing2
2025 Integrating large foundation models into multimodal named entity recognition with evidential fusion
Weide Liu, Xiaoyang Zhong, Jingwen Hou, Haozhe Huang, Wei Zhou 0021, Yuming Fang 0001
Neurocomputing3
2025 Opinion-unaware blind quality assessment of AI-generated omnidirectional images based on deep feature statistics
Xuelin Liu, Jiebin Yan, Yuming Fang 0001, Jingwen Hou
J. Vis. Commun. Image Represent.4
2025 Toward Transparent Deep Image Aesthetics Assessment With Tag-Based Content Descriptors
abstract
Deep learning approaches for Image Aesthetics Assessment (IAA) have shown promising results in recent years, but the internal mechanisms of these models remain unclear. Previous studies have demonstrated that image aesthetics can be predicted using semantic features, such as pre-trained object classification features. However, these semantic features are learned implicitly, and therefore, previous works have not elucidated what the semantic features are representing. In this work, we aim to create a more transparent deep learning framework for IAA by introducing explainable semantic features. To achieve this, we propose Tag-based Content Descriptors (TCDs), where each value in a TCD describes the relevance of an image to a human-readable tag that refers to a specific type of image content. This allows us to build IAA models from explicit descriptions of image contents. We first propose the explicit matching process to produce TCDs that adopt predefined tags to describe image contents. We show that a simple MLP-based IAA model with TCDs only based on predefined tags can achieve an SRCC of 0.767, which is comparable to most state-of-the-art methods. However, predefined tags may not be sufficient to describe all possible image contents that the model may encounter. Therefore, we further propose the implicit matching process to describe image contents that cannot be described by predefined tags. By integrating components obtained from the implicit matching process into TCDs, the IAA model further achieves an SRCC of 0.817, which significantly outperforms existing IAA methods. Both the explicit matching process and the implicit matching process are realized by the proposed TCD generator. To evaluate the performance of the proposed TCD generator in matching images with predefined tags, we also labeled 5101 images with photography-related tags to form a validation set. And experimental results show that the proposed TCD generator can meaningfully assign photography-related tags to images.
Jingwen Hou, Weisi Lin, Yuming Fang 0001, Haoning Wu 0001, Chaofeng Chen, Weide Liu
IEEE Trans. Image Process.1
2025 Diffusion-Based Facial Aesthetics Enhancement With 3D Structure Guidance
abstract
Facial Aesthetics Enhancement (FAE) aims to improve facial attractiveness by adjusting the structure and appearance of a facial image while preserving its identity as much as possible. Most existing methods adopted deep feature-based or score-based guidance for generation models to conduct FAE. Although these methods achieved promising results, they potentially produced excessively beautified results with lower identity consistency or insufficiently improved facial attractiveness. To enhance facial aesthetics with less loss of identity, we propose the Nearest Neighbor Structure Guidance based on Diffusion (NNSG-Diffusion), a diffusion-based FAE method that beautifies a 2D facial image with 3D structure guidance. Specifically, we propose to extract FAE guidance from a nearest neighbor reference face. To allow for less change of facial structures in the FAE process, a 3D face model is recovered by referring to both the matched 2D reference face and the 2D input face, so that the depth and contour guidance can be extracted from the 3D face model. Then the depth and contour clues can provide effective guidance to Stable Diffusion with ControlNet for FAE. Extensive experiments demonstrate that our method is superior to previous relevant methods in enhancing facial aesthetics while preserving facial identity.
Lisha Li, Jingwen Hou, Weide Liu, Yuming Fang 0001, Jiebin Yan
IEEE Trans. Image Process.2
2024 Q-Instruct: Improving Low-Level Visual Abilities for Multi-Modality Foundation Models
abstract
Multi-modality large language models (MLLMs), as represented by GPT-4V, have introduced a paradigm shift for visual perception and understanding tasks, that a variety of abilities can be achieved within one foundation model. While current MLLMs demonstrate primary low-level visual abilities from the identification of low-level visual attributes (e.g., clarity, brightness) to the evaluation on image quality, there's still an imperative to further improve the accuracy of MLLMs to substantially alleviate human burdens. To address this, we collect the first dataset consisting of human natural language feedback on low-level vision. Each feedback offers a comprehensive description of an image's low-level visual attributes, culminating in an overall quality assessment. The constructed Q-Pathway dataset includes 58K detailed human feedbacks on 18,973 multi-sourced images with diverse low-level appearance. To ensure MLLMs can adeptly handle diverse queries, we further propose a GPT-participated transformation to convert these feedbacks into a rich set of 200K instruction-response pairs, termed Q-Instruct. Experimental results indicate that the Q-Instruct consistently elevates various low-level visual capabilities across multiple base models. We anticipate that our datasets can pave the way for a future that foundation models can assist humans on low-level visual tasks.
Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Kaixin Xu, Chunyi Li 0001, Jingwen Hou, Guangtao Zhai, Geng Xue, Wenxiu Sun, Qiong Yan, Weisi Lin
CVPR9
2024 TOPIQ: A Top-Down Approach From Semantics to Distortions for Image Quality Assessment
abstract
Image Quality Assessment (IQA) is a fundamental task in computer vision that has witnessed remarkable progress with deep neural networks. Inspired by the characteristics of the human visual system, existing methods typically use a combination of global and local representations (i.e., multi-scale features) to achieve superior performance. However, most of them adopt simple linear fusion of multi-scale features, and neglect their possibly complex relationship and interaction. In contrast, humans typically first form a global impression to locate important regions and then focus on local details in those regions. We therefore propose a top-down approach that uses high-level semantics to guide the IQA network to focus on semantically important local distortion regions, named as TOPIQ. Our approach to IQA involves the design of a heuristic coarse-to-fine network (CFANet) that leverages multi-scale features and progressively propagates multi-level semantic information to low-level representations in a top-down manner. A key component of our approach is the proposed cross-scale attention mechanism, which calculates attention maps for lower level features guided by higher level features. This mechanism emphasizes active semantic regions for low-level distortions, thereby improving performance. TOPIQ can be used for both Full-Reference (FR) and No-Reference (NR) IQA. We use ResNet50 as its backbone and demonstrate that TOPIQ achieves better or competitive performance on most public FR and NR benchmarks compared with state-of-the-art methods based on vision transformers, while being much more efficient (with only ∼ 13% FLOPS of the current best FR method). Codes are released at https://github.com/chaofengc/IQA-PyTorch.
Chaofeng Chen, Jiadi Mo, Jingwen Hou, Haoning Wu 0001, Wenxiu Sun, Qiong Yan, Weisi Lin
IEEE Trans. Image Process.3
2023 Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives
abstract
The rapid increase in user-generated content (UGC) videos calls for the development of effective video quality assessment (VQA) algorithms. However, the objective of the UGC-VQA problem is still ambiguous and can be viewed from two perspectives: the $\color{Green}{\text{technical perspective}}$, measuring the perception of distortions; and the $\color{Blue}{\text{aesthetic perspective}}$, which relates to preference and recommendation on contents. To understand how these two perspectives affect overall subjective opinions in UGC-VQA, we conduct a large-scale subjective study to collect human quality opinions on the overall quality of videos as well as perceptions from aesthetic and technical perspectives. The collected Disentangled Video Quality Database (DIVIDE-3k) confirms that human quality opinions on UGC videos are universally and inevitably affected by both aesthetic and technical perspectives. In light of this, we propose the Disentangled Objective Video Quality Evaluator (DOVER) to learn the quality of UGC videos based on the two perspectives. The DOVER proves state-of-the-art performance in UGC-VQA under very high efficiency. With perspective opinions in DIVIDE-3k, we further propose DOVER++, the first approach to provide reliable clear-cut quality evaluations from a single aesthetic or technical perspective. Code at https://github.com/VQAssessment/DOVER.
Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin
ICCV5
2023 Exploring Opinion-Unaware Video Quality Assessment with Semantic Affinity Criterion
abstract
Recent learning-based video quality assessment (VQA) algorithms are expensive to implement due to the cost of data collection of human quality opinions, and are less robust across various scenarios due to the biases of these opinions. This motivates our exploration on opinion-unaware (a.k.a zero-shot) VQA approaches. Existing approaches only considers low-level naturalness in spatial or temporal domain, without considering impacts from high-level semantics. In this work, we introduce an explicit semantic affinity index for opinion-unaware VQA using text-prompts in the contrastive language-image pre-training (CLIP) model. We also aggregate it with different traditional low-level naturalness indexes through gaussian normalization and sigmoid rescaling strategies. Composed of aggregated semantic and technical metrics, the proposed Blind Unified Opinion-Unaware Video Quality Index via Semantic and Technical Metric Aggregation (BUONA-VISTA) outperforms existing opinion-unaware VQA methods by at least 20% improvements, and is more robust than opinion-aware approaches.
Haoning Wu 0001, Jingwen Hou, Chaofeng Chen, Erli Zhang 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin
ICME3
2023 Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted Approach
abstract
The proliferation of in-the-wild videos has greatly expanded the Video Quality Assessment (VQA) problem. Unlike early definitions that usually focus on limited distortion types, VQA on in-the-wild videos is especially challenging as it could be affected by complicated factors, including various distortions and diverse contents. Though subjective studies have collected overall quality scores for these videos, how the abstract quality scores relate with specific factors is still obscure, hindering VQA methods from more concrete quality evaluations (e.g. sharpness of a video). To solve this problem, we collect over two million opinions on 4,543 in-the-wild videos on 13 dimensions of quality-related factors, including in-capture authentic distortions (e.g. motion blur, noise, flicker), errors introduced by compression and transmission, and higher-level experiences on semantic contents and aesthetic issues (e.g. composition, camera trajectory), to establish the multi-dimensional Maxwell database. Specifically, we ask the subjects to label among a positive, a negative, and a neutral choice for each dimension. These explanation-level opinions allow us to measure the relationships between specific quality factors and abstract subjective quality ratings, and to benchmark different categories of VQA algorithms on each dimension, so as to more comprehensively analyze their strengths and weaknesses. Furthermore, we propose the MaxVQA, a language-prompted VQA approach that modifies vision-language foundation model CLIP to better capture important quality issues as observed in our analyses. The MaxVQA can jointly evaluate various specific quality factors and final quality scores with state-of-the-art accuracy on all dimensions, and superb generalization ability on existing datasets. Code and data available at https://github.com/VQAssessment/MaxVQA.
Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin
ACM Multimedia5
2023 Neighbourhood Representative Sampling for Efficient End-to-End Video Quality Assessment
abstract
The increased resolution of real-world videos presents a dilemma between efficiency and accuracy for deep Video Quality Assessment (VQA). On the one hand, keeping the original resolution will lead to unacceptable computational costs. On the other hand, existing practices, such as resizing or cropping, will change the quality of original videos due to difference in details or loss of contents, and are henceforth harmful to quality assessment. With obtained insight from the studies of spatial-temporal redundancy in the human visual system, visual quality around a neighbourhood has high probability to be similar, and this motivates us to investigate an effective quality-sensitive neighbourhood representative sampling scheme for VQA. In this work, we propose a unified scheme, spatial-temporal grid mini-cube sampling (St-GMS), and the resultant samples are namedfragments. In St-GMS, full-resolution videos are first divided into mini-cubes with predefined spatial-temporal grids, then the temporal-aligned quality representatives are sampled to compose the fragments that serve as inputs for VQA. In addition, we design the Fragment Attention Network (FANet), a network architecture tailored specifically for fragments. With fragments and FANet, the proposedFAST-VQAandFasterVQA(with an improved sampling scheme) achieves up to 1612× efficiency than the existing state-of-the-art, meanwhile achieving significantly better performance on all relevant VQA benchmarks.
Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Wenxiu Sun, Qiong Yan, Jinwei Gu, Weisi Lin
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 DisCoVQA: Temporal Distortion-Content Transformers for Video Quality Assessment
abstract
Compared with spatial counterparts, temporal relationships between frames and their influences on video quality assessment (VQA) are still relatively under-studied in existing works. These relationships lead to two important types of effects for video quality. Firstly, some meaningless temporal variations (such as shaking, flicker, and unsmooth scene transitions) cause temporal distortions that degrade quality of videos. Secondly, the human visual system often has different attention to frames with different contents, resulting in their different importance to the overall video quality. Based on prominent time-series modeling ability of transformers, we propose a novel and effective transformer-based VQA method to tackle these two issues. To better differentiate temporal variations and thus capture the temporal distortions, we design the Spatial-Temporal Distortion Extraction (STDE) module that extracts multi-level spatial-temporal features with a video swin transformer tiny (Swin-T) backbone and uses temporal difference layer to further capture these distortions. To tackle with temporal quality attention, we propose the encoder-decoder-like temporal content transformer (TCT). We also introduce the temporal sampling on features to reduce the input length for the TCT, so as to improve the learning effectiveness and efficiency of this module. Consisting of the STDE and the TCT, the proposed Temporal Distortion-Content Transformers for Video Quality Assessment (DisCoVQA) reaches state-of-the-art performance on several VQA benchmarks without any extra pre-training datasets and up to 10% better generalization ability than existing methods. We also conduct extensive ablation experiments to prove the effectiveness of each part in our proposed model, and provide visualizations to prove that the proposed modules achieve our intention on modeling these temporal issues. Our code is published athttps://github.com/QualityAssessment/DisCoVQA.
Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Wenxiu Sun, Qiong Yan, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.4
2023 Perceptual Quality Assessment of Enhanced Colonoscopy Images: A Benchmark Dataset and an Objective Method
abstract
In colonoscopy, the captured images are usually with low-quality appearance, such as non-uniform illumination, low contrast, etc., due to the specialized imaging environment, which may provide poor visual feedback and bring challenges to subsequent disease analysis. Many low-light image enhancement (LIE) algorithms have recently proposed to improve the perceptual quality. However, how to fairly evaluate the quality of enhanced colonoscopy images (ECIs) generated by different LIE algorithms remains a rarely-mentioned and challenging problem. In this study, we carry out a pioneering investigation on perceptual quality assessment of ECIs. Firstly, considering the lack of specific datasets, we collect 300 low-light images with diverse contents during the real-world colonoscopy and conduct rigorous subjective studies to compare the performance of 8 popular LIE methods, resulting in a benchmark dataset (named ECIQAD) for ECIs. Secondly, in view of the distinctive distortion characteristics of ECIs, we propose an effective no-reference Enhanced Colonoscopy Image Quality (ECIQ) method to automatically evaluate the perceptual quality of ECIs via analysis of brightness, contrast, colorfulness, naturalness, and noise. Extensive experiments on ECIQAD demonstrate the superiority of our proposed ECIQ method over 14 mainstream no-reference image quality assessment methods.
Guanghui Yue 0001, Tianwei Zhou, Jingwen Hou, Weide Liu, Long Xu 0001, Tianfu Wang 0001, Jun Cheng 0003
IEEE Trans. Circuits Syst. Video Technol.4
2023 Benchmarking Polyp Segmentation Methods in Narrow-Band Imaging Colonoscopy Images
abstract
In recent years, there has been significant progress in polyp segmentation in white-light imaging (WLI) colonoscopy images, particularly with methods based on deep learning (DL). However, little attention has been paid to the reliability of these methods in narrow-band imaging (NBI) data. NBI improves visibility of blood vessels and helps physicians observe complex polyps more easily than WLI, but NBI images often include polyps with small/flat appearances, background interference, and camouflage properties, making polyp segmentation a challenging task. This paper proposes a new polyp segmentation dataset (PS-NBI2K) consisting of 2,000 NBI colonoscopy images with pixel-wise annotations, and presents benchmarking results and analyses for 24 recently reported DL-based polyp segmentation methods on PS-NBI2K. The results show that existing methods struggle to locate polyps with smaller sizes and stronger interference, and that extracting both local and global features improves performance. There is also a trade-off between effectiveness and efficiency, and most methods cannot achieve the best results in both areas simultaneously. This work highlights potential directions for designing DL-based polyp segmentation methods in NBI colonoscopy images, and the release of PS-NBI2K aims to drive further development in this field.
Guanghui Yue 0001, Guibin Zhuo, Tianwei Zhou, Jingfeng Du, Weiqing Yan, Jingwen Hou, Weide Liu, Tianfu Wang 0001
IEEE J. Biomed. Health Informatics7
2023 Interaction-Matrix Based Personalized Image Aesthetics Assessment
abstract
Personalized image aesthetics assessment (IAA) aims to estimate aesthetic experiences subject to the preferences of individual users, contrary to generic IAA that estimates aesthetic experiences subject to average preferences. Most existing personalized IAA methods treat personalized aesthetic experiences as deviations from a generic aesthetic experience, and therefore, personalized IAA models are designed to build upon the prior knowledge on generic IAA. However, we propose that acquiring knowledge on generic IAA is not necessary for building a personalized IAA model. Instead of modeling personalized IAA on the basis of generic IAA, this work proposes to directly estimate personalized aesthetic experiences from the interactions between image contents and user preferences (i.e., preference-content interaction), where interaction-matrices representing preference-content interactions are constructed without needs for prior generic IAA knowledge. To this end, we construct interaction-matrices from content features constructed from pre-trained image classification features and latent preference features. To realize a robust interaction-matrix based personalized IAA model, we discuss in detail on different strategies for constructing interaction-matrices and estimating personalized aesthetic scores from the interaction-matrices. Besides the personalized IAA scenario, we further propose strategies to adapt the proposed personalized IAA model to different scenarios of generic IAA. Extensive experiments show that: 1) our method significantly outperforms 5 previous relevant personalized IAA methods on FLICKR-AES dataset, especially the methods that require generic IAA knowledge as the basis; 2) in terms of generic IAA, the proposed approach also outperforms 13 generic IAA methods on AVA dataset.
Jingwen Hou, Weisi Lin, Guanghui Yue 0001, Weide Liu, Baoquan Zhao
IEEE Trans. Multim.1
2022 Extreme Systematic Reviews: A Large Literature Screening Dataset to Support Environmental Policymaking
abstract
The United States Environmental Protection Agency (EPA) periodically releases Integrated Science Assessments (ISAs) that synthesize the latest research on each of six air pollutants to inform environmental policymaking. To guarantee the best possible coverage of relevant literature, EPA scientists spend months manually screening hundreds of thousands of references to identify a small proportion to be cited in an ISA. The challenge of extreme scale and the pursuit of maximum recall calls for effective machine-assisted approaches to reducing the time and effort required by the screening process. This work introduces the ISA literature screening dataset and the associated research challenges to the information and knowledge management community. Our pilot experiments show that combining multiple approaches in tackling this challenge is both promising and necessary. The dataset is available at https://catalog.data.gov/dataset/isa-literature-screening-dataset-v-1.
Jingwen Hou, Jean-Jacques Dubois, R. Byron Rice, Amanda Haddock, Yue Wang 0035
CIKM1
2022 FAST-VQA: Efficient End-to-End Video Quality Assessment with Fragment Sampling
Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin
ECCV (6)3
2022 Distilling Knowledge From Object Classification to Aesthetics Assessment
abstract
In this work, we point out that the major dilemma of image aesthetics assessment (IAA) comes from the abstract nature of aesthetic labels. That is, a vast variety of distinct contents can correspond to the same aesthetic label. On the one hand, during inference, the IAA model is required to relate various distinct contents to the same aesthetic label. On the other hand, when training, it would be hard for the IAA model to learn to distinguish different contents merely with the supervision from aesthetic labels, since aesthetic labels are not directly related to any specific content. To deal with this dilemma, we propose to distill knowledge on semantic patterns for a vast variety of image contents from multiple pre-trained object classification (POC) models to an IAA model. Expecting the combination of multiple POC models can provide sufficient knowledge on various image contents, the IAA model can easier learn to relate various distinct contents to a limited number of aesthetic labels. By supervising an end-to-end single-backbone IAA model with the distilled knowledge, the performance of the IAA model is significantly improved by 4.8% in SRCC compared to the version trained only with ground-truth aesthetic labels. On specific categories of images, the SRCC improvement brought by the proposed method can achieve up to 7.2%. Peer comparison also shows that our method outperforms 10 previous IAA methods.
Jingwen Hou, Henghui Ding, Weisi Lin, Weide Liu, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Content-Dependency Reduction With Multi-Task Learning In Blind Stitched Panoramic Image Quality Assessment
abstract
In this work, we investigate deep learning based solutions to blind quality assessment of stitched panoramic images (SPI). The main problem to tackle is that the ground truth data is usually insufficient. As a result, the learned model can easily overfit data with specific content. Because most distortions of SPIs lie within local regions, the problem cannot be alleviated by commonly-used patch-wise training, which assumes local quality equals global quality. We propose a multi-task learning strategy which encourages learned representation to be less dependent on image content. A siamese network with two weight-shared CNN branches is trained to simultaneously compare the quality of two images of the same scene and predict the quality score of each image. Since two images of the same scene are processed by the same CNN, the CNN tends to find their quality differences instead of content differences under the constraint of the quality ranking objective. Because two tasks share the same representations learned by the CNN, the regression task can be further benefited from the quality-sensitive representations. Extensive experiments demonstrate the effectiveness of the proposed model and its superiority over existing SPI quality assessment methods.
Jingwen Hou, Weisi Lin, Baoquan Zhao
ICIP1
2020 Attention-driven Unsupervised Image Retrieval for Beauty Products with Visual and Textual Clues
abstract
Beauty and personal care product retrieval (BPCR) aims to match a query image of an item to examples of the same item in a large database. The task is extremely challenging because a small number of ground-truth examples have to be found in a large search space. Previous works mostly search only with visual representations and have not made full use of the product descriptions. Since many noisy examples only have subtle visual differences comparing to the ground-truth examples (e.g. similar packaging but different brands) and those differences (e.g. product brands) are especially hard to be captured only by visual features, methods merely based on visual feature similarities can easily regard those noisy examples as examples of the same item in the query image. We notice that the product descriptions are good sources for capturing those subtle visual differences. Therefore, we propose a search method utilizing both images and product descriptions in this work. Before searching, we not only prepare attention-based visual features for each database image but also a textual index (TI) that matches each database example to other examples with similar product descriptions. During searching, the visual feature of the query image is firstly searched in the whole database and then searched in a subset obtained by looking up the TI. Finally, the second result is used to refine the initial result. Since the subset examples usually have similar properties (e.g. brands and type), the noisy examples in the initial result can be effectively replaced. We have experimentally proved the effectiveness of the proposed method on the validation set of the Perfect-500K dataset. Our team (NTU-Beauty) achieved the 3rd place in the leader board of the Grand Challenge of AI Meets Beauty in ACM Multimedia 2020. Our code is available at: https://github.com/jingwenh/2020-ai-meets-beauty_ntubeauty.git.
Jingwen Hou, Sijie Ji, Annan Wang
ACM Multimedia1
2020 Object-level Attention for Aesthetic Rating Distribution Prediction
abstract
We study the problem of image aesthetic assessment (IAA) and aim to automatically predict the image aesthetic quality in the form of discrete distribution, which is particularly important in IAA due to its nature of having possibly higher diversification of agreement for aesthetics. Previous works show the effectiveness of utilizing object-agnostic attention mechanisms to selectively concentrate on more contributive regions for IAA, e.g., attention is learned to weight pixels of input images when inferring aesthetic values. However, as suggested by some neuropsychology studies, the basic units of human attention are visual objects, i.e., the trace of human attention follows a series of objects. This inspires us to predict contributions of different regions at object level for better aesthetics evaluation. With our framework, region-of-interests (RoIs) are proposed by an object detector, and each RoI is associated with a regional feature vector. Then the contribution of each regional feature to the aesthetics prediction is adaptively determined. To the best of our knowledge, this is the first work modeling object-level attention for IAA and experimental results confirm the superiority of our framework over previous relevant methods.
Jingwen Hou, Sheng Yang 0006, Weisi Lin
ACM Multimedia1