VLDB 2026 Research / reviewers in the wild / expert
Zhaohui Che
dblp:167/9579
· DBLP profile ↗
20ranked-venue papers
8as first author
7since 2021 · last 2022
0000-0002-7323-4348ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Subjective and Objective Quality of Experience of Free Viewpoint VideosabstractFree viewpoint videos (FVVs) provide immersive experiences for end-users, and they have been applied in many applications, such as movies, sports, and TV shows. However, the development of quantifying the quality of experience (QoE) of FVVs is still relatively slow due to the high costs of data collection and limited public databases. In this paper, we conduct a comprehensive study on FVV QoE. First, we construct the largest, to the best of our knowledge, FVV QoE database called Youku-FVV from two complex real scenarios, i. e., entertainment and sports. Specifically, Youku-FVV originates from the videos captured by dozens of real cameras arranged annularly. We use these videos to generate virtual viewpoints, which make up FVVs together with real views. In constructing the FVV QoE database, we consider both internal and external influencing factors of QoE, which correspond to FVV generation and playback, respectively. Besides, we make an initial attempt to train an efficient no reference FVV QoE prediction model using this database, where several sparse frame sampling strategies are validated. And we demonstrate the feasibility of striving for the balance between effectiveness and efficiency of FVV QoE prediction. The proposed FVV QoE database and source codes are publicly available at https://github.com/QTJiebin/FVV_QoE. Jiebin Yan, Jing Li 0026, Yuming Fang 0001, Zhaohui Che, Xue Xia 0005, Yang Liu 0293 |
IEEE Trans. Image Process. | 4 |
| 2022 | SMGEA: A New Ensemble Adversarial Attack Powered by Long-Term Gradient MemoriesabstractDeep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of source models transfer to other target models and, thus, pose a security threat to black-box applications (when attackers have no access to the target models). Current transfer-based ensemble attacks, however, only consider a limited number of source models to craft an adversarial example and, thus, obtain poor transferability. Besides, recent query-based black-box attacks, which require numerous queries to the target model, not only come under suspicion by the target model but also cause expensive query cost. In this article, we propose a novel transfer-based black-box attack, dubbed serial-minigroup-ensemble-attack (SMGEA). Concretely, SMGEA first divides a large number of pretrained white-box source models into several "minigroups." For each minigroup, we design three new ensemble strategies to improve the intragroup transferability. Moreover, we propose a new algorithm that recursively accumulates the "long-term" gradient memories of the previous minigroup to the subsequent minigroup. This way, the learned adversarial information can be preserved, and the intergroup transferability can be improved. Experiments indicate that SMGEA not only achieves state-of-the-art black-box attack ability over several data sets but also deceives two online black-box saliency prediction systems in real world, i.e., DeepGaze-II (https://deepgaze.bethgelab.org/) and SALICON (http://salicon.net/demo/). Finally, we contribute a new code repository to promote research on adversarial attack and defense over ubiquitous pixel-to-pixel computer vision tasks. We share our code together with the pretrained substitute model zoo at https://github.com/CZHQuality/AAA-Pix2pix. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Xiongkuo Min, Guodong Guo, Patrick Le Callet |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Self-Conditioned Probabilistic Learning of Video RescalingabstractBicubic downscaling is a prevalent technique used to reduce the video storage burden or to accelerate the downstream processing speed. However, the inverse upscaling step is non-trivial, and the downscaled video may also deteriorate the performance of downstream tasks. In this paper, we propose a self-conditioned probabilistic framework for video rescaling to learn the paired downscaling and upscaling procedures simultaneously. During the training, we decrease the entropy of the information lost in the downscaling by maximizing its probability conditioned on the strong spatial-temporal prior information within the downscaled video. After optimization, the downscaled video by our framework preserves more meaningful information, which is beneficial for both the upscaling step and the downstream tasks, e.g., video action recognition task. We further extend the framework to a lossy video compression system, in which a gradient estimator for non-differential industrial lossy codecs is proposed for the end-to-end training of the whole system. Extensive experimental results demonstrate the superiority of our approach on video rescaling, video compression, and efficient action recognition tasks. Yuan Tian 0017, Guo Lu, Xiongkuo Min, Zhaohui Che, Guangtao Zhai, Guodong Guo |
ICCV | 4 |
| 2021 | Saliency4ASD: Challenge, dataset and tools for visual attention modeling for autism spectrum disorder
Jesús Gutiérrez 0001, Zhaohui Che, Guangtao Zhai, Patrick Le Callet |
Signal Process. Image Commun. | 2 |
| 2021 | Adversarial Attack Against Deep Saliency Models Powered by Non-Redundant PriorsabstractSaliency detection is an effective front-end process to many security-related tasks, e.g. automatic drive and tracking. Adversarial attack serves as an efficient surrogate to evaluate the robustness of deep saliency models before they are deployed in real world. However, most of current adversarial attacks exploit the gradients spanning the entire image space to craft adversarial examples, ignoring the fact that natural images are high-dimensional and spatially over-redundant, thus causing expensive attack cost and poor perceptibility. To circumvent these issues, this paper builds an efficient bridge between the accessible partially-white-box source models and the unknown black-box target models. The proposed method includes two steps: 1) We design a new partially-white-box attack, which defines the cost function in the compact hidden space to punish a fraction of feature activations corresponding to the salient regions, instead of punishing every pixel spanning the entire dense output space. This partially-white-box attack reduces the redundancy of the adversarial perturbation. 2) We exploit the non-redundant perturbations from some source models as the prior cues, and use an iterative zeroth-order optimizer to compute the directional derivatives along the non-redundant prior directions, in order to estimate the actual gradient of the black-box target model. The non-redundant priors boost the update of some "critical" pixels locating at non-zero coordinates of the prior cues, while keeping other redundant pixels locating at the zero coordinates unaffected. Our method achieves the best tradeoff between attack ability and perturbation redundancy. Finally, we conduct a comprehensive experiment to test the robustness of 18 state-of-the-art deep saliency models against 16 malicious attacks, under both of white-box and black-box settings, which contributes a new robustness benchmark to the saliency community for the first time. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Yuan Tian 0017, Guodong Guo, Patrick Le Callet |
IEEE Trans. Image Process. | 1 |
| 2021 | Quality Assessment of Free-Viewpoint Videos by Quantifying the Elastic Changes of Multi-Scale Motion TrajectoriesabstractVirtual viewpoints synthesis is an essential process for many immersive applications including Free-viewpoint TV (FTV). A widely used technique for viewpoints synthesis is Depth-Image-Based-Rendering (DIBR) technique. However, such technique may introduce challenging non-uniform spatial-temporal structure-related distortions. Most of the existing state-of-the-art quality metrics fail to handle these distortions, especially the temporal structure inconsistencies observed during the switch of different viewpoints. To tackle this problem, an elastic metric and multi-scale trajectory based video quality metric (EM-VQM) is proposed in this paper. Dense motion trajectory is first used as a proxy for selecting temporal sensitive regions, where local geometric distortions might significantly diminish the perceived quality. Afterwards, the amount of temporal structure inconsistencies and unsmooth viewpoints transitions are quantified by calculating 1) the amount of motion trajectory deformations with elastic metric and, 2) the spatial-temporal structural dissimilarity. According to the comprehensive experimental results on two FTV video datasets, the proposed metric outperforms the state-of-the-art metrics designed for free-viewpoint videos significantly and achieves a gain of 12.86% and 16.75% in terms of median Pearson linear correlation coefficient values on the two datasets compared to the best one, respectively. Suiyi Ling, Jing Li 0026, Zhaohui Che, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Image Process. | 3 |
| 2021 | Re-Visiting Discriminator for Blind Free-Viewpoint Image Quality AssessmentabstractAccurate measurement of perceptual quality is important for various immersive multimedia, which demand real-time quality control or quality-based bench-marking for relevant algorithms. For instance, virtual views rendering in Free-Viewpoint (FV) navigation scenarios is a typical case that introduces challenging distortions, particularly the ones around dis-occluded regions. Existing quality metrics, most of which are targeting for impairments caused by compression or network condition, fail to quantify such non-uniform structure-related distortions. Moreover, the lack of quality databases for such distortions makes it even more challenging to develop robust quality metrics. In this work, a Generative Adversarial Networks based No-Reference (NR) quality Metric, namely GANs-NRM, is proposed. We first present an approach to create masks mimicking dis-occlusions/textureless regions, which is applicable on large-scale 2D image databases publicly available in the computer vision domain. Using these synthetic data, we then train a GANs-based context renderer with the capability of rendering those masked regions. Since the naturalness of the rendered dis-occluded regions strongly relates to the perceptual quality, we assume that the discriminator of the trained GANs has an intrinsic ability for quality assessment. We thus use the features extracted from the discriminator to learn a Bag-of-Distortion-Word (BDW) codebook. We show that a quality predictor can be then well trained using only a small amount of subjective quality data for the FV views rendering. Moreover, in the proposed framework, the discriminator is also adapted as a distortion-detector to locate possible distorted regions. According to the experimental results, the proposed model outperforms significantly the state-of-the-art quality metrics. The corresponding context renderer also shows appealing visualized results over other rendering algorithms. Suiyi Ling, Jing Li 0026, Zhaohui Che, Wei Zhou 0021, Junle Wang, Patrick Le Callet |
IEEE Trans. Multim. | 3 |
| 2020 | A New Ensemble Adversarial Attack Powered by Long-Term Gradient MemoriesabstractDeep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of pre-trained source models can transfer to other new target models, thus pose a security threat to black-box applications (when the attackers have no access to the target models). Despite adopting diverse architectures and parameters, source and target models often share similar decision boundaries. Therefore, if an adversary is capable of fooling several source models concurrently, it can potentially capture intrinsic transferable adversarial information that may allow it to fool a broad class of other black-box target models. Current ensemble attacks, however, only consider a limited number of source models to craft an adversary, and obtain poor transferability. In this paper, we propose a novel black-box attack, dubbed Serial-Mini-Batch-Ensemble-Attack (SMBEA). SMBEA divides a large number of pre-trained source models into several mini-batches. For each single batch, we design 3 new ensemble strategies to improve the intra-batch transferability. Besides, we propose a new algorithm that recursively accumulates the “long-term” gradient memories of the previous batch to the following batch. This way, the learned adversarial information can be preserved and the inter-batch transferability can be improved. Experiments indicate that our method outperforms state-of-the-art ensemble attacks over multiple pixel-to-pixel vision tasks including image translation and salient region prediction. Our method successfully fools two online black-box saliency prediction systems including DeepGaze-II (Kummerer 2017) and SALICON (Huang et al. 2017). Finally, we also contribute a new repository to promote the research on adversarial attack and defense over pixel-to-pixel tasks: https://github.com/CZHQuality/AAA-Pix2pix. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Patrick Le Callet |
AAAI | 1 |
| 2020 | Few-Shot Pill RecognitionabstractPill image recognition is vital for many personal/public health-care applications and should be robust to diverse unconstrained real-world conditions. Most existing pill recognition models are limited in tackling this challenging few-shot learning problem due to the insufficient instances per category. With limited training data, neural network-based models have limitations in discovering most discriminating features, or going deeper. Especially, existing models fail to handle the hard samples taken under less controlled imaging conditions. In this study, a new pill image database, namely CURE, is first developed with more varied imaging conditions and instances for each pill category. Secondly, a W2-net is proposed for better pill segmentation. Thirdly, a Multi-Stream (MS) deep network that captures task-related features along with a novel two-stage training methodology are proposed. Within the proposed framework, a Batch All strategy that considers all the samples is first employed for the sub-streams, and then a Batch Hard strategy that considers only the hard samples mined in the first stage is utilized for the fusion network. By doing so, complex samples that could not be represented by one type of feature could be focused and the model could be forced to exploit other domain-related information more effectively. Experiment results show that the proposed model outperforms state-of-the-art models on both the National Institute of Health (NIH) and our CURE database. Suiyi Ling, Andreas Pastor, Jing Li 0026, Zhaohui Che, Junle Wang, Patrick Le Callet |
CVPR | 4 |
| 2020 | Self-supervised Motion Representation via Scattering Local Motion Cues
Yuan Tian 0017, Zhaohui Che, Wenbo Bao, Guangtao Zhai |
ECCV (14) | 2 |
| 2020 | How is Gaze Influenced by Image Transformations? Dataset and ModelabstractData size is the bottleneck for developing deep saliency models, because collecting eye-movement data is very time-consuming and expensive. Most of current studies on human attention and saliency modeling have used high-quality stereotype stimuli. In real world, however, captured images undergo various types of transformations. Can we use these transformations to augment existing saliency datasets? Here, we first create a novel saliency dataset including fixations of 10 observers over 1900 images degraded by 19 types of transformations. Second, by analyzing eye movements, we find that observers look at different locations over transformed versus original images. Third, we utilize the new data over transformed images, called data augmentation transformation (DAT), to train deep saliency models. We find that label-preserving DATs with negligible impact on human gaze boost saliency prediction, whereas some other DATs that severely impact human gaze degrade the performance. These label-preserving valid augmentation transformations provide a solution to enlarge existing saliency datasets. Finally, we introduce a novel saliency model based on generative adversarial networks (dubbed GazeGAN). A modified U-Net is utilized as the generator of the GazeGAN, which combines classic "skip connection" with a novel "center-surround connection" (CSC) module. Our proposed CSC module mitigates trivial artifacts while emphasizing semantic salient regions, and increases model nonlinearity, thus demonstrating better robustness against transformations. Extensive experiments and comparisons indicate that GazeGAN achieves state-of-the-art performance over multiple datasets. We also provide a comprehensive comparison of 22 saliency models on various transformed scenes, which contributes a new robustness benchmark to saliency community. Our code and dataset are available at. Zhaohui Che, Ali Borji, Guangtao Zhai, Xiongkuo Min, Guodong Guo, Patrick Le Callet |
IEEE Trans. Image Process. | 1 |
| 2019 | A dataset of eye movements for the children with autism spectrum disorderabstractSocial difficulties are the hallmark features of Autism Spectrum Disorder (ASD) and can lead to atypical visual attention towards stimuli. Eye movements encode rich information about attention and psychological factors of an individual, which could help to characterize the traits of ASD. Learning atypical eye movements of the individuals with ASD towards various stimuli is important and has many application scenarios. However, due to the lack of open datasets, research in this sense is still limited. In this work, we present an open dataset of eye movements of children with Autism Spectrum Disorder. It consists of 300 natural scene images and the corresponding eye movement data collected from 14 children with ASD and 14 healthy controls. In particular, fixation maps and scanpaths are available in the dataset. Based on this dataset, researchers could analyze the visual traits of children with ASD and design specialized visual attention models to promote research in related fields, as well as design specialized models to identify the individuals with ASD. The dataset can be accessed in http://doi.org/10.5281/zenodo.2647418 Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Zhaohui Che, Yi Fang 0009, Xiaokang Yang 0001, Jesús Gutiérrez 0001, Patrick Le Callet |
MMSys | 4 |
| 2018 | A Blind Quality Measure for Industrial 2D Matrix Symbols Using Shallow Convolutional Neural NetworkabstractIndustrial two-dimensional (2D) matrix symbols are ubiquitous throughout the automatic assembly lines. Most industrial 2D symbols are corrupted by various inevitable artifacts. State-of-the-art decoding algorithms are not able to directly handle low-quality symbols irrespective of problematic artifacts. Degraded symbols require appropriate preprocessing methods, such as morphology filtering, median filtering, or sharpening filtering, according to specific distortion type. In this paper, we first establish a database including 3000 industrial 2D symbols which are degraded by 6 types of distortions. Second, we utilize a shallow convolutional neural network (CNN) to identify the distortion type and estimate the quality grade for 2D symbols. Finally, we recommend an appropriate preprocessing method for low-quality symbol according to its distortion type and quality grade. Experimental results indicate that the proposed method outperforms state-of-the-art methods in terms of PLCC, SRCC and RMSE. It also promotes decoding efficiency at the cost of low extra time spent. Zhaohui Che, Guangtao Zhai, Jing Liu 0002, Ke Gu 0001, Patrick Le Callet, Jiantao Zhou 0001, Xianming Liu 0005 |
ICIP | 1 |
| 2018 | Learning to Predict where the Children with Asd LookabstractAs is known to us, people with Autism Spectrum Disorder (ASD) have atypical visual attention towards stimuli. Learning the visual attention of people especially, children, with ASD contribute to related research in the field of medicine and psychology. In this paper, we first construct a saliency prediction for children with autism (SPCA) database, which is the first of its kind and consists of 500 images and the corresponding eye tracking data collected from 13 different children with ASD. We compare the performance of five state-of-the-art deep neural networks (DNN)-based saliency prediction approaches with their original networks and the fine-tuned networks on our database. We predict the atypical visual attention of children with ASD for the first time and get the best saliency prediction results for individuals with ASD so far. Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Yi Fang 0009, Zhaohui Che, Xiaokang Yang 0001, Cheng Zhi, Hua Yang 0001 |
ICIP | 5 |
| 2018 | Adaptive Screen Content Image Enhancement Strategy using Layer-based SegmentationabstractThe ubiquitous screen content images (SCIs) play a significant role in various scenarios currently. However, most SCIs captured by consumer devices are frequently corrupted with distortions, especially contrast distortion. Unlike the natural images, SCIs are composed of text, graphics and natural scene pictures so that traditional image enhancement methods are not suitable for these compound images. Therefore, we innovatively proposed an adaptive strategy for enhancing SCIs in this paper. Firstly, we devised a segmentation method to divide SCI into text and pictorial regions. Next, the famous guided image filter (GIF) with big and small kernel sizes served as unsharpness masking for processing different regions adaptively. For verifying performance, the proposed method was tested on recently prevalent SCI datasets including SIQAD, and Webpage Dataset. Experimental results indicate that the proposed approach outperforms state-of-the-art methods in most SCIs with flat background. Zhaohui Che, Guangtao Zhai, Ke Gu 0001, Patrick Le Callet, Xianming Liu 0005, Deming Zhai, Xiao Gu 0001 |
ISCAS | 1 |
| 2018 | BAN, A Barcode Accurate Detection NetworkabstractBarcode detection has been elaborately investigated before. Traditional hand-crafted feature based methods achieve promising performances across some simple and high-quality test databases. However, they usually fall short as a part of barcode decoding pipeline when dealing with challenging scenarios, because of the low robustness and accuracy achieved by low-level features. In this paper, we present a barcode detection network pipeline, designed to generate accurate barcode regions with category labels. We also build a rich barcodes detection dataset-Mixed Barcode Dataset. Experimental results conducted on different barcode databases show the superiority of our method in terms of F-score and decoding accuracy compared to the state of arts. Yuan Tian 0017, Zhaohui Che, Guangtao Zhai |
VCIP | 2 |
| 2017 | Reduced-reference quality metric for screen content imageabstractWith the prevalence of digital products like cellphone, tablet and personal computer, the screen content image (SCI) consisting of text, graphic, and natural scene picture becomes a significant media in various communication scenarios. Consequently, we proposed a reduced-reference quality metric dedicated for SCI. The main contribution includes 2 aspects: 1) we innovatively proposed a layer-based segmentation method to divide SCI into text layer and pictorial layer; 2) we designed respective quality metrics dedicated for text and pictorial layers with a novel pooling strategy considering human visual saliency for SCI. Furthermore, exhaustive experimental results indicate that the proposed metric is highly comparative compared with state-of-the-art full-reference quality metrics. Zhaohui Che, Guangtao Zhai, Ke Gu 0001, Patrick Le Callet |
ICIP | 1 |
| 2016 | No-reference image quality assessment for photographic images of consumer deviceabstractIn this paper we study common, camera-specific kinds of distortions and propose a no-reference image quality assessment algorithm for photographic images produced by consumer devices. Those real consumer-type images, being different from simulated-distortion images, are with realistic artifacts and quality ranges. We find that the state-of-the-art no-reference image quality assessment approaches do not perform well on those photographic images, and propose an approach that achieves high prediction performance on a dataset of consumer-centric images. The proposed method, with no need for the original image, is able to reveal camera-specific problems and differentiate consumer cameras. Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Zhaohui Che |
ICASSP | 4 |
| 2016 | Closing the gap: Visual quality assessment considering viewing conditionsabstractMost of existing visual quality assessment algorithms are tested on standard databases that are created in controlled viewing conditions (e.g. display device, viewing distance and lighting). This implies that all the recoded subjective scores are only valid for the specific settings used in the database. However, with the prevalence of mobile devices, the practical viewing environments can significantly vary from moment to moment. It is our daily experience that the same image can look drastically different on dissimilar devices under changed viewing distance and/or lighting conditions. In other words, a gap exists between the eyes and the visual contents behind the screen in current research of quality assessment. Therefore, in this work, we perform subjective quality evaluation with varied actual viewing conditions. To make the research reproducible, we build a prototype system to record what the eyes really see from the screen and construct the viewing environment-changed image database. The database will be made available to the public. Meanwhile we design a dedicated effective environment-assessing algorithm. We believe that this work will benefit the research of visual quality assessment towards more practical applications. Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Zhaohui Che |
QoMEX | 4 |
| 2015 | A hierarchical saliency detection approach for bokeh imagesabstractBokeh is a popular photograph technique that aesthetically highlights the image subject by properly blurring the background contents. While a large number of bokeh images can be found in our albums, the impacts of bokeh on visual saliency has not been studied yet. Our study shows that traditional saliency models do not perform well on bokeh images in that the foreground/background cannot be efficiently differentiated. Therefore, in this paper we propose a hierarchical saliency model for bokeh images through combining local sharpness measure and foreground saliency detection. More specifically, first we use local sharpness feature as a clue to locate the foreground objects in bokeh image so as to get the first level saliency map. Then we compute the second level saliency map from the blurry background using a robust saliency model. In the third step, we can generate the final saliency map by unequally weighted pooling. In order to evaluate the models' performance quantitatively, we build up a bokeh image saliency database. We test the proposed model against traditional models on the bokeh image database. The results indicate that the proposed model systematically outperforms all traditional saliency models. Zhaohui Che, Guangtao Zhai, Xiongkuo Min |
MMSP | 1 |