EDBT 2026 Demo / reviewers in the wild / expert
Jing Li 0026
dblp:l/JingLi26
· DBLP profile ↗
34ranked-venue papers
11as first author
21since 2021 · last 2026
0000-0001-8645-9709ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 8 first-author · 17 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep Underwater Image Quality Assessment via Progressive Physics-Aware Multi-Prior CollaborationabstractUnderwater image quality assessment (UIQA) is a critical research area, challenged by underwater environments such as wavelength-dependent light attenuation, scattering, and non-uniform illumination. Existing deep learning-based UIQA methods often address these degradations in isolation, neglecting their complex interplay with human perception and lacking explicit modeling of underwater optical phenomena. To address this, we propose PhysIQ-Net, a novel framework that integrates physics-driven principles with progressive multi-prior interaction modeling through three key innovations: First, introduce dual physics-based decomposition that separates images into Backscatter, Transmission, Reflectance, and Illuminance components to capture distinct degradation mechanisms; Second, propose prior-guided dynamic filtering that adapts convolutional kernels to image-specific content using physical priors; and Third, propose physic-informed Cross-Domain Feature Interaction that enables bidirectional collaboration between color-aware and structure-aware representations to model their perceptual inter-dependencies. Extensive experiments across multiple benchmark datasets demonstrate that PhysIQ-Net significantly outperforms existing methods, with ablation studies validating each component’s contribution, providing a robust solution for UIQA. Zihan Zhou 0007, Jiaxue Lan, Yun Liang 0003, Jing Li 0026, Yong Xu 0007, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multi-view Clustering Based on Probabilistic Tensor RegressionabstractMulti-view clustering based on anchor graph and regression is widely used to deal with high dimensional and redundant data. However, most of these methods ignore the probabilistic characteristics of anchor graph, and the effective information in different views is not fully mined. To solve these problems, we propose a multi-view clustering method based on probabilistic tensor regression (MVCPTR). Specifically, we reinterpret the regression process of the anchor graph from the perspective of probability. By modeling the anchor graph as the transition probability from samples to anchors, we construct the implicit relationship between labels of samples and anchors. In order to further mine the complementary information of multi-view data, we extend the anchor graph matrix regression to tensor regression to achieve multi-level information fusion at the representational level and decision level, and impose the Schatten p-norm constraint on the anchor label tensor and the sample label tensor to realize the bi-clustering of the anchors and samples. A large number of experiments prove the effectiveness of our proposed algorithm. Yichen Bao, Yu Duan 0001, Jing Li 0026, Quanxue Gao |
ACM Multimedia | 4 |
| 2025 | Enhancing CNN-Based Blind Image Quality Assessment via Deep Cross-Layer Pattern EncodingabstractEvaluating image quality without reference images, known as blind image quality assessment (BIQA), is crucial for image communication. Recently, convolutional neural networks (CNNs) have emerged as a prominent BIQA approach due to their feature learning power. Usually, both high-level semantic information and low-level details significantly impact perceived visual quality. However, most existing CNN-based methods focus on high-level semantic information via aggregating features on top of the last convolutional layer into a global descriptor, neglecting the importance of shallow, low-level cues. To address this limitation, this paper proposes a novel approach that exploits local encoding and histogram-based pyramid pooling on crosslayer features produced by a CNN, achieving a joint local and global analysis. Specifically, we introduce a cross-layer pattern encoding model that characterizes features generated along convolutional layers via a soft histogram of local 3D binary patterns. This leads to a highly informative yet compact descriptor for score regression. By building this module into a ResNet backbone, we present an effective BIQA model demonstrating state-ofthe-art performance in extensive experiments on synthetic and authentic datasets. Zihan Zhou 0007, Yong Xu 0007, Yuhui Quan, Yun Liang 0003, Jing Li 0026, Patrick Le Callet |
IEEE Trans. Multim. | 5 |
| 2024 | Tensorized Label Learning on Anchor GraphabstractGraph-based multimedia data clustering has attracted much attention due to the impressive clustering performance for arbitrarily shaped multimedia data. However, existing graph-based clustering methods need post-processing to get labels for multimedia data with high computational complexity. Moreover, it is sub-optimal for label learning due to the fact that they exploit the complementary information embedded in data with different types pixel by pixel. To handle these problems, we present a novel label learning model with good interpretability for clustering. To be specific, our model decomposes anchor graph into the products of two matrices with orthogonal non-negative constraint to directly get soft label without any post-processing, which remarkably reduces the computational complexity. To well exploit the complementary information embedded in multimedia data, we introduce tensor Schatten p-norm regularization on the label tensor which is composed of soft labels of multimedia data. The solution can be obtained by iteratively optimizing four decoupled sub-problems, which can be solved more efficiently with good convergence. Experimental results on various datasets demonstrate the efficiency of our model. Jing Li 0026, Quanxue Gao, Qianqian Wang 0001, Wei Xia 0007 |
AAAI | 1 |
| 2024 | Label Learning Method Based on Tensor ProjectionabstractMulti-view clustering method based on anchor graph has been widely concerned due to its high efficiency and effectiveness. In order to avoid post-processing, most of the existing anchor graph-based methods learn bipartite graphs with connected components. However, such methods have high requirements on parameters, and in some cases it may not be possible to obtain bipartite graphs with clear connected components. To end this, we propose a label learning method based on tensor projection (LLMTP). Specifically, we project anchor graph into the label space through an orthogonal projection matrix to obtain cluster labels directly. Considering that the spatial structure information of multi-view data may be ignored to a certain extent when projected in different views separately, we extend the matrix projection transformation to tensor projection, so that the spatial structure information between views can be fully utilized. In addition, we introduce the tensor Schatten p-norm regularization to make the clustering label matrices of different views as consistent as possible. Extensive experiments have proved the effectiveness of the proposed method. Jing Li 0026, Quanxue Gao, Qianqian Wang 0001, Cheng Deng 0002, De-Yan Xie |
KDD | 1 |
| 2024 | Highly Efficient No-reference 4K Video Quality Assessment with Full-Pixel Covering Sampling and Training StrategyabstractDeep Video Quality Assessment (VQA) methods have shown impressive high-performance capabilities. Notably, no-reference (NR) VQA methods play a vital role in situations where obtaining reference videos is restricted or not feasible. Nevertheless, as more streaming videos are being created in ultra-high definition (e.g., 4K) to enrich viewers' experiences, the current deep VQA methods face unacceptable computational costs. Furthermore, the resizing, cropping, and local sampling techniques employed in these methods can compromise the details and content of original 4K videos, thereby negatively impacting quality assessment. In this paper, we propose a highly efficient and novel NR 4K VQA technology. Specifically, first, a novel data sampling and training strategy is proposed to tackle the problem of excessive resolution. This strategy allows the VQA Swin Transformer-based model to effectively train and make inferences using the full data of 4K videos on standard consumer-grade GPUs without compromising content or details. Second, a weighting and scoring scheme is developed to mimic the human subjective perception mode, which is achieved by considering the distinct impact of each sub-region within a 4K frame on the overall perception. Third, we incorporate the frequency domain information of video frames to better capture the details that affect video quality, consequently further improving the model's generalizability. To our knowledge, this is the first technology for the NR 4K VQA task. Thorough empirical studies demonstrate it not only significantly outperforms existing methods on a specialized 4K VQA dataset but also achieves state-of-the-art performance across multiple open-source NR video quality datasets. Xiaoheng Tan, Jiabin Zhang, Yuhui Quan, Jing Li 0026, Yajing Wu, Zilin Bian |
ACM Multimedia | 4 |
| 2024 | Deep Blind Image Quality Assessment Using Dynamic Neural Model With Dual-Order StatisticsabstractDeep convolutional neural networks (CNNs) have increasingly become a prominent method for blind image quality assessment (BIQA). The process of quality assessment typically involves feature extraction, average-based pooling, and quality regression. Based on this process, as well as the consensus that the visual quality of an image mainly relies on its content and distortions, this work improves CNNs for BIQA in two ways. First, considering the content-awareness of visual quality perception, we incorporate content-awareness via a dynamic filtering module to extract content-adaptive features and a dynamic regression module to learn content-adaptive perception rules based on local content and global semantics. Second, considering distortion-sensitivity in visual quality perception, we introduce second-order global variance pooling and combine it with global average pooling (GAP). First-order pooling methods like GAP are limited in distinguishing complex distortions that cause local degradation while preserving global features. Thus, pooling with dual-order statistics enables a more distortion-sensitive and discriminative global representation. These two improvements result in a content-adaptive BIQA model with a dual-order global pooling mechanism, improving generalization on diverse images with varying contents and distortion types. Extensive experiments on synthetic and authentic distortion datasets demonstrate state-of-the-art performance of the proposed approach. Zihan Zhou 0007, Jing Li 0026, Dexiang Zhong, Yong Xu 0007, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Efficient Anchor Graph Factorization for Multi-View ClusteringabstractDue to the excellent interpretability of non-negative matrix factorization (NMF), NMF-based multi-view clustering has attracted much attention for multi-media data analysis and processing. However, the existing clustering methods leverage NMF to cluster data matrix, resulting in high computational complexity. Moreover, they are sub-optimal to exploit the complementary information between views because they all measure the between-views error pixel by pixel. To tackle this problem, inspired by orthogonal NMF and anchor graph, we present an efficient anchor graph factorization model with orthogonal, non-negative, and tensor low-rank constraints. We use an anchor graph instead of a data matrix to get an indicator matrix without post-processing, which remarkably reduces the computational complexity. To exploit the between-views complementary information well, we introduce tensor Schatten$p$-norm regularization on the third tensor, composed of soft label matrices of views. The solution can be obtained by iteratively optimizing four decoupled sub-problems, which can be solved more efficiently with good convergence. Through experimental results on the six multi-view datasets, our approach ensures the enhancement of clustering performance while improving efficiency. Jing Li 0026, Qianqian Wang 0001, Ming Yang 0024, Quanxue Gao, Xinbo Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Orthogonal Non-negative Tensor Factorization based Multi-view ClusteringabstractMulti-view clustering (MVC) based on non-negative matrix factorization (NMF) and its variants have attracted much attention due to their advantages in clustering interpretability. However, existing NMF-based multi-view clustering methods perform NMF on each view respectively and ignore the impact of between-view. Thus, they can't well exploit the within-view spatial structure and between-view complementary information. To resolve this issue, we present orthogonal non-negative tensor factorization (Orth-NTF) and develop a novel multi-view clustering based on Orth-NTF with one-side orthogonal constraint. Our model directly performs Orth-NTF on the 3rd-order tensor which is composed of anchor graphs of views. Thus, our model directly considers the between-view relationship. Moreover, we use the tensor Schatten $p$-norm regularization as a rank approximation of the 3rd-order tensor which characterizes the cluster structure of multi-view data and exploits the between-view complementary information. In addition, we provide an optimization algorithm for the proposed method and prove mathematically that the algorithm always converges to the stationary KKT point. Extensive experiments on various benchmark datasets indicate that our proposed method is able to achieve satisfactory clustering performance. Jing Li 0026, Quanxue Gao, Qianqian Wang 0001, Ming Yang 0024, Wei Xia 0007 |
NeurIPS | 1 |
| 2023 | Low-rank discrete multi-view spectral clustering
Yu Yun, Jing Li 0026, Quanxue Gao, Ming Yang 0024, Xinbo Gao 0001 |
Neural Networks | 2 |
| 2022 | Saliency in Augmented RealityabstractWith the rapid development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary theory underlying AR is human visual confusion, which allows users to perceive the real-world scenes and augmented contents (virtual-world scenes) simultaneously by superimposing them together. To achieve good Quality of Experience (QoE), it is important to understand the interaction between two scenarios, and harmoniously display AR contents. However, studies on how this superimposition will influence the human visual attention are lacking. Therefore, in this paper, we mainly analyze the interaction effect between background (BG) scenes and AR contents, and study the saliency prediction problem in AR. Specifically, we first construct a Saliency in AR Dataset (SARD), which contains 450 BG images, 450 AR images, as well as 1350 superimposed images generated by superimposing BG and AR images in pair with three mixing levels. A large-scale eye-tracking experiment among 60 subjects is conducted to collect eye movement data. To better predict the saliency in AR, we propose a vector quantized saliency prediction method and generalize it for AR saliency prediction. For comparison, three benchmark methods are proposed and evaluated together with our proposed method on our SARD. Experimental results demonstrate the superiority of our proposed method on both of the common saliency prediction problem and the AR saliency prediction problem over benchmark methods. Our dataset and code are available at: https://github.com/DuanHuiyu/ARSaliency. Huiyu Duan, Wei Shen 0002, Xiongkuo Min, Danyang Tu, Jing Li 0026, Guangtao Zhai |
ACM Multimedia | 5 |
| 2022 | Image Quality Assessment: From Mean Opinion Score to Opinion Score DistributionabstractRecently, many methods have been proposed to predict the image quality which is generally described by the mean opinion score (MOS) of all subjective ratings given to an image. However, few efforts focus on predicting the opinion score distribution of the image quality ratings. In fact, the opinion score distribution reflecting subjective diversity, uncertainty, etc., can provide more subjective information about the image quality than a single MOS, which is worthy of in-depth study. In this paper, we propose a convolutional neural network based on fuzzy theory to predict the opinion score distribution of image quality. The proposed method consists of three main steps: feature extraction, feature fuzzification and fuzzy transfer. Specifically, we first use the pre-trained VGG16 without fully-connected layers to extract image features. Then, the extracted features are fuzzified by fuzzy theory, which is used to model epistemic uncertainty in the process of feature extraction. Finally, a fuzzy transfer network is used to predict the opinion score distribution of image quality by learning the mapping from epistemic uncertainty to the uncertainty existing in the image quality ratings. In addition, a new loss function is designed based on the subjective uncertainty of the opinion score distribution. Extensive experimental results prove the superior prediction performance of our proposed method. Xiongkuo Min, Yucheng Zhu, Jing Li 0026, Xiao-Ping Zhang 0002, Guangtao Zhai |
ACM Multimedia | 4 |
| 2022 | QoEVMA'22: 2nd Workshop on Quality of Experience (QoE) in Visual Multimedia ApplicationsabstractNowadays, people spend dramatically more time on watching videos through different devices. The advanced hardware technology and network allow for the increasing demands of users viewing experience. Thus, enhancing the Quality of Experience of end-users in advanced multimedia is the ultimate goal of service providers, as good services would attract more consumers. Quality assessment is thus important. The second workshop on "Quality of Experience (QoE) in visual multimedia applications" (QoEVMA'22) focuses on the QoE assessment of any visual multimedia applications both subjectively and objectively. The topics include 1) QoE assessment on different visual multimedia applications, including VoD for movies, dramas, variety shows, UGC on social networks, live streaming videos for gaming/shopping/social, etc. 2) QoE assessment for different video formats in multimedia services, including 2D, stereoscopic 3D, High Dynamic Range (HDR), Augmented Reality (AR), Virtual Reality (VR), 360, Free-Viewpoint Video(FVV), etc. 3) Key performance indicators (KPI) analysis for QoE. This summary gives a brief overview of the workshop, which took place on October 14, 2022 in Lisbon, Portugal, as a half-day workshop. The complete QOEVMA'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552469 Jing Li 0026, Patrick Le Callet, Xinbo Gao 0001, Zhi Li 0001, Wen Lu 0004, Junle Wang |
ACM Multimedia | 1 |
| 2022 | Subjective and Objective Quality of Experience of Free Viewpoint VideosabstractFree viewpoint videos (FVVs) provide immersive experiences for end-users, and they have been applied in many applications, such as movies, sports, and TV shows. However, the development of quantifying the quality of experience (QoE) of FVVs is still relatively slow due to the high costs of data collection and limited public databases. In this paper, we conduct a comprehensive study on FVV QoE. First, we construct the largest, to the best of our knowledge, FVV QoE database called Youku-FVV from two complex real scenarios, i. e., entertainment and sports. Specifically, Youku-FVV originates from the videos captured by dozens of real cameras arranged annularly. We use these videos to generate virtual viewpoints, which make up FVVs together with real views. In constructing the FVV QoE database, we consider both internal and external influencing factors of QoE, which correspond to FVV generation and playback, respectively. Besides, we make an initial attempt to train an efficient no reference FVV QoE prediction model using this database, where several sparse frame sampling strategies are validated. And we demonstrate the feasibility of striving for the balance between effectiveness and efficiency of FVV QoE prediction. The proposed FVV QoE database and source codes are publicly available at https://github.com/QTJiebin/FVV_QoE. Jiebin Yan, Jing Li 0026, Yuming Fang 0001, Zhaohui Che, Xue Xia 0005, Yang Liu 0293 |
IEEE Trans. Image Process. | 2 |
| 2022 | SMGEA: A New Ensemble Adversarial Attack Powered by Long-Term Gradient MemoriesabstractDeep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of source models transfer to other target models and, thus, pose a security threat to black-box applications (when attackers have no access to the target models). Current transfer-based ensemble attacks, however, only consider a limited number of source models to craft an adversarial example and, thus, obtain poor transferability. Besides, recent query-based black-box attacks, which require numerous queries to the target model, not only come under suspicion by the target model but also cause expensive query cost. In this article, we propose a novel transfer-based black-box attack, dubbed serial-minigroup-ensemble-attack (SMGEA). Concretely, SMGEA first divides a large number of pretrained white-box source models into several "minigroups." For each minigroup, we design three new ensemble strategies to improve the intragroup transferability. Moreover, we propose a new algorithm that recursively accumulates the "long-term" gradient memories of the previous minigroup to the subsequent minigroup. This way, the learned adversarial information can be preserved, and the intergroup transferability can be improved. Experiments indicate that SMGEA not only achieves state-of-the-art black-box attack ability over several data sets but also deceives two online black-box saliency prediction systems in real world, i.e., DeepGaze-II (https://deepgaze.bethgelab.org/) and SALICON (http://salicon.net/demo/). Finally, we contribute a new code repository to promote research on adversarial attack and defense over ubiquitous pixel-to-pixel computer vision tasks. We share our code together with the pretrained substitute model zoo at https://github.com/CZHQuality/AAA-Pix2pix. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Xiongkuo Min, Guodong Guo, Patrick Le Callet |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Decoupled IoU Regression for Object DetectionabstractNon-maximum suppression (NMS) is widely used in object detection pipelines for removing duplicated bounding boxes. The inconsistency between the confidence for NMS and the real localization confidence seriously affects detection performance. Prior works propose to predict Intersection-over-Union (IoU) between bounding boxes and corresponding ground-truths to improve NMS, while accurately predicting IoU is still a challenging problem. We argue that the complex definition of IoU and feature misalignment make it difficult to predict IoU accurately. In this paper, we propose a novel Decoupled IoU Regression (DIR) model to handle these problems. The proposed DIR decouples the traditional localization confidence metric IoU into two new metrics, Purity and Integrity. Purity reflects the proportion of the object area in the detected bounding box, and Integrity refers to the completeness of the detected object area. Separately predicting Purity and Integrity can divide the complex mapping between the bounding box and its IoU into two clearer mappings and model them independently. In addition, a simple but effective feature realignment approach is also introduced to make the IoU regressor work in a hindsight manner, which can make the target mapping more stable. The proposed DIR can be conveniently integrated with existing two-stage detectors and significantly improve their performance. Through a simple implementation of DIR with HTC, we obtain 51.3% AP on MS COCO benchmark, which outperforms previous methods and achieves state-of-the-art. Yan Gao 0017, Qimeng Wang, Xu Tang 0007, Jing Li 0026, Yao Hu 0002 |
ACM Multimedia | 6 |
| 2021 | Perceptual Quality Assessment of Internet VideosabstractWith the fast proliferation of online video sites and social media platforms, user, professionally and occupationally generated content (UGC, PGC, OGC) videos are streamed and explosively shared over the Internet. Consequently, it is urgent to monitor the content quality of these Internet videos to guarantee the user experience. However, most existing modern video quality assessment (VQA) databases only include UGC videos and cannot meet the demands for other kinds of Internet videos with real-world distortions. To this end, we collect 1,072 videos from Youku, a leading Chinese video hosting service platform, to establish the Internet video quality assessment database (Youku-V1K). A special sampling method based on several quality indicators is adopted to maximize the content and distortion diversities within a limited database, and a probabilistic graphical model is applied to recover reliable labels from noisy crowdsourcing annotations. Based on the properties of Internet videos originated from Youku, we propose a spatio-temporal distortion-aware model (STDAM). First, the model works blindly which means the pristine video is unnecessary. Second, the model is familiar with diverse contents by pre-training on the large-scale image quality assessment databases. Third, to measure spatial and temporal distortions, we introduce the graph convolution and attention module to extract and enhance the features of the input video. Besides, we leverage the motion information and integrate the frame-level features into video-level features via a bi-directional long short-term memory network. Experimental results on the self-built database and the public VQA databases demonstrate that our model outperforms the state-of-the-art methods and exhibits promising generalization ability. Jiahua Xu 0001, Jing Li 0026, Xingguang Zhou, Wei Zhou 0021, Baichao Wang, Zhibo Chen 0001 |
ACM Multimedia | 2 |
| 2021 | Adversarial Attack Against Deep Saliency Models Powered by Non-Redundant PriorsabstractSaliency detection is an effective front-end process to many security-related tasks, e.g. automatic drive and tracking. Adversarial attack serves as an efficient surrogate to evaluate the robustness of deep saliency models before they are deployed in real world. However, most of current adversarial attacks exploit the gradients spanning the entire image space to craft adversarial examples, ignoring the fact that natural images are high-dimensional and spatially over-redundant, thus causing expensive attack cost and poor perceptibility. To circumvent these issues, this paper builds an efficient bridge between the accessible partially-white-box source models and the unknown black-box target models. The proposed method includes two steps: 1) We design a new partially-white-box attack, which defines the cost function in the compact hidden space to punish a fraction of feature activations corresponding to the salient regions, instead of punishing every pixel spanning the entire dense output space. This partially-white-box attack reduces the redundancy of the adversarial perturbation. 2) We exploit the non-redundant perturbations from some source models as the prior cues, and use an iterative zeroth-order optimizer to compute the directional derivatives along the non-redundant prior directions, in order to estimate the actual gradient of the black-box target model. The non-redundant priors boost the update of some "critical" pixels locating at non-zero coordinates of the prior cues, while keeping other redundant pixels locating at the zero coordinates unaffected. Our method achieves the best tradeoff between attack ability and perturbation redundancy. Finally, we conduct a comprehensive experiment to test the robustness of 18 state-of-the-art deep saliency models against 16 malicious attacks, under both of white-box and black-box settings, which contributes a new robustness benchmark to the saliency community for the first time. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Yuan Tian 0017, Guodong Guo, Patrick Le Callet |
IEEE Trans. Image Process. | 5 |
| 2021 | Quality Assessment of Free-Viewpoint Videos by Quantifying the Elastic Changes of Multi-Scale Motion TrajectoriesabstractVirtual viewpoints synthesis is an essential process for many immersive applications including Free-viewpoint TV (FTV). A widely used technique for viewpoints synthesis is Depth-Image-Based-Rendering (DIBR) technique. However, such technique may introduce challenging non-uniform spatial-temporal structure-related distortions. Most of the existing state-of-the-art quality metrics fail to handle these distortions, especially the temporal structure inconsistencies observed during the switch of different viewpoints. To tackle this problem, an elastic metric and multi-scale trajectory based video quality metric (EM-VQM) is proposed in this paper. Dense motion trajectory is first used as a proxy for selecting temporal sensitive regions, where local geometric distortions might significantly diminish the perceived quality. Afterwards, the amount of temporal structure inconsistencies and unsmooth viewpoints transitions are quantified by calculating 1) the amount of motion trajectory deformations with elastic metric and, 2) the spatial-temporal structural dissimilarity. According to the comprehensive experimental results on two FTV video datasets, the proposed metric outperforms the state-of-the-art metrics designed for free-viewpoint videos significantly and achieves a gain of 12.86% and 16.75% in terms of median Pearson linear correlation coefficient values on the two datasets compared to the best one, respectively. Suiyi Ling, Jing Li 0026, Zhaohui Che, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Image Process. | 2 |
| 2021 | Re-Visiting Discriminator for Blind Free-Viewpoint Image Quality AssessmentabstractAccurate measurement of perceptual quality is important for various immersive multimedia, which demand real-time quality control or quality-based bench-marking for relevant algorithms. For instance, virtual views rendering in Free-Viewpoint (FV) navigation scenarios is a typical case that introduces challenging distortions, particularly the ones around dis-occluded regions. Existing quality metrics, most of which are targeting for impairments caused by compression or network condition, fail to quantify such non-uniform structure-related distortions. Moreover, the lack of quality databases for such distortions makes it even more challenging to develop robust quality metrics. In this work, a Generative Adversarial Networks based No-Reference (NR) quality Metric, namely GANs-NRM, is proposed. We first present an approach to create masks mimicking dis-occlusions/textureless regions, which is applicable on large-scale 2D image databases publicly available in the computer vision domain. Using these synthetic data, we then train a GANs-based context renderer with the capability of rendering those masked regions. Since the naturalness of the rendered dis-occluded regions strongly relates to the perceptual quality, we assume that the discriminator of the trained GANs has an intrinsic ability for quality assessment. We thus use the features extracted from the discriminator to learn a Bag-of-Distortion-Word (BDW) codebook. We show that a quality predictor can be then well trained using only a small amount of subjective quality data for the FV views rendering. Moreover, in the proposed framework, the discriminator is also adapted as a distortion-detector to locate possible distorted regions. According to the experimental results, the proposed model outperforms significantly the state-of-the-art quality metrics. The corresponding context renderer also shows appealing visualized results over other rendering algorithms. Suiyi Ling, Jing Li 0026, Zhaohui Che, Wei Zhou 0021, Junle Wang, Patrick Le Callet |
IEEE Trans. Multim. | 2 |
| 2021 | Image Quality Assessment Using Kernel Sparse CodingabstractOne key in image quality assessment (IQA) is the design of image representations that can capture the changes of image structures caused by distortions. Recent studies show that sparse coding has emerged as a promising approach to analyzing image structures for IQA. However, existing sparse-coding-based IQA approaches use linear coding models, which ignore the nonlinearities of manifolds of image patches and thus cannot analyze complex image structures well. To overcome such a weakness, in this paper, we introduce nonlinear sparse coding to IQA. A kernel dictionary construction scheme is proposed, which combines analytic dictionaries and learnable dictionaries to guarantee both the stability and effectiveness of kernel sparse coding in the context of IQA. Built upon the kernel dictionary construction, an effective full-reference IQA metric is developed. Benefiting from the considerations on nonlinearities during sparse coding, the proposed IQA metric not only characterizes image distortions better, but also achieves improvement on the consistency with subjective perception, when compared to the metrics built upon linear sparse coding. Such benefits are demonstrated with the experimental results on eight benchmark datasets in terms of common criteria. Zihan Zhou 0007, Jing Li 0026, Yuhui Quan, Ruotao Xu |
IEEE Trans. Multim. | 2 |
| 2020 | A New Ensemble Adversarial Attack Powered by Long-Term Gradient MemoriesabstractDeep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of pre-trained source models can transfer to other new target models, thus pose a security threat to black-box applications (when the attackers have no access to the target models). Despite adopting diverse architectures and parameters, source and target models often share similar decision boundaries. Therefore, if an adversary is capable of fooling several source models concurrently, it can potentially capture intrinsic transferable adversarial information that may allow it to fool a broad class of other black-box target models. Current ensemble attacks, however, only consider a limited number of source models to craft an adversary, and obtain poor transferability. In this paper, we propose a novel black-box attack, dubbed Serial-Mini-Batch-Ensemble-Attack (SMBEA). SMBEA divides a large number of pre-trained source models into several mini-batches. For each single batch, we design 3 new ensemble strategies to improve the intra-batch transferability. Besides, we propose a new algorithm that recursively accumulates the “long-term” gradient memories of the previous batch to the following batch. This way, the learned adversarial information can be preserved and the inter-batch transferability can be improved. Experiments indicate that our method outperforms state-of-the-art ensemble attacks over multiple pixel-to-pixel vision tasks including image translation and salient region prediction. Our method successfully fools two online black-box saliency prediction systems including DeepGaze-II (Kummerer 2017) and SALICON (Huang et al. 2017). Finally, we also contribute a new repository to promote the research on adversarial attack and defense over pixel-to-pixel tasks: https://github.com/CZHQuality/AAA-Pix2pix. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Patrick Le Callet |
AAAI | 5 |
| 2020 | Few-Shot Pill RecognitionabstractPill image recognition is vital for many personal/public health-care applications and should be robust to diverse unconstrained real-world conditions. Most existing pill recognition models are limited in tackling this challenging few-shot learning problem due to the insufficient instances per category. With limited training data, neural network-based models have limitations in discovering most discriminating features, or going deeper. Especially, existing models fail to handle the hard samples taken under less controlled imaging conditions. In this study, a new pill image database, namely CURE, is first developed with more varied imaging conditions and instances for each pill category. Secondly, a W2-net is proposed for better pill segmentation. Thirdly, a Multi-Stream (MS) deep network that captures task-related features along with a novel two-stage training methodology are proposed. Within the proposed framework, a Batch All strategy that considers all the samples is first employed for the sub-streams, and then a Batch Hard strategy that considers only the hard samples mined in the first stage is utilized for the fusion network. By doing so, complex samples that could not be represented by one type of feature could be focused and the model could be forced to exploit other domain-related information more effectively. Experiment results show that the proposed model outperforms state-of-the-art models on both the National Institute of Health (NIH) and our CURE database. Suiyi Ling, Andreas Pastor, Jing Li 0026, Zhaohui Che, Junle Wang, Patrick Le Callet |
CVPR | 3 |
| 2020 | A Probabilistic Graphical Model for Analyzing the Subjective Visual Quality Assessment Data from CrowdsourcingabstractThe swift development of the multimedia technology has raised dramatically the users' expectation on the quality of experience. To obtain the ground-truth perceptual quality for model training, subjective assessment is necessary. Crowdsourcing platform provides us a convenient and feasible way to run large-scale experiments. However, the obtained perceptual quality labels are generally noisy. In this paper, we propose a probabilistic graphical annotation model to infer the underlying ground truth and discovering the annotator's behavior. In the proposed model, the ground truth quality label is considered following a categorical distribution rather than a unique number, i.e., different reliable opinions on the perceptual quality are allowed. In addition, different annotator's behaviors in crowdsourcing are modeled, which allows us to identify the possibility that the annotator makes noisy labels during the test. The proposed model has been tested on both simulated data and real-world data, where it always shows superior performance than the other state-of-the-art models in terms of accuracy and robustness. Jing Li 0026, Suiyi Ling, Junle Wang, Patrick Le Callet |
ACM Multimedia | 1 |
| 2020 | QoEVMA'20: 1st Workshop on Quality of Experience (QoE) in Visual Multimedia ApplicationsabstractNowadays, people spend dramatically more time on watching videos through different devices. The advanced hardware technology and network allow for the increasing demands of users viewing experience. Thus, enhancing the Quality of Experience of end-users in advanced multimedia is the ultimate goal of service providers, as good services would attract more consumers. Quality assessment is thus important. The first workshop on "Quality of Experience (QoE) in visual multimedia applications" (QoEVMA'20) focuses on the QoE assessment of any visual multimedia applications both subjectively and objectively. The topics include 1)QoE assessment on different visual multimedia applications, including VoD for movies, dramas, variety shows, UGC on social networks, live streaming videos for gaming/shopping/social, etc. 2)QoE assessment for different video formats in multimedia services, including 2D, stereoscopic 3D, High Dynamic Range (HDR), Augmented Reality (AR), Virtual Reality (VR), 360, Free-Viewpoint Video(FVV), etc. 3)Key performance indicators (KPI) analysis for QoE. This summary gives a brief overview of the workshop, which took place at October 16, 2020 in Seattle (U.S.), as a half-day workshop. Xinbo Gao 0001, Patrick Le Callet, Jing Li 0026, Zhi Li 0001, Wen Lu 0004 |
ACM Multimedia | 3 |
| 2020 | Full-reference image quality metric for blurry images and compressed images using hybrid dictionary learning
Zihan Zhou 0007, Jing Li 0026, Yong Xu 0007, Yuhui Quan |
Neural Comput. Appl. | 2 |
| 2019 | Perceptual Representations of Structural Information in Images: Application to Quality Assessment of Synthesized View in FTV ScenarioabstractAs the immersive multimedia techniques like Free-viewpoint TV (FTV) develop at an astonishing rate, user's demand for high-quality immersive contents increases dramatically. Unlike traditional uniform artifacts, the distortions within immersive contents could be non-uniform structure-related and thus are challenging for commonly used quality metrics. Recent studies have demonstrated that the representation of visual features can be extracted from multiple levels of the hierarchy. Inspired by the hierarchical representation mechanism in the human visual system (HVS), in this paper, we explore to adopt structural representations to quantitatively measure the impact of such structure-related distortion on perceived quality in FTV scenario. More specifically, a bio-inspired full reference image quality metric is proposed based on 1) low-level contour descriptor; 2) mid-level contour category descriptor; and 3) task-oriented non-natural structure descriptor. The experimental results show that the proposed model outperforms significantly the state-of-the-art metrics. Suiyi Ling, Jing Li 0026, Patrick Le Callet, Junle Wang |
ICIP | 2 |
| 2019 | AccAnn: A New Subjective Assessment Methodology for Measuring Acceptability and Annoyance of Quality of ExperienceabstractUser expectations have a crucial impact on the levels of quality of experience (QoE) that they consider acceptable or satisfying. Measuring acceptability and annoyance has mainly been performed in separate or multi-step experiments without any control over participants' expectations. This paper introduces a simple methodology to obtain the information about both of the entities in a single step and compares several data processing strategies useful for results interpretation. A specifically designed subjective experiment, conducted on compressed videos, has shown that the multi-step procedures could be replaced by our proposed single-step approach, regardless of the viewing conditions, while the novel approach is significantly preferred by observers for its low time requirements and higher intuitiveness. The test has simultaneously proven that user expectations can be altered by the instructions and it is, therefore, possible to simulate different user profiles regardless of the participants' real habits. The acceptability/annoyance experimental results are also used to benchmark the state-of-the-art objective video quality metrics in predicting acceptability/annoyance of QoE. A case study on the determination of the threshold of acceptability/annoyance for objective quality metrics is conducted, which can be served as a guideline for video streaming service providers. Jing Li 0026, Lukas Krasula, Yoann Baveye, Zhi Li 0001, Patrick Le Callet |
IEEE Trans. Multim. | 1 |
| 2018 | Hybrid-MST: A Hybrid Active Sampling Strategy for Pairwise Preference AggregationabstractIn this paper we present a hybrid active sampling strategy for pairwise preference aggregation, which aims at recovering the underlying rating of the test candidates from sparse and noisy pairwise labeling. Our method employs Bayesian optimization framework and Bradley-Terry model to construct the utility function, then to obtain the Expected Information Gain (EIG) of each pair. For computational efficiency, Gaussian-Hermite quadrature is used for estimation of EIG. In this work, a hybrid active sampling strategy is proposed, either using Global Maximum (GM) EIG sampling or Minimum Spanning Tree (MST) sampling in each trial, which is determined by the test budget. The proposed method has been validated on both simulated and real-world datasets, where it shows higher preference aggregation ability than the state-of-the-art methods. Jing Li 0026, Rafal Mantiuk, Junle Wang, Suiyi Ling, Patrick Le Callet |
NeurIPS | 1 |
| 2018 | Quantifying the Influence of Devices on Quality of Experience for Video StreamingabstractThe Internet streaming is changing the way of watching videos for people. Traditional quality assessment on the cable/satellite broadcasting system mainly focused on the perceptual quality. Nowadays, this concept has been extended to Quality of Experience (QoE) which considers also the contextual factors, such as the environment, the display devices, etc. In this study, we focus on the influence of devices on QoE. A subjective experiment was conducted by using our proposed AccAnn methodology. The observers evaluated the QoE of the video sequences by considering their Acceptance and Annoyance. Two devices were used in this study, TV and Tablet. The experimental results showed that the device was a significant influence factor on QoE. In addition, we found that this influence varied with the QoE of the video sequences. To quantify this influence, the Eliminated-By-Aspects model was used. The results could be used for the training of a device-neutral objective QoE metric. For video streaming providers, the quantification results of the influence from devices could be used to optimize the selection of streaming content. On one hand it could satisfy the QoE expectations of the observers according to the used devices, on the other hand it could help to save the bitrates. Jing Li 0026, Lukas Krasula, Patrick Le Callet, Zhi Li 0001, Yoann Baveye |
PCS | 1 |
| 2018 | Improving the discriminability of standard subjective quality assessment methods: a case studyabstractSubjective assessment for image or video qualities is considered as the most reliable way to obtain the ground truth for the development of objective quality metrics, especially when leaded by Mean Opinion Score (MOS approaches). However, obtained MOS with standard protocols are noisy due to subject's personal characteristics, such as viewing experience, gender or profession, leading to uncertain ground truth driven by the number of panelists/subjects. The usual way to reduce uncertainty relies on raising this number. In this paper, we demonstrate how a recently introduced Maximum Likelihood Estimation (MLE) based quality recovery model can improve the discriminability of standard subjective quality assessment. Compared to straightforward MOS computation, we present a case study where one can save between 26% to 39% in terms of numbers of subjects at the same discriminability. Jing Li 0026, Patrick Le Callet |
QoMEX | 1 |
| 2017 | Visual Attention Modeling for Stereoscopic Video: A Benchmark and Computational ModelabstractIn this paper, we investigate the visual attention modeling for stereoscopic video from the following two aspects. First, we build one large-scale eye tracking database as the benchmark of visual attention modeling for stereoscopic video. The database includes 47 video sequences and their corresponding eye fixation data. Second, we propose a novel computational model of visual attention for stereoscopic video based on Gestalt theory. In the proposed model, we extract the low-level features, including luminance, color, texture, and depth, from discrete cosine transform coefficients, which are used to calculate feature contrast for the spatial saliency computation. The temporal saliency is calculated by the motion contrast from the planar and depth motion features in the stereoscopic video sequences. The final saliency is estimated by fusing the spatial and temporal saliency with uncertainty weighting, which is estimated by the laws of proximity, continuity, and common fate in Gestalt theory. Experimental results show that the proposed method outperforms the state-of-the-art stereoscopic video saliency detection models on our built large-scale eye tracking database and one other database (DML-ITRACK-3D). Yuming Fang 0001, Chi Zhang 0027, Jing Li 0026, Jianjun Lei 0001, Matthieu Perreira Da Silva, Patrick Le Callet |
IEEE Trans. Image Process. | 3 |
| 2012 | Analysis and improvement of a paired comparison method in the application of 3DTV subjective experimentabstractPaired comparison is a frequently used method in psychophysical studies. However, with the increase of the number of the stimuli, the number of comparisons increases exponentially. Square design is one of the balanced sub-set paired comparison methods which could reduce the number of comparisons while producing comparably precise results under some assumptions. However, when there are observation errors from observers' attentiveness, the square design would produce large estimation errors. Thus, an improved square design which is robust to observation errors is proposed. Using a Monte Carlo simulation, the proposed method is evaluated and shows improvement in efficiency. The original design is applied in a visual discomfort subjective test of 3DTV. In addition, both of the two designs are studied by utilizing our previous full comparison data. The test results showed that the proposed improved square design is more robust to observation errors. Another important finding is that the influence of the occurrence of some other stimuli on voting is significant. Whether the proposed method could reduce the prediction errors induced by it is still under study. Jing Li 0026, Marcus Barkowsky, Patrick Le Callet |
ICIP | 1 |
| 2010 | A new quality metric for compressed images based on DDCTabstractAs the performance-indicator of the image processing algorithms or systems, image quality assessment (IQA) has attracted the attention of many researchers. Aiming to the widely used compression standards, JPEG and JPEG2000, we propose a new no reference (NR) metric for compressed images to do IQA. This metric exploits the causes of distortion by JPEG and JPEG2000, employs the directional discrete cosine transform (DDCT) to obtain the detail and direction information of the images and incorporates with the visual perception to obtain the image quality index. Experimental results show that the proposed metric not only has outstanding performance on JPEG and JPEG2000 images, but also applicable to other types of artifacts. Wen Lu 0004, Jing Li 0026, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001 |
VCIP | 2 |