EDBT 2026 Demo / reviewers in the wild / expert
Suiyi Ling
dblp:193/4275
· DBLP profile ↗
26ranked-venue papers
11as first author
12since 2021 · last 2026
0000-0002-5306-9189ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 11 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Assessing the impact of central and peripheral obstructions on visual behavior: Insights from gaze-contingent eye-tracking studiesabstractVisual field loss, caused by conditions like glaucoma or macular degeneration, affects many people and impacts several life domains. This study contributes to the exploration of how people with visual field loss, such as central and peripheral scotomas, process visual stimuli in digital environments. Current visual attention models are based on experimental data obtained from individuals with normal vision, often overlooking those with limited vision. To address this issue, we compare gaze data from subjects viewing stimuli under two conditions: with and without visual-field masks of varying radii, with the main goal of understanding the role played by different components of vision in the overall grasping of visual information. We use metrics commonly employed for benchmarking saliency models as a means of assessing the similarity between data from each obstruction mask and the control stimuli, which could lead to the conclusion of whether foveal and peripheral vision contribute equally to natural vision or whether one of them stands out in information extraction. Novel saliency models could use this information to predict attention from visually-impaired individuals by possibly balancing these two sources of vision. Our results show a significantly higher similarity between control and central-scotoma saliency maps than between control and peripheral-scotoma data. Another statistical analysis shows no substantial learning effect or familiarity bias when participants revisit the same image under different conditions in the eye-tracking experiment. Finally, a difference-significance study reveals that different radii from central-scotoma conditions demonstrated no meaningful dispersion from each other. Claudio M. S. Coutinho, Maria C. O. Faria, Alexandre Bruckert, Suiyi Ling, Matthieu Perreira Da Silva, Ronaldo F. Zampolo, Patrick Le Callet |
Signal Process. Image Commun. | 4 |
| 2023 | Estimating Uncertainty On Video Quality MetricsabstractVideo Quality Metrics (VQM) are models used to predict the score that a user would give to the quality of a video visualization. They are widely used in video processing systems, for monitoring end-to-end quality or system troubleshooting for example. In these scenarios, the improvement is quantified based on a certain enhancement of a VQM score and trouble-detection is done based on a certain drop or threshold computed based on a VQM. Yet, whether such improvement or fault-detection is worth a significant increase in power consumption is questionable. Therefore, the goal of this work is to propose a method to predict the uncertainty of the quality metric. In this paper, we propose a framework to evaluate the confidence interval of a VQM for a given content using simple video features. We assess the performance of the framework by using the confidence intervals to predict if two videos are of similar or different quality and show that in most cases our approach performs better than just using a constant confidence interval. Patrick Le Callet, Suiyi Ling, Haixiong Wang, Ioannis Katsavounidis, Zafar Shahid, Cosmin Stejerean |
ICASSP | 3 |
| 2022 | Considering User Agreement in Learning to Predict the Aesthetic QualityabstractHow to robustly rank the aesthetic quality of given images has been a long-standing ill-posed topic. Such challenge stems mainly from the diverse subjective opinions of different observers about the varied types of content. There is a growing interest in estimating the user agreement by considering the standard deviation (σ) of the scores, instead of only predicting the mean aesthetic opinion score (µ). Nevertheless, when comparing a pair of contents, few studies consider how confident are we regarding the difference in the aesthetic scores. In this paper, we thus propose (1) a re-adapted multi-task attention network to predict both the mean opinion score and the standard deviation in an end-to-end manner; (2) a brand-new confidence interval ranking loss that encourages the model to focus on image-pairs that are less certain about the difference of their aesthetic scores. With such loss, the model is encouraged to learn the uncertainty of the content that is relevant to the diversity of observers’ opinions, i.e., user disagreement. Extensive experiments have demonstrated that the proposed multi-task aesthetic model achieves state-of-the-art performance on two different types of aesthetic datasets, i.e., AVA and TMGA. Suiyi Ling, Andreas Pastor, Junle Wang, Patrick Le Callet |
ICASSP | 1 |
| 2022 | Subjective And Objective Quality Assessment Of Mobile Gaming VideoabstractNowadays, with the vigorous expansion and development of gaming video streaming techniques and services, the expectation of users, especially the mobile phone users, for higher quality of experience is also growing swiftly. As most of the existing research focuses on traditional video streaming, there is a clear lack of both subjective study and objective quality models that are tailored for quality assessment of mobile gaming content. To this end, in this study, we first present a brand new Tencent Gaming Video dataset containing 1293 mobile gaming sequences encoded with three different codecs. Second, we propose an objective quality framework, namely Efficient hard-RAnk Quality Estimator (ERAQUE), that is equipped with (1) a novel hard pairwise ranking loss, which forces the model to put more emphasis on differentiating similar pairs; (2) an adapted model distillation strategy, which could be utilized to compress the proposed model efficiently without causing significant performance drop. Extensive experiments demonstrate the efficiency and robustness of our model. Shaoguo Wen, Suiyi Ling, Junle Wang, Yanqing Jing, Patrick Le Callet |
ICASSP | 2 |
| 2022 | BR-NPA: A non-parametric high-resolution attention model to improve the interpretability of attention
Tristan Gomez, Suiyi Ling, Thomas Fréour, Harold Mouchère |
Pattern Recognit. | 2 |
| 2022 | SMGEA: A New Ensemble Adversarial Attack Powered by Long-Term Gradient MemoriesabstractDeep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of source models transfer to other target models and, thus, pose a security threat to black-box applications (when attackers have no access to the target models). Current transfer-based ensemble attacks, however, only consider a limited number of source models to craft an adversarial example and, thus, obtain poor transferability. Besides, recent query-based black-box attacks, which require numerous queries to the target model, not only come under suspicion by the target model but also cause expensive query cost. In this article, we propose a novel transfer-based black-box attack, dubbed serial-minigroup-ensemble-attack (SMGEA). Concretely, SMGEA first divides a large number of pretrained white-box source models into several "minigroups." For each minigroup, we design three new ensemble strategies to improve the intragroup transferability. Moreover, we propose a new algorithm that recursively accumulates the "long-term" gradient memories of the previous minigroup to the subsequent minigroup. This way, the learned adversarial information can be preserved, and the intergroup transferability can be improved. Experiments indicate that SMGEA not only achieves state-of-the-art black-box attack ability over several data sets but also deceives two online black-box saliency prediction systems in real world, i.e., DeepGaze-II (https://deepgaze.bethgelab.org/) and SALICON (http://salicon.net/demo/). Finally, we contribute a new code repository to promote research on adversarial attack and defense over ubiquitous pixel-to-pixel computer vision tasks. We share our code together with the pretrained substitute model zoo at https://github.com/CZHQuality/AAA-Pix2pix. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Xiongkuo Min, Guodong Guo, Patrick Le Callet |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Seeing By Haptic Glance: Reinforcement Learning Based 3d Object RecognitionabstractHuman is able to conduct 3D recognition by a limited number of haptic contacts between the target object and his/her fingers without seeing the object. This capability is defined as ‘haptic glance’ in cognitive neuroscience. Most of the existing 3D recognition models were developed based on dense 3D data. Nonetheless, in many real-life use cases, where robots are used to collect 3D data by haptic exploration, only a limited number of 3D points could be collected. In this study, we thus focus on solving the intractable problem of how to obtain cognitively representative 3D key-points of a target object with limited interactions between the robot and the object. A novel reinforcement learning based framework is proposed, where the haptic exploration procedure (the agent iteratively predicts the next position for the robot to explore) is optimized simultaneously with the objective 3D recognition with actively collected 3D points. As the model is rewarded only when the 3D object is accurately recognized, it is driven to find the sparse yet efficient haptic-perceptua13D representation of the object. Experimental results show that our proposed model outperforms the state of the art models. Kévin Riou, Suiyi Ling, Guillaume Gallot, Patrick Le Callet |
ICIP | 2 |
| 2021 | Multi-Modal Aesthetic Assessment for Mobile Gaming ImageabstractWith the proliferation of various gaming technology, services, game styles, and platforms, multi-dimensional aesthetic assessment of the gaming contents is becoming more and more important for the gaming industry. Depending on the diverse needs of diversified game players, game designers, graphical developers, etc. in particular conditions, multi-modal aesthetic assessment is required to consider different aesthetic dimensions/perspectives. Since there are different underlying relationships between different aesthetic dimensions, e.g., between the ‘Colorfulness’ and ‘Color Harmony’, it could be advantageous to leverage effective information attached in multiple relevant dimensions. To this end, we solve this problem via multi-task learning. Our inclination is to seek and learn the correlations between different aesthetic relevant dimensions to further boost the generalization performance in predicting all the aesthetic dimensions. Therefore, the ‘bottleneck’ of obtaining good predictions with limited labeled data for one individual dimension could be unplugged by harnessing complementary sources of other dimensions, i.e., augment the training data indirectly by sharing training information across dimensions. According to experimental results, the proposed model outperforms state-of-the-art aesthetic metrics significantly in predicting four gaming aesthetic dimensions. Yejing Xie, Suiyi Ling, Andreas Pastor, Junle Wang, Junyu Dong, Patrick Le Callet |
MMSP | 3 |
| 2021 | The Effect of Temporal Sub-sampling on the Accuracy of Volumetric Video Quality AssessmentabstractVolumetric video content has attracted increasing research interests over the last decade, as it facilitates the integration of dynamic real world content in virtual environments. Point cloud is one of the most common alternatives to represent volumetric video content. Yet, such representation requires an enormous data storage and pose significant greater pressures on compression algorithms compared to the standard 2D video. This challenge has unleashed a new wave in the development of novel point cloud compression technologies, which need to be evaluated in terms of production quality. Due to the high dimensionality of the data, evaluating the performances of relevant coding algorithms can be time consuming. This puts a barrier on optimizing coding algorithms with complex, but perceptually accurate, objective quality metrics. In this study, we thus explore the possibility of reducing temporal-dimension of the content under-evaluation, i.e., temporal sub-sampling, for objective quality evaluation without sacrificing from the correlation with the subjective opinion. In addition, we exploit different temporal pooling methods to further make the quality evaluation procedure more efficient. In total 30 different objective quality metrics were tested on the the V-SENSE volumetric video quality database. According to experimental results, there is no need to employ full frame-rate (30 fps) assessment to reach the meaningful correlation for the considered quality metrics. These observations could be referred to reduce the computation complexity regarding the evaluation and optimization of the relevant compression algorithms. Ali Ak, Emin Zerman, Suiyi Ling, Patrick Le Callet, Aljoscha Smolic |
PCS | 3 |
| 2021 | Adversarial Attack Against Deep Saliency Models Powered by Non-Redundant PriorsabstractSaliency detection is an effective front-end process to many security-related tasks, e.g. automatic drive and tracking. Adversarial attack serves as an efficient surrogate to evaluate the robustness of deep saliency models before they are deployed in real world. However, most of current adversarial attacks exploit the gradients spanning the entire image space to craft adversarial examples, ignoring the fact that natural images are high-dimensional and spatially over-redundant, thus causing expensive attack cost and poor perceptibility. To circumvent these issues, this paper builds an efficient bridge between the accessible partially-white-box source models and the unknown black-box target models. The proposed method includes two steps: 1) We design a new partially-white-box attack, which defines the cost function in the compact hidden space to punish a fraction of feature activations corresponding to the salient regions, instead of punishing every pixel spanning the entire dense output space. This partially-white-box attack reduces the redundancy of the adversarial perturbation. 2) We exploit the non-redundant perturbations from some source models as the prior cues, and use an iterative zeroth-order optimizer to compute the directional derivatives along the non-redundant prior directions, in order to estimate the actual gradient of the black-box target model. The non-redundant priors boost the update of some "critical" pixels locating at non-zero coordinates of the prior cues, while keeping other redundant pixels locating at the zero coordinates unaffected. Our method achieves the best tradeoff between attack ability and perturbation redundancy. Finally, we conduct a comprehensive experiment to test the robustness of 18 state-of-the-art deep saliency models against 16 malicious attacks, under both of white-box and black-box settings, which contributes a new robustness benchmark to the saliency community for the first time. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Yuan Tian 0017, Guodong Guo, Patrick Le Callet |
IEEE Trans. Image Process. | 4 |
| 2021 | Quality Assessment of Free-Viewpoint Videos by Quantifying the Elastic Changes of Multi-Scale Motion TrajectoriesabstractVirtual viewpoints synthesis is an essential process for many immersive applications including Free-viewpoint TV (FTV). A widely used technique for viewpoints synthesis is Depth-Image-Based-Rendering (DIBR) technique. However, such technique may introduce challenging non-uniform spatial-temporal structure-related distortions. Most of the existing state-of-the-art quality metrics fail to handle these distortions, especially the temporal structure inconsistencies observed during the switch of different viewpoints. To tackle this problem, an elastic metric and multi-scale trajectory based video quality metric (EM-VQM) is proposed in this paper. Dense motion trajectory is first used as a proxy for selecting temporal sensitive regions, where local geometric distortions might significantly diminish the perceived quality. Afterwards, the amount of temporal structure inconsistencies and unsmooth viewpoints transitions are quantified by calculating 1) the amount of motion trajectory deformations with elastic metric and, 2) the spatial-temporal structural dissimilarity. According to the comprehensive experimental results on two FTV video datasets, the proposed metric outperforms the state-of-the-art metrics designed for free-viewpoint videos significantly and achieves a gain of 12.86% and 16.75% in terms of median Pearson linear correlation coefficient values on the two datasets compared to the best one, respectively. Suiyi Ling, Jing Li 0026, Zhaohui Che, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Image Process. | 1 |
| 2021 | Re-Visiting Discriminator for Blind Free-Viewpoint Image Quality AssessmentabstractAccurate measurement of perceptual quality is important for various immersive multimedia, which demand real-time quality control or quality-based bench-marking for relevant algorithms. For instance, virtual views rendering in Free-Viewpoint (FV) navigation scenarios is a typical case that introduces challenging distortions, particularly the ones around dis-occluded regions. Existing quality metrics, most of which are targeting for impairments caused by compression or network condition, fail to quantify such non-uniform structure-related distortions. Moreover, the lack of quality databases for such distortions makes it even more challenging to develop robust quality metrics. In this work, a Generative Adversarial Networks based No-Reference (NR) quality Metric, namely GANs-NRM, is proposed. We first present an approach to create masks mimicking dis-occlusions/textureless regions, which is applicable on large-scale 2D image databases publicly available in the computer vision domain. Using these synthetic data, we then train a GANs-based context renderer with the capability of rendering those masked regions. Since the naturalness of the rendered dis-occluded regions strongly relates to the perceptual quality, we assume that the discriminator of the trained GANs has an intrinsic ability for quality assessment. We thus use the features extracted from the discriminator to learn a Bag-of-Distortion-Word (BDW) codebook. We show that a quality predictor can be then well trained using only a small amount of subjective quality data for the FV views rendering. Moreover, in the proposed framework, the discriminator is also adapted as a distortion-detector to locate possible distorted regions. According to the experimental results, the proposed model outperforms significantly the state-of-the-art quality metrics. The corresponding context renderer also shows appealing visualized results over other rendering algorithms. Suiyi Ling, Jing Li 0026, Zhaohui Che, Wei Zhou 0021, Junle Wang, Patrick Le Callet |
IEEE Trans. Multim. | 1 |
| 2020 | A New Ensemble Adversarial Attack Powered by Long-Term Gradient MemoriesabstractDeep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of pre-trained source models can transfer to other new target models, thus pose a security threat to black-box applications (when the attackers have no access to the target models). Despite adopting diverse architectures and parameters, source and target models often share similar decision boundaries. Therefore, if an adversary is capable of fooling several source models concurrently, it can potentially capture intrinsic transferable adversarial information that may allow it to fool a broad class of other black-box target models. Current ensemble attacks, however, only consider a limited number of source models to craft an adversary, and obtain poor transferability. In this paper, we propose a novel black-box attack, dubbed Serial-Mini-Batch-Ensemble-Attack (SMBEA). SMBEA divides a large number of pre-trained source models into several mini-batches. For each single batch, we design 3 new ensemble strategies to improve the intra-batch transferability. Besides, we propose a new algorithm that recursively accumulates the “long-term” gradient memories of the previous batch to the following batch. This way, the learned adversarial information can be preserved and the inter-batch transferability can be improved. Experiments indicate that our method outperforms state-of-the-art ensemble attacks over multiple pixel-to-pixel vision tasks including image translation and salient region prediction. Our method successfully fools two online black-box saliency prediction systems including DeepGaze-II (Kummerer 2017) and SALICON (Huang et al. 2017). Finally, we also contribute a new repository to promote the research on adversarial attack and defense over pixel-to-pixel tasks: https://github.com/CZHQuality/AAA-Pix2pix. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Patrick Le Callet |
AAAI | 4 |
| 2020 | Few-Shot Pill RecognitionabstractPill image recognition is vital for many personal/public health-care applications and should be robust to diverse unconstrained real-world conditions. Most existing pill recognition models are limited in tackling this challenging few-shot learning problem due to the insufficient instances per category. With limited training data, neural network-based models have limitations in discovering most discriminating features, or going deeper. Especially, existing models fail to handle the hard samples taken under less controlled imaging conditions. In this study, a new pill image database, namely CURE, is first developed with more varied imaging conditions and instances for each pill category. Secondly, a W2-net is proposed for better pill segmentation. Thirdly, a Multi-Stream (MS) deep network that captures task-related features along with a novel two-stage training methodology are proposed. Within the proposed framework, a Batch All strategy that considers all the samples is first employed for the sub-streams, and then a Batch Hard strategy that considers only the hard samples mined in the first stage is utilized for the fusion network. By doing so, complex samples that could not be represented by one type of feature could be focused and the model could be forced to exploit other domain-related information more effectively. Experiment results show that the proposed model outperforms state-of-the-art models on both the National Institute of Health (NIH) and our CURE database. Suiyi Ling, Andreas Pastor, Jing Li 0026, Zhaohui Che, Junle Wang, Patrick Le Callet |
CVPR | 1 |
| 2020 | Towards Perceptually-Optimized Compression Of User Generated Content (UGC): Prediction Of UGC Rate-Distortion CategoryabstractHow to best evaluate the perceptual quality, and efficiently optimize the compression of User Generated Content (UGC) within an adaptive streaming system is becoming one of the most intractable challenges in the community. Rate-Distortion (R-D) characteristic based content analyses, which could be applied on the non-pristine originals, is inevitable to provide guidance in developing quality metrics and efficient compression system. To this end, we present a novel complete R-D category prediction system through the identification of discriminate features. To better understand the Rate-Distortion (R-D) behaviors of UGC, we first propose a Bjontegaard Delta (BD)-Rate, BD-Quality-based algorithm to categorize UGC. By using the predicted R-D related categories as ground-truth labels, we further identify features that characterize the R-D behaviors of UGC via a hierarchical feature selection framework. Finally, selected features are employed to predict the R-D category of under-test UGC. Comprehensive observations and results are summarized through extensive experiments. Suiyi Ling, Yoann Baveye, Patrick Le Callet, Jim Skinner, Ioannis Katsavounidis |
ICME | 1 |
| 2020 | A Probabilistic Graphical Model for Analyzing the Subjective Visual Quality Assessment Data from CrowdsourcingabstractThe swift development of the multimedia technology has raised dramatically the users' expectation on the quality of experience. To obtain the ground-truth perceptual quality for model training, subjective assessment is necessary. Crowdsourcing platform provides us a convenient and feasible way to run large-scale experiments. However, the obtained perceptual quality labels are generally noisy. In this paper, we propose a probabilistic graphical annotation model to infer the underlying ground truth and discovering the annotator's behavior. In the proposed model, the ground truth quality label is considered following a categorical distribution rather than a unique number, i.e., different reliable opinions on the perceptual quality are allowed. In addition, different annotator's behaviors in crowdsourcing are modeled, which allows us to identify the possibility that the annotator makes noisy labels during the test. The proposed model has been tested on both simulated data and real-world data, where it always shows superior performance than the other state-of-the-art models in terms of accuracy and robustness. Jing Li 0026, Suiyi Ling, Junle Wang, Patrick Le Callet |
ACM Multimedia | 2 |
| 2020 | Few-Shot Object Detection in Real Life: Case Study on Auto-HarvestabstractConfinement during COVID-19 has caused serious effects on agriculture all over the world. As one of the efficient solutions, mechanical harvest/auto-harvest that is based on object detection and robotic harvester becomes an urgent need. Within the auto-harvest system, robust few-shot object detection model is one of the bottlenecks, since the system is required to deal with new vegetable/fruit categories and the collection of large-scale annotated datasets for all the novel categories is expensive. There are many few-shot object detection models that were developed by the community. Yet whether they could be employed directly for real life agricultural applications is still questionable, as there is a context-gap between the commonly used training datasets and the images collected in real life agricultural scenarioas. To this end, in this study, we present a novel cucumber dataset and propose two data augmentation strategies that help to bridge the context-gap. Experimental results show that 1) the state-of-the-art few-shot object detection model performs poorly on the novel `cucumber' category; and 2) the proposed augmentation strategies outperform the commonly used ones. Kévin Riou, Suiyi Ling, Mathis Piquet, Vincent Truffault, Patrick Le Callet |
MMSP | 3 |
| 2019 | Perceptual Representations of Structural Information in Images: Application to Quality Assessment of Synthesized View in FTV ScenarioabstractAs the immersive multimedia techniques like Free-viewpoint TV (FTV) develop at an astonishing rate, user's demand for high-quality immersive contents increases dramatically. Unlike traditional uniform artifacts, the distortions within immersive contents could be non-uniform structure-related and thus are challenging for commonly used quality metrics. Recent studies have demonstrated that the representation of visual features can be extracted from multiple levels of the hierarchy. Inspired by the hierarchical representation mechanism in the human visual system (HVS), in this paper, we explore to adopt structural representations to quantitatively measure the impact of such structure-related distortion on perceived quality in FTV scenario. More specifically, a bio-inspired full reference image quality metric is proposed based on 1) low-level contour descriptor; 2) mid-level contour category descriptor; and 3) task-oriented non-natural structure descriptor. The experimental results show that the proposed model outperforms significantly the state-of-the-art metrics. Suiyi Ling, Jing Li 0026, Patrick Le Callet, Junle Wang |
ICIP | 1 |
| 2019 | Quality assessment for view synthesis using low-level and mid-level structural representation
Yu Zhou 0009, Leida Li, Suiyi Ling, Patrick Le Callet |
Signal Process. Image Commun. | 3 |
| 2018 | How to Learn the Effect of Non-Uniform Distortion on Perceived Visual Quality? Case Study Using Convolutional Sparse Coding for Quality Assessment of Synthesized ViewsabstractMachine learning has been attached greater attention in the field of quality assessment and related models have been developed by assuming that any region within images share the same quality score. However, this assumption may no longer stand when non-uniform distortions exist, e.g. geometric distortion within synthesized views in the case of free- view point video TV (FTV). In this paper, we explore to use Convolutional Sparse Coding (CSC), which computes a sparse representation for an entire image with the sum of a set of convolutions with dictionary filters instead of a linear combination of a set of dictionary atoms, to learn a visible codebook and propose a methodology to quantify the geometric distortion without using the reference. Experimental results show that the proposed no reference metric is the most visual friendly and reliable metric among the compared blind image metrics designed for synthesized images. Suiyi Ling, Patrick Le Callet |
ICIP | 1 |
| 2018 | No-Reference Quality Assessment for Stitched Panoramic Images Using Convolutional Sparse Coding and Compound Feature SelectionabstractImage stitching-composition of different viewpoint images to form a 360-degree panoramic image-is an essential component towards immersive applications like VR and AR. A no-reference (NR) quality metric specifically designed to evaluate stitched panoramic images is highly desirable when ground-truth reference images are not available. In this paper, we use Convolutional Sparse Coding (CSC) with a set of convolutional filters to locate stitching-specific distortions in a target image, and design trained kernels to quantify the compound effects of multiple distortion types in a local region. Specifically, our contributions are: i) a training database labeled with location information of the distortion regions is released; ii) a NR metric is proposed to accurately assess stitching-specific artifacts like ghosting using convolutional sparse coding; and iii) a novel sequential feature selection algorithm is proposed to quantify the aforementioned compound distortion effects. In extensive experiments, we show that the performance of our proposed NR metric is comparable to the state-of-the-art full-reference metrics designed for stitched images. Suiyi Ling, Gene Cheung, Patrick Le Callet |
ICME | 1 |
| 2018 | Hybrid-MST: A Hybrid Active Sampling Strategy for Pairwise Preference AggregationabstractIn this paper we present a hybrid active sampling strategy for pairwise preference aggregation, which aims at recovering the underlying rating of the test candidates from sparse and noisy pairwise labeling. Our method employs Bayesian optimization framework and Bradley-Terry model to construct the utility function, then to obtain the Expected Information Gain (EIG) of each pair. For computational efficiency, Gaussian-Hermite quadrature is used for estimation of EIG. In this work, a hybrid active sampling strategy is proposed, either using Global Maximum (GM) EIG sampling or Minimum Spanning Tree (MST) sampling in each trial, which is determined by the test budget. The proposed method has been validated on both simulated and real-world datasets, where it shows higher preference aggregation ability than the state-of-the-art methods. Jing Li 0026, Rafal Mantiuk, Junle Wang, Suiyi Ling, Patrick Le Callet |
NeurIPS | 4 |
| 2017 | Image quality assessment for free viewpoint video based on mid-level contours featureabstractFree view point video (FVV), which offers immersive experience to users with multiple views, is one of the new trends in advanced visual media. These new viewpoints are traditionally synthesized via depth image-based rendering(DIBR) and geometric distortions are therefore observed. Mid-level contours descriptors are capable of evaluating such edges incoherence among the synthesized images which common image quality metrics fail to capture. In this paper, we use the concept of ‘Sketch Token’, that is a mid-level contours descriptor, and introduce a novel metric for DIBR-synthesized image quality assessment by measuring how classes of contours change after synthesis. Experiments are conducted on the IRCCyN/IVC DIBR image database and the results show that the proposed metric achieves a correlation of 88.77% which is comparable to state-of-the-art metrics like MW-PSNR and MP-PSNR. Suiyi Ling, Patrick Le Callet |
ICME | 1 |
| 2017 | Image Quality Assessment for DIBR Synthesized Views using Elastic MetricabstractFrames of free viewpoint video (FVV) synthesized with depth image-based rendering (DIBR) mainly contains special local artifacts like geometric distortions, in which the shape of objects may be stretched/bent. Human observers tend to perceive such local severe deformations instead of consistent shifting artifacts that penalized by most of the existing metrics. Elastic metric is capable of measuring the difference in stretching or bending between two curves, and thus is suitable for evaluating such geometric distortions. In this paper, an elastic metric based image quality assessment (EM-IQA) scheme is proposed by first selecting local distortion regions and then quantifying the deformations of curves. According to the experimental results on the IRCCyn/IVC DIBR image database, the proposed EM-IQA outperforms the state of the art metrics designed for synthesized images and obtains a gain of 6.97% in pearson correlation compared to the second best performing MP-PSNRreduced. Suiyi Ling, Patrick Le Callet |
ACM Multimedia | 1 |
| 2017 | Quality assessment for synthesized view based on variable-length context treeabstractIn free viewpoint television (FTV) application scenario, views that synthesized with depth image-based rendering (DIBR) techniques mainly contain special artifacts like geometric distortions. These artifacts may affect the structure of images/videos by changing the global contour characteristics and thus are annoying for human observers. Context tree based contour coding scheme can be a good tool to measure such structure loss since the more geometric distortion there is in the synthesized view the lager the gap between the encoding cost of the reference and synthesized views. In this paper, we investigate whether such overall encoding cost can be related to perceptual annoyance reflecting in quality score and propose a variable-length context tree based image quality assessment (CT-IQA) scheme. This scheme quantify (1) the overall structure dissimilarity and (2) dissimilarities in various contour characteristics between the reference and synthesized views. The proposed metric is robust to global shifting artifact that is over penalized by traditional metrics. According to the experimental results on the IRCCyn/IVC DIBR image database, the performance of the proposed CT-IQA is promising. Suiyi Ling, Patrick Le Callet, Gene Cheung |
MMSP | 1 |
| 2016 | Effect of content features on short-term video quality in the visual peripheryabstractThe area outside our central field of vision, also referred to as the visual periphery, captures most information in a visual scene, although much less sensitive than the central Fovea. Vision studies in the past have stated that there is reduced sensitivity of texture, color, motion and flicker (temporal harmonic) perception in this area, that bears an interesting application in the domain of quality perception. In this work, we particularly analyze the perceived subjective quality of videos containing H.264/AVC transmission impairments, incident at various degrees of retinal eccentricities of observers. We relate the perceived drop in quality, to five basic types of features that are important from a perceptive standpoint: texture, color, flicker, motion trajectory distortions and also the semantic importance of the underlying regions. We are able to observe that the perceived drop in quality across the visual periphery, is closely related to the Cortical Magnification fall-off characteristics of the V1 cortical region. Additionally, we see that while object importance and low frequency spatial distortions are important indicators of quality in the central foveal region, temporal flicker and color distortions are the most important determinants of quality in the periphery. We therefore conclude that, although users are more forgiving of distortions they viewed peripherally, they are nevertheless not totally blind towards it: the effects of flicker and color distortions being particularly important. Yashas Rai, Ahmed Aldahdooh, Suiyi Ling, Marcus Barkowsky, Patrick Le Callet |
MMSP | 3 |