EDBT 2026 Demo / reviewers in the wild / expert
Sumei Li
dblp:159/9554
· DBLP profile ↗
63ranked-venue papers
13as first author
41since 2021 · last 2026
0000-0002-4793-3161ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 62 · 12 first-author · 41 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SFSIQA-Net: stereo image quality assessment based on selective feedback and transposed attention binocular fusion
Sumei Li, Mingxuan Xie |
Multim. Syst. | 1 |
| 2025 | EdgeStereoSR: A multi-task network with transformers for stereo image super-resolution considering edge prior
Anqi Liu 0005, Sumei Li, Yongli Chang, Yonghong Hou |
Signal Process. | 2 |
| 2024 | A Lightweight CNN and Spatial-Channel Transformer Hybrid Network for Image Super-ResolutionabstractTransformer-based methods have achieved excellent performance in single image super-resolution (SISR) due to their ability to model long-range dependency. However, most existing methods require huge computational resources, making it difficult to apply on mobile devices with limited computing and storage resources. In this paper, we propose a lightweight CNN and spatial-channel Transformer hybrid network (CSCTHN), which adopts spatial and channel self-attention alternately and leverages the local extraction capability of CNN. CSCTHN’s basic unit CNN and Transformer hybrid module (CTHM) comprises three key components: dual-branch interactive spatial self-attention block (DISSAB) for capturing spatial context with lower computational cost, channel self-attention block (CSAB) for capturing channel information and local feature enhancement block (LFEB) that utilizes CNN to extract local information. Extensive experiments demonstrate that CSCTHN is superior to the state-of-the-art methods in terms of reconstruction performance and model complexity (e.g.,31.33dB@Manga109 ×4 with only 706K parameters). Sumei Li, Xiaoxuan Chen, Peiming Lin |
ICME | 1 |
| 2024 | Top-Down Guidance Based ViT-CNN Network Considering Theme Information for Image Aesthetic AssessmentabstractImage Aesthetic Assessment (IAA) is a challenging task that is closely tied to human aesthetic experience. In this paper, inspired by Top-Down guidance and visual attention mechanism, we propose a Top-Down Guidance based ViT-CNN network considering theme information. Considering the guidance of global information on local information, we construct a two-stream network structure. It consists of Vision Transformer (ViT) and Convolutional Neural Network (CNN) streams. Meanwhile, a global and local feature attention guidance module (GLFAGM) is proposed to better realize the guidance from global features of ViT stream down to local features of CNN stream. In addition, considering the importance of theme information, the proposed network utilizes more comprehensive theme information as auxiliary information to achieve aesthetic assessment. To better utilize theme information, an attentionbased theme feature fusion module (ATFFM) is proposed to integrate theme features and visual features from CNN stream. The experimental results show that the proposed method achieves better performance and outperforms some state-of-the-art methods. Sumei Li, Xiaofei He 0011, Hangwei Liang |
ICME | 1 |
| 2024 | Multi-Scale and Multi-Patch Aggregation Network Based on Dual-Column Vision Fusion for Image Aesthetics AssessmentabstractAssessing image aesthetics requires a multi-level aesthetics representation and a comprehensive aesthetics vision. Therefore, providing fine-grained information and multi-scale information is of great significance in aesthetics assessment. The use of multi-scale information to process image aesthetics features and the multi-patch as input has become a common method in image aesthetics assessment (IAA). In this paper, we propose a multi-scale and multi-patch aggregation network based on dual-column vision fusion (MMANet) for IAA. First instead of random cropping, we proposed a visual saliency guided cropping (VS cropping) to get multi-patch. Second to effectively capture the characteristics between patches, we propose a multi-scale aesthetics patch fusion attention (MAPA) module based on the human visual stereoscopic imaging mechanism. In addition, a multi-patch feature aggregation (MFA) module for further fusing the features of multi-patch information for IAA is proposed. Extensive experiments on three public IAA databases demonstrate the superiority of the proposed MMANet model over the state-of-the-arts. Sumei Li, Hangwei Liang, Mingxuan Xie, Xiaofei He 0011 |
ICME | 1 |
| 2024 | I2GSRnet: Iterative Interaction Guidance Network for Stereo Image Super-ResolutionabstractStereo super-resolution (SR) aims to reconstruct high-resolution (HR) images from low-resolution (LR) stereo image pairs. However, existing methods ignore the connection between the interaction information at each stage in the image pair and acquire global information from the cross-view information with more computational complexity. In this work, we propose an iterative interaction guidance network (I2GSRnet) to obtain the rich cross-view information between image pairs. First, we propose an iterative interaction strategy to fully leverage the interaction guidance information between each iteration. Then, we design an adaptive bidirectional parallax attention module (ABPAM) to efficiently acquire parallax information horizontally and capture global information vertically with less complexity. Finally, we propose a distillation guidance module (DGM) to fully integrate cross-view information with intra-view information using multi-scale information guided by multiple receptive fields. Extensive experiments show that our approach achieves state-of-the-art performance on the KITTI2012, KITTI2015, Middlebury, and Flickr1024 datasets. Peiming Lin, Sumei Li, Zilin Zhao, Huilin Zhang |
ICME | 2 |
| 2024 | Pseudo-global strategy-based visual comfort assessment considering attention mechanism
Sumei Li, Huilin Zhang, Mingyue Zhou |
Multim. Syst. | 1 |
| 2024 | Multi-Scale Visual Perception Based Progressive Feature Interaction Network for Stereo Image Super-ResolutionabstractIn recent years, stereo image super-resolution based on convolutional neural network has been extensively researched and achieved impressive performance by introducing complementary information from another view. However, most existing methods still cannot fully capture both intra- and cross-view information due to the neglect of multi-scale information perception, multi-scale binocular alignment and the excitation of large scale to small scale in human vision system. And they generated blurry results due to the consideration of irrelevant information in search for cross-view information. To address these issues, we propose a multi-scale visual perception based progressive feature interaction network (MS-PFINet) for stereo image super-resolution. Specifically, to exploit comprehensive intra- and cross-view information for image reconstruction, we design a two-stream network with multi-branch structure to extract multi-scale features and progressively use cross-view interaction at larger scales to guide that at smaller scales. Moreover, to explore more proper and accurate cross-view information, we propose a feature transformer module (FTM) to search and transfer the most relevant features from another view by hard attention maps and soft attention maps, which are calculated by patch-wise similarity rather than pixel-wise. In addition, in order to encourage a more effective way to transfer texture features for the target view, we propose a perceptual texture matching loss to supervise the accuracy of feature transformer modules. Experimental results show that our proposed method is superior to the state-of-the-art methods in most cases. Anqi Liu 0005, Sumei Li, Yongli Chang, Yonghong Hou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Coarse-to-Fine Cross-View Interaction Based Accurate Stereo Image Super-Resolution NetworkabstractRecently, parallax attention based stereo image super-resolution (SR) methods, which can better explore cross-view information, have been widely studied. Despite the impressive performance of these methods, almost all of them calculate parallax attention maps at a single low resolution, which will lead to ambiguous stereo correspondence. Besides, the widely used parallax attention module (PAM) cannot handle the illuminance variations in stereo image pairs, and cannot distinguish the contribution of the captured cross-view features to the reconstruction of the target view. To this end, in this paper, we propose a coarse-to-fine cross-view interaction based network (C2FNet) to achieve more accurate cross-view information capturing. Firstly, in C2FNet, a coarse-to-fine cascaded parallax attention structure (C2F-CPAS), which conforms with the human visual mechanism, is constructed to gradually perform parallax attention from the low-resolution to high-resolution level. Thus, richer textures can be used to learn more reliable stereo correspondence. Meanwhile, a multi-level attention transfer loss is designed to further calibrate the accuracy of stereo correspondence at each level. Secondly, we propose a modified PAM (MPAM) to alleviate the limitations of common PAM so that illuminance-robust stereo correspondence can be learned and more important cross-view information can be selected. Extensive experimental results show that our proposed C2FNet outperforms the state-of-the-art methods on various datasets. Anqi Liu 0005, Sumei Li, Yongli Chang, Yonghong Hou |
IEEE Trans. Multim. | 2 |
| 2023 | Multi-Level Feature-Guided Stereoscopic Video Quality Assessment Based on Transformer and Convolutional Neural NetworkabstractStereoscopic video (3D video) has been increasingly applied in industry and entertainment. And the research of stereoscopic video quality assessment (SVQA) has become very important for promoting the development of stereoscopic video system. Many CNN-based models have emerged for SVQA task. However, these methods ignore the significance of the global information of the video frames for quality perception. In this paper, we propose a multi-level feature-fusion model based on Transformer and convolutional neural network (MFFTCNet) to assess the perceptual quality of the stereoscopic video. Firstly, we use global information from Transformer to guide local information from convolutional neural network (CNN). Moreover, we utilize low-level features in the CNN branch to guide high-level features. Besides, considering the binocular rivalry effect in the human vision system (HVS), we use 3D convolution to achieve rivalry fusion of binocular features. The proposed method is tested on two public stereoscopic video quality datasets. The result shows that this method correlates highly with human visual perception and outperforms state-of-the-art (SOTA) methods by a significant margin. Sumei Li |
ICME | 2 |
| 2023 | Cross-Level Attention Based Adaptive Feature Alignment Network for Arbitrary-Shaped Text DetectionabstractOver the past few years, segmentation-based text detection methods have achieved great progress in the scene text detection field. However, most of the existing methods with FPN tend to ignore the feature misalignment issues caused by semantic gap between features at different levels, leading to inaccurate segmentation prediction. Therefore, we propose an end-to-end trainable text detector to alleviate the above dilemma. Specifically, we propose an Adaptive Feature Transform Module (AFTM) to adaptively align features of different levels. Furthermore, a Cross-Level Attention (CLA) block is developed to capture the cross-level information, and selectively enhance element-wise features of multi-level feature maps. Experiments on four benchmark datasets, MSRA-TD500, Total-Text, CTW1500 and ICDAR2015, demonstrate that our proposed method has competitive performance and strong robustness. Sumei Li |
ICME | 2 |
| 2023 | Binocular Rivalry and Fusion Mechanisms Based No-Reference Stereoscopic Image Quality Assessment Considering Feedback GuidanceabstractWith the growing popularity of 3D content and virtual reality applications, effective no-reference stereoscopic image quality assessment (NR-SIQA) methods have become increasingly important. In this paper, we propose a convolutional neural network (CNN) SIQA model based on binocular rivalry and fusion mechanisms of the human visual system (HVS). In order to get a better representation of binocular information, we propose Two-Stage Enhanced Fusion Module (TSEFM) that consists of two stages for monocular features enhancement and binocular features fusion, respectively. Given the dynamical characteristics of binocular rivalry phenomenon, the proposed Content-Aware Binocular Rivalry Fusion Module (CABRFM) dynamically and adaptively adjusts its output based on the input content. Additionally, considering that feedback mechanism of HVS is indispensable and significant, we introduce feedback connections during feature aggregation to realize the guidance of high-level features to low-level features. Extensive experimental results demonstrate the superiority of our method over state-of-the-art metrics, showcasing its excellent performance. Haoxiang Chang, Sumei Li, Dongyue He |
VCIP | 2 |
| 2023 | Feature Reinforced and Adaptive Attention Guided Network for Multi-oriented Scene Text DetectionabstractThis paper proposes a text detection method for multi-oriented scene text detection based on feature reinforcement and adaptive text attention. Firstly, the method constructs different feature reinforcement modules for different levels of features. The high-level feature reinforcement module (HFRM) acquires multi-scale semantic information through the deformed Inception structure to enhance the expression ability of features; the low-level feature reinforcement module (LFRM) captures richer semantic information by constructing a residual-like connection structure to alleviate the semantic gap problem between different levels of features. Besides, the method also designs an adaptive text attention module (ATAM), which can adaptively adjust the attention weights according to the different input features, thus enhancing the fitting ability of the network. We carry out a series of experiments on ICDAR2015, TD500, and ICDAR2017-MLT to confirm the effectiveness of the proposed method. Sumei Li, Dongyue He |
VCIP | 2 |
| 2023 | Multi-scale Parallax Attention for Stereo Image Super-ResolutionabstractStereo image super-resolution (SR) has achieved great progress in recent years. However, the existing methods are unable to obtain rich cross-view information at a low computational cost. In addition, these methods treat each pixel equally when fusing the cross-view information with the intra-view information, resulting in non-robustness of the fused information. In this work, we propose a multi-scale parallax attention stereo super-resolution network (MPASSRnet) to address these problems. Firstly, we design a multi-scale parallax attention module (MSPAM), which computes the similarity between stereo images on multiple scale images based on the introduction of pixel-to-patch matching, thus acquiring multi-scale cross-view information at a low computational cost. Second, we propose a pixel attention fusion module (PAFM), which introduces spatial attention (SA) by calculating the correlation between cross-view information and intra-view information, and combines channel attention (CA) to make the fused information more robust. Finally, extensive experiments show that our method achieves state-of-the-art performance on the benchmark datasets. Zilin Zhao, Sumei Li |
VCIP | 2 |
| 2023 | Coarse-to-Fine Feedback Guidance Based Stereo Image Quality Assessment Considering Dominant Eye FusionabstractConsidering that the human brain always follows a coarse-to-fine (low-to-high spatial frequency) visual processing and fusion mechanism, we propose a coarse-to-fine feedback guidance based stereo image quality assessment (SIQA) network which considers a coarse-to-fine feedback guidance and adaptive dominant eye mechanism. The proposed network consists of two main sub-network streams, each of which has three branches to extract low, middle and high spatial frequency information in parallel. To better realize the guidance of the high-level features in the low spatial frequency branch to the low-level features in the high spatial frequency branch, an information feedback guidance module (IFGM) is proposed, which realizes a top-down guidance mechanism in each sub-network stream. Simultaneously, according to the theory of ocular dominance in human visual system (HVS), we design an adaptive bi-directional parallax-based binocular fusion module (BPBFM), which synthesizes two types of fusion feature by taking the left and right view features as dominant eye input. Furthermore, in order to obtain the better perceptual quality of stereo images, we design a weighted fusion strategy to weigh the quality scores from the two types of fusion features obtained by using an ensemble model with two multi-layer perceptrons (MLPs). The experimental results on four public stereo image datasets show that the proposed method is superior to the mainstream metrics and achieves an excellent performance. Yongli Chang, Sumei Li, Anqi Liu 0005, Wei Xiang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Cross-resolution feature attention network for image super-resolution
Sumei Li, Yongli Chang |
Vis. Comput. | 2 |
| 2022 | Quality Assessment of Screen Content Images Based on Multi-Pathway Convolutional Neural NetworkabstractIn this paper, considering the retinal structure of human eye, and the composition characteristics of screen content images (SCIs), a multi-pathway convolutional neural network (CNN) with picture-text competition is proposed for SCIs quality assessment. According to the visual mechanism of human retina, we design a retinal structure simulation module, which uses multiple parallel convolution pathways to simulate the parallel transmission of visual signals by bipolar cells and uses a multi-pathway feature fusion (MPFF) module to allocate the weight for each channel to simulate horizontal cells' regulation of the information transmission. In addition, we design an adaptive feature extraction and competition module (AFEC) to directly extract the features of textural and pictorial regions and distribute the weight. Furthermore, the attention module combined with deformable convolution and channel attention can accurately extract image edge features and reduce redundancy of information. Experimental results show that the proposed method is superior to the mainstream methods. Sumei Li |
VCIP | 2 |
| 2022 | Semantic Compensation Based Dual-Stream Feature Interaction Network for Multi-oriented Scene Text DetectionabstractDue to the various appearances of scene text instances and the disturbance of background, it is still a challenging task to design an effective and accurate text detector. To tackle this problem, in this paper we propose a novel dual-stream scene text detector considering semantic compensation and feature interaction. The detector extracts image features from two input images of different resolution, which improves its perceptive ability and contributes to detecting large and long texts. Specifically, we propose a Semantic Compensation Module (SCM) to aggregate features between the two streams, which compensates semantic information in features at each level via an attention mechanism. Moreover, we design a Feature Interaction Module (FIM) to obtain more expressive features. Experiments conducted on three benchmark datasets, ICDAR2015, MSRA-TD500 and ICDAR2017-MLT, demonstrate that our proposed method has competitive performance and strong robustness. Siyan Wang, Sumei Li |
VCIP | 2 |
| 2022 | No-reference Stereoscopic Image Quality Assessment Based on Parallel Multi-scale PerceptionabstractWith the rapid development of 3D technologies, effective no-reference stereoscopic image quality assessment (NR-SIQA) methods are in great demand. In this paper, we propose a parallel multi-scale feature extraction convolution neural network (CNN) model combined with novel binocular feature interaction consistent with human visual system (HVS). In order to simulate the characteristics of HVS sensing multi-scale information at the same time, parallel multi-scale feature extraction module (PMSFM) followed by compensation information is proposed. And modified convolutional block attention module (MCBAM) with less computational complexity is designed to generate visual attention maps for the multi-scale features extracted by the PMSFM. In addition, we employ cross-stacked strategy for multi-level binocular fusion maps and binocular disparity maps to simulate the hierarchical perception characteristics of HVS. Experimental results show that our method is superior to the state-of-the-art metrics and achieves an excellent performance. Sumei Li |
VCIP | 2 |
| 2022 | No Reference Stereoscopic Video Quality Assessment based on Human Vision SystemabstractIn this paper, we propose a no-reference stereoscopic video quality assessment (NR-SVQA) based on human vision system (HVS). Firstly, we build a frequency transform module (FTM), which maps spatial domain to frequency domain by cosine discrete transform (DCT), and selects important frequency components through channel attention mechanism. Secondly, we use dynamic convolution to regionally process the same input. Thirdly, we use convolutional long short term memory (Conv-LSTM) to extract spatio-temporal information rather than just temporal information. Finally, in order to better simulate the visual characteristics of human eyes, we build a optic chiasm module. The experiment results show that our method outperforms any other methods. Sumei Li |
VCIP | 2 |
| 2022 | Face Super Resolution based on Contrastive LearningabstractFace super resolution (FSR) is a sub-field of super resolution (SR), which is to reconstruct low resolution (LR) face image into high resolution (HR) face image. Recently, the FSR methods based on face prior have been proved to be effective in FSR on higher upscaling factors. However, existing prior guided methods mostly adopt supervised prior extraction models trained with labels. The performance of supervised prior extraction method mainly depends on the accuracy of label so that the implicit informations of data are not fully utilized. And in practical application, the label acquisition work is routine and laborious. Therefore, to solve these problems, this paper proposes a novel contrastive learning (CL) based FSR method, which is based on the iterative collaboration of image reconstruction network and contrastive learning network. In each iteration, the reconstruction network uses the priors generated by the contrastive learning network to assist the image reconstruction and generates higher-quality SR images. Then, the SR image will feed into contrastive learning network to obtain more accurate prior. In addition, a new contrastive learning constraint function is designed to extract the representation of the augmented facial image as a prior by analysing the principal component information of the image. Quantitative and qualitative experimental results show that the proposed method is superior to the most advanced FSR method in high-quality face images super resolution reconstruction. Sumei Li, Liqin Huang |
VCIP | 2 |
| 2022 | Stereo image quality assessment considering the difference of statistical feature in early visual pathway
Yongli Chang, Sumei Li, Anqi Liu 0005, Wei Xiang 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Multi-Scale Feature-Guided Stereoscopic Video Quality Assessment Based on 3d Convolutional Neural NetworkabstractWith the huge development of stereoscopic video technology, the research of stereoscopic video quality assessment (SVQA) has become very important for promoting the development of stereoscopic video system. These years, many SVQA methods based on convolutional neural network (CNN) have emerged. In this paper, we proposed a multi-scale feature-guided 3D convolutional neural network for SVQA which not only use 3D convolution to capture spatio-temporal features but also aggregate multi-scale information by a new multi-scale unit. Besides, we employ a multi-stage growing attention mechanism in this network to learn more critical deep semantic information. The proposed method is tested on two public stereoscopic video quality datasets, and the result shows that this method correlates highly with human visual perception and outperforms state-of-the-art methods by a large margin. Yingjie Feng, Sumei Li, Yongli Chang |
ICASSP | 2 |
| 2021 | Image Super-Resolution Using Multi-Resolution Attention NetworkabstractIn recent years, single image super-resolution based on convolution neural network (CNN) has been extensively researched. However, most CNN-based methods only focus on mining features at a single resolution, which will cause the loss of some useful information. Besides, most of them still have difficulty in training and obtaining high-quality images for large scale factors. To address these issues, we propose a multi-resolution attention network (MRAN), which progressively reconstructs images at large scale factors by aggregating features from multiple resolutions. Specially, a multi-resolution residual block (MRRB) is designed as basic block to specialize features at different resolutions and share information across different resolutions, improving the representation ability of features. Simultaneously, we design a resolution-wise attention block (RAB) to evaluate the importance of features from different resolutions, making the use of features more effective and enhancing feature fusion. Experimental results show that our proposed method is superior to the state-of-the-art methods. Sumei Li, Yongli Chang |
ICASSP | 2 |
| 2021 | No-Reference Stereoscopic Image Quality Assessment Based on the Human Visual SystemabstractStereoscopic image quality assessment (SIQA) is to predict the human perception quality of stereoscopic image pairs, which is more challenging than previous 2D image quality assessment due to the complicated binocular vision mechanism in the human visual system (HVS). Recently witnessed the significant progress of biotechnology and motivated by the deeper research on the HVS, we take a step to bridge the gap between HVS and SIQA by generalizing the optic chiasm algorithm and introducing biological vision fusion mechanism in our work. Firstly, a network structure is proposed in our work, which consists of an optic chiasm module for binocular information exchange and a multi-scale feature extraction module for binocular information fusion. Secondly, a mutual-perception attention fusion module is designed for simulating binocular fusion. In addition, we come up with an innovative data enhancement method. Experimental results show that the image quality assessment score obtained by the network is more consistent with human perception. Sumei Li, Yongli Chang |
ICASSP | 2 |
| 2021 | Quality Assessment of Screen Content Images Based on Convolutional Neural Network with Dual PathwaysabstractTo simulate the characteristics of perceiving things from binocular vision, a dual-pathway convolutional neural network (CNN) for quality assessment of screen content images (SCIs) is proposed. Considering the different sensitivity of retinal photoreceptor cells to RGB colors and the human visual attention mechanism, we employ a convolutional block attention module (CBAM) to weight the RGB channels and their spatial position on each channel. And 3D convolution considering inter-frame information is used to extract the correlation features between RGB channels. Moreover, because of the important role of optic chiasm in binocular vision, we design its simulation strategy in the proposed network. Furthermore, since the characteristics of multi-scale and multi-level are indispensable to perception of any objects in human visual system (HVS), a new multi-scale and multi-level feature fusion (MSMLFF) module is built to obtain perceptual features of different scales and levels. Experimental results show that the proposed method is superior to several mainstream SCIs metrics on publicly accessible databases. Yongli Chang, Sumei Li |
ICIP | 2 |
| 2021 | Two-Way Guided Super-Resolution Reconstruction Network Based on Gradient PriorabstractDeep convolutional neural networks (CNNs) have demonstrated remarkable progress on single image super-resolution (SISR). However, most networks do not make full use of the rich prior information of the image itself, and get smoother results. To solve this problem, we propose a two-way guided super-resolution reconstruction network based on gradient prior (TWGSR), which uses gradient branches to achieve two-way guidance and employs multi-level gradient loss constraints to reconstruct high-resolution images with more textures. In addition, we propose a multi-scale module (MSM), which adopts dilated convolution to aggregate feature information of different scales, and we add an error feedback module (EFM) to compensate for sampling errors to refine the extracted feature maps. Furthermore, we propose an improved cross residual-in-residual dense block (CRRDB) to enhance the feature extraction capability of the module. Experimental results show that our TWGSR achieves favorable performance against state-of-the-art methods. Sumei Li |
ICIP | 2 |
| 2021 | Enhanced Back Projection Network Based Stereo Image Super-Resolution Considering Parallax AttentionabstractRecent years have witnessed great advances in stereo image super-resolution (SR). However, the existing methods only consider the horizontal parallax when capturing the stereo correspondence, which is insufficient because the vertical parallax inevitably exists in stereo image pairs. To address this problem, we propose an enhanced back projection stereo SR network (EBPSSRnet) to make full use of the complementary information in stereo images for more accurate SR results. Specifically, we propose a relaxed parallax attention module (rePAM) to handle different stereo images with vertical and horizontal parallax. Then, an enhanced back projection block (EBPB) is developed to extract discriminative features for capturing the stereo correspondence and consolidate the best representation for reconstruction. Extensive experiments show that the proposed method achieves state-of-the-art performance on the Flickr1024, Middlebury, KITTI2012 and KITTI2015 datasets. Sumei Li |
ICIP | 2 |
| 2021 | Semantic-Compensated and Attention-Guided Network for Scene Text DetectionabstractRecently, scene text detection based on deep learning has witnessed rapid development. However, most previous methods suffer from (i) semantic information diluted, (ii) false positives in their detections. To alleviate the above dilemmas, we propose an end-to-end trainable text detector named Semantic-compensated and Attention-guided Network (SANet). It contains a Semantic Compensation Module (SCM) and a Text Attention Module (TAM). Specifically, SCM aims to alleviate the dilution of semantic information via compensating the semantic information to the features at each level in a top-down pathway. Furthermore, TAM is employed to encode the strong supervised information into convolutional features, where the text-related features are significantly enhanced by increasing the response gap between text and background. The experimental results on three benchmark datasets, ICDAR2015, MSRA-TD500, and ICDAR2017-MLT prove the effectiveness of our method. Yizhan Zhao, Sumei Li |
ICIP | 2 |
| 2021 | Stereoscopic Video Quality Assessment with Multi-level Binocular Fusion Network Considering Disparity and Multi-scale InformationabstractStereoscopic video quality assessment (SVQA) is of great importance to promote the development of the stereoscopic video industry. In this paper, we propose a three-branch multi-level binocular fusion convolutional neural network (MBFNet) which is highly consistent with human visual perception. Our network mainly includes three innovative structures. Firstly, we construct a multi-scale cross-dimension attention module (MSCAM) on the left and right branches to capture more critical semantic information. Then, we design a multi-level binocular fusion unit (MBFU) to fuse the features from left and right branches adaptively. Besides, a disparity compensation branch (DCB) containing an enhancement unit (EU) is added to provide disparity feature. The experimental results show that the proposed method is superior to other existing SVQA methods with state-of-the-art performance. Yingjie Feng, Sumei Li |
VCIP | 2 |
| 2021 | Binocular Visual Mechanism Guided No-Reference Stereoscopic Image Quality Assessment Considering Spatial SaliencyabstractIn recent years, with the popularization of 3D technology, stereoscopic image quality assessment (SIQA) has attracted extensive attention. In this paper, we propose a two-stage binocular fusion network for SIQA, which takes binocular fusion, binocular rivalry and binocular suppression into account to imitate the complex binocular visual mechanism in the human brain. Besides, to extract spatial saliency features of the left view, the right view, and the fusion view, saliency generating layers (SGLs) are applied in the network. The SGL apply multi-scale dilated convolution to emphasize essential spatial information of the input features. Experimental results on four public stereoscopic image databases demonstrate that the proposed method outperforms the state-of-the-art SIQA methods on both symmetrical and asymmetrical distortion stereoscopic images. Jinhui Feng, Sumei Li, Yongli Chang |
VCIP | 2 |
| 2021 | No-Reference Stereoscopic Image Quality Assessment Considering Binocular Disparity and Fusion CompensationabstractIn this paper, we propose an optimized dual stream convolutional neural network (CNN) considering binocular disparity and fusion compensation for no-reference stereoscopic image quality assessment (SIQA). Different from previous methods, we extract both disparity and fusion features from multiple levels to simulate hierarchical processing of the stereoscopic images in human brain. Given that the ocular dominance plays an important role in quality evaluation, the fusion weights assignment module (FWAM) is proposed to assign weight to guide the fusion of the left and the right features respectively. Experimental results on four public stereoscopic image databases show that the proposed method is superior to the state-of-the-art SIQA methods on both symmetrical and asymmetrical distortion stereoscopic images. Jinhui Feng, Sumei Li, Yongli Chang |
VCIP | 2 |
| 2021 | An Error Self-learning Semi-supervised Method for No-reference Image Quality AssessmentabstractIn recent years, deep learning has achieved significant progress in many respects. However, unlike other research fields with millions of labeled data such as image recognition, only several thousand labeled images are available in image quality assessment (IQA) field for deep learning, which heavily hinders the development and application for IQA. To tackle this problem, in this paper, we proposed an error self-learning semi-supervised method for no-reference (NR) IQA (ESSIQA), which is based on deep learning. We employed an advanced full reference (FR) IQA method to expand databases and supervise the training of network. In addition, the network outputs of expanding images were used as proxy labels replacing errors between subjective scores and objective scores to achieve error self-learning. Two weights of error back propagation were designed to reduce the impact of inaccurate outputs. The experimental results show that the proposed method yielded comparative effect. Yingjie Feng, Sumei Li, Sihan Hao |
VCIP | 2 |
| 2021 | Multi-Oriented Text Detection Network Based on Hybrid Feature Enhancement and Shallow Feature RefinementabstractScene text detection has achieved impressive progress over the past years. However, there are still two challenges for text detection. The first challenge is the limitation of receptive field on account of the large aspect ratio of texts. The second one is the loss of spatial information due to the long path between lower layers and topmost feature. To address the two problems, we propose an effective text detection network for multi-oriented text. In this paper, we introduce Hybrid feature enhancement module (HFEM) and Low-level feature refinement module (LFRM). HFEM is a multiple parallel branches module to enlarge receptive field and capture multi-scale information for better detection. LFRM is proposed to suppress background noise and strengthen feature propagation. What's more, low-level feature compensation mechanism preserves rich spatial information for the model. Experiments on datasets including ICDAR 2015, MSRA-TD500 and MLT-2017 validate that the proposed method is effective for multi-oriented text. We also provide ablation experiments on ICDAR 2015 to indicate the effectiveness of proposed modules in our network. Sumei Li, Yizhan Zhao |
VCIP | 2 |
| 2021 | Stereo Image Super-Resolution Based on Pixel-Wise Knowledge Distillation StrategyabstractIn stereo image super-resolution (SR), it is equally important to utilize intra-view and cross-view information. However, most existing methods only focus on the exploration of cross-view information and neglect the full mining of intra-view information, which limits the reconstruction performance of these methods. Since single image SR (SISR) methods are powerful in intra-view information exploitation, we propose to introduce the knowledge distillation strategy to transfer the knowledge of a SISR network (teacher network) to a stereo image SR network (student network). With the help of the teacher network, the student network can easily learn more intra-view information. Specifically, we propose pixel-wise distillation as the implementation method, which not only improves the intra-view information extraction ability of student network, but also ensures the effective learning of cross-view information. Moreover, we propose a lightweight student network named Adaptive Residual Feature Aggregation network (ARFAnet). Its main unit, the ARFA module, can aggregate informative residual features and produce more representative features for image reconstruction. Experimental results demonstrate that our teacher-student network achieves state-of-the-art performance on all benchmark datasets. Sumei Li |
VCIP | 2 |
| 2021 | No-Reference Stereoscopic Image Quality Assessment Based on The Visual Pathway of Human Visual SystemabstractWith the development of stereoscopic imaging technology, stereoscopic image quality assessment (SIQA) has gradually been more and more important, and how to design a method in line with human visual perception is full of challenges due to the complex relationship between binocular views. In this article, firstly, convolutional neural network (CNN) based on the visual pathway of human visual system (HVS) is built, which simulates different parts of visual pathway such as the optic chiasm, lateral geniculate nucleus (LGN), and visual cortex. Secondly, the two pathways of our method simulate the ‘what’ and ‘where’ visual pathway respectively, which are endowed with different feature extraction capabilities. Finally, we find a different application way for 3D-convolution, employing it fuse the information from left and right view, rather than just extracting temporal features in video. The experimental results show that our proposed method is more in line with subjective score and has good generalization. Sumei Li |
VCIP | 2 |
| 2021 | Multi-Dimension Aware Back Projection Network For Scene Text DetectionabstractRecently, scene text detection based on deep learning has progressed substantially. Nevertheless, most previous models with FPN are limited by the drawback of sample interpolation algorithms, which fail to generate high-quality up-sampled features. Accordingly, we propose an end-to-end trainable text detector to alleviate the above dilemma. Specifically, a Back Projection Enhanced Up-sampling (BPEU) block is proposed to alleviate the drawback of sample interpolation algorithms. It significantly enhances the quality of up-sampled features by employing back projection and detail compensation. Further-more, a Multi-Dimensional Attention (MDA) block is devised to learn different knowledge from spatial and channel dimensions, which intelligently selects features to generate more discriminative representations. Experimental results on three benchmarks, ICDAR2015, ICDAR2017- MLT and MSRA-TD500, demonstrate the effectiveness of our method. Yizhan Zhao, Sumei Li, Yongli Chang |
VCIP | 2 |
| 2021 | Two-stage Parallax Correction and Multi-stage Cross-view Fusion Network Based Stereo Image Super-ResolutionabstractStereo image super-resolution (SR) has achieved great progress in recent years. However, the two major problems of the existing methods are that the parallax correction is insufficient and the cross-view information fusion only occurs in the beginning of the network. To address these problems, we propose a two-stage parallax correction and a multi-stage cross-view fusion network for better stereo image SR results. Specially, the two-stage parallax correction module consists of horizontal parallax correction and refined parallax correction. The first stage corrects horizontal parallax by parallax attention. The second stage is based on deformable convolution to refine horizontal parallax and correct vertical parallax simultaneously. Then, multiple cascaded enhanced residual spatial feature transform blocks are developed to fuse cross-view information at multiple stages. Extensive experiments show that our method achieves state-of-the-art performance on the KITTI2012, KITTI2015, Middlebury and Flickr1024 datasets. Yijian Zheng, Sumei Li |
VCIP | 2 |
| 2021 | Deformable Convolution Based No-Reference Stereoscopic Image Quality Assessment Considering Visual Feedback MechanismabstractSimulation of human visual system (HVS) is very crucial for fitting human perception and improving assessment performance in stereoscopic image quality assessment (SIQA). In this paper, a no-reference SIQA method considering feedback mechanism and orientation selectivity of HVS is proposed. In HVS, feedback connections are indispensable during the process of human perception, which has not been studied in the existing SIQA models. Therefore, we design a new feedback module (FBM) to realize the guidance of the high-level region of visual cortex to the low-level region. In addition, given the orientation selectivity of primary visual cortex cells, a deformable feature extraction block is explored to simulate it, and the block can adaptively select the regions of interest. Meanwhile, retinal ganglion cells (RGCs) with different receptive fields have different sensitivities to objects of different sizes in the image. So a new multi receptive fields information extraction and fusion manner is realized in the network structure. Experimental results show that the proposed model is superior to the state-of-the-art no-reference SIQA methods and has excellent generalization ability. Mingyue Zhou, Sumei Li |
VCIP | 2 |
| 2021 | Quality assessment of screen content images based on multi-stage dictionary learning
Yongli Chang, Sumei Li |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Stereoscopic image quality assessment considering visual mechanism and multi-loss constraints
Sumei Li, Yongtian Han |
J. Vis. Commun. Image Represent. | 1 |
| 2020 | Lightweight Image Super-Resolution Reconstruction With Hierarchical Feature-Driven NetworkabstractDeep convolutional neural networks (CNNs) have emerged as powerful tool for single image super-resolution (SISR). However, enormous parameters hinder their real-world applications. To address this issue, we propose a lightweight hierarchical feature-driven network (HFDN) that can fully explore local and global hierarchical feature information. Specifically, we devise a hierarchical fuse module (HFM), which contains an adaptive dense unit (ADU) and an enhancement unit (EU), to effectively utilize local hierarchical information and encourage layer-wise feature reuse. Further, we introduce the channel and spatial attention mechanisms to emphasize informative details. In addition, we propose the multi-supervised reconstruction (MSR) strategy, which amplifies feature maps at different levels of network to exploit global hierarchical information and recover high-quality image. Experimental results on the benchmark datasets prove that the proposed network performs superior to the state-of-the-art methods both quantitatively and qualitatively. Sumei Li |
ICIP | 2 |
| 2020 | Channel Shuffle Reconstruction Network for Image Compressive SensingabstractWith the aim of improving the reconstruction quality for image compressive sensing, we propose a channel shuffle reconstruction network (CSRNet) by jointly optimize the sampling and the inverse reconstruction processes. Firstly, we build an initial reconstruction sub-network (IRSN) to adaptively learn the measurement matrix and generate a preliminary reconstructed image. Then, a deep channel shuffle sub-network (CSSN) is added to further improve the image quality. Specially, we combine the merits of the inverted residual structure with channel shuffle operation to propose an efficient channel shuffle block in CSSN. The inverted residual structure endows the network with more powerful feature extraction ability. The channel shuffle operation further promotes information interaction. Besides, we exploit the multi-scale convolution block to make better use of feature information at different scales. Experiments on the benchmark datasets demonstrate that the proposed network outperforms previous state-of-the-art algorithms with a comparable time complexity. Sumei Li, Renhe Liu |
ICIP | 2 |
| 2020 | Stereo Image Quality Assessment Considering the Asymmetry of Statistical Information in Early Visual PathwayabstractThe design of stereo image quality assessment (SIQA) methods cannot be well based on the biological theory of human vision, so the performance of many SIQA methods cannot achieve good consistency with the subjective perception. The research on the visual system tends to the dorsal and ventral pathways, which ignores the information asymmetry in the early visual pathways. It is worth noting that the ON and OFF receptive fields in retinal ganglion cells (RGCs) respond asymmetrically to the statistical features of images. Inspired by this, we propose a SIQA method based on monocular and binocular visual features, which takes into account the asymmetry of local contrast bright and dark features in early visual pathways. First, this paper extracts the response maps of ON and OFF cell in RGCs to left and right views respectively. And then the different information fusion modes of visual cortex are used to fuse the response maps information of left and right views. Final, monocular and binocular features were extracted and sent to support vector regression (SVR) for quality regression. Experimental results show that the proposed method is superior to several mainstream SIQA metrics on two publicly available databases. Yongli Chang, Sumei Li |
VCIP | 2 |
| 2020 | No-Reference Stereoscopic Image Quality Assessment Considering Multi-loss ConstraintsabstractIn this paper, a three-channel convolutional neural network (CNN) constrained by multiple loss functions is designed for stereoscopic image quality assessment (SIQA). Given that both monocular and binocular information are crucial for SIQA, we take the patches of left images, right images and difference images as the inputs of the three channels respectively. Since using the ground truth as the labels of image patches cannot accurately characterize their quality, we propose to individually label each image patch to preserve the quality difference among different regions and views. Moreover, the multi-loss structure is adopted in the proposed method to consider both local features and global features simultaneously, which can constrain the feature learning from multiple perspectives. And the additional adaptive loss weights make the multi-loss network more flexible and universal. The experimental results show that the proposed method is superior to other existing SIQA methods with state-of-the-art performance. Yongtian Han, Sumei Li, Guanghui Yue 0001, Yongli Chang |
VCIP | 2 |
| 2020 | A Weighted Mean Absolute Error Metric for Image Quality AssessmentabstractPixel-wise image quality assessment (IQA) algorithms, such as mean square error (MSE), mean absolute error (MAE) and peak signal-to-noise ratio (PSNR) correlate well with perceptual quality when dealing with images sharing the same distortion type but not well when processing images in different distortion types, which is inconsistent with human visual system (HVS). Although a large number of metrics based on image error has been proposed, there are still difficulties and limitations. To solve this problem, a full reference image quality assessment (FR-IQA) method based on MAE is proposed in this paper. The metric divides the image error (difference between distorted image and reference image) map into smooth region and texture-edge region, calculates their mean values respectively, and then gives them different weights considering the masking effect. The key innovation of this paper is to propose a distortion significance measurement, which is a visual quality coefficient that can effectively indicate the influence of different distortion types on perceptual quality and unify them with HVS. The segmented image error maps are weighted by the distortion significance coefficient. The experimental results on four largest benchmark databases show that the most of the distortions are successfully evaluated and the results are consistent with HVS. Sihan Hao, Sumei Li |
VCIP | 2 |
| 2020 | No-Reference Stereoscopic Image Quality Assessment Based on Convolutional Neural Network with A Long-Term Feature FusionabstractWith the rapid development of three-dimensional (3D) technology, the effective stereoscopic image quality assessment (SIQA) methods are in great demand. Stereoscopic image contains depth information, making it much more challenging in exploring a reliable SIQA model that fits human visual system. In this paper, a no-reference SIQA method is proposed, which better simulates binocular fusion and binocular rivalry. The proposed method applies convolutional neural network to build a dual-channel model and achieve a long-term process of feature extraction, fusion, and processing. What's more, both high and low frequency information are used effectively. Experimental results demonstrate that the proposed model outperforms the state-of-the-art no-reference SIQA methods and has a promising generalization ability. Sumei Li |
VCIP | 1 |
| 2020 | No-Reference Stereoscopic Image Quality Assessment Based On Visual Attention MechanismabstractIn this paper, we proposed an optimized model based on the visual attention mechanism(VAM) for no-reference stereoscopic image quality assessment (SIQA). A CNN model is designed based on dual attention mechanism (DAM), which includes channel attention mechanism and spatial attention mechanism. The channel attention mechanism can give high weight to the features with large contribution to final quality, and small weight to features with low contribution. The spatial attention mechanism considers the inner region of a feature, and different areas are assigned different weights according to the importance of the region within the feature. In addition, data selection strategy is designed for CNN model. According to VAM, visual saliency is applied to guide data selection, and a certain proportion of saliency patches are employed to fine tune the network. The same operation is performed on the test set, which can remove data redundancy and improve algorithm performance. Experimental results on two public databases show that the proposed model is superior to the state-of-the-art SIQA methods. Cross-database validation shows high generalization ability and high effectiveness of our model. Sumei Li, Yongli Chang |
VCIP | 1 |
| 2019 | Rethinking Temporal Structure Modeling Method for Temporal Action LocalizationabstractTemporal action localization in untrimmed videos is an important but difficult task. Difficulties are encountered in the application of existing methods when modeling temporal structures of videos. In the present study, we developed a novel method, referred to as Gemini Network, for effective modeling of temporal structures and achieving high-performance temporal action localization. The significant improvements afforded by the proposed method are attributable to three major factors. First, the developed network utilizes two sub-nets for effective modeling of temporal structures. Second, three parallel feature extraction pipelines are used to prevent interference between the extractions of different stage features. Third, the proposed method utilizes auxiliary supervision, with the auxiliary classifier losses affording additional constraints for improving the modeling capability of the network. As a demonstration of its effectiveness, the Gemini Network was used to achieve state-of-the-art temporal action localization performance on two challenging datasets, namely, THUMOS14 and ActivityNet. Jianxing Yang, Yuan Zhou 0006, Sumei Li |
ICIP | 4 |
| 2019 | An End-to-End Multi-Scale Residual Reconstruction Network for Image Compressive SensingabstractRecently, deep-learning based reconstruction methods have been proposed to improve recovery performance of compressive sensed image and overcome expensive time complexity drawbacks of iteration-based traditional algorithms. In this paper, we propose an end-to-end multi-scale residual convolutional neural network (CNN), dubbed MSRNet, to simulate image compressive sensing (CS) and inverse reconstruction process in real situation. In the reconstruction stage of MSRNet, we apply three parallel channels with different convolution kernel sizes to exploit different-scale feature information. Besides, residual learning is introduced to accelerate training process and enhance prediction accuracy of network. Moreover, different from generating CS measurements by random measurement matrix in previous methods, we integrate compressive sample process into MSRNet, which means measurement matrix can be adaptively learned by training the network. Experiments on benchmark datasets show our method outperforms other state-of-the-art algorithms by large margins and set a new level for CS reconstruction with competitive time complexity. Renhe Liu, Sumei Li, Chunping Hou |
ICIP | 2 |
| 2019 | Fast and Lightweight Image Super-Resolution Based on Dense Residuals Two-Channel NetworkabstractThe existing advanced super-resolution methods with deepening or widening network demand high computational resources and memory consumption. It is difficult to directly apply them in practice. Therefore, we propose a fast and lightweight two-channel end-to-end network with fewer parameters and low computational complexity in this paper. The shallow channel mainly restores the general outline of the image, while the deep channel mainly learns the high-frequency texture information. And the deep channel combines the dense block and residual connection. The dense block increases data flow of network, while the residual connection reduces the number of parameters and speeds up the convergence of network. Moreover, we propose an enhanced network using group convolution, which significantly reduces the parameters and computational complexity with slight performance loss. Our model has many fewer parameters and operations, and it is evaluated on different datasets, outperforming the current representative methods in accuracy and run time. Yonglian Shi, Sumei Li |
ICIP | 2 |
| 2019 | No-Reference Stereoscopic Image Quality Assessment Based on Local to Global Feature RegressionabstractIn this paper, we propose a two-channel deep convolutional neural network (DCNN) through local to global regression for no-reference (NR) stereoscopic images quality assessment (SIQA). Firstly, in most deep learning based methods, they use the given subjective mean opinion score (MOS) or the differential MOS (DMOS) value to adjust the network parameters. But it is unreasonable, especially for the asymmetrical distortion stereoscopic image. To alleviate the problem, we propose to use feature similarity index (FSIM) to provide pseudo labels for the left and right view respectively, named local regression, so that the left and right channel are trained better. Then, we use the given DMOS to finetune the locally trained model parameters, named global regression. So we achieve an end-to-end network to measure the stereoscopic image quality. The experimental results show that the proposed method is superior to other existing SIQA methods. Sumei Li, Jianwei Xue, Yongtian Han |
ICME | 1 |
| 2019 | Stereoscopic Image Quality Assessment Weighted Guidance by Disparity Map Using Convolutional Neural NetworkabstractIn this paper, we propose a new two-column dense Convolutional Neural Network (CNN) for stereoscopic image quality assessment. The input of one column is the cyclopean image which conforms to the binocular combination and rival mechanism in our brain. The input of other column is the disparity map which provides some compensation information for the cyclopean image. More importantly, we employ the features of disparity map to guide and weight the feature maps obtained from the cyclopean image, which is implemented by modifying the structure of Squeeze and Excitation block. This weighting strategy recalibrates the importance of feature maps extracted from cyclopean image. At the end of CNN, we combine the outputs from the two-column through 'Concat', and then process them to get the final quality score of the stereoscopic image. Experimental results demonstrate that the proposed method can achieve high consistent alignment with subjective assessment. Yixiu Ding, Sumei Li, Yongli Chang |
VCIP | 2 |
| 2019 | Stereo Image Quality Assessment Based on Sparse Binocular Fusion Convolution Neural NetworkabstractIn this paper, a sparse binocular fusion convolution neural network is proposed to evaluate the quality of stereo image. In order to simulate the long-term fusion and processing of the left and right views in the brain visual pathway, the network combines the two views of the stereo image four times, and the information processing is carried out through convolution together with the fusion operation. In addition, in order to overcome the computational-intensive and memory-intensive problems of convolution neural networks, a structural sparsity learning (SSL) method is used to regularize the proposed convolution neural network. The experimental results demonstrate that our proposed method performs effectively and efficiently. And the proposed method can achieve 2.0× speedups on LIVE I database and 2.3× speedup on LIVE II database on the basis of improved performance. Sumei Li |
VCIP | 1 |
| 2019 | No-Reference Stereoscopic Image Quality Assessment Based On Shuffle-Convolutional Neural NetworkabstractWith the development of stereoscopic imaging technology, stereoscopic image quality assessment (SIQA) has been gaining great attention. In this work, to find a better SIQA method conforming to the perceptual characteristics of our brain, we propose a two-channel convolutional neural network (CNN) based on shuffle unit, which is called SCNN, for no-reference SIQA. The shuffle unit is used to mix up the features extracted from the left and right views to complete information communication between the two views. Different from other SIQA methods, the four shuffle units among proposed model achieve the multiple binocular fusions while processing the left and right views. Moreover, the Shuffle v2 block before the global pooling layer further improves the accuracy of SCNN. In addition, it is worth noting that we employ decorrelated batch normalization (DBN) to obtain the better generalization ability. Experimental results demonstrate that the proposed model outperforms the state-of-the-art no-reference SIQA methods. Sumei Li, Chunping Hou |
VCIP | 1 |
| 2019 | A Progressive Network Based on Residual Multi-scale Aggregation for Image Super-ResolutionabstractIn recent years, with the rapid development of deep learning, single image super-resolution based on convolution neural network has achieved extensive research. However, most CNN-based method has difficulty in training and obtaining high quality images for large scale factors. To address these issues, we propose a network, which reconstructs HR images at large factors by progressively performing 2× SR on the input from the previous level. At each level, cascaded residual multi-scale aggregation blocks are used. The U-residual unit in it makes network simplifier and training easier without performance degradation. The multi-scale dilated unit in it provides more comprehensive information for image reconstruction. Before upsampling, the channel attention mechanism is adopted to recalibrate features. We train the network with two-stage training strategy which could accelerate the convergence and achieve better performance. Experiment results show that our proposed method is superior to the state-of-the-art methods on most datasets, especially on Urban100. Sumei Li |
VCIP | 2 |
| 2019 | Stereoscopic Video Quality Assessment Based on The Two-step-training Binocular Fusion NetworkabstractIn this paper, we propose a novel binocular fusion network for stereoscopic video quality assessment (SVQA). In this network, we construct a long-term fusion, competition, and processing process by simulating the long-term complex process of the whole visual pathway. And we employ a two-step-training strategy for this network, which solves the problem that the network is difficult to fit caused by using the same value to label the different quality regions and views of the same stereoscopic video. In the first step, we use the computed quality scores of different patches to train the local network, namely local regression. And then the global regression is performed by using MOS value based on the first step trained model. Besides, considering temporal information, we take spatiotemporal saliency feature flows as the inputs of the proposed network. The proposed method is tested on public stereoscopic video databases, and results show that our method outperforms any other methods. Sumei Li, Jianwei Xue, Yixiu Ding, Guanghui Yue 0001 |
VCIP | 2 |
| 2019 | No-Reference Stereoscopic Image Quality Assessment Based on Dilation ConvolutionabstractOver the years, with the popularization of 3D technology, the demands of accurate and efficient 3D image quality evaluation (SIQA) methods are increasing constantly. Due to the wide application of CNN, CNN-based SIQA methods emerge one after another. However, current methods only consider a single scale or resolution, and some CNN-based methods directly take left and right views as an input of the network ignoring the visual fusion mechanism. In this work, a multi-scale no-reference SIQA method is proposed based on dilation convolution neural network (DCNN). Different from other CNN-based SIQA methods, the proposed one uses dilation convolution to imitate different scale of information processing fields in the human brain. Instead of left or right image, the cyclopean image generated by a new method is used as the input of the network. Moreover, the proposed multi-scale unit significantly can reduce computational parameters and computational complexity. Experimental results on two public databases show that the proposed model is superior to the state-of-the-art no-reference SIQA methods. Sumei Li, Yongli Chang |
VCIP | 2 |
| 2019 | Adaptive Cyclopean Image-Based Stereoscopic Image-Quality Assessment Using Ensemble LearningabstractIn this paper, we proposed an effective 3-D image-quality assessment method based on an adaptive cyclopean image by using ensemble learning. Our cyclopean image is not only suitable for a symmetrical distortion image, but also especially suitable for an asymmetrical distortion image. This adaptivity of our cyclopean image can be attributed to the consideration of gain control and gain enhancement in a binocular rivalry visual mechanism. In addition, we use a salient map to modify our cyclopean image to let the salient area of our cyclopean become more attractive. As a result, we can get better results. To remove redundant information out from our cyclopean, the sparse representation is applied to extract essential features. Finally, to get better regression accuracy on extracted feature, we use ensemble learning to get the final quality score of a stereoscopic image. The ensemble learner can improve the regression accuracy by 2% than a single learner. Experimental results show that the proposed algorithm outperforms the state-of-the-art methods on two publicly available stereoscopic image-quality assessment databases LIVE I and LIVE II. Sumei Li, Yongli Chang |
IEEE Trans. Multim. | 1 |
| 2018 | Cyclopean Image Based Stereoscopic Image Quality Assessment by Using Sparse Representationabstract3D image quality assessment confronts more difficulties than 2D image quality assessment. In this paper a 3D image quality assessment metric based on sparse representation was proposed. The contributions of the proposed method mainly include the following points: a color cyclopean image is used to better simulate the process of image processing in human brain, which is also very suitable for evaluating the quality of asymmetric distortion image. Meanwhile, for during sparse reconstruction some important information will be lost, we use the corresponding color cyclopean image to do compensation before feature extracting. And the paper creatively extracts spatial and spectral entropy feature of the distortion color cyclopean image and the corresponding reconstruction cyclopean image, respectively. Finally, we uses SVR to evaluate the quality of stereoscopic image. Experimental results show that the proposed method is very much in line with human visual perception. Yongli Chang, Sumei Li, Chunping Hou |
ICIP | 2 |
| 2018 | Multiple Residual Learning Network for Single Image Super-ResolutionabstractDeep residual convolutional neural network (CNN) has recently achieved great success in image super-resolution (SR). Because residual learning accelerates convergence rate and eases the difficulty for reconstructing high-resolution (HR) image, these CNN models can achieve higher peak signal to noise ratio (PSNR) values with lower training cost. However, residual image used in present residual network still contains much high frequency information, which increases learning burden and limits learning ability of residual network. Moreover, training a very deep network faces many obstacles and costs too much time. In this paper, we propose a multiple residual learning network (MRLN), which not only further simplifies information complexity of residual image and improves the accuracy of residual network, but also obviously reduces time cost for training a very deep CNN. In MRLN, we use a shallow network formed by 30-layer convolutional layers as basic model and train it for multiple times. The output of previous basic model is used as the HR input of the next one. In this way, an extremely large CNN is converted into a series connection of shallow networks. Fig. 1 shows PSNR of recent state-of-the-art CNN models for scale factor 2 on Set5, our method performs better than other methods and set a new level for SR. Renhe Liu, Sumei Li, Chunping Hou, Guoqing Lei |
VCIP | 2 |
| 2018 | A two-channel convolutional neural network for image super-resolution
Sumei Li, Ru Fan, Guoqing Lei, Guanghui Yue 0001, Chunping Hou |
Neurocomputing | 1 |
| 2017 | Shallow and deep convolutional networks for image super-resolutionabstractA shallow and deep convolutional neural network is presented for the single-image super-resolution (SISR). The proposed method doesn't need hand-designed procedures, directly learning an end-to-end mapping between low-resolution (LR) and high-resolution (HR) images. The upsampling of the network by deconvolution leads to much more efficient and effective training, reducing the computational complexity of the overall SR operation. However, most existing methods based on CNNs for super resolution need preprocessing like bicubic interpolating LR images to the size of HR images. This method can restore more details by multi-scale manner, and has strong adaptability whether on images or videos. Our model is evaluated on different datasets, outperforming the existing methods in accuracy and visual impression. Ru Fan, Sumei Li, Guoqing Lei, Guanghui Yue 0001 |
ICIP | 2 |