EDBT 2026 Demo / reviewers in the wild / expert
Xiongli Chai
dblp:237/1571
· DBLP profile ↗
31ranked-venue papers
7as first author
30since 2021 · last 2026
0000-0002-4245-5391ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 20 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BigCounter: A Bidirectional-Guided Network With Scene-Semantics-Driven Fusion for RGB-Thermal Crowd CountingabstractAccurate crowd counting has become increasingly essential for public safety management and Internet of Video Things (IoVT) applications, driven by rapid population growth and urbanization. However, RGB-Thermal (RGB-T) crowd counting remains challenging due to poor recognition of small targets and degraded performance under extreme conditions such as low-light environments. To address these issues, we propose a bidirectional-guided network with scene-semantics driven fusion for RGB-T crowd counting (BigCounter) that enhances robustness and generalization in complex scenes. BigCounter comprises of three parallel branches: a primary branch, a dynamic illumination auxiliary enhancement branch (DIAEB), and a high-resolution auxiliary enhancement branch (HAEB), which respectively improve robustness under illumination variations and accuracy for small target detection. Moreover, a cross-layer scene-driven fusion module (CLSFM) and a cross-modal semantic-driven fusion module (CMSFM) are designed to strengthen structural consistency and explore semantic complementarity between modalities. Through multi-branch collaboration and semantic-aware fusion, BigCounter significantly enhances feature representation. Extensive experiments on two benchmark RGB-T datasets demonstrate that BigCounter achieves superior accuracy and generalization compared with state-of-the-art methods. Xiaomin Fan, Feng Shao 0001, Baoyang Mu, Xiongli Chai, Zhongjie Zhu, Zhiyi Mo |
IEEE Internet Things J. | 4 |
| 2026 | PU-TransMamba: A Hybrid Point Cloud Upsampling Framework With Detail-Aware Transformer and Spatially Coherent MambaabstractHigh-quality 3D point clouds are essential for high-fidelity perception in Internet of Things (IoT)-enabled intelligent systems. While Point Cloud Upsampling (PCU) is widely used to mitigate data sparsity, existing methods often struggle to balance the preservation of fine-grained local details with the maintenance of global topological consistency. Transformer-based approaches frequently suffer from excessive computational overhead and high-frequency detail loss, whereas emerging state space models like Mamba, despite their efficiency, inevitably sacrifice spatial coherence due to the 1D serialization of irregular 3D points. To address these critical bottlenecks, we introduce PU-TransMamba, a hybrid framework that synergistically leverages a Detail-Aware Transformer and a Spatially-Coherent Mamba. Each component is designed to resolve specific PCU limitations: a Complexity-Aware Bilateral Decoder is developed to adaptively recover sharp geometric edges by processing features across dual domains, while a Sequence-Aligned Mamba Encoder utilizes multiple spatial curvature descriptors to compensate for the spatial information loss inherent in serialization. Additionally, a Global Geometry Injector and a Local Neighbor Injector are designed to ensure structural integrity by infusing holistic skeletal priors and neighborhood context, respectively. To minimize feature discrepancies between the hybrid branches, we also propose a Self-Distillation Loss. Extensive experiments on five benchmark datasets demonstrate that PU-TransMamba outperforms state-of-the-art methods in both reconstruction accuracy and computational scalability. The results confirm its ability to recover intricate geometries, indicating significant potential for IoT-driven 3D perception and communication systems. Feng Shao 0001, Xiongli Chai, Hangwei Chen, Zhongjie Zhu, Zhiyi Mo |
IEEE Internet Things J. | 3 |
| 2026 | SGNet: A Structure-Guided Lightweight Network for VDT Salient Object DetectionabstractVisual-Depth-Thermal (VDT) salient object detection (SOD) aims to jointly exploit RGB, depth and thermal cues to segment the most visually significant regions. However, most existing VDT SOD models are heavy in parameters and computational cost, limiting their deployment on real-world and edge devices. To tackle this, we propose SGNet, a structure-guided lightweight network for efficient VDT SOD. Specifically, we design a lightweight Tri-modal Fusion Module (TFM) to integrate three modalities at the semantic level, and a Shared Structure Extraction Module (SSEM) to extract common structural information from depth and thermal modalities. A Structure Refine Module (SRM) further injects the extracted structure into the deepest semantic features, while a Multiscale Feature Refinement Module (MFRM) progressively decodes multi-level features under deep supervision to produce saliency maps with clear boundaries. Benefiting from these modules, SGNet achieves competitive performance on the VDT2048 benchmark with only 5.51 M parameters and a real-time speed of 120 FPS at 320 × 320 resolution, surpassing state-of-the-art methods while remaining deployment-friendly. Huizhi Wang, Feng Shao 0001, Xuebin Wei, Xiongli Chai, Hangwei Chen, Zhongjie Zhu |
IEEE Internet Things J. | 4 |
| 2026 | Joint luminance-chrominance learning for quality assessment of low-light image enhancement
Tuxin Guan, Qiuping Jiang, Xiongli Chai |
Pattern Recognit. | 3 |
| 2026 | USformer: A U-Shaped Structure Transformer for RGB-Thermal Semantic Segmentation and Traffic Scene UnderstandingabstractRecent advancements in multimodal approaches, particularly RGB-thermal (RGB-T) segmentation, have significantly promote the development of Intelligent Transportation Systems (ITS). However, existing methods still encounter challenges related to modality discrepancy and the effective integration of multi-scale features. To address these issues, we propose the U-shaped Structure Transformer (USformer) for RGB-T semantic segmentation. We improve the feature flow of existing methods by designing a novel U-shaped encoding network that integrates inter-layer fusion and cross-modal fusion. Specifically, our method introduces an inter-layer interaction mechanism that facilitates the iterative fusion of high-level semantic and low-level detail features. For each layer, our fusion process is divided into two stages: the Cross-Modal and -Scale Auxiliary (CMSA) module enforces distribution alignment across modalities and scales, while the Cross-Attention Feature Merger (CAFM) allows each modality to refine its own feature selection by employing a multi-head cross-attention mechanism. These modules effectively adapt and integrate well-established attention designs into our U-shaped encoding architecture, thereby achieving efficient multi-modal feature alignment and fusion. Finally, we utilize the Mask2Former decoder to aggregate the fused features from multiple layers and improve the segmentation across various object sizes and complex scenes. Extensive experiments on four RGB-T datasets demonstrate that our proposed USformer achieves state-of-the-art performance. Feng Shao 0001, Baoyang Mu, Xiongli Chai, Qiuping Jiang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | MDGINet: Multi-frequency Dynamic Guidance and Interaction network for image denoising
Feng Shao 0001, Hangwei Chen, Xiongli Chai, Qiuping Jiang |
Knowl. Based Syst. | 4 |
| 2025 | Art Comes From Life: Artistic Image Aesthetics Assessment via Attribute Knowledge AmalgamationabstractAssessing the aesthetic quality and visual appeal of artworks has become one of the hotspots in current research. The existing artistic image aesthetics assessment (AIAA) methods directly learn aesthetics from images, while ignoring the impact of variations in visual attributes on human aesthetic perception, which hampers the further development of AIAA. To address this issue, this paper presents a new AIAA method based on attribute knowledge amalgamation, named AKA-Net. Specifically, we initially learn common attribute aesthetic rules (e.g., composition and color) through pre-training on natural aesthetic images. Then, we devise a multi-model amalgamation strategy based on contrastive learning to transfer different types of prior attribute knowledge into a single target model, enabling flexible and efficient aesthetic prediction. Finally, an attribute-aware feature enhancement module (AFEM) is introduced to better establish the relationship between aesthetic quality and attribute knowledge. Experimental results on three public benchmark AIAA databases demonstrate that the proposed AKA-Net outperforms the state-of-the-art AIAA metrics. Hangwei Chen, Feng Shao 0001, Xiongli Chai, Baoyang Mu, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Adversarial Robust Salient Object Detection in Optical Remote Sensing Images With Implicit Feature EnhancementabstractDeep neural networks (DNNs) have achieved significant progress in optical remote sensing images salient object detection (ORSI-SOD) and are widely applied to various remote sensing image analysis tasks. However, few SOD models demonstrate robust performance under adversarial perturbations, which ultimately leads to a decline in detection accuracy. Moreover, most existing defense methods inject fixed Gaussian noise globally into the image. Although such approaches are easy to implement, they have several limitations in inaccurate uncertainty estimation and neglecting the unique characteristics of local salient regions. Furthermore, existing adversarial defense research rarely addresses the challenges specific to the ORSI-SOD task, leaving a gap in effective defense strategies. To tackle these issues, we propose a novel defense method, which enhances the adversarial robustness of ORSI-SOD models through implicit feature enhancement. The algorithm first proposes a two-stage strategy of reverse local noise search and forward global noise optimization, enhancing generalization ability by implicitly enhancing features to better simulate network uncertainty. Then, the algorithm proposes a global-guided texture information enhancement (GTIE) module for low-level features and a global-guided semantics information enhancement (GSIE) module for high-level features, focusing on strengthening low-level texture information and enhancing the model’s understanding of high-level contextual semantic features, respectively. This dual-module design effectively weakens the impact of adversarial noise, significantly improving the robustness and accuracy of object detection. Extensive experiments on three ORSI-SOD datasets demonstrate that our defense strategy better estimates the uncertainty, resulting in an average performance improvement of 23.2% in$F_{\beta } ^{\mathrm { max}}$and 34.1% in$E_{\xi } ^{\mathrm { max}}$across six ORSI-SOD models under five different adversarial attack methods. Our code will be released in the public repository athttps://github.com/kexi0714/IFe. Feng Shao 0001, Xiangchao Meng, Hangwei Chen, Xiongli Chai, Zhiyi Mo |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Cross-Projection Distilling Knowledge for Omnidirectional Image Quality Assessment
Huixin Hu, Feng Shao 0001, Hangwei Chen, Xiongli Chai, Qiuping Jiang |
IEEE Trans. Multim. | 4 |
| 2024 | Plain-PCQA: No-Reference Point Cloud Quality Assessment by Analysis of Plain Visual and Geometrical ComponentsabstractIn reviewing the research progress in Point Cloud Quality Assessment (PCQA), two main pathways have emerged, i.e., 2D projections and 3D point descriptors. The former primarily focuses on visual information, while the latter concentrates on crucial geometrical information in three-dimensional space. However, the current studies lack a thorough investigation of the impact of visual components and seldom pay special attention to plane-point fusion strategies. To comprehensively represent features and effectively tackle various types of impairments, we propose an end-to-end learning paradigm, only considering plain visual and geometrical factors called Plain-PCQA, for quantitatively evaluating objective metrics of 3D dense point clouds associated with human perception. Firstly, we explore a sophisticated preprocessing technique. The entire point clouds are packaged into six projections by moving virtual cameras, which can conveniently increase the visual samples during the training stage. Given the high resolution of the projected image, we have opted for a relatively lightweight network, namely ResNet-18, as the backbone to enable higher resolution input data. Five cropped patches from the projected image are collectively fed into this network. In light of the presence of some invalid information in the projections, a mask weight is devised to calculate the significance of each patch based on its effective informational content. Secondly, dual neural networks, comprising of a No-Reference (NR) branch and a Degraded-Reference (DR) branch, are designed with fundamental visual components to provide quantitative quality metrics. Specifically, the NR branch utilizes the feature output of each block in the Vision Transformer (ViT) model to obtain long-range low-level and high-level visual NR quality. The DR branch employs KLT (Karhunen-Loève Transform) to acquire the principal component information of an image as the macro-structural image, and then feeds the difference between input images and macro-structural images into a network for DR quality extraction. Thirdly, a Plane-Point Interaction Transformer (P2IT) is presented by incorporating texture and semantic features in 2D projections and geometrical features in 3D spaces to characterize the complete features with a connected 2D-3D feature representation. With these elaborately designed deep features, the proposed model can achieve competitive performances relying solely on plain visual and geometrical components. The experimental results demonstrate the potential of the proposed approach in multiple representative databases, which surpasses existing state-of-the-art methods significantly. Xiongli Chai, Feng Shao 0001, Baoyang Mu, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Blind Quality Evaluator of Light Field Images by Group-Based Representations and Multiple Plane-Oriented Perceptual CharacteristicsabstractDue to the emergency of multi-view cameras and commercial Light Field (LF) cameras, the demand of high-performance LF quality evaluator is of great significance for guiding LF acquisition, processing and application and further promoting the visual perceived quality of LF visualizations. However, LF Images (LFIs), as high-dimensional data, suffer from various quality degradations not only in the spatial domain but also in the angular domain. Therefore, it is of great challenge to predict LF quality accurately. An effective LF evaluator should be able to represent these heterogeneous artifacts. In this paper, we provide a novel No-Reference LF Quality Assessment Evaluator (NR LF-QAE) to tackle this problem. Firstly, to measure angular consistency among viewports, we utilize group-based representations to character information similarity of aligned view stacks. Secondly, to better describe the texture information of LFIs, unifying spatial-angular texture statistic measurement is performed via Local Binary Patterns from Three Orthogonal Planes (LBP-TOP). Thirdly, we design 3D Log-Gabor filters to extract LF global structure information in Sub-Aperture Images (SAIs) as spatial feature characterizations and 2D Log-Gabor filters are adopted to characterize ray direction/depth information in Epipolar Plane Images (EPIs) as angular feature characterizations. By comprehensive LF information analyses in angular consistency and spatial-angular feature extraction with texture and structure descriptors, experimental results demonstrate the superiority of the proposed NR LF-QAE over the state-of-the-art comparative models in predicting the quality of LFIs on three available benchmark databases. The code will be released athttps://github.com/zerosola/NR-LF-QAE. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xuejin Wang, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 1 |
| 2024 | SCFANet: Semantics and Context Feature Aggregation Network for 360° Salient Object DetectionabstractHow to solve the problem of geometric distortion is the key for salient object detection (SOD) in 360° omnidirectional images. Most of the current methods integrate global and local visual cues through the fusion of the 360° equirectangular images and corresponding 360° cube-map images. The fusion in a single level cannot effectively utilize the information between the 360° equirectangular images and corresponding 360° cube-map images. In this work, we innovatively propose a semantics and context feature aggregation network (SCFANet) by fully exploring the interactivity between the two projection data. Specifically, we use Vision Transformer (ViT) to capture global visual cues for 360° equirectangular images and Convolutional Neural Network (CNN) to capture local visual cues for 360° cube-map images. To achieve effective fusion of the two projection data, we design a semantic guidance module (SGM), in which semantic features are used to guide the information fusion of the 360° equirectangular images and corresponding 360° cube-map images at each level. Then, a context fusion module (CFM) containing one local input and two context inputs is designed to integrate multi-scale features, where the local input extracts its own multi-scale information, and the context inputs complements their fine details and location information. Finally, we use feature aggregation and refinement module (FARM) to aggregate semantics and context feature and adopt a deep supervision strategy for training. Extensive experiments on two public 360° datasets show that our SCFANet exhibits competitive performance compared to other state-of-the-art (SOTA) 360° salient object detection models. Feng Shao 0001, Xiongli Chai, Yo-Sung Ho |
IEEE Trans. Multim. | 4 |
| 2024 | Collaborative Learning and Style-Adaptive Pooling Network for Perceptual Evaluation of Arbitrary Style TransferabstractAlthough the research of arbitrary style transfer (AST) has achieved great progress in recent years, few studies pay special attention to the perceptual evaluation of AST images that are usually influenced by complicated factors, such as structure-preserving, style similarity, and overall vision (OV). Existing methods rely on elaborately designed hand-crafted features to obtain quality factors and apply a rough pooling strategy to evaluate the final quality. However, the importance weights between the factors and the final quality will lead to unsatisfactory performances by simple quality pooling. In this article, we propose a learnable network, named collaborative learning and style-adaptive pooling network (CLSAP-Net) to better address this issue. The CLSAP-Net contains three parts, i.e., content preservation estimation network (CPE-Net), style resemblance estimation network (SRE-Net), and OV target network (OVT-Net). Specifically, CPE-Net and SRE-Net use the self-attention mechanism and a joint regression strategy to generate reliable quality factors for fusion and weighting vectors for manipulating the importance weights. Then, grounded on the observation that style type can influence human judgment of the importance of different factors, our OVT-Net utilizes a novel style-adaptive pooling strategy guiding the importance weights of factors to collaboratively learn the final quality based on the trained CPE-Net and SRE-Net parameters. In our model, the quality pooling process can be conducted in a self-adaptive manner because the weights are generated after understanding the style type. The effectiveness and robustness of the proposed CLSAP-Net are well validated by extensive experiments on the existing AST image quality assessment (IQA) databases. Our code will be released at https://github.com/Hangwei-Chen/CLSAP-Net. Hangwei Chen, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Multi-layer and Multi-scale feature aggregation for DIBR-Synthesized image quality assessment
Xuejin Wang, Xiongli Chai, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | TCCL-Net: Transformer-Convolution Collaborative Learning Network for Omnidirectional Image Super-ResolutionabstractAs virtual reality and metaverse become more and more popular, the Omnidirectional Image (OI) has attracted extreme attention due to its immersive display characteristics. However, users only watch a portion of the content in a specific viewport extracted from a panoramic view, which will lead to a problem of resolution mismatch that requires High-Resolution (HR) for clear near-eye displays in viewports. Hence, it is necessary to exploit a Super Resolution (SR) solution for reconstructing Low-Resolution (LR) OIs. Different from 2D SR methods, the variation of pixel distributions along latitudes is a critical factor in designing an Omnidirectional Image Super-Resolution (OISR) scheme. In this paper, we put forward a novel end-to-end network with a Transformer and Convolution Collaborative Learning Network (TCCL-Net) for OISR. Firstly, Swin Transformer blocks and residual convolution blocks are employed to extract long-range and short-range dependencies, thereby digging into more rich and heterogeneous features from these two branches. Secondly, to better fuse these two features, cross-guided enhanced attention mechanisms are designed for bidirectional information enhancement onto both channel and spatial features . Thirdly, to alleviate nonuniformly pixel distributions across latitudes, we add an absolute positional encoding into Swin Transformer to represent patch weights at different positions and propose a tile-based panoramic reconstruction module to super-resolve various bands with different pixel sampling characteristics across latitudes. Experimental results on two available benchmark datasets demonstrate the superiority of the proposed approach over the state-of-the-art method in achieving OISR task. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Hongwei Ying |
Knowl. Based Syst. | 1 |
| 2023 | Modality-Induced Transfer-Fusion Network for RGB-D and RGB-T Salient Object DetectionabstractThe ability of capturing the complementary information of multi-modality data is critical to the development of multi-modality salient object detection (SOD). Most of existing studies attempt to integrate multi-modality information through various fusion strategies. However, most of these methods ignore the inherent differences in multi-modality data, resulting in poor performance when dealing with some challenging scenarios. In this paper, we propose a novel Modality-Induced Transfer-Fusion Network (MITF-Net) for RGB-D and RGB-T SOD by fully exploring the complementarity in multi-modality data. Specifically, we first deploy a modality transfer fusion (MTF) module to bridge the semantic gap between single and multi-modality data, and then mine the cross-modality complementarity based on point-to-point structural similarity information. Then, we design a cycle-separated attention (CSA) module to optimize the cross-layer information recurrently, and measure the effectiveness of cross-layer features through point-wise convolution-based multi-scale channel attention. Furthermore, we refine the boundaries in the decoding stage to obtain high-quality saliency maps with sharp boundaries. Extensive experiments on 13 RGB-D and RGB-T SOD datasets show that the proposed MITF-Net achieves a competitive and excellent performance. Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Quality Evaluation of Arbitrary Style Transfer: Subjective Study and Objective MetricabstractArbitrary neural style transfer is a vital topic with great research value and wide industrial application, which strives to render the structure of one image using the style of another. Recent researches have devoted great efforts on the task of arbitrary style transfer (AST) for improving the stylization quality. However, there are very few explorations about the quality evaluation of AST images, even it can potentially guide the design of different algorithms. In this paper, we first construct a new AST images quality assessment database (AST-IQAD), which consists 150 content-style image pairs and the corresponding 1200 stylized images produced by eight typical AST algorithms. Then, a subjective study is conducted on our AST-IQAD database, which obtains the subjective rating scores of all stylized images on the three subjective evaluations, i.e., content preservation (CP), style resemblance (SR), and overall vision (OV). To quantitatively measure the quality of AST image, we propose a new sparse representation-based method, which computes the quality according to the sparse feature similarity. Experimental results on our AST-IQAD have demonstrated the superiority of the proposed method. The dataset and source code will be released athttps://github.com/Hangwei-Chen/AST-IQAD-SRQE Hangwei Chen, Feng Shao 0001, Xiongli Chai, Yuese Gu, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Viewport-Sphere-Branch Network for Blind Quality Assessment of Stitched 360° Omnidirectional ImagesabstractCompared with conventional images/videos, omnidirectional data records rich information with higher resolution and wider Field-of-View. Moreover, the stitching distortions introduced in the panoramic content generation process make the quality assessment task more challenging. Targeting at designing an accurate and fast stitched 360° omnidirectional image quality evaluator, we propose a Viewport-Sphere-Branch Network (VSBNet) via dual-branch quality estimation. Specifically, for the viewport quality estimation, we extract distorted viewports around the stitching seams and conduct distortion rectification through a progressively complementary network to obtain pseudo-reference viewports. The qualitative and quantitative experiments validate that pseudo-reference viewports are reliable. Then, the differences between distorted and pseudo-reference viewports are quantified through transformer architecture to obtain quality scores of viewports. The introduction of pseudo-reference viewports can effectively improve the performance of the viewport quality prediction branch. To establish general scenario awareness and accurately evaluate the immersive experience, we extract feature representation through deformable convolutions to eliminate 2D-to-Sphere intrinsic sampling distortions and use multilayer perceptron to predict score of the whole sphere. The final prediction score is obtained by aggregating the quality scores from viewport and sphere branches. We evaluate the proposed VSBNet on two benchmark databases and results demonstrate that the combination of two branches can obtain more accurate results. Overall, our method is superior to existing full reference and no reference models designed for conventional images and 360° omnidirectional images. Chongzhen Tian, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Perceptual Quality Assessment of Cartoon ImagesabstractIn the animation industry, automatically predicting the quality of cartoon images based on the inputs of general distortions and color change is an urgent task, while the existing no-reference (NR) methods usually measure the perceptual quality of the natural images. In this paper, based on the observation that structure and color are the main factors affecting cartoon images quality, we proposed a new NR quality prediction metric for cartoon images, which fully takes gradient and color information into account. The experimental results on our newly constructed NBU-CIQAD dataset with color change and other existing cartoon image dataset demonstrate that the proposed method significantly outperforms existing no-references methods for the task of cartoon image quality assessment. The database and code will be released athttps://github.com/1010075746/NBU-CIQAD. Hangwei Chen, Xiongli Chai, Feng Shao 0001, Xuejin Wang, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2023 | Composition-Guided Neural Network for Image Cropping Aesthetic AssessmentabstractHow to explore the interaction between image aesthetic rules and crops is the key to finding views with good composition. Besides, it is subjective to evaluate candidate crops, which mainly depends on aesthetic knowledge, but it is not an easy task for people without extensive photography experience. However, existing methods mostly find good views by extracting general aesthetic features of crops without fully exploring the aesthetic rules. Motivated by this, we innovatively propose a composition-guided image cropping aesthetic assessment network (CGICAANet) for efficiently finding good crops and optimizing the cropping operation. Specifically, we adopt a direct and comprehensive composition pattern module, which adaptively mines suitable compositions for the images and emphasizes the dominant position of visual elements to contribute to optimizing the best crops in an interpretable way. Moreover, we designed a multi-task loss function to train the model. Particularly, to explore the commonality between predicted crops and labels, the complete intersection-over-union loss is adopted thoroughly considering the overlap area, central point distance and the consistency of aspect ratios for crops concurrently. Therefore, the predicted best crop can preserve the visual elements and have better composition. Experimental results with lightweight MobileNetV2 and ShuffleNetV2 as backbone networks demonstrate that our method can obtain comparable or better performance in terms of efficiency and accuracy. Shijia Ni, Feng Shao 0001, Xiongli Chai, Hangwei Chen, Yo-Sung Ho |
IEEE Trans. Multim. | 3 |
| 2022 | M2OVQA: Multi-space signal characterization and multi-channel information aggregation for quality assessment of compressed omnidirectional videos
Xiongli Chai, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Deep network based stereoscopic image quality assessment via binocular summing and differencing
Jinbin Hu 0002, Xuejin Wang, Xiongli Chai, Feng Shao 0001, Qiuping Jiang |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | Monocular and Binocular Interactions Oriented Deformable Convolutional Networks for Blind Quality Assessment of Stereoscopic Omnidirectional ImagesabstractStereoscopic omnidirectional content, as a novel visual media, has drawn wide attention in recent years due to its ability in providing strong immersive experience. Since Stereoscopic Omnidirectional Images (SOIs) involve the properties from panoramic and stereoscopic visual perception, it is very challenging to establish an efficient and effective visual quality evaluation model for SOIs. To better measure the user’s experience in virtual reality, we put forward a novel deep learning framework to assess the quality of SOIs in this paper. Firstly, the deformable convolutions instead of standard convolutions are adopted to ensure the invariant receptive fields of convolutional kernels on Equi-Rectangular Projection (ERP). Secondly, according to the stereoscopic property, we use binocular-difference information and a coarse-to-fine mechanism to construct the binocular feature extraction network. Thirdly, a three-channel network involving left-view, right-view and binocular-difference channels is presented to simulate the process of monocular and binocular interactions, in which independent quality labels are provided for each channel to reflect the individual effect of monocular and binocular visions on the whole visual quality. Finally, experimental results on two available benchmark databases demonstrate the superiority of the proposed metric over the state-of-the-art blind quality assessment models in predicting the quality of SOIs. Moreover, our model is efficient in computational cost as the feature extraction is directly applied on ERP images. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | CGMDRNet: Cross-Guided Modality Difference Reduction Network for RGB-T Salient Object DetectionabstractHow to explore the interaction between the RGB and thermal modalities is the key success of the RGB-T saliency object detection (SOD). Most of the existing methods integrate multi-modality information by designing various fusion strategies. However, the modality gap between the RGB and thermal features will lead to unsatisfactory performances by simple feature concatenation. To solve this problem, we innovatively propose a cross-guided modality difference reduction network (CGMDRNet) to achieve intrinsic consistency feature fusion via reducing the modality differences. Specifically, we design a modality difference reduction (MDR) module, which is embedded in each layer of the backbone network. The module uses a cross-guided strategy to reduce the modality difference between the RGB and thermal features. Then, a cross-attention fusion (CAF) module is designed to fuse cross-modality features with small modality differences. In addition, we use a transformer-based feature enhancement (TFE) module to enhance the high-level feature representation that contributes more to performance. Finally, the high-level features guide the fusion of low-level features to obtain a saliency map with clear boundaries. Extensive experiments on three public RGB-T datasets show that the proposed CGMDRNet achieves competitive performance compared with state-of-the-art (SOTA) RGB-T SOD models. Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | VSOIQE: A Novel Viewport-Based Stitched 360° Omnidirectional Image Quality EvaluatorabstractWith the rapid development of virtual reality (VR), 360° omnidirectional images and videos have drawn wide attention. However, the quality assessment of 360° omnidirectional images is a challenging task, especially when the panoramic image contains multiple stitching distortions. We propose a viewport-based stitched 360° omnidirectional image quality evaluator (VSOIQE), by first extracting the features of salient and stitching viewports, and then inferring the overall perceptual quality via multiple linear regression (MLR). Comprehensive image attributes including edge, color, shape and information entropy are considered in the framework. Experimental results on two benchmark databases demonstrate the superiority of the proposed metric over both the state-of-the-art quality models designed for 2D images and the quality models developed for 360° omnidirectional images. Chongzhen Tian, Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | A Blind Full Resolution Assessment Method for Pansharpened Images Based on Multistream Collaborative LearningabstractPansharpening aims to fuse a high spatial resolution (HR) panchromatic (PAN) image and a low spatial resolution (LR) multispectral (MS) image to obtain an HR-MS image. However, due to the lack of the real HR-MS reference image, determining pansharpened image quality at full resolution has always been a contentious issue in the community. We propose a blind full resolution assessment method for pansharpened images based on multi-stream collaborative learning. The proposed method designs a Siamese framework to collaboratively learn the spatial, spectral, and overall quality of the fused image. The parameters of the feature extraction layer in the spatial and spectral evaluation models are frozen for the overall evaluation model, improving accuracy and convergence speed. The proposed method was comprehensively tested and verified based on a large-scale data set consisting of 13620 fused images obtained by six pansharpening methods with four different thematic data sets. Furthermore, a large-scale subjective evaluation data set in which each of the 13620 fused images was assessed by 28 participants, was utilized to comprehensively valid the proposed method. The experimental results demonstrated the superior performance of the proposed method to other state-of-the-art quality assessments. Kedi Bao, Xiangchao Meng, Xiongli Chai, Feng Shao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | List-Wise Rank Learning for Stereoscopic Image Retargeting Quality AssessmentabstractStereoscopic imageretargeting (SIR) techniques attempt to display stereoscopic images on stereoscopic devices of various resolutions and aspect ratios to provide the users with better viewing experience. However, new quality perceptual problems emerge in the retargeted stereoscopic images generated by current SIR operators are quite different from those in the retargeted 2D images. In this paper, we dedicate to exploring the perceptual quality-related factors (e.g., shape preservation, object preservation and visual comfort.) of retargeted stereoscopic images, and propose a novel quality evaluation metric for SIR to achieve a more consistent evaluation with 3D perception and image degradation mechanism in the SIR process. Moreover, image quality features and 3D perceptual features are integrated into one representation for an overall perceptual quality prediction using a list-wise ranking approach, which gives priority to the ranking among the SIR results generated from the same stereoscopic source. Experimental results demonstrate that the proposed method outperforms most quality models developed for retargeted 2D/stereoscopic images. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Xiongli Chai, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 4 |
| 2021 | Stitched image quality assessment based on local measurement errors and global statistical properties
Chongzhen Tian, Xiongli Chai, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Quality assessment for color correction-based stitched images via bi-directional matching
Xuejin Wang, Xiongli Chai, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Roundness-Preserving Warping for Aesthetic Enhancement-Based Stereoscopic Image EditingabstractImage editing is an effective solution to adapt contents for different applications. In this paper, we present a roundness-preserving warping model for stereoscopic image editing, in which energy constraints from image quality energy, aesthetics energy and depth adaptation energy are involved in the framework to solve the optimization. Specifically, to preserve object roundness during warping, the relationship between object's shape and disparity is established and is applied for depth adaptation. Different from the existing stereoscopic image editing methods, the main innovations of our method are to achieve a tradeoff in balancing information loss and reducing semantic distortion while providing a novel death adaptation model for recomposition and retargeting applications. Experimental results demonstrate the effectiveness of our method in enhancing the aesthetics of stereoscopic images. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | MSTGAR: Multioperator-Based Stereoscopic Thumbnail Generation With Arbitrary ResolutionabstractAt present, thumbnail generation for 2D images has been extensively studied, but the research in thumbnail generation for stereoscopic images is still relatively lacking. This paper presents a novel thumbnail generation technology for stereoscopic images based on multioperator with the following innovations: 1) The warping technique is used to retarget a stereopair into six-scale resolutions with different contexts, and the disparity is uniformly adjusted to a certain value based on just noticeable depth difference (JNDD) model, which overcomes the issues that 3D perception in stereoscopic thumbnail is uncontrollable and the sense of depth disappears in low-resolution stereoscopic images. 2) The six-scale images are cropped via cropping network, and are optimized to a target resolution based on the designed image visual representation energy. As a result, our method has better visual effect than state-of-the-art methods in generating thumbnail for stereoscopic display. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Multim. | 1 |