VLDB 2026 Research / reviewers in the wild / expert
Shishun Tian
dblp:201/7429
· DBLP profile ↗
46ranked-venue papers
11as first author
41since 2021 · last 2026
0000-0002-7616-8382ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 9 first-author · 25 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MLWAC: A Modular, Low-coupling Waypoint-Angular Coordinated Network for visual navigation in unstructured environments
Yongdong Guo, Muxin Liao, Shishun Tian, Wenbin Zou, Chen Xu 0004 |
Knowl. Based Syst. | 6 |
| 2026 | No-Reference Quality Assessment of 3D Models Represented in Neural Radiance Fields and 3D Gaussian SplattingsabstractThe continuous breakthroughs in 3D reconstruction and rendering technologies have enabled synthesized 3D models to achieve exceptional realism. In particular, Neural Radiance Fields (NeRF) and 3D Gaussian Splattings (3DGS) have gained significant attention due to their impressive ability to deliver high-quality 3D models. However, research on the quality assessment of NeRF/3DGS models remains an underexplored area, hindering the development of relevant generation, compression, and transmission algorithms. To fill this gap, in this paper, we propose the first no-reference quality assessment metric for NeRF/3DGS models rendered in Processed Video Sequences (PVS). Considering the uniqueness and diversity of distortions introduced in NeRF/3DGS models, the core idea behind the proposed metric is to extract universal quality-aware features that are generalizable across various distortions, rather than targeting a specific one. Specifically, inspired by the fact that spatial distortions typically alter the statistical distributions, we first measure the spatial fidelity of the rendered PVS by analyzing the spatial explicit statistics of textural variation and naturalness. Then, motivated by the ability of implicit energy composition changes to reflect temporal distortions, we propose to evaluate the temporal consistency of the rendered PVS through inter-frame discrepancy energy and multi-frame motion energy in the Singular Value Decomposition (SVD) domain. Finally, the spatial explicit statistics and temporal implicit energy are combined as perceptual features to evaluate the quality of NeRF/3DGS models via Support Vector Regression (SVR). Extensive experimental results on three representative databases demonstrate the superiority of the proposed metric in various aspects, such as predictive accuracy, performance stability, and cross-database generalizability. The source code will be publicly available at https://github.com/ZhengyuZhang96/PVS-3DMQA. Yuhang Zhang 0011, Tiantian Zeng, Shishun Tian, Lu Zhang 0037 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Prior-Guided Test Time Adaptation for Blind Image Quality AssessmentabstractCurrent blind image quality assessment (BIQA) models usually lack adaptability to the test data with distribution shifts to the training data. This inspires an investigation into test time adaptation (TTA) methods to address distribution shifts between training and test data. However, existing methods mainly focus on simple feature alignment strategies, which may lead to incorrect knowledge generalization. To this issue, we propose a prior-guided test time adaptation (PGTA-IQA) for blind image quality assessment. Concretely, we extract the quality prior knowledge from the pre-trained BIQA model through clustering. The extracted quality prior knowledge forms the foundation for subsequent optimizations. These optimizations are carried out from two complementary perspectives: inter-cluster and intra-cluster. From the inter-cluster perspective, we propose a confident rank learning approach which consists of a relative quality matrix (RQM) and a confidence filtering strategy (CFS) to generate the high-confident quality rankings. From the intra-cluster perspective, we propose a selective feature alignment approach by only aligning the closest neighboring samples within the same cluster to reduce the impact of noisy labels. The experimental results demonstrate the effectiveness of the proposed approaches. Shishun Tian, Fangjie Hou, Guanghui Yue 0001, Yuanhao Gong, Wenbin Zou, Ting Su 0004 |
ICME | 1 |
| 2025 | Semantic-Guided Residual Learning for the Quality Assessment of Enhanced ImagesabstractImage enhancement algorithms are essential for improving visual quality but often introduce new distortions, highlighting the need for reliable image quality assessment (IQA). However, existing IQA methods typically focus on semantic information or distortion-prone regions while ignoring their interactions, resulting in unsatisfactory performance. To address this issue, we propose to integrate semantic information with edge residual learning and design a semantic-guided residual learning IQA framework tailored for enhanced images across diverse scenarios. Specifically, the proposed framework utilizes a covariance-guided encoder to extract semantic information, which is then enhanced using a semantic refinement module. The refined semantic information is subsequently utilized to guide edge residual feature learning in the decoder. Extensive experiments on multiple tasks such as deraining, dehazing, and low-light enhancement demonstrate that our method outperforms state-of-the-art approaches. Shishun Tian, Zhiwei Lan, Ting Su 0004, Xia Li 0006, Lu Zhang 0037 |
ICME | 1 |
| 2025 | Learning Content-enhanced Tokens for Domain Generalized Semantic SegmentationabstractVisual foundation models (VFMs) have demonstrated impressive generalization capabilities in computer vision tasks. Previous studies show that fine-tuning VFMs with learnable tokens can achieve better generalization performance than full-parameter fine-tuning. The problem we need to address is how to learn the tokens that focus on the content information while ignoring the influence of style. For this purpose, we propose a novel Dual-Branch Content-enhanced Token (DBCT) learning framework. Specifically, we construct a style-suppressing branch, which contains a Style-sensitive Channel Suppression (SCS) module to transform the frozen VFM features into style-suppressed features, enabling the learning of style-invariant tokens. In addition, to compensate for the content degradation caused by the style-suppressing branch, we introduce a content-preserving branch that directly takes the frozen VFM features as input to learn content-focused tokens. Meanwhile, we propose a Token-query Linking (TLink) strategy to connect the two sets of tokens with the queries in the decoder. Through extensive experiments, our method achieves advanced results on various benchmarks. Shishun Tian, Wenbin Zou, Yuanhao Gong, Guanghui Yue 0001, Ting Su 0004 |
MMAsia | 1 |
| 2025 | Distortion-Aware Network for Zero-Reference Retinal Image EnhancementabstractCaptured retinal images usually have quality issues, manifested as containing multiple distortions (e.g., low light and blurring). Low-quality images bring a challenge to the screening and diagnosis of ophthalmic diseases. Existing image enhancement methods typically neglect the analysis of distortions and require high-quality reference images for model learning, making them unsuitable for clinical applications. In this paper, we propose a Distortion-Aware Network (DANet) for retinal image enhancement in a zero-reference way. DANet consists of three parallel branches by incorporating atmospheric scattering theory, which decomposes the low-quality image into a clean image, a transmission map, and an atmospheric light map. The upper branch utilizes a dark channel prior module to estimate the atmospheric light map, and the middle branch uses a transmission map generation module to estimate the transmission map. In contrast, the lower branch uses a deblurring module and a low-light enhancement module to obtain a deblurred image and an illumination-enhanced image and fuses these two images using a fusion block to generate the final enhanced image. Taking into account the limited publicly available datasets, we curate two datasets for the retinal image enhancement task. Experimental results show that our DANet can greatly improve the visual quality of the image with good interpretability, achieving superior performance over seven state-of-the-art methods. Tianwei Zhou, Yuhang Feng, Shaoping Zhang, Linling Li, Guanghui Yue 0001, Shishun Tian, Tianfu Wang 0001 |
MMAsia | 6 |
| 2025 | Class-discriminative domain generalization for semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Rong You, Wenbin Zou, Xia Li 0006 |
Image Vis. Comput. | 2 |
| 2025 | A global reweighting approach for cross-domain semantic segmentation
Yuhang Zhang 0011, Shishun Tian, Muxin Liao, Guoguang Hua, Wenbin Zou, Chen Xu 0004 |
Signal Process. Image Commun. | 2 |
| 2025 | A New Benchmark Database and Objective Metric for Light Field Image Quality EvaluationabstractLight Field Image (LFI) records both angular and spatial information and provides immersive experiences for observers by rendering a scene from multiple perspectives. To cope with the resolution limitations of capture hardware, LFI angular reconstruction and spatial super-resolution are two widely-used methods, but they can also induce some special types of distortions, especially when two methods are adopted in combination. To this end, new challenges have been brought in assessing the quality of these distorted LFIs. In this paper, firstly, we conduct subjective experiments to evaluate the distorted LFI quality and present a novel perceptual quality assessment database with the associated subjective quality scores. Specifically, the proposed database focuses on the distortions introduced by deep learning-based LFI angular reconstruction and spatial super-resolution methods, individually and multiplely. Besides, in the case of multiple distortions, the adoption order of two distortions is taken into consideration. Further, our database presents three types of LFIs that suffer from distortions: real-world, dense synthesis, and sparse synthesis. As a result, the quality of distorted LFIs was subjectively assessed by 32 valid observers using the Pairwise Comparison (PC) protocol. Secondly, we develop a novel objective No-Reference (NR) metric for LFI quality evaluation, based on the features extracted from spatial gradients, angular-spatial statistics, and binocular disparity. Finally, a benchmark of the proposed metric and numerous state-of-the-art quality assessment metrics on the proposed database is presented. Experimental results demonstrate the superiority of the proposed metric over most existing metrics in various aspects. The proposed database and metric will be publicly available athttps://github.com/ZhengyuZhang96/IETR-LFI. Shishun Tian, Jinjia Zhou, Luce Morin, Lu Zhang 0037 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Class-Balanced Sampling and Discriminative Stylization for Domain Generalization Semantic SegmentationabstractExisting domain generalization semantic segmentation (DGSS) methods have achieved remarkable performance on unseen domains by generating stylized images to increase the diversity of training data. However, since the training data is usually class-imbalanced, uniform style randomization is unable to generate diverse minority classes. This means that models may overfit to the minority classes, resulting in suboptimal performance on the minority classes. In addition, the image-level style randomization may also corrupt the class-discriminative regions of objects, leading to a loss of the class-discriminative representation. To address these issues, a novel class-balanced sampling and discriminative stylization (CSDS) approach is proposed for DGSS. Specifically, first, a pixel-level class-balanced sampling (PCS) strategy is proposed to adaptively sample patches of the minority classes from the source domain images and paste the sampled patches on the input images. Unlike existing class sampling strategies that fix the minority classes, the PCS strategy dynamically determines the minority classes by estimating the class distribution after each sampling. Then, a class-discriminative style randomization (CSR) strategy is proposed to increase the style diversity of the sampled patches while preserving the class-discriminative regions. Finally, since the pasting positions of the sampled patches are uncertain, which may confuse the semantic relations between the classes, a semantic consistency constraint is proposed to ensure the learning of reliable semantic relations. Extensive experiments demonstrate that the proposed approach achieves superior performance compared to existing DGSS methods on multiple benchmarks. The source code has been released onhttps://github.com/seabearlmx/CSDS. Muxin Liao, Shishun Tian, Binbin Wei, Yuhang Zhang 0011, Wenbin Zou, Xia Li 0006 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Multi-Level Guided Discrepancy Learning for Source-Free Object Detection in Hazy ConditionsabstractHaze deteriorates the quality of captured images, which severely limits the accuracy of clean image-trained object detectors in hazy conditions. Source-free domain adaptation (SFDA) aims to adapt a clean source image-trained detector to the unlabeled hazy target domain without access to the clean source domain data. However, existing source-free object detection (SFOD) methods encounter two issues when leveraging pseudo labeling paradigm: 1) the large domain shift between clear and hazy images may introduce noises in pseudo labels, 2) the single confidence threshold-based methods may ignore valuable information of the low-confidence samples. To address these issues, we propose a multi-level guided discrepancy learning-based approach for SFOD in hazy conditions, named MGDL. Specifically, we first propose a differentiated enhancement module (DEM) to intentionally augment the diversity of data styles. It uses dehazing and random perturbation to generate reliable high-quality labels and strengthen the tolerance to haze-related factors. To enhance the consistency constraints for discrepancy learning, we propose a multi-level guiding strategy (MLGS) which consists of a dual label-level and an instance-level guidance. Considering that false negatives may dominate in noisy labels, we propose a dual label-level guiding (DLG) strategy to excavate comprehensive useful information from high-and low-confidence samples. Besides, an instance-level contrastive learning (ICL) approach is proposed to guide the model to focus on objects and make the model insensitive to style at the same time. Extensive experiments conducted on multiple datasets have demonstrated that our method achieves superior performance over the state-of-the-art SFOD methods. Shishun Tian, Tiantian Zeng, Wenbin Zou |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Prototypical Progressive Alignment and Reweighting for Generalizable Semantic SegmentationabstractGeneralizable semantic segmentation, aims to excel on unseen target domains, as a critical focus due to the widespread practical applications requiring high generalizability. Class-wise prototypes, which depict class-wise centroids, as a type of domain-invariant information are key to improving the model generalizability due to its stability and representativeness. However, this manner faces some challenges. First, the existing methods adopt a coarse prototypical alignment form, potentially compromising performance. Second, the naive prototype generally serves as the class centroid generated by an average operation from source data batches, risks source domain overfitting, and may be detrimentally impacted by unrelated source data. Third, from a broader perspective, rather than just from a prototypical alignment perspective, the existing methods treat all samples equally, which is against the conclusion that different source features have different adaptation difficulties. To tackle these issues, we propose a novel method for generalizable semantic segmentation called Prototypical Progressive Alignment and Reweighting (PPAR) depending on the strong generalized representation of the Contrastive Language-Image Pretraining (CLIP) model. In particular, we first define the Original Text Prototype (OTP) and Visual Text Prototype (VTP) generated by the CLIP model, laying the foundation for the subsequent effective alignment strategy. Then, we propose a prototypical progressive alignment strategy by an easy-to-difficult alignment form to reduce domain-variant information progressively instead of directly. Finally, we propose a prototypical reweighting learning strategy that estimates the importance of the source data and corrects its learning weight to alleviate the influence of unrelated source features, i.e. alleviate negative transfer. Moreover, we also offer a theoretical insight into our method and it shows that our method compiles well on the domain generalization theory. Extensive experiments on several popular datasets demonstrate that our PPAR method achieves superior performance, proving the effectiveness of our method. The source code will available at: https://github.com/Hectoor/PPAR Yuhang Zhang 0011, Muxin Liao, Shishun Tian, Wenbin Zou, Lu Zhang 0037, Chen Xu 0004 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Dual Residual-Guided Interactive Learning for the Quality Assessment of Enhanced ImagesabstractImage enhancement algorithms can facilitate computer vision tasks in real applications. However, various distortions may also be introduced by image enhancement algorithms. Therefore, the image quality assessment (IQA) plays a crucial role in accurately evaluating enhanced images to provide dependable feedback. Current enhanced IQA methods are mainly designed for single specific scenarios, resulting in limited performance in other scenarios. Besides, no-reference methods predict quality utilizing enhanced images alone, which ignores the existing degraded images that contain valuable information, are not reliable enough. In this work, we propose a degraded-reference image quality assessment method based on dual residual-guided interactive learning (DRGQA) for the enhanced images in multiple scenarios. Specifically, a global and local feature collaboration module (GLCM) is proposed to imitate the perception of observers to capture comprehensive quality-aware features by using convolutional neural networks (CNN) and Transformers in an interactive manner. Then, we investigate the structure damage and color shift distortions that commonly occur in the enhanced images and propose a dual residual-guided module (DRGM) to make the model concentrate on the distorted regions that are sensitive to human visual system (HVS). Furthermore, a distortion-aware feature enhancement module (DEM) is proposed to improve the representation abilities of features in deeper networks. Extensive experimental results demonstrate that our proposed DRGQA achieves superior performance with lower computational complexity compared to the state-of-the-art IQA methods. Shishun Tian, Tiantian Zeng, Wenbin Zou, Xia Li 0006 |
IEEE Trans. Multim. | 1 |
| 2024 | Layout Relationship Decoupling Framework for Multi-target Domain Adaptative Semantic Segmentation
Yuhang Zhang 0011, Cuixin Yang, Muxin Liao, Shishun Tian, Wenbin Zou, Chen Xu 0004 |
MMAsia | 4 |
| 2024 | Strip and asymmetric aggregation network for unstructured terrain segmentation in wild environments
Shishun Tian, Yuhang Zhang 0011, Muxin Liao, Guoguang Hua, Wenbin Zou |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | PDA: Progressive Domain Adaptation for Semantic Segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
Knowl. Based Syst. | 2 |
| 2024 | Considering representation diversity and prediction consistency for domain generalization semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
Knowl. Based Syst. | 2 |
| 2024 | Video Generalized Semantic Segmentation via Non-Salient Feature Reasoning and Consistency
Yuhang Zhang 0011, Muxin Liao, Shishun Tian, Rong You, Wenbin Zou, Chen Xu 0004 |
Knowl. Based Syst. | 4 |
| 2024 | MI-RPN: Integrating multi-modalities and multi-scales information for region proposal
Shishun Tian, Wenbin Zou, Xia Li 0006 |
Multim. Tools Appl. | 1 |
| 2024 | Fine-Grained Self-Supervision for Generalizable Semantic SegmentationabstractUnsupervised domain adaptative semantic segmentation is a powerful solution for the distribution shift problem between the source and target domains. However, such methods need specified target domain data that may be unavailable in actual applications due to excess expensive collection. Generalizable semantic segmentation as a new paradigm appears in recent research, which aims to generalize well on distinct unseen domains only using source domain data. The existing methods focus on learning domain-invariant features by using global distribution alignment strategies, which may lead to a decreased discriminability of the model. To cope with this challenge, we propose a fine-grained self-supervision (FGSS) framework for generalizable semantic segmentation that takes into account both discriminability and generalizability from the perspective of the intra-class relationship. The FGSS framework contains single-view and multi-view versions. In the single-view version, we propose a fine-grained self-supervision strategy to distinguish the sub-parts of the semantic class for better class discriminability. In the multi-view version, we propose a class prototype feature enhancement strategy to generate another view (i.e. another representation of the original representation). Then, we propose a multi-view mutual supervision loss to enforce consistency between different views and further enhance the generalizability of the model. Experimental results on five widely-used datasets, i.e., GTAV, SYNTHIA, BDD100K, Cityscapes, and Mapillary, demonstrate that our FGSS framework achieves superior performance compared to state-of-the-art methods. Yuhang Zhang 0011, Shishun Tian, Muxin Liao, Wenbin Zou, Chen Xu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Preserving Label-Related Domain-Specific Information for Cross-Domain Semantic SegmentationabstractUnsupervised domain adaptation semantic segmentation (UDASS) methods aim to learn domain-invariant information for alleviating the distribution shift problem between the source and target domains. However, ignoring the learning of domain-specific information that is label-related may limit the class discriminability on the target domain. We argue that a good representation for the UDASS task not only contains domain-invariant information but also preserves label-related domain-specific information. In this paper, a novel frequency spectrum domain adaptation approach via meta-learning (ML-FSDA) is proposed to achieve this goal for improving the class discriminability and generalization ability. ML-FSDA contains a frequency-spectrum meta-learning framework (FMF) and a class-aware domain-specific memory bank (CDMB). Specifically, first, inspired by the observation that the high-frequency component is consistent across different domains while the low-frequency component is much more domain-specific, the FMF aims to respectively learn label-related domain-specific and domain-invariant information from low-frequency and high-frequency images in a unified framework via the meta-learning strategy. Second, the CDMB is designed to preserve the label-related domain-specific information of each class in an external memory bank while the CDMB is updated in every iteration of the meta-training stage. Finally, the CDMB is utilized to embed the label-related domain-specific information into domain-invariant information at the class level during the meta-testing stage to enhance the class discriminability on the target domain. Extensive experiments demonstrate the effectiveness of ML-FSDA on two challenging cross-domain semantic segmentation benchmarks. Notably, for the GTA5 to Cityscapes task and the SYNTHIA to Cityscapes task, the proposed ML-FSDA achieves superior performance with 77.3% mIoU and 68.8% mIoU, respectively. The source code is released at https://github.com/seabearlmx/FSL. Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Calibration-Based Multi-Prototype Contrastive Learning for Domain Generalization Semantic Segmentation in Traffic ScenesabstractPrototypical contrastive learning (PCL) has been widely used to learn class-wise domain-invariant features for domain generalization semantic segmentation. These methods assume that the prototypes in different domains are invariant. However, the prototypes in different domains have discrepancies as well. First, the prototypes of the same class in different domains may be different. Second, the prototypes of different classes may be similar. To address these issues, a calibration-based multi-prototype contrastive learning (CMPCL) approach is proposed, which contains an uncertainty-guided multi-prototype contrastive learning (UMPCL) and a hard-weighted multi-prototype contrastive learning (HMPCL). Specifically, the UMPCL uses an uncertainty probability matrix, derived from element-wise discrepancies between the prototypes of the same class, to calibrate the weights of prototypes for alleviating the discrepancy between the prototypes of the same class in different domains. The HMPCL uses a hard-weighted matrix that is generated by the similarity between the prototypes of different classes, to calibrate the weights of the hard-aligned prototypes for alleviating the issue of similar prototypes between different classes, with hard-aligned prototypes referring to those exhibiting such similarity. Furthermore, since the learned class-wise domain-invariant features may overfit the prototype in the source domain, multi-prototype contrastive learning is used in the UMPCL and HMPCL to avoid this risk. Extensive experiments demonstrate that our approach achieves superior performance over current approaches on multiple benchmarks of domain generalization semantic segmentation. The source code has been released onhttps://github.com/seabearlmx/CMPCL. Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | "Where Does the Devil Lie?": Multimodal Multitask Collaborative Revision Network for Trusted Road SegmentationabstractRoad segmentation is an essential component of navigation systems. Although recent advancements in road segmentation, the occurrence of failure segmentations remains inevitable. For safety-critical tasks, e.g., navigation, knowing when and where road segmentation fails is crucial. In this paper, we propose a novel trusted road segmentation architecture, namely Multimodal Multitask Collaborative Revision Network (M2CRN), to improve the trust of road segmentation. Our approach incorporates two strategies to predict and rectify segmentation errors. Firstly, a joint learning framework is devised to generate road segmentation results while estimating failure segmentation masks. Secondly, the road segmentation branch is equipped with an Uncertainty-Aware Revision Module (UARM), which eliminates the error in road segmentation. Additionally, we suppress the response of error regions in the road segmentation branch with an innovative design, called Adaptive Soft Error Suppression (ASES). To validate our methods, extensive experiments are conducted on three benchmark road segmentation datasets. The results demonstrate significant performance improvements with a real-time inference speed of 33.3 FPS, reaffirming the soundness of our revision model. Guoguang Hua, Dalian Zheng, Shishun Tian, Wenbin Zou, Shenglan Liu 0001, Xia Li 0006 |
IEEE Trans. Multim. | 3 |
| 2023 | TRG-DQA: Texture Residual-Guided Dehazed Image Quality AssessmentabstractImage dehazing algorithms have emerged to solve the visual impairment caused by haze. It is important to establish dehazed image quality assessment (DQA) methods that can accurately evaluate the dehazed image quality and the performance of dehazing algorithms. However, classical image quality assessment (IQA) and most hand-crafted feature based DQA methods may not be able to adequately measure complex distortions of dehazed images. To address this issue, this paper proposes a Texture Residual-Guided Dehazed image Quality Assessment (TRG-DQA) method. Specifically, we first introduce a global and local feature extraction module employing a combination of the Transformer and convolutional neural networks (CNN) for extracting the comprehensive features. Considering that texture residual maps represent haze density and artifact distortion information, we propose a residual-guided module to guide the model for efficient learning. Additionally, to mitigate the information loss issue that occurs in deeper networks, a distortion-aware feature enhancement module is proposed. Extensive experiments on six DQA databases demonstrate the proposed TRG-DQA achieves superior performance among all the state-of-the-art methods. Tiantian Zeng, Lu Zhang 0037, Wenbin Zou, Xia Li 0006, Shishun Tian |
ICIP | 5 |
| 2023 | Blind Quality Assessment of Light Field Image Based on Spatio-Angular Textural VariationabstractLight Field Image Quality Assessment (LF-IQA) is vitally important to facilitate the development of immersive technologies. However, current state-of-the-art LF-IQA metrics still struggle to handle Light Field Image (LFI) with massive data in an efficient manner. To cope with this challenge, we propose a simple yet effective Blind LF-IQA metric based on Spatio-Angular Textural Variation, named SATV-BLiF. Given a distorted LFI, we first apply Local Binary Pattern (LBP) operator to measure the textural variation in the spatial and angular domains respectively. Then the generated spatial and angular textural matrices are merged and further transformed into statistical textural histogram features. Finally, Support Vector Regression (SVR) is employed to construct a nonlinear mapping function between the statistical textural histogram features and the perceptual quality score of the distorted LFI. Experimental results on three representative light field databases show that the proposed metric achieves state-of-the-art quality evaluation performance, while having much lower complexity than the existing No-Reference (NR) LF-IQA metrics. The code of the proposed SATV-BLiF metric is available at https://github.com/ZhengyuZhang96/SATV-BLiF. Shishun Tian, Wenbin Zou, Yuhang Zhang 0011, Luce Morin, Lu Zhang 0037 |
ICIP | 2 |
| 2023 | Need a dog for seeing eye? A Walk Viewpoint Dataset for Freespace Detection in Unstructured EnvironmentsabstractFreespace Detection (FD) is crucial for robust and safe autonomous navigation. However, existing datasets usually concentrate on structure road environments. The FD in unstructured environments, e.g., walk assistance for the visually-impaired, has been rarely investigated. In this paper, We propose a novel dataset called the Walk Viewpoint Dataset (WVD). Different from the previous datasets, we focus on the walk viewpoint, where FD can provide the potential for improving the walking of visually impaired people. The target regions of WVD are annotated with 20 categories by fine-grained labels, which consist of 3,737 images and depth images. Moreover, we propose a new annotation hierarchy, which allows different degrees of complexity and creates opportunities for new training methods. Finally, our study provides the statistical analysis of label characteristics and baseline analysis, which demonstrates its distinction compared to previous datasets. The dataset can be accessed through the project pages: http://www.sensingAI.com.cn. Wenbin Zou, Guoguang Hua, Guangxu Chen, Zaiyue He, Guangli Liu, Huakun Li, Shishun Tian |
ICME | 10 |
| 2023 | Calibration-based Dual Prototypical Contrastive Learning Approach for Domain Generalization Semantic SegmentationabstractPrototypical contrastive learning (PCL) has been widely used to learn class-wise domain-invariant features recently. These methods are based on the assumption that the prototypes, which are represented as the central value of the same class in a certain domain, are domain-invariant. Since the prototypes of different domains have discrepancies as well, the class-wise domain-invariant features learned from the source domain by PCL need to be aligned with the prototypes of other domains simultaneously. However, the prototypes of the same class in different domains may be different while the prototypes of different classes may be similar, which may affect the learning of class-wise domain-invariant features. Based on these observations, a calibration-based dual prototypical contrastive learning (CDPCL) approach is proposed to reduce the domain discrepancy between the learned class-wise features and the prototypes of different domains for domain generalization semantic segmentation. It contains an uncertainty-guided PCL (UPCL) and a hard-weighted PCL (HPCL). Since the domain discrepancies of the prototypes of different classes may be different, we propose an uncertainty probability matrix to represent the domain discrepancies of the prototypes of all the classes. The UPCL estimates the uncertainty probability matrix to calibrate the weights of the prototypes during the PCL. Moreover, considering that the prototypes of different classes may be similar in some circumstances, which means these prototypes are hard-aligned, the HPCL is proposed to generate a hard-weighted matrix to calibrate the weights of the hard-aligned prototypes during the PCL. Extensive experiments demonstrate that our approach achieves superior performance over current approaches on domain generalization segmentation tasks. The source code will be released at https://github.com/seabearlmx/CDPCL. Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
ACM Multimedia | 2 |
| 2023 | Channel Affinity Knowledge Distillation for Semantic SegmentationabstractIn recent years, convolutional neural networks have achieved significant success in computer vision tasks. However, the deployment of these algorithms remains challenging. Knowledge distillation (KD) as a type of important method enables a tiny model to extract helpful information from a large model. Most existing KD methods based semantic segmentation aim to align predicted maps in the spatial domain, but channel distillation also may help to improve segmentation performance. Additionally, pairwise pixel affinity provides efficiently structured reasoning for semantic segmentation. Motivated by these considerations, we propose a novel Channel Affinity KD (CAKD) framework for semantic segmentation that focuses on channel and cross-channel affinity relationship distillation to better align the distribution of the student and teacher models. Extensive experiments demonstrate that our proposed approach outperforms state-of-the-art KD methods on Cityscapes, Pascal VOC, and ADE20k datasets. Huakun Li, Yuhang Zhang 0011, Shishun Tian, Rong You, Wenbin Zou |
MMSP | 3 |
| 2023 | SPNet: An RGB-D Sequence Progressive Network for Road Semantic SegmentationabstractRoad semantic segmentation is an essential component of autonomous driving and blind navigation. Although many excellent RGB-based road semantic segmentation algorithms have been proposed, these methods may not detect correctly due to the lack of geometric information. Recently, RGB-D road semantic segmentation methods attract more research attention. However, the existing RGB-D methods ignore the impact of unknown noise in sensors. To solve this problem, we propose an RGB-D Sequence Progressive Network (SPNet) for road semantic segmentation. Specifically, we first propose a sequence-based RGB-D feature extractor to alleviate the effect of noise. Then, We propose a multi-modal feature fusion (MMFF) module to enhance the feature representation of multi-modal data by further alleviating the effect of noise. Finally, we propose a semantic flow prediction (SFP) module that aims to align the multi-modal features in the decoder. Extensive experiments are conducted on several challenging datasets, including KITTI and GMRP. Our method achieves an F-score of 97.21% on the KITTI official leaderboard and ranked third in the official leaderboard. Yuhang Zhang 0011, Guoguang Hua, Ruijing Long, Shishun Tian, Wenbin Zou |
MMSP | 5 |
| 2023 | Domain-invariant information aggregation for domain generalization semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
Neurocomputing | 2 |
| 2023 | A hybrid domain learning framework for unsupervised semantic segmentation
Yuhang Zhang 0011, Shishun Tian, Muxin Liao, Wenbin Zou, Chen Xu 0004 |
Neurocomputing | 2 |
| 2023 | Learning Shape-Invariant Representation for Generalizable Semantic SegmentationabstractSemantic segmentation assigns a category for each pixel and has achieved great success in a supervised manner. However, it fails to generalize well in new domains due to the domain gap. Domain adaptation is a popular way to solve this issue, but it needs target data and cannot handle unavailable domains. In domain generalization (DG), the model is trained without the target data and DG aims to generalize well in new unavailable domains. Recent works reveal that shape recognition is beneficial for generalization but still lack exploration in semantic segmentation. Meanwhile, the object shapes also exist a discrepancy in different domains, which is often ignored by the existing works. Thus, we propose a Shape-Invariant Learning (SIL) framework to focus on learning shape-invariant representation for better generalization. Specifically, we first define the structural edge, which considers both the object boundary and the inner structure of the object to provide more discrimination cues. Then, a shape perception learning strategy including a texture feature discrepancy reduction loss and a structural feature discrepancy enlargement loss is proposed to enhance the shape perception ability of the model by embedding the structural edge as a shape prior. Finally, we use shape deformation augmentation to generate samples with the same content and different shapes. Essentially, our SIL framework performs implicit shape distribution alignment at the domain-level to learn shape-invariant representation. Extensive experiments show that our SIL framework achieves state-of-the-art performance. Yuhang Zhang 0011, Shishun Tian, Muxin Liao, Guoguang Hua, Wenbin Zou, Chen Xu 0004 |
IEEE Trans. Image Process. | 2 |
| 2023 | EDDMF: An Efficient Deep Discrepancy Measuring Framework for Full-Reference Light Field Image Quality AssessmentabstractThe increasing demand for immersive experience has greatly promoted the quality assessment research of Light Field Image (LFI). In this paper, we propose an efficient deep discrepancy measuring framework for full-reference light field image quality assessment. The main idea of the proposed framework is to efficiently evaluate the quality degradation of distorted LFIs by measuring the discrepancy between reference and distorted LFI patches. Firstly, a patch generation module is proposed to extract spatio-angular patches and sub-aperture patches from LFIs, which greatly reduces the computational cost. Then, we design a hierarchical discrepancy network based on convolutional neural networks to extract the hierarchical discrepancy features between reference and distorted spatio-angular patches. Besides, the local discrepancy features between reference and distorted sub-aperture patches are extracted as complementary features. After that, the angular-dominant hierarchical discrepancy features and the spatial-dominant local discrepancy features are combined to evaluate the patch quality. Finally, the quality of all patches is pooled to obtain the overall quality of distorted LFIs. To the best of our knowledge, the proposed framework is the first patch-based full-reference light field image quality assessment metric based on deep-learning technology. Experimental results on four representative LFI datasets show that our proposed framework achieves superior performance as well as lower computational complexity compared to other state-of-the-art metrics. Shishun Tian, Wenbin Zou, Luce Morin, Lu Zhang 0037 |
IEEE Trans. Image Process. | 2 |
| 2023 | Multiple Relational Learning Network for Joint Referring Expression Comprehension and SegmentationabstractMulti-task learning is a successful learning framework which improves the performance of prediction models by leveraging knowledge among related tasks. Referring expression comprehension (REC) and segmentation (RES) are highly relevant tasks, which both are language-guided visual recognition tasks. However, their relations have not yet been fully exploited in previous works. In this paper, a Multiple Relational Learning Network (MRLN) is proposed for multi-task learning of REC and RES. First, a feature-feature interaction learning module is introduced to handle the complicated interactions among features. Moreover, we propose a feature-task dependence learning module, which associates the related features with target tasks. Furthermore, a task-task relationship learning module is designed, which captures the relationships among tasks automatically and guides the REC and RES fine-tuning adaptively. To verify our proposed approach, experiments are conducted on three benchmark datasets, i.e., RefCOCO, RefCOCO+, and RefCOCOg. Extensive experiments demonstrate that the multiple relationships are more appealing since it alleviates the prediction inconsistency issue in multi-task setup. In addition, the experimental results report the significant performance gains of MRLN over most existing methods, i.e., up to 83.46 % for REC and 63.62 % for RES over state-of-the-art methods, which demonstrate the validity and superiority of MRLN. Guoguang Hua, Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Wenbin Zou |
IEEE Trans. Multim. | 3 |
| 2022 | Deeblif: Deep Blind Light Field Image Quality Assessment by Extracting Angular and Spatial InformationabstractIn the era of immersive media, the high-dimensional Light Field Image (LFI) puts forward higher requirements for Light Field Image Quality Assessment (LF-IQA). However, currently most existing LF-IQA metrics still rely on sophisticated hand-crafted feature extraction, which fail to predict the quality of LFI accurately. In this paper, we propose a patch-based Deep Blind Light Field image quality assessment metric (abbreviated as DeeBLiF), by employing a two-stream Convolutional Neural Network (CNN) model specifically designed for extracting the angular and spatial information of LFI. Firstly, the spatio-angular patches are generated as input data, which effectively reflect the spatio-angular information of LFI. After that, a two-stream CNN model is exploited to extract the patch features and further predict the patch scores. Finally, all the patch scores are pooled into an overall quality score of LFI. Experimental results on the LFI dataset demonstrate that the proposed DeeBLiF outperforms the state-of-the-art LF-IQA metrics. The code will be publicly available at https://github.com/ZhengyuZhang96/DeeBLiF. Shishun Tian, Wenbin Zou, Luce Morin, Lu Zhang 0037 |
ICIP | 2 |
| 2022 | Edge-preserving Image Smoothing via Counting-weighted Total VariationabstractWe present a new counting-weighted total variation measure, which captures the consistency of gradient directions in local regions to distinguish edges and details. A novel optimization framework with the proposed counting-weighted total variation in the l1regularization term is then developed to realize the edge-preserving image smoothing. In order to solve the optimization problem, we adopt an iteratively re-weighted least square based algorithm. Experimental results demonstrate that the proposed method is capable of completing edge-preserving image smoothing, while avoiding blurring and over-sharpening the edge. The proposed method can also be used for the image detail enhancement without involving halos or gradient reversal artifacts, while achieve better quality scores in the comparison with other enhancement methods. Jiachao Dang, Yi Liu 0004, Wenjing Shuai, Cong Bai, Shishun Tian |
MMSP | 5 |
| 2022 | Exploring more concentrated and consistent activation regions for cross-domain semantic segmentation
Muxin Liao, Guoguang Hua, Shishun Tian, Yuhang Zhang 0011, Wenbin Zou, Xia Li 0006 |
Neurocomputing | 3 |
| 2022 | RGB-D Gate-guided edge distillation for indoor semantic segmentation
Wenbin Zou, Yingqing Peng, Shishun Tian, Xia Li 0006 |
Multim. Tools Appl. | 4 |
| 2021 | Quality assessment of DIBR-synthesized views: An overview
Shishun Tian, Lu Zhang 0037, Wenbin Zou, Xia Li 0006, Ting Su 0004, Luce Morin, Olivier Déforges |
Neurocomputing | 1 |
| 2021 | STA3D: Spatiotemporally attentive 3D network for video saliency prediction
Wenbin Zou, Shengkai Zhuo, Yi Tang 0008, Shishun Tian, Xia Li 0006, Chen Xu 0004 |
Pattern Recognit. Lett. | 4 |
| 2021 | SC-RPN: A Strong Correlation Learning Framework for Region ProposalabstractCurrent state-of-the-art two-stage detectors heavily rely on region proposals to guide the accurate detection for objects. In previous region proposal approaches, the interaction between different functional modules is correlated weakly, which limits or decreases the performance of region proposal approaches. In this paper, we propose a novel two-stage strong correlation learning framework, abbreviated as SC-RPN, which aims to set up stronger relationship among different modules in the region proposal task. Firstly, we propose a Light-weight IoU-Mask branch to predict intersection-over-union (IoU) mask and refine region classification scores as well, it is used to prevent high-quality region proposals from being filtered. Furthermore, a sampling strategy named Size-Aware Dynamic Sampling (SADS) is proposed to ensure sampling consistency between different stages. In addition, point-based representation is exploited to generate region proposals with stronger fitting ability. Without bells and whistles, SC-RPN achieves AR100014.5% higher than that of Region Proposal Network (RPN), surpassing all the existing region proposal approaches. We also integrate SC-RPN into Fast R-CNN and Faster R-CNN to test its effectiveness on object detection task, the experimental results achieve a gain of 3.2% and 3.8% in terms of mAP compared to the original ones. Wenbin Zou, Yingqing Peng, Canqun Xiang, Shishun Tian, Lu Zhang 0037 |
IEEE Trans. Image Process. | 5 |
| 2020 | Matrix Capsule Convolutional Projection for Deep Feature LearningabstractCapsule projection network (CapProNet) has shown its ability to obtain semantic information, and spatial structural information from the raw images. However, the vector capsule of CapProNet has limitations in representing semantic information due to ignoring local information. Besides, the number of trainable parameters also increases greatly with the dimension of the feature vector. To that end, we propose a matrix capsule convolution projection (MCCP) module by replacing the feature vector with a feature matrix, of which each column represents a local feature. The feature matrix is then convoluted by columns into capsule subspaces to decrease the number of trainable parameters effectively. Furthermore, the CapDetNet is designed to explore the structural information encoding of the MCCP module based on object detection task. Experimental results demonstrate that the proposed MCCP outperforms the baselines in image classification, and CapDetNet achieves the 2.3% performance gain in object detection. Canqun Xiang, Zhennan Wang 0001, Shishun Tian, Jianxin Liao, Wenbin Zou, Chen Xu 0004 |
IEEE Signal Process. Lett. | 3 |
| 2019 | A Benchmark of DIBR Synthesized View Quality Assessment Metrics on a New Database for Immersive Media ApplicationsabstractDepth-image-based rendering (DIBR) is a fundamental technology in several 3-D-related applications, such as free viewpoint video, virtual reality, and augmented reality. However, new challenges have also been brought in assessing the quality of DIBR-synthesized views since this process induces some new types of distortions, which are inherently different from the distortion caused by video coding. In this paper, we present a new DIBR-synthesized image database with the associated subjective scores. We also test the performances of the state-of-the-art objective quality metrics on this database. This paper focuses on the distortions only induced by different DIBR synthesis methods. Seven state-of-the-art DIBR algorithms, including inter-view synthesis and single-view-based synthesis methods, are considered in this database. The quality of synthesized views was assessed subjectively by 41 observers and objectively using 14 state-of-the-art objective metrics. Subjective test results show that the interview synthesis methods, having more input information, significantly outperform the single-view-based ones. Correlation results between the tested objective metrics and the subjective scores on this database reveal that further studies are still needed for a better objective quality metric dedicated to the DIBR-synthesized views. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
IEEE Trans. Multim. | 1 |
| 2018 | SC-IQA: Shift compensation based image quality assessment for DIBR-synthesized viewsabstractDepth-image-based-rendering (DIBR) has been used to generate the virtual views for Multi-view videos and Free-viewpoint videos. However, the quality assessment of DIBR-synthesized views is very challenging owing to the new types of distortions induced by inaccurate depth maps, dis-occlusions and image inpainting methods. There exist a large number of object shifts and geometric distortions in the synthesized view which the traditional 2D quality metrics may fail to assess. In this paper, we propose a shift compensation based image quality assessment metric (SC-IQA) for DIBR-synthesized views. Firstly, the global geometric shift is compensated roughly by an SURF + RANSAC homography approach. Then, a multi-resolution block matching method, which performs a more accurate matching, is used to precisely compensate the shift and penalize the local geometric distortion as well. In addition, a visual saliency map is also used as a weighting function. To calculate the final overall quality scores, only the worst blocks are utilized since the biggest distortions have the most effects on the overall perceptual quality. The results show that the proposed metric significantly outperforms the state-of-the-art synthesized view dedicated metrics and the conventional 2D IQA metrics. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
VCIP | 1 |
| 2018 | NIQSV+: A No-Reference Synthesized View Quality Assessment MetricabstractBenefiting from multi-view video plus depth and depth-image-based-rendering technologies, only limited views of a real 3-D scene need to be captured, compressed, and transmitted. However, the quality assessment of synthesized views is very challenging, since some new types of distortions, which are inherently different from the texture coding errors, are inevitably produced by view synthesis and depth map compression, and the corresponding original views (reference views) are usually not available. Thus the full-reference quality metrics cannot be used for synthesized views. In this paper, we propose a novel no-reference image quality assessment method for 3-D synthesized views (called NIQSV+). This blind metric can evaluate the quality of synthesized views by measuring the typical synthesis distortions: blurry regions, black holes, and stretching, with access to neither the reference image nor the depth map. To evaluate the performance of the proposed method, we compare it with four full-reference 3-D (synthesized view dedicated) metrics, five full-reference 2-D metrics, and three no-reference 2-D metrics. In terms of their correlations with subjective scores, our experimental results show that the proposed no-reference metric approaches the best of the state-of-the-art full reference and no-reference 3-D metrics; and outperforms the widely used no-reference and full-reference 2-D metrics significantly. In terms of its approximation of human ranking, the proposed metric achieves the best performance in the experimental test. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
IEEE Trans. Image Process. | 1 |
| 2017 | NIQSV: A no reference image quality assessment metric for 3D synthesized viewsabstractThe popularity of 3D applications, such as Free View-point TV (FTV) and Multi-view Video plus Depth (MVD), induces a heavy requirement of synthesized views. However, the quality assessment of synthesized views is very challenging because the corresponding original views (reference views) are usually not available at both encoder and decoder sides. In this paper, we propose a new no-reference quality assessment model to evaluate the quality of 3D synthesized views, called NIQSV (No-reference Image Quality assessment of Synthesized Views). This metric is based on the hypothesis that a good quality image is composed of flat areas (objects) separated by sharp edges, and the quality estimation involves only a set of simple morphological operators. NIQSV integrates the distortions of all the components, and then uses an edge image to weight the final distortions since the distortions of synthesized views mainly happen around object edges. The experimental results show that the proposed metric outperforms traditional 2D metrics and ranks among the best of dedicated 3D synthesized and full reference metrics. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
ICASSP | 1 |