EDBT 2026 Demo / reviewers in the wild / expert
Wei Zhou 0021
dblp:69/5011-21
· DBLP profile ↗
109ranked-venue papers
18as first author
90since 2021 · last 2026
0000-0003-3641-1429ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 83 · 14 first-author · 68 since 2021Artificial intelligence and machine learning · 17 · 13 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-author · 8 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal Inconsistency Guidance for Super-resolution Video Quality AssessmentabstractAs super-resolution (SR) techniques introduce unique distortions that fundamentally differ from those caused by traditional degradation processes (e.g., compression), there is an increasing demand for specialized video quality assessment (VQA) methods tailored to SR-generated content. One critical factor affecting perceived quality is temporal inconsistency, which refers to irregularities between consecutive frames. However, existing VQA approaches rarely quantify this phenomenon or explicitly investigate its relationship with human perception. Moreover, SR videos exhibit amplified inconsistency levels as a result of enhancement processes. In this paper, we propose Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment (TIG-SVQA) that underscores the critical role of temporal inconsistency in guiding the quality assessment of SR videos. We first design a perception-oriented approach to quantify frame-wise temporal inconsistency. Based on this, we introduce the Inconsistency Highlighted Spatial Module, which localizes inconsistent regions at both coarse and fine scales. Inspired by the human visual system, we further develop an Inconsistency Guided Temporal Module that performs progressive temporal feature aggregation: (1) a consistency-aware fusion stage in which a visual memory capacity block adaptively determines the information load of each temporal segment based on inconsistency levels, and (2) an informative filtering stage for emphasizing quality-related features. Extensive experiments on both single-frame and multi-frame SR video scenarios demonstrate that our method significantly outperforms state-of-the-art VQA approaches. Xiaoyuan Yang 0003, Weide Liu, Xin Jin 0014, Xu Jia 0012, Yukun Lai, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
AAAI | 9 |
| 2026 | Quality-Complexity Trade-offs for Sustainable Media Delivery
Hadi Amirpour, Christian Herglotz, Lingfeng Qu, Wei Zhou 0021, Christian Timmerer |
QoMEX | 4 |
| 2026 | Cross-Modal Interaction for Multi-Dimensional AI-Generated Image Quality Assessment
Minghao Zou, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
QoMEX | 5 |
| 2026 | DWCL: Dual-Weighted Contrastive Learning for robust multi-view clustering
Hanning Yuan, Lianhua Chi, Sijie Ruan, Wei Zhou 0021, Jinhui Pang, Xiaoshuai Hao |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | One Aligned LLM to Serve Them All: A Transfer Recipe for Training VLMs without Visual-Language Re-Alignment
Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Wei Zhou 0021, Dong Wang 0004, Huchuan Lu, You He 0002 |
Int. J. Comput. Vis. | 5 |
| 2026 | EHIN: Early-aware hierarchical interaction network for weakly-supervised referring image segmentation
Anqing Chen, Wanli Ma 0001, Weide Liu, Yakun Ju, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
Neurocomputing | 10 |
| 2026 | Deep learning-based point cloud upsampling: A survey of methodologies, performance comparisons, and noise robustness analysis
Yihang Yin, Li Yu 0004, Wei Zhou 0021, Moncef Gabbouj |
Neurocomputing | 3 |
| 2026 | BeatDance: Generating beat-consistent 3D dance with hierarchical spatial-temporal modeling
Xiaojian Shen, Dahu Shi, Jianrong Zhang, Yunzhi Zhuge, Zhiliang Wu, Guanghui Yue 0001, Wei Zhou 0021 |
Pattern Recognit. | 10 |
| 2026 | Robust low-light image enhancement in the wild via data synthesis and generative diffusion prior
Zhihua Wang 0002, Qinghua Lin, Weixia Zhang, Wei Zhou 0021 |
Pattern Recognit. | 5 |
| 2026 | Expressive Human Volumetric Video Generation With Rich TextabstractPlain text has become the dominant interactive interface for text-driven human volumetric video generation. However, its limited customization options hinder users from expressing motion effects with accuracy. For example, plain text struggles to specify continuous variables such as motion amplitude, speed, and joint trajectories with precision, and it fails to convey stylized motion characteristics. Additionally, crafting detailed textual prompts for complex motion sequences is cumbersome, while excessively long prompts strain text encoders. To address these limitations, we propose a rich text-based framework that supports font styles, sizes, and trajectory sketching. By extracting motion-related attributes from rich text, our method enables fine-grained control over motion styles, precise speed regulation, and accurate joint trajectory manipulation. These capabilities are realized through gradient-guided noise editing and ControlNet-based motion optimization, which operate within the latent motion diffusion process. Specifically, we design a unified gradient-guided adaptation mechanism to ensure that the generated motion video adheres strictly to the specified constraints. Furthermore, we introduce realism-oriented optimization for stylistic and joint-level control, refining motion synthesis at a granular level to produce smoother, more natural movements. We present multiple comparative evaluations showcasing volumetric video generation from both rich text and plain text. Through quantitative analysis, we demonstrate that our method surpasses strong plain-text baselines, producing expressive, customizable human volumetric motion videos. Guanghui Yue 0001, Wei Zhou 0021, Xudong Mao, Ruomei Wang 0001, Baoquan Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | DVLTA-VQA: Decoupled Vision-Language Modeling With Text-Guided Adaptation for Blind Video Quality AssessmentabstractInspired by the dual-stream (dorsal and ventral streams) theory of the human visual system (HVS), recent Video Quality Assessment (VQA) methods have integrated Contrastive Language-Image Pretraining (CLIP) to enhance semantic understanding. However, as CLIP is originally designed for images, it lacks the ability to adequately capture the temporal dynamics and motion perception (dorsal stream) inherent in videos. To address this limitation, we propose DVLTA-VQA (Decoupled Vision-Language Modeling with Text-Guided Adaptation), which decouples CLIP’s visual and textual components to better align with the NR-VQA pipeline. Specifically, we introduce a Video-Based Temporal CLIP module and a Temporal Context Module to explicitly model motion dynamics, effectively enhancing the dorsal stream representation. Complementing this, a Basic Visual Feature Extraction Module is employed to strengthen spatial detail analysis in the ventral stream. Furthermore, we propose a text-guided adaptive fusion strategy that leverages textual semantics to dynamically weight visual features, facilitating effective spatiotemporal integration. Extensive experiments on multiple public datasets demonstrate that the proposed method achieves state-of-the-art performance, significantly improving prediction accuracy and generalization capability. Li Yu 0004, Situo Wang, Wei Zhou 0021, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Pose-Guided Multi-Cue Explicit Query Construction for Disambiguating Human-Object InteractionsabstractHuman-Object Interaction (HOI) detection remains challenging due to the semantic ambiguity of interaction categories and the limited discriminability of their feature representations. Existing approaches often improve recognition by employing sophisticated models or auxiliary textual annotations. While effective in certain gains, these solutions incur additional computational or annotation costs and struggle to capture intrinsic interaction regularities. To address these issues, we propose Pose-Guided Multi-Cue Explicit Query Construction (PM-EQC), a unified Transformer-based framework that builds upon collaborative modeling of appearance, spatial, and pose cues for discriminative interaction reasoning. At its core, the Collaborative Multi-Cue Query Constructor (CM-CQC) jointly models dependencies among visual cues to generate explicit query embeddings. CM-CQC further incorporates a hierarchical pose contextualization mechanism: global body configurations adaptively guide attention to local critical joints, yielding fine-grained pose embeddings and more precise interaction disambiguation. Owing to its modular design, PM-EQC integrates seamlessly with diverse backbones and benefits from their advances. Extensive experiments on PhysLab, HICO-DET, and V-COCO datasets demonstrate that PM-EQC achieves state-of-the-art performance, and the code is publicly available at https://github.com/ZMHSDUST/ PM-EQC. Minghao Zou, Qingtian Zeng, Xue Zhang 0008, Guiyuan Yuan, Xiaoshuai Hao, Jun Liu 0036, Wei Zhou 0021 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Perception-Inspired Network for Stereo Image Quality AssessmentabstractExisting stereo image quality assessment (SIQA) methods generally have limitations in binocular fusion and fine-grained perception modeling. To address these issues, we propose a Perception-Inspired Network for SIQA that simulates binocular difference-guided fusion, high-frequency sensitivity, and hierarchical perception mechanisms of the human visual system (HVS). First, a difference-guided binocular fusion (DGBF) module is designed to mimic the binocular difference sensitivity mechanism, which exploits difference information at both the feature-level and image-level to optimize binocular fusion. Furthermore, the image distortion primarily affects the high-frequency components, which are critical for perceptual quality. To reflect this, we propose a high-frequency enhancement module (HFEM) to simulate the human eye's sensitivity to edge and texture distortions. Finally, to better achieve fine-grained perception modeling, we propose a hierarchical quality regression strategy that simulates the human perceptual process, from perceiving local details to forming a global quality judgment, thereby achieving a quality prediction more aligned with human subjective evaluation. Experimental results demonstrate that the proposed method outperforms mainstream approaches, achieving a PLCC of 0.9734 on the LIVE I database, and a PLCC of 0.9632 on the LIVE II database. Yongli Chang, Guanghui Yue 0001, Li Yu 0004, Yakun Ju, Hadi Amirpour, Moncef Gabbouj, Wei Zhou 0021 |
IEEE Trans. Image Process. | 8 |
| 2026 | Embodied Spatial Affordance: Spatial-Aware Affordance Learning for Embodied Navigation and ManipulationabstractEmbodied navigation and manipulation are fundamental capabilities for embodied agents operating in physical environments. A key challenge in this process is understanding the spatial context and the affordances of the environment, which involves recognizing how objects can be interacted with (object affordance) and identifying suitable locations for movement and object placement (free space affordance). While Vision-Language Models (VLMs) have shown promise in high-level task planning, their ability to translate reasoning into precise executable actions remains limited, particularly in image-based spatial understanding and precise affordance localization-a critical gap in image processing for robotics. To bridge this gap, we propose EspA, a novel image-to-keypoint model that leverages spatial-aware affordance learning to predict actionable affordances directly from 2D image inputs. Built on a hierarchical vision-language architecture, EspA jointly reasons about object affordances and free space affordances, enabling pixel-level localization of both types of interactions. Crucially, EspA translates language instructions into precise 2D affordance keypoints from observed images, which are then projected into 3D actionable coordinates using depth information. To support this unified affordance reasoning, we introduce the Embodied Spatial Affordance (ESA) dataset, which captures both object-centric interactions and free space contexts. By jointly modeling these affordances in a shared representation space, EspA overcomes the limitations of prior works that treat them independently. The dataset's fine-grained annotations enable our model to learn the intricate relationship between object functionality and spatial feasibility, significantly enhancing the spatial understanding in embodied tasks. Extensive experimental results demonstrate that EspA outperforms existing state-of-the-art Vision-Language Models (VLMs), both open-source and closed-source, in object and free space affordance prediction. Furthermore, it exhibits superior performance in real-world embodied navigation and manipulation experiments. Our work advances the field of image-based spatial reasoning by providing a scalable solution for translating high-level instructions into low-level actionable affordances. We believe this work paves the way for more robust and versatile embodied agents capable of effectively interacting with complex environments. The dataset, benchmark, and evaluation code will be publicly available to facilitate future research. Project website: https://embodied-spatial-affordance.github.io/. Xiaoshuai Hao, Yingbo Tang, Long Chen 0015, Wei Zhou 0021, Jungong Han, Wenbo Ding 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 5 |
| 2026 | Self-Supervised Unfolding Network With Shared Reflectance Learning for Low-Light Image EnhancementabstractRecently, incorporating Retinex theory with unfolding networks has attracted increasing attention in the low-light image enhancement field. However, existing methods have two limitations, i.e., ignoring the modeling of the physical prior of Retinex theory and relying on a large amount of paired data. To advance this field, we propose a novel self-supervised unfolding network, named S2UNet, for the LIE task. Specifically, we formulate a novel optimization model based on the principle that content-consistent images under different illumination should share the same reflectance. The model simultaneously decomposes two illumination-different images into a shared reflectance component and two independent illumination components. Due to the absence of the normal-light image, we process the low-light image with gamma correction to create the illumination-different image pair. Then, we translate this model into a multi-stage unfolding network, in which each stage alternately optimizes the shared reflectance component and the respective illumination components of the two images. During progressive multi-stage optimization, the network inherently encodes the reflectance consistency prior by jointly estimating an optimal reflectance across varying illumination conditions. Finally, considering the presence of noise in low-light images and to suppress noise amplification, we propose a self-supervised denoising mechanism. Extensive experiments on nine benchmark datasets demonstrate that our proposed S2UNet outperforms state-of-the-art unsupervised methods in terms of both quantitative metrics and visual quality, while achieving competitive performance compared to supervised methods. The source code will be available at https://github.com/J-Liu-DL/S2UNet. Jia Liu 0025, Yu Luo 0004, Guanghui Yue 0001, Jie Ling 0002, Chia-Wen Lin, Guangtao Zhai, Wei Zhou 0021 |
IEEE Trans. Image Process. | 8 |
| 2026 | Integrating SAM Supervision for 3D Weakly Supervised Point Cloud SegmentationabstractCurrent methods for 3D semantic segmentation propose training models with limited annotations to address the difficulty of annotating large, irregular, and unordered 3D point cloud data. They usually focus on the 3D domain only, without leveraging the complementary nature of 2D and 3D data. Besides, some methods extend original labels or generate pseudo labels to guide the training, but they often fail to fully use these labels or address the noise within them. Meanwhile, the emergence of comprehensive and adaptable foundation models has offered effective solutions for segmenting 2D data. Leveraging this advancement, we present a novel approach that maximizes the utility of sparsely available 3D annotations by incorporating segmentation masks generated by 2D foundation models. We further propagate the 2D segmentation masks into the 3D space by establishing geometric correspondences between 3D scenes and 2D views. We extend the highly sparse annotations to encompass the areas delineated by 3D masks, thereby substantially augmenting the pool of available labels. Furthermore, we apply confidence- and uncertainty-based consistency regularization on augmentations of the 3D point cloud and select the reliable pseudo labels, which are further spread on the 3D masks to generate more labels. This innovative strategy bridges the gap between limited 3D annotations and the powerful capabilities of 2D foundation models, ultimately improving the performance of 3D weakly supervised segmentation. Lechun You, Weide Liu, Xulei Yang, Jun Cheng 0003, Wei Zhou 0021, Bharadwaj Veeravalli, Guosheng Lin |
IEEE Trans. Image Process. | 6 |
| 2026 | StealthMark: Harmless and Stealthy Ownership Verification for Medical Segmentation via Uncertainty-Guided BackdoorsabstractAnnotating medical data for training AI models is often costly and limited due to the shortage of specialists with relevant clinical expertise. This challenge is further compounded by privacy and ethical concerns associated with sensitive patient information. As a result, well-trained medical segmentation models on private datasets constitute valuable intellectual property requiring robust protection mechanisms. Existing model protection techniques primarily focus on classification and generative tasks, while segmentation models-crucial to medical image analysis-remain largely underexplored. In this paper, we propose a novel, stealthy, and harmless method, StealthMark, for verifying the ownership of medical segmentation models under closed-box conditions. Our approach subtly modulates model uncertainty without altering the final segmentation outputs, thereby preserving the model's performance. To enable ownership verification, we incorporate model-agnostic explanation methods, e.g. LIME, to extract feature attributions from the model outputs. Under specific triggering conditions, these explanations reveal a distinct and verifiable watermark. We further design the watermark as a QR code to facilitate robust and recognizable ownership claims. We conducted extensive experiments across four medical imaging datasets (CMR dataset from UK Biobank, the SEG fundus dataset, the EchoNet echocardiography dataset, and the PraNet colonoscopy dataset) and five mainstream segmentation models. The results demonstrate the effectiveness, stealthiness, and harmlessness of our method on the original model's segmentation performance. For example, when applied to the SAM model, StealthMark consistently achieved attack success rates (ASR) above 95% across various datasets while maintaining less than a 1% drop in Dice and AUC scores-significantly outperforming backdoor-based watermarking methods and highlighting its strong potential for practical deployment. Our implementation code is made available at https://github.com/Qinkaiyu/StealthMark. Qinkai Yu, Chong Zhang 0006, Gaojie Jin, Tianjin Huang, Wei Zhou 0021, Xiao-Bo Jin, Bo Huang 0012, Yitian Zhao, Gregory Yoke Hong Lip, Yalin Zheng, Aline Villavicencio, Yanda Meng |
IEEE Trans. Image Process. | 5 |
| 2026 | Uncertainty-Guided Spatiotemporal Consistency Fusion Network for Infrared-Visible Video Fusion Under Extremely Low-Light ConditionsabstractInfrared-visible video fusion under extremely low-light conditions is critically important yet remains underexplored, largely due to the scarcity of high-quality datasets and challenges posed by spatiotemporal uncertainty and modality bias. To address the dataset shortage, we built a dataset of 4,739 infrared and visible registration video pairs captured under extremely low-light conditions, spanning 5 scene types and 17 subcategories. Further, we proposed an Uncertainty-guided Spatiotemporal Consistency Fusion Network, termed USCFNet, for the infrared-visible video fusion. At each layer of the encoder, an Entropy-Gated SpatioTemporal Attention (EGSTA) module is introduced to capture temporal instability and spatial reliability variations through entropy-aware attention modulation, thereby enhancing feature spatiotemporal consistency. The refined infrared and visible features are then fused via a Difference-Guided Fusion (DGF) module, which adaptively exploits their content and edge differences to improve structural integrity and detail clarity. By progressively connecting DGF modules from shallow to deep layers, the network achieves the synergistic fusion of shallow textures and deep semantics. Subsequently, the output of the last DGF module is fused with the modality features of the last layer through a hierarchical mixture-of-experts fusion module. This module enables the balanced integration of modality information while preserving fine local details. Finally, the fusion feature is fed into the decoder to produce the final fused video. Extensive experiments on our dataset and two public datasets show that USCFNet outperforms competing methods, achieving lower distortion and stronger spatiotemporal consistency. The source code and dataset are available at https://github.com/Zhaocheng1/ELVID. Cheng Zhao 0003, Tianyun Song, Zhiliang Wu, Tianfu Wang 0001, Moncef Gabbouj, Guanghui Yue 0001, Bai Ying Lei, Wei Zhou 0021 |
IEEE Trans. Image Process. | 8 |
| 2026 | KSIQA: A Knowledge-Sharing Model for No-Reference Image Quality AssessmentabstractNo-reference image quality assessment (NR-IQA) aims to quantitatively measure human perception of visual quality without comparing a distorted image to a reference. Despite recent advances, existing NR-IQR approaches often demonstrate insufficient ability to capture perceptual cues in the absence of a reference, limiting their generalisability across diverse and complex real-world image degradations. These limitations hinder their ability to match the reliability of full-reference IQA (FR-IQA) counterparts. A key challenge, therefore, is to enable NR-IQA models to emulate the reference-aware reasoning exhibited by humans and FR-IQA methods. To address this challenge, we propose a novel NR-IQA model based on a knowledge-sharing (KS) strategy to simulate this capability and predict image quality more effectively. Specifically, we designate an FR-IQA model as the teacher and an NR-IQA model as the student. Unlike conventional knowledge distillation (KD), our proposed architecture enables the NR-IQA student and FR-IQA teacher to share a decoder rather than being independent models. Furthermore, the student model contains a Mental Imagery Generation (MIG) module to learn mental imagery as the reference. To fully exploit local and global information, we adopt a vision transformer (ViT) branch and a convolutional neural network branch for feature extraction (FE). Finally, a quality-aware regressor (QAR) combined with deep ordinal regression is constructed to infer the quality score. Experiments show that our proposed NR-IQA model, KSIQA, has class-leading performance against current no-reference (NR) techniques across widespread benchmark datasets. Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Ying Chen 0011, Roger M. Whitaker, Walter Colombo, Hantao Liu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2026 | DADA++: Dual Alignment Domain Adaptation for Unsupervised Video-Text RetrievalabstractVideo-text retrieval aims at returning the most semantically relevant videos given a textual query, which is a thriving topic in both computer vision and natural language processing communities. This article focuses on a more challenging task, i.e., Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), wherein training and testing data come from different distributions. Previous approaches are mostly derived from classification-based domain adaptation methods, which are neither multi-modal nor suitable for retrieval tasks. They merely alleviate the domain shift while overlooking the pairwise misalignment issue in the target domain, i.e., there exist no semantic relationships between target videos and texts. While Foundation Models like CLIP perform well in in-domain video-text retrieval, their effectiveness significantly drops during domain shifts due to this lack of alignment. To tackle this, we propose a novel method named D ual A lignment D omain A daptation ( DADA ++). Specifically, we first introduce cross-modal semantic embedding to generate discriminative source features in a joint embedding space. Besides, we utilize cross-modal domain adaptations to balance the minimization of domain shift in a smooth manner. Furthermore, we empirically identify the pairwise misalignment in the target domain, and thus propose the i ntegrated D ual A lignment C onsistency (iDAC). The proposed iDAC adaptively aligns the video-text pairs, which are more likely to be relevant in the target domain, by verifying their cross-modal semantic proximity reciprocally in both hard and soft manners. This enables positive pairs to increase progressively while potentially aligning noisy pairs throughout the training procedure. We also provide insights into the functionality of DADA ++ through the lens of Foundation Models, explaining its superiority in a theoretical way. Compared with state-of-the-art methods, DADA ++ achieves 9.4% and 8.5% relative improvements on R@1 under the settings of TGIF \(\rightarrow\) MSR-VTT and TGIF \(\rightarrow\) MSVD, respectively, demonstrating its superior performance. Xiaoshuai Hao, Yunfeng Diao, Rong Yin 0001, Guangyin Jin, Jing Zhang 0037, Wanqian Zhang, Wei Zhou 0021 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2026 | Rethinking the Effect of Unimodal Labels in Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis aims to comprehensively understand human sentiment by integrating diverse modalities, such as text, audio, and vision. To improve modality complementarity, the recent Multimodal Multi-task Learning (MML) framework employs joint training of unimodal and multimodal sentiment analysis tasks using sub-annotations of modality. In this work, we further draw attention to the observation that integrating unimodal tasks may introduce conflicting task information, negatively affecting the multimodal task performance. Motivated by this issue, we propose the Multimodal Task Correlation-aware Learning (MTCL) framework to leverage beneficial task correlations and suppress harmful ones. Specifically, MTCL introduces a Correlation-Adaptive Training (CAT) strategy to learn a task-relation aware unimodal encoder for each modality. First, in order to distinguish whether a sample contains conflicting information, CAT strategy incorporates a Dual-Branch Contrast (DBC) module which divides the training set into a beneficial subset and a harmful subset. Based on this division, CAT strategy proposes an adaptive training loss to guide the model in understanding nuanced multitask correlations. The adaptive training loss has two components: (1) For the beneficial subset, a contrastive loss is utilized to improve the model’s ability to extract complementary representations. (2) For the harmful subset, we apply a task-correction loss to mitigate the negative interference caused by harmful task associations. With the CAT strategy, our framework can effectively distinguish beneficial and harmful task correlations to extract distinctive and robust unimodal representations. The superiority of MTCL is verified via extensive experiments on several multimodal video sentiment analysis benchmarks. Our work is publicly available at https://github.com/tiggers23/MTCL . Tianrui Li 0001, Baiyu Lu, Junlin Fang, Desheng Zheng, Wei Zhou 0021, Weide Liu, Fengmao Lv |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2026 | A New Semi-Supervised Video Anomaly Detection Baseline in Lack of Anomalous SamplesabstractVideo anomaly detection (VAD) has been widely studied for its important applications in multimedia community. Recently, many Weakly Supervised VAD (WS-VAD) methods have been proposed, which tend to treat VAD as a classification task through multiple instance learning and result in the need to collect sufficient anomaly classes and samples to be used for training a classifier. However, anomaly events tend to be open-set and rare in real-world applications, so we often have difficulty collecting all anomaly classes and enough sample anomalies, which is a difficult situation for WS-VAD to cope with. To this end, we consider to treat VAD as an out-of-distribution detection task rather than a classification task and propose a simple but effective semi-supervised baseline method. First, we leverage the powerful zero-shot capability of large visual language models to generate summary text descriptions for videos and extract visual features as intermediates for subsequent use. Next, we use a text encoder to extract language features and combine them with visual features to obtain robust multimodal features. Finally, we introduce an out-of-distribution detection method that learns the center of normality in multimodal space from normal and unlabeled samples, while deviating abnormal samples from the center to cope with the scarcity of abnormal samples. To implement our baseline method, we also provide a new semi-supervised dataset by reorganizing an existing benchmark, which is the first available dataset in the VAD community that provides trimmed videos consisting of complete abnormal events. Experiments demonstrate that our method performs more robustly when fewer anomaly classes and anomaly samples collected. Mengyang Zhao 0002, Haiyang Yu 0004, Teng Fu 0001, Yang Liu 0246, Wei Zhou 0021, Bin Li 0015, Xiangyang Xue 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | AdaptSAM: Adaptive SAM for Cross-Domain Few-Shot Medical Image SegmentationabstractDeep networks excel in medical image segmentation with large annotated datasets but struggle with generalization to unseen out-of-domain data, a common challenge in clinical settings. Although the Segment Anything Model (SAM) demonstrates strong prompt-driven generalization in natural image segmentation, its application to clinical segmentation, especially in cross-domain and few-shot scenarios, faces limitations due to insufficient domain-specific features, weak structural fidelity, and prompt instability. To address these challenges, we propose AdaptSAM, an innovative framework to improve cross-domain few-shot medical image segmentation through three key components: Frequency-aware, Semantically-aligned, and Promptadaptive strategies. The Frequency-domain Multi-scale Feature Enhancement (FFE) module extracts frequency-aware features and enriches them with context-sensitive semantics, mitigating domain-specific feature loss. The enhanced features are fused with the Hierarchical Semantic Refinement (HSR) module, utilizing high-level semantic activations to recalibrate shallow-layer features, improving fine-structure fidelity and boundary preservation. Additionally, the Dynamic Curriculum Prompt (DCP) mechanism adjusts prompt box sizes during training, guiding the model to learn object-boundary interactions and background context in a coarse-to-fine manner, thereby aligning prompts with target domain features and enhancing segmentation robustness across domain shifts. AdaptSAM outperforms state-of-the-art methods on ten public medical segmentation datasets, achieving 6.22 % higher Dice than SAM-based methods and 10.37 % higher than fully supervised domain adaptation, showcasing superior cross-domain generalization with just one labeled sample. Code is provided at: https://github.com/Ggllllllll/AdaptSAM. Wei Zhou 0021, Guilin Guan, Qifeng Yan, Yitian Zhao |
BIBM | 1 |
| 2025 | FE-CLIP: Frequency Enhanced CLIP Model for Zero-Shot Anomaly Detection and Segmentation
Qi Chu 0001, Bin Liu 0016, Wei Zhou 0021, Nenghai Yu |
ICCV | 4 |
| 2025 | MSPoint-Gait: Multi-Scale Point Cloud Analysis for 3D Gait Recognition via Cross-Modal LearningabstractRecent advances in LiDAR technology have enabled privacy-preserving gait recognition using 3D point cloud data. However, existing approaches struggle with the inherent challenges of point cloud processing and understanding such as spatial sparsity, irregular sampling, and complex temporal dynamics. In this paper, we present MSPoint-Gait, a novel framework that addresses these challenges through multi-scale analysis and cross-modal learning. At the core of our framework lies a Depth-Aware Attention Module (DAAM) that leverages rich 3D geometric information to generate attention-weighted depth representations, enabling fine-grained feature extraction from point cloud sequences. We further introduce a Multi-Scale Spatio-Temporal (MSST) network that hierarchically captures both local and global gait patterns through adaptive convolution kernels across multiple spatial and temporal scales. These components are unified through a novel cross-modal learning strategy that effectively bridges the semantic gap between raw point clouds and structured depth representations. The proposed frame-work achieves state-of-the-art performance on the challenging SUSTech1K dataset, with 91.9% Rank-1 and 98.0% Rank-5 accuracy, demonstrating significant improvements over existing methods across various walking conditions and viewpoints. Xinzhu Li, Yikun Chen, Guanghui Yue 0001, Wei Zhou 0021, Ruomei Wang 0001, Xudong Mao, Juepeng Zheng, Fan Zhou 0001, Ziqi Qiu, Baoquan Zhao |
ICME | 5 |
| 2025 | MCSMoG: Multi-Conditional Diffusion for Stylized Motion Generation with Parametric ControlabstractStylized human motion synthesis remains a fundamental challenge in computer animation and graphics, with a wide spectrum of applications spanning gaming, film production, virtual reality, and beyond. While recent advances in text-driven motion generation have shown promise, existing approaches face critical limitations including the inability to maintain consistent trajectory control, the lack of fine-grained stylization intensity adjustment, and inadequate generalization across diverse motion styles. To address these challenges, We introduce MCSMoG, a novel framework for controllable stylized motion synthesis through multi-conditional guidance. First, a new Multi-Conditional Motion Latent Diffusion (MC-MLD) model is proposed to introduce additional trajectory guidance and achieve trajectory decoupling. Second, we develop a Style and Non-Style Feature Fusion Module that dynamically blends motion features through an adjustable parameter, providing control over stylization intensity. Third, we integrate MotionCLIP as our style encoder, enhancing the model’s generalization capability across diverse and unseen motion styles. Extensive experiments conducted on the combined HumanML3D and 100STYLE datasets demonstrate that our approach outperforms state-of-the-art methods, achieving a 4.6% reduction in FID scores and a 4.1% increase in motion diversity. User studies further confirm the superiority of our method in style fidelity, semantic consistency, and motion naturalness. Xinzhu Li, Guanghui Yue 0001, Wei Zhou 0021, Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
ICME | 5 |
| 2025 | Multi-Attribute Continual Learning for Blind Image Quality AssessmentabstractBlind image quality assessment (BIQA) has evolved into a critical task in visual computing, requiring effective evaluation across multiple quality attributes such as brightness, sharpness, contrast, and colorfulness. Traditional BIQA methods based on single-task learning often suffer from catastrophic forgetting and struggle to generalize across diverse IQ attributes. To address these challenges, we propose a novel Multi-attribute Continual Learning framework, MaC-BIQA, which integrates Gated Attention Mechanism and Knowledge Graph Embedding (KGE) with the Learning without Forgetting (LwF) approach. Specifically, the Gated Attention Mechanism dynamically adjusts attention distribution by focusing on task-specific key regions, while the integration of Knowledge Graph Embedding (KGE) complements LwF by improving the understanding of inter-task relationships, ensuring that relevant information from previous tasks is preserved more effectively, even as the model learns new tasks. Together, they enhance the model’s adaptability and robustness in handling complex multi-attribute tasks, effectively mitigating catastrophic forgetting and significantly improving overall performance in multi-task environments. Extensive experiments on the KonIQ-10K and SPAQ datasets show that our method significantly reduces forgetting rates and improves robustness and generalization across multiple quality attributes. This study presents a key technical contribution by addressing the limitations of catastrophic forgetting and offering a scalable, adaptive solution for real-world multi-attribute BIQA. Yunhao Luo 0005, Jinming Liu 0001, Wei Zhou 0021, Xin Jin 0014 |
ISCAS | 3 |
| 2025 | CLIP-DQA: Blindly Evaluating Dehazed Images from Global and Local Perspectives Using CLIPabstractBlind dehazed image quality assessment (BDQA), which aims to accurately predict the visual quality of dehazed images without any reference information, is essential for the evaluation, comparison, and optimization of image dehazing algorithms. Existing learning-based BDQA methods have achieved remarkable success, while the small scale of DQA datasets limits their performance. To address this issue, in this paper, we propose to adapt Contrastive Language-Image Pre-Training (CLIP), pre-trained on large-scale image-text pairs, to the BDQA task. Specifically, inspired by the fact that the human visual system understands images based on hierarchical features, we take global and local information of the dehazed image as the input of CLIP. To accurately map the input hierarchical information of dehazed images into the quality score, we tune both the vision branch and language branch of CLIP with prompt learning. Experimental results on two authentic DQA datasets demonstrate that our proposed approach, named CLIP-DQA, achieves more accurate quality predictions over existing BDQA methods. The code is available at https://github.com/JunFu1995/CLIP-DQA. Yirui Zeng, Jun Fu 0007, Hadi Amirpour, Huasheng Wang, Guanghui Yue 0001, Hantao Liu, Ying Chen 0011, Wei Zhou 0021 |
ISCAS | 8 |
| 2025 | Parameterized Diffusion Optimization Enabled Autoregressive Ordinal Regression for Diabetic Retinopathy Grading
Qinkai Yu, Wei Zhou 0021, Hantao Liu, Yanyu Xu 0001, Meng Wang 0038, Yitian Zhao, Huazhu Fu, Xujiong Ye, Yalin Zheng, Yanda Meng |
MICCAI (15) | 2 |
| 2025 | MCHM25: Multimedia Computing for Health and MedicineabstractRecent years have witnessed an unprecedented growth of multimodal data in healthcare, ranging from distributed sensors and medical imaging devices (MRI, CT, X-rays) to digital health platforms that integrate audio, video, 3D geometry, and clinical text. The increasing availability of such data presents significant opportunities for computer-aided diagnosis and intelligent healthcare solutions, yet also poses substantial challenges in multimodal integration, large-scale analysis, and real-world deployment. The 2nd International Workshop on Multimedia Computing for Health and Medicine (MCHM'25), held in conjunction with ACM Multimedia 2025, focuses on advanced multimedia computing techniques, including mobile and hardware solutions, for tackling real-world problems in healthcare. The workshop brings together researchers and practitioners in multimedia computing, artificial intelligence, and medicine to explore emerging methods, applications, and systems that have a direct impact on human health. Wei Zhou 0021, Hadi Amirpour, Li Yu 0004, Jungong Han, Richang Hong, Paul L. Rosin |
ACM Multimedia | 1 |
| 2025 | Perceptual Visual Quality Assessment in Multimedia CommunicationabstractThe rapid expansion of multimedia services, such as video streaming, video conferencing, virtual reality, and cloud gaming, makes maintaining and evaluating high perceptual visual quality essential for user experience and system competitiveness. However, visual content can degrade at multiple stages, including acquisition, compression, transmission, enhancement, and display, where suboptimal enhancement may also introduce artifacts and reduce perceived quality. The core challenge is to reliably measure and predict this perceived quality so that it can be maintained or improved. Perceptual Visual Quality Assessment (PVQA) addresses this by evaluating visual quality from the perspective of human subjects, through subjective studies and objective prediction models. Beyond humans, recent work also extends PVQA to machines and robots, where the goal is to preserve downstream task performance (e.g., segmentation accuracy and planning success) under distortions or bandwidth constraints. This tutorial provides a concise, practice-oriented overview of PVQA: fundamentals and human vision considerations; image and video quality assessment; methods for immersive/3D media; opportunities and challenges in the era of foundation models and GenAI; perceptual optimization loops that close the gap between assessment and decisions in coding, streaming, and embodied perception; and domain applications. Finally, we summarize the key concepts, toolchains, and future opportunities for PVQA to be used in modern multimedia communication. Wei Zhou 0021, Hadi Amirpour |
ACM Multimedia | 1 |
| 2025 | A Spatial Relationship Aware Dataset for RoboticsabstractRobotic task planning in real-world environments requires not only object recognition but also a nuanced understanding of spatial relationships between objects. We present a spatial-relationship-aware dataset of nearly 1,000 robot-acquired indoor images, annotated with object attributes, positions, and detailed spatial relationships. Captured using a Boston Dynamics Spot robot and labelled with a custom annotation tool, the dataset reflects complex scenarios with similar or identical objects and intricate spatial arrangements. We benchmark six state-of-the-art scene-graph generation models on this dataset, analysing their inference speed and relational accuracy. Our results highlight significant differences in model performance and demonstrate that integrating explicit spatial relationships into foundation models, such as ChatGPT 4o, substantially improves their ability to generate executable, spatially-aware plans for robotics. The dataset and annotation tool are publicly available at https://github.com/PengPaulWang/SpatialAwareRobotDataset, supporting further research in spatial reasoning for robotics. Peng Wang 0076, Minh Huy Pham, Wei Zhou 0021 |
ACM Multimedia | 4 |
| 2025 | Compressed Feature Quality Assessment: Dataset and BaselinesabstractThe widespread deployment of large models in resource-constrained environments has underscored the need for efficient transmission of intermediate feature representations. In this context, feature coding, which compresses features into compact bitstreams, becomes a critical component for scenarios involving feature transmission, storage, and reuse. However, this compression process inevitably introduces semantic degradation that is difficult to quantify with traditional metrics. To address this, we formalize the research problem of Compressed Feature Quality Assessment (CFQA), aiming to evaluate the semantic fidelity of compressed features. To advance CFQA research, we propose the first benchmark dataset, comprising 300 original features and 12000 compressed features derived from three vision tasks and four feature codecs. Task-specific performance degradation is provided as true semantic distortion for evaluating CFQA metrics. We systematically assess three widely used metrics -- MSE, cosine similarity, and Centered Kernel Alignment (CKA) -- in terms of their ability to capture semantic degradation. Our findings demonstrate the representativeness of the proposed dataset while underscoring the need for more sophisticated metrics capable of measuring semantic distortion in compressed features. This work advances the field by establishing a foundational benchmark and providing a critical resource for the community to explore CFQA. To foster further research, we release the dataset and all associated source code at https://github.com/chansongoal/Compressed-Feature-Quality-Assessment. Changsheng Gao, Wei Zhou 0021, Guosheng Lin, Weisi Lin |
ACM Multimedia | 2 |
| 2025 | EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion AssessmentabstractThe furnishing of multi-modal large language models (MLLMs) has led to the emergence of numerous benchmark studies, particularly those evaluating their perception and understanding capabilities. Among these, understanding image-evoked emotions aims to enhance MLLMs' empathy, with significant applications such as human-machine interaction and advertising recommendations. However, current evaluations of this MLLM capability remain coarse-grained, and a systematic and comprehensive assessment is still lacking. To this end, we introduce EEmo-Bench, a novel benchmark dedicated to the analysis of the evoked emotions in images across diverse content categories. Our core contributions include: 1) Regarding the diversity of the evoked emotions, we adopt an emotion ranking strategy and employ the Valence-Arousal-Dominance (VAD) as emotional attributes for emotional assessment. In line with this methodology, 1,960 images are collected and manually annotated. 2) We design four tasks to evaluate MLLMs' ability to capture the evoked emotions by single images and their associated attributes: Perception, Ranking, Description, and Assessment. Additionally, image-pairwise analysis is introduced to investigate the model's proficiency in performing joint and comparative analysis. In total, we collect 6,773 question-answer pairs and perform a thorough assessment on 19 commonly-used MLLMs. The results indicate that while some proprietary and large-scale open-source MLLMs achieve promising overall performance, the analytical capabilities in certain evaluation dimensions remain suboptimal. Our EEmo-Bench paves the path for further research aimed at enhancing the comprehensive perceiving and understanding capabilities of MLLMs concerning image-evoked emotions, which is crucial for machine-centric emotion perception and understanding. Our code and benchmark datasets are available at https://github.com/workerred/EEmo-Bench. Lancheng Gao, Ziheng Jia, Yunhao Zeng, Wei Sun 0029, Wei Zhou 0021, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 6 |
| 2025 | SDART: Spatial Dart AR Simulation with Hand-Tracked InputabstractWe present a physics-driven 3D dart-throwing interaction system for Apple Vision Pro (AVP), developed using Unity 6 engine and running in augmented reality (AR) mode on the device. The system utilizes the PolySpatial and Apple's ARKit software development kits (SDKs) to ensure hand input and tracking in order to intuitively spawn, grab, and throw virtual darts similar to real darts. The application benefits from physics simulations alongside the innovative no-controller input system of AVP to manipulate objects realistically in an unbounded spatial volume. By implementing spatial distance measurement, scoring logic, and recording user performance, this project enables user studies on quality of experience in interactive experiences. To evaluate the perceived quality and realism of the interaction, we conducted a subjective study with 10 participants using a structured questionnaire. The study measured various aspects of the user experience, including visual and spatial realism, control fidelity, depth perception, immersiveness, and enjoyment. Results indicate high mean opinion scores (MOS) across key dimensions. Milad Ghanbari, Wei Zhou 0021, Cosmin Stejerean, Christian Timmerer, Hadi Amirpour |
ACM Multimedia | 2 |
| 2025 | SVD: Spatial Video DatasetabstractStereoscopic video has long been the subject of research due to its ability to deliver immersive three-dimensional content to a wide range of applications. The dual-view format inherently provides binocular disparity cues that enhance depth perception and realism, making it indispensable for fields such as telepresence, 3D mapping, and robotic vision. Until recently, however, end-to-end pipelines for capturing, encoding, and viewing high-quality stereoscopic video were neither widely accessible nor optimized for consumer-grade devices. Today's smartphones, such as the iPhone Pro, and modern Head-Mounted Displays (HMDs) like the Apple Vision Pro, offer built-in support for stereoscopic video capture, hardware-accelerated encoding, and seamless playback on devices like the Apple Vision Pro and Meta Quest 3, which require minimal user intervention. Apple refers to this streamlined workflow as spatial Video. Making the full stereoscopic video process available to everyone has made new applications possible. Despite these advances, there remains a notable absence of publicly available datasets that include the complete spatial video pipeline on consumer platforms, hindering reproducibility and comparative evaluation of emerging algorithms. Mohammad Hossein Izadimehr, Milad Ghanbari, Guodong Chen 0004, Wei Zhou 0021, Xiaoshuai Hao, Mallesham Dasari, Christian Timmerer, Hadi Amirpour |
ACM Multimedia | 4 |
| 2025 | DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait RecognitionabstractRobust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent literature, these 2D representations fall short of capturing sufficient cues that can be exploited to handle viewpoint variations, and capture finer and meaningful details of gait. In this paper, we introduce a novel framework, termed DepthGait, that incorporates RGB-derived depth maps and silhouettes for enhanced gait recognition. Specifically, apart from the 2D silhouette representation of the human body, the proposed pipeline explicitly estimates depth maps from a given RGB image sequence and uses them as a new modality to capture discriminative features inherent in human locomotion. In addition, a novel multi-scale and cross-level fusion scheme has also been developed to bridge the modality gap between depth maps and silhouettes. Extensive experiments on standard benchmarks demonstrate that the proposed DepthGait achieves state-of-the-art performance compared to peer methods and attains an impressive mean rank-1 accuracy on the challenging datasets. Xinzhu Li, Juepeng Zheng, Yikun Chen, Xudong Mao, Guanghui Yue 0001, Wei Zhou 0021, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
ACM Multimedia | 6 |
| 2025 | VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question AnsweringabstractCross-video question answering presents significant challenges beyond traditional single-video understanding, particularly in establishing meaningful connections across video streams and managing the complexity of multi-source information retrieval. We introduce VideoForest, a novel framework that addresses these challenges through person-anchored hierarchical reasoning, enabling effective cross-video understanding without requiring end-to-end training. VideoForest integrates three key innovations: 1) a human-anchored feature extraction mechanism that employs ReID and tracking algorithms to establish robust spatiotemporal relationships across multiple video sources; 2) a multi-granularity spanning tree structure that hierarchically organizes visual content around person-level trajectories; and 3) a multi-agent reasoning framework that efficiently traverses this hierarchical structure to answer complex queries. To evaluate our method, we develop CrossVideoQA, a comprehensive benchmark specifically designed for person-centric cross-video analysis. Experimental results demonstrate VideoForest's superior performance in cross-video reasoning tasks, achieving 71.93% accuracy in person recognition, 83.75% in behavior analysis, and 51.67% in summarization and reasoning. Yiran Meng, Junhong Ye, Wei Zhou 0021, Guanghui Yue 0001, Xudong Mao, Ruomei Wang 0001, Baoquan Zhao |
ACM Multimedia | 3 |
| 2025 | MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World ScenariosabstractRapid advances in Artificial Intelligence Generated Content (AIGC) have enabled increasingly sophisticated face forgeries, posing a significant threat to social security. However, current Deepfake detection methods are limited by constraints in existing datasets, which lack the diversity necessary in real-world scenarios. Specifically, these data sets fall short in four key areas: unknown of advanced forgery techniques, variability of facial scenes, richness of real data, and degradation of real-world propagation. To address these challenges, we propose the Multi-dimensional Face Forgery Image (MFFI ) dataset, tailored for real-world scenarios. MFFI enhances realism based on four strategic dimensions: 1) Wider Forgery Methods; 2) Varied Facial Scenes; 3) Diversified Authentic Data; 4) Multi-level Degradation Operations. MFFI integrates 50 different forgery methods and contains 1024K image samples. Benchmark evaluations show that MFFI outperforms existing public datasets in terms of scene complexity, cross-domain generalization capability, and detection difficulty gradients. These results validate the technical advance and practical utility of MFFI in simulating real-world conditions. The dataset and additional details are publicly available at https://github.com/inclusionConf/MFFI. Changtao Miao, Weiwei Feng, Qi Chu 0001, Jianshu Li, Yunfeng Diao, Wei Zhou 0021, Joey Tianyi Zhou, Xiaoshuai Hao |
ACM Multimedia | 10 |
| 2025 | Evaluating Perceptual Color Preferences in Smartphone Photography: Dataset and ChallengesabstractInternational audience Zhihua Wang 0002, Weixia Zhang, Wei Zhou 0021, Xiaohong Liu 0001, Guangtao Zhai, Patrick Le Callet |
ACM Multimedia | 3 |
| 2025 | PhysLab: A Benchmark Dataset for Multi-Granularity Visual Parsing of Physics ExperimentsabstractVisual parsing of images and videos is critical for a wide range of real-world applications. However, progress in this field is constrained by limitations of existing datasets: (1) limited annotation diversity, which limits the support for diverse vision tasks within a unified dataset; (2) insufficient coverage of domains, particularly a lack of datasets tailored for educational scenarios; and (3) a lack of explicit procedural guidance, with weak logical rules and insufficient representation of a structured task process. To address these gaps, we introduce PhysLab, the first dataset that captures students conducting complex physics experiments. The dataset includes four representative experiments that feature diverse scientific instruments and rich human-object interaction (HOI) patterns. PhysLab comprises 620 long-form videos and provides multi-granularity annotations that support a variety of vision tasks, including action recognition, object detection, HOI analysis, etc. We establish baselines and perform extensive evaluations to highlight key challenges in the parsing of procedural educational videos. We expect PhysLab to serve as a valuable resource for advancing comprehensive visual parsing, facilitating intelligent classroom systems, and fostering closer integration among computer vision, multimedia, and educational technologies. The dataset and the evaluation toolkit are publicly available at https://github.com/ZMH-SDUST/PhysLab. Minghao Zou, Qingtian Zeng, Yongping Miao, Hantao Liu, Wei Zhou 0021 |
ACM Multimedia | 7 |
| 2025 | VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment: Methods and ResultsabstractThis paper reviews the VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment. The competition aims to foster innovation in both subjective and objective VQA techniques tailored to live broadcasting videos, addressing the unique challenges posed by live streaming impairments while emphasizing the evaluation of QoE. The grand challenge used live broadcasting database LBVD which consists of 1013 videos focusing on distortion in live broadcasting videos. The competition had 14 participants and 5 teams submitted valid solutions for the final testing phase. The proposed solutions have shown significant progress in areas such as combining traditional feature engineering with deep learning models, achieved state-of-the-art performances for LBVD. Team ATHENA-Live-QoE and Team HZX Force tied for the first position. The dataset can be found at https://github.com/cpf0079/LBVD. Wenqi Fei, Yuhua Zhang, MohammadAli Hamidi, Hadi Amirpour, Erjia Xiao, Zhenjie Su, Hao Cheng 0015, Yu Liu 0023, Wei Zhou 0021, Yanbiao Ma, Renjing Xu, Long Chen 0015, Xiaoshuai Hao, Yipo Huang, Tushar Shinde |
VCIP | 12 |
| 2025 | STACK: Spatial Tower Assembly using Controlled KineticsabstractThis paper presents a block stacking simulation developed for Apple Vision Pro (AVP) using Unity’s PolySpatial framework, designed to study both depth perception in spatial computing and physics comprehension of user-driven kinetic controls in augmented reality (AR). The simulation offers two interactive modes: a tower assembly mode and a removal mode. Each game session includes four stages with the virtual table positioned at various distances to observe user adaptation across varying virtual depths. User input is captured through eye tracking and hand tracking, and block behavior is handled by real-time physics simulation, which includes collision response, gravity, and mass-based interactions. The system supports two physics configurations: raw Unity physics and a modified variant with adjusted material and rigidbody parameters for improved stability and realism. It utilizes spatial computing features such as world anchoring to preserve spatial consistency and depth perception through stereoscopic rendering and dynamic shadows, so that users can better judge the spatial coordinates between virtual blocks and their physical surroundings. The simulation is intended to evaluate how 3D spatial rendering and physically realistic interactions contribute to immersion and task performance in AR environments. To assess user performance, the system records key interaction metrics to support analysis of learning progression, control accuracy, and adaptability across varying distances and physics configurations. This work contributes to the understanding of spatial and physics-based interaction design in AR and may inform future applications in education, simulation, and spatial gaming. Milad Ghanbari, Hadi Amirpour, Christian Timmerer, Mohammad Hossein Izadimehr, Wei Zhou 0021, Cosmin Stejerean |
VCIP | 5 |
| 2025 | Frequency-Aware Native Resolution Assessment of 8K Omnidirectional ImagesabstractOmnidirectional images (ODIs) serve as fundamental visual medium for presenting virtual reality (VR) contents, supporting fully immersive experiences through 360-degree scene representation. Typically, a high pixel density is essential for visual quality in VR environments, which in turn requires sufficiently high-resolution imagery to achieve. However, capturing native high-resolution ODIs requires expensive omnidirectional cameras with large sensors (e.g., Insta360 TITAN). An alternative approach is to use low-resolution cameras to acquire original images and then enhance their resolution via super-resolution algorithms. In this work, we explore whether super-resolution ODIs can be easily distinguished from native high-resolution ODIs at 8K scale. To this end, we firstly construct the Native Resolution Assessment of 8K Omnidirectional Images (NRA- 8KODI) dataset, whose native 8K ODIs are collected with an Insta360 TITAN camera and 8K super-resolution images are generated from SOTA open-sourced algorithms. Recognizing high-frequency signals are essential for differentiating non-native 8K ODIs, a frequency-aware model is designed to capture high-frequency details. Specially, to maintain high-frequency details kept in high-resolutions while reduce computational costs brought by high-resolutions, we propose a frequency-aware compressor module to suppress feature channels dominated by low-frequency details. Finally, our model achieves 97.2% accuracy in detecting non-native 8K ODIs, implying that super-resolution for ODIs can still be improved for visual experience in VR applications. Jingwen Hou, Zengliang Li, Jiebin Yan, Weide Liu, Yuming Fang 0001, Wei Zhou 0021 |
VCIP | 6 |
| 2025 | CVBench: Benchmarking and Comparing Video Generation with Large Multimodal ModelsabstractLarge multimodal models (LMMs) have revolutionized both text-to-video (T2V) generation and video-to-text (V2T) interpretation. However, despite these advancements, issues such as imperfect perceptual quality and inconsistent text-video alignment continue to limit the practical deployment of AI-generated videos (AIGVs). Consequently, there is a pressing need for a reliable benchmark and automatic evaluation framework tailored for AIGVs. To this end, we propose CVBench, the largest and most comprehensive dataset for Comparative Video Benchmarking, including 60K video pairs generated by 30 state-of-the-art T2V models and 600K pairwise comparisons annotated with over 1.7 million human judgments from perspectives of both perceptual quality and text-video correspondence. This dataset enables bidirectional benchmarking and evaluation of both T2V generation models and V2T interpretation models. Based on CVBench, we propose VComp, a novel LMM-based evaluation metric that captures fine-grained quality differences from multiple perspectives for pairwise comparison at both the instance level and model level. Extensive experiments show that VComp achieves state-of-the-art alignment with human preferences. Both the CVBench dataset and VComp metric will be available at https://github.com/IntMeGroup/CVBench. Huiyu Duan, Yuke Xing, Wei Zhou 0021, Guangtao Zhai, Xiongkuo Min |
VCIP | 4 |
| 2025 | Is there a relationship between Mean Opinion Score (MOS) and Just Noticeable Difference (JND)?abstractEvaluating perceived video quality is essential for ensuring high Quality of Experience (QoE) in modern streaming applications. While existing subjective datasets and Video Quality Metrics (VQMs) cover a broad quality range, many practical use cases—especially for premium users—focus on high-quality scenarios requiring finer granularity. Just Noticeable Difference (JND) has emerged as a key concept for modeling perceptual thresholds in these high-end regions and plays an important role in perceptual bitrate ladder construction. However, the relationship between JND and the more widely used Mean Opinion Score (MOS) remains unclear. In this paper, we conduct a Degradation Category Rating (DCR) subjective study based on an existing JND dataset to examine how MOS corresponds to the 75% Satisfied User Ratio (SUR) points of the 1stand 2ndJNDs. We find that while MOS values at JND points generally align with theoretical expectations (e.g., 4.75 for the 75% SUR of the 1stJND), the reverse mapping—from MOS to JND—is ambiguous due to overlapping confidence intervals across PVS indices. Statistical significance analysis further shows that DCR studies with limited participants may not detect meaningful differences between reference and JND videos. Hadi Amirpour, Wei Zhou 0021, Patrick Le Callet |
VCIP | 3 |
| 2025 | Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai |
Sci. China Inf. Sci. | 24 |
| 2025 | Physically-guided open vocabulary segmentation with weighted patched alignment loss
Weide Liu, Jieming Lou, Wei Zhou 0021, Jun Cheng 0003, Xulei Yang |
Neurocomputing | 4 |
| 2025 | Integrating large foundation models into multimodal named entity recognition with evidential fusion
Weide Liu, Xiaoyang Zhong, Jingwen Hou, Haozhe Huang, Wei Zhou 0021, Yuming Fang 0001 |
Neurocomputing | 6 |
| 2025 | CLIP-DQA V2: Exploring CLIP for Dehazed Image Quality Assessment From a Fragment-Level Perspective
Yirui Zeng, Jun Fu 0007, Guanghui Yue 0001, Hantao Liu, Wei Zhou 0021 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Adaptive Spatiotemporal Graph Transformer Network for Action Quality AssessmentabstractLong video action quality assessment (AQA) aims to evaluate the performance of long-term actions depicted in a video and produce an overall assessment for action quality. A video of long-term actions often contains more complicated temporal and spatial information than that of short-term actions. However, existing approaches that segment a video into individual clips for independent analysis potentially disrupt the narrative flow and diminish contextual details within and across clips, impeding comprehensive video understanding. To address this challenge, we propose an adaptive spatiotemporal graph transformer network (ASGTN) that combines multiple graph structures and transformer attention mechanisms to capture both local and global contextual information within and across clips in a long video. Specifically, the adaptive spatiotemporal graph (ASG) combines a spatial graph branch, designed to enrich the local nuanced spatiotemporal relations within an individual clip, and a temporal graph branch, tailored to dynamically learn the semantic context across different clips. Furthermore, a transformer encoder is integrated to amplify the global dependencies across clips in the entire video. This structure is designed to preserve narrative coherence and maintain essential contextual details in video-level features. Finally, we employ a level-focused decoder to predict the action quality score distribution. Experiments demonstrate that our model achieves state-of-the-art results on popular AQA datasets. Our code is available athttps://github.com/jiangliu5/ASGTN_AQA. Huasheng Wang, Wei Zhou 0021, Katarzyna Stawarz, Padraig Corcoran, Ying Chen 0011, Hantao Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Perception-Oriented Bidirectional Attention Network for Image Super-Resolution Quality AssessmentabstractMany super-resolution (SR) algorithms have been proposed to increase image resolution. However, full-reference (FR) image quality assessment (IQA) metrics for comparing and evaluating different SR algorithms are limited. In this work, we propose the Perception-oriented Bidirectional Attention Network (PBAN) for image SR FR-IQA, which is composed of three modules: an image encoder module, a perception-oriented bidirectional attention (PBA) module, and a quality prediction module. First, we encode the input images for feature representations. Inspired by the characteristics of the human visual system, we then construct the perception-oriented PBA module. Specifically, different from existing attention-based SR IQA methods, we conceive a Bidirectional Attention to bidirectionally construct visual attention to distortion, which is consistent with the generation and evaluation processes of SR images. To further guide the quality assessment towards the perception of distorted information, we propose Grouped Multi-scale Deformable Convolution, enabling the proposed method to adaptively perceive distortion. Moreover, we design Sub-information Excitation Convolution to direct visual perception to both sub-pixel and sub-channel attention. Finally, the quality prediction module is exploited to integrate quality-aware features and regress quality scores. Extensive experiments demonstrate that our proposed PBAN outperforms state-of-the-art quality assessment methods. Xiaoyuan Yang 0003, Guanghui Yue 0001, Jun Fu 0007, Qiuping Jiang, Xu Jia 0012, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
IEEE Trans. Image Process. | 9 |
| 2025 | Subjective and Objective Quality Assessment of Colonoscopy VideosabstractCaptured colonoscopy videos usually suffer from multiple real-world distortions, such as motion blur, low brightness, abnormal exposure, and object occlusion, which impede visual interpretation. However, existing works mainly investigate the impacts of synthesized distortions, which differ from real-world distortions greatly. This research aims to carry out an in-depth study for colonoscopy Video Quality Assessment (VQA). In this study, we advance this topic by establishing both subjective and objective solutions. Firstly, we collect 1,000 colonoscopy videos with typical visual quality degradation conditions in practice and construct a multi-attribute VQA database. The quality of each video is annotated by subjective experiments from five distortion attributes (i.e., temporal-spatial visibility, brightness, specular reflection, stability, and utility), as well as an overall perspective. Secondly, we propose a Distortion Attribute Reasoning Network (DARNet) for automatic VQA. DARNet includes two streams to extract features related to spatial and temporal distortions, respectively. It adaptively aggregates the attribute-related features through a multi-attribute association module to predict the quality score of each distortion attribute. Motivated by the observation that the rating behaviors for all attributes are different, a behavior guided reasoning module is further used to fuse the attribute-aware features, resulting in the overall quality. Experimental results on the constructed database show that our DARNet correlates well with subjective ratings and is superior to nine state-of-the-art methods. Guanghui Yue 0001, Jingfeng Du, Tianwei Zhou, Wei Zhou 0021, Weisi Lin |
IEEE Trans. Medical Imaging | 5 |
| 2025 | VQM4HAS: A Real-Time Quality Metric for HEVC Videos in HTTP Adaptive StreamingabstractIn HTTP Adaptive Streaming (HAS), a video is encoded at various bitrate-resolution pairs, collectively known as the bitrate ladder, allowing users to select the most suitable representation based on their network conditions. Optimizing this set of pairs to enhance the Quality of Experience (QoE) requires accurately measuring the quality of these representations. VMAF and ITU-T's P.1204.3 are highly reliable metrics for assessing the quality of representations in HAS. However, in practice, using these metrics for optimization is often impractical for live streaming applications due to their high computational costs and the large number of bitrate-resolution pairs in the bitrate ladder that need to be evaluated. To address their high complexity, our paper introduces a new method calledVQM4HAS, which extractslow-complexityfeatures, including ($i$) video complexity features, ($ii$) frame-level encoding statistics logged during the encoding process, and ($iii$) lightweight video quality metrics. These extracted features are then fed into a regression model to predict VMAF or P.1204.3. TheVQM4HASmodel is designed to operate on a per bitrate-resolution pair, per-resolution, and cross-representation basis, optimizing quality predictions across different scenarios. Our experimental results demonstrate thatVQM4HASachieves a high correlation with VMAF and P.1204.3, with Pearson correlation coefficients (PCC) ranging from 0.95 to 0.96 for VMAF and 0.97 to 0.99 for P.1204.3, depending on the resolution. Despite achieving a high correlation with VMAF and P.1204.3,VQM4HASexhibits significantly less complexity than both metrics, with 98% and 99% less complexity for VMAF and P.1204.3, respectively, making it suitable for live streaming scenarios. We also conduct a feature importance analysis to further reduce the complexity of the proposed method. Furthermore, we evaluate the effectiveness of our method by using it to predict subjective quality scores. The results show thatVQM4HASachieves a higher correlation with subjective scores at various resolutions despite its minimal complexity. The source code is available athttps://github.com/cd-athena/VQM4HAS. Hadi Amirpour, Wei Zhou 0021, Patrick Le Callet, Christian Timmerer |
IEEE Trans. Multim. | 3 |
| 2025 | No-Reference Point Cloud Quality Assessment via Graph Convolutional NetworkabstractThree-dimensional (3D) point cloud, as an emerging visual media format, is increasingly favored by consumers as it can provide more realistic visual information than two-dimensional (2D) data. Similar to 2D plane images and videos, point clouds inevitably suffer from quality degradation and information loss through multimedia communication systems. Therefore, automatic point cloud quality assessment (PCQA) is of critical importance. In this work, we propose a novel no-reference PCQA method by using a graph convolutional network (GCN) to characterize the mutual dependencies of multi-view 2D projected image contents. The proposed GCN-based PCQA (GC-PCQA) method contains three modules, i.e., multi-view projection, graph construction, and GCN-based quality prediction. First, multi-view projection is performed on the test point cloud to obtain a set of horizontally and vertically projected images. Then, a perception-consistent graph is constructed based on the spatial relations among different projected images. Finally, reasoning on the constructed graph is performed by GCN to characterize the mutual dependencies and interactions between different projected images, and aggregate feature information of multi-view projected images for final quality prediction. Experimental results on two publicly available benchmark databases show that our proposed GC-PCQA can achieve superior performance than state-of-the-art quality assessment metrics. Qiuping Jiang, Wei Zhou 0021, Feng Shao 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2025 | Unsupervised Low-Light Image Enhancement With Self-Paced LearningabstractLow-light image enhancement (LIE) aims to restore images taken under poor lighting conditions, thereby extracting more information and details to robustly support subsequent visual tasks. While past deep learning (DL)-based techniques have achieved certain restoration effects, these existing methods treat all samples equally, ignoring the fact that difficult samples may be detrimental to the network's convergence at the initial training stages of network training. In this paper, we introduce a self-paced learning (SPL)-based LIE method named SPNet, which consists of three key components: the feature extraction module (FEM), the low-light image decomposition module (LIDM), and a pre-trained denoise module. Specifically, for a given low-light image, we first input the image, its pseudo-reference image, and its histogram-equalized version into the FEM to obtain preliminary features. Second, to avoid ambiguities during the early stages of training, these features are then adaptively fused via an SPL strategy and processed for retinex decomposition via LIDM. Third, we enhance the network performance by constraining the gradient prior relationship between the illumination components of the images. Finally, a pre-trained denoise module reduces noise inherent in LIE. Extensive experiments on nine public datasets reveal that the proposed SPNet outperforms eight state-of-the-art DL-based methods in both qualitative and quantitative evaluations and outperforms three conventional methods in quantitative assessments. Yu Luo 0004, Xuanrong Chen, Jie Ling 0002, Chao Huang 0001, Wei Zhou 0021, Guanghui Yue 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Subjective Quality Assessment of Thermal Infrared ImagesabstractThermal infrared images (TIIs) can be distorted by multiple factors, resulting in noise, low contrast, limited dynamic range, and fuzziness, which greatly impede their usefulness. It is crucial to evaluate the quality of TIIs. Unfortunately, there have been very few attempts to study this problem. In this study, we collected 1,000 authentically distorted TIIs using thermal infrared acquisition equipment and conducted strict subjective experiments to obtain a thermal infrared image quality assessment (IQA) database. Each image’s quality score was obtained under strict scoring rules. Finally, we investigated the feasibility of several no-reference (NR) IQA methods in quality assessment of TIIs. We found that existing NR-IQA methods achieve ordinary performance in such a task, and there is an urgent need to develop a specific IQA methods for TIIs. The findings together with the constructed database are expected to pave the way for the development of more advanced IQA methods for further development of this field. Guanghui Yue 0001, Jinxia Zhang, Zhaofei Xu, Shuigen Wang, Tianwei Zhou, Yuanhao Gong, Wei Zhou 0021 |
ICIP | 8 |
| 2024 | LiSD: An Efficient Multi-Task Learning Framework For Lidar Segmentation and DetectionabstractWith the rapid proliferation of autonomous driving, there has been a heightened focus on the research of lidar-based 3D semantic segmentation and object detection methodologies, aiming to ensure the safety of traffic participants. In recent decades, learning-based approaches have emerged, demonstrating remarkable performance gains in comparison to conventional algorithms. However, the segmentation and detection tasks have traditionally been examined in isolation to achieve the best precision. To this end, we propose an efficient multi-task learning framework named LiSD which can address both segmentation and detection tasks, aiming to optimize the overall performance. Our proposed LiSD is a voxel-based encoder-decoder framework that contains a hierarchical feature collaboration module and a holistic information aggregation module. Different integration methods are adopted to keep sparsity in segmentation while densifying features for query initialization in detection. Besides, cross-task information is utilized in an instance-aware refinement module to obtain more accurate predictions. Experimental results on the nuScenes dataset and Waymo Open Dataset demonstrate the effectiveness of our proposed model. It is worth noting that LiSD achieves the state-of-the-art performance of $83.3 \% \mathrm{mIoU}$ on the nuScenes segmentation benchmark for lidar-only methods. Jiahua Xu 0001, Si Zuo, Chenfeng Wei, Wei Zhou 0021 |
ICIP | 4 |
| 2024 | Deep Bi-directional Attention Network for Image Super-Resolution Quality AssessmentabstractThere has emerged a growing interest in exploring efficient quality assessment algorithms for image super-resolution (SR). However, employing deep learning techniques, especially dual-branch algorithms, to automatically evaluate the visual quality of SR images remains challenging. Existing SR image quality assessment (IQA) metrics based on two-stream networks lack interactions between branches. To address this, we propose a novel full-reference IQA (FR-IQA) method for SR images. Specifically, producing SR images and evaluating how close the SR images are to the corresponding HR references are separate processes. Based on this consideration, we construct a deep Bidirectional Attention Network (BiAtten-Net) that dynamically deepens visual attention to distortions in both processes, which aligns well with the human visual system (HVS). Experiments on public SR quality databases demonstrate the superiority of our proposed BiAtten-Net over state-of-the-art quality assessment methods. In addition, the visualization results and ablation study show the effectiveness of bi-directional attention. Xiaoyuan Yang 0003, Jun Fu 0007, Guanghui Yue 0001, Wei Zhou 0021 |
ICME | 5 |
| 2024 | Perceptual Crack Detection for Rendered 3D Textured MeshesabstractRecent years have witnessed many advancements in the applications of 3D textured meshes. As the demand continues to rise, evaluating the perceptual quality of this new type of media content becomes crucial for quality assurance and optimization purposes. Different from traditional image quality assessment, crack is an annoying artifact specific to rendered 3D meshes that severely affects their perceptual quality. In this work, we make one of the first attempts to propose a novel Perceptual Crack Detection (PCD) method for detecting and localizing crack artifacts in rendered meshes. Specifically, motivated by the characteristics of the human visual system (HVS), we adopt contrast and Laplacian measurement modules to characterize crack artifacts and differentiate them from other undesired artifacts. Extensive experiments on large-scale public datasets of 3D textured meshes demonstrate effectiveness and efficiency of the proposed PCD method in correct localization and detection of crack artifacts. Moreover, to quantify the performance of the proposed detection method and validate its effectiveness, we propose a simple yet effective weighting mechanism to incorporate the resulting crack map into classical quality assessment (QA) models, which creates significant performance improvement in predicting the perceptual image quality when tested on public datasets of static 3D textured meshes. A software release of the proposed method is publicly available at: https://github.com/arshafiee/crack-detection-VVM Armin Shafiee Sarvestani, Wei Zhou 0021, Zhou Wang 0001 |
QoMEX | 2 |
| 2024 | Neural image re-exposure
Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Dong Wang 0004, Lihe Zhang, Bolun Zheng, Wei Zhou 0021, Huchuan Lu |
Comput. Vis. Image Underst. | 7 |
| 2024 | Bayesian graph convolutional network for traffic prediction
Jun Fu 0007, Wei Zhou 0021, Zhibo Chen 0001 |
Neurocomputing | 2 |
| 2024 | Vision-Language Consistency Guided Multi-Modal Prompt Learning for Blind AI Generated Image Quality AssessmentabstractRecently, textual prompt tuning has shown inspirational performance in adapting Contrastive Language-Image Pre-training (CLIP) models to natural image quality assessment. However, such uni-modal prompt learning method only tunes the language branch of CLIP models. This is not enough for adapting CLIP models to AI generated image quality assessment (AGIQA) since AGIs visually differ from natural images. In addition, the consistency between AGIs and user input text prompts, which correlates with the perceptual quality of AGIs, is not investigated to guide AGIQA. In this letter, we propose vision-language consistency guided multi-modal prompt learning for blind AGIQA, dubbed CLIP-AGIQA. Specifically, we introduce learnable textual and visual prompts in language and vision branches of CLIP models, respectively. Moreover, we design a text-to-image alignment quality prediction task, whose learned vision-language consistency knowledge is used to guide the optimization of the above multi-modal prompts. Experimental results on two public AGIQA datasets demonstrate that the proposed method outperforms state-of-the-art quality assessment models. Jun Fu 0007, Wei Zhou 0021, Qiuping Jiang, Hantao Liu, Guangtao Zhai |
IEEE Signal Process. Lett. | 2 |
| 2024 | Boundary Refinement Network for Colorectal Polyp Segmentation in Colonoscopy ImagesabstractPrecise polyp segmentation is vitally essential for detection and diagnosis of early colorectal cancer. Recent advances in artificial intelligence have brought infinite possibilities for this task. However, polyps usually vary greatly in shape and size and contain ambiguous boundary, bringing tough challenges to precise segmentation. In this letter, we introduce a novel Boundary Refinement Network (BRNet) for polyp segmentation. To be specific, we first introduce a boundary generation module (BGM) to generate boundary map by fusing both low-level spatial details and high-level concepts. Then, we utilize the boundary-guided refinement module to refine the polyp-aware features at each layer with the help of boundary cues from the BGM and the prediction from the adjacent high layer. Through top-down deep supervision, our BRNet can localize the polyp regions accurately with clear boundary. Extensive experiments are carried out on five datasets, and the results indicate the effectiveness of our BRNet over seven recently reported methods. Guanghui Yue 0001, Yuanyan Li, Wenchao Jiang, Wei Zhou 0021, Tianwei Zhou |
IEEE Signal Process. Lett. | 4 |
| 2024 | Dynamic Hypergraph Convolutional Network for No-Reference Point Cloud Quality AssessmentabstractWith the rapid advancement of three-dimensional (3D) sensing technology, point cloud has emerged as one of the most important approaches for representing 3D data. However, quality degradation inevitably occurs during the acquisition, transmission, and process of point clouds. Therefore, point cloud quality assessment (PCQA) with automatic visual quality perception is particularly critical. In the literature, the graph convolutional networks (GCNs) have achieved certain performance in point cloud-related tasks. However, they cannot fully characterize the nonlinear high-order relationship of such complex data. In this paper, we propose a novel no-reference (NR) PCQA method with hypergraph learning. Specifically, a dynamic hypergraph convolutional network (DHCN) composing of a projected image encoder, a point group encoder, a dynamic hypergraph generator, and a perceptual quality predictor, is devised. First, a projected image encoder and a point group encoder are used to extract feature representations from projected images and point groups, respectively. Then, using the feature representations obtained by the two encoders, dynamic hypergraphs are generated during each iteration, aiming to constantly update the interactive information between the vertices of hypergraphs. Finally, we design the perceptual quality predictor to conduct quality reasoning on the generated hypergraphs. By leveraging the interactive information among hypergraph vertices, feature representations are well aggregated, resulting in a notable improvement in the accuracy of quality pediction. Experimental results on several point cloud quality assessment databases demonstrate that our proposed DHCN can achieve state-of-the-art performance. The code will be available at:https://github.com/chenwuwq/DHCN. Qiuping Jiang, Wei Zhou 0021, Long Xu 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Blind Image Quality Assessment via Adaptive Graph AttentionabstractRecent advancements in blind image quality assessment (BIQA) are primarily propelled by deep learning technologies. While leveraging transformers can effectively capture long-range dependencies and contextual details in images, the significance of local information in image quality assessment can be undervalued. To address this challenging problem, we propose a novel feature enhancement framework tailored for BIQA. Specifically, we devise an Adaptive Graph Attention (AGA) module to simultaneously augment both local and contextual information. It not only refines the post-transformer features into an adaptive graph, facilitating local information enhancement, but also exploits interactions amongst diverse feature channels. The proposed technique can better reduce redundant information introduced during feature updates compared to traditional convolution layers, streamlining the self-updating process for feature maps. Experimental results show that our proposed model outperforms state-of-the-art BIQA models in predicting the perceived quality of images. The code of the model will be made publicly available. Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Hantao Liu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Dual-Constraint Coarse-to-Fine Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) is an important yet challenging task, with great application values in industrial defect detection, medical care, etc. The challenges mainly come from the high intrinsic similarities between target objects and background. In this paper, inspired by the biological studies that object detection consists of two steps, i.e., search and identification, we propose a novel framework, named DCNet, for accurate COD. DCNet explores candidate objects and extra object-related edges through two constraints (object area and boundary) and detects camouflaged objects in a coarse-to-fine manner. Specifically, we first exploit an area-boundary decoder (ABD) to obtain initial region cues and boundary cues simultaneously by fusing multi-level features of the backbone. Then, an area search module (ASM) is embedded into each level of the backbone to adaptively search coarse regions of objects with the assistance of region cues from the ABD. After the ASM, an area refinement module (ARM) is utilized to identify fine regions of objects by fusing adjacent-level features with the guidance of boundary cues. Through the deep supervision strategy, DCNet can finally localize the camouflaged objects precisely. Extensive experiments on three benchmark COD datasets demonstrate that our DCNet is superior to 12 state-of-the-art COD methods. In addition, DCNet shows promising results on two COD-related tasks, i.e., industrial defect detection and polyp segmentation. Guanghui Yue 0001, Houlu Xiao, Hai Xie, Tianwei Zhou, Wei Zhou 0021, Weiqing Yan, Baoquan Zhao, Tianfu Wang 0001, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Perceptual Depth Quality Assessment of Stereoscopic Omnidirectional ImagesabstractDepth perception plays an essential role in the viewer experience for immersive virtual reality (VR) visual environments. However, previous research investigations in the depth quality of 3D/stereoscopic images are rather limited, and in particular, are largely lacking for 3D viewing of 360-degree omnidirectional content. In this work, we make one of the first attempts to develop an objective quality assessment model named depth quality index (DQI) for efficient no-reference (NR) depth quality assessment of stereoscopic omnidirectional images. Motivated by the perceptual characteristics of the human visual system (HVS), the proposed DQI is built upon multi-color-channel, adaptive viewport selection, and interocular discrepancy features. Experimental results demonstrate that the proposed method outperforms state-of-the-art image quality assessment (IQA) and depth quality assessment (DQA) approaches in predicting the perceptual depth quality when tested using both single-viewport and omnidirectional stereoscopic image databases. Furthermore, we demonstrate that combining the proposed depth quality model with existing IQA methods significantly boosts the performance in predicting the overall quality of 3D omnidirectional images. Wei Zhou 0021, Zhou Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Blind Quality Assessment of Dense 3D Point Clouds with Structure Guided ResamplingabstractObjective quality assessment of three-dimensional (3D) point clouds is essential for the development of immersive multimedia systems in real-world applications. Despite the success of perceptual quality evaluation for 2D images and videos, blind/no-reference metrics are still scarce for 3D point clouds with large-scale irregularly distributed 3D points. Therefore, in this article, we propose an objective point cloud quality index with Structure Guided Resampling (SGR) to automatically evaluate the perceptually visual quality of dense 3D point clouds. The proposed SGR is a general-purpose blind quality assessment method without the assistance of any reference information. Specifically, considering that the human visual system is highly sensitive to structure information, we first exploit the unique normal vectors of point clouds to execute regional pre-processing that consists of keypoint resampling and local region construction. Then, we extract three groups of quality-related features, including (1) geometry density features, (2) color naturalness features, and (3) angular consistency features. Both the cognitive peculiarities of the human brain and naturalness regularity are involved in the designed quality-aware features that can capture the most vital aspects of distorted 3D point clouds. Extensive experiments on several publicly available subjective point cloud quality databases validate that our proposed SGR can compete with state-of-the-art full-reference, reduced-reference, and no-reference quality assessment algorithms. Wei Zhou 0021, Qi Yang 0003, Qiuping Jiang, Guangtao Zhai, Weisi Lin |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | In-vehicle Performance and Distraction for Midair and Touch Directional GesturesabstractWe compare the performance and level of distraction of expressive directional gesture input in the context of in-vehicle system commands. Center console touchscreen swipes and midair swipe-like movements are tested in 8-directions, with 8-button touchscreen tapping as a baseline. Participants use these input methods for intermittent target selections while performing the Lane Change Task in a virtual driving simulator. Input performance is measured with time and accuracy, cognitive load with deviation of lane position and speed, and distraction from frequency of off-screen glances. Results show midair gestures were less distracting and faster, but with lower accuracy. Touchscreen swipes and touchscreen tapping are comparable across measures. Our work provides empirical evidence for vehicle interface designers and manufacturers considering midair or touch directional gestures for centre console input. Arman Hafizi, Jay Henderson, Ali Neshati, Wei Zhou 0021, Edward Lank, Daniel Vogel 0001 |
CHI | 4 |
| 2023 | Subjective Quality Assessment of Enhanced Retinal ImagesabstractMany retinal images sometimes suffer from uneven illumination, which influences the analysis and diagnosis of retinal diseases. To improve the image quality of those retinal images, one feasible solution is to utilize low-light image enhancement (LIE) algorithms. However, how to evaluate the perceptual quality of enhanced retinal images (ERIs) generated by different LIE algorithms remains a challenging problem. In this paper, we conduct subjective experiments to investigate the quality assessment of ERIs. First, we collect 250 retinal images with the authentic low-light distortion, and then adopt eight LIE algorithms to produce 2000 ERIs. Second, a subjective experiment is conducted, resulting in the proposed Enhanced Retinal Image Quality Assessment Database (ERIQAD). Finally, we test some well-known no reference image quality assessment (NR IQA) methods on our proposed ERIQAD. Experimental results demonstrate that existing mainstream NR IQA methods merely achieve ordinary performance to predict the perceptual quality of ERIs. Guanghui Yue 0001, Shaoping Zhang, Tianwei Zhou, Wei Zhou 0021 |
ICIP | 6 |
| 2023 | Blind Omnidirectional Image Quality Assessment: Integrating Local Statistics and Global SemanticsabstractOmnidirectional image quality assessment (OIQA) aims to predict the perceptual quality of omnidirectional images that cover the whole 180×360° viewing range of the visual environment. Here we propose a blind/no-reference OIQA method named Local Statistics and Global Semantics metric (LSGS) that bridges the gap between low-level statistics and high-level semantics of omnidirectional images. Specifically, statistic and semantic features are extracted in separate paths from multiple local viewports and the hallucinated global omnidirectional image, respectively. A quality regression along with a weighting process is then followed that maps the extracted quality-aware features to a perceptual quality prediction. Experimental results demonstrate that the proposed LSGS method offers highly competitive performance against state-of-the-art methods. Wei Zhou 0021, Zhou Wang 0001 |
ICIP | 1 |
| 2023 | Reduced-Reference Quality Assessment of Point Clouds via Content-Oriented Saliency ProjectionabstractMany dense 3D point clouds have been exploited to represent visual objects instead of traditional images or videos. To evaluate the perceptual quality of various point clouds, in this letter, we propose a novel and efficient Reduced-Reference quality metric for point clouds, which is based on Content-oriented sAliency Projection (RR-CAP). Specifically, we make the first attempt to simplify reference and distorted point clouds into projected saliency maps with a downsampling operation. Through this process, we tackle the issue of transmitting large-volume original point clouds to end-users for quality assessment. Then, motivated by the characteristics of the human visual system (HVS), the objective quality scores of distorted point clouds are produced by combining content-oriented similarity and statistical correlation measurements. Finally, extensive experiments are conducted on SJTU-PCQA and WPC databases. The experiment results demonstrate that our proposed algorithm outperforms existing reduced-reference and no-reference quality metrics, and significantly reduces the performance gap between state-of-the-art full-reference quality assessment methods. In addition, we show the performance variation of each proposed technical component by ablation tests. Wei Zhou 0021, Guanghui Yue 0001, Ruizeng Zhang, Yipeng Qin, Hantao Liu |
IEEE Signal Process. Lett. | 1 |
| 2023 | LIQA: Lifelong Blind Image Quality AssessmentabstractThe image distortions are complex and dynamically changing in the real-world scenario, due to the fast development of the image processing system. The blind image quality assessment (BIQA) models may encounter the challenge of processing images with distortion types never seen before deployment. However, existing BIQA models generally cannot evolve with unseen distortion types adaptively, which greatly limits the deployment and application of BIQA models in real-world scenarios. To address this problem, we propose a novel Lifelong blind Image Quality Assessment (LIQA) approach, targeting to achieve the lifelong learning of BIQA. Without accessing to previous training data, our proposed LIQA can not only learn new knowledge, but also mitigate the catastrophic forgetting of learned knowledge. Specifically, we adopt the Split-and-Merge distillation strategy to train a single-head network that makes task-agnostic predictions. In the split stage, we first employ a distortion-specific generator to generate pseudo features of each previously seen distortion. Then, we utilize an auxiliary multi-head regression network to keep the response of each distortion. In the merge stage, we replay the pseudo features and use the pseudo labels generated by the auxiliary multi-head network to distill the knowledge of the multiple heads, which can build the final regression single head. Extensive experiments demonstrate that LIQA can perform well in handling both inner-dataset distortion shift and cross-dataset distortion shift. More importantly, our model can achieve stable performance even if the task sequences are long. Jianzhao Liu, Wei Zhou 0021, Xin Li 0082, Jiahua Xu 0001, Zhibo Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | GraphIQA: Learning Distortion Graph Representations for Blind Image Quality AssessmentabstractA good distortion representation is crucial for the success of deep blind image quality assessment (BIQA). However, most previous methods do not effectively model the relationship between distortions or the distribution of samples with the same distortion type but different distortion levels. In this work, we start from the analysis of the relationship between perceptual image quality and distortion-related factors, such as distortion types and levels. Then, we propose a Distortion Graph Representation (DGR) learning framework for IQA, named GraphIQA, in which each distortion is represented as a graph,i.e., DGR. One can distinguish distortion types by learning the contrast relationship between these different DGRs, and can infer the ranking distribution of samples from different levels in a DGR. Specifically, we develop two sub-networks to learn the DGRs: a) Type Discrimination Network (TDN) that aims to embed DGR into a compact code for better discriminating distortion types and learning the relationship between types; b) Fuzzy Prediction Network (FPN) that aims to extract the distributional characteristics of the samples in a DGR and predicts fuzzy degrees based on a Gaussian prior. Experiments show that our GraphIQA achieves state-of-the-art performance on many benchmark datasets of both synthetic and authentic distortions. Simeng Sun, Tao Yu 0012, Jiahua Xu 0001, Wei Zhou 0021, Zhibo Chen 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | RTN: Reinforced Transformer Network for Coronary CT Angiography Vessel-level Image Quality Assessment
Yiting Lu, Jun Fu 0007, Xin Li 0082, Wei Zhou 0021, Sen Liu 0001, Wei Wu 0021, Congfu Jia, Zhibo Chen 0001 |
MICCAI (1) | 4 |
| 2022 | Adaptive Hypergraph Convolutional Network for No-Reference 360-degree Image Quality AssessmentabstractIn no-reference 360-degree image quality assessment (NR 360IQA), graph convolutional networks (GCNs), which model interactions between viewports through graphs, have achieved impressive performance. However, prevailing GCN-based NR 360IQA methods suffer from three main limitations. First, they only use high-level features of the distorted image to regress the quality score, while the human visual system scores the image based on hierarchical features. Second, they simplify complex high-order interactions between viewports in a pairwise fashion through graphs. Third, in the graph construction, they only consider the spatial location of the viewport, ignoring its content characteristics. Accordingly, to address these issues, we propose an adaptive hypergraph convolutional network for NR 360IQA, denoted as AHGCN. Specifically, we first design a multi-level viewport descriptor for extracting hierarchical representations from viewports. Then, we model interactions between viewports through hypergraphs, where each hyperedge connects two or more viewports. In the hypergraph construction, we build a location-based hyperedge and a content-based hyperedge for each viewport. Experimental results on two public 360IQA databases demonstrate that our proposed approach has a clear advantage over state-of-the-art full-reference and no-reference IQA models. Jun Fu 0007, Chen Hou, Wei Zhou 0021, Jiahua Xu 0001, Zhibo Chen 0001 |
ACM Multimedia | 3 |
| 2022 | Quality Assessment of Image Super-Resolution: Balancing Deterministic and Statistical FidelityabstractThere has been a growing interest in developing image super-resolution (SR) algorithms that convert low-resolution (LR) to higher resolution images, but automatically evaluating the visual quality of super-resolved images remains a challenging problem. Here we look at the problem of SR image quality assessment (SR IQA) in a two-dimensional (2D) space of deterministic fidelity (DF) versus statistical fidelity (SF). This allows us to better understand the advantages and disadvantages of existing SR algorithms, which produce images at different clusters in the 2D space of (DF, SF). Specifically, we observe an interesting trend from more traditional SR algorithms that are typically inclined to optimize for DF while losing SF, to more recent generative adversarial network (GAN) based approaches that by contrast exhibit strong advantages in achieving high SF but sometimes appear weak at maintaining DF. Furthermore, we propose an uncertainty weighting scheme based on content-dependent sharpness and texture assessment that merges the two fidelity measures into an overall quality prediction named the Super Resolution Image Fidelity (SRIF) index, which demonstrates superior performance against state-of-the-art IQA models when tested on subject-rated datasets. Wei Zhou 0021, Zhou Wang 0001 |
ACM Multimedia | 1 |
| 2022 | A brief survey on adaptive video streaming quality assessment
Wei Zhou 0021, Xiongkuo Min, Qiuping Jiang |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | No-Reference Quality Assessment for 360-Degree Images by Analysis of Multifrequency Information and Local-Global Naturalnessabstract360-degree/omnidirectional images (OIs) have received remarkable attention due to the increasing applications of virtual reality (VR). Compared to conventional 2D images, OIs can provide more immersive experiences to consumers, benefiting from the higher resolution and plentiful field of views (FoVs). Moreover, observing OIs is usually in a head-mounted display (HMD) without references. Therefore, an efficient blind quality assessment method, which is specifically designed for 360-degree images, is urgently desired. In this paper, motivated by the characteristics of the human visual system (HVS) and the viewing process of VR visual content, we propose a novel and effective no-reference omnidirectional image quality assessment (NR OIQA) algorithm by MultiFrequency Information and Local-Global Naturalness (MFILGN). Specifically, inspired by the frequency-dependent property of the visual cortex, we first decompose the projected equirectangular projection (ERP) maps into wavelet subbands by using discrete Haar wavelet transform (DHWT). Then, the entropy intensities of low-frequency and high-frequency subbands are exploited to measure the multifrequency information of OIs. In addition to considering the global naturalness of ERP maps, owing to the browsed FoVs, we extract the natural scene statistics (NSS) features from each viewport image as the measure of local naturalness. With the proposed multifrequency information measurement and local-global naturalness measurement, we utilize support vector regression (SVR) as the final image quality regressor to train the quality evaluation model from visual quality-related features to human ratings. To our knowledge, the proposed model is the first no-reference quality assessment method for 360-degree images that combines multifrequency information and image naturalness. Experimental results on two publicly available OIQA databases demonstrate that our proposed MFILGN outperforms state-of-the-art full-reference (FR) and NR approaches. Wei Zhou 0021, Jiahua Xu 0001, Qiuping Jiang, Zhibo Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Elbow-Anchored Interaction: Designing Restful Mid-Air InputabstractWe designed a mid-air input space for restful interactions on the couch. We observed people gesturing in various postures on a couch and found that posture affects the choice of arm motions when no constraints are imposed by a system. Study participants that sat with the arm rested were more likely to use the forearm and wrist, as opposed to the whole arm. We investigate how a spherical input space, where forearm angles are mapped to screen coordinates, can facilitate restful mid-air input in multiple postures. We present two controlled studies. In the first, we examine how a spherical space compares with a planar space in an elbow-anchored setup, with a shoulder-level input space as baseline. In the second, we examine the performance of a spherical input space in four common couch postures that set unique constraints to the arm. We observe that a spherical model that captures forearm movement facilitates comfortable input across different seated postures. Rafael Veras, Gaganpreet Singh, Farzin Farhadi-Niaki, Ritesh Udhani, Parth Pradeep Patekar, Wei Zhou 0021, Pourang Irani, Wei Li 0002 |
CHI | 6 |
| 2021 | Leveraging CD Gain for Precise Barehand Video Timeline Browsing on Smart Displays
Futian Zhang, Sachi Mizobuchi, Wei Zhou 0021, Taslim Arefin Khan, Wei Li 0002, Edward Lank |
INTERACT (4) | 3 |
| 2021 | Deep Multi-Scale Features Learning for Distorted Image Quality AssessmentabstractImage quality assessment (IQA) aims to estimate human perception based image visual quality. Although existing deep neural networks (DNNs) have shown significant effectiveness for tackling the IQA problem, it still needs to improve the DNN- based quality assessment models by exploiting efficient multi- scale features. In this paper, motivated by the human visual system (HVS) combining multi-scale features for perception, we propose to use pyramid features learning to build a DNN with hierarchical multi-scale features for distorted image quality prediction. Our model is based on both residual maps and distorted images in luminance domain, where the proposed network contains spatial pyramid pooling and feature pyramid from the network structure. Our proposed network is optimized in a deep end-to-end supervision manner. To validate the effectiveness of the proposed method, extensive experiments are conducted on four widely-used image quality assessment databases, demonstrating the superiority of our algorithm. Wei Zhou 0021, Zhibo Chen 0001 |
ISCAS | 1 |
| 2021 | Perceptual Quality Assessment of Internet VideosabstractWith the fast proliferation of online video sites and social media platforms, user, professionally and occupationally generated content (UGC, PGC, OGC) videos are streamed and explosively shared over the Internet. Consequently, it is urgent to monitor the content quality of these Internet videos to guarantee the user experience. However, most existing modern video quality assessment (VQA) databases only include UGC videos and cannot meet the demands for other kinds of Internet videos with real-world distortions. To this end, we collect 1,072 videos from Youku, a leading Chinese video hosting service platform, to establish the Internet video quality assessment database (Youku-V1K). A special sampling method based on several quality indicators is adopted to maximize the content and distortion diversities within a limited database, and a probabilistic graphical model is applied to recover reliable labels from noisy crowdsourcing annotations. Based on the properties of Internet videos originated from Youku, we propose a spatio-temporal distortion-aware model (STDAM). First, the model works blindly which means the pristine video is unnecessary. Second, the model is familiar with diverse contents by pre-training on the large-scale image quality assessment databases. Third, to measure spatial and temporal distortions, we introduce the graph convolution and attention module to extract and enhance the features of the input video. Besides, we leverage the motion information and integrate the frame-level features into video-level features via a bi-directional long short-term memory network. Experimental results on the self-built database and the public VQA databases demonstrate that our model outperforms the state-of-the-art methods and exhibits promising generalization ability. Jiahua Xu 0001, Jing Li 0026, Xingguang Zhou, Wei Zhou 0021, Baichao Wang, Zhibo Chen 0001 |
ACM Multimedia | 4 |
| 2021 | Image Super-Resolution Quality Assessment: Structural Fidelity Versus Statistical NaturalnessabstractSingle image super-resolution (SISR) algorithms reconstruct high-resolution (HR) images with their low-resolution (LR) counterparts. It is desirable to develop image quality assessment (IQA) methods that can not only evaluate and compare SISR algorithms, but also guide their future development. In this paper, we assess the quality of SISR generated images in a two-dimensional (2D) space of structural fidelity versus statistical naturalness. This allows us to observe the behaviors of different SISR algorithms as a tradeoff in the 2D space. Specifically, SISR methods are traditionally designed to achieve high structural fidelity but often sacrifice statistical naturalness, while recent generative adversarial network (GAN) based algorithms tend to create more natural-looking results but lose significantly on structural fidelity. Furthermore, such a 2D evaluation can be easily fused to a scalar quality prediction. Interestingly, we find that a simple linear combination of a straightforward local structural fidelity and a global statistical naturalness measures produce surprisingly accurate predictions of SISR image quality when tested using public subject-rated SISR image datasets. Code of the proposed SFSN model is publicly available at https://github.con/weizhou-geek/SFSN. Wei Zhou 0021, Zhou Wang 0001, Zhibo Chen 0001 |
QoMEX | 1 |
| 2021 | Perceptual Evaluation of Pre-processing for Video TranscodingabstractRecently, the pre-processed video transcoding has attracted wide attention and has been increasingly used in practical applications for improving the perceptual experience and saving transmission resources. However, very few works have been conducted to evaluate the performance of pre-processing methods. In this paper, we select the source (SRC) videos and various pre-processing approaches to construct the first Pre-processed and Transcoded Video Database (PTVD). Then, we conduct the subjective experiment, showing that compared with the video sent to the codec directly at the same bitrate, the appropriate pre-processing methods indeed improve the perceptual quality. Finally, existing image/video quality metrics are evaluated on our database. The results indicate that the performance of the existing image/video quality assessment (IQA/VQA) approaches remain to be improved. We will make our database publicly available soon. Shiyu Huang 0002, Ziyuan Luo, Jiahua Xu 0001, Wei Zhou 0021, Zhibo Chen 0001 |
VCIP | 4 |
| 2021 | Corrections to "Blind quality assessment for image superresolution using deep two-stream convolutional networks"
Wei Zhou 0021, Qiuping Jiang, Yuwang Wang, Zhibo Chen 0001, Weiping Li 0003 |
Inf. Sci. | 1 |
| 2021 | Analyzing Midair Object Pointing Mappings for Smart Display InputabstractOne common task when controlling smart displays is the manipulation of menu items. Given current examples of smart displays that support distant bare hand control, in this paper we explore menu item selection tasks with three different mappings of barehand movement to target selection. Through a series of experiments, we demonstrate that Positional mapping is faster than other mappings when the target is visible but requires many clutches in large targeting spaces. Rate-based mapping is, in contrast, preferred by participants due to its perceived lower effort, despite being slightly harder to learn initially. Tradeoffs in the design of target selection in smart tv displays are discussed. Futian Zhang, Sachi Mizobuchi, Wei Zhou 0021, Edward Lank |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2021 | Blind Omnidirectional Image Quality Assessment With Viewport Oriented Graph Convolutional NetworksabstractQuality assessment of omnidirectional images has become increasingly urgent due to the rapid growth of virtual reality applications. Different from traditional 2D images and videos, omnidirectional contents can provide consumers with freely changeable viewports and a larger field of view covering the 360°×180°spherical surface, which makes the objective quality assessment of omnidirectional images more challenging. In this paper, motivated by the characteristics of the human vision system (HVS) and the viewing process of omnidirectional contents, we propose a novel Viewport oriented Graph Convolution Network (VGCN) for blind omnidirectional image quality assessment (IQA). Generally, observers tend to give the subjective rating of a 360-degree image after passing and aggregating different viewports information when browsing the spherical scenery. Therefore, in order to model the mutual dependency of viewports in the omnidirectional image, we build a spatial viewport graph. Specifically, the graph nodes are first defined with selected viewports with higher probabilities to be seen, which is inspired by the HVS that human beings are more sensitive to structural information. Then, these nodes are connected by spatial relations to capture interactions among them. Finally, reasoning on the proposed graph is performed via graph convolutional networks. Moreover, we simultaneously obtain global quality using the entire omnidirectional image without viewport sampling to boost the performance according to the viewing experience. Experimental results demonstrate that our proposed model outperforms state-of-the-art full-reference and no-reference IQA metrics on two public omnidirectional IQA databases. Jiahua Xu 0001, Wei Zhou 0021, Zhibo Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Re-Visiting Discriminator for Blind Free-Viewpoint Image Quality AssessmentabstractAccurate measurement of perceptual quality is important for various immersive multimedia, which demand real-time quality control or quality-based bench-marking for relevant algorithms. For instance, virtual views rendering in Free-Viewpoint (FV) navigation scenarios is a typical case that introduces challenging distortions, particularly the ones around dis-occluded regions. Existing quality metrics, most of which are targeting for impairments caused by compression or network condition, fail to quantify such non-uniform structure-related distortions. Moreover, the lack of quality databases for such distortions makes it even more challenging to develop robust quality metrics. In this work, a Generative Adversarial Networks based No-Reference (NR) quality Metric, namely GANs-NRM, is proposed. We first present an approach to create masks mimicking dis-occlusions/textureless regions, which is applicable on large-scale 2D image databases publicly available in the computer vision domain. Using these synthetic data, we then train a GANs-based context renderer with the capability of rendering those masked regions. Since the naturalness of the rendered dis-occluded regions strongly relates to the perceptual quality, we assume that the discriminator of the trained GANs has an intrinsic ability for quality assessment. We thus use the features extracted from the discriminator to learn a Bag-of-Distortion-Word (BDW) codebook. We show that a quality predictor can be then well trained using only a small amount of subjective quality data for the FV views rendering. Moreover, in the proposed framework, the discriminator is also adapted as a distortion-detector to locate possible distorted regions. According to the experimental results, the proposed model outperforms significantly the state-of-the-art quality metrics. The corresponding context renderer also shows appealing visualized results over other rendering algorithms. Suiyi Ling, Jing Li 0026, Zhaohui Che, Wei Zhou 0021, Junle Wang, Patrick Le Callet |
IEEE Trans. Multim. | 4 |
| 2020 | Learning Disentangled Feature Representation for Hybrid-Distorted Image Restoration
Xin Li 0082, Xin Jin 0014, Sen Liu 0001, Yaojun Wu 0001, Tao Yu 0012, Wei Zhou 0021, Zhibo Chen 0001 |
ECCV (29) | 7 |
| 2020 | LIRA: Lifelong Image Restoration from Unknown Blended Distortions
Jianzhao Liu, Xin Li 0082, Wei Zhou 0021, Sen Liu 0001, Zhibo Chen 0001 |
ECCV (18) | 4 |
| 2020 | Deep Local and Global Spatiotemporal Feature Aggregation for Blind Video Quality AssessmentabstractIn recent years, deep learning has achieved promising success for multimedia quality assessment, especially for image quality assessment (IQA). However, since there exist more complex temporal characteristics in videos, very little work has been done on video quality assessment (VQA) by exploiting powerful deep convolutional neural networks (DCNNs). In this paper, we propose an efficient VQA method named Deep SpatioTemporal video Quality assessor (DeepSTQ) to predict the perceptual quality of various distorted videos in a no-reference manner. In the proposed DeepSTQ, we first extract local and global spatiotemporal features by pre-trained deep learning models without fine-tuning or training from scratch. The composited features consider distorted video frames as well as frame difference maps from both global and local views. Then, the feature aggregation is conducted by the regression model to predict the perceptual video quality. Finally, experimental results demonstrate that our proposed DeepSTQ outperforms state-of-the-art quality assessment algorithms. Wei Zhou 0021, Zhibo Chen 0001 |
VCIP | 1 |
| 2020 | Blind quality assessment for image superresolution using deep two-stream convolutional networks
Wei Zhou 0021, Qiuping Jiang, Yuwang Wang, Zhibo Chen 0001, Weiping Li 0003 |
Inf. Sci. | 1 |
| 2020 | No-Reference Light Field Image Quality Assessment Based on Spatial-Angular MeasurementabstractLight field image quality assessment (LFI-QA) is a significant and challenging research problem. It helps to better guide light field acquisition, processing and applications. However, only a few objective models have been proposed and none of them completely consider intrinsic factors affecting the LFI quality. In this paper, we propose a No-Reference Light Field image Quality Assessment (NR-LFQA) scheme, where the main idea is to quantify the LFI quality degradation through evaluating the spatial quality and angular consistency. We first measure the spatial quality deterioration by capturing the naturalness distribution of the light field cyclopean image array, which is formed when human observes the LFI. Then, as a transformed representation of LFI, the Epipolar Plane Image (EPI) contains the slopes of lines and involves the angular information. Therefore, EPI is utilized to extract the global and local features from LFI to measure angular consistency degradation. Specifically, the distribution of gradient direction map of EPI is proposed to measure the global angular consistency distortion in the LFI. We further propose the weighted local binary pattern to capture the characteristics of local angular consistency degradation. Extensive experimental results on four publicly available LFI quality datasets demonstrate that the proposed method outperforms state-of-the-art 2D, 3D, multi-view, and LFI quality assessment algorithms. Likun Shi, Wei Zhou 0021, Zhibo Chen 0001, Jinglin Zhang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Tensor Oriented No-Reference Light Field Image Quality AssessmentabstractLight field image (LFI) quality assessment is becoming more and more important, which helps to better guide the acquisition, processing and application of immersive media. However, due to the inherent high dimensional characteristics of LFI, the LFI quality assessment turns into a multi-dimensional problem that requires consideration of the quality degradation in both spatial and angular dimensions. Therefore, we propose a novel Tensor oriented No-reference Light Field image Quality evaluator (Tensor-NLFQ) based on tensor theory. Specifically, since the LFI is regarded as a low-rank 4D tensor, the principal components of four oriented sub-aperture view stacks are obtained via Tucker decomposition. Then, the Principal Component Spatial Characteristic (PCSC) is designed to measure the spatial-dimensional quality of LFI considering its global naturalness and local frequency properties. Finally, the Tensor Angular Variation Index (TAVI) is proposed to measure angular consistency quality by analyzing the structural similarity distribution between the first principal component and each view in the view stack. Extensive experimental results on four publicly available LFI quality databases demonstrate that the proposed Tensor-NLFQ model outperforms state-of-the-art 2D, 3D, multi-view, and LFI quality assessment algorithms. Wei Zhou 0021, Likun Shi, Zhibo Chen 0001, Jinglin Zhang 0003 |
IEEE Trans. Image Process. | 1 |
| 2019 | Unsupervised Single Image Deraining with Self-Supervised ConstraintsabstractMost existing single image deraining methods require learning supervised models from a large set of paired synthetic training data, which limits their generality and practicality in real-world multimedia applications. Besides, due to lack of labeled-supervised constraints, directly applying existing unsupervised frameworks to the image deraining task will suffer from low-quality recovery. Therefore, we propose an Unsupervised Deraining Generative Adversarial Network (UD-GAN) to tackle above problems by introducing self-supervised constraints from the intrinsic statistics of unpaired rainy and clean images. Specifically, we design two collaboratively optimized modules, namely Rain Guidance Module (RGM) and Background Guidance Module (BGM), to take full advantage of rainy image characteristics. UD-GAN outperforms state-of-the-art methods on various benchmarking datasets in both quantitative and qualitative comparisons. Xin Jin 0014, Zhibo Chen 0001, Wei Zhou 0021 |
ICIP | 5 |
| 2019 | How do you Perceive Differently from an AI - A Database for Semantic Distortion MeasurementabstractArtificial intelligence (AI) is enabling the automated analysis of large amounts of image/video data, boosting the speed of multimedia data processing remarkably. Meanwhile, Image Quality Assessment (IQA) plays an important role in developing automatic analysis methods. To ensure the effectiveness of AI, images in multimedia applications should be considered for visual examination by both human and machine. Therefore, it is significant to understand the differences between human's and AI's perception of semantic distortion. However, little work has been done due to the lack of data from human on the semantic level. In this paper, we first propose a semantic database (SID) based on the surveillance scenarios, by collecting subjective average recognition rates of 3 semantic targets (face, pedestrian, license plate) with 3 types of distortion (JPEG Compression, BPG Compression, Motion Blur). Then, we present a detailed analysis of how human and AI perceive semantic distortion differently. Experimental results show that AI is stronger in tolerance to distortion than human beings on average, while weaker at generalization and stability. It is also implied in the experiments that existing IQA methods are not effective enough at judging the semantic distortion. Shuxin Zhao, Jiahua Xu 0001, Yongquan Hu, Wei Zhou 0021, Sen Liu 0001, Zhibo Chen 0001 |
ISCAS | 4 |
| 2019 | Hand-Over-Face Input Sensing for Interaction with Smartphones through the Built-in CameraabstractThis paper proposes using face as a touch surface and employing hand-over-face (HOF) gestures as a novel input modality for interaction with smartphones, especially when touch input is limited. We contribute InterFace, a general system framework that enables the HOF input modality using advanced computer vision techniques. As an examplar of the usage of this framework, we demonstrate the feasibility and usefulness of HOF with an Android application for improving single-user and group selfie-taking experience through providing appearance customization in real-time. In a within-subjects study comparing HOF against touch input for single-user interaction, we found that HOF input led to significant improvements in accuracy and perceived workload, and was preferred by the participants. Qualitative results of an observational study also demonstrated the potential of HOF input modality to improve the user experience in multi-user interactions. Based on the lessons learned from our studies, we propose a set of potential applications of HOF to support smartphone interaction. We envision that the affordances provided by the this modality can expand the mobile interaction vocabulary and facilitate scenarios where touch input is limited or even not possible. Mona Hosseinkhani Loorak, Wei Zhou 0021, Ha Trinh, Jian Zhao 0010, Wei Li 0002 |
MobileHCI | 2 |
| 2019 | No-Reference Light Field Image Quality Assessment Based on Micro-Lens ImageabstractLight field image quality assessment (LF-IQA) plays a significant role due to its guidance to Light Field (LF) contents acquisition, processing and application. The LF can be represented as 4-D signal, and its quality depends on both angular consistency and spatial quality. However, few existing LF-IQA methods concentrate on effects caused by angular inconsistency. Especially, no-reference methods lack effective utilization of 2D angular information. In this paper, we focus on measuring the 2-D angular consistency for LF-IQA. The Micro-Lens Image (MLI) refers to the angular domain of the LF image, which can simultaneously record the angular information in both horizontal and vertical directions. Since the MLI contains 2D angular information, we propose a No-Reference Light Field image Quality assessment model based on MLI (LF-QMLI). Specifically, we first utilize Global Entropy Distribution (GED) and Uniform Local Binary Pattern descriptor (ULBP) to extract features from the MLI, and then pool them together to measure angular consistency. In addition, the information entropy of SubAperture Image (SAI) is adopted to measure spatial quality. Extensive experimental results show that LF-QMLI achieves the state-of-the-art performance. Ziyuan Luo, Wei Zhou 0021, Likun Shi, Zhibo Chen 0001 |
PCS | 2 |
| 2019 | Quality Assessment of Stereoscopic 360-degree Images from Multi-viewportsabstractObjective quality assessment of stereoscopic panoramic images becomes a challenging problem owing to the rapid growth of 360-degree contents. Different from traditional 2D image quality assessment (IQA), more complex aspects are involved in 3D omnidirectional IQA, especially unlimited field of view (FoV) and extra depth perception, which brings difficulty to evaluate the quality of experience (QoE) of 3D omnidirectional images. In this paper, we propose a multi-viewport based full-reference stereo 360 IQA model. Due to the freely changeable viewports when browsing in the head-mounted display, our proposed approach processes the image inside FoV rather than the projected one such as equirectangular projection (ERP). In addition, since overall QoE depends on both image quality and depth perception, we utilize the features estimated by the difference map between left and right views which can reflect disparity. The depth perception features along with binocular image qualities are employed to further predict the overall QoE of 3D 360 images. The experimental results on our public Stereoscopic OmnidirectionaL Image quality assessment Database (SOLID) show that the proposed method achieves a significant improvement over some well-known IQA metrics and can accurately reflect the overall QoE of perceived images. Jiahua Xu 0001, Ziyuan Luo, Wei Zhou 0021, Zhibo Chen 0001 |
PCS | 3 |
| 2019 | Dual-Stream Interactive Networks for No-Reference Stereoscopic Image Quality AssessmentabstractThe goal of objective stereoscopic image quality assessment (SIQA) is to predict the human perceptual quality of stereoscopic/3D images automatically and accurately. Compared with traditional 2D image quality assessment, the quality assessment of stereoscopic images is more challenging because of complex binocular vision mechanisms and multiple quality dimensions. In this paper, inspired by the hierarchical dual-stream interactive nature of the human visual system, we propose a stereoscopic image quality assessment network (StereoQA-Net) for no-reference stereoscopic image quality assessment. The proposed StereoQA-Net is an end-to-end dual-stream interactive network containing left and right view sub-networks, where the interaction of the two sub-networks exists in multiple layers. We evaluate our method on the LIVE stereoscopic image quality databases. The experimental results show that our proposed StereoQA-Net outperforms state-of-the-art algorithms on both symmetrically and asymmetrically distorted stereoscopic image pairs of various distortion types. In a more general case, the proposed StereoQA-Net can effectively predict the perceptual quality of local regions. In addition, cross-dataset experiments also demonstrate the generalization ability of our algorithm. Wei Zhou 0021, Zhibo Chen 0001, Weiping Li 0003 |
IEEE Trans. Image Process. | 1 |
| 2018 | A Decomposed Dual-Cross Generative Adversarial Network for Image Rain Removal
Xin Jin 0014, Zhibo Chen 0001, Jiale Chen 0001, Wei Zhou 0021, Chaowei Shan |
BMVC | 5 |
| 2018 | SDM: Semantic Distortion Measurement for Video EncryptionabstractSemantic information is important in video encryption. However, existing image quality assessment (IQA) methods, such as the peak signal to noise ratio (PSNR), are still widely applied to measure the encryption security. Generally, these traditional IQA methods aim to evaluate the image quality from the perspective of visual signal rather than semantic information. In this paper, we propose a novel semantic-level full-reference image quality assessment (FR-IQA) method named Semantic Distortion Measurement (SDM) to measure the degree of semantic distortion for video encryption. Then, based on a semantic saliency dataset, we verify that the proposed SDM method outperforms state-of-the-art algorithms. Furthermore, we construct a Region Of Semantic Saliency (ROSS) video encryption system to demonstrate the effectiveness of our proposed SDM method in the practical application. Yongquan Hu, Wei Zhou 0021, Shuxin Zhao, Zhibo Chen 0001, Weiping Li 0003 |
FG | 2 |
| 2018 | Perceptual Evaluation of Light Field ImageabstractRecently, light field image has attracted wide attention. However, much less work has been conducted on the perceptual evaluation of light field image. In this work, we create the first windowed 5 degree of freedom light field image database (Win5-LID) based on stereoscopic display, which provides windowed 5 DOF experience and all the depth cues of light field image. The database consists of light field images with representative compression and reconstruction artifacts. We assume that the light field quality is not only affected by sub-views quality but also depth cues. Picture quality and overall quality are then evaluated and the results validate our assumption. Finally, the performance of existing image quality metrics is analyzed on our database. The results indicate that the performance of the state-of-the-art image quality metrics remains to be improved. Likun Shi, Shengyang Zhao, Wei Zhou 0021, Zhibo Chen 0001 |
ICIP | 3 |
| 2018 | Visual Comfort Assessment for Stereoscopic Image RetargetingabstractIn recent years, visual comfort assessment (VCA) for 3D/stereoscopic content has aroused extensive attention. However, much less work has been done on the perceptual evaluation of stereoscopic image retargeting. In this paper, we first build a Stereoscopic Image Retargeting Database (SIRD), which contains source images and retargeted images produced by four typical stereoscopic retargeting methods. Then, the subjective experiment is conducted to assess four aspects of visual distortion, i.e. visual comfort, image quality, depth quality and the overall quality. Furthermore, we propose a Visual Comfort Assessment metric for Stereoscopic Image Retargeting (VCA-SIR). Based on the characteristics of stereoscopic retargeted images, the proposed model introduces novel features like disparity range, boundary disparity as well as disparity intensity distribution into the assessment model. Experimental results demonstrate that VCA-SIR can achieve high consistency with subjective perception. Wei Zhou 0021, Zhibo Chen 0001 |
ISCAS | 2 |
| 2018 | Augmented Coarse-to-Fine Video Frame Synthesis with Semantic Loss
Xin Jin 0014, Zhibo Chen 0001, Sen Liu 0001, Wei Zhou 0021 |
PRCV (1) | 4 |
| 2018 | Blind Stereoscopic Video Quality Assessment: From Depth Perception to Overall ExperienceabstractStereoscopic video quality assessment (SVQA) is a challenging problem. It has not been well investigated on how to measure depth perception quality independently under different distortion categories and degrees, especially exploit the depth perception to assist the overall quality assessment of 3D videos. In this paper, we propose a new depth perception quality metric (DPQM) and verify that it outperforms existing metrics on our published 3D video extension of High Efficiency Video Coding (3D-HEVC) video database. Furthermore, we validate its effectiveness by applying the crucial part of the DPQM to a novel blind stereoscopic video quality evaluator (BSVQE) for overall 3D video quality assessment. In the DPQM, we introduce the feature of auto-regressive prediction-based disparity entropy (ARDE) measurement and the feature of energy weighted video content measurement, which are inspired by the free-energy principle and the binocular vision mechanism. In the BSVQE, the binocular summation and difference operations are integrated together with the fusion natural scene statistic measurement and the ARDE measurement to reveal the key influence from texture and disparity. Experimental results on three stereoscopic video databases demonstrate that our method outperforms state-of-the-art SVQA algorithms for both symmetrically and asymmetrically distorted stereoscopic video pairs of various distortion types. Zhibo Chen 0001, Wei Zhou 0021, Weiping Li 0003 |
IEEE Trans. Image Process. | 2 |
| 2016 | 3D-HEVC visual quality assessment: Database and bitstream modelabstractVisual Quality Assessment of 3D/stereoscopic video (3D VQA) is significant for both quality monitoring and optimization of the existing 3D video services. In this paper, we build a 3D video database based on the latest 3D-HEVC video coding standard, to investigate the relationship among video quality, depth quality, and overall quality of experience (QoE) of 3D/stereoscopic video. We also analyze the pivotal factors to the video and depth qualities. Moreover, we develop a No-Reference 3D-HEVC bitstream-level objective video quality assessment model, which utilizes the key features extracted from the 3D video bitstreams to assess the perceived quality of the stereoscopic video. The model is verified to be effective on our database as compared with widely used 2D Full-Reference quality metrics as well as a state-of-the-art 3D FR pixel-level video quality metric. Wei Zhou 0021, Ning Liao, Zhibo Chen 0001, Weiping Li 0003 |
QoMEX | 1 |