EDBT 2026 Demo / reviewers in the wild / expert
Zihan Zhou 0007
dblp:00/6525-7
· DBLP profile ↗
13ranked-venue papers
7as first author
12since 2021 · last 2026
0009-0002-9375-1905ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ObjectDiff: An object-centric diffusion policy with modality-specific conditioning for robot manipulation
Yong Xu 0007, Zhiyu Wei, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002 |
Knowl. Based Syst. | 4 |
| 2026 | Nighttime image dehazing via a physics-aware dynamic neural model with progressive contrastive regularization
Yun Liang 0003, Xinjie Xiao, Zihan Zhou 0007, Lianghui Li, Yuhui Quan |
Pattern Recognit. | 3 |
| 2026 | Deep Underwater Image Quality Assessment via Progressive Physics-Aware Multi-Prior CollaborationabstractUnderwater image quality assessment (UIQA) is a critical research area, challenged by underwater environments such as wavelength-dependent light attenuation, scattering, and non-uniform illumination. Existing deep learning-based UIQA methods often address these degradations in isolation, neglecting their complex interplay with human perception and lacking explicit modeling of underwater optical phenomena. To address this, we propose PhysIQ-Net, a novel framework that integrates physics-driven principles with progressive multi-prior interaction modeling through three key innovations: First, introduce dual physics-based decomposition that separates images into Backscatter, Transmission, Reflectance, and Illuminance components to capture distinct degradation mechanisms; Second, propose prior-guided dynamic filtering that adapts convolutional kernels to image-specific content using physical priors; and Third, propose physic-informed Cross-Domain Feature Interaction that enables bidirectional collaboration between color-aware and structure-aware representations to model their perceptual inter-dependencies. Extensive experiments across multiple benchmark datasets demonstrate that PhysIQ-Net significantly outperforms existing methods, with ablation studies validating each component’s contribution, providing a robust solution for UIQA. Zihan Zhou 0007, Jiaxue Lan, Yun Liang 0003, Jing Li 0026, Yong Xu 0007, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Self-Correcting Robot Manipulation via Gaussian-Splatted ForesightabstractLanguage-conditioned robotic manipulation in unstructured environments presents significant challenges for intelligent robotic systems. However, due to partial observation or imprecise action prediction, failure may be unavoidable for learned policies. Moreover, operational failures can lead to the robotic arm entering an untrained state, potentially causing destructive results. Consequently, the ability to detect and self-correct failures is crucial for the development of practical robotic systems. To address this challenge, we propose a foresight-driven failure detection and self-correction module for robot manipulation. By leveraging 3D Gaussian Splatting, we represent the current scene with multiple Gaussians. Subsequently, we train a prediction network to forecast the Gaussian representation of future scenes conditioned on planned actions. Failure is detected when the predicted future significantly deviates from the real observation after action execution. In such cases, the end-effector rolls back to the previous action to avoid an untrained state. Integrating this approach with the PerACT framework, we develop a self-correcting robot manipulation policy. Evaluations on ten RLBench tasks with 166 variations demonstrate the superior performance of the proposed method, which outperforms state-of-the-art methods by 12.0% success rate on average. Shaohui Pan, Yong Xu 0007, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002, Zhu Liang Yu |
AAAI | 4 |
| 2025 | Rethinking 3D Robotic Perception: Elastic Voxel Representation with Splatting DistillationabstractLanguage-guided robotic manipulation is advancing rapidly with Vision-Language-Action (VLA) models, yet faces fundamental challenges in 3D perception. This paper addresses two critical challenges: the scale elasticity requirement for simultaneously processing coarse environmental context and fine manipulation details, and the scarcity of action-annotated training data. We present Splat-Actor, a novel robotic manipulation framework that introduces two key innovations. First, we develop an elastic voxel encoder that combines multi-scale processing with selective tokenization, enabling efficient 3D spatial reasoning while adaptively focusing on informative regions. Second, we propose a depth-constrained feature distillation framework that leverages Gaussian Splatting to bridge 2D and 3D representations, transferring rich semantic features from pre-trained vision models to enhance 3D understanding. Extensive experiments across 10 manipulation tasks with 166 variations demonstrate that Splat-Actor achieves a 6.8% improvement over state-of-the-art methods while maintaining the computational efficiency. Shaohui Pan, Yong Xu 0007, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002, Zhu Liang Yu, Patrick Le Callet |
ICME | 4 |
| 2025 | Enhancing CNN-Based Blind Image Quality Assessment via Deep Cross-Layer Pattern EncodingabstractEvaluating image quality without reference images, known as blind image quality assessment (BIQA), is crucial for image communication. Recently, convolutional neural networks (CNNs) have emerged as a prominent BIQA approach due to their feature learning power. Usually, both high-level semantic information and low-level details significantly impact perceived visual quality. However, most existing CNN-based methods focus on high-level semantic information via aggregating features on top of the last convolutional layer into a global descriptor, neglecting the importance of shallow, low-level cues. To address this limitation, this paper proposes a novel approach that exploits local encoding and histogram-based pyramid pooling on crosslayer features produced by a CNN, achieving a joint local and global analysis. Specifically, we introduce a cross-layer pattern encoding model that characterizes features generated along convolutional layers via a soft histogram of local 3D binary patterns. This leads to a highly informative yet compact descriptor for score regression. By building this module into a ResNet backbone, we present an effective BIQA model demonstrating state-ofthe-art performance in extensive experiments on synthetic and authentic datasets. Zihan Zhou 0007, Yong Xu 0007, Yuhui Quan, Yun Liang 0003, Jing Li 0026, Patrick Le Callet |
IEEE Trans. Multim. | 1 |
| 2024 | Spatial Adaptive Filter Network With Scale-Sharing Convolution for Image DemoiréingabstractRemoving moiré patterns is a challenging task as it is a spatially varying degradation that varies in shape, color and scale. Existing image restoration models often rely on static convolutional neural networks (CNNs)-based architectures, and hence potentially suboptimal for addressing the diverse manifestations of moiré patterns across different images and spatial positions. To this end, we propose a spatially adaptive neural network for image demoiréing. This network introduces a dual-branch filter prediction module engineered to predict pixel-wise adaptive filters that can process moiré patterns of varying orientations and color-shift issues. To further tackle the challenge presented by scale variability, a scale-sharing convolution module is proposed, utilizing pixel-wise adaptive filters with multiple dilations to handle moiré patterns of different sizes but similar shapes effectively. Upon extensive evaluations of three benchmark datasets, our model consistently outperforms existing methods, yielding a PSNR improvement of over 0.37dB across all evaluated datasets and providing additional benefits in terms of model size. Yong Xu 0007, Zhiyu Wei, Ruotao Xu, Zihan Zhou 0007, Zhu Liang Yu |
IEEE Signal Process. Lett. | 4 |
| 2024 | Deep Blind Image Quality Assessment Using Dynamic Neural Model With Dual-Order StatisticsabstractDeep convolutional neural networks (CNNs) have increasingly become a prominent method for blind image quality assessment (BIQA). The process of quality assessment typically involves feature extraction, average-based pooling, and quality regression. Based on this process, as well as the consensus that the visual quality of an image mainly relies on its content and distortions, this work improves CNNs for BIQA in two ways. First, considering the content-awareness of visual quality perception, we incorporate content-awareness via a dynamic filtering module to extract content-adaptive features and a dynamic regression module to learn content-adaptive perception rules based on local content and global semantics. Second, considering distortion-sensitivity in visual quality perception, we introduce second-order global variance pooling and combine it with global average pooling (GAP). First-order pooling methods like GAP are limited in distinguishing complex distortions that cause local degradation while preserving global features. Thus, pooling with dual-order statistics enables a more distortion-sensitive and discriminative global representation. These two improvements result in a content-adaptive BIQA model with a dual-order global pooling mechanism, improving generalization on diverse images with varying contents and distortion types. Extensive experiments on synthetic and authentic distortion datasets demonstrate state-of-the-art performance of the proposed approach. Zihan Zhou 0007, Jing Li 0026, Dexiang Zhong, Yong Xu 0007, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | CDINet: Content Distortion Interaction Network for Blind Image Quality AssessmentabstractPerceptual image quality is related to content and distortion. Distortion classification is a common way to learn distortion information. How to extract distortion information consistent with human perception is a problem to be solved. Besides, the joint effect on image quality caused by the interplay of content and distortion has not been fully studied. In this paper, a novel Content Distortion Interaction Network (CDINet) is proposed for blind image quality assessment. Distortion representation are guided by content representation to learn quality-aware representation. CDINet consists of four components: a Distortion-Aware Module (DAM), a Content-Aware Module (CAM), an Asymmetric Content-Distortion Interaction (ACDI) module, and a quality regression module. The content representation and distortion representation are extracted respectively and fused interactively in CDINet. Specifically, with the assistance of image restoration, distortion representation consistent with human perception is learned. To further improve the ability in distortion representation, the DAM is used to construct the differences between the distorted image and its reference image. The proposed ACDI module enables the interaction of content and distortion representations to occur at different levels with less computational cost. Since the proposed CDINet considers the joint impact on image quality caused by the interplay of content and distortion, the predicted image qualities highly align with human perception. Comprehensive experiments on 8 benchmark datasets demonstrate that the proposed CDINet effectively extracts quality-aware representation, achieving state-of-the-art performance in evaluating both synthetically and authentically distorted images. Limin Zheng, Yu Luo 0004, Zihan Zhou 0007, Jie Ling 0002, Guanghui Yue 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Deep Blind Image Quality Assessment Using Dual-Order StatisticsabstractDeep convolutional neural networks (CNNs) have become a promising approach to blind image quality assessment (BIQA). Existing CNN-based BIQA methods often employ global average pooling (GAP) to aggregate feature maps into a fixed-size representation for regression, so as to handle input images with varying sizes. However, GAP is only capable of extracting the first-order statistics of feature distributions, which is ineffective for distinguishing complex distortions that cause local degradation or preserve global features. To tackle this problem, we introduce the second-order global covariance pooling (GCP) for aggregating feature maps, leading to a more distortion-sensitive and more discriminative global representation. By incorporating GCP and GAP into a ResNet backbone, we propose an effective deep model for BIQA. The experimental results on five BIQA benchmark datasets, including both the synthetic and authentic ones, have demon-strated the excellent performance of the proposed method. Zihan Zhou 0007, Yong Xu 0007, Yuhui Quan, Ruotao Xu |
ICME | 1 |
| 2022 | No-Reference Image Quality Assessment Using Dynamic Complex-Valued Neural ModelabstractDeep convolutional neural networks (CNNs) have become a promising approach to no-reference image quality assessment (NR-IQA). This paper aims at improving the power of CNNs for NR-IQA in two aspects. Firstly, motivated by the deep connection between complex-valued transforms and human visual perception, we introduce complex-valued convolutions and phase-aware activations beyond traditional real-valued CNNs, which improves the accuracy of NR-IQA without bringing noticeable additional computational costs. Secondly, considering the content-awareness of visual quality perception, we include a dynamic filtering module for better extracting content-aware features, which predicts features based on both local content and global semantics. These two improvements lead to a complex-valued content-aware neural NR-IQA model with good generalization. Extensive experiments on both synthetically and authentically distorted data have demonstrated the state-of-the-art performance of the proposed approach. Zihan Zhou 0007, Yong Xu 0007, Ruotao Xu, Yuhui Quan |
ACM Multimedia | 1 |
| 2021 | Image Quality Assessment Using Kernel Sparse CodingabstractOne key in image quality assessment (IQA) is the design of image representations that can capture the changes of image structures caused by distortions. Recent studies show that sparse coding has emerged as a promising approach to analyzing image structures for IQA. However, existing sparse-coding-based IQA approaches use linear coding models, which ignore the nonlinearities of manifolds of image patches and thus cannot analyze complex image structures well. To overcome such a weakness, in this paper, we introduce nonlinear sparse coding to IQA. A kernel dictionary construction scheme is proposed, which combines analytic dictionaries and learnable dictionaries to guarantee both the stability and effectiveness of kernel sparse coding in the context of IQA. Built upon the kernel dictionary construction, an effective full-reference IQA metric is developed. Benefiting from the considerations on nonlinearities during sparse coding, the proposed IQA metric not only characterizes image distortions better, but also achieves improvement on the consistency with subjective perception, when compared to the metrics built upon linear sparse coding. Such benefits are demonstrated with the experimental results on eight benchmark datasets in terms of common criteria. Zihan Zhou 0007, Jing Li 0026, Yuhui Quan, Ruotao Xu |
IEEE Trans. Multim. | 1 |
| 2020 | Full-reference image quality metric for blurry images and compressed images using hybrid dictionary learning
Zihan Zhou 0007, Jing Li 0026, Yong Xu 0007, Yuhui Quan |
Neural Comput. Appl. | 1 |