EDBT 2026 Demo / reviewers in the wild / expert
Kaiwei Zhang
dblp:236/8413
· DBLP profile ↗
36ranked-venue papers
11as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 7 first-author · 22 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement LearningabstractVideo quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainability, which restrict their applicability in real-world scenarios. To address these challenges, we propose VQAThinker, a reasoning-based VQA framework that leverages large multimodal models (LMMs) with reinforcement learning to jointly model video quality understanding and scoring, emulating human perceptual decision-making. Specifically, we adopt group relative policy optimization (GRPO), a rule-guided reinforcement learning algorithm that enables reasoning over video quality under score-level supervision, and introduce three VQA-specific rewards: (1) a bell-shaped regression reward that increases rapidly as the prediction error decreases and becomes progressively less sensitive near the ground truth; (2) a pairwise ranking reward that guides the model to correctly determine the relative quality between video pairs; and (3) a temporal consistency reward that encourages the model to prefer temporally coherent videos over their perturbed counterparts. Extensive experiments demonstrate that VQAThinker achieves state-of-the-art performance on both in-domain and OOD VQA benchmarks, showing strong generalization for video quality scoring. Furthermore, evaluations on video quality understanding tasks validate its superiority in distortion attribution and quality description compared to existing explainable VQA models and LMMs. These findings demonstrate that reinforcement learning offers an effective pathway toward building generalizable and explainable VQA models solely with score-level supervision. Linhan Cao, Wei Sun 0029, Weixia Zhang, Jun Jia, Kaiwei Zhang, Dandan Zhu 0001, Guangtao Zhai, Xiongkuo Min |
AAAI | 6 |
| 2026 | SalDiff-DTM: A Novel Dual-Temporal Modulated Diffusion Model for Omnidirectional Images Scanpath PredictionabstractScanpath prediction in omnidirectional images (ODIs) serves as a critical component for optimizing foveated rendering efficiency and enhancing interactive quality in virtual reality systems. However, existing scanpath prediction methods for ODIs still suffer from fundamental limitations: (1) inadequate modeling and capturing of long-range temporal dependencies in fixation regions, and (2) suboptimal integration of spatial and temporal visual features, ultimately compromising prediction performance. To address these limitations, we propose a novel Dual-Temporal Modulated Diffusion model for Omnidirectional Images Scanpath Prediction, named SalDiff-DTM model, to effectively generate realistic human eye viewing trajectories. Specifically, to effectively model spatial relationships, we propose a novel Dual-Graph Convolutional Network (Dual-GCN) module that simultaneously captures semantic-level and image-level correlations. By integrating both local spatial details and global contextual information across the internal temporal dimension, this module achieves comprehensive and robust modeling of spatial relationships. To further enhance the modeling of temporal dependencies inherent in diverse fixation patterns, we introduce TABiMamba (Temporal-Aware BiLSTM-Mamba), a dedicated module that synergistically combines the contextual sensitivity of BiLSTM with the long-range sequence modeling capabilities of Mamba. This design facilitates deep information flow and context-aware sequential reasoning, thereby enabling high-fidelity capture of intricate temporal correlations. Inspired by the progressive refinement mechanism of diffusion models in various generative tasks, we propose a saliency-guided diffusion module that formulates the prediction problem as a conditional generative process, iteratively yielding accurate and perceptually plausible scanpaths. Extensive experiments demonstrate that SalDiff-DTM significantly outperforms state-of-the-art models, paving the way for future advancements in eye-tracking technologies and cognitive modeling. Xiaohui Kong, Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min |
AAAI | 4 |
| 2026 | One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving FrameworkabstractQi Jia, Ye Shen, Xiujie Song, Kaiwei Zhang, Shibo Wang, Dun Pei, Xiangyang Zhu, Guangtao Zhai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ye Shen, Xiujie Song, Kaiwei Zhang, Dun Pei, Guangtao Zhai |
ACL (1) | 4 |
| 2026 | Assessing Personality Consistency in Large Language Models: A Psychometric Framework for Human-Centric Quality of Experience
Yitian Kou, Dandan Zhu 0001, Wei Sun 0029, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai |
QoMEX | 4 |
| 2026 | Developing Evolving Adaptability in Biological Intelligence: A Novel Biologically-Inspired Continual Learning Model for Video Saliency PredictionabstractIn the era of deep learning, video saliency prediction task still remains major challenge due to the issue of catastrophic forgetting during feature learning. Most prior works commonly employ generative replay strategies to generate pseudo-samples from previous tasks, enabling them to recall the data distribution. However, scaling up generative replay to accommodate class-incremental and task-incremental settings poses challenges, as generated data with low quality can severely deteriorate performance. Additionally, existing advances mainly focus on preserving memory stability to alleviate catastrophic forgetting, but they remain difficult to flexibly adapt to incremental changes in dynamic scenes. To achieve a better balance between memory stability and learning plasticity, we propose a novel biologically-inspired continual learning (BICL) model tailored to effectively predict human attention in dynamic scenes while mitigate catastrophic forgetting. In particular, inspired by the function of the hippocampus in the human neural system, we elaborately design a visual saliency memory bank module to explicitly store and retrieve representative features from previous tasks. Furthermore, drawing inspiration from the Drosophila $\gamma$γMB system, we propose an active forgetting strategy equipped with multiple parallel adaptive learner modules, which can appropriately attenuate old memories in parameter distribution to enhance learning plasticity to adapt to new tasks, and accordingly to ensure compatibility among multiple learners. Notably, without compromising the performance of old tasks, our proposed model can achieve a better trade-off between memory stability and learning plasticity. Through extensive experiments on several benchmark datasets, our model not only enhances performance in task-incremental settings, but also potentially provides deep insights into neurological adaptive mechanisms. Dandan Zhu 0001, Kaiwei Zhang, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Subjective and Objective Quality Assessment of Display Content VideosabstractDisplay quality assessment plays a crucial role in evaluating the performance of display devices. However, existing video quality assessment methods primarily target compression-related distortions, failing to capture display-specific degradations including definition loss, color distortions, and motion artifacts that critically affect user subjective experiences during video playback. To address these limitations, we develop a specialized video dataset, namely Video Displaying Quality Assessment Dataset (VDQA), constructed using a DSLR camera with standardized parameter optimization of exposure settings (aperture, ISO sensitivity, and shutter speed). VDQA comprises 250 high-resolution video clips covering diverse content categories, providing a robust foundation for evaluating display devices across multiple quality dimensions. Additionally, we propose a deep learning-based model specifically designed for display quality assessment that employs three complementary pathways to independently evaluate definition, color fidelity, and motion quality. The model integrates Canny edge detection for explicit sharpness measurement, a color attention mechanism to enhance sensitivity to display color reproduction characteristics, and temporal modeling for motion artifact assessment. Experimental results demonstrate that the proposed model achieves superior performance in reflecting user subjective experiences for display content videos compared to state-of-the-art methods, with significant improvements in both color fidelity assessment and definition evaluation. Fangfang Lu, Huiqun Yu, Kaiwei Zhang, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | MEScan360: A Memory-Enhanced Scanpath Prediction Model for Omnidirectional Images
Dandan Zhu 0001, Kaiwei Zhang, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D GraphicsabstractTextured meshes significantly enhance the realism and detail of objects by mapping intricate texture details onto the geometric structure of 3D models. This advancement is valuable across various applications, including entertainment, education, and industry. While traditional mesh saliency studies focus on non-textured meshes, our work explores the complexities introduced by detailed texture patterns. We present a new dataset for textured mesh saliency, created through an innovative eye-tracking experiment in a six degrees of freedom (6-DOF) VR environment. This dataset addresses the limitations of previous studies by providing comprehensive eye-tracking data from multiple viewpoints, thereby advancing our understanding of human visual behavior and supporting more accurate and effective 3D content creation. Our proposed model predicts saliency maps for textured mesh surfaces by treating each triangular face as an individual unit and assigning a saliency density value to reflect the importance of each local surface region. The model incorporates a texture alignment module and a geometric extraction module, combined with an aggregation module to integrate texture and geometry for precise saliency prediction. We believe this approach will enhance the visual fidelity of geometric processing while ensuring computational efficiency, essential for real-time rendering and high-detail applications such as VR and gaming. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
AAAI | 1 |
| 2025 | Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured MeshesabstractMesh saliency enhances the adaptability of 3D vision by identifying and emphasizing regions that naturally attract visual attention. To investigate the interaction between geometric structure and texture in shaping visual attention, we establish a comprehensive mesh saliency dataset, which is the first to systematically capture the differences in saliency distribution under both textured and non-textured visual conditions. Furthermore, we introduce mesh Mamba, a unified saliency prediction model based on a state space model (SSM), designed to adapt across various mesh types. Mesh Mamba effectively analyzes the geometric structure of the mesh while seamlessly incorporating texture features into the topological framework, ensuring coherence throughout appearance-enhanced modeling. More importantly, by sub-graph embedding and a bidirectional SSM, the model enables global context modeling for both local geometry and texture, preserving the topological structure and improving the understanding of visual details and structural complexity. Through extensive theoretical and empirical validation, our model not only improves performance across various mesh types but also demonstrates high scalability and versatility, particularly through cross validations of various visual features. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
CVPR | 1 |
| 2025 | A Novel Framework for Realistic 3D Scene Regeneration with Graph of ThoughtsabstractIn embodied intelligence applications, highly realistic 3D scenes lay the foundation for perception and decision-making, while 3D scene regeneration creates more coherent and personalized virtual spaces, facilitating more efficient task adaptation and agent training. To address this, we propose a reasoning framework based on the Graph of Thoughts (GoT), which enhances the prompting capabilities of large language models (LLM) and integrates a synergistic mechanism of retrospective memory and feedback loops into the regeneration process. During the initial generation phase, we retain the Holodeck paradigm, combining LLM-driven scene design inferences with the spatial layout of 3D assets from Objaverse. In the regeneration phase, dynamic feedback loops trigger backtracking of reasoning memory to adjust relevant elements according to evolving requirements, while maintaining stability and consistency in unrelated elements, ensuring the scene’s overall coherence. We conduct both subjective and objective experiments to validate the effectiveness of this framework, demonstrating significant improvements in 3D scene generation. Yitian Kou, Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
ICME | 2 |
| 2025 | MEScan360: A Memory-Enhanced Scanpath Prediction Model for Omnidirectional ImagesabstractScanpath prediction for omnidirectional images (ODIs) aims to capture the dynamic human visual attention. However, the complicated gaze behavior and inevitable projection distortion make scanpath prediction in ODIs extremely challenging. Most existing models neither capture the long-term dependencies across visual states nor fully incorporate historical memory information, leading to limited performance. To this end, we propose MEScan360, a memory-enhanced scanpath prediction model for ODIs. We introduce two key innovations: long-term memory storage unit and memory interaction module. These two components establish a more explicit link between past visual information and current visual inputs, thereby significantly enhancing the performance of scanpath prediction. Furthermore, a robust feature extraction module is designed to extract semantic feature precisely from distorted ODIs with a more lightweight structure. Extensive experiments on several benchmark datasets demonstrate that our proposed model achieves competitive performance in both accuracy and efficiency. Dandan Zhu 0001, Kaiwei Zhang, Fei Jiang 0006, Guangtao Zhai |
ICME | 3 |
| 2025 | A Dataset and Method for Assessing the Quality of Display DevicesabstractThis paper proposes a novel objective quality assessment method for display devices, aiming to comprehensively evaluate their overall quality, with a particular focus on the user’s subjective experience. By constructing a dedicated video dataset and integrating a no-reference video quality assessment (NR VQA) model, we directly learn spatiotemporal features related to quality from video frames and introduce a color attention mechanism to enhance the model’s sensitivity to color distortions. The model consists of feature extraction, quality regression, and quality pooling modules, capable of extracting quality-aware features from spatial, spatiotemporal, and color domains, and determining the final quality score using a temporal averaging pooling strategy. Experimental results demonstrate that the proposed model performs excellently in evaluating the overall quality of display devices, showing high practical applicability. The main contributions of this paper include constructing a video dataset, proposing a new model, and introducing a color attention mechanism, which effectively enhances the comprehensive evaluation of the subjective quality of display devices, filling the gap in display device quality assessment. Haoyang Ni, Kaiwei Zhang, Fangfang Lu, Xiongkuo Min, Guangtao Zhai |
ISCAS | 3 |
| 2025 | Sculpting molecules in text-3D space: a flexible substructure aware framework for text-oriented molecular optimizationabstractThe integration of deep learning, particularly AI-Generated Content, with high-quality data derived from ab initio calculations has emerged as a promising avenue for transforming the landscape of scientific research. However, the challenge of designing molecular drugs or materials that incorporate multi-modality prior knowledge remains a critical and complex undertaking. Specifically, achieving a practical molecular design necessitates not only meeting the diversity requirements but also addressing structural and textural constraints with various symmetries outlined by domain experts. In this article, we present an innovative approach to tackle this inverse design problem by formulating it as a multi-modality guidance optimization task. Our proposed solution involves a textural-structure alignment symmetric diffusion framework for the implementation of molecular optimization tasks, namely 3DToMolo. 3DToMolo aims to harmonize diverse modalities including textual description features and graph structural features, aligning them seamlessly to produce molecular structures adhere to specified symmetric structural and textural constraints by experts in the field. Experimental trials across three guidance optimization settings have shown a superior hit optimization performance compared to state-of-the-art methodologies. Moreover, 3DToMolo demonstrates the capability to discover potential novel molecules, incorporating specified target substructures, without the need for prior knowledge. This work not only holds general significance for the advancement of deep learning methodologies but also paves the way for a transformative shift in molecular design strategies. 3DToMolo creates opportunities for a more nuanced and effective exploration of the vast chemical space, opening new frontiers in the development of molecular entities with tailored properties and functionalities. Kaiwei Zhang, Yange Lin, Guangcheng Wu, Yuxiang Ren, Xuecang Zhang, Weitao Du |
BMC Bioinform. | 1 |
| 2025 | Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai |
Sci. China Inf. Sci. | 17 |
| 2025 | ScanDTM: A Novel Dual-Temporal Modulation Scanpath Prediction Model for Omnidirectional ImagesabstractScanpath prediction for omnidirectional images aims to effectively simulate the human visual perception mechanism to generate dynamic realistic fixation trajectories. However, the majority of scanpath prediction methods for omnidirectional images are still in their infancy as they fail to accurately capture the time-dependency of viewing behavior and suffer from sub-optimal performance along with limited generalization capability. A desirable solution should achieve a better trade-off between prediction performance and generalization ability. To this end, we propose a novel dual-temporal modulation scanpath prediction (ScanDTM) model for omnidirectional images. Such a model is designed to effectively capture long-range time-dependencies between various fixation regions across both internal and external time dimensions, thereby generating more realistic scanpaths. In particular, we design a Dual Graph Convolutional Network (Dual-GCN) module comprising a semantic-level GCN and an image-level GCN. This module servers as a robust visual encoder that captures spatial relationships among various object regions within an image and fully utilizes similar images as complementary information to capture similarity relations across relevant images. Notably, the proposed Dual-GCN focuses on modeling temporal correlations from both local and global perspectives within the internal time dimension. Furthermore, drawing inspiration from the promising generalization capabilities of diffusion models across various generative tasks, we introduce a novel diffusion-guided saliency module. This module formulates the prediction issue as a conditional generative process for the saliency map, utilizing extracted semantic-level and image-level visual features as conditions. With the well-designed diffusion-guided saliency module, our proposed ScanDTM model acting as an external temporal modulator, we can progressively refine the generated scanpath from the noisy map. We conduct extensive experiments on several benchmark datasets, and the results demonstrate that our ScanDTM model significantly outperforms other competitors. Meanwhile, when applied to tasks such as saliency prediction and image quality assessment, our ScanDTM model consistently achieves superior generalization performance. Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | How Does Audio Influence Visual Attention in Omnidirectional Videos? Database and ModelabstractUnderstanding and predicting viewer attention in omnidirectional videos (ODVs) is crucial for enhancing user engagement in virtual and augmented reality applications. Although both audio and visual modalities are essential for saliency prediction in ODVs, the joint exploitation of these two modalities has been limited, primarily due to the absence of large-scale audio-visual saliency databases and comprehensive analyses. This paper comprehensively investigates audio-visual attention in ODVs from both subjective and objective perspectives. Specifically, we first introduce a new audio-visual saliency database for omnidirectional videos, termed AVS-ODV database, containing 162 ODVs and corresponding eye movement data collected from 60 subjects under three audio modes including mute, mono, and ambisonics. Based on the constructed AVS-ODV database, we perform an in-depth analysis of how audio influences visual attention in ODVs. To advance the research on audio-visual saliency prediction for ODVs, we further establish a new benchmark based on the AVS-ODV database by testing numerous state-of-the-art saliency models, including visual-only models and audio-visual models. In addition, given the limitations of current models, we propose an innovative omnidirectional audio-visual saliency prediction network (OmniAVS), which is built based on the U-Net architecture, and hierarchically fuses audio and visual features from the multimodal aligned embedding space. Extensive experimental results demonstrate that the proposed OmniAVS model outperforms other state-of-the-art models on both ODV AVS prediction and traditional AVS prediction tasks. The AVS-ODV database and the OmniAVS model are available at: https://github.com/IntMeGroup/AVS-ODV. Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Xilei Zhu, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Image Process. | 3 |
| 2025 | From Haziness to Clarity: A Novel Iterative Memory-Retrospective Emergence Model for Omnidirectional Image Saliency PredictionabstractTo achieve saliency prediction in omnidirectional images (ODIs), the majority of prior works typically adopt the convolutional neural networks (CNNs)-based saliency models to extract semantic features to predict prominent regions in ODIs. Albeit achieving substantially performance gains, these works all employed purely visual computing paradigms and ignore to explore the nature of human visual attention mechanisms. In other words, existing saliency prediction works for ODIs are insufficient to capture the biological characteristics of the visual attention mechanism in the human brain. To establish a more explicit link between saliency prediction performance and brain-like visual attention mechanism, we simulate the mechanism of human retrospective memory in neuropsychology and propose IMRE model, a novel iterative memory-retrospective emergence model can predict and infer the salient features by recalling previously learned information. In IMRE model, we introduce four key modules to simulate the visual attention mechanism for predicting human fixations in the human brain. Firstly, the visual stimulus response module is designed to effectively extract semantic features and capture the intricate relationship between these features, acting as the human visual cortex. Secondly, the retrospective integration module serves to distill valuable information from a fuzzy memory ensemble, resembling the role of the basal ganglia in the neural system. Thirdly, the memory bank module explicitly records and stores subconscious response information and learned knowledge, acting like the hippocampus in neural system. Lastly, the prospective inference module accurately infers saliency maps from the refined useful information, resembling the role of the prefrontal cortex. During prediction, we utilize the introduced memory bank to retrieve and recall previously learned information, which simulates the process of memory emergence from haziness to clarity. Such a process aligns with the retrospective memory mechanism of the human brain. To validate the superiority of the proposed model in ODIs saliency prediction tasks, we conduct extensive experiments on two benchmark datasets. Experiments show impressive performances that IMRE model outperforms other state-of-the-art methods across all benchmark datasets. Importantly, experiments also highlight the IMRE model's ability to trace back to specific instances during prediction, thereby reducing model inference costs and enhancing interpretability. Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Explain Vision Focus: Blending Human Saliency Into Synthetic Face ImagesabstractSynthetic faces have been extensively researched and applied in various fields, such as face parsing and recognition. Compared to real face images, synthetic faces engender more controllable and consistent experimental stimuli due to the ability to precisely merge expression animations onto the facial skeleton. Accordingly, we establish an eye-tracking database with 780 synthetic face images and fixation data collected from 22 participants. The use of synthetic images with consistent expressions ensures reliable data support for exploring the database and determining the following findings: (1) A correlation study between saliency intensity and facial movement reveals that the variation of attention distribution within facial regions is mainly attributed to the movement of the mouth. (2) A categorized analysis of different demographic factors demonstrates that the bias towards salient regions aligns with differences in some demographic categories of synthetic characters. In practice, inference of facial saliency distribution is commonly used to predict the regions of interest for facial video-related applications. Therefore, we propose a benchmark model that accurately predicts saliency maps, closely matching the ground truth annotations. This achievement is made possible by utilizing channel alignment and progressive summation for feature fusion, along with the incorporation of Sinusoidal Position Encoding. The ablation experiment also demonstrates the effectiveness of our proposed model. We hope that this paper will contribute to advancing the photorealism of generative digital humans. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Huiyu Duan, Guangtao Zhai |
IEEE Trans. Multim. | 1 |
| 2025 | Elevating Mesh Saliency in VR: Introducing a Novel Prediction Network and DatasetabstractIn computer graphics, polygon meshes stand out as a popular representation providing effective delineation of delicate textures and complex geometries. When dealing with geometric processing tasks for critical regions of the mesh, it is necessary to consider the human visual perception related to saliency. Therefore, we establish a novel mesh saliency dataset, facilitated by a more comprehensive gathering pipeline of eye-tracking from subjects observing mesh models at arbitrary viewpoints in a virtual reality space with six degrees of freedom. Additionally, we propose a mesh saliency prediction model that accurately infers visual attention density maps for complex and irregular mesh surfaces. This model integrates surface curvature and triangular face shape information from multi-scale neighboring ranges as local geometric features, while also leveraging surface spatial positioning as a global feature. Our work aims to preserve critical areas and minimize visual loss in saliency-driven tasks such as mesh simplification, rendering, and texturing. We believe that our research can offer valuable insights for human-centered mesh computation applications. Kaiwei Zhang, Mohan He, Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Unified Approach to Mesh Saliency: Evaluating Textured and Non-Textured Meshes Through VR and Multifunctional PredictionabstractMesh saliency aims to empower artificial intelligence with strong adaptability to highlight regions that naturally attract visual attention. Existing advances primarily emphasize the crucial role of geometric shapes in determining mesh saliency, but it remains challenging to flexibly sense the unique visual appeal brought by the realism of complex texture patterns. To investigate the interaction between geometric shapes and texture features in visual perception, we establish a comprehensive mesh saliency dataset, capturing saliency distributions for identical 3D models under both non-textured and textured conditions. Additionally, we propose a unified saliency prediction model applicable to various mesh types, providing valuable insights for both detailed modeling and realistic rendering applications. This model effectively analyzes the geometric structure of the mesh while seamlessly incorporating texture features into the topological framework, ensuring coherence throughout appearance-enhanced modeling. Through extensive theoretical and empirical validation, our approach not only enhances performance across different mesh types, but also demonstrates the model's scalability and generalizability, particularly through cross-validation of various visual features. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Q-Refine: A Perceptual Quality Refiner for AI-Generated ImageabstractWith the rapid evolution of the Text-to-Image (T2I) model in recent years, their unsatisfactory generation result has become a challenge. However, uniformly refining AI-Generated Images (AIGIs) of different qualities not only limited optimization capabilities for low-quality AIGIs but also brought negative optimization to high-quality AIGIs. To address this issue, a quality-award refiner named Q-Refine is proposed. Based on the preference of the Human Visual System (HVS), Q-Refine uses the Image Quality Assessment (IQA) metric to guide the refining process for the first time, and modify images of different qualities through three adaptive pipelines. Experimental data shows that for mainstream T2I models, Q-Refine can perform effective optimization to AIGIs of different qualities. It can be a general refiner to optimize AIGIs from both fidelity and aesthetic quality levels, thus expanding the application of the T2I generation models. The code is released on https://github.com/Q-Future/Q-Refine. Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Kaiwei Zhang, Lei Bai 0001, Xiaohong Liu 0001, Xiongkuo Min, Weisi Lin, Guangtao Zhai |
ICME | 5 |
| 2024 | PrefIQA: Human Preference Learning for AI-generated Image Quality AssessmentabstractDespite recent advancements in generative models, the variation in image quality remains a significant concern. To tackle this issue, we propose PrefIQA, an effective human preference learning metric, which can better evaluate the quality of AI-generated images. PrefIQA consists of two units, namely Feature Extraction Unit and Feature Fusion Unit. In Feature Extraction Unit, we introduce a prompt-segmentation module to divide prompts into multiple phrases, enabling a more detailed evaluation of the alignment between images and texts. In Feature Fusion Unit, we introduce a modality-fusion module, which effectively mixes text features and image features to improve the overall performance. In the experiment part, extensive experiments are conducted, demonstrating that PrefIQA surpasses existing text-to-image alignment metrics. We believe that PrefIQA’s proposal would facilitate researches on AI-generated image quality assessment, and make a valuable contribution to the field of text-to-image generation. Hengjian Gao, Kaiwei Zhang, Wei Sun 0029, Chunyi Li 0001, Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ISCAS | 2 |
| 2024 | PAPS-OVQA: Projection-Aware Patch Sampling for Omnidirectional Video Quality AssessmentabstractIn immersive multimedia systems, the perceptual quality model of omnidirectional video is indispensable. However, to cope with its resolution that is several times higher than ordinary video, the existing omnidirectional video quality assessment (OVQA) models require extremely high computational complexity and usually need to transcode the projection into a certain format. Therefore, to assess the perceptual quality of omnidirectional video effectively, we propose Projection-Aware Patch Sampling (PAPS)-OVQA to process its three common projection formats simultaneously while resizing high-resolution video into patches sampled from uniform grids and finally apply Fragment Attention Network (FANet) to perform quality regression. As a result, we avoid the overhead computational cost of projection transcoding and reduce the complexity of the quality model greatly. Experimental data show that PAPS-OVQA guarantees good performance while retaining high efficiency under different projection formats. Chunyi Li 0001, Haoning Wu 0001, Kaiwei Zhang, Lei Bai 0001, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
ISCAS | 4 |
| 2024 | ACIQA: A Dataset and Method for Assessing the Imaging Quality of Automotive CamerasabstractThe imaging quality of automotive cameras is crucial in complex driving environments. Therefore, it is essential to conduct subjective experiments that can realistically reflect drivers’ evaluation of the imaging quality of automotive cameras in real traffic scenarios. To accurately assess the imaging quality of automotive cameras, this paper proposes a no-reference quality assessment method with quality scores that are highly consistent with human subjective perception. Initially, this study constructs a new image quality assessment dataset and then obtains the subjective scores of image quality through subjective experiments. The dataset is constructed by using a variety of realistic props to simulate scene elements that might be captured by an automotive camera and are captured using a wide range of cameras with different sensor types, lens focus, and viewing angles, resulting in a dataset of diverse images. The objective quality assessment method proposed in this paper consists of an object detection network and a multi-branch quality evaluation network. The object detection network is responsible for identifying and classifying scene elements, while the multi-branch quality evaluation network performs feature extraction and score regression on various types of elements to effectively evaluate the imaging quality of the automotive cameras. In the experiments, this no-reference quality assessment method is tested on our built dataset, and the results show that the proposed method exhibits the best performance compared with the state-of-the-art image quality assessment methods. Haoyang Ni, Kaiwei Zhang, Ziheng Jia, Fangfang Lu, Xiongkuo Min, Guangtao Zhai |
VCIP | 3 |
| 2024 | Synergetic Assessment of Quality and Aesthetic: Approach and Comprehensive Benchmark DatasetabstractQuantifications of image quality and aesthetic have been regarded as two independent fields in computer vision. Generally, image quality assessment aims at measuring image distortions and image aesthetic is judged by commonly established photography rules. However, either measuring image quality or aesthetic alone is not sufficient to qualitatively rank images. Therefore, this paper puts forward the synergetic assessment of quality and aesthetic to help understand the subjective human preferences of digital pictures more comprehensively. Specifically, considering that the images of existing benchmark datasets are only labeled with single attribute, we first establish a new dataset which contains 9042 real-world images with the corresponding human rated pair-wise quality-aesthetic scores. Previously, these images are only labeled with aesthetic score, and we evaluate the subjective quality score of them, so that it can make up the lack of image dataset with double attributes. Moreover, since the existing methods are mostly designed for individual attribute prediction. We then propose a two-stream learning network to assess both quality and aesthetic of images in parallel. This network follows the top-down perception mechanism which learns from both fined grained details and holistic image layout simultaneously. Furthermore, we introduce a Channel-Diversity loss, which can be deployed in grouped convolution operation, and can constrain channels to be mutually exclusive across the spatial dimensions. To some extent, this contributes to spotlight different local discriminative regions with a finer granularity. Finally, experiments demonstrate that our method outperforms the state-of-the-art methods on our established benchmark dataset and other benchmark datasets in terms of image quality and aesthetic assessment. We hope this paper could serve as a potent reference and be useful for future research on the study of image ranking. Both the benchmark dataset and the code will be publicly available to facilitate further research. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Zhongpai Gao, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Unified Audio-Visual Saliency Model for Omnidirectional Videos With Spatial AudioabstractSpatial audio is a crucial component of omnidirectional videos (ODVs), which can provide an immersive experience by enabling viewers to perceive sound sources in all directions. However, most visual attention modeling works for ODVs focus only on visual cues, and audio modality is rather rarely considered. Additionally, the existing audio-visual saliency models for ODVs lack spatial audio location-awareness (i.e. sound source location-agnostic) and audio content attributes discriminability (i.e. audio content attributes-agnostic). To this end, we propose a novel audio-visual perception saliency (AVPS) model with spatial audio location-awareness and audio content attributes-adaptive to efficiently address the problem of fixation prediction in ODVs. Specifically, we first utilize the improved group equivariant convolutional neural network (G-CNN) with eidetic 3D LSTM (E3D-LSTM) to extract spatial-temporal visual features. Then we perceive sound source locations by computing the audio energy map (AEM) of the audio information in ODVs. Subsequently, we introduce SoundNet to extract audio features with multiple attributes. Finally, we develop an audio-visual feature fusion module to adaptively integrate spatial-temporal visual features and spatial auditory information to generate the final audio-visual saliency map. Extensive experiments in three audio modalities validate the effectiveness of the proposed model. Meanwhile, the performance of the proposed model is superior to the other 10 state-of-the-art saliency models. Dandan Zhu 0001, Kaiwei Zhang, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Audio-Visual Saliency for Omnidirectional Videos
Xilei Zhu, Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Li Chen 0021, Xiongkuo Min, Guangtao Zhai |
ICIG (5) | 5 |
| 2023 | Rumor Detection with Diverse Counterfactual EvidenceabstractThe growth in social media has exacerbated the threat of fake news to individuals and communities. This draws increasing attention to developing efficient and timely rumor detection methods. The prevailing approaches resort to graph neural networks (GNNs) to exploit the post-propagation patterns of the rumor-spreading process. However, these methods lack inherent interpretation of rumor detection due to the black-box nature of GNNs. Moreover, these methods suffer from less robust results as they employ all the propagation patterns for rumor detection. In this paper, we address the above issues with the proposed Diverse Counterfactual Evidence framework for Rumor Detection (DCE-RD). Our intuition is to exploit the diverse counterfactual evidence of an event graph to serve as multi-view interpretations, which are further aggregated for robust rumor detection results. Specifically, our method first designs a subgraph generation strategy to efficiently generate different subgraphs of the event graph. We constrain the removal of these subgraphs to cause the change in rumor detection results. Thus, these subgraphs naturally serve as counterfactual evidence for rumor detection. To achieve multi-view interpretation, we design a diversity loss inspired by Determinantal Point Processes (DPP) to encourage diversity among the counterfactual evidence. A GNN-based rumor detection model further aggregates the diverse counterfactual evidence discovered by the proposed DCE-RD to achieve interpretable and robust rumor detection results. Extensive experiments on two real-world datasets show the superior performance of our method. Our code is available at https://github.com/Vicinity111/DCE-RD. Kaiwei Zhang, Junchi Yu, Haichao Shi, Jian Liang 0001, Xiaoyu Zhang 0002 |
KDD | 1 |
| 2023 | Audio-visual aligned saliency model for omnidirectional video with implicit neural representation learning
Dandan Zhu 0001, Xuan Shao, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
Appl. Intell. | 3 |
| 2023 | Human attention based movie summarization: Dataset and baseline modelabstractA movie summarization model can automatically edit a condensed version of a movie by selecting keyframes . Some previous works have proposed some movie summarizers based on traditional methods or recent neural networks and achieved some progress. Despite the demonstrated successes, there are some limitations: (1) previous works mainly resort to hand-crafted heuristics and most of them are unsupervised; (2) currently there is no publicly suitable dataset available for the supervised movie summarization; (3) existing works only focus on the movies themselves while neglecting the audiences, who have the most to say in which part of the movie is more attractive. To break through the aforementioned limitations, we establish a movie summarization dataset Movie50 and propose a novel human attention based annotation pipeline. Furthermore, we propose the A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better simulate human attention as well as exploit more plentiful information. The network is designed, trained end-to-end, and evaluated on the public dataset and our dataset. Extensive experiments demonstrate the superiority of the proposed method. Defang Zhao, Dandan Zhu 0001, Xiongkuo Min, Jiaomin Yue, Kaiwei Zhang, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
Neurocomputing | 5 |
| 2023 | Decoupled dynamic group equivariant filter for saliency prediction on omnidirectional image
Dandan Zhu 0001, Kaiwei Zhang, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
Neurocomputing | 2 |
| 2023 | Implicit Neural Representation Learning for Hyperspectral Image Super-ResolutionabstractHyperspectral image (HSI) super-resolution (SR) without additional auxiliary image remains a constant challenge due to its high-dimensional spectral patterns, where learning an effective spatial and spectral representation is a fundamental issue. Recently, implicit neural representations (INRs) are making strides as a novel and effective representation, especially in the reconstruction task. Therefore, in this work, we propose a novel HSI reconstruction model based on INR which represents HSI by a continuous function mapping a spatial coordinate to its corresponding spectral radiance values. In particular, as a specific implementation of INR, the parameters of the parametric model are predicted by a hypernetwork that operates on feature extraction using a convolution network. It makes the continuous functions map the spatial coordinates to pixel values in a content-aware manner. Moreover, periodic spatial encoding is deeply integrated with the reconstruction procedure, which makes our model capable of recovering more high-frequency details. To verify the efficacy of our model, we conduct experiments on three HSI datasets (CAVE, NUS, and NTIRE2018). Experimental results show that the proposed model can achieve competitive reconstruction performance in comparison with the state-of-the-art methods. In addition, we provide an ablation study on the effect of individual components of our model. We hope this article could serve as a potent reference for future research. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Implicit Neural Representation Learning for Hyperspectral Image Super-ResolutionabstractHyperspectral image (HSI) super-resolution without additional auxiliary image remains a constant challenge due to its high-dimensional spectral patterns, where learning an effective spatial and spectral representation is a fundamental issue. Recently, Implicit Neural Representations (INRs) are making strides as a novel and effective representation, especially in the reconstruction task. Therefore, in this work, we propose a novel HSI reconstruction model based on INR which represents HSI by a continuous function mapping a spatial coordinate to its corresponding spectral radiance values. In particular, as a specific implementation of INR, the parameters of parametric model are predicted by a hypernetwork. It makes the continuous functions map the spatial coordinates to pixel values in a content-aware manner. Moreover, periodic spatial encoding are deeply integrated with the reconstruction procedure, which makes our model capable of recovering more high frequency details. Experimental results on CAVE, NUS, and NTIRE2018 datasets demonstrate the superiority of our model. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
ICME | 1 |
| 2022 | Human Attention Based Movie Summarization: Dataset and Baseline ModelabstractThe movie summarization model can automatically edit a condensed and succinct version of the movie by selecting the keyframes. Previous works mainly resort to hand-crafted heuristics and most of them are unsupervised. Supervised movie summarization is a new research field and, there is currently no publicly suitable dataset available. Moreover, existing works only focus on the movies themselves while neglecting the audiences, who have the most say in which part of the movie is more attractive. To deal with the aforementioned limitations, we establish a human attention based movie summarization dataset Movie50. Specifically, we explore the human attention variations when watching videos and have the following findings: (1) The attention of humans is concentrated when watching keyframes. (2) The attention of humans is distracted when watching non-keyframes. Inspired by these findings, we collect the eye fixations of 20 participants when watching 50 movies and propose a novel human attention based annotation pipeline. In addition, we introduce A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better model human attention as well as exploit more plentiful information. Extensive experiments demonstrate the superiority of the proposed method. Defang Zhao, Dandan Zhu 0001, Xiongkuo Min, Jiaomin Yue, Kaiwei Zhang, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 5 |
| 2021 | Efficient Construction of Public-Key Matrices in Lattice-Based Cryptography: Chaos Strikes Again
Kaiwei Zhang, Ailun Ma, Shanxiang Lyu, Shuting Lou |
ISPEC | 1 |
| 2019 | High Joint Spectral-Spatial Resolution Imaging via Nanostructured Random Broadband FilteringabstractIt is a challenge to acquire a snapshot image of very high resolutions in both spectral and spatial domain via a single short exposure. In this setting one cannot trade time for spectral resolution, such as via spectral bands scanning. Cameras of color filter arrays (CFA) (e.g., the Bayer mosaic) cannot obtain high spectral resolution. To overcome these difficulties, we propose a new multispectral imaging system that makes random linear broadband measurements of the spectrum via a nanostructured multispectral filter array (MSFA). These MS-FA random measurements can be used by sparsity-based recovery algorithms to achieve much higher spectral resolution than conventional CFA cameras, without sacrificing spatial resolution. The key innovation is to jointly exploit both spatial and spectral sparsity properties that are inherent to spectral reflectance of natural objects. Experimental results establish the superior performance of the proposed multispectral imaging system over existing ones. Xiaolin Wu 0001, Dahua Gao, Kaiwei Zhang |
ICIP | 4 |