VLDB 2026 Research / reviewers in the wild / expert
Dandan Zhu 0001
dblp:09/11267-1
· DBLP profile ↗
72ranked-venue papers
16as first author
64since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 10 first-author · 38 since 2021Artificial intelligence and machine learning · 29 · 7 first-author · 25 since 2021Computer networks · 9 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement LearningabstractVideo quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainability, which restrict their applicability in real-world scenarios. To address these challenges, we propose VQAThinker, a reasoning-based VQA framework that leverages large multimodal models (LMMs) with reinforcement learning to jointly model video quality understanding and scoring, emulating human perceptual decision-making. Specifically, we adopt group relative policy optimization (GRPO), a rule-guided reinforcement learning algorithm that enables reasoning over video quality under score-level supervision, and introduce three VQA-specific rewards: (1) a bell-shaped regression reward that increases rapidly as the prediction error decreases and becomes progressively less sensitive near the ground truth; (2) a pairwise ranking reward that guides the model to correctly determine the relative quality between video pairs; and (3) a temporal consistency reward that encourages the model to prefer temporally coherent videos over their perturbed counterparts. Extensive experiments demonstrate that VQAThinker achieves state-of-the-art performance on both in-domain and OOD VQA benchmarks, showing strong generalization for video quality scoring. Furthermore, evaluations on video quality understanding tasks validate its superiority in distortion attribution and quality description compared to existing explainable VQA models and LMMs. These findings demonstrate that reinforcement learning offers an effective pathway toward building generalizable and explainable VQA models solely with score-level supervision. Linhan Cao, Wei Sun 0029, Weixia Zhang, Jun Jia, Kaiwei Zhang, Dandan Zhu 0001, Guangtao Zhai, Xiongkuo Min |
AAAI | 7 |
| 2026 | SalDiff-DTM: A Novel Dual-Temporal Modulated Diffusion Model for Omnidirectional Images Scanpath PredictionabstractScanpath prediction in omnidirectional images (ODIs) serves as a critical component for optimizing foveated rendering efficiency and enhancing interactive quality in virtual reality systems. However, existing scanpath prediction methods for ODIs still suffer from fundamental limitations: (1) inadequate modeling and capturing of long-range temporal dependencies in fixation regions, and (2) suboptimal integration of spatial and temporal visual features, ultimately compromising prediction performance. To address these limitations, we propose a novel Dual-Temporal Modulated Diffusion model for Omnidirectional Images Scanpath Prediction, named SalDiff-DTM model, to effectively generate realistic human eye viewing trajectories. Specifically, to effectively model spatial relationships, we propose a novel Dual-Graph Convolutional Network (Dual-GCN) module that simultaneously captures semantic-level and image-level correlations. By integrating both local spatial details and global contextual information across the internal temporal dimension, this module achieves comprehensive and robust modeling of spatial relationships. To further enhance the modeling of temporal dependencies inherent in diverse fixation patterns, we introduce TABiMamba (Temporal-Aware BiLSTM-Mamba), a dedicated module that synergistically combines the contextual sensitivity of BiLSTM with the long-range sequence modeling capabilities of Mamba. This design facilitates deep information flow and context-aware sequential reasoning, thereby enabling high-fidelity capture of intricate temporal correlations. Inspired by the progressive refinement mechanism of diffusion models in various generative tasks, we propose a saliency-guided diffusion module that formulates the prediction problem as a conditional generative process, iteratively yielding accurate and perceptually plausible scanpaths. Extensive experiments demonstrate that SalDiff-DTM significantly outperforms state-of-the-art models, paving the way for future advancements in eye-tracking technologies and cognitive modeling. Xiaohui Kong, Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min |
AAAI | 3 |
| 2026 | RLSLM: A Hybrid Framework Combining Reinforcement Learning and a Rule-based Social Locomotion Model for Socially-aware NavigationabstractNavigating human-populated environments without causing discomfort is a critical capability for socially-aware agents. While rule-based approaches offer interpretability through predefined psychological principles, they often lack generalizability and flexibility. Conversely, data-driven methods can learn complex behaviors from large-scale datasets, but are typically inefficient, opaque, and difficult to align with human intuitions. To bridge this gap, we propose RLSLM, a hybrid Reinforcement Learning framework that integrates a rule-based Social Locomotion Model, grounded in empirical behavioral experiments, into the reward function of a reinforcement learning framework. The social locomotion model generates an orientation-sensitive social comfort field that quantifies human comfort across space, enabling socially aligned navigation policies with minimal training. RLSLM then jointly optimizes mechanical energy and social comfort, allowing agents to avoid intrusions into personal or group space. A human-agent interaction experiment using an immersive VR-based setup demonstrates that RLSLM outperforms state-of-the-art rule-based models in user experience. Ablation and sensitivity analyses further show the model’s significantly improved interpretability over conventional data-driven methods. This work presents a scalable, human-centered methodology that effectively integrates cognitive science and machine learning for real-world social navigation. Yitian Kou, Yihe Gu, Dandan Zhu 0001, Shu-Guang Kuai |
AAAI | 4 |
| 2026 | CLIP2Pose: Frozen CLIP as Semantic Guide for Domain Adaptive Pose EstimationabstractUnsupervised domain adaptive pose estimation is a fundamental yet challenging task due to the need to transfer from labeled synthetic data to unlabeled real data. Nevertheless, the underlying pose semantics, which are governed by spatial structure, remain largely consistent across domains. This observation motivates the use of vision-language models, which provide domain-invariant representations that align well with high-level semantic concepts. Motivated by this, we propose CLIP2Pose, a novel framework that leverages the semantic robustness of frozen CLIP encoders to facilitate cross-domain generalization. We first introduce a semantic-driven prompt mechanism that encodes structural priors, domain-specific appearance, and instance-level context into the image representation. This guides the model to focus on semantically meaningful and structurally relevant features. Next, we propose a semantic modulation module that adaptively refines visual features by conditioning them on prompt-derived embeddings, enhancing alignment between semantics and visual patterns. To further bridge the modality and domain gaps, we design a directional alignment loss that encourages consistent structural reasoning across both vision and language representations. Extensive experiments on domain adaptive human body and hand pose benchmarks show that CLIP2Pose achieves state-of-the-art performance. Fei Jiang 0006, Dandan Zhu 0001, Jinxin Shi, Aimin Zhou |
AAAI | 3 |
| 2026 | Assessing Personality Consistency in Large Language Models: A Psychometric Framework for Human-Centric Quality of Experience
Yitian Kou, Dandan Zhu 0001, Wei Sun 0029, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai |
QoMEX | 2 |
| 2026 | LEIQ-Assessor: Multi-Dimensional Quality Assessment of Low-Light Enhanced Images via Multi-Task Learning
Wei Sun 0029, Yanwei Jiang, Dandan Zhu 0001, Jinqiu Sang, Jikai Xu, Weixia Zhang, Guangtao Zhai |
QoMEX | 3 |
| 2026 | QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu |
QoMEX | 13 |
| 2026 | URSD: A uncertainty-resistant semi-supervised approach for object detection
Fei Jiang 0006, Dandan Zhu 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Beyond catastrophic forgetting: A continual learning-driven multi-modal fusion model for saliency prediction in dynamic scenes
Jiaqi Wang 0003, Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai |
Expert Syst. Appl. | 3 |
| 2026 | Developing Evolving Adaptability in Biological Intelligence: A Novel Biologically-Inspired Continual Learning Model for Video Saliency PredictionabstractIn the era of deep learning, video saliency prediction task still remains major challenge due to the issue of catastrophic forgetting during feature learning. Most prior works commonly employ generative replay strategies to generate pseudo-samples from previous tasks, enabling them to recall the data distribution. However, scaling up generative replay to accommodate class-incremental and task-incremental settings poses challenges, as generated data with low quality can severely deteriorate performance. Additionally, existing advances mainly focus on preserving memory stability to alleviate catastrophic forgetting, but they remain difficult to flexibly adapt to incremental changes in dynamic scenes. To achieve a better balance between memory stability and learning plasticity, we propose a novel biologically-inspired continual learning (BICL) model tailored to effectively predict human attention in dynamic scenes while mitigate catastrophic forgetting. In particular, inspired by the function of the hippocampus in the human neural system, we elaborately design a visual saliency memory bank module to explicitly store and retrieve representative features from previous tasks. Furthermore, drawing inspiration from the Drosophila $\gamma$γMB system, we propose an active forgetting strategy equipped with multiple parallel adaptive learner modules, which can appropriately attenuate old memories in parameter distribution to enhance learning plasticity to adapt to new tasks, and accordingly to ensure compatibility among multiple learners. Notably, without compromising the performance of old tasks, our proposed model can achieve a better trade-off between memory stability and learning plasticity. Through extensive experiments on several benchmark datasets, our model not only enhances performance in task-incremental settings, but also potentially provides deep insights into neurological adaptive mechanisms. Dandan Zhu 0001, Kaiwei Zhang, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Surveillance Facial Image Quality Assessment: A Multi-Dimensional Dataset and Lightweight ModelabstractSurveillance facial images are often captured under unconstrained conditions, resulting in severe quality degradation due to factors such as low resolution, motion blur, occlusion, and poor lighting. Although recent face restoration techniques applied to surveillance cameras can significantly enhance visual quality, they often compromise fidelity (i.e., identity-preserving features), which directly conflicts with the primary objective of surveillance images -- reliable identity verification. Existing facial image quality assessment (FIQA) predominantly focus on either visual quality or recognition-oriented evaluation, thereby failing to jointly address visual quality and fidelity, which are critical for surveillance applications. To bridge this gap, we propose the first comprehensive study on surveillance facial image quality assessment (SFIQA), targeting the unique challenges inherent to surveillance scenarios. Specifically, we first construct SFIQA-Bench, a multi-dimensional quality assessment benchmark for surveillance facial images, which consists of 5,004 surveillance facial images captured by three widely deployed surveillance cameras in real-world scenarios. A subjective experiment is conducted to collect six dimensional quality ratings, including noise, sharpness, colorfulness, contrast, fidelity and overall quality, covering the key aspects of SFIQA. Furthermore, we propose SFIQA-Assessor, a lightweight multi-task FIQA model that jointly exploits complementary facial views through cross-view feature interaction, and employs learnable task tokens to guide the unified regression of multiple quality dimensions. The experiment results on the proposed dataset show that our method achieves the best performance compared with the state-of-the-art general image quality assessment (IQA) and FIQA methods, validating its effectiveness for real-world surveillance applications. Yanwei Jiang, Wei Sun 0029, Yingjie Zhou 0003, Yuqin Cao, Jun Jia, Sijing Wu, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | MEScan360: A Memory-Enhanced Scanpath Prediction Model for Omnidirectional Images
Dandan Zhu 0001, Kaiwei Zhang, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Pose as a Modality: A Psychology-Inspired Network for Personality Recognition with a New Multimodal DatasetabstractIn recent years, predicting Big Five personality traits from multimodal data has received significant attention in artificial intelligence (AI). However, existing computational models often fail to achieve satisfactory performance. Psychological research has shown a strong correlation between pose and personality traits, yet previous research has largely ignored pose data in computational models. To address this gap, we develop a novel multimodal dataset that incorporates full-body pose data. The dataset includes video recordings of 287 participants completing a virtual interview with 36 questions, along with self-reported Big Five personality scores as labels. To effectively utilize this multimodal data, we introduce the Psychology-Inspired Network (PINet), which consists of three key modules: Multimodal Feature Awareness (MFA), Multimodal Feature Interaction (MFI), and Psychology-Informed Modality Correlation Loss (PIMC Loss). The MFA module leverages the Vision Mamba Block to capture comprehensive visual features related to personality, while the MFI module efficiently fuses the multimodal features. The PIMC Loss, grounded in psychological theory, guides the model to emphasize different modalities for different personality dimensions. Experimental results show that the PINet outperforms several state-of-the-art baseline models. Furthermore, the three modules of PINet contribute almost equally to the model’s overall performance. Incorporating pose data significantly enhances the model’s performance, with the pose modality ranking mid-level in importance among the five modalities. These findings address the existing gap in personality-related datasets that lack full-body pose data and provide a new approach for improving the accuracy of personality prediction models, highlighting the importance of integrating psychological insights into AI frameworks. Bin Tang 0010, Keqi Pan, Miao Zheng, Jialu Sui, Dandan Zhu 0001, Cheng-Long Deng, Shu-Guang Kuai |
AAAI | 6 |
| 2025 | Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D GraphicsabstractTextured meshes significantly enhance the realism and detail of objects by mapping intricate texture details onto the geometric structure of 3D models. This advancement is valuable across various applications, including entertainment, education, and industry. While traditional mesh saliency studies focus on non-textured meshes, our work explores the complexities introduced by detailed texture patterns. We present a new dataset for textured mesh saliency, created through an innovative eye-tracking experiment in a six degrees of freedom (6-DOF) VR environment. This dataset addresses the limitations of previous studies by providing comprehensive eye-tracking data from multiple viewpoints, thereby advancing our understanding of human visual behavior and supporting more accurate and effective 3D content creation. Our proposed model predicts saliency maps for textured mesh surfaces by treating each triangular face as an individual unit and assigning a saliency density value to reflect the importance of each local surface region. The model incorporates a texture alignment module and a geometric extraction module, combined with an aggregation module to integrate texture and geometry for precise saliency prediction. We believe this approach will enhance the visual fidelity of geometric processing while ensuring computational efficiency, essential for real-time rendering and high-detail applications such as VR and gaming. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
AAAI | 2 |
| 2025 | BIDP: Brain-Inspired Dual-Process CNN-Transformer for Salient Object Detection
Wenyi Wu, Chen Liao, Qiangqiang Zhou, Dandan Zhu 0001, Xinping Rao |
CGI (2) | 4 |
| 2025 | Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured MeshesabstractMesh saliency enhances the adaptability of 3D vision by identifying and emphasizing regions that naturally attract visual attention. To investigate the interaction between geometric structure and texture in shaping visual attention, we establish a comprehensive mesh saliency dataset, which is the first to systematically capture the differences in saliency distribution under both textured and non-textured visual conditions. Furthermore, we introduce mesh Mamba, a unified saliency prediction model based on a state space model (SSM), designed to adapt across various mesh types. Mesh Mamba effectively analyzes the geometric structure of the mesh while seamlessly incorporating texture features into the topological framework, ensuring coherence throughout appearance-enhanced modeling. More importantly, by sub-graph embedding and a bidirectional SSM, the model enables global context modeling for both local geometry and texture, preserving the topological structure and improving the understanding of visual details and structural complexity. Through extensive theoretical and empirical validation, our model not only improves performance across various mesh types but also demonstrates high scalability and versatility, particularly through cross validations of various visual features. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
CVPR | 2 |
| 2025 | FDFRL: Credit Card Fraud Detection Based on Federated Reinforcement Learning
Kun Zhu 0024, Dandan Zhu 0001 |
ICANN (4) | 4 |
| 2025 | A Novel Framework for Realistic 3D Scene Regeneration with Graph of ThoughtsabstractIn embodied intelligence applications, highly realistic 3D scenes lay the foundation for perception and decision-making, while 3D scene regeneration creates more coherent and personalized virtual spaces, facilitating more efficient task adaptation and agent training. To address this, we propose a reasoning framework based on the Graph of Thoughts (GoT), which enhances the prompting capabilities of large language models (LLM) and integrates a synergistic mechanism of retrospective memory and feedback loops into the regeneration process. During the initial generation phase, we retain the Holodeck paradigm, combining LLM-driven scene design inferences with the spatial layout of 3D assets from Objaverse. In the regeneration phase, dynamic feedback loops trigger backtracking of reasoning memory to adjust relevant elements according to evolving requirements, while maintaining stability and consistency in unrelated elements, ensuring the scene’s overall coherence. We conduct both subjective and objective experiments to validate the effectiveness of this framework, demonstrating significant improvements in 3D scene generation. Yitian Kou, Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
ICME | 3 |
| 2025 | LD2Scan: A Lightweight Dual-Temporal Constrained Scanpath Prediction Model for Omnidirectional ImagesabstractPredicting scanpaths in omnidirectional images (ODIs) is essential for simulating human gaze behaviors. However, current methods often struggle with long-term dependencies and exhibit high complexity, which limits their efficiency and scalability. To tackle these challenges, we propose LD2Scan, a lightweight diffusion-based model specifically designed for scanpath prediction in ODIs. It employs Efficient Equivariant (E4) convolution to enhance feature extraction from distorted ODIs while improving computational performance, thereby reducing resource demands. LD2Scan utilizes a dual-graph convolutional network (GCN) to enforce internal time constraints between fixations, integrating semantic-level GCN for sequential fixation modeling and image-level GCN to capture relationships across different images, enriching contextual information. We formulate the scanpath prediction issue as a conditional generation task, refining noisy scanpaths using features encoded by the dual-GCN and robust E4-processed features. Experimental results on several benchmark datasets demonstrate that LD2Scan outperforms existing methods in terms of both accuracy and efficiency. Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai |
ICME | 3 |
| 2025 | MEScan360: A Memory-Enhanced Scanpath Prediction Model for Omnidirectional ImagesabstractScanpath prediction for omnidirectional images (ODIs) aims to capture the dynamic human visual attention. However, the complicated gaze behavior and inevitable projection distortion make scanpath prediction in ODIs extremely challenging. Most existing models neither capture the long-term dependencies across visual states nor fully incorporate historical memory information, leading to limited performance. To this end, we propose MEScan360, a memory-enhanced scanpath prediction model for ODIs. We introduce two key innovations: long-term memory storage unit and memory interaction module. These two components establish a more explicit link between past visual information and current visual inputs, thereby significantly enhancing the performance of scanpath prediction. Furthermore, a robust feature extraction module is designed to extract semantic feature precisely from distorted ODIs with a more lightweight structure. Extensive experiments on several benchmark datasets demonstrate that our proposed model achieves competitive performance in both accuracy and efficiency. Dandan Zhu 0001, Kaiwei Zhang, Fei Jiang 0006, Guangtao Zhai |
ICME | 2 |
| 2025 | HandNet: Occlusion-robust 3D hand mesh reconstruction with prior information
Fei Jiang 0006, Dandan Zhu 0001, Aimin Zhou |
Knowl. Based Syst. | 3 |
| 2025 | Robust long-tailed recognition with distribution-aware adversarial example generation
Bo Li 0126, Yongqiang Yao, Jingru Tan, Dandan Zhu 0001, Ruihao Gong, Ye Luo 0004 |
Neural Networks | 4 |
| 2025 | ScanDTM: A Novel Dual-Temporal Modulation Scanpath Prediction Model for Omnidirectional ImagesabstractScanpath prediction for omnidirectional images aims to effectively simulate the human visual perception mechanism to generate dynamic realistic fixation trajectories. However, the majority of scanpath prediction methods for omnidirectional images are still in their infancy as they fail to accurately capture the time-dependency of viewing behavior and suffer from sub-optimal performance along with limited generalization capability. A desirable solution should achieve a better trade-off between prediction performance and generalization ability. To this end, we propose a novel dual-temporal modulation scanpath prediction (ScanDTM) model for omnidirectional images. Such a model is designed to effectively capture long-range time-dependencies between various fixation regions across both internal and external time dimensions, thereby generating more realistic scanpaths. In particular, we design a Dual Graph Convolutional Network (Dual-GCN) module comprising a semantic-level GCN and an image-level GCN. This module servers as a robust visual encoder that captures spatial relationships among various object regions within an image and fully utilizes similar images as complementary information to capture similarity relations across relevant images. Notably, the proposed Dual-GCN focuses on modeling temporal correlations from both local and global perspectives within the internal time dimension. Furthermore, drawing inspiration from the promising generalization capabilities of diffusion models across various generative tasks, we introduce a novel diffusion-guided saliency module. This module formulates the prediction issue as a conditional generative process for the saliency map, utilizing extracted semantic-level and image-level visual features as conditions. With the well-designed diffusion-guided saliency module, our proposed ScanDTM model acting as an external temporal modulator, we can progressively refine the generated scanpath from the noisy map. We conduct extensive experiments on several benchmark datasets, and the results demonstrate that our ScanDTM model significantly outperforms other competitors. Meanwhile, when applied to tasks such as saliency prediction and image quality assessment, our ScanDTM model consistently achieves superior generalization performance. Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | From Haziness to Clarity: A Novel Iterative Memory-Retrospective Emergence Model for Omnidirectional Image Saliency PredictionabstractTo achieve saliency prediction in omnidirectional images (ODIs), the majority of prior works typically adopt the convolutional neural networks (CNNs)-based saliency models to extract semantic features to predict prominent regions in ODIs. Albeit achieving substantially performance gains, these works all employed purely visual computing paradigms and ignore to explore the nature of human visual attention mechanisms. In other words, existing saliency prediction works for ODIs are insufficient to capture the biological characteristics of the visual attention mechanism in the human brain. To establish a more explicit link between saliency prediction performance and brain-like visual attention mechanism, we simulate the mechanism of human retrospective memory in neuropsychology and propose IMRE model, a novel iterative memory-retrospective emergence model can predict and infer the salient features by recalling previously learned information. In IMRE model, we introduce four key modules to simulate the visual attention mechanism for predicting human fixations in the human brain. Firstly, the visual stimulus response module is designed to effectively extract semantic features and capture the intricate relationship between these features, acting as the human visual cortex. Secondly, the retrospective integration module serves to distill valuable information from a fuzzy memory ensemble, resembling the role of the basal ganglia in the neural system. Thirdly, the memory bank module explicitly records and stores subconscious response information and learned knowledge, acting like the hippocampus in neural system. Lastly, the prospective inference module accurately infers saliency maps from the refined useful information, resembling the role of the prefrontal cortex. During prediction, we utilize the introduced memory bank to retrieve and recall previously learned information, which simulates the process of memory emergence from haziness to clarity. Such a process aligns with the retrospective memory mechanism of the human brain. To validate the superiority of the proposed model in ODIs saliency prediction tasks, we conduct extensive experiments on two benchmark datasets. Experiments show impressive performances that IMRE model outperforms other state-of-the-art methods across all benchmark datasets. Importantly, experiments also highlight the IMRE model's ability to trace back to specific instances during prediction, thereby reducing model inference costs and enhancing interpretability. Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Explain Vision Focus: Blending Human Saliency Into Synthetic Face ImagesabstractSynthetic faces have been extensively researched and applied in various fields, such as face parsing and recognition. Compared to real face images, synthetic faces engender more controllable and consistent experimental stimuli due to the ability to precisely merge expression animations onto the facial skeleton. Accordingly, we establish an eye-tracking database with 780 synthetic face images and fixation data collected from 22 participants. The use of synthetic images with consistent expressions ensures reliable data support for exploring the database and determining the following findings: (1) A correlation study between saliency intensity and facial movement reveals that the variation of attention distribution within facial regions is mainly attributed to the movement of the mouth. (2) A categorized analysis of different demographic factors demonstrates that the bias towards salient regions aligns with differences in some demographic categories of synthetic characters. In practice, inference of facial saliency distribution is commonly used to predict the regions of interest for facial video-related applications. Therefore, we propose a benchmark model that accurately predicts saliency maps, closely matching the ground truth annotations. This achievement is made possible by utilizing channel alignment and progressive summation for feature fusion, along with the incorporation of Sinusoidal Position Encoding. The ablation experiment also demonstrates the effectiveness of our proposed model. We hope that this paper will contribute to advancing the photorealism of generative digital humans. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Huiyu Duan, Guangtao Zhai |
IEEE Trans. Multim. | 2 |
| 2025 | Elevating Mesh Saliency in VR: Introducing a Novel Prediction Network and DatasetabstractIn computer graphics, polygon meshes stand out as a popular representation providing effective delineation of delicate textures and complex geometries. When dealing with geometric processing tasks for critical regions of the mesh, it is necessary to consider the human visual perception related to saliency. Therefore, we establish a novel mesh saliency dataset, facilitated by a more comprehensive gathering pipeline of eye-tracking from subjects observing mesh models at arbitrary viewpoints in a virtual reality space with six degrees of freedom. Additionally, we propose a mesh saliency prediction model that accurately infers visual attention density maps for complex and irregular mesh surfaces. This model integrates surface curvature and triangular face shape information from multi-scale neighboring ranges as local geometric features, while also leveraging surface spatial positioning as a global feature. Our work aims to preserve critical areas and minimize visual loss in saliency-driven tasks such as mesh simplification, rendering, and texturing. We believe that our research can offer valuable insights for human-centered mesh computation applications. Kaiwei Zhang, Mohan He, Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Audio-Visual Saliency Prediction Model with Implicit Neural RepresentationabstractWith the remarkable advancement of deep learning techniques and the wide availability of large-scale datasets, the performance of audio-visual saliency prediction has been drastically improved. Actually, audio-visual saliency prediction is still at an early exploration stage due to the spatial-temporal signal complexity and dynamic continuity of video content. To our knowledge, most existing audio-visual saliency prediction approaches usually represent videos as 3D grid of RGB values using discrete convolutional neural networks (CNNs), which inevitably incurs video content-agnostic and ignores the dynamic continuity issues. This article proposes a novel parametric audio-visual saliency (PAVS) model with implicit neural representation (INR) to address the aforementioned problems. Specifically, by using the proposed parametric neural network, we can effectively encode the space-time coordinates of video frames into corresponding saliency values, which can significantly enhance the compact feature representation ability. Meanwhile, a parametric feature fusion method is developed to achieve intrinsic interactions between audio and visual information streams, which can adaptively fuse audio and visual features to obtain competitive performance. Notably, without resorting to any specific audio-visual feature fusion strategy, the proposed PAVS model outperforms other state-of-the-art saliency methods by a large margin. Dandan Zhu 0001, Kun Zhu 0024, Guangtao Zhai, Xiaokang Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Unified Approach to Mesh Saliency: Evaluating Textured and Non-Textured Meshes Through VR and Multifunctional PredictionabstractMesh saliency aims to empower artificial intelligence with strong adaptability to highlight regions that naturally attract visual attention. Existing advances primarily emphasize the crucial role of geometric shapes in determining mesh saliency, but it remains challenging to flexibly sense the unique visual appeal brought by the realism of complex texture patterns. To investigate the interaction between geometric shapes and texture features in visual perception, we establish a comprehensive mesh saliency dataset, capturing saliency distributions for identical 3D models under both non-textured and textured conditions. Additionally, we propose a unified saliency prediction model applicable to various mesh types, providing valuable insights for both detailed modeling and realistic rendering applications. This model effectively analyzes the geometric structure of the mesh while seamlessly incorporating texture features into the topological framework, ensuring coherence throughout appearance-enhanced modeling. Through extensive theoretical and empirical validation, our approach not only enhances performance across different mesh types, but also demonstrates the model's scalability and generalizability, particularly through cross-validation of various visual features. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | E2DAS: An Efficient Equivariant Dynamic Aggregation Saliency Model for Omnidirectional Images
Dandan Zhu 0001, Kun Zhu 0024, Guangtao Zhai, Xiaokang Yang 0001 |
ICPR (3) | 3 |
| 2024 | TDiffSal: Text-Guided Diffusion Saliency Prediction Model for Images
Dandan Zhu 0001, Kun Zhu 0024, Guangtao Zhai |
ICPR (8) | 3 |
| 2024 | IAPCP: An Effective Cross-Project Defect Prediction Model via Intra-Domain Alignment and Programming-Based Distribution AdaptationabstractCross‐project defect prediction (CPDP) aims to identify defect‐prone software instances in one project (target) using historical data collected from other software projects (source), which can help maintainers allocate limited testing resources reasonably. Unfortunately, the feature distribution discrepancy between the source and target projects makes it challenging to transfer the matching feature representation and severely hinders CPDP performance. Besides, existing CPDP models require an intensively expensive and time‐consuming process to tune a lot of parameters. To address the above limitations, we propose an effective CPDP model named IAPCP based on distribution adaptation in this study, which consists of two stages: correlation alignment and intra‐domain programming. Correlation alignment first calculates the covariance matrices of the source and target projects and then erases some features of the source project (i.e., whitening operation) and employs the features of the target project (i.e., target covariance) to fill the source project, thereby well aligning the source and target feature distributions and reducing the distribution discrepancy across projects. Intra‐domain programming can directly learn a nonparametric linear transfer defect predictor with strong discriminative capacity by solving a probabilistic annotation matrix (PAM) based on the adjusted features of the source project. The model does not require model selection and parameter tuning. Extensive experiments on a total of 82 cross‐project pairs from 16 software projects demonstrate that IAPCP can achieve competitive CPDP effectiveness and efficiency compared with multiple state‐of‐the‐art baseline models. Kun Zhu 0024, Dandan Zhu 0001 |
IET Softw. | 3 |
| 2024 | BEVSOC: Self-Supervised Contrastive Learning for Calibration-Free BEV 3-D Object Detectionabstract3D object detection based on multi-view cameras and bird’s-eye view (BEV) representation is a key task for autonomous driving, as it enables the perception systems to understand the surrounding scenes. However, most existing BEV representation methods rely on the projection matrix of camera intrinsic and extrinsic parameters, which requires a complex and time-consuming calibration process that may introduce errors and degrade the detection performance. Moreover, the calibration results may vary due to environmental changes and affect the stability of the detection system. To address this problem, we propose a calibration-free 3D object detection method that leverages a group-equivariant convolutional network to extract features from multi-view images and a projection network module to learn the implicit 3D-to-2D projection relationship for obtaining BEV representation. Furthermore, we employ contrastive learning to pre-train the projection network module without using manually annotated data. By exploiting the multi-view camera data through contrastive learning, our proposed method eliminates the need for tedious calibration, avoids calibration errors, and reduces the dependence on a large amount of annotated data for calibration-free 3D object detection. We evaluate our method on the nuScenes dataset and demonstrate its competitive performance. Our method improves the stability and reliability of 3D object detection in long-term autonomous driving. Yongqing Chen, Nanyu Li, Dandan Zhu 0001, Charles Zhou, Zhuhua Hu, Yong Bai 0002, Jun Yan 0009 |
IEEE Internet Things J. | 3 |
| 2024 | WSBCV: A data-driven cross-version defect model via multi-objective optimization and incremental representation learning
Kun Zhu 0024, Weiping Ding 0001, Dandan Zhu 0001 |
Inf. Sci. | 4 |
| 2024 | IMDAC: A robust intelligent software defect prediction model via multi-objective optimization and end-to-end hybrid deep learning networksabstractAbstract Software defect prediction (SDP) aims to build an effective prediction model for historical defect data from software repositories by some specialized techniques or algorithms, and predict the defect proneness of new software modules. Nevertheless, the complex internal intrinsic structure hidden behind the defect data makes it challenging for the built prediction model to capture the most expressive defect feature representations, and largely limits the SDP performance. Fortunately, artificial intelligence is interacting closely with humans and provides powerful intelligent technical support for addressing these SDP issues. In this article, we propose a robust intelligent SDP model called IMDAC based on deep learning and soft computing techniques. This model has three main advantages: (1) an effective deep generative network—InfoGAN (information maximizing GANs) is employed to conduct data augmentation, namely generating sufficient defect instances and achieving defect class balance simultaneously. (2) Select the fewest representative feature subset for the minimum error via an advanced multi‐objective optimization approach—MSEA (multi‐stage evolutionary algorithm). (3) Build a powerful end‐to‐end deep defect predictor by hybrid deep learning techniques—DAE (Denoising AutoEncoder) and CNN (convolutional neural network), which can not only reconstruct a clean “repaired” input with strong robustness and generalization capabilities via DAE, but also learn the abstract deep semantic features with strong discriminating capability via CNN. Experimental results verify the superiority and robustness of the IMDAC model across 15 software projects. Kun Zhu 0024, Changjun Jiang 0002, Dandan Zhu 0001 |
Softw. Pract. Exp. | 4 |
| 2024 | Synergetic Assessment of Quality and Aesthetic: Approach and Comprehensive Benchmark DatasetabstractQuantifications of image quality and aesthetic have been regarded as two independent fields in computer vision. Generally, image quality assessment aims at measuring image distortions and image aesthetic is judged by commonly established photography rules. However, either measuring image quality or aesthetic alone is not sufficient to qualitatively rank images. Therefore, this paper puts forward the synergetic assessment of quality and aesthetic to help understand the subjective human preferences of digital pictures more comprehensively. Specifically, considering that the images of existing benchmark datasets are only labeled with single attribute, we first establish a new dataset which contains 9042 real-world images with the corresponding human rated pair-wise quality-aesthetic scores. Previously, these images are only labeled with aesthetic score, and we evaluate the subjective quality score of them, so that it can make up the lack of image dataset with double attributes. Moreover, since the existing methods are mostly designed for individual attribute prediction. We then propose a two-stream learning network to assess both quality and aesthetic of images in parallel. This network follows the top-down perception mechanism which learns from both fined grained details and holistic image layout simultaneously. Furthermore, we introduce a Channel-Diversity loss, which can be deployed in grouped convolution operation, and can constrain channels to be mutually exclusive across the spatial dimensions. To some extent, this contributes to spotlight different local discriminative regions with a finer granularity. Finally, experiments demonstrate that our method outperforms the state-of-the-art methods on our established benchmark dataset and other benchmark datasets in terms of image quality and aesthetic assessment. We hope this paper could serve as a potent reference and be useful for future research on the study of image ranking. Both the benchmark dataset and the code will be publicly available to facilitate further research. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Zhongpai Gao, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Unified Audio-Visual Saliency Model for Omnidirectional Videos With Spatial AudioabstractSpatial audio is a crucial component of omnidirectional videos (ODVs), which can provide an immersive experience by enabling viewers to perceive sound sources in all directions. However, most visual attention modeling works for ODVs focus only on visual cues, and audio modality is rather rarely considered. Additionally, the existing audio-visual saliency models for ODVs lack spatial audio location-awareness (i.e. sound source location-agnostic) and audio content attributes discriminability (i.e. audio content attributes-agnostic). To this end, we propose a novel audio-visual perception saliency (AVPS) model with spatial audio location-awareness and audio content attributes-adaptive to efficiently address the problem of fixation prediction in ODVs. Specifically, we first utilize the improved group equivariant convolutional neural network (G-CNN) with eidetic 3D LSTM (E3D-LSTM) to extract spatial-temporal visual features. Then we perceive sound source locations by computing the audio energy map (AEM) of the audio information in ODVs. Subsequently, we introduce SoundNet to extract audio features with multiple attributes. Finally, we develop an audio-visual feature fusion module to adaptively integrate spatial-temporal visual features and spatial auditory information to generate the final audio-visual saliency map. Extensive experiments in three audio modalities validate the effectiveness of the proposed model. Meanwhile, the performance of the proposed model is superior to the other 10 state-of-the-art saliency models. Dandan Zhu 0001, Kaiwei Zhang, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Hidden Barcode in Sub-Images with Invisible Locating MarkerabstractThe prevalence of the Internet of Things (IoT) has led to the widespread adoption of 2D barcodes as a means of offline-to-online communication. Whereas, 2D barcodes are not ideal for publicity materials, due to their space-consuming nature. Recent works have proposed 2D image barcodes that contain invisible codes or hyperlinks to transmit hidden information from offline to online. However, these methods undermine the purpose of the codes being invisible, due to the the requirement of markers to locate them. The conference version of this work has presents a novel imperceptible information embedding framework for display or print-camera scenarios, which includes not only hiding and recvoery but also locating and correcting. With the assistance of learned invisible markers, hidden codes can be rendered truly imperceptible. A highly effective multi-stage training scheme is proposed to achieve high visual fidelity and retrieval resiliency, wherein information is concealed in a sub-region rather than the entire image. However, our conference version does not address the optimal sub-region for hiding, which is crucial when dealing with local region concealment problems. In this paper extension, we consider human perceptual characteristics and introduce an optimal hiding region recommendation algorithm that comprehensively incorporates Just Noticeable Difference (JND) and visual saliency factors into consideration. Extensive experiments demonstrate superior visual quality and robustness compared to state-of-the-art methods. With the assistance of our proposed hiding region recommendation algorithm, concealed information becomes even less visible than the results of our conference version without compromising robustness. Jun Jia, Zhongpai Gao, Yiwei Yang 0007, Wei Sun 0029, Dandan Zhu 0001, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Learning Golf Swing Key Events from Gaussian Soft Labels Using Multi-Scale Temporal MLPFormerabstractA complete golf swing includes several key events. The standardization of poses in each key event is directly related to the hitting effect. Thus, it is meaningful for the players to analyze their poses, especially at key frames, so as to improve swing performances. With the rapid development of deep learning techniques in computer vision, we are able to detect key frames during a golf swing. In this paper, we propose a framework to recognize key events in golf swing based on pure monocular video data. To achieve this, we have combined attention mechanism in the backbone network to extract concise features and leveraged the transformer structure to fuse multi-scale temporal information to enhance the feature representation. Besides, we also introduce Gaussian kernels into the label generation process, which can effectively solve the problem of ambiguity in detecting key events within their neighbouring similar frames. Notably, our method achieves an average recognition accuracy of 83.4% (+7.3% compared with SwingNet) for eight golf swing events on GoIfDB dataset. Yanting Zhang 0001, Fuyu Tu, Zijian Wang 0010, Dandan Zhu 0001 |
IJCNN | 5 |
| 2023 | GET: group equivariant transformer for person detection of overhead fisheye images
Yongqing Chen, Dandan Zhu 0001, Nanyu Li, Yong Bai 0002 |
Appl. Intell. | 2 |
| 2023 | Audio-visual aligned saliency model for omnidirectional video with implicit neural representation learning
Dandan Zhu 0001, Xuan Shao, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
Appl. Intell. | 1 |
| 2023 | Human attention based movie summarization: Dataset and baseline modelabstractA movie summarization model can automatically edit a condensed version of a movie by selecting keyframes . Some previous works have proposed some movie summarizers based on traditional methods or recent neural networks and achieved some progress. Despite the demonstrated successes, there are some limitations: (1) previous works mainly resort to hand-crafted heuristics and most of them are unsupervised; (2) currently there is no publicly suitable dataset available for the supervised movie summarization; (3) existing works only focus on the movies themselves while neglecting the audiences, who have the most to say in which part of the movie is more attractive. To break through the aforementioned limitations, we establish a movie summarization dataset Movie50 and propose a novel human attention based annotation pipeline. Furthermore, we propose the A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better simulate human attention as well as exploit more plentiful information. The network is designed, trained end-to-end, and evaluated on the public dataset and our dataset. Extensive experiments demonstrate the superiority of the proposed method. Defang Zhao, Dandan Zhu 0001, Xiongkuo Min, Jiaomin Yue, Kaiwei Zhang, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
Neurocomputing | 2 |
| 2023 | Decoupled dynamic group equivariant filter for saliency prediction on omnidirectional image
Dandan Zhu 0001, Kaiwei Zhang, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
Neurocomputing | 1 |
| 2023 | Application of QR Code Watermarking and Encryption in the Protection of Data Privacy of Intelligent Mouth-Opening TrainerabstractQuick response (QR) codes are widely used in offline to online channels to transfer information from promotional materials to mobile devices. Self-service medical equipment can record the data of each test, so the use of QR codes can realize the data exchange between patients and doctors, medical institutions, and self-service medical equipment, and create a medical information platform for health files. However, since anyone can easily read the information in the QR code, it is not conducive to the protection of patient privacy. Therefore, we propose a QR code encryption and decryption model based on robust digital watermarking. We implement digital watermarking through the generative adversarial networks and increase the robustness of the watermark by adding noise to the model. At the same time, we encrypt and decrypt the QR code information through advanced encryption standards. Experimental results show that the proposed method can well protect the privacy of patients without affecting the data acquisition by patients and doctors. Jiannan Liu, Jun Jia, Dandan Zhu 0001, Guangtao Zhai |
IEEE Internet Things J. | 5 |
| 2023 | Implicit Neural Representation Learning for Hyperspectral Image Super-ResolutionabstractHyperspectral image (HSI) super-resolution (SR) without additional auxiliary image remains a constant challenge due to its high-dimensional spectral patterns, where learning an effective spatial and spectral representation is a fundamental issue. Recently, implicit neural representations (INRs) are making strides as a novel and effective representation, especially in the reconstruction task. Therefore, in this work, we propose a novel HSI reconstruction model based on INR which represents HSI by a continuous function mapping a spatial coordinate to its corresponding spectral radiance values. In particular, as a specific implementation of INR, the parameters of the parametric model are predicted by a hypernetwork that operates on feature extraction using a convolution network. It makes the continuous functions map the spatial coordinates to pixel values in a content-aware manner. Moreover, periodic spatial encoding is deeply integrated with the reconstruction procedure, which makes our model capable of recovering more high-frequency details. To verify the efficacy of our model, we conduct experiments on three HSI datasets (CAVE, NUS, and NTIRE2018). Experimental results show that the proposed model can achieve competitive reconstruction performance in comparison with the state-of-the-art methods. In addition, we provide an ablation study on the effect of individual components of our model. We hope this article could serve as a potent reference for future research. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | RIVIE: Robust Inherent Video Information EmbeddingabstractImagine an interesting situation when watching a movie, we can scan the screen using our smartphones to get some extra information about this movie such as the cast, the release date, the movie's homepage, etc. Our prospect is a world where each video contains invisible information that can be delivered to us through mobile devices with cameras. This paper proposes the first deep learning-based information hiding method for videos to achieve information transmission from screens to cameras. Compared with hiding information in single images, the methods for videos need to maintain visual quality in both spatial and temporal domains. Furthermore, the training of video models builds on a large video dataset, which needs much more computational resources than training models for images. To reduce the computational complexity, we propose to simulate data on-the-fly to generate simulated sequences from single images. Then, we use the simulated data to train a spatio-temporal generator that hides information in videos while maintaining visual quality. During training, a temporal loss function based on the simulated data is exploited to ensure the temporal consistency of generated videos. After embedding, we use a decoder to recover the hidden information. To simulate the imaging pipeline from screens to cameras in the real world, we insert a distortion network between the generator and decoder. The distortion network is based on differentiable 3D rendering to cover possible distortions introduced in the procedure of camera imaging. Experimental results show that the hidden information in videos can be extracted by cameras without impacting the visual quality. Our work can be applied to many fields, such as advertisement, entertainment, and education. Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Menghan Hu, Guangtao Zhai |
IEEE Trans. Multim. | 3 |
| 2023 | SLAM for Indoor Parking: A Comprehensive Benchmark Dataset and a Tightly Coupled Semantic FrameworkabstractFor the task of autonomous indoor parking, various Visual-Inertial Simultaneous Localization And Mapping (SLAM) systems are expected to achieve comparable results with the benefit of complementary effects of visual cameras and the Inertial Measurement Units. To compare these competing SLAM systems, it is necessary to have publicly available datasets, offering an objective way to demonstrate the pros/cons of each SLAM system. However, the availability of such high-quality datasets is surprisingly limited due to the profound challenge of the groundtruth trajectory acquisition in the Global Positioning Satellite denied indoor parking environments. In this article, we establish BeVIS, a large-scale Be nchmark dataset with V isual (front-view), I nertial and S urround-view sensors for evaluating the performance of SLAM systems developed for autonomous indoor parking, which is the first of its kind where both the raw data and the groundtruth trajectories are available. In BeVIS, the groundtruth trajectories are obtained by tracking artificial landmarks scattered in the indoor parking environments, whose coordinates are recorded in a surveying manner with a high-precision Electronic Total Station. Moreover, the groundtruth trajectories are comprehensively evaluated in terms of two respects, the reprojection error and the pose volatility, respectively. Apart from BeVIS, we propose a novel tightly coupled semantic SLAM framework, namely VIS SLAM -2, leveraging V isual (front-view), I nertial, and S urround-view sensor modalities, specially for the task of autonomous indoor parking. It is the first work attempting to provide a general form to model various semantic objects on the ground. Experiments on BeVIS demonstrate the effectiveness of the proposed VIS SLAM -2. Our benchmark dataset BeVIS is publicly available at https://shaoxuan92.github.io/BeVIS . Xuan Shao, Ying Shen 0005, Lin Zhang 0014, Shengjie Zhao 0001, Dandan Zhu 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Toward Visual Behavior and Attention Understanding for Augmented 360 Degree VideosabstractAugmented reality (AR) overlays digital content onto reality. In an AR system, correct and precise estimations of user visual fixations and head movements can enhance the quality of experience by allocating more computational resources for analyzing, rendering, and 3D registration on the areas of interest. However, there is inadequate research to help in understanding the visual explorations of the users when using an AR system or modeling AR visual attention. To bridge the gap between the saliency prediction on real-world scenes and on scenes augmented by virtual information, we construct the ARVR saliency dataset. The virtual reality (VR) technique is employed to simulate the real-world. Annotations of object recognition and tracking as augmented contents are blended into omnidirectional videos. The saliency annotations of head and eye movements for both original and augmented videos are collected and together constitute the ARVR dataset. We also design a model that is capable of solving the saliency prediction problem in AR. Local block images are extracted to simulate the viewport and offset the projection distortion. Conspicuous visual cues in the local block images are extracted to constitute the spatial features. The optical flow information is estimated as an important temporal feature. We also consider the interplay between virtual information and reality. The composition of the augmentation information is distinguished, and the joint effects of adversarial augmentation and complementary augmentation are estimated. The Markov chain is constructed with block images as graph nodes. In the determination of the edge weights, both the characteristics of the viewing behaviors and the visual saliency mechanisms are considered. The order of importance for block images is estimated through the state of equilibrium of the Markov chain. Extensive experiments are conducted to demonstrate the effectiveness of the proposed method. Yucheng Zhu, Xiongkuo Min, Dandan Zhu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Ke Gu 0001, Jiantao Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | A Novel Lightweight Audio-visual Saliency Model for VideosabstractAudio information has not been considered an important factor in visual attention models regardless of many psychological studies that have shown the importance of audio information in the human visual perception system. Since existing visual attention models only utilize visual information, their performance is limited but also requires high-computational complexity due to the limited information available. To overcome these problems, we propose a lightweight audio-visual saliency (LAVS) model for video sequences. To the best of our knowledge, this article is the first trial to utilize audio cues for an efficient deep-learning model for the video saliency estimation. First, spatial-temporal visual features are extracted by the lightweight receptive field block (RFB) with the bidirectional ConvLSTM units. Then, audio features are extracted by using an improved lightweight environment sound classification model. Subsequently, deep canonical correlation analysis (DCCA) aims at capturing the correspondence between audio and spatial-temporal visual features, thus obtaining a spatial-temporal auditory saliency. Lastly, the spatial-temporal visual and auditory saliency are fused to obtain the audio-visual saliency map. Extensive comparative experiments and ablation studies validate the performance of the LAVS model in terms of effectiveness and complexity. Dandan Zhu 0001, Xuan Shao, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Learning Invisible Markers for Hidden Codes in Offline-to-online PhotographyabstractQR (quick response) codes are widely used as an offline-to-online channel to convey information (e.g., links) from publicity materials (e.g., display and print) to mobile devices. However, QR codes are not favorable for taking up valuable space of publicity materials. Recent works propose invisible codes/hyperlinks that can convey hidden information from offline to online. However, they require markers to locate invisible codes, which fails the purpose of invisible codes to be visible because of the markers. This paper proposes a novel invisible information hiding architecture for display/print-camera scenarios, consisting of hiding, locating, correcting, and recovery, where invisible markers are learned to make hidden codes truly invisible. We hide information in a sub-image rather than the entire image and include a localization module in the end-to-end framework. To achieve both high visual quality and high recovering robustness, an effective multi-stage training strategy is proposed. The experimental results show that the proposed method outperforms the state-of-the-art information hiding methods in both visual quality and robustness. In addition, the automatic localization of hidden codes significantly reduces the time of manually correcting geometric distortions for photos, which is a revolutionary innovation for information hiding in mobile applications. Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
CVPR | 3 |
| 2022 | Attribute-Guided Fashion Image Retrieval by Iterative Similarity LearningabstractImage retrieval methods in the fashion field mainly take advantage of query images that reflect user needs, without considering additional keywords that users can provide to specify the attributes in their interests. To achieve the fine-grained fashion retrieval, we propose an iterative similarity learning network (ISLN) for attribute-guided image retrieval, which takes a query image and a specified attribute as input, and outputs other images with the same or similar attribute values. The core of the network is the iterative similarity learning module, which leverages the aggressive learning ability of the deep neural network (DNN) to focus on the area of interest and extract a more accurate feature embedding during the learning process of image and text semantic mapping. Extensive experiments on FashionAI and DARN (+8.33% and +10.73% in mAP) datasets show that ISLN performs better than the state-of-the-art methods in fine-grained similarity retrieval tasks. Cairong Yan, Yanting Zhang 0001, Yongquan Wan, Dandan Zhu 0001 |
ICME | 5 |
| 2022 | Implicit Neural Representation Learning for Hyperspectral Image Super-ResolutionabstractHyperspectral image (HSI) super-resolution without additional auxiliary image remains a constant challenge due to its high-dimensional spectral patterns, where learning an effective spatial and spectral representation is a fundamental issue. Recently, Implicit Neural Representations (INRs) are making strides as a novel and effective representation, especially in the reconstruction task. Therefore, in this work, we propose a novel HSI reconstruction model based on INR which represents HSI by a continuous function mapping a spatial coordinate to its corresponding spectral radiance values. In particular, as a specific implementation of INR, the parameters of parametric model are predicted by a hypernetwork. It makes the continuous functions map the spatial coordinates to pixel values in a content-aware manner. Moreover, periodic spatial encoding are deeply integrated with the reconstruction procedure, which makes our model capable of recovering more high frequency details. Experimental results on CAVE, NUS, and NTIRE2018 datasets demonstrate the superiority of our model. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
ICME | 2 |
| 2022 | Human Attention Based Movie Summarization: Dataset and Baseline ModelabstractThe movie summarization model can automatically edit a condensed and succinct version of the movie by selecting the keyframes. Previous works mainly resort to hand-crafted heuristics and most of them are unsupervised. Supervised movie summarization is a new research field and, there is currently no publicly suitable dataset available. Moreover, existing works only focus on the movies themselves while neglecting the audiences, who have the most say in which part of the movie is more attractive. To deal with the aforementioned limitations, we establish a human attention based movie summarization dataset Movie50. Specifically, we explore the human attention variations when watching videos and have the following findings: (1) The attention of humans is concentrated when watching keyframes. (2) The attention of humans is distracted when watching non-keyframes. Inspired by these findings, we collect the eye fixations of 20 participants when watching 50 movies and propose a novel human attention based annotation pipeline. In addition, we introduce A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better model human attention as well as exploit more plentiful information. Extensive experiments demonstrate the superiority of the proposed method. Defang Zhao, Dandan Zhu 0001, Xiongkuo Min, Jiaomin Yue, Kaiwei Zhang, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 2 |
| 2022 | Learning bi-grained cross-correlation siamese networks for visual tracking
Defang Zhao, Dandan Zhu 0001, Jia Shuai |
Appl. Intell. | 3 |
| 2022 | Software defect prediction based on stacked sparse denoising autoencoders and enhanced extreme learning machineabstractAbstract Software defect prediction is an important software quality assurance technique. Nevertheless, the prediction performance of the constructed model is easily susceptible to irrelevant or redundant features in the software projects and is not predominant enough. To address these two issues, a novel defect prediction model called SSEPG based on Stacked Sparse Denoising AutoEncoders (SSDAE) and Extreme Learning Maching (ELM) optimised by Particle Swarm Optimisation (PSO) and another complementary Gravitational Search Algorithm (GSA) are proposed in this paper, which has two main merits: (1) employ a novel deep neural network – SSDAE to extract new combined features, which can effectively learn the robust deep semantic feature representation. (2) integrate strong exploitation capacity of PSO with strong exploration capability of GSA to optimise the input weights and hidden layer biases of ELM, and utilise the superior discriminability of the enhanced ELM to predict the defective modules. The SSDAE is compared with eleven state‐of‐the‐art feature extraction methods in effect and efficiency, and the SSEPG model is compared with multiple baseline models that contain five classic defect predictors and three variants across 24 software defect projects. The experimental results exhibit the superiority of the SSDAE and the SSEPG on six evaluation metrics. Shi Ying 0002, Kun Zhu 0024, Dandan Zhu 0001 |
IET Softw. | 4 |
| 2022 | IVKMP: A robust data-driven heterogeneous defect model based on deep representation optimization learning
Kun Zhu 0024, Shi Ying 0002, Weiping Ding 0001, Dandan Zhu 0001 |
Inf. Sci. | 5 |
| 2022 | Dynamic clustering based contextual combinatorial multi-armed bandit for online recommendationabstractRecommender systems still face a trade-off between exploring new items to maximize user satisfaction and exploiting those already interacted with to match user interests. This problem is widely recognized as the exploration/exploitation (EE) dilemma, and the multi-armed bandit (MAB) algorithm has proven to be an effective solution. As the scale of users and items in real-world application scenarios increases, their purchase interactions become sparser. Then three issues need to be investigated when building MAB-based recommender systems. First, large-scale users and sparse interactions increase the difficulty of user preference mining. Second, traditional bandits model items as arms and cannot deal with ever-growing items effectively. Third, widely used Bernoulli-based reward mechanisms only feedback 0 or 1, ignoring rich implicit feedback such as behaviors like click and add-to-cart. To address these problems, we propose an algorithm named Dynamic Clustering based Contextual Combinatorial Multi-Armed Bandits (DC3MAB), which consists of three configurable key components. Specifically, a dynamic user clustering strategy enables different users in the same cluster to cooperate in estimating the expected rewards of arms. A dynamic item partitioning approach based on collaborative filtering significantly reduces the scale of arms and produces a recommendation list instead of one item to provide diversity. In addition, a multi-class reward mechanism based on fine-grained implicit feedback helps better capture user preferences. Extensive empirical experiments on three real-world datasets demonstrate the superiority of our proposed DC3MAB over state-of-the-art bandits (On average, +75.8% in F1 and +54.3% in cumulative reward). The source code is available at https://github.com/HaixHan/DC3MAB. Cairong Yan, Haixia Han, Yanting Zhang 0001, Dandan Zhu 0001, Yongquan Wan |
Knowl. Based Syst. | 4 |
| 2022 | Cross-Modal Prostate Cancer Segmentation via Self-Attention DistillationabstractThe automatic and accurate segmentation of the prostate cancer from the multi-modal magnetic resonance images is of prime importance for the disease assessment and follow-up treatment plan. However, how to use the multi-modal image features more efficiently is still a challenging problem in the field of medical image segmentation. In this paper, we develop a cross-modal self-attention distillation network by fully exploiting the encoded information of the intermediate layers from different modalities, and the generated attention maps of different modalities enable the model to transfer significant and discriminative information that contains more details. Moreover, a novel spatial correlated feature fusion module is further employed for learning more complementary correlation and non-linear information of different modality images. We evaluate our model in five-fold cross-validation on 358 MRI images with biopsy confirmed. Without bells and whistles, our proposed network achieves state-of-the-art performance on extensive experiments. Xiaoang Shen, Yudong Zhang 0001, Ye Luo 0004, Jihao Luo, Dandan Zhu 0001, Hanmei Yang, Binghui Zhao |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | A Lightweight Saliency Prediction Model for Omnidirectional ImagesabstractAt present, most high-performing saliency prediction models for omnidirectional images (ODIs) depend on deeper or wider convolutional neural networks (CNNs), benefiting from their superior feature representation capability but suffering from high computational costs. To address this issue, we propose a novel lightweight saliency prediction model to predict the eye fixations on ODIs. Specifically, our proposed model consists of three modules: a lightweight feature representation module, a supervised attention module, and a dynamic convolution aggregation module. Different from the existing saliency prediction models, our proposed model is the first to introduce the dynamic convolution into the saliency prediction and aggregate multiple parallel convolution kernels dynamically based on their attention. Such a dynamic convolution operation is not only computationally efficient (small kernel size), but also increases the feature representation capability since these convolution kernels are aggregated in a non-linear manner via attention. Experimental results on two benchmark datasets show that our model is lightweight and outperforms other state-of-the-art methods. Dandan Zhu 0001, Yongqing Chen, Defang Zhao, Xiongkuo Min, Qiangqiang Zhou, Shaobo Yu, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 1 |
| 2021 | Lavs: A Lightweight Audio-Visual Saliency Prediction ModelabstractAudio information is essential for guiding human attention and visual perception, which has been verified by many comprehensive psychological studies. However, the audio modality has been rather neglected in modeling visual attention, most of the current visual attention models heavily depend on visual information. Additionally, current existing high-performing visual attention models rely on deeper convolution neural networks (CNNs), benefiting from their extraordinary feature learning ability but incurring high computational cost. To this end, we propose a novel lightweight audio-visual saliency (LAVS) model to efficiently address the problem of fixation prediction in videos. To the best of our knowledge, our proposed model constitutes the first attempt to exploit a lightweight network and combines the visual and audio cues to perform saliency estimation in videos. Specifically, our proposed model consists of four modules, which are spatial-temporal visual saliency estimation module, audio features extraction module, source sound localization module, and audio-visual saliency fusion module. Extensive experiments across datasets validate the effectiveness and real-time performance of the proposed LAVS model, which outperforms the other state-of-the-art methods. Dandan Zhu 0001, Defang Zhao, Xiongkuo Min, Tian Han 0001, Qiangqiang Zhou, Shaobo Yu, Yongqing Chen, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 1 |
| 2021 | Inter-Observer Visual Congruency in Video-ViewingabstractThere are individual differences in human visual attention between observers when viewing the same scene. Inter-observer visual congruency (IOVC) describes the dispersion between different people's visual attention areas when they observe the same stimulus. Research on the IOVC of video is interesting but lacking. In this paper, we first introduce the measurement to calculate the IOVC of video. And an eye-tracking experiment is conducted in a realistic movie-watching environment to establish a movie scene dataset. Then we propose a method to predict the IOVC of video, which employs a dual-channel network to extract and integrate content and optical flow features. The effectiveness of the proposed prediction model is validated on our dataset. And the correlation between inter-observer congruency and video emotion is analyzed. Jiaomin Yue, Dandan Zhu 0001, Xiongkuo Min, Xiao-Ping Zhang 0002, Guangtao Zhai |
VCIP | 3 |
| 2021 | RANSP: Ranking attention network for saliency prediction on omnidirectional images
Dandan Zhu 0001, Yongqing Chen, Xiongkuo Min, Yucheng Zhu, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
Neurocomputing | 1 |
| 2021 | WGNCS: A robust hybrid cross-version defect model via multi-objective optimization and deep enhanced feature representation
Shi Ying 0002, Weiping Ding 0001, Kun Zhu 0024, Dandan Zhu 0001 |
Inf. Sci. | 5 |
| 2021 | Software defect prediction based on enhanced metaheuristic feature selection optimization and a hybrid deep neural network
Kun Zhu 0024, Shi Ying 0002, Dandan Zhu 0001 |
J. Syst. Softw. | 4 |
| 2021 | Towards multi-scale deep features learning with correlation metric for person re-identification
Dandan Zhu 0001, Qiangqiang Zhou, Tian Han 0001, Yongqing Chen, Defang Zhao, Xiaokang Yang 0001 |
Knowl. Based Syst. | 1 |
| 2020 | Ransp: Ranking Attention Network For Saliency Prediction On Omnidirectional ImagesabstractVarious convolutional neural network (CNN)-based methods have shown the ability to boost the performance of saliency prediction on omnidirectional images (ODIs). However, these methods are limited by sub-optimal accuracy, because not all the features extracted by the CNN model are not useful for the final fine-grained saliency prediction. Features are redundant and have negative impact on the final fine-grained saliency prediction. To tackle this problem, we propose a novel Ranking Attention Network for saliency prediction (RANSP) of head fixations on ODIs. Specifically, the part-guided attention (PA) module and channel-wise feature (CF) extraction module are integrated in a unified framework and are trained in an end-to-end manner for fine-grained saliency prediction. To better utilize the channel-wise feature map, we further propose a new Ranking Attention Module (RAM), which automatically ranks and selects these maps based on scores for fine-grained saliency prediction. Extensive experiments are conducted to show the effectiveness of our method for saliency prediction of ODIs. Dandan Zhu 0001, Yongqing Chen, Tian Han 0001, Defang Zhao, Yucheng Zhu, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 1 |
| 2020 | Saliency Prediction on Omnidirectional Images with Brain-Like Shallow Neural NetworkabstractDeep feedforward convolutional neural networks (CNNs) perform well in the saliency prediction of omnidirectional images (ODIs), and have become the leading class of candidate models of the visual processing mechanism in the primate ventral stream. These CNNs have evolved from shallow network architecture to extremely deep and branching architecture to achieve superb performance in various vision tasks, yet it is unclear how brain-like they are. In particular, these deep feedforward CNNs are difficult to mapping to ventral stream structure of the brain visual system due to their vast number of layers and missing biologically-important connections, such as recurrence. To tackle this issue, some brain-like shallow neural networks are introduced. In this paper, we propose a novel brain-like network model for saliency prediction of head fixations on ODIs. Specifically, our proposed model consists of three modules: a CORnet-S module, a template feature extraction module and a ranking attention module (RAM). The CORnet-S module is a lightweight artificial neural network (ANN) with four anatomically mapped areas (V1, V2, V4 and IT) and it can simulate the visual processing mechanism of ventral visual stream in the human brain. The template features extraction module is introduced to extract attention maps of ODIs and provide guidance for the feature ranking in the following RAM module. The RAM module is used to rank and select features that are important for fine-grained saliency prediction. Extensive experiments have validated the effectiveness of the proposed model in predicting saliency maps of ODIs, and the proposed model outperforms other state-of-the-art methods with similar scale. Dandan Zhu 0001, Yongqing Chen, Xiongkuo Min, Defang Zhao, Yucheng Zhu, Qiangqiang Zhou, Xiaokang Yang 0001, Tian Han 0001 |
ICPR | 1 |
| 2020 | Within-project and cross-project just-in-time defect prediction based on denoising autoencoder and convolutional neural networkabstractJust‐in‐time defect prediction is an important and useful branch in software defect prediction. At present, deep learning is a research hotspot in the field of artificial intelligence, which can combine basic defect features into deep semantic features and make up for the shortcomings of machine learning algorithms. However, the mainstream deep learning techniques have not been applied yet in just‐in‐time defect prediction. Therefore, the authors propose a novel just‐in‐time defect prediction model named DAECNN‐JDP based on denoising autoencoder and convolutional neural network in this study, which has three main advantages: (i) Different weights for the position vector of each dimension feature are set, which can be automatically trained by adaptive trainable vector. (ii) Through the training of denoising autoencoder, the input features that are not contaminated by noise can be obtained, thus learning more robust feature representation. (iii) The authors leverage a powerful representation‐learning technique, convolution neural network, to construct the basic change features into the abstract deep semantic features. To evaluate the performance of the DAECNN‐JDP model, they conduct extensive within‐project and cross‐project defect prediction experiments on six large open source projects. The experimental results demonstrate that the superiority of DAECNN‐JDP on five evaluation metrics. Kun Zhu 0024, Shi Ying 0002, Dandan Zhu 0001 |
IET Softw. | 4 |
| 2018 | Spatial Pyramid Dilated Network for Pulmonary Nodule Malignancy ClassificationabstractLung cancer has been the most prevalent cancer in the world and an effective way to diagnose the cancer at the early stage is to detect the pulmonary nodule by computer-aided system. However, the size of the pulmonary nodules varies and the one with small diameter is generally one of the most difficult cases to diagnose. Under this condition, traditional convolution network based nodule classification methods fail to achieve satisfied result due to the miss of tiny but vital features by the pooling operation. To tackle this problem, we propose a novel 3D spatial pyramid dilated convolution network to classify the malignancy of the pulmonary nodules. Instead of using the pooling layers, we utilize the 3D dilated convolution to capture and preserve more detailed characteristic information of the nodules. Moreover, a multiple receptive field fusion strategy is applied to extract the multi-scale features from the nodule CT images. Extensive experimental results show that our model achieves a better result with an accuracy of 88.6% which outperforms other state-of-the-art methods. Ye Luo 0004, Dandan Zhu 0001, Yixuan Xu 0002, Yunxin Sun 0001 |
ICPR | 3 |
| 2018 | Salient object detection via a local and global method based on deep residual network
Dandan Zhu 0001, Ye Luo 0004, Xuan Shao, Qiangqiang Zhou, Laurent Itti |
J. Vis. Commun. Image Represent. | 1 |
| 2017 | Saliency prediction based on new deep multi-layer convolution neural networkabstractRecent advances in saliency detection have utilized deep learning to obtain high-level features to detect salient regions. These advances have demonstrated superior results over previous works that utilize hand-crafted low-level features for saliency detection. In this paper, we propose a new multilayer Convolutional Neural Network (CNN) model to learn high-level features for saliency detection. Compared to other methods, our method presents two merits. First, when performing features extraction, apart from the convolution and pooling step in our method, we add Restricted Boltzmann Machine (RBM) into the CNN framework to obtain more accurate features in intermediate step. Second, in order to deal with case of non-linear classification, we add the Deep Belief Network (DBN) classifier at the end of this model to classify the salient and non-salient regions. Quantitative and qualitative experiments on three benchmark datasets demonstrate that our method performs favorably against the state-of-the-art methods. Dandan Zhu 0001, Ye Luo 0004, Xuan Shao, Laurent Itti |
ICIP | 1 |
| 2017 | Scanpath Prediction Based on High-Level Features and Memory Bias
Xuan Shao, Ye Luo 0004, Dandan Zhu 0001, Laurent Itti |
ICONIP (3) | 3 |
| 2017 | Deep Salient Object Detection via Hierarchical Network Learning
Dandan Zhu 0001, Ye Luo 0004, Xuan Shao, Laurent Itti |
ICONIP (3) | 1 |