VLDB 2026 Research / reviewers in the wild / expert
Xiaodan Zhang 0005
dblp:29/2631-5
· DBLP profile ↗
10ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0001-9808-8159ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal Emotion-Aligned Cognitive Networks for Image Aesthetic AssessmentabstractImage aesthetic assessment (IAA) is a challenging task due to the subjectivity and abstraction of aesthetic perception. Psychological studies reveal that aesthetic experiences often trigger emotional responses, while comment texts directly reflect people’s expressions of aesthetics and emotions. However, existing multimodal IAA methods neglect the alignment between modalities. To address this, we propose a multimodal emotion-alignment cognitive network (MEC-Net) for IAA, employing strategies of emotion alignment, subjective–objective interaction, and multimodal fusion. First, an emotion alignment module is introduced to align image and text modalities using emotional stimuli, enhancing the consistency of heterogeneous modal features. Then, a subjective and objective representation module is proposed to extract multi-source information from text and images separately. Next, a subjective-objective interactive LSTM (SO-LSTM) is designed to capture the deep interaction between images and text in aesthetic understanding. Finally, an dynamic multimodal fusion (DMF) based on low-rank decomposition is proposed to integrate subjective, objective, and subjective-objective interactive modal features for aesthetic distribution prediction. Extensive experiments and qualitative analysis on image aesthetic benchmarks indicate that the proposed MEC-Net outperforms the state-of-the-art on three IAA tasks. Further, we increase emotion classification task-driven evaluation metrics to verify the strong generalizability of the proposed MEC-Net. Xixi Nie, Shixin Huang, Jiawei Luo 0002, Xiaodan Zhang 0005, Leida Li, Hongchun Qu, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | ASMCC-Diff: Arbitrary Size Multi-Condition Controllable Chinese Landscape Painting Generation with Diffusion ModelsabstractThanks to the emergence of generative models, Chinese Landscape Painting Generation (CLPG) has garnered increasing attention. However, existing works are primarily limited to relying on text control conditions, lacking more fine-grained control over spatial layout and style. Additionally, they are limited to a fixed size and aspect ratios. But different landscape scenes require different sizes to appear more balanced and natural. Thus, a question arises: Is it possible to design a model that can be controlled by multiple conditions (text, style and sketches) while generating images with various sizes? In this paper, we explore this issue and propose ASMCC-Diff. Specifically, it consists of two modules, i.e. multi-condition controlled image generation module and arbitrary size up-scaling module. The critical insights of multi-condition controlled image generation module are to embed multiple conditions with distinct priorities. The sketch condition serves as the primary flow, guiding the overall structure, while the text and style conditions act as auxiliary components, injected into the diffusion model via a cross-attention module. Additionally, to avoid conflicts between the semantics of style images and text, we use Q-former to separate the semantic and stylistic information of the reference image. For the arbitrary size upscaling module, we first truncate the generation process, up-sample the image to the specified size, and then continue the generation. Furthermore, we introduce a new Chinese landscape painting database that supports multiple conditions, facilitating further research. Experimental results demonstrate the superior performance of our proposed model. The code and dataset will be released. Xiaodan Zhang 0005, Xiteng Zhang, Jianqiang Yan, Qiyao Hu, Jinye Peng 0001, Qiannan Duan |
ECAI | 1 |
| 2025 | Mixture-of-Modality-Experts for Unified Image Aesthetic Assessment with Multi-Level AdaptationabstractMulti-modal image aesthetic assessment (MIAA) has gained significant progress, by predicting aesthetic based on both an image and its text comments. However, most MIAA methods are not applicable, when there are no text comments available. To combat this challenge, we propose a unified image aesthetic assessment (IAA) framework, termed AesFormer, by using mixtures of vision-language Transformers. Specially, AesFormer first learns aligned image-text representations through contrastive learning, and uses a vision-language head for MIAA prediction. Afterward, we propose a multi-level adaptation (MLA) method to adapt the learned MIAA model to the case without text comments, and use another vision head for vison-only IAA (VIAA) prediction. Extensive experimental results show that AesFormer significantly outperforms previous methods in both MIAA and VIAA tasks, on diverse benchmarking datasets. Our code has been released at: https://github.com/AiArt-Gao/AesFormer Fei Gao 0006, Xiaodan Zhang 0005, Lihuo He, Nannan Wang 0001 |
ICME | 4 |
| 2024 | Confidence-based dynamic cross-modal memory network for image aesthetic assessment
Xiaodan Zhang 0005, Jinye Peng 0001, Xinbo Gao 0001, Bo Hu 0008 |
Pattern Recognit. | 1 |
| 2023 | BMI-Net: A Brain-inspired Multimodal Interaction Network for Image Aesthetic AssessmentabstractImage aesthetic assessment (IAA) has drawn wide attention in recent years as more and more users post images and texts on the Internet to share their views. The intense subjectivity and complexity of IAA make it extremely challenging. Text triggers the subjective expression of human aesthetic experience based on human implicit memory, so incorporating the textual information and identifying the relationship with the image is of great importance for IAA. However, IAA with the image as input fails to fully consider subjectivity, while existing multimodal IAA ignores the interrelationship among modalities. To this end, we propose a brain-inspired multimodal interaction network (BMI-Net) that simulates how the association area of the cerebral cortex processes sensory stimuli. In particular, the knowledge integration LSTM (KI-LSTM) is proposed to learn the image-text interaction relation. The proposed scalable multimodal fusion (SMF) based on low-rank decomposition fuses image, text and interaction modalities to predict the aesthetic distribution. Extensive experiments show that the proposed BMI-Net outperforms existing state-of-the-art methods on three IAA tasks. Xixi Nie, Bo Hu 0008, Xinbo Gao 0001, Leida Li, Xiaodan Zhang 0005, Bin Xiao 0002 |
ACM Multimedia | 5 |
| 2021 | MSCAN: Multimodal Self-and-Collaborative Attention Network for image aesthetic prediction tasks
Xiaodan Zhang 0005, Xinbo Gao 0001, Lihuo He, Wen Lu 0004 |
Neurocomputing | 1 |
| 2021 | Beyond Vision: A Multimodal Recurrent Attention Convolutional Neural Network for Unified Image Aesthetic Prediction TasksabstractOver the past few years, image aesthetic prediction has attracted increasing attention because of its wide applications, such as image retrieval, photo album management and aesthetic-driven image enhancement. However, previous studies in this area only achieve limited success because 1) they primarily depend on visual features and ignore textual information. 2) they tend to focus equally on to each part of images and ignore the selective attention mechanism. This paper overcomes these limitations by proposing a novel multimodal recurrent attention convolutional neural network (MRACNN). More specifically, the MRACNN consists of two streams: the vision stream and the language stream. The former employs the recurrent attention network to tune out irrelevant information and focuses on some key regions to extract visual features. The latter utilizes the Text-CNN to capture the high-level semantics of user comments. Finally, a multimodal factorized bilinear (MFB) pooling approach is used to achieve effective fusion of textual and visual features. Extensive experiments demonstrate that the proposed MRACNN significantly outperforms state-of-the-art methods for unified aesthetic prediction tasks: (i) aesthetic quality classification; (ii) aesthetic score regression; and (iii) aesthetic score distribution prediction. Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He, Jie Li 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Fusion global and local deep representations with neural attention for aesthetic quality assessmentabstractIn recent years, deep-learning based aesthetics assessment methods have shown promising results. However, existing methods can only achieve limited success because 1) most of the methods take one fixed-size patch as the training example, which loses the fine grained details and the holistic layout information, and 2) most of the methods ignore ordinal issues in image aesthetic assessment, i.e. image scored 5.3 is more likely to be in the high quality class than image scored 4.5. To address these challenges, we presents a novel convolutional networks with two branches to encode global and local features . The first branch not only captures the spatial layout information but also feedbacks the top-down neural attention. The second branch selects the important attended region to extract the fine details features. A sobel-based attention layer is integrated with the second branch to enhance fine details encoding. Regarding the second problem, we combine the strength of classification approach and regression approach by a multi-task learning framework. Extensive experiments on challenging Aesthetic and Visual Analysis (AVA) dataset and Photo.net dataset indicate the effectiveness of the proposed method. Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He |
Signal Process. Image Commun. | 1 |
| 2019 | A Gated Peripheral-Foveal Convolutional Neural Network for Unified Image Aesthetic PredictionabstractLearning fine-grained details is a key issue in image aesthetic assessment. Most of the previous methods extract the fine-grained details via random cropping strategy, which may undermine the integrity of semantic information. Extensive studies show that humans perceive fine-grained details with a mixture of foveal vision and peripheral vision. Fovea has the highest possible visual acuity and is responsible for seeing the details. The peripheral vision is used for perceiving the broad spatial scene and selecting the attended regions for the fovea. Inspired by these observations, we propose a gated peripheral-foveal convolutional neural network. It is a dedicated double-subnet neural network (i.e., a peripheral subnet and a foveal subnet). The former aims to mimic the functions of peripheral vision to encode the holistic information and provide the attended regions. The latter aims to extract fine-grained features on these key regions. Considering that the peripheral vision and foveal vision play different roles in processing different visual stimuli, we further employ a gated information fusion network to weigh their contributions. The weights are determined through the fully connected layers followed by a sigmoid function. We conduct comprehensive experiments on the standard Aesthetic Visual Analysis (AVA) dataset and Photo.net dataset for unified aesthetic prediction tasks: 1) aesthetic quality classification; 2) aesthetic score regression; and 3) aesthetic score distribution prediction. The experimental results demonstrate the effectiveness of the proposed method. Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He |
IEEE Trans. Multim. | 1 |
| 2018 | Dominant vanishing point detection in the wild with application in composition analysis
Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He, Qi Liu 0054 |
Neurocomputing | 1 |