VLDB 2026 Research / reviewers in the wild / expert
Xixi Nie
dblp:295/0047
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | M2CR: A primary liver cancer diagnosis system with multimodal multitask collaborative reasoning
Shixin Huang, Jiawei Luo 0002, Xiaoyu Wan, Xixi Nie |
Expert Syst. Appl. | 4 |
| 2026 | Multimodal Emotion-Aligned Cognitive Networks for Image Aesthetic AssessmentabstractImage aesthetic assessment (IAA) is a challenging task due to the subjectivity and abstraction of aesthetic perception. Psychological studies reveal that aesthetic experiences often trigger emotional responses, while comment texts directly reflect people’s expressions of aesthetics and emotions. However, existing multimodal IAA methods neglect the alignment between modalities. To address this, we propose a multimodal emotion-alignment cognitive network (MEC-Net) for IAA, employing strategies of emotion alignment, subjective–objective interaction, and multimodal fusion. First, an emotion alignment module is introduced to align image and text modalities using emotional stimuli, enhancing the consistency of heterogeneous modal features. Then, a subjective and objective representation module is proposed to extract multi-source information from text and images separately. Next, a subjective-objective interactive LSTM (SO-LSTM) is designed to capture the deep interaction between images and text in aesthetic understanding. Finally, an dynamic multimodal fusion (DMF) based on low-rank decomposition is proposed to integrate subjective, objective, and subjective-objective interactive modal features for aesthetic distribution prediction. Extensive experiments and qualitative analysis on image aesthetic benchmarks indicate that the proposed MEC-Net outperforms the state-of-the-art on three IAA tasks. Further, we increase emotion classification task-driven evaluation metrics to verify the strong generalizability of the proposed MEC-Net. Xixi Nie, Shixin Huang, Jiawei Luo 0002, Xiaodan Zhang 0005, Leida Li, Hongchun Qu, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | MLR-NET: An Arbitrary Skew Angle Detection Algorithm for Complex Layout Document Images
Peisen Wang, Xixi Nie, Chunyi Guo, Kaijiang Li |
PRCV (7) | 3 |
| 2024 | MRAM: Multi-scale Regional Attribute-weighting via Meta-learning for Personalized Image Aesthetics Assessment
Xixi Nie, Shixin Huang, Xinbo Gao 0001, Jiawei Luo 0002 |
Knowl. Based Syst. | 1 |
| 2024 | A Self-Supervised Network-Based Smoke Removal and Depth Estimation for Monocular Endoscopic VideosabstractIn minimally invasive surgery videos, label-free monocular laparoscopic depth estimation is challenging due to smoke. For this reason, we propose a self-supervised collaborative network-based depth estimation method with smoke-removal for monocular endoscopic video, which is decomposed into two steps of smoke-removal and depth estimation. In the first step, we develop a de-endoscopic smoke for cyclic GAN (DS-cGAN) to mitigate the smoke components at different concentrations. The designed generator network comprises sharpened guide encoding module (SGEM), residual dense bottleneck module (RDBM) and refined upsampling convolution module (RUCM), which restores more detailed organ edges and tissue structures. In the second step, high resolution residual U-Net (HRR-UNet) consisting of a DepthNet and two PoseNets is designed to improve the depth estimation accuracy, and adjacent frames are used for camera self-motion estimation. In particular, the proposed method requires neither manual labeling nor patient computed tomography scans during the training and inference phases. Experimental studies on the laparoscopic data set of the Hamlyn Centre show that our method can effectively achieve accurate depth information after net smoking in real surgical scenes while preserving the blood vessels, contours and textures of the surgical site. The experimental results demonstrate that the proposed method outperforms existing state-of-the-art methods in effectiveness and achieves a frame rate of 94.45fps in real time, making it a promising clinical application. Xinbo Gao 0001, Hongying Meng, Xixi Nie |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | BMI-Net: A Brain-inspired Multimodal Interaction Network for Image Aesthetic AssessmentabstractImage aesthetic assessment (IAA) has drawn wide attention in recent years as more and more users post images and texts on the Internet to share their views. The intense subjectivity and complexity of IAA make it extremely challenging. Text triggers the subjective expression of human aesthetic experience based on human implicit memory, so incorporating the textual information and identifying the relationship with the image is of great importance for IAA. However, IAA with the image as input fails to fully consider subjectivity, while existing multimodal IAA ignores the interrelationship among modalities. To this end, we propose a brain-inspired multimodal interaction network (BMI-Net) that simulates how the association area of the cerebral cortex processes sensory stimuli. In particular, the knowledge integration LSTM (KI-LSTM) is proposed to learn the image-text interaction relation. The proposed scalable multimodal fusion (SMF) based on low-rank decomposition fuses image, text and interaction modalities to predict the aesthetic distribution. Extensive experiments show that the proposed BMI-Net outperforms existing state-of-the-art methods on three IAA tasks. Xixi Nie, Bo Hu 0008, Xinbo Gao 0001, Leida Li, Xiaodan Zhang 0005, Bin Xiao 0002 |
ACM Multimedia | 1 |
| 2023 | Reduced-reference image deblurring quality assessment based on multi-scale feature enhancement and aggregation
Bo Hu 0008, Shuaijian Wang, Xinbo Gao 0001, Leida Li, Ji Gan, Xixi Nie |
Neurocomputing | 6 |
| 2023 | User-Guided Personalized Image Aesthetic Assessment Based on Deep Reinforcement LearningabstractPersonalized image aesthetic assessment (PIAA) has recently become a hot topic due to its wide applications, such as photography, film, television, e-commerce, fashion design, and so on. This task is more seriously affected by subjective factors and samples provided by users. In order to acquire precise personalized aesthetic distribution by small amount of samples, we propose a novel user-guided personalized image aesthetic assessment framework. This framework leverages user interactions to retouch and rank images for aesthetic assessment based on deep reinforcement learning (DRL), and generates personalized aesthetic distribution that is more in line with the aesthetic preferences of different users. It mainly consists of two stages. In the first stage, personalized aesthetic ranking is generated by interactive image enhancement and manual ranking, meanwhile, two policy networks will be trained. These two networks will be trained iteratively and alternatively to facilitate the final personalized aesthetic assessment. In the second stage, these modified images are labeled with aesthetic attributes by one style-specific classifier, and then the personalized aesthetic distribution is generated based on the multiple aesthetic attributes of these images, which conforms to the aesthetic preference of users better. Compared with other existing methods, our approach has achieved new state-of-the-art in the task of personalized image aesthetic assessment on the public AVA and FLICKR-AES datasets. Pei Lv, Jianqi Fan, Xixi Nie, Weiming Dong, Xiaoheng Jiang, Bing Zhou 0003, Mingliang Xu 0001, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2023 | MLNet: A Multi-Domain Lightweight Network for Multi-Focus Image FusionabstractExisting multi-focus image fusion (MFIF) methods are difficult to achieve satisfactory results in both fusion performance and rate simultaneously. The spatial domain methods are hard to determine the focus/defocus boundary (FDB), and the transform domain methods are likely to damage the content information of the source images. Moreover, the deep learning-based MFIF methods are usually confronted with low rate due to complex models and enormous learnable parameters. To address these issues, we propose a multi-domain lightweight network (MLNet) for MFIF, which can achieve competitive results in both performance and rate. The proposed MLNet mainly includes three modules, namely focus extraction (FE), focus measure (FM) and image fusion (IF). In the interpretable FE module, the image features extracted by discrete cosine transform-based convolution (DCTConv) and local binary pattern-based convolution (LBPConv) are concatenated and fed into the FM module. DCTConv based on transform domain takes DCT coefficients to construct a fixed convolution kernel without parameter learning, which can effectively capture the high/low frequency content of the image. LBPConv based on spatial domain can achieve structure features and gradient information from source images. In the FM module, a 3-layer 1 × 1 convolution with a few learnable parameters is employed to generate the initial decision map, which has the properties of flexible input. The fused image is obtained by the IF module according to the final decision map. In terms of quantitative and qualitative evaluations, extensive experiments validate that the proposed method outperforms existing state-of-the-art methods on three public datasets. In addition, the proposed MLNet contains only 0.01 M parameters, which is 0.2% of the first CNN-based MFIF method [25]. Xixi Nie, Bo Hu 0008, Xinbo Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | A focus measure in discrete cosine transform domain for multi-focus image fast fusion
Xixi Nie, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
Neurocomputing | 1 |