EDBT 2026 Demo / reviewers in the wild / expert
Xiangxian Li
dblp:305/3204
· DBLP profile ↗
17ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0001-6638-2361ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating VR Motion Sickness Through Multi-sensory Simulation of Wind Sensation (MSSWS): A Vestibular-Visual Synchronization ApproachabstractUsers are more likely to experience visually induced motion sickness (VIMS) during passive motion in virtual reality (VR), particularly in passive virtual driving scenarios. To solve this challenge, we propose a Multi-sensory Simulation of Wind Sensation (MSSWS) method to alleviate VIMS symptoms by enhancing the user’s sense of embodiment (SoE). The method utilizes low-fidelity airflow simulation to align vestibular perception with visual, auditory, and tactile cues. Then, we developed an interactive wind simulation helmet that can generate low-fidelity airflow along four directional axes around the user’s head, and implemented a passive virtual motorcycle riding environment with synchronized visual and auditory feedback matching the airflow patterns. A user experiment was conducted to systematically evaluate the effectiveness of MSSWS in enhancing SoE and reducing VIMS under various conditions, including helmet activation states, speed variations, and movement directions, with additional validation conducted in a Cave Automatic Virtual Environment (CAVE). Experimental results show that MSSWS significantly enhances sense of presence and SoE while substantially reducing VIMS during passive navigation. Notably, MSSWS shows significantly greater effects on VIMS in high-risk conditions characterized by rapid speed changes and multi-directional movement. Yuan Yue, Chao Zhou 0012, Tangjun Qu, Yan Hu 0003, Juan Liu 0008, Tianren Luo, Xiangxian Li, Yulong Bian |
VR | 8 |
| 2025 | Causal Inference over Visual-Semantic-Aligned Graph for Image ClassificationabstractIncorporating tagging information to regularize the representation learning of images usually leads to improved performance in image classification by aligning the visual features with the textual ones of higher discriminative power. Existing methods typically follow the predictive approach, which uses tags as the semantic labels for visual input to make predictions. However, they typically face the problem of handling the heterogeneity between modalities. In order to learn accurate visual-semantic mapping, this paper presents a visual-semantic causal association modeling framework termed VSCNet. It aligns visual regions with tags, uses a pre-learned hierarchy of visual and semantic exemplars to refine tag predictions and constructs an augmented heterogeneous graph to perform causal intervention. Specifically, the fine-grained visual-semantic alignment (FVA) module adaptively locates the semantic-intensive regions corresponding to tags. The heterogeneous association refinement (HAR) module associates the visual regions, semantic elements and pre-learned visual prototypes in a heterogeneous graph to filter the error predictions and enrich the information. The causal inference with graphical masking (CIM) module applies self-learned masks to discover the causal nodes and edges in the heterogeneous graph to address the spurious association, forming robust causal representations. Experimental results from two benchmarking datasets show that VSCNet effectively builds the visual-semantic associations from images and leads to better performance than the state-of-the-art methods with enriched predictive information. Lei Meng 0001, Xiangxian Li, Xiaoshuo Yan, Haokai Ma, Zhuang Qi, Xiangxu Meng |
AAAI | 2 |
| 2025 | TongueBCI: An Interaction Method Based on EEG Signals from Tongue Movement Direction
Dingming Tan, Zifeng Ni, Baiqiao Zhang, Chao Zhou 0012, Tianshuo Bai, Juan Liu 0008, Xiangxian Li, Yulong Bian |
ICXR | 7 |
| 2025 | MAGNeT: Multimodal Adaptive Gaussian Networks for Intent Inference in Moving Target Selection across Complex ScenariosabstractMoving target selection in multimedia interactive systems faces unprecedented challenges as users increasingly interact across diverse, dynamic contexts-from live streaming in moving vehicles to VR gaming in varying environments. Existing approaches rely on probabilistic models that relate endpoint distribution to target properties (size, speed). However, these methods require substantial training data for each new context and lack transferability across scenarios, limiting their practical deployment in diverse multimedia environments where rich multimodal contextual information is readily available. This paper introduces MAGNeT (Multimodal Adaptive Gaussian Networks), which addresses these problems by combining classical statistical modeling with context-aware multimodal method. MAGNeT dynamically fuses pre-fitted Ternary-Gaussian models from various scenarios based on real-time contextual cues, enabling effective adaptation with minimal training data while preserving model interpretability. We take experiments on self-constructed 2D and 3D moving target selection datasets under in-vehicle vibration conditions. Extensive experiments demonstrate that MAGNeT achieves lower error rates with few-shot samples, by applying context-aware fusion of Gaussian experts from multi-factor conditions. Xiangxian Li, Baiqiao Zhang, Yijia Ma, Xianhui Cao, Juan Liu 0008, Yulong Bian, Jin Huang 0009, Chenglei Yang |
ACM Multimedia | 1 |
| 2025 | Assessing Dynamic Flow Experience from EEG Signals: A Processing-based Approach
Shilong Liu 0002, Chaorui Tong, Zelu Liu, Xiangxian Li, Chao Zhou 0012, Juan Liu 0008, Yulong Bian |
UIST | 4 |
| 2025 | Enhancing Recognition of Stereotyped Movements in ASD Children Through Action Pattern Mining and Multi-Channel FusionabstractStereotyped movements play a crucial role in diagnosing Autism Spectrum Disorder (ASD). However, recognizing them poses challenges, due to limited data availability and the movements' specificity and varying duration. To support in-depth analysis of ASD children's movements, we constructed the ACSA653 dataset, comprising 653 videos across six classes of stereotyped movements. This dataset surpasses existing ones in both scale and category. To improve the recognition of stereotyped movements, we propose APMFNet, a model that integrates three modules: Visual Motion Learning (VML), Skeleton Relation Mining (SRM), and Multi-channel Fusion (MF). The VML module focuses on extracting spatial and motion information from RGB and optical-flow sequences. The SRM module effectively mines essential motion patterns associated with stereotyped movements through cross-modal graph. The MF module fuses multi-modal information through cross-modality attention to facilitate decision-making. Tested on ACSA653, APMFNet outperforms current state-of-the-art methods, suggesting its potential to identify stable patterns of stereotyped movements in children with ASD. Baiqiao Zhang, Yanran Yuan, Xiangxian Li, Weiying Liu, Wenxin Yao, Yulong Bian, Juan Liu 0008 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Unifying Visual and Semantic Feature Spaces with Diffusion Models for Enhanced Cross-Modal Alignment
Yuze Zheng, Xiangxian Li, Xiangxu Meng, Lei Meng 0001 |
ICANN (6) | 3 |
| 2024 | Cross-modal learning using privileged information for long-tailed image classificationabstractThe prevalence of long-tailed distributions in real-world data often results in classification models favoring the dominant classes, neglecting the less frequent ones. Current approaches address the issues in long-tailed image classification by rebalancing data, optimizing weights, and augmenting information. However, these methods often struggle to balance the performance between dominant and minority classes because of inadequate representation learning of the latter. To address these problems, we introduce descriptional words into images as cross-modal privileged information and propose a cross-modal enhanced method for long-tailed image classification, referred to as CMLTNet. CMLTNet improves the learning of intraclass similarity of tail-class representations by cross-modal alignment and captures the difference between the head and tail classes in semantic space by cross-modal inference. After fusing the above information, CMLTNet achieved an overall performance that was better than those of benchmark long-tailed and cross-modal learning methods on the long-tailed cross-modal datasets, NUS-WIDE and VireoFood-172. The effectiveness of the proposed modules was further studied through ablation experiments. In a case study of feature distribution, the proposed model was better in learning representations of tail classes, and in the experiments on model attention, CMLTNet has the potential to help learn some rare concepts in the tail class through mapping to the semantic space. Xiangxian Li, Yuze Zheng, Haokai Ma, Zhuang Qi, Xiangxu Meng, Lei Meng 0001 |
Comput. Vis. Media | 1 |
| 2023 | Unsupervised Segmentation of Haze Regions as Hard Attention for Haze Classification
Haokai Ma, Xiangxian Li, Zhuang Qi, Xiangxu Meng, Lei Meng 0001 |
ICIG (4) | 3 |
| 2023 | A Multi-View Co-Learning Method for Multimodal Sentiment AnalysisabstractExisting works on multimodal sentiment analysis have focused on learning more discriminative unimodal sentiment information or improving multimodal fusion methods to enhance modal complementarity. However, practical results of these methods have been limited owing to the problems of insufficient intra-modal representation and inter-modal noise. To alleviate this problem, we propose a multi-view co-learning method (MVATF) for video sentiment analysis. First, we propose a multi-view features extraction module to capture more perspectives from a single modality. Second, we propose a two-level fusion sentiment enhancement strategy that uses hierarchical attentive learning fusion and a multi-task learning fusion module to achieve co-learning to effectively filter inter-modal noise for better multimodal sentiment fusion features. Experimental results on the CH-SIMS, CMU-MOSI and MOSEI datasets show that the proposed method outperforms the state-of-the-art methods. Wenxiu Geng, Yulong Bian, Xiangxian Li |
ICME | 3 |
| 2023 | Cross-Modal Content Inference and Feature Enrichment for Cold-Start RecommendationabstractMultimedia recommendation aims to fuse the multi-modal information of items for feature enrichment to improve the recommendation performance. However, existing methods typically introduce multi-modal information based on collaborative information to improve the overall recommendation precision, while failing to explore its cold-start recommendation performance. Meanwhile, these above methods are only applicable when such multi-modal data is available. To address this problem, this paper proposes a recommendation framework, named Cross-modal Content Inference and Feature Enrichment Recommendation (CIERec), which exploits the multi-modal information to improve its cold-start recommendation performance. Specifically, CIERec first introduces image annotation as the privileged information to help guide the mapping of unified features from the visual space to the semantic space in the training phase. And then CIERec enriches the content representation with the fusion of collaborative, visual, and cross-modal inferred representations, so as to improve its cold-start recommendation performance. Experimental results on two real-world datasets show that the content representations learned by CIERec are able to achieve superior cold-start recommendation performance over existing visually-aware recommendation algorithms. More importantly, CIERec can consistently achieve significant improvements with different conventional visually-aware backbones, which verifies its universality and effectiveness. Haokai Ma, Zhuang Qi, Xinxin Dong, Xiangxian Li, Yuze Zheng, Xiangxu Meng, Lei Meng 0001 |
IJCNN | 4 |
| 2023 | Multi-channel Attentive Weighting of Visual Frames for Multimodal Video ClassificationabstractMultimodal video classification aims to incorporate semantic information to regularize the visual representation learning of videos. Conventional methods typically focus on analyzing all information extracted from different modals rather than key information. However, they usually face the problem of handling the redundant video frames of little categorical information. To address this problem, this paper proposes a novel approach that employs multi-channel weighting of visual frames to mitigate the interference of redundant information. Specifically, the proposed algorithm, termed MCA-WF, includes two main modules, where the multi-channel attentive weighting of video frames (McAW) module performs the multi-granularity and multi-channel frame weighting mechanism based on visual self-attention, contrastive attention and cross-modal attention constraints to filter visual noise and redundant information. The visual frame selection (VFS) module explores the combination of multi-channel attention mechanisms to select the key visual information in the video. Experiments were conducted on MSR-VTT and ActivityNet Captions datasets in terms of performance comparison, ablation study, in-depth analysis, and case studies. The results verified that MCA-WF can notice the key information in the classification and effectively improve the ability of information complementation and integration between modals, which leads to better performance than the state-of-the-art methods. Zhuang Qi, Xiangxian Li, Xiangxu Meng, Lei Meng 0001 |
IJCNN | 3 |
| 2023 | A Dual-branch Enhanced Multi-task Learning Network for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis is a complex research problem. Firstly, current multimodal approaches fail to adequately consider the intricate multi-level correspondence between modalities and the unique contextual information within each modality; secondly, cross-modal fusion methods for inter-modal fusion somewhat weaken the mode-specific internal features, which is a limitation of the traditional single-branch model. To this end, we proposes a dual-branch enhanced multi-task learning network (DBEM), a new architecture that considers both the multiple dependencies of sequences and the heterogeneity of multimodal data, for better multimodal sentiment analysis. The global-local branch takes into account the intra-modal dependencies of different length time subsequences and aggregates global and local features to enrich the feature diversity. The cross-refine branch considers the difference in information density of different modalities and adopts coarse-to-fine fusion learning to model the inter-modal dependencies. Coarse-grained fusion achieves low-level feature reinforcement of audio and visual modalities, and fine-grained fusion improves the ability to integrate information complementarity between different levels of modalities. Finally, multi-task learning is carried out to improve the generalization and performance of the model based on the enhanced fusion features obtained from the dual-branch network. Compared with the single branch network (SBEM, variant of DBEM model) and SOTA methods, the experimental results on the two datasets CH-SIMS and CMU-MOSEI validate the effectiveness of the DBEM model. Wenxiu Geng, Xiangxian Li, Yulong Bian |
ICMR | 2 |
| 2023 | Class-level Structural Relation Modeling and Smoothing for Visual Representation LearningabstractRepresentation learning for images has been advanced by recent progress in more complex neural models such as the Vision Transformers and new learning theories such as the structural causal models. However, these models mainly rely on the classification loss to implicitly regularize the class-level data distributions, and they may face difficulties when handling classes with diverse visual patterns. We argue that the incorporation of the structural information between data samples may improve this situation. To achieve this goal, this paper presents a framework termed Class-level Structural Relation Modeling and Smoothing for Visual Representation Learning (CSRMS), which includes the Class-level Relation Modelling, Class-aware Graph Sampling, and Relational Graph-Guided Representation Learning modules to model a relational graph of the entire dataset and perform class-aware smoothing and regularization operations to alleviate the issue of intra-class visual diversity and inter-class similarity. Specifically, the Class-level Relation Modelling module uses a clustering algorithm to learn the data distributions in the feature space and identify three types of class-level sample relations for the training set; Class-aware Graph Sampling module extends typical training batch construction process with three strategies to sample dataset-level sub-graphs; and Relational Graph-Guided Representation Learning module employs a graph convolution network with knowledge-guided smoothing operations to ease the projection from different visual patterns to the same class. Experiments demonstrate the effectiveness of structured knowledge modelling for enhanced representation learning and show that CSRMS can be incorporated with any state-of-the-art visual representation learning models for performance gains. The source codes and demos have been released at https://github.com/czt117/CSRMS. Zitan Chen, Zhuang Qi, Xiao Cao, Xiangxian Li, Xiangxu Meng, Lei Meng 0001 |
ACM Multimedia | 4 |
| 2023 | Class-aware Convolution and Attentive Aggregation for Image ClassificationabstractDeep learning has been proven to be effective in image classification tasks. However, existing methods may face difficulties in distinguishing complex images due to the distraction caused by diverse image content. To overcome this challenge, we propose a class-aware convolution and attentive aggregation (CA-Net) framework that improves the effectiveness of representation learning and reduces the influence of irrelevant background. CA-Net includes three main modules: the discrete representation learning (DRL) module that uses a group learning method to learn discriminative representations, the class-aware score of discrete representation (CSDR) module that infers class-aware scores to generate weights for representation learners, and the class-aware representation fusion module(CRF) that aggregates class-aware representations using the class-aware scores as a guide. Our experimental results on three benchmarking datasets show that CA-Net improves the performance of state-of-the-art backbones and enhances feature extraction robustness. Zitan Chen, Zhuang Qi, Xiangxian Li, Lei Meng 0001, Xiangxu Meng |
MMAsia | 3 |
| 2022 | Personalized User Interface Elements Recommendation System
Hao Liu 0026, Xiangxian Li, Wei Gai, Jingbo Zhou 0003, Chenglei Yang |
CGI | 2 |
| 2022 | Unsupervised Contrastive Masking for Visual Haze ClassificationabstractHaze classification has gained much attention recently as a cost-effective solution for air quality monitoring. Different from conventional image classification tasks, it requires the classifier to capture the haze patterns of different severity degrees. Existing efforts typically focus on the extraction of effective haze features, such as the dark channel and deep features. However, it is observed that the light-haze images are often mis-classified due to the presence of diverse background scenes. To address this issue, this paper presents an unsupervised contrastive masking (UCM) algorithm to segment the haze regions without any supervision, and develops a dual-channel model-agnostic framework, termed magnifier neural network (MagNet), to effectively use the segmented haze regions to enhance the learning of haze features by conventional deep learning models. Specifically, MagNet employs the haze regions to provide the pixel- and feature-level visual information via three strategies, including Input Augmentation, Network Constraint, and Feature Enhancement, which work as a soft-attention regularizer to alleviates the trade-off between capturing the global scene information and the local information in the haze regions. Experiments were conducted on two datasets in terms of performance comparison, parameter estimation, ablation studies, and case studies, and the results verified that UCM can accurately and rapidly segment the haze regions, and the proposed three strategies of MagNet consistently improve the performance of the state-of-the-art deep learning backbones. Haokai Ma, Xiangxian Li, Zhuang Qi, Lei Meng 0001, Xiangxu Meng |
ICMR | 3 |