EDBT 2026 Demo / reviewers in the wild / expert
Songpei Xu
dblp:317/7181
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Residuals: A Progressive Semantic-Preserving Quantization Approach for Recommendation
Liwen Xiao, Songpei Xu, Da Guo, Yintao Ren, Dongjing Wang, Chuanjiang Luo |
DASFAA (6) | 3 |
| 2026 | Focal-RegionFace: Generating Fine-Grained Multi-attribute Descriptions for Arbitrarily Selected Face Focal RegionsabstractFacial analysis is a fundamental problem in vision–language research, with important applications in affective computing. However, existing methods primarily focus on global facial attributes or single-dimension analysis, lacking fine-grained, interpretable multi-attribute modeling of arbitrary local facial regions. We introduce FaceFocalDesc, a new problem that aims to generate and recognize multi-attribute natural language descriptions for arbitrarily selected facial regions. The target attributes include facial action units, emotional states, and age. We argue that explicit region-level modeling enables more controllable and interpretable facial understanding. To support this task, we construct a new dataset with region-level annotations and corresponding language descriptions. We further propose Focal-RegionFace, a vision–language model fine-tuned from Qwen2.5-VL, which progressively refines its focus on localized facial features through multi-stage training. Experiments show that Focal-RegionFace achieves state-of-the-art performance on the proposed benchmark under both standard and newly introduced metrics, demonstrating its effectiveness in fine-grained region-focused facial analysis. Kaiwen Zheng 0002, Junchen Fu, Songpei Xu, Yaoqin He, Joemon M. Jose, Hu Han 0001, Xuri Ge |
ICMR | 3 |
| 2025 | Progressive Semantic Residual Quantization for Multimodal-Joint Interest Modeling in Music RecommendationabstractIn music recommendation systems, multimodal interest learning is pivotal, which allows the model to capture nuanced preferences, including textual elements such as lyrics and various musical attributes such as different instruments and melodies. Recently, methods that incorporate multimodal content features through semantic IDs have achieved promising results. However, existing methods suffer from two critical limitations: 1) intra-modal semantic degradation, where residual-based quantization processes gradually decouple discrete IDs from original content semantics, leading to semantic drift; and 2) inter-modal modeling gaps, where traditional fusion strategies either overlook modal-specific details or fail to capture cross-modal correlations, hindering comprehensive user interest modeling. To address these challenges, we propose a novel multimodal recommendation framework with two stages. In the first stage, our Progressive Semantic Residual Quantization (PSRQ) method generates modal-specific and modal-joint semantic IDs by explicitly preserving the prefix semantic feature. In the second stage, to model multimodal interest of users, a Multi-Codebook Cross-Attention (MCCA) network is designed to enable the model to simultaneously capture modal-specific interests and perceive cross-modal correlations. Extensive experiments on multiple real-world datasets demonstrate that our framework outperforms state-of-the-art baselines. This framework has been deployed on one of China's largest music streaming platforms, and online A/B tests confirm significant improvements in commercial metrics, underscoring its practical value for industrial-scale recommendation systems. Tianpei Ouyang, Dongjing Wang, Yintao Ren, Songpei Xu, Da Guo, Chuanjiang Luo |
CIKM | 6 |
| 2025 | Climber: Toward Efficient Scaling Laws for Large Recommendation ModelsabstractTransformer-based generative models have achieved remarkable success across domains with various scaling law manifestations. However, our extensive experiments reveal persistent challenges when applying Transformer to recommendation systems: (1) Transformer scaling is not ideal with increased computational resources, due to structural incompatibilities with recommendation-specific features such as multi-source data heterogeneity; (2) critical online inference latency constraints (tens of milliseconds) that intensify with longer user behavior sequences and growing computational demands. We propose Climber, an efficient recommendation framework comprising two synergistic components: the model architecture for efficient scaling and the co-designed acceleration techniques. Our proposed model adopts two core innovations: (1) multi-scale sequence extraction that achieves a time complexity reduction by a constant factor, enabling more efficient scaling with sequence length; (2) dynamic temperature modulation adapting attention distributions to the multi-scenario and multi-behavior patterns. Complemented by acceleration techniques, Climber achieves a 5.15× throughput gain without performance degradation by adopting a ''single user, multiple item'' batched processing and memory-efficient Key-Value caching. Songpei Xu, Da Guo, Xianwen Guo, Bin Huang 0012, Guanlin Wu, Chuanjiang Luo |
CIKM | 1 |
| 2025 | Double-Filter: Efficient Fine-tuning of Pre-trained Vision-Language Models via Patch&Layer FilteringabstractIn this paper, we present a novel approach, termed Double-Filter,to “slim down” the fine-tuning process of vision-language pre-trained (VLP) models via filtering redundancies in feature inputs and architectural components. We enhance the fine-tuning process using two approaches. First, we develop a new patch selection method incorporating image patch filtering through background and foreground separation, followed by a refined patch selection process. Second, we design a genetic algorithm to eliminate redundant fine-grained architecture layers, improving the efficiency and effectiveness of the model. The former makes patch selection semantics more comprehensive, improving inference efficiency while ensuring semantic representation. The latter’s fine-grained layer filter removes architectural redundancy to the extent possible and mitigates the impact on performance. Experimental results demonstrate that the proposed Double-Filter achieves superior efficiency of model fine-tuning and maintains competitive performance compared with the advanced efficient fine-tuning methods on three downstream tasks, VQA, NLVR and Retrieval. In addition, it has been proven to be effective under METER and ViLT VLP models. Yaoqin He, Junchen Fu, Kaiwen Zheng 0002, Songpei Xu, Fuhai Chen, Jie Li 0052, Joemon M. Jose, Xuri Ge |
ICML | 4 |
| 2025 | HandSolo: A Mid-Air Hand Pose Interaction Method Based on Disentangled Degrees-of-Hand-FreedomabstractThis study aims to utilise mid-air hand-pose movements to implement various interactive controls, e.g. dial and slider controlling, through independent low-dimensional embeddings. Towards this, we develop a novel adjustable hand-pose space disentanglement approach for a learnable VAE-based high-to-low dimensional embedding model (HandSolo). It disentangles the latent embeddings intomultiple independent one- or two-dimensional embedding spaces, enabling independent control. HandSolo allows multi-dimensional settings and multi-DOF combinations, providing a new paradigm for flexible and extensible hand-pose interaction systems. Additionally, to exploit model potential and make user interaction comfortable, we propose a visual interaction evaluation strategy (VIEs) to help system designers understand model capability and user habits. Finally, we provide an example virtual interaction system that integrates various virtual interaction objects, showing how our innovations improve their interaction capabilities. Experimental user studies demonstrate the effectiveness of our embedding-disentanglement designs, including discovery experiment (n=4) for VIEs, inspiration experiment (n=4) for approach extensibility, and exploration experiment (n=8) for the virtual interaction system. Songpei Xu, Xuri Ge, Chaitanya Kaul, Roderick Murray-Smith |
ACM Multimedia | 1 |
| 2025 | Hire: Hybrid-Modal Interaction with Multiple Relational Enhancements for Image-Text MatchingabstractImage-Text Matching (ITM) is a fundamental problem in computer vision. The key issue lies in jointly learning the visual and textual representation to estimate their similarity accurately. Most existing methods focus on feature enhancement within modality or feature interaction across modalities, which, however, neglects the contextual information of the object representation based on the inter-object relationships that match the corresponding sentences with rich contextual semantics. In this article, we propose a Hybrid-modal Interaction with multiple Relational Enhancements (termed Hire ) for ITM, which correlates the intra- and inter-modal semantics between objects and words with implicit and explicit relationship modeling. In particular, the explicit intra-modal spatial-semantic graph-based reasoning network is designed to improve the contextual representation of visual objects with salient spatial and semantic relational connectivities, guided by the explicit relationships of the objects’ spatial positions and their scene graph. We use implicit relationship modeling for potential relationship interactions before explicit modeling to improve the fault tolerance of explicit relationship detection. Then the visual and textual semantic representations are refined jointly via inter-modal interactive attention and cross-modal alignment. To correlate the context of objects with the textual context, we further refine the visual semantic representation via cross-level object-sentence and word-image-based interactive attention. Extensive experiments validate that the proposed hybrid-modal interaction with implicit and explicit modeling is more beneficial for ITM. And the proposed Hire obtains new state-of-the-art results on MS-COCO and Flickr30K benchmarks. Xuri Ge, Fuhai Chen, Songpei Xu, Fuxiang Tao, Jie Wang 0072, Joemon M. Jose |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | HpEIS: Learning Hand Pose Embeddings for Multimedia Interactive SystemsabstractWe present a novel Hand-pose Embedding Interactive System (HpEIS) as a virtual sensor, which maps users’ flexible hand poses to a two-dimensional visual space using a Variational Autoencoder (VAE) trained on a variety of hand poses. HpEIS enables visually interpretable and guidable support for user explorations in multimedia collections, using only a camera as an external hand pose acquisition device. We identify general usability issues associated with system stability and smoothing requirements through pilot experiments with expert and inexperienced users. We then design stability and smoothing improvements, including hand-pose data augmentation, an anti-jitter regularisation term added to loss function, stabilising post-processing for movement turning points and smoothing post-processing based on One Euro Filters. In target selection experiments (n=12), we evaluate HpEIS by measures of task completion time and the final distance to target points, with and without the gesture guidance window condition. Experimental responses indicate that HpEIS provides users with a learnable, flexible, stable and smooth mid-air hand movement interaction experience. Songpei Xu, Xuri Ge, Chaitanya Kaul, Roderick Murray-Smith |
ICME | 1 |
| 2024 | 3SHNet: Boosting image-sentence retrieval via visual semantic-spatial self-highlighting
Xuri Ge, Songpei Xu, Fuhai Chen, Jie Wang 0072, Shan An, Joemon M. Jose |
Inf. Process. Manag. | 2 |
| 2024 | MGRR-Net: Multi-level Graph Relational Reasoning Network for Facial Action Unit DetectionabstractThe Facial Action Coding System (FACS) encodes the action units (AUs) in facial images, which has attracted extensive research attention due to its wide use in facial expression analysis. Many methods that perform well on automatic facial action unit (AU) detection primarily focus on modeling various AU relations between corresponding local muscle areas or mining global attention–aware facial features; however, they neglect the dynamic interactions among local-global features. We argue that encoding AU features just from one perspective may not capture the rich contextual information between regional and global face features, as well as the detailed variability across AUs, because of the diversity in expression and individual characteristics. In this article, we propose a novel Multi-level Graph Relational Reasoning Network (termed MGRR-Net ) for facial AU detection. Each layer of MGRR-Net performs a multi-level (i.e., region-level, pixel-wise, and channel-wise level) feature learning. On the one hand, the region-level feature learning from the local face patch features via graph neural network can encode the correlation across different AUs. On the other hand, pixel-wise and channel-wise feature learning via graph attention networks (GAT) enhance the discrimination ability of AU features by adaptively recalibrating feature responses of pixels and channels from global face features. The hierarchical fusion strategy combines features from the three levels with gated fusion cells to improve AU discriminative ability. Extensive experiments on DISFA and BP4D AU datasets show that the proposed approach achieves superior performance than the state-of-the-art methods. Xuri Ge, Joemon M. Jose, Songpei Xu, Xiao Liu 0040, Hu Han 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | A Transfer Learning-Based Method for Personalized State of Health Estimation of Lithium-Ion BatteriesabstractState of health (SOH) estimation of lithium-ion batteries (LIBs) is of critical importance for battery management systems (BMSs) of electronic devices. An accurate SOH estimation is still a challenging problem limited by diverse usage conditions between training and testing LIBs. To tackle this problem, this article proposes a transfer learning-based method for personalized SOH estimation of a new battery. More specifically, a convolutional neural network (CNN) combined with an improved domain adaptation method is used to construct an SOH estimation model, where the CNN is used to automatically extract features from raw charging voltage trajectories, while the domain adaptation method named maximum mean discrepancy (MMD) is adopted to reduce the distribution difference between training and testing battery data. This article extends MMD from classification tasks to regression tasks, which can therefore be used for SOH estimation. Three different datasets with different charging policies, discharging policies, and ambient temperatures are used to validate the effectiveness and generalizability of the proposed method. The superiority of the proposed SOH estimation method is demonstrated through the comparison with direct model training using state-of-the-art machine learning methods and several other domain adaptation approaches. The results show that the proposed transfer learning-based method has wide generalizability as well as a positive precision improvement. Guijun Ma, Songpei Xu, Tao Yang 0003, Zhenbang Du, Limin Zhu 0001, Han Ding 0001, Ye Yuan 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Continuous Interaction with A Smart Speaker via Low-Dimensional Embeddings of Dynamic Hand PoseabstractThis paper presents a new continuous interaction strategy with visual feedback of hand pose and mid-air gesture recognition and control for a smart music speaker, which utilizes only 2 video frames to recognize gestures. Frame-based hand pose features from MediaPipe Hands, containing 21 landmarks, are embedded into a 2 dimensional pose space by an autoencoder. The corresponding space for interaction with the music content is created by embedding high-dimensional music track profiles to a compatible two-dimensional embedding. A PointNet-based model is then applied to classify gestures which are used to control the device interaction or explore music spaces. By jointly optimising the autoencoder with the classifier, we manage to learn a more useful embedding space for discriminating gestures. We demonstrate the functionality of the system with experienced users selecting different musical moods by varying their hand pose. Songpei Xu, Chaitanya Kaul, Xuri Ge, Roderick Murray-Smith |
ICASSP | 1 |
| 2023 | Cross-modal Semantic Enhanced Interaction for Image-Sentence RetrievalabstractImage-sentence retrieval has attracted extensive research attention in multimedia and computer vision due to its promising application. The key issue lies in jointly learning the visual and textual representation to accurately estimate their similarity. To this end, the mainstream schema adopts an object-word based attention to calculate their relevance scores and refine their interactive representations with the attention features, which, however, neglects the context of the object representation on the inter-object relationship that matches the predicates in sentences. In this paper, we propose a Cross-modal Semantic Enhanced Interaction method, termed CMSEI for image-sentence retrieval, which correlates the intra- and inter-modal semantics be-tween objects and words. In particular, we first design the intra-modal spatial and semantic graphs based reasoning to enhance the semantic representations of objects guided by the explicit relationships of the objects’ spatial positions and their scene graph. Then the visual and textual semantic representations are refined jointly via the inter-modal interactive attention and the cross-modal alignment. To correlate the context of objects with the textual context, we further refine the visual semantic representation via the cross-level object-sentence and word-image based interactive attention. Experimental results on seven standard evaluation metrics show that the proposed CMSEI outperforms the state-of-the-art and the alternative approaches on MS-COCO and Flickr30K benchmarks. Xuri Ge, Fuhai Chen, Songpei Xu, Fuxiang Tao, Joemon M. Jose |
WACV | 3 |