Xiaohao Liu

dblp:331/1422 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Integrating Symbolic and Waveform Music Into Large Language Models
Teng Tu 0002, Xiaohao Liu, Yunshan Ma 0002, Ji Qi 0003, Tat-Seng Chua
MMM (2)2
2026 Principled Multimodal Representation Learning
abstract
Multimodal representation learning seeks to create a unified representation space by integrating diverse data modalities to improve multimodal understanding. Traditional methods often depend on pairwise contrastive learning, which relies on a predefined anchor modality, restricting alignment across all modalities. Recent advances have investigated the simultaneous alignment of multiple modalities, yet several challenges remain, such as limitations imposed by fixed anchor points and instability arising from optimizing the product of singular values. To address the challenges, in this paper, we propose Principled Multimodal Representation Learning (PMRL), a novel framework that achieves simultaneous alignment of multiple modalities without anchor dependency in a more stable manner. Specifically, grounded in the theoretical insight that full alignment corresponds to a rank-1 Gram matrix, PMRL optimizes the dominant singular value of the representation matrix to align modalities along a shared leading direction. We propose a softmax-based loss function that treats singular values as logits to prioritize the largest singular value. Besides, instance-wise contrastive regularization on the leading eigenvectors maintains inter-instance separability and prevents representation collapse. Extensive experiments across diverse tasks demonstrate PMRL's superiority compared to baseline methods.
Xiaohao Liu, Xiaobo Xia, See-Kiong Ng, Tat-Seng Chua
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Singular Value Decomposition Based Indoor Localization Using Small Scale Crowd Sensing Data
abstract
Traditional crowd sensing based indoor localization methods rely on large scale pre-collected fingerprint data to construct a radio map with cumbersome prior preparation. However, when they lack floor plan information or only have a little of data is willing to share, the tracking accuracy degrades significantly. In this paper, we propose a singular value decomposition (SVD) track matching scheme to obtain an effective radio map based on small scale crowd sensing data, which is a non-learning based system (SVD-CSP). SVD-CSP fuses received signal strength indicator (RSSI), inertial measurement unit (IMU), and magnetic field strength to label surrounding WiFi access points as marker points. The proposed scheme uses SVD method to directly compute the rotation matrix and displacement vector among the crowd sensing trajectories and attain the reliable tracks. The radio map is constructed and users are tracked according to our developed bidirectional Bayesian filter, which contains forward filter and reverse filter. The density-based spatial clustering of applications with noise (DBSCAN) is embedded within the forward filter to improve the radio map quality. Meanwhile, the reverse filter fuses pedestrian dead reckoning (PDR) and radio map-based localization to track users. Experimental results demonstrate that SVD-CSP can achieve robust localization using extremely sparse crowd trajectories (e.g., 4 trajectories in a 648 m2scenario, 30 trajectories in a 2856 m2scenario) without deep learning training or infrastructure knowledge.
Xiaohao Liu, Yubin Zhao, Xiaofan Li 0001, Huaming Wu, Cheng-Zhong Xu 0001
IEEE Trans. Mob. Comput.1
2025 WTB-YOLO: Wind Turbine Blade Defect Detection with Scale Sensitivity and Cascade Structure
abstract
The wind turbine blade represents a pivotal element within the context of wind power generation systems, with its performance and reliability exerting a direct influence on the overall efficiency and safety of the system. The identification of defects in wind turbine blades is of considerable economic and safety significance. The existing YOLOv8 target detector is still inadequate in terms of loss function in the process of detecting minor defects in wind turbine blades, which affects the learning efficiency and model performance. Therefore, a new loss function is proposed, which is realised by combining scale sensitivity and cascade structure. The proposed method, implemented in the YOLOv8 detector, markedly enhances both the accuracy of detection and the computational efficiency. Moreover, wind turbine blade defect detection is typically implemented in edge devices with constrained performance. To address this issue, a lightweight convolution network is introduced, reducing the number of parameters by 60% to 1.12M, thereby achieving a balance between performance and efficiency. The experimental results on the DTU and the datasets constructed by our own research team demonstrate improvements in detection accuracy by 2.28% and 1.8%, respectively.
Xiaohao Liu, Ruolong Wang, Bing Jia, Baoqi Huang, Xiaoyue Zhu
CSCWD1
2025 COS-SLAM: Coordinate Attention Semantic SLAM with Pixel-to-Line Transformer
Handong Shen, Lingyu Liang, Xiaohao Liu, Xinchao Li, Guoxi Sun, Shuangping Huang
ICIG (2)4
2025 LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential Recommendation
abstract
Sequential recommendation aims to predict users' future interactions by modeling collaborative filtering (CF) signals from historical behaviors of similar users or items. Traditional sequential recommenders predominantly rely on ID-based embeddings, which capture CF signals through high-order co-occurrence patterns. However, these embeddings depend solely on past interactions, lacking transferable knowledge to generalize to unseen domains. Recent advances in large language models (LLMs) have motivated text-based recommendation approaches that derive item representations from textual descriptions. While these methods enhance generalization, they fail to encode CF signals-i.e., latent item correlations and preference patterns-crucial for effective recommendation. We argue that an ideal embedding model should seamlessly integrate CF signals with rich semantic representations to improve both in-domain and out-of-domain recommendation performance. To this end, we propose LLM2Rec, a novel embedding model tailored for sequential recommendation, integrating the rich semantic understanding of LLMs with CF awareness. Our approach follows a two-stage training framework: (1) Collaborative Supervised Fine-tuning, which adapts LLMs to infer item relationships based on historical interactions, and (2) Item-level Embedding Modeling, which refines these specialized LLMs into structured item embedding models that encode both semantic and collaborative information. Extensive experiments on real-world datasets demonstrate that LLM2Rec effectively improves recommendation quality across both in-domain and out-of-domain settings. Our findings highlight the potential of leveraging LLMs to build more robust, generalizable embedding models for sequential recommendation. Our codes are available at: https://github.com/HappyPointer/LLM2Rec.
Yingzhi He, Xiaohao Liu, An Zhang 0003, Yunshan Ma 0002, Tat-Seng Chua
KDD (2)2
2025 Fine-tuning Multimodal Large Language Models for Product Bundling
abstract
Recent advances in product bundling have leveraged multimodal information through sophisticated encoders, but remain constrained by limited semantic understanding and a narrow scope of knowledge. Therefore, some attempts employ In-context Learning (ICL) to explore the potential of large language models (LLMs) for their extensive knowledge and complex reasoning abilities. However, these efforts are inadequate in understanding mulitmodal data and exploiting LLMs' knowledge for product bundling. To bridge the gap, we introduce Bundle-MLLM, a novel framework that fine-tunes LLMs through a hybrid item tokenization approach within a well-designed optimization strategy. Specifically, we integrate textual, media, and relational data into a unified tokenization, introducing a soft separation token to distinguish between textual and non-textual tokens. Additionally, a streamlined yet powerful multimodal fusion module is employed to embed all non-textual features into a single, informative token, significantly boosting efficiency. To tailor product bundling tasks for LLMs, we reformulate the task as a multiple-choice question with candidate items as options. We further propose a progressive optimization strategy that fine-tunes LLMs for disentangled objectives: learning bundle patterns and enhancing multimodal semantic understanding specific to product bundling. Extensive experiments demonstrate that our approach outperforms a range of state-of-the-art (SOTA) methods. Codes are available at https://github.com/Xiaohao-Liu/Bundle-MLLM
Xiaohao Liu, Zhulin Tao, Yunshan Ma 0002, Yinwei Wei, Tat-Seng Chua
KDD (1)1
2025 Towards Modality Generalization: A Benchmark and Prospective Analysis
Xiaohao Liu, Xiaobo Xia, See-Kiong Ng, Tat-Seng Chua
ACM Multimedia1
2025 Continual Multimodal Contrastive Learning
abstract
Multimodal Contrastive Learning (MCL) advances in aligning different modalities and generating multimodal representations in a joint space. By leveraging contrastive learning across diverse modalities, large-scale multimodal data enhances representational quality. However, a critical yet often overlooked challenge remains: multimodal data is rarely collected in a single process, and training from scratch is computationally expensive. Instead, emergent multimodal data can be used to optimize existing models gradually, \textit{i.e.}, models are trained on a sequence of modality pair data. We define this problem as Continual Multimodal Contrastive Learning (CMCL), an underexplored yet crucial research direction at the intersection of multimodal and continual learning. In this paper, we formulate CMCL through two specialized principles of stability and plasticity. We theoretically derive a novel optimization-based method, which projects updated gradients from dual sides onto subspaces where any gradient is prevented from interfering with the previously learned knowledge. Two upper bounds provide theoretical insights on both stability and plasticity in our solution. Beyond our theoretical contributions, we conduct experiments on multiple datasets by comparing our method against advanced continual learning baselines. The empirical results further support our claims and demonstrate the efficacy of our method. Our codes are available at https://github.com/Xiaohao-Liu/CMCL.
Xiaohao Liu, Xiaobo Xia, See-Kiong Ng, Tat-Seng Chua
NeurIPS1
2025 L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models
abstract
Large language models (LLMs) have achieved notable progress. Despite their success, next-token prediction (NTP), the dominant method for LLM training and inference, is constrained in both contextual coverage and inference efficiency due to its inherently sequential process. To overcome these challenges, we propose leap multi-token prediction~(L-MTP), an innovative token prediction method that extends the capabilities of multi-token prediction (MTP) by introducing a leap-based mechanism. Unlike conventional MTP, which generates multiple tokens at adjacent positions, L-MTP strategically skips over intermediate tokens, predicting non-sequential ones in a single forward pass. This structured leap not only enhances the model's ability to capture long-range dependencies but also enables a decoding strategy specially optimized for non-sequential leap token generation, effectively accelerating inference. We theoretically demonstrate the benefit of L-MTP in improving inference efficiency. Experiments across diverse benchmarks validate its merit in boosting both LLM performance and inference speed. The source code is available at https://github.com/Xiaohao-Liu/L-MTP.
Xiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang, Xianzhi Yu, Xiu Su, Shuo Yang 0006, See-Kiong Ng, Tat-Seng Chua
NeurIPS1
2024 Leveraging Identity-Specific Facial Contours for Enhanced Heart Rate Estimation in Remote Photoplethysmography
abstract
Remote photoplethysmography (rPPG) estimates heart rate by capturing blood volume pulse (BVP) signals from subtle pixel variations in video frames. This study presents a novel approach that leverages facial physiological characteristics to improve heart rate estimation. Specifically, we introduce a method for extracting Identity-Specific Facial Contours (ISFCs) and utilize a self-learning combination mechanism that directs the model’s attention to these ISFCs, resulting in significant improvements in both accuracy and robustness. Our rFaceNet model effectively extracts ISFCs from temporally normalized frames using a Temporal Compressor Unit (TCU) and refines focus on relevant facial regions via a Cross-Task Feature Combiner (CTFC). Through meticulous training, rFaceNet significantly enhances the quality and interpretability of facial physiological signals when compared to previous approaches. Moreover, our method outperforms State-of-the-Art (SOTA) models across various heart rate estimation benchmarks.
Dali Zhu, Hualin Zeng, Xiaohao Liu, Jiaqi Zheng 0008
BIBM4
2024 Collaborative Denoising Shilling Attack for Recommendation Systems
abstract
Recommendation systems, while enhancing user experiences and increasing platform revenues, are vulnerable to shilling attacks due to their ubiquity and inherent openness. These attacks manipulate recommendations by strategically injecting fake profiles, often capitalizing on noisy data inconsistencies to produce imperceptible, attack-aware profiles. Existing research largely emphasizes heuristic and deep neural network (DNN) methodologies, overlooking this noise-impact aspect. In this context, we present the Collaborative Denoising Shilling Attack (CDSA) model. CDSA integrates the collaborative information bottleneck (CIB) to distill latent features, optimizing mutual information between denoised inputs and imperceptible, attack-aware features. The multi-feature learning component refines these features via multiple objectives, enhancing the information bottleneck representation. These objectives are sequentially optimized, following a hierarchical framework. Our empirical tests affirm CDSA’s superiority over traditional attack strategies and cutting-edge methods, such as AUSH [1] and GSPAttack [2].
Xiaohao Liu, Hualin Zeng
CSCWD1
2024 Pest-YOLO: A Lightweight Pest Detection Model Based on Multi-level Feature Fusion
Xiaoyue Zhu, Bing Jia, Baoqi Huang, Xiaohao Liu, Winston Khoon Guan Seah
ICIC (4)5
2024 LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
abstract
As online music consumption increasingly shifts towards playlist-based listening, the task of playlist continuation, in which an algorithm suggests songs to extend a playlist in a personalized and musically cohesive manner, has become vital to the success of music streaming services. Currently, many existing playlist continuation approaches rely on collaborative filtering methods to perform their recommendations. However, such methods will struggle to recommend songs that lack interaction data, an issue known as the cold-start problem. Current approaches to this challenge design complex mechanisms for extracting relational signals from sparse collaborative signals and integrating them into content representations. However, these approaches leave content representation learning out of scope and utilize frozen, pre-trained content models that may not be aligned with the distribution or format of a specific musical setting. Furthermore, even the musical state-of-the-art content modules are either (1) incompatible with the cold-start setting or (2) unable to effectively integrate cross-modal and relational signals. In this paper, we introduce LARP, a multi-modal cold-start playlist continuation model, to effectively overcome these limitations. LARP is a three-stage contrastive learning framework that integrates both multi-modal and relational signals into its learned representations. Our framework uses increasing stages of task-specific abstraction: within-track (language-audio) contrastive loss, track-track contrastive loss, and track-playlist contrastive loss. Experimental results on two publicly available datasets demonstrate the efficacy of LARP over uni-modal and multi-modal models for playlist continuation in a cold-start setting. Finally, this work pioneers the perspective of addressing cold-start recommendation via relational representation learning. Code and dataset are released at: https://github.com/Rsalganik1123/LARP/
Rebecca Salganik, Xiaohao Liu, Yunshan Ma 0002, Jian Kang 0008, Tat-Seng Chua
KDD2
2024 Leveraging Multimodal Features and Item-level User Feedback for Bundle Construction
abstract
Automatic bundle construction is a crucial prerequisite step in various bundle-aware online services. Previous approaches are mostly designed to model the bundling strategy of existing bundles. However, it is hard to acquire large-scale well-curated bundle dataset, especially for those platforms that have not offered bundle services before. Even for platforms with mature bundle services, there are still many items that are included in few or even zero bundles, which give rise to sparsity and cold-start challenges in the bundle construction models. To tackle these issues, we target at leveraging multimodal features, item-level user feedback signals, and the bundle composition information, to achieve a comprehensive formulation of bundle construction. Nevertheless, such formulation poses two new technical challenges: 1) how to learn effective representations by unifying multiple features optimally, and 2) how to address the problems of modality missing, noise, and sparsity problems induced by the incomplete query bundles. In this work, to address these technical challenges, we propose a Contrastive Learning-enhanced Hierarchical Encoder method (CLHE). Specifically, we use self-attention modules to combine the multimodal and multi-item features, and then leverage both item- and bundle-level contrastive learning to enhance the representation learning, thus to counter the modality missing, noise, and sparsity problems. Extensive experiments on four datasets in two application domains demonstrate that our method outperforms a list of SOTA methods. The code and dataset are available at https://github.com/Xiaohao-Liu/CLHE.
Yunshan Ma 0002, Xiaohao Liu, Yinwei Wei, Zhulin Tao, Xiang Wang 0010, Tat-Seng Chua
WSDM2
2023 Strategy-aware Bundle Recommender System
abstract
A bundle is a group of items that provides improved services to users and increased profits for sellers. However, locating the desired bundles that match the users' tastes still challenges us, due to the sparsity issue. Despite the remarkable performance of existing approaches, we argue that they seldom consider the bundling strategy (i.e., how the items within a bundle are associated with each other) in the bundle recommendation, resulting in the suboptimal user and bundle representations for their interaction prediction. Therefore, we propose to model the strategy-aware user and bundle representations for the bundle recommendation.
Yinwei Wei, Xiaohao Liu, Yunshan Ma 0002, Xiang Wang 0010, Liqiang Nie, Tat-Seng Chua
SIGIR2
2023 Self-Supervised Learning for Multimedia Recommendation
abstract
Learning representations for multimedia content is critical for multimedia recommendation. Current representation learning methods roughly fall into two groups: (1) using the historical interactions to create ID embeddings of users and items, and (2) treating multi-modal data as the side information of items to enrich their ID embeddings. Each user-item interaction offers the supervisory signal to optimize the representation learning by the traditional supervised learning paradigm. Due to the overlook of the multi-modal patterns ($e.g.$, co-occurrence of visual, acoustic, textual features in micro-videos a user saw before, and her behavioral features) hidden in the data, these methods are insufficient to create powerful representations and obtain satisfactory recommendation accuracy. To capture multi-modal patterns in the data itself, we go beyond the supervised learning paradigm, and incorporate the idea of self-supervised learning (SSL) into multimedia recommendation. Specifically, SSL consists of two components: (1) data augmentation upon multi-modal contents, where we design three operators — feature dropout (FD), feature masking (FM), feature fine and coarse spaces (FAC) — to generate multiple views of individual items; and (2) contrastive learning, which differentiates the views of an item from the others’ to distill additional supervisory signals. Clearly, SSL enables us to explore and exhibit the underlying relations among modalities, thereby resulting in powerful representations. We denote the generic framework by Self-supervised Learning-guided Multimedia Recommendation (SLMRec). Extensive experiments are performed on three real-world datasets, showing that SLMRec achieves significant improvements over several state-of-the-art baselines like LightGCN [1], MMGCN [2]. Further analysis shows how SSL affects recommendation performance.
Zhulin Tao, Xiaohao Liu, Yewei Xia, Xiang Wang 0010, Lifang Yang, Xianglin Huang, Tat-Seng Chua
IEEE Trans. Multim.2
2022 EliMRec: Eliminating Single-modal Bias in Multimedia Recommendation
abstract
The main idea of multimedia recommendation is to introduce the profile content of multimedia documents as an auxiliary, so as to endow recommenders with generalization ability and gain better performance. However, recent studies using non-uniform datasets roughly fuse single-modal features into multi-modal features and adopt the strategy of directly maximizing the likelihood of user preference scores, leading to the single-modal bias. Owing to the defect in architecture, there is still room for improvement for recent multimedia recommendation.
Xiaohao Liu, Zhulin Tao, Jiahong Shao, Lifang Yang, Xianglin Huang
ACM Multimedia1