Baolin Liu 0001

dblp:71/2948-1 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0003-0975-2316ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 13 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 EGFE-Net: An edge-guided and feature elimination network for small object detection
Baolin Liu 0001
Expert Syst. Appl.2
2025 LLM-Based Keyphrase-Augmented Framework for Semantic Relevance Assessment in E-Commerce
Songyan Liu, Tianshu Wu, PengjieWang, Jian Xu 0015, Bo Zheng 0007, Baolin Liu 0001
DASFAA (6)11
2025 Interpretable Word Representation Learning Framework for Modeling Semantic Relevance in E-commerce
Tianshu Wu, Pengjie Wang 0002, Jian Xu 0015, Bo Zheng 0007, Baolin Liu 0001
DASFAA (6)9
2025 An fMRI-based auditory decoding framework combined with convolutional neural network for predicting the semantics of real-life sounds from brain activity
Mingqian Zhao, Baolin Liu 0001
Appl. Intell.2
2025 TFG-Net: A Text Feature-Guided Network for Small Traffic Sign Detection
abstract
Detecting small signs in complex real-world environments remains challenging due to limited feature information and interference from other objects. In this article, we propose a novel text feature-guided network (TFG-Net) to improve the performance of the small signs detection not only enhancing the feature information of small signs but also avoiding the influence of other objects. As the name suggests, TFG-Net incorporates a text detection branch, which extracts additional textual features from the signs and supplies them to the object detection branch. Furthermore, the object detection branch of TFG-Net optimizes the backbone network's output structure by merging deep features and introducing a high-resolution feature layer. Finally, a fusion method that enhances both overall and local features is proposed to fully integrate detailed and semantic information. Experimental results display that our TFG-Net reaches the highest mean average precision (mAP) of 92.5% on the public datasets Tsinghua-Tencent 100K (TT100K), 83.7% on CCTSDB2021, and 79.1% on DFG, surpassing current state-of-the-art object detectors.
Renxin Chu, Baolin Liu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 A multimodal fusion-based deep learning framework combined with local-global contextual TCNs for continuous emotion recognition from videos
Congbao Shi, Baolin Liu 0001
Appl. Intell.3
2024 Semantic-guided spatio-temporal attention for few-shot action recognition
Baolin Liu 0001
Appl. Intell.2
2024 TFRS: A task-level feature rectification and separation method for few-shot video action recognition
Yanfei Qin, Baolin Liu 0001
Neural Networks2
2024 An fMRI-based visual decoding framework combined with two-stage learning and asynchronous iterative update strategy
Baolin Liu 0001
Pattern Anal. Appl.3
2023 PS-SA: An Efficient Self-Attention via Progressive Sampling for User Behavior Sequence Modeling
abstract
As the self-attention mechanism offers powerful capabilities for capturing sequential relationships, it has become increasingly popular to use it for modeling user behavior sequences in recommender systems. However, the self-attention mechanism has a quadratic computational complexity of O(n^2), as it conducts interactions among all item pairs in the sequence. This can lead to expensive model training and slow inference speeds, which may hinder practical deployment. To this end, we pursue to develop alternative approaches to improve the efficiency of the self-attention mechanism. We observe that the attention scores calculated from each item interacting with other items (including itself) are sparse, indicating that there are limited valuable item pairs (with non-zero attention weight) that contribute to the final output. This motivates us to develop effective strategies for discerning valuable items and computing attention scores solely for these items, thereby minimizing the consumption of unnecessary computations. Herein, we present a novel Progressive Sampling-based Self-Attention (PS-SA) mechanism, which utilizes a learnable progressive sampling strategy to identify the most valuable items. Subsequently, we solely utilize these selected items to produce the final output. Experiments on academic and production datasets demonstrate PS-SA could still achieve promising results while reducing computational costs. It is notable that we have successfully deployed it on Alibaba display advertising system, resulting in a 2.6% CTR and 1.3% RPM increase.
Jiacen Hu, Zhangming Chan, Yu Zhang 0176, Shuguang Han, Siyuan Lou, Baolin Liu 0001, Han Zhu 0001, Yuning Jiang 0001, Jian Xu 0015, Bo Zheng 0007
CIKM6
2023 Capturing Conversion Rate Fluctuation during Sales Promotions: A Novel Historical Data Reuse Approach
abstract
Conversion rate (CVR) prediction is one of the core components in online recommender systems, and various approaches have been proposed to obtain accurate and well-calibrated CVR estimation. However, we observe that a well-trained CVR prediction model often performs sub-optimally during sales promotions. This can be largely ascribed to the problem of the data distribution shift, in which the conventional methods no longer work. To this end, we seek to develop alternative modeling techniques for CVR prediction. Observing similar purchase patterns across different promotions, we propose reusing the historical promotion data to capture the promotional conversion patterns. Herein, we propose a novel Historical Data Reuse (HDR) approach that first retrieves historically similar promotion data and then fine-tunes the CVR prediction model with the acquired data for better adaptation to the promotion mode. HDR consists of three components: an automated data retrieval module that seeks similar data from historical promotions, a distribution shift correction module that re-weights the retrieved data for better aligning with the target promotion, and a TransBlock module that quickly fine-tunes the original model for better adaptation to the promotion mode. Experiments conducted with real-world data demonstrate the effectiveness of HDR, as it improves both ranking and calibration metrics to a large extent. HDR has also been deployed on the display advertising system in Alibaba, bringing a lift of 9% RPM and 16% CVR during Double 11 Sales in 2022.
Zhangming Chan, Yu Zhang 0176, Shuguang Han, Xiang-Rong Sheng, Siyuan Lou, Jiacen Hu, Baolin Liu 0001, Yuning Jiang 0001, Jian Xu 0015, Bo Zheng 0007
KDD8
2023 OTDE: optimal transport distribution enhancement for few-shot video recognition
Yanfei Qin, Baolin Liu 0001
Appl. Intell.2
2023 Construction of the brain-inspired computing model verified by spatiotemporal correspondence between the hierarchical computation of the model and the complex multi-stage processing of the human brain during facial expression recognition
Qianyi Zhang, Baolin Liu 0001
Appl. Intell.2
2022 KDM: A knowledge-guided and data-driven method for few-shot video action recognition
Yanfei Qin, Baolin Liu 0001
Neurocomputing2
2021 Reconstruction of natural images from evoked brain activity with a dictionary-based invertible encoding procedure
Chao Li 0052, Baolin Liu 0001, Jianguo Wei
Neurocomputing2
2020 Visual Encoding and Decoding of the Human Brain Based on Shared Features
abstract
Using a convolutional neural network to build visual encoding and decoding models of the human brain is a good starting point for the study on relationship between deep learning and human visual cognitive mechanism. However, related studies have not fully considered their differences. In this paper, we assume that only a portion of neural network features is directly related to human brain signals, which we call shared features. In the encoding process, we extract shared features from the lower and higher layers of the neural network, and then build a non-negative sparse map to predict brain activities. In the decoding process, we use back-propagation to reconstruct visual stimuli, and use dictionary learning and a deep image prior to improve the robustness and accuracy of the algorithm. Experiments on a public fMRI dataset confirm the rationality of the encoding models, and comparing with a recently proposed method, our reconstruction results obtain significantly higher accuracy.
Chao Li 0052, Baolin Liu 0001, Jianguo Wei
IJCAI2
2018 Decoding natural images from evoked brain activities using encoding models with invertible mapping
Chao Li 0052, Junhai Xu, Baolin Liu 0001
Neural Networks3