Xiaofeng Pan

dblp:220/4010 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TuningIQA: Fine-Grained Blind Image Quality Assessment for Livestreaming Camera Tuning
abstract
Livestreaming has become increasingly prevalent in modern visual communication, where automatic camera quality tuning is essential for delivering superior user Quality of Experience (QoE). Such tuning requires accurate blind image quality assessment (BIQA) to guide parameter optimization decisions. Unfortunately, the existing BIQA models typically only predict an overall coarse-grained quality score, which cannot provide fine-grained perceptual guidance for precise camera parameter tuning. To bridge this gap, we first establish FGLive-10K, a comprehensive fine-grained BIQA database containing 10,185 high-resolution images captured under varying camera parameter configurations across diverse livestreaming scenarios. The dataset features 50,925 multi-attribute quality annotations and 19,234 fine-grained pairwise preference annotations. Based on FGLive-10K, we further develop TuningIQA, a fine-grained BIQA metric for livestreaming camera tuning, which integrates human-aware feature extraction and graph-based camera parameter fusion. Extensive experiments and comparisons demonstrate that TuningIQA significantly outperforms state-of-the-art BIQA methods in both score regression and fine-grained quality ranking, achieving superior performance when deployed for livestreaming camera tuning.
Xiangfei Sheng, Zhichao Duan 0002, Xiaofeng Pan, Yipo Huang, Zhichao Yang 0013, Pengfei Chen 0003, Leida Li
AAAI3
2026 Fine-grained Image Quality Assessment for Perceptual Image Restoration
abstract
Recent years have witnessed remarkable achievements in perceptual image restoration (IR), creating an urgent demand for accurate image quality assessment (IQA), which is essential for both performance comparison and algorithm optimization. Unfortunately, the existing IQA metrics exhibit inherent weakness for IR task, particularly when distinguishing fine-grained quality differences among restored images. To address this dilemma, we contribute the first-of-its-kind fine-grained image quality assessment dataset for image restoration, termed FGRestore, comprising 18,408 restored images across six common IR tasks. Beyond conventional scalar quality scores, FGRestore was also annotated with 30,886 fine-grained pairwise preferences. Based on FGRestore, a comprehensive benchmark was conducted on the existing IQA metrics, which reveal significant inconsistencies between score-based IQA evaluations and the fine-grained restoration quality. Motivated by these findings, we further propose FGResQ, a new IQA model specifically designed for image restoration, which features both coarse-grained score regression and fine-grained quality ranking. Extensive experiments and comparisons demonstrate that FGResQ significantly outperforms state-of-the-art IQA metrics.
Xiangfei Sheng, Xiaofeng Pan, Zhichao Yang 0013, Pengfei Chen 0003, Leida Li
AAAI2
2026 S3CD: A Self-Supervised Semantic Change Detection Method by Mining Transition Patterns and Consistency in Remote Sensing Images
abstract
Semantic change detection (SCD) endeavors to identify land-cover changes from multitemporal remote sensing images, providing essential information for various applications. Nevertheless, conventional supervised SCD methods necessitate extensive pixel-level annotations, limiting their applicability. The capability of self-supervised methods to learn feature representations with large amounts of unlabeled data and minimal annotation, and to achieve superior performance, has made them one of the hot topics in remote sensing. However, most self-supervised methods in remote sensing are primarily designed to learn general semantic representations of images, which limits their effectiveness for tasks like SCD that require the analysis of complex semantic transformations. To address this, we propose a multistage, multitask, and multilevel self-supervised network, named S3CD, that learns semantic changes from bi-temporal remote sensing images across scene, pixel, and prototype levels in two stages. In particular, in Stage 2, the network enhances the robustness of SCD by learning semantic consistency within the semantic stable categories across different temporal and capturing the temporal patterns of semantic change categories. We evaluate S3CD on two widely used remote sensing change detection (CD) datasets, where it outperformed state-of-the-art self-supervised and supervised SCD methods. Notably, in the binary CD (BCD) task (i.e., detecting the locations of changes), S3CD also outperforms most supervised learning methods. Therefore, this approach facilitates the application of self-supervised learning in the field of remote sensing CD.
Jiayi Li 0001, Xiaofeng Pan, Xin Huang 0002
IEEE Trans. Cybern.3
2025 Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning
abstract
Recent works of music representation learning mainly focus on learning acoustic music representations with unlabeled audios or further attempt to acquire multi-modal music representations with scarce annotated audio-text pairs. They either ignore the language semantics or rely on labeled audio datasets that are difficult and expensive to create. Moreover, merely modeling semantic space usually fails to achieve satisfactory performance on music recommendation tasks since the user preference space is ignored. In this paper, we propose a novel Hierarchical Two-stage Contrastive Learning (HTCL) method that models similarity from the semantic perspective to the user perspective hierarchically to learn a comprehensive music representation bridging the gap between semantic and user preference spaces. We devise a scalable audio encoder and leverage a pre-trained BERT model as the text encoder to learn audio-text semantics via large-scale contrastive pre-training. Further, we explore a simple yet effective way to exploit interaction data from our online music platform to adapt the semantic space to user preference space via contrastive fine-tuning, which differs from previous works that follow the idea of collaborative filtering. As a result, we obtain a powerful audio encoder that not only distills language semantics from the text encoder but also models similarity in user preference space with the integrity of semantic space preserved. Experimental results on both music semantic and recommendation tasks confirm the effectiveness of our method.
Xiaofeng Pan, Jing Chen 0073, Haitong Zhang, Menglin Xing, Jiayi Wei, Xuefeng Mu, Zhongqian Xie
ICMR1
2024 A Cross-Angle Propagation Network for Built-Up Area Extraction by Fusing Spatial-Spectral-Angular Features From the ZY-3 Multiview Satellite Imagery: Dataset and Analysis of China's 41 Major Cities
abstract
Obtaining timely and reliable built-up area (BUA) information across extensive geographical zones holds crucial significance for understanding environmental change and human activities. BUAs often exhibit detailed textures and structures in high-resolution imagery but also present strong heterogeneity. Current methods for BUA extraction primarily relied on planar information from single-view imagery, struggling to effectively capture the 3-D attributes of urban landscapes. Therefore, to address this challenge, this article proposes a cross-angle propagation network (CAPNet) based on multiview remote sensing stereo observation imagery. Our contributions are threefold: 1) we propose the cross-angle fusion module (CAFM) to exploit BUA’s complementary spatial-spectral-angular context across different viewing angles. This module leverages attention mechanisms for the automated acquisition of multiangle feature representation learning from diverse angle combinations. 2) We propose a multiangular propagation decoder (MAPD) that pioneers the exploration of gradually propagating multiangle disparity information through bidirectional-adjacent feature fusion across hierarchical levels. 3) We construct a large-scale, high-resolution multiview BUA (MVBA) dataset over China’s 41 major cities based on the ZY-3 satellites. Extensive experiment results on MVBA and the public WV-3 multiview semantic stereo datasets verify CAPNet’s superiority to existing state-of-the-art (SOTA) models, on preserving overall BUA shape, edge, and internal structures. The dataset and the source code of CAPNet will be publicly available athttps://github.com/zuo-ux/Cross-Angle-Propagation-Network.
Renxiang Zuo, Xin Huang 0002, Jiayi Li 0001, Xiaofeng Pan
IEEE Trans. Geosci. Remote. Sens.4
2023 DPAN: Dynamic Preference-based and Attribute-aware Network for Relevant Recommendations
abstract
In e-commerce platforms, the relevant recommendation is a unique scenario providing related items for a trigger item that users are interested in. However, users' preferences for the similarity and diversity of recommendation results are dynamic and vary under different conditions. Moreover, individual item-level diversity is too coarse-grained since all recommended items are related to the trigger item. Thus, the two main challenges are to learn fine-grained representations of similarity and diversity and capture users' dynamic preferences for them under different conditions. To address these challenges, we propose a novel method called the Dynamic Preference-based and Attribute-aware Network (DPAN) for predicting Click-Through Rate (CTR) in relevant recommendations. Specifically, based on Attribute-aware Activation Values Generation (AAVG), Bi-dimensional Compression-based Re-expression (BCR) is designed to obtain similarity and diversity representations of user interests and item information. Then Shallow and Deep Union-based Fusion (SDUF) is proposed to capture users' dynamic preferences for the diverse degree of recommendation results according to various conditions. DPAN has demonstrated its effectiveness through extensive offline experiments and online A/B testing, resulting in a significant 7.62% improvement in CTR. Currently, DPAN has been successfully deployed on our e-commerce platform serving the primary traffic for relevant recommendations.
Yingmin Su, Xiaofeng Pan, Yufeng Wang 0001, Nan Xu 0018, Chengjun Mao, Bo Cao 0007
CIKM3
2023 FAN: Fatigue-Aware Network for Click-Through Rate Prediction in E-commerce Recommendation
Naiyin Liu, Xiaofeng Pan, Yingmin Su, Chengjun Mao
DASFAA (4)3
2023 MOEF: Modeling Occasion Evolution in Frequency Domain for Promotion-Aware Click-Through Rate Prediction
Xiaofeng Pan, Yibin Shen, Jing Zhang 0037, Hong Wen 0002, Chengjun Mao
DASFAA (2)1
2022 MetaCVR: Conversion Rate Prediction via Meta Learning in Small-Scale Recommendation Scenarios
abstract
Different from large-scale platforms such as Taobao and Amazon, CVR modeling in small-scale recommendation scenarios is more challenging due to the severe Data Distribution Fluctuation (DDF) issue. DDF prevents existing CVR models from being effective since 1) several months of data are needed to train CVR models sufficiently in small scenarios, leading to considerable distribution discrepancy between training and online serving; and 2) e-commerce promotions have significant impacts on small scenarios, leading to distribution uncertainty of the upcoming time period. In this work, we propose a novel CVR method named MetaCVR from a perspective of meta learning to address the DDF issue. Firstly, a base CVR model which consists of a Feature Representation Network (FRN) and output layers is designed and trained sufficiently with samples across months. Then we treat time periods with different data distributions as different occasions and obtain positive and negative prototypes for each occasion using the corresponding samples and the pre-trained FRN. Subsequently, a Distance Metric Network (DMN) is devised to calculate the distance metrics between each sample and all prototypes to facilitate mitigating the distribution uncertainty. At last, we develop an Ensemble Prediction Network (EPN) which incorporates the output of FRN and DMN to make the final CVR prediction. In this stage, we freeze the FRN and train the DMN and EPN with samples from recent time period, therefore effectively easing the distribution discrepancy. To the best of our knowledge, this is the first study of CVR prediction targeting the DDF issue in small-scale recommendation scenarios. Experimental results on real-world datasets validate the superiority of our MetaCVR and online A/B test also shows our model achieves impressive gains of 11.92% on PCVR and 8.64% on GMV.
Xiaofeng Pan, Jing Zhang 0037, Keren Yu, Hong Wen 0002, Chengjun Mao
SIGIR1
2022 MAT: Motion-aware multi-object tracking
Shoudong Han, Piao Huang, En Yu, Donghaisheng Liu, Xiaofeng Pan
Neurocomputing6