Songlin Fan

dblp:312/7838 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-4906-6660ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MDiff4STR: Mask Diffusion Model for Scene Text Recognition
abstract
Mask Diffusion Models (MDMs) have recently emerged as a promising alternative to auto-regressive models (ARMs) for vision-language tasks, owing to their flexible balance of efficiency and accuracy. In this paper, for the first time, we introduce MDMs into the Scene Text Recognition (STR) task. We show that vanilla MDM lags behind ARMs in terms of accuracy, although it improves recognition efficiency. To bridge this gap, we propose MDiff4STR, a Mask Diffusion model enhanced with two key improvement strategies tailored for STR. Specifically, we identify two key challenges in applying MDMs to STR: noising gap between training and inference, and overconfident predictions during inference. Both significantly hinder the performance of MDMs. To mitigate the first issue, we develop six noising strategies that better align training with inference behavior. For the second, we propose a token-replacement noise mechanism that provides a non-mask noise type, encouraging the model to reconsider and revise overly confident but incorrect predictions. We conduct extensive evaluations of MDiff4STR on both standard and challenging STR benchmarks, covering diverse scenarios including irregular, artistic, occluded, and Chinese text, as well as whether the use of pretraining. Across these settings, MDiff4STR consistently outperforms popular STR models, surpassing state-of-the-art ARMs in accuracy, while maintaining fast inference with only three denoising steps. Code: https://github.com/Topdu/OpenOCR.
Yongkun Du, Miaomiao Zhao, Songlin Fan, Zhineng Chen, Caiyan Jia, Yu-Gang Jiang 0001
AAAI3
2025 VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
abstract
Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics for video editing are still notably absent. To address this, we introduce VE-Bench, a benchmark suite tailored to the assessment of text-driven video editing. This suite includes VE-Bench DB, a video quality assessment (VQA) database for video editing. VE-Bench DB encompasses a diverse set of source videos featuring various motions and subjects, along with multiple distinct editing prompts, editing results from 8 different models, and the corresponding Mean Opinion Scores (MOS) from 24 human annotators. Based on VE-Bench DB, we further propose VE-Bench QA, a quantitative human-aligned measurement for the text-driven video editing task. In addition to the aesthetic, distortion, and other visual quality indicators that traditional VQA methods emphasize, VE-Bench QA focuses on the text-video alignment and the relevance modeling between source and edited videos. It introduces a new assessment network for video editing that attains superior performance in alignment with human preferences.To the best of our knowledge, VE-Bench introduces the first quality assessment dataset for video editing and proposes an effective subjective-aligned quantitative metric for this domain. All models, data, and code will be publicly available to the community.
Shangkun Sun, Songlin Fan, Wenxu Gao, Wei Gao 0003
AAAI3
2025 Stochasticity-aware No-Reference Point Cloud Quality Assessment
abstract
The evolution of point cloud processing algorithms necessitates an accurate assessment for their quality. Previous works consistently regard point cloud quality assessment (PCQA) as a MOS regression problem and devise a deterministic mapping, ignoring the stochasticity in generating MOS from subjective tests. This work presents the first probabilistic architecture for no-reference PCQA, motivated by the labeling process of existing datasets. The proposed method can model the quality judging stochasticity of subjects through a tailored conditional variational autoencoder (CVAE) and produces multiple intermediate quality ratings. These intermediate ratings simulate the judgments from different subjects and are then integrated into an accurate quality prediction, mimicking the generation process of a ground truth MOS. Specifically, our method incorporates a Prior Module, a Posterior Module, and a Quality Rating Generator, where the former two modules are introduced to model the judging stochasticity in subjective tests, while the latter is developed to generate diverse quality ratings. Extensive experiments indicate that our approach outperforms previous cutting-edge methods by a large margin and exhibits gratifying crossdataset robustness. Codes are available at https://git.openi.org.cn/OpenPointCloud/nrpcqa.
Songlin Fan, Wei Gao 0003, Zhineng Chen, Ge Li 0002, Qicheng Wang
IJCAI1
2025 Deep Learning-Based Point Cloud Compression: An In-Depth Survey and Benchmark
abstract
With the maturity of 3D capture technology, the explosive growth of point cloud data has burdened the storage and transmission process. Traditional hybrid point cloud compression (PCC) tools relying on handcrafted priors have limited compression performance and are increasingly weak in addressing the burden induced by data growth. Recently, deep learning-based PCC methods have been introduced to continue to push the PCC performance boundary. With the thriving of deep PCC, the community urgently demands a systematic overview to conclude the past progress and present future research directions. In this paper, we have a detailed review that covers popular point cloud datasets, algorithm evolution, benchmarking analysis, and future trends. Concretely, we first introduce several widely-used PCC datasets according to their major properties. Then the algorithm evolution of existing studies on deep PCC, including lossy ones and lossless ones proposed for various point cloud types, is reviewed. Apart from academic studies, we also investigate the development of relevant international standards (i.e., MPEG standards and JPEG standards). To help have an in-depth understanding of the advance of deep PCC, we select a representative set of methods and conduct extensive experiments on multiple datasets. Comprehensive benchmarking comparisons and analysis reveal the pros and cons of previous methods. Finally, based on the profound analysis, we highlight the challenges and future trends of deep learning-based PCC, paving the way for further study.
Wei Gao 0003, Liang Xie 0013, Songlin Fan, Ge Li 0002, Shan Liu 0001, Wen Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Point-MPP: Point Cloud Self-Supervised Learning From Masked Position Prediction
abstract
Masked autoencoding has gained momentum for improving fine-tuning performance in many downstream tasks. However, it tends to focus on low-level reconstruction details, lacking high-level semantics and resulting in weak transfer capability. This article presents a novel jigsaw puzzle solver inspired by the idea that predicting the positions of disordered point cloud patches provides more semantic information, similar to how children learn by solving jigsaw puzzles. Our method adopts the mask-then-predict paradigm, erasing the positions of selected point patches rather than their contents. We first partition input point clouds into irregular patches and randomly erase the positions of some patches. Then, a Transformer-based model is used to learn high-level semantic features and regress the positions of the masked patches. This approach forces the model to focus on learning transfer-robust semantics while paying less attention to low-level details. To tie the predictions within the encoding space, we further introduce a consistency constraint on their latent representations to encourage the encoded features to contain more semantic cues. We demonstrate that a standard Transformer backbone with our pretraining scheme can capture discriminative point cloud semantic information. Furthermore, extensive experiments indicate that our method outperforms the previous best competitor across six popular downstream vision tasks, achieving new state-of-the-art performance. Codes will be available at https://git.openi.org.cn/OpenPointCloud/Point-MPP.
Songlin Fan, Wei Gao 0003, Ge Li 0002
IEEE Trans. Neural Networks Learn. Syst.1
2024 PDNet: Parallel Dual-branch Network for Point Cloud Geometry Compression and Analysis
abstract
Integrating compression with analysis for point clouds poses a formidable challenge due to the inherent tension between the primary goals of compression for a compact representation and analysis for rich semantic retention. To alleviate this gap and maximize the practical requirements, we introduce a Parallel Dual-branch Network (PDNet) for lossy point cloud geometry compression, whose outputs are also analysis-friendly. The proposed method uses a novel Transformer-based encoder-decoder framework to incorporate local and global attention for point cloud latent representation computation. Specifically, the encoder comprises a Multi-scale Local-Global Feature extraction (MLGF) block to capture compact local and global latent features. The decoding and the hyper-prior modules employ a Transformer with No Position Embedding (TNPE) block and a Multilayer Perceptron (MLP) layer to reconstruct point clouds accurately. Furthermore, our method allows simultaneous point cloud analysis based on the compressed bitstream, such as point cloud classification. Experimental results demonstrate that our PDNet achieves nearly a 40% BD-Rate gain compared to G-PCC and other point-based compression counterparts. Besides, a 26% accuracy improvement in instance classification is observed compared to reconstructed point cloud classification.
Liang Xie 0013, Wei Gao 0003, Songlin Fan, Zhaojian Yao
DCC3
2023 Screen-based 3D Subjective Experiment Software
abstract
Recently, widespread 3D graphics (e.g., point clouds and meshes) have drawn considerable efforts from academia and industry to assess their perceptual quality by conducting subjective experiments. However, lacking a handy software for 3D subjective experiments complicates the construction of 3D graphics quality assessment datasets, thus hindering the prosperity of relevant fields. In this paper, we develop a powerful platform with which users can flexibly design their 3D subjective methodologies and build high-quality datasets, easing a broad spectrum of 3D graphics subjective quality study. To accurately illustrate the perceptual quality differences of 3D stimuli, our software can simultaneously render the source stimulus and impaired stimulus and allows both stimuli to respond synchronously to viewer interactions. Compared with amateur 3D visualization tool-based or image/video rendering-based schemes, our approach embodies typical 3D applications while minimizing cognitive overload during subjective experiments. We organized a subjective experiment involving 40 participants to verify the validity of the proposed software. Experimental analyses demonstrate that subjective tests on our software can produce reasonable subjective quality scores of 3D models. All resources in this paper can be found at https://openi.pcl.ac.cn/OpenDatasets/3DQA.
Songlin Fan, Wei Gao 0003
ACM Multimedia1
2023 A Thorough Benchmark and a New Model for Light Field Saliency Detection
abstract
Compared with current RGB or RGB-D saliency detection datasets, those for light field saliency detection often suffer from many defects, e.g., insufficient data amount and diversity, incomplete data formats, and rough annotations, thus impeding the prosperity of this field. To settle these issues, we elaborately build a large-scale light field dataset, dubbed PKU-LF, comprising 5,000 light fields and covering diverse indoor and outdoor scenes. Our PKU-LF provides all-inclusive representation formats of light fields and offers a unified platform for comparing algorithms utilizing different input formats. For sparking new vitality in saliency detection tasks, we present many unexplored scenarios (such as underwater and high-resolution scenes) and the richest annotations (such as scribble annotations, bounding boxes, object-/instance-level annotations, and edge annotations), on which many potential attention modeling tasks can be investigated. To facilitate the development of saliency detection, we systematically evaluate and analyze 16 representative 2D, 3D, and 4D methods on four existing datasets and the proposed dataset, furnishing a thorough benchmark. Furthermore, tailored to the distinct structural characteristics of light fields, a novel symmetric two-stream architecture (STSA) network is proposed to predict the saliency of light fields more accurately. Specifically, our STSA incorporates a focalness interweavement module (FIM) and three partial decoder modules (PDM). The former is designed to efficiently establish long-range dependencies across focal slices, while the latter aims to effectively aggregate the extracted coadjutant features in a mutual-enhancement way. Extensive experiments demonstrate that our method can significantly outperform the competitors.
Wei Gao 0003, Songlin Fan, Ge Li 0002, Weisi Lin
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Salient Object Detection for Point Clouds
Songlin Fan, Wei Gao 0003, Ge Li 0002
ECCV (28)1
2021 No-Reference Deep Quality Assessment of Compressed Light Field Images
abstract
Unlike traditional 2D image quality assessment, the structural relationship among sub-aperture images (SAIs) is an essential factor affecting the quality evaluation of light field (LF) images, where the labeled datasets are also not sufficient for improving learning performances. To solve these problems, we present a novel deep neural network-based approach to accurately predict the quality of compressed LF images without pristine images. Two modules dubbed SAI-Fusion and Global Context Perception (GCP) are proposed to obtain the relationship among SAIs. For effective training, we compress LF images from EPFL and HCI datasets and propose a ranking-based method to generate pseudo-labels as equivalents of Mean Opinion Score (MOS), i.e., Ranking-MOS. Therefore, we can pre-train our quality assessment network on compressed LF images with Ranking-MOS, and then fine-tune the model at small-scale datasets with real labels. Experiments demonstrate that the proposed method achieves state-of-the-art performance on compressed LF images of Win5-LID dataset.
Zixuan Guo 0002, Wei Gao 0003, Haiqiang Wang, Junle Wang, Songlin Fan
ICME5