Linlong Fan

dblp:355/7878 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-6350-7875ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 58% 3D vision · 25% Segmentation and scene understanding · 9%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model › image restoration
diffusion-based image restoration
0.912025
Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders · NeurIPS 2025
Image and video processing › super-resolution
image super-resolution
0.912025
Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders · NeurIPS 2025
Image and video processing › super-resolution › image super-resolution
real-world image super-resolution
0.912025
Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders · NeurIPS 2025
Computer vision › 3D vision
3d object recognition
0.812024
Beyond Viewpoint: Robust 3D Object Recognition Under Arbitrary Views Through Joint Multi-part Representation · ECCV (52) 2024
Computer vision › Segmentation and scene understanding › image segmentation
scene text segmentation
0.312025
Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

text-aware attention · 1.7joint segmentation decoders · 1.7multi-part representation learning · 0.8
YearPublicationVenuePosition
2025 Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders
abstract
The introduction of generative models has significantly advanced image super-resolution (SR) in handling real-world degradations. However, they often incur fidelity-related issues, particularly distorting textual structures. In this paper, we introduce a novel diffusion-based SR framework, namely TADiSR, which integrates text-aware attention and joint segmentation decoders to recover not only natural details but also the structural fidelity of text regions in degraded real-world images. Moreover, we propose a complete pipeline for synthesizing high-quality images with fine-grained full-image text masks, combining realistic foreground text regions with detailed background content. Extensive experiments demonstrate that our approach substantially enhances text legibility in super-resolved images, achieving state-of-the-art performance across multiple evaluation metrics and exhibiting strong generalization to real-world scenarios. Our code is available at [here](https://github.com/mingcv/TADiSR).
Linlong Fan, Yiyan Luo, Yuhang Yu, Xiaojie Guo 0001, Qingnan Fan
NeurIPS2
2024 Beyond Viewpoint: Robust 3D Object Recognition Under Arbitrary Views Through Joint Multi-part Representation
Linlong Fan, Yanqi Ge, Wen Li 0001, Lixin Duan
ECCV (52)1
2023 Multi-View Token Clustering and Fusion for 3D Object Recognition and Retrieval
abstract
3D object recognition has received extensive attention in recent years. Many existing methods tackle the task by rendering 3D objects from multiple views. However, most multi-view recognition methods do not utilize fine-grained information from different views, which is found to be crucial for improving 3D object representation in the multi-view setting. In this paper, we propose a transformer-based method, referred to as MVCFormer, for multi-view feature clustering and fusion. MVCFormer clusters semantically similar tokens at the same stages and selects representative fine-grained features, which helps to eliminate feature redundancy and remove cluttered backgrounds and make the selected features more diverse. On the other hand, our model also integrates selected features from all stages to obtain a discriminative 3D object representation by a cross-attention fusion method. Extensive experiments on benchmark datasets (e.g., ModelNet40, ModelNet10, ShapeNetCore55, and RGBD) clearly demonstrate the effectiveness of our proposed MVCFormer over existing baselines.
Linlong Fan, Yanqi Ge, Wen Li 0001, Lixin Duan
ICME1