Baolin Liu 0002

dblp:71/2948-2 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0007-1313-7486ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Rendering · 56% Virtual and augmented reality · 27% Image and video processing · 16%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Virtual and augmented reality
3d display
0.912025
EYE3: Turn Anything into Naked-Eye 3D · ICCV 2025
Rendering › image-based rendering
light field display rendering
0.812024
DirectL: Efficient Radiance Fields Rendering for 3D Light Field Displays · ACM Trans. Graph. 2024
Rendering
neural rendering
0.812024
DirectL: Efficient Radiance Fields Rendering for 3D Light Field Displays · ACM Trans. Graph. 2024
Rendering › neural rendering
radiance field rendering
0.812024
DirectL: Efficient Radiance Fields Rendering for 3D Light Field Displays · ACM Trans. Graph. 2024
Machine learning › Generative modeling
diffusion model
0.712023
DocDiff: Document Enhancement via Residual Diffusion Models · ACM Multimedia 2023
Machine learning › Generative modeling › diffusion model › diffusion model architecture
residual diffusion model
0.712023
DocDiff: Document Enhancement via Residual Diffusion Models · ACM Multimedia 2023
Image and video processing › image enhancement
document image enhancement
0.712023
DocDiff: Document Enhancement via Residual Diffusion Models · ACM Multimedia 2023
Virtual and augmented reality › 3d display › stereoscopic display
autostereoscopic display
0.212024
DirectL: Efficient Radiance Fields Rendering for 3D Light Field Displays · ACM Trans. Graph. 2024

Methods — techniques the papers use, named apart from their topics

residual refinement · 1.3diffusion model · 1.3subpixel repurposing · 0.8optimized rendering pipeline · 0.8interleaved ray mapping · 0.8
YearPublicationVenuePosition
2026 ELF: Edit anything for light field displays
Baolin Liu 0002, Zongyuan Yang, Yingde Song, Yongping Xiong
Pattern Recognit.1
2025 EYE3: Turn Anything into Naked-Eye 3D
Yingde Song, Zongyuan Yang, Baolin Liu 0002, Yongping Xiong, Sai Chen, Lan Yi, Zhaohe Zhang, Xunbo Yu
ICCV3
2025 TextDiff: Enhancing scene text image super-resolution with mask-guided residual diffusion models
Baolin Liu 0002, Zongyuan Yang, Chinwai Chiu, Yongping Xiong
Pattern Recognit.1
2024 GDB: Gated Convolutions-based Document Binarization
Zongyuan Yang, Baolin Liu 0002, Yongping Xiong, Guibin Wu
Pattern Recognit.2
2024 FAT: Field-Aware Transformer for Point Cloud Segmentation With Adaptive Attention Fields
abstract
Point cloud segmentation is crucial for various industrial applications, such as autonomous driving and robotics. Recent developments underscore the significant potential of transformer models in this field. However, existing attention mechanisms apply the same feature learning paradigm for all points equally, ignoring the considerable size differences among objects in a scene. To rectify this, we introduce the field-aware transformer (FAT), engineered to tailor effective receptive fields to objects of varying sizes. Our FAT achieves field-aware learning through two primary components: the multigranularity attention (MGA) scheme and the reattention module. The MGA scheme is proficient in aggregating tokens from distant areas while preserving multiscale features within each attention layer. The reattention module dynamically adjusts the attention scores to the fine- and coarse-grained features output by MGA for each point. Extensive experimental results underscore the effectiveness and efficiency of our FAT, which delivers state-of-the-art performance on both the stanford 3D indoor scene dataset (S3DIS) and ScanNetV2 datasets.
Junjie Zhou 0001, Baolin Liu 0002, Yongping Xiong, Chinwai Chiu, Xiangyang Gong
IEEE Trans. Ind. Informatics2
2024 DirectL: Efficient Radiance Fields Rendering for 3D Light Field Displays
abstract
Autostereoscopic display technology, despite decades of development, has not achieved extensive application, primarily due to the daunting challenge of three-dimensional (3D) content creation for non-specialists. The emergence of Radiance Field as an innovative 3D representation has markedly revolutionized the domains of 3D reconstruction and generation, simplifying 3D content creation for common users and broadening the applicability of Light Field Displays (LFDs). However, the combination of these two technologies remains largely unexplored. The standard paradigm to create optimal content for parallax-based light field displays demands rendering at least 45 slightly shifted views preferably at high resolution per frame, a substantial hurdle for real-time rendering. We introduce DirectL, a novel rendering paradigm for Radiance Fields on autostereoscopic displays with lenticular lens. By thoroughly analyzing the interleaved mapping of spatial rays to screen sub-pixels, we accurately render only the light rays entering the human eye and propose subpixel repurposing to significantly reduce the pixel count required for rendering. Tailored for the two predominant radiance fields---Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS), we propose corresponding optimized rendering pipelines that directly render the light field images instead of multi-view images, achieving state-of-the-art rendering speeds on autostereoscopic displays. Extensive experiments across various autostereoscopic displays and user visual perception assessments demonstrate that DirectL accelerates rendering by up to 40 times compared to the standard paradigm without sacrificing visual quality. Its rendering process-only modification allows seamless integration into subsequent radiance field tasks. Finally, we incorporate DirectL into diverse applications, showcasing the stunning visual experiences and the synergy between Light Field Displays and Radiance Fields, which reveals the immense potential for application prospects. DirectL Project Homepage: direct-l.github.io
Zongyuan Yang, Baolin Liu 0002, Yingde Song, Lan Yi, Yongping Xiong, Zhaohe Zhang, Xunbo Yu
ACM Trans. Graph.2
2023 DocDiff: Document Enhancement via Residual Diffusion Models
abstract
Removing degradation from document images not only improves their visual quality and readability, but also enhances the performance of numerous automated document analysis and recognition tasks. However, existing regression-based methods optimized for pixel-level distortion reduction tend to suffer from significant loss of high-frequency information, leading to distorted and blurred text edges. To compensate for this major deficiency, we propose DocDiff, the first diffusion-based framework specifically designed for diverse challenging document enhancement problems, including document deblurring, denoising, and removal of watermarks and seals. DocDiff consists of two modules: the Coarse Predictor (CP), which is responsible for recovering the primary low-frequency content, and the High-Frequency Residual Refinement (HRR) module, which adopts the diffusion models to predict the residual (high-frequency information, including text edges), between the ground-truth and the CP-predicted image. DocDiff is a compact and computationally efficient model that benefits from a well-designed network architecture, an optimized training loss objective, and a deterministic sampling process with short time steps. Extensive experiments demonstrate that DocDiff achieves state-of-the-art (SOTA) performance on multiple benchmark datasets, and can significantly enhance the readability and recognizability of degraded document images. Furthermore, our proposed HRR module in pre-trained DocDiff is plug-and-play and ready-to-use, with only 4.17M parameters. It greatly sharpens the text edges generated by SOTA deblurring methods without additional joint training. Available codes: https://github.com/Royalvice/DocDiff https://github.com/Royalvice/DocDiff.
Zongyuan Yang, Baolin Liu 0002, Yongping Xiong, Lan Yi, Guibin Wu, Junjie Zhou 0001
ACM Multimedia2