Hongtao Fu

dblp:324/8795 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-6692-0913ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Segmentation and scene understanding · 35% Deep learning architectures and training · 31% Learning paradigms · 19%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
dense prediction
1.222023
Learning to Upsample by Learning to Sample · ICCV 2023
SAPA: Similarity-Aware Point Affiliation for Feature Upsampling · NeurIPS 2022
Machine learning › Learning paradigms › data balancing
oversampling
1.222023
Learning to Upsample by Learning to Sample · ICCV 2023
FADE: Fusing the Assets of Decoder and Encoder for Task-Agnostic Upsampling · ECCV (27) 2022
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling
0.712023
Masked Image Residual Learning for Scaling Deeper Vision Transformers · NeurIPS 2023
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.712023
Masked Image Residual Learning for Scaling Deeper Vision Transformers · NeurIPS 2023
Machine learning › Deep learning architectures and training › neural network layer design
feature upsampling
0.612022
SAPA: Similarity-Aware Point Affiliation for Feature Upsampling · NeurIPS 2022
Machine learning › Deep learning architectures and training
encoder-decoder architecture
0.312025
FADE: A Task-Agnostic Upsampling Operator for Encoder-Decoder Architectures · Int. J. Comput. Vis. 2025
Computer vision › Image recognition and object detection
object detection
0.212023
Masked Image Residual Learning for Scaling Deeper Vision Transformers · NeurIPS 2023
Computer vision › Segmentation and scene understanding
semantic segmentation
0.212023
Masked Image Residual Learning for Scaling Deeper Vision Transformers · NeurIPS 2023
Computer vision › 3D vision
depth estimation
0.212022
SAPA: Similarity-Aware Point Affiliation for Feature Upsampling · NeurIPS 2022
Image and video processing
image matting
0.212022
SAPA: Similarity-Aware Point Affiliation for Feature Upsampling · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

task-agnostic upsampling · 1.4similarity-aware kernel generation · 1.1point affiliation · 1.1self-supervised pretraining · 0.7residual learning · 0.7point sampling · 0.7dynamic convolution · 0.7
YearPublicationVenuePosition
2025 FADE: A Task-Agnostic Upsampling Operator for Encoder-Decoder Architectures
Hao Lu 0003, Wenze Liu, Hongtao Fu, Zhiguo Cao 0001
Int. J. Comput. Vis.3
2023 Learning to Upsample by Learning to Sample
abstract
We present DySample, an ultra-lightweight and effective dynamic upsampler. While impressive performance gains have been witnessed from recent kernel-based dynamic upsamplers such as CARAFE, FADE, and SAPA, they introduce much workload, mostly due to the time-consuming dynamic convolution and the additional sub-network used to generate dynamic kernels. Further, the need for high-res feature guidance of FADE and SAPA somehow limits their application scenarios. To address these concerns, we bypass dynamic convolution and formulate upsampling from the perspective of point sampling, which is more resource-efficient and can be easily implemented with the standard built-in function in PyTorch. We first showcase a naive design, and then demonstrate how to strengthen its upsampling behavior step by step towards our new upsampler, DySample. Compared with former kernel-based dynamic upsamplers, DySample requires no customized CUDA package and has much fewer parameters, FLOPs, GPU memory, and latency. Besides the light-weight characteristics, DySample outperforms other upsamplers across five dense prediction tasks, including semantic segmentation, object detection, instance segmentation, panoptic segmentation, and monocular depth estimation. Code is available at https://github.com/tiny-smart/dysample.
Wenze Liu, Hao Lu 0003, Hongtao Fu, Zhiguo Cao 0001
ICCV3
2023 Masked Image Residual Learning for Scaling Deeper Vision Transformers
abstract
Deeper Vision Transformers (ViTs) are more challenging to train. We expose a degradation problem in deeper layers of ViT when using masked image modeling (MIM) for pre-training. To ease the training of deeper ViTs, we introduce a self-supervised learning framework called $\textbf{M}$asked $\textbf{I}$mage $\textbf{R}$esidual $\textbf{L}$earning ($\textbf{MIRL}$), which significantly alleviates the degradation problem, making scaling ViT along depth a promising direction for performance upgrade. We reformulate the pre-training objective for deeper layers of ViT as learning to recover the residual of the masked image. We provide extensive empirical evidence showing that deeper ViTs can be effectively optimized using MIRL and easily gain accuracy from increased depth. With the same level of computational complexity as ViT-Base and ViT-Large, we instantiate $4.5{\times}$ and $2{\times}$ deeper ViTs, dubbed ViT-S-54 and ViT-B-48. The deeper ViT-S-54, costing $3{\times}$ less than ViT-Large, achieves performance on par with ViT-Large. ViT-B-48 achieves 86.2\% top-1 accuracy on ImageNet. On one hand, deeper ViTs pre-trained with MIRL exhibit excellent generalization capabilities on downstream tasks, such as object detection and semantic segmentation. On the other hand, MIRL demonstrates high pre-training efficiency. With less pre-training time, MIRL yields competitive performance compared to other approaches.
Guoxi Huang, Hongtao Fu, Adrian G. Bors
NeurIPS2
2023 SIERRA: A robust bilateral feature upsampler for dense prediction
Hongtao Fu, Wenze Liu, Zhiguo Cao 0001, Hao Lu 0003
Comput. Vis. Image Underst.1
2022 FADE: Fusing the Assets of Decoder and Encoder for Task-Agnostic Upsampling
Hao Lu 0003, Wenze Liu, Hongtao Fu, Zhiguo Cao 0001
ECCV (27)3
2022 SAPA: Similarity-Aware Point Affiliation for Feature Upsampling
abstract
We introduce point affiliation into feature upsampling, a notion that describes the affiliation of each upsampled point to a semantic cluster formed by local decoder feature points with semantic similarity. By rethinking point affiliation, we present a generic formulation for generating upsampling kernels. The kernels encourage not only semantic smoothness but also boundary sharpness in the upsampled feature maps. Such properties are particularly useful for some dense prediction tasks such as semantic segmentation. The key idea of our formulation is to generate similarity-aware kernels by comparing the similarity between each encoder feature point and the spatially associated local region of decoder features. In this way, the encoder feature point can function as a cue to inform the semantic cluster of upsampled feature points. To embody the formulation, we further instantiate a lightweight upsampling operator, termed Similarity-Aware Point Affiliation (SAPA), and investigate its variants. SAPA invites consistent performance improvements on a number of dense prediction tasks, including semantic segmentation, object detection, depth estimation, and image matting. Code is available at: https://github.com/poppinace/sapa
Hao Lu 0003, Wenze Liu, Zixuan Ye, Hongtao Fu, Zhiguo Cao 0001
NeurIPS4