VLDB 2026 Research / reviewers in the wild / expert
Hongtao Fu
dblp:324/8795
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-6692-0913ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Segmentation and scene understanding · 35% Deep learning architectures and training · 31% Learning paradigms · 19% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 10 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
dense prediction |
1.2 | 2 | 2023 | Learning to Upsample by Learning to Sample · ICCV 2023 SAPA: Similarity-Aware Point Affiliation for Feature Upsampling · NeurIPS 2022 |
Machine learning › Learning paradigms › data balancing
oversampling |
1.2 | 2 | 2023 | Learning to Upsample by Learning to Sample · ICCV 2023 FADE: Fusing the Assets of Decoder and Encoder for Task-Agnostic Upsampling · ECCV (27) 2022 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling |
0.7 | 1 | 2023 | Masked Image Residual Learning for Scaling Deeper Vision Transformers · NeurIPS 2023 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.7 | 1 | 2023 | Masked Image Residual Learning for Scaling Deeper Vision Transformers · NeurIPS 2023 |
Machine learning › Deep learning architectures and training › neural network layer design
feature upsampling |
0.6 | 1 | 2022 | SAPA: Similarity-Aware Point Affiliation for Feature Upsampling · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
encoder-decoder architecture |
0.3 | 1 | 2025 | FADE: A Task-Agnostic Upsampling Operator for Encoder-Decoder Architectures · Int. J. Comput. Vis. 2025 |
Computer vision › Image recognition and object detection
object detection |
0.2 | 1 | 2023 | Masked Image Residual Learning for Scaling Deeper Vision Transformers · NeurIPS 2023 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.2 | 1 | 2023 | Masked Image Residual Learning for Scaling Deeper Vision Transformers · NeurIPS 2023 |
Computer vision › 3D vision
depth estimation |
0.2 | 1 | 2022 | SAPA: Similarity-Aware Point Affiliation for Feature Upsampling · NeurIPS 2022 |
Image and video processing
image matting |
0.2 | 1 | 2022 | SAPA: Similarity-Aware Point Affiliation for Feature Upsampling · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
task-agnostic upsampling · 1.4similarity-aware kernel generation · 1.1point affiliation · 1.1self-supervised pretraining · 0.7residual learning · 0.7point sampling · 0.7dynamic convolution · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FADE: A Task-Agnostic Upsampling Operator for Encoder-Decoder Architectures
Hao Lu 0003, Wenze Liu, Hongtao Fu, Zhiguo Cao 0001 |
Int. J. Comput. Vis. | 3 |
| 2023 | Learning to Upsample by Learning to SampleabstractWe present DySample, an ultra-lightweight and effective dynamic upsampler. While impressive performance gains have been witnessed from recent kernel-based dynamic upsamplers such as CARAFE, FADE, and SAPA, they introduce much workload, mostly due to the time-consuming dynamic convolution and the additional sub-network used to generate dynamic kernels. Further, the need for high-res feature guidance of FADE and SAPA somehow limits their application scenarios. To address these concerns, we bypass dynamic convolution and formulate upsampling from the perspective of point sampling, which is more resource-efficient and can be easily implemented with the standard built-in function in PyTorch. We first showcase a naive design, and then demonstrate how to strengthen its upsampling behavior step by step towards our new upsampler, DySample. Compared with former kernel-based dynamic upsamplers, DySample requires no customized CUDA package and has much fewer parameters, FLOPs, GPU memory, and latency. Besides the light-weight characteristics, DySample outperforms other upsamplers across five dense prediction tasks, including semantic segmentation, object detection, instance segmentation, panoptic segmentation, and monocular depth estimation. Code is available at https://github.com/tiny-smart/dysample. Wenze Liu, Hao Lu 0003, Hongtao Fu, Zhiguo Cao 0001 |
ICCV | 3 |
| 2023 | Masked Image Residual Learning for Scaling Deeper Vision TransformersabstractDeeper Vision Transformers (ViTs) are more challenging to train. We expose a degradation problem in deeper layers of ViT when using masked image modeling (MIM) for pre-training.
To ease the training of deeper ViTs, we introduce a self-supervised learning framework called $\textbf{M}$asked $\textbf{I}$mage $\textbf{R}$esidual $\textbf{L}$earning ($\textbf{MIRL}$), which significantly alleviates the degradation problem, making scaling ViT along depth a promising direction for performance upgrade. We reformulate the pre-training objective for deeper layers of ViT as learning to recover the residual of the masked image.
We provide extensive empirical evidence showing that deeper ViTs can be effectively optimized using MIRL and easily gain accuracy from increased depth.
With the same level of computational complexity as ViT-Base and ViT-Large, we instantiate $4.5{\times}$ and $2{\times}$ deeper ViTs, dubbed ViT-S-54 and ViT-B-48.
The deeper ViT-S-54, costing $3{\times}$ less than ViT-Large, achieves performance on par with ViT-Large.
ViT-B-48 achieves 86.2\% top-1 accuracy on ImageNet.
On one hand, deeper ViTs pre-trained with MIRL exhibit excellent generalization capabilities on downstream tasks, such as object detection and semantic segmentation. On the other hand, MIRL demonstrates high pre-training efficiency. With less pre-training time, MIRL yields competitive performance compared to other approaches. Guoxi Huang, Hongtao Fu, Adrian G. Bors |
NeurIPS | 2 |
| 2023 | SIERRA: A robust bilateral feature upsampler for dense prediction
Hongtao Fu, Wenze Liu, Zhiguo Cao 0001, Hao Lu 0003 |
Comput. Vis. Image Underst. | 1 |
| 2022 | FADE: Fusing the Assets of Decoder and Encoder for Task-Agnostic Upsampling
Hao Lu 0003, Wenze Liu, Hongtao Fu, Zhiguo Cao 0001 |
ECCV (27) | 3 |
| 2022 | SAPA: Similarity-Aware Point Affiliation for Feature UpsamplingabstractWe introduce point affiliation into feature upsampling, a notion that describes the affiliation of each upsampled point to a semantic cluster formed by local decoder feature points with semantic similarity. By rethinking point affiliation, we present a generic formulation for generating upsampling kernels. The kernels encourage not only semantic smoothness but also boundary sharpness in the upsampled feature maps. Such properties are particularly useful for some dense prediction tasks such as semantic segmentation. The key idea of our formulation is to generate similarity-aware kernels by comparing the similarity between each encoder feature point and the spatially associated local region of decoder features. In this way, the encoder feature point can function as a cue to inform the semantic cluster of upsampled feature points. To embody the formulation, we further instantiate a lightweight upsampling operator, termed Similarity-Aware Point Affiliation (SAPA), and investigate its variants. SAPA invites consistent performance improvements on a number of dense prediction tasks, including semantic segmentation, object detection, depth estimation, and image matting. Code is available at: https://github.com/poppinace/sapa Hao Lu 0003, Wenze Liu, Zixuan Ye, Hongtao Fu, Zhiguo Cao 0001 |
NeurIPS | 4 |