VLDB 2026 Research / reviewers in the wild / expert
Henrique Morimitsu
dblp:01/8671
· DBLP profile ↗
7ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0001-9455-8571ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
3D vision · 42% Efficient and distributed learning · 32% Vision and language · 18% | |
| Computer graphics and multimedia
2 papers |
Image and video processing · 100% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › motion estimation
optical flow |
1.5 | 2 | 2024 | RAPIDFlow: Recurrent Adaptable Pyramids with Iterative Decoding for Efficient Optical Flow Estimation · ICRA 2024 Recurrent Partial Kernel Network for Efficient Optical Flow Estimation · AAAI 2024 |
Computer vision › 3D vision › motion estimation
optical flow |
0.9 | 1 | 2025 | DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid Framework · CVPR 2025 |
Machine learning › Efficient and distributed learning › on-device inference
embedded inference |
0.8 | 1 | 2024 | RAPIDFlow: Recurrent Adaptable Pyramids with Iterative Decoding for Efficient Optical Flow Estimation · ICRA 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | RAPIDFlow: Recurrent Adaptable Pyramids with Iterative Decoding for Efficient Optical Flow Estimation · ICRA 2024 |
Computer vision › Vision and language
image-text retrieval |
0.7 | 1 | 2023 | CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023 |
Computer vision › Vision and language
vision-language pretraining |
0.7 | 1 | 2023 | CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023 |
Computer vision › 3D vision
3d reconstruction |
0.4 | 1 | 2020 | A Unified Framework for Piecewise Semantic Reconstruction in Dynamic Scenes via Exploiting Superpixel Relations · ICRA 2020 |
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction |
0.4 | 1 | 2020 | A Unified Framework for Piecewise Semantic Reconstruction in Dynamic Scenes via Exploiting Superpixel Relations · ICRA 2020 |
Computer vision › Segmentation and scene understanding › instance segmentation
semantic instance segmentation |
0.4 | 1 | 2020 | A Unified Framework for Piecewise Semantic Reconstruction in Dynamic Scenes via Exploiting Superpixel Relations · ICRA 2020 |
Computer vision › 3D vision
structure from motion |
0.4 | 1 | 2020 | A Unified Framework for Piecewise Semantic Reconstruction in Dynamic Scenes via Exploiting Superpixel Relations · ICRA 2020 |
Computer vision › 3D vision
depth estimation |
0.4 | 1 | 2019 | Monocular Piecewise Depth Estimation in Dynamic Scenes by Exploiting Superpixel Relations · ICCV 2019 |
Computer vision › 3D vision › depth estimation › video depth estimation
dynamic scene depth estimation |
0.4 | 1 | 2019 | Monocular Piecewise Depth Estimation in Dynamic Scenes by Exploiting Superpixel Relations · ICCV 2019 |
Computer vision › 3D vision › depth estimation
monocular depth estimation |
0.4 | 1 | 2019 | Monocular Piecewise Depth Estimation in Dynamic Scenes by Exploiting Superpixel Relations · ICCV 2019 |
Machine learning › Efficient and distributed learning › model compression
lightweight neural network |
0.2 | 1 | 2024 | Recurrent Partial Kernel Network for Efficient Optical Flow Estimation · AAAI 2024 |
Computer vision › Vision and language
image captioning |
0.2 | 1 | 2023 | CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.2 | 1 | 2023 | CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023 |
Methods — techniques the papers use, named apart from their topics
separable large kernels · 1.5recurrent neural network · 1.5recurrent network · 1.5partial kernel convolution · 1.5feature pyramid · 1.51d convolution · 1.5dual-pyramid architecture · 0.9adaptive downscaling · 0.9pre-ranking and ranking · 0.7knowledge distillation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid FrameworkabstractOptical flow estimation is essential for video processing tasks, such as restoration and action recognition. The quality of videos is constantly increasing, with current standards reaching 8K resolution. However, optical flow methods are usually designed for low resolution and do not generalize to large inputs due to their rigid architectures. They adopt downscaling or input tiling to reduce the input size, causing a loss of details and global information. There is also a lack of optical flow benchmarks to judge the actual performance of existing methods on high-resolution samples. Previous works only conducted qualitative high-resolution evaluations on hand-picked samples. This paper fills this gap in optical flow estimation in two ways. We propose DPFlow, an adaptive optical flow architecture capable of generalizing up to 8K resolution inputs while trained with only low-resolution samples. We also introduce Kubric-NK, a new benchmark for evaluating optical flow methods with input resolutions ranging from 1K to 8K. Our high-resolution evaluation pushes the boundaries of existing methods and reveals new insights about their generalization capabilities. Extensive experimental results show that DPFlow achieves state-of-the-art results on the MPI-Sintel, KITTI 2015, Spring, and other high-resolution benchmarks. The code and dataset are available at https://github.com/hmorimitsu/ptlflow/tree/main/ptlflow/models/dpflow. Henrique Morimitsu, Xiaobin Zhu 0001, Roberto Marcondes Cesar Junior, Xiangyang Ji, Xu-Cheng Yin |
CVPR | 1 |
| 2024 | Recurrent Partial Kernel Network for Efficient Optical Flow EstimationabstractOptical flow estimation is a challenging task consisting of predicting per-pixel motion vectors between images. Recent methods have employed larger and more complex models to improve the estimation accuracy. However, this impacts the widespread adoption of optical flow methods and makes it harder to train more general models since the optical flow data is hard to obtain. This paper proposes a small and efficient model for optical flow estimation. We design a new spatial recurrent encoder that extracts discriminative features at a significantly reduced size. Unlike standard recurrent units, we utilize Partial Kernel Convolution (PKConv) layers to produce variable multi-scale features with a single shared block. We also design efficient Separable Large Kernels (SLK) to capture large context information with low computational cost. Experiments on public benchmarks show that we achieve state-of-the-art generalization performance while requiring significantly fewer parameters and memory than competing methods. Our model ranks first in the Spring benchmark without finetuning, improving the results by over 10% while requiring an order of magnitude fewer FLOPs and over four times less memory than the following published method without finetuning. The code is available at github.com/hmorimitsu/ptlflow/tree/main/ptlflow/models/rpknet. Henrique Morimitsu, Xiaobin Zhu 0001, Xiangyang Ji, Xu-Cheng Yin |
AAAI | 1 |
| 2024 | RAPIDFlow: Recurrent Adaptable Pyramids with Iterative Decoding for Efficient Optical Flow EstimationabstractExtracting motion information from videos with optical flow estimation is vital in multiple practical robot applications. Current optical flow approaches show remarkable accuracy, but top-performing methods have high computational costs and are unsuitable for embedded devices. Although some previous works have focused on developing low-cost optical flow strategies, their estimation quality has a noticeable gap with more robust methods. In this paper, we develop a novel method to efficiently estimate high-quality optical flow in embedded devices. Our proposed RAPIDFlow model combines efficient NeXt1D convolution blocks with a fully recurrent structure based on feature pyramids to decrease computational costs without significantly impacting estimation accuracy. The adaptable recurrent encoder produces multi-scale features with a single shared block, which allows us to adjust the pyramid length at inference time and make it more robust to changes in input size. Also, it enables our model to offer multiple tradeoffs between accuracy and speed to suit different applications. Experiments using a Jetson Orin NX embedded system on the MPI-Sintel and KITTI public benchmarks show that RAPIDFlow outperforms previous approaches by significant margins at faster speeds. Our code is available at https://github.com/hmorimitsu/ptlflow/tree/main/ptlflow/models/rapidflow. Henrique Morimitsu, Xiaobin Zhu 0001, Roberto Marcondes Cesar Junior, Xiangyang Ji, Xu-Cheng Yin |
ICRA | 1 |
| 2023 | CCMB: A Large-scale Chinese Cross-modal BenchmarkabstractVision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets and downstream datasets with Chinese corpus remain largely unexplored. In this work, we build a large-scale high-quality Chinese Cross-Modal Benchmark named CCMB for the research community, which contains the currently largest public pre-training dataset Zero and five human-annotated fine-tuning datasets for downstream tasks. Zero contains 250 million images paired with 750 million text descriptions, plus two of the five fine-tuning datasets are also currently the largest ones for Chinese cross-modal downstream tasks. Along with the CCMB, we also develop a VLP framework named R2D2, applying a pre-Ranking + Ranking strategy to learn powerful vision-language representations and a two-way distillation method (i.e., target-guided Distillation and feature-guided Distillation) to further enhance the learning capability. With the Zero and the R2D2 VLP framework, we achieve state-of-the-art performance on twelve downstream datasets from five broad categories of tasks including image-text retrieval, image-text matching, image caption, text-to-image generation, and zero-shot image classification. The datasets, models, and codes are available at https://github.com/yuxie11/R2D2 Chunyu Xie, Heng Cai, Jincheng Li 0002, Fanjing Kong, Jianfei Song, Henrique Morimitsu, Lin Yao 0003, Xiangzheng Zhang, Dawei Leng, Baochang Zhang 0001, Xiangyang Ji, Yafeng Deng |
ACM Multimedia | 7 |
| 2020 | A Unified Framework for Piecewise Semantic Reconstruction in Dynamic Scenes via Exploiting Superpixel RelationsabstractThis paper presents a novel framework for dense piecewise semantic reconstruction in dynamic scenes containing complex background and moving objects via exploiting superpixel relations. We utilize two kinds of superpixel relations: motion relations and spatial relations, each having three subcategories: coplanar, hinge, and crack. Spatial relations provide constraints on the spatial locations of neighboring superpixels and thus can be used to reconstruct dynamic scenes. However, spatial relations can not be estimated directly with epipolar geometry due to moving objects in dynamic scenes. We synthesize the results of semantic instance segmentation and motion relations to estimate spatial relations. Given consecutive frames, we mainly develop our method in five main stages: preprocessing, motion estimation, superpixel relation analysis, reconstruction and refinement. Extensive experiments on various datasets demonstrate that our method outperforms competitors in reconstruction quality. Furthermore, our method presents a feasible way to incorporate semantic information in Structure-from-Motion (SFM) based reconstruction pipelines. Yan Di, Henrique Morimitsu, Zhiqiang Lou, Xiangyang Ji |
ICRA | 2 |
| 2019 | Monocular Piecewise Depth Estimation in Dynamic Scenes by Exploiting Superpixel RelationsabstractIn this paper, we propose a novel and specially designed method for piecewise dense monocular depth estimation in dynamic scenes. We utilize spatial relations between neighboring superpixels to solve the inherent relative scale ambiguity (RSA) problem and smooth the depth map. However, directly estimating spatial relations is an ill-posed problem. Our core idea is to predict spatial relations based on the corresponding motion relations. Given two or more consecutive frames, we first compute semi-dense (CPM) or dense (optical flow) point matches between temporally neighboring images. Then we develop our method in four main stages: superpixel relations analysis, motion selection, reconstruction, and refinement. The final refinement process helps to improve the quality of the reconstruction at pixel level. Our method does not require per-object segmentation, template priors or training sets, which ensures flexibility in various applications. Extensive experiments on both synthetic and real datasets demonstrate that our method robustly handles different dynamic situations and presents competitive results to the state-of-the-art methods while running much faster than them. Henrique Morimitsu, Shan Gao 0003, Xiangyang Ji |
ICCV | 2 |
| 2017 | Exploring structure for long-term tracking of multiple objects in sports videos
Henrique Morimitsu, Isabelle Bloch, Roberto Marcondes Cesar Junior |
Comput. Vis. Image Underst. | 1 |