VLDB 2026 Research / reviewers in the wild / expert
Azin Jahedi
dblp:324/3043
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-3956-0761ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reviving Unsupervised Optical Flow: Concept Reevaluation, Multi-Scale Advances and Full Open-Source ReleaseabstractUnsupervised optical flow methods have become more popular in the last decade, enabling the training of models across domains without ground truth data. Although RAFT and its successors have achieved significant success in the supervised settings, many unsupervised approaches continue to use older backbones such as PWC-Net. One reason for this architectural stagnation is that the current RAFT-based SOTA approach has proven challenging for the community to reproduce. In this paper, we revive and advance unsupervised optical flow: First, we introduce Sun-RAFT: a simple unsupervised RAFT. Second, building on Sun-RAFT, we present Muun-RAFT: a novel multi-scale unsupervised RAFT, where we propose a gradual context-based upsampling to refine the flow, further improving both accuracy and preservation of details. Third, we reexamine previously advised unsupervised strategies to identify effective training settings. In terms of results, both our methods demonstrate strong generalization capabilities and set a new SOTA for unsupervised two-frame approaches on MPI-Sintel, with Muun-RAFT surpassing even the current multi-frame SOTA by up to 28%. Finally, we open-source our PyTorch code, enabling further developments in the field: https://cv-stuttgart.github.io/Reviving-Unsupervised-OpticalFlow. Azin Jahedi, Marc Rivinius, Noah Berenguel Senn, Andrés Bruhn |
WACV | 1 |
| 2025 | MS-RAFT-3D: A Multi-Scale Architecture for Recurrent Image-Based Scene FlowabstractAlthough multi-scale concepts have recently proven useful for recurrent network architectures in the field of optical flow and stereo, they have not been considered for image-based scene flow so far. Hence, based on a single-scale recurrent scene flow backbone, we develop a multi-scale approach that generalizes successful hierarchical ideas from optical flow to image-based scene flow. By considering suitable concepts for the feature and the context encoder, the overall coarse-to-fine framework and the training loss, we succeed to design a scene flow approach that outperforms the current state of the art on KITTI and Spring by 8.7% (3.89 vs. 4.26) and 65.8% (9.13 vs. 26.71), respectively. Our code is available at https://github.com/cv-stuttgart/MS-RAFT-3D. Jakob Schmid, Azin Jahedi, Noah Berenguel Senn, Andrés Bruhn |
ICIP | 2 |
| 2024 | CCMR: High Resolution Optical Flow Estimation via Coarse-to-Fine Context-Guided Motion ReasoningabstractAttention-based motion aggregation concepts have recently shown their usefulness in optical flow estimation, in particular when it comes to handling occluded regions. However, due to their complexity, such concepts have been mainly restricted to coarse-resolution single-scale approaches that fail to provide the detailed outcome of high-resolution multi-scale networks. In this paper, we hence propose CCMR: a high-resolution coarse-to-fine approach that leverages attention-based motion grouping concepts to multi-scale optical flow estimation. CCMR relies on a hierarchical two-step attention-based context-motion grouping strategy that first computes global multi-scale context features and then uses them to guide the actual motion grouping. As we iterate both steps over all coarse-to-fine scales, we adapt cross covariance image transformers to allow for an efficient realization while maintaining scale-dependent properties. Experiments and ablations demonstrate that our efforts of combining multi-scale and attention-based concepts pay off. By providing highly detailed flow fields with strong improvements in both occluded and non-occluded regions, our CCMR approach not only outperforms both the corresponding single-scale attention-based and multi-scale attention-free baselines by up to 23.0% and 21.6%, respectively, it also achieves state-of-the-art results, ranking first on KITTI 2015 and second on MPI Sintel Clean and Final. Code and trained models are available at https://github.com/cv-stuttgart/CCMR. Azin Jahedi, Maximilian Luz, Marc Rivinius, Andrés Bruhn |
WACV | 1 |
| 2024 | MS-RAFT+: High Resolution Multi-Scale RAFTabstractAbstract Hierarchical concepts have proven useful in many classical and learning-based optical flow methods regarding both accuracy and robustness. In this paper we show that such concepts are still useful in the context of recent neural networks that follow RAFT’s paradigm refraining from hierarchical strategies by relying on recurrent updates based on a single-scale all-pairs transform. To this end, we introduce MS-RAFT+: a novel recurrent multi-scale architecture based on RAFT that unifies several successful hierarchical concepts. It employs a coarse-to-fine estimation to enable the use of finer resolutions by useful initializations from coarser scales. Moreover, it relies on RAFT’s correlation pyramid that allows to consider non-local cost information during the matching process. Furthermore, it makes use of advanced multi-scale features that incorporate high-level information from coarser scales. And finally, our method is trained subject to a sample-wise robust multi-scale multi-iteration loss that closely supervises each iteration on each scale, while allowing to discard particularly difficult samples. In combination with an appropriate mixed-dataset training strategy, our method performs favorably. It not only yields highly accurate results on the four major benchmarks (KITTI 2015, MPI Sintel, Middlebury and VIPER), it also allows to achieve these results with a single model and a single parameter setting. Our trained model and code are available at https://github.com/cv-stuttgart/MS_RAFT_plus . Azin Jahedi, Maximilian Luz, Marc Rivinius, Lukas Mehl, Andrés Bruhn |
Int. J. Comput. Vis. | 1 |
| 2023 | Spring: A High-Resolution High-Detail Dataset and Benchmark for Scene Flow, Optical Flow and StereoabstractWhile recent methods for motion and stereo estimation recover an unprecedented amount of details, such highly detailed structures are neither adequately reflected in the data of existing benchmarks nor their evaluation methodology. Hence, we introduce Spring - a large, high-resolution, high-detail, computer-generated benchmark for scene flow, optical flow, and stereo. Based on rendered scenes from the open-source Blender movie “Spring”, it provides photo-realistic HD datasets with state-of-the-art visual effects and ground truth training data. Furthermore, we provide a website to upload, analyze and compare results. Using a novel evaluation methodology based on a super-resolved UHD ground truth, our Spring benchmark can assess the quality of fine structures and provides further detailed performance statistics on different image regions. Regarding the number of ground truth frames, Spring is 60× larger than the only scene flow benchmark, KITTI 2015, and 15× larger than the well-established MPI Sintel optical flow benchmark. Initial results for recent methods on our benchmark show that estimating fine details is indeed challenging, as their accuracy leaves significant room for improvement. The Spring benchmark and the corresponding datasets are available at http://spring-benchmark.org. Lukas Mehl, Jenny Schmalfuss, Azin Jahedi, Yaroslava Nalivayko, Andrés Bruhn |
CVPR | 3 |
| 2023 | M-FUSE: Multi-frame Fusion for Scene Flow EstimationabstractRecently, neural network for scene flow estimation show impressive results on automotive data such as the KITTI benchmark. However, despite of using sophisticated rigidity assumptions and parametrizations, such networks are typically limited to only two frame pairs which does not allow them to exploit temporal information. In our paper we address this shortcoming by proposing a novel multi-frame approach that considers an additional preceding stereo pair. To this end, we proceed in two steps: Firstly, building upon the recent RAFT-3D approach, we develop an improved two-frame baseline by incorporating an advanced stereo method. Secondly, and even more importantly, exploiting the specific modeling concepts of RAFT-3D, we propose a U-Net architecture that performs a fusion of forward and backward flow estimates and hence allows to integrate temporal information on demand. Experiments on the KITTI benchmark do not only show that the advantages of the improved baseline and the temporal fusion approach complement each other, they also demonstrate that the computed scene flow is highly accurate. More precisely, our approach ranks second overall and first for the even more challenging foreground objects, in total outperforming the original RAFT-3D method by more than 16%. Code is available at https://github.com/cv-stuttgart/M-FUSE. Lukas Mehl, Azin Jahedi, Jenny Schmalfuss, Andrés Bruhn |
WACV | 2 |
| 2022 | Multi-Scale Raft: Combining Hierarchical Concepts for Learning-Based Optical Flow EstimationabstractMany classical and learning-based optical flow methods rely on hierarchical concepts to improve both accuracy and robustness. However, one of the currently most successful approaches – RAFT – hardly exploits such concepts. In this work, we show that multi-scale ideas are still valuable. More precisely, using RAFT as a baseline, we propose a novel multi-scale neural network that combines several hierarchical concepts within a single estimation framework. These concepts include (i) a partially shared coarse-to-fine architecture, (ii) multi-scale features, (iii) a hierarchical cost volume and (iv) a multi-scale multi-iteration loss. Experiments on MPI Sintel and KITTI clearly demonstrate the benefits of our approach. They show not only substantial improvements compared to RAFT, but also state-of-the-art results – in particular in non-occluded regions. Code will be available at https://github.com/cv-stuttgart/MS_RAFT. Azin Jahedi, Lukas Mehl, Marc Rivinius, Andrés Bruhn |
ICIP | 1 |