Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Anlan Qiu

dblp:388/3563 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 62% Vision and language · 38%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › stereo vision
motion stereo
1.012026
Match Stereo Videos via Bidirectional Alignment · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision › stereo vision
stereo matching
1.012026
Match Stereo Videos via Bidirectional Alignment · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Vision and language › 3d vision and language
3d question answering
0.912025
Hypo3D: Exploring Hypothetical Reasoning in 3D · ICML 2025
Computer vision › 3D vision
3d scene understanding
0.912025
Hypo3D: Exploring Hypothetical Reasoning in 3D · ICML 2025
Computer vision › Vision and language
visual question answering
0.912025
Hypo3D: Exploring Hypothetical Reasoning in 3D · ICML 2025

Methods — techniques the papers use, named apart from their topics

bidirectional alignment · 1.0BiDAStabilizer · 1.0vision-language foundation model · 0.9
YearPublicationVenuePosition
2026 Match Stereo Videos via Bidirectional Alignment
abstract
Video stereo matching is the task of estimating consistent disparity maps from rectified stereo videos. There is considerable scope for improvement in both datasets and methods within this area. Recent learning-based methods often focus on optimizing performance for independent stereo pairs, leading to temporal inconsistencies in videos. Existing video methods typically employ sliding window operation over time dimension, which can result in low-frequency oscillations corresponding to the window size. To address these challenges, we propose a bidirectional alignment mechanism for adjacent frames as a fundamental operation. Building on this, we introduce a novel video processing framework, BiDAStereo, and a plugin stabilizer network, BiDAStabilizer, compatible with general image-based methods. Regarding datasets, current synthetic object-based and indoor datasets are commonly used for training and benchmarking, with a lack of outdoor nature scenarios. To bridge this gap, we present a realistic synthetic dataset and benchmark focused on natural scenes, along with a real-world dataset captured by a stereo camera in diverse urban scenes for qualitative evaluation. Extensive experiments on in-domain, out-of-domain, and robustness evaluation demonstrate the contribution of our methods and datasets, showcasing improvements in prediction quality and achieving state-of-the-art results on various commonly used benchmarks.
Junpeng Jing, Ye Mao, Anlan Qiu, Krystian Mikolajczyk
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Hypo3D: Exploring Hypothetical Reasoning in 3D
abstract
The rise of vision-language foundation models marks an advancement in bridging the gap between human and machine capabilities in 3D scene reasoning. Existing 3D reasoning benchmarks assume real-time scene accessibility, which is impractical due to the high cost of frequent scene updates. To this end, we introduce Hypothetical 3D Reasoning, namely Hypo3D, a benchmark designed to evaluate models’ ability to reason without access to real-time scene data. Models need to imagine the scene state based on a provided change description before reasoning. Hypo3D is formulated as a 3D Visual Question Answering (VQA) benchmark, comprising 7,727 context changes across 700 indoor scenes, resulting in 14,885 question-answer pairs. An anchor-based world frame is established for all scenes, ensuring consistent reference to a global frame for directional terms in context changes and QAs. Extensive experiments show that state-of-the-art foundation models struggle to reason effectively in hypothetically changed scenes. This reveals a substantial performance gap compared to humans, particularly in scenarios involving movement changes and directional reasoning. Even when the change is irrelevant to the question, models often incorrectly adjust their answers. The code and dataset are publicly available at: https://matchlab-imperial.github.io/Hypo3D.
Ye Mao, Weixun Luo, Junpeng Jing, Anlan Qiu, Krystian Mikolajczyk
ICML4