EDBT 2026 Demo / reviewers in the wild / expert
Yunzhe Xu
dblp:317/5488
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-9448-4525ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 54% Robot navigation and mapping · 26% Reinforcement learning · 20% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-and-language navigation |
2.7 | 3 | 2026 | Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation · IEEE Trans. Pattern Anal. Mach. Intell. 2026 FLAME: Learning to Navigate with Multimodal LLM in Urban Environments · AAAI 2025 Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation · AAAI 2025 |
Machine learning › Reinforcement learning › partially observable reinforcement learning
memory-based navigation |
1.0 | 1 | 2026 | Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Robotics › Robot navigation and mapping
embodied navigation |
0.9 | 1 | 2025 | FLAME: Learning to Navigate with Multimodal LLM in Urban Environments · AAAI 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | FLAME: Learning to Navigate with Multimodal LLM in Urban Environments · AAAI 2025 |
Robotics › Robot navigation and mapping › mobile robot navigation › outdoor navigation
urban navigation |
0.9 | 1 | 2025 | FLAME: Learning to Navigate with Multimodal LLM in Urban Environments · AAAI 2025 |
Image and video processing › image restoration › transform-domain image restoration
frequency-domain image restoration |
0.9 | 1 | 2025 | From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral Perspective · CVPR 2025 |
Image and video processing
image restoration |
0.9 | 1 | 2025 | From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral Perspective · CVPR 2025 |
Image and video processing › image restoration
ultra-high-definition image restoration |
0.9 | 1 | 2025 | From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral Perspective · CVPR 2025 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.3 | 1 | 2026 | Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Methods — techniques the papers use, named apart from their topics
world model · 1.0imagination-guided retrieval · 1.0hybrid memory · 1.0three-phase tuning · 0.9spectral decomposition · 0.9pre-training · 0.9multimodal LLM · 0.9kolmogorov-arnold network · 0.9instruction tuning · 0.9imaginative memory system · 0.9frequency-windowed network · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) requires agents to follow natural language instructions through environments, with memory-persistent variants demanding progressive improvement through accumulated experience. Existing approaches for memory-persistent VLN face critical limitations: they lack effective memory access mechanisms, instead relying on entire memory incorporation or fixed-horizon lookup, and predominantly store only environmental observations while neglecting navigation behavioral patterns that encode valuable decision-making strategies. We present Memoir, which employs imagination as a retrieval mechanism grounded by explicit memory: a world model imagines future navigation states as queries to selectively retrieve relevant environmental observations and behavioral histories. The approach comprises: 1) a language-conditioned world model that imagines future states serving dual purposes: encoding experiences for storage and generating retrieval queries; 2) Hybrid Viewpoint-Level Memory that anchors both observations and behavioral patterns to viewpoints, enabling hybrid retrieval; and 3) an experience-augmented navigation model that integrates retrieved knowledge through specialized encoders. Extensive evaluation across diverse memory-persistent VLN benchmarks with 10 distinct testing scenarios demonstrates Memoir's effectiveness: significant improvements across all scenarios, with 5.4% SPL gains on IR2R over the best memory-persistent baseline, accompanied by $8.3\times$8.3× training speedup and 74% inference memory reduction. The results validate that predictive retrieval of both environmental and behavioral memories enables more effective navigation, with analysis indicating substantial headroom (73.3% vs 93.4% upper bound) for this imagination-guided paradigm. Yunzhe Xu, Yiyuan Pan, Zhe Liu 0022 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language NavigationabstractHumans navigate unfamiliar environments using episodic simulation and episodic memory, which facilitate a deeper understanding of the complex relationships between environments and objects. Developing an imaginative memory system inspired by human mechanisms can enhance the navigation performance of embodied agents in unseen environments. However, existing Vision-and-Language Navigation (VLN) agents lack a memory mechanism of this kind. To address this, we propose a novel architecture that equips agents with a reality-imagination hybrid memory system. This system enables agents to maintain and expand their memory through both imaginative mechanisms and navigation actions. Additionally, we design tailored pre-training tasks to develop the agent's imaginative capabilities. Our agent can imagine high-fidelity RGB images for future scenes, achieving state-of-the-art results in a Success rate weighted by Path Length (SPL). Yiyuan Pan, Yunzhe Xu, Zhe Liu 0022, Hesheng Wang 0001 |
AAAI | 2 |
| 2025 | FLAME: Learning to Navigate with Multimodal LLM in Urban EnvironmentsabstractLarge Language Models (LLMs) have demonstrated potential in Vision-and-Language Navigation (VLN) tasks, yet current applications face challenges. While LLMs excel in general conversation scenarios, they struggle with specialized navigation tasks, yielding suboptimal performance compared to specialized VLN models. We introduce FLAME (FLAMingo-Architected Embodied Agent), a novel Multimodal LLM-based agent and architecture designed for urban VLN tasks that efficiently handles multiple observations. Our approach implements a three-phase tuning technique for effective adaptation to navigation tasks, including single perception tuning for street view description, multiple perception tuning for route summarization, and end-to-end training on VLN datasets. The augmented datasets are synthesized automatically. Experimental results demonstrate FLAME's superiority over existing methods, surpassing state-of-the-art methods by a 7.3% increase in task completion on Touchdown dataset. This work showcases the potential of Multimodal LLMs (MLLMs) in complex navigation tasks, representing an advancement towards applications of MLLMs in the field of embodied intelligence. Yunzhe Xu, Yiyuan Pan, Zhe Liu 0022, Hesheng Wang 0001 |
AAAI | 1 |
| 2025 | From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral PerspectiveabstractUltra-high-definition (UHD) image restoration faces significant challenges due to its high resolution, complex content, and intricate details. To cope with these challenges, we analyze the restoration process in depth through a progressive spectral perspective, and deconstruct the complex UHD restoration problem into three progressive stages: zero-frequency enhancement, low-frequency restoration, and high-frequency refinement. Building on this insight, we propose a novel framework, ERR, which comprises three collaborative sub-networks: the zero-frequency enhancer (ZFE), the low-frequency restorer (LFR), and the high-frequency refiner (HFR). Specifically, the ZFE integrates global priors to learn global mapping, while the LFR restores low-frequency information, emphasizing reconstruction of coarse-grained content. Finally, the HFR employs our designed frequency-windowed kolmogorov-arnold networks (FW-KAN) to refine textures and details, producing high-quality image restoration. Our approach significantly outperforms previous UHD methods across various tasks, with extensive ablation studies validating the effectiveness of each component. The code is available at here. Zhizhou Chen, Yunzhe Xu, Enxuan Gu, Zili Yi, Ying Tai |
CVPR | 3 |
| 2025 | Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual NavigationabstractVisual navigation is a fundamental problem in embodied AI, yet practical deployments demand long-horizon planning capabilities to address multi-objective tasks. A major bottleneck is data scarcity: policies learned from limited data often overfit and fail to generalize OOD. Existing neural network-based agents typically increase architectural complexity that paradoxically become counterproductive in the small-sample regime. This paper introduce NeuRO, a integrated learning-to-optimize framework that tightly couples perception networks with downstream task-level robust optimization. Specifically, NeuRO addresses core difficulties in this integration: (i) it transforms noisy visual predictions under data scarcity into convex uncertainty sets using Partially Input Convex Neural Networks (PICNNs) with conformal calibration, which directly parameterize the optimization constraints; and (ii) it reformulates planning under partial observability as a robust optimization problem, enabling uncertainty-aware policies that transfer across environments. Extensive experiments on both unordered and sequential multi-object navigation tasks demonstrate that NeuRO establishes SoTA performance, particularly in generalization to unseen environments. Our work thus presents a significant advancement for developing robust, generalizable autonomous agents. Yiyuan Pan, Yunzhe Xu, Zhe Liu 0022, Hesheng Wang 0001 |
NeurIPS | 2 |
| 2025 | UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality DatasetabstractUltra-high-resolution (UHR) text-to-image (T2I) generation has seen notable progress. However, two key challenges remain : 1) the absence of a large-scale high-quality UHR T2I dataset, and (2) the neglect of tailored training strategies for fine-grained detail synthesis in UHR scenarios. To tackle the first challenge, we introduce \textbf{UltraHR-100K}, a high-quality dataset of 100K UHR images with rich captions, offering diverse content and strong visual fidelity. Each image exceeds 3K resolution and is rigorously curated based on detail richness, content complexity, and aesthetic quality. To tackle the second challenge, we propose a frequency-aware post-training method that enhances fine-detail generation in T2I diffusion models. Specifically, we design (i) \textit{Detail-Oriented Timestep Sampling (DOTS)} to focus learning on detail-critical denoising steps, and (ii) \textit{Soft-Weighting Frequency Regularization (SWFR)}, which leverages Discrete Fourier Transform (DFT) to softly constrain frequency components, encouraging high-frequency detail preservation. Extensive experiments on our proposed UltraHR-eval4K benchmarks demonstrate that our approach significantly improves the fine-grained detail quality and overall fidelity of UHR image generation. The code is available at \href{https://github.com/NJU-PCALab/UltraHR-100k}{here}. Chen Zhao 0002, En Ci, Yunzhe Xu, Tiehan Fan, Shanyan Guan, Yanhao Ge, Jian Yang 0003, Ying Tai |
NeurIPS | 3 |
| 2024 | RealWeb: A Benchmark for Universal Instruction Following in Realistic Web Services NavigationabstractTraditional methods of interacting with web pages, such as clicking and scrolling, greatly hinder users, especially those with disabilities and the elderly from conveniently accessing web services. By following user instructions, automatic web service navigation agents accomplish complex tasks on the websites, which is a natural interactions with web services. To study this task, previous works constructed simple web pages within simulated environments, but the realistic websites are far more intricate in true environments. Moreover, existing methods for this task rely on manually collecting human demonstrations on the given websites, which is time-consuming and labor-intensive, and reduce the generalization ability of service agents to unseen websites. Thus, we construct the first Chinese multimodal benchmark for web services navigation under the realistic settings: across domains and without human demonstrations. Our benchmark comprises a dataset (RealWeb) and a baseline method (WeServe). RealWeb consists of 40 real-world websites, 110 pages, and 11,739 language instructions. To detect and understand the feasible operations of pages in the visual mode, the screenshot of each page is annotated with 5 critical areas in RealWeb. WeServeis a multi-modal framework for web services navigation that combines visual and textual information, enabling universal navigation on any web pages with a success rate of 68.61%. Our dataset and codes are available at https://gitee.com/plabrolin/real-web. Shiyun Xiong, Dianbo Sui, Yunzhe Xu, Zhiying Tu |
ICWS | 4 |