Niranjan Sujay

dblp:393/0338 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Robot navigation and mapping · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
embodied navigation
0.912025
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos · CVPR 2025
Robotics › Robot navigation and mapping
visual navigation
0.912025
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos · CVPR 2025
Robotics › Robot navigation and mapping › mobile robot navigation › outdoor navigation
urban navigation
0.312025
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos · CVPR 2025

Methods — techniques the papers use, named apart from their topics

imitation learning · 0.9action supervision extraction · 0.9
YearPublicationVenuePosition
2025 CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
abstract
Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progress, existing visual navigation methods struggle in map-free or off-street settings, limiting the deployment of autonomous agents like last-mile delivery robots. To overcome these obstacles, we propose a scalable, data-driven approach for human-like urban navigation by training agents on thousands of hours of in-the-wild city walking and driving videos sourced from the web. We introduce a simple and scalable data processing pipeline that extracts action supervision from these videos, enabling large-scale imitation learning without costly annotations. Our model learns sophisticated navigation policies to handle diverse challenges and critical scenarios. Experimental results show that training on large-scale, diverse datasets significantly enhances navigation performance, surpassing current methods. This work shows the potential of using abundant online video data to develop robust navigation policies for embodied agents in dynamic urban settings.
Xinhao Liu 0003, Jintong Li, Niranjan Sujay, Juexiao Zhang, John Abanes, Chen Feng 0002
CVPR4