Teng Wang 0006

dblp:49/9-6 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0002-1802-0435ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 First Learn, Then Review: Human-Like Continual Learning for Cross-View Geo-Localization with Limited Field of View
abstract
This paper addresses cross-view geo-localization in real-world scenarios, where the field-of-view (FoV) is restricted and the orientation is unknown for ground-view images. This task is extremely challenging due to the huge domain gap. Existing methods typically treat tasks with different FoVs as independent tasks. These approaches not only require separate retraining for each FoV, but also neglect the strong correlations between different FoVs, leading to poor performance under extremely limited FoV. To overcome these limitations, we propose HCL-Geo, a framework follows human-like continual learning paradigm of "first learn, then review" for geo-localization: in the first "learn" stage, tasks are presented to the model in an easy-to-hard sequence to enable gradual learning and knowledge retention, so that their natural correlations could be exploited to facilitate knowledge transfer. In the second "review" stage, expert modules are incorporated to efficiently handle tasks with varying FoVs. This approach eliminates the need for retraining separate models and demonstrates state-of-the-art performance across different FoVs with strong generalization capabilities. Remarkably, the recall rate@top-1 improves from 49.1% to 68.3% and from 24.6% to 34.3% respectively on CVUSA and CVACT benchmarks with 70° FoV.
Daikun Liu, Teng Wang 0006
AAAI4
2026 A modern look at simplicity bias in image classification tasks
Xiaoguang Chang, Teng Wang 0006, Changyin Sun 0001
Neural Networks2
2026 Window-to-window BEV representation learning for limited FoV cross-view geo-localization
Daikun Liu, Lingquan Meng, Teng Wang 0006, Changyin Sun 0001
Neural Networks4
2026 ImagineNav++: Prompting Vision-Language Models as Embodied Navigator Through Scene Imagination
abstract
Visual navigation is a fundamental capability for autonomous home-assistance robots, enabling the execution of long-horizon tasks such as object search. While recent methods have leveraged Large Language Models (LLMs) to incorporate commonsense reasoning and improve exploration efficiency, their planning processes remain constrained by textual representations, which cannot adequately capture spatial occupancy or scene geometry-critical factors for informed navigation decisions. In this work, we explore whether Vision-Language Models (VLMs) can achieve mapless visual navigation using only onboard RGB/RGB-D streams, unlocking their potential for spatial perception and planning. We achieve this by developing the imagination-powered navigation framework ImagineNav++, which imagines the future observation images at valuable robot views and translates the complex navigation planning process into a rather simple best-view image selection problem for VLMs. Specifically, we first introduce a future-view imagination module, which distills human navigation preferences to generate semantically meaningful candidate viewpoints with high exploration potential. These imagined future views then serve as visual prompts for the VLM to identify the most informative viewpoint. To maintain spatial consistency, we develop a selective foveation memory mechanism, which hierarchically integrates keyframe observations through a sparse-to-dense framework, thereby constructing a compact yet comprehensive memory for long-term spatial reasoning. This integrated approach effectively transforms the challenging goal-oriented navigation problem into a series of tractable point-goal navigation tasks. Extensive experiments on open-vocabulary object and instance navigation benchmarks demonstrate that our ImagineNav++ achieves SOTA performance in mapless setting, even surpassing most cumbersome map-based methods, revealing the importance of scene imagination and scene memory in VLM-based spatial reasoning.
Teng Wang 0006, Xinxin Zhao, Wenzhe Cai, Changyin Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Enhancing scene graph generation via semantic-aligned masked vision-and-language pre-training
Xiaoguang Chang, Teng Wang 0006, Lele Xu, Changyin Sun 0001
Vis. Comput.2
2025 EDCFlow: Exploring Temporally Dense Difference Maps for Event-based Optical Flow Estimation
abstract
Recent learning-based methods for event-based optical flow estimation utilize cost volumes for pixel matching but suffer from redundant computations and limited scalability to higher resolutions for flow refinement. In this work, we take advantage of the complementarity between temporally dense feature differences of adjacent event frames and cost volume and present a lightweight event-based optical flow network (EDCFlow) to achieve high-quality flow estimation at a higher resolution. Specifically, an attention-based multi-scale temporal feature difference layer is developed to capture diverse motion patterns at high resolution in a computation-efficient manner. An adaptive fusion of high-resolution difference motion features and low-resolution correlation motion features is performed to enhance motion representation and model generalization. Notably, EDCFlow can serve as a plug-and-play refinement module for RAFT-like event-based methods to enhance flow details. Extensive experiments demonstrate that EDCFlow achieves better performance with lower complexity compared to existing methods, offering superior generalization. Codes and models will be available at here.
Daikun Liu, Teng Wang 0006, Changyin Sun 0001
CVPR3
2025 ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
abstract
Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve exploration efficiency. However, the planning process of LLMs is limited within texts and it is difficult to represent the spatial occupancy and geometry layout only by texts. Both are important for making rational navigation decisions. In this work, we seek to unleash the spatial perception and planning ability of Vision-Language Models (VLMs), and explore whether the VLM, with only on-board camera captured RGB/RGB-D stream inputs, can efficiently finish the visual navigation tasks in a mapless manner. We achieve this by developing the imagination-powered navigation framework ImagineNav, which imagines the future observation images at valuable robot views and translates the complex navigation planning process into a rather simple best-view image selection problem for VLM. To generate appropriate candidate robot views for imagination, we introduce the Where2Imagine module, which is distilled to align with human navigation habits. Finally, to reach the VLM preferred views, an off-the-shelf point-goal navigation policy is utilized. Empirical experiments on the challenging open-vocabulary object navigation benchmarks demonstrates the superiority of our proposed system.
Xinxin Zhao, Wenzhe Cai, Likun Tang, Teng Wang 0006
ICLR4
2025 Discovering Intrinsic Subgoals for Vision- and-Language Navigation via Hierarchical Reinforcement Learning
abstract
Vision-and-language navigation requires an agent to navigate in a photo-realistic environment by following natural language instructions. Mainstream methods employ imitation learning (IL) to let the agent imitate the behavior of the teacher. The trained model will overfit the teacher's biased behavior, resulting in poor model generalization. Recently, researchers have sought to combine IL and reinforcement learning (RL) to overcome overfitting and enhance model generalization. However, these methods still face the problem of expensive trajectory annotation. We propose a hierarchical RL-based method-discovering intrinsic subgoals via hierarchical (DISH) RL-which overcomes the generalization limitations of current methods and gets rid of expensive label annotations. First, the high-level agent (manager) decomposes the complex navigation problem into simple intrinsic subgoals. Then, the low-level agent (worker) uses an intrinsic subgoal-driven attention mechanism for action prediction in a smaller state space. We place no constraints on the semantics that subgoals may convey, allowing the agent to autonomously learn intrinsic, more generalizable subgoals from navigation tasks. Furthermore, we design a novel history-aware discriminator (HAD) for the worker. The discriminator incorporates historical information into subgoal discrimination and provides the worker with additional intrinsic rewards to alleviate the reward sparsity. Without labeled actions, our method provides supervision for the worker in the form of self-supervision by generating subgoals from the manager. The final results of multiple comparison experiments on the Room-to-Room (R2R) dataset show that our DISH can significantly outperform the baseline in accuracy and efficiency.
Teng Wang 0006, Lele Xu, Zichen He, Changyin Sun 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 DGMem: learning visual navigation policy without any labels by dynamic graph memory
Wenzhe Cai, Teng Wang 0006, Guangran Cheng, Lele Xu, Changyin Sun 0001
Appl. Intell.2
2024 Voxel-Based Multi-Scale Transformer Network for Event Stream Processing
abstract
Event cameras are bio-inspired dynamic vision sensors that are superior to frame-based cameras in terms of low power consumption, high dynamic range, and high temporal resolution in computer vision tasks. Recent advances in voxel-based representation learning have successfully exploited the sparsity of events with low computational complexity, but face challenges in extracting spatio-temporal features within voxels and representative global dependencies between voxels, thus limiting their representation power. In this work, towards a better trade-off between accuracy and computation overhead, we propose a novel voxel-based multi-scale transformer network (VMST-Net) to process event streams. Specifically, VMST-Net projects events within voxels into multi-channel frames along the time axis, such that 2D convolutions could be leveraged to encode spatio-temporal features in voxels. Then, VMST-Net utilizes a novel multi-scale multi-head self-attention (MSMHSA) mechanism with a multi-scale fusion (MSF) module that allows different heads within each layer to attend different scale 3D neighborhoods to adaptively aggregate the coarse-to-fine voxel features with little computational costs and parameters. Moreover, to model effective global features while saving computations, we aggregate features in a local-to-global manner by enlarging the coverage of 3D neighborhoods as the network gets deeper. Extensive experimental results on benchmark datasets demonstrate that our model advances state-of-the-art accuracy with low model complexity and computational complexity in all three visual tasks, including object classification, action recognition, and human pose estimation.
Daikun Liu, Teng Wang 0006, Changyin Sun 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 DeHi: A Decoupled Hierarchical Architecture for Unaligned Ground-to-Aerial Geo-Localization
abstract
Ground-to-aerial (G2A) geo-localization remains extremely challenging due to the drastic appearance and geometry differences between ground and aerial views, especially when their relative orientation is unknown. In this paper, we focus on the challenging problem of unaligned G2A geo-localization, where the query ground-level image is not perfectly orientation-aligned with respect to reference aerial imagery. We cast this problem as a metric embedding task and propose a decoupled hierarchical (DeHi) architecture to progressively learn meaningful multi-grained features. Specifically, DeHi first leverages CNN to extract high-level semantic features, and then introduces a novel orthogonally factorized transformer model consisting of part-level and global transformer encoders to learn part-level and global feature descriptors sequentially. For the purpose of enhancing representation power, cross-level connections are introduced to enrich part-level and global descriptors by CNN features, and the pooled part-level descriptor is combined with the global descriptor to construct the final query representation. Furthermore, such a decoupled hierarchical architecture allows for incorporating multi-level deep supervision. We introduce two part-level losses combined with one cross-level loss to complement the widely used global retrieval loss. Extensive experiments on standard benchmark datasets show significant boosting in recall rates compared with the previous state-of-the-art. Remarkably, DeHi improves the recall rate @top-1 from 78.59% to 82.38% (+3.79%) and from 72.91% to 77.94% (+5.03%) on CVUSA and CVACT datasets, respectively, under random orientation misalignments. Besides, DeHi maintains competitive inference efficiency with less parameters compared to existing transformer-based methods.
Teng Wang 0006, Jiawen Li 0006, Changyin Sun 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 LANDMARK: language-guided representation enhancement framework for scene graph generation
Xiaoguang Chang, Teng Wang 0006, Shaowei Cai 0002, Changyin Sun 0001
Appl. Intell.2
2023 Towards better generalization in quadrotor landing using deep reinforcement learning
Teng Wang 0006, Zichen He, Wenzhe Cai, Changyin Sun 0001
Appl. Intell.2
2023 UAV target following in complex occluded environments with adaptive multi-modal fusion
Lele Xu, Teng Wang 0006, Wenzhe Cai, Changyin Sun 0001
Appl. Intell.2
2023 Learning a World Model With Multitimescale Memory Augmentation
abstract
Model-based reinforcement learning (RL) is regarded as a promising approach to tackle the challenges that hinder model-free RL. The success of model-based RL hinges critically on the quality of the predicted dynamic models. However, for many real-world tasks involving high-dimensional state spaces, current dynamics prediction models show poor performance in long-term prediction. To that end, we propose a novel two-branch neural network architecture with multi-timescale memory augmentation to handle long-term and short-term memory differently. Specifically, we follow previous works to introduce a recurrent neural network architecture to encode history observation sequences into latent space, characterizing the long-term memory of agents. Different from previous works, we view the most recent observations as the short-term memory of agents and employ them to directly reconstruct the next frame to avoid compounding error. This is achieved by introducing a self-supervised optical flow prediction structure to model the action-conditional feature transformation at pixel level. The reconstructed observation is finally augmented by the long-term memory to ensure semantic consistency. Experimental results show that our approach is able to generate visually-realistic long-term predictions in DeepMind maze navigation games, and outperforms the prevalent state-of-the-art methods in prediction accuracy by a large margin. Furthermore, we also evaluate the usefulness of our world model by using the predicted frames to drive an imagination-augmented exploration strategy to improve the model-free RL controller.
Wenzhe Cai, Teng Wang 0006, Changyin Sun 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 A coarse-to-fine approach for dynamic-to-static image translation
Teng Wang 0006, Lin Wu 0007, Changyin Sun 0001
Pattern Recognit.1
2021 Multi-Modal Visual Place Recognition in Dynamics-Invariant Perception Space
abstract
Visual place recognition is one of the essential and challenging problems in the fields of robotics. In this letter, we for the first time explore the use of multi-modal fusion of semantic and visual modalities in dynamics-invariant space to improve place recognition in dynamic environments. We achieve this by first designing a novel deep learning architecture to generate the static semantic segmentation and recover the static image directly from the corresponding dynamic image. We then innovatively leverage the spatial-pyramid-matching model to encode the static semantic segmentation into feature vectors. In parallel, the static image is encoded using the popular Bag-of-words model. On the basis of the above multi-modal features, we finally measure the similarity between the query image and target landmark by the joint similarity of their semantic and visual codes. Extensive experiments demonstrate the effectiveness and robustness of the proposed approach for place recognition in dynamic environments.
Lin Wu 0007, Teng Wang 0006, Changyin Sun 0001
IEEE Signal Process. Lett.2
2021 Attention-Based Road Registration for GPS-Denied UAS Navigation
abstract
Matching and registration between aerial images and prestored road landmarks are critical techniques to enhance unmanned aerial system (UAS) navigation in the global positioning system (GPS)-denied urban environments. Current registration processes typically consist of two separate stages of road extraction and road registration. These two-stage registration approaches are time-consuming and less robust to noise. To that end, in this article, we, for the first time, investigate the problem of end-to-end Aerial-Road registration. Using deep learning, we develop a novel attention-based neural network architecture for Aerial-Road registration. In this model, we construct two-branch neural networks with shared weights to map two input images into a common embedding space. Besides, considering that road features are sparsely distributed in images, we incorporate a novel multibranch attention module to filter out false descriptor matches from the indiscriminative background in order to improve registration accuracy. Finally, the results from extensive experiments show that compared with state-of-the-art approaches, the mean absolute errors of our approach in rotation angle and the translations in the x - and y -directions are reduced down by a factor of 1.24, 1.38, and 1.44, respectively. Furthermore, as a byproduct, our experimental results prove the feasibility of a neural network multitask learning approach to simultaneously achieve accurate Aerial-Road matching and registration, thus providing an efficient and accurate UAS geolocalization.
Teng Wang 0006, Arun K. Somani, Changyin Sun 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Attention-based face alignment: A solution to speed/accuracy trade-off
Teng Wang 0006, Xinjie Tong, Wenzhe Cai
Neurocomputing1
2020 Aerial-DEM geolocalization for GPS-denied UAS navigation
Teng Wang 0006, Arun K. Somani
Mach. Vis. Appl.1
2016 Characterization of mountain drainage patterns for GPS-denied UAS navigation augmentation
Teng Wang 0006, Koray Çelik, Arun K. Somani
Mach. Vis. Appl.1
2014 Characterization of Mountain Drainage Patterns for GPS-Denied UAS Navigation Augmentation
abstract
We present a novel approach to use mountain drainage patterns for GPS-Denied navigation of small unmanned aerial systems (UAS such as Scan Eagle), utilizing a down-looking fixed focus monocular imager. We leverage the analogy between mountain drainage patterns, human arteriograms, and human fingerprints. We match local drainage patterns to GPU-rendered parallax occlusion maps of offline geo-registered radar returns (GRRR) in real-time. We represent a given mountain area with a set of spatially distributed minutiae of drainage patterns so that conventional minutiae-based fingerprint matching approaches can be used. We use medical arteriography processing techniques to extract these patterns. The minutiae-based representation of mountains is achieved by exposing mountain ridges/valleys with a series of filters and then extracting mountain minutiae from these ridges/valleys. Effectiveness of minutiae-based mountain representation method is experimentally validated with no human interaction and no human-made geographic objects. Our research was in part funded by Rockwell Collins.
Teng Wang 0006, Koray Çelik, Arun K. Somani
ICPR1