Shuo Wang 0015

dblp:63/1591-15 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-6720-1646ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction
abstract
Reconstructing dense geometry for dynamic scenes from a monocular video is a critical yet challenging task. Recent memory-based methods enable efficient online reconstruction, but they fundamentally suffer from a Memory Demand Dilemma: The memory representation faces an inherent conflict between the long-term stability required for static structures and the rapid, high-fidelity detail retention needed for dynamic motion. This conflict forces existing methods into a compromise, leading to either geometric drift in static structures or blurred, inaccurate reconstructions of dynamic objects. To address this dilemma, we propose Mem4D, a novel framework that decouples the modeling of static geometry and dynamic motion. Guided by this insight, we design a dual-memory architecture: 1) The Transient Dynamics Memory (TDM) focuses on capturing high-frequency motion details from recent frames, enabling accurate and fine-grained modeling of dynamic content; 2) The Persistent Structure Memory (PSM) compresses and preserves long-term spatial information, ensuring global consistency and drift-free reconstruction for static elements. By alternating queries to these specialized memories, Mem4D simultaneously maintains static geometry with global consistency and reconstructs dynamic elements with high fidelity. Experiments on challenging benchmarks demonstrate that our method achieves state-of-the-art or competitive performance while maintaining high efficiency.
Shuo Wang 0015, Peng Wang 0106, Yongcai Wang, Zhaoxin Fan, Tianbao Zhang, Jianrong Tao, Yeying Jin, Deying Li 0001
AAAI2
2026 MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
abstract
Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on Vision-Language Action (VLA) models achieve strong results with monocular input, yet they still lag behind methods using panoramic RGB-D information. We present MonoDream, a lightweight VLA framework that enables monocular agents to learn a Unified Navigation Representation (UNR). This shared feature representation jointly aligns navigation-relevant visual semantics (e.g., global layout, depth, and future cues) and language-grounded action intent, enabling more reliable action prediction. MonoDream further introduces Latent Panoramic Dreaming (LPD) tasks to supervise the UNR, which train the model to predict latent features of panoramic RGB and depth observations at both current and future steps based on only monocular input. Experiments on multiple VLN benchmarks show that MonoDream consistently improves monocular navigation performance and significantly narrows the gap with panoramic-based agents.
Shuo Wang 0015, Yongcai Wang, Zhaoxin Fan, Maiyue Chen, Kaihui Wang, Zhizhong Su, Yeying Jin, Deying Li 0001
AAAI1
2026 Fusion Information Bottleneck-Driven Vision Transformer Model for Diagnosis of Retinal Diseases via Internet of Medical Things
abstract
Ocular diseases are among the leading causes of visual impairment and blindness worldwide, and early, accurate diagnosis is crucial for preventing disease progression. In recent years, deep learning methods based on fundus images have achieved remarkable progress in intelligent diagnosis, yet they still face challenges such as severe overfitting, redundant features, and insufficient capability in fine-grained lesion recognition. These issues become more pronounced in the Internet of Medical Things (IoMT) scenario, further limiting the generalization and practicality of the models. To address this, this paper proposes a Vision Transformer model incorporating Information Bottleneck and Contrastive Learning mechanisms (FIBCL-ViT) for multi-class fundus disease diagnosis. Using ViT as the backbone, the model introduces a variational information bottleneck module to compress redundant features and highlight task-relevant information, while integrating a contrastive learning strategy to enhance feature discriminability, thereby improving the model’s generalization ability and robustness. During training, the model jointly optimizes cross-entropy loss, contrastive loss, and the information bottleneck regularization term to achieve efficient learning. Experimental results demonstrate that the proposed FIBCL-ViT model outperforms mainstream comparison methods across multiple public fundus image datasets, achieving superior performance in terms of accuracy, AUC, and other metrics. For the 7:3 training–testing split, the proposed method achieves an accuracy of 0.9376, a precision of 0.9384, a recall of 0.9374, and an AUC of 0.9902 on public fundus image datasets. This study provides a novel solution for automated screening and remote intelligent diagnosis of ophthalmic diseases, showing strong potential for clinical applications.
Shuo Wang 0015, Bingshuo Li, Hanyi Ren, Xinji Yang, Yongcai Wang, Yueyue Li
IEEE Internet Things J.4
2026 ThinkMatter: Panoramic-Aware Instructional Semantics for Monocular Vision-and-Language Navigation
abstract
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an embodied robot to navigate the target destination following the natural language instruction. Most existing methods use panoramic RGB-D cameras for 360° observation of environments. However, these methods struggle in real-world applications because of the higher cost of panoramic RGB-D cameras. This paper studies a low-cost and practical VLN-CE setting, e.g., using monocular cameras of limited field of view, which means "Look Less" for visual observations and environment semantics. In this paper, we propose a ThinkMatter framework for monocular VLN-CE, where we motivate monocular robots to "Think More" by 1) generating novel views and 2) integrating instruction semantics. Specifically, we achieve the former by the proposed 3DGS-based panoramic generation to render novel views at each step, based on past observation collections. We achieve the latter by the proposed enhancement of the occupancy-instruction semantics, which integrates the spatial semantics of occupancy maps with the textual semantics of language instructions. These operations promote monocular robots with wider environment perceptions as well as transparent semantic connections with the instruction. Both extensive experiments in the simulators and real-world environments demonstrate the effectiveness of ThinkMatter, providing a promising practice for real-world navigation.
Guangzhao Dai, Shuo Wang 0015, Hao Zhao 0002, Bin Zhu 0006, Qianru Sun, Xiangbo Shu
IEEE Trans. Image Process.2
2026 Dust to Tower: Prior-Driven Coarse-to-Fine Photo-Realistic Scene Reconstruction From Sparse Uncalibrated Images
abstract
Photo-realistic scene reconstruction from sparse-view, uncalibrated images is highly required in practice. Although some successes have been made, existing methods are either Sparse-View but require accurate camera parameters (i.e., intrinsic and extrinsic), or SfM-free but need densely captured images. This paper proposes Dust to Tower (D2T), a novel coarse-to-fine framework to address the coupled difficulty. The key idea is to explicitly narrow down the solution space and then introduce reliable supervision at novel viewpoints without resorting to expensive diffusion-based view synthesis. To do this, we first introduce a Coarse Construction Module (CCM) which exploits a fast Multi-View Stereo model to initialize a 3D Gaussian Splatting (3DGS) and recover initial camera poses. To refine the 3D model at novel viewpoints, we introduce Confidence-Aware Depth Alignment (CADA), which aligns a monocular inverse-depth prior to the reliable regions of the coarse depth using DUSt3R confidence, producing sharp and scale-consistent depth maps for accurate warping. We further propose Warped Image-Guided Inpainting (WIGI), which converts the accurate warped views into multi-view-consistent pseudo supervision via elaborate warping and inpainting process. Experiments on three benchmark datasets show that D2T achieves superior novel view synthesis quality and pose accuracy over ten representative baselines, while keeping high efficiency.
Yongcai Wang, Zhaoxin Fan, Shuo Wang 0015, Deying Li 0001, Lun Luo, Minhang Wang, Hongyuan Zhang 0001, Xuelong Li 0001
IEEE Trans. Vis. Comput. Graph.5
2025 MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing
abstract
Deep visual odometry has demonstrated great advancements by learning-to-optimize technology. This approach heavily relies on the visual matching across frames. However, ambiguous matching in challenging scenarios leads to significant errors in geometric modeling and bundle adjustment optimization, which undermines the accuracy and robustness of pose estimation. To address this challenge, this paper proposes MambaVO, which conducts robust initialization, Mamba-based sequential matching refinement, and smoothed training to enhance the matching quality and improve the pose estimation. Specifically, the new frame is matched with the closest keyframe in the maintained Point-Frame Graph (PFG) via the semi-dense based Geometric Initialization Module (GIM). Then the initialized PFG is processed by a proposed Geometric Mamba Module (GMM), which exploits the matching features to refine the overall inter-frame matching. The refined PFG is finally processed by differentiable BA to optimize the poses and the map. To deal with the gradient variance, a Trending-Aware Penalty (TAP) is proposed to smooth training and enhance convergence and stability. A loop closure module is finally applied to enable MambaVO++. On public benchmarks, MambaVO and MambaVO++ demonstrate SOTA performance, while ensuring real-time running.
Shuo Wang 0015, Yongcai Wang, Zhaoxin Fan, Jian Zhao 0006, Deying Li 0001
CVPR1
2025 Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation
abstract
Vision-Language Navigation is a critical task for developing embodied agents that can follow natural language instructions to navigate in complex real-world environments. Recent advances by finetuning large pretrained models have significantly improved generalization and instruction grounding compared to traditional approaches. However, the role of reasoning strategies in navigation—an action-centric, long-horizon task—remains underexplored, despite Chain-of-Thought reasoning's demonstrated success in static tasks like question answering and visual reasoning. To address this gap, we conduct the first systematic evaluation of reasoning strategies for VLN, including No-Think (direct action prediction), Pre-Think (reason before action), and Post-Think (reason after action). Surprisingly, our findings reveal the Inference-time Reasoning Collaps issue, where inference-time reasoning degrades navigation accuracy, highlighting the challenges of integrating reasoning into VLN. Based on this insight, we propose Aux-Think, a framework that trains models to internalize structured reasoning patterns through CoT supervision during training, while preserving No-Think inference for efficient action prediction. To support this framework, we release R2R-CoT-320k, a large-scale Chain-of-Thought annotated dataset. Empirically, Aux-Think significantly reduces training effort without compromising performance.
Shuo Wang 0015, Yongcai Wang, Maiyue Chen, Kaihui Wang, Zhizhong Su, Deying Li 0001, Zhaoxin Fan
NeurIPS1
2025 A Geometric and Hypothesis-Based Method for Low-Overlap, Sparse, and Featureless Point Set Matching
abstract
This article proposes a general solution for point set matching that effectively addresses the challenges of low-overlap, sparse, or featureless point set matching (LSFPM). Unlike previous methods that mainly rely on feature or neighborhood similarity that often fail under such difficult conditions, this work proposes a Geometry-based Point Matching (GPM) method. GPM first introduces two geometric concepts: the “Structural Element” (SE) and the “Superstructural Element” (SSE), both of which are constructed based on local geometric structures. The SSE is an enhanced version of the SE. A descriptor for the SE, called the SE Descriptor (SED), is designed to encode the SE and facilitate an efficient geometry-based similarity metric. We demonstrate that the cosine similarity of SEDs is invariant to scale and rotation. Subsequently, a SE Matching Maximization (SEMM) problem is formulated to identify a size-penalized SSE set that maximizes the sum of similarities. This problem is efficiently solved using the proposed SEMM-MCMC (Markov Chain Monte Carlo) algorithm. The matched SSEs then vote on corresponding point matches, generating high-confidence one-to-one matches, low-confidence one-to-one matches and one-to-many matches. Finally, the InferMatch algorithm is proposed to jointly assess low-confidence one-to-one point set matching while simultaneously distinguishing one-to-many point set matching. The GPM approach can also complement other feature-based and motion-based methods. It has been extensively validated on both synthetic and real datasets, demonstrating its versatility in addressing various point set matching problems, and is not limited to the LSFPM problem. Extensive experiments on diverse datasets, including SPair-71k, UAVDT, VisDrone2021-MOT, and SparseMatch, further demonstrate the robustness and versatility of GPM. The proposed approach significantly improves matching performance under challenging conditions and effectively addresses key limitations of existing point set matching methods.
Xuewei Bai, Yongcai Wang, Peng Wang 0106, Chunxu Li, Shuo Wang 0015, Deying Li 0001
ACM Trans. Sens. Networks5
2024 RoCo: Robust Cooperative Perception By Iterative Object Matching and Pose Adjustment
abstract
Collaborative autonomous driving with multiple vehicles usually requires the data fusion from multiple modalities. To ensure effective fusion, the data from each individual modality shall maintain a reasonably high quality. However, in collaborative perception, the quality of object detection based on a modality is highly sensitive to the relative pose errors among the agents. It leads to feature misalignment and significantly reduces collaborative performance. To address this issue, we propose RoCo, a novel unsupervised framework to conduct iterative object matching and agent pose adjustment. To the best of our knowledge, our work is the first to model the pose correction problem in collaborative perception as an object matching task, which reliably associates common objects detected by different agents. On top of this, we propose a graph optimization process to adjust the agent poses by minimizing the alignment errors of the associated objects, and the object matching is re-done based on the adjusted agent poses. This process is carried out iteratively until convergence. Experimental study on both simulated and real-world datasets demonstrates that the proposed framework RoCo consistently outperforms existing relevant methods in terms of the collaborative object detection performance, and exhibits highly desired robustness when the pose information of agents is with high-level noise. Ablation studies are also provided to show the impact of its key parameters and components. The code is released at https://github.com/HuangZhe885/RoCo.
Shuo Wang 0015, Yongcai Wang, Deying Li 0001, Lei Wang 0001
ACM Multimedia2
2024 GSLAMOT: A Tracklet and Query Graph-based Simultaneous Locating, Mapping, and Multiple Object Tracking System
Shuo Wang 0015, Yongcai Wang, Yongyu Guo, Xuewei Bai, Deying Li 0001
ACM Multimedia1
2024 InferLoc: Hypothesis-Based Joint Edge Inference and Localization in Sparse Sensor Networks
abstract
Ranging-based localization is a fundamental problem in the Internet of Things and unmanned aerial vehicle networks. However, the nodes’ limited-ranging scope and users’ broad coverage purpose inevitably cause network sparsity or subnetwork sparsity. The performances of existing localization algorithms are extremely unsatisfactory in sparse networks. A crucial way to deal with the sparsity is to exploit the hidden knowledge provided by the unmeasured edges, which inspires this work to propose a hypothesis-based Joint Edge Inference and Localization algorithm called InferLoc . InferLoc mines the Unmeasured but Inferable Edges (UIEs). Each UIE is an unmeasured edge, but it is restricted through other edges in the network to be inside a rigid component, so it has only a limited number of possible lengths. We propose an efficient method to detect UIEs and geometric approaches to infer possible lengths for UIEs in 2D and 3D networks. The inferred possible lengths of UIEs are then treated as multiple hypotheses to determine the node locations and the lengths of UIEs simultaneously through a joint graph optimization process. In the joint graph optimization model, to make the 0/1 decision variables for hypotheses selection differentiable, differentiable functions are proposed to relax the 0/1 selections, and rounding is applied to select the final length after the optimization converges. We also prove the condition when a UIE can contribute to sparse localization. Extensive experiments show remarkably better accuracy and efficiency performances of InferLoc than the state-of-the-art network localization algorithms. In particular, it reduces the localization errors by more than 90% and speeds up the convergence time more than 100 times than that of the widely used G2O-based methods in sparse networks.
Xuewei Bai, Yongcai Wang, Haodi Ping, Xiaojia Xu, Deying Li 0001, Shuo Wang 0015
ACM Trans. Sens. Networks6
2023 ColSLAM: A Versatile Collaborative SLAM System for Mobile Phones Using Point-Line Features and Map Caching
abstract
Over the past years, augmented reality (AR) based on mobile phones has gained great attention. When multiple phones are used in AR applications, collaborative simultaneous localization and mapping (SLAM) is considered one of the enabling technologies, i.e., multiple mobile phones complete the localization and mapping through collaboration. However, the state-of-the-art collaborative SLAM systems not only suffer from the delays introduced by a high-complexity graph optimization problem, but also may exhibit varying levels of accuracy across dissimilar environments or different types of mobile devices. In this paper, we propose a scalable and robust collaborative SLAM system, point-line-based Collaborative SLAM (ColSLAM). Technically, ColSLAM includes two innovative features that help achieve satisfactory scalability and robustness. First, a mapping cacher (MC) is designed for each agent on the server, which uses global keyframes to detect loop closures, updates the cached local map, and quickly responds to the agent's pose drifts. With MC, each agent's local pose is corrected using global knowledge in real-time. Secondly, to improve the robustness performance, ColSLAM employs point-line-fusion-based Visual Inertial Odometry (VIO), point-line-fusion-based NetVLAD loop detection, and an enhanced geometric verification and relative pose calculation method called PNPL. Empirical evaluations based on the EuRoc dataset and real degenerate environments demonstrate that ColSLAM outperforms the existing collaborative SLAM systems in terms of accuracy, robustness, and scalability.
Yongcai Wang, Yongyu Guo, Shuo Wang 0015, Xuewei Bai, Qiang Ye 0001, Deying Li 0001
ACM Multimedia4
2023 Communication Efficient, Distributed Relative State Estimation in UAV Networks
abstract
Distributed estimation of 6-DOF relative states, including three-dimensional relative poses and three-dimensional relative positions, is a key problem in UAV (Unmanned Aerial Vehicle) networks, which generally requires vision-involved iterative state estimation. How to achieve communication efficiency is a crucial challenge considering the large volume of vision data. This paper jointly considers the communication efficiency, latency, and accuracy for distributed relative state estimation involving vision data in UAV networks. The key is to solve a distributed graph optimization problem, which includes two key steps: (1) local graph construction and node state initialization in an initialization phase, and (2) iterative state update and communication with neighbors until convergence in online iteration phase. A communication efficient, Locating Then Informing (LTI) initialization scheme is proposed, which is run only once by each node to initialize each node’s local graph and initial states. For online iteration, a RIPPLE-like distributed state iteration scheme is proposed. It inherits the advantages of traditional sequential and parallel methods while avoiding their drawbacks. It enables nodes’ states to converge quickly using fewer rounds of communications. The communication costs for the initialization and online iteration processes are analyzed theoretically. Extensive evaluations use synthetic data generated by AirSim (a widely used UAV network simulation platform) and real-world data are presented. The results show that the proposed method provides accuracy comparable to the centralized graph optimization method and significantly outperforms the other distributed methods in terms of accuracy, communication cost, and latency.
Shuo Wang 0015, Yongcai Wang, Xuewei Bai, Deying Li 0001
IEEE J. Sel. Areas Commun.1
2022 AirBirds: A Large-scale Challenging Dataset for Bird Strike Prevention in Real-world Airports
Hongyu Sun 0006, Yongcai Wang, Peng Wang 0106, Deying Li 0001, Shuo Wang 0015
ACCV (5)8