Sitong Mao

dblp:204/0152 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-2490-2896ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 7 since 2021Systems, architecture and hardware · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ASCENT: Autonomous Skill Learning Toward Complex Embodied Tasks With Foundation Models
abstract
Collecting data from simulated scenarios for training robotic skills provides a safer and more controllable alternative to real-world environments. However, it demands considerable effort, including the manual construction of simulation environments, the careful design of tasks, and the challenge of obtaining effective trajectories. These limitations hinder the efficiency of data collection from simulated scenarios. In this paper, we leverage the prior knowledge of Large Language Models (LLMs) and Large Multimodal Models (LMMs) to generate simulated scenarios and embodied tasks. We introduce a novel framework, ASCENT (Autonomous Skill learning toward Complex Embodied tasks with fouNdaTion models), designed to efficiently accomplish these tasks and generate trajectory data. ASCENT features a fully autonomous skill learning mechanism based on AI agent. During task training, the AI agent identifies suitable atomic skills from an atomic skill library to either directly complete the task or serve as an initial policy for further training. Newly acquired atomic skills are subsequently added to the library. To address training failures and enhance efficiency, the AI agent uses an LLM to automatically optimize the skill training process based on feedback received from simulations. Experimental results indicate that the number of training steps required for learning new tasks can be reduced by up to 65.9 %.
Yuecheng Liu, Junyi Dong, Sitong Mao, Hesheng Wang 0001, Weigang Wu, Shunbo Zhou
ICRA5
2025 Ms. NAMI: Multimodal Semantic Navigation on Relative Metric Intention Graph
abstract
Embodied navigation in unknown environments presents the significant challenge of integrating tasks with multimodal goals into a unified framework. In this paper, we propose the Multimodal Semantic Navigation on Relative Metric Intention Graph (Ms. NAMI), a framework that integrates various navigation tasks with multimodal goals based on a relative topo-metric intention graph. A reinforcement learning based policy with a concise action space, consisting of frontier nodes and intention nodes, is designed to guide the agent to select reasonable sub-goals. A sparse reward design is introduced to reduce bias during training. Additionally, several engineering optimizations are implemented to enhance overall performance. The experimental results indicate that our method can achieve robust navigation performance in a variety of unknown environments.
Shichao Zhai, Yuxiang Cui, Shuhao Ye, Sitong Mao, Shunbo Zhou, Rong Xiong, Yue Wang 0020
ICRA5
2025 PanopticSplatting: End-to-End Panoptic Gaussian Splatting
abstract
Open-vocabulary panoptic reconstruction is a challenging task for simultaneous scene reconstruction and understanding. Recently, methods have been proposed for 3D scene understanding based on Gaussian splatting. However, these methods are multi-staged, suffering from the accumulated errors and the dependence of hand-designed components. To streamline the pipeline and achieve global optimization, we propose PanopticSplatting, an end-to-end system for open-vocabulary panoptic reconstruction. Our method introduces query-guided Gaussian segmentation with local cross attention, lifting 2D instance masks without cross-frame association in an end-to-end way. The local cross attention within view frustum effectively reduces the training memory, making our model more accessible to large scenes with more Gaussians and objects. In addition, to address the challenge of noisy labels in 2D pseudo masks, we propose label blending to promote consistent 3D segmentation with less noisy floaters, as well as label warping on 2D predictions which enhances multi-view coherence and segmentation accuracy. Our method demonstrates strong performances in 3D scene panoptic reconstruction on the ScanNet-V2 and ScanNet++ datasets, compared with both NeRF-based and Gaussian-based panoptic reconstruction methods. Moreover, PanopticSplatting can be easily generalized to numerous variants of Gaussian splatting, and we demonstrate its robustness on different Gaussian base models.
Changjian Jiang, Sitong Mao, Shunbo Zhou, Rui Fan 0001, Rong Xiong, Yue Wang 0020
IROS4
2024 SCALE: Self-Correcting Visual Navigation for Mobile Robots via Anti-Novelty Estimation
abstract
Although visual navigation has been extensively studied using deep reinforcement learning, online learning for real-world robots remains a challenging task. Recent work directly learned from offline dataset to achieve broader generalization in the real-world tasks, which, however, faces the out-of-distribution (OOD) issue and potential robot localization failures in a given map for unseen observation. This significantly drops the success rates and even induces collision. In this paper, we present a self-correcting visual navigation method, SCALE, that can autonomously prevent the robot from the OOD situations without human intervention. Specifically, we develop an image-goal conditioned offline reinforcement learning method based on implicit Q-learning (IQL). When facing OOD observation, our novel localization recovery method generates the potential future trajectories by learning from the navigation affordance, and estimates the future novelty via random network distillation (RND). A tailored cost function searches for the candidates with the least novelty that can lead the robot to the familiar places. We collect offline data and conduct evaluation experiments in three real-world urban scenarios. Experiment results show that SCALE outperforms the previous state-of-the-art methods for open-world navigation with a unique capability of localization recovery, significantly reducing the need for human intervention. Code is available at https://github.com/KubeEdge4Robotics/ScaleNav.
Yuecheng Liu, Yuzheng Zhuang, Sitong Mao, Shunbo Zhou
ICRA4
2024 Scale Disparity of Instances in Interactive Point Cloud Segmentation
abstract
Interactive point cloud segmentation has become a pivotal task for understanding 3D scenes, enabling users to guide segmentation models with simple interactions such as clicks, therefore significantly reducing the effort required to tailor models to diverse scenarios and new categories. However, in the realm of interactive segmentation, the meaning of instance diverges from that in instance segmentation, because users might desire to segment instances of both thing and stuff categories that vary greatly in scale. Existing methods have focused on thing categories, neglecting the segmentation of stuff categories and the difficulties arising from scale disparity. To bridge this gap, we propose ClickFormer, an innovative interactive point cloud segmentation model that accurately segments instances of both thing and stuff categories. We propose a query augmentation module to augment click queries by a global query sampling strategy, thus maintaining consistent performance across different instance scales. Additionally, we employ global attention in the query-voxel transformer to mitigate the risk of generating false positives, along with several other network structure improvements to further enhance the model’s segmentation performance. Experiments demonstrate that ClickFormer outperforms existing interactive point cloud segmentation methods across both indoor and outdoor datasets, providing more accurate segmentation results with fewer user clicks in an open-world setting. Project page: https://sites.google.com/view/clickformer/
Chenrui Han, Yili Liu, Sitong Mao, Shunbo Zhou, Rong Xiong, Yue Wang 0020
IROS5
2024 PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic Reconstruction
abstract
Panoptic reconstruction is a challenging task in 3D scene understanding. However, most existing methods heavily rely on pre-trained semantic segmentation models and known 3D object bounding boxes for 3D panoptic segmentation, which is not available for in-the-wild scenes. In this paper, we propose a novel zero-shot panoptic reconstruction method from RGB-D images of scenes. For zero-shot segmentation, we leverage open-vocabulary instance segmentation, but it has to face partial labeling and instance association challenges. We tackle both challenges by propagating partial labels with the aid of dense generalized features and building a 3D instance graph for associating 2D instance IDs. Specifically, we exploit partial labels to learn a classifier for generalized semantic features to provide complete labels for scenes with dense distilled features. Moreover, we formulate instance association as a 3D instance graph segmentation problem, allowing us to fully utilize the scene geometry prior and all 2D instance masks to infer global unique pseudo 3D instance ID. Our method outperforms state-of-the-art methods on the indoor dataset ScanNet V2 and the outdoor dataset KITTI-360, demonstrating the effectiveness of our graph segmentation method and reconstruction network.
Yili Liu, Chenrui Han, Sitong Mao, Shunbo Zhou, Rong Xiong, Yiyi Liao, Yue Wang 0020
IROS4
2023 A Novel Framework for Adaptive Quadruped Robot Locomotion Learning in Uncertain Environments
Bin Guo 0001, Kaixing Zhao, Ruonan Xu, Sicong Liu 0005, Sitong Mao, Shunbo Zhou, Qiaobo Xu, Zhiwen Yu 0001
GPC (2)6
2021 Network Together: Node Classification via Cross-Network Deep Network Embedding
abstract
Network embedding is a highly effective method to learn low-dimensional node vector representations with original network structures being well preserved. However, existing network embedding algorithms are mostly developed for a single network, which fails to learn generalized feature representations across different networks. In this article, we study a cross-network node classification problem, which aims at leveraging the abundant labeled information from a source network to help classify the unlabeled nodes in a target network. To succeed in such a task, transferable features should be learned for nodes across different networks. To this end, a novel cross-network deep network embedding (CDNE) model is proposed to incorporate domain adaptation into deep network embedding in order to learn label-discriminative and network-invariant node vector representations. On the one hand, CDNE leverages network structures to capture the proximities between nodes within a network, by mapping more strongly connected nodes to have more similar latent vector representations. On the other hand, node attributes and labels are leveraged to capture the proximities between nodes across different networks by making the same labeled nodes across networks have aligned latent vector representations. Extensive experiments have been conducted, demonstrating that the proposed CDNE model significantly outperforms the state-of-the-art network embedding algorithms in cross-network node classification.
Xiao Shen 0001, Quanyu Dai, Sitong Mao, Korris Fu-Lai Chung, Kup-Sze Choi
IEEE Trans. Neural Networks Learn. Syst.3
2020 Cross-Network Learning With Fuzzy Labels for Seed Selection and Graph Sparsification in Influence Maximization
abstract
To maximize the influence across multiple heterogeneous networks, we propose an innovative cross-network learning model to study the influence maximization problem from two perspectives, namely, seed selection and graph sparsification. On one hand, we consider seed selection as a cross-network node prediction task, by leveraging the greedy seed selection knowledge prelearned in a smaller source network, to heuristically select the nodes most likely to act as seed for the target networks. On the other hand, we consider graph sparsification as a cross-network edge prediction problem, by adapting the influence propagation knowledge previously acquired in the source network to remove the edges least likely to contribute to influence propagation in the target networks. To address domain discrepancy, a fuzzy self-learning algorithm is proposed to iteratively train the prediction model by leveraging not only the fully labeled data in the source network, but also the most confident predicted instances with their predicted fuzzy labels in the target network. With such fuzzy labels, we can differentiate the confident levels of predictions generated by different self-training iterations, thus lowering the negative effects caused by less confident predictions. The performance of the proposed model is benchmarked with the popular influence maximization algorithms for seed selection; and also competed with several graph sparsification algorithms for inactive edge prediction. Experimental results on the real-world datasets show that the proposed cross-network learning model can achieve a good tradeoff between the efficiency and effectiveness of the influence maximization task in the target networks.
Xiao Shen 0001, Sitong Mao, Korris Fu-Lai Chung
IEEE Trans. Fuzzy Syst.2
2018 Deep Domain Adaptation Based on Multi-layer Joint Kernelized Distance
abstract
Domain adaptation refers to the learning scenario where a model learned from the source data is applied on the target data which have the same categories but different distributions. In information retrieval, there exist application scenarios like cross domain recommendation characterized similarly. In this paper, by utilizing deep features extracted from the deep networks, we proposed to compute the multi-layer joint kernelized mean distance between the k th target data predicted as the i th category and all the source data of the j th category $d_ij ^k$. Then, target data $T_m$ that are most likely to belong to the i th category can be found by calculating the relative distance $d_ii ^k/\sum_j d_ij ^k$. By iteratively adding $T_m$ to the training data, the finetuned deep model can adapt on the target data progressively. Our results demonstrate that the proposed method can achieve a better performance compared to a number of state-of-the-art methods.
Sitong Mao, Xiao Shen 0001, Korris Fu-Lai Chung
SIGIR1
2017 Leveraging Cross-Network Information for Graph Sparsification in Influence Maximization
abstract
When tackling large-scale influence maximization (IM) problem, one effective strategy is to employ graph sparsification as a pre-processing step, by removing a fraction of edges to make original networks become more concise and tractable for the task. In this work, a Cross-Network Graph Sparsification (CNGS) model is proposed to leverage the influence backbone knowledge pre-detected in a source network to predict and remove the edges least likely to contribute to the influence propagation in the target networks. Experimental results demonstrate that conducting graph sparsification by the proposed CNGS model can obtain a good trade-off between efficiency and effectiveness of IM, i.e., existing IM greedy algorithms can run more efficiently, while the loss of influence spread can be made as small as possible in the sparse target networks.
Xiao Shen 0001, Korris Fu-Lai Chung, Sitong Mao
SIGIR3