Yingdong Hu

dblp:219/8916 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-6564-4099ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Computer networks · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dynamics-Aware Gaussian Splatting Streaming Toward Fast On-the-Fly 4D Reconstruction
abstract
The recent development of 3D Gaussian splatting (3DGS) has led to great interest in 4D dynamic spatial reconstruction. Existing approaches mainly rely on full-length multi-view videos, while there has been limited exploration of online reconstruction methods that enable on-the-fly training and per-timestep streaming. Current 3DGS-based streaming methods treat the Gaussian primitives uniformly and constantly renew the densified Gaussians. Thus, they overlook the difference between dynamic and static features and neglect the temporal continuity of the scene. To address these limitations, we propose a novel pipeline for iterative streamable 4D dynamic spatial reconstruction. It comprises three stages: a selective inheritance stage that retains priors from previous timesteps to preserve the temporal continuity, a dynamics-aware shift stage that distinguishes dynamic and static primitives and employs distinct strategies to optimize their movements, and an error-guided densification stage that efficiently identifies Gaussians requiring densification to accommodate emerging objects. Our method achieves state-of-the-art performance in online 4D reconstruction, demonstrating compact storage, the fastest on-the-fly training speed, and superior representation quality.
Zhening Liu 0001, Yingdong Hu, Jiawei Shao, Zehong Lin, Jun Zhang 0004
IEEE Trans. Vis. Comput. Graph.2
2026 Joint User Scheduling and Multi-Domain Resource Allocation for Terrestrial and Non-Terrestrial Networks Integration
abstract
Efficient resource utilization is vital for terrestrial and non-terrestrial networks (TN-NTN) integration. However, different spatio-temporal resource scales in TN and NTN networks pose challenges for joint resource allocation. To tackle this problem, we propose a joint user scheduling and multi-domain resource allocation scheme in the downlink network, to improve coverage for ground users (GUs). Specifically, the scheme is designed in two time-scales, including large-scale satellite beam-hopping (i.e., spatial resource allocation) and small-scale time-frequency resource allocation. For beam-hopping, we first analyze the coverage of terrestrial base stations (TBS) for GUs, and accordingly propose a joint design of user scheduling and beam-hopping. For time-frequency resource allocation, to cope with the complexity induced by multi-domain and multi-scale resources, we propose a two-step approach which first obtains a preliminary allocation with worst-case co-frequency interference assumption, then employs the genetic algorithm to re-allocate redundant resources, thereby increasing the proportion of successfully served GUs. We evaluate the performance of the proposed scheme through simulations with different user demands and network service settings. Results show that the proposed scheme provides considerable improvement over existing schemes, which can efficiently reduce co-frequency interference and provide service for more GUs.
Yingdong Hu, Ye Li 0004, Jue Wang 0006, Ruifeng Gao, Sheng Wu 0001, Tony Q. S. Quek
IEEE Trans. Wirel. Commun.1
2026 Neural Representation for Wireless Radiation Field Reconstruction: A 3D Gaussian Splatting Approach
abstract
Wireless channel modeling plays a pivotal role in designing, analyzing, and optimizing wireless communication systems. Nevertheless, developing an effective channel modeling approach has been a long-standing challenge. This issue has been escalated due to denser network deployment, larger antenna arrays, and broader bandwidth in next-generation networks. To address this challenge, we put forth WRF-GS, a novel framework for channel modeling based on wireless radiation field (WRF) reconstruction using 3D Gaussian splatting (3D-GS). WRF-GS employs 3D Gaussian primitives and neural networks to capture the interactions between the environment and radio signals, enabling efficient WRF reconstruction and visualization of the propagation characteristics. The reconstructed WRF can then be used to synthesize the spatial spectrum for comprehensive wireless channel characterization. While WRF-GS demonstrates remarkable effectiveness, it faces limitations in capturing high-frequency signal variations caused by complex multipath effects. To overcome these limitations, we propose WRF-GS+, an enhanced framework that integrates electromagnetic wave physics into the neural network design. WRF-GS+ leverages deformable 3D Gaussians to model both static and dynamic components of the WRF, significantly improving its ability to characterize signal variations. In addition, WRF-GS+ accelerates the splatting process by simplifying the 3D-GS modeling operation and reducing sample complexity. Experimental results demonstrate that both WRF-GS and WRF-GS+ outperform baselines for spatial spectrum synthesis, including ray tracing and other deep-learning approaches. Notably, WRF-GS+ achieves state-of-the-art performance in the received signal strength indication (RSSI) and channel state information (CSI) prediction tasks, surpassing existing methods by more than 0.7 dB and 3.36 dB, respectively. The code is available at https://github.com/wenchaozheng/WRF-GSplus.
Chaozheng Wen, Jingwen Tong, Yingdong Hu, Zehong Lin, Jun Zhang 0004
IEEE Trans. Wirel. Commun.3
2025 Data Scaling Laws in Imitation Learning for Robotic Manipulation
abstract
Data scaling has revolutionized fields like natural language processing and computer vision, providing models with remarkable generalization capabilities. In this paper, we investigate whether similar data scaling laws exist in robotics, particularly in robotic manipulation, and whether appropriate data scaling can yield single-task robot policies that can be deployed zero-shot for any object within the same category in any environment. To this end, we conduct a comprehensive empirical study on data scaling in imitation learning. By collecting data across numerous environments and objects, we study how a policy’s generalization performance changes with the number of training environments, objects, and demonstrations. Throughout our research, we collect over 40,000 demonstrations and execute more than 15,000 real-world robot rollouts under a rigorous evaluation protocol. Our findings reveal several intriguing results: the generalization performance of the policy follows a roughly power-law relationship with the number of environments and objects. The diversity of environments and objects is far more important than the absolute number of demonstrations; once the number of demonstrations per environment or object reaches a certain threshold, additional demonstrations have minimal effect. Based on these insights, we propose an efficient data collection strategy. With four data collectors working for one afternoon, we collect sufficient data to enable the policies for two tasks to achieve approximately 90\% success rates in novel environments with unseen objects.
Fanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen, Jiacheng You, Yang Gao 0029
ICLR2
2025 WRF-GS: Wireless Radiation Field Reconstruction with 3D Gaussian Splatting
Chaozheng Wen, Jingwen Tong, Yingdong Hu, Zehong Lin, Jun Zhang 0004
INFOCOM3
2025 Self-supervised vision transformers for semantic segmentation
Xianfan Gu, Yingdong Hu, Chuan Wen, Yang Gao 0029
Comput. Vis. Image Underst.2
2024 Imitation Learning from Observation with Automatic Discount Scheduling
abstract
Humans often acquire new skills through observation and imitation. For robotic agents, learning from the plethora of unlabeled video demonstration data available on the Internet necessitates imitating the expert without access to its action, presenting a challenge known as Imitation Learning from Observation (ILfO). A common approach to tackle ILfO problems is to convert them into inverse reinforcement learning problems, utilizing a proxy reward computed from the agent's and the expert's observations. Nonetheless, we identify that tasks characterized by a progress dependency property pose significant challenges for such approaches; in these tasks, the agent needs to initially learn the expert's preceding behaviors before mastering the subsequent ones. Our investigation reveals that the main cause is that the reward signals assigned to later steps hinder the learning of initial behaviors. To address this challenge, we present a novel ILfO framework that enables the agent to master earlier behaviors before advancing to later ones. We introduce an Automatic Discount Scheduling (ADS) mechanism that adaptively alters the discount factor in reinforcement learning during the training phase, prioritizing earlier rewards initially and gradually engaging later rewards only when the earlier behaviors have been mastered. Our experiments, conducted on nine Meta-World tasks, demonstrate that our method significantly outperforms state-of-the-art methods across all tasks, including those that are unsolvable by them. Our code is available at https://il-ads.github.io.
Weijun Dong, Yingdong Hu, Chuan Wen, Zhao-Heng Yin, Chongjie Zhang, Yang Gao 0029
ICLR3
2024 CoPa: General Robotic Manipulation through Spatial Constraints of Parts with Foundation Models
abstract
Foundation models pre-trained on web-scale data are shown to encapsulate extensive world knowledge beneficial for robotic manipulation in the form of task planning. However, the actual physical implementation of these plans often relies on task-specific learning methods, which require significant data collection and struggle with generalizability. In this work, we introduce Robotic Manipulation through Spatial Constraints of Parts (CoPa), a novel framework that leverages the common sense knowledge embedded within foundation models to generate a sequence of 6-DoF end-effector poses for open-world robotic manipulation. Specifically, we decompose the manipulation process into two phases: task-oriented grasping and task-aware motion planning. In the task-oriented grasping phase, we employ foundation vision-language models (VLMs) to select the object’s grasping part through a novel coarse-to-fine grounding mechanism. During the task-aware motion planning phase, VLMs are utilized again to identify the spatial geometry constraints of task-relevant object parts, which are then used to derive post-grasp poses. We also demonstrate how CoPa can be seamlessly integrated with existing robotic planning algorithms to accomplish complex, long-horizon tasks. Our comprehensive real-world experiments show that CoPa possesses a fine-grained physical understanding of scenes, capable of handling open-set instructions and objects with minimal prompt engineering and without additional training. Project page: copa-2024.github.io
Haoxu Huang, Fanqi Lin, Yingdong Hu, Yang Gao 0029
IROS3
2024 UAV Data Collection With Deep Reinforcement Learning for Grant-Free IoT
abstract
The utilization of unmanned aerial vehicles (UAVs) for efficient data collection has gained considerable attention. In this paper, we examine a scenario involving grant-free access from Internet of Things (IoT) devices, where the random access may cause packet collision, stemming from multiple devices concurrently transmitting data. To address this issue, we propose a deep reinforcement learning-based collision avoidance (DRL-CA) approach for UAV data collection, which optimizes the UAV trajectory. The approach assists UAVs in identifying and maximizing the acquisition of device packet in an environment characterized by probabilistic packet transmission and potential collisions among device packets while ensuring a timely arrival at the destination. Through simulations, our proposed method effectively mitigates unnecessary conflicts among device packets while achieving the optimization objective.
Jiale Zhong, Yingdong Hu, Ye Li 0004, Ruifeng Gao, Jue Wang 0006
WCNC2
2024 Transparent RIS: Wireless Coverage Enhancement via Region-Oriented Passive Beamforming
abstract
We investigate a new deployment form of reflective intelligent surface (RIS), which aims at enhancing the quality of service of a main communication system in a target region, while without the need of changing its transmission protocol and scheme (i.e., the RIS is “transparent” to the main system). To this end, we mathematically formulate a coverage enhancement problem, where a RIS is used transparently in the sense that the BS can be unaware of its existence, while the minimum channel link strength, measured from every BS antenna to any point in the target region, can be maximized. The formulated problem is non-convex with mixed discrete-continuous variables. To tackle this challenge, we recast it into a convex feasibility problem via spatial sampling and semi-definite relaxation. Based on a derived analytical upper bound on the link strength difference between any two location points, we further characterize the coverage-similarity region of a given location, and accordingly propose an improved spatial sampling scheme for efficient implementation. Simulation results show that the proposed transparent RIS design achieves better coverage performance than benchmark schemes. More importantly, it can effectively improve the communication performance without affecting the transmission scheme originally adopted by the main communication system.
Jue Wang 0006, Yingdong Hu, Ye Li 0004, Ruifeng Gao, Jun Zhang 0023, Yu Han 0004, Shi Jin 0002
IEEE Trans. Wirel. Commun.2
2023 For Pre-Trained Vision Models in Motor Control, Not All Policy Learning Methods are Created Equal
abstract
In recent years, increasing attention has been directed to leveraging pre-trained vision models for motor control. While existing works mainly emphasize the importance of this pre-training phase, the arguably equally important role played by downstream policy learning during control-specific fine-tuning is often neglected. It thus remains unclear if pre-trained vision models are consistent in their effectiveness under different control policies. To bridge this gap in understanding, we conduct a comprehensive study on 14 pre-trained vision models using 3 distinct classes of policy learning methods, including reinforcement learning (RL), imitation learning through behavior cloning (BC), and imitation learning with a visual reward function (VRF). Our study yields a series of intriguing results, including the discovery that the effectiveness of pre-training is highly dependent on the choice of the downstream policy learning algorithm. We show that conventionally accepted evaluation based on RL methods is highly variable and therefore unreliable, and further advocate for using more robust methods like VRF and BC. To facilitate more universal evaluations of pre-trained models and their policy learning methods in the future, we also release a benchmark of 21 tasks across 3 different environments alongside our work.
Yingdong Hu, Renhao Wang, Li Erran Li, Yang Gao 0029
ICML1
2023 Policy Contrastive Imitation Learning
abstract
Adversarial imitation learning (AIL) is a popular method that has recently achieved much success. However, the performance of AIL is still unsatisfactory on the more challenging tasks. We find that one of the major reasons is due to the low quality of AIL discriminator representation. Since the AIL discriminator is trained via binary classification that does not necessarily discriminate the policy from the expert in a meaningful way, the resulting reward might not be meaningful either. We propose a new method called Policy Contrastive Imitation Learning (PCIL) to resolve this issue. PCIL learns a contrastive representation space by anchoring on different policies and uses a smooth cosine-similarity-based reward to encourage imitation learning. Our proposed representation learning objective can be viewed as a stronger version of the AIL objective and provide a more meaningful comparison between the agent and the policy. From a theoretical perspective, we show the validity of our method using the apprenticeship learning framework. Furthermore, our empirical evaluation on the DeepMind Control suite demonstrates that PCIL can achieve state-of-the-art performance. Finally, qualitative results suggest that PCIL builds a smoother and more meaningful representation space for imitation learning.
Jialei Huang, Zhao-Heng Yin, Yingdong Hu, Yang Gao 0029
ICML3
2023 Low-Complexity Streaming Forward Erasure Correction for Non-Terrestrial Networks
abstract
As the 6G network is evolving towards a space-air-ground integrated scale with ubiquitous long-distance non-terrestrial network (NTN) links, packet-level streaming forward erasure correction (FEC), which can achieve low end-to-end in-order delivery delay over lossy links with long propagation delay, has drawn increasing interest. However, the existing streaming FEC has a problem that full-length encoding windows (EWs) including all non-acknowledged source packets are used when generating repair packets, which incurs high computational cost when the link’s bandwidth-delay product is large. To address the problem, this paper proposes a new low-complexity streaming FEC design, where a mixture of short and full-length EWs are used. We propose a novel method to analyze the decoding window width observed by arriving repair packets, which is based on the analysis of the busy period of a virtual queue using renewal theory. Later, using the analysis as the key enabler, a design problem is formulated and solved to optimize parameters including the EW width and the fraction of short-length repair packets such that the computational cost is reduced. Evaluations using real-life code implementations show that the proposed design can significantly reduce the computational cost, while maintaining the key benefits of the original streaming FEC.
Ye Li 0004, Yingdong Hu, Ruifeng Gao, Jue Wang 0006, Sheng Wu 0001
IEEE Trans. Commun.3
2022 Semantic-Aware Fine-Grained Correspondence
Yingdong Hu, Renhao Wang, Yang Gao 0029
ECCV (31)1
2019 On proactive eavesdropping using anti-relay-selection jamming in multi-relay communication systems
Yingdong Hu, Ruifeng Gao, Ye Li 0004, Shibing Zhang
Sci. China Inf. Sci.1
2018 Novel spectrum sensing and access in cognitive radio networks
Shibing Zhang, Yingdong Hu, Zhihua Bao
Sci. China Inf. Sci.2