Yuncong Yang

dblp:24/10763 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
abstract
Constructing compact and informative 3D scene representations is essential for effective embodied exploration and reasoning, especially in complex environments over extended periods. Existing representations, such as object-centric 3D scene graphs, oversimplify spatial relationships by modeling scenes as isolated objects with restrictive textual relationships, making it difficult to address queries requiring nuanced spatial understanding. Moreover, these representations lack natural mechanisms for active exploration and memory management, hindering their application to lifelong autonomy. In this work, we propose 3DMem, a novel 3D scene memory framework for embodied agents. 3D-Mem employs informative multi-view images, termed Memory Snapshots, to capture rich visual information of explored regions. It further integrates frontier-based exploration by introducing Frontier Snapshots—glimpses of unexplored areas—enabling agents to make decisions by considering both known and potential new information. To support lifelong memory in active exploration settings, we present an incremental construction pipeline for 3D-Mem, as well as a memory retrieval technique for memory management. Experimental results on three benchmarks demonstrate that 3D-Mem significantly enhances agents’ exploration and reasoning capabilities in 3D environments, highlighting its potential for advancing applications in embodied AI.
Yuncong Yang, Jiachen Zhou 0003, Peihao Chen, Yilun Du, Chuang Gan 0001
CVPR1
2025 MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
abstract
Spatial reasoning in 3D space is central to human cognition and indispensable for embodied tasks such as navigation and manipulation. However, state-of-the-art vision–language models (VLMs) struggle frequently with tasks as simple as anticipating how a scene will look after an egocentric motion: they perceive 2D images but lack an internal model of 3D dynamics. We therefore propose SpatialNavigator, a test-time scaling framework that grants a VLM with this missing capability by coupling it to a controllable world model based on video diffusion. The VLM iteratively sketches a concise camera trajectory, while the world model synthesizes the corresponding view at each step. The VLM then reasons over this multi-view evidence gathered during the interactive exploration. Without any fine-tuning, our SpatialNavigator achieves an average 7.7\% performance boost on the representative spatial reasoning benchmark SAT, showing that pairing VLMs with world models for test-time scaling offers a simple, plug-and-play route to robust 3D reasoning. Meanwhile, our method also improves upon the test-time inference VLMs trained through reinforcement learning, which demonstrates the potential of our method that utilizes world models for test-time scaling.
Yuncong Yang, Jiageng Liu, Reuben Tan, Yilun Du, Chuang Gan 0001
NeurIPS1
2025 Learning 3D Persistent Embodied World Models
abstract
The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing work has explored how to construct such world models using video models, they are often myopic in nature, without any memory of a scene not captured by currently observed images, preventing agents from making consistent long-horizon plans in complex environments where many parts of the scene are partially observed. We introduce a new persistent embodied world model with an explicit memory of previously generated content, enabling much more consistent long-horizon simulation. During generation time, our video diffusion model predicts RGB-D video of the future observations of the agent. This generation is then aggregated into a persistent 3D map of the environment. By conditioning the video model on this 3D spatial map, we illustrate how this enables video world models to faithfully simulate both seen and unseen parts of the world. Finally, we illustrate the efficacy of such a world model in downstream embodied applications, enabling effective planning and policy learning.
Yilun Du, Yuncong Yang, Peihao Chen, Dit-Yan Yeung, Chuang Gan 0001
NeurIPS3
2025 Learning-Based Privacy-Preserving Graph Publishing Against Sensitive Link Inference Attacks
abstract
Publishing graph data is widely desired to enable a variety of structural analyses and downstream tasks. However, it also potentially poses severe privacy leakage, as attackers may leverage the released graph data to launch attacks and precisely infer private information such as the existence of hidden sensitive links in the graph. Prior studies on privacy-preserving graph data publishing relied on heuristic graph modification strategies and it is difficult to determine the graph with the optimal privacy–utility trade-off for publishing. In contrast, we propose the first privacy-preserving graph structure learning framework against sensitive link inference attacks, named PPGSL, which can automatically learn a graph with the optimal privacy–utility trade-off. The PPGSL operates by first simulating a powerful surrogate attacker conducting sensitive link attacks on a given graph. It then trains a parameterized graph to defend against the simulated adversarial attacks while maintaining the favorable utility of the original graph. To learn the parameters of both parts of the PPGSL, we introduce a secure iterative training protocol. It can enhance privacy preservation and ensure stable convergence during the training process, as supported by the theoretical proof. Additionally, we incorporate multiple acceleration techniques to improve the efficiency of the PPGSL in handling large-scale graphs. The experimental results confirm that the PPGSL achieves state-of-the-art privacy–utility trade-off performance and effectively thwarts various sensitive link inference attacks.
Yucheng Wu 0002, Yuncong Yang, Xiao Han 0001, Leye Wang, Junjie Wu 0002
IEEE Trans. Inf. Forensics Secur.2
2024 Privacy-Preserving Network Embedding Against Private Link Inference Attacks
abstract
Network embedding represents network nodes by a low-dimensional informative vector. While it is generally effective for various downstream tasks, it may leak some private information of networks, such as hidden private links. In this work, we address a novel problem ofprivacy-preserving network embedding against private link inference attacks. Basically, we propose to perturb the original network by adding or removing links, and expect the embedding generated on the perturbed network can leak little information about private links but hold high utility for various downstream tasks. Towards this goal, we first propose general measurements to quantify privacy gain and utility loss incurred by candidate network perturbations; we then design aPrivacy-PreservingNetworkEmbedding (i.e., PPNE) framework to identify the optimal perturbation solution with the best privacy-utility trade-off in an iterative way. Furthermore, we propose many techniques to accelerate PPNE and ensure its scalability. For instance, as the skip-gram embedding methods including DeepWalk and LINE can be seen as matrix factorization with closed-form embedding results, we devise efficient privacy gain and utility loss approximation methods to avoid the repetitive time-consuming embedding training for every candidate network perturbation in each iteration. Experiments on real-life network datasets (with up to millions of nodes) verify that PPNE outperforms baselines by sacrificing less utility and obtaining higher privacy protection.
Xiao Han 0001, Yuncong Yang, Leye Wang, Junjie Wu 0002
IEEE Trans. Dependable Secur. Comput.2
2024 HyObscure: Hybrid Obscuring for Privacy-Preserving Data Publishing
abstract
Minimizing privacy leakage while ensuring data utility is a critical problem in a privacy-preserving data publishing task, from which data holders can boost platform engagements or enlarge data values. Most prior research concerned only with either privacy-insensitive or exact private data and resorts to a single obscuring method to achieve a privacy-utility tradeoff, which is inadequate for real-life hybrid data especially when facing machine learning-based inference attacks. This work takes a pilot study on privacy-preserving data publishing when both widely adopted generalization and obfuscation operations are employed for privacy-heterogeneous data protection. Specifically, we first propose novel measures for privacy and utility values quantification and formulate the hybrid privacy-preserving data obscuring problem to account for the joint effect of generalization and obfuscation. We then design a novel protection mechanism called HyObscure, which decomposes the original problem into three sub-problems to cross-iteratively optimize the hybrid operations for maximum privacy protection under a certain data utility guarantee. The convergence of the iterative process and the privacy leakage bound of HyObscure are also provided in theory. Extensive experiments demonstrate that HyObscure significantly outperforms a variety of state-of-the-art baseline methods when facing various inference attacks in different scenarios.
Xiao Han 0001, Yuncong Yang, Junjie Wu 0002, Hui Xiong 0001
IEEE Trans. Knowl. Data Eng.2
2023 TempCLR: Temporal Alignment Representation with Contrastive Learning
Yuncong Yang, Jiawei Ma, Shiyuan Huang 0001, Long Chen 0016, Xudong Lin 0003, Guangxing Han, Shih-Fu Chang
ICLR1
2022 Few-Shot End-to-End Object Detection via Constantly Concentrated Encoding Across Heads
Jiawei Ma, Guangxing Han, Shiyuan Huang 0001, Yuncong Yang, Shih-Fu Chang
ECCV (26)4
2022 MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge
abstract
Autonomous agents have made great strides in specialist domains like Atari games and Go. However, they typically learn tabula rasa in isolated environments with limited and manually conceived objectives, thus failing to generalize across a wide spectrum of tasks and capabilities. Inspired by how humans continually learn and adapt in the open world, we advocate a trinity of ingredients for building generalist agents: 1) an environment that supports a multitude of tasks and goals, 2) a large-scale database of multimodal knowledge, and 3) a flexible and scalable agent architecture. We introduce MineDojo, a new framework built on the popular Minecraft game that features a simulation suite with thousands of diverse open-ended tasks and an internet-scale knowledge base with Minecraft videos, tutorials, wiki pages, and forum discussions. Using MineDojo's data, we propose a novel agent learning algorithm that leverages large pre-trained video-language models as a learned reward function. Our agent is able to solve a variety of open-ended tasks specified in free-form language without any manually designed dense shaping reward. We open-source the simulation suite, knowledge bases, algorithm implementation, and pretrained models (https://minedojo.org) to promote research towards the goal of generally capable embodied agents.
Linxi Fan, Guanzhi Wang, Yunfan Jiang 0001, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, Anima Anandkumar
NeurIPS5
2014 Uniform color space based facial complexion recognition for Traditional Chinese Medicine
abstract
Face diagnosis of Traditional Chinese Medicine (TCM) is carried out by observing the facial complexion to obtain the disease diagnostic results. Color space based on human visual system will be more conducive to facial complexion recognition, which is more suitable to measure and distinguish facial complexion. Uniform color space based facial complexion recognition for TCM is proposed in this paper, which include: (1) the skin blocks in the human facial region are extracted by locating the eye position and mouth corner accurately; (2) the statistical characteristic of color histogram and the characteristic of aberration chromatic in Lab color space are introduced to extract the facial complexion feature; (3) the support vector machine (SVM) is used to evaluate the performance of facial complexion recognition. The experimental results show the proposed complexion feature can achieve good performance, with the facial complexion recognition rate up to 81%.
Jing Zhang 0023, Chao Wang 0017, Li Zhuo 0001, Yuncong Yang
ICARCV4
2014 Human Facial Complexion Recognition of Traditional Chinese Medicine Based on Uniform Color Space
abstract
Face diagnosis of Traditional Chinese Medicine (TCM) is carried out by observing the human facial complexion to obtain the disease diagnostic results. The morbidity of the organs can be revealed by the human facial complexion, so the color space based on human visual system will be more conducive to facial complexion recognition. It is much suitable to measure and distinguish facial complexion by uniform Lab color space, as it has the characteristic of isometry and high resolving power. First, the skin blocks in the human facial region are extracted by locating the eye position and mouth corner accurately. Second, the statistical characteristic of color histogram and the characteristic of aberration chromatic in Lab color space are introduced to extract the facial complexion feature. At last, the support vector machine (SVM) is used to evaluate the performance of facial complexion recognition. The experimental results show the complexion feature proposed in this paper can achieve the better performance, with the facial complexion recognition rate up to 81%.
Li Zhuo 0001, Yuncong Yang, Jing Zhang 0023
Int. J. Pattern Recognit. Artif. Intell.2
2013 An approach of bag-of-words based on visual attention model for pornographic images recognition in compressed domain
Jing Zhang 0023, Lei Sui, Li Zhuo 0001, Yuncong Yang
Neurocomputing5
2011 Regions of Interest Extraction Based on Visual Saliency in Compressed Domain
abstract
Recently bag-of-words (BoW) model having been widely used in textual information processing has been extended into many tasks in visual domain such as image classification, scene analysis, image annotation and image retrieval, namely bag-of-visual-words (BoVW) model. Therefore, it is essential to create an effective visual vocabulary. Most of existing approaches create visual vocabularies from image in pixel domain, which requires extra processing time in decompressed images, since most images are stored in compressed format. In this paper we propose to create a visual vocabulary based on Scale Invariant Feature Transform(SIFT) descriptor in compressed domain with the following three steps, (1) constructing low-resolution images in compressed domain, (2) extracting SIFT descriptor from low-resolution images, and (3) creating a visual vocabulary based on extracted SIFT descriptors. In order to evaluate the performance of the visual words, experiments have been conducted on identifying pornographic images. Experimental results indicate that the proposed method can recognize pornographic images accurately with much reduced computational time.
Lei Sui, Jing Zhang 0023, Li Zhuo 0001, Yuncong Yang
ISM4