Yirong Yang

dblp:35/4213 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0008-6154-1915ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Robot navigation and mapping · 24% Image recognition and object detection · 20% 3D vision · 18%
Databases, data mining, and information retrieval
2 papers
Data mining · 70% Information retrieval · 15% Indexing and storage engines · 15%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
1.012026
TAS-DAQ: Task-Adaptive Sparse Prediction With Dense Query Auxiliary Supervisory for Efficient 3D Object Detection · IEEE Trans. Multim. 2026
Natural language and speech › Language models and text generation
instruction following
1.012026
UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human Trajectories · AAAI 2026
Robotics › Robot navigation and mapping › visual navigation
language-guided navigation
1.012026
UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human Trajectories · AAAI 2026
Computer vision › Image recognition and object detection
object detection
1.012026
Active Style-Content Dual-Branch Domain Adaptation for Semi-Supervised SAR Object Detection · IEEE Trans. Image Process. 2026
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles
1.012026
TAS-DAQ: Task-Adaptive Sparse Prediction With Dense Query Auxiliary Supervisory for Efficient 3D Object Detection · IEEE Trans. Multim. 2026
Computer vision › 3D vision › 3d object detection
query-based 3d object detection
1.012026
TAS-DAQ: Task-Adaptive Sparse Prediction With Dense Query Auxiliary Supervisory for Efficient 3D Object Detection · IEEE Trans. Multim. 2026
Computer vision › Image recognition and object detection › object detection › remote sensing object detection
SAR object detection
1.012026
Active Style-Content Dual-Branch Domain Adaptation for Semi-Supervised SAR Object Detection · IEEE Trans. Image Process. 2026
Machine learning › Transfer learning and domain adaptation › domain adaptation › low-resource domain adaptation
semi-supervised domain adaptation
1.012026
Active Style-Content Dual-Branch Domain Adaptation for Semi-Supervised SAR Object Detection · IEEE Trans. Image Process. 2026
Machine learning › Learning paradigms › lifelong learning
anti-forgetting adaptation
0.912025
C-NAV: Towards Self-Evolving Continual Object Navigation in Open World · NeurIPS 2025
Robotics › Robot navigation and mapping
object goal navigation
0.912025
C-NAV: Towards Self-Evolving Continual Object Navigation in Open World · NeurIPS 2025
Robotics › Robot navigation and mapping
visual navigation
0.912025
C-NAV: Towards Self-Evolving Continual Object Navigation in Open World · NeurIPS 2025
Computer vision › Image recognition and object detection › image classification
remote sensing image classification
0.312026
Active Style-Content Dual-Branch Domain Adaptation for Semi-Supervised SAR Object Detection · IEEE Trans. Image Process. 2026
Robotics › Autonomous driving
urban scene understanding
0.312026
UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human Trajectories · AAAI 2026
Data mining › pattern mining › tree mining
frequent subtree mining
0.122005
Mining Closed and Maximal Frequent Subtrees from Databases of Labeled Rooted Trees · IEEE Trans. Knowl. Data Eng. 2005
Indexing and Mining Free Trees · ICDM 2003
Data mining
pattern mining
0.122005
Mining Closed and Maximal Frequent Subtrees from Databases of Labeled Rooted Trees · IEEE Trans. Knowl. Data Eng. 2005
Indexing and Mining Free Trees · ICDM 2003
Information retrieval
indexing
0.012003
Indexing and Mining Free Trees · ICDM 2003
Indexing and storage engines
tree index
0.012003
Indexing and Mining Free Trees · ICDM 2003

Methods — techniques the papers use, named apart from their topics

web-scale video annotation · 1.0temporal fusion · 1.0sparse queries · 1.0navigation policy learning · 1.0image fusion · 1.0feature alignment · 1.0active sampling · 1.0BEV query · 1.0feature distillation · 0.9adaptive sampling · 0.9heuristic ordering · 0.1enumeration tree pruning · 0.1tree indexing · 0.0canonical string · 0.0canonical form · 0.0
YearPublicationVenuePosition
2026 UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human Trajectories
abstract
Navigating complex urban environments using natural language instructions poses significant challenges for embodied agents, including noisy language instructions, ambiguous spatial references, diverse landmarks, and dynamic street scenes. Current visual navigation methods are typically limited to simulated or off-street environments, and often rely on precise goal formats, such as specific coordinates or images. This limits their effectiveness for autonomous agents like last-mile delivery robots navigating unfamiliar cities. To address these limitations, we introduce UrbanNav, a scalable framework that trains embodied agents to follow free-form language instructions in diverse urban settings. Leveraging web-scale city walking videos, we develop an scalable annotation pipeline that aligns human navigation trajectories with language instructions grounded in real-world landmarks. UrbanNav encompasses over 1,500 hours of navigation data and 3 million instruction-trajectory-landmark triplets, capturing a wide range of urban scenarios. Our model learns robust navigation policies to tackle complex urban scenarios, demonstrating superior spatial reasoning, robustness to noisy instructions, and generalization to unseen urban settings. Experimental results show that UrbanNav significantly outperforms existing methods, highlighting the potential of large-scale web video data to enable language-guided, real-world urban navigation for embodied agents.
Yanghong Mei, Yirong Yang, Longteng Guo, Qunbo Wang, Ming-Ming Yu, Xingjian He, Wenjun Wu 0001, Jing Liu 0001
AAAI2
2026 VQ-SSR: Offline Safe Reinforcement Learning via Discrete Skill Quantization
Diyuan Hou, Yirong Yang, Jiazhi Zhang, Wenjun Wu 0001
ICIC (2)2
2026 Active Style-Content Dual-Branch Domain Adaptation for Semi-Supervised SAR Object Detection
abstract
Synthetic Aperture Radar (SAR) images offer unique advantages in all-weather, all-day remote sensing, but the high acquisition costs and time-consuming annotation processes limit their widespread implementation. Semi-supervised domain adaptation leverages abundant annotated optical images and a small number of labeled SAR images to achieve great performance on SAR images. However, existing semi-supervised domain adaptation object detection methods typically select SAR domain labeled samples randomly, making it difficult to fully exploit the valuable information and distinctive features inherent in the target domain data. Moreover, there is a significant style and content gap between optical and SAR images, and previous methods have not adapted to them in a task-specific manner. To this end, this paper proposes an active style-content dual-branch domain adaptation method specifically designed for semi-supervised object detection in SAR images. The proposed approach employs Task-aware Active Sampling (TAS) module to select the most valuable SAR samples, addressing inefficiencies in random sampling. Also, we employ a dual-branch framework to address the style and content gaps between optical and SAR images. Multi-layer Feature Alignment (MFA) module ensures style alignment by maintaining consistent feature representations across different visual styles, while Gaussian-SAM Image Fusion (G-SIF) module is employed to integrate content from the source domain into the target domain, effectively bridging the gap between optical and SAR images. Extensive experiments on multiple ship and aircraft datasets demonstrate the exceptional generalization capabilities of our proposed model.
Xi Yang 0011, Quantao Xie, Yirong Yang, Nannan Wang 0001
IEEE Trans. Image Process.3
2026 Implicitly Defined Material Decomposition Estimator and Learned Physics-Informed Neural Proxy for Photon Counting CT
abstract
Photon counting detector-based CT (PCCT) systems provide spectral count measurements, enabling material decomposition (MD) for quantitative imaging. Maximum-likelihood estimation (MLE) for MD offers asymptotically unbiased and efficient (minimum variance) results but is usually solved iteratively, making the entire process computationally expensive and time-consuming. Conversely, representative empirical methods relying on calibration aim to construct a direct measurement-decomposition conversion, which can be fast but may suffer from bias or noise amplification. In this work, we show that the iterative MLE method implicitly defines the functional mapping from measurements to estimates, i.e., MD results, and the corresponding mean and noise yield analytical approximation forms from the Implicit Function Theorem. From this perspective, we demonstrate that it is possible to distill knowledge from the implicit function defined by the iterative MLE, i.e., finding the explicit proxy, by leveraging universal approximators such as neural networks and the derivative-aware Sobolev Training paradigm. We show that the proposed method, namely Proxy MD, is both computationally efficient (providing >200 times speedup) and approaches the performance of iterative MLE. Thus it outperforms conventional empirical methods, enabling high-quality real-time quantitative spectral imaging. Furthermore, we also demonstrate that the theoretical Jacobian analysis provides new perspectives in making iterative MD differentiable, enabling differentiable PCCT quantitative imaging and corresponding cross-domain end-to-end training and optimization. The code has been made available at: https://github.com/senwang320/ProxyMD_Demo.
Yirong Yang, Fredrik Grönberg, Grant M. Stevens, Adam S. Wang
IEEE Trans. Medical Imaging2
2026 TAS-DAQ: Task-Adaptive Sparse Prediction With Dense Query Auxiliary Supervisory for Efficient 3D Object Detection
abstract
Detecting 3D objects from surround-view images focuses on capturing the spatio-temporal positions of the surrounding environment, serving as a pivotal capability for vision-centric autonomous driving and robotics. While existing approaches primarily employ either dense BEV queries or sparse 3D queries, both paradigms have inherent limitations: dense queries suffer from redundant feature interactions and optimization conflicts, while sparse queries rely on high-quality initialization and struggle with error propagation in complex scenarios. To address these challenges, we proposeTAS-DAQ, a novel two-stage framework that synergizes dense and sparse query strategies. In Stage I, we generate geometry-aware coarse queries through the BEV feature providing robust initialization, thereby ensuring robust query initialization with explicit 3D priors. Stage II introduces a learnable Query Bank with temporal fusion to iteratively refine sparse queries by capturing discriminative instance features across views and frames. Moreover, considering the optimization conflicts caused by redundant query interactions in dense paradigms, we introduce adaptive query aggregation in the query bank that dynamically prioritizes high-confidence queries from BEV features, effectively addressing query error propagation while enhancing instance-level representation consistency. Extensive experiments on the nuScenes R50 benchmark demonstrate state-of-the-art performance, achieving56.9 % NDSand46.1% mAP.
Yirong Yang, Qunbo Wang, Longteng Guo, Ruyi Ji, Ming-Ming Yu, Wenjun Wu 0001, Jing Liu 0001
IEEE Trans. Multim.1
2025 C-NAV: Towards Self-Evolving Continual Object Navigation in Open World
abstract
Embodied agents are expected to perform object navigation in dynamic, open-world environments. However, existing approaches typically rely on static trajectories and a fixed set of object categories during training, overlooking the real-world requirement for continual adaptation to evolving scenarios. To facilitate related studies, we introduce the continual object navigation benchmark, which requires agents to acquire navigation skills for new object categories while avoiding catastrophic forgetting of previously learned knowledge. To tackle this challenge, we propose C-Nav, a continual visual navigation framework that integrates two key innovations: (1) A dual-path anti-forgetting mechanism, which comprises feature distillation that aligns multi-modal inputs into a consistent representation space to ensure representation consistency, and feature replay that retains temporal features within the action decoder to ensure policy consistency. (2) An adaptive sampling strategy that selects diverse and informative experiences, thereby reducing redundancy and minimizing memory overhead. Extensive experiments across multiple model architectures demonstrate that C-Nav consistently outperforms existing approaches, achieving superior performance even compared to baselines with full trajectory retention, while significantly lowering memory requirements. The code will be publicly available at \url{https://bigtree765.github.io/C-Nav-project}.
Mingming Yu, Fei Zhu 0004, Wenzhuo Liu, Yirong Yang, Qunbo Wang, Wenjun Wu 0001, Jing Liu 0001
NeurIPS4
2025 Emulating Low-Dose PCCT Image Pairs With Independent Noise for Self-Supervised Spectral Image Denoising
abstract
Photon counting CT (PCCT) acquires spectral measurements and enables generation of material decomposition (MD) images that provide distinct advantages in various clinical situations. However, noise amplification is observed in MD images, and denoising is typically applied. Clean or high-quality references are rare in clinical scans, often making supervised learning (Noise2Clean) impractical. Noise2Noise is a self-supervised counterpart, using noisy images and corresponding noisy references with zero-mean, independent noise. PCCT counts transmitted photons separately, and raw measurements are assumed to follow a Poisson distribution in each energy bin, providing the possibility to create noise-independent pairs. The approach is to use binomial selection to split the counts into two low-dose scans with independent noise. We prove that the reconstructed spectral images inherit the noise independence from counts domain through noise propagation analysis and also validated it in numerical simulation and experimental phantom scans. The method offers the flexibility to split measurements into desired dose levels while ensuring the reconstructed images share identical underlying features, thereby strengthening the model's robustness for input dose levels and capability of preserving fine details. In both numerical simulation and experimental phantom scans, we demonstrated that Noise2Noise with binomial selection outperforms other common self-supervised learning methods based on different presumptive conditions.
Yirong Yang, Grant M. Stevens, Zhye Yin, Adam S. Wang
IEEE Trans. Medical Imaging2
2020 PointSpherical: Deep Shape Context for Point Cloud Learning in Spherical Coordinates
abstract
We propose Spherical Hierarchical modeling of 3D point cloud. Inspired by Shape Context, we design a receptive field on each 3D point by placing a spherical coordinate on it. We sample points using the furthest point method and creating overlapping balls of points. We divide the space into radial, polar angular, and azimuthal angular bins on which we form a Spherical Hierarchy for each ball. We apply 1x1 CNN convolution on points to start the initial feature extraction. Repeated 3D CNN and max-pooling over the Spherical bins propagate contextual information until all the information is condensed in the center bin. Extensive experiments on five datasets strongly evidence that our method outperforms current models on various Point Cloud Learning tasks, including 2D/3D shape classification, 3D part segmentation, and 3D semantic segmentation.
Bin Fan 0001, Yongcheng Liu, Yirong Yang, Jianbo Shi, Chunhong Pan, Huiwen Xie
ICPR4
2020 Deep Space Probing for Point Cloud Analysis
abstract
3D points distribute in a continuous 3D space irregularly, thus directly adapting 2D image convolution to 3D points is not an easy job. Previous works often artificially divide the space into regular grids, yet it could be suboptimal to learn geometry. In this paper, we propose SPCNN, namely, Space Probing Convolutional Neural Network, which naturally generalizes image CNN to deal with point clouds. The key idea of SPCNN is learning to probe the 3D space in an adaptive manner. Specifically, we define a pool of learnable convolutional weights, and let each point in the local region learn to choose a suitable convolutional weight from the pool. This is achieved by constructing a geometry guided index-mapping function that implicitly establishes a correspondence between convolutional weights and some local regions in the neighborhood (Fig. 1). In this way, the index-mapping function learns to adaptively partition nearby space for local geometry pattern recognition. With this convolution as a basic operator, SPCNN, a hierarchical architecture can be developed for effective point cloud analysis. Extensive experiments on challenging benchmarks across three tasks demonstrate that SPCNN achieves the state-of-the-art or has competitive performance.
Yirong Yang, Bin Fan 0001, Yongcheng Liu, Jiyong Zhang 0001, Xin Liu 0027, Xinyu Cai, Shiming Xiang, Chunhong Pan
ICPR1
2005 Canonical forms for labelled trees and their applications in frequent subtree mining
Yun Chi, Yirong Yang, Richard R. Muntz
Knowl. Inf. Syst.2
2005 Mining Closed and Maximal Frequent Subtrees from Databases of Labeled Rooted Trees
abstract
Tree structures are used extensively in domains such as computational biology, pattern recognition, XML databases, computer networks, and so on. One important problem in mining databases of trees is to find frequently occurring subtrees. Because of the combinatorial explosion, the number of frequent subtrees usually grows exponentially with the size of frequent subtrees and, therefore, mining all frequent subtrees becomes infeasible for large tree sizes. We present CMTreeMiner, a computationally efficient algorithm that discovers only closed and maximal frequent subtrees in a database of labeled rooted trees, where the rooted trees can be either ordered or unordered. The algorithm mines both closed and maximal frequent subtrees by traversing an enumeration tree that systematically enumerates all frequent subtrees. Several techniques are proposed to prune the branches of the enumeration tree that do not correspond to closed or maximal frequent subtrees. Heuristic techniques are used to arrange the order of computation so that relatively expensive computation is avoided as much as possible. We study the performance of our algorithm through extensive experiments, using both synthetic data and data sets from real applications. The experimental results show that our algorithm is very efficient in reducing the search space and quickly discovers all closed and maximal frequent subtrees.
Yun Chi, Yirong Yang, Richard R. Muntz
IEEE Trans. Knowl. Data Eng.3
2005 Correction to "Mining Closed and Maximal Frequent Subtrees from Databases of Labeled Rooted Trees"
Yun Chi, Yirong Yang, Richard R. Muntz
IEEE Trans. Knowl. Data Eng.3
2004 Application of continuous Hopfield network to solve the TSP
abstract
Traveling salesman problem (TSP) is a classic of difficult optimization problem. It is simple to describe, mathematically well characterized. But the actual best solution to TSP is computationally very hard, called a NP-complete problem. In this paper, continuous Hopfield network (CHN) is applied to solve TSP. The energy function to be minimized is determined both by constraints for a valid solution and by total length of touring path. Setting of parameters in energy function is crucial to the convergence and performance of the network. The role of each parameter is analyzed and criteria for choosing these parameters are described. Iterative computation algorithm of CHN is given. Computer simulation is conducted for 6-, City TSP. Some simulation results, such as convergence curve, iteration count, computation time, are used to evaluate this method.
Helei Wu, Yirong Yang
ICARCV2
2004 CMTreeMiner: Mining Both Closed and Maximal Frequent Subtrees
Yun Chi, Yirong Yang, Richard R. Muntz
PAKDD2
2004 HybridTreeMiner: An Efficient Algorithm for Mining Frequent Rooted Trees and Free Trees Using Canonical Form
Yun Chi, Yirong Yang, Richard R. Muntz
SSDBM2
2003 Indexing and Mining Free Trees
abstract
Tree structures are used extensively in domains such as computational biology, pattern recognition, computer networks, and so on. We present an indexing technique for free trees and apply this indexing technique to the problem of mining frequent subtrees. We first define a novel representation, the canonical form, for rooted trees and extend the definition to free trees. We also introduce another concept, the canonical string, as a simpler representation for free trees in their canonical forms. We then apply our tree indexing technique to the frequent subtree mining problem and present FreeTreeMiner, a computationally efficient algorithm that discovers all frequently occurring subtrees in a database of free trees. We study the performance and the scalability of our algorithms through extensive experiments based on both synthetic data and datasets from two real applications: a dataset of chemical compounds and a dataset of Internet multicast trees.
Yun Chi, Yirong Yang, Richard R. Muntz
ICDM2