VLDB 2026 Research / reviewers in the wild / expert
Yuran Yang
dblp:327/3762
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2026
0009-0004-5292-2469ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Research on Memory Algorithm Based on Time Series Model and Reinforcement LearningabstractABSTRACT Spaced repetition is a highly effective method of memorization that helps learners to remember large amounts of content efficiently. This paper presents a spaced repetition framework integrating time‐series modelling with reinforcement learning. We propose the GLD‐HLR model, which utilizes Discrete Cosine Transform (DCT) to decouple multi‐scale temporal features in the frequency domain and a Legendre Projection Unit (LPU) to represent continuous memory trajectories via orthogonal basis functions. This architecture significantly reduces computational complexity while enhancing responsiveness to non‐linear memory changes. Furthermore, a PPO‐MMC algorithm is developed to optimize review intervals within a continuous state space. By achieving joint learning of memory prediction and policy scheduling, the framework effectively minimizes review costs while maximizing long‐term retention. This paper validated through ablation and comparative experiments that the mean absolute error (MAE) of the GLD‐HLR model's recall probability predictions remained below 0.03, achieving at least a 4% reduction compared to the LSTM‐HLR model. The mean absolute percentage error (MAPE) for half‐life predictions was below 0.2, which is smaller than the prediction errors of all other models. The PPO‐MMC algorithm achieved a cumulative number of words learned (WTL) exceeding 8000 within 1000 days, with the number of words memorized at the target half‐life (THR) surpassing 7000. This indicates that the algorithm can efficiently help learners master a large number of vocabulary words within a limited time frame and achieve long‐term retention. Long Shao, Yaxiu Qiao, Yuran Yang |
Expert Syst. J. Knowl. Eng. | 4 |
| 2025 | MCOP: Multi-UAV Collaborative Occupancy PredictionabstractUnmanned Aerial Vehicle (UAV) swarm systems necessitate efficient collaborative perception mechanisms for diverse operational scenarios. Current Bird's Eye View (BEV)-based approaches exhibit two main limitations: bounding-box representations fail to capture complete semantic and geometric information of the scene, and their performance significantly degrades when encountering undefined or occluded objects. To address these limitations, we propose a novel multi-UAV collaborative occupancy prediction framework. Our framework effectively preserves 3D spatial structures and semantics through integrating a Spatial-Aware Feature Encoder and Cross-Agent Feature Integration. To enhance efficiency, we further introduce Altitude-Aware Feature Reduction to compactly represent scene information, along with a Dual-Mask Perceptual Guidance mechanism to adaptively select features and reduce communication overhead. Due to the absence of suitable benchmark datasets, we extend three datasets for evaluation: two virtual datasets (Air-to-Pred-Occ and UAV3D-Occ) and one real-world dataset (GauUScene-Occ). Experiments results demonstrate that our method achieves state-of-the-art accuracy, significantly outperforming existing collaborative methods while reducing communication overhead to only a fraction of previous approaches. Zefu Lin, Xiaojuan Jin, Yuran Yang, Lue Fan, Zhaoxiang Zhang 0001 |
ICCV | 4 |
| 2025 | TC-Light: Temporally Coherent Generative Rendering for Realistic World TransferabstractIllumination and texture rerendering are critical dimensions for world-to-world transfer, which is valuable for applications including sim2real and real2real visual data scaling up for embodied AI. Existing techniques generatively re-render the input video to realize the transfer, such as video relighting models and conditioned world generation models. Nevertheless, these models are predominantly limited to the domain of training data (e.g., portrait) or fall into the bottleneck of temporal consistency and computation efficiency, especially when the input video involves complex dynamics and long durations. In this paper, we propose **TC-Light**, a novel paradigm characterized
by the proposed two-stage post optimization mechanism. Starting from the video preliminarily relighted by an inflated video relighting model, it optimizes appearance embedding in the first stage to align global illumination. Then it optimizes the proposed canonical video representation, i.e., **Unique Video Tensor (UVT)**, to align fine-grained texture and lighting in the second stage. To comprehensively evaluate performance, we also establish a long and highly dynamic video benchmark. Extensive experiments show that our method enables physically plausible re-rendering results with superior temporal coherence and low computation cost. The code and video demos are available at our [Project Page](https://dekuliutesla.github.io/tclight/). Yang Liu 0347, Chuanchen Luo, Zimo Tang, Yingyan Li, Yuran Yang, Yuanyong Ning, Lue Fan, Junran Peng, Zhaoxiang Zhang 0001 |
NeurIPS | 5 |
| 2024 | Fully Data-Driven Pseudo Label Estimation for Pointly-Supervised Panoptic SegmentationabstractThe core of pointly-supervised panoptic segmentation is estimating accurate dense pseudo labels from sparse point labels to train the panoptic head. Previous works generate pseudo labels mainly based on hand-crafted rules, such as connecting multiple points into polygon masks, or assigning the label information of labeled pixels to unlabeled pixels based on the artificially defined traversing distance. The accuracy of pseudo labels is limited by the quality of the hand-crafted rules (polygon masks are rough at object contour regions, and the traversing distance error will result in wrong pseudo labels). To overcome the limitation of hand-crafted rules, we estimate pseudo labels with a fully data-driven pseudo label branch, which is optimized by point labels end-to-end and predicts more accurate pseudo labels than previous methods. We also train an auxiliary semantic branch with point labels, it assists the training of the pseudo label branch by transferring semantic segmentation knowledge through shared parameters. Experiments on Pascal VOC and MS COCO demonstrate that our approach is effective and shows state-of-the-art performance compared with related works. Codes are available at https://github.com/BraveGroup/FDD. Jing Li 0112, Junsong Fan, Yuran Yang, Shuqi Mei, Jun Xiao 0005, Zhaoxiang Zhang 0001 |
AAAI | 3 |
| 2024 | MemoNav: Working Memory Model for Visual NavigationabstractImage-goal navigation is a challenging task that requires an agent to navigate to a goal indicated by an image in unfamiliar environments. Existing methods utilizing diverse scene memories suffer from inefficient exploration since they use all historical observations for decision-making without considering the goal-relevant fraction. To address this limitation, we present MemoNav, a novel memory model for image-goal navigation, which utilizes a working memory-inspired pipeline to improve navigation performance. Specifically, we employ three types of navigation memory. The node features on a map are stored in the short-term memory (STM), as these features are dynamically updated. A forgetting module then retains the informative STM fraction to increase efficiency. We also introduce long-term memory (LTM) to learn global scene representations by progressively aggregating STM features. Subsequently, a graph attention module encodes the retained STM and the LTM to generate working memory (WM) which contains the scene features essential for efficient navigation. The synergy among these three memory types boosts navigation performance by enabling the agent to learn and leverage goal-relevant scene features within a topological map. Our evaluation on multi-goal tasks demonstrates that MemoNav significantly outperforms previous methods across all difficulty levels in both Gibson and Matterport3D scenes. Qualitative results further illustrate that MemoNav plans more efficient routes. Xu Yang 0004, Yuran Yang, Shuqi Mei, Zhaoxiang Zhang 0001 |
CVPR | 4 |
| 2024 | OpenSatMap: A Fine-grained High-resolution Satellite Dataset for Large-scale Map ConstructionabstractIn this paper, we propose OpenSatMap, a fine-grained, high-resolution satellite dataset for large-scale map construction. Map construction is one of the foundations of the transportation industry, such as navigation and autonomous driving. Extracting road structures from satellite images is an efficient way to construct large-scale maps. However, existing satellite datasets provide only coarse semantic-level labels with a relatively low resolution (up to level 19), impeding the advancement of this field. In contrast, the proposed OpenSatMap (1) has fine-grained instance-level annotations; (2) consists of high-resolution images (level 20); (3) is currently the largest one of its kind; (4) collects data with high diversity. Moreover, OpenSatMap covers and aligns with the popular nuScenes dataset and Argoverse 2 dataset to potentially advance autonomous driving technologies. By publishing and maintaining the dataset, we provide a high-quality benchmark for satellite-based map construction and downstream tasks like autonomous driving. Hongbo Zhao 0006, Lue Fan, Yuntao Chen, Yuran Yang, Xiaojuan Jin, Gaofeng Meng, Zhaoxiang Zhang 0001 |
NeurIPS | 5 |
| 2023 | Informative Data Mining for One-shot Cross-Domain Semantic SegmentationabstractContemporary domain adaptation offers a practical solution for achieving cross-domain transfer of semantic segmentation between labelled source data and unlabeled target data. These solutions have gained significant popularity; however, they require the model to be retrained when the test environment changes. This can result in unbearable costs in certain applications due to the time-consuming training process and concerns regarding data privacy. One-shot domain adaptation methods attempt to overcome these challenges by transferring the pre-trained source model to the target domain using only one target data. Despite this, the referring style transfer module still faces issues with computation cost and over-fitting problems. To address this problem, we propose a novel framework called Informative Data Mining (IDM) that enables efficient one-shot domain adaptation for semantic segmentation. Specifically, IDM provides an uncertainty-based selection criterion to identify the most informative samples, which facilitates quick adaptation and reduces redundant training. We then perform a model adaptation method using these selected samples, which includes patch-wise mixing and prototype-based information maximization to update the model. This approach effectively enhances adaptation and mitigates the overfitting problem. In general, we provide empirical evidence of the effectiveness and efficiency of IDM. Our approach outperforms existing methods and achieves a new state-of-the-art one-shot performance of 56.7%/55.4% on the GTA5/SYNTHIA to Cityscapes adaptation tasks, respectively. The code will be released at https://github.com/yxiwang/IDM. Yuxi Wang 0001, Jian Liang 0001, Jun Xiao 0005, Shuqi Mei, Yuran Yang, Zhaoxiang Zhang 0001 |
ICCV | 5 |
| 2023 | SSF: Accelerating Training of Spiking Neural Networks with Stabilized Spiking FlowabstractSurrogate gradient (SG) is one of the most effective approaches for training spiking neural networks (SNNs). While assisting SNNs to achieve classification performance comparable to artificial neural networks, SG suffers from the problem of time-consuming training, preventing it from efficient learning. In this paper, we formally analyze the backward process of classic SG and find that the membrane accumulation through time leads to exponential growth of training time. With this discovery, we propose Stabilized Spiking Flow (SSF), a simple yet effective approach to accelerate training of SG-based SNNs. For each spiking neuron, SSF averages its input and output activations over time to yield stabilized input and output, respectively. Then, instead of back propagating all errors that are related to current neuron and inherently entangled in time domain, the auxiliary gradient is directly propagated from the stabilized output to input through a devised relationship mapping. Additionally, SSF method is suitable to different neuron models. Extensive experiments on both static and neuromorphic datasets demonstrate that SNNs trained with SSF approach can achieve performance comparable to the original counterparts, while reducing the training time significantly. In particular, SSF speeds up the training process of state-of-the-art SNN models up to 10× when time steps equal to 80. Zengjie Song, Yuxi Wang 0001, Jun Xiao 0005, Yuran Yang, Shuqi Mei, Zhaoxiang Zhang 0001 |
ICCV | 5 |
| 2023 | Coarse Mask Guided Interactive Object SegmentationabstractInteractive object segmentation aims to produce object masks with user interactions, such as clicks, bounding boxes, and scribbles. Click point is the most popular interactive cue for its efficiency, and related deep learning methods have attracted lots of interest in recent years. Most works encode click points as gaussian maps and concatenate them with images as the model's input. However, the spatial and semantic information of gaussian maps would be noised through multiple convolution layers and won't be fully exploited by top layers for mask prediction. To pass click information to top layers exactly and efficiently, we propose a coarse mask guided model (CMG) which predicts coarse masks with a coarse module to guide the object mask prediction. Specifically, the coarse module encodes user clicks as query features and enriches their semantic information with backbone features through transformer layers, coarse masks are generated based on the enriched query feature and fed into CMG's decoder. Benefiting from the efficiency of transformer, CMG's coarse module and decoder module are lightweight and computationally efficient, making the interaction process more smooth. Experiments on several segmentation benchmarks demonstrate the effectiveness of our method, and we get new state-of-the-art results compared with previous works. Jing Li 0112, Junsong Fan, Yuxi Wang 0001, Yuran Yang, Zhaoxiang Zhang 0001 |
IEEE Trans. Image Process. | 4 |