Yadang Chen

dblp:75/8876 · DBLP profile ↗
← Back
38ranked-venue papers
21as first author
29since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 14 first-author · 18 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Select prompting with chain-of-thought paired with large language models
Xun Che, Wenjia Wu, Yadang Chen, Luanjuan Jiang, Qianmu Li
Expert Syst. Appl.3
2026 Implicit Alignment with Complementary Information for Text-based Person Re-identification
Guoqing Zhang 0002, Yadang Chen, Le Sun 0002, Yulin Cao, Yuhui Zheng
Knowl. Based Syst.3
2026 Target-agnostic common attributes learning for few-shot semantic segmentation
Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu
Pattern Recognit.1
2026 Enhanced image retrieval: Leveraging multi-head attention & multi-scale descriptors and hybrid aggregation feature indexing
Wenbin Yu 0002, Yadang Chen, Na Yin, Alex X. Liu
Signal Process. Image Commun.5
2026 Boosting Video Object Segmentation With Discriminative Core Features and Adaptive Position Refinement
Yadang Chen, Guolong Li, Yuhui Zheng, Bin Sheng 0001, Zhi-Xin Yang 0001, Enhua Wu
IEEE Trans. Circuits Syst. Video Technol.1
2026 Learnable Object Queries for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment unseen-category objects given only a few annotated samples. Although significant progress has been made in the field of FSS, selecting an appropriate feature matching method remains a challenge. Traditional prototype-based methods can preserve high-level semantic features, but they tend to lose detailed information. On the other hand, pixel-level comparison methods retain fine-grained details but are vulnerable to distractors and noise, leading to poor robustness. To address these issues, this paper proposes a target-agnostic object-based method. Specifically, we propose a set of learnable "object queries" to extract object features, which preserve both high-level semantic information and fine-grained details. Additionally, during the training phase, we exploit the prior knowledge of foreground and background embedded in the samples to enhance the model's performance. In the inference phase, the model utilizes both the support set and the learned prior knowledge to perform segmentation tasks, mitigating the data distribution bias caused by limited samples. Extensive experiments on benchmark datasets demonstrate that our method outperforms state-of-the-art approaches in both accuracy and robustness. Code is available at https://github.com/wenbo456/OTBNet.
Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu
IEEE Trans. Image Process.1
2026 Refinement on Both Foreground and Background Prototype for Few-Shot Segmentation
abstract
Although few-shot segmentation (FSS) methods have achieved remarkable results, there remain challenges associated with the limited number of support samples.i)The objects in support and query images may have substantially different appearances even though they belong to the same category, which is known as the prototype bias problem.ii)Most methods neglect the background information, especially the query background during the inference stage. To address these problems, we propose DPRNet, a novel network with dual branch of foreground and background prototype refinement modules. Specifically, we first present a Variational Feature Semantic Enhancement (VFSE) module, in which we refine the object prototype with a variational autoencoder and word-text labels. In this way, the biased class-wise prototype caused by the limited support samples can be aligned, achieving better performance. Second, we design a Background Prototype Refinement (BPR) module that effectively explores the potential information in the background for both the support and query images. More importantly, it is designed to generate online predictions of the query background during the training stage to fully mimic the inference stage. These advancements enhance the robustness and generalizability of our method, and the results of experiments demonstrate its effectiveness. In the 1-way 5-shot setting on PASCAL-$5^{i}$, our method achieves a mean-IoU improvement of 1.59% over the competing method.
Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu
IEEE Trans. Multim.1
2026 Multi-granularity cross-modal representation for occlusion-invariant group re-identification
Jiangxiangyu Lou, Xiaoshu Sun, Runtao Liu, Yadang Chen
Vis. Comput.7
2026 Hybrid token learning with bidirectional attention for few-shot semantic segmentation
Liting Lei, Yadang Chen, Jianlin Qiu
Vis. Comput.3
2025 Motion path planning method based on map reference points
Yadang Chen, Yangsong Li, Na Yin, Xiaolin Cen
Multim. Tools Appl.4
2025 Adaptive set-level metric for few-Shot image classification
Yadang Chen, Jin Wang 0005, Zhi-Xin Yang 0001
Neural Networks1
2024 Space-time Reinforcement Network for Video Object Segmentation
abstract
Recently, video object segmentation (VOS) networks typically use memory-based methods: for each query frame, the mask is predicted by space-time matching to memory frames. Despite these methods having superior performance, they suffer from two issues: 1) Challenging data can destroy the space-time coherence between adjacent video frames. 2) Pixel-level matching will lead to undesired mismatching caused by the noises or distractors. To address the aforementioned issues, we first propose to generate an auxiliary frame between adjacent frames, serving as an implicit short-temporal reference for the query one. Next, we learn a prototype for each video object and prototype-level matching can be implemented between the query and memory. The experiment demonstrated that our network outperforms the state-of-the-art method on the DAVIS 2017, achieving a ℐ&ℱ score of 86.4%, and attains a competitive result 85.0% on YouTube VOS 2018. In addition, our network exhibits a high inference speed of 32+ FPS.
Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
ICME1
2024 A Transformer-Based Adaptive Prototype Matching Network for Few-Shot Semantic Segmentation
Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu
IJCAI2
2024 Eliciting knowledge from language models with automatically generated continuous prompts
Yadang Chen, Duolin Wang, Dichao Li
Expert Syst. Appl.1
2024 Detecting fake information with knowledge-enhanced AutoPrompt
Xun Chen 0001, Yadang Chen, Qianmu Li
Neural Comput. Appl.3
2024 Cross-domain few-shot learning based on feature adaptive distillation
Dingwei Zhang, Yadang Chen, Dichao Li, Chuanyan Hao
Neural Comput. Appl.3
2024 Learning self-target knowledge for few-shot segmentation
Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
Pattern Recognit.1
2024 Prototype-wise self-knowledge distillation for few-shot segmentation
Yadang Chen, Chenchen Wei, Chuhan Lu
Signal Process. Image Commun.1
2024 Cross Modulation and Region Contrast Learning Network for Few-Shot Medical Image Segmentation
abstract
Few-shot segmentation is an important approach for mitigating issues related to data annotation and model generalization. However, when applied to medical images, two challenges arise. Firstly, the varying shapes and sizes of different organs result in intra-class variation, and current solutions often introduce a large number of parameters. Secondly, background information similar to the target organ category often leads to inter-class distractors. To address these challenges, this paper proposes a framework consisting of two modules. Among that, using the Cross Modulation Module, the connections between support and query images are explored, thereby enhancing their correlation and reducing the impact of intra-class variation. In the other hand, the Regional Prototype Contrast Module is used to introduce triplet loss to increase separation between prototypes of similar classes in the background and foreground, thereby mitigating the impact of inter-class distractors. Extensive organ segmentation experiments using abdominal magnetic resonance imaging (MRI) and computed tomography (CT) datasets demonstrate that the proposed model meets state-of-the-art performance benchmarks.
Kangting Tang, Shanjie Wang, Yadang Chen
IEEE Signal Process. Lett.3
2024 Boosting Video Object Segmentation via Robust and Efficient Memory Network
abstract
Recently, memory-based methods have exhibited remarkable performance in Video Object Segmentation (VOS) by employing non-local pixel-wise matching between the query and memory. Nevertheless, these methods suffer from two limitations: 1) Non-local pixel-wise matching can result in the incorrect segmentation of background distractor objects, and 2) memory features with substantial temporal redundancy consume significant computing resources and reduce the inference speed. To address the limitations, we first propose a local attention mechanism to suppress background features, and we introduce a novel training framework based on contrast learning to ensure the network learns reliable and robust pixel-wise correspondence between query and memory. We adaptively determine whether to update the memory based on the variation of foreground objects. Next, we propose a dynamic memory bank, which utilizes a lightweight and differentiable soft modulation gate to determine the number of memory features to remove along the temporal dimension. This allows efficient and flexible management of memory features. Our network achieves competitive results (e.g., 92.1% on DAVIS 2016 val, 87.6%/81.3% on DAVIS 2017 val/test, 87.0% on YouTube-VOS 2018 val) compared with the state-of-the-art methods while maintaining a faster inference speed of 25+FPS. Moreover, our network demonstrates a favorable balance between performance and speed when dealing with the long-time video dataset.
Yadang Chen, Dingwei Zhang, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu, Haixing Zhao
IEEE Trans. Circuits Syst. Video Technol.1
2024 Dual Branch Multi-Level Semantic Learning for Few-Shot Segmentation
abstract
Few-shot semantic segmentation aims to segment novel-class objects in a query image with only a few annotated examples in support images. Although progress has been made recently by combining prototype-based metric learning, existing methods still face two main challenges. First, various intra-class objects between the support and query images or semantically similar inter-class objects can seriously harm the segmentation performance due to their poor feature representations. Second, the latent novel classes are treated as the background in most methods, leading to a learning bias, whereby these novel classes are difficult to correctly segment as foreground. To solve these problems, we propose a dual-branch learning method. The class-specific branch encourages representations of objects to be more distinguishable by increasing the inter-class distance while decreasing the intra-class distance. In parallel, the class-agnostic branch focuses on minimizing the foreground class feature distribution and maximizing the features between the foreground and background, thus increasing the generalizability to novel classes in the test stage. Furthermore, to obtain more representative features, pixel-level and prototype-level semantic learning are both involved in the two branches. The method is evaluated on PASCAL-5i1-shot, PASCAL-5i5-shot, COCO-20i1-shot, and COCO-20i5-shot, and extensive experiments show that our approach is effective for few-shot semantic segmentation despite its simplicity.
Yadang Chen, Ren Jiang, Yuhui Zheng, Bin Sheng 0001, Zhi-Xin Yang 0001, Enhua Wu
IEEE Trans. Image Process.1
2023 Robust and Efficient Memory Network for Video Object Segmentation
abstract
This paper proposes a Robust and Efficient Memory Network, referred to as REMN, for studying semi-supervised video object segmentation (VOS). Memory-based methods have recently achieved outstanding VOS performance by performing non-local pixel-wise matching between the query and memory. However, these methods have two limitations. 1) Non-local matching could cause distractor objects in the background to be incorrectly segmented. 2) Memory features with high temporal redundancy consume significant computing resources. For limitation 1, we introduce a local attention mechanism that tackles the background distraction by enhancing the features of foreground objects with the previous mask. For limitation 2, we first adaptively decide whether to update the memory features depending on the variation of foreground objects to reduce temporal redundancy. Second, we employ a dynamic memory bank, which uses a lightweight and differentiable soft modulation gate to decide how many memory features need to be removed in the temporal dimension. Experiments demonstrate that our REMN achieves state-of-the-art results on DAVIS 2017, with a $\mathcal{J}\& \mathcal{F}$ score of 86.3% and on YouTube-VOS 2018, with a $\mathcal{G}$ over mean of 85.5%. Furthermore, our network shows a high inference speed of 25+ FPS and uses relatively few computing resources.
Yadang Chen, Dingwei Zhang, Zhi-Xin Yang 0001, Enhua Wu
ICME1
2023 Spatial constraint for efficient semi-supervised video object segmentation
Yadang Chen, Chuanjun Ji, Zhi-Xin Yang 0001, Enhua Wu
Comput. Vis. Image Underst.1
2023 Global video object segmentation with spatial constraint module
abstract
We present a lightweight and efficient semi-supervised video object segmentation network based on the space-time memory framework. To some extent, our method solves the two difficulties encountered in traditional video object segmentation: one is that the single frame calculation time is too long, and the other is that the current frame’s segmentation should use more information from past frames. The algorithm uses a global context (GC) module to achieve high-performance, real-time segmentation. The GC module can effectively integrate multi-frame image information without increased memory and can process each frame in real time. Moreover, the prediction mask of the previous frame is helpful for the segmentation of the current frame, so we input it into a spatial constraint module (SCM), which constrains the areas of segments in the current frame. The SCM effectively alleviates mismatching of similar targets yet consumes few additional resources. We added a refinement module to the decoder to improve boundary segmentation. Our model achieves state-of-the-art results on various datasets, scoring 80.1% on YouTube-VOS 2018 and a $${\cal J}{\rm{\& }}{\cal F}$$ score of 78.0% on DAVIS 2017, while taking 0.05 s per frame on the DAVIS 2016 validation dataset.
Yadang Chen, Duolin Wang, Zhi-Xin Yang 0001, Enhua Wu
Comput. Vis. Media1
2023 Video object segmentation through semantic visual words matching
Chuanyan Hao, Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
Multim. Tools Appl.2
2023 Spatio-temporal compression for semi-supervised video object segmentation
Chuanjun Ji, Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
Vis. Comput.2
2022 Fast target-aware learning for few-shot video object segmentation
Yadang Chen, Chuanyan Hao, Zhi-Xin Yang 0001, Enhua Wu
Sci. China Inf. Sci.1
2022 Meta-transfer-adjustment learning for few-shot learning
Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
J. Vis. Commun. Image Represent.1
2021 A new nonlocal means based framework for mixed noise removal
Jielin Jiang, Jian Yang 0003, Zhi-Xin Yang 0001, Yadang Chen, Lei Luo 0001
Neurocomputing5
2020 Higher-order potentials for video object segmentation in bilateral space
Chuanyan Hao, Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
Neurocomputing2
2019 Multilevel Model for Video Object Segmentation Based on Supervision Optimization
abstract
In this work, we present a supervised object segmentation algorithm for unconstrained video. Instead of arbitrarily picking a few frames for manual labeling, as in many existing supervised methods, the proposed method selects frames in a more reasonable manner, called supervision optimization. For this, we formulate a principled objective function by inferring the propagation error from appearance and motion clues. After this, we construct a multilevel segmentation model, which consists of low-level and high-level features. On the low level, image pixels are used for a more accurate estimation of motion and segmentation. On the high level, image segments are considered for a more semantic classification of the foreground and background. By integrating these in one segmentation graph, the result can be further improved by leveraging the knowledge from both levels. In experiments, the proposed approach is evaluated by different measures, and the results on a benchmark demonstrate the effectiveness in comparison with other state-of-the-art algorithms.
Yadang Chen, Chuanyan Hao, Alex X. Liu, Enhua Wu
IEEE Trans. Multim.1
2019 Appearance-consistent Video Object Segmentation Based on a Multinomial Event Model
abstract
In this study, we propose an effective and efficient algorithm for unconstrained video object segmentation, which is achieved in a Markov random field (MRF). In the MRF graph, each node is modeled as a superpixel and labeled as either foreground or background during the segmentation process. The unary potential is computed for each node by learning a transductive SVM classifier under supervision by a few labeled frames. The pairwise potential is used for the spatial-temporal smoothness. In addition, a high-order potential based on the multinomial event model is employed to enhance the appearance consistency throughout the frames. To minimize this intractable feature, we also introduce a more efficient technique that simply extends the original MRF structure. The proposed approach was evaluated in experiments with different measures and the results based on a benchmark demonstrated its effectiveness compared with other state-of-the-art algorithms.
Yadang Chen, Chuanyan Hao, Alex X. Liu, Enhua Wu
ACM Trans. Multim. Comput. Commun. Appl.1
2018 Efficient frame-sequential label propagation for video object segmentation
Yadang Chen, Chuanyan Hao, Wen Wu 0001, Enhua Wu
Multim. Tools Appl.1
2016 Robust dense reconstruction by range merging based on confidence estimation
Yadang Chen, Chuanyan Hao, Wen Wu 0001, Enhua Wu
Sci. China Inf. Sci.1
2015 Image completion with perspective constraint based on a single image
Chuanyan Hao, Yadang Chen, Wen Wu 0001, Enhua Wu
Sci. China Inf. Sci.2
2015 An iterated randomized search algorithm for large-scale texture synthesis and manipulations
Chuanyan Hao, Yadang Chen, Wen Wu 0001, Enhua Wu
Vis. Comput.2
2013 Secure P2P topology based on a multidimensional DHT space mapping
Zhixin Sun, Bingqing Luo, Yadang Chen, Kai Bu
Sci. China Inf. Sci.3
2013 Live accurate and dense reconstruction from a handheld camera
abstract
ABSTRACT We present a method to make an accurate and dense reconstruction from the input of video captured by a free moving handheld camera in real time. By the method firstly, the positions of the camera and sparse 3D points are estimated by simultaneous localization mapping. Then the depth maps of selected reference frames are computed from corresponding camera bundles. Lastly a novel linear algorithm is also proposed to integrate all the depth maps into dense meshes partially. The main contributions of this paper are in the following points: the reference frames and corresponding camera bundles are able to be selected automatically, then accurate and smooth depth maps are generated in real time, and the depth maps are merged into a dense mesh by using a linear algorithm based on the error clouds optimization. Our algorithm is implemented on dual CPU and graphics processing unit in a parallel framework for improving the performance. Copyright © 2013 John Wiley & Sons, Ltd.
Yadang Chen, Chuanyan Hao, Zhongmou Cai, Wen Wu 0001, Enhua Wu
Comput. Animat. Virtual Worlds1