EDBT 2026 Demo / reviewers in the wild / expert
Dongchao Wen
dblp:230/2355
· DBLP profile ↗
13ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0001-7311-1842ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 10 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Exemplar Prompt Learning via Bi-directional Visual-Semantic Alignment for Multi-Object TrackingabstractRecent multi-object tracking (MOT) approaches increasingly leverage pre-trained CLIP models to boost cross-domain generalization. A common strategy uses a predefined TrackBook—a closed-set of visual concepts—as textual prompts to guide learning of domain-invariant representations. However, these fixed prompts lack adaptive context, causing limited generalization. To address this limitation, this paper introduces a robust Exemplar Prompt Learning (EPL) framework via Bi-directional Visual-Semantic Alignment (BiVSA), termed EPL-MOT, which augments textual prompts with instance-aware contextual information derived during tracking. Specifically, an EPL module is designed to dynamically enrich textual prompts with contextual cues, enabling instance-specific adaptation without inducing category shift. Furthermore, a BiVSA module is proposed to deepen cross-modal interaction by incorporating bidirectional learnable prompts into both textual and visual branches. This facilitates progressive integration of global semantic features with local visual structures, resulting in a more effectively aligned visual-semantic space. Finally, to enhance robustness against distractors, a Category-guided Detection Query Generator (CDQG) is constructed, which incorporates base-class textual information to suppress irrelevant targets. Comprehensive evaluations on MOT17 and MOT20 demonstrate that the proposed EPL-MOT achieves competitive performance across both in-domain and cross-domain settings. Lingyan Liang, Gang Dong, Dongchao Wen, Kaihua Zhang 0001 |
ICMR | 4 |
| 2025 | Continuously Learning Video-level Object Tokens for Robust UAV trackingabstractDue to the dynamic changes in flight motion and viewpoint, the objects in unmanned aerial vehicle (UAV) tracking scenarios often suffer from drastic appearance variations. Existing UAV trackers often leverage a frame-level matching mechanism, which measures the appearance similarity between the object template and the search frame. The drastic object appearance variations degrade the learned model, leading to drift issue. To this end, this paper presents a video-level UAV tracking framework that focuses on Continuously Learning (CL) effective and efficient spatio-temporal object tokens for robust tracking, dubbed as CLTrack. Specifically, the CLTrack first learns a series of spatio-temporal object tokens via a dynamic filtering module (DFM), which encodes more consensus object appearance information from each frame. Afterwards, a spatio-temporal enhancement module (STEM) is designed via cascading a temporal and a spatial attention to fully interact with the selected tokens with stable long-range spatio-temporal context information of the tracked object. Finally, to ensure the learned model encodes the rich context information without catastrophic forgetting, a video-level tracking loss is designed to supervise feature learning from the whole video frames. Extensive experiments on three UAV benchmarks including UAV123, DTB70 and VisDrone2018 demonstrate that the proposed CLTrack achieves state-of-the-art performance. Shenglong Hu, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001 |
ICASSP | 5 |
| 2025 | Easy-to-hard Instance-level Feature Fusion for Co-saliency DetectionabstractExisting leading deep learning-based Co-saliency Detection (CoD) methods often learn the consensus features from the input image group without considering the complexity of each image. Despite the demonstrated success, the input images may contain hard samples with high complexity, e.g., those containing distractors that have similar appearance but different semantics to the co-salient objects. This is prone to mislead the learned model to treat these distractors as co-salient objects, leading to classification ambiguity. To address this issue, this paper presents an easy-to-hard instance-level feature Fusion framework for CoD, termed E2HCoD. The E2HCoD exploits the instance-level co-salient object consensus cues from the easy samples as reliable guidance to accurately fuse the co-salient object features in the hard samples. First, we design a Feature Filtering Module (FFM) that evaluates image complexity by integrating entropy, variance, texture, and edge density cues, allowing the model to select the easy samples with relatively easy backgrounds. Then, we develop an Easy-instance Embedding Branch (EEB), which accurately segments the co-salient object masks from the easy samples as the instance-level guidance to learn the accurate co-salient object consensus cues. Then, with the consensus knowledge from the easy samples as guidance, we construct an Easy-instance guided Fusion Branch (EFB), which fully interacts with the consensus features from the hard samples via a cross-attention mechanism, yielding the refined features that highlight the co-salient objects while suppressing the distractors. Finally, the refined features are fed into the decoder, generating a high-quality CoD prediction. Extensive experiments demonstrate that the proposed E2HCoD achieves state-of-the-art performance on CoSal2015, CoCA, and CoSOD3k. Chuang Ding, Zhidong Han, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001 |
ICASSP | 5 |
| 2025 | Spatio-Semantic Prompt guided Adaptive Segment Anything for Remote Sensing Change DetectionabstractExisting leading remote sensing change detection (RSCD) often takes a semantic-agnostic learning paradigm, which uses a binary ground-truth mask as supervision for model training. Despite the demonstrated success, due to the intrinsic characteristic of extremely complicated scene changes in RS images, this paradigm is prone to be misled by irrelevant semantic category changes, leading to a noisy CD mask prediction. To address this issue, this paper presents a Spatio-Semantic Prompt (SSP) guided adaptive Segment Anything Model (SAM) for RSCD, dubbed as SSP-SAM. The SSP-SAM introduces sparse textual and dense mask prompts into SAM to encode the task-specific semantic knowledge for RSCD. Specifically, we first encode the powerful textual semantic knowledge using Contrastive Language-Image Pre-training (CLIP) to determine the desired change semantic category. Then, we design a spatial dense prompt module that yields an attention map as prompt features to further refine the desired changed regions. Subsequently, we fine-tune the SAM through an adaptor to integrate the spatial-semantic prompt cues, yielding a coarse CD mask prediction. Finally, guided by the coarse CD mask, a multi-scale mask attention mechanism is adopted to learn the refined semantic representations of the changed targets, predicting the accurate CD mask. Extensive experiments on a variety of benchmark datasets demonstrate that the proposed SSP-SAM achieves state-of-the-art performance. Shenglong Hu, Zhidong Han, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001 |
ICASSP | 5 |
| 2025 | Unsupervised evaluation for out-of-distribution detection
Jiani Hu, Dongchao Wen, Weihong Deng |
Pattern Recognit. | 3 |
| 2024 | Enhancing Generalization Of Invisible Facial Privacy Cloak Via Gradient AccumulationabstractThe blooming of social media and face recognition (FR) systems has increased people’s concern about privacy and security. A new type of adversarial privacy cloak (class-universal) can be applied to all the images of regular users, to prevent malicious FR systems from acquiring their identity information. In this work, we discover the optimization dilemma in the existing methods – the local optima problem in large-batch optimization and the gradient information elimination problem in small-batch optimization. To solve these problems, we propose Gradient Accumulation (GA) to aggregate multiple small-batch gradients into a one-step iterative gradient to enhance the gradient stability and reduce the usage of quantization operations. Experiments show that our proposed method achieves high performance on the Privacy-Commons dataset against black-box face recognition models. Xuannan Liu, Yaoyao Zhong, Weihong Deng, Hongzhi Shi, Xingchen Cui, Yunfeng Yin, Dongchao Wen |
ICASSP | 7 |
| 2022 | Dynamic Training Data Dropout for Robust Deep Face RecognitionabstractLearning with noise is a practically challenging problem in deep face recognition. Despite the success of large margin softmax loss functions, these methods are designed for clean face databases. Considering the inevitable noise in the large scale databases, we first analyze the performance of noise in the training databases. For noise-robust deep face recognition, we propose a dynamic training data dropout (DTDD) method to dynamically filter the noise in the training database and gradually form a stable refined database for model learning. Specifically, we leverage the information provided by the model predictions of accumulated training epochs, which can distinguish regular samples and noise effectively and accurately. The proposed DTDD method is easy and stable for implementation, and can be combined with existing state-of-the-art loss functions and network architectures. Extensive experiments on CASIA-WebFace, VGGFace2, and MS-Celeb-1 M databases empirically demonstrate that our proposed method can robustly train deep face recognition models in the presence of label noise and low quality images. Yaoyao Zhong, Weihong Deng, Han Fang 0002, Jiani Hu, Dongyue Zhao, Dongchao Wen |
IEEE Trans. Multim. | 7 |
| 2021 | Augmented Face Representation Learning via Transitive DistillationabstractThe wild face of large variations is hard to recognize in unconstrained scenarios. To tackle this issue, existing works synthesize and augment the variation-specific faces for recognition. However, directly feeding generated samples results in negative transfer, because the feature spaces are shifted compared with normal samples. Instead, we propose a transitive distillation network (TDNet) that introduces a transitive domain to transfer cross-variation representations, which alleviates the negative influence of synthesized data. Specifically, data of diverse variations are firstly synthesized. Then we construct distributions from different variations as teachers to distill student. The negative transfer is mitigated by adopting adaptor as a bridge to break large domain distance. To handle faces of different quality, we propose a novel strategy to define easy and hard samples, which are utilized to select specific transitive status. Meanwhile, bilateral classification with curriculum learning is proposed to improve confidence of synthesized data gradually, enhancing the robustness of representation learning. Experiments show that our method achieves superiority on unconstrained face benchmarks such as IJB-C and SCface, while maintaining competence on general test sets. Han Fang 0002, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen |
FG | 7 |
| 2021 | Adaptive Label Noise Cleaning with Meta-Supervision for Deep Face RecognitionabstractThe training of a deep face recognition system usually faces the interference of label noise in the training data. However, it is difficult to obtain a high-precision cleaning model to remove these noises. In this paper, we propose an adaptive label noise cleaning algorithm based on meta-learning for face recognition datasets, which can learn the distribution of the data to be cleaned and make automatic adjustments based on class differences. It first learns re-liable cleaning knowledge from well-labeled noisy data, then gradually transfers it to the target data with meta-supervision to improve performance. A threshold adapter module is also proposed to address the drift problem in transfer learning methods. Extensive experiments clean two noisy in-the-wild face recognition datasets and show the effectiveness of the proposed method to reach state-of-the-art performance on the IJB-C face recognition benchmark. Yaobin Zhang, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen |
ICCV | 7 |
| 2021 | FTAFace: Context-enhanced Face Detector with Fine-grained Task AttentionabstractIn face detection, it is a common strategy to treat samples differently according to their difficulty for balancing training data distribution. However, we observe that widely used sampling strategies, such as OHEM and Focal loss, can lead to the performance imbalance between different tasks (e.g., classification and localization). Through analysis, we point out that, due to the driving of classification information, these sample-based strategies are difficult to coordinate the attention of different tasks during the training, thus leading to the above imbalance. Accordingly, we first confirm this by shifting the attention from the sample level to the task level. Then, we propose a fine-grained task attention method, a.k.a FTA, including inter-task importance and intra-task importance, which adaptively adjusts the attention of each item in the task from both global and local perspectives, so as to achieve finer optimization. In addition, we introduce transformer as a feature enhancer to assist our convolution network, and propose a context enhancement transformer, a.k.a CET, to mine the spatial relationship in the features towards more robust feature representation. Extensive experiments on WiderFace and FDDB benchmarks demonstrate that our method significantly boosts the baseline performance by 2.7%, 2.3% and 4.9% on easy, medium and hard validation sets respectively. Furthermore, the proposed FTAFace-light achieves higher accuracy than the state-of-the-art and reduces the amount of computation by 28.9%. Dongchao Wen, Wei Tao 0001, Lingxiao Yin, Tse-Wei Chen 0001, Tadayuki Ito, Kinya Osa, Masami Kato |
ACM Multimedia | 2 |
| 2021 | SFace: Sigmoid-Constrained Hypersphere Loss for Robust Face RecognitionabstractDeep face recognition has achieved great success due to large-scale training databases and rapidly developing loss functions. The existing algorithms devote to realizing an ideal idea: minimizing the intra-class distance and maximizing the inter-class distance. However, they may neglect that there are also low quality training images which should not be optimized in this strict way. Considering the imperfection of training databases, we propose that intra-class and inter-class objectives can be optimized in a moderate way to mitigate overfitting problem, and further propose a novel loss function, named sigmoid-constrained hypersphere loss (SFace). Specifically, SFace imposes intra-class and inter-class constraints on a hypersphere manifold, which are controlled by two sigmoid gradient re-scale functions respectively. The sigmoid curves precisely re-scale the intra-class and inter-class gradients so that training samples can be optimized to some degree. Therefore, SFace can make a better balance between decreasing the intra-class distances for clean examples and preventing overfitting to the label noise, and contributes more robust deep face recognition models. Extensive experiments of models trained on CASIA-WebFace, VGGFace2, and MS-Celeb-1M databases, and evaluated on several face recognition benchmarks, such as LFW, MegaFace and IJB-C databases, have demonstrated the superiority of SFace. Yaoyao Zhong, Weihong Deng, Jiani Hu, Dongyue Zhao, Dongchao Wen |
IEEE Trans. Image Process. | 6 |
| 2020 | Fully Supervised and Guided Distillation for One-Stage Detectors
Dongchao Wen, Junjie Liu 0003, Wei Tao 0001, Tse-Wei Chen 0001, Kinya Osa, Masami Kato |
ACCV (3) | 2 |
| 2020 | Global-Local GCN: Large-Scale Label Noise Cleansing for Face RecognitionabstractIn the field of face recognition, large-scale web-collected datasets are essential for learning discriminative representations, but they suffer from noisy identity labels, such as outliers and label flips. It is beneficial to automatically cleanse their label noise for improving recognition accuracy. Unfortunately, existing cleansing methods cannot accurately identify noise in the wild. To solve this problem, we propose an effective automatic label noise cleansing framework for face recognition datasets, FaceGraph. Using two cascaded graph convolutional networks, FaceGraph performs global-to-local discrimination to select useful data in a noisy environment. Extensive experiments show that cleansing widely used datasets, such as CASIA-WebFace, VGGFace2, MegaFace2, and MS-Celeb-1M, using the proposed method can improve the recognition performance of state-of-the-art representation learning methods like Arcface. Further, we cleanse massive self-collected celebrity data, namely MillionCelebs, to provide 18.8M images of 636K identities. Training with the new data, Arcface surpasses state-of-the-art performance by a notable margin to reach 95.62% TPR at 1e-5 FPR on the IJB-C benchmark. Yaobin Zhang, Weihong Deng, Jiani Hu, Dongyue Zhao, Dongchao Wen |
CVPR | 7 |