Dongsheng An

dblp:173/5382 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-6765-2578ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AuthGuard: Generalizable Deepfake Detection via Language Guidance
abstract
Existing deepfake detection techniques struggle to keep-up with the ever-evolving novel, unseen forgeries methods. This limitation stems from their reliance on statistical artifacts learned during training, which are often tied to specific generation processes that may not be representative of samples from new, unseen deepfake generation methods encountered at test time. We propose that incorporating language guidance can improve deepfake detection generalization by integrating human-like commonsense reasoning – such as recognizing logical inconsistencies and perceptual anomalies – alongside statistical cues. To achieve this, we train an expert deepfake vision encoder by combining discriminative classification with image-text contrastive learning, where the text is generated by generalist MLLMs using few-shot prompting. This allows the encoder to extract both language-describable, commonsense deepfake artifacts and statistical forgery artifacts from pixel-level distributions. To further enhance robustness, we integrate data uncertainty learning into vision-language contrastive learning, mitigating noise in image-text supervision. Our expert vision encoder seamlessly interfaces with an LLM, further enabling more generalized and interpretable deepfake detection while also boosting accuracy. The resulting framework, AuthGuard, achieves state-of-the-art deepfake detection accuracy in both in-distribution and out-of-distribution settings, achieving AUC gains of 6.15% on the DFDC dataset and 16.68% on the DF40 dataset. Additionally, AuthGuard significantly enhances deepfake reasoning, improving performance by 24.69% on the DDVQA dataset.
Guangyu Shen, Tianchen Zhao, Zheng Zhang 0001, Dongsheng An, Zhuowen Tu, Yifan Xing
WACV6
2026 Image compression using optimal transport mapping based on ranking visual saliency
Dongsheng An, Xianfeng Gu, Xiaoyin Xu, Min Zhang 0069
Pattern Recognit.2
2025 Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
abstract
This work presents a simple yet effective workflow for automatically scaling instruction-following data to elicit pixel-level grounding capabilities of VLMs under complex instructions. In particular, we address five critical real-world challenges in text-instruction-based grounding: hallucinated references, multi-object scenarios, reasoning, multi-granularity, and part-level references. By leveraging knowledge distillation from a pre-trained teacher model, our approach generates high-quality instruction-response pairs linked to existing pixel-level annotations, minimizing the need for costly human annotation. The resulting dataset, Ground-V, captures rich object localization knowledge and nuanced pixel-level referring expressions. Experiment results show that models trained on Ground-V exhibit substantial improvements across diverse grounding tasks. Specifically, incorporating Ground-V during training directly achieve an average accuracy boost of 4.4% for LISA and a 7.9% for PSALM across six benchmarks on the gIoU metric. It also sets new state-of-the-art results on standard benchmarks such as RefCOCO/+/g. Notably, on gRefCOCO, we achieve an N-Acc of 83.3%, exceeding the previous state-of-the-art by more than 20%.
Yongshuo Zong, Dongsheng An, Linghan Xu, Zhuowen Tu, Yifan Xing, Onkar Dabeer
CVPR3
2025 Salient Concept-Aware Generative Data Augmentation
abstract
Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts. This challenge arises because representations in the synthesis process often become entangled with non-essential input image attributes such as environmental contexts, creating conflicts with text prompts intended to modify these elements. To address this, we propose a personalized image generation framework that uses a salient concept-aware image embedding model to reduce the influence of irrelevant visual details during the synthesis process, thereby maintaining intuitive alignment between image and text inputs. By generating images that better preserve class-discriminative features with additional controlled variations, our framework effectively enhances the diversity of training datasets and thereby improves the robustness of downstream models. Our approach demonstrates superior performance across eight fine-grained vision datasets, outperforming state-of-the-art augmentation methods with averaged classification accuracy improvements by 0.73\% and 6.5\% under conventional and long-tail settings, respectively.
Tianchen Zhao, Xuanbai Chen, Dongsheng An, Zhuowen Tu, Yifan Xing
NeurIPS5
2024 Learning for Transductive Threshold Calibration in Open-World Recognition
abstract
In deep metric learning for visual recognition, the calibration of distance thresholds is crucial for achieving desired model performance in the true positive rates (TPR) or true negative rates (TNR). However, calibrating this threshold presents challenges in open-world scenarios, where the test classes can be entirely disjoint from those encountered during training. We define the problem of finding distance thresholds for a trained embedding model to achieve target performance metrics over unseen open-world test classes as open-world threshold calibration. Existing posthoc threshold calibration methods, reliant on inductive inference and requiring a calibration dataset with a similar distance distribution as the test data, often prove ineffective in open-world scenarios. To address this, we introduce OpenGCN, a Graph Neural Network-based transductive threshold calibration method with enhanced adaptability and robustness. OpenGCN learns to predict pairwise connectivity for the unlabeled test instances embedded in a graph to determine its TPR and TNR at various distance thresholds, allowing for transductive inference of the distance thresholds which also incorporates test-time information. Extensive experiments across open-world visual recognition benchmarks validate OpenGCN's superiority over existing posthoc calibration methods for open-world threshold calibration.
Dongsheng An, Tianjun Xiao, Tong He 0002, Qingming Tang, Ying Nian Wu, Joseph Tighe, Yifan Xing
CVPR2
2023 WordGesture-GAN: Modeling Word-Gesture Movement with Generative Adversarial Network
abstract
Word-gesture production models that can synthesize word-gestures are critical to the training and evaluation of word-gesture keyboard decoders. We propose WordGesture-GAN, a conditional generative adversarial network that takes arbitrary text as input to generate realistic word-gesture movements in both spatial (i.e., (x, y) coordinates of touch points) and temporal (i.e., timestamps of touch points) dimensions. WordGesture-GAN introduces a Variational Auto-Encoder to extract and embed variations of user-drawn gestures into a Gaussian distribution which can be sampled to control variation in generated gestures. Our experiments on a dataset with 38k gesture samples show that WordGesture-GAN outperforms existing gesture production models including the minimum jerk model [37] and the style-transfer GAN [31, 32] in generating realistic gestures. Overall, our research demonstrates that the proposed GAN structure can learn variations in user-drawn gestures, and the resulting WordGesture-GAN can generate word-gesture movement and predict the distribution of gestures. WordGesture-GAN can serve as a valuable tool for designing and evaluating gestural input systems.
Jeremy Chu, Dongsheng An, Yan Ma 0006, Wenzhe Cui, Shumin Zhai, Xianfeng Gu, Xiaojun Bi 0001
CHI2
2023 Volumetric Optimal Transportation by Fast Fourier Transform
Na Lei, Dongsheng An, Min Zhang 0069, Xiaoyin Xu, Xianfeng Gu
ICLR2
2022 Efficient Optimal Transport Algorithm by Accelerated Gradient Descent
abstract
Optimal transport (OT) plays an essential role in various areas like machine learning and deep learning. However, computing discrete optimal transport plan for large scale problems with adequate accuracy and efficiency is still highly challenging. Recently, methods based on the Sinkhorn algorithm add an entropy regularizer to the prime problem and get a trade off between efficiency and accuracy. In this paper, we propose a novel algorithm to further improve the efficiency and accuracy based on Nesterov's smoothing technique. Basically, the non-smooth c-transform of the Kantorovich potential is approximated by the smooth Log-Sum-Exp function, which finally smooths the original non-smooth Kantorovich dual functional. The smooth Kantorovich functional can be optimized by the fast proximal gradient algorithm (FISTA) efficiently. Theoretically, the computational complexity of the proposed method is lower than current estimation of the Sinkhorn algorithm in terms of the precision. Empirically, compared with the Sinkhorn algorithm, our experimental results demonstrate that the proposed method achieves faster convergence and better accuracy with the same parameter.
Dongsheng An, Na Lei, Xiaoyin Xu, Xianfeng Gu
AAAI1
2022 Approximate Discrete Optimal Transport Plan with Auxiliary Measure Method
Dongsheng An, Na Lei, Xianfeng Gu
ECCV (23)1
2022 Image Compression Based on Importance Using Optimal Mass Transportation Map
abstract
Demand for efficient image transmission and storage is increasing rapidly because of the continuing growth of multimedia technology and VR and AR applications. In this paper, we proposed an image compression method based on the recognition of importance of regions in images. As not all the information in an image is equally useful, we can identify important regions in an image for high fidelity compression and accept a comparatively more lossy compression about less important regions of the image. First, we segment images to two parts, namely, foreground and background, where the foreground represents the more important component and the background is of less importance. Second, we apply optimal mass transportation mapping in a GAN (generative adversarial network) framework to both the foreground and background to magnify the foreground and shrink the background while keeping the shape and total image area unchanged. As a result, in the processed image, the ratio of foreground to background is larger than the corrresponding ratio in the original image. This ratio is controllable in our process, giving users the ability to control the degree of compression. The GAN-processed image is then used for compression. To restore the image, we apply a GAN model to the compressed image and recover the ratio of foreground and background using an optimal mass transportation map. Test results show that our method is highly effective in reconstructing detail of important components in compressed images while achieving a high compression ratio.
Dongsheng An, Yingjie Feng, Xianfeng Gu, Xiaoyin Xu, Min Zhang 0069
ICIP2
2022 End-to-End Evidential-Efficient Net for Radiomics Analysis of Brain MRI to Predict Oncogene Expression and Overall Survival
Yingjie Feng, Jun Wang 0039, Dongsheng An, Xianfeng Gu, Xiaoyin Xu, Min Zhang 0069
MICCAI (3)3
2021 Learning Deep Latent Variable Models by Short-Run MCMC Inference With Optimal Transport Correction
abstract
Learning latent variable models with deep top-down architectures typically requires inferring the latent variables for each training example based on the posterior distribution of these latent variables. The inference step typically relies on either time-consuming long-run Markov chain Monte Carlo (MCMC) sampling or a separate inference model for variational learning. In this paper, we propose to use a shortrun MCMC, such as a short-run Langevin dynamics, as an approximate flow-based inference engine. The bias existing in the output distribution of the non-convergent short-run Langevin dynamics is corrected by the optimal transport (OT), which aims at transforming the biased distribution produced by the finite-step MCMC to the prior distribution with a minimum transport cost. Our experiments not only verify the effectiveness of the OT correction for the short-run MCMC, but also demonstrate that the latent variable model trained by the proposed strategy performs better than the variational auto-encoder (VAE) in terms of image reconstruction/generation and anomaly detection.
Dongsheng An, Jianwen Xie, Ping Li 0001
CVPR1
2020 AE-OT-GAN: Training GANs from Data Specific Latent Distribution
Dongsheng An, Min Zhang 0069, Xin Qi 0011, Na Lei, Xianfeng Gu
ECCV (26)1
2020 Ae-OT: a New Generative Model based on Extended Semi-discrete Optimal transport
Dongsheng An, Na Lei, Zhongxuan Luo, Shing-Tung Yau, Xianfeng Gu
ICLR1
2016 Fast and High Quality Highlight Removal From a Single Image
abstract
Specular reflection exists widely in photography and causes the recorded color deviating from its true value, thus, fast and high quality highlight removal from a single nature image is of great importance. In spite of the progress in the past decades in highlight removal, achieving wide applicability to the large diversity of nature scenes is quite challenging. To handle this problem, we propose an analytic solution to highlight removal based on an L2chromaticity definition and corresponding dichromatic model. Specifically, this paper derives a normalized dichromatic model for the pixels with identical diffuse color: a unit circle equation of projection coefficients in two subspaces that are orthogonal to and parallel with the illumination, respectively. In the former illumination orthogonal subspace, which is specular-free, we can conduct robust clustering with an explicit criterion to determine the cluster number adaptively. In the latter, illumination parallel subspace, a property called pure diffuse pixels distribution rule helps map each specular-influenced pixel to its diffuse component. In terms of efficiency, the proposed approach involves few complex calculation, and thus can remove highlight from high resolution images fast. Experiments show that this method is of superior performance in various challenging cases.
Jin-Li Suo, Dongsheng An, Xiangyang Ji, Haoqian Wang, Qionghai Dai
IEEE Trans. Image Process.2