Duo Peng

dblp:175/3967 · DBLP profile ↗
← Back
22ranked-venue papers
12as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 10 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 12 since 2021Systems, architecture and hardware · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unleashing the Power of Text-to-Image Diffusion Models for Category-Agnostic Pose Estimation
abstract
Category-Agnostic Pose Estimation (CAPE) aims to detect keypoints of unseen object categories in a few-shot setting, where the scarcity of labeled data poses significant challenges to generalization. In this work, we propose Prompt Pose Matching (PPM), a novel framework that unleashes the power of off-the-shelf text-to-image diffusion models for CAPE. PPM learns pseudo prompts from few-shot examples via the text-to-image diffusion model. These learned pseudo prompts capture semantic information of keypoints, which can then be used to locate the same type of keypoints from images. To provide prompts with representative initialization, we introduce a category-agnostic pre-training strategy to capture the foreground prior shared across categories and keypoints. To support the reliable prompt pre-training, we propose a Foreground-Aware Region Aggregation (FARA) module to provide robust and consistent supervision signal. Based on the foreground prior, a Foreground-Guided Attention Refinement (FGAR) module is further proposed to reinforce cross-attention responses for accurate keypoint localization. For efficiency, a Prompt Ensemble Inference (PEI) scheme enables joint keypoint prediction. Unlike previous methods that highly rely on base-category annotated data, our PPM framework can operate in a base-category-free setting while retaining strong performance. Code will be available at: https://github.com/DuoPeng-CVer/Prompt-Pose-Matching.
Duo Peng, Zhengbo Zhang, Ping Hu 0001, Qiuhong Ke, De Wen Soh, Mohammed Bennamoun, Jun Liu 0036
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Driver Interaction Intent Prediction With Dynamic Scene Semantic Fusion for Intelligent Cockpit
abstract
Driver interaction intent prediction is one of the core technologies enabling intelligent cockpits to transition from passive response to proactive service. Existing intent prediction methods mainly model for single service functions and are deficient in generalization of cross-functional prediction and generalized temporal prediction models predominantly neglect the spatio-temporal (ST) causal relationships between dynamic evolution of scenes and interaction intents, resulting in compromised prediction efficacy within complex environments. To address the above limitations, we propose a driver interaction intent prediction framework with scene semantic fusion in intelligent cockpits. The methodology involves three key stages: first, extracting temporal patterns from cockpit interaction sequences via LSTM networks, second, converting driving dynamics data into structured scene semantic text through parameter efficient fine-tuning of large language models (LLM), and third, establishing collaborative representations of interaction behavioral features and scene information based on scene-guided attention fusion mechanism to achieve accurate driver interaction intent prediction. Experimental results demonstrate that our method has superior performance compared with other methods, with Precision, Recall, F1-score, and Accuracy values of 0.8416, 0.7842, 0.8045, and 0.7948, respectively. In particular, the introduction of scene semantic information can effectively trace the motivating source of the driver interaction intent generation and effectively enhance the interpretability of the model. This research establishes a technical foundation for implementing personalized proactive services in intelligent cockpits.
Hongyu Hu, Zhiwen Wei, Duo Peng, Changjia Tian, Rui Zhao 0021, Fei Gao 0020
IEEE Trans. Ind. Informatics4
2025 Training-free Dense-Aligned Diffusion Guidance for Modular Conditional Image Synthesis
abstract
Conditional image synthesis is a crucial task with broad applications, such as artistic creation and virtual reality. However, current generative methods are often task-oriented with a narrow scope, handling a restricted condition with constrained applicability. In this paper, we propose a novel approach that treats conditional image synthesis as the modular combination of diverse fundamental condition units. Specifically, we divide conditions into three primary units: text, layout, and drag. To enable effective control over these conditions, we design a dedicated alignment module for each. For the text condition, we introduce a Dense Concept Alignment (DCA) module, which achieves dense visual-text alignment by drawing on diverse textual concepts. For the layout condition, we propose a Dense Geometry Alignment (DGA) module to enforce comprehensive geometric constraints that preserve the spatial configuration. For the drag condition, we introduce a Dense Motion Alignment (DMA) module to apply multi-level motion regularization, ensuring that each pixel follows its desired trajectory without visual artifacts. By flexibly inserting and combining these alignment modules, our framework enhances the model’s adaptability to diverse conditional generation tasks and greatly expands its application range. Extensive experiments demonstrate the superior performance of our framework across a variety of conditions, including textual description, segmentation mask (bounding box), drag manipulation, and their combinations. Code is available at https://github.com/ZixuanWang0525/DADG
Duo Peng, Feng Chen 0047, Yinjie Lei
CVPR2
2025 Visual Prompting for One-shot Controllable Video Editing without Inversion
abstract
One-shot controllable video editing (OCVE) is an important yet challenging task, aiming to propagate user edits that are made – using any image editing tool – on the first frame of a video to all subsequent frames, while ensuring content consistency between edited frames and source frames. To achieve this, prior methods employ DDIM inversion to transform source frames into latent noise, which is then fed into a pre-trained diffusion model, conditioned on the user-edited first frame, to generate the edited video. However, the DDIM inversion process accumulates errors, which hinder the latent noise from accurately reconstructing the source frames, ultimately compromising content consistency in the generated edited frames. To overcome it, our method eliminates the need for DDIM inversion by performing OCVE through a novel perspective based on visual prompting. Furthermore, inspired by consistency models that can perform multi-step consistency sampling to generate a sequence of content-consistent images, we propose a content consistency sampling (CCS) to ensure content consistency between the generated edited frames and the source frames. Moreover, we introduce a temporal-content consistency sampling (TCS) based on Stein Variational Gradient Descent to ensure temporal consistency across the edited frames. Extensive experiments validate the effectiveness of our approach.
Zhengbo Zhang, Duo Peng, Joo-Hwee Lim, Zhigang Tu 0001, De Wen Soh, Lin Geng Foo
CVPR3
2025 Unified Prompt Attack Against Text-to-Image Generation Models
abstract
Text-to-Image (T2I) models have advanced significantly, but their growing popularity raises security concerns due to their potential to generate harmful images. To address these issues, we propose UPAM, a novel framework to evaluate the robustness of T2I models from an attack perspective. Unlike prior methods that focus solely on textual defenses, UPAM unifies the attack on both textual and visual defenses. Additionally, it enables gradient-based optimization, overcoming reliance on enumeration for improved efficiency and effectiveness. To handle cases where T2I models block image outputs due to defenses, we introduce Sphere-Probing Learning (SPL) to enable optimization even without image results. Following SPL, our model bypasses defenses, inducing the generation of harmful content. To ensure semantic alignment with attacker intent, we propose Semantic-Enhancing Learning (SEL) for precise semantic control. UPAM also prioritizes the naturalness of adversarial prompts using In-context Naturalness Enhancement (INE), making them harder for human examiners to detect. Additionally, we address the issue of iterative queries-common in prior methods and easily detectable by API defenders-by introducing Transferable Attack Learning (TAL), allowing effective attacks with minimal queries. Extensive experiments validate UPAM's superiority in effectiveness, efficiency, naturalness, and low query detection rates.
Duo Peng, Qiuhong Ke, Mark He Huang, Ping Hu 0001, Jun Liu 0036
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Harnessing Text-to-Image Diffusion Models for Category-Agnostic Pose Estimation
Duo Peng, Zhengbo Zhang, Ping Hu 0001, Qiuhong Ke, David K. Y. Yau, Jun Liu 0036
ECCV (13)1
2024 Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers
Zhengbo Zhang, Duo Peng, Hossein Rahmani 0001, Jun Liu 0036
ECCV (28)3
2024 UPAM: Unified Prompt Attack in Text-to-Image Generation Models Against Both Textual Filters and Visual Checkers
abstract
Text-to-Image (T2I) models have raised security concerns due to their potential to generate inappropriate or harmful images. In this paper, we propose UPAM, a novel framework that investigates the robustness of T2I models from the attack perspective. Unlike most existing attack methods that focus on deceiving textual defenses, UPAM aims to deceive both textual and visual defenses in T2I models. UPAM enables gradient-based optimization, offering greater effectiveness and efficiency than previous methods. Given that T2I models might not return results due to defense mechanisms, we introduce a Sphere-Probing Learning (SPL) scheme to support gradient optimization even when no results are returned. Additionally, we devise a Semantic-Enhancing Learning (SEL) scheme to finetune UPAM for generating target-aligned images. Our framework also ensures attack stealthiness. Extensive experiments demonstrate UPAM's effectiveness and efficiency.
Duo Peng, Qiuhong Ke, Jun Liu 0036
ICML1
2024 Unsupervised Domain Adaptation via Domain-Adaptive Diffusion
abstract
Unsupervised Domain Adaptation (UDA) is quite challenging due to the large distribution discrepancy between the source domain and the target domain. Inspired by diffusion models which have strong capability to gradually convert data distributions across a large gap, we consider to explore the diffusion technique to handle the challenging UDA task. However, using diffusion models to convert data distribution across different domains is a non-trivial problem as the standard diffusion models generally perform conversion from the Gaussian distribution instead of from a specific domain distribution. Besides, during the conversion, the semantics of the source-domain data needs to be preserved to classify correctly in the target domain. To tackle these problems, we propose a novel Domain-Adaptive Diffusion (DAD) module accompanied by a Mutual Learning Strategy (MLS), which can gradually convert data distribution from the source domain to the target domain while enabling the classification model to learn along the domain transition process. Consequently, our method successfully eases the challenge of UDA by decomposing the large domain gap into small ones and gradually enhancing the capacity of classification model to finally adapt to the target domain. Our method outperforms the current state-of-the-arts by a large margin on three widely used UDA datasets.
Duo Peng, Qiuhong Ke, Arulmurugan Ambikapathi, Yasin Yazici, Yinjie Lei, Jun Liu 0036
IEEE Trans. Image Process.1
2023 Diffusion-based Image Translation with Label Guidance for Domain Adaptive Semantic Segmentation
abstract
Translating images from a source domain to a target domain for learning target models is one of the most common strategies in domain adaptive semantic segmentation (DASS). However, existing methods still struggle to preserve semantically-consistent local details between the original and translated images. In this work, we present an innovative approach that addresses this challenge by using sourcedomain labels as explicit guidance during image translation. Concretely, we formulate cross-domain image translation as a denoising diffusion process and utilize a novel Semantic Gradient Guidance (SGG) method to constrain the translation process, conditioning it on the pixel-wise source labels. Additionally, a Progressive Translation Learning (PTL) strategy is devised to enable the SGG method to work reliably across domains with large gaps. Extensive experiments demonstrate the superiority of our approach over state-of-the-art methods.
Duo Peng, Ping Hu 0001, Qiuhong Ke, Jun Liu 0036
ICCV1
2023 Joint Attribute and Model Generalization Learning for Privacy-Preserving Action Recognition
abstract
Privacy-Preserving Action Recognition (PPAR) aims to transform raw videos into anonymous ones to prevent privacy leakage while maintaining action clues, which is an increasingly important problem in intelligent vision applications. Despite recent efforts in this task, it is still challenging to deal with novel privacy attributes and novel privacy attack models that are unavailable during the training phase. In this paper, from the perspective of meta-learning (learning to learn), we propose a novel Meta Privacy-Preserving Action Recognition (MPPAR) framework to improve both generalization abilities above (i.e., generalize to *novel privacy attributes* and *novel privacy attack models*) in a unified manner. Concretely, we simulate train/test task shifts by constructing disjoint support/query sets w.r.t. privacy attributes or attack models. Then, a virtual training and testing scheme is applied based on support/query sets to provide feedback to optimize the model's learning toward better generalization. Extensive experiments demonstrate the effectiveness and generalization of the proposed framework compared to state-of-the-arts.
Duo Peng, Qiuhong Ke, Ping Hu 0001, Jun Liu 0036
NeurIPS1
2023 Localization algorithm for anisotropic wireless sensor networks based on the adaptive chaotic slime mold algorithm
Duo Peng, Yuwei Gao
Neural Comput. Appl.1
2023 Deep Hypersphere Feature Regularization for Weakly Supervised RGB-D Salient Object Detection
abstract
We propose a weakly supervised approach for salient object detection from multi-modal RGB-D data. Our approach only relies on labels from scribbles, which are much easier to annotate, compared with dense labels used in conventional fully supervised setting. In contrast to existing methods that employ supervision signals on the output space, our design regularizes the intermediate latent space to enhance discrimination between salient and non-salient objects. We further introduce a contour detection branch to implicitly constrain the semantic boundaries and achieve precise edges of detected salient objects. To enhance the long-range dependencies among local features, we introduce a Cross-Padding Attention Block (CPAB). Extensive experiments on seven benchmark datasets demonstrate that our method not only outperforms existing weakly supervised methods, but is also on par with several fully-supervised state-of-the-art models. Code is available at https://github.com/leolyj/DHFR-SOD.
Munawar Hayat, Duo Peng, Yinjie Lei
IEEE Trans. Image Process.4
2023 Dear-Net: Learning Diversities for Skeleton-Based Early Action Recognition
abstract
Early actionrecognition, i.e., recognizing an action before it is fully performed, is a challenging and important task. Existing works mainly focus on deterministic early action recognition outputting only a single class, and ignore the uncertainty and diversity that essentially exist in this task. Intuitively, when only the early portion of the action is observed, there could be multiple possibilities of the full action, as diversified actions can share almost identical early segments in many scenarios. Thus taking uncertainties and diversities into account, and outputting multiple plausible predictions, instead of a single one, can be important for the sake of authenticity and requirement of many practical applications. To this end, we propose a novel Diversified Early Action Recognition Network (Dear-Net) that is capable of outputting multiple reasonable action classes for each partial sequence by utilizing mode conversion. Specifically, we introduce an effective action diversity learning strategy to drive our network towards predicting diverse and reasonable results, in which each learnable action class is matched with the most suitable mode. Meanwhile, the collapsed modes which fail to receive any action class, are also considered in this strategy in order to ensure diversity. Moreover, we design a sequence decoder within our network to capture latent global information for better early action recognition. It provides a feasible scheme for weakly-supervised setting in which the Dear-Net leverages unlabelled data to improve performance. Experimental results on three challenging datasets clearly show the effectiveness of our approach.
Rui Wang 0108, Jun Liu 0036, Qiuhong Ke, Duo Peng, Yinjie Lei
IEEE Trans. Multim.4
2022 Semantic-Aware Domain Generalized Segmentation
abstract
Deep models trained on source domain lack generalization when evaluated on unseen target domains with different data distributions. The problem becomes even more pro-nounced when we have no access to target domain samples for adaptation. In this paper, we address domain generalized semantic segmentation, where a segmentation model is trained to be domain-invariant without using any target domain data. Existing approaches to tackle this problem standardize data into a unified distribution. We argue that while such a standardization promotes global normalization, the resulting features are not discriminative enough to get clear segmentation boundaries. To enhance separation between categories while simultaneously promoting domain invariance, we propose a framework including two novel modules: Semantic-Aware Normalization (SAN) and Semantic-Aware Whitening (SAW). Specifically, SAN focuses on category-level center alignment between features from different image styles, while SAW enforces distributed alignment for the already center-aligned features. With the help of SAN and SAW, we encourage both intra-category compactness and inter-category separability. We validate our approach through extensive experiments on widely-used datasets (i.e. GTAV, SYNTHIA, Cityscapes, Mapillary and BDDS). Our approach shows significant improvements over existing state-of-the-art on various backbone networks. Code is available at https://github.com/leolyj/SAN-SAW
Duo Peng, Yinjie Lei, Munawar Hayat, Yulan Guo, Wen Li 0001
CVPR1
2022 Proximity-Distance Mapping and Jaya Optimization Algorithm Based on Localization for Wireless Sensor Network
abstract
Aiming at the problem of large location errors of traditional ranging-free algorithms in Wireless Sensor Network (WSN), a novel node location algorithm based on proximity-distance mapping (PDM) and Jaya optimization was proposed. In this algorithm, proximity and Euclidean distance are extracted from the relationship of anchor nodes to construct a mapping matrix by using the idea of PDM. It is calculated by using the mapping matrix that the estimated distance from the unknown node to the anchor node can be used for the subsequent calculations. After the estimated distance is obtained, the Jaya optimization algorithm is imported to calculate the location of the unknown one. To accelerate the convergence and enhance the accuracy of the algorithm, the idea of a boundary box is used to limit the initial feasible region of unknown nodes. The experiment results show that the PDM–Jaya algorithm has better positioning accuracy than the original PDM in the same condition.
Duo Peng, Yuwei Gao
Int. J. Pattern Recognit. Artif. Intell.1
2021 Sparse-to-dense Feature Matching: Intra and Inter domain Cross-modal Learning in Domain Adaptation for 3D Semantic Segmentation
abstract
Domain adaptation is critical for success when confronting with the lack of annotations in a new domain. As the huge time consumption of labeling process on 3D point cloud, domain adaptation for 3D semantic segmentation is of great expectation. With the rise of multi-modal datasets, large amount of 2D images are accessible besides 3D point clouds. In light of this, we propose to further leverage 2D data for 3D domain adaptation by intra and inter domain cross modal learning. As for intra-domain cross modal learning, most existing works sample the dense 2D pixel-wise features into the same size with sparse 3D point-wise features, resulting in the abandon of numerous useful 2D features. To address this problem, we propose Dynamic sparse-to-dense Cross Modal Learning (DsCML) to increase the sufficiency of multi-modality information interaction for domain adaptation. For inter-domain cross modal learning, we further advance Cross Modal Adversarial Learning (CMAL) on 2D and 3D data which contains different semantic content aiming to promote high-level modal complementarity. We evaluate our model under various multi-modality domain adaptation settings including day-to-night, country-to-country and dataset-to-dataset, brings large improvements over both uni-modal and multi-modal domain adaptation methods on all settings. Code is available at https://github.com/leolyj/DsCML
Duo Peng, Yinjie Lei, Wen Li 0001, Yulan Guo
ICCV1
2021 Hierarchical Paired Channel Fusion Network for Street Scene Change Detection
abstract
Street Scene Change Detection (SSCD) aims to locate the changed regions between a given street-view image pair captured at different times, which is an important yet challenging task in the computer vision community. The intuitive way to solve the SSCD task is to fuse the extracted image feature pairs, and then directly measure the dissimilarity parts for producing a change map. Therefore, the key for the SSCD task is to design an effective feature fusion method that can improve the accuracy of the corresponding change maps. To this end, we present a novel Hierarchical Paired Channel Fusion Network (HPCFNet), which utilizes the adaptive fusion of paired feature channels. Specifically, the features of a given image pair are jointly extracted by a Siamese Convolutional Neural Network (SCNN) and hierarchically combined by exploring the fusion of channel pairs at multiple feature levels. In addition, based on the observation that the distribution of scene changes is diverse, we further propose a Multi-Part Feature Learning (MPFL) strategy to detect diverse changes. Based on the MPFL strategy, our framework achieves a novel approach to adapt to the scale and location diversities of the scene change regions. Extensive experiments on three public datasets (i.e., PCD, VL-CMU-CD and CDnet2014) demonstrate that the proposed framework achieves superior performance which outperforms other state-of-the-art methods with a considerable margin.
Yinjie Lei, Duo Peng, Qiuhong Ke, Haifeng Li 0007
IEEE Trans. Image Process.2
2021 Global and Local Texture Randomization for Synthetic-to-Real Semantic Segmentation
abstract
Semantic segmentation is a crucial image understanding task, where each pixel of image is categorized into a corresponding label. Since the pixel-wise labeling for ground-truth is tedious and labor intensive, in practical applications, many works exploit the synthetic images to train the model for real-word image semantic segmentation, i.e., Synthetic-to-Real Semantic Segmentation (SRSS). However, Deep Convolutional Neural Networks (CNNs) trained on the source synthetic data may not generalize well to the target real-world data. To address this problem, there has been rapidly growing interest in Domain Adaption technique to mitigate the domain mismatch between the synthetic and real-world images. Besides, Domain Generalization technique is another solution to handle SRSS. In contrast to Domain Adaption, Domain Generalization seeks to address SRSS without accessing any data of the target domain during training. In this work, we propose two simple yet effective texture randomization mechanisms, Global Texture Randomization (GTR) and Local Texture Randomization (LTR), for Domain Generalization based SRSS. GTR is proposed to randomize the texture of source images into diverse unreal texture styles. It aims to alleviate the reliance of the network on texture while promoting the learning of the domain-invariant cues. In addition, we find the texture difference is not always occurred in entire image and may only appear in some local areas. Therefore, we further propose a LTR mechanism to generate diverse local regions for partially stylizing the source images. Finally, we implement a regularization of Consistency between GTR and LTR (CGL) aiming to harmonize the two proposed mechanisms during training. Extensive experiments on five publicly available datasets (i.e., GTA5, SYNTHIA, Cityscapes, BDDS and Mapillary) with various SRSS settings (i.e., GTA5/SYNTHIA to Cityscapes/BDDS/Mapillary) demonstrate that the proposed method is superior to the state-of-the-art methods for domain generalization based SRSS.
Duo Peng, Yinjie Lei, Lingqiao Liu, Jun Liu 0036
IEEE Trans. Image Process.1
2019 Deep learning based mobile data offloading in mobile edge computing systems
Xianlong Zhao, Qimei Chen, Duo Peng, Hao Jiang 0010, Xianze Xu, Xinzhuo Shuang
Future Gener. Comput. Syst.4
2017 Delay analysis of MSW-ARQ system based on wireless multimedia services
Suoping Li, Zufang Dou, Duo Peng, Yongqiang Zhou
Multim. Tools Appl.4
2016 Analysis of dual-hop and multiple relays cooperative truncated ARQ with relay selection in WSNs
Suoping Li, Yongqiang Zhou, Duo Peng, Zufang Dou
Acta Informatica3