Wenlong Deng

dblp:248/7433 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0002-0545-6384ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs
abstract
Wenlong Deng, Qi Zeng, Jiaming Zhang, Minghui Chen, Zixin Ding, Christos Thrampoulidis, Boying Gong, Xiaoxiao Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wenlong Deng, Zixin Ding, Christos Thrampoulidis, Boying Gong
ACL (1)1
2025 Can Textual Gradient Work in Federated Learning?
abstract
Recent studies highlight the promise of LLM-based prompt optimization, especially with TextGrad, which automates ``differentiation'' via texts and backpropagates textual feedback provided by LLMs. This approach facilitates training in various real-world applications that do not support numerical gradient propagation or loss calculation. It opens new avenues for optimization in decentralized, resource-constrained environments, suggesting that users of black-box LLMs (e.g., ChatGPT) could enhance components of LLM agentic systems (such as prompt optimization) through collaborative paradigms like federated learning (FL). In this paper, we systematically explore the potential and challenges of incorporating textual gradient into FL. Our contributions are fourfold. **Firstly**, we introduce a novel FL paradigm, Federated Textual Gradient (FedTextGrad), that allows FL clients to upload their locally optimized prompts derived from textual gradients, while the FL server aggregates the received prompts through text summarization. Unlike traditional FL frameworks, which are designed for numerical aggregation, FedTextGrad is specifically tailored for handling textual data, expanding the applicability of FL to a broader range of problems that lack well-defined numerical loss functions. **Secondly**, building on this design, we conduct extensive experiments to explore the feasibility of federated textual gradients. Our findings highlight the importance of properly tuning key factors (e.g., local steps) in FL training to effectively integrate textual gradients. **Thirdly**, we highlight a major challenge in federated textual gradient aggregation: retaining essential information from distributed prompt updates. Concatenation often produces prompts that exceed the LLM API’s context window, while summarization can degrade performance by generating overly condensed or complex text that lacks key context. **Last but not least**, in response to this issue, we improve the vanilla variant of FedTextGrad by providing actionable guidance to the LLM when summarizing client prompts by leveraging the Uniform Information Density principle. Such a design reduces the complexity of the aggregated global prompt, thereby better incentivizing the LLM's reasoning ability. Through this principled study, we enable the adoption of textual gradients in FL for optimizing LLMs, identify important issues, and pinpoint future directions, thereby opening up a new research area that warrants further investigation.
Ruinan Jin, Wenlong Deng, Yuanyuan Chen 0012, Han Yu 0001, Xiaoxiao Li 0001
ICLR3
2025 DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
abstract
Storing open-source fine-tuned models separately introduces redundancy and increases response times in applications utilizing multiple models. Delta-parameter pruning (DPP), particularly the random drop and rescale (DARE) method proposed by Yu et al., addresses this by pruning the majority of delta parameters—the differences between fine-tuned and pre-trained model weights—while typically maintaining minimal performance loss. However, DARE fails when either the pruning rate or the magnitude of the delta parameters is large. We highlight two key reasons for this failure: (1) an excessively large rescaling factor as pruning rates increase, and (2) high mean and variance in the delta parameters. To push DARE’s limits, we introduce DAREx (DARE the eXtreme), which features two algorithmic improvements: (1) DAREx-q, a rescaling factor modification that significantly boosts performance at high pruning rates (e.g., > 30% on COLA and SST2 for encoder models, with even greater gains in decoder models), and (2) DAREx-L2, which combines DARE with AdamR, an in-training method that applies appropriate delta regularization before DPP. We also demonstrate that DAREx-q can be seamlessly combined with vanilla parameter-efficient fine-tuning techniques like LoRA and can facilitate structural DPP. Additionally, we revisit the application of importance-based pruning techniques within DPP, demonstrating that they outperform random-based methods when delta parameters are large. Through this comprehensive study, we develop a pipeline for selecting the most appropriate DPP method under various practical scenarios.
Wenlong Deng, Yize Zhao, Vala Vakilian, Christos Thrampoulidis
ICLR1
2025 GMValuator: Similarity-based Data Valuation for Generative Models
abstract
Data valuation plays a crucial role in machine learning. Existing data valuation methods, mainly focused on discriminative models, overlook generative models that have gained attention recently. In generative models, data valuation measures the impact of training data on generated datasets. Very few existing attempts at data valuation methods designed for deep generative models either concentrate on specific models or lack robustness in their outcomes. Moreover, efficiency still reveals vulnerable shortcomings. We formulate the data valuation problem in generative models from a similarity matching perspective to bridge the gaps. Specifically, we introduce Generative Model Valuator (GMValuator), the first training-free and model-agnostic approach to providing data valuation for image generation tasks. It empowers efficient data valuation through our innovative similarity matching module, calibrates biased contributions by incorporating image quality assessment, and attributes credits to all training samples based on their contributions to the generated samples. Additionally, we introduce four evaluation criteria for assessing data valuation methods in generative models. GMValuator is extensively evaluated on benchmark and high-resolution datasets and various mainstream generative architectures to demonstrate its effectiveness. Our code is available at: https://github.com/ubc-tea/GMValuator.
Jiaxi Yang 0003, Wenlong Deng, Benlin Liu, Yangsibo Huang
ICLR2
2025 On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
abstract
Reinforcement learning (RL) has become popular in enhancing the reasoning capabilities of large language models (LLMs), with Group Relative Policy Optimization (GRPO) emerging as a widely used algorithm in recent systems. Despite GRPO's widespread adoption, we identify a previously unrecognized phenomenon we term Lazy Likelihood Displacement (LLD), wherein the likelihood of correct responses marginally increases or even decreases during training. This behavior mirrors a recently discovered misalignment issue in Direct Preference Optimization (DPO), attributed to the influence of negative gradients. We provide a theoretical analysis of GRPO’s learning dynamic, identifying the source of LLD as the naive penalization of all tokens in incorrect responses with the same strength. To address this, we develop a method called NTHR, which downweights penalties on tokens contributing to the LLD. Unlike prior DPO-based approaches, NTHR takes advantage of GRPO’s group-based structure, using correct responses as anchors to identify influential tokens. Experiments on math reasoning benchmarks demonstrate that NTHR effectively mitigates LLD, yielding consistent performance gains across models ranging from 0.5B to 3B parameters.
Wenlong Deng, Muchen Li, Danica J. Sutherland, Christos Thrampoulidis
NeurIPS1
2025 Corrigendum to "LESS: Label-efficient multi-scale learning for cytological whole slide image screening" [Medical Image Analysis 94 (2024): 103109]
Beidi Zhao, Wenlong Deng, Zi Han Li, Zuhua Gao
Medical Image Anal.2
2025 Multi-Source Joint Adaptive Distribution With Online Transfer Learning for Cross-Domain Fault Diagnosis
abstract
Methods based on transfer learning have achieved rich research results in the field of intelligent diagnosis of mechanical devices. However, current transfer learning methods typically require the source and target domains to be known in advance, heavily relying on historical data, which fails to meet the requirements of practical applications. Therefore, this study proposes a multisource online transfer learning with joint adaptive distribution selection (MSOTL-JADS) algorithm for real-time diagnosis of online samples. First, during the offline phase, a metric function is designed to extract sub-domain feature information in the multisource domain scenario, facilitating accurate knowledge transfer. Second, the multisource weight distribution parameters obtained based on the distribution distance are dynamically selected for the multisource domains to achieve differentiated transfer in the domain. Third, in the online phase, the pretrained offline model is integrated using online input target samples to construct an online diagnostic model and fine-tuning the weight parameters to match the model accuracy. Finally, several online transfer diagnostic tasks were constructed using two public datasets and one self-constructed dataset. The experimental results demonstrate that the proposed MSOTL-JADS outperforms the comparison methods in terms of performance.
Wenlong Deng, Sisindisiwe Nomalanga Ncube, Ruotong Ming, Chaoqun Duan, Yi Qin 0004, Jun Luo 0006, Huayan Pu
IEEE Trans. Reliab.2
2024 Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
abstract
Vision Transformers (ViT) and Visual Prompt Tuning (VPT) achieve state-of-the-art performance with improved efficiency in various computer vision tasks. This suggests a promising paradigm shift of adapting pretrained ViT models to Federated Learning (FL) settings. However, the challenge of data heterogeneity among FL clients presents a significant hurdle in effectively deploying ViT models. Existing Generalized FL (GFL) and Personalized FL (PFL) methods have limitations in balancing performance across both global and local data distributions. In this paper, we present a novel algorithm, SGPT, that integrates GFL and PFL approaches by employing a unique combination of both shared and group-specific prompts. This design enables SGPT to capture both common and group-specific features. A key feature of SGPT is its prompt selection module, which facilitates the training of a single global model capable of automatically adapting to diverse local client data distributions without the need for local fine-tuning. To effectively train the prompts, we utilize block coordinate descent (BCD), learning from common feature information (shared prompts), and then more specialized knowledge (group prompts) iteratively. Theoretically, we justify that learning the proposed prompts can reduce the gap between global and local performance. Empirically, we conduct experiments on both label and feature heterogeneity settings in comparison with state-of-the-art baselines, along with extensive ablation studies, to substantiate the superior performance of SGPT.
Wenlong Deng, Christos Thrampoulidis, Xiaoxiao Li 0001
CVPR1
2024 Debiased Noise Editing on Foundation Models for Fair Medical Image Classification
Ruinan Jin, Wenlong Deng, Xiaoxiao Li 0001
MICCAI (10)2
2024 LESS: Label-efficient multi-scale learning for cytological whole slide image screening
Beidi Zhao, Wenlong Deng, Zi Han (henry) Li, Zuhua Gao, Xiaoxiao Li 0001
Medical Image Anal.2
2023 Community-Aware Transformer for Autism Prediction in fMRI Connectome
Anushree Bannadabhavi, Wenlong Deng, Rex Ying
MICCAI (8)3
2023 Duplex adversarial domain discriminative network for cross-domain partial transfer fault diagnosis
Wenlong Deng, Chaoqun Duan, Yi Qin 0004, Jun Luo 0006, Huayan Pu
Knowl. Based Syst.2
2020 Joint Human Pose Estimation and Stereo 3D Localization
abstract
We present an end-to-end trainable Neural Network architecture for stereo imaging that jointly locates and estimates human body poses in 3D. Our method defines a 2D pose for each human in a stereo pair of images and uses a correlation layer with a composite field to associate each left-right pair of joints. In absence of a stereo pose dataset, we show that we can train our method with synthetic data only and test it on real-world images (i.e., our training stage is domain invariant). Our method is particularly suitable for autonomous vehicles. We achieve state-of-the-art results for the 3D localization task on the challenging real-world KITTI dataset while running four times faster.
Wenlong Deng, Lorenzo Bertoni, Sven Kreiss, Alexandre Alahi
ICRA1