VLDB 2026 Research / reviewers in the wild / expert
Dong-Dong Wu
dblp:323/9534
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Learning paradigms · 49% Efficient and distributed learning · 17% Trustworthy machine learning · 16% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 100% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms › weakly supervised learning
partial label learning |
2.2 | 3 | 2025 | Realistic Evaluation of Deep Partial-Label Learning Algorithms · ICLR 2025 Distilling Reliable Knowledge for Instance-Dependent Partial Label Learning · AAAI 2024 Revisiting Consistency Regularization for Deep Partial Label Learning · ICML 2022 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial transferability |
0.9 | 1 | 2025 | A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1 · NeurIPS 2025 |
Machine learning › Learning theory
model selection |
0.9 | 1 | 2025 | Realistic Evaluation of Deep Partial-Label Learning Algorithms · ICLR 2025 |
Machine learning › Learning paradigms
weakly supervised learning |
0.9 | 1 | 2025 | Realistic Evaluation of Deep Partial-Label Learning Algorithms · ICLR 2025 |
Security and privacy of machine learning › adversarial attack › multimodal adversarial attack
vision-language model attack |
0.9 | 1 | 2025 | A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1 · NeurIPS 2025 |
Machine learning › Learning paradigms › weakly supervised learning › partial label learning
instance-dependent partial label learning |
0.8 | 1 | 2024 | Distilling Reliable Knowledge for Instance-Dependent Partial Label Learning · AAAI 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.8 | 1 | 2024 | Distilling Reliable Knowledge for Instance-Dependent Partial Label Learning · AAAI 2024 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation |
0.8 | 1 | 2024 | Distilling Reliable Knowledge for Instance-Dependent Partial Label Learning · AAAI 2024 |
Security and privacy of machine learning
model stealing |
0.8 | 1 | 2024 | Efficient Model Stealing Defense with Noise Transition Matrix · CVPR 2024 |
Security and privacy of machine learning › model stealing
model stealing defense |
0.8 | 1 | 2024 | Efficient Model Stealing Defense with Noise Transition Matrix · CVPR 2024 |
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization |
0.6 | 1 | 2022 | Revisiting Consistency Regularization for Deep Partial Label Learning · ICML 2022 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.6 | 1 | 2022 | Revisiting Consistency Regularization for Deep Partial Label Learning · ICML 2022 |
Computer vision › Image recognition and object detection
image classification |
0.3 | 1 | 2025 | Realistic Evaluation of Deep Partial-Label Learning Algorithms · ICLR 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2025 | A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1 · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
data augmentation |
0.2 | 1 | 2022 | Revisiting Consistency Regularization for Deep Partial Label Learning · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
random cropping · 1.7local-aggregated perturbation · 1.7embedding space alignment · 1.7model selection criteria · 0.9benchmark construction · 0.9representation refinement · 0.8rectification · 0.8noise transition matrix · 0.8bi-level optimization · 0.8consistency regularization · 0.6conformal label distribution · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Realistic Evaluation of Deep Partial-Label Learning AlgorithmsabstractPartial-label learning (PLL) is a weakly supervised learning problem in which
each example is associated with multiple candidate labels and only one is the
true label. In recent years, many deep PLL algorithms have been developed to
improve model performance. However, we find that some early developed
algorithms are often underestimated and can outperform many later algorithms
with complicated designs. In this paper, we delve into the empirical
perspective of PLL and identify several critical but previously overlooked
issues. First, model selection for PLL is non-trivial, but has never been
systematically studied. Second, the experimental settings are highly
inconsistent, making it difficult to evaluate the effectiveness of the
algorithms. Third, there is a lack of real-world image datasets that can be
compatible with modern network architectures. Based on these findings, we
propose PLENCH, the first Partial-Label learning bENCHmark to systematically
compare state-of-the-art deep PLL algorithms. We investigate the model
selection problem for PLL for the first time, and propose novel model selection
criteria with theoretical guarantees. We also create Partial-Label CIFAR-10
(PLCIFAR10), an image dataset of human-annotated partial labels collected from
Amazon Mechanical Turk, to provide a testbed for evaluating the performance of
PLL algorithms in more realistic scenarios. Researchers can quickly and
conveniently perform a comprehensive and fair evaluation and verify the
effectiveness of newly developed algorithms based on PLENCH. We hope that
PLENCH will facilitate standardized, fair, and practical evaluation of PLL
algorithms in the future. Wei Wang 0373, Dong-Dong Wu, Jindong Wang 0001, Gang Niu 0001, Min-Ling Zhang, Masashi Sugiyama |
ICLR | 2 |
| 2025 | A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1abstractDespite promising performance on open-source large vision-language models (LVLMs), transfer-based targeted attacks often fail against closed-source commercial LVLMs. Analyzing failed adversarial perturbations reveals that the learned perturbations typically originate from a uniform distribution and lack clear semantic details, resulting in unintended responses. This critical absence of semantic information leads commercial black-box LVLMs to either ignore the perturbation entirely or misinterpret its embedded semantics, thereby causing the attack to fail. To overcome these issues, we propose to refine semantic clarity by encoding explicit semantic details within local regions, thus ensuring the capture of finer-grained features and inter-model transferability, and by concentrating modifications on semantically rich areas rather than applying them uniformly. To achieve this, we propose *a simple yet highly effective baseline*: at each optimization step, the adversarial image is cropped randomly by a controlled aspect ratio and scale, resized, and then aligned with the target image in the embedding space. While the naive source-target matching method has been utilized before in the literature, we are the first to provide a tight analysis, which establishes a close connection between perturbation optimization and semantics. Experimental results confirm our hypothesis. Our adversarial examples crafted with local-aggregated perturbations focused on crucial regions exhibit surprisingly good transferability to commercial LVLMs, including GPT-4.5, GPT-4o, Gemini-2.0-flash, Claude-3.5/3.7-sonnet, and even reasoning models like o1, Claude-3.7-thinking and Gemini-2.0-flash-thinking. Our approach achieves success rates exceeding 90\% on GPT-4.5, 4o, and o1, significantly outperforming all prior state-of-the-art attack methods with lower $\ell_1/\ell_2$ perturbations. Our optimized adversarial examples under different configurations are available at https://huggingface.co/datasets/MBZUAI-LLM/M-Attack_AdvSamples and our training code at https://github.com/VILA-Lab/M-Attack. Xiaohan Zhao, Dong-Dong Wu, Jiacheng Cui |
NeurIPS | 3 |
| 2024 | Distilling Reliable Knowledge for Instance-Dependent Partial Label LearningabstractPartial label learning (PLL) refers to the classification task where each training instance is ambiguously annotated with a set of candidate labels. Despite substantial advancements in tackling this challenge, limited attention has been devoted to a more specific and realistic setting, denoted as instance-dependent partial label learning (IDPLL). Within this contex, the assignment of partial labels depends on the distinct features of individual instances, rather than being random. In this paper, we initiate an exploration into a self-distillation framework for this problem, driven by the proven effectiveness and stability of this framework. Nonetheless, a crucial shortfall is identified: the foundational assumption central to IDPLL, involving what we term as partial label knowledge stipulating that candidate labels should exhibit superior confidence compared to non-candidates, is not fully upheld within the distillation process. To address this challenge, we introduce DIRK, a novel distillation approach that leverages a rectification process to DIstill Reliable Knowledge, while concurrently preserves informative fine-grained label confidence. In addition, to harness the rectified confidence to its fullest potential, we propose a knowledge-based representation refinement module, seamlessly integrated into the DIRK framework. This module effectively transmits the essence of similarity knowledge from the label space to the feature space, thereby amplifying representation learning and subsequently engendering marked improvements in model performance. Experiments and analysis on multiple datasets validate the rationality and superiority of our proposed approach. Dong-Dong Wu, Dengbao Wang, Min-Ling Zhang |
AAAI | 1 |
| 2024 | Efficient Model Stealing Defense with Noise Transition MatrixabstractWith the escalating complexity and investment cost of training deep neural networks, safeguarding them from unauthorized usage and intellectual property theft has become imperative. Especially the rampant misuse of prediction APIs to replicate models without access to the original data or architecture poses grave security threats. Diverse defense strategies have emerged to address these vulnerabilities, yet these defenses either incur heavy inference overheads or assume idealized attack scenarios. To address these challenges, we revisit the utilization of noise transition matrix as an efficient perturbation technique, which injects noise into predicted posteriors in a linear manner and integrates seamlessly into existing systems with minimal overhead, for model stealing defense. Provably, with such perturbed posteriors, the attacker's cloning process degrades into learning from noisy data. Toward optimizing the noise transition matrix, we proposed a novel bi-level optimization training framework, which performs fidelity on the victim model while the surrogate model adversarially. Comprehensive experimental results demonstrate that our method effectively thwarts model stealing attacks and achieves minimal utility tradeoffs, outperforming existing state-of-the-art defenses. Dong-Dong Wu, Chilin Fu, Weichang Wu, Wenwen Xia, Jun Zhou 0011, Min-Ling Zhang |
CVPR | 1 |
| 2022 | Revisiting Consistency Regularization for Deep Partial Label LearningabstractPartial label learning (PLL), which refers to the classification task where each training instance is ambiguously annotated with a set of candidate labels, has been recently studied in deep learning paradigm. Despite advances in recent deep PLL literature, existing methods (e.g., methods based on self-training or contrastive learning) are confronted with either ineffectiveness or inefficiency. In this paper, we revisit a simple idea namely consistency regularization, which has been shown effective in traditional PLL literature, to guide the training of deep models. Towards this goal, a new regularized training framework, which performs supervised learning on non-candidate labels and employs consistency regularization on candidate labels, is proposed for PLL. We instantiate the regularization term by matching the outputs of multiple augmentations of an instance to a conformal label distribution, which can be adaptively inferred by the closed-form solution. Experiments on benchmark datasets demonstrate the superiority of the proposed method compared with other state-of-the-art methods. Dong-Dong Wu, Dengbao Wang, Min-Ling Zhang |
ICML | 1 |