VLDB 2026 Research / reviewers in the wild / expert
Yunshu Wu
dblp:339/0713
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2024
0009-0002-4537-8800ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Generative modeling · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 50% Knowledge graphs · 50% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model
denoising |
0.8 | 1 | 2024 | Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
parallel sampling |
0.8 | 1 | 2024 | Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training · NeurIPS 2024 |
Knowledge graphs › knowledge graph alignment
entity alignment |
0.6 | 1 | 2022 | TENALIGN: Joint Tensor Alignment and Coupled Factorization · ICDM 2022 |
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis
tensor factorization |
0.6 | 1 | 2022 | TENALIGN: Joint Tensor Alignment and Coupled Factorization · ICDM 2022 |
Methods — techniques the papers use, named apart from their topics
self-supervised learning · 0.8log-likelihood ratio · 0.8contrastive training · 0.8embedding-based matching · 0.6coupled tensor factorization · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive TrainingabstractDiffusion models learn to denoise data and the trained denoiser is then used to generate new samples from the data distribution.
In this paper, we revisit the diffusion sampling process and identify a fundamental cause of sample quality degradation: the denoiser is poorly estimated in regions that are far Outside Of the training Distribution (OOD), and the sampling process inevitably evaluates in these OOD regions.
This can become problematic for all sampling methods, especially when we move to parallel sampling which requires us to initialize and update the entire sample trajectory of dynamics in parallel, leading to many OOD evaluations.
To address this problem, we introduce a new self-supervised training objective that differentiates the levels of noise added to a sample, leading to improved OOD denoising performance. The approach is based on our observation that diffusion models implicitly define a log-likelihood ratio that distinguishes distributions with different amounts of noise, and this expression depends on denoiser performance outside the standard training distribution.
We show by diverse experiments that the proposed contrastive diffusion training is effective for both sequential and parallel settings, and it improves the performance and speed of parallel samplers significantly. Code for our paper can be found at https://github.com/yunshuwu/ContrastiveDiffusionLoss Yunshu Wu, Yingtao Luo, Xianghao Kong, Evangelos E. Papalexakis, Greg Ver Steeg |
NeurIPS | 1 |
| 2023 | High-Performance and Flexible Parallel Algorithms for Semisort and Related ProblemsabstractSemisort is a fundamental algorithmic primitive widely used in the design and analysis of efficient parallel algorithms. It takes input as an array of records and a function extracting a key per record, and reorders them so that records with equal keys are contiguous. Since many applications only require collecting equal values, but not fully sorting the input, semisort is broadly applicable, e.g., in string algorithms, graph analytics, and geometry processing, among many other domains. However, despite dozens of recent papers that use semisort in their theoretical analysis and the existence of an asymptotically optimal parallel semisort algorithm, most implementations of these parallel algorithms choose to implement semisort by using comparison or integer sorting in practice, due to potential performance issues in existing semisort implementations. Xiaojun Dong 0001, Yunshu Wu, Laxman Dhulipala, Yan Gu 0001, Yihan Sun 0001 |
SPAA | 2 |
| 2022 | TENALIGN: Joint Tensor Alignment and Coupled FactorizationabstractMultimodal datasets represented as tensors oftentimes share some of their modes. However, even though there may exist a one-to-one (or perhaps partial) correspondence between the coupled modes, such correspondence/alignment may not be given, especially when integrating datasets from disparate sources. This is a very important problem, broadly termed as entity alignment or matching, and subsets of the problem such as graph matching have been extremely popular in the recent years. In order to solve this problem, current work computes the alignment based on existing embeddings of the data. This can be problematic if our end goal is the joint analysis of the two datasets into the same latent factor space: the embeddings computed separately per dataset may yield a suboptimal alignment, and if such an alignment is used to subsequently compute the joint latent factors, the computation will similarly be plagued by compounding errors incurred by the imperfect alignment. In this work, we are the first to define and solve the problem of joint tensor alignment and factorization into a shared latent space. By posing this as a unified problem and solving for both tasks simultaneously, we observe that the both alignment and factorization tasks benefit each other resulting in superior performance compared to two-stage approaches. We extensively evaluate our proposed method TENALIGN and conduct a thorough sensitivity and ablation analysis. We demonstrate that TENALIGN significantly outperforms baseline approaches where embedding and matching happen separately. Yunshu Wu, Uday Singh Saini, Jia Chen 0002, Evangelos E. Papalexakis |
ICDM | 1 |