VLDB 2026 Research / reviewers in the wild / expert
Yudi Su
dblp:354/3610
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0006-4938-0683ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Representation and self-supervised learning · 25% Information extraction and text analysis · 21% Trustworthy machine learning · 21% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › topic model
neural topic model |
1.3 | 2 | 2023 | Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource Regimes · NeurIPS 2023 Bayesian Progressive Deep Topic Model with Knowledge Informed Textual Data Coarsening Process · ICML 2023 |
Natural language and speech › Information extraction and text analysis
topic model |
1.3 | 2 | 2023 | Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource Regimes · NeurIPS 2023 Bayesian Progressive Deep Topic Model with Knowledge Informed Textual Data Coarsening Process · ICML 2023 |
Machine learning › Representation and self-supervised learning
adaptive representation |
0.9 | 1 | 2025 | Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability
concept-based models |
0.9 | 1 | 2025 | Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification · CVPR 2025 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
0.9 | 1 | 2025 | Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification · CVPR 2025 |
Machine learning › Efficient and distributed learning › model compression
embedding compression |
0.9 | 1 | 2025 | Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation · ICML 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification · CVPR 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation · ICML 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification · CVPR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding |
0.9 | 1 | 2025 | Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation · ICML 2025 |
Computer vision › Vision and language › image captioning
image caption evaluation |
0.8 | 1 | 2024 | HICEScore: A Hierarchical Metric for Image Captioning Evaluation · ACM Multimedia 2024 |
Computer vision › Vision and language
image captioning |
0.8 | 1 | 2024 | HICEScore: A Hierarchical Metric for Image Captioning Evaluation · ACM Multimedia 2024 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
embedding adaptation |
0.7 | 1 | 2023 | Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource Regimes · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.7 | 1 | 2023 | Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource Regimes · NeurIPS 2023 |
Information retrieval › retrieval models › neural retrieval
embedding-based retrieval |
0.3 | 1 | 2025 | Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 2.4matryoshka representation learning · 1.7contrastive learning · 1.7hierarchical scoring · 1.5large language model · 0.9domain descriptor orthogonality regularizer · 0.9autoencoding · 0.9auto-encoding · 0.9variational inference · 0.7knowledge-informed coarsening · 0.7graph-enhanced decoder · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAIR-Net: Reliability-Aware Information Routing for Robust Multimodal Object Detection Under Modality DegradationabstractMultimodal remote sensing combines optical and synthetic aperture radar (SAR) imagery to improve perception, yet real deployments face spatially varying degradations (e.g., clouds, low light, sensor interference) that can corrupt fusion. To make robustness measurable, we introduce a controlled mixed-severity setting in which only the optical stream is synthetically cloud-degraded while SAR remains intact, providing a standardized testbed for evaluating multimodal detection under modality imbalance. We further present CAIR-Net, a reliability–aware information routing network that follows adenoise-then-fuseprinciple: a Local Reliability Modulation (LRM) module learns soft, spatial reliability maps to suppress degraded regionsbeforecross-modal interaction, and a Global Information Selection Mechanism (GISM) performs confidence-aware expert routing across optical, fused, and SAR experts. On the mixed-severity benchmark, CAIR-Net consistently outperforms strong unimodal and fusion baselines and exhibits a substantially smaller performance drop under severe clouds (only a 7.3% AP reduction versus drops exceeding 25% for representative alternatives). These results indicate that explicit reliability modeling and quality-guided routing provide a practical path toward robust multimodal detection when one modality is partially or nearly completely occluded. Yudi Su, Jialei Ni, Tiansheng Wen, Hongwei Liu 0001, Hongtao Su, Bo Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image ClassificationabstractConcept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequently undermine the model generalization capabilities, and prevent the model from being used in high-stake applications. In this paper, we propose a novel Language-guided Concept-Erasing (LanCE) framework. In particular, we empirically demonstrate that pre-trained vision-language models (VLMs) can approximate distinct visual domain shifts via domain descriptors while prompting large Language Models (LLMs) can easily simulate a wide range of descriptors of unseen visual domains. Then, we introduce a novel plug-in domain descriptor orthogonality (DDO) regularizer to mitigate the impact of these domain-specific concepts on the final predictions. Notably, the DDO regularizer is agnostic to the design of concept-based models and we integrate it into several prevailing models. Through evaluation of domain generalization on four standard benchmarks and three newly introduced benchmarks, we demonstrate that DDO can significantly improve the out-of-distribution (OOD) generalization over the previous state-of-the-art concept-based models. Our code is available at https://github.com/joeyz0z/LanCE. Zequn Zeng, Yudi Su, Tiansheng Wen, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001, Jiawei Ma |
CVPR | 2 |
| 2025 | Beyond Matryoshka: Revisiting Sparse Coding for Adaptive RepresentationabstractMany large-scale systems rely on high-quality deep representations (embeddings) to facilitate tasks like retrieval, search, and generative modeling. Matryoshka Representation Learning (MRL) recently emerged as a solution for adaptive embedding lengths, but it requires full model retraining and suffers from noticeable performance degradations at short lengths. In this paper, we show that sparse coding offers a compelling alternative for achieving adaptive representation with minimal overhead and higher fidelity. We propose Contrastive Sparse Representation (CSR), a method that specifies pre-trained embeddings into a high-dimensional but selectively activated feature space. By leveraging lightweight autoencoding and task-aware contrastive objectives, CSR preserves semantic quality while allowing flexible, cost-effective inference at different sparsity levels. Extensive experiments on image, text, and multimodal benchmarks demonstrate that CSR consistently outperforms MRL in terms of both accuracy and retrieval speed—often by large margins—while also cutting training time to a fraction of that required by MRL. Our results establish sparse coding as a powerful paradigm for adaptive representation learning in real-world applications where efficiency and fidelity are both paramount. Code is available at this URL. Tiansheng Wen, Yifei Wang 0001, Zequn Zeng, Zhong Peng, Yudi Su, Bo Chen 0001, Hongwei Liu 0001, Stefanie Jegelka, Chenyu You |
ICML | 5 |
| 2024 | HICEScore: A Hierarchical Metric for Image Captioning EvaluationabstractImage captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to evaluate descriptive captions with abundant visual details produced by advanced multimodal large language models, due to their heavy reliance on limited human-annotated references. In contrast, previous reference-free metrics have been proven effective via CLIP cross-modality similarity. Nonetheless, CLIP-based metrics, constrained by their solution of global image-text compatibility, often have a deficiency in detecting local textual hallucinations and are insensitive to small visual objects. Besides, their single-scale designs are unable to provide an interpretable evaluation process such as pinpointing the position of caption mistakes and identifying visual regions that have not been described. To move forward, we propose a novel reference-free metric for image captioning evaluation, dubbed Hierarchical Image Captioning Evaluation Score (HICE-S). By detecting local visual regions and textual phrases, HICE-S builds an interpretable hierarchical scoring mechanism, breaking through the barriers of the single-scale structure of existing reference-free metrics. Comprehensive experiments indicate that our proposed metric achieves the SOTA performance on several benchmarks, outperforming existing reference-free metrics like CLIP-S and PAC-S, and reference-based metrics like METEOR and CIDEr. Moreover, several case studies reveal that the assessment process of HICE-S on detailed captions closely resembles interpretable human judgments.Our code is available at https://github.com/joeyz0z/HICE. Zequn Zeng, Hao Zhang 0050, Tiansheng Wen, Yudi Su, Zhengjue Wang, Bo Chen 0001 |
ACM Multimedia | 5 |
| 2023 | Bayesian Progressive Deep Topic Model with Knowledge Informed Textual Data Coarsening ProcessabstractDeep topic models have shown an impressive ability to extract multi-layer document latent representations and discover hierarchical semantically meaningful topics.However, most deep topic models are limited to the single-step generative process, despite the fact that the progressive generative process has achieved impressive performance in modeling image data. To this end, in this paper, we propose a novel progressive deep topic model that consists of a knowledge-informed textural data coarsening process and a corresponding progressive generative model. The former is used to build multi-level observations ranging from concrete to abstract, while the latter is used to generate more concrete observations gradually. Additionally, we incorporate a graph-enhanced decoder to capture the semantic relationships among words at different levels of observation. Furthermore, we perform a theoretical analysis of the proposed model based on the principle of information theory and show how it can alleviate the well-known "latent variable collapse" problem. Finally, extensive experiments demonstrate that our proposed model effectively improves the ability of deep topic models, resulting in higher-quality latent document representations and topics. Zhibin Duan, Yudi Su, Yishi Xu, Bo Chen 0001, Mingyuan Zhou |
ICML | 3 |
| 2023 | Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource RegimesabstractEmbedding-based neural topic models have turned out to be a superior option for low-resourced topic modeling. However, current approaches consider static word embeddings learnt from source tasks as general knowledge that can be transferred directly to the target task, discounting the dynamically changing nature of word meanings in different contexts, thus typically leading to sub-optimal results when adapting to new tasks with unfamiliar contexts. To settle this issue, we provide an effective method that centers on adaptively generating semantically tailored word embeddings for each task by fully exploiting contextual information. Specifically, we first condense the contextual syntactic dependencies of words into a semantic graph for each task, which is then modeled by a Variational Graph Auto-Encoder to produce task-specific word representations. On this basis, we further impose a learnable Gaussian mixture prior on the latent space of words to efficiently learn topic representations from a clustering perspective, which contributes to diverse topic discovery and fast adaptation to novel tasks. We have conducted a wealth of quantitative and qualitative experiments, and the results show that our approach comprehensively outperforms established topic models. Yishi Xu, Yudi Su, Zhibin Duan, Bo Chen 0001, Mingyuan Zhou |
NeurIPS | 3 |