Yudi Su

dblp:354/3610 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0006-4938-0683ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Representation and self-supervised learning · 25% Information extraction and text analysis · 21% Trustworthy machine learning · 21%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › topic model
neural topic model
1.322023
Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource Regimes · NeurIPS 2023
Bayesian Progressive Deep Topic Model with Knowledge Informed Textual Data Coarsening Process · ICML 2023
Natural language and speech › Information extraction and text analysis
topic model
1.322023
Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource Regimes · NeurIPS 2023
Bayesian Progressive Deep Topic Model with Knowledge Informed Textual Data Coarsening Process · ICML 2023
Machine learning › Representation and self-supervised learning
adaptive representation
0.912025
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation · ICML 2025
Machine learning › Trustworthy machine learning › interpretability
concept-based models
0.912025
Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification · CVPR 2025
Machine learning › Transfer learning and domain adaptation
domain generalization
0.912025
Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification · CVPR 2025
Machine learning › Efficient and distributed learning › model compression
embedding compression
0.912025
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation · ICML 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification · CVPR 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation · ICML 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.912025
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation · ICML 2025
Computer vision › Vision and language › image captioning
image caption evaluation
0.812024
HICEScore: A Hierarchical Metric for Image Captioning Evaluation · ACM Multimedia 2024
Computer vision › Vision and language
image captioning
0.812024
HICEScore: A Hierarchical Metric for Image Captioning Evaluation · ACM Multimedia 2024
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
embedding adaptation
0.712023
Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource Regimes · NeurIPS 2023
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.712023
Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource Regimes · NeurIPS 2023
Information retrieval › retrieval models › neural retrieval
embedding-based retrieval
0.312025
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation · ICML 2025

Methods — techniques the papers use, named apart from their topics

vision-language model · 2.4matryoshka representation learning · 1.7contrastive learning · 1.7hierarchical scoring · 1.5large language model · 0.9domain descriptor orthogonality regularizer · 0.9autoencoding · 0.9auto-encoding · 0.9variational inference · 0.7knowledge-informed coarsening · 0.7graph-enhanced decoder · 0.7
YearPublicationVenuePosition
2026 CAIR-Net: Reliability-Aware Information Routing for Robust Multimodal Object Detection Under Modality Degradation
abstract
Multimodal remote sensing combines optical and synthetic aperture radar (SAR) imagery to improve perception, yet real deployments face spatially varying degradations (e.g., clouds, low light, sensor interference) that can corrupt fusion. To make robustness measurable, we introduce a controlled mixed-severity setting in which only the optical stream is synthetically cloud-degraded while SAR remains intact, providing a standardized testbed for evaluating multimodal detection under modality imbalance. We further present CAIR-Net, a reliability–aware information routing network that follows adenoise-then-fuseprinciple: a Local Reliability Modulation (LRM) module learns soft, spatial reliability maps to suppress degraded regionsbeforecross-modal interaction, and a Global Information Selection Mechanism (GISM) performs confidence-aware expert routing across optical, fused, and SAR experts. On the mixed-severity benchmark, CAIR-Net consistently outperforms strong unimodal and fusion baselines and exhibits a substantially smaller performance drop under severe clouds (only a 7.3% AP reduction versus drops exceeding 25% for representative alternatives). These results indicate that explicit reliability modeling and quality-guided routing provide a practical path toward robust multimodal detection when one modality is partially or nearly completely occluded.
Yudi Su, Jialei Ni, Tiansheng Wen, Hongwei Liu 0001, Hongtao Su, Bo Chen 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification
abstract
Concept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequently undermine the model generalization capabilities, and prevent the model from being used in high-stake applications. In this paper, we propose a novel Language-guided Concept-Erasing (LanCE) framework. In particular, we empirically demonstrate that pre-trained vision-language models (VLMs) can approximate distinct visual domain shifts via domain descriptors while prompting large Language Models (LLMs) can easily simulate a wide range of descriptors of unseen visual domains. Then, we introduce a novel plug-in domain descriptor orthogonality (DDO) regularizer to mitigate the impact of these domain-specific concepts on the final predictions. Notably, the DDO regularizer is agnostic to the design of concept-based models and we integrate it into several prevailing models. Through evaluation of domain generalization on four standard benchmarks and three newly introduced benchmarks, we demonstrate that DDO can significantly improve the out-of-distribution (OOD) generalization over the previous state-of-the-art concept-based models. Our code is available at https://github.com/joeyz0z/LanCE.
Zequn Zeng, Yudi Su, Tiansheng Wen, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001, Jiawei Ma
CVPR2
2025 Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
abstract
Many large-scale systems rely on high-quality deep representations (embeddings) to facilitate tasks like retrieval, search, and generative modeling. Matryoshka Representation Learning (MRL) recently emerged as a solution for adaptive embedding lengths, but it requires full model retraining and suffers from noticeable performance degradations at short lengths. In this paper, we show that sparse coding offers a compelling alternative for achieving adaptive representation with minimal overhead and higher fidelity. We propose Contrastive Sparse Representation (CSR), a method that specifies pre-trained embeddings into a high-dimensional but selectively activated feature space. By leveraging lightweight autoencoding and task-aware contrastive objectives, CSR preserves semantic quality while allowing flexible, cost-effective inference at different sparsity levels. Extensive experiments on image, text, and multimodal benchmarks demonstrate that CSR consistently outperforms MRL in terms of both accuracy and retrieval speed—often by large margins—while also cutting training time to a fraction of that required by MRL. Our results establish sparse coding as a powerful paradigm for adaptive representation learning in real-world applications where efficiency and fidelity are both paramount. Code is available at this URL.
Tiansheng Wen, Yifei Wang 0001, Zequn Zeng, Zhong Peng, Yudi Su, Bo Chen 0001, Hongwei Liu 0001, Stefanie Jegelka, Chenyu You
ICML5
2024 HICEScore: A Hierarchical Metric for Image Captioning Evaluation
abstract
Image captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to evaluate descriptive captions with abundant visual details produced by advanced multimodal large language models, due to their heavy reliance on limited human-annotated references. In contrast, previous reference-free metrics have been proven effective via CLIP cross-modality similarity. Nonetheless, CLIP-based metrics, constrained by their solution of global image-text compatibility, often have a deficiency in detecting local textual hallucinations and are insensitive to small visual objects. Besides, their single-scale designs are unable to provide an interpretable evaluation process such as pinpointing the position of caption mistakes and identifying visual regions that have not been described. To move forward, we propose a novel reference-free metric for image captioning evaluation, dubbed Hierarchical Image Captioning Evaluation Score (HICE-S). By detecting local visual regions and textual phrases, HICE-S builds an interpretable hierarchical scoring mechanism, breaking through the barriers of the single-scale structure of existing reference-free metrics. Comprehensive experiments indicate that our proposed metric achieves the SOTA performance on several benchmarks, outperforming existing reference-free metrics like CLIP-S and PAC-S, and reference-based metrics like METEOR and CIDEr. Moreover, several case studies reveal that the assessment process of HICE-S on detailed captions closely resembles interpretable human judgments.Our code is available at https://github.com/joeyz0z/HICE.
Zequn Zeng, Hao Zhang 0050, Tiansheng Wen, Yudi Su, Zhengjue Wang, Bo Chen 0001
ACM Multimedia5
2023 Bayesian Progressive Deep Topic Model with Knowledge Informed Textual Data Coarsening Process
abstract
Deep topic models have shown an impressive ability to extract multi-layer document latent representations and discover hierarchical semantically meaningful topics.However, most deep topic models are limited to the single-step generative process, despite the fact that the progressive generative process has achieved impressive performance in modeling image data. To this end, in this paper, we propose a novel progressive deep topic model that consists of a knowledge-informed textural data coarsening process and a corresponding progressive generative model. The former is used to build multi-level observations ranging from concrete to abstract, while the latter is used to generate more concrete observations gradually. Additionally, we incorporate a graph-enhanced decoder to capture the semantic relationships among words at different levels of observation. Furthermore, we perform a theoretical analysis of the proposed model based on the principle of information theory and show how it can alleviate the well-known "latent variable collapse" problem. Finally, extensive experiments demonstrate that our proposed model effectively improves the ability of deep topic models, resulting in higher-quality latent document representations and topics.
Zhibin Duan, Yudi Su, Yishi Xu, Bo Chen 0001, Mingyuan Zhou
ICML3
2023 Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource Regimes
abstract
Embedding-based neural topic models have turned out to be a superior option for low-resourced topic modeling. However, current approaches consider static word embeddings learnt from source tasks as general knowledge that can be transferred directly to the target task, discounting the dynamically changing nature of word meanings in different contexts, thus typically leading to sub-optimal results when adapting to new tasks with unfamiliar contexts. To settle this issue, we provide an effective method that centers on adaptively generating semantically tailored word embeddings for each task by fully exploiting contextual information. Specifically, we first condense the contextual syntactic dependencies of words into a semantic graph for each task, which is then modeled by a Variational Graph Auto-Encoder to produce task-specific word representations. On this basis, we further impose a learnable Gaussian mixture prior on the latent space of words to efficiently learn topic representations from a clustering perspective, which contributes to diverse topic discovery and fast adaptation to novel tasks. We have conducted a wealth of quantitative and qualitative experiments, and the results show that our approach comprehensively outperforms established topic models.
Yishi Xu, Yudi Su, Zhibin Duan, Bo Chen 0001, Mingyuan Zhou
NeurIPS3