VLDB 2026 Research / reviewers in the wild / expert
Yi-Kai Zhang
dblp:330/8964
· DBLP profile ↗
8ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Efficient and distributed learning · 24% Trustworthy machine learning · 18% Learning paradigms · 13% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 2 | 2025 | Wings: Learning Multimodal LLMs without Text-only Forgetting · NeurIPS 2024 ZooProbe: A Data Engine for Evaluating, Exploring, and Evolving Large-scale Training Data for Multimodal LLMs · ICLR 2025 |
Machine learning › Efficient and distributed learning › inference efficiency
LLM routing |
0.9 | 1 | 2025 | Capability Instruction Tuning · AAAI 2025 |
Machine learning › Learning theory
model selection |
0.9 | 1 | 2025 | Capability Instruction Tuning · AAAI 2025 |
Machine learning › Efficient and distributed learning › data curation
training data curation |
0.9 | 1 | 2025 | ZooProbe: A Data Engine for Evaluating, Exploring, and Evolving Large-scale Training Data for Multimodal LLMs · ICLR 2025 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.8 | 1 | 2024 | Wings: Learning Multimodal LLMs without Text-only Forgetting · NeurIPS 2024 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.8 | 1 | 2024 | Wings: Learning Multimodal LLMs without Text-only Forgetting · NeurIPS 2024 |
Machine learning › Learning paradigms › continual learning
class-incremental learning |
0.7 | 1 | 2023 | Few-Shot Class-Incremental Learning via Training-Free Prototype Calibration · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › debiasing
debiased representation learning |
0.7 | 1 | 2023 | Learning Debiased Representations via Conditional Attribute Interpolation · CVPR 2023 |
Machine learning › Trustworthy machine learning
fairness |
0.7 | 1 | 2023 | Learning Debiased Representations via Conditional Attribute Interpolation · CVPR 2023 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.7 | 1 | 2023 | Few-Shot Class-Incremental Learning via Training-Free Prototype Calibration · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
model ranking |
0.7 | 1 | 2023 | Model Spider: Learning to Rank Pre-Trained Models Efficiently · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model reuse
model zoo |
0.7 | 1 | 2023 | Model Spider: Learning to Rank Pre-Trained Models Efficiently · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation › pre-trained models
pre-trained model selection |
0.7 | 1 | 2023 | Model Spider: Learning to Rank Pre-Trained Models Efficiently · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › prototype learning
prototype calibration |
0.7 | 1 | 2023 | Few-Shot Class-Incremental Learning via Training-Free Prototype Calibration · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
model zoo · 0.9model capability encoder · 0.9heuristic quality ranking · 0.9capability instruction tuning · 0.9a* search · 0.9token-wise routing · 0.8low-rank residual attention · 0.8x2-model · 0.7metric learning · 0.7conditional attribute interpolation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Capability Instruction TuningabstractLarge Language Models (LLMs) have demonstrated human-like instruction-following abilities, particularly those exceeding 100 billion parameters. The combined capability of some smaller, resource-friendly LLMs can address most of the instructions that larger LLMs excel at. In this work, we explore how to route the best-performing LLM for each instruction to achieve better overall performance. We develop a new paradigm, constructing capability instructions with model capability representation, user instruction, and performance inquiry prompts to assess the performance. To learn from capability instructions, we introduce a new end-to-end framework called Model Selection with Aptitude Test (Model-SAT), which generates positive and negative samples based on what different models perform well or struggle with. Model-SAT uses a model capability encoder that extends its model representation to a lightweight LLM. Our experiments show that Model-SAT understands the performance dimensions of candidate models and provides the probabilities of their capability to handle various instructions. Additionally, during deployment, a new model can quickly infer its aptitude test results across 50 tasks, each with 20 shots. Model-SAT performs state-of-the-art model routing without candidate inference and in real-world new model-released scenarios. Yi-Kai Zhang, De-Chuan Zhan, Han-Jia Ye |
AAAI | 1 |
| 2025 | ZooProbe: A Data Engine for Evaluating, Exploring, and Evolving Large-scale Training Data for Multimodal LLMsabstractMultimodal Large Language Models (MLLMs) are thriving through continuous fine-tuning by LLMs. Driven by the law that "scale is everything", MLLMs expand their training sets during version iterations. In this paper, we propose a large-scale training data engine built around an evaluating-exploring-evolving (E3) loop. Evaluating the data provides insights into its characteristics. Exploring quality rules helps identify which data enhances training. Together, these processes facilitate the systematic evolution of new, high-quality data. With the E3 loop, we introduce ZooProbe, an efficient data engine for MLLMs. First, the problem of data expansion is formalized as a tree of sampling and growth. ZooProbe introduces a small-scale model *zoo* to obtain comprehensive evaluations for child datasets. From multiple perspectives, visual, textual, and multimodal models cover over 50 dimensions of intrinsic and meta attributes, such as object and topic distribution, and higher-level properties, like annotation quality and scene complexity. ZooProbe constructs based on A$^\star$ search, modeling the heuristic function as a quality estimate from data evaluation results. It dynamically explores the rule of data quality based on the model state of the *probe* datasets. Additionally, it evolves new targeted data with identified high-quality rules. We also develop an extra heuristic quality ranker with the data utilized and discarded during the expansion. Our experiments show that ZooProbe significantly breaks the scaling law in multimodal instruction fine-tuning at scales of 260$k$ and below.
ZooProbe generates high-quality data that accelerates MLLM training and enhances performance, automating the evolution of large-scale training data. Yi-Kai Zhang, Shiyin Lu, De-Chuan Zhan, Han-Jia Ye |
ICLR | 1 |
| 2025 | Let the LLM Stick to Its Strengths: Learning to Route Economical LLMabstractRecently, test-time scaling of Large Language Models (LLMs) has emerged as a practical alternative to parameter and data scaling. Reasoning tasks often require large-scale, RLVR-based LLMs, while more economical LLMs can handle simpler tasks. Routing an LLM tailored to *suitability* (*i.e.*, capability and cost) ensures usability and efficiency. We introduce LLMRec, which routes the most suitable LLM to the user query without pre-inference on the candidate LLM zoo. It pioneeringly reframes the LLM routing problem as a comprehensive recommendation system (RecSys) task. Our core insight is that an LLM's suitability for a query is a complex, latent signal equal to user-item preference. LLMRec systematically engineers features for candidate LLMs (intrinsic attributes and capability distributions), queries (general semantics and meta-dimensional info), and context (inference type, cost budgets). It also incorporates behavioral features to learn high-order interactions. LLMRec is designed to generalize to out-of-domain datasets and adapt to new LLMs as the model zoo evolves. We define the metric with the Pareto frontier under user-specified cost budgets. Across six datasets, LLMRec achieves an average cost reduction of over 38% while maintaining accuracy and consistently outperforming baselines in converging toward the Pareto frontier. Yi-Kai Zhang, Shiyin Lu, Weihua Luo, De-Chuan Zhan, Han-Jia Ye |
NeurIPS | 1 |
| 2024 | Wings: Learning Multimodal LLMs without Text-only ForgettingabstractMultimodal large language models (MLLMs), initiated with a trained LLM, first align images with text and then fine-tune on multimodal mixed inputs. However, during the continued training, the MLLM catastrophically forgets the text-only instructions that the initial LLM masters. In this paper, we present Wings, a novel MLLM that excels in both text-only and multimodal instructions. By examining attention across layers of MLLM, we find that *text-only forgetting* is related to the attention shifts from pre-image to post-image text. From that, we construct an additional Low-Rank Residual Attention (LoRRA) block that acts as the "modality learner" to expand the learnable space and compensate for the attention shift. The complementary learners, like "wings" on either side, are connected in parallel to each layer's attention block. The LoRRA mirrors the structure of attention but utilizes low-rank connections to ensure efficiency. Initially, image and text inputs are aligned with visual learners operating alongside the main attention, balancing focus on visual elements. Later, textual learners are integrated with token-wise routing, blending the outputs of both modality learners collaboratively. Our experimental results demonstrate that Wings outperforms equally-scaled MLLMs in both text-only and visual question-answering tasks. Wings with *compensation of learners* addresses text-only forgetting during visual modality expansion in general MLLMs. Yi-Kai Zhang, Shiyin Lu, Yanqing Ma, Weihua Luo, Kaifu Zhang, De-Chuan Zhan, Han-Jia Ye |
NeurIPS | 1 |
| 2023 | Learning Debiased Representations via Conditional Attribute InterpolationabstractAn image is usually described by more than one attribute like “shape” and “color”. When a dataset is biased, i.e., most samples have attributes spuriously correlated with the target label, a Deep Neural Network (DNN) is prone to make predictions by the “unintended” attribute, especially if it is easier to learn. To improve the generalization ability when training on such a biased dataset, we propose a X2-model to learn debiased representations. First, we design a x-shape pattern to match the training dynamics of a DNN and find Intermediate Attribute Samples (IASs) — samples near the attribute decision boundaries, which indicate how the value of an attribute changes from one extreme to another. Then we rectify the representation with a X-structured metric learning objective. Conditional interpolation among IASs eliminates the negative effect of periph-eral attributes and facilitates retaining the intra-class compactness. Experiments show that X2-modellearns debiased representation effectively and achieves remarkable improvements on various datasets. Code is available at: https://github.com/ZhangYikaii/chi-square Yi-Kai Zhang, Qi-Wei Wang, De-Chuan Zhan, Han-Jia Ye |
CVPR | 1 |
| 2023 | Few-Shot Class-Incremental Learning via Training-Free Prototype CalibrationabstractReal-world scenarios are usually accompanied by continuously appearing classes with scare labeled samples, which require the machine learning model to incrementally learn new classes and maintain the knowledge of base classes. In this Few-Shot Class-Incremental Learning (FSCIL) scenario, existing methods either introduce extra learnable components or rely on a frozen feature extractor to mitigate catastrophic forgetting and overfitting problems. However, we find a tendency for existing methods to misclassify the samples of new classes into base classes, which leads to the poor performance of new classes. In other words, the strong discriminability of base classes distracts the classification of new classes. To figure out this intriguing phenomenon, we observe that although the feature extractor is only trained on base classes, it can surprisingly represent the *semantic similarity* between the base and *unseen* new classes. Building upon these analyses, we propose a *simple yet effective* Training-frEE calibratioN (TEEN) strategy to enhance the discriminability of new classes by fusing the new prototypes (i.e., mean features of a class) with weighted base prototypes. In addition to standard benchmarks in FSCIL, TEEN demonstrates remarkable performance and consistent improvements over baseline methods in the few-shot learning scenario. Code is available at: https://github.com/wangkiw/TEEN Qi-Wei Wang, Da-Wei Zhou 0001, Yi-Kai Zhang, De-Chuan Zhan, Han-Jia Ye |
NeurIPS | 3 |
| 2023 | Model Spider: Learning to Rank Pre-Trained Models EfficientlyabstractFiguring out which Pre-Trained Model (PTM) from a model zoo fits the target task is essential to take advantage of plentiful model resources. With the availability of numerous heterogeneous PTMs from diverse fields, efficiently selecting the most suitable one is challenging due to the time-consuming costs of carrying out forward or backward passes over all PTMs. In this paper, we propose Model Spider, which tokenizes both PTMs and tasks by summarizing their characteristics into vectors to enable efficient PTM selection. By leveraging the approximated performance of PTMs on a separate set of training tasks, Model Spider learns to construct representation and measure the fitness score between a model-task pair via their representation. The ability to rank relevant PTMs higher than others generalizes to new tasks. With the top-ranked PTM candidates, we further learn to enrich task repr. with their PTM-specific semantics to re-rank the PTMs for better selection. Model Spider balances efficiency and selection ability, making PTM selection like a spider preying on a web. Model Spider exhibits promising performance across diverse model zoos, including visual models and Large Language Models (LLMs). Code is available at https://github.com/zhangyikaii/Model-Spider. Yi-Kai Zhang, Ting-Ji Huang, Yao-Xiang Ding 0001, De-Chuan Zhan, Han-Jia Ye |
NeurIPS | 1 |
| 2022 | Audio-Visual Generalized Few-Shot Learning with Prototype-Based Co-Adaptation
Yi-Kai Zhang, Da-Wei Zhou 0001, Han-Jia Ye, De-Chuan Zhan |
INTERSPEECH | 1 |