VLDB 2026 Research / reviewers in the wild / expert
Yuhang Ge
dblp:340/3648
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-2522-1753ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning › data preprocessing
data cleaning |
0.9 | 1 | 2025 | KnowTrans: Boosting Transferability of Data Preparation LLMs via Knowledge Augmentation · ICDE 2025 |
Data integration and cleaning
data preprocessing |
0.9 | 1 | 2025 | KnowTrans: Boosting Transferability of Data Preparation LLMs via Knowledge Augmentation · ICDE 2025 |
Data integration and cleaning › missing data
missing value imputation |
0.9 | 1 | 2025 | KnowTrans: Boosting Transferability of Data Preparation LLMs via Knowledge Augmentation · ICDE 2025 |
Methods — techniques the papers use, named apart from their topics
selective knowledge concentration · 0.9large language model · 0.9knowledge augmentation · 0.9automatic knowledge bridging · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KnowTrans: Boosting Transferability of Data Preparation LLMs via Knowledge AugmentationabstractData Preparation (DP), which involves tasks such as data cleaning, imputation and integration, is a fundamental process in data-driven applications. Recently, Large Language Models (LLMs) fine-tuned for DP tasks, i.e., DP-LLMs, have achieved state-of-the-art performance. However, transferring DP-LLMs to novel datasets and tasks typically requires a substantial amount of labeled data, which is impractical in many real-world scenarios. To address this, we propose a knowledge augmentation framework for data preparation, dubbed KNOWTRANS. This framework allows DP-LLMs to be transferred to novel datasets and tasks with a few data points, significantly decreasing the dependence on extensive labeled data. KNOWTRANS comprises two components: Selective Knowledge Concentration and Automatic Knowledge Bridging. The first component re-uses knowledge from previously learned tasks, while the second automatically integrates additional knowledge from external sources. Extensive experiments on 13 datasets demonstrate the effectiveness of KNOWTRANS. KNOWTRANS boosts the performance of the state-of-the-art DP-LLM, Jellyfish-7B, by an average of 4.93%, enabling it to outperform both GPT-4 and GPT-4o. Yuhang Ge, Fengyu Li, Yuren Mao, Congcong Ge, Zhaoqiang Chen, Yunjun Gao |
ICDE | 1 |
| 2025 | A survey on LoRA of large language modelsabstractAbstract Low-Rank Adaptation (LoRA), which updates the dense neural network layers with pluggable low-rank matrices, is one of the best performed parameter efficient fine-tuning paradigms. Furthermore, it has significant advantages in cross-task generalization and privacy-preserving. Hence, LoRA has gained much attention recently, and the number of related literature demonstrates exponential growth. It is necessary to conduct a comprehensive overview of the current progress on LoRA. This survey categorizes and reviews the progress from the perspectives of (1) downstream adaptation improving variants that improve LoRA’s performance on downstream tasks; (2) cross-task generalization methods that mix multiple LoRA plugins to achieve cross-task generalization; (3) efficiency-improving methods that boost the computation-efficiency of LoRA; (4) data privacy-preserving methods that use LoRA in federated learning; (5) application. Besides, this survey also discusses the future directions in this field. Yuren Mao, Yuhang Ge, Yijiang Fan, Yu Mi, Zhonghao Hu, Yunjun Gao |
Frontiers Comput. Sci. | 2 |
| 2024 | Learning shared and non-redundant label-specific features for partial multi-label classification
Yizhang Zou, Xuegang Hu, Pei-Pei Li 0001, Yuhang Ge |
Inf. Sci. | 4 |
| 2022 | Multi-label Learning with Data Self-augmentation
Yuhang Ge, Xuegang Hu, Pei-Pei Li 0001, Haobo Wang 0001, Junbo Zhao 0002 |
ICONIP (4) | 1 |