VLDB 2026 Research / reviewers in the wild / expert
Peng Tan 0002
dblp:69/1700-2
· DBLP profile ↗
6ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-3749-9266ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Efficient and distributed learning · 72% Transfer learning and domain adaptation · 14% Representation and self-supervised learning · 14% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 57% Machine learning and data management · 43% |
Topics — the 4 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model reuse |
2.2 | 4 | 2026 | Tabular Learnwares Can Be Repurposed for Seemingly Irrelevant New Tasks · AAAI 2026 Beimingwu: A Learnware Dock System · KDD 2024 Handling Learnwares from Heterogeneous Feature Spaces with Explicit Label Exploitation · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model reuse
learnware |
1.7 | 2 | 2026 | Tabular Learnwares Can Be Repurposed for Seemingly Irrelevant New Tasks · AAAI 2026 Handling Learnwares Developed from Heterogeneous Feature Spaces without Auxiliary Data · IJCAI 2023 |
Data mining › predictive modeling › classification
tabular data classification |
1.0 | 1 | 2026 | Tabular Learnwares Can Be Repurposed for Seemingly Irrelevant New Tasks · AAAI 2026 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
unified embedding |
0.8 | 1 | 2024 | Handling Learnwares from Heterogeneous Feature Spaces with Explicit Label Exploitation · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
reduced kernel mean embedding · 2.2probability-based model reuse · 2.0model specification · 1.5pseudo-label exploitation · 0.8conditional distribution comparison · 0.8matrix factorization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tabular Learnwares Can Be Repurposed for Seemingly Irrelevant New TasksabstractThe learnware paradigm aims to help users solve new tasks by reusing existing models rather than starting from scratch. A learnware consists of a model and the specification describing its capabilities. Numerous learnwares are accommodated by the learnware dock system. When users solve tasks with the system, learnwares that fully match the user task are often scarce or unavailable. This paper focuses on tabular classification tasks and explores reusing learnwares for new user tasks with significantly different feature and label spaces, leveraging the potential of numerous existing specialized tabular models developed for various tasks. Under the learnware paradigm, we find that tabular learnwares that seem semantically irrelevant can sometimes be beneficial for new user tasks. The proposed method relies solely on model-predicted probabilities and does not require gradient information, making it applicable to a wide range of tabular models. Experiments suggest that tabular learnwares can be reused beyond their original purpose across heterogeneous tasks. Peng Tan 0002, Zhi-Hao Tan, Zhi-Hua Zhou |
AAAI | 1 |
| 2024 | Beimingwu: A Learnware Dock SystemabstractThe learnware paradigm proposed by Zhou [40] aims to enable users to leverage numerous existing high-performing models instead of building machine learning models from scratch.This paradigm envisions that: Any developer worldwide can submit their well-trained models spontaneously into a learnware dock system (formerly known as learnware market).The system uniformly generates a specification for each model to form a learnware and accommodates it.As the key component, a specification should represent the capabilities of the model while preserving developer's original data.Based on the specifications, the learnware dock system can identify and assemble existing learnwares for users to solve new machine learning tasks.Recently, based on reduced kernel mean embedding (RKME) specification, a series of studies have shown the effectiveness of the learnware paradigm theoretically and empirically.However, the realization of a learnware dock system is still missing and remains a big challenge.This paper proposes Beimingwu, the first open-source learnware dock system, providing foundational support for future research.The system provides implementations and extensibility for the entire process of learnware paradigm, including the submitting, usability testing, organization, identification, deployment, and reuse of learnwares.Utilizing Beimingwu, the model development for new user tasks can be significantly streamlined, thanks to integrated architecture and engine design, specifying unified learnware structure and scalable APIs, and the integration of various algorithms for learnware identification and reuse.Notably, this is possible even for users with limited data and minimal expertise in machine learning, without compromising the raw data's security.The system facilitates the future research implementations in learnware-related algorithms and systems, and lays the ground for hosting a vast array of learnwares and establishing a learnware ecosystem.The system is fully open-source and we expect the research community Zhi-Hao Tan, Jian-Dong Liu, Xiaodong Bi, Peng Tan 0002, Qin-Cheng Zheng, Hai-Tian Liu, Xiao-Chuan Zou, Yang Yu 0001, Zhi-Hua Zhou |
KDD | 4 |
| 2024 | Handling Learnwares from Heterogeneous Feature Spaces with Explicit Label ExploitationabstractThe learnware paradigm aims to help users leverage numerous existing high-performing models instead of starting from scratch, where a learnware consists of a well-trained model and the specification describing its capability. Numerous learnwares are accommodated by a learnware dock system. When users solve tasks with the system, models that fully match the task feature space are often rare or even unavailable. However, models with heterogeneous feature space can still be helpful. This paper finds that label information, particularly model outputs, is helpful yet previously less exploited in the accommodation of heterogeneous learnwares. We extend the specification to better leverage model pseudo-labels and subsequently enrich the unified embedding space for better specification evolvement. With label information, the learnware identification can also be improved by additionally comparing conditional distributions. Experiments demonstrate that, even without a model explicitly tailored to user tasks, the system can effectively handle tasks by leveraging models from diverse feature spaces. Peng Tan 0002, Hai-Tian Liu, Zhi-Hao Tan, Zhi-Hua Zhou |
NeurIPS | 1 |
| 2024 | Towards enabling learnware to handle heterogeneous feature spaces
Peng Tan 0002, Zhi-Hao Tan, Yuan Jiang 0001, Zhi-Hua Zhou |
Mach. Learn. | 1 |
| 2023 | Handling Learnwares Developed from Heterogeneous Feature Spaces without Auxiliary DataabstractThe learnware paradigm proposed by Zhou [2016] devotes to constructing a market of numerous well-performed models, enabling users to solve problems by reusing existing efforts rather than starting from scratch. A learnware comprises a trained model and the specification which enables the model to be adequately identified according to the user's requirement. Previous studies concentrated on the homogeneous case where models share the same feature space based on Reduced Kernel Mean Embedding (RKME) specification. However, in real-world scenarios, models are typically constructed from different feature spaces. If such a scenario can be handled by the market, all models built for a particular task even with different feature spaces can be identified and reused for a new user task. Generally, this problem would be easier if there were additional auxiliary data connecting different feature spaces, however, obtaining such data in reality is challenging. In this paper, we present a general framework for accommodating heterogeneous learnwares without requiring additional auxiliary data. The key idea is to utilize the submitted RKME specifications to establish the relationship between different feature spaces. Additionally, we give a matrix factorization-based implementation and propose the overall procedure for constructing and exploiting the heterogeneous learnware market. Experiments on real-world tasks validate the efficacy of our method. Peng Tan 0002, Zhi-Hao Tan, Yuan Jiang 0001, Zhi-Hua Zhou |
IJCAI | 1 |
| 2020 | Multi-label optimal margin distribution machine
Zhi-Hao Tan, Peng Tan 0002, Yuan Jiang 0001, Zhi-Hua Zhou |
Mach. Learn. | 2 |