Peng Tan 0002

dblp:69/1700-2 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-3749-9266ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 72% Transfer learning and domain adaptation · 14% Representation and self-supervised learning · 14%
Databases, data mining, and information retrieval
2 papers
Data mining · 57% Machine learning and data management · 43%

Topics — the 4 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model reuse
2.242026
Tabular Learnwares Can Be Repurposed for Seemingly Irrelevant New Tasks · AAAI 2026
Beimingwu: A Learnware Dock System · KDD 2024
Handling Learnwares from Heterogeneous Feature Spaces with Explicit Label Exploitation · NeurIPS 2024
Machine learning › Efficient and distributed learning › model reuse
learnware
1.722026
Tabular Learnwares Can Be Repurposed for Seemingly Irrelevant New Tasks · AAAI 2026
Handling Learnwares Developed from Heterogeneous Feature Spaces without Auxiliary Data · IJCAI 2023
Data mining › predictive modeling › classification
tabular data classification
1.012026
Tabular Learnwares Can Be Repurposed for Seemingly Irrelevant New Tasks · AAAI 2026
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
unified embedding
0.812024
Handling Learnwares from Heterogeneous Feature Spaces with Explicit Label Exploitation · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

reduced kernel mean embedding · 2.2probability-based model reuse · 2.0model specification · 1.5pseudo-label exploitation · 0.8conditional distribution comparison · 0.8matrix factorization · 0.7
YearPublicationVenuePosition
2026 Tabular Learnwares Can Be Repurposed for Seemingly Irrelevant New Tasks
abstract
The learnware paradigm aims to help users solve new tasks by reusing existing models rather than starting from scratch. A learnware consists of a model and the specification describing its capabilities. Numerous learnwares are accommodated by the learnware dock system. When users solve tasks with the system, learnwares that fully match the user task are often scarce or unavailable. This paper focuses on tabular classification tasks and explores reusing learnwares for new user tasks with significantly different feature and label spaces, leveraging the potential of numerous existing specialized tabular models developed for various tasks. Under the learnware paradigm, we find that tabular learnwares that seem semantically irrelevant can sometimes be beneficial for new user tasks. The proposed method relies solely on model-predicted probabilities and does not require gradient information, making it applicable to a wide range of tabular models. Experiments suggest that tabular learnwares can be reused beyond their original purpose across heterogeneous tasks.
Peng Tan 0002, Zhi-Hao Tan, Zhi-Hua Zhou
AAAI1
2024 Beimingwu: A Learnware Dock System
abstract
The learnware paradigm proposed by Zhou [40] aims to enable users to leverage numerous existing high-performing models instead of building machine learning models from scratch.This paradigm envisions that: Any developer worldwide can submit their well-trained models spontaneously into a learnware dock system (formerly known as learnware market).The system uniformly generates a specification for each model to form a learnware and accommodates it.As the key component, a specification should represent the capabilities of the model while preserving developer's original data.Based on the specifications, the learnware dock system can identify and assemble existing learnwares for users to solve new machine learning tasks.Recently, based on reduced kernel mean embedding (RKME) specification, a series of studies have shown the effectiveness of the learnware paradigm theoretically and empirically.However, the realization of a learnware dock system is still missing and remains a big challenge.This paper proposes Beimingwu, the first open-source learnware dock system, providing foundational support for future research.The system provides implementations and extensibility for the entire process of learnware paradigm, including the submitting, usability testing, organization, identification, deployment, and reuse of learnwares.Utilizing Beimingwu, the model development for new user tasks can be significantly streamlined, thanks to integrated architecture and engine design, specifying unified learnware structure and scalable APIs, and the integration of various algorithms for learnware identification and reuse.Notably, this is possible even for users with limited data and minimal expertise in machine learning, without compromising the raw data's security.The system facilitates the future research implementations in learnware-related algorithms and systems, and lays the ground for hosting a vast array of learnwares and establishing a learnware ecosystem.The system is fully open-source and we expect the research community
Zhi-Hao Tan, Jian-Dong Liu, Xiaodong Bi, Peng Tan 0002, Qin-Cheng Zheng, Hai-Tian Liu, Xiao-Chuan Zou, Yang Yu 0001, Zhi-Hua Zhou
KDD4
2024 Handling Learnwares from Heterogeneous Feature Spaces with Explicit Label Exploitation
abstract
The learnware paradigm aims to help users leverage numerous existing high-performing models instead of starting from scratch, where a learnware consists of a well-trained model and the specification describing its capability. Numerous learnwares are accommodated by a learnware dock system. When users solve tasks with the system, models that fully match the task feature space are often rare or even unavailable. However, models with heterogeneous feature space can still be helpful. This paper finds that label information, particularly model outputs, is helpful yet previously less exploited in the accommodation of heterogeneous learnwares. We extend the specification to better leverage model pseudo-labels and subsequently enrich the unified embedding space for better specification evolvement. With label information, the learnware identification can also be improved by additionally comparing conditional distributions. Experiments demonstrate that, even without a model explicitly tailored to user tasks, the system can effectively handle tasks by leveraging models from diverse feature spaces.
Peng Tan 0002, Hai-Tian Liu, Zhi-Hao Tan, Zhi-Hua Zhou
NeurIPS1
2024 Towards enabling learnware to handle heterogeneous feature spaces
Peng Tan 0002, Zhi-Hao Tan, Yuan Jiang 0001, Zhi-Hua Zhou
Mach. Learn.1
2023 Handling Learnwares Developed from Heterogeneous Feature Spaces without Auxiliary Data
abstract
The learnware paradigm proposed by Zhou [2016] devotes to constructing a market of numerous well-performed models, enabling users to solve problems by reusing existing efforts rather than starting from scratch. A learnware comprises a trained model and the specification which enables the model to be adequately identified according to the user's requirement. Previous studies concentrated on the homogeneous case where models share the same feature space based on Reduced Kernel Mean Embedding (RKME) specification. However, in real-world scenarios, models are typically constructed from different feature spaces. If such a scenario can be handled by the market, all models built for a particular task even with different feature spaces can be identified and reused for a new user task. Generally, this problem would be easier if there were additional auxiliary data connecting different feature spaces, however, obtaining such data in reality is challenging. In this paper, we present a general framework for accommodating heterogeneous learnwares without requiring additional auxiliary data. The key idea is to utilize the submitted RKME specifications to establish the relationship between different feature spaces. Additionally, we give a matrix factorization-based implementation and propose the overall procedure for constructing and exploiting the heterogeneous learnware market. Experiments on real-world tasks validate the efficacy of our method.
Peng Tan 0002, Zhi-Hao Tan, Yuan Jiang 0001, Zhi-Hua Zhou
IJCAI1
2020 Multi-label optimal margin distribution machine
Zhi-Hao Tan, Peng Tan 0002, Yuan Jiang 0001, Zhi-Hua Zhou
Mach. Learn.2