Jinfeng Peng

dblp:226/6795 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0005-6443-6692ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Device Type Identification with Deep Metric Learning
Xinyu Yin, Fan Shi 0003, Chengxi Xu, Jinfeng Peng, Jiatang Zhao
ICIC (4)4
2026 Beyond Up and Down: Analyzing Temporal Phenomena in IPv6 Probing Responses
Yudong Lian, Jinfeng Peng, Chengxi Xu, Fan Shi 0003, Min Zhang 0054, Jiatang Zhao, Quming Peng
IWQoS2
2026 TNet: Efficient IPv6 active network discovery
Jiatang Zhao, Fan Shi 0003, Chengxi Xu, Jinfeng Peng, Mingyi Ge, Min Zhang 0054
Comput. Networks4
2025 An Effective and Secure Federated Multi-View Clustering Method with Information-Theoretic Perspective
abstract
Recently, federated multi-view clustering (FedMVC) has gained attention for its ability to mine complementary clustering structures from multiple clients without exposing private data. Existing methods mainly focus on addressing the feature heterogeneity problem brought by views on different clients and mitigating it using shared client information. Although these methods have achieved performance improvements, the information they choose to share, such as model parameters or intermediate outputs, inevitably raises privacy concerns. In this paper, we propose an Effective and Secure Federated Multi-view Clustering method, ESFMC, to alleviate the dilemma between privacy protection and performance improvement. This method leverages the information-theoretic perspective to split the features extracted locally by clients, retaining sensitive information locally and only sharing features that are highly relevant to the task. This can be viewed as a form of privacy-preserving information sharing, reducing privacy risks for clients while ensuring that the server can mine high-quality global clustering structures. Theoretical analysis and extensive experiments demonstrate that the proposed method more effectively mitigates the trade-off between privacy protection and performance improvement compared to state-of-the-art methods.
Xinyue Chen 0004, Jinfeng Peng, Xiaorong Pu, Yang Yang 0002, Yazhou Ren 0001
ICML2
2025 GARF+: self-supervised and interpretable data cleaning with sequence generative adversarial networks
Jinfeng Peng, Hanghai Cui, Derong Shen, Nan Tang 0001, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001
VLDB J.1
2024 GARF: A Self-supervised Data Cleaning System with SeqGAN
abstract
High-quality data is essential for data science and machine learning applications, but unfortunately, real-world data often contains significant amounts of errors, such as typos, missing values, and data inconsistencies. Despite all the efforts in cleaning data using either logical or learning-based methods, in practice, data cleaning still requires high human cost, for either manually providing data repairing rules or preparing labeled datasets for training machine learning models. In this paper, we introduce GARF, a novel data cleaning system based on sequence generative adversarial networks (SeqGAN). One key information GARF tries to learn is data repair rules. To automatically extracts data repair rules from dirty data, GARF employs a SeqGAN to capture the dependency relationships, and converts the information learned by machine to interpretable data repair rules for humans. Additionally, considering that both generated rules and data may not be fully trusted, GARF provides a co-cleaning process to iteratively update inaccurate rules and repair dirty data until there is no tuple violating rules. We have implemented and deployed GARF as an open-sourced system, and demonstrated its usability on data cleaning in real-world scenarios.
Jinfeng Peng, Hanghai Cui, Derong Shen, Yue Kou, Tiezheng Nie, Tianlong Guo
CIKM1
2024 RLclean: An unsupervised integrated data cleaning framework based on deep reinforcement learning
Jinfeng Peng, Derong Shen, Tiezheng Nie, Yue Kou
Inf. Sci.1
2022 Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks
abstract
We study the problem of self-supervised and interpretable data cleaning, which automatically extracts interpretable data repair rules from dirty data. In this paper, we propose a novel framework, namely Garf, based on sequence generative adversarial networks (SeqGAN). One key information Garf tries to capture is data repair rules (for example, if the city is "Dothan", then the county should be "Houston"). Garf employs a SeqGAN consisting of a generator G and a discriminator D that trains G to learn the dependency relationships ( e.g. , given a city value "Dothan" as input, the county can be determined as "Houston"). After training, the generator G can be used to generate data repair rules, but may contain both trusted and untrusted rules, especially when learning from dirty data. To mitigate this problem, Garf further updates the learned relationships with another discriminator D' to iteratively improve the quality of both rules and data. Garf takes advantages of both logical and learning-based methods, which allow cleaning dirty data with high interpretability and have no requirements for prior knowledge and training data. Extensive experiments on real-world and synthetic datasets demonstrate the effectiveness of Garf. Garf achieves new state-of-the-art data cleaning result with high accuracy, through learning from dirty datasets without human supervision.
Jinfeng Peng, Derong Shen, Nan Tang 0001, Tieying Liu, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001
Proc. VLDB Endow.1
2018 Fraud Detection of Medical Insurance Employing Outlier Analysis
abstract
Fraud detection is an important issue in the area of data science, and it has a lot of practical applications in related fields, such as business, health, and environment. Most traditional methods detect fraud based on rulemaking. Unfortunately, it is not always useful in the medical field since the boundary of fraud detection is vague. As a result, outlier detection is a promising method. This paper develops an outlier detection method of analyzing the correlation of patients to detect fraud. We construct a heterogeneous information network which bridges the medicines used and diseases of patients. In light of the network, we calculate the correlation score of different patients and design a discriminant rule. Through the discriminating rule, fraudulent patients represented by the abnormal nodes can be found. Our experiments use real medical insurance data sets and the results confirm that our method is accurate and effective.
Jinfeng Peng, Qingzhong Li, Hui Li 0048, Lei Liu 0003, Zhongmin Yan, Shidong Zhang
CSCWD1