Zhengfeng Zhang

dblp:30/10145 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Representation and self-supervised learning · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
Enhancing Unsupervised Sentence Embeddings via Knowledge-Driven Data Augmentation and Gaussian-Decayed Contrastive Learning · ACL (1) 2025
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding
0.912025
Enhancing Unsupervised Sentence Embeddings via Knowledge-Driven Data Augmentation and Gaussian-Decayed Contrastive Learning · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

knowledge graph · 0.9gaussian-decayed contrastive learning · 0.9data augmentation · 0.9
YearPublicationVenuePosition
2026 Cluster-based prototypical contrastive learning for unsupervised sentence embedding
Peichao Lai, Ruiqing Wang, Jiayong Li, Ruixiong Fang, Zhengfeng Zhang, Qingwei Lyu
Eng. Appl. Artif. Intell.5
2025 Enhancing Unsupervised Sentence Embeddings via Knowledge-Driven Data Augmentation and Gaussian-Decayed Contrastive Learning
abstract
Recently, using large language models (LLMs) for data augmentation has led to considerable improvements in unsupervised sentence embedding models. However, existing methods encounter two primary challenges: limited data diversity and high data noise. Current approaches often neglect fine-grained knowledge, such as entities and quantities, leading to insufficient diversity. Besides, unsupervised data frequently lacks discriminative information, and the generated synthetic samples may introduce noise. In this paper, we propose a pipeline-based data augmentation method via LLMs and introduce the Gaussian-decayed gradient-assisted Contrastive Sentence Embedding (GCSE) model to enhance unsupervised sentence embeddings. To tackle the issue of low data diversity, our pipeline utilizes knowledge graphs (KGs) to extract entities and quantities, enabling LLMs to generate more diverse samples. To address high data noise, the GCSE model uses a Gaussian-decayed function to limit the impact of false hard negative samples, enhancing the model’s discriminative capability. Experimental results show that our approach achieves state-of-the-art performance in semantic textual similarity (STS) tasks, using fewer data samples and smaller LLMs, demonstrating its efficiency and robustness across various models.
Peichao Lai, Zhengfeng Zhang, Wentao Zhang 0001, Fangcheng Fu, Bin Cui 0001
ACL (1)2
2024 NCSE: Neighbor Contrastive Learning for Unsupervised Sentence Embeddings
abstract
Unsupervised sentence embedding methods based on contrastive learning have gained attention for effectively representing sentences in natural language processing. Retrieving additional samples via a nearest-neighbor approach can enhance the model’s ability to learn relevant semantics and distinguish sentences. However, previous related research mainly focused on retrieving neighboring samples within a single batch range or global range, which makes the model possibly unable to capture effective semantic information or incurs excessive time cost. Furthermore, previous methods use retrieved neighbor samples as hard negatives. We argue that nearest neighbor samples contain relevant semantic information, and treating them as hard negatives risks losing valuable semantic knowledge. In this work, we introduce Neighbor Contrastive learning for unsupervised Sentence Embeddings(NCSE), which combines contrastive learning with the nearest-neighbor approach. Specifically, we create a candidate set to store sentence embeddings across multiple batches. Retrieving the candidate set can ensure sufficient samples, making it easier for the model to learn relevant semantics. Using retrieved nearest neighbor samples as positives and applying the self-attention mechanism to aggregate the sample and its neighbors encourages the model to learn relevant semantics from multiple neighbors. Experiments on the semantic text similarity task demonstrate our method’s effectiveness in sentence embedding learning.
Zhengfeng Zhang, Peichao Lai, Ruiqing Wang, Feiyang Ye 0002
IJCNN1
2024 Span-Based Chinese Few-Shot NER with Contrastive and Prompt Learning
Feiyang Ye 0002, Peichao Lai, Sanhe Yang, Zhengfeng Zhang
NLPCC (2)4
2023 Embedding expert demonstrations into clustering buffer for effective deep reinforcement learning
abstract
As one of the most fundamental topics in reinforcement learning (RL), sample efficiency is essential to the deployment of deep RL algorithms. Unlike most existing exploration methods that sample an action from different types of posterior distributions, we focus on the policy sampling process and propose an efficient selective sampling approach to improve sample efficiency by modeling the internal hierarchy of the environment. Specifically, we first employ clustering methods in the policy sampling process to generate an action candidate set. Then we introduce a clustering buffer for modeling the internal hierarchy, which consists of on-policy data, off-policy data, and expert data to evaluate actions from the clusters in the action candidate set in the exploration stage. In this way, our approach is able to take advantage of the supervision information in the expert demonstration data. Experiments on six different continuous locomotion environments demonstrate superior reinforcement learning performance and faster convergence of selective sampling. In particular, on the LGSVL task, our method can reduce the number of convergence steps by 46.7% and the convergence time by 28.5%. Furthermore, our code is open-source for reproducibility. The code is available at https://github.com/Shihwin/SelectiveSampling .
Shihmin Wang, Binqi Zhao, Zhengfeng Zhang, Junping Zhang, Jian Pu
Frontiers Inf. Technol. Electron. Eng.3
2021 Development of a Long-Term Dataset of China Surface Urban Heat Island for Policy Making: Spatio-Temporal Characteristics
abstract
The buffer algorithm's urban heat island intensity is challenging to analyze urban heat islands' socioeconomic drivers, which undoubtedly increases the difficulty of making urban heat island mitigate policy. To address this question, in this study, we combine administrative borders data with satellite remote sensing data to comprehensively depict the 8-day urban heat islands intensity of 286 cities in China from 2001–2018 and analyze Spatio-temporal characteristics. We find that 90.7% of cities have urban heat islands during the daytime, becoming 91.6% at nighttime. There is a significant spatial clustering effect for both nighttime and daytime urban heat islands, and the temporal trend shows that urban heat islands have a greater degree of mitigation at nighttime compared to daytime. This study extends the methodology for characterizing urban heat islands and provides a China surface urban heat island (CSUHI) dataset for future interdisciplinary research.
Lu Niu, Zhong Peng, Ronglin Tang, Zhengfeng Zhang
IGARSS4
2020 EV charging bidding by multi-DQN reinforcement learning in electricity auction market
Yang Zhang 0097, Zhengfeng Zhang, Qingyu Yang 0003, Dou An, Donghe Li, Ce Li 0001
Neurocomputing2
2013 On the Security of an Efficient Attribute-Based Signature
Dengguo Feng, Zhengfeng Zhang, Liwu Zhang
NSS3