Wenjun Peng 0001

dblp:166/6129-1 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-2392-4946ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Synthetic Data Powers Product Retrieval for Long-tail Knowledge-Intensive Queries in E-commerce Search
abstract
Product retrieval is the backbone of e-commerce search: for each user query, it identifies a high-recall candidate set from billions of items, laying the foundation for high-quality ranking and user experience. Despite extensive optimization for mainstream queries, existing systems still struggle with long-tail queries, especially knowledge-intensive ones. These queries exhibit diverse linguistic patterns, often lack explicit purchase intent, and require domain-specific knowledge reasoning for accurate interpretation. They also suffer from a shortage of reliable behavioral logs, which makes such queries a persistent challenge for retrieval optimization.
Gui Ling, Weiyuan Li, Wenjun Peng 0001, Xingxian Liu, Dongshuai Li, Fuyu Lv, Dan Ou, Haihong Tang
SIGIR4
2025 SoMORE: Social Context-Aware MLLM for Video Character Search
Xin Kou, Wenjun Peng 0001, Tong Xu 0001
KSEM (2)2
2024 Large language models for generative information extraction: a survey
abstract
Abstract Information Extraction (IE) aims to extract structural knowledge from plain natural language texts. Recently, generative Large Language Models (LLMs) have demonstrated remarkable capabilities in text understanding and generation. As a result, numerous works have been proposed to integrate LLMs for IE tasks based on a generative paradigm. To conduct a comprehensive systematic review and exploration of LLM efforts for IE tasks, in this study, we survey the most recent advancements in this field. We first present an extensive overview by categorizing these works in terms of various IE subtasks and techniques, and then we empirically analyze the most advanced methods and discover the emerging trend of IE tasks with LLMs. Based on a thorough review conducted, we identify several insights in technique and promising research directions that deserve further exploration in future studies. We maintain a public repository and consistently update related works and resources on GitHub (LLM4IE repository).
Derong Xu, Wei Chen 0156, Wenjun Peng 0001, Chao Zhang 0096, Tong Xu 0001, Xiangyu Zhao 0001, Xian Wu 0001, Yefeng Zheng 0001, Yang Wang 0001, Enhong Chen
Frontiers Comput. Sci.3
2023 Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark
abstract
Wenjun Peng, Jingwei Yi, Fangzhao Wu, Shangxi Wu, Bin Bin Zhu, Lingjuan Lyu, Binxing Jiao, Tong Xu, Guangzhong Sun, Xing Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Wenjun Peng 0001, Jingwei Yi, Fangzhao Wu, Shangxi Wu, Bin B. Zhu, Lingjuan Lyu, Binxing Jiao, Tong Xu 0001, Guangzhong Sun, Xing Xie 0001
ACL (1)1
2023 Social Context-aware GCN for Video Character Search via Scene-prior Enhancement
abstract
With the increasing demand for intelligent services of online video platforms, video character search task has attracted wide attention to support downstream applications like fine-grained retrieval and summarization. However, traditional solutions only focus on visual or coarse-grained social information and thus cannot perform well when facing complex scenes, such as changing camera view or character posture. Along this line, we leverage social information and scene context as prior knowledge to solve the problem of character search in complex scenes. Specifically, we propose a scene-prior-enhanced framework, named SoCoSearch. We first integrate multimodal clues for scene context to estimate the prior probability of social relationships, and then capture characters’ co-occurrence to generate an enhanced social context graph. Afterwards, we design a social context-aware GCN framework to achieve feature passing between characters to obtain robust representation for the character search task. Extensive experiments have validated the effectiveness of SoCoSearch in various metrics.
Wenjun Peng 0001, Weidong He, Derong Xu, Tong Xu 0001, Chen Zhu 0003, Enhong Chen
ICME1
2023 Are GPT Embeddings Useful for Ads and Recommendation?
Wenjun Peng 0001, Derong Xu, Tong Xu 0001, Jianjin Zhang, Enhong Chen
KSEM (4)1
2022 IHGNN: Interactive Hypergraph Neural Network for Personalized Product Search
abstract
A good personalized product search (PPS) system should not only focus on retrieving relevant products, but also consider user personalized preference. Recent work on PPS mainly adopts the representation learning paradigm, e.g., learning representations for each entity (including user, product and query) from historical user behaviors (aka. user-product-query interactions). However, we argue that existing methods do not sufficiently exploit the crucial collaborative signal, which is latent in historical interactions to reveal the affinity between the entities. Collaborative signal is quite helpful for generating high-quality representation, exploiting which would benefit the representation learning of one node from its connected nodes.
Dian Cheng, Jiawei Chen 0007, Wenjun Peng 0001, Wenqin Ye, Fuyu Lv, Xiaoyi Zeng, Xiangnan He 0001
WWW3