VLDB 2026 Research / reviewers in the wild / expert
Wen Pu
dblp:76/6838
· DBLP profile ↗
10ranked-venue papers
2as first author
2since 2021 · last 2026
0009-0007-9386-2628ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 90% Data mining · 10% | |
| Artificial intelligence
2 papers |
Probabilistic and Bayesian machine learning · 80% Representation and self-supervised learning · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › document retrieval › domain-specific retrieval
job search |
1.0 | 1 | 2026 | Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn · SIGIR 2026 |
Information retrieval › retrieval models
semantic representation |
1.0 | 1 | 2026 | Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn · SIGIR 2026 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
partition function approximation |
0.2 | 1 | 2015 | A Deterministic Partition Function Approximation for Exponential Random Graph Models · IJCAI 2015 |
Computational social science and digital humanities
social network analysis |
0.1 | 1 | 2012 | Identifying Bullies with a Computer Game · AAAI 2012 |
Data mining
probabilistic graphical models |
0.1 | 1 | 2012 | Identifying Bullies with a Computer Game · AAAI 2012 |
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
document representation |
0.1 | 1 | 2007 | Local Word Bag Model for Text Categorization · ICDM 2007 |
Data mining › text mining
text classification |
0.1 | 1 | 2007 | Local Word Bag Model for Text Categorization · ICDM 2007 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › relational model
statistical network models |
0.1 | 1 | 2015 | A Deterministic Partition Function Approximation for Exponential Random Graph Models · IJCAI 2015 |
Methods — techniques the papers use, named apart from their topics
representation learning · 1.0probabilistic model · 0.3classifier comparison · 0.3support vector machine · 0.1pyramid match kernel · 0.1hierarchical clustering · 0.1bag-of-words · 0.1TF-IDF · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn
Baofen Zheng, Jianqiang Shen, Benjamin Le, Wen Pu, Neha Saraf, Alice Leung, Qianqi Shen, Liangjie Hong, Jingwei Wu |
SIGIR | 6 |
| 2024 | Learning Links for Adaptable and Explainable RetrievalabstractWeb-scale search systems typically tackle the scalability challenge with a two-step paradigm: retrieval and ranking. The retrieval step, also known as candidate selection, often involves extracting entities, creating an inverted index, and performing term matching for retrieval. Such traditional methods require manual and time-consuming development of retrieval models. In this paper, we propose a framework for constructing a graph that integrates human knowledge with user activity data analysis. The learned links are utilized for retrieval purposes. The model is easy to explain, debug, and tune. The system implementation is straightforward and can directly leverage existing inverted index systems. We applied this retrieval framework to enhance the job search and recommendation systems on a large professional networking portal, resulting in significant performance improvements. Jianqiang Shen, Yuchin Juan, Ping Liu 0002, Wen Pu, Qianqi Shen, Liangjie Hong |
CIKM | 4 |
| 2020 | Impression Pacing for Jobs Marketplace at LinkedInabstractThe goal of Jobs Marketplace at LinkedIn is to match members to promoted job postings such that both job posters' ROI is optimized (amount of money spent per job clicks and applications) and the members are presented with relevant jobs that they are interested in and qualified for. This is achieved via a first-price auction mechanism where each job provides a bid for the member that comes to the job recommendations page. This bid depends on the match of the member to the job, as well as the daily budget that remains for the job, and its capability to spend it via clicks (e.g. some jobs might have more demand and have it easier to spend their budgets via clicks than others). In such a scheme, budget pacing, i.e. the capability of a job to spend its daily budget evenly, or according to a preset plan, is extremely important towards efficient utilization of its budget via reaching a higher number of candidates, and obey a variety of spending plans optimizing for different events such as clicks and applications. Sahin Cem Geyik, Luthfur Chowdhury, Florian Raudies, Wen Pu, Jianqiang Shen |
CIKM | 4 |
| 2015 | A Deterministic Partition Function Approximation for Exponential Random Graph Models
Wen Pu, Jaesik Choi, Yunseong Hwang, Eyal Amir |
IJCAI | 1 |
| 2012 | Identifying Bullies with a Computer GameabstractCurrent computer involvement in adolescent social networks (youth between the ages of 11 and 17) provides new opportunities to study group dynamics, interactions amongst peers, and individual preferences. Nevertheless, most of the research in this area focuses on efficiently retrieving information that is explicit in large social networks (e.g., properties of the graph structure), but not on how to use the dynamics of the virtual social network to discover latent characteristics of the real-world social network. In this paper, we present the analysis of a game designed to take advantage of the familiarity of adolescents with online social networks, and describe how the data generated by the game can be used to identify bullies in 5th grade classrooms. We present a probabilistic model of the game and using the in-game interactions of the players (i.e., content of chat messages) infer their social role within their classroom (either a bully or non-bully). The evaluation of our model is done by using previously collected data from psychological surveys on the same 5th grade population and by comparing the performance of the new model with off-the-shelf classifiers. Juan Fernando Mancilla-Caceres, Wen Pu, Eyal Amir, Dorothy Espelage |
AAAI | 2 |
| 2008 | An online approach based on locally weighted learning for short-term traffic flow predictionabstractTraffic flow prediction is a basic function of Intelligent Transportation System. Due to the complexity of traffic phenomenon, most existing methods build complex models such as neural networks for traffic flow prediction. As a model may lose effect with time lapse, it is important to update the model on line. However, the high computational cost of maintaining a complex model puts great challenge for model updating. The high computation cost lies in two aspects: computation of complex model coefficients and huge amount training data for it. In this paper, we propose to use a nonparametric approach based on locally weighted learning to predict traffic flow. Our approach incrementally incorporates new data to the model and is computationally efficient, which makes it suitable for online model updating and predicting. In addition, we adopt wavelet analysis to extract the periodic characteristic of the traffic data, which is then used for the input of the prediction model instead of the raw traffic flow data. The primary experiments on real data demonstrate the effectiveness and efficiency of our approach. Meng Shuai, Kunqing Xie, Wen Pu, Guojie Song, Xiujun Ma |
GIS | 3 |
| 2007 | Local Word Bag Model for Text CategorizationabstractMany text processing applications adopted the bag of words (BOW) model representation of documents, in which each document is represented as a vector of weighted terms or n-grams, and then the cosine distance between two vectors is used as the similarity measurement. Although the great success in information retrieval and text categorization, the conventional BOW model ignores the detailed local text information, i.e. the co-occurrence pattern of words at sentence or paragraph level. In this paper, we propose a novel approach to represent a document as a set of local tf-idf vectors, or what we called local word bags (LWB). By encapsulating local information distributed around a document into multiple LWBs, we can measure the similarity of two documents via the partial match of their corresponding local bags. To perform the matching efficiently, we introduce the local word bag kernel (LWB kernel), a variant of VG-Pyramid match kernel. The new kernel enables the discriminative machine learning methods like SVM to compute the partial matching between two sets of LWBs in linear time after an one time hierarchical clustering procedure over all local bags at the initialization stage. Experiments on real world datasets demonstrate the effectiveness of our new approach. Wen Pu, Ning Liu 0001, Shuicheng Yan, Jun Yan 0001, Kunqing Xie, Zheng Chen 0001 |
ICDM | 1 |
| 2005 | A spatio-temporal aquarium for visual exploration on geographic phenomena
Xiujun Ma, Kunqing Xie, Cuo Cai, Wen Pu |
IGARSS | 6 |
| 2005 | Detecting spatio-temporal outliers in climate dataset: a method studyabstractOutlier detecting is one of the most important data analysis technologies in data mining, which can be used to discover anomalous phenomena in huge dataset. Many literatures on spatial outlier detecting and time series outlier detecting have appeared, while the area of spatio-temporal outliers considering both spatial and temporal dimensions has still rarely been touched. Defining outliers in traditional dataset is more explicit because the data structure we need to focus on is very straightforward (e.g., a spatial point or a transaction record). However, it is much more difficult to give outlier a definite characterization in spatio-temporal lattice data, since there are so many data structures we can pay attention to. With the aim of detecting useful and meaningful outliers in climate dataset, we introduce a formalized way to define outliers in spatio-temporal lattice data, in which the importance of clarifying basic data structure (we call it basic element in our paper) is stressed. As a case study, we define two kinds of spatio-temporal outliers based on a global climate dataset, according to the three aspects we propose in defining an outlier. The introduction of basic element and the formulation of outlier definition process make it easier and clearer to define meaningful outliers. Thus outlier detecting in spatio-temporal lattice data will provide us with really interesting and useful knowledge. Kunqing Xie, Xiujun Ma, Xingxing Jin, Wen Pu |
IGARSS | 5 |
| 2005 | Adaptive sampling for selectivity estimation in spatial databaseabstractSpatial sampling is a significant part of query processing in spatial database. In this paper, an adaptive data-driven sampling for selectivity estimation is proposed in spatial database. This technique presents an efficient sampling for spatial data, especially for two-dimensional line and polygon data, and it make the sample size fit the limit of time and memory, or a user-defined parameter. The data-driven sampling technique is compared with various techniques on different type of datasets in our experimental study, and it out outperforms the other techniques over a broad range of query workloads and datasets. Xiujun Ma, Kunqing Xie, Huibin Zhang, Wen Pu |
IGARSS | 6 |