VLDB 2026 Research / reviewers in the wild / expert
Jingyuan Yang 0001
dblp:174/2233-1
· DBLP profile ↗
13ranked-venue papers
3as first author
3since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Data mining · 47% Recommender systems · 24% Web and social media mining · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Smart cities and intelligent transportation · 58% Computational social science and digital humanities · 42% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 87% Representation and self-supervised learning · 13% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
spatiotemporal data mining |
0.6 | 1 | 2022 | Exploiting Interpretable Patterns for Flow Prediction in Dockless Bike Sharing Systems · IEEE Trans. Knowl. Data Eng. 2022 |
Data mining › clustering › high-dimensional clustering
subspace clustering |
0.6 | 1 | 2022 | Exploiting Interpretable Patterns for Flow Prediction in Dockless Bike Sharing Systems · IEEE Trans. Knowl. Data Eng. 2022 |
Recommender systems › domain-specific recommendation
marketing recommendation |
0.5 | 2 | 2018 | A Unified View of Social and Temporal Modeling for B2B Marketing Campaign Recommendation · IEEE Trans. Knowl. Data Eng. 2018 Exploiting Temporal and Social Factors for B2B Marketing Campaign Recommendations · ICDM 2015 |
Smart cities and intelligent transportation
bike sharing systems |
0.4 | 2 | 2022 | Station Site Optimization in Bike Sharing Systems · ICDM 2015 Exploiting Interpretable Patterns for Flow Prediction in Dockless Bike Sharing Systems · IEEE Trans. Knowl. Data Eng. 2022 |
Web and social media mining › social network analysis
edge weight prediction |
0.4 | 1 | 2019 | Dynamic Talent Flow Analysis with Deep Sequence Prediction Modeling · IEEE Trans. Knowl. Data Eng. 2019 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.3 | 1 | 2018 | CADEN: A Context-Aware Deep Embedding Network for Financial Opinions Mining · ICDM 2018 |
Natural language and speech › Information extraction and text analysis
text classification |
0.3 | 1 | 2018 | CADEN: A Context-Aware Deep Embedding Network for Financial Opinions Mining · ICDM 2018 |
Data mining › predictive modeling
classification |
0.3 | 1 | 2018 | Incomplete Label Uncertainty Estimation for Petition Victory Prediction with Dynamic Features · ICDM 2018 |
Recommender systems
graph-based recommendation |
0.3 | 1 | 2018 | A Unified View of Social and Temporal Modeling for B2B Marketing Campaign Recommendation · IEEE Trans. Knowl. Data Eng. 2018 |
Data mining › structured data mining › graph mining
community detection |
0.2 | 1 | 2016 | Talent Circle Detection in Job Transition Networks · KDD 2016 |
Data mining › structured data mining
graph mining |
0.2 | 1 | 2016 | Talent Circle Detection in Job Transition Networks · KDD 2016 |
Knowledge graphs
link prediction |
0.2 | 1 | 2015 | Exploiting Temporal and Social Factors for B2B Marketing Campaign Recommendations · ICDM 2015 |
Graph data management
temporal graph |
0.2 | 1 | 2015 | Exploiting Temporal and Social Factors for B2B Marketing Campaign Recommendations · ICDM 2015 |
Data mining › predictive modeling › forecasting
multi-step forecasting |
0.1 | 1 | 2019 | Dynamic Talent Flow Analysis with Deep Sequence Prediction Modeling · IEEE Trans. Knowl. Data Eng. 2019 |
Data mining › time series analysis
time series forecasting |
0.1 | 1 | 2019 | Dynamic Talent Flow Analysis with Deep Sequence Prediction Modeling · IEEE Trans. Knowl. Data Eng. 2019 |
Machine learning › Representation and self-supervised learning
text embedding |
0.1 | 1 | 2018 | CADEN: A Context-Aware Deep Embedding Network for Financial Opinions Mining · ICDM 2018 |
Web and social media mining
social network analysis |
0.1 | 1 | 2018 | A Unified View of Social and Temporal Modeling for B2B Marketing Campaign Recommendation · IEEE Trans. Knowl. Data Eng. 2018 |
Web and social media mining › social network analysis
professional social network analysis |
0.1 | 1 | 2016 | Talent Circle Detection in Job Transition Networks · KDD 2016 |
Smart cities and intelligent transportation
demand prediction |
0.1 | 1 | 2015 | Station Site Optimization in Bike Sharing Systems · ICDM 2015 |
Methods — techniques the papers use, named apart from their topics
graph regularized sparse representation · 1.1graph laplacian · 1.1uncertainty estimation · 0.7expectation-maximization · 0.7recurrent neural network · 0.4deep sequence prediction · 0.4user embedding · 0.3probabilistic graphical model · 0.3low-rank graph reconstruction · 0.3deep embedding network · 0.3attention mechanism · 0.3normalized discounted cumulative gain optimization · 0.2network embedding · 0.2genetic algorithm · 0.2artificial neural network · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Iterative Prediction-and-Optimization for E-Logistics Distribution Network DesignabstractThe emergence of online retailers has brought new opportunities to the design of their distribution networks. Notably, for online retailers that do not operate offline stores, their target customers are more sensitive to the quality of logistic services, such as delivery speed and reliability. This paper is motivated by a leading online retailer for cosmetic products on Taobao.com that aimed to improve its logistics efficiency by redesigning its centralized distribution network into a multilevel one. The multilevel distribution network consists of a layer of primary facilities to hold stocks from suppliers and transshipment and a layer of secondary facilities to provide last-mile delivery. There are two major challenges of designing such a facility network. First, online customers can respond significantly to the change of logistics efficiency with the redesigned network, thereby rendering the network optimized under the original demand distribution suboptimal. Second, because online retailers have relatively small sales volumes and are very flexible in choosing facility locations, the facility candidate set can be large, causing the facility location optimization challenging to solve. To this end, we propose an iterative prediction-and-optimization strategy for distribution network design. Specifically, we first develop an artificial neural network (ANN) to predict customer demands, factoring in the logistic service quality given the network and the city-level purchasing power based on demographic statistics. Then, a mixed integer linear programming (MILP) model is formulated to choose facility locations with minimum transportation, facility setup, and package processing costs. We further develop an efficient two-stage heuristic for computing high-quality solutions to the MILP model, featuring an agglomerative hierarchical clustering algorithm and an expectation and maximization algorithm. Subsequently, the ANN demand predictor and two-stage heuristic are integrated for iterative network design. Finally, using a real-world data set, we validate the demand prediction accuracy and demonstrate the mutual interdependence between the demand and network design. Summary of Contribution: We propose an iterative prediction-and-optimization algorithm for multilevel distribution network design for e-logistics and evaluate its operational value for online retailers. We address the issue of the interplay between distribution network design and the demand distribution using an iterative framework. Further, combining the idea in operational research and data mining, our paper provides an end-to-end solution that can provide accurate predictions of online sales distribution, subsequently solving large-scale optimization problems for distribution network design problems. Weiwei Chen 0003, Jingyuan Yang 0001, Hui Xiong 0001 |
INFORMS J. Comput. | 3 |
| 2022 | Cross-Lingual Knowledge Transferring by Structural Correspondence and Space TransferabstractThe cross-lingual sentiment analysis (CLSA) aims to leverage label-rich resources in the source language to improve the models of a resource-scarce domain in the target language, where monolingual approaches based on machine learning usually suffer from the unavailability of sentiment knowledge. Recently, the transfer learning paradigm that can transfer sentiment knowledge from resource-rich languages, for example, English, to resource-poor languages, for example, Chinese, has gained particular interest. Along this line, in this article, we propose semisupervised learning with SCL and space transfer (ssSCL-ST), a semisupervised transfer learning approach that makes use of structural correspondence learning as well as space transfer for cross-lingual sentiment analysis. The key idea behind ssSCL-ST, at a high level, is to explore the intrinsic sentiment knowledge in the target-lingual domain and to reduce the loss of valuable knowledge due to the knowledge transfer via semisupervised learning. ssSCL-ST also features in pivot set extension and space transfer, which helps to enhance the efficiency of knowledge transfer and improve the classification accuracy in the target language domain. Extensive experimental results demonstrate the superiority of ssSCL-ST to the state-of-the-art approaches without using any parallel corpora. Deqing Wang 0001, Junjie Wu 0002, Jingyuan Yang 0001, Baoyu Jing, Xiaonan He, Hui Zhang 0028 |
IEEE Trans. Cybern. | 3 |
| 2022 | Exploiting Interpretable Patterns for Flow Prediction in Dockless Bike Sharing SystemsabstractUnlike the traditional dock-based systems, dockless bike-sharing systems are more convenient for users in terms of flexibility. However, the flexibility of these dockless systems comes at the cost of management and operation complexity. Indeed, the imbalanced and dynamic use of bikes leads to mandatory rebalancing operations, which impose a critical need for effective bike traffic flow prediction. While efforts have been made in developing traffic flow prediction models, existing approaches lack interpretability, and thus have limited value in practical deployment. To this end, we propose an Interpretable Bike Flow Prediction (IBFP) framework, which can provide effective bike flow prediction with interpretable traffic patterns. Specifically, by dividing the urban area into regions according to flow density, we first model the spatio-temporal bike flows between regions with graph regularized sparse representation, where graph Laplacian is used as a smooth operator to preserve the commonalities of the periodic data structure. Then, we extract traffic patterns from bike flows using subspace clustering with sparse representation to construct interpretable base matrices. Moreover, the bike flows can be predicted with the interpretable base matrices and learned parameters. Finally, experimental results on real-world data show the advantages of the IBFP method for flow prediction in dockless bike sharing systems. In addition, the interpretability of our flow pattern exploitation is further illustrated through a case study where IBFP provides valuable insights into bike flow analysis. Jingjing Gu, Qiang Zhou 0007, Jingyuan Yang 0001, Yanchi Liu, Fuzhen Zhuang, Yanchao Zhao, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2019 | Adaptively Transfer Category-Classifier for Handwritten Chinese Character Recognition
Yongchun Zhu, Fuzhen Zhuang, Jingyuan Yang 0001, Qing He 0003 |
PAKDD (1) | 3 |
| 2019 | NeuO: Exploiting the sentimental bias between ratings and reviews with neural networks
Yuanbo Xu, Yongjian Yang 0001, En Wang, Fuzhen Zhuang, Jingyuan Yang 0001, Hui Xiong 0001 |
Neural Networks | 6 |
| 2019 | Dynamic Talent Flow Analysis with Deep Sequence Prediction ModelingabstractTalent flow analysis is a process for analyzing and modeling the flows of employees into and out of targeted organizations, regions, or industries. A clear understanding of talent flows is critical for many applications, such as human resource planning, brain drain monitoring, and future workforce forecasting. However, existing studies on talent flow analysis are either qualitative or limited by coarse level quantitative modeling. To this end, in this paper, we provide a fine-grained data-driven approach to model the dynamics and evolving nature of talent flows by leveraging the rich information available in job transition networks. Specifically, we first investigate how to enrich the sparse talent flow data by exploiting the correlations between the stock price movement and the talent flows of public companies. Then, we formalize the talent flow modeling problem as to predict the increments of the edge weights in the dynamic job transition network. In this way, the problem is transformed into a multi-step time series forecasting problem. A deep sequence prediction model is developed based on the recurrent neural network model, which consumes multiple input sources derived from dynamic job transition networks. Finally, experimental results on real-world data show that the proposed model outperforms other benchmark models in terms of prediction accuracy. The results also indicate that the proposed model can provide reasonable performance even if the historical talent flow data are not completely available. Huang Xu 0001, Zhiwen Yu 0001, Jingyuan Yang 0001, Hui Xiong 0001, Hengshu Zhu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2018 | Incomplete Label Uncertainty Estimation for Petition Victory Prediction with Dynamic FeaturesabstractIt is important for decision-makers to effectively and proactively differentiate the significance of various public concerns, and address them with optimal strategy under the limited resources. Online Petition Platforms (OPPs) are replacing traditional social and market surveys for the advantages of low financial cost and high-fidelity social indicators. Despite benefits from OPPs, the raw information from millions of petition signers can easily overwhelm decision makers. In addition, spatio-temporal and semantic dissemination patterns increase the complexity of such OPP data. These two aspects show the necessity of a framework that learns from all available data, which is encoded by dynamic representation of features, to predict whether a petition will successfully lead to a change by decision makers. To build such framework, we need to overcome several challenges including: 1) missing values in dynamic features; 2) strong uncertainty in petition prediction; 3) unknown labels for ongoing petitions and 4) Scalability regarding increasing features and petitions. To address these difficulties simultaneously, we propose a novel chain-structure Multi-task Learning framework with Uncertainty Estimation (MLUE) to predict potentially victorious petitions, which facilitates the process of decision making. Specifically, we divide data into different Increasing Feature Blocks (IFBs) according to missing patterns. Besides, we propose a novel criterion to estimate uncertainty in order to label petitions as early as possible. To handle the challenge of scalability, we present an Expectation-Maximization (EM)-based algorithm to optimize the non-convex objective function accurately and efficiently. Various experiments on six petition datasets demonstrate that our MLUE outperformed other baselines by a large margin. Andreas Züfle, Jingyuan Yang 0001, Liang Zhao 0002 |
ICDM | 4 |
| 2018 | CADEN: A Context-Aware Deep Embedding Network for Financial Opinions MiningabstractFollowing the recent advances of artificial intelligence, financial text mining has gained new potential to benefit theoretical research with practice impacts. An essential research question for financial text mining is how to accurately identify the actual financial opinions (e.g., bullish or bearish) behind words in plain text. Traditional methods mainly consider this task as a text classification problem with solutions based on machine learning algorithms. However, most of them rely heavily on the hand-crafted features extracted from the text. Indeed, a critical issue along this line is that the latent global and local contexts of the financial opinions usually cannot be fully captured. To this end, we propose a context-aware deep embedding network for financial text mining, named CADEN, by jointly encoding the global and local contextual information. Especially, we capture and include an attitude-aware user embedding to enhance the performance of our model. We validate our method with extensive experiments based on a real-world dataset and several state-of-the-art baselines for investor sentiment recognition. Our results show a consistently superior performance of our approach for identifying the financial opinions from texts of different formats. Liang Zhang 0031, Keli Xiao, Hengshu Zhu, Chuanren Liu, Jingyuan Yang 0001, Bo Jin 0001 |
ICDM | 5 |
| 2018 | A Unified View of Social and Temporal Modeling for B2B Marketing Campaign RecommendationabstractBusiness to Business (B2B) marketing aims at meeting the needs of other businesses instead of individual consumers, and thus entails management of more complex business needs than consumer marketing. The buying processes of the business customers involve series of different marketing campaigns providing multifaceted information about the products or services. While most existing studies focus on individual consumers, little has been done to guide business customers due to the dynamic and complex nature of these business buying processes. To this end, in this paper, we focus on providing a unified view of social and temporal modeling for B2B marketing campaign recommendation. Along this line, we first exploit the temporal behavior patterns in the B2B buying processes and develop a marketing campaign recommender system. Specifically, we start with constructing a temporal graph as the knowledge representation of the buying process of each business customer. Temporal graph can effectively extract and integrate the campaign order preferences of individual business customers. It is also worth noting that our system is backward compatible since the participating frequency used in conventional static recommender systems is naturally embedded in our temporal graph. The campaign recommender is then built in a low-rank graph reconstruction framework based on probabilistic graphical models. Our framework can identify the common graph patterns and predict missing edges in the temporal graphs. In addition, since business customers very often have different decision makers from the same company, we also incorporate social factors, such as community relationships of the business customers, for further improving overall performances of the missing edge prediction and recommendation. Finally, we have performed extensive empirical studies on real-world B2B marketing data sets and the results show that the proposed method can effectively improve the quality of the campaign recommendations for challenging B2B marketing tasks. Jingyuan Yang 0001, Chuanren Liu, Mingfei Teng, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Buyer targeting optimization: A unified customer segmentation perspectiveabstractIn marketing analytics, customer segmentation (clustering) divides a customer base into groups of similar individuals, while buyer targeting (classification) identifies promising customers. Both customer segmentation and buyer targeting help the business to improve marketing performances by allocating resources to the most profitable customers. Due to the heterogeneity across the customer groups, some studies have been made on combining the tasks of customer segmentation and buyer targeting for tailored marketing strategies. However, these efforts usually combine these two tasks in a simple step-by-step approach. It is still unclear how to implement these two tasks in a more integrated and optimized way, which is the research objective of this paper. Specifically, we formulate customer segmentation and buyer targeting as a unified optimization problem. Then, the customer segments are adaptively realized during the targeting optimization process. In this way, the integrated approach not only improves the buyer targeting performances but also provides a new perspective of segmentation based on the buying decision preferences of the customers. The unified customer segmentation and buyer targeting method not only quantifies the purchase tendency of a specific customer but also characterizes the buying decision behaviors at the segment level. We also develop an efficient K-Classifiers Segmentation algorithm to solve the unified optimization problem. Moreover, we show that the customer segmentation based on the buying decision preferences can also be consistent with the features on customer profiles. Finally, we have performed the extensive experiments on several real-world Business to Business (B2B) marketing data sets. The results show that our approach offers not only more accurate targeting of promising customers but also meaningful customer segmentation solutions with interpretable buying decision preferences for each customer segment. Jingyuan Yang 0001, Chuanren Liu, Mingfei Teng, March Liao, Hui Xiong 0001 |
IEEE BigData | 1 |
| 2016 | Talent Circle Detection in Job Transition NetworksabstractWith the high mobility of talent, it becomes critical for the recruitment team to find the right talent from the right source in an efficient manner. The prevalence of Online Professional Networks (OPNs), such as LinkedIn, enables the new paradigm for talent recruitment and job search. However, the dynamic and complex nature of such talent information imposes significant challenges to identify prospective talent sources from large-scale professional networks. Therefore, in this paper, we propose to create a job transition network where vertices stand for organizations and a directed edge represents the talent flow between two organizations for a time period. By analyzing this job transition network, it is able to extract talent circles in a way such that every circle includes the organizations with similar talent exchange patterns. Then, the characteristics of these talent circles can be used for talent recruitment and job search. To this end, we develop a talent circle detection model and design the corresponding learning method by maximizing the Normalized Discounted Cumulative Gain (NDCG) of inferred probability for the edge existence based on edge weights. Then, the identified circles will be labeled by the representative organizations as well as keywords in job descriptions. Moreover, based on these identified circles, we develop a talent exchange prediction method for talent recommendation. Finally, we have performed extensive experiments on real-world data. The results show that, our method can achieve much higher modularity when comparing to the benchmark approaches, as well as high precision and recall for talent exchange prediction. Huang Xu 0001, Zhiwen Yu 0001, Jingyuan Yang 0001, Hui Xiong 0001, Hengshu Zhu |
KDD | 3 |
| 2015 | Station Site Optimization in Bike Sharing SystemsabstractBike sharing systems, aiming at providing the missing links in the public transportation systems, are becoming popular in urban cities. In an ideal bike sharing network, the station locations are usually selected in a way that there are balanced pick-ups and drop-offs among stations. This can help avoid expensive re-balancing operations and maintain high user satisfaction. However, it is a challenging task to develop such an efficient bike sharing system with appropriate station locations. Indeed, the bike station demand is influenced by multiple factors of surrounding environment and complex public transportation networks. Limited efforts have been made to develop demand-and-balance prediction models for bike sharing systems by considering all these factors. To this end, in this paper, we propose a bike sharing network optimization approach by considering multiple influential factors. The goal is to enhance the quality and efficiency of the bike sharing service by selecting the right station locations. Along this line, we first extract fine-grained discriminative features from human mobility data, point of interests (POI), as well as station network structures. Then, prediction models based on Artificial Neural Networks (ANN) are developed for predicting station demand and balance. In addition, based on the learned patterns of station demand and balance, a genetic algorithm based optimization model is built to choose a set of stations from a large number of candidates in a way such that the station usage is maximized and the number of unbalanced stations is minimized. Finally, the extensive experimental results on the NYC CitiBike sharing system show the advantages of our approach for optimizing the station site allocation in terms of the bike usage as well as the required re-balancing efforts. Meng Qu, Weiwei Chen 0003, Jingyuan Yang 0001, Hui Xiong 0001, Hao Zhong 0002, Yanjie Fu |
ICDM | 5 |
| 2015 | Exploiting Temporal and Social Factors for B2B Marketing Campaign RecommendationsabstractBusiness to Business (B2B) marketing aims at meeting the needs of other businesses instead of individual consumers. In B2B markets, the buying processes usually involve series of different marketing campaigns providing necessary information to multiple decision makers with different interests and motivations. The dynamic and complex nature of these processes imposes significant challenges to analyze the process logs for improving the B2B marketing practice. Indeed, most of the existing studies only focus on the individual consumers in the markets, such as movie/product recommender systems. In this paper, we exploit the temporal behavior patterns in the buying processes of the business customers and develop a B2B marketing campaign recommender system. Specifically, we first propose the temporal graph as the temporal knowledge representation of the buying process of each business customer. The key idea is to extract and integrate the campaign order preferences of the customer using the temporal graph. We then develop the low-rank graph reconstruction framework to identify the common graph patterns and predict the missing edges in the temporal graphs. We show that the prediction of the missing edges is effective to recommend the marketing campaigns to the business customers during their buying processes. Moreover, we also exploit the community relationships of the business customers to improve the performances of the graph edge predictions and the marketing campaign recommendations. Finally, we have performed extensive empirical studies on real-world B2B marketing data sets and the results show that the proposed method can effectively improve the quality of the campaign recommendations for challenging B2B marketing tasks. Jingyuan Yang 0001, Chuanren Liu, Mingfei Teng, Hui Xiong 0001, March Liao, Vivian Zhu |
ICDM | 1 |