VLDB 2026 Research / reviewers in the wild / expert
Yusen Xia
dblp:40/4308
· DBLP profile ↗
10ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-2360-5574ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Theory of computation · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Variational Bayesian Semi-Supervised Keyword ExtractionabstractThe expansion of textual data, stemming from various sources such as online product reviews and scholarly publications on scientific discoveries, has created a significant demand for the extraction of succinct yet comprehensive information. While many methods have been proposed for automatic keyword extraction in unsupervised and fully supervised settings, effectively leveraging a partial list of known keywords, such as author-specified keywords or Twitter hashtags, remains under-explored. This work aims to enhance both the effectiveness and scalability of semi-supervised keyword extraction. We propose a novel variational Bayesian semi-supervised (VBSS) method that builds upon recent Bayesian advancement in the field, replacing computationally expensive posterior sampling with variational inference and data augmentation. This leads to closed-form updates and substantial speedups, particularly for long texts. Our numerical results show that the VBSS method not only improves performance on longer texts but also offers better control over false discovery rates compared to state-of-the-art keyword extraction techniques. Yaofang Hu, Yichen Cheng, Yusen Xia, Xinlei Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | A variational Bayesian approach for multimodal multi-instance classification
Yaofang Hu, Yichen Cheng, Yusen Xia, Xinlei Wang 0001 |
Pattern Recognit. | 3 |
| 2024 | Predicting digital product performance with team composition features derived from a graph network
Houping Xiao, Yusen Xia, Aaron Baird |
Decis. Support Syst. | 2 |
| 2024 | A Deep Learning and Image Processing Pipeline for Object Characterization in Firm OperationsabstractGiven the abundance of images related to operations that are being captured and stored, it behooves firms to innovate systems using image processing to improve operational performance that refers to any activity that can save labor cost. In this paper, we use deep learning techniques, combined with classic image/signal processing methods, to propose a pipeline to solve certain types of object counting and layer characterization problems in firm operations. Using data obtained by us through a collaborative effort with real manufacturers, we demonstrate that the proposed pipeline method is able to achieve higher than 93% accuracy in layer and log counting. Theoretically, our study conceives, constructs, and evaluates proof of concept of a novel pipeline method in characterizing and quantifying the number of defined items with images, which overcomes the limitations of methods based only on deep learning or signal processing. Practically, our proposed method can help firms significantly reduce labor costs and/or improve quality and inventory control by recording the number of products in real time, more accurately and with minimal up-front technological investment. The codes and data are made publicly available online through the INFORMS Journal on Computing GitHub site. History: Accepted by Ram Ramesh, Area Editor for Data Science and Machine Learning. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.0260 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2022.0260 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Alireza Aghasi, Arun Rai, Yusen Xia |
INFORMS J. Comput. | 3 |
| 2023 | Bayesian multitask learning for medicine recommendation based on online patient reviewsabstractMOTIVATION: We propose a drug recommendation model that integrates information from both structured data (patient demographic information) and unstructured texts (patient reviews). It is based on multitask learning to predict review ratings of several satisfaction-related measures for a given medicine, where related tasks can learn from each other for prediction. The learned models can then be applied to new patients for drug recommendation. This is fundamentally different from most recommender systems in e-commerce, which do not work well for new customers (referred to as the cold-start problem). To extract information from review texts, we employ both topic modeling and sentiment analysis. We further incorporate variable selection into the model via Bayesian LASSO, which aims to filter out irrelevant features. To our best knowledge, this is the first Bayesian multitask learning method for ordinal responses. We are also the first to apply multitask learning to medicine recommendation. The sample code and data are made available at GitHub: https://github.com/thrushcyc-github/BMull. RESULTS: We evaluate the proposed method on two sets of drug reviews involving 17 depression/high blood pressure-related drugs. Overall, our method performs better than existing benchmark methods in terms of accuracy and AUC (area under the receiver operating characteristic curve). It is effective even with a small sample size and only a few available features, and more robust to possible noninformative covariates. Due to our model explainability, insights generated from our model may work as a useful reference for doctors. In practice, however, a final decision should be carefully made by combining the information from the proposed recommender with doctors' domain knowledge and past experience. AVAILABILITY AND IMPLEMENTATION: The sample code and data are publicly available at GitHub: https://github.com/thrushcyc-github/BMull. Yichen Cheng, Yusen Xia, Xinlei Wang 0001 |
Bioinform. | 2 |
| 2023 | A Bayesian Semisupervised Approach to Keyword Extraction with Only Positive and Unlabeled DataabstractIn the era of big data, people benefit from the existence of tremendous amounts of information. However, availability of said information may pose great challenges. For instance, one big challenge is how to extract useful yet succinct information in an automated fashion. As one of the first few efforts, keyword extraction methods summarize an article by identifying a list of keywords. Many existing keyword extraction methods focus on the unsupervised setting, with all keywords assumed unknown. In reality, a (small) subset of the keywords may be available for a particular article. To use such information, we propose a rigorous probabilistic model based on a semisupervised setup. Our method incorporates the graph-based information of an article into a Bayesian framework via an informative prior so that our model facilitates formal statistical inference, which is often absent from existing methods. To overcome the difficulty arising from high-dimensional posterior sampling, we develop two Markov chain Monte Carlo algorithms based on Gibbs samplers and compare their performance using benchmark data. We use a false discovery rate (FDR)-based approach for selecting the number of keywords, whereas the existing methods use ad hoc threshold values. Our numerical results show that the proposed method compared favorably with state-of-the-art methods for keyword extraction. History: Accepted by Ramaswamy Ramesh, Area Editor for Data Science and Machine Learning. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.1283 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2021.0234 ) at ( http://dx.doi.org/10.5281/zenodo.7348935 ). Guanshen Wang, Yichen Cheng, Yusen Xia, Qiang Ling 0001, Xinlei Wang 0001 |
INFORMS J. Comput. | 3 |
| 2022 | Hyper-clustering enhanced spatio-temporal deep learning for traffic and demand prediction in bike-sharing systems
Shengjie Zhao 0001, Kai Zhao 0011, Yusen Xia, Wenzhen Jia |
Inf. Sci. | 3 |
| 2021 | Supervised t-Distributed Stochastic Neighbor Embedding for Data Visualization and ClassificationabstractWe propose a novel supervised dimension-reduction method called supervised t-distributed stochastic neighbor embedding (St-SNE) that achieves dimension reduction by preserving the similarities of data points in both feature and outcome spaces. The proposed method can be used for both prediction and visualization tasks with the ability to handle high-dimensional data. We show through a variety of data sets that when compared with a comprehensive list of existing methods, St-SNE has superior prediction performance in the ultrahigh-dimensional setting in which the number of features p exceeds the sample size n and has competitive performance in the p ≤ n setting. We also show that St-SNE is a competitive visualization tool that is capable of capturing within-cluster variations. In addition, we propose a penalized Kullback–Leibler divergence criterion to automatically select the reduced-dimension size k for St-SNE. Summary of Contribution: With the fast development of data collection and data processing technologies, high-dimensional data have now become ubiquitous. Examples of such data include those collected from environmental sensors, personal mobile devices, and wearable electronics. High-dimensionality poses great challenges for data analytics routines, both methodologically and computationally. Many machine learning algorithms may fail to work for ultrahigh-dimensional data, where the number of the features p is (much) larger than the sample size n. We propose a novel method for dimension reduction that can (i) aid the understanding of high-dimensional data through visualization and (ii) create a small set of good predictors, which is especially useful for prediction using ultrahigh-dimensional data. Yichen Cheng, Xinlei Wang 0001, Yusen Xia |
INFORMS J. Comput. | 3 |
| 2010 | Managing Supply Uncertainties Through Bayesian Information UpdateabstractRecently, firms have experienced severe disasters that caused major supply disruptions. In this paper we study the strategies of dual sourcing and inventory management of a manufacturer facing disrupted supplies. Although abundant research has been conducted in this field, researchers rarely address the problem in the presence of information asymmetry or imperfection, which occurs because unstable supplies are often highly volatile and unpredictable in early stages. Without accurate and prompt forecasts of upstream supplies, it is difficult for a manufacturer to manage the disruption risks in an optimal manner. Here, a Bayesian model is proposed to dynamically update the knowledge of supply risks, which uses Dirichlet prior distributions to achieve mathematical tractability in Bayesian updating. Optimal-sourcing strategies are studied under this framework. Simulation results show that the proposed approach is effective in cost reduction and robust in reacting to imperfect or incomplete initial knowledge of disruptions. Min Chen 0014, Yusen Xia, Xinlei Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2010 | Modeling the Relationship Between EDI Implementation and Firm Performance Improvement With Neural NetworksabstractThis paper examines a number of electronic data interchange (EDI) usage and implementation factors and their role in improving a firm's efficiency, productivity and competitiveness. Unlike other studies in the literature that use exclusively linear models, we apply nonlinear neural networks to model the relationship between performance improvement and a set of predictor variables of EDI usage and supply chain coordination activities. A variable selection method is employed to identify key factors to predict a firm's operational excellence due to EDI implementation. In addition, a bootstrap resampling scheme is used to evaluate the robustness of the results. Guoqiang Peter Zhang, Craig A. Hill, Yusen Xia, Faming Liang |
IEEE Trans Autom. Sci. Eng. | 3 |