VLDB 2026 Research / reviewers in the wild / expert
Weijing Wang
dblp:55/7249
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LogLessen: A Representation Learning Approach to Log Compression via Variational Autoencoders
Xinghui Gan, Ningjiang Chen, Weijing Wang |
ICIC (9) | 3 |
| 2026 | Quantification of the tumor microenvironment and prognostic analysis in colorectal cancer based on CAMuTILS
Tuyu Li, Lingfeng Yang, Qilai Zhang, Yibo Jin, Shiqiang Han, Nianxiang Zhan, Weijing Wang, Yukang Zeng, Yanhong Ji |
Artif. Intell. Medicine | 8 |
| 2025 | Sac-lstm: optimal resource allocation based on user intent in computility networks
Yingying Zheng, Ningjiang Chen, Yin Yin, Zizhan Huang, Weijing Wang, Xinghui Gan |
J. Supercomput. | 6 |
| 2024 | On the Evaluation of Large Language Models in Unit Test GenerationabstractUnit testing is an essential activity in software development for verifying the correctness of software components. However, manually writing unit tests is challenging and time-consuming. The emergence of Large Language Models (LLMs) offers a new direction for automating unit test generation. Existing research primarily focuses on closed-source LLMs (e.g., ChatGPT and CodeX) with fixed prompting strategies, leaving the capabilities of advanced open-source LLMs with various prompting settings unexplored. Particularly, open-source LLMs offer advantages in data privacy protection and have demonstrated superior performance in some tasks. Moreover, effective prompting is crucial for maximizing LLMs' capabilities. In this paper, we conduct the first empirical study to fill this gap, based on 17 Java projects, five widely-used open-source LLMs with different structures and parameter sizes, and comprehensive evaluation metrics. Our findings highlight the significant influence of various prompt factors, show the performance of open-source LLMs compared to the commercial GPT-4 and the traditional Evosuite, and identify limitations in LLM-based unit test generation. We then derive a series of implications from our study to guide future research and practical use of LLM-based unit test generation. Lin Yang 0030, Shutao Gao, Weijing Wang, Bo Wang 0050, Qihao Zhu, Xiao Chu, Guangtai Liang, Qianxiang Wang, Junjie Chen 0003 |
ASE | 4 |
| 2023 | A particle swarm optimization routing scheme for wireless sensor networks
Guoxiang Tong, Shushu Zhang, Weijing Wang, Guisong Yang |
CCF Trans. Pervasive Comput. Interact. | 3 |
| 2023 | Understanding and predicting incident mitigation time
Weijing Wang, Junjie Chen 0003, Lin Yang 0030, Hongyu Zhang 0002 |
Inf. Softw. Technol. | 1 |
| 2023 | Exploring Better Black-Box Test Case Prioritization via Log AnalysisabstractTest case prioritization (TCP) has been widely studied in regression testing, which aims to optimize the execution order of test cases so as to detect more faults earlier. TCP has been divided into white-box test case prioritization (WTCP) and black-box test case prioritization (BTCP) . WTCP can achieve better prioritization effectiveness by utilizing source code information, but is not applicable in many practical scenarios (where source code is unavailable, e.g., outsourced testing). BTCP has the benefit of not relying on source code information, but tends to be less effective than WTCP. That is, both WTCP and BTCP suffer from limitations in the practical use. To improve the practicability of TCP, we aim to explore better BTCP, significantly bridging the effectiveness gap between BTCP and WTCP. In this work, instead of statically analyzing test cases themselves in existing BTCP techniques, we conduct the first study to explore whether this goal can be achieved via log analysis. Specifically, we propose to mine test logs produced during test execution to more sufficiently reflect test behaviors, and design a new BTCP framework (called LogTCP), including log pre-processing, log representation, and test case prioritization components. Based on the LogTCP framework, we instantiate seven log-based BTCP techniques by combining different log representation strategies with different prioritization strategies. We conduct an empirical study to explore the effectiveness of LogTCP. Based on 10 diverse open-source Java projects from GitHub, we compared LogTCP with three representative BTCP techniques and four representative WTCP techniques. Our results show that all of our LogTCP techniques largely perform better than all the BTCP techniques in average fault detection, to the extent that they become competitive to the WTCP techniques. That demonstrates the great potential of logs in practical TCP. Junjie Chen 0003, Weijing Wang, Meng Wang 0002, Xiang Chen 0005, Jianmin Wang 0015 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2021 | Semi-supervised Log-based Anomaly Detection via Probabilistic Label EstimationabstractWith the growth of software systems, logs have become an important data to aid system maintenance. Log-based anomaly detection is one of the most important methods for such purpose, which aims to automatically detect system anomalies via log analysis. However, existing log-based anomaly detection approaches still suffer from practical issues due to either depending on a large amount of manually labeled training data (supervised approaches) or unsatisfactory performance without learning the knowledge on historical anomalies (unsupervised and semi-supervised approaches). In this paper, we propose a novel practical log-based anomaly detection approach, PLELog, which is semi-supervised to get rid of time-consuming manual labeling and incorporates the knowledge on historical anomalies via probabilistic label estimation to bring supervised approaches' superiority into play. In addition, PLELog is able to stay immune to unstable log data via semantic embedding and detect anomalies efficiently and effectively by designing an attention-based GRU neural network. We evaluated PLELog on two most widely-used public datasets, and the results demonstrate the effectiveness of PLELog, significantly outperforming the compared approaches with an average of 181.6% improvement in terms of F1-score. In particular, PLELog has been applied to two real-world systems from our university and a large corporation, further demonstrating its practicability Lin Yang 0030, Junjie Chen 0003, Weijing Wang, Jiajun Jiang, Xuyuan Dong, Wenbin Zhang 0010 |
ICSE | 4 |
| 2021 | How Long Will it Take to Mitigate this Incident for Online Service Systems?abstractOnline service systems may encounter a large number of incidents, which should be mitigated as soon as possible to minimize the service disruption time and ensure high service availability. The ability to predict TTM (Time To Mitigation) of incidents can help service teams better organize the mainte-nance efforts. Although there are many traditional bug-fixing time prediction methods, we find that there are not readily available for incident- TTM prediction due to the characteristics of incidents. To better understand how incidents are mitigated, we conduct the first empirical study of incident TTM on 20 large-scale online service systems in Microsoft. We investigate the time distribution in the main stages of the incident life cycle and explore factors affecting TTM. Based on our empirical findings, we propose TTMPred, a deep-learning-based approach for incident- TTM prediction in a continuous triage scenario. Our model designs a two-level attention-based bidirectional GRU model to capture both the semantic information in text data and the temporal information in incremental discussions. And based on a novel continuous loss function, it builds a regression model to achieve accurate TTM prediction as much as possible at each time point of prediction. Our experiments on four large-scale online service systems in Microsoft show that TTMPred is effective and significantly outperforms the compared approaches. For example, TTMPred improves the state-of-the-art regression-based approach by 25.66% on average in terms of MAE (Mean Absolute Error). Weijing Wang, Junjie Chen 0003, Lin Yang 0030, Hongyu Zhang 0002, Pu Zhao 0004, Bo Qiao 0001, Yu Kang 0006, Qingwei Lin, Saravanakumar Rajmohan, Feng Gao 0022, Zhangwei Xu, Yingnong Dang, Dongmei Zhang 0001 |
ISSRE | 1 |
| 2021 | Joint Representations of Texts and Labels with Compositional Loss for Short Text ClassificationabstractShort text classification is an important foundation for natural language processing (NLP) tasks. Though, the text classification based on deep language models (DLMs) has made a significant headway, in practical applications however, some texts are ambiguous and hard to classify in multi-class classification especially, for short texts whose context length is limited. The mainstream method improves the distinction of ambiguous text by adding context information. However, these methods rely only the text representation, and ignore that the categories overlap and are not completely independent of each other. In this paper, we establish a new general method to solve the problem of ambiguous text classification by introducing label embedding to represent each category, which makes measurable difference between the categories. Further, a new compositional loss function is proposed to train the model, which makes the text representation closer to the ground-truth label and farther away from others. Finally, a constraint is obtained by calculating the similarity between the text representation and label embedding. Errors caused by ambiguous text can be corrected by adding constraints to the output layer of the model. We apply the method to three classical models and conduct experiments on six public datasets. Experiments show that our method can effectively improve the classification accuracy of the ambiguous texts. In addition, combining our method with BERT, we obtain the state-of-the-art results on the CNT dataset. Weijing Wang |
J. Web Eng. | 2 |
| 2019 | Bi-Dimensional Representation of Patients for Diagnosis PredictionabstractPrevious work on learning representation for patients from Electronic Health Records (EHRs) has succeeded in assisting medical diagnosis. Three perspectives of records, patient symptoms, medical treatments and diagnosis codes, are consisted in EHRs. However, existing approaches on patient representation learning take one perspective of patient symptoms and medical treatments into consideration, which miss out the latent correlations between them. Actually, based on the sequence of hospital visits, physical symptoms and associated treatments together affect the diagnosis and recovery of patients. In this paper, we propose Patient2vec, a novel model to learn the bi-dimensional representation for patients by jointly extracting features from physical symptoms and medical treatments. We introduce RNN model into Patient2vec to learn the sequential context-aware features of visits. The learned representations are then fed into a classifier to diagnosis prediction. Experiments on public dataset through multi-classification tasks indicate that Patient2vec achieves up to 76% improvement in area under the ROC curve (AUC) on average, demonstrating that our method significantly outperforms single dimension representation for patients. Weijing Wang, Chenkai Guo, Jing Xu 0008, Ao Liu 0007 |
COMPSAC (2) | 1 |
| 2019 | Systematic Comprehension for Developer Reply in Mobile System ForumabstractReview-based software development has become increasingly prevalent in recent years. Existing efforts aiming at either informative evaluation or sentiment analysis are mainly from the perspective of the reviewers, while neglecting the attitude and behavior of the developers. Such efforts inevitably suffer from recommendation bias in practice, and thus benefit little for the improvement of user reviews.In this paper, we attempt to bridge the gap between user review and developer reply, and conduct a systematic study for review reply in development forums, especially in Chinese mobile system forums. To this end, we concentrate on three research questions: 1) should a targeted review be replied; 2) how long time it should be replied; 3) does traditional review analysis help to pursue a reply for certain review? To answer such questions, given certain review datasets, we perform a systematical study including the following three stages: 1) a binary classification for reply behavior prediction, 2) a regression for prediction of reply time, 3) a systematic factor study for the relationship between traditional review analysis and reply performance. To enhance the accuracy of prediction and analysis, we proposed a CNN-based weak-supervision analysis framework, which exploits manifold techniques from NLP and deep learning. We validate our approach via extensive comparison experiments. The results show that our analysis framework is effective. More importantly, we have uncovered several interesting findings, which provide valuable guidance for further review improvement and recommendation. Chenkai Guo, Weijing Wang, Naipeng Dong, Quanqi Ye, Jing Xu 0008 |
SANER | 2 |
| 2012 | Study on Web Text Feature Selection Based on Rough Set
Xianghua Lu, Weijing Wang |
ICIC (1) | 2 |