VLDB 2026 Research / reviewers in the wild / expert
Wei Zhang 0249
dblp:10/4661-249
· DBLP profile ↗
9ranked-venue papers
9as first author
4since 2021 · last 2022
0000-0001-6123-2786ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Towards addressing unauthorized sharing of subscriptions
Wei Zhang 0249, Chris Challis |
Appl. Intell. | 1 |
| 2021 | Virtual-SRE For Monitoring Large Scale Time-series DataabstractMonitoring online time-series data and providing real-time alerts are crucial for a variety of industry applications. But doing it at a large scale is a challenging task. In this paper, we introduce Virtual-SRE, an anomaly detection framework which utilizes human knowledge and is suitable for large-scale applications. The proposed approach starts without any human labeling. Instead, human knowledge is obtained through feedbacks. A joint data and model optimization process is proposed to ensure that accurate models can be built. This makes it easily applicable to large-scale and heterogeneous tasks while enjoying the benefits of supervised learning. The proposed framework has been deployed to monitor thousands of time-series data streams and obtained favorable results. Wei Zhang 0249, Chris Challis |
IEEE BigData | 1 |
| 2021 | Learning User Preferences Without FeedbacksabstractRecommending relevant data is vital for helping users to navigate through the ocean of data. We developed a service that learns user preferences through natural user interactions, without asking for user feedbacks, so users are not distracted from their regular workflow. Our approach has few parameters and very low time and space complexities, making it suitable for large scale applications. We demonstrate through experiments how it converges to user preferences and adapts to user behavior changes. Wei Zhang 0249, Chris Challis |
DSAA | 1 |
| 2021 | Weakly Supervised Anomaly Detection for Streaming DataabstractAnomaly detection for time-series data is an important component for monitoring tasks. Due to the massive and diverse types of streaming sources, it is impossible to manually label training data for large-scale applications. So existing solutions are based on unsupervised learning. However, each time-series data has its own domain specific property. Without incorporating human knowledge, unsupervised learning methods cannot match the performance of supervised learning. We introduce an anomaly detection solution which utilizes human knowledge and is suitable for large-scale applications. The proposed approach starts without any human labeling. Instead, human knowledge is obtained through feedbacks. We iteratively refine the training data to ensure that accurate models can be built. This makes it easily applicable to large-scale and heterogeneous tasks while enjoying the benefits of supervised learning. Wei Zhang 0249, Chris Challis |
ISM | 1 |
| 2020 | Efficient Bug Triage For Industrial EnvironmentsabstractBug triage is an important task for software maintenance, especially in the industrial environment, where timely bug fixing is critical for customer experience. This process is usually done manually and often takes significant time. In this paper, we propose a machine-learning-based solution to address the problem efficiently. We argue that in the industrial environment, it is more suitable to assign bugs to software components (then to responsible developers) than to developers directly. Because developers can change their roles in industry, they may not oversee the same software module as before. We also demonstrate experimentally that assigning bugs to components rather than developers leads to much higher accuracy. Our solution is based on text-projection features extracted from bug descriptions. We use a Deep Neural Network to train the classification model. The proposed solution achieves state-of-the-art performance based on extensive experiments using multiple data sets. Moreover, our solution is computationally efficient and runs in near real-time. Wei Zhang 0249 |
ICSME | 1 |
| 2020 | Automatic Identification of Account Sharing for Video Streaming Services
Wei Zhang 0249, Chris Challis |
IEA/AIE | 1 |
| 2019 | Software Component Prediction for Bug ReportsabstractIn a software life cycle, bugs could happen at any time. Assigning bugs to relevant components/developers is a crucial task for software development. It is also a tough and resource consuming job. First, there are many components in a complex system and it is hard to understand their interactions and identify the root cause. Second, the list of components keeps growing for actively developed products and it is not easy to catch all updates. This task also faces several challenges from the machine learning point of view: 1) the ground truth is mixed with multiple levels of labels; 2) the data are severely imbalanced. 3). concept drift as future bugs are unlikely to come from the same distribution as the historical data. In this paper, we present a machine learning based solution for the bug assignment problem. We build component classifiers using a multi-layer Neural Network, based on features that were learned from data directly. A hierarchical classification framework is proposed to address the mixed label problem and improve the prediction accuracy. We also introduce a recency based sampling procedure to alleviate the data imbalance and concept drift problem. Our solution can easily accommodate new data and handle continuous system development/update. Wei Zhang 0249, Chris Challis |
ACML | 1 |
| 2018 | Efficient Feature Selection Framework for Digital Marketing Applications
Wei Zhang 0249, Shiladitya Bose, Said Kobeissi, Scott Tomko, Chris Challis |
PAKDD (3) | 1 |
| 2017 | Adaptive Sampling Scheme for Learning in Severely Imbalanced Large Scale DataabstractImbalanced data poses a serious challenge for many machine learning and data mining applications. It may significantly affect the performance of learning algorithms. In digital marketing applications, events of interest (positive instances for building predictive models) such as click and purchase are rare. A retail website can easily receive a million visits every day, yet only a small percentage of visits lead to purchase. The large amount of raw data and the small percentage of positive instances make it challenging to build decent predictive models in a timely fashion. In this paper, we propose an adaptive sampling strategy to deal with this problem. It efficiently returns high quality training data, ensures system responsiveness and improves predictive performances. Wei Zhang 0249, Said Kobeissi, Scott Tomko, Chris Challis |
ACML | 1 |