VLDB 2026 Research / reviewers in the wild / expert
Wei Fu 0002
dblp:26/4472-2
· DBLP profile ↗
11ranked-venue papers
3as first author
2since 2021 · last 2022
0000-0001-9118-9563ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Predicting health indicators for open source projects (using hyperparameter optimization)
Tianpei Xia, Wei Fu 0002, Tim Menzies |
Empir. Softw. Eng. | 2 |
| 2021 | How to "DODGE" Complex Software AnalyticsabstractMachine learning techniques applied to software engineering tasks can be improved by hyperparameter optimization, i.e., automatic tools that find good settings for a learner's control parameters. We show that such hyperparameter optimization can be unnecessarily slow, particularly when the optimizers waste time exploring “redundant tunings”, i.e., pairs of tunings which lead to indistinguishable results. By ignoring redundant tunings,DODGE($\mathcal {E})$E), a tuning tool, runs orders of magnitude faster, while also generating learners with more accurate predictions than seen in prior state-of-the-art approaches. Amritanshu Agrawal, Wei Fu 0002, Xipeng Shen, Tim Menzies |
IEEE Trans. Software Eng. | 2 |
| 2018 | 500+ times faster than deep learning: a case study exploring faster methods for text mining stackoverflowabstractDeep learning methods are useful for high-dimensional data and are becoming widely used in many areas of software engineering. Deep learners utilizes extensive computational power and can take a long time to train- making it difficult to widely validate and repeat and improve their results. Further, they are not the best solution in all domains. For example, recent results show that for finding related Stack Overflow posts, a tuned SVM performs similarly to a deep learner, but is significantly faster to train. Suvodeep Majumder, Nikhila Balaji, Katie Brey, Wei Fu 0002, Tim Menzies |
MSR | 4 |
| 2018 | Data-driven search-based software engineeringabstractThis paper introduces Data-Driven Search-based Software Engineering (DSE), which combines insights from Mining Software Repositories (MSR) and Search-based Software Engineering (SBSE). While MSR formulates software engineering problems as data mining problems, SBSE reformulate Software Engineering (SE) problems as optimization problems and use meta-heuristic algorithms to solve them. Both MSR and SBSE share the common goal of providing insights to improve software engineering. The algorithms used in these two areas also have intrinsic relationships. We, therefore, argue that combining these two fields is useful for situations (a) which require learning from a large data source or (b) when optimizers need to know the lay of the land to find better solutions, faster. Vivek Nair, Amritanshu Agrawal, Wei Fu 0002, George Mathew, Tim Menzies, Leandro L. Minku, Markus Wagner 0007, Zhe Yu 0002 |
MSR | 4 |
| 2018 | Applications of psychological science for actionable analyticsabstractAccording to psychological scientists, humans understand models that most match their own internal models, which they characterize as lists of "heuristic"s (i.e. lists of very succinct rules). One such heuristic rule generator is the Fast-and-Frugal Trees (FFT) preferred by psychological scientists. Despite their successful use in many applied domains, FFTs have not been applied in software analytics. Accordingly, this paper assesses FFTs for software analytics. Wei Fu 0002, Rahul Krishna, Tim Menzies |
ESEC/SIGSOFT FSE | 2 |
| 2018 | What is wrong with topic modeling? And how to fix it using search-based software engineering
Amritanshu Agrawal, Wei Fu 0002, Tim Menzies |
Inf. Softw. Technol. | 2 |
| 2018 | Heterogeneous Defect PredictionabstractMany recent studies have documented the success of cross-project defect prediction (CPDP) to predict defects for new projects lacking in defect data by using prediction models built by other projects. However, most studies share the same limitations: it requires homogeneous data; i.e., different projects must describe themselves using the same metrics. This paper presents methods for heterogeneous defect prediction (HDP) that matches up different metrics in different projects. Metric matching for HDP requires a “large enough” sample of distributions in the source and target projects-which raises the question on how large is “large enough” for effective heterogeneous defect prediction. This paper shows that empirically and theoretically, “large enough” may be very small indeed. For example, using a mathematical model of defect prediction, we identify categories of data sets were as few as 50 instances are enough to build a defect prediction model. Our conclusion for this work is that, even when projects use different metric sets, it is possible to quickly transfer lessons learned about defect prediction. Jaechang Nam, Wei Fu 0002, Sunghun Kim 0001, Tim Menzies, Lin Tan 0001 |
IEEE Trans. Software Eng. | 2 |
| 2017 | Easy over hard: a case study on deep learningabstractWhile deep learning is an exciting new technique, the benefits of this method need to be assessed with respect to its computational cost. This is particularly important for deep learning since these learners need hours (to weeks) to train the model. Such long training time limits the ability of (a)~a researcher to test the stability of their conclusion via repeated runs with different random seeds; and (b)~other researchers to repeat, improve, or even refute that original work. Wei Fu 0002, Tim Menzies |
ESEC/SIGSOFT FSE | 1 |
| 2017 | Revisiting unsupervised learning for defect predictionabstractCollecting quality data from software projects can be time-consuming and expensive. Hence, some researchers explore "unsupervised" approaches to quality prediction that does not require labelled data. An alternate technique is to use "supervised" approaches that learn models from project data labelled with, say, "defective" or "not-defective". Most researchers use these supervised models since, it is argued, they can exploit more knowledge of the projects. Wei Fu 0002, Tim Menzies |
ESEC/SIGSOFT FSE | 1 |
| 2016 | Too much automation? the bellwether effect and its implications for transfer learningabstractTransfer learning: is the process of translating quality predictors learned in one data set to another. Transfer learning has been the subject of much recent research. In practice, that research means changing models all the time as transfer learners continually exchange new models to the current project. This paper offers a very simple bellwether transfer learner. Given N data sets, we find which one produce the best predictions on all the others. This bellwether data set is then used for all subsequent predictions (or, until such time as its predictions start failing-- at which point it is wise to seek another bellwether). Bellwethers are interesting since they are very simple to find (just wrap a for-loop around standard data miners). Also, they simplify the task of making general policies in SE since as long as one bellwether remains useful, stable conclusions for N data sets can be achieved just by reasoning over that bellwether. From this, we conclude (1) this bellwether method is a useful (and very simple) transfer learning method; (2) bellwethers are a baseline method against which future transfer learners should be compared; (3) sometimes, when building increasingly complex automatic methods, researchers should pause and compare their supposedly more sophisticated method against simpler alternatives. Rahul Krishna, Tim Menzies, Wei Fu 0002 |
ASE | 3 |
| 2016 | Tuning for software analytics: Is it really necessary?
Wei Fu 0002, Tim Menzies, Xipeng Shen |
Inf. Softw. Technol. | 1 |