EDBT 2026 Demo / reviewers in the wild / expert
Bo Zhou 0010
dblp:65/3628-10
· DBLP profile ↗
21ranked-venue papers
4as first author
1since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 2 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Empirical software engineering · 87% Software maintenance and evolution · 13% | |
| Human-computer interaction and pervasive computing
1 paper |
Ubiquitous computing and smart environments · 100% |
Topics — the 2 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering › developer studies
developer behavior |
0.2 | 1 | 2015 | scvRipper: Video Scraping Tool for Modeling Developers' Behavior Using Interaction Data · ICSE (2) 2015 |
Empirical software engineering
developer studies |
0.2 | 1 | 2015 | Tracking and Analyzing Cross-Cutting Activities in Developers' Daily Work (N) · ASE 2015 |
Methods — techniques the papers use, named apart from their topics
video scraping · 0.2screen capture · 0.2computer vision · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Unlocking the Power of Diversity in Index Tuning for Cluster Databases
Haitian Hang, Xiu Tang, Bo Zhou 0010, Jianling Sun |
DEXA (2) | 3 |
| 2020 | Power Trading Model for Distributed Power Generation Systems Based on Consortium BlockchainsabstractDistributed power generation systems based on clean energy experience a large development worldwide. However, existing centralized power trading models cannot support direct transactions between distributed power generation systems and power consumers. Existing centralized power trading models suffer from multi-agent untrustworthiness and risk of data tampering. In addition, the direct transactions involve a large number of producers and consumers, thus requiring high efficiency of transaction processing in power trading models. To tackle with the problems, we propose a weakly centralized power trading model based on consortium blockchains in this paper. The trading model facilitate the power trading process by using smart contracts on consortium blockchains. In addition, we propose an approach based on the Hotelling theory to determine the pricing for transaction in our power trading model. The experimental results show that our power trading model provides a stable throughput as the request rate of transactions increases. The power trading model proposed in this paper is more practical than those based on public blockchains. Jiabei Zhang, Fuquan Yuan, Bo Zhou 0010, Haiming Zhou |
Internetware | 4 |
| 2017 | Extracting and analyzing time-series HCI data from screen-captured task videos
Lingfeng Bao, Jing Li 0034, Zhenchang Xing, Xinyu Wang 0001, Xin Xia 0001, Bo Zhou 0010 |
Empir. Softw. Eng. | 6 |
| 2015 | scvRipper: Video Scraping Tool for Modeling Developers' Behavior Using Interaction DataabstractScreen-capture tool can record a user's interaction with software and application content as a stream of screenshots which is usually stored in certain video format. Researchers have used screen-captured videos to study the programming activities that the developers carry out. In these studies, screen-captured videos had to be manually transcribed to extract software usage and application content data for the study purpose. This paper presents a computer-vision based video scraping tool (called scvRipper) that can automatically transcribe a screen-captured video into time-series interaction data according to the analyst's need. This tool can address the increasing need for automatic behavioral data collection methods in the studies of human aspects of software engineering. Lingfeng Bao, Jing Li 0034, Zhenchang Xing, Xinyu Wang 0001, Bo Zhou 0010 |
ICSE (2) | 5 |
| 2015 | Tracking and Analyzing Cross-Cutting Activities in Developers' Daily Work (N)abstractDevelopers use many software applications to process large amounts of diverse information in their daily work. The information is usually meaningful beyond the context of an application that manages it. However, as different applications function independently, developers have to manually track, correlate and re-find cross-cutting information across separate applications. We refer to this difficulty as information fragmentation problem. In this paper, we present ActivitySpace, an interapplication activity tracking and analysis framework for tackling information fragmentation problem in software development. ActivitySpace can monitor the developer's activity in many applications at a low enough level to obviate application-specific support while accounting for the ways by which low-level activity information can be effectively aggregated to reflect the developer's activity at higher-level of abstraction. A system prototype has been implemented on Microsoft Windows. Our preliminary user study showed that the ActivitySpace system is promising in supporting interapplication information needs in developers' daily work. Lingfeng Bao, Zhenchang Xing, Xinyu Wang 0001, Bo Zhou 0010 |
ASE | 4 |
| 2015 | Reverse engineering time-series interaction data from screen-captured videosabstractIn recent years the amount of research on human aspects of software engineering has increased. Many studies use screen-capture software (e.g., Snagit) to record developers' behavior as they work on software development tasks. The recorded task videos capture direct information about which activities the developers carry out with which content and in which applications during the task. Such behavioral data can help researchers and practitioners understand and improve software engineering practices from human perspective. However, extracting time-series interaction data (software usage and application content) from screen-captured videos requires manual transcribing and coding of videos, which is tedious and error-prone. In this paper we present a computer-vision based video scraping technique to automatically reverse-engineer time-series interaction data from screen-captured videos. We report the usefulness, effectiveness and runtime performance of our video scraping technique using a case study of the 29 hours task videos of 20 developers in the two development tasks. Lingfeng Bao, Jing Li 0034, Zhenchang Xing, Xinyu Wang 0001, Bo Zhou 0010 |
SANER | 5 |
| 2015 | Automatic, high accuracy prediction of reopened bugs
Xin Xia 0001, David Lo 0001, Emad Shihab, Xinyu Wang 0001, Bo Zhou 0010 |
Autom. Softw. Eng. | 5 |
| 2015 | Dual analysis for recommending developers to resolve bugsabstractAbstract Bug resolution refers to the activity that developers perform to diagnose, fix, test, and document bugs during software development and maintenance. Given a bug report, we would like to recommend the set of bug resolvers that could potentially contribute their knowledge to fix it. We refer to this problem as developer recommendation for bug resolution. In this paper, we propose a new and accurate method named DevRec for the developer recommendation problem. DevRec is a composite method that performs two kinds of analysis: bug reports based analysis (BR‐Based analysis) and developer based analysis (D‐Based analysis). We evaluate our solution on five large bug report datasets including GNU Compiler Collection, OpenOffice, Mozilla, Netbeans, and Eclipse containing a total of 107,875 bug reports. We show that DevRec could achieve recall@5 and recall@10 scores of 0.4826–0.7989, and 0.6063–0.8924, respectively. The results show that DevRec on average improves recall@5 and recall@10 scores of Bugzie by 57.55% and 39.39%, outperforms DREX by 165.38% and 89.36%, and outperforms NonTraining by 212.39% and 168.01%, respectively. Moreover, we evaluate the stableness of DevRec with different parameters, and the results show that the performance of DevRec is stable for a wide range of parameters. Copyright © 2015 John Wiley & Sons, Ltd. Xin Xia 0001, David Lo 0001, Xinyu Wang 0001, Bo Zhou 0010 |
J. Softw. Evol. Process. | 4 |
| 2014 | Automated Configuration Bug Report Prediction Using Text MiningabstractConfiguration bugs are one of the dominant causes of software failures. Previous studies show that a configuration bug could cause huge financial losses in a software system. The importance of configuration bugs has attracted various research studies, e.g., To detect, diagnose, and fix configuration bugs. Given a bug report, an approach that can identify whether the bug is a configuration bug could help developers reduce debugging effort. We refer to this problem as configuration bug reports prediction. To address this problem, we develop a new automated framework that applies text mining technologies on the natural-language description of bug reports to train a statistical model on historical bug reports with known labels (i.e., Configuration or non-configuration), and the statistical model is then used to predict a label for a new bug report. Developers could apply our model to automatically predict labels of bug reports to improve their productivity. Our tool first applies feature selection techniques (e.g., Information gain and Chi-square) to pre-process the textual information in bug reports, and then applies various text mining techniques (e.g., Naive Bayes, SVM, naive Bayes multinomial) to build statistical models. We evaluate our solution on 5 bug report datasets including accumulo, activemq, camel, flume, and wicket. We show that naive Bayes multinomial with information gain achieves the best performance. On average across the 5 projects, its accuracy, configuration F-measure and non-configuration F-measure are 0.811, 0.450, and 0.880, respectively. We also compare our solution with the method proposed by Arshad et al. The results show that our proposed approach that uses naive Bayes multinomial with information gain on average improves accuracy, configuration F-measure and non-configuration F-measure scores of Arshad et al.'s method by 8.34%, 103.7%, and 4.24%, respectively. Xin Xia 0001, David Lo 0001, Weiwei Qiu, Xingen Wang, Bo Zhou 0010 |
COMPSAC | 5 |
| 2014 | Build Predictor: More Accurate Missed Dependency Prediction in Build Configuration FilesabstractSoftware build system (e.g., Make) plays an important role in compiling human-readable source code into an executable program. One feature of build system such as make-based system is that it would use a build configuration file (e.g., Make file) to record the dependencies among different target and source code files. However, sometimes important dependencies would be missed in a build configuration file, which would cause additional debugging effort to fix it. In this paper, we propose a novel algorithm named Build Predictor to mine the missed dependncies. We first analyze dependencies in a build configuration file (e.g., Make file), and establish a dependency graph which captures various dependencies in the build configuration file. Next, considering that a build configuration file is constructed based on the source code dependency relationship, we establish a code dependency graph (code graph). Build Predictor is a composite model, which combines both dependency graph and code graph, to achieve a high prediction performance. We collected 7 build configuration files from various open source projects, which are Zlib, putty, vim, Apache Portable Runtime (APR), memcached, nginx, and Tengine, to evaluate the effectiveness of our algorithm. The experiment results show that compared with the state-of-the-art link prediction algorithms used by Xia et al., our Build Predictor achieves the best performance in predicting the missed dependencies. Bo Zhou 0010, Xin Xia 0001, David Lo 0001, Xinyu Wang 0001 |
COMPSAC | 1 |
| 2014 | Automatic Defect Categorization Based on Fault Triggering ConditionsabstractDue to the complexity of software systems, defects are inevitable. Understanding the types of defects could help developers to adopt measures in current and future software releases. In practice, developers often categorize defects into various types. One common categorization is based on fault triggers of defects. Fault trigger is a set of conditions which activate a defect (i.e., Fault) and propagate the defect into a failure. In general, there are two types of defect based fault triggering conditions, Bohrbug and Mandelbug. Bohrbug refers to a bug which can be easily isolated, and its activation and error propagation is simple. Mandelbug refers to a bug whose activation and/or error propagation is complex (e.g., A time lag between the fault activation and the failure occurrence). With these category labels, developers can better perform post-mortem analysis to identify common characteristic of the defects, and design specific fault-tolerance mechanisms. However, in most software systems, these category labels are often unavailable. To address this problem, in this paper, we propose a text mining solution which categorize defects into fault trigger categories by analyzing the natural-language description of bug reports. A previous study shows that Mandelbug is more complex and needs more time to be fixed. Thus, to better identify Mandelbugs, we propose a novel Fuzzy Set based Feature Selection algorithm named USES, which selects the features (i.e., Terms) which have high ability to distinguish Mandelbugs from Bohrbugs. USES first caches a set of terms based on their fuzzy affinity scores to Bohrbug or Mandelbug. Next, it iterates many times, and in each iteration, it selects a subset of terms, and builds a classifier on these terms. USES selects the classifier and the terms which could achieve the best performance on a training data. We evaluate our solution on 4 datasets including Linux, Mysql, Apache HTTPD, and AXIS containing a total of 809 bug reports. We show that USES with naive Bayes multinomial achieves the best performance, it achieves Mandelbug F-measure scores of 0.298 - 0.615. We also compare USES with other baseline approaches. The results show that USES on average improves Mandelbug F-measure scores of the best performing baseline by 12.3%. Xin Xia 0001, David Lo 0001, Xinyu Wang 0001, Bo Zhou 0010 |
ICECCS | 4 |
| 2014 | Towards more accurate content categorization of API discussionsabstractNowadays, software developers often discuss the usage of various APIs in online forums. Automatically assigning pre-defined semantic categorizes to API discussions in these forums could help manage the data in online forums, and assist developers to search for useful information. We refer to this process as content categorization of API discussions. To solve this problem, Hou and Mo proposed the usage of naive Bayes multinomial, which is an effective classification algorithm. Bo Zhou 0010, Xin Xia 0001, David Lo 0001, Cong Tian 0001, Xinyu Wang 0001 |
ICPC | 1 |
| 2014 | Efficient Points-To Analysis for Partial Call Graph Construction
Zhiyuan Wan, Bo Zhou 0010, Yuanhong Shen |
SEKE | 2 |
| 2013 | Software Internationalization and Localization: An Industrial ExperienceabstractSoftware internationalization and localization are important steps in distributing and deploying software to different regions of the world. Internationalization refers to the process of reengineering a system such that it could support various languages and regions without further modification. Localization refers to the process of adapting an internationalized software for a specific language or region. Due to various reasons, many large legacy systems did not consider internationalization and localization at the early stage of development. In this paper, we present our experience on, and propose a process along with tool supports for software internationalization and localization. We reengineer a large legacy commercial financial system called PAM of State Street Corporation, which is written in C/C++, containing 30 different modules, and more than 5 millions of lines of source code. We propose a source code ranker that recovers important source code to be analyzed. Based on this code, we extract general patterns of the source code that need to be reengineered for internationalization. We divide the patterns into 2 categories: convertible patterns and suspicious patterns. To locate the source code that need to be modified, we develop an automated tool I18nLocator, that consumes these patterns and outputs the locations that match the patterns. The source codes matching the convertible patterns are automatically converted, and those matching the suspicious patterns are converted by developers considering the context of the corresponding codes. For localization, we extract hard-coded strings, translate them, and store them into resource data files. Out of the 504 thousands of lines of source code that are modified using our proposed approach, we can automatically modify 79.76% of them, saving much valuable developers' time. The quality of the resultant system is also good. The number of bugs per lines of code modified found during user acceptance test and deployment to the production environment is 0.000218 bugs/LOC. Xin Xia 0001, David Lo 0001, Xinyu Wang 0001, Bo Zhou 0010 |
ICECCS | 5 |
| 2013 | Geographic Location-Based Network-Aware QoS Prediction for Service CompositionabstractQoS-aware service composition intends to maximize the global QoS of a composite service while selecting candidate services from different providers with local and global QoS constraints. With more and more candidate services emerging from all over the world, the network delays often greatly impact the performance of the composite service, which are usually not easy to be collected before the composition. One remedy is to predict them for the composition. However, new issues occur in predicting network delay for the composition, including prediction accuracy and on-demand measures to new services, which affect the performance of network-aware composite services. To solve these critical challenges, in this paper, we take advantage of the geographic location information of candidate services. We propose a network-aware QoS (NQoS) model for the composite service. Based on that, we present a novel geographic location-based NQoS prediction approach before composition, and a NQoS re-prediction approach during the execution of the composite service. Extensive experiments are conducted on the real-world dataset collected from PlanetLab. Comparative experiment results reveal our approach facilitates to improve the prediction accuracy and predictability of the NQoS values, and increase global NQoS of the composite service while ensuring its reliability constraints. Yuanhong Shen, Jianke Zhu, Xinyu Wang 0001, Xiaohu Yang 0001, Bo Zhou 0010 |
ICWS | 6 |
| 2013 | Tag recommendation in software information sitesabstractNowadays, software engineers use a variety of online media to search and become informed of new and interesting technologies, and to learn from and help one another. We refer to these kinds of online media which help software engineers improve their performance in software development, maintenance and test processes as software information sites. It is common to see tags in software information sites and many sites allow users to tag various objects with their own words. Users increasingly use tags to describe the most important features of their posted contents or projects. In this paper, we propose TagCombine, an automatic tag recommendation method which analyzes objects in software information sites. TagCombine has 3 different components: 1. multilabel ranking component which considers tag recommendation as a multi-label learning problem; 2. similarity based ranking component which recommends tags from similar objects; 3. tag-term based ranking component which considers the relationship between different terms and tags, and recommends tags after analyzing the terms in the objects. We evaluate TagCombine on 2 software information sites, StackOverflow and Freecode, which contain 47,668 and 39,231 text documents, respectively, and 437 and 243 tags, respectively. Experiment results show that for StackOverflow, our TagCombine achieves recall@5 and recall@10 scores of 0.5964 and 0.7239, respectively; For Freecode, it achieves recall@5 and recall@10 scores of 0.6391 and 0.7773, respectively. Moreover, averaging over StackOverflow and Freecode results, we improve TagRec proposed by Al-Kofahi et al. by 22.65% and 14.95%, and the tag recommendation method proposed by Zangerle et al. by 18.5% and 7.35% for recall@5 and recall@10 scores. Xin Xia 0001, David Lo 0001, Xinyu Wang 0001, Bo Zhou 0010 |
MSR | 4 |
| 2011 | Ranking in Co-effecting Multi-object/Link Types NetworksabstractResearch on link based object ranking attracts increasing attention these years, which also brings computer science research and business marketing brand-new concepts, opportunities as well as a great deal of challenges. With prosperity of web pages search engine and widely use of social networks, recent graph-theoretic ranking approaches have achieved remarkable successes although most of them are focus on homogeneous networks studying. Previous study on co-ranking methods tries to divide heterogeneous networks into multiple homogeneous sub-networks and ties between different sub-networks. This paper proposes an efficient topic biased ranking method for bringing order to co-effecting heterogeneous networks among authors, papers and accepted institutions (journals/conferences) within one single random surfer. This new method aims to update ranks for different types of objects (author, paper, journals/conferences) at each random walk. Bo Zhou 0010, Manna Wu, Xin Xia 0001 |
ICTAI | 1 |
| 2010 | A Tree-Based Reliability Model for Composite Web Service with Common-Cause Failures
Bo Zhou 0010, Keting Yin, Honghong Jiang, Aleksander J. Kavs |
GPC | 1 |
| 2010 | Model Based Load Testing of Web ApplicationsabstractIn this paper, a usage model is proposed to simulate users' behaviors realistically in load testing of web applications, and another relevant workload model is proposed to help generate realistic load for load testing. It also demonstrates an eclipse-based load testing tool “Load Testing Automation Framework (LTAF)” which is based on these two models and can perform load testing of web applications easily and automatically. Furthermore, these models and tools were successfully applied into a representative web-based system from a big Corporation. Xingen Wang, Bo Zhou 0010 |
ISPA | 2 |
| 2009 | Software testing sizing in incremental development: A case studyabstractThe paper presents a case study on software testing sizing of 20 project releases delivered by a technical team through incremental development. Different sizing metrics including KLOC, development effort, test case number and function sizes are reviewed and evaluated. The results show that KLOC doesn't behave well, development effort surprisingly act as a good measure, test case number is an excellent testing sizing method, and existing function size ways are not suitable in our case. In the meanwhile, several sizing evaluation models are proposed during the discussion. Xiaochun Zhu, Bo Zhou 0010 |
ESEM | 2 |
| 2008 | Estimate Test Execution Effort at an Early Stage: An Empirical StudyabstractSoftware testing is becoming more and more important as it is a widely used activity to ensure software quality. Test execution becomes an activity in the critical path of a project. In this case, early estimation of test execution effort can benefit both tester managers and software projects. This paper reports an empirical study on early test execution effort estimation. In the study, we propose an approach which mainly consists of two parts: test case number prediction from use cases and test effort estimation based on the test suite execution vector model which combines test case number, test execution complexity and its tester together. For each part, we evaluate it with the data of real projects from a financial software company. Xiaochun Zhu, Bo Zhou 0010 |
CW | 2 |