Jing Jiang 0005

dblp:68/1974-5 · DBLP profile ↗
← Back
38ranked-venue papers
15as first author
13since 2021 · last 2026
0000-0002-6582-0654ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 26 · 8 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Peer-aided repairer: empowering large language models to repair advanced student assignments
Qianhui Zhao, Li Zhang 0029, Fang Liu 0032, Yang Liu 0003, Jing Jiang 0005, Ge Li 0001, Zian Sun, Zhong-Qi Li, Yuchi Ma
Empir. Softw. Eng.8
2026 Facilitating Wise Decision-Making for Bounty Backers in Open Source Software Communities
abstract
Bounty programs have become a pivotal incentive mechanism in open-source software (OSS) communities, attracting contributors by offering monetary rewards for task completion. Despite their long-standing implementation, the optimal utilization of this mechanism from the perspective of backers (individuals or entities funding bounties) remains insufficiently understood, hindering its refinement and broader adoption. To bridge this gap, we conduct a mixed-methods study analyzing 10,561 bounty issues fromGitcoin, their linkedGitHubdevelopment data, and surveys from 46 bounty backers. We investigate three core decision-making dimensions: (1) why backers use bounties and the actual outcomes, (2) what issues backers prioritize, and (3) how bounty amounts are set. Our findings reveal that backers primarily seek to enhance developer engagement, project visibility, and task efficiency. However, the actual outcomes often diverge from expectations: although bounty issues have a higher resolution rate (+12%) than non-bounty issues, they also introduce systemic challenges, such as delayed resolutions (+33 days) and difficulties in engaging new developers. Notably, backers tend to prioritize feature-related, intermediate-complexity tasks with short completion timelines, while showing relatively less interest in overly simplistic or highly specialized work. Reward allocation follows a nuanced approach: lower bounties target beginner-friendly tasks, while higher rewards are reserved for advanced skills or multi-week commitments. However, backers often lack systematic methods to calibrate rewards, leading to frequent bounty adjustments. To enable data-driven decision-making, we propose a bounty recommendation predictor that uses empirical factors to predict appropriate bounty amount. By synthesizing these insights, our study offers OSS communities actionable strategies to refine bounty programs, balancing short-term productivity with long-term ecosystem sustainability.
Xin Tan 0003, Xianjun Ni, Yuxia Zhang, Jing Jiang 0005, Minghui Zhou 0001, Li Zhang 0029
IEEE Trans. Software Eng.5
2025 Explainable Fault Localization for Programming Assignments via LLM-Guided Annotation
abstract
Providing timely and personalized guidance for students’ programming assignments, particularly by indicating fine-grained error locations with explanations, offers significant practical value for helping students complete assignments and enhance their learning outcomes. In recent years, various automated Fault Localization (FL) techniques, particularly those leveraging Large Language Models (LLMs), have demonstrated promising results in identifying errors in programs. However, existing fault localization techniques face challenges when applied to educational contexts. Most approaches operate at the method level without explanatory feedback, resulting in granularity too coarse for students who need actionable insights to identify and fix their errors. While some approaches attempt line-level fault localization, they often depend on predicting line numbers directly in numerical form, which is ill-suited to LLMs. To address these challenges, we propose FLAME, a fine-grained, explainable Fault Localization method tailored for programming assignments via LLM-guided Annotation and Model Ensemble. FLAME leverages rich contextual information specific to programming assignments to guide LLMs in identifying faulty code lines. Instead of directly predicting line numbers, we prompt the LLM to annotate faulty code lines with detailed explanations, enhancing both localization accuracy and educational value. To further improve reliability, we introduce a weighted multi-model voting strategy that aggregates results from multiple LLMs to determine the suspiciousness of each code line. Extensive experimental results demonstrate that FLAME outperforms state-of-the-art fault localization baselines on programming assignments, successfully localizing 207 more faults at top-1 over the best-performing baseline. Beyond educational contexts, FLAME also generalizes effectively to general-purpose software codebases, outperforming all baselines on the Defects4J benchmark.
Fang Liu 0032, Tianze Wang, Li Zhang 0029, Jing Jiang 0005, Zian Sun
ASE5
2025 Measuring and Mining Community Evolution in Developer Social Networks with Entropy-Based Indices
abstract
This work presents four novel entropy-based indices for measuring the community evolution of developer social networks (DSNs) in open source software (OSS) projects. The proposed indices offer a quantitative measure of community split, shrink, merge, and expand events. The indices have proven properties like monotonicity, and they have defined maximum and minimum values that signify meaningful scenarios. These indices can be combined to describe complex community evolution events such as emergence and extinction. Expanding upon these indices, this research proposes a novel machine learning approach, leveraging shapelet mining, to unearth representative patterns of community evolution. The results from real-world OSS projects show that these indices effectively capture various community evolution behaviors with a 94.1% accuracy compared to existing work. They also predict OSS team productivity with a 0.718 accuracy. With the shapelet mining and learning framework, the indices can identify patterns of community evolution and predict the survival of OSS projects with 93% accuracy 3 months before the projects’ last observed commits. The findings highlight the potential of these entropy-based indices for understanding OSS project status and predicting future trends, which are valuable for supporting future research on DSNs and OSS communities.
Jierui Zhang, Liang Wang 0006, Ying Li 0114, Jing Jiang 0005, Tao Wang 0158, XianPing Tao
ACM Trans. Softw. Eng. Methodol.4
2024 FastFixer: An Efficient and Effective Approach for Repairing Programming Assignments
abstract
Providing personalized and timely feedback for student's programming assignments is useful for programming education. Automated program repair (APR) techniques have been used to fix the bugs in programming assignments, where the Large Language Models (LLMs) based approaches have shown promising results. Given the growing complexity of identifying and fixing bugs in advanced programming assignments, current fine-tuning strategies for APR are inadequate in guiding the LLM to identify bugs and make accurate edits during the generative repair process. Furthermore, the autoregressive decoding approach employed by the LLM could potentially impede the efficiency of the repair, thereby hindering the ability to provide timely feedback. To tackle these challenges, we propose FastFixer, an efficient and effective approach for programming assignment repair. To assist the LLM in accurately identifying and repairing bugs, we first propose a novel repair-oriented fine-tuning strategy, aiming to enhance the LLM's attention towards learning how to generate the necessary patch and its associated context. Furthermore, to speed up the patch generation, we propose an inference acceleration approach that is specifically tailored for the program repair task. The evaluation results demonstrate that FastFixer obtains an overall improvement of 20.46% in assignment fixing when compared to the state-of-the-art baseline. Considering the repair efficiency, FastFixer achieves a remarkable inference speedup of 16.67× compared to the autoregressive decoding algorithm.
Fang Liu 0032, Qianhui Zhao, Jing Jiang 0005, Li Zhang 0029, Zian Sun, Ge Li 0001, Zhong-Qi Li, Yuchi Ma
ASE4
2024 Understanding Real-Time Collaborative Programming: A Study of Visual Studio Live Share
abstract
Real-time collaborative programming (RCP) entails developers working simultaneously, regardless of their geographic locations. RCP differs from traditional asynchronous online programming methods, such as Git or SVN, where developers work independently and update the codebase at separate times. Although various real-time code collaboration tools (e.g., Visual Studio Live Share , Code with Me , and Replit ) have kept emerging in recent years, none of the existing studies explicitly focus on a deep understanding of the processes or experiences associated with RCP. To this end, we combine interviews and an e-mail survey with the users of Visual Studio Live Share , aiming to understand (i) the scenarios, (ii) the requirements, and (iii) the challenges when developers participate in RCP. We find that developers participate in RCP in 18 different scenarios belonging to six categories, e.g., pair programming , group debugging , and code review . However, existing users’ attitudes toward the usefulness of the current RCP tools in these scenarios were significantly more negative than the expectations of potential users. As for the requirements, the most critical category is live editing , followed by the need for sharing terminals to enable hosts and guests to run commands and see the results, as well as focusing and following , which involves “following” the host’s edit location and “focusing” the guests’ attention on the host with a notification. Under these categories, we identify 17 requirements, but most of them are not well supported by current tools. In terms of challenges, we identify 19 challenges belonging to seven categories. The most severe category of challenges is lagging followed by permissions and conflicts . The above findings indicate that the current RCP tools and even collaborative environment need to be improved greatly and urgently. Based on these findings, we discuss the recommendations for different stakeholders, including practitioners, tool designers, and researchers.
Xin Tan 0003, Xinyue Lv, Jing Jiang 0005, Li Zhang 0029
ACM Trans. Softw. Eng. Methodol.3
2022 The Influence of Sponsorship on Open-Source Software Developers' Activities on GitHub
abstract
Studies on the OSS communities have shown that financial supports are critical to OSS developers and projects to maintain their progress and sustainability. However, there were few developers being paid directly for maintaining OSS projects in the past. The GitHub Sponsors program that brings financial supports to the general OSS developers in GitHub-the world's largest OSS platform may make a difference on this situation in the future. In this paper, we present a data set on GitHub Sponsors and conduct a data-driven study to analyze the participants of the program and the impact of sponsorships to developers' activities and their projects' outcomes and qualities. The results of our survey suggest that most developers state they will contribute more with sponsorships and provide some privilege for their sponsors. And through quantitative study, we find that developers make more contributions on GitHub after they got/offered sponsorships. Moreover, gaining sponsorship also has a weakly positive impact on developers' collaborators that did not get sponsorship. And not only developers, but their own or contributed projects also can be motivate by sponsorships. Our findings are useful to the community by understanding the impact of sponsorships on users' activities and projects' progress and sustainability, and helping the managers to improve the current financial support mechanism.
Liang Wang 0006, Hao Hu 0001, Jing Jiang 0005, Hongyu Kuang, XianPing Tao
COMPSAC4
2022 A method for identifying references between projects in GitHub
Baochuan Liu, Li Zhang 0029, Jing Jiang 0005, Liang Wang 0006
Sci. Comput. Program.3
2022 How Developers Modify Pull Requests in Code Review
abstract
In pull-based development process, contributors submit their code to open-source projects by pull requests, which are accepted or rejected by reviewers. Contributors may modify their code, which causes several iterations of code review process, and makes code reviews time-consuming for both contributors and reviewers. In this article, we set out to study pull request modifications in a code review process. We collect nine projects on GitHub with 104 307 pull requests, and investigate pull request modifications through analyzing added commits after pull requests’ submission. By studying four research questions, we conclude our major findings as follow. First, 34.56$\%$of collected pull requests have modifications. Pull requests with modifications have longer lifetime but higher pass rates. Second, we conclude eight modification types indicating why pull requests are modified. Third, we propose a novel method called MClassify to automatically classify pull request modifications, which achieves the accuracy of 0.807. Fourth, various modification types affect code review differently from the perspective of lifetime and pass rate. Pull requests with source control system management modifications have the longest lifetime. These findings enable developers and researchers to understand a pull-based code review process better and make improvements.
Jing Jiang 0005, Jiangfeng Lv, Jiateng Zheng, Li Zhang 0029
IEEE Trans. Reliab.1
2021 Predicting accepted pull requests in GitHub
Jing Jiang 0005, Jiateng Zheng, Yun Yang 0001, Li Zhang 0029
Sci. China Inf. Sci.1
2021 Hot question prediction in Stack Overflow
abstract
Abstract Stack Overflow is a very popular programming question and answer community. Some questions become hot, and receive high views, which are of widespread concern to developers. Finding hot questions early can give priority to recommend potential hot questions to answers, thereby shortening the response time. Besides, the hot question prediction is also helpful for making advertising plan, planning advertising campaigns and estimating costs. Therefore, it is important to predict hot questions. The authors propose the VSAF method which analyses the V iew amount changes, A nswer amount changes and S core changes soon after questions' creation based on F ully convolutional neural network. The performance of the VSAF method based on a training set and two different test sets has been evaluated. The training set has 1600 hot questions and 1600 cold questions. The random test set has 381 hot questions and 2819 cold questions, while the balanced test set has 400 hot questions and 400 cold questions. The experimental results show that using the balanced test set, VSAF achieves Accuracy, F 1 hot and F 1 cold of 80%, 77.77% and 81.81%, which outperforms the baseline approach by 25.59%, 21.52% and 29.04%, respectively. Using the random test set for evaluation, VSAF achieves Accuracy , F 1 hot and F 1 cold of 84.91%, 53.96% and 90.97%, which outperforms the baseline approach by 31.83%, 84.16% and 19.35%, respectively. The VSAF method significantly outperforms the state‐of‐the‐art approach on hot question prediction.
Lixian Zhao, Li Zhang 0029, Jing Jiang 0005
IET Softw.3
2021 Recommending tags for pull requests in GitHub
Jing Jiang 0005, Qiudi Wu, Xin Xia 0001, Li Zhang 0029
Inf. Softw. Technol.1
2021 image2emmet: Automatic code generation from web user interface image
abstract
Abstract Web development usually follows with analyzing the functionality, designing the user interface (UI) prototype, implementing the UI by front‐end (FE) developers and implementing the REpresentational State Transfer (RESTful) application programming interface (API) by back‐end (BE) programmers. Unfortunately, web development is a tedious, cumbersome, and time‐consuming task, which makes it a challenge for the FE programmers to work in an efficient way. In this paper, we propose an approach, image2emmet, to assist FE programmers in implementing the UI. First, we collect HyperText Markup Language, Cascading Style Sheets (HTML‐CSS) dataset in an automatic and efficient way. The HTML‐CSS dataset used for model training consists of HTML‐CSS code and its display images. Second, the faster region‐based convolutional neural network (CNN) (R‐CNN) is utilized to detect the UI component. Finally, we build a model combining CNN and long short‐term memory (LSTM) to transform the UI component into the HTML‐CSS code. The empirical study demonstrates that image2emmet can achieve a precision of 80% on the UI component detection and 60% on the transformation of UI component into HTML‐CSS code.
Lili Bo, Xiaobing Sun 0001, Bin Li 0006, Jing Jiang 0005
J. Softw. Evol. Process.5
2019 Detecting Duplicate Questions in Stack Overflow via Deep Learning Approaches
abstract
Stack Overflow is a popular question and answer website based on the software programming. Different users often ask the same questions in different ways, resulting in a large number of duplicate questions in Stack Overflow. Generally, the users with high reputation manually analyze and mark duplicate questions, which is time consuming and low efficiency. Therefore, the automatic duplicate question detection approach is demanded. We first investigate the application of deep learning models to software engineering task. Then, three deep learning models (i.e., CNN, RNN and LSTM) are applied to demonstrate whether they are effective to duplicate question detection task in Stack Overflow. In this paper, we explore three deep learning approaches DQ-CNN, DQ-RNN and DQ-LSTM based on CNN, RNN and LSTM to detect duplicate questions. The effectiveness of DQ-CNN, DQ-RNN and DQ-LSTM is evaluated by six different question groups. The experimental results show that DQ-LSTM outperforms DupPredictor, Dupe, DupePredictorRep-T and DupeRep in terms of recall-rate@5, recall-rate@10 and recall-rate@20 except for Ruby question group.
Li Zhang 0029, Jing Jiang 0005
APSEC3
2019 IEA: an answerer recommendation approach on stack overflow
Li Zhang 0029, Jing Jiang 0005
Sci. China Inf. Sci.3
2019 A first look at unfollowing behavior on GitHub
Jing Jiang 0005, David Lo 0001, Yun Yang 0001, Li Zhang 0029
Inf. Softw. Technol.1
2019 Who should make decision on this pull request? Analyzing time-decaying relationships and file similarities for integrator prediction
Jing Jiang 0005, David Lo 0001, Jiateng Zheng, Xin Xia 0001, Yun Yang 0001, Li Zhang 0029
J. Syst. Softw.1
2018 Predicting Which Pull Requests Will Get Reopened in GitHub
abstract
In GitHub, integrators inspect submitted code changes, make evaluation decision, and close pull requests. However, some pull requests may get reopened for further modification and code review. It is important to predict reopened pull requests immediately after pull requests' first close, and help integrators reopen pull requests in time. If pull requests are reopened a long time after their close, they may cause conflicts with newly submitted pull requests, add software maintenance cost, and increase burden for already busy developers. To the best of our knowledge, we present the first look at predicting reopened pull requests in GitHub. We propose an approach DTPre which is an automatic predictor of reopened pull requests based on Decision Tree classifier. DTPre mainly analyzes code features of modified changes, review features during evaluation, and developer feature of contributors. We evaluate the effectiveness of DTPre on 7 Open Source projects containing 100,622 pull requests. Experimental results show that DTPre has high performances by achieving a precision of 95.53%, recall of 99.01% and F1-measure of 97.23% on average. In comparison with predictors based on neural network, naïve Bayes, logistic regression and SVM, DTPre based on decision tree improves F-1 measures by 41.76%, 59.45%, 42.25% and 9.98% on average across 7 projects.
Abdillah Mohamed, Li Zhang 0029, Jing Jiang 0005, Ahmed Ktob
APSEC3
2018 Recommending frequently encountered bugs
abstract
Developers introduce bugs during software development which reduce software reliability. Many of these bugs are commonly occurring and have been experienced by many other developers. Informing developers, especially novice ones, about commonly occurring bugs in a domain of interest (e.g., Java), can help developers comprehend program and avoid similar bugs in the future. Unfortunately, information about commonly occurring bugs are not readily available. To address this need, we propose a novel approach named RFEB which recommends frequently encountered bugs (FEBugs) that may affect many other developers. RFEB analyzes Stack Overflow which is the largest software engineering-specific Q&A communities. Among the plenty of questions posted in Stack Overflow, many of them provide the descriptions and solutions of different kinds of bugs. Unfortunately, the search engine that comes with Stack Overflow is not able to identify FEBugs well. To address the limitation of the search engine of Stack Overflow, we propose RFEB which is an integrated and iterative approach that considers both relevance and popularity of Stack Overflow questions to identify FEBugs. To evaluate the performance of RFEB, we perform experiments on a dataset from Stack Overflow which contains more than ten million posts. We compared our model with Stack Overflow's search engine on 10 domains, and the experiment results show that RFEB achieves the average NDCG10 score of 0.96, which improves Stack Overflow's search engine by 20%.
Yun Zhang 0011, David Lo 0001, Xin Xia 0001, Jing Jiang 0005, Jianling Sun
ICPC4
2018 Timing Analysis for Microkernel-based Real-Time Embedded System
abstract
Currently, more and more application-specific operating systems (ASOS) are applied in real-time embedded systems.With the development of microkernel technique, the ASOS is usually customized based on the microkernel using the configurable policy, which has various alternatives.In the design of the real-time embedded system (RTES) based on such ASOS, evaluating its timing performance at the early design stage is helpful to guide the designer towards choosing the most appropriate policy.However, the existing works lack a uniform approach to support analyzing the various alternatives of the configured policy.To solve this problem, this paper presents a general-purpose timing analysis approach for the ASOS-based RTES.In the analysis, a timing analysis tree is proposed to characterize the tasks and the ASOS in the RTES.Then, each of the alternative policies in the ASOS is refined by the uniform execution rules in the tree.Finally, the task's response time under the various alternative policies is analyzed by a traversal of the timing analysis tree using a uniform way.In the case study, we take the scheduling policy as an example to show the use of our approach on a real-life robot controller system.
Rongfei Xu, Li Zhang 0029, Ning Ge 0002, Jing Jiang 0005
SEKE4
2018 Empirical Research in Software Engineering - A Literature Survey
Li Zhang 0029, Jia-Hao Tian, Jing Jiang 0005, Yi-Jun Liu, Meng-Yuan Pu, Tao Yue 0002
J. Comput. Sci. Technol.3
2018 An approach for optimized feature selection in large-scale software product lines
Xiaoli Lian, Li Zhang 0029, Jing Jiang 0005, William Goss
J. Syst. Softw.3
2017 Why and how developers fork what from whom in GitHub
Jing Jiang 0005, David Lo 0001, Jia-Huan He, Xin Xia 0001, Pavneet Singh Kochhar, Li Zhang 0029
Empir. Softw. Eng.1
2017 Understanding inactive yet available assignees in GitHub
Jing Jiang 0005, David Lo 0001, Fuli Feng, Li Zhang 0029
Inf. Softw. Technol.1
2017 Who should comment on this pull request? Analyzing attributes for more accurate commenter recommendation in pull-based development
Jing Jiang 0005, Yun Yang 0001, Jia-Huan He, Xavier Blanc 0001, Li Zhang 0029
Inf. Softw. Technol.1
2017 CSLabel: An Approach for Labelling Mobile App Reviews
Li Zhang 0029, Xin-Yue Huang, Jing Jiang 0005, Ya-Kun Hu
J. Comput. Sci. Technol.3
2016 Long-Term Active Integrator Prediction in the Evaluation of Code Contributions
abstract
In open source software (OSS) projects, integrators are given high-level access to repositories so that they could maintain and manage projects.Although integrators play a critical role in evaluating code changes for OSS projects, they may be short-term active.Long-term active integrators keep in evaluating code update submission and managing responses from contributors.In order to survive and succeed, OSS projects need to attract and retain long-term active integrators.To assist OSS projects to retain active integrators, we propose a method called LTAPredict to predict whether integrators will be longterm active in the evaluation of code contributions.LTAPredict collects activity data of integrators, extracts a rich set of features, and makes prediction via machine learning techniques.We perform experiments on 37 popular projects, containing a total of 1,073 integrators.Results show that based on the Decision Tree, LTAPredict achieves the accuracy as 0.829, the precision as 0.81, the recall as 0.827 and the F1 as 0.818.Meanwhile, we evaluate the feature importance to identify the most significant indicators of long-term active integrators.We observe that whether integrators becoming long-term active is associated with the number of active months and social distance with contributors in their first year as integrators.These findings assist OSS projects to identify potential long-term active integrators and adopt better strategies to retain them in the evaluation of code contributions.
Jing Jiang 0005, Fuli Feng, Xiaoli Lian, Li Zhang 0029
SEKE1
2015 CoreDevRec: Automatic Core Member Recommendation for Contribution Evaluation
Jing Jiang 0005, Jia-Huan He, Xue-Yuan Chen
J. Comput. Sci. Technol.1
2015 Understanding Sybil Groups in the Wild
Jing Jiang 0005, Zifei Shan, Xiao Wang 0018, Li Zhang 0029, Yafei Dai
J. Comput. Sci. Technol.1
2014 For User-Driven Software Evolution: Requirements Elicitation Derived from Mining Online Reviews
Haibin Ruan, Li Zhang 0029, Philip Lew, Jing Jiang 0005
PAKDD (2)5
2014 Forwarding Links without Browsing Links in Online Social Networks
Jing Jiang 0005, Xiao Wang 0018, Li Zhang 0029, Yafei Dai
SEKE1
2013 QoS-Based Service Composition under Various QoS Requirements
abstract
There are various approaches for QoS-based web service composition. However, most of them are concerned about the algorithms of service compositions while ignoring the flexibility and expressiveness of users to set QoS constraints. Most of them assume that users could specify accurate QoS constraints easily. In reality, it is often not true especially for non-expert users. Users may not know the exact ranges or values of QoS requirements for their tasks. They may just want service composition solutions based on current QoS technical levels or want composition solutions in cost priority or quality priority. To deal with this non-clarity and variety of QoS requirements, in this paper, we propose a multi-strategic approach of service composition. This approach provides four strategies and aids to help users complete their QoS constraints and find optimization web service composition solutions.
Li Zhang 0029, Jing Jiang 0005
APSEC (1)3
2013 Understanding latent interactions in online social networks
abstract
Popular online social networks (OSNs) like Facebook and Twitter are changing the way users communicate and interact with the Internet. A deep understanding of user interactions in OSNs can provide important insights into questions of human social behavior and into the design of social platforms and applications. However, recent studies have shown that a majority of user interactions on OSNs are latent interactions , that is, passive actions, such as profile browsing, that cannot be observed by traditional measurement techniques. In this article, we seek a deeper understanding of both active and latent user interactions in OSNs. For quantifiable data on latent user interactions, we perform a detailed measurement study on Renren, the largest OSN in China with more than 220 million users to date. All friendship links in Renren are public, allowing us to exhaustively crawl a connected graph component of 42 million users and 1.66 billion social links in 2009. Renren also keeps detailed, publicly viewable visitor logs for each user profile. We capture detailed histories of profile visits over a period of 90 days for users in the Peking University Renren network and use statistics of profile visits to study issues of user profile popularity, reciprocity of profile visits, and the impact of content updates on user popularity. We find that latent interactions are much more prevalent and frequent than active events, are nonreciprocal in nature, and that profile popularity is correlated with page views of content rather than with quantity of content updates. Finally, we construct latent interaction graphs as models of user browsing behavior and compare their structural properties, evolution, community structure, and mixing times against those of both active interaction graphs and social graphs.
Jing Jiang 0005, Christo Wilson, Xiao Wang 0018, Wenpeng Sha, Peng Huang 0005, Yafei Dai, Ben Y. Zhao
ACM Trans. Web1
2011 Towards more accurate retrieval of duplicate bug reports
abstract
In a bug tracking system, different testers or users may submit multiple reports on the same bugs, referred to as duplicates, which may cost extra maintenance efforts in triaging and fixing bugs. In order to identify such duplicates accurately, in this paper we propose a retrieval function (REP) to measure the similarity between two bug reports. It fully utilizes the information available in a bug report including not only the similarity of textual content in summary and description fields, but also similarity of non-textual fields such as product, component, version, etc. For more accurate measurement of textual similarity, we extend BM25F - an effective similarity formula in information retrieval community, specially for duplicate report retrieval. Lastly we use a two-round stochastic gradient descent to automatically optimize REP for specific bug repositories in a supervised learning manner. We have validated our technique on three large software bug repositories from Mozilla, Eclipse and OpenOffice. The experiments show 10-27% relative improvement in recall rate@k and 17-23% relative improvement in mean average precision over our previous model. We also applied our technique to a very large dataset consisting of 209,058 reports from Eclipse, resulting in a recall rate@k of 37-71% and mean average precision of 47%.
Chengnian Sun, David Lo 0001, Siau-Cheng Khoo, Jing Jiang 0005
ASE4
2010 A discriminative model approach for accurate duplicate bug report retrieval
abstract
Bug repositories are usually maintained in software projects. Testers or users submit bug reports to identify various issues with systems. Sometimes two or more bug reports corre-spond to the same defect. To address the problem with du-plicate bug reports, a person called a triager needs to man-ually label these bug reports as duplicates, and link them to their ”master ” reports for subsequent maintenance work. However, in practice there are considerable duplicate bug re-ports sent daily; requesting triagers to manually label these bugs could be highly time consuming. To address this issue, recently, several techniques have be proposed using various similarity based metrics to detect candidate duplicate bug reports for manual verification. Au-tomating triaging has been proved challenging as two reports of the same bug could be written in various ways. There is still much room for improvement in terms of accuracy of du-plicate detection process. In this paper, we leverage recent advances on using discriminative models for information re-trieval to detect duplicate bug reports more accurately. We have validated our approach on three large software bug repositories from Firefox, Eclipse, and OpenOffice. We show that our technique could result in 17–31%, 22–26%, and 35– 43 % relative improvement over state-of-the-art techniques in OpenOffice, Firefox, and Eclipse datasets respectively using commonly available natural language information only.
Chengnian Sun, David Lo 0001, Xiaoyin Wang, Jing Jiang 0005, Siau-Cheng Khoo
ICSE (1)4
2010 Understanding latent interactions in online social networks
abstract
Popular online social networks (OSNs) like Facebook and Twitter are changing the way users communicate and interact with the Internet. A deep understanding of user interactions in OSNs can provide important insights into questions of human social behavior, and into the design of social platforms and applications. However, recent studies have shown that a majority of user interactions on OSNs are latent interactions, passive actions such as profile browsing that cannot be observed by traditional measurement techniques. In this paper, we seek a deeper understanding of both visible and latent user interactions in OSNs. For quantifiable data on latent user interactions, we perform a detailed measurement study on Renren, the largest OSN in China with more than 150 million users to date. All friendship links in Renren are public, allowing us to exhaustively crawl a connected graph component of 42 million users and 1.66 billion social links in 2009. Renren also keeps detailed visitor logs for each user profile, and counters for each photo and diary/blog entry. We capture detailed histories of profile visits over a period of 90 days for more than 61,000 users in the Peking University Renren network, and use statistics of profile visits to study issues of user profile popularity, reciprocity of profile visits, and the impact of content updates on user popularity. We find that latent interactions are much more prevalent and frequent than visible events, non-reciprocal in nature, and that profile popularity are uncorrelated with the frequency of content updates. Finally, we construct latent interaction graphs as models of user browsing behavior, and compare their structural properties against those of both visible interaction graphs and social graphs.
Jing Jiang 0005, Christo Wilson, Xiao Wang 0018, Peng Huang 0005, Wenpeng Sha, Yafei Dai, Ben Y. Zhao
Internet Measurement Conference1
2010 A multiple user sharing behaviors based approach for fake file detection in P2P environments
Jing Jiang 0005, Qinyuan Feng, Peng Huang 0005, Yafei Dai
Sci. China Inf. Sci.1
2009 User behavior modeling in peer-to-peer file sharing networks: Dissecting download and removal actions
abstract
User behavior models are important for building realistic simulation environment for research on P2P multimedia file-sharing systems. In this paper, we build a user download behavior model and a user removal behavior model, which can describe important user characteristics that are not captured in the existing models. The proposed download behavior model incorporates retry behavior, and the removal behavior model integrates free-riding, file usage, and file removal. Based on two-year real user logs, we derive the range of all the model parameters and generate many interesting observations. To validate the proposed models, we compare several models in a case study on the number of living replicas in P2P file-sharing systems. The results demonstrate the accuracy, usability, and advantage of the proposed models.
Qinyuan Feng, Yan Lindsay Sun, Jing Jiang 0005, Yafei Dai
ICASSP4