VLDB 2026 Research / reviewers in the wild / expert
Xinjun Mao
dblp:74/4827 · also XinJun Mao
· DBLP profile ↗
89ranked-venue papers
6as first author
42since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 50 · 2 first-author · 25 since 2021Artificial intelligence and machine learning · 26 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 3 · 1 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Evaluating LLMs in ROS robotic software code generation
Xinjun Mao, Tanghaoran Zhang, Zhiqun Xiao |
Empir. Softw. Eng. | 2 |
| 2026 | Rethinking Obscured Sub-Optimality in Analytic Learning for Exemplar-Free Class-Incremental LearningabstractExemplar-free Class-Incremental Learning (EFCIL) poses a significant challenge in mitigating catastrophic forgetting, due to the absence of exemplars. Recently, analytic learning-based methods propose a recursive alignment procedure to execute EFCIL in a phase-invariant manner and show state-of-the-art performance. However, they heavily rely on a frozen feature extractor trained with the initial dataset to avoid the misalignment between feature and label spaces, ignoring the importance of acquiring generalizable features across incremental tasks for performance improvement. To tackle this, we rethink the obscured sub-optimality of analytic learning-based methods, particularly through empirical reevaluation, and then introduce the Multi-head analytic learning (Muheal) approach. Muheal forms the multi-head model with a delicate feature extractor, thereby introducing a feature optimization procedure and a forgetting compensation module to balance the learning and forgetting. Specifically, within the feature optimization procedure, the feature extractor seeks to learn more generalizable features in a self-supervised manner using the fully-connected classification head. An analytic learning-based classification head follows to align the feature-label space. Additionally, we employ the compensation module to generate and align pseudo-features with a replicated analytic head, thus preventing overfitting and testing. Comprehensive experiments on several benchmark datasets have demonstrated that Muheal significantly outperforms existing state-of-the-art EFCIL methods and is comparable, if not superior, to methods that use replay techniques. Zijian Gao, Kele Xu, Xingxing Zhang 0001, Huiping Zhuang, Tianjiao Wan, Bo Ding 0001, Xinjun Mao, Huaimin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Towards Mitigation of False Negatives in Text-to-Image Person Re-IdentificationabstractText-to-image person re-identification (TIReID) aims to retrieve semantically related images from a large gallery given a text query. Most existing TIReID methods train the model on an ideal assumption that positive (negative) samples are semantically correlated (uncorrelated) on the visual and textual modal. However, we observe that incorrect annotation is unavoidable and ambiguous textual description is often taken as the input in practice, which leads to false negatives that ruin the feature aligning between modals for deviated identification. This work presents a novel False Negative Mitigation (FNM) method by identifying potential false negatives through distribution differences and mitigating them with dedicated loss. Specifically, the false negative mitigation loss is designed to adaptively adjust the optimization margin for potential false negatives in the latent space. Moreover, a Dual-Level Feature Representation method is introduced to leverage both global and local features during the mitigation, whereas a Momentum Contractive (MoC) module is plugged to enrich the training data for accurate similarity distribution estimation. We conduct extensive experiments on three public benchmark datasets and the results demonstrate that our FNM method achieves State-of-The-Art (SoTA) performance at all evaluation metrics. Ruigeng Zeng, Wentao Ma 0003, Tongqing Zhou, Siqi Wang 0001, Xinjun Mao, Jie Liu 0002 |
IEEE Trans. Multim. | 6 |
| 2025 | Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental LearningabstractLogit-based knowledge distillation (KD) is commonly used to mitigate catastrophic forgetting in class-incremental learning (CIL) caused by data distribution shifts. However, the strict match of logit values between student and teacher models conflicts with the cross-entropy (CE) loss objective of learning new classes, leading to significant recency bias (i.e. unfairness). To address this issue, we rethink the overlooked limitations of KD-based methods through empirical analysis. Inspired by our findings, we introduce a plug-and-play pre-process method that normalizes the logits of both the student and teacher across all classes, rather than just the old classes, before distillation. This approach allows the student to focus on both old and new classes, capturing intrinsic inter-class relations from the teacher. By doing so, our method avoids the inherent conflict between KD and CE, maintaining fairness between old and new classes. Additionally, recognizing that overconfident teacher predictions can hinder the transfer of inter-class relations (i.e., dark knowledge), we extend our method to capture intra-class relations among different instances, ensuring fairness within old classes. Our method integrates seamlessly with existing logit-based KD approaches, consistently enhancing their performance across multiple CIL benchmarks without incurring additional training costs. Zijian Gao, Shanhao Han, Xingxing Zhang 0001, Kele Xu, Dulan Zhou, Xinjun Mao, Yong Dou, Huaimin Wang 0001 |
AAAI | 6 |
| 2025 | Knowledge Memorization and Rumination for Pre-trained Model-based Class-Incremental LearningabstractClass-Incremental Learning (CIL) enables models to continuously learn new classes while mitigating catastrophic forgetting. Recently, Pre-Trained Models (PTMs) have greatly enhanced CIL performance, even when fine-tuning is limited to the first task. This advantage is particularly beneficial for CIL methods that freeze the feature extractor after first-task fine-tuning, such as analytic learning-based approaches using a least squares solution-based classification head to acquire knowledge recursively. In this work, we revisit the analytical learning approach combined with PTMs and identify its limitations in adapting to new classes, leading to sub-optimal performance. To address this, we propose the Momentum-based Analytical Learning (MoAL) approach. MoAL achieves robust knowledge memorization via an analytical classification head and improves adaptivity to new classes through momentum-based adapter weight interpolation, leading to forgetting outdated knowledge. Importantly, we introduce a knowledge rumination mechanism that leverages refined adaptivity, allowing the model to revisit and reinforce old knowledge, thereby improving performance on old classes. MoAL facilitates the acquisition of new knowledge and consolidates old knowledge, achieving a win-win outcome between plasticity and stability. Extensive experiments on various incremental settings show MoAL’s state-of-the-art performance1. Zijian Gao, Wangwang Jia, Xingxing Zhang 0001, Dulan Zhou, Kele Xu, Yong Dou, Xinjun Mao, Huaimin Wang 0001 |
CVPR | 8 |
| 2025 | False Negatives Consensus Suppression for Text-to-Image Person Re-identificatioabstractText-Image Person Re-identification (TIReID) aims to retrieve the relevant pedestrian images according to the given textual query. Recent methods typically achieve this goal through image-text contrastive learning, which assumes that only paired images and texts from the same pedestrian are considered positive samples. However, we observe that there exist negative samples, termed false negatives, that are highly semantically related to the anchor in practice. Training with these false negatives may adversely affect feature representation learning and semantic alignment between modalities. This work proposed a false negative detection and suppression (FNDS) method to mitigate their adverse impact. Our FNCD consists of a False Negative Consensus Detection (FNCD) mechanism and an Adaptive False Negative Suppression (AFNS) method. FNCD combines dual-grained detection to consensually identify potential false negatives, while AFNS assigns adaptive weights to the false negative similarities for more robust suppression. Extensive experiments conducted on three public benchmark datasets demonstrate the effectiveness of the proposed method. Ruigeng Zeng, Wentao Ma 0003, Xinjun Mao, Jie Liu 0002 |
ICME | 4 |
| 2025 | Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt EngineeringabstractCode snippet adaptation is a fundamental activity in the software development process. Unlike code generation, code snippet adaptation is not a “free creation”, which requires developers to tailor a given code snippet in order to fit specific requirements and the code context. Recently, large language models (LLMs) have confirmed their effectiveness in the code generation task with promising results. However, their performance on code snippet adaptation, a reuse-oriented and context-dependent code change prediction task, is still unclear. To bridge this gap, we conduct an empirical study to investigate the performance and issues of LLMs on the adaptation task. We first evaluate the adaptation performances of three popular LLMs and compare them to the code generation task. Our result indicates that their adaptation ability is weaker than generation, with a nearly 15% decrease on pass@1 and more context-related errors. By manually inspecting 200 cases, we further investigate the causes of LLMs' sub-optimal performance, which can be classified into three categories, i.e., Unclear Requirement, Requirement Misalignment and Context Misapplication. Based on the above empirical research, we propose an interactive prompting approach to eliciting LLMs' ability on the adaptation task. Specifically, we enhance the prompt by enriching the context and decomposing the task, which alleviates context misapplication and improves requirement understanding. Besides, we enable LLMs' reflection by requiring them to interact with a human or a LLM counselor, compensating for unclear requirement. Our experimental result reveals that our approach greatly improve LLMs' adaptation performance. The best-performing Human-LLM interaction successfully solves 159 out of the 202 identified defects and improves the pass@1 and pass@5 by over 40% compared to the initial instruction-based prompt. Considering human efforts, we suggest multi-agent interaction as a trade-off, which can achieve comparable performance with excellent generalization ability. We deem that our approach could provide methodological assistance for autonomous code snippet reuse and adaptation with LLMs. Tanghaoran Zhang, Yue Yu 0001, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Yao Lu 0003, Zhang Zhang 0005 |
ICSE | 3 |
| 2025 | Understanding the Faults in Serverless Computing Based Applications: An Empirical StudyabstractServerless computing is a novel cloud computing paradigm that enables developers to develop, deploy, and run applications in the cloud without complex and error-prone cloud resource management. However, its characteristics also introduce new types of faults (e.g., faults due to insufficient computing resource allocation) and challenges to serverless computing-based applications (abbreviated as serverless applications). While prior studies have highlighted that serverless developers encounter various challenges, no attempts have been made to understand the faults in serverless applications. These faults may cause catastrophic consequences such as application crash, thereby hindering the further spread of serverless computing. We aim in this paper to understand the symptoms, root causes, and fix patterns of faults in serverless applications. To this end, we conduct an empirical study investigating developers' issues on GitHub and posts on Stack Overflow (SO). We first identify 546 real-world serverless-related faults from GitHub and SO. Then, we manually analyze and construct taxonomies of the symptoms, root causes, and fix patterns for these faults, respectively. Our study leads to the first taxonomy for symptoms of serverlessrelated faults, covering 5 categories and 21 subcategories. The findings of our study inform that the Permission Denied error is the most common type ($\mathbf{1 0. 8 1 \%}$) of faults. Furthermore, the Incorrect Code Logic is the main cause ($\mathbf{1 7. 9 5 \%}$) behind the faults. Furthermore, we summarize 15 fix patterns that can resolve$\mathbf{7 3. 6 3 \%}$of faults in this study. Based on the results, we provide actionable implications that can potentially facilitate research and assist developers in improving the development of serverless applications. Finally, we implement a knowledge-based Q&A tool named SafHelper to help developers understand and fix faults. Changrong Xie, Yang Zhang 0026, Xinjun Mao, Kang Yang 0001, Tanghaoran Zhang |
ICSME | 3 |
| 2025 | Large Language Models Are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence TasksabstractPre-trained code models are essential for various code intelligence tasks. Yet, their effectiveness is heavily influenced by the quality of the pre-training dataset, particularly human-written reference comments, which usually serve as a bridge between the programming language and natural language. One significant challenge is that such comments could become inconsistent with the corresponding code as the software evolves, leading to suboptimal model performance. Large language models (LLMs) have demonstrated superior capabilities in generating high-quality code comments. This work investigates whether substituting original human-written comments with LLM-generated ones can improve pre-training datasets for more effective pretrained code models. As existing reference-based metrics cannot evaluate the quality of human-written reference comments themselves, to enable direct comparison between LLM-generated and human reference comments, we introduce two auxiliary tasks as novel reference-free metrics, including code-comment inconsistency detection and semantic code search. Experimental results show that LLM-generated comments exhibit superior semantic consistency with the code compared to human-written reference comments. Our manual evaluation also corroborates this conclusion, which indicates the potential of utilizing LLMs to enhance the quality of the pre-training dataset. Based on this finding, we rebuilt the CodeSearchNet dataset with LLM-generated comments and re-pre-trained the CodeT5 model. Evaluations on multiple code intelligence tasks demonstrate that models pretrained by LLM-enhanced data outperform their counterparts (pre-trained by original human reference comments data) on code summarization, code generation, and code translation tasks. This research validates the feasibility of rebuilding the pre-training dataset by LLMs to advance code intelligence tasks. It advocates rethinking the reliance on human reference comments for coderelated tasks. Kang Yang 0001, Xinjun Mao, Shangwen Wang, Yanlin Wang 0001, Tanghaoran Zhang, Bo Lin 0011, Yihao Qin, Zhang Zhang 0005, Yao Lu 0003, Kamal Al-Sabahi |
ICPC | 2 |
| 2025 | AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet AdaptationabstractRecent advancements in large language models (LLMs) have automated various software engineering tasks, with benchmarks emerging to evaluate their capabilities. However, for adaptation, a critical activity during code reuse, there is no benchmark to assess LLMs’ performance, leaving their practical utility in this area unclear. To fill this gap, we propose AdaptEval, a benchmark designed to evaluate LLMs on code snippet adaptation. Unlike existing benchmarks, AdaptEval incorporates the following three distinctive features: First, practical context. Tasks in AdaptEval are derived from developers’ practices, preserving rich contextual information from Stack Overflow and GitHub communities. Second, multi-granularity annotation. Each task is annotated with requirements at both task and adaptation levels, supporting the evaluation of LLMs across diverse adaptation scenarios. Third, fine-grained evaluation. AdaptEval includes a two-tier testing framework combining adaptation-level and function-level tests, which enables evaluating LLMs’ performance across various individual adaptations. Based on AdaptEval, we conduct the first empirical study to evaluate six instruction-tuned LLMs and especially three reasoning LLMs on code snippet adaptation. Experimental results demonstrate that AdaptEval enables the assessment of LLMs’ adaptation capabilities from various perspectives. It also provides critical insights into their current limitations, particularly their struggle to follow explicit instructions. We hope AdaptEval can facilitate further investigation and enhancement of LLMs’ capabilities in code snippet adaptation, supporting their real-world applications. Tanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yao Lu 0003, Zhang Zhang 0005, Kang Yang 0001, Yue Yu 0001 |
ASE | 2 |
| 2025 | Improving API Knowledge Comprehensibility: A Context-Dependent Entity Detection and Context Completion Approach Using LLMabstractExtracting API knowledge from Stack Overflow has become a crucial way to assist developers in using APIs. Existing research has primarily focused on extracting relevant API-related knowledge at the sentence level to enhance API documentation. However, this level of extraction can lead to a loss of crucial context, especially when sentences contain context-dependent entities (i.e., whose understanding requires reference to the surrounding context) that may hinder developers' understanding. To investigate this issue, we conducted an empirical study of 384 Stack Overflow posts and found that (1) approximately one-third of API functionality sentences contain context-dependent entities, and (2) these entities fall into two categories: Referential ContextDependent Entities and Local Variable Context-Dependent Entities. In response, we developed a novel method, CEDCC, which combines an entity filtering strategy informed by insights from our empirical study, with a large language model (LLM) to construct coreference chains for detecting context-dependent entities. Additionally, it employs a step-by-step approach with the LLM to complete the necessary context for understanding these entities. To evaluate CEDCC, we constructed a dataset of 1,023 API knowledge sentences, including 567 context-dependent entities and their required contexts. The results demonstrate the effectiveness of CEDCC in accurately detecting contextdependent entities and completing context tasks, achieving an F1score of 0.865 and a BERTScore of 0.373, significantly surpassing the baseline methods. Human evaluations further confirmed that CEDCC effectively improves the comprehensibility of API knowledge sentences. Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Tanghaoran Zhang, Xunhui Zhang |
SANER | 2 |
| 2025 | ConflictLens: an LLM-Based Method for Detecting Semantic Merge ConflictsabstractSemantic conflicts in branch merging occur when merged code violates specifications from one or both branches.These conflicts are often subtle and can lead to serious runtime errors such as crashes or data corruption.Existing detection methods fail to achieve both high precision and recall: Static analysis-based methods ensure high recall but lack precision, whereas dynamic execution-based methods provide better precision but struggle with recall due to limited test coverage.To better understand such conflicts, we first conduct an empirical study on a real-world merge dataset and identify four common conflict patterns.These patterns reveal key characteristics of semantic conflicts and serve as guidance for automated detection.Based on these insights, we propose ConflictLens, a two-stage LLM-based method that combines static analysis and dynamic execution to balance precision and recall.First, LLMs are guided by few-shot and chain-of-thought prompting using the patterns to localize conflicts statically.Then, conflicts are dynamically verified with LLM-generated targeted tests, refined through execution feedback.Evaluated on 85 real-world merge scenarios, ConflictLens achieves 0.91 precision and 0.76 recall, outperforming static and dynamic baselines.Ablation studies demonstrate the contribution and synergy of each component.Cross-LLM evaluations confirm robustness, with DeepSeek-R1 performing best and cost-efficient models like GPT-4o-Mini still competitive. Longfei Sun, Yao Lu 0003, Xinjun Mao, Tanghaoran Zhang, Zhang Zhang 0005 |
SEKE | 3 |
| 2025 | Decoding Serverless Security: Exploring Developer Challenges and Solutions from Stack OverflowabstractServerless computing is gaining increasing attention from developers due to its simplicity of infrastructure management.With the widespread adoption of this paradigm, the security of serverless computing (abbreviated as serverless security) has become a key concern.The security responsibility model of serverless computing is different from traditional architecture, which may introduce new security-related challenges (e.g., potential attack surface increase due to eventdriven architecture).While prior studies have investigated specific problems related to serverless security, no attempts have been made to explore the serverless security challenges discussed in the developer community.In this paper, we aim to gain a better understanding of serverless security-related challenges on Stack Overflow (SO).To this end, we manually analyze 472 serverless security-related questions to construct a taxonomy of challenges and then summarize common solutions for these challenges.Moreover, we employ a series of heuristics to gauge the popularity, difficulty, and expertise status of these challenges.Our study groups serveless security-related challenges into five categories and summarize 20 common solutions for them.Our analysis informs that Authentication is the most common challenge type (32.47%) and also the most difficult to address.Based on the results, we provide actionable implications that can facilitate research and help developers enhance the security of serverless applications. Changrong Xie, Yang Zhang 0026, Xinjun Mao |
SEKE | 3 |
| 2025 | FMCC-RT: a scalable and fine-grained all-reduce algorithm for large-scale SMP clusters
Jintao Peng, Jie Liu 0002, Jianbin Fang, Zhiquan Lai, Bo Yang 0023, Chunye Gong, Xinjun Mao, Guo Mao, Jie Ren 0007 |
Sci. China Inf. Sci. | 9 |
| 2025 | ROS package search for robot software development: a knowledge graph-based approach
Xinjun Mao, Shuo Yang 0005, Menghan Wu 0001, Zhang Zhang 0005 |
Frontiers Comput. Sci. | 2 |
| 2025 | Hierarchical knowledge-guided reasoning for text-based person re-identification
Ruigeng Zeng, Wentao Ma 0003, Tongqing Zhou, Shan Zhao 0002, Xinjun Mao, Jie Liu 0002 |
Neural Networks | 5 |
| 2025 | CARLDA: An Approach for Stack Overflow API Mention Recognition Driven by Context and LLM-Based Data AugmentationabstractABSTRACT The recognition of Application Programming Interface (API) mentions in software‐related texts is vital for extracting API‐related knowledge, providing deep insights into API usage and enhancing productivity efficiency. Previous research identifies two primary technical challenges in this task: (1) differentiating APIs from common words and (2) identifying morphological variants of standard APIs. While deep learning‐based methods have demonstrated advancements in addressing these challenges, they rely heavily on high‐quality labeled data, leading to another significant data‐related challenge: (3) the lack of such high‐quality data due to the substantial effort required for labeling. To overcome these challenges, this paper proposes a context‐aware API recognition method named CARLDA. This approach utilizes two key components, namely, Bidirectional Encoder Representations from Transformers (BERT) and Bidirectional Long Short‐Term Memory (BiLSTM), to extract context at both the word and sequence levels, capturing syntactic and semantic information to address the first challenge. For the second challenge, it incorporates a character‐level BiLSTM with an attention mechanism to grasp global character‐level context, enhancing the recognition of morphological features of APIs. To address the third challenge, we developed specialized data augmentation techniques using large language models (LLMs) to tackle both in‐library and cross‐library data shortages. These techniques generate a variety of labeled samples through targeted transformations (e.g., replacing tokens and restructuring sentences) and hybrid augmentation strategies (e.g., combining real‐world and generated data while applying style rules to replicate authentic programming contexts). Given the uncertainty about the quality of LLM‐generated samples, we also developed sample selection algorithms to filter out low‐quality samples (i.e., incomplete or incorrectly labeled samples). Moreover, specific datasets have been constructed to evaluate CARLDA's ability to address the aforementioned challenges. Experimental results demonstrate that (1) CARLDA significantly enhances F1 by 11.0% and the Matthews correlation coefficient (MCC) by 10.0% compared to state‐of‐the‐art methods, showing superior overall performance and effectively tackling the first two challenges, and (2) LLM‐based data augmentation techniques successfully yield high‐quality labeled data and effectively alleviate the third challenge. Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Tanghaoran Zhang, Yao Lu 0003 |
J. Softw. Evol. Process. | 2 |
| 2024 | An Empirical Study of Cross-Project Pull Request Recommendation in GitHubabstractAs a core contribution merge mechanism in distributed collaborative development, pull requests contain valuable knowledge of code evolution and issue resolution. With the co-evolution of multiple projects in a software ecosystem, relevant and similar issues can arise across different projects. Leveraging existing solutions in pull requests (PRs) through cross-project pull request recommendation (CPR) can enrich context knowledge and improve the efficiency of issue resolution. However, the characteristics of CPR and its effectiveness in the process of issue resolution still remain unclear. To bridge this gap, we conduct an empirical study of the CPR on GitHub. We first extract 4,445 CPRs from 2,500 open source projects and quantitatively analyze the characteristics of CPR. Then we conduct a qualitative analysis of sampled CPR cases to understand the influence of CPR. We also use a regression model to explore the impact of CPRs on issue resolution. Our main findings are as follows: (1) Experienced contributors in target projects make most of the CPRs and their CPRs are more timely than inexperienced contributors; (2) In CPR dataset, bugs constitute the largest proportion of target issue types, followed by enhancements, features and questions; (3) Nearly half of the CPRs are accepted by issue participants; (4) A greater number of the CPRs contribute indirectly to solving the target issue by offering solutions and contextual information, rather than providing appropriate code that can be directly applied to the issue; (5) Most of CPR-related factors have a significant impact on issue resolution delay. Among these, recommendation latency has the most significant impact, followed by the type of recommender. Our work has important insights into CPR and offers important guidance for developers on recommending cross-project PRs to resolve the mushrooming issues. Wenyu Xu, Yao Lu 0003, Xunhui Zhang, Tanghaoran Zhang, Bo Lin 0011, Xinjun Mao |
APSEC | 6 |
| 2024 | CAREER: Context-Aware API Recognition with Data Augmentation for API Knowledge ExtractionabstractThe recognition of Application Programming Interface (API) mentions in the software-related texts is a prerequisite task for extracting API-related knowledge. Previous studies have demonstrated the superiority of deep learning-based methods in accomplishing this task. However, such techniques still meet their bottlenecks due to their inability to effectively handle the following three challenges: (1) differentiating APIs from common words; (2) identifying APIs in morphological variants of the standard APIs; and (3) the lack of high-quality labeled data for training. To overcome these challenges, this paper proposes a context-aware API recognition method named CAREER. This approach utilizes two key components, namely Bidirectional Encoder Representations from Transformers (BERT) and Bi-directional Long Short-Term Memory (BiLSTM), to extract context information at both the word-level and sequence-level. This strategic combination empowers the method to dynamically capture both syntactic and semantic information, effectively addressing the first challenge. To tackle the second challenge, CAREER introduces a character-level BiLSTM component, enriched with an attention mechanism. This enables the model to grasp character-level global context information, thereby enhancing the recognition of morphological attributes within API mentions. Furthermore, to address the third challenge, the paper introduces three data augmentation techniques aimed at generating new data samples. Accompanying these techniques is a novel sample selection algorithm designed to screen out high-quality instances. This dual-pronged approach effectively mitigates the requirement for data labeling. Experiments demonstrate that CAREER significantly improves F1-score by 11.0% compared with state-of-the-art methods. We also construct specific datasets to assess CAREER's capacity to tackle the aforementioned challenges. Results confirm that (1) CAREER significantly outperforms baseline methods in addressing the first and second challenges, and (2) with the aid of data augmentation techniques and sample selection algorithms, high-quality samples can be generated to improve the performance, and alleviate the third challenge. Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Yao Lu 0003 |
ICPC | 2 |
| 2024 | Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language TasksabstractContinual learning (CL) empowers pre-trained vision-language (VL) models to efficiently adapt to a sequence of downstream tasks. However, these models often encounter challenges in retaining previously acquired skills due to parameter shifts and limited access to historical data. In response, recent efforts focus on devising specific frameworks and various replay strategies, striving for a typical learning-forgetting trade-off. Surprisingly, both our empirical research and theoretical analysis demonstrate that the stability of the model in consecutive zero-shot predictions serves as a reliable indicator of its anti-forgetting capabilities for previously learned tasks.
Motivated by these insights, we develop a novel replay-free CL method named ZAF (Zero-shot Antidote to Forgetting), which preserves acquired knowledge through a zero-shot stability regularization applied to wild data in a plug-and-play manner. To enhance efficiency in adapting to new tasks and seamlessly access historical models, we introduce a parameter-efficient EMA-LoRA neural architecture based on the Exponential Moving Average (EMA). ZAF utilizes new data for low-rank adaptation (LoRA), complemented by a zero-shot antidote on wild data, effectively decoupling learning from forgetting. Our extensive experiments demonstrate ZAF's superior performance and robustness in pre-trained models across various continual VL concept learning tasks, achieving leads of up to 3.70\%, 4.82\%, and 4.38\%, along with at least a 10x acceleration in training speed on three benchmarks, respectively. Additionally, our zero-shot antidote significantly reduces forgetting in existing models by at least 6.37\%. Our code is available at https://github.com/Zi-Jian-Gao/Stabilizing-Zero-Shot-Prediction-ZAF. Zijian Gao, Xingxing Zhang 0001, Kele Xu, Xinjun Mao, Huaimin Wang 0001 |
NeurIPS | 4 |
| 2024 | Unveiling the Dynamics of Extrinsic Motivations in Shaping Future Experts' Contributions to Developer Q&A CommunitiesabstractDeveloper question and answering communities rely on experts to provide helpful answers. However, these communities face a shortage of experts. To cultivate more experts, the community needs to quantify and analyze the rules of the influence of extrinsic motivations on the ongoing contributions of those developers who can become experts in the future (potential experts). Currently, there is a lack of potential expert‐centred research on community incentives. To address this gap, we propose a motivational impact model with self‐determination theory‐based hypotheses to explore the impact of five extrinsic motivations (badge, status, learning, reputation, and reciprocity) for potential experts. We develop a status‐based timeline partitioning method to count information on the sustained contributions of potential experts from Stack Overflow data and propose a multifactor assessment model to examine the motivational impact model to determine the relationship between potential experts’ extrinsic motivations and sustained contributions. Our results show that (i) badge and reciprocity promote the continuous contributions of potential experts while reputation and status reduce their contributions; (ii) status significantly affects the impact of reciprocity on potential experts’ contributions; (iii) the difference in the influence of extrinsic motivations on potential experts and active developers lies in the influence of reputation, learning, and status and its moderating effect. Based on these findings, we recommend that community managers identify potential experts early and optimize reputation and status incentives to incubate more experts. Yi Yang 0004, Xinjun Mao, Menghan Wu 0001 |
IET Softw. | 2 |
| 2024 | Less confidence, less forgetting: Learning with a humbler teacher in exemplar-free Class-Incremental learning
Zijian Gao, Kele Xu, Huiping Zhuang, Li Liu 0036, Xinjun Mao, Bo Ding 0001, Huaimin Wang 0001 |
Neural Networks | 5 |
| 2024 | How Do Developers Adapt Code Snippets to Their Contexts? An Empirical Study of Context-Based Code Snippet AdaptationsabstractReusing code snippets from online programming Q&A communities has become a common development practice, in which developers often need to adapt code snippets to their code contexts to satisfy their own programming needs. However, how developers make these code adaptations based on contexts is still unclear. To bridge this gap, we first conduct a semi-structured interview of 21 developers to investigate their adaptation practices and perceived challenges during this process. The result suggests that code snippet adaptation is a challenging and exhausting task for developers, as they should tailor the snippets to guarantee their correctness and quality with laborious work. We also note that developers all resort to their intra-file context to complete adaptations, which motivates us to further study how developers performed context-based adaptations (CAs) in real scenarios. To this end, we conduct a quantitative study on an adaptation dataset comprising 300 code snippet reuse cases with 1,384 adaptations from Stack Overflow to GitHub. For each adaptation, we manually annotate its intention and relationship with the context. Based on our annotated data, we employ frequent itemset mining to obtain four CA patterns from our dataset, includingFortification,Code Wiring,Attribute-izationandParameterization. Our main findings reveal that: (1) more than half of the code snippet reuse cases include CAs and 23.3% of the adaptations are CAs; (2) more than half of the CAs are corrective adaptations and variable is the primary adapted language construct; (3) attribute is the most frequently utilized context and 88% of the local contexts are within the nearest 10 LOCs; and (4) CAs towards different intentions are repetitive, which are useful for automatic adaptation. Overall, our study provides valuable insights into code snippet adaptation and has important implications for research, practice, and tool design. Tanghaoran Zhang, Yao Lu 0003, Yue Yu 0001, Xinjun Mao, Yang Zhang 0026 |
IEEE Trans. Software Eng. | 4 |
| 2023 | Understanding Developers' Contribution Motivation in Stack Overflow: A Systematic ReviewabstractMotivations are vital for developers to choose and remain in developer question and answering (Q&A) communities to maintain the prosperity of these communities. Hence, they have been investigated by many researchers in the past decade. However, a systematic research overview has yet to be performed in this domain, leaving a fragmented picture of developers' motivations for participating in these communities. To fill this gap, this work (1) groups developer motivations from multiple data sources by a deep text clustering method, (2) statistically identifies determining factors of developers, and (3) analyzes the regulation styles of the motivation determinants based on the self-determination theory. Taking Stack Overflow as the research context, this study collected 134 motivations from 41 related research. With the motivations collected, this study took them into seventeen clusters. The centres of eight were considered the determining factors: reputation, badge, getting information, learning, reciprocity, helping others, obtaining privileges, and fun. These motivational determinants fall under five regulation styles, from intrinsic to extrinsic. The findings presented a systematic picture of developers' motivations for participating in Stack Overflow and can help researcher and practitioners optimize strategies to stimulate developers to build sustainable online communities. Yi Yang 0004, Xinjun Mao |
APSEC | 2 |
| 2023 | MUSE: A Multi-Feature Semantic Fusion Method for ROS Node Search Based on Knowledge GraphabstractReusing ROS components, specifically ROS Nodes, is crucial for improving the efficiency and quality of robotic software development. However, developers face challenges in finding the desired ROS Nodes for reuse due to scattered organization of ROS Nodes and the ambiguity in their ROS Node name. To address these challenges, this paper proposes a MUlti-feature SEmantic fusion method (MUSE) that leverages a domain-specific ROS knowledge graph for searching ROS Nodes. Firstly, a large dataset is constructed, comprising code files and textual descriptions related to ROS Nodes obtained from GitHub and ROS Wiki. Secondly, an in-depth analysis of user queries regarding the reuse of ROS Nodes is conducted, leading to the selection of multiple features that provide a comprehensive representation of ROS Node semantics, including Function, Hardware, Input, and Output. Subsequently, a knowledge graph of ROS Nodes is developed based on the dataset, incorporating the selected features. This knowledge graph effectively organizes scattered knowledge and resolves the issue of diverse mentions through entity disambiguation and resolution. To eliminate the semantic gap between the descriptions of features mentioned in user queries and the entities in the knowledge graph, a pretrained transformer-based model was used to measure the multi-feature semantic similarity between user queries and ROS Nodes knowledge. Finally, we employ a linear regression model to integrate the multi-feature knowledge between user queries and ROS Nodes knowledge. The proposed method has shown a 20% improvement in performance on NDCG@1 compared to other ROS Node search methods. Further evaluations highlight the effectiveness of each feature incorporated in the knowledge graph, as well as the significance of each parameter within the regression model. These findings underscore the robustness of this research in optimizing the reuse of ROS Nodes and facilitating the development of robotics software. Xinjun Mao, Tanghaoran Zhang, Zhang Zhang 0005 |
APSEC | 2 |
| 2023 | Complementary Learning System Based Intrinsic Reward in Reinforcement LearningabstractDeep reinforcement learning has achieved encouraging performance in many realms. However, one of its primary challenges is the sparsity of extrinsic rewards, which is still far from solved. Complementary learning system theory suggests that effective human learning relies on two complementary learning systems utilizing short-term and long-term memories. Inspired by the fact that humans evaluate curiosity by comparing current observations with historical information, we propose a novel intrinsic reward, namely CLS-IR, which aims to address the problems caused by sparse extrinsic rewards. Specifically, we train a self-supervised predictive model with short-term and long-term memories via exponential moving averages. We employ the information gain between the two memories as the intrinsic reward, which does not incur additional training costs but leads to better exploration. To investigate the effectiveness of CLS-IR, we conduct extensive experimental evaluations; the results demonstrate that CLS-IR can achieve state-of-the-art performance on Atari games and DeepMind Control Suite. Zijian Gao, Kele Xu, Hongda Jia, Tianjiao Wan, Bo Ding 0001, Xinjun Mao, Huaimin Wang 0001 |
ICASSP | 7 |
| 2023 | An Integrated Approach to Predicting the Influence of Reputation Mechanisms on Q&A Communities
Yi Yang 0004, Xinjun Mao, Menghan Wu 0001 |
ICCBR | 2 |
| 2023 | An Extensive Study of the Structure Features in Transformer-based Code Semantic SummarizationabstractTransformers are now widely utilized in code intelligence tasks. To better fit highly structured source code, various structure information is passed into Transformer, such as positional encoding and abstract syntax tree (AST) based structures. However, it is still not clear how these structural features affect code intelligence tasks, such as code summarization. Addressing this problem is of vital importance for designing Transformer-based code models. Existing works are keen to introduce various structural information into Transformers while lacking persuasive analysis to reveal their contributions and interaction effects. In this paper, we conduct an empirical study of frequently-used code structure features for code representation, including two types of position encoding features and AST-based structure features. We propose a couple of probing tasks to detect how these structure features perform in Transformer and conduct comprehensive ablation studies to investigate how these structural features affect code semantic summarization tasks. To further validate the effectiveness of code structure features in code summarization tasks, we assess Transformer models equipped with these code structure features on a structural dependent summarization dataset. Our experimental results reveal several findings that may inspire future study: (1) there is a conflict between the influence of the absolute positional embeddings and relative positional embeddings in Transformer; (2) AST-based code structure features and relative position encoding features show a strong correlation and much contribution overlap for code semantic summarization tasks indeed exists between them; (3) Transformer models still have space for further improvement in explicitly understanding code structure information. Kang Yang 0001, Xinjun Mao, Shangwen Wang, Yihao Qin, Tanghaoran Zhang, Yao Lu 0003, Kamal Al-Sabahi |
ICPC | 2 |
| 2023 | An Effective Method for Constructing Knowledge Graph to Search Reusable ROS Nodes (S)abstractDeveloping robot software is difficult for most software engineers as it requires multi-discipline knowledge such as robotics, AI, and software engineering.Robot Operating Systems (ROS) provides a software development framework and lots of reusable ROS Nodes that encapsulate various robotics functions, which can simplify robot software development in terms of software reuse.However, searching and reusing required ROS Nodes from thousands of ROS Nodes is still challenging due to the scattered distribution of ROS Node information and the need for adequate search methods.In this paper, we present an effective method to construct a ROS Node knowledge graph in support of searching and reusing ROS Nodes.Our method uses multiple data sources, including open-source ROS software in Github and ROS wiki community.We extract two-tuple functional information and task-related noun phrases from the ROS Node description and ROS communication interactions from the ROS Node source code.The constructed ROS Node knowledge graph (RNKG) contains 14,065 entities and 15,767 relations.It provides rich semantic information to comprehensively and precisely describe ROS Nodes, their services, and related interaction topics and messages. Xinjun Mao, Sun Bo, Tanghaoran Zhang, Shuo Yang 0005 |
SEKE | 2 |
| 2022 | Towards An Efficient Searching Approach of ROS Message by Knowledge GraphabstractThe Robot Operating System (ROS) has become the most popular robot development framework in the last few years, which has loosely coupled structure and provides remote communications between different component nodes. The ROS messages are critical to bridge the communication channels and clearly define the data structures. The developers can use the standardized or user-customized ROS message types to construct a communication channel between two component nodes uniquely. However, it becomes increasingly difficult for developers to find the required ROS message type from thousands of diverse ROS message types in ROS-based robotic software development. Finding the proper ROS message type is a non-trivial task because developers may hardly know the exact names of required ROS messages but only has a rough knowledge of the task domain features. To tackle this challenge, we construct a novel ROS Message Knowledge Graph (RMKG) with 4543 entities and 14320 relationships, including all ROS message types and message packages. We take the shortest path algorithm to search ROS message in RMKG by searching with ROS message feature or ROS message package and visualize the subgraph structure of the search results. Moreover, we develop a ROS message package library that supports fuzzy queries to find the required message package. A comprehensive evaluation of RMKG shows the high accuracy of our knowledge construction approach. A user study indicates that RMKG is promising in helping developers find suitable ROS message types for robotics software development tasks. An effect evaluation of message package fuzzy query shows the good effects of our fuzzy query method under different situations. Sun Bo, Xinjun Mao, Shuo Yang 0005 |
COMPSAC | 2 |
| 2022 | The Maintenance of Top Frameworks and Libraries Hosted on GitHub: An Empirical StudyabstractThe number of repositories on GitHub is huge and growing rapidly.However, most repositories are inactive, while active maintenance is essential in choosing a project.In this paper, we study the maintenance of top (i.e., most starred) frameworks and libraries hosted on GitHub, for they can be widely reused and in critical positions in the dependency network, so their maintenance status is significant.Furthermore, their maintenance practices may inspire other projects to thrive on the collaborative development platform.By investigating their adoption of recommended Open Source Software (OSS) maintenance practices and recent maintenance activities, and the association between maintenance status and usage, we find that: (1) more than 20% of the top frameworks and libraries have no commit for more than one year; (2) Some maintenance practices (e.g., codes of conduct) have relatively low adoption rates, while continuous integration has a high adoption rate of around 80%; (3) the maintenance status may have an effect on the usage frequency. Xinjun Mao, Zhang Zhang 0005 |
SEKE | 2 |
| 2022 | On the Way to Microservices: Exploring Problems and Solutions from Online Q&A CommunityabstractMicroservice architecture is a dominant architectural style in SaaS industry, which helps to develop a single application as a collection of independent, well-defined, and inter-communicating services. The number of microservice-related questions in Q&Awebsites, such as Stack Overflow, has expanded substantially over the last years. Due to its increasing popularity, it is essential to understand the existing problems that microservice developers face in practices as well as the potential solutions to these problems. Such an investigation of problems and solutions is vital for long-term, impactful, and qualified research and practices in microservice community. Unfortunately, we currently know relatively little about such knowledge. To fill this gap, we conduct a large-scale in-depth empirical study on 17,522 Stack Overflow microservice-related posts. Our analysis leads to the first taxonomy of microservice-related topics based on the software development process. By analyzing the characteristics of the accepted answers, we find that there are fewer experts in the microservice than other domains, and such a phenomenon is most significant with respect to the microservice design phase. Furthermore, we perform manual analysis on 6,013 answers accepted by developers and distill 47 general solution strategies for different microservice-related problems, 22 of which are proposed for the first time. For instance, several problems inherent in the delivery phase can be lessened by referring to external sources like GitHub code examples. Our findings can therefore facilitate research and development on emerging microservice systems. Menghan Wu 0001, Yang Zhang 0026, Shangwen Wang, Zhang Zhang 0005, Xin Xia 0001, Xinjun Mao |
SANER | 7 |
| 2022 | Recommending Base Image for Docker Containers based on Deep Configuration ComprehensionabstractDocker containers are being widely used in large-scale industrial environments. In practice, developers must manually specify the base image in the dockerfile in the process of container creation. However, finding the proper base image is a nontrivial task because manually searching is time-consuming and easily leads to the use of unsuitable base images, especially for newcomers. There is still a lack of automatic approaches for recommending related base image for developers through dockerfile configuration. To tackle this problem, this paper makes the first attempt to propose a neural network approach named DCCimagerec which is based on deep configuration comprehension. It aims to use the structural configuration features of dockerfile extracted by AST and path-attention model to recommend potentially suitable base image. The evaluation experiments based on about 83,000 dockerfiles show that DCCimagerec outperforms multiple baselines, improving Precision by 7.5%-67.5%, Recall by 6.2%-106.6%, and F1 by 7.5%-150.2%. Yinyuan Zhang, Yang Zhang 0026, Xinjun Mao, Yiwen Wu 0001, Bo Lin 0011, Shangwen Wang |
SANER | 3 |
| 2022 | Towards a behavior tree-based robotic software architecture with adjoint observation schemes for robotic software development
Shuo Yang 0005, Xinjun Mao, Yao Lu 0003 |
Autom. Softw. Eng. | 2 |
| 2022 | HAF: a hybrid annotation framework based on expert knowledge and learning technique
Yue Yu 0001, Tao Wang 0006, Gang Yin, Xinjun Mao, Huaimin Wang 0001 |
Sci. China Inf. Sci. | 5 |
| 2022 | FENSE: A feature-based ensemble modeling approach to cross-project just-in-time defect prediction
Tanghaoran Zhang, Yue Yu 0001, Xinjun Mao, Yao Lu 0003, Huaimin Wang 0001 |
Empir. Softw. Eng. | 3 |
| 2022 | ForkXplorer: an approach of fork summary generation
Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003 |
Frontiers Comput. Sci. | 2 |
| 2022 | A Concentrated Time-Frequency Method for Reservoir Detection Using Adaptive Synchrosqueezing TransformabstractThe synchrosqueezing transform (SST) is an effective technique to concentrate the time-frequency (TF) energy and to retrieve the components of a non-stationary multicomponent signal. Therefore, it has been widely used to process and interpret seismic data. However, due to the fixed window width of the short-time Fourier transform (STFT), the STFT-based SST (FSST) is not well suitable for the characterization of the oil and gas reservoirs with varying layer thickness. Here, an adaptive SST is employed to characterize the features of the seismic signals for identifying the oil reservoirs with varying thickness. Overall, the proposed method is based on STFT with time-varying windows. For the local harmonic wave approximation and well-separated condition of a non-stationary multicomponent signal, the adaptive STFT is windowed with the Gaussian function. Compared with the conventional FSST, the adaptive FSST (AFSST) provides the TF concentration with better concentration and separates the components more accurately. In this work, both synthetic model and field seismic data are applied to validate the AFSST method, demonstrating that the AFSST method can precisely distinguish the stratigraphic characteristics. Xinjun Mao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Motivation Under Gamification: An Empirical Study of Developers' Motivations and Contributions in Stack OverflowabstractTo encourage developers' volunteer contributions, modern programming question and answer (Q&A) sites like Stack Overflow (SO) employ gamified incentive mechanisms such as reputation and badges. Understanding developers' motivations in the presence of gamification and the relationship between their motivations and behavioral outcomes is crucial for community building and designing good incentive mechanisms. Grounded on self-determination theory, we conducted a survey with 938 developers who participate in SO to understand their participation motivations and incentive perceptions. By connecting the survey responses with the SO data, we quantitatively analyzed how the developers' motivations and satisfaction of needs relate to their effort and contribution quality. Our main findings are as follows: (1) despite the presence of gamified incentive mechanisms, developers are mainly motivated by intrinsic motivation to participate in SO; (2) developers who have strong motivations to gain gamification rewards are associated with higher intrinsic and integrated motivations, while developers with more development experiences are less motivated by the gamified incentives; (3) both extrinsic motivations (in terms of career prospects) and intrinsic motivations (regarding self-improvement and helping others) can motivate developers to make high-quantity and high-quality contributions; and (4) high-level satisfaction of needs for competency and autonomy has a positive effect on developers making high-quantity and high-quality contributions and addressing difficult problems. Based on these findings, we discuss implications for developer motivation and gamification in the crowdsourcing context and for the mechanism design of gamified crowdsourced platforms. Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Zude Li, Tao Wang 0006, Gang Yin, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 2 |
| 2021 | Towards Adjoint Sensing and Acting Schemes and Interleaving Task Planning for Robust Robot PlanabstractRobots operating in open environments expect to have robust plans to achieve tasks successfully under environment uncertainties. However, both partial observability and dynamics of environment states have significantly decreased the robustness of task achievement, making robot task planning much more challenging. The partially observable states require the robot to obtain observations for optimally acting of the task goal. Also, state dynamics expects the robot to continuously observe surroundings for acting safely. Both challenges practically demand the purposeful and tight interactions between robot state-changing actuating actions and sensor-based observation actions. This paper proposes a novel model of Adjoint Sensing and Acting (ASA) that explicitly defines two parallel and sequential interaction schemes between actuating and observation actions, as well as an extended Behavior Tree for a concrete implementation of above schemes. We further propose an interleaving task planning approach for planning ASA-style plans, which integrates a deliberative POMDP planner for pursuing task goals, and a reactive Behavior Tree executive for fast responding to unexpected events. We experimentally demonstrate that ASA interaction schemes are practical and applicable to model and plan the open environment robot tasks. The plans from the interleaving task planning approach are both reactive in run-time response and efficient in task achievement. Shuo Yang 0005, Xinjun Mao, Huaiyu Xiao, Yuanzhou Xue |
ICRA | 2 |
| 2021 | An Efficient ROS Package Searching Approach Powered By Knowledge GraphabstractOver the past several years, the Robot Operating System (ROS), has grown from a small research project into the most popular framework for robotics development.It offers a core set of software for operating robots that can be extended by creating or using existing packages, making it possible to program robotic software that can be reused on different hardware platforms.With thousands of packages available per stable distribution, encapsulating algorithms, sensor drivers, etc., it is the de facto middleware for robotics.However, finding the proper ROS package is a nontrivial task because ROS packages involve different functions and even with the same function, there are different ROS packages for different tasks.So it is timeconsuming for developers to find suitable ROS packages for given task, especially for newcomers.To tackle this challenge, we build a ROS package knowledge graph, ROSKG, including the basic information of ROS packages and ROS package characteristics extracted from text descriptions, to comprehensively and precisely characterize ROS packages.Based on ROSKG, we support ROS packages search with specific task description or attributes as input.A comprehensive evaluation of ROSKG shows the high accuracy of our knowledge construction approach.A user study shows that ROSKG is promising in helping developers find suitable ROS packages for robotics software development tasks. Xinjun Mao, Yinyuan Zhang, Shuo Yang 0005 |
SEKE | 2 |
| 2021 | Detecting Duplicate Contributions in Pull-Based Model Combining Textual and Change Similarities
Yue Yu 0001, Tao Wang 0006, Gang Yin, Xinjun Mao, Huaimin Wang 0001 |
J. Comput. Sci. Technol. | 5 |
| 2020 | Software Engineering for Autonomous Robot: Challenges, Progresses and OpportunitiesabstractSoftware engineering for autonomous robot (SE4AR) is an emerging interdisciplinary field with the aims to support the development, running and evolution of software systems for autonomous robot - a safety and mission critical cyber-physical system. Such field has recently gained increasing attentions from both academia and industry and made rapid progresses in the past years. However, many challenges pose on the SE4AR due to the specific features and complexities of autonomous robot software (ARS) and several open problems should be tackled in the future researches. This paper aims to present a comprehensive investigation on the researches and practices of SE4AR. Our contributions are three-fold as follows: an in-depth analysis on the development challenges, a systematic review on the current progresses, and an open discussion of weaknesses in current researches and opportunities in future researches. Xinjun Mao |
APSEC | 1 |
| 2020 | An Empirical Study on the Influence of Social Interactions for the Acceptance of Answers in Stack OverflowabstractIn knowledge-sharing communities like Stack Overflow (SO), users can post questions, give answers and choose one answer as an accepted answer. The accepted answers will be important references for users when they encounter similar questions. Essentially, posting questions and giving answers is an interactive process occurring among community users, and choosing accepted answers is actually a decision-making process involving multiple factors. Previous works examined the impact on this decision process from the user, question and answer viewpoints. Social interactions between the questioners and answerers, although being popular according to our pre-analysis, have never been considered as a factor that can influence the decisions. To fill this gap, this paper first proposes a comprehensive answer acceptance model that integrates the answer features established by social interactions as well as information of users, questions and answers. We then divide social interactions into two stages and propose a method to calculate the relationship between the questioner and the answerer by analyzing these social interactions. Finally, we investigate the influence of social interactions for the acceptance of answers by performing logistic regression analysis. The results reveal several findings: (1) social-based features explain 16.6 % of the variance explained together, indicating that social interactions have significant and important effects on the acceptance of answers; (2) social interactions that occur after the answer is posted are more influential than these occur before the answer is posted. Based on the findings, we further conduct an online study of 132 SO users, and the respondents report that social interactions have a greater impact on the acceptance of answers than other judgments of answers such as upvotes, downvotes and not accepting answers. Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003, Shangwen Wang, Jinyu Lu |
APSEC | 2 |
| 2020 | Gathering GitHub OSS Requirements from Q&A Community: an Empirical StudyabstractCross-community collaboration can exploit the expertise and knowledges of crowds in different communities. Recently increasing users in open source software (OSS) community like GitHub attempt to gather software requirements from question and answer (Q&A) communities such as Stack Overflow (SO). In order to investigate this emerging cross-community collaboration phenomenon, the paper presents an exploratory study on cross-community requirements gathering of OSS projects in GitHub. We manually sample 3266 practice cases and quantitatively analyze the popularity of the phenomenon, the characteristics of the gathered requirements, and cross-community collaboration behaviors of users. Some important findings are obtained: more than half of the requirements gathered from SO are enhancements and the majority of the gathered requirements are non-functional requirements. In addition, OSS developers can directly obtain related solutions and contributions of the gathered requirements from SO in the gathering process. Yao Lu 0003, Xinjun Mao |
ICECCS | 3 |
| 2020 | Haste Makes Waste: An Empirical Study of Fast Answers in Stack OverflowabstractModern programming question & answer (Q&A) sites such as Stack Overflow (SO) employ gamified mechanisms to stimulate volunteers' contributions. To maximize the chances of winning gamification rewards such as reputation and badges, a portion of users race to post answers as quickly as possible (i.e., fast answers or FAs), which makes SO the fastest Q&A site; however, this behavior may affect the contribution quality as well. In this paper, we report on a large-scale, mixed-methods empirical study of the gamification-influenced FA phenomenon in SO. We first quantitatively investigate the popularity of the phenomenon and user behaviors regarding FAs. Then, we study the quality of FAs by using regression modeling and qualitatively analyzing 300 instances of FAs. Our main findings reveal that more than 70% and 90% of FAs are not edited by the answerers and other users, respectively, and that later incoming answers have lower chances of being voted on and accepted. Notably, we find that the answer length, code snippets length, and readability of FAs are significantly lower than those of non-fast answers. Although FAs have higher crowd assessment scores, they have no relationship with acceptance from the perspective of asker assessment, and a considerable portion of FAs solve the problem by interacting with the asker in the comments. These results help us better understand the effects of reward-based gamification on crowdsourced software engineering communitites and provide implications for designers of gamified systems. Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Tao Wang 0006, Zude Li |
ICSME | 2 |
| 2020 | Exploring the Dependency Network of Docker Containers: Structure, Diversity, and RelationshipabstractContainer technologies are being widely used in large scale production cloud environments, of which Docker has become the de-facto industry standard. As a key step, containers need to define their dependent base image, which makes complex dependencies exist in a large number of containers. Prior studies have shown that references between software packages could form technical dependencies, thus forming a dependency network. However, little is known about the details of docker container dependency networks. In this paper, we perform an empirical study on the dependency network of docker containers from more than 120,000 dockerfiles. We construct the container dependency network and analyze its network structure. Further, we focus on the Top-100 dominant containers and investigate their subnetworks, including diversity and relationships. Our findings help to characterize and understand the container dependencies in the docker community and motivate the need for developing container dependency management tools. Yinyuan Zhang, Yang Zhang 0026, Yiwen Wu 0001, Yao Lu 0003, Tao Wang 0006, Xinjun Mao |
Internetware | 6 |
| 2020 | Who Should Close the Questions: Recommending Voters for Closing Questions Based on Tags
Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003, Jinyu Lu |
SEKE | 2 |
| 2020 | Exploring CQA User Contributions and Their Influence on Answer Distribution
Yi Yang 0004, Xinjun Mao, Zixi Xu, Yao Lu 0003 |
SEKE | 2 |
| 2020 | Towards an Extended POMDP Planning Approach with Adjoint Action Model for Robotic TaskabstractIn real-world environments, robotic task planning is expected to handle both partial observability and unexpected dynamics of the environment. A robust plan for the task requires the robot's observation actions to concurrently run with the task actions, to observe and adapt to environmental changes. The Partially Observable Markov Decision Process (POMDP) has been widely applied for planning under partially observable domains. For realistic robotic tasks, however, the POMDP model and planning algorithm are quite restrictive and unrealistic. One limitation is that task actions are modelled as atomic entities that only have endpoint effects, with no conditions specified at arbitrary points during task action execution. Also, the observation is obtained only after each task action execution, with no intermediate observations and decision-making during task action execution. To mitigate the limitations of POMDP planning, this paper first proposes an Adjoint Action Model (AAM) that explicitly defines the continuous interaction between robot's observation and task actions. Then we extend the POMDP task action model with intermediate invariant conditions which specifies the runtime properties of action execution. Finally, we propose the AAM-extended POMDP planning approach which handles observation action planning and task replanning for task action execution. We experimentally demonstrate that the plan from our proposed approach is more effective and robust to cope with the environment dynamics, comparing with the standard POMDP planning approach. Shuo Yang 0005, Xinjun Mao, Wanwei Liu |
SMC | 2 |
| 2020 | Improving students' programming quality with the continuous inspection process: a social coding perspective
Yao Lu 0003, Xinjun Mao, Tao Wang 0006, Gang Yin, Zude Li |
Frontiers Comput. Sci. | 2 |
| 2020 | Automatic Voter Recommendation Method for Closing Questions in Stack OverflowabstractStack Overflow is the most popular programming question and answer community that continuously receives a large number of questions every day. To ensure the quality of questions, the community grants privileges for the moderators and a group of experienced users to review the quality of questions and close the low-quality ones (e.g. duplicate or irrelevant questions). The review process is a typical crowdsourcing job that relies on users’ volunteer participation, and the current practices of closing questions in Stack Overflow face two aspects of challenges: (1) an obvious increase in both the absolute number and the percentage of “closed” questions; (2) a considerable decrease in participation willingness of experienced users to close questions. In order to solve the problem, we present a novel model of user willingness for reviewing and voting questions by incorporating four types of user activity history, including questions, answers, comments and votes of closing questions. Then we propose an automatic recommendation method based on the model to assign experienced users proper questions, to utilize the forces of them to close questions. The evaluation shows that the successful recommendation probability in the top 5, top 10, top 20, top 30, top 40, top 50 users are 48.23%, 58.93%, 68.83%, 74.27%, 78.13% and 81%, respectively. Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003, Jinyu Lu, Yue Yu 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2020 | A trajectory privacy-preserving scheme based on a dual-K mechanism for continuous location-based services
Shaobo Zhang 0001, Xinjun Mao, Kim-Kwang Raymond Choo, Tao Peng 0011, Guojun Wang 0001 |
Inf. Sci. | 2 |
| 2019 | A field-based service management and discovery method in multiple clouds context
Shuai Zhang 0048, Xinjun Mao, Fu Hou, Peini Liu |
Frontiers Comput. Sci. | 2 |
| 2018 | Accompanying Observation Modes and Software Architecture for Autonomous Robot SoftwareabstractTo support robust task execution in open environment, autonomous robots (AR) should include reactive capabilities to cope with the dynamics and uncertainties from real-world environment.The uncertainties pose great challenges for robots being sensitive to the environmental changes and flexible to adjust self-behaviors.To this end, this paper aims to improve the sensing and acting capabilities of autonomous robots by novel behavioral theories, observation modes and software architectures.Specifically, this paper has three main contribution: (1) presents an accompanying model that specifies a novel accompanying pattern for interacting robot behaviors;(2) proposes four types of accompanying observation modes that coordinate multiple robot sensing behaviors; (3) proposes a concrete multi-agent software architecture that implements aforementioned accompanying model and accompanying observation modes.To demonstrate the applicability and validity of our accompanying modes and MAS-based software architecture, this paper conducts a case study to implement a domestic service example, which requires the robot to run in a highly dynamic environment and can adapt its behaviors to unexpected situations. Xinjun Mao, Shuo Yang 0005 |
SEKE | 2 |
| 2018 | Towards Reference Architecture for a Multi-layer Controlled Self-adaptive Microservice SystemabstractWith the features of high distribution in deployment and independence in running, the microservice systems that operate in heterogeneous infrastructures and open Internet environment are expected to be self-adaptive to adapt to various changes of both operating contexts and application requirements.This requires the adaptability of the microservice systems to be diverse and flexible, and independent of implementation technologies and platforms.This paper presents a reference architecture for self-adaptive microservice systems with the abilities of multi-layer controlled self-adaptations, including infrastructure-controlled layer and application-controlled layer.Such reference architecture presents a blueprint to cope with diverse changes from different levels in microservice systems and supports the interactions between layers.We have implemented a practical platform called SAMSP based on the reference architecture and Kubernetes and evaluated our approach using a sample.The experimental results are promising, and demonstrate the feasibility and effectiveness of our proposed reference architecture. Peini Liu, Xinjun Mao, Shuai Zhang 0048, Fu Hou |
SEKE | 2 |
| 2018 | A Self-Adaptation Framework of Microservice Systems (S)abstractMicroservice has been more and more applied to build software systems in industry field and research in academic field.And software systems are increasingly expected to dynamically self-adapt to accommodate resource variability, changing user needs, and system faults.Compared with traditional software systems, microservice systems have some characteristics, such as the highly self-contained components and the dynamic running instance, which pose challenges to traditional self-adaptation methods.Therefore, it needs to propose corresponding techniques and methods to cope with the characteristics of microservice systems.This paper analyzes the special selfadaptive requirements of microservice systems, proposes a microservice reference model, which describes basic elements and their relationships of microservice systems.Then we present a microservice system self-adaptation framework MSSAF to support the self-adaptation of microservice systems.We illustrate the feasibility and effectiveness of our approach in the context of a microservice system case. Shuai Zhang 0048, Xinjun Mao, Peini Liu, Fu Hou |
SEKE | 2 |
| 2018 | Towards Dynamic Evolution of Runtime Variability Based on Computational ReflectionabstractGiven the frequently changing nature of the user requirements and environments in software systems, runtime variability in today’s software systems should be capable of evolving during execution. Computational reflection is required to facilitate accessing and customizing runtime variability during this evolution process. However, realizing this computational reflection includes various practical complexities since the runtime variability is typically neither explicitly represented in software systems nor changeable during runtime. To address this problem, this paper proposes a software architecture to support computational reflection of runtime variability, along with a corresponding causal-connection mechanism to realize the introspection and intercession (i.e. representing runtime variability model, and adding, removing, replacing variability elements and their relations). The proposed software architecture consists of a meta level that represents runtime variability model using objectification, and a base level that organizes and manipulates the implementation of variability elements via reconfiguration. The causal-connection mechanism integrated in our proposed model is designed to synchronize the representation and the implementation. Further, we developed a Reflective Runtime Variability Framework (R2VF) to support the development and operation of the systems with the reflection of runtime variability. The effectiveness and applicability of our approach has been evaluated by applying R2VF to Personal Data Resource Network. Xinjun Mao |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2018 | Internal quality assurance for external contributions in GitHub: An empirical investigationabstractAbstract For popular open‐source software projects, there are always a large number of worldwide developers who have been glued to making code contributions, while most of these developers play the role of casual contributors because of their very limited code commits. The frequent turnover of such a group of developers and the wide variations in their coding experiences challenge the project management on code and quality. This paper aims to investigate the status quo of internal quality assurance for external contributions in social coding sites. We first conducted a case study of 21 popular GitHub projects to estimate the code quality of the casual contributors. The quantitative results show that the casual contributors introduced greater quantity and severity of code quality issues than the main contributors; the developers who contribute to different projects as main and casual contributors did not perform significantly differently in terms of their code quality. On the basis of these findings, we further conducted a survey of 81 developers on GitHub to understand their practices on internal quality assurance. The qualitative results expose some limitations of present internal quality control for external contributions in GitHub. Finally, we discuss an alternative quality management paradigm: Continuous Inspection for industrial practices. Yao Lu 0003, Xinjun Mao, Zude Li, Yang Zhang 0026, Tao Wang 0006, Gang Yin |
J. Softw. Evol. Process. | 2 |
| 2017 | A Dual-Loop Control Model and Software Framework for Autonomous Robot SoftwareabstractAutonomous robot software is an extremely complex system that drives robots to operate in an open and dynamic environment to accomplish tasks. This requires such software to be designed with effective control models to support continuous and flexible feedback. To this end, this paper presents a dualloop control model called D-SMPA and corresponding software framework that adopts several design principles to support the development and running of autonomous robot. Our proposed model explicitly abstracts the robot behaviors into observation-oriented and task-oriented types, and provides a dual-loop control to coordinate these behaviors. Control mechanisms in D-SMPA are proposed to support a tight integration and accompanying cooperation between observation and task behaviors so that the task accomplishment can get required feedback of sensors. Our approach aims to enrich the capabilities of autonomous robot software by obtaining on-demand feedback from multiple sources, and at the same time to simplify software development by separating complex behaviors of robots. We present a multi-agent software framework called AutoRobot to implement the proposed control model and provide tools to support the development of autonomous robot software. Our case study validates the effectiveness and applicability of our proposed approach and framework. Xinjun Mao, Shuo Yang 0005 |
APSEC | 2 |
| 2017 | The Accompanying Behavior Model and Implementation Architecture of Autonomous Robot SoftwareabstractAutonomous robots are increasingly applied in realworld environments, and expect to execute plans robustly to accomplish the assigned tasks in the presence of dynamics and uncertainties of the changing environment. The robustness of plan execution requires the robot to keep aware of plan execution status and to adapt the plan towards possible execution contingencies. Such requirements pose a great challenge for autonomous robot software in terms of the abstraction model over robot behavior patterns. Conventional abstraction models for robot behaviors generally follow the sense-model-planact and behavior-based paradigms, which show limitations in tight integration with sensory inputs and tracking execution traces of robot plans. This paper proposes an accompanying behavior model that considers robot behaviors as task-oriented and observation-based types with diverse aims, and develops the run-time mechanisms to facilitate collaboration between two types of behaviors. Additionally, we implement the model by the multi-agent approach which develops the robot software as a multi-agent system. To demonstrate the feasibility and applicability of proposed model, we conduct a case study by implementing a typical example of service scenarios, e.g., a robot that autonomously picks up and drops off dishes for remote guests in the open and dynamic environment. Shuo Yang 0005, Xinjun Mao, Jiangtao Xue, Zixi Xu |
APSEC | 2 |
| 2017 | Deadlock Prevention in Rendezvous Generation for On-demand Inter-robot Resource Delivery
Yin Chen 0003, Xinjun Mao, Fu Hou |
ICAART (2) | 2 |
| 2017 | A survey of agent-oriented programming from software engineering perspectiveabstractAgent-oriented programming (AOP) represents a novel programming paradigm that adopts concepts and technologies of multi-agent system to implement software. It has gained great attentions of researchers and practitioners from both artificial intelligence field and software engineering field. Dozens of AOP languages have been proposed in the past two decades. However the acceptance and adoption of AOP in software engineering community remain limited and the current practices of applying AOP do not convince such paradigm has extensively exploited its technical advantages and potentials. The experiences and practices of programming language researches in software engineering field can give us some inspirations to get out of the dilemma. The paper aims at providing a survey of AOP from software engineering perspectives, including its research history and the state-of-the-art of researches on agent-oriented programming concepts and models, languages, CASE tools and running manners. We investigate how current AOP studies satisfy design principles of programming language like simplicity, regularity, maintainability, expressiveness, reliability and efficiency, discuss several challenges of AOP researches and practices, and point out some directions for future studies. Xinjun Mao, Qiuzhen Wang |
Web Intell. | 1 |
| 2016 | Does the Role Matter? An Investigation of the Code Quality of Casual Contributors in GitHubabstractFor popular Open Source Software (OSS) projects there are always a large number of worldwide developers who have been glued to making code contributions, while most of these developers play the role of casual contributors due to their very limited code commits (for fixing defects and enhancing features, casually). The frequent turnover of such group of casual developers and the wide variations among their coding experiences challenge the project management on code and quality.This paper describes a case study which aims to estimate the quality of code made by casual contributors in 21 popular GitHub projects. The results of this case study show that: (1) casual contributors introduced greater quantity and severity of Code Quality Issues (CQIs) than main contributors; (2) developers who contribute in different projects as main and casual contributors didn't perform statistically differently in terms of code quality; (3) casual contributors who have few project stars introduced more CQIs than those who have many. Furthermore, the paper lists the CQI categories which are most frequently introduced by casual contributors in the investigated projects. These findings provide valuable insights into code quality in the OSS context, and can guide OSS developers in improving the quality of the code contributions. Yao Lu 0003, Xinjun Mao, Zude Li, Yang Zhang 0026, Tao Wang 0006, Gang Yin |
APSEC | 2 |
| 2016 | A Field-Based Model for Representing Dynamic and Evolving Features of Cloud ServicesabstractIn cloud computing context services present the dynamic and evolving features, which greatly affects the accuracy and efficiency of service discovery. It is the prerequisite for supporting service discovery to capture and represent these features during the process of services organization and management. This paper proposes a field-based model to describe the dynamic and evolving features of cloud services in both qualitative and quantitative way, which is inspired by the Bohr atom model. The concept of energy level in Bohr model is used to represent the services status and demarcate cloud services, electrons jumping mechanism in Bohr model is used to depict services' dynamic and evolving features and analyze how services status are changed from one energy level to another. The field model of services provides abstractions to classify a set of services according to their status and mechanism to explain their changes and demarcation according to their potential energy variation. The concept of user acceptable services region is proposed to represent the search scope and method of user service discovery request in cloud services field model. The algorithms to generate field model of services and form user acceptable services region are designed, with which field-based service discovery algorithm is proposed. Based on QWS dataset, we conduct a series of experiments to evaluate and validate the effectiveness of field-based service model for organizing and discovering cloud services. Fu Hou, Xinjun Mao |
ICPADS | 2 |
| 2016 | Combining re-allocating and re-scheduling for dynamic multi-robot task allocationabstractMulti-robot systems (MRS) working in open and dynamic environments are expected to deal with uncertain arrival of new tasks and environment changes, by repeatedly adapting the current task allocation and schedule, in order to maintain its performance (e.g., total utility, balance, etc.). This paper presents an adaptive approach to multi-robot task allocation (MRTA), which combines two adaptive measures corresponding to different levels of a MRS: (1) re-allocating at inter-robot level, for balancing task allocation, and improving total utility of the MRS, and (2) re-scheduling at intra-robot level, for maintaining each robot's utility against the influence of both re-allocating and environment changes. Our approach is expected to have significantly higher adaptation power than both re-allocating only and re-scheduling only cases. An experiment is conducted to evaluate our approach's capability of improving balance and total utility of the MRS, under different environment settings and different combinations of re-allocating and re-scheduling. Yin Chen 0003, Xinjun Mao, Fu Hou, Qiuzhen Wang, Shuo Yang 0005 |
SMC | 2 |
| 2016 | Cross-clouds services autonomic management approach based on self-organizing multi-agent technologyabstractSummary With the rapid development and application of cloud computing, there exist plenty of clouds that are distributed on the open Internet, decentralized in the management, evolving with various services providing diverse functionalities and QoS. Moreover, because of the potential correlativity of cloud services and the dynamic of tenants' requirements, these services in multiple clouds are expected to be effectively managed in an autonomic and transparent way so that tenants can obtain continuous and efficient services. To this end, this paper proposes an approach based on self‐organizing multi‐agent system to achieving cross‐clouds services management, including the service provision at the tenant‐end and services aggregation at the cloud‐end. In this approach, the clouds services are managed by a series of autonomous agents that are capable of autonomously accessing managed services. They can interact with each other to obtain macro‐level services aggregation in terms of self‐organization to adapt to the changes of both tenants' requirements and cloud services. The paper details the architecture, mechanisms, and algorithms to implement the aggregation and provision of services in cross‐clouds. We also develop relevant cross‐clouds services management platform called as CCloudMan, with which several experiments based on the public data sets have been conducted and the experimental results show the efficiency and usability of our proposed approach. Copyright © 2016 John Wiley & Sons, Ltd. Fu Hou, Xinjun Mao |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | A Lightweight Social Computing Approach to Emergency Management Policy SelectionabstractIn order to select effective policies for emergency management in a timely manner, this paper proposes an agile and lightweight social computing approach to facilitating policy selection, evaluation, and adjustment relative to emergency management in both quantitative and qualitative ways. The approach consists of three components represented as PZE: 1) (P) emergency management policy selecting; 2) (Z) modeling artificial societies with the zombie-city model (a general and formal artificial society model); and 3) (E) policy evaluation. The formal specification of the zombie-city model and rigorous expressions of scenarios enable rigorous description and formal reasoning of an artificial society. A feedback loop of this approach supports the iterative adjustment of emergency management policies and the creation of more effective policies. This approach is verified by applying it to a case of an infectious disease transmission with quantitative evaluations, qualitative reasoning and analysis, and iterative adjustments. Results indicate effective emergency management policies can be established with the approach in an iterative way. In contrast with existing research, our proposed approach offers the benefits of being simple, general, rapidly adaptive to changes, and low cost. Mingsheng Tang, Haibin Zhu 0001, Xinjun Mao |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2015 | Towards Realisation of Evolvable Runtime Variability in Internet-Based Service Systems via Dynamical Software UpdateabstractToday's Internet-based service systems tend to run in open environments and try to satisfy varying requirements. In the context, the changes of requirements and environments can emerge at any point of their life cycle, and their runtime variability is supposed to be evolvable to deal with the changes. In other words, the number, type or attribute of variability elements (variation points and their associated variants) in the systems is regarded to be changeable, especially at runtime. How to realise evolvable runtime variability in software systems is a challenge in the community of software engineering. Software architecture is supposed to support updating software without interrupting service if runtime variability is to be changed. Besides, the relevant mechanisms have to be provided to achieve dynamic update. Thus, we propose a dynamically reconfigurable reference architecture to construct the systems with evolvable software variability. The resultant system established using our approach consists of variability units and their containers. Variability units embody the variability elements that can be changed through local relay, i.e., starting a new version of variability unit to take place of older ones. Containers are mainly used to carry out the dynamic updating for variability units and realise the functionality invocation among different variability units. They can also be added, removed or replaced without shutting down the entire system. We analyse the requirements and scenarios of evolvable runtime variability in the case of Personal Data Resource Network, and further show the effectiveness and applicability of our approach by constructing the system using our class library and solving the issues proposed in the case. Xinjun Mao |
Internetware | 2 |
| 2015 | An Agent-Based Artificial Society Approach to Analyzing Social PropagationabstractSocial propagation issue that appears in both the real world and the cyberspace recently gains great attentions in several communities. It typically is involved with a great number of autonomous participated individuals and their local interactions, and may arise global emergent behaviors or properties. It is therefore necessary to investigate what emergence may occur and how to influence the emergence in the social propagation. This paper proposes an agent-based artificial society approach to analyzing social propagation issue, which includes AVI (Agent-Virus-Interaction) model for modeling social propagation and the systematic method to investigate the emergence. The AVI model that borrows ideas from multiagent and organization metaphors has three core concepts: agent, virus and interaction, representing the propagation carrier, content, and channel, respectively. Based on the AVI model, a systematic method for analyzing social propagation that consists of five analysis activities including modeling artificial society, programming simulation codes, designing experiments, performing simulations, and analyzing results, is proposed. We study a case of an infectious virus propagation to illustrate the approach and show its usability and effectiveness. Besides, we design and conduct a number of experiments to analyze what/how factors influence the infectious virus propagation. Finally, we discuss some potential extensions to the AVI model for satisfying further requirements of social propagation applications. Mingsheng Tang, Xinjun Mao |
SMC | 2 |
| 2015 | The Roadmap and Challenges of Robot Programming LanguagesabstractGreat attentions have been put on the programming technologies to construct robot software in both academic research and industry due to the increasingly wide applications of robots in various areas, and potential challenges resulting from the complexity of robot software. In the past years, diverse programming technologies and languages have been designed to support the development of robot software. However, with robot applications and requirements change, to develop robot software remains a great challenge, especially when robots are widely used in open environment and expected to provide better and friendly services for human beings. This paper aims at analyzing the technical requirements for designing robot programming language, presenting the roadmap of the robot programming language and discussing its trends and potential challenges. Shuo Yang 0005, Xinjun Mao, Binbin Ge |
SMC | 2 |
| 2014 | Towards realisation of evolvable runtime variability in internet-based service systems via dynamical software updateabstractToday’s Internet-based service systems tend to run in open environments and try to satisfy varying requirements. In the context, the changes of requirements and environments can emerge at any point of their life cycle, and their runtime variability is supposed to be evolvable to deal with the changes. In other words, the number, type or attribute of variability elements (variation points and their associated variants) in the systems is regarded to be changeable, especially at runtime. How to realise evolvable runtime variability in software systems is a challenge in the community of software engineering. Software architecture is supposed to support updating software without interrupting service if runtime variability is to be changed. Besides, the relevant mechanisms have to be provided to achieve dynamic update. Thus, we propose a dynamically reconfigurable reference architecture to construct the systems with evolvable software variability. The resultant system established using our approach consists of variability units and their containers. Variability units embody the variability elements that can be changed through local relay, i.e., starting a new version of variability unit to take place of older ones. Containers are mainly used to carry out the dynamic updating for variability units and realise the functionality invocation among different variability units. They can also be added, removed or replaced without shutting down the entire system. We analyse the requirements and scenarios of evolvable runtime variability in the case of Personal Data Resource Network, and further show the effectiveness and applicability of our approach by constructing the system using our class library and solving the issues proposed in the case. Xinjun Mao |
Internetware | 2 |
| 2014 | Change and Role as First-Class Abstractions for Realising Dynamic Evolution
Yin Chen 0003, Xinjun Mao |
SEKE | 2 |
| 2014 | Policy evaluation and analysis of choosing whom to tweet information on social mediaabstractChoosing whom to tweet information to promote information spread on social media is an interesting and significant topic for both academic and industrial areas. Effective policies to choose whom to tweet information on social media should make more users to be aware of the information and to be willing to retweet it. Aiming at investigating how information spreads on social media, this paper proposes an interest-based dissemination model to depict the information spread process. The interest is introduced as the foundation of users' rationality to follow other users and retweet information. Meanwhile, this paper adopts artificial society and a lightweight social computing method to evaluate and analyze the effectiveness of policies, in which users on social media are modelled as agents with various social relationships and interests. Based on PZE approach, several policies for information promotion on social media have been made, and simulations are undertaken quantitatively to evaluate their effectiveness in various scenarios. The experimental results reveal that the effectiveness of a policy is tightly related to two kinds of users' rational behaviors: retweeting information and following other users. Mingsheng Tang, Xinjun Mao, Shuqiang Yang, Haibin Zhu 0001 |
SMC | 2 |
| 2014 | Organization-based agent-oriented programming: model, mechanisms, and language
Cuiyun Hu, Xinjun Mao |
Frontiers Comput. Sci. | 2 |
| 2013 | Extending autonomic architecture for constructing internetware systemabstractThe scale and complexity of modern software systems keep increasing, especially in the context of Internet. An emerging software paradigm named Internetware was proposed to handle openness, dynamism of software systems in the context of Internet. An Internetware system is composed of self-adaptive and cooperative entities, which impliedly requires autonomy. How to construct a system equipped with these features brings challenge to developers. Autonomic computing that intends to tackle the self-management issues of complex system is helpful to deal with the development and evolution of Internetware software, but the original model of autonomic element in the area can only be used to make a system autonomic. Hence, in this paper, we extend the model of autonomic elements as a basic construct to support the design and construction of Internetware software. Then we compose an architecture for Internetware entity out of extended elements to support constructing Internetware systems. The architecture describe the essential structure of an Internetware entity and make it and most of its modules performing self-adaptiveness and cooperativity possible, which makes our works different from most of those related. A case study is conducted to illustrate our approach and show its applicability. Xinjun Mao |
Internetware | 2 |
| 2012 | An approach to modelling city-scale artificial society based-on organization metaphorabstractComplexity issues in the real world cannot be completely solved by experiments on the real society. ACP approach (Artificial society for modelling, Computations experiment for analysis and Parallel execution for control) has been proposed to solve to eliminate the gap of micro-to-macro emergent behaviours. Artificial society modelling is the consitituent of the ACP approach. However, there are no widely accepted and well-established methods to conduct and standardize artificial society modelling. This paper describes the main aspects of artificial society modelling, including agent population, environment and event. Based on the social organization metaphor, it presents an artificial society modelling approach, including the modelling process. This approach can effectively aid and guide artificial society modelling. Mingsheng Tang, Xinjun Mao, Xueyan Tan |
SMC | 2 |
| 2011 | Programming Dynamics of Multi-Agent Systems
Cuiyun Hu, Xinjun Mao |
PRIMA | 2 |
| 2011 | A Survey of Software Engineering for Self-Organization Systems
Xinjun Mao, Cuiyun Hu, Junwen Yin, Jiang Cao |
SEKE | 2 |
| 2011 | Capability as Requirement MetaphorabstractRequirement Engineering (RE) has become an attractive field in both industry and academic. Many RE approaches have been presented in the past years to support eliciting, modeling, analyzing and specifying requirements of system to be built. However, requirement characteristics of kinds of systems like large scale software intensive systems pose several issues to requirements analysis and therefore challenge the extant RE approaches. This paper investigates a number of important metaphors in RE and proposes a novel RE approach that adopts capability as requirement metaphor. We discuss the requirements challenges coming from the changes of system-to-be and argue the necessity to introduce new abstraction and technology into RE to deal with the problems. The notions of capability and the reason to adopt capability as requirement metaphor are analyzed. The meta-model and framework of capability-based requirement engineering is proposed. A case is also studied in order to illustrate our approach. Capability as new abstraction in RE provides a new way to represent, analyze and tradeoff requirements. Jiang Cao, Xinjun Mao, Huining Yan, Yushi Huang, Huaimin Wang 0001, Xicheng Lu |
TrustCom | 2 |
| 2011 | Design Pattern for Self-Organization Multi-agent Systems Based on PolicyabstractSelf-organization is an important characteristic for multi-agent systems. However developing self-organization multi-agent systems in a repeated and effective way is still a great challenge. It is well known that design patterns are the practice and knowledge about solutions for recurring problems, and reusing design pattern can significantly improve the software development quality and efficiency. On the other hand, current self-organization applications are still lack of effective way to handle the unpredicted environment and changing user requirements. In this paper, we take policy as the basic principle for self-organization multi-agent systems to solve this problem, and propose a policy based design pattern for the self-organization multi-agent systems. Xinjun Mao, Cuiyun Hu |
TrustCom | 2 |
| 2009 | SADE: A Development Environment for Adaptive Multi-Agent Systems
Menggao Dong, Xinjun Mao, Junwen Yin, Zhiming Chang, Zhichang Qi |
PRIMA | 2 |
| 2008 | Towards a Formal Model for Reconfigurable Software Architectures by BigraphsabstractWith the spread of the Internet and software evolution in complex intensive systems, software architecture often need be reconfigured during runtime to adapt variable environments and design objectives. To deal with reconfigurable software architectures, the formal method should be presented to describe software architectures and express their changes so that these changes on the evolutions of software architectures could be reasoned about. However, current formal methods for reconfigurable software architectures are difficult to represent hierarchy and model context-aware systems. In this paper, we use and extend bigraph as a formal method to describe reconfigurable software architecture. By providing graphic elements and term languages, extended bigraphs can survey static and dynamic architectures easily. Then we represent basic architectural operations based on extended bigraphs, through a case describe reconfigurations with constraints and context-aware information by reaction rules, and illustrate how to check the properties to satisfy design requirements by BiLog. Zhiming Chang, Xinjun Mao, Zhichang Qi |
WICSA | 2 |
| 2007 | Integrating Agent Technology and SIP Technology to Develop Telecommunication Applications with JadexT
Xinjun Mao, Huocheng Wu |
PRIMA | 1 |
| 2007 | Engineering Adaptive Multi-Agent Systems with ODAM Methodology
Xinjun Mao, Jianming Zhao, Ji Wang 0001 |
PRIMA | 1 |
| 2007 | An Approach based on Bigraphical Reactive Systems to Check Architectural Instance Conforming to its StyleabstractWith the spread of the Internet and software evolution in complex intensive systems, software architecture often need be reconfigured during run time in dynamic, heterogeneous environments in order to satisfy design objectives, which poses new problems such as, does the architecture of a system conform to the given architectural style? Existing formal methods for the conformance check are either obscure to be understood, or inadequate to express parameters, global conditions, and so on. In this paper, we present an approach to check architectural instance conforming to its style based on bigraphical reactive systems (BRSs). We extend bigraph and Sigma-sorted BRS to describe architectural instance and architectural style respectively, and provide an approach to support the conformance check. The approach not only provides a visual and formal mechanism to specify architectural instances and styles, but also enriches the capability to model evolving systems and deal with parametric reaction rules, which are excellent over other existing formal methods naturally. An important theorem the changing bigraphs always preserve the constraints defined by Sigma-sorted BRS if the initial bigraph and reaction rules do is proved and a conformance algorithm is presented. Two cases are studied in order to illustrate the effectiveness of our approach. Zhiming Chang, Xinjun Mao, Zhichang Qi |
TASE | 2 |
| 2006 | The Dynamic Casteship Mechanism for Modeling and Designing Adaptive Agents
Xinjun Mao, Zhiming Chang, Lijun Shan, Hong Zhu 0002, Ji Wang 0001 |
SEKE | 1 |
| 2005 | An OO-based Design Model of Software AgentabstractHow to implement software agent is a key problem for developing agent-oriented programming languages and tools. Aiming to solve this problem based on the OO technology, this paper first discusses differences between agent and object, and then puts forward an implemental architecture of software agent based on the OO technology and a simplified improvement of the BDI agent model. An OO-based design framework of software agent is also proposed by using the POAD method finally. These research results are helpful to clarify how to suitably extend the OO method for solving the key problem Jianxing Li, Xinjun Mao, Yao Shu |
PDCAT | 2 |
| 2003 | Generating Test Oracle for Role Binding in Multi-Agent SystemsabstractMultiagent systems (MAS) are typical distributed systems. It is an open problem for generating test oracles for MAS. A role-based MAS model is proposed, together with some basic definitions on soft gene, role and agent. A method to generate test oracle for such a MAS is brought forward, which may address the checking of dynamic binding between agents and roles. RoboCup simulation football team case is studied to illustrate the new MAS model and the corresponding test oracle generating method. Qi Yan 0001, Xinjun Mao, Zhichang Qi |
APSEC | 3 |