EDBT 2026 Demo / reviewers in the wild / expert
Yongji Wang 0002
dblp:06/771-2
· DBLP profile ↗
44ranked-venue papers
1as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 10 · 5 since 2021Databases, data management, data science and information retrieval · 8 · 4 since 2021Software engineering, systems software and programming languages · 6Systems, architecture and hardware · 5Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Security and privacy · 3Computer networks · 2 · 1 since 2021Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Private-library-oriented code generation with large language models
Daoguang Zan, Bei Chen 0008, Yongshun Gong, Junzhi Cao, Fengji Zhang, Bingchao Wu, Bei Guan, Yilong Yin, Yongji Wang 0002 |
Knowl. Based Syst. | 9 |
| 2024 | A GAN-Based Data Poisoning Framework Against Anomaly Detection in Vertical Federated LearningabstractIn vertical federated learning (VFL), commercial entities collaboratively train a model while preserving data privacy. However, a malicious participant's poisoning attack may degrade the performance of this collaborative model. The main challenge in achieving the poisoning attack is the absence of access to the server-side top model, leaving the malicious par-ticipant without a clear target model. To address this challenge, we introduce an innovative end-to-end poisoning framework P-GAN. Specifically, the malicious participant initially employs semi-supervised learning to train a surrogate target model. Subsequently, this participant employs a GAN-based method to produce adversarial perturbations to degrade the surrogate target model's performance. Finally, the generator is obtained and tailored for VFL poisoning. Besides, we develop an anomaly detection algorithm based on a deep auto-encoder (DAE), offering a robust defense mechanism to VFL scenarios. Through extensive experiments, we evaluate the efficacy of P-GAN and DAE, and further analyze the factors that influence their performance. Daoguang Zan, Wei Li 0326, Bei Guan, Yongji Wang 0002 |
ICC | 5 |
| 2024 | FIA-TE: Feature Inference Attack on Decision Tree Ensembles in Vertical Federated LearningabstractVertical federated learning (VFL) enables multiple parties to collaboratively train a model while preserving privacy. However, recent studies have raised concerns about the susceptibility of VFL models, including those using logistic regression and neural networks, to feature inference attacks. Meanwhile, the non-differentiable characteristics of decision tree ensembles make conducting such attacks impractical. To address this challenge, we introduce a feature inference attack framework FIA-TE tailored for decision tree ensembles, including gradient boosted decision trees (GBDT) and random forest. Specifically, we distill the knowledge from trees into neural networks by leaf embedding and structure distillation to create a targeted model for the inference attack. We then employ a generative model based on the deconvolutional network for capturing correlation features and reconstructing the target features. Through extensive experiments on table and image data, we evaluate the effectiveness of our framework and provide an analysis of potential influencing factors. Daoguang Zan, Wei Li 0326, Bei Guan, Yongji Wang 0002 |
ICME | 5 |
| 2023 | Large Language Models Meet NL2Code: A SurveyabstractDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Wang Yongji, Jian-Guang Lou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Daoguang Zan, Bei Chen 0008, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Yongji Wang 0002, Jian-Guang Lou |
ACL (1) | 7 |
| 2023 | Hierarchical and Contrastive Representation Learning for Knowledge-Aware RecommendationabstractIncorporating knowledge graph into recommendation is an effective way to alleviate data sparsity. Most existing knowledge-aware methods usually perform recursive embedding propagation by enumerating graph neighbors. However, the number of nodes’ neighbors grows exponentially as the hop number increases, forcing the nodes to be aware of vast neighbors under this recursive propagation for distilling the high-order semantic relatedness. This may induce more harmful noise than useful information into recommendation, leading the learned node representations to be indistinguishable from each other, that is, the well-known over-smoothing issue. To relieve this issue, we propose a Hierarchical and CONtrastive representation learning framework for knowledge-aware recommendation named HiCON. Specifically, for avoiding the exponential expansion of neighbors, we propose a hierarchical message aggregation mechanism to interact separately with low-order neighbors and meta-path-constrained high-order neighbors. Moreover, we also perform cross-order contrastive learning to enforce the representations to be more discriminative. Extensive experiments on three datasets show the remarkable superiority of HiCON over state-of-the-art approaches. The code is available now1. Bingchao Wu, Yangyuxuan Kang, Daoguang Zan, Bei Guan, Yongji Wang 0002 |
ICME | 5 |
| 2023 | We Are Not So Similar: Alleviating User Representation Collapse in Social RecommendationabstractIntegrating social relations into recommendation is an effective way to mitigate data sparsity. Most social recommendation methods encode user representations from a unified graph that includes user-user and user-item relations. Due to the enriched relations on this graph, a large fraction of users are aware of each other within only a few hops, and the user representations generated by existing methods may encode the information received from a large number of neighbors. Thus, many user representations are enforced to be too similar, which hinders modeling fine-grained user interest. Here, we name this phenomenon as user representation collapse. To address this problem, in this paper we propose a robust user representation learning method named RobustSR with social regularization and multi-view contrastive learning, which aim to enhance the model’s awareness of relation informativeness and the discriminativeness of user representations, respectively. Concretely, the social regularization mechanism encourages the model to learn from the relation importance weights derived from graph topologies, which helps recognize important observed relations meanwhile mining potential useful relations. To enhance the discriminativeness of user representations, we further perform multi-view contrastive learning between collaborative and social-enhanced user representations. Extensive experiments on four benchmark datasets show that RobustSR effectively alleviates user representation collapse and improves recommendation performance. Our code is deposited at https://github.com/paulpig/RobustSR. Bingchao Wu, Yangyuxuan Kang, Bei Guan, Yongji Wang 0002 |
ICMR | 4 |
| 2022 | Enhancing Sequential Recommendation via Decoupled Knowledge Graphs
Bingchao Wu, Chenglong Deng, Bei Guan, Yongji Wang 0002, Yuxuan Kangyang |
ESWC | 4 |
| 2022 | CERT: Continual Pre-training on Sketches for Library-oriented Code GenerationabstractCode generation is a longstanding challenge, aiming to generate a code snippet based on a natural language description. Usually, expensive text-code paired data is essential for training a code generation model. Recently, thanks to the success of pre-training techniques, large language models are trained on large unlabelled code corpora and perform well in generating code. In this paper, we investigate how to leverage an unlabelled code corpus to train a model for library-oriented code generation. Since it is a common practice for programmers to reuse third-party libraries, in which case the text-code paired data are harder to obtain due to the huge number of libraries. We observe that library-oriented code snippets are more likely to share similar code sketches. Hence, we present CERT with two steps: a sketcher generates the sketch, then a generator fills the details in the sketch. Both the sketcher and generator are continually pre-trained upon a base model using unlabelled data. Also, we carefully craft two benchmarks to evaluate library-oriented code generation named PandasEval and NumpyEval. Experimental results have shown the impressive performance of CERT. For example, it surpasses the base model by an absolute 15.67% improvement in terms of pass@1 on PandasEval. Our work is available at https://github.com/microsoft/PyCodeGPT. Daoguang Zan, Bei Chen 0008, Dejian Yang, Zeqi Lin, Bei Guan, Yongji Wang 0002, Weizhu Chen, Jian-Guang Lou |
IJCAI | 7 |
| 2022 | Complex Question Answering over Incomplete Knowledge Graph as N-ary Link PredictionabstractThe Question Answering over Knowledge Graph (KGQA) task seeks entities (answers) from the Knowledge Graph (KG) in order to answer natural language questions. In practice, KG is often incomplete, with numerous missing links and nodes. With such an incomplete KG, it is tricky to use the semantics inside the KG to get the golden answers, particularly for complex questions. Some current efforts concentrate on using external corpora to overcome KG sparsity; however, identifying and obtaining the corpora is challenging. Other types of work aim to leverage the pre-trained embeddings to resolve the issue but perform slightly worse on complex questions involving numerous triple facts in KG. To address the aforementioned problems, we present a framework CAPKGQA, which transforms Complex KGQA into an n-Ary link Prediction task capable of explicitly modeling complex questions. Furthermore, previous methods also suffer from incomplete KG throughout the candidate answer generation phase. Therefore, we devise an embedding-based retrieval strategy to extract more reliable candidate answers from incomplete KG. Extensive experiments reveal that our approach beats the state-of-the-art models on incomplete and complex KGQA tasks by a significant margin. Daoguang Zan, Kun Zhou 0002, Wei Wu 0014, Wayne Xin Zhao, Bingchao Wu, Bei Guan, Yongji Wang 0002 |
IJCNN | 9 |
| 2022 | S2QL: Retrieval Augmented Zero-Shot Question Answering over Knowledge Graph
Daoguang Zan, Yuanmeng Yan, Wei Wu 0014, Bei Guan, Yongji Wang 0002 |
PAKDD (3) | 7 |
| 2021 | Fed-EINI: An Efficient and Interpretable Inference Framework for Decision Tree Ensembles in Vertical Federated LearningabstractVertical federated learning has a great potential of driving a great variety of business cooperation among enterprises in many fields. In machine learning, decision tree ensembles such as gradient boosting decision trees (GBDT) and random forest are widely applied powerful models with high interpretability and modeling efficiency. However, state-of-art framework for decision tree ensembles in vertical federated learning frameworks adapt anonymous features to avoid possible data breaches, makes the interpretability of the model compromised.To address this issue in the inference process, in this paper, we firstly make a problem analysis about the necessity of disclosure meanings of feature to Guest Party in vertical federated learning. We protect data privacy and allow the disclosure of feature meaning by concealing decision paths and adapt a communication-efficient secure computation method for inference outputs. The advantages of Fed-EINI will be demonstrated through both theoretical analysis and extensive numerical results. We improve the interpretability of the model by disclosing the meaning of features while ensuring efficiency and accuracy. Shuai Zhou 0001, Bei Guan, Hao Fao, Yongji Wang 0002 |
IEEE BigData | 7 |
| 2019 | An Efficient Approach for Mitigating Covert Storage Channel Attacks in Virtual Machines by the Anti-Detection Criterion
Nasro Min-Allah, Bei Guan, Yuqi Lin, JingZheng Wu, Yongji Wang 0002 |
J. Comput. Sci. Technol. | 6 |
| 2018 | A novel anti-detection criterion for covert storage channel threat estimation
Changyou Zhang, Yongji Wang 0002 |
Sci. China Inf. Sci. | 5 |
| 2018 | An interleaved depth-first search method for the linear optimization problem with disjunctive constraints
Yinrun Lyu, Changyou Zhang, Dacheng Qu, Nasro Min-Allah, Yongji Wang 0002 |
J. Glob. Optim. | 6 |
| 2017 | Solving linear optimization over arithmetic constraint formula
Yinrun Lyu, JingZheng Wu, Changyou Zhang, Nasro Min-Allah, Jamal Alhiyafi, Yongji Wang 0002 |
J. Glob. Optim. | 8 |
| 2016 | An Efficient Approach for Solving Optimization over Linear Arithmetic Constraints
JingZheng Wu, Yinrun Lv, Yongji Wang 0002 |
J. Comput. Sci. Technol. | 4 |
| 2016 | Designing and Modeling of Covert Channels in Operating SystemsabstractCovert channels are widely considered as a major risk of information leakage in various operating systems, such as desktop, cloud, and mobile systems. The existing works of modeling covert channels have mainly focused on using finite state machines (FSMs) and their transforms to describe the process of covert channel transmission. However, a FSM is rather an abstract model, where information about the shared resource, synchronization, and encoding/decoding cannot be presented in the model, making it difficult for researchers to realize and analyze the covert channels. In this paper, we use the high-level Petri Nets (HLPN) to model the structural and behavioral properties of covert channels. We use the HLPN to model the classic covert channel protocol. Moreover, the results from the analysis of the HLPN model are used to highlight the major shortcomings and interferences in the protocol. Furthermore, we propose two new covert channel models, namely: (a) two channel transmission protocol (TCTP) model and (b) self-adaptive protocol (SAP) model. The TCTP model circumvents the mutual inferences in encoding and synchronization operations; whereas the SAP model uses sleeping time and redundancy check to ensure correct transmission in an environment with strong noise. To demonstrate the correctness and usability of our proposed models in heterogeneous environments, we implement the TCTP and SAP in three different systems: (a) Linux, (b) Xen, and (c) Fiasco.OC. Our implementation also indicates the practicability of the models in heterogeneous, scalable and flexible environments. Yuqi Lin, Saif Ur Rehman Malik, Kashif Bilal, Qiusong Yang, Yongji Wang 0002, Samee Ullah Khan |
IEEE Trans. Computers | 5 |
| 2015 | POSTER: biTheft: Stealing Your Secrets by Bidirectional Covert Channel Communication with Zero-Permission Android ApplicationabstractAndroid has 81.5% of the smartphone market now, and it is also suffering from the explosive growth of malicious applications (or apps). These apps steal users' secret data and transmit it out of the phones. By analyzing the required permissions and the abnormal behaviors, some malicious apps may be easily detected. However, in this paper, we present a bidirectional covert channel in Android, named biTheft, which steals secrets and privacies covertly without any permission. biTheft firstly collects secret data from a set of unprotected shared resources in Android system. Then, it analyzes and infers secrets from the data. With the Intent mechanism, biTheft transmits secrets by legally launching some activities of other apps without requiring any permission itself. biTheft also monitors the usages and statuses of the shared resources to receive commands from remote server. We implement a biTheft scenario, and demonstrate that some types of secrets can be stolen and transmitted out. With pre-agreement, biTheft dynamically adjusts according with the remote server commands. Comparing with the traditional covert channels, biTheft is more practical in the real world scenarios. JingZheng Wu, Mutian Yang, Zhifei Wu, Tianyue Luo, Yongji Wang 0002 |
CCS | 6 |
| 2014 | Mobile robots' modular navigation controller using spiking neural networks
Xiuqing Wang, Zeng-Guang Hou, Feng Lv, Min Tan 0001, Yongji Wang 0002 |
Neurocomputing | 5 |
| 2014 | C2Detector: a covert channel detection framework in cloud computingabstractABSTRACT Cloud computing is becoming increasingly popular because of the dynamic deployment of computing service. Another advantage of cloud is that data confidentiality is protected by the cloud provider with the virtualization technology. However, a covert channel can break the isolation of the virtualization platform and leak confidential information without letting it known by virtual machines. In this paper, the threat model of covert channels is analyzed. The channels are classified into three categories, and only the category that is new to cloud computing is concerned, for example, CPU load‐based, cache‐based, and shared memory‐based covert channels. The covert channel scenario is modeled into an error‐corrected four‐state automaton, and two error‐corrected algorithms are designed. A new detection framework termed C2Detector is presented. C2Detector includes a captor located in the hypervisor and a two‐phase synthesis algorithm implemented as Markov and Bayesian detectors. A prototype of C2Detector is implemented on Xen hypervisor, and its performance of detecting the covert channels is demonstrated. The experiment results show that C2Detector can detect the three types of the covert channels with an acceptable false positive rate by using a pessimistic threshold. Moreover, C2Detector is a plug‐in framework and can be easily extended. It is believed that new covert channels can be detected by C2Detector in the future. Copyright © 2013 John Wiley & Sons, Ltd. JingZheng Wu, Liping Ding, Nasro Min-Allah, Samee Ullah Khan, Yongji Wang 0002 |
Secur. Commun. Networks | 6 |
| 2014 | CIVSched: A Communication-Aware Inter-VM Scheduling Technique for Decreased Network Latency between Co-Located VMsabstractServer consolidation in cloud computing environments makes it possible for multiple servers or desktops to run on a single physical server for high resource utilization, low cost, and reduced energy consumption. However, the scheduler in the virtual machine monitor (VMM), such as Xen credit scheduler, is agnostic about the communication behavior between the guest operating systems (OS). The aforementioned behavior leads to increased network communication latency in consolidated environments. In particular, the CPU resources management has a critical impact on the network latency between co-located virtual machines (VMs) when there are CPUand I/O-intensive workloads running simultaneously. This paper presents the design and implementation of a communication-aware inter-VM scheduling (CIVSched) technique that takes into account the communication behavior between inter-VMs running on the same virtualization platform. The CIVSched technique inspects the network packets transmitted between local co-resident domains to identify the target VM and process that will receive the packets. Thereafter, the target VM and process are preferentially scheduled by the VMM and the guest OS. The cooperation of these two schedulers makes the network packets to be timely received by the target application. Experimental results on the Xen virtualization platform depict that the CIVSched technique can reduce the average response time of network traffic by approximately 19 percent for the highly consolidated environment, while keeping the inherent fairness of the VMM scheduler. Bei Guan, JingZheng Wu, Yongji Wang 0002, Samee Ullah Khan |
IEEE Trans. Cloud Comput. | 3 |
| 2013 | Vulnerability Detection of Android System in Fuzzing CloudabstractThe rapid growth of Android system has encountered enormous security challenges. The vulnerabilities caused by the limited security models, coarse permission system and code flaws lead to private information leakage, deny of service, potential costs, etc. To detect these vulnerabilities, some analysis and security testing methods have been presented. However, most of these methods focus on certain aspects, for example, applications, permission, or capability leakage. In this paper, we propose a new detection paradigm named Fuzzing Cloud to detect vulnerabilities in Android system. Firstly, the architecture of fuzzing cloud is introduced, and the fuzzing nodes are investigated. Then, each layer of the Android system is decomposed into separated modules, and the fuzzing test cases are created with the endless capacity of processing power and storage in fuzzing cloud. Finally, the prototype of fuzzing cloud has been implemented, and some separated modules have been tested. The experiment results show that some vulnerabilities can be detected by the fuzzing cloud. It is also believed that after small extension, fuzzing cloud can detect vulnerabilities in other systems. JingZheng Wu, Mutian Yang, Zhifei Wu, Yongji Wang 0002 |
IEEE CLOUD | 5 |
| 2013 | CIVSched: Communication-aware Inter-VM Scheduling in Virtual Machine Monitor Based on the ProcessabstractServer consolidation in Cloud Computing makes it possible for multiple servers or desktops to run on one physical server to get high resource utilization, low cost and less energy consumption. However, the scheduler in virtual machine monitor (VMM) is agnostic about the communication behavior between the guest operating systems. It leads to inefficient network communication in consolidated environment. In particular, the CPU resource management has a critical impact on the network latency between co-resident virtual machines (VMs) when there are CPU-bound and I/O-bound workloads existing simultaneously. It brings a negative impact on latency-sensitive VMs. In this paper, we present the design and implementation of CIVSched scheduling to make the VMM aware of the communication behavior between two inter-VMs running on the same virtual platform. CIVSched inspects the network packets transmitted between local domains and find the destination VM and the target process inside that will receive the packets. Then, the destination VM and the target process are preferentially scheduled by VMM scheduler and guest OS scheduler respectively. The cooperation of these two schedulers makes the network packets received by the target application timely. Experimental results show that the CIVSched scheduling can reduce the average response time of network traffic by up to 18% for the highly consolidated environment while keeping the fairness of the VMM scheduler. Bei Guan, Liping Ding, Yongji Wang 0002 |
CCGRID | 4 |
| 2013 | Analysis of the Key Factors for Software Quality in Crowdsourcing Development: An Empirical Study on TopCoder.comabstractCrowdsourcing is a distributed problem-solving and production model. It takes advantage of the internet technology, helps enterprises save cost and improve efficiency. However, uncertain quality is a significant challenge for crowdsourcing. On the basis of the existing literatures, this paper proposes 23 software quality factors from two aspects: platform and project. By using multiple regression analysis on the data of one of the most successful software crowdsourcing platforms TopCoder.com, this paper analyzes the impact of the factors on software quality and identifies six key factors, including the average quality score of the platform, the number of contemporary projects, the length of component document, the number of registered developers, the maximum rating of submitted developers, and the design score. According to the result, this paper suggests four aspects for enterprises to improve software quality: choosing the prosperous period of platform to post a project, reducing the scale of projects, attracting more and higher skillful developers to participate, and improving software design score. Junchao Xiao, Yongji Wang 0002, Qing Wang 0001 |
COMPSAC | 3 |
| 2013 | A survey on resource allocation in high performance distributed computing systems
Hameed Hussain, Saif Ur Rehman Malik, Abdul Hameed, Samee Ullah Khan, Gage Bickler, Nasro Min-Allah, Muhammad Bilal Qureshi, Yongji Wang 0002, Nasir Ghani, Joanna Kolodziej, Albert Y. Zomaya, Cheng-Zhong Xu 0001, Pavan Balaji, Abhinav Vishnu, Frédéric Pinel, Johnatan E. Pecero, Dzmitry Kliazovich, Pascal Bouvry, Hongxiang Li 0001, Lizhe Wang 0001, Dan Chen 0001, Ammar Rayes |
Parallel Comput. | 9 |
| 2012 | XenPump: A New Method to Mitigate Timing Channel in Cloud ComputingabstractCloud computing security has become the focus in information security, where much attention has been drawn to the user privacy leakage. Although isolation and some other security policies have been provided to protect the security of cloud computing, confidential information can be still stolen by timing channels without being detected. In this paper, a new method named XenPump is presented aiming to mitigate the threat of the timing channels by adding latency. XenPump is designed as a module located in hypervisor, monitoring the hypercalls used by the timing channels and adding latencies to lower the threat into an acceptable level. The prototype of XenPump has been implemented in Xen virtualization platform, and the performance is evaluated by the shared memory based timing channel. The experiment results show that XenPump can mitigate the threat of the timing channel by interrupting both the capacity and transmission accuracy. It is believed that after small extension, XenPump can mitigate the incoming timing channels. JingZheng Wu, Liping Ding, Yuqi Lin, Nasro Min-Allah, Yongji Wang 0002 |
IEEE CLOUD | 5 |
| 2012 | Improved Linear Analysis on Block Cipher MULTI2
Yi Lu 0002, Liping Ding, Yongji Wang 0002 |
CANS | 3 |
| 2012 | A Target-Reaching Controller for Mobile Robots Using Spiking Neural Networks
Xiuqing Wang, Zeng-Guang Hou, Feng Lv, Min Tan 0001, Yongji Wang 0002 |
ICONIP (4) | 5 |
| 2012 | Integrated Heuristic for Hardware/Software Co-design on Reconfigurable DevicesabstractHardware/Software (HW/SW) partitioning and scheduling are essential to the embedded systems. In this paper, a hybrid algorithm derived from Tabu Search and Simulated Annealing is proposed for solving the HW/SW partitioning problem. The virtual hardware resource is set to implement the customized Tabu Search. Earliest-Deadline-First strategy is introduced to describe the reconfiguration of FPGA. Moreover, an algorithm combining the Breadth-First-Search with Depth-First-Search is proposed for HW/SW tasks scheduling to fit for the feature of reconfigurable systems. Experimental results show that, the proposed algorithms produce better performance than the previoPus methods cited in this paper. Peng Liu 0045, Jigang Wu, Yongji Wang 0002 |
PDCAT | 3 |
| 2012 | Optimal task execution times for periodic tasks using nonlinear constrained optimization
Nasro Min-Allah, Samee Ullah Khan, Yongji Wang 0002 |
J. Supercomput. | 3 |
| 2011 | Identification and Evaluation of Sharing Memory Covert Timing Channel in Xen Virtual MachinesabstractVirtualization technology is the basis of cloud computing, and the most important property of virtualization is isolation. Isolation guarantees security between virtual machines. However, covert channel breaks the isolation and leaks sensitive message covertly. In this paper, we formally model the isolation into noninterference, and define that all the transmission channels violating noninterference are covert channels. With this definition, we present an identification method based on information flow. This method first compiles the source code into a more structured equivalent code with LLVM. And then a search algorithm is proposed to obtain the shared resources and the operational processes in the equivalent code. A new covert channel termed sharing memory covert timing channel (SMCTC) is identified from Xen source code. We construct channel scenario for SMCTC, and evaluate its threat with the metrics of channel capacity and transmission accuracy. The results show that SMCTC is much more threatened than CPU load based and cache based covert channels etc. JingZheng Wu, Liping Ding, Yongji Wang 0002 |
IEEE CLOUD | 3 |
| 2009 | PP-HAS: A Task Priority Based Preemptive Human Resource Scheduling Method
Lizi Xie, Qing Wang 0001, Junchao Xiao, Yongji Wang 0002 |
SEKE | 4 |
| 2008 | Mining Individual Performance Indicators in Collaborative Development Using Software RepositoriesabstractA better understanding of the individual developers¿ performance has been shown to result in benefits such as improved project estimation accuracy and enhanced software quality assurance. However, new challenges of distinguishing the individual activities involved in software evolution arise when considering collaborative development environments. Since software repositories such as version control systems (VCS) and bug tracking systems (BTS) are available for most software projects and hold a detailed and rich record of the historical development information, this paper presents our experiences mining individual performance indicators in collaborative development environments by using these repositories. The base of our key idea is to identify the complexity metrics (in the code base) and field defects (from bug tracking system) at individual-level by incorporating the historical data from version control system. We also remotely measure and analyze these indicators mined from a libre project jEdit, which involves around one hundred developer. The results show that these indicators are feasible and instructive in the understanding of the individual performance. Yongji Wang 0002, Junchao Xiao |
APSEC | 2 |
| 2008 | Incremental probabilistic latent semantic analysis for automatic question recommendationabstractWith the fast development of web 2.0, user-centric publishing and knowledge management platforms, such as Wiki, Blogs, and Q & A systems attract a large number of users. Given the availability of the huge amount of meaningful user generated content, incremental model based recommendation techniques can be employed to improve users' experience using automatic recommendations. In this paper, we propose an incremental recommendation algorithm based on Probabilistic Latent Semantic Analysis (PLSA). The proposed algorithm can consider not only the users' long-term and short-term interests, but also users' negative and positive feedback. We compare the proposed method with several baseline methods using a real-world Question & Answer website called Wenda. Experiments demonstrate both the effectiveness and the efficiency of the proposed methods. Hu Wu 0001, Yongji Wang 0002, Xiang Cheng 0001 |
RecSys | 2 |
| 2007 | Mining Software Repositories to Understand the Performance of Individual DevelabstractVersion control information can be enhanced with the data from defect tracking systems and archived communications. All the information sources together lead to an extensive data set, which can help individual developers to drive changes to personal development behavior. However, the formats of and access to these information vary considerably across different software repositories that complicates the integration of these data. Furthermore, the raw information obtained from repositories is too large to provide deep insight into the software evolution at the developer level, and hence poses difficulties for researchers in carrying out a quantitative data analysis In this paper, we outline our experiences mining and merging the individual-related data from some popular software repositories. Then we adopt a multi- criteria analysis model, data envelopment analysis (DEA) to analyze the collected data based on the predefined metrics. Yongji Wang 0002 |
COMPSAC (1) | 2 |
| 2007 | Revisiting Fixed Priority Techniques
Nasro Min-Allah, Yongji Wang 0002, Jiansheng Xing, Junxiang Liu |
EUC | 2 |
| 2007 | AttributeNets: An Incremental Learning Method for Interpretable Classification
Hu Wu 0001, Yongji Wang 0002, Xiaoyong Huai |
PAKDD | 2 |
| 2006 | ARIMAmmse: An Improved ARIMA-basedabstractProductivity is a critical performance index of process resources. As successive history productivity data tends to be auto-correlated, time series prediction method based on auto-regressive integrated moving average (ARIMA) model was introduced into software productivity prediction by Humphrey et al. In this paper, a variant of their prediction method named ARIMAmmse is proposed. This variant formulates the ARIMA parameter estimation issue as a minimum mean square error (MMSE) based constrained optimization problem. The ARIMA model is used to describe constraints of the parameter estimation problem, while MMSE is used as the objective function of the constrained optimization problem. According to the optimization theory, ARIMAmmse will definitely achieve a higher MMSE prediction precision than Humphrey et al's which is based on the Yule-Walk estimation technique. Two comparative experiments are also presented. The experimental results further confirm the theoretical superiority of ARIMAmmse Yongji Wang 0002, Qing Wang 0001, Fengdi Shu, Haitao Zeng |
COMPSAC (2) | 2 |
| 2006 | Difficulties in Estimating Available BandwidthabstractAvailable bandwidth estimation is very important for network operators, users, and bandwidth-sensitive applications. In the last 15 years, a large number of techniques and software tools have been introduced to estimate the available bandwidth actively. Many of them were verified in simulation and over a limited number of Internet paths. However, none of them have been widely used because there is still great uncertainty in their accuracy over the Internet at large. Based on the analysis of 11 well-known tools and our own experience in developing one of them, we present a comprehensive analysis of the fundamental difficulties with most of the existing tools. Yongji Wang 0002, Xiaoyong Huai |
ICC | 2 |
| 2006 | BSR: a statistic-based approach for establishing and refining software process performance baselineabstractHigh-level process management is quantitative management. The Process Performance Baseline (PPB) of process or subprocess under statistical management is the most important concept. It is the basis of process control and improvement. The existing methods for establishing process baseline are too coarse-grained or have some limitation, which lead to inaccurate or ineffective quantitative management. In this paper, we propose an approach called BSR (Baseline-Statistic-Refinement) for establishing and refining software process performance baseline, and present the experience result to validate its effectiveness for quantitative process management. Qing Wang 0001, Nan Jiang 0001, Lang Gou, Mingshu Li 0001, Yongji Wang 0002 |
ICSE | 6 |
| 2006 | QSCM: Engineering QoS in Web-Based Software Configuration Management SystemabstractThe conventional Software Configuration Management (SCM) tools have shown their weakness in large and complex software systems with the development of modern software. This paper provides a better SCM tool with high availability to facilitate software development process. A novel approach named Quality of Service (QoS)-based SCM (QSCM) is presented, which introduces QoS guarantee techniques at the application level to web-based SCM system in order to improve SCM system's performance. QSCM classifies users together with SCM activities and treats with requests according to their priority. When the server is overloaded, it stops treatment with new requests so that to avoid from crash. The work described here is a first step towards this goal. Yongji Wang 0002 |
Web Intelligence | 2 |
| 2005 | Mining Quantitative Associations in Large Database
Chenyong Hu, Yongji Wang 0002, Benyu Zhang, Qiang Yang 0001, Qing Wang 0001, Jinhui Zhou, Yun Yan |
APWeb | 2 |
| 2005 | Learning quantifiable associations via principal sparse non-negative matrix factorization
Chenyong Hu, Benyu Zhang, Yongji Wang 0002, Shuicheng Yan, Zheng Chen 0001, Qing Wang 0001, Qiang Yang 0001 |
Intell. Data Anal. | 3 |
| 2005 | A Generalized Real-Time Obstacle Avoidance Method Without the Cspace Calculation
Yongji Wang 0002, Matthew P. Cartmell, Qiu-Ming Tao |
J. Comput. Sci. Technol. | 1 |