VLDB 2026 Research / reviewers in the wild / expert
Feng Xu 0007
dblp:03/2611-7
· DBLP profile ↗
62ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0003-3347-7510ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 24 · 5 since 2021Artificial intelligence and machine learning · 23 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 18 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 4 since 2021Security and privacy · 3Human-computer interaction and ubiquitous computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robustness evaluation and enhancement of LLMs in code generation: an empirical study
Senrong Xu, Yuan Yao 0001, Yibin Shen, Ping Yu 0011, Feng Xu 0007, Xiaoxing Ma |
Empir. Softw. Eng. | 8 |
| 2025 | Simulate, Refine and Integrate: Strategy Synthesis for Efficient SMT SolvingabstractSatisfiability Modulo Theories (SMT) solvers are crucial in many applications, yet their performance is often a bottleneck. This paper introduces SIRISMT, a novel framework that employs machine learning techniques for the automatic synthesis of efficient SMT-solving strategies. Specifically, SIRISMT targets at Z3 and consists of three key stages. First, given a set of training SMT formulas, SIRISMT simulates the solving process by leveraging reinforcement learning to guide its exploration within the strategy space. Next, SIRISMT refines the collected strategies by pruning redundant tactics and generating augmented strategies based on the subsequence structure of the learned strategies. These refined strategies are then fed back into the reinforcement learning model. Finally, the refined and optimized strategies are integrated into one strategy, which can be directly plugged into modern SMT solvers. Extensive evaluations show the superior performance of SIRISMT over the baseline methods. For example, compared to the default Z3, it solves 26.8% more formulas and achieves up to an 86.3% improvement in the Par-2 score on benchmark datasets. Additionally, we show that the synthesized strategy can improve the code coverage by up to 11.8% in a downstream symbolic execution benchmark. Bingzhe Zhou, Hannan Wang, Yuan Yao 0001, Taolue Chen 0001, Feng Xu 0007, Xiaoxing Ma |
IJCAI | 5 |
| 2025 | Exploiting Booster Pass Chain for Compiler Phase OrderingabstractThe phase ordering problem, which aims to find suitable pass sequences for a given program on a target architecture, is critical in compiler optimization.One key challenge of this problem lies in the complex interplay among different passes within the vast optimization space of possible pass sequences.To better explore the interplay among passes, this paper proposes a new concept called booster pass chain (BPC), and presents a novel approach that identifies and leverages the BPCs to optimize the code size.Specifically, a BPC is a sequence of passes with positive interplay that, when presented as a whole, may exhibit significant optimization effects for certain programs.We then propose an iterative algorithm to extract BPCs, based on which we build a candidate set of pass sequences.For a given program, we also train a neural network to predict the suitable pass sequences from the candidate set.Experimental evaluations on 16 datasets containing 6,186 programs demonstrate the effectiveness of the proposed approach.That is, the candidate set achieves an average of 9.9% improvement compared to the LLVM -Oz flag in code size reduction, and selecting the top-3 pass sequences using the neural network predictor achieves 6.9% improvement.Our code and results are available at https://github.com/SoftWiser-group/EBPC4CPO. Yihan Chen 0008, Huanhuan Chen 0005, Yuan Yao 0001, Ping Yu 0011, Feng Xu 0007, Xiaoxing Ma |
Internetware | 5 |
| 2025 | Comprehend, Imitate, and then Update: Unleashing the Power of LLMs in Test Suite EvolutionabstractSoftware testing plays a crucial role in software engineering, ensuring the reliability and correctness of evolving systems. Well-maintained test suites are essential for ensuring software quality. However, in modern development cycles that emphasize rapid feature iteration, the co-evolution of test suites often lags behind, leading to more appearance of obsolete tests. To this end, automated approaches for updating obsolete test code have been proposed, and recent approaches have achieved the state-of-the-art performance with the support of large language models (LLMs). This paper presents COMMITUP, a new approach that leverages LLMs to effectively automate method-level obsolete test code updates. COMMITUP mimics how humans solve the problem, first comprehending the code modifications, searching for similar examples to imitate, and finally performing the update. We evaluate COMMITUP on a curated dataset from real-world Java projects. The results demonstrate the superior performance of COMMITUP, achieving 96.4%, 94.4%, 93.1% success rates for generating compilable, runtime failure-free, and full coverage updates, respectively. We believe our study can provide new insight into LLM-based test code update. The dataset and code are available at https://github.com/SoftWiser-group/CommitUp. Tangzhi Xu, Jianhan Liu, Yuan Yao 0001, Cong Li 0003, Feng Xu 0007, Xiaoxing Ma |
ASE | 5 |
| 2025 | Detecting and Untangling Composite Commits via Attributed Graph Modeling
Sheng-Bin Xu, Yuan Yao 0001, Feng Xu 0007 |
J. Comput. Sci. Technol. | 4 |
| 2024 | Inspecting Prediction Confidence for Detecting Black-Box Backdoor AttacksabstractBackdoor attacks have been shown to be a serious security threat against deep learning models, and various defenses have been proposed to detect whether a model is backdoored or not. However, as indicated by a recent black-box attack, existing defenses can be easily bypassed by implanting the backdoor in the frequency domain. To this end, we propose a new defense DTInspector against black-box backdoor attacks, based on a new observation related to the prediction confidence of learning models. That is, to achieve a high attack success rate with a small amount of poisoned data, backdoor attacks usually render a model exhibiting statistically higher prediction confidences on the poisoned samples. We provide both theoretical and empirical evidence for the generality of this observation. DTInspector then carefully examines the prediction confidences of data samples, and decides the existence of backdoor using the shortcut nature of backdoor triggers. Extensive evaluations on six backdoor attacks, four datasets, and three advanced attacking types demonstrate the effectiveness of the proposed defense. Yuan Yao 0001, Feng Xu 0007, Miao Xu 0001, Shengwei An, Ting Wang 0006 |
AAAI | 3 |
| 2024 | On the Heterophily of Program Graphs: A Case Study of Graph-based Type InferenceabstractTreating programs as graphs and employing graph learning techniques to analyze them have been widely adopted in many software engineering tasks. A recent progress in this vein is to apply graph neural networks (GNNs) to model program graphs, which is built upon the homophily assumption, i.e., similar nodes tend to connect each other. However, this assumption is not always valid in program graphs, as various edges such as AST edges and token occurrence edges may connect dissimilar nodes with quite different properties. Such phenomenon is termed as the heterophily of program graphs. In this paper, we propose a new heterophily-aware graph convolutional network (HAGCN) to better handle the heterophilic program graphs. Specifically, we first introduce the subtraction operation into the message passing mechanism of GNNs, which allows HAGCN to push apart dissimilar nodes in the representation space. Then, HAGCN separately encodes each type of edges, and uses a global relation-aware attention mechanism to fuse messages from different edge types. Moreover, we also theoretically analyze the expressive power of HAGCN from the perspective of convolution filters and contrast the differences between HAGCN and other GNNs. Finally, we take type inference as an example to evaluate the effectiveness of the proposed approach. Experimental results demonstrate that HAGCN significantly outperforms the existing non-heterophilic competitors, as well as the existing state-of-the-art graph-based type inference approaches. Senrong Xu, Jiamei Shen, Yuan Yao 0001, Ping Yu 0011, Feng Xu 0007, Xiaoxing Ma |
Internetware | 6 |
| 2023 | Data Quality Matters: A Case Study of Obsolete Comment DetectionabstractMachine learning methods have achieved great success in many software engineering tasks. However, as a data-driven paradigm, how would the data quality impact the effectiveness of these methods remains largely unexplored. In this paper, we explore this problem under the context of just-in-time obsolete comment detection. Specifically, we first conduct data cleaning on the existing benchmark dataset, and empirically observe that with only 0.22% label corrections and even 15.0% fewer data, the existing obsolete comment detection approaches can achieve up to 10.7% relative accuracy improvement. To further mitigate the data quality issues, we propose an adversarial learning framework to simultaneously estimate the data quality and make the final predictions. Experimental evaluations show that this adversarial learning framework can further improve the relative accuracy by up to 18.1% compared to the state-of-the-art method. Although our current results are from the obsolete comment detection problem, we believe that the proposed two-phase solution, which handles the data quality issues through both the data aspect and the algorithm aspect, is also generalizable and applicable to other machine learning based software engineering tasks. Shengbin Xu, Yuan Yao 0001, Feng Xu 0007, Tianxiao Gu, Jingwei Xu 0001, Xiaoxing Ma |
ICSE | 3 |
| 2023 | Hybrid API Migration: A Marriage of Small API Mapping Models and Large Language ModelsabstractAPI migration is an essential step for code migration between libraries or programming languages, and it is a challenging task as it requires detailed comprehension of both source and target APIs. The existing work either recommends mapped API names only and requires developers to select specific parameters and return value, or uses encoder-decoder models to directly “translate” the source API code into the target API code without considering the characteristics of APIs. In this paper, we propose a hybrid approach that combines small API mapping models with Large Language Models (LLMs). Specifically, the small API mapping model is employed to embed API semantics through their usages and declarations, enabling accurate inference of API mappings across different libraries and programming languages. The inferred mappings are subsequently used as part of the prompts to guide LLMs to generate the target API code corresponding to the source API code. Experimental evaluations demonstrate the effectiveness of our approach in comparison to existing approaches w.r.t. both cross-library and cross-language API migration. Bingzhe Zhou, Shengbin Xu, Yuan Yao 0001, Minxue Pan, Feng Xu 0007, Xiaoxing Ma |
Internetware | 6 |
| 2023 | On the Vulnerability of Graph Learning-based Collaborative FilteringabstractGraph learning-based collaborative filtering (GLCF), which is built upon the message-passing mechanism of graph neural networks (GNNs), has received great recent attention and exhibited superior performance in recommender systems. However, although GNNs can be easily compromised by adversarial attacks as shown by the prior work, little attention has been paid to the vulnerability of GLCF. Questions like can GLCF models be just as easily fooled as GNNs remain largely unexplored. In this article, we propose to study the vulnerability of GLCF. Specifically, we first propose an adversarial attack against CLCF. Considering the unique challenges of attacking GLCF, we propose to adopt the greedy strategy in searching for the local optimal perturbations and design a reasonable attacking utility function to handle the non-differentiable ranking-oriented metrics. Next, we propose a defense to robustify GCLF. The defense is based on the observation that attacks usually introduce suspicious interactions into the graph to manipulate the message-passing process. We then propose to measure the suspicious score of each interaction and further reduce the message weight of suspicious interactions. We also give a theoretical guarantee of its robustness. Experimental results on three benchmark datasets show the effectiveness of both our attack and defense. Senrong Xu, Liangyue Li, Zenan Li, Yuan Yao 0001, Feng Xu 0007, Zulong Chen, Hanghang Tong |
ACM Trans. Inf. Syst. | 5 |
| 2022 | An Invisible Black-Box Backdoor Attack Through Frequency Domain
Yuan Yao 0001, Feng Xu 0007, Shengwei An, Hanghang Tong, Ting Wang 0006 |
ECCV (13) | 3 |
| 2022 | Untangling Composite Commits by Attributed Graph ClusteringabstractDuring software development, it is considered to be a best practice if each commit represents one distinct concern, such as fixing a bug or adding a new feature. However, developers may not always follow this practice and sometimes tangle multiple concerns into a single composite commit. This makes automatic commit untangling a necessary task, and recent approaches mainly untangle commits via applying graph clustering on the code dependency graph. In this paper, we propose a new commit untangling approach, ComUnt, to decompose the composite commits into atomic ones. Different from existing approaches, ComUnt is built upon the observation that both the textual content of code statements and the dependencies between code statements contain useful semantic information so as to better comprehend the committed code changes. Based on this observation, ComUnt first constructs an attributed graph for each commit, where code statements and various code dependencies are modeled as nodes and edges, respectively, and the textual body of code statements are maintained as node attributes. It then conducts attributed graph clustering on the constructed graph. The used attributed graph clustering algorithm can simultaneously encode both graph structure and node attributes so as to better separate the code changes into clusters with distinct concerns. We evaluate our approach on nine C# projects, and the experimental result shows that ComUnt improves the state-of-the-art by 7.8% in terms of untangling accuracy, and meanwhile it is more than 6 times faster. Shengbin Xu, Yuan Yao 0001, Feng Xu 0007 |
Internetware | 4 |
| 2022 | Combining Code Context and Fine-grained Code Difference for Commit Message GenerationabstractGenerating natural language messages for source code changes is an essential task in software development and maintenance. Existing solutions mainly treat a piece of code difference as natural language, and adopt seq2seq learning to translate it into a commit message. The basic assumption of such solutions lies in the naturalness hypothesis, i.e., source code written by programming languages is to some extent similar to natural language text. However, compared with natural language, source code also bears syntactic regularities. In this paper, we propose to simultaneously model the naturalness and syntactic regularities of source code changes for commit message generation. Specifically, to model syntactic regularities, we first enlarge the input with additional context information, i.e., the code statements that have dependency with the variables in the code difference, and then extract the paths in the corresponding ASTs. Moreover, to better model code difference, we align the two versions of code before and after the committed code change at token level, and annotate their differences with fine-grained edit operations. The context and difference are simultaneously encoded in a learning framework to generate the commit messages. We collected from GitHub a large dataset containing 480 Java projects with over 160k commits, and the experimental results demonstrate the effectiveness of the proposed approach. Shengbin Xu, Yuan Yao 0001, Feng Xu 0007, Tianxiao Gu, Hanghang Tong |
Internetware | 3 |
| 2022 | Fair Representation Learning: An Alternative to Mutual InformationabstractLearning fair representations is an essential task to reduce bias in data-oriented decision making. It protects minority subgroups by requiring the learned representations to be independent of sensitive attributes. To achieve independence, the vast majority of the existing work primarily relaxes it to the minimization of the mutual information between sensitive attributes and learned representations. However, direct computation of mutual information is computationally intractable, and various upper bounds currently used either are still intractable or contradict the utility of the learned representations. In this paper, we introduce distance covariance as a new dependence measure into fair representation learning. By observing that sensitive attributes (e.g., gender, race, and age group) are typically categorical, the distance covariance can be converted to a tractable penalty term without contradicting the utility desideratum. Based on the tractable penalty, we propose FairDisCo, a variational method to learn fair representations. Experiments demonstrate that FairDisCo outperforms existing competitors for fair representation learning. Zenan Li, Yuan Yao 0001, Feng Xu 0007, Xiaoxing Ma, Miao Xu 0001, Hanghang Tong |
KDD | 4 |
| 2022 | Structure Meets Sequences: Predicting Network of Co-evolving SequencesabstractCo-evolving sequences are ubiquitous in a variety of applications, where different sequences are often inherently inter-connected with each other. We refer to such sequences, together with their inherent connections modeled as a structured network, as network of co-evolving sequences (NoCES). Typical NoCES applications include road traffic monitoring, company revenue prediction, motion capture, etc. To date, it remains a daunting challenge to accurately model NoCES due to the coupling between network structure and sequences. In this paper, we propose to modeling \pname\ with the aim of simultaneously capturing both the dynamics and the interplay between network structure and sequences. Specifically, we propose a joint learning framework to alternatively update the network representations and sequence representations as the sequences evolve over time. A unique feature of our framework lies in that it can deal with the case when there are co-evolving sequences on both network nodes and edges. Experimental evaluations on four real datasets demonstrate that the proposed approach (1) outperforms the existing competitors in terms of prediction accuracy, and (2) scales linearly w.r.t. the sequence length and the network size. Yaojing Wang, Yuan Yao 0001, Feng Xu 0007, Yada Zhu, Hanghang Tong |
WSDM | 3 |
| 2022 | Auditing Network Embedding: An Edge Influence Based ApproachabstractLearning node representations in a network has a wide range of applications. Most of the existing work focuses on improving the performance of the learned node representations by designing advanced network embedding models. In contrast to these work, this article aims to provide some understanding of the rationale behind the existing network embedding models, e.g.,whya given embedding algorithm outputs the specific node representations andhowthe resulting node representations relate to the structure of the input network. In particular, we propose to discern the edge influence for two widely-studied classes of network embedding models, i.e., skip-gram based models and graph neural networks. We provide algorithms to effectively and efficiently quantify the edge influence on node representations, and further identify high-influential edges by exploiting the linkage between edge influence and network structure. Experimental evaluations are conducted on real datasets showing that: 1) in terms of quantifying edge influence, the proposed method is significantly faster (up to$2,000\times$) than straightforward methods with little quality loss, and 2) in terms of identifying high-influential edges, the identified edges by the proposed method have a significant impact in the context of downstream prediction task and adversarial attacking. Yaojing Wang, Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | Enhancing Context-Based Meta-Reinforcement Learning Algorithms via An Efficient Task Encoder (Student Abstract)abstractMeta-Reinforcement Learning (meta-RL) algorithms enable agents to adapt to new tasks from small amounts of exploration, based on the experience of similar tasks. Recent studies have pointed out that a good representation of a task is key to the success of off-policy context-based meta-RL. Inspired by contrastive methods in unsupervised representation learning, we propose a new method to learn the task representation based on the mutual information between transition tuples in a trajectory and the task embedding. We also propose a new estimation for task similarity based on Q-function, which can be used to form a constraint on the distribution of the encoded task variables, making the task encoder encode the task variables more effective on new tasks. Experiments on meta-RL tasks show that the newly proposed method outperforms existing meta-RL algorithms. Feng Xu 0007, Shengyi Jiang, Zongzhang Zhang, Yang Yu 0001, Ming Li 0005, Dong Li 0016, Wulong Liu |
AAAI | 1 |
| 2021 | Cross-modal Domain Adaptation for Cost-Efficient Visual Reinforcement LearningabstractIn visual-input sim-to-real scenarios, to overcome the reality gap between images rendered in simulators and those from the real world, domain adaptation, i.e., learning an aligned representation space between simulators and the real world, then training and deploying policies in the aligned representation, is a promising direction. Previous methods focus on same-modal domain adaptation. However, those methods require building and running simulators that render high-quality images, which can be difficult and costly. In this paper, we consider a more cost-efficient setting of visual-input sim-to-real where only low-dimensional states are simulated. We first point out that the objective of learning mapping functions in previous methods that align the representation spaces is ill-posed, prone to yield an incorrect mapping. When the mapping crosses modalities, previous methods are easier to fail. Our algorithm, Cross-mOdal Domain Adaptation with Sequential structure (CODAS), mitigates the ill-posedness by utilizing the sequential nature of the data sampling process in RL tasks. Experiments on MuJoCo and Hand Manipulation Suite tasks show that the agents deployed with our method achieve similar performance as it has in the source domain, while those deployed with previous methods designed for same-modal domain adaptation suffer a larger performance gap. Xiong-Hui Chen, Shengyi Jiang, Feng Xu 0007, Zongzhang Zhang, Yang Yu 0001 |
NeurIPS | 3 |
| 2021 | Regret Minimization Experience Replay in Off-Policy Reinforcement LearningabstractIn reinforcement learning, experience replay stores past samples for further reuse. Prioritized sampling is a promising technique to better utilize these samples. Previous criteria of prioritization include TD error, recentness and corrective feedback, which are mostly heuristically designed. In this work, we start from the regret minimization objective, and obtain an optimal prioritization strategy for Bellman update that can directly maximize the return of the policy. The theory suggests that data with higher hindsight TD error, better on-policiness and more accurate Q value should be assigned with higher weights during sampling. Thus most previous criteria only consider this strategy partially. We not only provide theoretical justifications for previous criteria, but also propose two new methods to compute the prioritization weight, namely ReMERN and ReMERT. ReMERN learns an error network, while ReMERT exploits the temporal ordering of states. Both methods outperform previous prioritized sampling algorithms in challenging RL benchmarks, including MuJoCo, Atari and Meta-World. Xu-Hui Liu, Zhenghai Xue, Jing-Cheng Pang, Shengyi Jiang, Feng Xu 0007, Yang Yu 0001 |
NeurIPS | 5 |
| 2021 | Unsupervised Attributed Network Embedding via Cross FusionabstractAttributed network embedding aims to learn low dimensional node representations by combining both the network's topological structure and node attributes. Most of the existing methods either propagate the attributes over the network structure or learn the node representations by an encoder-decoder framework. However, propagation based methods tend to prefer network structure to node attributes, whereas encoder-decoder methods tend to ignore the longer connections beyond the immediate neighbors. In order to address these limitations while enjoying the best of the two worlds, we design cross fusion layers for unsupervised attributed network embedding. Specifically, we first construct two separate views to handle network structure and node attributes, and then design cross fusion layers to allow flexible information exchange and integration between the two views. The key design goals of the cross fusion layers are three-fold: 1) allowing critical information to be propagated along the network structure, 2) encoding the heterogeneity in the local neighborhood of each node during propagation, and 3) incorporating an additional node attribute channel so that the attribute information will not be overshadowed by the structure view. Extensive experiments on three datasets and three downstream tasks demonstrate the effectiveness of the proposed method. Guosheng Pan, Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
WSDM | 4 |
| 2020 | Bringing Order to Network Embedding: A Relative Ranking based ApproachabstractNetwork embedding aims to automatically learn the node representations in networks. The basic idea of network embedding is to first construct a network to describe the neighborhood context for each node, and then learn the node representations by designing an objective function to preserve certain properties of the constructed context network. The vast majority of the existing methods, explicitly or implicitly, follow a pointwise design principle. That is, the objective can be decomposed into the summation of the certain goodness function over each individual edge of the context network. In this paper, we propose to go beyond such pointwise approaches, and introduce the ranking-oriented design principle for network embedding. The key idea is to decompose the overall objective function into the summation of a goodness function over a set of edges to collectively preserve their relative rankings on the context network. We instantiate the ranking-oriented design principle by two new network embedding algorithms, including a pairwise network embedding method PaWine which optimizes the relative weights of edge pairs, and a listwise method LiWine which optimizes the relative weights of edge lists. Both proposed algorithms bear a linear time complexity, making themselves scalable to large networks. We conduct extensive experimental evaluations on five real datasets with a variety of downstream learning tasks, which demonstrate that the proposed approaches consistently outperform the existing methods. Yaojing Wang, Guosheng Pan, Yuan Yao 0001, Hanghang Tong, Hongxia Yang, Feng Xu 0007, Jian Lu 0001 |
CIKM | 6 |
| 2020 | Trading Personalization for Accuracy: Data Debugging in Collaborative FilteringabstractCollaborative filtering has been widely used in recommender systems. Existing work has primarily focused on improving the prediction accuracy mainly via either building refined models or incorporating additional side information, yet has largely ignored the inherent distribution of the input rating data. In this paper, we propose a data debugging framework to identify overly personalized ratings whose existence degrades the performance of a given collaborative filtering model. The key idea of the proposed approach is to search for a small set of ratings whose editing (e.g., modification or deletion) would near-optimally improve the recommendation accuracy of a validation set. Experimental results demonstrate that the proposed approach can significantly improve the recommendation accuracy. Furthermore, we observe that the identified ratings significantly deviate from the average ratings of the corresponding items, and the proposed approach tends to modify them towards the average. This result sheds light on the design of future recommender systems in terms of balancing between the overall accuracy and personalization. Yuan Yao 0001, Feng Xu 0007, Miao Xu 0001, Hanghang Tong |
NeurIPS | 3 |
| 2020 | Enhancing supervised bug localization with metadata and stack-trace
Yaojing Wang, Yuan Yao 0001, Hanghang Tong, Xuan Huo, Ming Li 0005, Feng Xu 0007, Jian Lu 0001 |
Knowl. Inf. Syst. | 6 |
| 2019 | An Integral Tag Recommendation Model for Textual ContentabstractRecommending suitable tags for online textual content is a key building block for better content organization and consumption. In this paper, we identify three pillars that impact the accuracy of tag recommendation: (1) sequential text modeling meaning that the intrinsic sequential ordering as well as different areas of text might have an important implication on the corresponding tag(s) , (2) tag correlation meaning that the tags for a certain piece of textual content are often semantically correlated with each other, and (3) content-tag overlapping meaning that the vocabularies of content and tags are overlapped. However, none of the existing methods consider all these three aspects, leading to a suboptimal tag recommendation. In this paper, we propose an integral model to encode all the three aspects in a coherent encoder-decoder framework. In particular, (1) the encoder models the semantics of the textual content via Recurrent Neural Networks with the attention mechanism, (2) the decoder tackles the tag correlation with a prediction path, and (3) a shared embedding layer and an indicator function across encoder-decoder address the content-tag overlapping. Experimental results on three realworld datasets demonstrate that the proposed method significantly outperforms the existing methods in terms of recommendation accuracy. Shijie Tang, Yuan Yao 0001, Suwei Zhang, Feng Xu 0007, Tianxiao Gu, Hanghang Tong, Jian Lu 0001 |
AAAI | 4 |
| 2019 | Hashtag Recommendation for Photo Sharing ServicesabstractHashtags can greatly facilitate content navigation and improve user engagement in social media. Meaningful as it might be, recommending hashtags for photo sharing services such as Instagram and Pinterest remains a daunting task due to the following two reasons. On the endogenous side, posts in photo sharing services often contain both images and text, which are likely to be correlated with each other. Therefore, it is crucial to coherently model both image and text as well as the interaction between them. On the exogenous side, hashtags are generated by users and different users might come up with different tags for similar posts, due to their different preference and/or community effect. Therefore, it is highly desirable to characterize the users’ tagging habits. In this paper, we propose an integral and effective hashtag recommendation approach for photo sharing services. In particular, the proposed approach considers both the endogenous and exogenous effects by a content modeling module and a habit modeling module, respectively. For the content modeling module, we adopt the parallel co-attention mechanism to coherently model both image and text as well as the interaction between them; for the habit modeling module, we introduce an external memory unit to characterize the historical tagging habit of each user. The overall hashtag recommendations are generated on the basis of both the post features from the content modeling module and the habit influences from the habit modeling module. We evaluate the proposed approach on real Instagram data. The experimental results demonstrate that the proposed approach significantly outperforms the state-of-theart methods in terms of recommendation accuracy, and that both content modeling and habit modeling contribute significantly to the overall recommendation accuracy. Suwei Zhang, Yuan Yao 0001, Feng Xu 0007, Hanghang Tong, Jian Lu 0001 |
AAAI | 3 |
| 2019 | DeepIntent: Deep Icon-Behavior Learning for Detecting Intention-Behavior Discrepancy in Mobile AppsabstractMobile apps have been an indispensable part in our daily life. However, there exist many potentially harmful apps that may exploit users' privacy data, e.g., collecting the user's information or sending messages in the background. Keeping these undesired apps away from the market is an ongoing challenge. While existing work provides techniques to determine what apps do, e.g., leaking information, little work has been done to answer, are the apps' behaviors compatible with the intentions reflected by the app's UI? In this work, we explore the synergistic cooperation of deep learning and program analysis as the first step to address this challenge. Specifically, we focus on the UI widgets that respond to user interactions and examine whether the intentions reflected by their UIs justify their permission uses. We present DeepIntent, a framework that uses novel deep icon-behavior learning to learn an icon-behavior model from a large number of popular apps and detect intention-behavior discrepancies. In particular, DeepIntent provides program analysis techniques to associate the intentions (i.e., icons and contextual texts) with UI widgets' program behaviors, and infer the labels (i.e., permission uses) for the UI widgets based on the program behaviors, enabling the construction of a large-scale high-quality training dataset. Based on the results of the static analysis, DeepIntent uses deep learning techniques that jointly model icons and their contextual texts to learn an icon-behavior model, and detects intention-behavior discrepancies by computing the outlier scores based on the learned model. We evaluate DeepIntent on a large-scale dataset (9,891 benign apps and 16,262 malicious apps). With 80% of the benign apps for training and the remaining for evaluation, DeepIntent detects discrepancies with AUC scores 0.8656 and 0.8839 on benign apps and malicious apps, achieving 39.9% and 26.1% relative improvements over the state-of-the-art approaches. Shengqu Xi, Shao Yang, Xusheng Xiao, Yuan Yao 0001, Yayuan Xiong, Fengyuan Xu, Haoyu Wang 0001, Peng Gao 0008, Zhuotao Liu, Feng Xu 0007, Jian Lu 0001 |
CCS | 10 |
| 2019 | Discerning Edge Influence for Network EmbeddingabstractNetwork embedding, which learns the low-dimensional representations of nodes, has gained significant research attention. Despite its superior empirical success, often measured by the prediction performance of downstream tasks (e.g., multi-label classification), it is unclear \em why a given embedding algorithm outputs the specific node representations, and \em how the resulting node representations relate to the structure of the input network. In this paper, we propose to discern the edge influence as the first step towards understanding skip-gram basd network embedding methods. For this purpose, we propose an auditing framework Near, whose key part includes two algorithms (Near-add \ and Near-del ) to effectively and efficiently quantify the influence of each edge. Based on the algorithms, we further identify high-influential edges by exploiting the linkage between edge influence and the network structure. Experimental results demonstrate that the proposed algorithms (Near-add \ and Near-del ) are significantly faster (up to $2,000\times$) than straightforward methods with little quality loss. Moreover, the proposed framework can efficiently identify the most influential edges for network embedding in the context of downstream prediction task and adversarial attacking. Yaojing Wang, Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
CIKM | 4 |
| 2019 | Commit Message Generation for Source Code ChangesabstractCommit messages, which summarize the source code changes in natural language, are essential for program comprehension and software evolution understanding. Unfortunately, due to the lack of direct motivation, commit messages are sometimes neglected by developers, making it necessary to automatically generate such messages. State-of-the-art adopts learning based approaches such as neural machine translation models for the commit message generation problem. However, they tend to ignore the code structure information and suffer from the out-of-vocabulary issue. In this paper, we propose CoDiSum to address the above two limitations. In particular, we first extract both code structure and code semantics from the source code changes, and then jointly model these two sources of information so as to better learn the representations of the code changes. Moreover, we augment the model with copying mechanism to further mitigate the out-of-vocabulary issue. Experimental evaluations on real data demonstrate that the proposed approach significantly outperforms the state-of-the-art in terms of accurately generating the commit messages. Shengbin Xu, Yuan Yao 0001, Feng Xu 0007, Tianxiao Gu, Hanghang Tong, Jian Lu 0001 |
IJCAI | 3 |
| 2019 | Speedup Automatic Program Repair Using Dynamic Software Updating: An Empirical StudyabstractA typical generate-and-validate automatic program repair (APR) tool needs to repeatedly run the same test suite to validate each generated patch. This procedure is expensive when the number of patches is huge. Additionally, to scale to large programs, a program repair tool has to consider a small patch space in practice and thus may sacrifice the capability to find potential correct repairs. In this work, we propose to speed up automatic program repair to mitigate the above issues. One the one hand, we found that restarting processes to load patched code consumes the majority of total validation time. This problem is even severe when the program is running in a managed runtime such as Java virtual machine (JVM). On the other hand, dynamic software updating (DSU) can load and execute new code without restarting. To this end, we propose to use DSU techniques to speed up automatic program repair and present an empirical study in this paper. Within our study, DSU can bring up to 66.3 times speedup in comparison with the traditional restart approach. However, DSU may not be able to handle all patches and can also incur unknown side effects that lead to inconsistent validation results. We then further study the feasibility and consistency of applying DSU to speed up APR. Our results show that 1) less than 1% patches cannot be dynamically updated using the builtin DSU ability of JVM, and 2) DSU based validation leads to potentially harmful inconsistency in only 16 of 1,897,518 patches. Rongxun Guo, Tianxiao Gu, Yuan Yao 0001, Feng Xu 0007, Xiaoxing Ma |
Internetware | 4 |
| 2019 | Bug Triaging Based on Tossing Sequence Modeling
Shengqu Xi, Yuan Yao 0001, Xusheng Xiao, Feng Xu 0007, Jian Lu 0001 |
J. Comput. Sci. Technol. | 4 |
| 2019 | Dual-regularized one-class collaborative filtering with implicit feedback
Yuan Yao 0001, Hanghang Tong, Guo Yan, Feng Xu 0007, Xiang Zhang 0001, Boleslaw K. Szymanski, Jian Lu 0001 |
World Wide Web | 4 |
| 2018 | Bug Localization via Supervised Topic ModelingabstractBug tracking systems, which help to track the reported software bugs, have been widely used in software development and maintenance. In these systems, recognizing relevant source files among a large number of source files for a given bug report is a time-consuming and labor-intensive task for software developers. To tackle this problem, information retrieval methods have been widely used to capture either the textual similarities or the semantic similarities between bug reports and source files. However, these two types of similarities are usually considered separately and the historical bug fixings are largely ignored by the existing methods. In this paper, we propose a supervised topic modeling method (STMLOCATOR) for automatically locating the relevant source files for a given bug report. In particular, the proposed model is built upon three key observations. First, supervised modeling can effectively make use of the existing fixing histories. Second, certain words in bug reports tend to appear multiple times in their relevant source files. Third, longer source files tend to have more bugs. By integrating the above three observations, the proposed STMLOCATOR utilizes historical fixings in a supervised way and learns both the textual similarities and semantic similarities between bug reports and source files. We further consider a special type of bug reports with stack-traces in bug reports, and propose a variant of STMLOCATOR to tailor for such bug reports. Experimental evaluations on three real data sets demonstrate that the proposed STMLOCATOR can achieve up to 23.6% improvement in terms of prediction accuracy over its best competitors, and scales linearly with the size of the data. Moreover, the proposed variant further improves STMLOCATOR by up to 76.2% on those bug reports with stack-traces. Yaojing Wang, Yuan Yao 0001, Hanghang Tong, Xuan Huo, Feng Xu 0007, Jian Lu 0001 |
ICDM | 6 |
| 2018 | An Effective Approach for Routing the Bug Reports to the Right FixersabstractRouting the bug reports to potential fixers (i.e., bug triaging), is an integral step in software development and maintenance. However, manually inspecting and assigning bug reports is tedious and time-consuming, especially in those software projects that have a large amount of bug reports and developers. To make bug triaging more efficient, many machine learning and information retrieval based approaches have been proposed to automatically assign bug reports for suitable developers to fix. However, these techniques typically ignore two important facts in bug fixing. First, for some bug reports, the bug reporter himself/herself is one of the developers in the project, and he/she is likely to fix his/her reported bugs in the future. Second, for some bug reports, there may be a tossing sequence which contains several developers from the first potential fixer to the last actual fixer. Such tossing sequences encode valuable information such as the dependency of developers for the bug triaging task. To make use of the above facts, we propose a sequence to sequence model named SeqTriage to automatically route a given bug report to its responsible fixer. Evaluation results on three different open-source projects show that the proposed approach has significantly improved the accuracy of bug triaging compared with the state-of-the-art approaches (20% at best and 5% at least). Shengqu Xi, Yuan Yao 0001, Xusheng Xiao, Feng Xu 0007, Jian Lu 0001 |
Internetware | 4 |
| 2018 | Team Expansion in Collaborative Environments
Yuan Yao 0001, Guibing Guo, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
PAKDD (3) | 5 |
| 2018 | Guiding supervised topic modeling for content based tag recommendation
Shengqu Xi, Yuan Yao 0001, Feng Xu 0007, Hanghang Tong, Jian Lu 0001 |
Neurocomputing | 4 |
| 2017 | Exploring Metadata in Bug Reports for Bug LocalizationabstractInformation retrieval methods have been proposed to help developers locate related buggy source files for a given bug report. The basic assumption of these methods is that the bug description in a bug report should be textually similar to its buggy source files. However, the metadata (such as the component and version information) in bug reports is largely ignored by these methods. In this paper, we propose to explore the metadata for the bug localization task. In particular, we first apply a generative model to locate buggy source files based on the bug descriptions, and then propose to add the available metadata in bug reports into the localization process. Experimental evaluations on several software projects indicate that the metadata is useful to improve the localization accuracy and that the proposed bug localization method outperforms several existing methods. Yuan Yao 0001, Yaojing Wang, Feng Xu 0007, Jian Lu 0001 |
APSEC | 4 |
| 2017 | Scalable Algorithms for CQA Post Voting PredictionabstractCommunity Question Answering (CQA) sites, such as Stack Overflow and Yahoo! Answers, have become very popular in recent years. These sites contain rich crowdsourcing knowledge contributed by the site users in the form of questions and answers, and these questions and answers can satisfy the information needs of more users. In this article, we aim at predicting the voting scores of questions/answers shortly after they are posted in the CQA sites. To accomplish this task, we identify three key aspects that matter with the voting of a post, i.e., the non-linear relationships between features and output, the question and answer coupling, and the dynamic fashion of data arrivals. A family of algorithms are proposed to model the above three key aspects. Some approximations and extensions are also proposed to scale up the computation. We analyze the proposed algorithms in terms of optimality, correctness, and complexity. Extensive experimental evaluations conducted on two real data sets demonstrate the effectiveness and efficiency of our algorithms. Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2017 | Version-Aware Rating Prediction for Mobile App RecommendationabstractWith the great popularity of mobile devices, the amount of mobile apps has grown at a more dramatic rate than ever expected. A technical challenge is how to recommend suitable apps to mobile users. In this work, we identify and focus on a unique characteristic that exists in mobile app recommendation—that is, an app usually corresponds to multiple release versions. Based on this characteristic, we propose a fine-grain version-aware app recommendation problem. Instead of directly learning the users’ preferences over the apps, we aim to infer the ratings of users on a specific version of an app. However, the user-version rating matrix will be sparser than the corresponding user-app rating matrix, making existing recommendation methods less effective. In view of this, our approach has made two major extensions. First, we leverage the review text that is associated with each rating record; more importantly, we consider two types of version-based correlations. The first type is to capture the temporal correlations between multiple versions within the same app, and the second type of correlation is to capture the aggregation correlations between similar apps. Experimental results on a large dataset demonstrate the superiority of our approach over several competitive methods. Yuan Yao 0001, Wayne Xin Zhao, Yaojing Wang, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
ACM Trans. Inf. Syst. | 5 |
| 2016 | Tag2Word: Using Tags to Generate Words for Content Based Tag RecommendationabstractTag recommendation is helpful for the categorization and searching of online content. Existing tag recommendation methods can be divided into collaborative filtering methods and content based methods. In this paper, we put our focus on the content based tag recommendation due to its wider applicability. Our key observation is the tag-content co-occurrence, i.e., many tags have appeared multiple times in the corresponding content. Based on this observation, we propose a generative model (Tag2Word), where we generate the words based on the tag-word distribution as well as the tag itself. Experimental evaluations on real data sets demonstrate that the proposed method outperforms several existing methods in terms of recommendation accuracy, while enjoying linear scalability. Yuan Yao 0001, Feng Xu 0007, Hanghang Tong, Jian Lu 0001 |
CIKM | 3 |
| 2015 | MATAR: Keywords Enhanced Multi-label Learning for Tag Recommendation
Yuan Yao 0001, Feng Xu 0007, Jian Lu 0001 |
APWeb | 3 |
| 2015 | ConRec: A Software Framework for Context-Aware Recommendation Based on Dynamic and Personalized ContextabstractContextual information is proven helpful to recommender system. And context-aware recommender system(CARS) has been applied in various applications. To improve the accuracy of context-aware recommendation and make recommender application development easier, we develop a lightweight software framework named ConRec, which introduces a dynamic context oriented approach to extend traditional reduction based recommender. This framework takes the dynamic nature of context into full consideration from different aspects to get better recommendation result. The dynamism of context exists in the process of context modeling, the computation of context weight and the handling of newly emergent context. In ConRec, context is dynamically modeled by clustering similar context values into one set automatically, rather than statically predefined by domain experts. Users' preferences to different types of context are explicitly measured through context weighting function based on real dataset. Moreover, ConRec supports incrementally adding new type of context to recommendation process, which reduces much cost of re-building the whole recommender model. Based on our improved reduction-based algorithm, ConRec is built as a highly scalable and reusable software framework for developing context-aware recommender applications. Finally, we evaluate our proposed approach on public datasets and get more accurate recommendation than traditional methods. Ping Yu 0004, Chun Cao, Feng Xu 0007, Jian Lu 0001 |
COMPSAC | 4 |
| 2015 | Detecting Buggy Files based on Bug Reports: A Random Walk Based ApproachabstractDuring an Open Source Software (OSS) maintenance, bug localization is a laborsome and time-consuming work for geographically-separated developers. It is desirable to automatically identify related buggy files once a bug report is submitted. Existing work has proposed information retrieval techniques for this problem. However, these proposals tend to neglect the inherent structure in bug localization and they are largely dependent on the comments (annotations) in source files which may be unavailable. In this paper, we propose a random walk based approach to detecting buggy source files based on bug reports. In particular, we separately process source files and bug reports to make it less sensitive to comments, and then apply random walk to capture the inherent structure in bug localization. Experimental evaluations on three real-world open-source projects demonstrate that the proposed approach can outperform several existing methods and that it is less dependent on the comments in source files. Yaojing Wang, Feng Xu 0007, Yuan Yao 0001 |
Internetware | 2 |
| 2015 | RIT: Enhancing Recommendation with Inferred Trust
Guo Yan, Yuan Yao 0001, Feng Xu 0007, Jian Lu 0001 |
PAKDD (2) | 3 |
| 2015 | Detecting high-quality posts in community question answering sites
Yuan Yao 0001, Hanghang Tong, Tao Xie 0001, Leman Akoglu, Feng Xu 0007, Jian Lu 0001 |
Inf. Sci. | 5 |
| 2014 | Joint voting prediction for questions and answers in CQAabstractCommunity Question Answering (CQA) sites have become valuable repositories that host a massive volume of human knowledge. How can we detect a high-value answer which clears the doubts of many users? Can we tell the user if the question s/he is posting would attract a good answer? In this paper, we aim to answer these questions from the perspective of the voting outcome by the site users. Our key observation is that the voting score of an answer is strongly positively correlated with that of its question, and such correlation could be in turn used to boost the prediction performance. Armed with this observation, we propose a family of algorithms to jointly predict the voting scores of questions and answers soon after they are posted in the CQA sites. Experimental evaluations demonstrate the effectiveness of our approaches. Yuan Yao 0001, Hanghang Tong, Tao Xie 0001, Leman Akoglu, Feng Xu 0007, Jian Lu 0001 |
ASONAM | 5 |
| 2014 | Dual-Regularized One-Class Collaborative FilteringabstractCollaborative filtering is a fundamental building block in many recommender systems. While most of the existing collaborative filtering methods focus on explicit, multi-class settings (e.g., 1-5 stars in movie recommendation), many real-world applications actually belong to the one-class setting where user feedback is implicitly expressed (e.g., views in news recommendation and video recommendation). The main challenges in such one-class setting include the ambiguity of the unobserved examples and the sparseness of existing positive examples. Yuan Yao 0001, Hanghang Tong, Guo Yan, Feng Xu 0007, Xiang Zhang 0001, Boleslaw K. Szymanski, Jian Lu 0001 |
CIKM | 4 |
| 2014 | Predicting long-term impact of CQA posts: a comprehensive viewpointabstractCommunity Question Answering (CQA) sites have become valuable platforms to create, share, and seek a massive volume of human knowledge. How can we spot an insightful question that would inspire massive further discussions in CQA sites? How can we detect a valuable answer that benefits many users? The long-term impact (e.g., the size of the population a post benefits) of a question/answer post is the key quantity to answer these questions. In this paper, we aim to predict the long-term impact of questions/answers shortly after they are posted in the CQA sites. In particular, we propose a family of algorithms for the prediction problem by modeling three key aspects, i.e., non-linearity, question/answer coupling, and dynamics. We analyze our algorithms in terms of optimality, correctness, and complexity. We conduct extensive experimental evaluations on two real CQA data sets to demonstrate the effectiveness and efficiency of our algorithms. Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
KDD | 3 |
| 2014 | Exploring Review Content for Recommendation via Latent Factor Model
Yuan Yao 0001, Feng Xu 0007, Jian Lu 0001 |
PRICAI | 3 |
| 2014 | A Parallel Approach to Link Sign Prediction in Large-Scale Online Social NetworksabstractAnalyzing the underlying social network is very important for the development of online applications. Owing to the increasingly growing size of these networks, parallel techniques play important roles in many network analysis tasks. In this paper, we explore the link sign prediction problem in large-scale online social networks, and propose a parallel approach, called PLSP, to solve the problem. Specifically, we first extract a set of features that serve as a base for prediction. Experiments on several real datasets show that these features outperform those proposed by existing methods in predictive accuracy. Next, we present two speedup strategies, i.e. dataset division and feature selection, to shorten the training time. Experimental evaluations show that our parallel approach is much faster than the traditional non-parallel method and achieves higher predictive accuracy than other methods at the same time. Jiufeng Zhou, Lixin Han, Yuan Yao 0001, Xiaoqin Zeng, Feng Xu 0007 |
Comput. J. | 5 |
| 2014 | Multi-Aspect + Transitivity + Bias: An Integral Trust Inference ModelabstractInferring the pair-wise trust relationship is a core building block for many real applications. State-of-the-art approaches for such trust inference mainly employ the transitivity property of trust by propagating trust along connected users, but largely ignore other important properties such as trust bias, multi-aspect, etc. In this paper, we propose a new trust inference model to integrate all these important properties. To apply the model to both binary and continuous inference scenarios, we further propose a family of effective and efficient algorithms. Extensive experimental evaluations on real data sets show that our method achieves significant improvement over several existing benchmark approaches, for both quantifying numerical trustworthiness scores and predicting binary trust/distrust signs. In addition, it enjoys linear scalability in both time and space. Yuan Yao 0001, Hanghang Tong, Xifeng Yan, Feng Xu 0007, Jian Lu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2013 | Enhancing trustworthiness evaluation in internetware with similarity and non-negative constraintsabstractInternetware is envisioned as a new software paradigm where software developers usually need to interact with unknown partners as well as the software entities developed by them. To reduce uncertainty and boost collaborations in such setting, it is important to provide trustworthiness evaluation mechanisms so that trustworthy partners/entities can be easily found. In this work, we propose a novel trustworthiness evaluation mechanism by enhancing existing mechanisms with similarity and non-negative constraints. To be specific, we first extend an existing multi-aspect trust inference model by incorporating the non-negative constraint. One of the advantages of such constraint is its strong interpretability. Second, we incorporate similarity into two neighborhood models borrowed from recommender systems. When computing similarity, we make use of the intermediate results from the first step. Finally, these models are combined under a machine learning framework. To show the effectiveness of our method, we conduct experiments on a real data-set. The results show that: both our non-negativity extension and similarity computation improve the evaluation accuracy of the original methods, and the combined method outperforms several state-of-the-art methods. Guo Yan, Feng Xu 0007, Yuan Yao 0001, Jian Lu 0001 |
Internetware | 2 |
| 2013 | MATRI: a multi-aspect and transitive trust inference modelabstractTrust inference, which is the mechanism to build new pair-wise trustworthiness relationship based on the existing ones, is a fundamental integral part in many real applications, e.g., e-commerce, social networks, peer-to-peer networks, etc. State-of-the-art trust inference approaches mainly employ the transitivity property of trust by propagating trust along connected users (a.k.a. trust propagation), but largely ignore other important properties, e.g., prior knowledge, multi-aspect, etc. Yuan Yao 0001, Hanghang Tong, Xifeng Yan, Feng Xu 0007, Jian Lu 0001 |
WWW | 4 |
| 2013 | SelfTrust: leveraging self-assessment for trust inference in Internetware
Yuan Yao 0001, Feng Xu 0007, Yongli Ren, Hanghang Tong, Jian Lu 0001 |
Sci. China Inf. Sci. | 2 |
| 2012 | Subgraph Extraction for Trust Inference in Social NetworksabstractTrust inference is an essential task in many real world applications. Most of the existing inference algorithms suffer from the scalability issue, making themselves computationally costly, or even infeasible, for the graphs with more than thousands of nodes. In addition, the inference result, which is typically an abstract, numerical trustworthiness score, might be difficult for the end-user to interpret. In this paper, we propose sub graph extraction to address these challenges. The core of the proposed method consists of two stages: path selection and component induction. The outputs of both stages can be used as an intermediate step to speed up a variety of existing trust inference algorithms. Our experimental evaluations on real graphs show that the proposed method can accelerate existing trust inference algorithms, while maintaining high accuracy. In addition, the extracted sub graph provides an intuitive way to interpret the resulting trustworthiness score. Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
ASONAM | 3 |
| 2012 | A group recommendation approach for service selectionabstractThere are more and more services that fulfill similar functionality, such as image service provided by Flickr, Picasa and Facebook. Which should be adopted to construct our software system in the open, dynamic and non-deterministic Internet environment is a key problem. Earlier work[15, 9] analyze this problem from the point view of QoS and established generic and extensible QoS computation framework for service selection. However those framework are almost designed for individuals. As social network emerges and gets widespread, people tend to be more connected and self-organize themselves into groups. Benefits of all members should be considered when we select service for group. In this article, we propose a revised group recommendation algorithm which takes advantage of collaborative filtering technology for service selection. As the experiment demonstrates, our algorithm exhibits high accuracy. Feng Xu 0007, Yuan Yao 0001, Jian Lu 0001 |
Internetware | 2 |
| 2010 | An internetware based approach to building web page integration applications for mobile devicesabstractMobile devices are more and more popular in recent years. As a result, there're huge requests of mobile applications, especially those integrated with multiple information. However, on one hand, most of the mobile applications at present just contain some certain kinds of information and they cannot adapt to the rapid change of users' requirements, either. On the other hand, to build these applications, it's usually time consuming and there are not enough resource components with programmable interfaces. In this paper, we propose an approach based on Internerware to building web page integration applications for mobile device. We introduce a framework that provides abundant internet-programmable interfaces, a flexible integration mechanism to meet the users' rapid changing requirements and a reliable mechanism that guarantees the quality of the referred resources effectively. With this framework, we can rapidly build an application that integrates all the information according to users' requirement. Tianwei Sun, Feng Xu 0007, Jian Lu 0001 |
Internetware | 2 |
| 2009 | Internetware: a shift of software paradigmabstractInternetware is envisioned as a new software paradigm for resource integration and sharing in the open, dynamic and autonomous network environment. In this paper we discuss our visions and explorations of this new paradigm, with focus placed on the methodological perspective. A set of enabling techniques on flexible coordination of autonomous services, automatic adaptation to changing environment and trust management-based assurance of dependability are proposed to help the development of Internetware applications. Jian Lu 0001, Xiaoxing Ma, Yu Huang 0002, Chun Cao, Feng Xu 0007 |
Internetware | 5 |
| 2009 | A dynamic trust network based simulation framework for reputation-based service selectionabstractService-oriented computing is a promising approach to software system construction by selecting and composing autonomous services under the open, dynamic and non-deterministic Internet environment. Appropriate selection of high quality services used in from numerous candidates declaring similar functionalities is crucial to the overall quality of the composed system. As authority centers are not generally available in open environments, reputation-based mechanisms must be adopted to evaluate services. With more and more reputation systems proposed in the literature, there is an increasing need to evaluate and compare them objectively and systematically with a common controlled experiment of trust network environment. In this paper we propose a general simulation framework based on dynamic trust network for this purpose. Especially, the framework is capable to simulate the dynamic evolutions of the trust network, in addition to the static snapshots of the trust relationships. With this framework, some case studies are made to evaluate the effectiveness of several representative reputation mechanisms, and some interesting characters are revealed. Feng Xu 0007, Yuan Yao 0001, Jian Lu 0001 |
Internetware | 2 |
| 2009 | A Broker-Assisting Trust and Reputation System Based on Artificial Neural NetworkabstractDue to the dynamic and anonymous nature of open environments, it is critically important for agents to identify trustful cooperators which work consistently as they claim. In the e-services and e-commerce communities, trust and reputation systems are applied broadly as one kind of decision support systems, and aim to cope with the consistency problems caused by uncertain trust relationships. However, challenges still exist: on the one hand, we require more flexible trust computation models to satisfy various personal requirements since agents in these communities are heterogeneous; on the other hand, trust and reputation systems calculate the trustworthiness of agents based on the agents' past behavior. The open environments are dynamic, agents are anonymous and the records about agents' past behavior are distributed in the environments, so agents have to search the required records through the environments due to their lack of valid information. Thus, efficient, scalable and effective information collection strategies are required to address these issues. In this paper we present a distributed trust and reputation system to cope with the challenges. We propose a novel and flexible trust computation model based on artificial neural networks. With the advantages of ANN, our trust model tunes the parameters automatically to adapt to various personal requirements. We propose a broker-assisting information collection strategy based on clustering method. With the support of brokers, subcommunities are managed by reputation mechanism in an efficient and scalable way and help their members collect information with high quality. We show the performance of our trust and reputation system by simulation. Bo Zong, Feng Xu 0007, Jun Jiao, Jian Lu 0001 |
SMC | 2 |
| 2009 | A Trust-Based Approach to Estimating the Confidence of the Software System in Open Environments
Feng Xu 0007 |
J. Comput. Sci. Technol. | 1 |
| 2007 | A Trust Evolution Model for P2P Networks
Ye Tao 0012, Ping Yu 0004, Feng Xu 0007, Jian Lu 0001 |
ATC | 4 |
| 2006 | Toward Trust Management in Autonomic and Coordination Applications
Feng Xu 0007, Ye Tao 0012, Chun Cao, Jian Lu 0001 |
ATC | 2 |