VLDB 2026 Research / reviewers in the wild / expert
Yuan Yao 0001
dblp:25/4120-1
· DBLP profile ↗
85ranked-venue papers
11as first author
42since 2021 · last 2026
0000-0002-6913-6542ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 29 · 9 first-author · 8 since 2021Software engineering, systems software and programming languages · 24 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 2 since 2021Security and privacy · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robustness evaluation and enhancement of LLMs in code generation: an empirical study
Senrong Xu, Yuan Yao 0001, Yibin Shen, Ping Yu 0011, Feng Xu 0007, Xiaoxing Ma |
Empir. Softw. Eng. | 3 |
| 2025 | An Embarrassingly Simple but Effective Knowledge-enhanced RecommenderabstractKnowledge graphs (KG) have demonstrated significant potential in recommender systems by providing complementary semantic information that is typically absent in user-item interaction graphs (IG). While contrastive learning has emerged as a powerful paradigm for integrating these dual information sources, we identify a critical limitation in existing approaches: current methods fail to effectively balance the contrastive views derived from IG and KG, often resulting in performance degradation compared to using IG alone. To address this fundamental challenge, we propose SimKGCL, a novel contrastive learning framework that introduces a simple yet principled solution -- cross-view, layer-wise fusion between IG and KG representations prior to contrastive learning. This design ensures effective knowledge transfer while maintaining the discriminative power of contrastive objectives. Comprehensive experiments across three real-world benchmarks demonstrate that our approach not only consistently outperforms existing methods but also achieves remarkable efficiency gains. Our code is available through this link: https://figshare.com/articles/conference_contribution/SimKGCL/22783382. Haibo Ye, Yuan Yao 0001, Xinjie Li 0007 |
CIKM | 3 |
| 2025 | LoRA Decompose: Serving Fine-Tuned Models into LoRA-LikeabstractLarge language models (LLMs) achieve remarkable performance across diverse tasks but face increasing GPU-memory demands due to the growing variety and complexity of downstream tasks. Efficient inference has thus become essential, especially for resource-limited settings. In this paper, we propose LoRA Decompose, a novel compression approach based on a key insight: instruction-fine-tuned models share a common pretrained-like base component and differ primarily through low-rank, LoRA-like delta components. Leveraging this observation, we reformulate the inference problem as a constrained optimization task that jointly identifies a shared low-rank structure across multiple models, significantly reducing their memory footprints. We solve this optimization efficiently using a custom-designed block coordinate descent algorithm, converging quickly within a few iterations. Empirical experiments with Llama-2 7B and 13B models demonstrate that our method achieves a remarkable >32x GPU memory reduction while preserving task accuracy, allowing substantial efficiency gains for practical deployment. Yibo Han, Tangzhi Xu, Zenan Li, Youshan Miao, Yuan Yao 0001, Ningyi Xu |
ECAI | 6 |
| 2025 | Proving Olympiad Inequalities by Synergizing LLMs and Symbolic ReasoningabstractLarge language models (LLMs) can prove mathematical theorems formally by generating proof steps (\textit{a.k.a.} tactics) within a proof system. However, the space of possible tactics is vast and complex, while the available training data for formal proofs is limited, posing a significant challenge to LLM-based tactic generation. To address this, we introduce a neuro-symbolic tactic generator that synergizes the mathematical intuition learned by LLMs with domain-specific insights encoded by symbolic methods. The key aspect of this integration is identifying which parts of mathematical reasoning are best suited to LLMs and which to symbolic methods. While the high-level idea of neuro-symbolic integration is broadly applicable to various mathematical problems, in this paper, we focus specifically on Olympiad inequalities (Figure~1). We analyze how humans solve these problems and distill the techniques into two types of tactics: (1) scaling, handled by symbolic methods, and (2) rewriting, handled by LLMs. In addition, we combine symbolic tools with LLMs to prune and rank the proof goals for efficient proof search. We evaluate our framework on 161 challenging inequalities from multiple mathematics competitions, achieving state-of-the-art performance and significantly outperforming existing LLM and symbolic approaches without requiring additional training data. Zenan Li, Yuan Yao 0001, Xujie Si, Kaiyu Yang, Xiaoxing Ma |
ICLR | 5 |
| 2025 | Simulate, Refine and Integrate: Strategy Synthesis for Efficient SMT SolvingabstractSatisfiability Modulo Theories (SMT) solvers are crucial in many applications, yet their performance is often a bottleneck. This paper introduces SIRISMT, a novel framework that employs machine learning techniques for the automatic synthesis of efficient SMT-solving strategies. Specifically, SIRISMT targets at Z3 and consists of three key stages. First, given a set of training SMT formulas, SIRISMT simulates the solving process by leveraging reinforcement learning to guide its exploration within the strategy space. Next, SIRISMT refines the collected strategies by pruning redundant tactics and generating augmented strategies based on the subsequence structure of the learned strategies. These refined strategies are then fed back into the reinforcement learning model. Finally, the refined and optimized strategies are integrated into one strategy, which can be directly plugged into modern SMT solvers. Extensive evaluations show the superior performance of SIRISMT over the baseline methods. For example, compared to the default Z3, it solves 26.8% more formulas and achieves up to an 86.3% improvement in the Par-2 score on benchmark datasets. Additionally, we show that the synthesized strategy can improve the code coverage by up to 11.8% in a downstream symbolic execution benchmark. Bingzhe Zhou, Hannan Wang, Yuan Yao 0001, Taolue Chen 0001, Feng Xu 0007, Xiaoxing Ma |
IJCAI | 3 |
| 2025 | Exploiting Booster Pass Chain for Compiler Phase OrderingabstractThe phase ordering problem, which aims to find suitable pass sequences for a given program on a target architecture, is critical in compiler optimization.One key challenge of this problem lies in the complex interplay among different passes within the vast optimization space of possible pass sequences.To better explore the interplay among passes, this paper proposes a new concept called booster pass chain (BPC), and presents a novel approach that identifies and leverages the BPCs to optimize the code size.Specifically, a BPC is a sequence of passes with positive interplay that, when presented as a whole, may exhibit significant optimization effects for certain programs.We then propose an iterative algorithm to extract BPCs, based on which we build a candidate set of pass sequences.For a given program, we also train a neural network to predict the suitable pass sequences from the candidate set.Experimental evaluations on 16 datasets containing 6,186 programs demonstrate the effectiveness of the proposed approach.That is, the candidate set achieves an average of 9.9% improvement compared to the LLVM -Oz flag in code size reduction, and selecting the top-3 pass sequences using the neural network predictor achieves 6.9% improvement.Our code and results are available at https://github.com/SoftWiser-group/EBPC4CPO. Yihan Chen 0008, Huanhuan Chen 0005, Yuan Yao 0001, Ping Yu 0011, Feng Xu 0007, Xiaoxing Ma |
Internetware | 3 |
| 2025 | Comprehend, Imitate, and then Update: Unleashing the Power of LLMs in Test Suite EvolutionabstractSoftware testing plays a crucial role in software engineering, ensuring the reliability and correctness of evolving systems. Well-maintained test suites are essential for ensuring software quality. However, in modern development cycles that emphasize rapid feature iteration, the co-evolution of test suites often lags behind, leading to more appearance of obsolete tests. To this end, automated approaches for updating obsolete test code have been proposed, and recent approaches have achieved the state-of-the-art performance with the support of large language models (LLMs). This paper presents COMMITUP, a new approach that leverages LLMs to effectively automate method-level obsolete test code updates. COMMITUP mimics how humans solve the problem, first comprehending the code modifications, searching for similar examples to imitate, and finally performing the update. We evaluate COMMITUP on a curated dataset from real-world Java projects. The results demonstrate the superior performance of COMMITUP, achieving 96.4%, 94.4%, 93.1% success rates for generating compilable, runtime failure-free, and full coverage updates, respectively. We believe our study can provide new insight into LLM-based test code update. The dataset and code are available at https://github.com/SoftWiser-group/CommitUp. Tangzhi Xu, Jianhan Liu, Yuan Yao 0001, Cong Li 0003, Feng Xu 0007, Xiaoxing Ma |
ASE | 3 |
| 2025 | A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM ReasoningabstractTest-time scaling seeks to improve the reasoning performance of large language models (LLMs) by adding computational resources. A prevalent approach within the field is *sampling-based test-time scaling methods*, which enhance reasoning by generating multiple reasoning paths for a given input during inference. However, despite its practical success, the theoretical foundations remain underexplored. In this paper, we provide the first theoretical framework for analyzing sampling-based test-time scaling methods, grounded in the perspective of confidence estimation. Based on the framework, we analyze two dominant paradigms: self-consistency and perplexity, and reveal key limitations: self-consistency suffers from high estimation error while perplexity exhibits substantial modeling error and possible degradation of the estimation error convergence. To address these limitations, we introduce RPC, a hybrid method that leverages our theoretical insights through two key components: *Perplexity Consistency* and *Reasoning Pruning*. *Perplexity Consistency* combines the strengths of self-consistency and perplexity, boosting the convergence rate of estimation error from linear to exponential while preserving model error. *Reasoning Pruning* prevents degradation by eliminating low-probability reasoning paths.
Both theoretical analysis and empirical results across seven benchmark datasets demonstrate that RPC has a strong potential for reducing reasoning error. Notably, RPC achieves reasoning performance comparable to self-consistency while not only enhancing confidence reliability but also reducing sampling costs by 50%. The code and resources are available at https://wnjxyk.github.io/RPC. Zhi Zhou 0007, Tan Yuhao, Zenan Li, Yuan Yao 0001, Lan-Zhe Guo, Yufeng Li 0008, Xiaoxing Ma |
NeurIPS | 4 |
| 2025 | NexuSym: Marrying symbolic path finders with large language models
Ping Yu 0011, Yi Qin 0002, Yanyan Jiang 0001, Yuan Yao 0001, Xiaoxing Ma |
Autom. Softw. Eng. | 5 |
| 2025 | Detecting and Untangling Composite Commits via Attributed Graph Modeling
Sheng-Bin Xu, Yuan Yao 0001, Feng Xu 0007 |
J. Comput. Sci. Technol. | 3 |
| 2024 | Inspecting Prediction Confidence for Detecting Black-Box Backdoor AttacksabstractBackdoor attacks have been shown to be a serious security threat against deep learning models, and various defenses have been proposed to detect whether a model is backdoored or not. However, as indicated by a recent black-box attack, existing defenses can be easily bypassed by implanting the backdoor in the frequency domain. To this end, we propose a new defense DTInspector against black-box backdoor attacks, based on a new observation related to the prediction confidence of learning models. That is, to achieve a high attack success rate with a small amount of poisoned data, backdoor attacks usually render a model exhibiting statistically higher prediction confidences on the poisoned samples. We provide both theoretical and empirical evidence for the generality of this observation. DTInspector then carefully examines the prediction confidences of data samples, and decides the existence of backdoor using the shortcut nature of backdoor triggers. Extensive evaluations on six backdoor attacks, four datasets, and three advanced attacking types demonstrate the effectiveness of the proposed defense. Yuan Yao 0001, Feng Xu 0007, Miao Xu 0001, Shengwei An, Ting Wang 0006 |
AAAI | 2 |
| 2024 | PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language ModelsabstractHaoran Li, Dadi Guo, Donghao Li, Wei Fan, Qi Hu, Xin Liu, Chunkit Chan, Duanyi Yao, Yuan Yao, Yangqiu Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Haoran Li 0003, Dadi Guo, Wei Fan 0001, Xin Liu 0039, Chunkit Chan, Duanyi Yao, Yuan Yao 0001, Yangqiu Song |
ACL (1) | 9 |
| 2024 | Putting APIs in the Right Order with Gated Graph Neural NetworksabstractAPI plays an important role in modern software development. Automatic API recommendation has been studied for years to facilitate developers' learning process of APIs. Previous approaches mainly use statistical models and collab-orative filtering (CF) techniques to mine API usage patterns for recommendation. Despite the encouraging results, they still struggle to obtain the accurate embeddings of the client methods and called APIs. Prior studies generally formulate the process of API call interactions as undirected graph structure, neglecting the order in which the API invocations appear, thus fail to seize the rich relationship and complex transitions of API calls. To transcend the limitations, we propose a novel method, namely PARO, to predict the next API invocations using gated graph neural networks (GGNNs). In our proposed method, the API call sequences are modeled as directed graphs, thus the GNN models prone to capture features such as the partial order and complex transitions between API invocations. Besides, we also learn the text attribute representations of API invocations and client methods through word embedding, which further corroborates the semantic and lexical similarities between them. We conduct experimental evaluations on a large number of Java projects extracted from Github and Maven Central. Results show that our approach outperforms the state-of-the-art by a large margin, in terms of Hit@N and MRR@N. Ling Wan, Ping Yu 0011, Yuan Yao 0001 |
APSEC | 3 |
| 2024 | Towards Open Domain Text-Driven Synthesis of Multi-person Motions
Mengyi Shan, Lu Dong 0004, Yutao Han, Yuan Yao 0001, Ifeoma Nwogu, Guo-Jun Qi, Mitch Hill |
ECCV (65) | 4 |
| 2024 | On the Heterophily of Program Graphs: A Case Study of Graph-based Type InferenceabstractTreating programs as graphs and employing graph learning techniques to analyze them have been widely adopted in many software engineering tasks. A recent progress in this vein is to apply graph neural networks (GNNs) to model program graphs, which is built upon the homophily assumption, i.e., similar nodes tend to connect each other. However, this assumption is not always valid in program graphs, as various edges such as AST edges and token occurrence edges may connect dissimilar nodes with quite different properties. Such phenomenon is termed as the heterophily of program graphs. In this paper, we propose a new heterophily-aware graph convolutional network (HAGCN) to better handle the heterophilic program graphs. Specifically, we first introduce the subtraction operation into the message passing mechanism of GNNs, which allows HAGCN to push apart dissimilar nodes in the representation space. Then, HAGCN separately encodes each type of edges, and uses a global relation-aware attention mechanism to fuse messages from different edge types. Moreover, we also theoretically analyze the expressive power of HAGCN from the perspective of convolution filters and contrast the differences between HAGCN and other GNNs. Finally, we take type inference as an example to evaluate the effectiveness of the proposed approach. Experimental results demonstrate that HAGCN significantly outperforms the existing non-heterophilic competitors, as well as the existing state-of-the-art graph-based type inference approaches. Senrong Xu, Jiamei Shen, Yuan Yao 0001, Ping Yu 0011, Feng Xu 0007, Xiaoxing Ma |
Internetware | 4 |
| 2024 | Datactive: Data Fault Localization for Object Detection SystemsabstractObject detection (OD) models are seamlessly integrated into numerous intelligent software systems, playing a crucial role in various tasks. These models are typically constructed upon humanannotated datasets, whose quality can greatly affect their performance and reliability. Erroneous and inadequate annotated datasets can induce classification/localization inaccuracies during deployment, precipitating security breaches or traffic accidents that inflict property damage or even loss of life. Therefore, ensuring and improving data quality is a crucial issue for the reliability of the object detection system. This paper introduces Datactive, a data fault localization technique for object detection systems. Datactive is designed to locate various types of data faults including mislocalization and missing objects, without utilizing the prediction of object detection models trained on dirty datasets. To achieve this, we first construct foreground-only and background-included datasets via data disassembling strategies, and then employ a robust learning method to train classifiers using disassembled datasets. Based on the classifier predictions, Datactive produces a unified suspiciousness score for both foreground annotations and image backgrounds. It allows testers to easily identify and correct faulty or missing annotations with minimal effort. To validate the effectiveness, we conducted experiments on three datasets with 6 baselines, and demonstrated the superiority of Datactive from various aspects. We also explored Datactive's ability to find natural data faults and its application in both training and evaluation scenarios. Yining Yin, Yang Feng 0003, Shihao Weng, Yuan Yao 0001, Jia Liu 0015 |
ISSTA | 4 |
| 2024 | LLM Meets Bounded Model Checking: Neuro-symbolic Loop Invariant InferenceabstractLoop invariant inference, a key component in program verification, is a challenging task due to the inherent undecidability and complex loop behaviors in practice. Recently, machine learning based techniques have demonstrated impressive performance in generating loop invariants automatically. However, these methods highly rely on the labeled training data, and are intrinsically random and uncertain, leading to unstable performance. In this paper, we investigate a synergy of large language models (LLMs) and bounded model checking (BMC) to address these issues. The key observation is that, although LLMs may not be able to return the correct loop invariant in one response, they usually can provide all individual predicates of the correct loop invariant in multiple responses. To this end, we propose a "query-filter-reassemble" strategy, namely, we first leverage the language generation power of LLMs to produce a set of candidate invariants, where training data is not needed. Then, we employ BMC to identify valid predicates from these candidate invariants, which are assembled to produce new candidate invariants and checked by off-the-shelf SMT solvers. The feedback is incorporated into the prompt for the next round of LLM querying. We expand the existing benchmark of 133 programs to 316 programs, providing a more comprehensive testing ground. Experimental results demonstrate that our approach significantly outperforms the state-of-the-art techniques, successfully generating 309 loop invariants out of 316 cases, whereas the existing baseline methods are only able to tackle 219 programs at best. The code is publicly available at https://github.com/SoftWiser-group/LaM4Inv.git. Guangyuan Wu, Weining Cao, Yuan Yao 0001, Hengfeng Wei, Taolue Chen 0001, Xiaoxing Ma |
ASE | 3 |
| 2024 | Neuro-Symbolic Data Generation for Math ReasoningabstractA critical question about Large Language Models (LLMs) is whether their apparent deficiency in mathematical reasoning is inherent, or merely a result of insufficient exposure to high-quality mathematical data. To explore this, we developed an automated method for generating high-quality, supervised mathematical datasets. The method carefully mutates existing math problems, ensuring both diversity and validity of the newly generated problems. This is achieved by a neuro-symbolic data generation framework combining the intuitive informalization strengths of LLMs, and the precise symbolic reasoning of math solvers along with projected Markov chain Monte Carlo sampling in the highly-irregular symbolic space.
Empirical experiments demonstrate the high quality of data generated by the proposed method, and that the LLMs, specifically LLaMA-2 and Mistral, when realigned with the generated data, surpass their state-of-the-art counterparts. Zenan Li, Zhi Zhou 0007, Yuan Yao 0001, Yufeng Li 0008, Chun Cao, Xiaoxing Ma |
NeurIPS | 3 |
| 2024 | Denoised Graph Collaborative Filtering via Neighborhood Similarity and Dynamic ThresholdingabstractGraph collaborative filtering (GCF) has achieved great success in recommender systems due to its ability in mining high-order collaborative signals from historical user-item interactions. However, GCF's performance could be severely affected by the intrinsic noise within the user-item interactions. To this end, several denoised GCF frameworks have been proposed, whose heart is to estimate and handle the reliability of existing interactions. However, most of them suffer from two limitations: 1) the reliability computation itself is noisy, and 2) the reliability threshold is difficult to determine. To address the two limitations, in this paper, we propose a newNeighborhood-informedDenoising framework NiDen for GCF. Specifically, for an existing user-item interaction, NiDen first estimates its reliability by employing the neighborhood information of the user and the item, and then determines whether the interaction is noisy or not via a dynamic thresholding strategy. After that, NiDen mitigates the negative impact of noise by both structure denoising and sample re-weighting. We instantiate NiDen on two representative GCF models and conduct extensive experiments on four widely-used datasets. The results show that NiDen achieves the best performance compared to the existing denoising methods, especially on datasets with heavy noise. Haibo Ye, Yuan Yao 0001, Sheng-Jun Huang |
IEEE Trans. Big Data | 3 |
| 2023 | An Embarrassingly Simple Backdoor Attack on Self-supervised LearningabstractAs a new paradigm in machine learning, self-supervised learning (SSL) is capable of learning high-quality representations of complex data without relying on labels. In addition to eliminating the need for labeled data, research has found that SSL improves the adversarial robustness over supervised learning since lacking labels makes it more challenging for adversaries to manipulate model predictions. However, the extent to which this robustness superiority generalizes to other types of attacks remains an open question.We explore this question in the context of backdoor attacks. Specifically, we design and evaluate Ctrl, an embarrassingly simple yet highly effective self-supervised backdoor attack. By only polluting a tiny fraction of training data (≤ 1%) with indistinguishable poisoning samples, Ctrl causes any trigger-embedded input to be misclassified to the adversary's designated class with a high probability (≥ 99%) at inference time. Our findings suggest that SSL and supervised learning are comparably vulnerable to backdoor attacks. More importantly, through the lens of Ctrl, we study the inherent vulnerability of SSL to backdoor attacks. With both empirical and analytical evidence, we reveal that the representation invariance property of SSL, which benefits adversarial robustness, may also be the very reason making SSL highly susceptible to backdoor attacks. Our findings also imply that the existing defenses against supervised backdoor attacks are not easily retrofitted to the unique vulnerability of SSL. Code is available at: https://github.com/meet-cjli/CTRL Changjiang Li, Ren Pang, Zhaohan Xi, Tianyu Du, Shouling Ji, Yuan Yao 0001, Ting Wang 0006 |
ICCV | 6 |
| 2023 | Softened Symbol Grounding for Neuro-symbolic Systems
Zenan Li, Yuan Yao 0001, Taolue Chen 0001, Jingwei Xu 0001, Chun Cao, Xiaoxing Ma, Jian Lu 0001 |
ICLR | 2 |
| 2023 | Learning with Logical Constraints but without Shortcut Satisfaction
Zenan Li, Zehua Liu, Yuan Yao 0001, Jingwei Xu 0001, Taolue Chen 0001, Xiaoxing Ma, Jian Lu 0001 |
ICLR | 3 |
| 2023 | Lightweight Approaches to DNN Regression Error Reduction: An Uncertainty Alignment PerspectiveabstractRegression errors of Deep Neural Network (DNN) models refer to the case that predictions were correct by the old-version model but wrong by the new-version model. They frequently occur when upgrading DNN models in production systems, causing disproportionate user experience degradation. In this paper, we propose a lightweight regression error reduction approach with two goals: 1) requiring no model retraining and even data, and 2) not sacrificing the accuracy. The proposed approach is built upon the key insight rooted in the unmanaged model uncertainty, which is intrinsic to DNN models, but has not been thoroughly explored especially in the context of quality assurance of DNN models. Specifically, we propose a simple yet effective ensemble strategy that estimates and aligns the two models' uncertainty. We show that a Pareto improvement that reduces the regression errors without compromising the overall accuracy can be guaranteed in theory and largely achieved in practice. Comprehensive experiments with various representative models and datasets confirm that our approaches significantly outperform the state-of-the-art alternatives. Zenan Li, Maorun Zhang, Jingwei Xu 0001, Yuan Yao 0001, Chun Cao, Taolue Chen 0001, Xiaoxing Ma, Jian Lu 0001 |
ICSE | 4 |
| 2023 | Data Quality Matters: A Case Study of Obsolete Comment DetectionabstractMachine learning methods have achieved great success in many software engineering tasks. However, as a data-driven paradigm, how would the data quality impact the effectiveness of these methods remains largely unexplored. In this paper, we explore this problem under the context of just-in-time obsolete comment detection. Specifically, we first conduct data cleaning on the existing benchmark dataset, and empirically observe that with only 0.22% label corrections and even 15.0% fewer data, the existing obsolete comment detection approaches can achieve up to 10.7% relative accuracy improvement. To further mitigate the data quality issues, we propose an adversarial learning framework to simultaneously estimate the data quality and make the final predictions. Experimental evaluations show that this adversarial learning framework can further improve the relative accuracy by up to 18.1% compared to the state-of-the-art method. Although our current results are from the obsolete comment detection problem, we believe that the proposed two-phase solution, which handles the data quality issues through both the data aspect and the algorithm aspect, is also generalizable and applicable to other machine learning based software engineering tasks. Shengbin Xu, Yuan Yao 0001, Feng Xu 0007, Tianxiao Gu, Jingwei Xu 0001, Xiaoxing Ma |
ICSE | 2 |
| 2023 | Hybrid API Migration: A Marriage of Small API Mapping Models and Large Language ModelsabstractAPI migration is an essential step for code migration between libraries or programming languages, and it is a challenging task as it requires detailed comprehension of both source and target APIs. The existing work either recommends mapped API names only and requires developers to select specific parameters and return value, or uses encoder-decoder models to directly “translate” the source API code into the target API code without considering the characteristics of APIs. In this paper, we propose a hybrid approach that combines small API mapping models with Large Language Models (LLMs). Specifically, the small API mapping model is employed to embed API semantics through their usages and declarations, enabling accurate inference of API mappings across different libraries and programming languages. The inferred mappings are subsequently used as part of the prompts to guide LLMs to generate the target API code corresponding to the source API code. Experimental evaluations demonstrate the effectiveness of our approach in comparison to existing approaches w.r.t. both cross-library and cross-language API migration. Bingzhe Zhou, Shengbin Xu, Yuan Yao 0001, Minxue Pan, Feng Xu 0007, Xiaoxing Ma |
Internetware | 4 |
| 2023 | Neuro-symbolic Learning Yielding Logical ConstraintsabstractNeuro-symbolic systems combine the abilities of neural perception and logical reasoning. However, end-to-end learning of neuro-symbolic systems is still an unsolved challenge. This paper proposes a natural framework that fuses neural network training, symbol grounding, and logical constraint synthesis into a coherent and efficient end-to-end learning process. The capability of this framework comes from the improved interactions between the neural and the symbolic parts of the system in both the training and inference stages. Technically, to bridge the gap between the continuous neural network and the discrete logical constraint, we introduce a difference-of-convex programming technique to relax the logical constraints while maintaining their precision. We also employ cardinality constraints as the language for logical constraint learning and incorporate a trust region method to avoid the degeneracy of logical constraint in learning. Both theoretical analyses and empirical evaluations substantiate the effectiveness of the proposed framework. Zenan Li, Yunpeng Huang, Yuan Yao 0001, Jingwei Xu 0001, Taolue Chen 0001, Xiaoxing Ma, Jian Lu 0001 |
NeurIPS | 4 |
| 2023 | Dynamic Data Fault Localization for Deep Neural NetworksabstractRich datasets have empowered various deep learning (DL) applications, leading to remarkable success in many fields. However, data faults hidden in the datasets could result in DL applications behaving unpredictably and even cause massive monetary and life losses. To alleviate this problem, in this paper, we propose a dynamic data fault localization approach, namely DFauLo, to locate the mislabeled and noisy data in the deep learning datasets. DFauLo is inspired by the conventional mutation-based code fault localization, but utilizes the differences between DNN mutants to amplify and identify the potential data faults. Specifically, it first generates multiple DNN model mutants of the original trained model. Then it extracts features from these mutants and maps them into a suspiciousness score indicating the probability of the given data being a data fault. Moreover, DFauLo is the first dynamic data fault localization technique, prioritizing the suspected data based on user feedback, and providing the generalizability to unseen data faults during training. To validate DFauLo, we extensively evaluate it on 26 cases with various fault types, data types, and model structures. We also evaluate DFauLo on three widely-used benchmark datasets. The results show that DFauLo outperforms the state-of-the-art techniques in almost all cases and locates hundreds of different types of real data faults in benchmark datasets. Yining Yin, Yang Feng 0003, Shihao Weng, Yuan Yao 0001, Zhenyu Chen 0001 |
ESEC/SIGSOFT FSE | 5 |
| 2023 | ImU: Physical Impersonating Attack for Face Recognition System with Natural Style ChangesabstractThis paper presents a novel physical impersonating attack against face recognition systems. It aims at generating consistent style changes across multiple pictures of the attacker under different conditions and poses. Additionally, the style changes are required to be physically realizable by make-up and can induce the intended misclassification. To achieve the goal, we develop novel techniques to embed multiple pictures of the same physical person to vectors in the StyleGAN’s latent space, such that the embedded latent vectors have some implicit correlations to make the search for consistent style changes feasible. Our digital and physical evaluation results show our approach can allow an outsider attacker to successfully impersonate the insiders with consistent and natural changes. Shengwei An, Yuan Yao 0001, Qiuling Xu, Shiqing Ma, Guanhong Tao 0001, Siyuan Cheng 0005, Kaiyuan Zhang 0002, Yingqi Liu, Guangyu Shen, Ian Kelk, Xiangyu Zhang 0001 |
SP | 2 |
| 2023 | MUSENET: Multi-Scenario Learning for Repeat-Aware Personalized RecommendationabstractPersonalized recommendation has been instrumental in many real applications. Despite the great progress, the underlying multi-scenario characteristics (e.g., users may behave differently under different scenarios) are largely ignored by existing recommender systems. Intuitively, modeling different scenarios properly could significantly improve the recommendation accuracy, and some existing work has explored this direction. However, these work assumes the scenarios are explicitly given, and thus becomes less effective when such information is unavailable. To complicate things further, proper scenario modeling from data is challenging and the recommendation models may easily overfit to some scenarios. In this paper, we propose a multi-scenario learning framework, MUSENET, for personalized recommendation. The key idea of MUSENET is to learn multiple implicit scenarios from the user behaviors, with a careful design inspired by the causal interpretation of recommender systems to avoid the overfitting issue. Additionally, since users' repeat consumptions account for a large part of the user behavior data on many e-commerce platforms, a repeat-aware mechanism is integrated to handle users' repurchase intentions within each scenario. Comprehensive experimental results on both industrial and public datasets demonstrate the effectiveness of the proposed approach compared with the state-of-the-art methods. Senrong Xu, Liangyue Li, Yuan Yao 0001, Zulong Chen, Hanghang Tong |
WSDM | 3 |
| 2023 | On the Vulnerability of Graph Learning-based Collaborative FilteringabstractGraph learning-based collaborative filtering (GLCF), which is built upon the message-passing mechanism of graph neural networks (GNNs), has received great recent attention and exhibited superior performance in recommender systems. However, although GNNs can be easily compromised by adversarial attacks as shown by the prior work, little attention has been paid to the vulnerability of GLCF. Questions like can GLCF models be just as easily fooled as GNNs remain largely unexplored. In this article, we propose to study the vulnerability of GLCF. Specifically, we first propose an adversarial attack against CLCF. Considering the unique challenges of attacking GLCF, we propose to adopt the greedy strategy in searching for the local optimal perturbations and design a reasonable attacking utility function to handle the non-differentiable ranking-oriented metrics. Next, we propose a defense to robustify GCLF. The defense is based on the observation that attacks usually introduce suspicious interactions into the graph to manipulate the message-passing process. We then propose to measure the suspicious score of each interaction and further reduce the message weight of suspicious interactions. We also give a theoretical guarantee of its robustness. Experimental results on three benchmark datasets show the effectiveness of both our attack and defense. Senrong Xu, Liangyue Li, Zenan Li, Yuan Yao 0001, Feng Xu 0007, Zulong Chen, Hanghang Tong |
ACM Trans. Inf. Syst. | 4 |
| 2023 | Towards Robust Neural Graph Collaborative Filtering via Structure Denoising and Embedding PerturbationabstractNeural graph collaborative filtering has received great recent attention due to its power of encoding the high-order neighborhood via the backbone graph neural networks. However, their robustness against noisy user-item interactions remains largely unexplored. Existing work on robust collaborative filtering mainly improves the robustness by denoising the graph structure, while recent progress in other fields has shown that directly adding adversarial perturbations in the embedding space can significantly improve the model robustness. In this work, we propose to improve the robustness of neural graph collaborative filtering via both denoising in the structure space and perturbing in the embedding space. Specifically, in the structure space, we measure the reliability of interactions and further use it to affect the message propagation process of the backbone graph neural networks; in the embedding space, we add in-distribution perturbations by mimicking the behavior of adversarial attacks and further combine it with contrastive learning to improve the performance. Extensive experiments have been conducted on four benchmark datasets to evaluate the effectiveness and efficiency of the proposed approach. The results demonstrate that the proposed approach outperforms the recent neural graph collaborative filtering methods especially when there are injected noisy interactions in the training data. Haibo Ye, Xinjie Li 0007, Yuan Yao 0001, Hanghang Tong |
ACM Trans. Inf. Syst. | 3 |
| 2022 | An Invisible Black-Box Backdoor Attack Through Frequency Domain
Yuan Yao 0001, Feng Xu 0007, Shengwei An, Hanghang Tong, Ting Wang 0006 |
ECCV (13) | 2 |
| 2022 | DescribeCtx: Context-Aware Description Synthesis for Sensitive Behaviors in Mobile AppsabstractWhile mobile applications (i.e., apps) are becoming capable of handling various needs from users, their increasing access to sensitive data raises privacy concerns. To inform such sensitive behaviors to users, existing techniques propose to automatically identify explanatory sentences from app descriptions; however, many sensitive behaviors are not explained in the corresponding app descriptions. There also exist general techniques that translate code to sentences. However, these techniques lack the vocabulary to explain the uses of sensitive data and fail to consider the context (i.e., the app functionalities) of the sensitive behaviors. To address these limitations, we propose DescribeCtx, a context-aware description synthesis approach that trains a neural machine translation model using a large set of popular apps, and generates app-specific descriptions for sensitive behaviors. Specifically, DescribeCtx encodes three heterogeneous sources as input, i.e., vocabularies provided by privacy policies, behavior summary provided by the call graphs in code, and contextual information provided by GUI texts. Our evaluations on 1,262 Android apps show that, compared with existing baselines, DescribeCtx produces more accurate descriptions (24.96 in BLEU) and achieves higher user ratings with respect to the reference sentences manually identified in the app descriptions. Shao Yang, Yuehan Wang, Yuan Yao 0001, Haoyu Wang 0001, Yanfang Ye 0001, Xusheng Xiao |
ICSE | 3 |
| 2022 | Untangling Composite Commits by Attributed Graph ClusteringabstractDuring software development, it is considered to be a best practice if each commit represents one distinct concern, such as fixing a bug or adding a new feature. However, developers may not always follow this practice and sometimes tangle multiple concerns into a single composite commit. This makes automatic commit untangling a necessary task, and recent approaches mainly untangle commits via applying graph clustering on the code dependency graph. In this paper, we propose a new commit untangling approach, ComUnt, to decompose the composite commits into atomic ones. Different from existing approaches, ComUnt is built upon the observation that both the textual content of code statements and the dependencies between code statements contain useful semantic information so as to better comprehend the committed code changes. Based on this observation, ComUnt first constructs an attributed graph for each commit, where code statements and various code dependencies are modeled as nodes and edges, respectively, and the textual body of code statements are maintained as node attributes. It then conducts attributed graph clustering on the constructed graph. The used attributed graph clustering algorithm can simultaneously encode both graph structure and node attributes so as to better separate the code changes into clusters with distinct concerns. We evaluate our approach on nine C# projects, and the experimental result shows that ComUnt improves the state-of-the-art by 7.8% in terms of untangling accuracy, and meanwhile it is more than 6 times faster. Shengbin Xu, Yuan Yao 0001, Feng Xu 0007 |
Internetware | 3 |
| 2022 | Combining Code Context and Fine-grained Code Difference for Commit Message GenerationabstractGenerating natural language messages for source code changes is an essential task in software development and maintenance. Existing solutions mainly treat a piece of code difference as natural language, and adopt seq2seq learning to translate it into a commit message. The basic assumption of such solutions lies in the naturalness hypothesis, i.e., source code written by programming languages is to some extent similar to natural language text. However, compared with natural language, source code also bears syntactic regularities. In this paper, we propose to simultaneously model the naturalness and syntactic regularities of source code changes for commit message generation. Specifically, to model syntactic regularities, we first enlarge the input with additional context information, i.e., the code statements that have dependency with the variables in the code difference, and then extract the paths in the corresponding ASTs. Moreover, to better model code difference, we align the two versions of code before and after the committed code change at token level, and annotate their differences with fine-grained edit operations. The context and difference are simultaneously encoded in a learning framework to generate the commit messages. We collected from GitHub a large dataset containing 480 Java projects with over 160k commits, and the experimental results demonstrate the effectiveness of the proposed approach. Shengbin Xu, Yuan Yao 0001, Feng Xu 0007, Tianxiao Gu, Hanghang Tong |
Internetware | 2 |
| 2022 | ADEPT: A Testing Platform for Simulated Autonomous DrivingabstractEffective quality assurance methods for autonomous driving systems ADS have attracted growing interests recently. In this paper, we report a new testing platform ADEPT, aiming to provide practically realistic and comprehensive testing facilities for DNN-based ADS. ADEPT is based on the virtual simulator CARLA and provides numerous testing facilities such as scene construction, ADS importation, test execution and recording, etc. In particular, ADEPT features two distinguished test scenario generation strategies designed for autonomous driving. First, we make use of real-life accident reports from which we leverage natural language processing to fabricate abundant driving scenarios. Second, we synthesize physically-robust adversarial attacks by taking the feedback of ADS into consideration and thus are able to generate closed-loop test scenarios. The experiments confirm the efficacy of the platform. Zhuheng Sheng, Jingwei Xu 0001, Taolue Chen 0001, Junjun Zhu, Yuan Yao 0001, Xiaoxing Ma |
ASE | 7 |
| 2022 | Fair Representation Learning: An Alternative to Mutual InformationabstractLearning fair representations is an essential task to reduce bias in data-oriented decision making. It protects minority subgroups by requiring the learned representations to be independent of sensitive attributes. To achieve independence, the vast majority of the existing work primarily relaxes it to the minimization of the mutual information between sensitive attributes and learned representations. However, direct computation of mutual information is computationally intractable, and various upper bounds currently used either are still intractable or contradict the utility of the learned representations. In this paper, we introduce distance covariance as a new dependence measure into fair representation learning. By observing that sensitive attributes (e.g., gender, race, and age group) are typically categorical, the distance covariance can be converted to a tractable penalty term without contradicting the utility desideratum. Based on the tractable penalty, we propose FairDisCo, a variational method to learn fair representations. Experiments demonstrate that FairDisCo outperforms existing competitors for fair representation learning. Zenan Li, Yuan Yao 0001, Feng Xu 0007, Xiaoxing Ma, Miao Xu 0001, Hanghang Tong |
KDD | 3 |
| 2022 | MIRROR: Model Inversion for Deep LearningNetwork with High Fidelity
Guanhong Tao 0001, Qiuling Xu, Yingqi Liu, Guangyu Shen, Shengwei An, Jingwei Xu 0001, Xiangyu Zhang 0001, Yuan Yao 0001 |
NDSS | 8 |
| 2022 | A Deep Learning Dataloader with Shared Data PreparationabstractExecuting a family of Deep Neural Networks (DNNs) training jobs on the same or similar datasets in parallel is typical in current deep learning scenarios. It is time-consuming and resource-intensive because each job repetitively prepares (i.e., loads and preprocesses) the data independently, causing redundant consumption of I/O and computations. Although the page cache or a centralized cache component can alleviate the redundancies by reusing the data prep work, each job's data sampled uniformly at random presents a low sampling locality in the shared dataset that causes the heavy cache thrashing. Prior work tries to solve the problem by enforcing all training jobs iterating over the dataset in the same order and requesting each data in lockstep, leading to strong constraints: all jobs must have the same dataset and run simultaneously. In this paper, we propose a dependent sampling algorithm (DSA) and domain-specific cache policy to relax the constraints. Besides, a novel tree data structure is designed to efficiently implement DSA. Based on the proposed technologies, we implemented a prototype system, named Joader, which can share data prep work as long as the datasets share partially. We evaluate the proposed Joader in practical scenarios, showing a greater versatility and superiority over training speed improvement (up to 500% in ResNet18). Jingwei Xu 0001, Guochang Wang, Yuan Yao 0001, Zenan Li, Chun Cao, Hanghang Tong |
NeurIPS | 4 |
| 2022 | Structure Meets Sequences: Predicting Network of Co-evolving SequencesabstractCo-evolving sequences are ubiquitous in a variety of applications, where different sequences are often inherently inter-connected with each other. We refer to such sequences, together with their inherent connections modeled as a structured network, as network of co-evolving sequences (NoCES). Typical NoCES applications include road traffic monitoring, company revenue prediction, motion capture, etc. To date, it remains a daunting challenge to accurately model NoCES due to the coupling between network structure and sequences. In this paper, we propose to modeling \pname\ with the aim of simultaneously capturing both the dynamics and the interplay between network structure and sequences. Specifically, we propose a joint learning framework to alternatively update the network representations and sequence representations as the sequences evolve over time. A unique feature of our framework lies in that it can deal with the case when there are co-evolving sequences on both network nodes and edges. Experimental evaluations on four real datasets demonstrate that the proposed approach (1) outperforms the existing competitors in terms of prediction accuracy, and (2) scales linearly w.r.t. the sequence length and the network size. Yaojing Wang, Yuan Yao 0001, Feng Xu 0007, Yada Zhu, Hanghang Tong |
WSDM | 2 |
| 2022 | Auditing Network Embedding: An Edge Influence Based ApproachabstractLearning node representations in a network has a wide range of applications. Most of the existing work focuses on improving the performance of the learned node representations by designing advanced network embedding models. In contrast to these work, this article aims to provide some understanding of the rationale behind the existing network embedding models, e.g.,whya given embedding algorithm outputs the specific node representations andhowthe resulting node representations relate to the structure of the input network. In particular, we propose to discern the edge influence for two widely-studied classes of network embedding models, i.e., skip-gram based models and graph neural networks. We provide algorithms to effectively and efficiently quantify the edge influence on node representations, and further identify high-influential edges by exploiting the linkage between edge influence and network structure. Experimental evaluations are conducted on real datasets showing that: 1) in terms of quantifying edge influence, the proposed method is significantly faster (up to$2,000\times$) than straightforward methods with little quality loss, and 2) in terms of identifying high-influential edges, the identified edges by the proposed method have a significant impact in the context of downstream prediction task and adversarial attacking. Yaojing Wang, Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Unsupervised Attributed Network Embedding via Cross FusionabstractAttributed network embedding aims to learn low dimensional node representations by combining both the network's topological structure and node attributes. Most of the existing methods either propagate the attributes over the network structure or learn the node representations by an encoder-decoder framework. However, propagation based methods tend to prefer network structure to node attributes, whereas encoder-decoder methods tend to ignore the longer connections beyond the immediate neighbors. In order to address these limitations while enjoying the best of the two worlds, we design cross fusion layers for unsupervised attributed network embedding. Specifically, we first construct two separate views to handle network structure and node attributes, and then design cross fusion layers to allow flexible information exchange and integration between the two views. The key design goals of the cross fusion layers are three-fold: 1) allowing critical information to be propagated along the network structure, 2) encoding the heterogeneity in the local neighborhood of each node during propagation, and 3) incorporating an additional node attribute channel so that the attribute information will not be overshadowed by the structure view. Extensive experiments on three datasets and three downstream tasks demonstrate the effectiveness of the proposed method. Guosheng Pan, Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
WSDM | 2 |
| 2020 | Bringing Order to Network Embedding: A Relative Ranking based ApproachabstractNetwork embedding aims to automatically learn the node representations in networks. The basic idea of network embedding is to first construct a network to describe the neighborhood context for each node, and then learn the node representations by designing an objective function to preserve certain properties of the constructed context network. The vast majority of the existing methods, explicitly or implicitly, follow a pointwise design principle. That is, the objective can be decomposed into the summation of the certain goodness function over each individual edge of the context network. In this paper, we propose to go beyond such pointwise approaches, and introduce the ranking-oriented design principle for network embedding. The key idea is to decompose the overall objective function into the summation of a goodness function over a set of edges to collectively preserve their relative rankings on the context network. We instantiate the ranking-oriented design principle by two new network embedding algorithms, including a pairwise network embedding method PaWine which optimizes the relative weights of edge pairs, and a listwise method LiWine which optimizes the relative weights of edge lists. Both proposed algorithms bear a linear time complexity, making themselves scalable to large networks. We conduct extensive experimental evaluations on five real datasets with a variety of downstream learning tasks, which demonstrate that the proposed approaches consistently outperform the existing methods. Yaojing Wang, Guosheng Pan, Yuan Yao 0001, Hanghang Tong, Hongxia Yang, Feng Xu 0007, Jian Lu 0001 |
CIKM | 3 |
| 2020 | Trading Personalization for Accuracy: Data Debugging in Collaborative FilteringabstractCollaborative filtering has been widely used in recommender systems. Existing work has primarily focused on improving the prediction accuracy mainly via either building refined models or incorporating additional side information, yet has largely ignored the inherent distribution of the input rating data. In this paper, we propose a data debugging framework to identify overly personalized ratings whose existence degrades the performance of a given collaborative filtering model. The key idea of the proposed approach is to search for a small set of ratings whose editing (e.g., modification or deletion) would near-optimally improve the recommendation accuracy of a validation set. Experimental results demonstrate that the proposed approach can significantly improve the recommendation accuracy. Furthermore, we observe that the identified ratings significantly deviate from the average ratings of the corresponding items, and the proposed approach tends to modify them towards the average. This result sheds light on the design of future recommender systems in terms of balancing between the overall accuracy and personalization. Yuan Yao 0001, Feng Xu 0007, Miao Xu 0001, Hanghang Tong |
NeurIPS | 2 |
| 2020 | Enhancing supervised bug localization with metadata and stack-trace
Yaojing Wang, Yuan Yao 0001, Hanghang Tong, Xuan Huo, Ming Li 0005, Feng Xu 0007, Jian Lu 0001 |
Knowl. Inf. Syst. | 2 |
| 2020 | Fast discrete factorization machine for personalized item recommendation
Shilin Qu, Guibing Guo, Yuan Liu 0002, Yuan Yao 0001, Wei Wei 0002 |
Knowl. Based Syst. | 4 |
| 2019 | An Integral Tag Recommendation Model for Textual ContentabstractRecommending suitable tags for online textual content is a key building block for better content organization and consumption. In this paper, we identify three pillars that impact the accuracy of tag recommendation: (1) sequential text modeling meaning that the intrinsic sequential ordering as well as different areas of text might have an important implication on the corresponding tag(s) , (2) tag correlation meaning that the tags for a certain piece of textual content are often semantically correlated with each other, and (3) content-tag overlapping meaning that the vocabularies of content and tags are overlapped. However, none of the existing methods consider all these three aspects, leading to a suboptimal tag recommendation. In this paper, we propose an integral model to encode all the three aspects in a coherent encoder-decoder framework. In particular, (1) the encoder models the semantics of the textual content via Recurrent Neural Networks with the attention mechanism, (2) the decoder tackles the tag correlation with a prediction path, and (3) a shared embedding layer and an indicator function across encoder-decoder address the content-tag overlapping. Experimental results on three realworld datasets demonstrate that the proposed method significantly outperforms the existing methods in terms of recommendation accuracy. Shijie Tang, Yuan Yao 0001, Suwei Zhang, Feng Xu 0007, Tianxiao Gu, Hanghang Tong, Jian Lu 0001 |
AAAI | 2 |
| 2019 | Hashtag Recommendation for Photo Sharing ServicesabstractHashtags can greatly facilitate content navigation and improve user engagement in social media. Meaningful as it might be, recommending hashtags for photo sharing services such as Instagram and Pinterest remains a daunting task due to the following two reasons. On the endogenous side, posts in photo sharing services often contain both images and text, which are likely to be correlated with each other. Therefore, it is crucial to coherently model both image and text as well as the interaction between them. On the exogenous side, hashtags are generated by users and different users might come up with different tags for similar posts, due to their different preference and/or community effect. Therefore, it is highly desirable to characterize the users’ tagging habits. In this paper, we propose an integral and effective hashtag recommendation approach for photo sharing services. In particular, the proposed approach considers both the endogenous and exogenous effects by a content modeling module and a habit modeling module, respectively. For the content modeling module, we adopt the parallel co-attention mechanism to coherently model both image and text as well as the interaction between them; for the habit modeling module, we introduce an external memory unit to characterize the historical tagging habit of each user. The overall hashtag recommendations are generated on the basis of both the post features from the content modeling module and the habit influences from the habit modeling module. We evaluate the proposed approach on real Instagram data. The experimental results demonstrate that the proposed approach significantly outperforms the state-of-theart methods in terms of recommendation accuracy, and that both content modeling and habit modeling contribute significantly to the overall recommendation accuracy. Suwei Zhang, Yuan Yao 0001, Feng Xu 0007, Hanghang Tong, Jian Lu 0001 |
AAAI | 2 |
| 2019 | DeepIntent: Deep Icon-Behavior Learning for Detecting Intention-Behavior Discrepancy in Mobile AppsabstractMobile apps have been an indispensable part in our daily life. However, there exist many potentially harmful apps that may exploit users' privacy data, e.g., collecting the user's information or sending messages in the background. Keeping these undesired apps away from the market is an ongoing challenge. While existing work provides techniques to determine what apps do, e.g., leaking information, little work has been done to answer, are the apps' behaviors compatible with the intentions reflected by the app's UI? In this work, we explore the synergistic cooperation of deep learning and program analysis as the first step to address this challenge. Specifically, we focus on the UI widgets that respond to user interactions and examine whether the intentions reflected by their UIs justify their permission uses. We present DeepIntent, a framework that uses novel deep icon-behavior learning to learn an icon-behavior model from a large number of popular apps and detect intention-behavior discrepancies. In particular, DeepIntent provides program analysis techniques to associate the intentions (i.e., icons and contextual texts) with UI widgets' program behaviors, and infer the labels (i.e., permission uses) for the UI widgets based on the program behaviors, enabling the construction of a large-scale high-quality training dataset. Based on the results of the static analysis, DeepIntent uses deep learning techniques that jointly model icons and their contextual texts to learn an icon-behavior model, and detects intention-behavior discrepancies by computing the outlier scores based on the learned model. We evaluate DeepIntent on a large-scale dataset (9,891 benign apps and 16,262 malicious apps). With 80% of the benign apps for training and the remaining for evaluation, DeepIntent detects discrepancies with AUC scores 0.8656 and 0.8839 on benign apps and malicious apps, achieving 39.9% and 26.1% relative improvements over the state-of-the-art approaches. Shengqu Xi, Shao Yang, Xusheng Xiao, Yuan Yao 0001, Yayuan Xiong, Fengyuan Xu, Haoyu Wang 0001, Peng Gao 0008, Zhuotao Liu, Feng Xu 0007, Jian Lu 0001 |
CCS | 4 |
| 2019 | Discerning Edge Influence for Network EmbeddingabstractNetwork embedding, which learns the low-dimensional representations of nodes, has gained significant research attention. Despite its superior empirical success, often measured by the prediction performance of downstream tasks (e.g., multi-label classification), it is unclear \em why a given embedding algorithm outputs the specific node representations, and \em how the resulting node representations relate to the structure of the input network. In this paper, we propose to discern the edge influence as the first step towards understanding skip-gram basd network embedding methods. For this purpose, we propose an auditing framework Near, whose key part includes two algorithms (Near-add \ and Near-del ) to effectively and efficiently quantify the influence of each edge. Based on the algorithms, we further identify high-influential edges by exploiting the linkage between edge influence and the network structure. Experimental results demonstrate that the proposed algorithms (Near-add \ and Near-del ) are significantly faster (up to $2,000\times$) than straightforward methods with little quality loss. Moreover, the proposed framework can efficiently identify the most influential edges for network embedding in the context of downstream prediction task and adversarial attacking. Yaojing Wang, Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
CIKM | 2 |
| 2019 | Practical GUI testing of Android applications via model abstraction and refinementabstractThis paper introduces a new, fully automated modelbased approach for effective testing of Android apps. Different from existing model-based approaches that guide testing with a static GUI model (i.e., the model does not evolve its abstraction during testing, and is thus often imprecise), our approach dynamically optimizes the model by leveraging the runtime information during testing. This capability of model evolution significantly improves model precision, and thus dramatically enhances the testing effectiveness compared to existing approaches, which our evaluation confirms.We have realized our technique in a practical tool, APE. On 15 large, widely-used apps from the Google Play Store, APE outperforms the state-of-the-art Android GUI testing tools in terms of both testing coverage and the number of detected unique crashes. To further demonstrate APE's effectiveness and usability, we conduct another evaluation of APE on 1,316 popular apps, where it found 537 unique crashes. Out of the 38 reported crashes, 13 have been fixed and 5 have been confirmed. Tianxiao Gu, Chengnian Sun, Xiaoxing Ma, Chun Cao, Chang Xu 0001, Yuan Yao 0001, Qirun Zhang, Jian Lu 0001, Zhendong Su 0001 |
ICSE | 6 |
| 2019 | Commit Message Generation for Source Code ChangesabstractCommit messages, which summarize the source code changes in natural language, are essential for program comprehension and software evolution understanding. Unfortunately, due to the lack of direct motivation, commit messages are sometimes neglected by developers, making it necessary to automatically generate such messages. State-of-the-art adopts learning based approaches such as neural machine translation models for the commit message generation problem. However, they tend to ignore the code structure information and suffer from the out-of-vocabulary issue. In this paper, we propose CoDiSum to address the above two limitations. In particular, we first extract both code structure and code semantics from the source code changes, and then jointly model these two sources of information so as to better learn the representations of the code changes. Moreover, we augment the model with copying mechanism to further mitigate the out-of-vocabulary issue. Experimental evaluations on real data demonstrate that the proposed approach significantly outperforms the state-of-the-art in terms of accurately generating the commit messages. Shengbin Xu, Yuan Yao 0001, Feng Xu 0007, Tianxiao Gu, Hanghang Tong, Jian Lu 0001 |
IJCAI | 2 |
| 2019 | Speedup Automatic Program Repair Using Dynamic Software Updating: An Empirical StudyabstractA typical generate-and-validate automatic program repair (APR) tool needs to repeatedly run the same test suite to validate each generated patch. This procedure is expensive when the number of patches is huge. Additionally, to scale to large programs, a program repair tool has to consider a small patch space in practice and thus may sacrifice the capability to find potential correct repairs. In this work, we propose to speed up automatic program repair to mitigate the above issues. One the one hand, we found that restarting processes to load patched code consumes the majority of total validation time. This problem is even severe when the program is running in a managed runtime such as Java virtual machine (JVM). On the other hand, dynamic software updating (DSU) can load and execute new code without restarting. To this end, we propose to use DSU techniques to speed up automatic program repair and present an empirical study in this paper. Within our study, DSU can bring up to 66.3 times speedup in comparison with the traditional restart approach. However, DSU may not be able to handle all patches and can also incur unknown side effects that lead to inconsistent validation results. We then further study the feasibility and consistency of applying DSU to speed up APR. Our results show that 1) less than 1% patches cannot be dynamically updated using the builtin DSU ability of JVM, and 2) DSU based validation leads to potentially harmful inconsistency in only 16 of 1,897,518 patches. Rongxun Guo, Tianxiao Gu, Yuan Yao 0001, Feng Xu 0007, Xiaoxing Ma |
Internetware | 3 |
| 2019 | Bug Triaging Based on Tossing Sequence Modeling
Shengqu Xi, Yuan Yao 0001, Xusheng Xiao, Feng Xu 0007, Jian Lu 0001 |
J. Comput. Sci. Technol. | 2 |
| 2019 | Dual-regularized one-class collaborative filtering with implicit feedback
Yuan Yao 0001, Hanghang Tong, Guo Yan, Feng Xu 0007, Xiang Zhang 0001, Boleslaw K. Szymanski, Jian Lu 0001 |
World Wide Web | 1 |
| 2018 | Bug Localization via Supervised Topic ModelingabstractBug tracking systems, which help to track the reported software bugs, have been widely used in software development and maintenance. In these systems, recognizing relevant source files among a large number of source files for a given bug report is a time-consuming and labor-intensive task for software developers. To tackle this problem, information retrieval methods have been widely used to capture either the textual similarities or the semantic similarities between bug reports and source files. However, these two types of similarities are usually considered separately and the historical bug fixings are largely ignored by the existing methods. In this paper, we propose a supervised topic modeling method (STMLOCATOR) for automatically locating the relevant source files for a given bug report. In particular, the proposed model is built upon three key observations. First, supervised modeling can effectively make use of the existing fixing histories. Second, certain words in bug reports tend to appear multiple times in their relevant source files. Third, longer source files tend to have more bugs. By integrating the above three observations, the proposed STMLOCATOR utilizes historical fixings in a supervised way and learns both the textual similarities and semantic similarities between bug reports and source files. We further consider a special type of bug reports with stack-traces in bug reports, and propose a variant of STMLOCATOR to tailor for such bug reports. Experimental evaluations on three real data sets demonstrate that the proposed STMLOCATOR can achieve up to 23.6% improvement in terms of prediction accuracy over its best competitors, and scales linearly with the size of the data. Moreover, the proposed variant further improves STMLOCATOR by up to 76.2% on those bug reports with stack-traces. Yaojing Wang, Yuan Yao 0001, Hanghang Tong, Xuan Huo, Feng Xu 0007, Jian Lu 0001 |
ICDM | 2 |
| 2018 | An Effective Approach for Routing the Bug Reports to the Right FixersabstractRouting the bug reports to potential fixers (i.e., bug triaging), is an integral step in software development and maintenance. However, manually inspecting and assigning bug reports is tedious and time-consuming, especially in those software projects that have a large amount of bug reports and developers. To make bug triaging more efficient, many machine learning and information retrieval based approaches have been proposed to automatically assign bug reports for suitable developers to fix. However, these techniques typically ignore two important facts in bug fixing. First, for some bug reports, the bug reporter himself/herself is one of the developers in the project, and he/she is likely to fix his/her reported bugs in the future. Second, for some bug reports, there may be a tossing sequence which contains several developers from the first potential fixer to the last actual fixer. Such tossing sequences encode valuable information such as the dependency of developers for the bug triaging task. To make use of the above facts, we propose a sequence to sequence model named SeqTriage to automatically route a given bug report to its responsible fixer. Evaluation results on three different open-source projects show that the proposed approach has significantly improved the accuracy of bug triaging compared with the state-of-the-art approaches (20% at best and 5% at least). Shengqu Xi, Yuan Yao 0001, Xusheng Xiao, Feng Xu 0007, Jian Lu 0001 |
Internetware | 2 |
| 2018 | Team Expansion in Collaborative Environments
Yuan Yao 0001, Guibing Guo, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
PAKDD (3) | 2 |
| 2018 | Guiding supervised topic modeling for content based tag recommendation
Shengqu Xi, Yuan Yao 0001, Feng Xu 0007, Hanghang Tong, Jian Lu 0001 |
Neurocomputing | 3 |
| 2017 | Exploring Metadata in Bug Reports for Bug LocalizationabstractInformation retrieval methods have been proposed to help developers locate related buggy source files for a given bug report. The basic assumption of these methods is that the bug description in a bug report should be textually similar to its buggy source files. However, the metadata (such as the component and version information) in bug reports is largely ignored by these methods. In this paper, we propose to explore the metadata for the bug localization task. In particular, we first apply a generative model to locate buggy source files based on the bug descriptions, and then propose to add the available metadata in bug reports into the localization process. Experimental evaluations on several software projects indicate that the metadata is useful to improve the localization accuracy and that the proposed bug localization method outperforms several existing methods. Yuan Yao 0001, Yaojing Wang, Feng Xu 0007, Jian Lu 0001 |
APSEC | 2 |
| 2017 | Personalized travel mode detection with smartphone sensorsabstractDetecting the travel modes such as walking and driving a car is an important task for user behavior understanding as well as transportation planning and management. Existing solutions for this task mainly train a generic classifier for all users although the walking or driving behaviors may differ greatly from one user to another. In this paper, we propose to build a personalized travel mode detection method. In particular, the proposed method can be divided into two stages. First, for a given target user, it applies user similarity computation to borrow data from a set of pre-collected data for transfer learning. Second, it estimates the data distribution in feature space, and uses it to reweight the borrowed data so as to minimize the model loss with respect to the target user. Experimental evaluations on real travel data show that the proposed method outperforms the generic method and the transfer learning method with kernel mean matching in terms of prediction accuracy. Xing Su 0002, Yuan Yao 0001, Qing He 0011, Hanghang Tong |
IEEE BigData | 2 |
| 2017 | HoORaYs: High-order Optimization of Rating Distance for Recommender SystemsabstractLatent factor models have become a prevalent method in recommender systems, to predict users' preference on items based on the historical user feedback. Most of the existing methods, explicitly or implicitly, are built upon the first-order rating distance principle, which aims to minimize the difference between the estimated and real ratings. In this paper, we generalize such first-order rating distance principle and propose a new latent factor model (HoORaYs) for recommender systems. The core idea of the proposed method is to explore high-order rating distance, which aims to minimize not only (i) the difference between the estimated and real ratings of the same (user, item) pair (i.e., the first-order rating distance), but also (ii) the difference between the estimated and real rating difference of the same user across different items (i.e., the second-order rating distance). We formulate it as a regularized optimization problem, and propose an effective and scalable algorithm to solve it. Our analysis from the geometry and Bayesian perspectives indicate that by exploring the high-order rating distance, it helps to reduce the variance of the estimator, which in turns leads to better generalization performance (e.g., smaller prediction error). We evaluate the proposed method on four real-world data sets, two with explicit user feedback and the other two with implicit user feedback. Experimental results show that the proposed method consistently outperforms the state-of-the-art methods in terms of the prediction accuracy. Jingwei Xu 0001, Yuan Yao 0001, Hanghang Tong, XianPing Tao, Jian Lu 0001 |
KDD | 2 |
| 2017 | RaPare: A Generic Strategy for Cold-Start Rating Prediction ProblemabstractIn recent years, recommender system is one of indispensable components in many e-commerce websites. One of the major challenges that largely remains open is the cold-start problem, which can be viewed as a barrier that keeps the cold-start users/items away from the existing ones. In this paper, we aim to break through this barrier for cold-start users/items by the assistance of existing ones. In particular, inspired by the classic Elo Rating System, which has been widely adopted in chess tournaments, we propose a novel rating comparison strategy (RAPARE) to learn the latent profiles of cold-start users/items. The centerpiece of our RAPARE is to provide a fine-grained calibration on the latent profiles of cold-start users/items by exploring the differences between cold-start and existing users/items. As a generic strategy, our proposed strategy can be instantiated into existing methods in recommender systems. To reveal the capability of RAPARE strategy, we instantiate our strategy on two prevalent methods in recommender systems, i.e., the matrix factorization based and neighborhood based collaborative filtering. Experimental evaluations on five real data sets validate the superiority of our approach over the existing methods in cold-start scenario. Jingwei Xu 0001, Yuan Yao 0001, Hanghang Tong, XianPing Tao, Jian Lu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Scalable Algorithms for CQA Post Voting PredictionabstractCommunity Question Answering (CQA) sites, such as Stack Overflow and Yahoo! Answers, have become very popular in recent years. These sites contain rich crowdsourcing knowledge contributed by the site users in the form of questions and answers, and these questions and answers can satisfy the information needs of more users. In this article, we aim at predicting the voting scores of questions/answers shortly after they are posted in the CQA sites. To accomplish this task, we identify three key aspects that matter with the voting of a post, i.e., the non-linear relationships between features and output, the question and answer coupling, and the dynamic fashion of data arrivals. A family of algorithms are proposed to model the above three key aspects. Some approximations and extensions are also proposed to scale up the computation. We analyze the proposed algorithms in terms of optimality, correctness, and complexity. Extensive experimental evaluations conducted on two real data sets demonstrate the effectiveness and efficiency of our algorithms. Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2017 | Version-Aware Rating Prediction for Mobile App RecommendationabstractWith the great popularity of mobile devices, the amount of mobile apps has grown at a more dramatic rate than ever expected. A technical challenge is how to recommend suitable apps to mobile users. In this work, we identify and focus on a unique characteristic that exists in mobile app recommendation—that is, an app usually corresponds to multiple release versions. Based on this characteristic, we propose a fine-grain version-aware app recommendation problem. Instead of directly learning the users’ preferences over the apps, we aim to infer the ratings of users on a specific version of an app. However, the user-version rating matrix will be sparser than the corresponding user-app rating matrix, making existing recommendation methods less effective. In view of this, our approach has made two major extensions. First, we leverage the review text that is associated with each rating record; more importantly, we consider two types of version-based correlations. The first type is to capture the temporal correlations between multiple versions within the same app, and the second type of correlation is to capture the aggregation correlations between similar apps. Experimental results on a large dataset demonstrate the superiority of our approach over several competitive methods. Yuan Yao 0001, Wayne Xin Zhao, Yaojing Wang, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2016 | Tag2Word: Using Tags to Generate Words for Content Based Tag RecommendationabstractTag recommendation is helpful for the categorization and searching of online content. Existing tag recommendation methods can be divided into collaborative filtering methods and content based methods. In this paper, we put our focus on the content based tag recommendation due to its wider applicability. Our key observation is the tag-content co-occurrence, i.e., many tags have appeared multiple times in the corresponding content. Based on this observation, we propose a generative model (Tag2Word), where we generate the words based on the tag-word distribution as well as the tag itself. Experimental evaluations on real data sets demonstrate that the proposed method outperforms several existing methods in terms of recommendation accuracy, while enjoying linear scalability. Yuan Yao 0001, Feng Xu 0007, Hanghang Tong, Jian Lu 0001 |
CIKM | 2 |
| 2016 | QUINT: On Query-Specific Optimal NetworksabstractMeasuring node proximity on large scale networks is a fundamental building block in many application domains, ranging from computer vision, e-commerce, social networks, software engineering, disaster management to biology and epidemiology. The state of the art (e.g., random walk based methods) typically assumes the input network is given a priori, with the known network topology and the associated edge weights. A few recent works aim to further infer the optimal edge weights based on the side information. This paper generalizes the challenge in multiple dimensions, aiming to learn optimal networks for node proximity measures. First (optimization scope), our proposed formulation explores a much larger parameter space, so that it is able to simultaneously infer the optimal network topology and the associated edge weights. This is important as a noisy or missing edge could greatly mislead the network node proximity measures. Second (optimization granularity), while all the existing works assume one common optimal network, be it given as the input or learned by the algorithms, exists for all queries, our method performs optimization at a much finer granularity, essentially being able to infer an optimal network that is specific to a given query. Third (optimization efficiency), we carefully design our algorithms with a linear complexity wrt the neighborhood size of the user preference set. We perform extensive empirical evaluations on a diverse set of 10+ real networks, which show that the proposed algorithms (1) consistently outperform the existing methods on all six commonly used metrics; (2) empirically scale sub-linearly to billion-scale networks and (3) respond in a fraction of a second. Liangyue Li, Yuan Yao 0001, Jie Tang 0001, Wei Fan 0001, Hanghang Tong |
KDD | 2 |
| 2016 | Large-Scale Off-Target Identification Using Fast and Accurate Dual Regularized One-Class Collaborative Filtering and Its Application to Drug RepurposingabstractTarget-based screening is one of the major approaches in drug discovery. Besides the intended target, unexpected drug off-target interactions often occur, and many of them have not been recognized and characterized. The off-target interactions can be responsible for either therapeutic or side effects. Thus, identifying the genome-wide off-targets of lead compounds or existing drugs will be critical for designing effective and safe drugs, and providing new opportunities for drug repurposing. Although many computational methods have been developed to predict drug-target interactions, they are either less accurate than the one that we are proposing here or computationally too intensive, thereby limiting their capability for large-scale off-target identification. In addition, the performances of most machine learning based algorithms have been mainly evaluated to predict off-target interactions in the same gene family for hundreds of chemicals. It is not clear how these algorithms perform in terms of detecting off-targets across gene families on a proteome scale. Here, we are presenting a fast and accurate off-target prediction method, REMAP, which is based on a dual regularized one-class collaborative filtering algorithm, to explore continuous chemical space, protein space, and their interactome on a large scale. When tested in a reliable, extensive, and cross-gene family benchmark, REMAP outperforms the state-of-the-art methods. Furthermore, REMAP is highly scalable. It can screen a dataset of 200 thousands chemicals against 20 thousands proteins within 2 hours. Using the reconstructed genome-wide target profile as the fingerprint of a chemical compound, we predicted that seven FDA-approved drugs can be repurposed as novel anti-cancer therapies. The anti-cancer activity of six of them is supported by experimental evidences. Thus, REMAP is a valuable addition to the existing in silico toolbox for drug target identification, drug repurposing, phenotypic screening, and side effect prediction. The software and benchmark are available at https://github.com/hansaimlim/REMAP. Hansaim Lim, Aleksandar Poleksic, Yuan Yao 0001, Hanghang Tong, Di He 0003, Luke Zhuang, Patrick Meng, Lei Xie 0006 |
PLoS Comput. Biol. | 3 |
| 2015 | MATAR: Keywords Enhanced Multi-label Learning for Tag Recommendation
Yuan Yao 0001, Feng Xu 0007, Jian Lu 0001 |
APWeb | 2 |
| 2015 | Ice-Breaking: Mitigating Cold-Start Recommendation Problem by Rating Comparison
Jingwei Xu 0001, Yuan Yao 0001, Hanghang Tong, XianPing Tao, Jian Lu 0001 |
IJCAI | 2 |
| 2015 | Detecting Buggy Files based on Bug Reports: A Random Walk Based ApproachabstractDuring an Open Source Software (OSS) maintenance, bug localization is a laborsome and time-consuming work for geographically-separated developers. It is desirable to automatically identify related buggy files once a bug report is submitted. Existing work has proposed information retrieval techniques for this problem. However, these proposals tend to neglect the inherent structure in bug localization and they are largely dependent on the comments (annotations) in source files which may be unavailable. In this paper, we propose a random walk based approach to detecting buggy source files based on bug reports. In particular, we separately process source files and bug reports to make it less sensitive to comments, and then apply random walk to capture the inherent structure in bug localization. Experimental evaluations on three real-world open-source projects demonstrate that the proposed approach can outperform several existing methods and that it is less dependent on the comments in source files. Yaojing Wang, Feng Xu 0007, Yuan Yao 0001 |
Internetware | 3 |
| 2015 | RIT: Enhancing Recommendation with Inferred Trust
Guo Yan, Yuan Yao 0001, Feng Xu 0007, Jian Lu 0001 |
PAKDD (2) | 2 |
| 2015 | Detecting high-quality posts in community question answering sites
Yuan Yao 0001, Hanghang Tong, Tao Xie 0001, Leman Akoglu, Feng Xu 0007, Jian Lu 0001 |
Inf. Sci. | 1 |
| 2014 | Joint voting prediction for questions and answers in CQAabstractCommunity Question Answering (CQA) sites have become valuable repositories that host a massive volume of human knowledge. How can we detect a high-value answer which clears the doubts of many users? Can we tell the user if the question s/he is posting would attract a good answer? In this paper, we aim to answer these questions from the perspective of the voting outcome by the site users. Our key observation is that the voting score of an answer is strongly positively correlated with that of its question, and such correlation could be in turn used to boost the prediction performance. Armed with this observation, we propose a family of algorithms to jointly predict the voting scores of questions and answers soon after they are posted in the CQA sites. Experimental evaluations demonstrate the effectiveness of our approaches. Yuan Yao 0001, Hanghang Tong, Tao Xie 0001, Leman Akoglu, Feng Xu 0007, Jian Lu 0001 |
ASONAM | 1 |
| 2014 | Dual-Regularized One-Class Collaborative FilteringabstractCollaborative filtering is a fundamental building block in many recommender systems. While most of the existing collaborative filtering methods focus on explicit, multi-class settings (e.g., 1-5 stars in movie recommendation), many real-world applications actually belong to the one-class setting where user feedback is implicitly expressed (e.g., views in news recommendation and video recommendation). The main challenges in such one-class setting include the ambiguity of the unobserved examples and the sparseness of existing positive examples. Yuan Yao 0001, Hanghang Tong, Guo Yan, Feng Xu 0007, Xiang Zhang 0001, Boleslaw K. Szymanski, Jian Lu 0001 |
CIKM | 1 |
| 2014 | Predicting long-term impact of CQA posts: a comprehensive viewpointabstractCommunity Question Answering (CQA) sites have become valuable platforms to create, share, and seek a massive volume of human knowledge. How can we spot an insightful question that would inspire massive further discussions in CQA sites? How can we detect a valuable answer that benefits many users? The long-term impact (e.g., the size of the population a post benefits) of a question/answer post is the key quantity to answer these questions. In this paper, we aim to predict the long-term impact of questions/answers shortly after they are posted in the CQA sites. In particular, we propose a family of algorithms for the prediction problem by modeling three key aspects, i.e., non-linearity, question/answer coupling, and dynamics. We analyze our algorithms in terms of optimality, correctness, and complexity. We conduct extensive experimental evaluations on two real CQA data sets to demonstrate the effectiveness and efficiency of our algorithms. Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
KDD | 1 |
| 2014 | Exploring Review Content for Recommendation via Latent Factor Model
Yuan Yao 0001, Feng Xu 0007, Jian Lu 0001 |
PRICAI | 2 |
| 2014 | A Parallel Approach to Link Sign Prediction in Large-Scale Online Social NetworksabstractAnalyzing the underlying social network is very important for the development of online applications. Owing to the increasingly growing size of these networks, parallel techniques play important roles in many network analysis tasks. In this paper, we explore the link sign prediction problem in large-scale online social networks, and propose a parallel approach, called PLSP, to solve the problem. Specifically, we first extract a set of features that serve as a base for prediction. Experiments on several real datasets show that these features outperform those proposed by existing methods in predictive accuracy. Next, we present two speedup strategies, i.e. dataset division and feature selection, to shorten the training time. Experimental evaluations show that our parallel approach is much faster than the traditional non-parallel method and achieves higher predictive accuracy than other methods at the same time. Jiufeng Zhou, Lixin Han, Yuan Yao 0001, Xiaoqin Zeng, Feng Xu 0007 |
Comput. J. | 3 |
| 2014 | Multi-Aspect + Transitivity + Bias: An Integral Trust Inference ModelabstractInferring the pair-wise trust relationship is a core building block for many real applications. State-of-the-art approaches for such trust inference mainly employ the transitivity property of trust by propagating trust along connected users, but largely ignore other important properties such as trust bias, multi-aspect, etc. In this paper, we propose a new trust inference model to integrate all these important properties. To apply the model to both binary and continuous inference scenarios, we further propose a family of effective and efficient algorithms. Extensive experimental evaluations on real data sets show that our method achieves significant improvement over several existing benchmark approaches, for both quantifying numerical trustworthiness scores and predicting binary trust/distrust signs. In addition, it enjoys linear scalability in both time and space. Yuan Yao 0001, Hanghang Tong, Xifeng Yan, Feng Xu 0007, Jian Lu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2013 | Enhancing trustworthiness evaluation in internetware with similarity and non-negative constraintsabstractInternetware is envisioned as a new software paradigm where software developers usually need to interact with unknown partners as well as the software entities developed by them. To reduce uncertainty and boost collaborations in such setting, it is important to provide trustworthiness evaluation mechanisms so that trustworthy partners/entities can be easily found. In this work, we propose a novel trustworthiness evaluation mechanism by enhancing existing mechanisms with similarity and non-negative constraints. To be specific, we first extend an existing multi-aspect trust inference model by incorporating the non-negative constraint. One of the advantages of such constraint is its strong interpretability. Second, we incorporate similarity into two neighborhood models borrowed from recommender systems. When computing similarity, we make use of the intermediate results from the first step. Finally, these models are combined under a machine learning framework. To show the effectiveness of our method, we conduct experiments on a real data-set. The results show that: both our non-negativity extension and similarity computation improve the evaluation accuracy of the original methods, and the combined method outperforms several state-of-the-art methods. Guo Yan, Feng Xu 0007, Yuan Yao 0001, Jian Lu 0001 |
Internetware | 3 |
| 2013 | MATRI: a multi-aspect and transitive trust inference modelabstractTrust inference, which is the mechanism to build new pair-wise trustworthiness relationship based on the existing ones, is a fundamental integral part in many real applications, e.g., e-commerce, social networks, peer-to-peer networks, etc. State-of-the-art trust inference approaches mainly employ the transitivity property of trust by propagating trust along connected users (a.k.a. trust propagation), but largely ignore other important properties, e.g., prior knowledge, multi-aspect, etc. Yuan Yao 0001, Hanghang Tong, Xifeng Yan, Feng Xu 0007, Jian Lu 0001 |
WWW | 1 |
| 2013 | SelfTrust: leveraging self-assessment for trust inference in Internetware
Yuan Yao 0001, Feng Xu 0007, Yongli Ren, Hanghang Tong, Jian Lu 0001 |
Sci. China Inf. Sci. | 1 |
| 2012 | Subgraph Extraction for Trust Inference in Social NetworksabstractTrust inference is an essential task in many real world applications. Most of the existing inference algorithms suffer from the scalability issue, making themselves computationally costly, or even infeasible, for the graphs with more than thousands of nodes. In addition, the inference result, which is typically an abstract, numerical trustworthiness score, might be difficult for the end-user to interpret. In this paper, we propose sub graph extraction to address these challenges. The core of the proposed method consists of two stages: path selection and component induction. The outputs of both stages can be used as an intermediate step to speed up a variety of existing trust inference algorithms. Our experimental evaluations on real graphs show that the proposed method can accelerate existing trust inference algorithms, while maintaining high accuracy. In addition, the extracted sub graph provides an intuitive way to interpret the resulting trustworthiness score. Yuan Yao 0001, Hanghang Tong, Feng Xu 0007, Jian Lu 0001 |
ASONAM | 1 |
| 2012 | A group recommendation approach for service selectionabstractThere are more and more services that fulfill similar functionality, such as image service provided by Flickr, Picasa and Facebook. Which should be adopted to construct our software system in the open, dynamic and non-deterministic Internet environment is a key problem. Earlier work[15, 9] analyze this problem from the point view of QoS and established generic and extensible QoS computation framework for service selection. However those framework are almost designed for individuals. As social network emerges and gets widespread, people tend to be more connected and self-organize themselves into groups. Benefits of all members should be considered when we select service for group. In this article, we propose a revised group recommendation algorithm which takes advantage of collaborative filtering technology for service selection. As the experiment demonstrates, our algorithm exhibits high accuracy. Feng Xu 0007, Yuan Yao 0001, Jian Lu 0001 |
Internetware | 3 |
| 2009 | A dynamic trust network based simulation framework for reputation-based service selectionabstractService-oriented computing is a promising approach to software system construction by selecting and composing autonomous services under the open, dynamic and non-deterministic Internet environment. Appropriate selection of high quality services used in from numerous candidates declaring similar functionalities is crucial to the overall quality of the composed system. As authority centers are not generally available in open environments, reputation-based mechanisms must be adopted to evaluate services. With more and more reputation systems proposed in the literature, there is an increasing need to evaluate and compare them objectively and systematically with a common controlled experiment of trust network environment. In this paper we propose a general simulation framework based on dynamic trust network for this purpose. Especially, the framework is capable to simulate the dynamic evolutions of the trust network, in addition to the static snapshots of the trust relationships. With this framework, some case studies are made to evaluate the effectiveness of several representative reputation mechanisms, and some interesting characters are revealed. Feng Xu 0007, Yuan Yao 0001, Jian Lu 0001 |
Internetware | 3 |