VLDB 2026 Research / reviewers in the wild / expert
Xiaojun Ye 0001
dblp:79/5188-1 · also Xiao-jun Ye 0001
· DBLP profile ↗
17ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-9780-4827ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VFLAIR-LLM: A Comprehensive Framework and Benchmark for Split Learning of LLMsabstractWith the advancement of Large Language Models (LLMs), LLM applications have expanded into a growing number of fields. However, users with data privacy concerns face limitations in directly utilizing LLM APIs, while private deployments incur significant computational demands. This creates a substantial challenge in achieving secure LLM adaptation under constrained local resources. To address this issue, collaborative learning methods, such as Split Learning (SL), offer a resource-efficient and privacy-preserving solution for adapting LLMs to private domains. In this study, we introduce VFLAIR-LLM (available at https://github.com/FLAIR-THU/VFLAIR-LLM), an extensible and lightweight split learning framework for LLMs, enabling privacy-preserving LLM inference and fine-tuning in resource-constrained environments. Our library provides two LLM partition settings, supporting three task types and 18 datasets. In addition, we provide standard modules for implementing and evaluating attacks and defenses. We benchmark 5 attacks and 9 defenses under various Split Learning for LLM(SL-LLM) settings, offering concrete insights and recommendations on the choice of model partition configurations, defense strategies, and relevant hyperparameters for real-world applications. Zixuan Gu, Qiufeng Fan, Yang Liu 0165, Xiaojun Ye 0001 |
KDD (2) | 5 |
| 2023 | A White-Box Testing for Deep Neural Networks Based on Neuron CoverageabstractWith the introduction of neuron coverage as a testing criterion for deep neural networks (DNNs), covering more neurons to detect more internal logic of DNNs became the main goal of many research studies. While some works had made progress, some new challenges for testing methods based on neuron coverage had been proposed, mainly as establishing better neuron selection and activation strategies influenced not only obtaining higher neuron coverage, but also more testing efficiency, validating testing results automatically, labeling generated test cases to extricate manual work, and so on. In this article, we put forward Test4Deep, an effective white-box testing DNN approach based on neuron coverage. It is based on a differential testing framework to automatically verify inconsistent DNNs' behavior. We designed a strategy that can track inactive neurons and constantly triggered them in each iteration to maximize neuron coverage. Furthermore, we devised an optimization function that guided the DNN under testing to deviate predictions between the original input and generated test data and dominated unobservable generation perturbations to avoid manually checking test oracles. We conducted comparative experiments with two state-of-the-art white-box testing methods DLFuzz and DeepXplore. Empirical results on three popular datasets with nine DNNs demonstrated that compared to DLFuzz and DeepXplore, Test4Deep, on average, exceeded by 32.87% and 35.69% in neuron coverage, while reducing 58.37% and 53.24% testing time, respectively. In the meantime, Test4Deep also produced 58.37% and 53.24% more test cases with 23.81% and 98.40% fewer perturbations. Even compared with the two highest neuron coverage strategies of DLFuzz, Test4Deep still enhanced neuron coverage by 4.34% and 23.23% and achieved 94.48% and 85.67% higher generation time efficiency. Furthermore, Test4Deep could improve the accuracy and robustness of DNNs by merging generated test cases and retraining. Jing Yu 0028, Shukai Duan 0001, Xiaojun Ye 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Traffic Sign Detection and Recognition in Multiimages Using a Fusion Model With YOLO and VGG NetworkabstractThe detection and recognition of traffic signs is an important topic in intelligent transportation systems. The automatic detection and recognition of traffic signs during driving is the basis for realizing the unmanned driving. Therefore, the work on the detection and recognition of traffic signs has a potential value and application prospect. In the traditional detection and recognition methods, they often detect and recognize traffic signs image by image. In this case, only the information of the current image is used, and the relationship between the image sequences is not considered. To end this issue, we propose a novel model that can use the relationship in multi-images to detect and recognize traffic signs in a driving video sequence quickly and accurately. The model proposed in this paper is a fusion model based on YOLO-V3 and VGG19 network. Finally, we test this proposed model on a public dataset and compare it to the baseline method, and results show that this proposed model achieves accuracy over 90% and outperforms the baseline method for all types of traffic signs in different conditions. Thus, we can conclude this proposed model is efficient and accurate. Jing Yu 0028, Xiaojun Ye 0001, Qiang Tu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Efficient Transition Adjacency Relation Computation for Process Model SimilarityabstractMany activities in business process management, such as process retrieval, process mining, and process integration, need to determine the similarity between business processes. Along with many other relational behavior semantics, Transition Adjacency Relation (abbr. TAR) has been proposed as a kind of behavioral gene of process models and a useful perspective for process similarity measurement. In this article we explain why it is still relevant and necessary to improve TAR or pTAR (i.e., projected TAR) computation efficiency and put forward a novel approach for TAR computation based on Petri net unfolding. This approach not only improves the efficiency of TAR computation, but also enables the long-expected combined usage of TAR and Behavior Profiles (abbr. BP) in process model similarity estimation. Jisheng Pei, Lijie Wen 0001, Xiaojun Ye 0001, Akhil Kumar 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2021 | Network Embedding With Completely-Imbalanced LabelsabstractNetwork embedding, aiming to project a network into a low-dimensional space, is increasingly becoming a focus of network research. Semi-supervised network embedding takes advantage of labeled data, and has shown promising performance. However, existing semi-supervised methods would get unappealing results in thecompletely-imbalancedlabel setting where some classes have no labeled nodes at all. To alleviate this, we propose two novel semi-supervised network embedding methods. The first one is a shallow method named RSDNE. Specifically, to benefit from the completely-imbalanced labels, RSDNE guarantees both intra-class similarity and inter-class dissimilarity in an approximate way. The other method is RECT which is a new class of graph neural networks. Different from RSDNE, to benefit from the completely-imbalanced labels, RECT explores the class-semantic knowledge. This enables RECT to handle networks with node features and multi-label setting. Experimental results on several real-world datasets demonstrate the superiority of the proposed methods. Zheng Wang 0045, Xiaojun Ye 0001, Chaokun Wang, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Estimating Global Completeness of Event Logs: A Comparative StudyabstractEvent logs are the basis of process mining techniques and tools that extract process behavior information for better understanding and optimization of business processes. While it has been widely realized that the degree of completeness of event logs may largely determine the effectiveness of these techniques, how to estimate the completeness of event logs has not yet been fully addressed. This is mainly because ground-truth process models are usually unknown. To attack this problem, we pay a closer look to several concepts and implicit assumptions in the log completeness estimation problem and characterize it as a special case of the species estimation problem in the field of statistics. Although species estimation is still an open problem, a number of statistic models and techniques with approximate solutions have been available. To investigate the relevance of these methods for event log completeness estimation, we have designed and conducted a wide scope of empirical study and quantitative experiments on both real-world and synthesized event logs to compare the performance of these methods. In addition, the completeness estimation of several important and widely used real-world events logs are reported for the first time together with some best practice experience learned through this research. Jisheng Pei, Lijie Wen 0001, Hedong Yang, Jianmin Wang 0001, Xiaojun Ye 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2016 | A semantic-aware data generator for ETL workflowsabstractSummary Extract, transform, and load (ETL) processes organized as workflows play an important role in the future data integration for cloud services. ETL designers/administrators need testing data set that is aware of semantics of ETL workflow workloads to evaluate their developed ETL systems. Populating testing ETL systems with meaningful workload data is a difficult task. In this paper, we propose a semantic‐aware data generator for ETL workflows. With given ETL workflow models and workload characterizations, the generator is able to generate synthetic data that capture the semantics of ETL activities. This is carried out by a three‐staged approach. First, we derive expected cardinalities of all the source, intermediate, and target data sets involved in the ETL workflow model with some user‐specified cardinality requirements. Then, with the concept of symbolic test, symbolic data instead of concrete data involved in ETL activities are generated, and semantics of the ETL workflow models are transformed to various constraints over these symbols. At last, concrete data are derived on the basis of resolving constraints. Our generator may facilitate ETL workload test case generation for ETL toolkit performance and function evaluations as well as ETL workflow solution benchmarking. Copyright © 2013 John Wiley & Sons, Ltd. Naiqiao Du, Xiaojun Ye 0001, Jianmin Wang 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | A Creative Lifecycle Model for Internetware Based Software DevelopmentabstractWhile culture being the "software" controlling human mind, computer software development becomes one of the most creative activities that human undertake since the civilisation began. The only limitation in software creation is human imagination, and that limit is often self-imposed. The "Internetware", refers to a software paradigm, aims to satisfy the need of human kind using Internet as an integrated development and execution platform. Such software systems composed of entities distributed through the Internetwork, allowing connections that would be impossible or difficult to make otherwise. This paper gear towards the tasks for the Internetware is to accommodate creativity. In particular, we propose a six-step approach to Internetware development, an approach suggests essential difference to traditional methods. The six steps in this approach are search, ideation, specification, coding, testing, and operation. For each of these steps, we suggest a set of techniques to carry out the step in practice. The proposed developmental model asks researchers and practitioners of the Internetware paradigm to better understand human creativity and to formulate an algorithmic perspective on creative behavior in human, to design programs that can enhance human creativity without necessarily being creative themselves. Lin Liu 0001, Jianmin Wang 0001, Xiaojun Ye 0001 |
Internetware | 3 |
| 2014 | Requirements model driven adaption and evolution of Internetware
Lin Liu 0001, Jianmin Wang 0001, Xiaojun Ye 0001, Xiaodong Liu 0002 |
Sci. China Inf. Sci. | 4 |
| 2013 | Image Tag Completion via Image-Specific and Tag-Specific Linear Sparse ReconstructionsabstractThough widely utilized for facilitating image management, user-provided image tags are usually incomplete and insufficient to describe the whole semantic content of corresponding images, resulting in performance degradations in tag-dependent applications and thus necessitating effective tag completion methods. In this paper, we propose a novel scheme denoted as LSR for automatic image tag completion via image-specific and tag-specific Linear Sparse Reconstructions. Given an incomplete initial tagging matrix with each row representing an image and each column representing a tag, LSR optimally reconstructs each image (i.e. row) and each tag (i.e. column) with remaining ones under constraints of sparsity, considering image-image similarity, image-tag association and tag-tag concurrence. Then both image-specific and tag-specific reconstruction values are normalized and merged for selecting missing related tags. Extensive experiments conducted on both benchmark dataset and web images well demonstrate the effectiveness of the proposed LSR. Zijia Lin, Guiguang Ding, Mingqing Hu, Jianmin Wang 0001, Xiaojun Ye 0001 |
CVPR | 5 |
| 2013 | A multidimensional data model for TPC-DS benchmarkingabstractTo evaluate the performance of OLAP engines, adapted benchmarks have been developed in the industry. In this paper, we propose a multidimensional model with cube structure for data analytic system benchmarking based on TPC-DS. A number of analytic operations corresponding with SQL templates are given with MDX forms. Finally, we demonstrate the effectiveness of proposed model on big data stores by experiments of SQL and MDX queries on relational DBMSs and NoSQL databases. Xiaojun Ye 0001, Jianmin Wang 0001 |
Internetware | 2 |
| 2012 | Approaches on Getting Workflow Task Execution Number of TimesabstractCloud computing has generated significant interest in the business processing community, where workflow process in a business context can be organized by integrating web-based services or hosted activities. The budget for capital expenses and operating expenses on cloud computing based workflow systems can be made previous the real deployment when expenses of involved tasks (cloud services) and expected task execution number of times (expnum) are available. To get expnum, we propose two approaches, which are flow-balance and divide-and-conquer. We also analyze workflow control-flow patterns to help in getting expnum. Case study and simulations have verified the ability of our approaches. Naiqiao Du, Xiaojun Ye 0001, Jianmin Wang 0001 |
APSCC | 2 |
| 2012 | Protecting data confidentiality in cloud systemsabstractTo achieve a trustworthy cloud data service, there is a need to both provide the right services from a security engineering perspective, as well as to allows specific types of computations to be carried out on encrypted cloud data. However, traditional encryption solutions can't be used to process outsourcing encrypted data hosting to an untrusted cloud provider. A novel encryption scheme, called fully homomorphic encryption (FHE), could afford the circuit ability over encrypted data without decrypting it. In this paper, we deliver a universal construction framework for fully homomorphic encryption schemes. At first, this framework initializes a somewhat homomorphic encryption scheme based on the concept of metric space in abstract algebra which encodes the plaintext into a offset vector and generates a ciphertext by adding the offset vector to a random eigenvector in the metric space. As an abelian group, the ring is closed under addition and multiplication, this abstract algebra assume the metric space could forma ring and the eigenvectors belong to an ideal of this ring, then this framework could achieve homomorphism by having the scheme live in rings. We also deduce some well-known fully homomorphic schemes from the construction framework, and propose a prototype with an FHE encryption proxy to solve confidentiality problems in cloud systems. At last, we show the performance of FHE with some experiments, and speed the performance of fully homomorphic encryption up with cloud computing (parallel computing, distribute computing, etc.). We also discuss some opening issues and directions for future fully homomorphic encryption researches. Xiaojun Ye 0001, Jianmin Wang 0001 |
Internetware | 2 |
| 2009 | A configurable benchmark test management frameworkabstractBy looking over the performance benchmarks, we found that test systems are becoming more and more complex [1]. It is expensive for academe to implement every benchmark systems from scratch, as well as the customized benchmarks, to measure particular system under test (SUT). We propose a configurable benchmark test management framework, which gives a structure intended to serve as a support or guide for the building of testbed benchmark systems that expands the reference architecture into something useful in internetware components benchmark and real system performance evaluation. Naiqiao Du, Xiaojun Ye 0001, Jianmin Wang 0001 |
Internetware | 2 |
| 2009 | Trusted resource dissemination in Internetware systemsabstractTrusted resource dissemination over Internet is still an open problem. However, it is crucial to the success of Internet-based or Internetware systems. We propose an approach that supports trusted resource dissemination between Internetware nodes. In this paper, we focus on an evidence-based trustworthiness-assurance mechanism in the approach. An architecture and a protocol that enforces this approach in Internetware systems are also described. Zude Li, Xiaojun Ye 0001, Jianmin Wang 0001 |
Internetware | 2 |
| 2009 | Requirements-driven Internetware services evaluationabstractIn the services era, evaluation of existing services and planning for new services according to user requirements are key activities that need systematic support. This paper proposes a requirements-driven evaluation framework for Internetware-based services, with respect to both their functionality and risk. In particular, we offer an account of how to model these requirements, how to derive from them a space of service functionality alternatives, and how to select among these alternatives on the basis of desired qualities. In essence, the selection of service functionality is framed as a satisfaction problem for requirements; while service risk is addressed as an analysis of failure rates. We use a typical logistics example scenario to illustrate the proposed framework. Wenting Ma, Lin Liu 0001, Xiaojun Ye 0001, Jianmin Wang 0001, John Mylopoulos |
Internetware | 3 |
| 2006 | FGAC-QD: Fine-Grained Access Control Model Based on Query Decomposition Strategy
Guoqiang Zhan, Zude Li, Xiaojun Ye 0001, Jianmin Wang 0001 |
TrustBus | 3 |