VLDB 2026 Research / reviewers in the wild / expert
Jie Wang 0006
dblp:29/5259-6
· DBLP profile ↗
43ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0003-1857-5569ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 4 since 2021Software engineering, systems software and programming languages · 8 · 1 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Computer networks · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Hallucinations in Large Language Models via Causal ReasoningabstractLarge language models (LLMs) exhibit logically inconsistent hallucinations that appear coherent yet violate reasoning principles, with recent research suggesting an inverse relationship between causal reasoning capabilities and such hallucinations. However, existing reasoning approaches in LLMs, such as Chain-of-Thought (CoT) and its graph-based variants, operate at the linguistic token level rather than modeling the underlying causal relationships between variables, lacking the ability to represent conditional independencies or satisfy causal identification assumptions. To bridge this gap, we introduce causal-DAG construction and reasoning (CDCR-SFT), a supervised fine-tuning framework that trains LLMs to explicitly construct variable-level directed acyclic graph (DAG) and then perform reasoning over it. Moreover, we present a dataset comprising 25,368 samples (CausalDR), where each sample includes an input question, explicit causal DAG, graph-based reasoning trace, and validated answer. Experiments on four LLMs across eight tasks show that CDCR-SFT improves the causal reasoning capability with the state-of-the-art 95.33% accuracy on CLADDER (surpassing human performance of 94.8% for the first time) and reduces the hallucination on HaluEval with 10% improvements. It demonstrates that explicit causal structure modeling in LLMs can effectively mitigate logical inconsistencies in LLM outputs. Yuangang Li 0002, Yiqing Shen 0003, Yi Nian, Jiechao Gao, Ziyi Wang 0012, Chenxiao Yu, Li Li 0006, Jie Wang 0006, Xiyang Hu, Yue Zhao 0016 |
AAAI | 8 |
| 2025 | Enhancing Interpretability in Self-Training with Tsetlin Machines for Mitigating Noisy Pseudo-Labels
Jiechao Gao, Rohan Kumar Yadav, Jie Wang 0006 |
IEEE Big Data | 4 |
| 2025 | Federated Neural Architecture Search with Model-Agnostic Meta Learning
Jiechao Gao, Jie Wang 0006 |
IEEE Big Data | 3 |
| 2025 | Ontology-based Adaptive Knowledge System (OAKS): Adaptive and Consistent Knowledge Acquisition through LLMs for Diverse User BackgroundsabstractAlthough most of the research on large language models (LLMs) focuses on their development and validation against datasets, significant gaps remain in their application to real-world knowledge-intensive tasks. This research addresses key challenges in using LLMs for extracting and synthesizing knowledge from unstructured sources, focusing on applications where the validity and consistency of the results are critical. We propose Ontology-based Adaptive Knowledge System (OAKS) as a holistic approach to manage the complexities of acquiring unstructured knowledge, varying user expertise, and dynamic query formulation. This research provides practical value for enabling the domain user community to leverage their technical documentation and expertise and accelerate ongoing working projects through improved literature review and cross-disciplinary insight discovery. Validated through empirical studies, our findings offer insight into best practices for the deployment of OAKS, bridging the gaps between AI capabilities and real-world needs in knowledge acquisition and research. Muran Yu, Jie Wang 0006, Yirong Chen, Michael D. Lepech, Ying Liu 0039, Kincho H. Law |
COMPSAC | 2 |
| 2025 | ZION: A Practical Confidential Virtual Machine Architecture on Commodity RISC-V ProcessorsabstractTrusted Execution Environments (TEEs) provide robust hardware-based isolation to mitigate data breaches and privacy risks. Confidential Virtual Machines (confidential VMs or CVMs) extend these capabilities by using VMs as their execution abstraction, offering superior compatibility over process-based TEEs like Intel SGX. The rising demand for Confidential VMs has spurred innovations from major chip manufacturers, such as AMD SEV, Intel TDX, and Arm CCA, and their integration into leading cloud platforms, including AWS, Azure, and Google Cloud. On the RISC-V platform, however, existing TEE architectures rely on process-level abstractions or custom hardware, leading to limited compatibility and scalability.This paper presents ZION, a confidential VM architecture for commodity RISC-V hardware that operates without custom extensions. ZION ensures security, flexibility, and efficiency through a short-path CVM mode and a secure vCPU mechanism for protecting and efficiently updating vCPU states, enhancing context-switching performance. It combines Physical Memory Protection (PMP) with paging for scalable memory isolation, employs a hierarchical memory structure for efficient management, and introduces a split-page-table-based mechanism for secure memory sharing with virtio devices. Evaluations show ZION achieves under 5% overhead in real-world applications, demonstrating its practicality. Jie Wang 0006, Juan Wang 0006, Yinqian Zhang |
DAC | 1 |
| 2025 | MPF: A Multi-Noise Perception Framework to Enhance Online Map Matching AlgorithmsabstractMap matching is crucial to facilitating location-based services, and recent advancements in map matching have demonstrated excellent performance with high-quality data. However, the use of low-precision devices often introduces high measurement noise, and the slow update rate of maps may result in errors in digital maps. Consequently, multiple types of noise significantly impact the performance of map matching algorithms. To tackle this issue, this paper presents a novel multi-noise perception framework, named MPF, aiming to enhance the performance and robustness of existing map matching algorithms. The main challenge lies in detecting anomalies during map matching, identifying the root causes, and devising appropriate solutions. Firstly, we propose a matching quality assessment (MQA) method that assesses abnormal variance in matching probability. Secondly, we introduce a multiple noise discrimination (MND) mechanism to effectively differentiate between measurement noise and map errors. Thirdly, we present a missing segment generation (MSG) scheme that dynamically fills in map gaps to prevent significant detours. To validate the effectiveness of MPF, we conduct experiments using real-world taxi trajectories from four cities, covering a total distance of 79,670.6 km. MPF is compare with seven online map matching algorithms and is used to optimize their performance. The experiments show that MPF outperforms the top baselines by 15.6%-26.9% and enhances their performance by 18.7%-38.2%. Hanwen Hu, Shiyou Qian, Jian Cao 0001, Yirong Chen, Jie Wang 0006 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | PREFER: A Pre-trained Model Recommendation Framework for Edge Computing Enabled Traffic Flow PredictionabstractThe recent years have witnessed a surge in the development of traffic flow prediction methods, often deployed on cloud platforms to offer predictive services for entire transportation networks. However, the processes of training and executing a model for the entire traffic network are both time-consuming and computationally expensive. As a result, the utilization of edge servers for local sub-network prediction services has gained prominence. Nevertheless, training prediction models for numerous sub-networks within the extensive traffic network remains a time-intensive and computing resource-consuming task. To tackle this challenge, this article introduces the Pre-trained model REcommendation Framework for Edge computing enabled tRaffic flow prediction (PREFER). PREFER trains a set of traffic flow prediction models on selected sub-networks, then recommends optimal pre-trained models for edge servers. The recommendation is specifically based on performance prediction, integrating neural collaborative filtering and traffic flow characteristics. Experiments conducted on real datasets reveal that the pre-trained models recommended by PREFER perform close to the actual optimal ones and significantly outperform existing recommendation algorithms. Qiqi Cai, Jian Cao 0001, Yirong Chen, Shiyou Qian, Liangxiao Yuan, Jie Wang 0006 |
ACM Trans. Knowl. Discov. Data | 6 |
| 2024 | Multi-agent reinforcement learning for vehicular task offloading with multi-step trajectory prediction
Yanmin Zhu 0006, Chunyang Wang 0001, Jian Cao 0001, Yirong Chen, Jie Wang 0006 |
CCF Trans. Pervasive Comput. Interact. | 6 |
| 2023 | SvTPM: SGX-Based Virtual Trusted Platform Modules for Cloud ComputingabstractVirtual Trusted Platform Modules (vTPMs) are widely used in commercial cloud platforms (e.g., VMware Cloud, Google Cloud, and Microsoft Azure) to provide virtual root-of-trust and security services for virtual machines. Unfortunately, current state-of-the-art vTPM implementations for cloud computing cannot provide strong protection for vTPMs at run-time and suffer from poor performance under binding vTPMs to a physical TPM. In this paper, we propose SvTPM, an SGX-based virtual trusted platform module, which provides complete life cycle protection of vTPMs in the cloud and does not rely on the physical TPM. SvTPM provides strong isolation protection so malicious cloud tenants or even cloud administrators cannot access vTPM's private keys or any other sensitive data. In this paper, we implement a prototype of SvTPM, which identifies and solves a couple of critical security challenges for vTPM protection with SGX, such as NVRAM rollback attacks, NVRAM binding attacks, and vTPM rollback attacks. SvTPM also shows how to establish trust between vTPM and SGX Platform. Our performance evaluation shows that the NVRAM launch time of SvTPM is$1700\times$faster than vTPM built upon hardware TPM. In TPM standard command evaluation, we find that SvTPM incurs negligible performance overhead while providing strong isolation and protection. To our knowledge, SvTPM is the first practical work to solve the critical security challenges of securing vTPM using SGX. Juan Wang 0006, Jie Wang 0006, Chengyang Fan, Fei Yan 0008, Yueqiang Cheng, Yinqian Zhang, Mengda Yang, Hongxin Hu |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | AMM: An Adaptive Online Map Matching AlgorithmabstractOnline map matching is essential for some location-based services, such as car navigation. However, due to GPS measurement errors and/or the lack of sufficient information, the performance of most existing online algorithms will degrade in the increasingly complex traffic environment. In this paper, we propose an adaptive online map matching algorithm called AMM. The basic idea is that AMM should be able to calibrate GPS observation data for various measurement errors and under complex urban conditions. First, we establish a collaborative evaluation model between GPS points and candidate points to effectively filter low-quality GPS measurement points, dynamically set the weights of different features and comprehensively select the best candidate points. Second, we propose a retrospective correction mechanism to correct the previous matching results when more information is available, which will help improve the accuracy of future GPS points. Furthermore, we define parameter self-tuning rules for AMM to enhance its portability by avoiding time-consuming parameter tuning steps. We conduct extensive experiments to evaluate the performance of AMM on real vehicle trajectory datasets. The experiment results show that AMM outperforms its counterparts by up to 32% in terms of accuracy and its performance in different traffic conditions is more stable. Hanwen Hu, Shiyou Qian, Jingchao Ouyang, Jian Cao 0001, Jie Wang 0006, Yirong Chen |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | @ME: A Fine-grained Route Recommendation System to Grab Impatient PassengersabstractData analysis reveals that passengers can only endure a few minutes before taking a taxi. However, most existing route recommendation systems are not adequate to satisfy impatient customers due to two shortcomings: inaccurate demand forecast and the lack of an efficient supply-demand balance mechanism. In this paper, we propose a recommendation system called @ Me, which aims to dispatch vacant taxis to the vicinity of potential customers at the right time. To achieve minute-level demand forecasting, we ensemble a contextualized spatial-temporal network (CSTN) with an LSTM network to optimize prediction accuracy. In addition, we characterize the attractiveness of the road grid to vacant taxis as a force model, on which a taxi scheduling algorithm is proposed to dynamically balance supply and demand. Extensive experiments on real datasets clearly indicate that our method is superior to the selected baselines. Vacant taxis that follow the routes suggested by @ME can catch more impatient customers in a shorter cruising time. The 7 -day experimental results on the Manhattan dataset show that @Me can carry an additional 48,332 passengers and increase drivers' revenue by $570,317. Hanwen Hu, Yirong Chen, Jingchao Ouyang, Shiyou Qian, Jian Cao 0001, Jie Wang 0006, Michael D. Lepech |
IJCNN | 7 |
| 2022 | Multi-Agent Deep Reinforcement Learning For Real-World Traffic Signal Controls - A Case StudyabstractIncreasing traffic congestion leads to significant costs, whereby poorly configured signaled intersections are a common bottleneck and root cause. Traditional traffic signal control (TSC) systems employ rule-based or heuristic methods to decide signal timings, while adaptive TSC solutions utilize a traffic-actuated control logic to increase their adaptability to real-time traffic changes. However, such systems are expensive to deploy and are often not flexible enough to adequately adapt to the volatility of today’s traffic dynamics. More recently, this problem became a frontier topic in the domain of deep reinforcement learning (DRL) and enabled the development of multi-agent DRL approaches that can operate in environments with several agents present, such as traffic systems with multiple signaled intersections. However, many of these proposed approaches were validated using artificial traffic grids. This paper presents a case study, where real-world traffic data from the town of Lemgo in Germany is used to create a realistic road model within VISSIM. A multi-agent DRL setup, comprising multiple independent deep Q-networks, is applied to the simulated traffic network. Traditional rule-based signal controls, modeled in LISA+ and currently employed in the real world at the studied intersections, are integrated into the traffic model and serve as a performance baseline. The performance evaluation indicates a significant reduction of traffic congestion when using the RL-based signal control policy over the conventional TSC approach with LISA+. Consequently, this paper reinforces the applicability of RL concepts in the domain of TSC engineering by employing a highly realistic traffic model. Maxim Friesen, Tian Tan 0015, Jürgen Jasperneite, Jie Wang 0006 |
INDIN | 4 |
| 2022 | ProcGuard: Process Injection Behaviours Detection Using Fine-grained Analysis of API Call Chain with Deep LearningabstractNew malware increasingly adopts novel fileless techniques to evade detection from antivirus programs. Process injection is one of the most popular fileless attack techniques. This technique makes malware more stealthy by writing malicious code into memory space and reusing the name and port of the host process. It is difficult for traditional security software to detect and intercept process injections due to the stealthiness of its behavior. We propose a novel framework called ProcGuard for detecting process injection behaviors. This framework collects sensitive function call information of typical process injection. Then we perform a fine-grained analysis of process injection behavior based on the function call chain characteristics of the program, and we also use the improved RCNN network to enhance API analysis on the tampered memory segments. We combine API analysis with deep learning to determine whether a process injection attack has been executed. We collect a large number of malicious samples with process injection behavior and construct a dataset for evaluating the effectiveness of ProcGuard. The experimental results demonstrate that it achieves an accuracy of 81.58% with a lower false-positive rate compared to other systems. In addition, we also evaluate the detection time and runtime performance loss metrics of ProcGuard, both of which are improved compared to previous detection tools. Juan Wang 0006, Chenjun Ma, Huanyu Yuan, Jie Wang 0006 |
TrustCom | 5 |
| 2021 | CBPCS: A Cache-block-based Service Process Caching Strategy to Accelerate the Execution of Service ProcessesabstractWith the development of cloud computing and the advent of the Web 2.0 era, composing a set of Web services as a service process is becoming a common practice to provide more functional services. However, a service process involves multiple service invocations over the network, which incurs a huge time cost and could become a bottleneck to performance. To accelerate its execution, we propose an engine-side cache-block-based service process caching strategy (CBPCS). It is based on, and derives its advantages from, three key ideas. First, the invocation of Web services embodies semantics, which enables the application of semantic-based caching. Second, cache blocks are identified from a service process, and each block is equipped with a separate cache so that the time overhead of service invocation and caching can be minimized. Third, a replacement strategy is introduced taking into account time and space factors to manage the space allocation for a process with multiple caches. The algorithms and methods used in CBPCS are introduced in detail. Moreover, how CBPCS can be applied to multiple service process models is also investigated. Finally, CBPCS is validated via comparison experiments, which shows the considerable improvements of CBPCS over other strategies. Jian Cao 0001, Tingjie Jia, Shiyou Qian, Haiyan Zhao 0002, Jie Wang 0006 |
ACM Trans. Web | 5 |
| 2020 | Multi-Agent Context-Aware Dynamic-Scheduling for Large-scale Processing NetworksabstractOptimally scheduling jobs in processing networks to meet multiple objectives for economic considerations and operational efficiencies has been a hot topic. However, most data-driven methods are rarely applied. One of the critical reasons is that the scheduling policy generated by these methods tends to bias toward the specific environment. In order to better deal with the discrepancy, this work in progress paper presents a context-aware dynamic scheduling method (CADS) that can adaptively select a specific policy based on the on-demand context. The CADS has two components: 1) The evaluation module that evaluates the performances of the policies that are learned from each operational context; 2) The decision-making module that maintains the knowledge of each policy’s performance under each context and constructs the weighted best-fit policy based on the identified context. The promising preliminary result using numerical simulation that demonstrates the effectiveness of CADS is presented. CADS outperforms traditional scheduling methods in various kinds of processing network environments. Shuhui Qu, Yirong Chen, Jürgen Jasperneite, Michael D. Lepech, Jie Wang 0006 |
ETFA | 5 |
| 2020 | Cooperative Deep Reinforcement Learning for Large-Scale Traffic Grid Signal ControlabstractExploiting reinforcement learning (RL) for traffic congestion reduction is a frontier topic in intelligent transportation research. The difficulty in this problem stems from the inability of the RL agent simultaneously monitoring multiple signal lights when taking into account complicated traffic dynamics in different regions of a traffic system. Such challenge is even more outstanding when forming control decisions on a large-scale traffic grid, where the RL action space grows exponentially with the number of intersections within the traffic grid. In this paper, we tackle such a problem by proposing a cooperative deep reinforcement learning (Coder) framework. The intuition behind Coder is to decompose the original difficult RL task as a number of subproblems with relatively easy RL goals. Accordingly, we implement Coder with multiple regional agents and a centralized global agent. Each regional agent learns its own RL policy and value functions over a small region with limited actions. Then, the centralized global agent hierarchically aggregates RL achievements from different regional agents and forms the final Q -function over the entire large-scale traffic grid. The experimental investigations demonstrate that the proposed Coder could reduce on average 30% congestions in terms of the number of waiting vehicles during high density traffic flows in simulations. Tian Tan 0003, Feng Bao 0002, Yue Deng 0001, Alex Jin, Qionghai Dai, Jie Wang 0006 |
IEEE Trans. Cybern. | 6 |
| 2020 | Multi-Agent Deep Reinforcement Learning for Large-Scale Traffic Signal ControlabstractReinforcement learning (RL) is a promising data-driven approach for adaptive traffic signal control (ATSC) in complex urban traffic networks, and deep neural networks further enhance its learning power. However, the centralized RL is infeasible for large-scale ATSC due to the extremely high dimension of the joint action space. The multi-agent RL (MARL) overcomes the scalability issue by distributing the global control to each local RL agent, but it introduces new challenges: now, the environment becomes partially observable from the viewpoint of each local agent due to limited communication among agents. Most existing studies in MARL focus on designing efficient communication and coordination among traditional Q-learning agents. This paper presents, for the first time, a fully scalable and decentralized MARL algorithm for the state-of-the-art deep RL agent, advantage actor critic (A2C), within the context of ATSC. In particular, two methods are proposed to stabilize the learning procedure, by improving the observability and reducing the learning difficulty of each local agent. The proposed multi-agent A2C is compared against independent A2C and independent Q-learning algorithms, in both a large synthetic traffic grid and a large real-world traffic network of Monaco city, under simulated peak-hour traffic dynamics. The results demonstrate its optimality, robustness, and sample efficiency over the other state-of-the-art decentralized MARL algorithms. Jie Wang 0006, Lara Codeca, Zhaojian Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Dynamic scheduling in modern processing systems using expert-guided distributed reinforcement learningabstractIn the environment of modern processing systems, one topic of great interest is how to optimally schedule (i.e., allocate) jobs with different requirements for the systems to meet various objectives. Methods using distributed reinforcement learning (DRL) have recently achieved great success with large-scale dynamic scheduling problems. However, most DRL methods require a huge amount of computational time and a large amount of data for the DRL agents for a control policy. Meanwhile, various scheduling experts have already developed several scheduling policies (i.e., dispatching rules) that can successfully fulfill different objectives with acceptable performance for a processing system based on their understandings of the processing systems characteristics. In this paper, we propose to learn from experts to reduce the learning and searching costs of a good policy in large-scale dynamic scheduling problems. In the learning process, our DRL agents select the experts who have better performance in the scheduling environment, observe the experts' actions, and learn a scheduling policy guided by these experts' demonstrations. Our realistic simulations results demonstrate that this expert-guided DRL (EGDRL) approach outperforms DRL methods without expert guidance, as well as some other reinforcement learning from demonstration (RLfD) methods in several systems. To the best of our knowledge, our research is one of the first works that incorporates existing expert policies to guide the learning of optimal policies for large-scale dynamic scheduling problems. Shuhui Qu, Jie Wang 0006, Jürgen Jasperneite |
ETFA | 2 |
| 2019 | A fast and anti-matchability matching algorithm for content-based publish/subscribe systems
Shiyou Qian, Jian Cao 0001, Weichao Mao, Yanmin Zhu 0006, Jiadi Yu, Minglu Li 0001, Jie Wang 0006 |
Comput. Networks | 7 |
| 2018 | Dynamic scheduling in large-scale stochastic processing networks for demand-driven manufacturing using distributed reinforcement learningabstractIn the area of demand-driven manufacturing systems, one critical problem currently is how to dynamically and optimally schedule, i.e., allocate, actual jobs with different customer requirements in large-scale manufacturing systems in order to meet the various objectives. Scheduling experts have proposed several scheduling methods with acceptable performance for manufacturing systems based on their understanding of the systems' characteristics. However, the problem remains challenging due to the complicated composition of the multiple objectives of the system, the complex system dynamics, constraints, and the extremely high computational cost for large-scale manufacturing systems. In this paper, we apply a stochastic processing network that can capture the stochasticity and dynamics of discrete manufacturing systems. We then propose a data-driven, distributed reinforcement learning (DRL) method so that little information about the system dynamics is required, and the learning and search costs of a scheduling policy with high production performance in a processing system can be reduced, thus this method is capable of scaling to large-scale processing systems. In particular, we first use a stochastic processing network, i.e., a queueing model, to represent the production processes in a typical discrete manufacturing system so that it can be simulated. We then decompose the reinforcement learning into local processes. Each local processs agent can make decisions locally by assigning indices to jobs based on each job's real-time information (index policy). Because of this distributed learning characteristic and index policy, our approach is much more scalable and efficient than either centralized methods or traditional decentralized reinforcement learning methods. Based on our simulation, we find our approach can achieve higher production performance than other heuristics, past decentralized reinforcement learning methods, or centralized methods in the stochastic processing networks with different scales. Shuhui Qu, Jie Wang 0006, Jürgen Jasperneite |
ETFA | 2 |
| 2017 | Recommendations Based on Comprehensively Exploiting the Latent Factors Hidden in Items' Ratings and ContentabstractTo improve the performance of recommender systems in a practical manner, several hybrid approaches have been developed by considering item ratings and content information simultaneously. However, most of these hybrid approaches make recommendations based on aggregating different recommendation techniques using various strategies, rather than considering joint modeling of the item’s ratings and content, and thus fail to detect many latent factors that could potentially improve the performance of the recommender systems. For this reason, these approaches continue to suffer from data sparsity and do not work well for recommending items to individual users. A few studies try to describe a user’s preference by detecting items’ latent features from content-description texts as compensation for the sparse ratings. Unfortunately, most of these methods are still generally unable to accomplish recommendation tasks well for two reasons: (1) they learn latent factors from text descriptions or user--item ratings independently, rather than combining them together; and (2) influences of latent factors hidden in texts and ratings are not fully explored. In this study, we propose a probabilistic approach that we denote as latent random walk (LRW) based on the combination of an integrated latent topic model and random walk (RW) with the restart method, which can be used to rank items according to expected user preferences by detecting both their explicit and implicit correlative information, in order to recommend top-ranked items to potentially interested users. As presented in this article, the goal of this work is to comprehensively discover latent factors hidden in items’ ratings and content in order to alleviate the data sparsity problem and to improve the performance of recommender systems. The proposed topic model provides a generative probabilistic framework that discovers users’ implicit preferences and items’ latent features simultaneously by exploiting both ratings and item content information. On the basis of this probabilistic framework, RW can predict a user’s preference for unrated items by discovering global latent relations. In order to show the efficiency of the proposed approach, we test LRW and other state-of-the-art methods on three real-world datasets, namely, CAMRa2011, Yahoo!, and APP. The experiments indicate that our approach outperforms all comparative methods and, in addition, that it is less sensitive to the data sparsity problem, thus demonstrating the robustness of LRW for recommendation tasks. Jian Cao 0001, Jie Wang 0006, Shiyou Qian |
ACM Trans. Knowl. Discov. Data | 3 |
| 2017 | A Reliable and Efficient Distributed Service Composition Approach in Pervasive EnvironmentsabstractThe global, ubiquitous usage of smart handsets and diverse wireless communication tools calls for a meticulous reexamination of complex and dynamic service componentization and remote invocation. In order to satisfy ever-increasing service requirements and enrich users' experiences, efficient service composition approaches, which leverage the computing resources on nearby devices to form an on-demand composite service, should be developed. This is especially true for situations that are confronted with limited local computing capacity and device mobility. For any mobile pervasive environment, execution reliability and latency of the composite service are major concerns that impact users' satisfaction. In this paper, we propose a novel three-staged approach which takes reliability and latency into account to solve a distributed service composition efficiently. First, the graph of the functional process description is decomposed into multiple path structures through a graph-traversing algorithm. Second, messages are forwarded among the network nodes (i.e., intelligent handsets) to search for the sub-solutions for these path structures. Finally, an efficient combinatorial optimization algorithm computes the optimal service composition by the selection from these sub-solutions. This approach is validated extensively in static and mobile environments, and the results show the effectiveness and outperformance of this approach over existing approaches. Jian Cao 0001, Jie Wang 0006 |
IEEE Trans. Mob. Comput. | 3 |
| 2017 | Improving the Quality of Recommendations for Users and Items in the Tail of DistributionabstractShort-head and long-tail distributed data are widely observed in the real world. The same is true of recommender systems (RSs), where a small number of popular items dominate the choices and feedback data while the rest only account for a small amount of feedback. As a result, most RS methods tend to learn user preferences from popular items since they account for most data. However, recent research in e-commerce and marketing has shown that future businesses will obtain greater profit from long-tail selling. Yet, although the number of long-tail items and users is much larger than that of short-head items and users, in reality, the amount of data associated with long-tail items and users is much less. As a result, user preferences tend to be popularity-biased. Furthermore, insufficient data makes long-tail items and users more vulnerable to shilling attack. To improve the quality of recommendations for items and users in the tail of distribution, we propose a coupled regularization approach that consists of two latent factor models: C-HMF, for enhancing credibility, and S-HMF, for emphasizing specialty on user choices. Specifically, the estimates learned from C-HMF and S-HMF recurrently serve as the empirical priors to regularize one another. Such coupled regularization leads to the comprehensive effects of final estimates, which produce more qualitative predictions for both tail users and tail items. To assess the effectiveness of our model, we conduct empirical evaluations on large real-world datasets with various metrics. The results prove that our approach significantly outperforms the compared methods. Liang Hu 0004, Longbing Cao, Jian Cao 0001, Zhiping Gu, Guandong Xu, Jie Wang 0006 |
ACM Trans. Inf. Syst. | 6 |
| 2016 | Learning adaptive dispatching rules for a manufacturing process system by using reinforcement learning approachabstractAdvanced manufacturing, which employs the latest information and communication technologies to facilitate interconnected, efficient, and adaptive manufacturing systems, has become a prominent research topic in both academics and industry in recent years. One critical aspect of advanced manufacturing is how to match various complex job requirements with a manufacturer's real-time processing capabilities and its other resources within a job shop to optimally schedule manufacturing processes with multiple objectives. In general, a manufacturing scheduling problem is a non-deterministic, polynomial-time (NP)-hard problem. Due to its complexity, scheduling problems present a number of challenges to find the best possible solutions. In order to deal with a dynamic job shop scheduling problem-particularly, a dispatching problem-for a manufacturing system that is able to handle multiple product types through multi-stages and multi-machines with dynamic orders, stochastic processing time and setup time. Based on a generic framework, this research develops a solution by using a reinforcement learning-based scheduling approach that can adaptively update the production schedule by utilizing the real-time product and process events information during executions. More specifically, we first propose a framework that describes the process of dealing with a complex scheduling problem. Next, to learn the dispatching pattern, we formally define the scheduling problem through the construction of objective functions and related constraints for the manufacturing tasks. We then apply a reinforcement learning approach that incorporates the real-time production environment to generate optimal policies under various manufacturing-process conditions. When tested under different objectives and constraint conditions, our results demonstrate that the proposed learning-based method provides better performance than most common dispatching rules. Shuhui Qu, Jie Wang 0006, Govil Shivani |
ETFA | 2 |
| 2016 | Group Recommendations Based on Comprehensive Latent Relationship DiscoveryabstractIn recent years, due to an increasing overload of information on the Internet, there are many scenarios where Recommender Systems (RSs) are employed to provide suggestions to user groups. However, most proposed approaches of group recommendations simply aggregate individual ratings or individual prediction results, rather than comprehensively investigating the hidden correlative information between members and the group, which results in inferior recommendation performance. In this paper, we propose a new approach, RWR-UTM, for group recommendations based on the combination of an integrated probabilistic topic model - a User Topic Model (UTM) and the Random Walk with Restart (RWR) method. The UTM provides a latent framework of users, groups, and items by exploiting both the users' preference profiles and the items' content information, which together can describe group interests and item features in a more complete manner. This latent framework is then combined with RWR to predict the preference degrees of groups to unrated items by detecting comprehensive latent relationships. In particular, we devised two group-based recommendation algorithms on the basis of different recommendation strategies. Finally, we conducted experiments to evaluate our approach and compare it with other state-of-the-art approaches using the real-world CAMRa2011 data-set. The results demonstrate the advantage of our approach over comparative ones. Jian Cao 0001, Jie Wang 0006, Jing He 0004 |
ICWS | 3 |
| 2016 | An Optimal Engine Component Placement Strategy for Cloud Workflow ServiceabstractWorkflows have been used to represent a variety of applications that involve coordinating a set of business services or scientific services, which are generally geographically distributed. With the development of cloud computing, a workflow engine can be deployed as a cloud service, responsible for executing customers' workflow instances. In a cloud workflow service, workflow engine components can be placed into different cloud regions. Thus, one challenging problem that arises is how to select the appropriate cloud regions to place the workflow engine components in order to efficiently execute a service workflow instance. Because this is a typical nondeterministic polynomial-time hard (NP-hard) problem, we propose a heuristic algorithm to select the regions where to place workflow engine components in an optimal and efficient way, with the objective of reducing the execution time of the service workflow instance. The experimental results prove that our proposed algorithm has higher performance than other approaches in terms of the solution quality and the running speed. Yan Yao 0001, Jian Cao 0001, Yusheng Jiang, Jie Wang 0006 |
ICWS | 4 |
| 2015 | A centralized reinforcement learning approach for proactive scheduling in manufacturingabstractDue to rapid development of information and communications technology (ICT) and the impetus for more effective, efficient and adaptive manufacturing, the concept of ICT based advanced manufacturing has increasingly become a prominent research topic across academia and industry during recent years. One critical aspect of advanced manufacturing is how to incorporate real time information and then optimally schedule manufacturing processes with multiple objectives. Due to its complexity and the need for adaptation, the manufacturing scheduling problem presents challenges for utilizing advanced ICT and thus calls for new approaches. The paper proposes a centralized reinforcement learning approach for optimally scheduling of a manufacturing system of multi-stage processes and multiple machines for multiple types of products. The approach, which employs learning and control algorithms to enable real time cooperation of each processing unit inside the system, is able to adaptively respond to dynamic scheduling changes. More specifically, we first formally define the scheduling problem through the construction of an objective function and related heuristic constraints for the underlying manufacturing tasks. Next, to effectively deal with the problem we defined, we maintain a distributed weighted vector to capture the cooperative pattern of massive action space and apply the reinforcement-learning approach to achieve the optimal policies for a set of processing machines according to a real time production environment, including dynamic requests for various products. Numerical experiments demonstrate that compared to different heuristic methods and multi-agent algorithms, the proposed centralized reinforcement learning method can provide more reliable solutions for the scheduling problem. Shuhui Qu, Jie Wang 0006, James O. Leckie, Weiwen Jian |
ETFA | 3 |
| 2015 | A Parallel Approach for Service Composition with Complex Structures in Pervasive EnvironmentsabstractThe composition of multiple services that are deployed on smart devices generally incurs significant communication overheads, especially when an optimized composition is pursued. A parallel approach for service composition is proposed for obtaining the optimal solution with minimum executing time efficiently. In this approach, the request for a composite service is represented by a function graph which is decomposed into multiple path-structured sub graphs firstly. Then messages are sent among the service nodes to search for the corresponding sub-solutions. Finally, a Branch-and-Bound strategy is applied to generate the optimal solution over these sub-solutions. The experiments show our optimization strategies can reduce the time and communication cost considerably under different experimental settings. Jian Cao 0001, Jie Wang 0006 |
ICWS | 3 |
| 2015 | Progressive online aggregation in a distributed stream system
Dingyu Yang, Jian Cao 0001, Sai Wu, Jie Wang 0006 |
J. Syst. Softw. | 4 |
| 2015 | H-Tree: An Efficient Index Structurefor Event Matching in Content-BasedPublish/Subscribe SystemsabstractContent-based publish/subscribe systems have been employed to deal with complex distributed information flows in many applications. It is well recognized that event matching is a fundamental component of such large-scale systems. Event matching searches a space which is composed of all subscriptions. As the scale and complexity of a system grows, the efficiency of event matching becomes more critical to system performance. However, most existing methods suffer significant performance degradation when the system has large numbers of both subscriptions and their component constraints. In this paper, we present Hash Tree (H-Tree), a highly efficient index structure for event matching. H-Tree is a hash table in nature that is a combination of hash lists and hash chaining. A hash list is built up on an indexed attribute by realizing novel overlapping divisions of the attribute's value domain, providing more efficient space consumption. Multiple hash lists are then combined into a hash tree. The basic idea behind H-Tree is that matching efficiencies are improved when the search space is substantially reduced by pruning most of the subscriptions that are not matched. We have implemented H-Tree and conducted extensive experiments in different settings. Experimental results demonstrate that H-Tree has better performance than its counterparts by a large margin. In particular, the matching speed is faster by three orders of magnitude than its counterparts when the numbers of both subscriptions and their component constraints are huge. Shiyou Qian, Jian Cao 0001, Yanmin Zhu 0006, Minglu Li 0001, Jie Wang 0006 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2013 | A Social-Aware Service Recommendation Approach for Mashup CreationabstractMashup is a user-centric approach to create value-added new services by utilizing and recombining existing service components. However, as services become increasingly more spontaneous and prevalent on the Internet, finding suitable services from which to develop a mashup based on users' explicit and implicit requirements remains a daunting task. Several approaches already exist for recommending specific services for users but they are limited to proposing only services with similar functionality. In order to recommend a set of suitable services for a general mashup based on users' functional specifications, a novel social-aware service recommendation approach, where multi-dimensional social relationships among potential users, topics, mashups, and services are described by a coupled matrix model, is proposed in this paper. Accordingly, a factorization algorithm is designed to predict unobserved relationships, and as a result, a comprehensive service recommendation model can be readily constructed. Experimental results for a realistic mashup data set indicate that the proposed approach outperforms other state-of-the-art methods. Wenxing Xu, Jian Cao 0001, Liang Hu 0004, Jie Wang 0006, Minglu Li 0001 |
ICWS | 4 |
| 2013 | Cross-Domain Collaborative Filtering via Bilinear Multilevel Analysis
Liang Hu 0004, Jian Cao 0001, Guandong Xu, Jie Wang 0006, Zhiping Gu, Longbing Cao |
IJCAI | 4 |
| 2013 | H-Tree: An efficient index structure for event matching in publish/subscribe systems
Shiyou Qian, Jian Cao 0001, Yanmin Zhu 0006, Minglu Li 0001, Jie Wang 0006 |
Networking | 5 |
| 2013 | An event view specification approach for Supporting Service process collaborationabstractABSTRACT Designing and implementing an interoperable and flexible service process collaboration strategy is one of key issues for business to business integrations. To better support service process collaboration, an event view model is proposed, which is composed of a set of event types and their dependency relationships. It provides a general and flexible way to define a public view of a service process model and serves as the basis for defining service process collaboration protocols. In the paper, the basic concepts and a system framework for event‐based service process collaboration are first introduced. The definitions of event and the dependency relationships among event types are then presented. Especially, how to identify dependency relationships among composite event types is studied in detail. After discussing the definition of event view and its specifying approach, a procedure for transforming a BPEL process model into an event model and deriving dependencies among events is given. Finally, a case study is presented, and some implementation issues for defining and publishing an event view are discussed. Copyright © 2013 John Wiley & Sons, Ltd. Jian Cao 0001, Jie Wang 0006, Haiyan Zhao 0002, Minglu Li 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | A pattern fusion model for multi-step-ahead CPU load prediction
Dingyu Yang, Jian Cao 0001, Jiwen Fu, Jie Wang 0006, Jianmei Guo |
J. Syst. Softw. | 4 |
| 2013 | A service process optimization method based on model refinement
Jian Cao 0001, Jie Wang 0006, Haiyan Zhao 0002 |
J. Supercomput. | 2 |
| 2012 | A Service Intermediary Agent Framework for Web Service IntegrationabstractBecause web services and agent technology can help each other, to integrate them together is becoming a trend. Most of current research works focus on developing an integration framework to provide gateways that can connect the agent and web service worlds. In the paper, the service intermediary agent is proposed as a value-added professional service provider, which can manage a group of inherently related Web services. It organizes Web services as a set of plans and reacts to the requests by selecting and executing the appropriate plans. In order to process the incoming requests intelligently, the semantic goal structure is employed to facilitate plan organizing and plan selecting. The structure of the service intermediary agent and its working process are given. The semantic goal structure together with the deliberation process is given. The case study and experiments are also provided. Jian Cao 0001, Jie Wang 0006, Liang Hu 0004, Rujie Lai |
APSCC | 2 |
| 2012 | Fuzzy-QFD approach based decision support model for licensor selection
Yi-Kai Juan, Jie Wang 0006, Kai-Meng Li, Colin Ong |
Expert Syst. Appl. | 3 |
| 2010 | Service Process Collaboration Based on Event View ModelabstractIn order to better support service collaboration for complex tasks, an event view model is proposed to facilitate the activation and integration of individual service processes. An event view is composed of a set of event types and their dependency relationships, it provides a public view of a service process model and serves as the basis for defining process collaboration protocols. After discussing the definition of event view and the service process collaboration framework based on event view, its construction approach and a procedure for transforming a service process model into an event model are proposed. Finally, a case study is presented and some implement issues are discussed. Jian Cao 0001, Jie Wang 0006 |
APSCC | 2 |
| 2010 | A dynamically self-configurable service process engine
Jian Cao 0001, Haiyan Zhao 0002, Minglu Li 0001, Jie Wang 0006 |
World Wide Web | 4 |
| 2007 | An Ontology-Based Framework for Building Adaptable Knowledge Management Systems
Jianmei Guo, Jie Wang 0006 |
KSEM | 4 |
| 2006 | An interactive service customization model
Jian Cao 0001, Jie Wang 0006, Kincho H. Law, Shensheng Zhang, Minglu Li 0001 |
Inf. Softw. Technol. | 2 |
| 2004 | A dynamically reconfigurable system based on workflow and service agents
Jian Cao 0001, Jie Wang 0006, Shensheng Zhang, Minglu Li 0001 |
Eng. Appl. Artif. Intell. | 2 |