VLDB 2026 Research / reviewers in the wild / expert
Rubén Ruiz
dblp:20/165
· DBLP profile ↗
44ranked-venue papers
0as first author
19since 2021 · last 2025
0000-0003-3295-3888ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 7 since 2021Human-computer interaction and ubiquitous computing · 10 · 3 since 2021Systems, architecture and hardware · 9 · 2 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Theory of computation · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Supplier Selection and Material Sourcing With Multiuncertainties in Cloud Manufacturing Using Reinforcement LearningabstractCompared to traditional manufacturing, there are several unique characteristics in cloud manufacturing (CMfg): more candidate suppliers, more material types, more supply modes, and broader geographically distributed suppliers. These characteristics lead to a huge set of candidate supply plans with several uncertainties in logistics. It is a great challenge to effectively and efficiently select appropriate suppliers for each type of material. In this article, we consider a SSMS problem with multiuncertainties in logistics to minimize the total cost of a CMfg manufacturing enterprise. The delivery time is stochastic along with the consideration of stochastic disruption and loss in logistics. Based on state-and-transition modeling, a stochastic dynamic programming model is developed for the problem under study. By integrating proximal policy optimization (PPO) with recurrent neural network (RNN) and expectation model (EM), a stochastic optimization method PPO-REM is proposed to minimize the cost by effectively selecting suppliers and intelligently making make-or-buy decisions under uncertainties. The proposed method is evaluated by comparing to other reinforcement learning (RL) methods and existing methods for similar problems over a comprehensive set of numerical experiments with some real-world data. Experimental results show that the proposed PPO-REM converges faster with a higher reward than other RL methods, and it costs the least as compared to existing methods for similar problems. Zhongyi Chen, Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Scheduling Workflows With Limited Budget to Cloud Server and Serverless ResourcesabstractServerless functions (SFs) and on-demand virtual machines (VMs) are common cloud resources for scientific workflow applications, which are widespread in many fields. SFs are paid by actual running time with higher unit costs and higher resource utilization than VMs which are paid by billing time units. Generally, each application is executed on a limited budget. In this article, we study the challenging cloud workflow scheduling problem with a limited budget to minimize makespan in a hybridization of SFs and on-demand VMs for which the BCWS (Budget Constrained Workflow Scheduling) algorithm is proposed. Methods are developed to determine the task execution order, rent cloud resources and map tasks to resources respectively. Together with initial schedule construction and schedule improvement policies, these procedures are repeatedly applied in BCWS. The proposed algorithm is evaluated by comparing it to existing algorithms for similar problems over a comprehensive set of workflow instances. Experimental results show that the proposed algorithm significantly reduces the makespan with a hybrid configuration of VMs and SFs compared to the server only or the serverless only configurations and outperforms the compared algorithms which are the best existing ones for similar problems. Xiaoping Li 0001, Long Chen 0021, Rubén Ruiz |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | Smart Offloading Computation-intensive & Delay-intensive Tasks of Real-time Workflows in Mobile Edge ComputingabstractIn MEC, many deadline-constrained real-time work-flows with computation-intensive and/or delay-sensitive tasks are common in intelligent mobile devices (MDs). Though a task can be executed by either the local MD or an MEC server, the tasks of each work-flow are constrained by complex precedences, and real-time task offloading is somewhat tricky. In this paper, we consider the task offloading problem for stochastic work-flows with soft deadline constraints to minimize total tardiness and proposed an online RL-based offloading algorithm. In the algorithm, realtime tasks are dynamically partitioned into partial precedences in terms of which real-time RL states are constructed. Adaptive offloading actions are developed to determine task execution sequences for different states to optimize total tardiness. Experimental results show that the proposed online offloading algorithm outperforms the compared ones. Haihong Zhu, Xiaoping Li 0001, Long Chen 0021, Rubén Ruiz |
ICWS | 4 |
| 2023 | Mixed-Integer Programming vs. Constraint Programming for Shop Scheduling Problems: New Results and OutlookabstractConstraint programming (CP) has been recently in the spotlight after new CP-based procedures have been incorporated into state-of-the-art solvers, most notably the CP Optimizer from IBM. Classical CP solvers were only capable of guaranteeing the optimality of a solution, but they could not provide bounds for the integer feasible solutions found if interrupted prematurely due to, say, time limits. New versions, however, provide bounds and optimality guarantees, effectively making CP a viable alternative to more traditional mixed-integer programming (MIP) models and solvers. We capitalize on these developments and conduct a computational evaluation of MIP and CP models on 12 select scheduling problems. 1 We carefully chose these 12 problems to represent a wide variety of scheduling problems that occur in different service and manufacturing settings. We also consider basic and well-studied simplified problems. These scheduling settings range from pure sequencing (e.g., flow shop and open shop) or joint assignment-sequencing (e.g., distributed flow shop and hybrid flow shop) to pure assignment (i.e., parallel machine) scheduling problems. We present MIP and CP models for each variant of these problems and evaluate their performance over 17 relevant and standard benchmarks that we identified in the literature. The computational campaign encompasses almost 6,623 experiments and evaluates the MIP and CP models along five dimensions of problem characteristics, objective function, decision variables, input parameters, and quality of bounds. We establish the areas in which each one of these models performs well and recognize their conceivable reasons. The obtained results indicate that CP sets new limits concerning the maximum problem size that can be solved using off-the-shelf exact techniques. History: Accepted by Pascal Van Hentenryck, Area Editor for Computational Modeling: Methods & Analysis. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.1287 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2021.0326 ) at ( http://dx.doi.org/10.5281/zenodo.7541223 ). B. Naderi 0001, Rubén Ruiz, Vahid Roshanaei |
INFORMS J. Comput. | 2 |
| 2023 | An incremental learning evolutionary algorithm for many-objective optimization with irregular Pareto fronts
Mingjing Wang, Xiaoping Li 0001, Long Chen 0021, Huiling Chen 0001, Rubén Ruiz |
Inf. Sci. | 6 |
| 2023 | Feature Selection With Maximal Relevance and Minimal Supervised RedundancyabstractFeature selection (FS) for classification is crucial for large-scale images and bio-microarray data using machine learning. It is challenging to select informative features from high-dimensional data which generally contains many irrelevant and redundant features. These features often impede classifier performance and misdirect classification tasks. In this article, we present an efficient FS algorithm to improve classification accuracy by taking into account both the relevance of the features and the pairwise features correlation in regard to class labels. Based on conditional mutual information and entropy, a new supervised similarity measure is proposed. The supervised similarity measure is connected with feature redundancy minimization evaluation and then combined with feature relevance maximization evaluation. A new criterion max-relevance and min-supervised-redundancy (MRMSR) is introduced and theoretically proved for FS. The proposed MRMSR-based method is compared to seven existing FS approaches on several frequently studied public benchmark datasets. Experimental results demonstrate that the proposal is more effective at selecting informative features and results in better competitive classification performance. Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Cybern. | 3 |
| 2023 | Failure-Aware Elastic Cloud Workflow SchedulingabstractWith an increasing complexity and functionality in cloud data centers, fault tolerance becomes an essential requirement for tasks executed in clouds, especially for workflows with task precedences. Hosts and network devices are the main physical components in a cloud data center. The PB (Primary-Backup) model is a desirable approach to fault tolerance. Many PB-based workflow scheduling algorithms have been proposed for host faults. However, only a few studies focus on cloud workflow scheduling considering network device faults. This paper analyzes the fault-tolerant properties for scheduling dependent tasks and migrating VMs based on the PB model, considering both host and network device faults in a cloud data center. A failure-aware elastic cloud workflow scheduling algorithm is designed for both host and network device fault tolerance. Additionally, an elastic resource provisioning mechanism is proposed and incorporated into the proposed algorithm to improve resource utilization. Performance evaluations on both randomly generated and real-world workflows show that the proposal effectively improves resource utilization while guaranteeing fault tolerance. Guangshun Yao, Xiaoping Li 0001, Qian Ren, Rubén Ruiz |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | Scheduling multi-tenant cloud workflow tasks with resource reliability
Xiaoping Li 0001, Dongyuan Pan, Rubén Ruiz |
Sci. China Inf. Sci. | 4 |
| 2022 | Task Scheduling for Spark Applications With Data Affinity on Heterogeneous ClustersabstractThe Internet of Things (IoT)-enabled applications use sensors and actuators to collect big data, which are processed by big data models, e.g., Spark. Generally, data processing tasks are precedence constrained and the computation results are transmitted to other IoT devices. In this article, we consider the Spark workflow problem of scheduling tasks with data affinity to heterogeneous servers to minimize the maximum completion time. In a Spark instance, jobs are precedence constrained and stages for each job are also precedence constrained. There are a large number of topological stage orders. A balance between task execution times, determined by heterogeneous servers, and transmission times caused by data affinity is difficult to achieve. A scheduling optimization algorithm framework is proposed, which consists of five components: 1) temporal parameter calculation; 2) ready stage adding; 3) task sequencing; 4) resource allocation; and 5) schedule improvement. Strategies for each component are developed. The algorithmic components are statistically calibrated over a comprehensive set of instances. The proposed algorithm is compared to two modified classic algorithms for similar problems on typical scientific workflow instances. The experimental results demonstrate the effectiveness of the proposal for the considered problem. Zhang Xiaodong, Xiaoping Li 0001, Houan Du, Rubén Ruiz |
IEEE Internet Things J. | 4 |
| 2022 | A referenced iterated greedy algorithm for the distributed assembly mixed no-idle permutation flowshop scheduling problem with the total tardiness criterion
Yuanzhen Li, Quan-Ke Pan, Rubén Ruiz, Hongyan Sang |
Knowl. Based Syst. | 3 |
| 2022 | A Survey on Sparse Learning Models for Feature SelectionabstractFeature selection is important in both machine learning and pattern recognition. Successfully selecting informative features can significantly increase learning accuracy and improve result comprehensibility. Various methods have been proposed to identify informative features from high-dimensional data by removing redundant and irrelevant features to improve classification accuracy. In this article, we systematically survey existing sparse learning models for feature selection from the perspectives of individual sparse feature selection and group sparse feature selection, and analyze the differences and connections among various sparse learning models. Promising research directions and topics on sparse learning models are analyzed. Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Cybern. | 3 |
| 2022 | MapReduce Task Scheduling in Heterogeneous Geo-Distributed Data CentersabstractDifferent data transmission times, processing times which are difficult to predict and node-dependent access times make MapReduce task scheduling rather complex. In this article, we consider the problem of scheduling MapReduce tasks to heterogeneous geo-distributed data centers to minimize the total tardiness. A new architecture is constructed to analyze data in the considered scenario. We model distinct data transmission levels, inter- and intra- data centers and heterogeneity of nodes mathematically. An algorithm framework is proposed to schedule MapReduce tasks to heterogeneous nodes in geographically distributed data centers. The proposed algorithm is suitable for both Hadoop MRv1 and MRv2. In terms of the number of idle containers detected in each heartbeat, the same number of tasks are selected from a sorted job sequence. For the map and reduce phases, two measurements are developed with data locality and completion time, respectively, based on which the classical Hungarian algorithm is adopted to optimally assign selected tasks to corresponding idle containers. Components and parameters of the proposal are statistically calibrated over a large set of random instances. A comparison of the proposed algorithm to existing methods for similar problems is carried out. Experimental results demonstrate the proposal is effective for the considered problem. Xiaoping Li 0001, Fuchao Chen, Rubén Ruiz, Jie Zhu 0002 |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | Energy-Aware Cloud Workflow Applications Scheduling With Geo-Distributed DataabstractElectricity prices differ during different time periods and change from place to place. Cloud workflow applications often require geo-distributed data which is transmitted among heterogeneous servers in intra- and inter- data centers. Such varying electricity prices and data transmission time bring great challenges when optimizing the energy cost for scheduling tasks in workflow applications to heterogeneous servers in cloud data centers. In this article, we minimize the total electricity cost in a deadline constrained energy-aware workflow scheduling problem with data being geographically distributed across data centers. A scheduling algorithm is proposed. Strategies are developed to sequence workflow applications, divide deadlines and sort tasks. An adaptive local search method is presented to improve solutions during the search process which dynamically balances intensification using neighborhood structures of increasing size. Components and parameter values are statistically calibrated over a comprehensive set of random instances. The proposed algorithm is compared to modified classical algorithms for similar problems. Experimental results demonstrate the effectiveness of the proposal for the considered problem. Xiaoping Li 0001, Rubén Ruiz, Jie Zhu 0002 |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | Energy Utilization Task Scheduling for MapReduce in Heterogeneous ClustersabstractNowadays, energy costs are the most important factor in cloud computing. Therefore, the implementation of energy-aware task scheduling methods is of utmost importance. A task scheduling framework considering deadlines, data locality and resource utilization is proposed to save on energy costs in heterogeneous clusters. The framework consists of task list construction, task scheduling and slot list updating. In terms of deadline constraints, number of job slots allocated and possible processing times of jobs, a new job sequence is proposed to construct an reasonable task list. Tasks are scheduled to promising slots from their rack-local servers, cluster-local servers and remote servers in the produced task scheduling, which greatly improves data locality. After the assignment among tasks and slots, an update of available slots in clusters is proposed not only to find available slots but also to improve server resource utilization using fuzzy logic with the available number of slots according to current CPU, memory and bandwidth utilization. Experimental results show that the proposed heuristic results in lower energy consumption than the adapted existing algorithms with a variable total number of slots. Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | A Hybrid Fault-Tolerant Scheduling for Deadline-Constrained Tasks in Cloud SystemsabstractAmong multiple fault-tolerant strategies, resubmission, and replication are fundamental and widely recognized in distributed computing systems. In recent years, many algorithms based on replication or resubmission have been proposed. However, few of them consider these two techniques together, especially in Cloud systems. In this article, we propose a Hybrid Fault-Tolerant Scheduling Algorithm (HFTSA) for independent tasks with deadlines by integrating the above techniques in virtualized Cloud systems. During the task scheduling process, HFTSA selects fault-tolerant strategies from resubmission and replication for each accepted task based on the characteristics of both task and Cloud resources and then reserves suitable resources. During the task execution process, HFTSA adopts an online adjustment scheme for fault-tolerant strategies of some tasks if necessary while providing an online scheduling scheme for faults. Moreover, an elastic resource provisioning mechanism is designed and incorporated into HFTSA to dynamically adjust the provided resources to improve resource utilization. Experiments on a real cloud platform and a simulated platform are conducted to verify the effectiveness of the proposed HFTSA. The results demonstrate that HFTSA can provide an efficient fault-tolerant scheduling strategy for deadline-constrained tasks with high resource utilization and performs better than corresponding competitors. Guangshun Yao, Qian Ren, Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Serv. Comput. | 5 |
| 2021 | Scheduling Microservice-based Workflows to Containers in On-demand Cloud ResourcesabstractThough microservices process and communicate with lightweight mechanisms, finer tasks result in much more complicated precedence constraints. Different tasks have distinct resource requirements and different VMs (Virtual Machine) have various configurations and prices. In this paper, we consider the problem of scheduling microservice tasks of workflow applications to containers configured on on-demand VMs to minimize the total rental cost. The problem is mathematically modelled using integer programming and an algorithm framework is proposed. For dynamic available containers and resource requirements, a task scheduling heuristic is presented for scheduling precedence-constrained or independent tasks to available containers. All parameters and components of the proposed algorithm framework are statistically calibrated by the Analysis of Variance technique on a large number of random instances. Performance of the the proposed algorithm is verified over a lot of instances. Xiaoping Li 0001, Rubén Ruiz |
CSCWD | 3 |
| 2021 | Hybrid Resource Provisioning for Cloud Workflows with Malleable and Rigid TasksabstractIn cloud computing, reserved and on-demand instances are generally provided by service providers. Hybridization of the two alternatives can considerably save costs when renting resources from the cloud. However, it is a big challenge to determine the appropriate amount of reserved and on-demand resources in terms of users’ requirements. In this paper, the workflow scheduling problem with both reserved and on-demand instances is considered. The objective is to minimize the total rental cost under deadline constrains. The considered problem is mathematically modeled. A multiple sequence-based earliest finish time method is proposed to construct schedules for the workflows. Four different rules are used to generate initial task allocation sequences. Types and quantities of resources are determined by a free time block-based schedule construction mechanism. New sequences are generated by a variable neighborhood search method. Experimental and statistical analyses and results demonstrate that the proposed algorithm algorithm generates considerable cost savings when compared to the algorithms with only on-demand or reserved instances. Long Chen 0021, Xiaoping Li 0001, Yucheng Guo, Rubén Ruiz |
IEEE Trans. Cloud Comput. | 4 |
| 2021 | Multi-Queue Request Scheduling for Profit Maximization in IaaS CloudsabstractIn cloud computing, service providers rent heterogeneous servers from cloud providers, i.e., Infrastructure as a Service (IaaS), to meet requests of consumers. The heterogeneity of servers and impatience of consumers pose great challenges to service providers for profit maximization. In this article, we transform this problem into a multi-queue model where the optimal expected response time of each queue is theoretically analyzed. A multi-queue request scheduling algorithm framework is proposed to maximize the total profit of service providers, which consists of three components: request stream splitting, requests allocation, and server assignment. A request stream splitting algorithm is designed to split the arriving requests to minimize the response time in the multi-queue system. An allocation algorithm, which adopts a one-step improvement strategy, is developed to further optimize the response time of the requests. Furthermore, an algorithm is developed to determine the appropriate number of required servers of each queue. After statistically calibrating parameters and algorithm components over a comprehensive set of random instances, the proposed algorithms are compared with the state-of-the-art over both simulated and real-world instances. The results indicate that the proposed multi-queue request scheduling algorithm outperforms the other algorithms with acceptable computational time. Shuang Wang 0012, Xiaoping Li 0001, Quan Z. Sheng, Rubén Ruiz, Amin Beheshti |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | Group Scheduling With Nonperiodical Maintenance and Deteriorating EffectsabstractIn this paper, we consider single-machine group scheduling with nonperiodical maintenance and deteriorating effects. Nonperiodical maintenance, which has unfixed maintaining interval or the number of jobs in each group is unfixed, results in a variable number of groups. Deteriorating effects lead to longer processing times of which the deterioration index depends on job grouping. This problem is of significance in different production settings and is much more difficult than and general that other simpler single-machine group scheduling problems. Making use of historical processing times, we construct the actual processing time model for jobs. We prove that the problem under study is NP-hard. By transforming the optimization objective, properties are discovered and two batch-based heuristics are presented for small size problems. To further improve the effectiveness for large size problems, an iterated greedy algorithm is proposed being its main advantages simplicity and effectiveness. The proposed methods are evaluated over a large number of random instances with calibrated parameters and components. Comprehensive computational and statistical analyses demonstrate the superiority of the methods proposed over adapted existing approaches. Xiaoping Li 0001, Rubén Ruiz, Haihong Zhu |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Allocating MapReduce workflows with deadlines to heterogeneous servers in a cloud data center
Xiaoping Li 0001, Rubén Ruiz, Hanchuan Xu |
Serv. Oriented Comput. Appl. | 3 |
| 2020 | Performance Analysis for Heterogeneous Cloud Servers Using Queueing TheoryabstractIn this article, we consider the problem of selecting appropriate heterogeneous servers in cloud centers for stochastically arriving requests in order to obtain an optimal tradeoff between the expected response time and power consumption. Heterogeneous servers with uncertain setup times are far more common than homogenous ones. The heterogeneity of servers and stochastic requests pose great challenges in relation to the tradeoff between the two conflicting objectives. Using the Markov decision process, the expected response time of requests is analyzed in terms of a given number of available candidate servers. For a given system availability, a binary search method is presented to determine the number of servers selected from the candidates. An iterative improvement method is proposed to determine the best servers to select for the considered objectives. After evaluating the performance of the system parameters on the performance of algorithms using the analysis of variance, the proposed algorithm and three of its variants are compared over a large number of random and real instances. The results indicate that proposed algorithm is much more effective than the other four algorithms within acceptable CPU times. Shuang Wang 0012, Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Computers | 3 |
| 2020 | Scheduling Periodical Multi-Stage Jobs With Fuzziness to Elastic Cloud ResourcesabstractWe investigate a workflow scheduling problem with stochastic task arrival times and fuzzy task processing times and due dates. The problem is common in many real-time and workflow-based applications, where tasks with fixed stage number and linearly dependency are executed on scalable cloud resources with multiple price options. The challenges lie in proposing effective, stable, and robust algorithms under stochastic and fuzzy tasks. A triangle fuzzy number-based model is formulated. Two metrics are explored: the cost and the degree of satisfaction. An iterated heuristic framework is proposed to periodically schedule tasks, which consists of a task collection and a fuzzy task scheduling phases. Two task collection strategies are presented and two task prioritization strategies are employed. In order to achieve a high satisfaction degree, deadline constraints are defined at both job and task levels. By designing delicate experiments and applying sophisticated statistical techniques, experimental results show that the proposed algorithm is more effective and robust than the two existing methods. Jie Zhu 0002, Xiaoping Li 0001, Rubén Ruiz, Wei Li 0058, Haiping Huang, Albert Y. Zomaya |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Resource Renting for Periodical Cloud Workflow ApplicationsabstractCloud computing is a new resource provisioning mechanism, which represents a convenient way for users to access different computing resources. Periodical workflow applications commonly exist in scientific and business analysis, among many other fields. One of the most challenging problems is to determine the right amount of resources for multiple periodical workflow applications. In this paper, the periodical workflow applications scheduling problem with total renting cost minimization is considered. The novelty of this work relies precisely on this objective function, which is more realistic in practice than the more commonly considered makespan minimization. An integer programming model is constructed for the problem under study. A Precedence Tree based Heuristic (PTH) is developed which considers three types of initial schedule construction methods. Based on the initial schedule, two improvement procedures are presented. The proposed methods are compared with existing algorithms for the related makespan based multiple workflow scheduling problem. Experimental and statistical results demonstrate the effectiveness and efficiency of the proposed algorithm. Long Chen 0021, Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Serv. Comput. | 3 |
| 2019 | Cost Minimization for Service Providers with Impatient Consumers in Cloud ComputingabstractIn this paper, we consider the cost minimization problem for scheduling stochastic service requests to heterogenous servers in cloud computing. Service requests are impatient with different maximizing waiting time. Using queuing theory, a queuing system model is constructed. An algorithm framework is proposed to minimize the cost. The actual expected waiting time of service requests is analyzed. The rejection probability of the system is obtained. Comparing the rejection probability of the system to a given system availability, suitable servers are selected to minimize the cost. Based on the proposed framework, algorithms with different components are compared. Experimental results show that the algorithm with the mixed server selection strategy outperforms the others on efficiency. Shuang Wang 0012, Xiaoping Li 0001, Rubén Ruiz |
CSCWD | 3 |
| 2019 | Feature Selection via Adaptive Spectral Clustering based on Joint Mutual InformationabstractFeature selection plays an important role in big data mining and pattern recognition, which includes selecting a subset of the most informative features that produces compatible results as the original entire set of features. A new similarity measure is proposed in terms of information theory, in particular joint mutual information between features with respect to class labels. Based on the similarity measure, an adaptive spectral clustering that could determine cluster number adaptively is proposed. Furthermore, we propose a new feature selection algorithm called Adaptive Spectral Clustering based on Joint Mutual Information (ASC-JMI) which can effectively and efficiently deal with both irrelevant and redundant features, and select a high-quality feature subset. The experimental results on five high-dimensional benchmark datasets demonstrate that the ASC-JMI is effective in selecting the informative features and obtains competitive classification performance. Xiaoping Li 0001, Rubén Ruiz |
CSCWD | 3 |
| 2019 | Resource Provisioning for Task-Batch Based Workflows with Deadlines in Public CloudsabstractTo meet the dynamic workload requirements in widespread task-batch based workflow applications, it is important to design algorithms for DAG-based platforms (such as Dryad, Spark and Pegasus) to rent virtual machines from public clouds dynamically. In terms of depths and functionalities, tasks of different task-batches are merged into task-units. A unit-aware deadline division method is investigated for properly dividing workflow deadlines to task deadlines so as to minimize the utilization of rented intervals. A rule-based task scheduling method is presented for allocating tasks to time slots of rented Virtual Machines (VMs) with a task right shifting operation and a weighted priority composite rule. A Unit-aware Rule-based Heuristic (URH) is proposed for elastically provisioning VMs to task-batch based workflows to minimize the rental cost in DAG-based cloud platforms. Effectiveness of the proposed URH methods is verified by comparing them against two adapted existing algorithms for similar problems on some realistic workflows. Zhicheng Cai, Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Cloud Comput. | 3 |
| 2019 | Weighted General Group Lasso for Gene Selection in Cancer ClassificationabstractRelevant gene selection is crucial for analyzing cancer gene expression datasets including two types of tumors in cancer classification. Intrinsic interactions among selected genes cannot be fully identified by most existing gene selection methods. In this paper, we propose a weighted general group lasso (WGGL) model to select cancer genes in groups. A gene grouping heuristic method is presented based on weighted gene co-expression network analysis. To determine the importance of genes and groups, a method for calculating gene and group weights is presented in terms of joint mutual information. To implement the complex calculation process of WGGL, a gene selection algorithm is developed. Experimental results on both random and three cancer gene expression datasets demonstrate that the proposed model achieves better classification performance than two existing state-of-the-art gene selection methods. Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Cybern. | 3 |
| 2018 | A Fast Algorithm for Finding the Bi-objective Shortest Path in Complicated NetworksabstractThe bi-objective shortest path problem exists in many practical applications with complex networks. It is much time-consuming for searching all non-dominated solutions, especially for large size problems. However, only a small number of them are crucial for users' decision making. In this paper, the Pareto front is segmented by grids. The grids are determined according to the two requirements given by users. Non-dominated solutions are assumed to be similar and any of them meets the user's requirements, i.e., finding only one solution is necessary for each grid. A fast algorithm is proposed for finding the bi-objective shortest path in such user-driven problems. Experimental results illustrate the efficiency and effectiveness of the proposed algorithm. Xiaoping Li 0001, Rubén Ruiz |
CSCWD | 3 |
| 2018 | Price forecasting for spot instances in Cloud computing
Zhicheng Cai, Xiaoping Li 0001, Rubén Ruiz, Qianmu Li |
Future Gener. Comput. Syst. | 3 |
| 2018 | Idle block based methods for cloud workflow scheduling with preemptive and non-preemptive tasks
Long Chen 0021, Xiaoping Li 0001, Rubén Ruiz |
Future Gener. Comput. Syst. | 3 |
| 2018 | An iterated greedy heuristic for no-wait flow shops with sequence dependent setup times, learning and forgetting effects
Xiaoping Li 0001, Rubén Ruiz, Shaochun Sui |
Inf. Sci. | 3 |
| 2018 | An Iterated Greedy Heuristic for Mixed No-Wait Flowshop ProblemsabstractThe mixed no-wait flowshop problem with both wait and no-wait constraints has many potential real-life applications. The problem can be regarded as a generalization of the traditional permutation flowshop and the no-wait flowshop. In this paper, we study, for the first time, this scheduling setting with makespan minimization. We first propose a mathematical model and then we design a speed-up makespan calculation procedure. By introducing a varying number of destructed jobs, a modified iterated greedy algorithm is proposed for the considered problem which consists of four components: 1) initialization solution construction; 2) destruction; 3) reconstruction; and 4) local search. To further improve the intensification and efficiency of the proposal, insertion is performed on some neighbor jobs of the best position in a sequence during the initialization, solution construction, and reconstruction phases. After calibrating parameters and components, the proposal is compared with five existing algorithms for similar problems on adapted Taillard benchmark instances. Experimental results show that the proposal always obtains the best performance among the compared methods. Xiaoping Li 0001, Rubén Ruiz, Shaochun Sui |
IEEE Trans. Cybern. | 3 |
| 2018 | Scheduling Stochastic Multi-Stage Jobs to Elastic Hybrid Cloud ResourcesabstractWe consider a special workflow scheduling problem in a hybrid-cloud-based workflow management system in which tasks are linearly dependent, compute-intensive, stochastic, deadline-constrained and executed on elastic and distributed cloud resources. This kind of problems closely resemble many real-time and workflow-based applications. Three optimization objectives are explored: number, usage time and utilization of rented VMs. An iterated heuristic framework is presented to schedule jobs event by event which mainly consists of job collecting and event scheduling. Two job collecting strategies are proposed and two timetabling methods are developed. The proposed methods are calibrated through detailed designs of experiments and sound statistical techniques. With the calibrated components and parameters, the proposed algorithm is compared to existing methods for related problems. Experimental results show that the proposal is robust and effective for the problems under study. Jie Zhu 0002, Xiaoping Li 0001, Rubén Ruiz, Xiaolong Xu 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | Cloud Workflow Scheduling with Deadlines and Time Slot AvailabilityabstractAllocating service capacities in cloud computing is based on the assumption that they are unlimited and can be used at any time. However, available service capacities change with workload and cannot satisfy users' requests at any time from the cloud provider's perspective because cloud services can be shared by multiple tasks. Cloud service providers provide available time slots for new user's requests based on available capacities. In this paper, we consider workflow scheduling with deadline and time slot availability in cloud computing. An iterated heuristic framework is presented for the problem under study which mainly consists of initial solution construction, improvement, and perturbation. Three initial solution construction strategies, two greedy- and fair-based improvement strategies and a perturbation strategy are proposed. Different strategies in the three phases result in several heuristics. Experimental results show that different initial solution and improvement strategies have different effects on solution qualities. Xiaoping Li 0001, Lihua Qian, Rubén Ruiz |
IEEE Trans. Serv. Comput. | 3 |
| 2018 | Methods for Scheduling Problems Considering Experience, Learning, and Forgetting EffectsabstractWorkers with different levels of experience and knowledge have different effects on job processing times. By taking into account 1) the sum-of-processing-time; 2) the job-position; and 3) the experience of workers, a more general learning model is introduced for scheduling problems. We show that this model generalizes existing ones and brings the consideration of learning and forgetting effects closer to reality. We demonstrate that some single machine scheduling problems are polynomially solvable under this general model. Considering the forgetting effect caused by the idle time on the second machine, we construct a learning-forgetting model for the two-machine permutation flow shop scheduling problem with makespan minimization. A branch-and-bound method and four heuristics are presented to find optimal and approximate solutions, respectively. The proposed heuristics are evaluated over a large number of randomly generated instances. Experimental results show that the proposed heuristics are effective and efficient. Xiaoping Li 0001, Yulu Jiang, Rubén Ruiz |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2017 | Cloud workflow scheduling with on-demand and spot block instancesabstractCloud computing enables users to access different resources conveniently based on the `pay-as-you-go' model. However, the unit cost of these on-demand instances are usually high. The spot instances provide a dynamic and cheaper manner for renting resources from the cloud. However, failures are often occurred due to the fluctuations of the price of the spot instance. It is a big challenge to determine the appropriate amounts of spot and on-demand resources in terms of users' requirements. In this paper, the workflow scheduling problem with both spot and on-demand instances is considered. The objective is to minimize the total renting cost under deadline constrains. An idle time block-based method is proposed to construct schedules for workflow applications. Schedules are improved by a forward and backward moving mechanism. Experimental and statistical results demonstrate the effectiveness of the proposed algorithm over a lot of tests with different sizes. Long Chen 0021, Xiaoping Li 0001, Rubén Ruiz |
CSCWD | 3 |
| 2017 | Trust constrained workflow scheduling in cloud computingabstractIn cloud environments, trust is necessary because services have the characteristics of uncertainty, dynamic, false or fraudulent which usually make users difficult to obtain desired services. In this paper, we consider trust-oriented workflow scheduling with temporal constraints including service setup times and workflow deadlines. The considered problem is mathematically modeled. A behavior-based trust model is established to assess trusts of services. An iterative adjustment heuristic framework is proposed which consists of initial solution construction, solution sets generation and adjustment. Three heuristic algorithms are developed and compared with the best existing methods for similar workflow scheduling problems. Experimental results demonstrate effectiveness of the proposal. Xiaoping Li 0001, Taoyong Ding, Rubén Ruiz |
SMC | 4 |
| 2017 | A delay-based dynamic scheduling algorithm for bag-of-task workflows with stochastic task execution times in clouds
Zhicheng Cai, Xiaoping Li 0001, Rubén Ruiz, Qianmu Li |
Future Gener. Comput. Syst. | 3 |
| 2017 | An Exact Algorithm for the Shortest Path Problem With Position-Based Learning EffectsabstractThe shortest path problems (SPPs) with learning effects (SPLEs) have many potential and interesting applications. However, at the same time they are very complex and have not been studied much in the literature. In this paper, we show that learning effects make SPLEs completely different from SPPs. An adapted A* (AA*) is proposed for the SPLE problem under study. Though global optimality implies local optimality in SPPs, it is not the case for SPLEs. As all subpaths of potential shortest solution paths need to be stored during the search process, a search graph is adopted by AA* rather than a search tree used by A*. Admissibility of AA* is proven. Monotonicity and consistency of the heuristic functions of AA* are redefined and the corresponding properties are analyzed. Consistency/monotonicity relationships between the heuristic functions of AA* and those of A* are explored. Their impacts on efficiency of searching procedures are theoretically analyzed and experimentally evaluated. Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2016 | Scheduling Stochastic Multi-stage Jobs on Elastic Computing Services in Hybrid CloudsabstractIn this paper, we consider the widespread multi-stage job scheduling problem (e.g., in big data processed by MapReduce) in which jobs arrive at hybrid cloud systems stochastically. The objective is to minimize the number of elastic computing instances. Along with hard deadlines of jobs, the problem under study is NP-hard in strong sense. In terms of initial job priorities, timetables are constructed by adjusting job priorities adaptively and generating feasible schedules iteratively. Job sequences are generated by two simple dispatching rules. A fast local search heuristic and a rescheduling process are developed for improving the obtained sequences. Experimental results show that the proposed heuristics improve the utilization of computing resources effectively while meeting the cloud service quality requirements. Jie Zhu 0002, Xiaoping Li 0001, Rubén Ruiz, Xiaolong Xu 0002, Yi Zhang 0009 |
ICWS | 3 |
| 2016 | Heuristics for periodical batch job scheduling in a MapReduce computing framework
Xiaoping Li 0001, Tianze Jiang, Rubén Ruiz |
Inf. Sci. | 3 |
| 2014 | Simple constructive heuristics for the Distributed Assembly Permutation Flowshop Scheduling Problem with sequence dependent setup timesabstractIn this paper we consider a realistic production setting, referred to the Distributed Assembly Permutation Flowshop Scheduling Problem or DAPFSP in short. The problem consists of two stages, production and assembly. The paper is an extension of the recent work by Hatami et al. [6], that the additional consideration of sequence-dependent setup times (SDST) is added on both production and assembly stages. The resulting problem is very close to the reality of multi-factory scheduling. The production stage as the first stage is composed of f identical flowshop production factories where each factory is a flowshop that produce the different jobs. In the assembly stage, all produced jobs are assembled into final products. The objective is to minimize the overall makespan. As this problem is known to be NP-hard, we propose two simple constructive heuristics to solve the problem. In this paper we detail the proposed heuristics and carry out a comprehensive computational and statistical experiment to analyze their performance. Sara Hatami, Rubén Ruiz, Carlos Andrés-Romano |
CoDIT | 2 |
| 2012 | Efficient Waste Collection by Means of Assignment Problems
Kostanca Katragjini, Federico Perea, Rubén Ruiz |
ICORES | 3 |
| 2008 | A Review and Evaluation of Multiobjective Algorithms for the Flowshop Scheduling ProblemabstractThis paper contains a complete and updated review of the literature for multiobjective flowshop problems, which are among the most studied environments in the scheduling research area. No previous comprehensive reviews exist in the literature. Papers about lexicographical, goal programming, objective weighting, and Pareto approaches have been reviewed. Exact, heuristic, and metaheuristic methods have been surveyed. Furthermore, a complete computational evaluation is also carried out. A total of 23 different algorithms including both flowshop-specific methods as well as general multiobjective optimization approaches have been tested under three different two-criteria combinations with a comprehensive benchmark. All methods have been studied under recent state-of-the-art quality measures. Parametric and nonparametric statistical testing is profusely employed to support the observed performance of the compared methods. As a result, we have identified the best-performing methods from the literature, which along with the review, constitutes a reference work for further research. Gerardo Minella, Rubén Ruiz, Michele Ciavotta |
INFORMS J. Comput. | 2 |