Aibo Song

dblp:24/3937 · also Ai-Bo Song · DBLP profile ↗
← Back
44ranked-venue papers
5as first author
17since 2021 · last 2025
0000-0003-4447-7305ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 2 since 2021Computer networks · 7 · 4 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Exploring Multimodal Relation Extraction of Hierarchical Tabular Data with Multi-task Learning
abstract
Relation Extraction (RE) is a key task in table understanding, aiming to extract semantic relations between columns.However, complex tables with hierarchical headers are hard to obtain high-quality textual formats (e.g., Markdown) for input under practical scenarios like webpage screenshots and scanned documents, while table images are more accessible and intuitive.Besides, existing works overlook the need of mining relations among multiple columns rather than just the semantic relation between two specific columns in real-world practice.In this work, we explore utilizing Multimodal Large Language Models (MLLMs) to address RE in tables with complex structures.We creatively extend the concept of RE to include calculational relations, enabling multi-task learning of both semantic and calculational RE for mutual reinforcement.Specifically, we reconstruct table images into graph structure based on neighboring nodes to extract graph-level visual features.Such feature enhancement alleviates the insensitivity of MLLMs to the positional information within table images.We then propose a Chain-of-Thought distillation framework with self-correction mechanism to enhance MLLMs' reasoning capabilities without increasing parameter scale.Our method significantly outperforms most baselines on wide datasets.Additionally, we release a benchmark dataset for calculational RE in complex tables.
Aibo Song, Jingyi Qiu, Jiahui Jin 0001, Tianbo Zhang, Xiaolin Fang 0001
ACL (1)2
2025 Graph Representation-Aware Online Aggregations over Knowledge Graph
Jingyi Qiu, Aibo Song, Tongwei Liu, Tianbo Zhang
KSEM (2)2
2025 Enhancing Self-Supervised Fine-Grained Video Object Tracking with Dynamic Memory Prediction
abstract
Successful video analysis relies on accurate recognition of pixels across frames, and frame reconstruction methods based on video correspondence learning are popular due to their efficiency. Existing frame reconstruction methods, while efficient, neglect the value of direct involvement of multiple reference frames for reconstruction and decision-making aspects, especially in complex situations such as occlusion or fast movement. In this paper, we introduce a Dynamic Memory Prediction (DMP) framework that innovatively utilizes multiple reference frames to concisely and directly enhance frame reconstruction. Its core component is a Reference Frame Memory Engine that dynamically selects frames based on object pixel features to improve tracking accuracy. In addition, a Bidirectional Target Prediction Network is built to utilize multiple reference frames to improve the robustness of the model. Through experiments, our algorithm outperforms the state-of-the-art self-supervised techniques on two fine-grained video object tracking tasks: object segmentation and keypoint tracking.
Changrui Dai, Aibo Song, Xiaolin Fang 0001
ICMR3
2024 CQP-RFFI: Injecting a Communication-Quality Preserving RF Fingerprint for Wi-Fi Device Identification
abstract
Recently, there has been an emerging radio frequency fingerprint identification (RFFI) technology that enhances fingerprint distinguishability by deliberately injecting I/Q imbalance into the device’s Wi-Fi baseband signal. Due to the additional injection of I/Q imbalance, this approach inevitably impacts the communication quality between devices, as it reduces the accuracy of channel estimation. To address this issue, we propose injecting the I/Q imbalance into a short training field (STF) instead of the entire baseband signal. Our findings indicate that this method can effectively preserve the quality of the original wireless communication. Building upon this, we introduce a fingerprinting scheme called CQP-RFFI that generates distinguishable fingerprints for a set of devices by injecting appropriate I/Q imbalance into the STF. Leveraging the short-term invariance of the channel, we design a practical I/Q imbalance extraction method based on the communication-quality preserving injection. Moreover, we design an optimal assignment method for I/Q imbalance to maximize the distinguishability of RF fingerprints for all devices. Finally, we implement the CQP-RFFI solution and conduct experiments in real-world scenarios. The experimental results demonstrate that CQP-RFFI achieves 96% precision, recall, and F1-score, and can consistently maintain good communication quality.
Xiaolin Gu, Wenjia Wu, Yusen Zhou, Aibo Song, Ming Yang 0001, Zhen Ling 0001, Junzhou Luo
IWQoS4
2024 TEA-RFFI: Temperature adjusted radio frequency fingerprint-based smartphone identification
Xiaolin Gu, Wenjia Wu, Yusen Zhou, Aibo Song, Ming Yang 0001, Zhen Ling 0001, Junzhou Luo
Comput. Networks4
2024 Relation-oriented few-shot knowledge graph prototype networks
Yingying Xue, Aibo Song, Jiahui Jin 0001, Jingyi Qiu, Xiaolin Fang 0001, Xiaorui Zhai
Neurocomputing2
2024 RF-TESI: Radio Frequency Fingerprint-based Smartphone Identification under Temperature Variation
abstract
Radio frequency fingerprint identification (RFFI) is a promising technique for smartphone identification. However, we find that the temperature of the RF front end in smartphones can significantly impact the RF features, including the carrier frequency offset (CFO) and statistical RF features. The unstable RF features caused by temperature changes can negatively affect the performance of state-of-the-art RFFI approaches. To this end, we propose the RF-TESI solution for smartphone identification under temperature variation. First, we construct a dataset by extracting temperature and RF features. In the dataset, the extracted temperature values constitute a set of temperature values and each registered temperature value corresponds to a group of RF features. Next, we evaluate the distinctiveness of RF features across smartphones to select the most suitable RF fingerprint. Then, we train multiple random forest models, each tagged with a registered temperature. In addition, because there are still many temperatures out of the temperature set, we design an RF fingerprint estimation method to estimate RF fingerprints at unregistered temperatures. Finally, the experiments show RF-TESI demonstrates satisfactory performance under different scenarios, taking into account variations in temperature, time and position. Besides, our proposed approach is better than all state-of-the-art approaches in smartphone identification.
Xiaolin Gu, Wenjia Wu, Aibo Song, Ming Yang 0001, Zhen Ling 0001, Junzhou Luo
ACM Trans. Sens. Networks3
2024 Matching Tabular Data to Knowledge Graph with Effective Core Column Set Discovery
abstract
Matching tabular data to a knowledge graph (KG) is critical for understanding the semantic column types, column relationships, and entities of a table. Existing matching approaches rely heavily on core columns that represent primary subject entities on which other columns in the table depend. However, discovering these core columns before understanding the table’s semantics is challenging. Most prior works use heuristic rules, such as the leftmost column, to discover a single core column, while an insightful discovery of the core column set that accurately captures the dependencies between columns is often overlooked. To address these challenges, we introduce Dependency-aware Core Column Set Discovery ( DaCo ), an iterative method that uses a novel rough matching strategy to identify both inter-column dependencies and the core column set. Additionally, DaCo can be seamlessly integrated with pre-trained language models, as proposed in the optimization module. Unlike other methods, DaCo does not require labeled data or contextual information, making it suitable for real-world scenarios. In addition, it can identify multiple core columns within a table, which is common in real-world tables. We conduct experiments on six datasets, including five datasets with single core columns and one dataset with multiple core columns. Our experimental results show that DaCo outperforms existing core column set detection methods, further improving the effectiveness of table understanding tasks.
Jingyi Qiu, Aibo Song, Jiahui Jin 0001, Jiaoyan Chen 0001, Xiaolin Fang 0001, Tianbo Zhang
ACM Trans. Web2
2023 A Multi-granularity Knowledge-enhanced Sequential Recommendation
abstract
Knowledge graph has been widely utilized to provide knowledge for the sequential recommendation. But their divergence produces new challenges to appropriately capture and represent items’ knowledge for recommendation. Additionally, the complex temporal features also increase the difficulty to model the user-item interaction sequences. The key issues remaining to be settled are how to mine adequate information from KG to capture items’ knowledge, and how to model the evolving user interests accurately. Facing the above problems, we propose a multi-task learning framework MLKT to learn the embedding of items and users’ dynamic preferences in both a knowledge-aware and time-aware way. Different from the previous works, we not only consider the timestamps of the interactions in the sequence, but also capture its evolution from multiple perspectives by designing a multi-scale temporal attention mechanism. Additionally, we represent each item’s features in different granularity by learning from neighbors and the category information in KG. Furthermore, to make the KG embeddings to be more suitable for the recommendation, a cross-component is designed to efficiently achieve the item information communication between KG and user-item interactions. Extensive experiments demonstrate the significant improvements of MLKT than the baseline methods on two recommendation tasks in three datasets.
Yingying Xue, Aibo Song, Jibin Sun
CSCWD4
2023 Dependency-Aware Core Column Discovery for Table Understanding
Jingyi Qiu, Aibo Song, Jiahui Jin 0001, Tianbo Zhang, Jingyi Ding, Xiaolin Fang 0001, Jianguo Qian
ISWC2
2023 Intra- and inter-semantic with multi-scale evolving patterns for dynamic graph learning
Yingying Xue, Aibo Song, Xiaolin Fang 0001, Jiahui Jin 0001, Xiangguo Sun, Yingxue Zhang 0008
Knowl. Based Syst.2
2022 TeRFF: Temperature-aware Radio Frequency Fingerprinting for Smartphones
abstract
In recent years, radio frequency (RF) fingerprinting has attracted more and more attention. Many different types of RF fingerprints have been proposed, such as carrier frequency offset (CFO), sampling frequency offset and error vector magnitude. Among them, the CFO fingerprint is recognized as a promising RF fingerprint. However, for commonly used smartphones, we find that its CFO fingerprint is unstable, because the temperature of crystal oscillator varies greatly and large fluctuations of temperature significantly affect its CFO fingerprint. Therefore, the solutions of CFO-based fingerprinting will no longer be effective for smartphones if the temperature of crystal oscillator is not involved. To this end, we propose a more reliable and applicable CFO-based fingerprinting approach called temperature-aware radio frequency fingerprinting (TeRFF). First, we construct a dataset by extracting crystal oscillator's temperature and the corresponding CFO value on multiple smartphones over a period. In the dataset, the extracted temperature values constitute a set of temperature values, and each registered temperature value corresponds to a group of CFO samples. On this basis, we train multiple Naive Bayes models, each tagged with a registered temperature value. Moreover, since there are many temperature values which are not in the temperature set, we design a CFO estimation method to estimate the CFO fingerprint at the unregistered temperature. Finally, the experimental results demonstrate that our proposed solution TeRFF makes the CFO fingerprinting still effective for smartphone identification, and its performance is better than other existing RF fingerprinting schemes.
Xiaolin Gu, Wenjia Wu, Naixuan Guo, Aibo Song, Ming Yang 0001, Zhen Ling 0001, Junzhou Luo
SECON5
2022 Identifying critical nodes in power networks: A group-driven framework
Aibo Song, Yingying Xue, Jiahui Jin 0001
Expert Syst. Appl.2
2022 FedAda: Fast-convergent adaptive federated learning in heterogeneous mobile edge computing environment
Jinghui Zhang 0001, Jiahui Jin 0001, Aibo Song, Wei Zhao 0023, Liangsheng Wen
World Wide Web7
2021 802.11ac Device Identification based on MAC Frame Analysis
abstract
In Wi-Fi networks, devices can be identified by physical features or MAC layer features, and the solutions of device identification can be used to enhance device authentication. Since 802.11ac Standard has been widely applied in Wi-Fi devices in recent years, the traditional identification methods designed for 802.11b/g/n devices will be no longer applicable. Therefore, it is necessary to design the corresponding 802.11ac device identification method. Compared with the physical feature-based method, the MAC layer-based method has advantages of low cost and easy deployment, so it has attracted more and more researchers' attention. In this paper, we use the fields from 802.11ac MAC frame as fingerprints. Through the analysis of 802.11ac MAC frame, a preprocessing method of the frame is proposed to mask strong and easy-to-modified identifiers. Then to overcome the difficulties caused by random changes in field values, we propose a device identification method based on the deep learning to select features automatically. Compared with the previous one using the transmitting rate as a feature, our method does not spend much time capturing packets in the device identification stage and has better performance whose average precision and recall exceed 99%.
Xiaolin Gu, Wenjia Wu, Zhouguo Chen, Aibo Song, Zhen Ling 0001, Ming Yang 0001
CSCWD4
2021 BSDP: A Novel Balanced Spark Data Partitioner
abstract
As a memory-based distributed big data computing framework, Spark has been widely used in big data processing systems. However, during the execution of Spark, due to the imbalance of input data distribution and the shortage of existing data partitioners in Spark, it is easy to cause partition skew problem and reduce the execution efficiency of Spark. Aiming at this problem, this paper proposes a balanced Spark data partitioner called BSDP (Balanced Spark Data Partitioner). By deeply analyzing the partitioning characteristics of Shuffle intermediate data, the Spark Shuffle intermediate data equalization partitioning model is established. The model aims to minimize the partition skew and find a Shuffle intermediate data equalization partitioning strategy. Based on the model, this paper designs and implements a data equalization partitioning algorithm of BSDP. This algorithm transforms the Shuffle intermediate data equalization partitioning problem into a classic List-Scheduling task scheduling problem, effectively realizes the balanced partitioning of Shuffle intermediate data. The experiment verifies that the BSDP can effectively realize the balanced partitioning of the Shuffle intermediate data and improve the execution efficiency of Spark.
Aibo Song, Bowen Peng, Jingyi Qiu, Yingying Xue, Mingyang Du
ICPADS1
2021 Relation-based multi-type aware knowledge graph embedding
Yingying Xue, Jiahui Jin 0001, Aibo Song, Yingxue Zhang 0008
Neurocomputing3
2020 A Topology-Adaptive Deep Model for Power System Multifault Diagnosis
abstract
Quickly identifying faulty sections is tremendously important for power systems, yet challenging due to handling the variations of complex alarm patterns. Existing works have focused on finding fault section clues solely from alarm information (and ignoring power system topology information). So they are only sensitive to alarms from power systems with pre-assumed topology structures, and encounter difficulties when a system's topology changes. To adapt to unknown or varying system topologies, here we present a Topology-Adaptive Deep Model (TADM) for power system multifault diagnosis. TADM mines the underlying mapping from alarm and topology information to each section's fault status. It consists of a deep iterative network (DIN), a one-layer fully connected network (FCN), and section-wise multifault diagnosis (SWMD) subnetwork. TADM first models a fault power system as a graph, from which DIN iteratively integrates the alarm and topology information in the region from each node to its T -hop neighbors, and learns their local correlation. Limited to T 's size, FCN then combines all local correlations to determine the global correlation between alarm and topology information across the entire power system. To implement multifault diagnosis, learned local and global correlations serve as topology-related fault representations for input as an SWMD (to predict all sections' fault states one by one). A comprehensive experimental study demonstrates that TADM outperforms state-of-the-art models in both multifault diagnosis and adapting to system topologies. The source code of the TADM is available onlline1.
Aibo Song, Jiahui Jin 0001, Mingyu Zhai, Yingying Xue
SMC2
2020 Intelligent online catastrophe assessment and preventive control via a stacked denoising autoencoder
Mingyu Zhai, Jiahui Jin 0001, Aibo Song, Jikeng Lin, Zhiang Wu 0001, Yixin Zhao
Neurocomputing4
2020 Efficiently Translating Complex SQL Query to MapReduce Jobflow on Cloud
abstract
MapReduce is a widely-used programming model in cloud environment for parallel processing large-scale data sets. The combination of the high-level language with a SQL-to-MapReduce translator allows programmers to code using SQL-like declarative language, so that each program can afterwards be complied into a MapReduce jobflow automatically. This way is helpful to narrow the gap between non-professional users and cloud platforms, and thus significantly improve the usability of the cloud. Although a number of translators have been developed, the auto-generated MapReduce programs still suffered from extremely inefficiency. In this paper, we present an efficient Cost-Aware SQL-to-MapReduce Translator (CAT). CAT has two notable features. First, it defines two intra-SQL correlations: Generalized Job Flow Correlation (GJFC) and Input Correlation (IC), based on which a set of looser merging rules are introduced. Thus, both Top-Down (TD) and Bottom-Up (BU) merging strategies are proposed and integrated into CAT simultaneously. Second, it adopts a cost estimation model for MapReduce jobflows to guide the selection of a more efficient MapReduce jobflows auto-generated by TD and BU merging strategies. Finally, comparative experiments on TPC-H benchmark demonstrate the effectiveness and scalability of CAT.
Zhiang Wu 0001, Aibo Song, Jie Cao 0001, Junzhou Luo, Lu Zhang 0030
IEEE Trans. Cloud Comput.2
2019 Query optimization Approach with Shuffle Intermediate Cache Layer for Spark SQL
abstract
Spark SQL is a big data processing tool for structured data query and analysis. However, due to the execution of Spark SQL, there are multiple times to write intermediate data to the disk, which reduces the execution efficiency of Spark SQL. Targeting on the existing issues, we design and implement an intermediate data cache layer between the underlying file system and the upper Spark core to reduce the cost of random disk I/O. By using the query pre-analysis module, we can dynamically adjust the capacity of cache layer for different queries. And the allocation module can allocate proper memory for each node in cluster. This paper develops the SSO (Spark SQL optimizer) module and integrates it into the original Spark system to achieve the above functions. This paper compares the query performance with the existing Spark SQL by experiment data generated by TPC-H tool. The experimental results show that the SSO module can effectively improve the query efficiency, reduce the disk I/O cost and make full use of the cluster memory resources.
Mingyu Zhai, Aibo Song, Jingyi Qiu, Xuechun Ji, Qingxi Wu
IPCCC2
2019 Graph partition-based data and task co-scheduling of scientific workflow in geo-distributed datacenters
abstract
Summary Most large‐scale scientific workflows take place in multiple collaborative datacenters for access to community‐wide resources, while adhering to each datacenter's non‐uniform resource limits. However, moving both initial input datasets with predetermined locations and intermediate datasets needing placement decisions across geo‐distributed datacenters hinders efficient execution of large‐scale data‐intensive scientific workflows. Thus, scientific workflow's data and task co‐scheduling deal with situations such as pre‐placed initial input datasets, placement of intermediate datasets and each datacenter's non‐uniform computation and storage constraint, while minimizing the cross‐datacenter data transfer. Since this scheduling problem is known to be NP‐hard, here, we propose a novel approach, based on the multilevel graph coarsening and uncoarsening framework, together with a specialized hybrid genetic algorithm having distinctive graph partition driven features of repair and local improvement, for scheduling data‐intensive scientific workflows in geo‐distributed datacenters and optimizing the cross‐datacenter data transfer volume. Extensive simulations, based on four real‐world workflow traces, show that our algorithm significantly reduces the overall geo‐distributed data transfer and demonstrate its effectiveness.
Jinghui Zhang 0001, Jiahui Jin 0001, Aibo Song
Concurr. Comput. Pract. Exp.5
2019 A local random walk model for complex networks based on discriminative feature combinations
Aibo Song, Zhiang Wu 0001, Mingyu Zhai, Junzhou Luo
Expert Syst. Appl.1
2019 A data-locality-aware task scheduler for distributed social graph queries
Jiahui Jin 0001, Junzhou Luo, Mingyang Du, Yongcheng Dang, Jinghui Zhang 0001, Aibo Song
Future Gener. Comput. Syst.7
2018 Query Optimization Approach with Middle Storage Layer for Spark SQL
abstract
Currently, Spark SQL cannot optimize the multi-query tasks: tasks provided by batch processing are translated into different Spark jobs, and these jobs cannot share input data. To solve this problem, this paper explores the optimization of translating an SQL query into Spark jobs via two strategies: 1) A Middle Storage Layer is added between the persistent file system and the Spark core to address the data sharing problem among multiple queries. 2) When loading data into the Middle Storage Layer, the use of a selection strategy based on the cost model, in which the input data are distributed to multiple proper nodes for processing to achieve efficient utilization of cluster resources, is proposed. Based on this exploration, we develop a system called QOMS(Query Optimization approach with Middle Storage layer for Spark SQL), which offers better performance in improving query speed compared to Spark SQL with respect to the TPC-H benchmark.
Aibo Song, Mingyu Zhai, Yingying Xue, Mingyang Du, Yutong Wan
CSCWD1
2016 CAT: A Cost-Aware Translator for SQL-query workflow to MapReduce jobflow
Aibo Song, Zhiang Wu 0001, Junzhou Luo
Data Knowl. Eng.1
2014 A Sampling-Based Hybrid Approximate Query Processing System in the Cloud
abstract
Sampling-based approximate query processing method provides the way, in which the users can save their time and resources for 'Big Data' analytical applications, if the estimated results can satisfy the accuracy expectation earlier before a long wait for the final accurate results. Online aggregation (OLA) is such an attractive technology to respond aggregation queries by calculating approximate results with the confidence interval getting tighter over time. It has been built into the MapReuduce-based cloud system for big data analytics, which allows users to monitor the query progress and save money by killing the computation earlier once sufficient accuracy has been obtained. Unfortunately, there exists a major obstacle that is the estimation failure of OLA affects the OLA performance, which is resulted from the biased sample set that violates the unbiased assumption of OLA sampling. To handle this problem, we first propose a hybrid approximate query processing model to improve the overall OLA performance, where a dynamic scheme switching mechanism is deliberately designed to switch unpromising OLA queries into the bootstrap scheme for further processing, avoiding the whole dataset scanning resulted from the OLA estimation failure. In addition, we also present a progressive estimation method to reduce the false positive ratio of our dynamic scheme switching mechanism. Moreover, we have implemented our hybrid approximate query processing system in Hadoop, and conducted extensive experiments on the TPC-H benchmark for skewed data distribution. Our results demonstrate that our hybrid system can produce acceptable approximate results within a time period one order of magnitude shorter compared to the original OLA over Hadoop.
Yuxiang Wang 0001, Junzhou Luo, Aibo Song, Fang Dong 0001
ICPP3
2014 OATS: online aggregation with two-level sharing strategy in cloud
Yuxiang Wang 0001, Junzhou Luo, Aibo Song, Fang Dong 0001
Distributed Parallel Databases3
2013 Partition-Based Online Aggregation with Shared Sampling in the Cloud
Yuxiang Wang 0001, Junzhou Luo, Aibo Song, Fang Dong 0001
J. Comput. Sci. Technol.3
2012 Improving Online Aggregation Performance for Skewed Data Distribution
Yuxiang Wang 0001, Junzhou Luo, Aibo Song, Jiahui Jin 0001, Fang Dong 0001
DASFAA (1)3
2012 An effective data aggregation based adaptive long term CPU load prediction mechanism on computational grid
Fang Dong 0001, Junzhou Luo, Aibo Song, Jiuxin Cao, Jun Shen 0001
Future Gener. Comput. Syst.3
2011 BAR: An Efficient Data Locality Driven Task Scheduling Algorithm for Cloud Computing
abstract
Large scale data processing is increasingly common in cloud computing systems like MapReduce, Hadoop, and Dryad in recent years. In these systems, files are split into many small blocks and all blocks are replicated over several servers. To process files efficiently, each job is divided into many tasks and each task is allocated to a server to deals with a file block. Because network bandwidth is a scarce resource in these systems, enhancing task data locality(placing tasks on servers that contain their input blocks) is crucial for the job completion time. Although there have been many approaches on improving data locality, most of them either are greedy and ignore global optimization, or suffer from high computation complexity. To address these problems, we propose a heuristic task scheduling algorithm called Balance-Reduce(BAR), in which an initial task allocation will be produced at first, then the job completion time can be reduced gradually by tuning the initial task allocation. By taking a global view, BAR can adjust data locality dynamically according to network state and cluster workload. The simulation results show that BAR is able to deal with large problem instances in a few seconds and outperforms previous related algorithms in term of the job completion time.
Jiahui Jin 0001, Junzhou Luo, Aibo Song, Fang Dong 0001, Runqun Xiong
CCGRID3
2011 Load-aware based adaptive rescheduling mechanism for workflow application
abstract
In order to integrate the massive distributed resources to accomplish the complex engineering applications cooperatively, workflow scheduling is an important aspect. However, as the available computing power of Grid resources is changing dynamically, static scheduling scheme will lead to low performance in real Grid environment. Therefore, the rescheduling mechanism should be taken into consideration. Although a few relevant mechanisms have been proposed in recent year, as they do not consider the essence of dynamic feature and the relevant algorithms are too simple, they can not obtain a good enough result yet. To address these problems, a load-aware based adaptive rescheduling mechanism for DAG application called LAR is proposed. Therein, during application running, the load exception will be detected and the execution state of application will be judged to decide whether the rescheduling process should be triggered. And in rescheduling stage, an effective rescheduling algorithm which utilizes the latest prediction information is present. The simulation results show that our mechanism can outperform the relevant algorithms in NRSL, and can effectively solve the performance decreasing problem in real Grid environment.
Fang Dong 0001, Junzhou Luo, Aibo Song, Jiuxin Cao
CSCWD3
2011 QoS Preference-Aware Replica Selection Strategy Using MapReduce-Based PGA in Data Grids
abstract
Data replication is an important technique to reduce access latency and bandwidth consumption in Grid environment. As one of the major functions of data replication, replica selection determines the best replica according to some specific criteria in Data Grid environment, where the data resources are limited and Grid users compete for these resources. In this paper, we focus mainly on a novel QoS preference-aware replica selection strategy which will meet individual QoS sensitivity (IQS) constraints for different users/applications. We first present a framework that characterize QoS properties of replica services and establish its mathematical model by introducing quantification methods. In order to deal with the IQS constraints and to perceive Grid users' QoS preferences accurately, we propose a QoS preference acquisition algorithm based on Analytic Hierarchy Process (AHP). We then design and implement a novel effective and efficient parallel genetic algorithm (PGA) based on Map Reduce paradigm for optimizing the objective function which corresponds to the optimal replica. Simulation results show that our strategy has a better performance in validity as well as scalability, and the optimal replica can always be obtained for Grid users with different IQS constraints under Data Grid environments that vary in system loads, scheduling strategies and user types.
Runqun Xiong, Junzhou Luo, Aibo Song, Bo Liu 0004, Fang Dong 0001
ICPP3
2010 Data Aggregation based Adaptive Long term load Prediction mechanism in Grid environment
abstract
In recent years, as a popular technique to support CSCW, Grid computing is becoming more and more attractive. Hereinto, as the CPU load information can guide task scheduling process greatly, the long-term CPU load prediction becomes a very hot research field and has been widely studied. However, as the prediction errors will be accumulated gradually and meanwhile the relevant parameters' optimal values may change dynamically with the variance of load series, the previous prediction algorithms usually can not obtain good prediction accuracy when the length of prediction interval is quite large. To address these feature, a Data Aggregation based Adaptive Long term load Prediction mechanism called DA2LP is proposed in this paper. Therein, in order to reduce the number of prediction step and increase the amount of useful input load information, the data aggregation concept is introduced to integrate with AR model. Meanwhile, with the observation and analysis of the relevant parameters' impact on prediction accuracy in our prediction model, an adaptive parameter selection mechanism is proposed, where the optimal relevant parameters can be adapted automatically to enhance prediction accuracy during the prediction process. The experiments show that our proposed mechanism can outperform significantly the previous prediction methods in mean square error (MSE) for long term load prediction.
Fang Dong 0001, Junzhou Luo, Aibo Song, Jiuxin Cao
CSCWD4
2010 SLA-Based Resource Co-Allocation in Multi-Cluster Grid
abstract
Resource co-allocation is a crucial but challenging problem for Grid Computing. With the emergence of WS-Resource Framework and Open Grid Services Architecture, resource co-allocation is commonly associated with a service level agreement (SLA) to determine the achieved QoS level. In this paper, we present an approach for resource co-allocation in Multi-cluster Grid that maximizes the user satisfaction degree while satisfying the QoS requirements defined in a SLA. The QoS metrics considered in this paper include deadline, cluster availability, service reliability and budget. Experimental results are presented to show the effectiveness of our approach.
Wei Wang 0089, Junzhou Luo, Aibo Song, Fang Dong 0001
GLOBECOM3
2010 Resource Load Based Stochastic DAGs Scheduling Mechanism for Grid Environment
abstract
The dynamic feature is one of the most important differences between Grid and traditional heterogeneous distributed systems, thus the most significant challenge for task scheduling in Grid environment is how to relieve the resource performance dynamism effectively. However, the existing schedule algorithms usually suppose that computation or communication times are deterministic and static, thus they will lead to bad performance in the practical Grid environment. To address this problem, a mechanism which is used to estimate the probability distribution of task execution time based on resource load is proposed. And then a Resource Load based Stochastic DAGs Scheduling algorithm for Grid environments is introduced. The simulation results show that our mechanism can achieve a significant improvement in several metrics (such as normalized real schedule length) and can relieve the influence brought by the dynamic nature of Grid effectively.
Fang Dong 0001, Junzhou Luo, Aibo Song, Jiahui Jin 0001
HPCC3
2010 A novel task scheduling algorithm based on dynamic critical path and effective duplication for pervasive computing environment
abstract
Abstract In order to effectively utilize massive heterogeneous resources and provide transparent computing capability to upper applications, task scheduling as the key issue of pervasive computing system becomes significantly important. Previous proposed priority and duplication based task scheduling algorithms, which can be applied in pervasive computing environment, usually have following limitations: critical path cannot be calculated accurately while neglecting the effect of resource availability in scheduling; in duplication based resource allocation stage, duplications without restriction would lead to some negative effects on final schedule length (SL). For the purpose of solving these problems, a novel task scheduling algorithm based on dynamic critical path (DCP) and effective duplication, called DCPED, is presented in this paper. In DCPED, a more accurate DCP calculation method which takes resource availability into account is introduced. Meanwhile an effective task duplication strategy is proposed to eliminate ineffective duplications and make an optimized schedule result by using space compression technique and dynamic critical path length (DCPL) based evaluation technique respectively. Finally, simulation results show that DCPED can outperform previous algorithms significantly in NSL and speedup rate metrics. Especially, it is very effective for utilizing computing resources and scheduling the fine‐grain and large‐scale workflow applications in pervasive computing system. Copyright © 2008 John Wiley & Sons, Ltd.
Junzhou Luo, Fang Dong 0001, Jiuxin Cao, Aibo Song
Wirel. Commun. Mob. Comput.4
2008 An improved DHT-based Grid Information services architecture
abstract
Fundamental problems that confront traditional centralized Grid information services are low query efficiency and single points of failure. This paper presents an improved DHT-based Grid information services architecture, that resolves these problems well. The architecture assigns each resource node and information server an m-bit identifier by hashing attribute values. Every node makes connection with successor information server according to their identifiers to organize Virtual Organization(VO). Resource query is highly efficient with locating identifier. Information servers are fully decentralized, solving single points of failure. Range query is a difficult problem in DHT but an indispensable part in Grid information services. We improve our architecture with tree data structure and Path Caching Schemes to support range query. Results from theory analysis and simulations show that DHT-based architecture is highly efficient, robust, load balanced and scalable.
Junzhou Luo, Aibo Song
CSCWD3
2008 A Scalable and Adaptive Distributed Service Discovery Mechanism in SOC Environments
Junzhou Luo, Aibo Song
NPC3
2008 Grid Service Discovery Based on Cross-VO Service Domain Model
Jingya Zhou, Junzhou Luo, Aibo Song
NPC3
2007 A Trust Degree Based Access Control for Multi-domains in Grid Environment
abstract
The grid security focuses on implementation of safe access for resource of different domains in dynamic grid environment. Trust as an important factor in grid security is increasingly applied to management of security. But the research of application with trust to access control is rare and coarse. In this paper, we propose the concept of trust degree which is the measurement of trust and combine it with access control framework. A fine-granularity access control model has been realized in a single domain and the trust degree based access control framework accomplishes the work to access resource across multi-domains. Conversion between domains is correct and effective. Simulation results present that the access control model is practicable and credible.
Xudong Ni, Junzhou Luo, Aibo Song
CSCWD3
2005 A semantic access control model for grid services
abstract
Grid computing is appropriate for supporting cooperative work Many designers and engineers from different companies or institutions can dynamically form a virtual organization for a given design task. In order to protect each company's sensitive data and services, access control is therefore necessary and important. In this paper, we present a new approach to authorize and administrate access requests, in which the requests are obliged to negotiate with a policy enforcement point in order to gain access to the target grid service. The new access control model will exploit semantic Web technology, and use machine reasoning about the messages and policies at a semantic level.
Junzhou Luo, Aibo Song
CSCWD (1)3
2004 Discovering User Profiles for Web Personalized Recommendation
Aibo Song, Mao-Xian Zhao, Zuo-Peng Liang
J. Comput. Sci. Technol.1