Liudong Zuo

dblp:137/0591 · DBLP profile ↗
← Back
3ranked-venue papers in the field
2as first author
2since 2021 · last 2024
0000-0001-9360-1064ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3 (2 first)
YearPublicationVenuePosition
2024 ℓ1, 2-Norm and CUR Decomposition based Sparse Online Active Learning for Data Streams with Streaming Features
abstract
Aiming at learning from a sequence of data instances over time, online learning has attracted increasing attention in the big data era. As two important variants, sparse online learning has been extensively explored by facilitating sparse constraints for online models such as truncated gradient, ℓ1-norm regularization, ℓ1-ball projection, and regularized dual averaging; while online active learning aims to build an online prediction model with a limited number of labeled instances, deploying the so called query strategies to select informative instances over time. However, most existing studies consider sparse online learning or online active learning with fixed feature spaces, whereby in real practice the features may be dynamically evolved over time. To the end, we propose a novel unified one-pass online learning framework named OASF for simultaneously online active learning and sparse online learning tailored for data streams described by open feature spaces, where new features can emerge constantly, and old features may be vanished over various time spans. Specifically, we technically develop an effective online CUR matrix decomposition based on the ℓ1,2mixed norm constraint for simultaneously selecting important up-to-date samples in a sliding window and facilitating stable and meaningful features in open feature spaces over time. If the loss function is simultaneously Lipschitz and convex, a sub-linear regret bound of our proposed algorithm is guaranteed with. Extensive experiments that are conducted with multiple streaming datasets have demonstrated the effectiveness of the proposed OASF compared with state-of-the-art online active learning and sparse online learning methods.
Zhong Chen 0003, Yi He 0007, Di Wu 0056, Liudong Zuo, Keren Li, Wenbin Zhang 0002, Zhiqiang Deng
IEEE Big Data4
2024 Optimal Scheduling Algorithms for Cost-Effective Bandwidth Reservation in HPNs
abstract
Vast amounts of data are continually being produced in various scientific fields. Once these large datasets are generated, they often require rapid transfer over long distances using bandwidth reservation services provided by high-performance networks (HPNs) dedicated to collaborative data storage and analysis. The primary goal of data transfer is typically to achieve the earliest completion time (ECT), but users may also seek to minimize financial costs associated with the transfer. Balancing these differing requirements can be challenging. In this paper, we explore the trade-off between ECT and cost in data transfers that use bandwidth reservation on variable paths with fixed bandwidth within dedicated HPNs. Our investigation focuses on two types of bandwidth reservation requests (BRRs) and their scheduling: (i) minimizing data transfer cost while meeting a data transfer deadline, and (ii) achieving ECT while meeting a specified maximum cost. To optimize the scheduling of both types of BRRs, we propose two novel algorithms and conduct extensive simulations to demonstrate their effectiveness and efficiency.
Liudong Zuo, Pan Lai, Zhong Chen 0003
IEEE Big Data1
2016 Bandwidth provision strategies for reliable data movements in dedicated networks
abstract
Sheer volumes of data are being generated in extreme-scale distributed scientific applications, and need to be transferred remotely in fast, predictable and reliable way for data storage and analysis purpose. Reserving bandwidth along selected paths in high-performance networks (HPNs) has proved to be an effective way to satisfy the high-demanding performance requirements of such data transfer. However, node and link failures within the HPNs potentially degrade the quality of data transfer. In this paper, we focus on the scheduling of two generic types of bandwidth reservation requests concerning data transfer reliability: (i) to achieve the highest data transfer reliability under a given data transfer deadline, and (ii) to achieve the earliest data transfer completion time while satisfying a given data transfer reliability requirement. Poisson distribution is used to model the node and failures within the HPNs, and two periodic bandwidth reservation algorithms with rigorous optimality proofs are proposed.
Liudong Zuo, Mengxia Zhu
IEEE BigData1