Yubo Zhuang

dblp:312/6900 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0009-0004-8145-2523ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Optimization for machine learning · 62% Probabilistic and Bayesian machine learning · 38%
Databases, data mining, and information retrieval
2 papers
Data mining · 100%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
1.322024
Statistically Optimal K-means Clustering via Nonnegative Low-rank Semidefinite Programming · ICLR 2024
Wasserstein $K$-means for clustering probability distributions · NeurIPS 2022
Data mining › clustering
k-means clustering
1.322024
Statistically Optimal K-means Clustering via Nonnegative Low-rank Semidefinite Programming · ICLR 2024
Wasserstein $K$-means for clustering probability distributions · NeurIPS 2022
Machine learning › Optimization for machine learning › non-convex optimization
burer-monteiro factorization
0.812024
Statistically Optimal K-means Clustering via Nonnegative Low-rank Semidefinite Programming · ICLR 2024
Machine learning › Optimization for machine learning
non-convex optimization
0.812024
Statistically Optimal K-means Clustering via Nonnegative Low-rank Semidefinite Programming · ICLR 2024
Mathematical optimization
semidefinite programming
0.812024
Statistically Optimal K-means Clustering via Nonnegative Low-rank Semidefinite Programming · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning › clustering
model-based clustering
0.712023
Likelihood Adjusted Semidefinite Programs for Clustering Heterogeneous Data · ICML 2023
Machine learning › Optimization for machine learning › convex relaxation
semidefinite programming relaxation
0.712023
Likelihood Adjusted Semidefinite Programs for Clustering Heterogeneous Data · ICML 2023
Mathematical optimization › convex relaxation
semidefinite relaxation
0.612022
Wasserstein $K$-means for clustering probability distributions · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

semidefinite programming relaxation · 2.3burer-monteiro factorization · 2.3semidefinite programming · 1.8non-negative matrix factorization · 1.5wasserstein barycenter · 1.1optimal transport · 1.1nonnegative matrix factorization · 0.8likelihood-based inference · 0.7expectation-maximization · 0.7
YearPublicationVenuePosition
2025 A Fast Consensus Algorithm of Large-Scale Heterogeneous Dynamic IoT Nodes for DAG-Based Blockchain
abstract
To solve issues of frequent network topology changes to slow consensus speed, and low security in consensus algorithms within heterogeneous dynamic IoT systems, we propose a fast consensus algorithm of large-scale heterogeneous dynamic IoT nodes for DAG-based blockchain, LSHD_DAG. Firstly, an adaptive regional division method for heterogeneous dynamic nodes is proposed, which monitors nodes’ state and counts within each region, adjusting the boundaries of adjacent regions. Secondly, an event-triggered master-slave DAG chain design is proposed. When transaction information pertains to a single region, the transaction is processed as a local transaction to construct a regional DAG slave chain. If transaction information involves multiple regions, it is organized into a multi-regional transaction set and incorporated into the global DAG master chain. Next, a node grade identification mechanism that balances state and reputation factors is proposed. This mechanism adopts the LOF algorithm to classify or dynamically update node grades by considering node state and reputation evaluation results, assigning them identity permissions. Finally, a weighted voting consensus based on the master-slave DAG chain is introduced. Nearby honest nodes with stable movement are selected as consensus nodes, which perform multi-regional weighted voting consensus on transactions and transaction sets. This process concludes with Two-Way Pegging with multi-signature on the master-slave DAG, ensuring secure and rapid consensus for data in heterogeneous and dynamic IoT environments. The experimental results show that no matter how the number of nodes changes, LSHD_DAG can improve transaction throughput, and reduce latency and communication overhead, outperforming the CDBFT, DAG-D, and Avalanche.
Yourong Chen, Yubo Zhuang, Yidan Guo, Zhen Hong
IEEE Internet Things J.2
2024 Statistically Optimal K-means Clustering via Nonnegative Low-rank Semidefinite Programming
abstract
$K$-means clustering is a widely used machine learning method for identifying patterns in large datasets. Recently, semidefinite programming (SDP) relaxations have been proposed for solving the $K$-means optimization problem, which enjoy strong statistical optimality guarantees. However, the prohibitive cost of implementing an SDP solver renders these guarantees inaccessible to practical datasets. In contrast, nonnegative matrix factorization (NMF) is a simple clustering algorithm widely used by machine learning practitioners, but it lacks a solid statistical underpinning and theoretical guarantees. In this paper, we consider an NMF-like algorithm that solves a nonnegative low-rank restriction of the SDP-relaxed $K$-means formulation using a nonconvex Burer--Monteiro factorization approach. The resulting algorithm is as simple and scalable as state-of-the-art NMF algorithms while also enjoying the same strong statistical optimality guarantees as the SDP. In our experiments, we observe that our algorithm achieves significantly smaller mis-clustering errors compared to the existing state-of-the-art while maintaining scalability.
Yubo Zhuang, Richard Y. Zhang 0001
ICLR1
2024 A Large-Scale Node Lightweight Consensus Algorithm of Blockchain for Internet of Things
abstract
In order to solve the problems in the blockchain consensus algorithm, such as slow consensus speed, high resource consumption, and low consensus security due to excessive large-scale node data volume and Byzantine attacks in the Internet of Things (IoT) system, we propose a large-scale node lightweight consensus algorithm of blockchain for IoT (LNLCA). First, a transaction set construction mechanism is proposed to improve consensus efficiency and security. The mechanism packages multiple data monitored by IoT nodes into data transactions, and builds the transaction set. Then, it constructs nonconflict subsets and conflict subsets. Second, an adjacent parent node sampling and response mechanism is proposed to reduce the consumption of communication resources. The new transaction set selects nearby parent nodes based on the time distance summation method and collaborates to batch verify each transaction in the transaction set. Finally, an efficient consistent consensus based on transaction set directed acyclic graph (DAG) is proposed to quickly vote for each transaction in the transaction set, thereby constructing the transaction set DAG for batch uploading to the chain, and achieving a secure and lightweight consensus on IoT data. The experimental results show that no matter how the number of Byzantine nodes changes, LNLCA can improve transaction throughput, and reduce transaction delay and communication overhead, which outperforms the credit-delegated Byzantine fault tolerance, Avalanche, and Hashgraph.
Yubo Zhuang, Yourong Chen, Xudong Zhang 0003, Tiaojuan Ren, Muhammad Alam 0002, Zhen Hong
IEEE Internet Things J.1
2024 Efficient and Secure Blockchain Consensus Algorithm for Heterogeneous Industrial Internet of Things Nodes Based on Double-DAG
abstract
With the Industrial Internet of Things (IIoT) continuing to expand, lots of data collection, exchange, and authentication generated from an increasing number of access devices is required with heterogeneity, multidimension, and multiobjective networks as its characteristics. However, traditional IIoT systems are vulnerable to security challenges, such as data leakage, theft, and tampering. As one of the most promising solutions, blockchain has played an essential role in ensuring security and transparency in the IIoT. But there are still some challenges that prevent the secure and effective implementation of blockchain-based IIoT systems in consensus security, consensus efficiency, and consensus application. To address these problems, we propose an effective security blockchain consensus algorithm for heterogeneous IIoT nodes aiming to defend against the consensus attack and improve consensus efficiency. First, we design a blockchain-based IIoT system architecture. Then, we present an identity authentication and transformation protocol to defend against consensus attacks. Furthermore, we introduce a method for constructing communication directed acyclic graphs (DAGs) and transaction set DAGs to enhance transaction throughput. Based on these two DAGs, we propose an efficient and security consensus algorithm (DAG-D). DAG-D employs transaction sets instead of single transactions or blocks, leveraging communication DAG propagation to swiftly confirm transaction set DAGs based on parent transactions for associated confirmation. Experimental results show that our proposed DAG-D outperforms DAG-M, DAG-Avalanche, and DAG-CoDAG, regarding transaction throughput, transaction latency, and communication overhead.
Yourong Chen, Yubo Zhuang, Kelei Miao, Seyed Amin Pouriyeh
IEEE Trans. Ind. Informatics3
2023 Likelihood Adjusted Semidefinite Programs for Clustering Heterogeneous Data
abstract
Clustering is a widely deployed unsupervised learning tool. Model-based clustering is a flexible framework to tackle data heterogeneity when the clusters have different shapes. Likelihood-based inference for mixture distributions often involves non-convex and high-dimensional objective functions, imposing difficult computational and statistical challenges. The classic expectation-maximization (EM) algorithm is a computationally thrifty iterative method that maximizes a surrogate function minorizing the log-likelihood of observed data in each iteration, which however suffers from bad local maxima even in the special case of the standard Gaussian mixture model with common isotropic covariance matrices. On the other hand, recent studies reveal that the unique global solution of a semidefinite programming (SDP) relaxed $K$-means achieves the information-theoretically sharp threshold for perfectly recovering the cluster labels under the standard Gaussian mixture model. In this paper, we extend the SDP approach to a general setting by integrating cluster labels as model parameters and propose an iterative likelihood adjusted SDP (iLA-SDP) method that directly maximizes the exact observed likelihood in the presence of data heterogeneity. By lifting the cluster assignment to group-specific membership matrices, iLA-SDP avoids centroids estimation -- a key feature that allows exact recovery under well-separateness of centroids without being trapped by their adversarial configurations. Thus iLA-SDP is less sensitive than EM to initialization and more stable on high-dimensional data. Our numeric experiments demonstrate that iLA-SDP can achieve lower mis-clustering errors over several widely used clustering methods including $K$-means, SDP and EM algorithms.
Yubo Zhuang
ICML1
2022 Sketch-and-lift: scalable subsampled semidefinite program for K-means clustering
abstract
Semidefinite programming (SDP) is a powerful tool for tackling a wide range of computationally hard problems such as clustering. Despite the high accuracy, semidefinite programs are often too slow in practice with poor scalability on large (or even moderate) datasets. In this paper, we introduce a linear time complexity algorithm for approximating an SDP relaxed K-means clustering. The proposed sketch-and-lift (SL) approach solves an SDP on a subsampled dataset and then propagates the solution to all data points by a nearest-centroid rounding procedure. It is shown that the SL approach enjoys a similar exact recovery threshold as the K-means SDP on the full dataset, which is known to be information-theoretically tight under the Gaussian mixture model. The SL method can be made adaptive with enhanced theoretic properties when the cluster sizes are unbalanced. Our simulation experiments demonstrate that the statistical accuracy of the proposed method outperforms state-of-the-art fast clustering algorithms without sacrificing too much computational efficiency, and is comparable to the original K-means SDP with substantially reduced runtime.
Yubo Zhuang
AISTATS1
2022 Wasserstein $K$-means for clustering probability distributions
abstract
Clustering is an important exploratory data analysis technique to group objects based on their similarity. The widely used $K$-means clustering method relies on some notion of distance to partition data into a fewer number of groups. In the Euclidean space, centroid-based and distance-based formulations of the $K$-means are equivalent. In modern machine learning applications, data often arise as probability distributions and a natural generalization to handle measure-valued data is to use the optimal transport metric. Due to non-negative Alexandrov curvature of the Wasserstein space, barycenters suffer from regularity and non-robustness issues. The peculiar behaviors of Wasserstein barycenters may make the centroid-based formulation fail to represent the within-cluster data points, while the more direct distance-based $K$-means approach and its semidefinite program (SDP) relaxation are capable of recovering the true cluster labels. In the special case of clustering Gaussian distributions, we show that the SDP relaxed Wasserstein $K$-means can achieve exact recovery given the clusters are well-separated under the $2$-Wasserstein metric. Our simulation and real data examples also demonstrate that distance-based $K$-means can achieve better classification performance over the standard centroid-based $K$-means for clustering probability distributions and images.
Yubo Zhuang
NeurIPS1