Zerui Tao

dblp:296/4527 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2025
0009-0003-9230-721XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 25% Optimization for machine learning · 23% Efficient and distributed learning · 18%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
tensor decomposition
0.922025
Undirected Probabilistic Model for Tensor Decomposition · NeurIPS 2023
Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models · ICCV 2025
Machine learning › Generative modeling › diffusion model
controllable generation
0.912025
Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models · ICCV 2025
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.912025
Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models · ICCV 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models · ICCV 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models · ICCV 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.812024
Efficient Nonparametric Tensor Decomposition for Binary and Count Data · AAAI 2024
Machine learning › Learning theory
generalization bounds
0.812024
Generalization Analysis of Stochastic Weight Averaging with General Sampling · ICML 2024
Machine learning › Optimization for machine learning › mini-batch sampling
sampling without replacement
0.812024
Generalization Analysis of Stochastic Weight Averaging with General Sampling · ICML 2024
Machine learning › Optimization for machine learning
stochastic gradient descent
0.812024
Generalization Analysis of Stochastic Weight Averaging with General Sampling · ICML 2024
Machine learning › Optimization for machine learning › iterate averaging
stochastic weight averaging
0.812024
Generalization Analysis of Stochastic Weight Averaging with General Sampling · ICML 2024
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis
tensor factorization
0.812024
Efficient Nonparametric Tensor Decomposition for Binary and Count Data · AAAI 2024
Machine learning › Generative modeling
energy-based model
0.712023
Undirected Probabilistic Model for Tensor Decomposition · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent factor model
0.712023
Undirected Probabilistic Model for Tensor Decomposition · NeurIPS 2023
Mathematical optimization
metaheuristic optimization
0.612022
Permutation Search of Tensor Network Structures via Local Sampling · ICML 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
noise contrastive estimation
0.212023
Undirected Probabilistic Model for Tensor Decomposition · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

variational inference · 1.5stochastic natural gradient · 1.5pólya-gamma augmentation · 1.5tensor decomposition · 0.9low-rank adaptation · 0.9knowledge distillation · 0.9uniform stability · 0.8mathematical induction · 0.8noise contrastive estimation · 0.7deep energy-based model · 0.7metaheuristic search · 0.6local sampling · 0.6
YearPublicationVenuePosition
2025 Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models
abstract
Parameter-Efficient Fine-Tuning (PEFT) of text-to-image models has become an increasingly popular technique with many applications. Among the various PEFT methods, Low-Rank Adaptation (LoRA) and its variants have gained significant attention due to their effectiveness, enabling users to fine-tune models with limited computational resources. However, the approximation gap between the low-rank assumption and desired fine-tuning weights prevents the simultaneous acquisition of ultra-parameter-efficiency and better performance. To reduce this gap and further improve the power of LoRA, we propose a new PEFT method that combines two classes of adaptations, namely, transform and residual adaptations. In specific, we first apply a full-rank and dense transform to the pre-trained weight. This learnable transform is expected to align the pre-trained weight as closely as possible to the desired weight, thereby reducing the rank of the residual weight. Then, the residual part can be effectively approximated by more compact and parameter-efficient structures, with a smaller approximation error. To achieve ultra-parameter-efficiency in practice, we design highly flexible and effective tensor decompositions for both the transform and residual adaptations. Additionally, popular PEFT methods such as DoRA can be summarized under this transform plus residual adaptation scheme. Experiments are conducted on fine-tuning Stable Diffusion models in subject-driven and controllable generation. The results manifest that our method can achieve better performances and parameter efficiency compared to LoRA and several baselines.
Zerui Tao, Yuhta Takida, Naoki Murata, Qibin Zhao, Yuki Mitsufuji
ICCV1
2025 Adversarial guided diffusion models for adversarial purification
Guang Lin 0002, Zerui Tao, Toshihisa Tanaka 0001, Qibin Zhao
Neural Networks2
2024 Efficient Nonparametric Tensor Decomposition for Binary and Count Data
abstract
In numerous applications, binary reactions or event counts are observed and stored within high-order tensors. Tensor decompositions (TDs) serve as a powerful tool to handle such high-dimensional and sparse data. However, many traditional TDs are explicitly or implicitly designed based on the Gaussian distribution, which is unsuitable for discrete data. Moreover, most TDs rely on predefined multi-linear structures, such as CP and Tucker formats. Therefore, they may not be effective enough to handle complex real-world datasets. To address these issues, we propose ENTED, an Efficient Nonparametric TEnsor Decomposition for binary and count tensors. Specifically, we first employ a nonparametric Gaussian process (GP) to replace traditional multi-linear structures. Next, we utilize the Pólya-Gamma augmentation which provides a unified framework to establish conjugate models for binary and count distributions. Finally, to address the computational issue of GPs, we enhance the model by incorporating sparse orthogonal variational inference of inducing points, which offers a more effective covariance approximation within GPs and stochastic natural gradient updates for nonparametric models. We evaluate our model on several real-world tensor completion tasks, considering binary and count datasets. The results manifest both better performance and computational advantages of the proposed model.
Zerui Tao, Toshihisa Tanaka 0001, Qibin Zhao
AAAI1
2024 Generalization Analysis of Stochastic Weight Averaging with General Sampling
abstract
Stochastic weight averaging (SWA) method has empirically proven its advantages compared to stochastic gradient descent (SGD). Despite it is widespread used, theoretical investigations have been limited, particularly in scenarios beyond the ideal setting of convex and sampling with replacement. However, non-convex cases and sampling without replacement are very practical in real-world applications. The main challenges under the above settings are two-folds: (i) All the historical gradient information introduced by SWA is considered, while the analysis of SGD using the tool of uniform stability requires only to bound the current gradient. (ii) The $(1+\alpha\beta)$-expansion property causes the boundary of each gradient step dependent on the previous step, making the boundary of each historical gradient in SWA nested and the theoretical analysis even harder. To address the theoretical challenges, we adopt mathematical induction to find a recursive representation that bounds the gradient at each step. Based on this, we establish stability bounds supporting sampling with and without replacement in the non-convex setting. Furthermore, the derived generalization bounds of SWA are sharper than SGD. At last, experimental results on several benchmarks verify our theoretical results.
Li Shen 0008, Zerui Tao, Shuaida He, Dacheng Tao
ICML3
2024 Nonparametric tensor ring decomposition with scalable amortized inference
Zerui Tao, Toshihisa Tanaka 0001, Qibin Zhao
Neural Networks1
2023 Scalable Bayesian Tensor Ring Factorization for Multiway Data Analysis
Zerui Tao, Toshihisa Tanaka 0001, Qibin Zhao
ICONIP (1)1
2023 Undirected Probabilistic Model for Tensor Decomposition
abstract
Tensor decompositions (TDs) serve as a powerful tool for analyzing multiway data. Traditional TDs incorporate prior knowledge about the data into the model, such as a directed generative process from latent factors to observations. In practice, selecting proper structural or distributional assumptions beforehand is crucial for obtaining a promising TD representation. However, since such prior knowledge is typically unavailable in real-world applications, choosing an appropriate TD model can be challenging. This paper aims to address this issue by introducing a flexible TD framework that discards the structural and distributional assumptions, in order to learn as much information from the data. Specifically, we construct a TD model that captures the joint probability of the data and latent tensor factors through a deep energy-based model (EBM). Neural networks are then employed to parameterize the joint energy function of tensor factors and tensor entries. The flexibility of EBM and neural networks enables the learning of underlying structures and distributions. In addition, by designing the energy function, our model unifies the learning process of different types of tensors, such as static tensors and dynamic tensors with time stamps. The resulting model presents a doubly intractable nature due to the presence of latent tensor factors and the unnormalized probability function. To efficiently train the model, we derive a variational upper bound of the conditional noise-contrastive estimation objective that learns the unnormalized joint probability by distinguishing data from conditional noises. We show advantages of our model on both synthetic and several real-world datasets.
Zerui Tao, Toshihisa Tanaka 0001, Qibin Zhao
NeurIPS1
2022 Permutation Search of Tensor Network Structures via Local Sampling
abstract
Recent works put much effort into tensor network structure search (TN-SS), aiming to select suitable tensor network (TN) structures, involving the TN-ranks, formats, and so on, for the decomposition or learning tasks. In this paper, we consider a practical variant of TN-SS, dubbed TN permutation search (TN-PS), in which we search for good mappings from tensor modes onto TN vertices (core tensors) for compact TN representations. We conduct a theoretical investigation of TN-PS and propose a practically-efficient algorithm to resolve the problem. Theoretically, we prove the counting and metric properties of search spaces of TN-PS, analyzing for the first time the impact of TN structures on these unique properties. Numerically, we propose a novel meta-heuristic algorithm, in which the searching is done by randomly sampling in a neighborhood established in our theory, and then recurrently updating the neighborhood until convergence. Numerical results demonstrate that the new algorithm can reduce the required model size of TNs in extensive benchmarks, implying the improvement in the expressive power of TNs. Furthermore, the computational cost for the new algorithm is significantly less than that in (Li and Sun, 2020).
Chao Li 0013, Junhua Zeng, Zerui Tao, Qibin Zhao
ICML3
2021 Bayesian Latent Factor Model for Higher-order Data
abstract
Latent factor models are canonical tools to learn low-dimensional and linear embedding of original data. Traditional latent factor models are based on low-rank matrix factorization of covariance matrices. However, for higher-order data with multiple modes, i.e., tensors, this simple treatment fails to take into account the mode-specific relations. This ignorance leads to inefficiency in analysis of complex structures as well as poor data compression ability. In this paper, unlike covariance matrices, we investigate high-order covariance tensor directly by exploiting tensor ring (TR) format and propose the Bayesian TR latent factor model, which can represent complex multi-linear correlations and achieves efficient data compression. To overcome the difficulty of finding the optimal TR-ranks and simultaneously imposing sparsity on loading coefficients, a multiplicative Gamma process (MGP) prior is adopted to automatically infer the ranks and obtain sparsity. Then, we establish an efficient parameter-expanded EM algorithm to learn the maximum a posteriori (MAP) estimate of model parameters. Finally, we evaluate our model on covariance estimation, latent factor learning and image inpainting problems.
Zerui Tao, Toshihisa Tanaka 0001, Qibin Zhao
ACML1
2021 Tensor Decomposition Via Core Tensor Networks
abstract
Tensor decomposition (TD) has shown promising performance in image completion and denoising. Existing methods always aim to decompose one tensor into latent factors or core tensors by optimizing a particular cost function based on a specific tensor model. These algorithms iteratively learn the optima from random initialization given any individual tensor, resulting in slow convergence and low efficiency. In this paper, we propose an efficient TD algorithm that aims to learn a global mapping from input tensors to latent core tensors, under the assumption that the mappings of multiple tensors might be shared or highly correlated. To this end, we train a deep neural network (DNN) to model the global mapping and then apply it to decompose a newly given tensor with high efficiency. Furthermore, the initial values of DNN are learned based on meta-learning methods. By leveraging the pretrained core tensor DNN, our proposed method enables us to perform TD efficiently and accurately. Experimental results demonstrate the significant improvements of our method over other TD methods in terms of speed and accuracy.
Jianfu Zhang 0003, Zerui Tao, Liqing Zhang 0001, Qibin Zhao
ICASSP2