VLDB 2026 Research / reviewers in the wild / expert
Maolin Che
dblp:174/8695
· DBLP profile ↗
15ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0001-9956-062XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 5 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient frequent directions algorithms for approximate decomposition of matrices and higher-order tensorsabstractIn the framework of the FD (frequent directions) algorithm, we first develop two efficient algorithms for low-rank matrix approximations under the embedding matrices composed of the product of any SpEmb (sparse embedding) matrix and any standard Gaussian matrix, or any SpEmb matrix and any SRHT (subsampled randomized Hadamard transform) matrix. The theoretical results are also achieved based on the bounds of singular values of standard Gaussian matrices and the theoretical results for SpEmb and SRHT matrices. With a given Tucker-rank, we then obtain several efficient FD-based randomized variants of T-HOSVD (the truncated high-order singular value decomposition) and ST-HOSVD (sequentially T-HOSVD), which are two common algorithms for computing the approximate Tucker decomposition of any tensor with a given Tucker-rank. We also consider efficient FD-based randomized algorithms for computing the approximate TT (tensor-train) decomposition of any tensor with a given TT-rank. Finally, we illustrate the efficiency and accuracy of these algorithms using synthetic and real-world matrix (and tensor) data. Maolin Che, Yimin Wei 0001, Hong Yan 0001 |
J. Mach. Learn. Res. | 1 |
| 2026 | The CUR Decomposition of Self-Attention Matrices in Vision TransformersabstractTransformers have achieved great success in natural language processing and computer vision. The core and basic technique of transformers is the self-attention mechanism. The vanilla self-attention mechanism has quadratic complexity, which limits its applications to vision tasks. Most of the existing linear self-attention mechanisms will sacrifice performance to some extent to reduce complexity. In this paper, we propose a novel linear approximation of the vanilla self-attention mechanism named CURSA to achieve both high performance and low complexity at the same time. CURSA is based on the CUR decomposition to decompose the multiplication of large matrices into the multiplication of several small matrices to achieve almost linear complexity. Experiment results of CURSA in image classification tasks, semantic segmentation tasks, object detection tasks, and long-range arena show that it outperforms state-of-the-art self-attention mechanisms with better data efficiency, faster speed, and higher accuracy. Chong Wu 0007, Maolin Che, Hong Yan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | ELFATT: Efficient Linear Fast Attention for Vision TransformersabstractThe attention mechanism is the key to the success of transformers in different machine learning tasks. However, the quadratic complexity with respect to the sequence length of the vanilla softmax-based attention mechanism becomes the major bottleneck for the application of long sequence tasks, such as vision tasks. Although various efficient linear attention mechanisms have been proposed, they need to sacrifice performance to achieve high efficiency. What's more, memory-efficient methods, such as FlashAttention-1-3, still have quadratic computation complexity which can be further improved. In this paper, we propose a novel efficient linear fast attention (ELFATT) mechanism to achieve low memory input/output operations, linear computational complexity, and high performance at the same time. ELFATT offers 4-7x speedups over the vanilla softmax-based attention mechanism in high-resolution vision tasks without losing performance. ELFATT is FlashAttention friendly. Using FlashAttention-2 acceleration, ELFATT still offers 2-3x speedups over the vanilla softmax-based attention mechanism on high-resolution vision tasks without losing performance. Even in some non-vision tasks of long-range arena, ELFATT still achieves leading performance and offers 1.2-2.3x speedups over FlashAttention-2. Even on edge GPUs, ELFATT still offers 1.6x to 2.0x speedups compared to state-of-the-art attention mechanisms in various power modes from 5W to 60W. Furthermore, ELFATT can be used to enhance and accelerate diffusion tasks directly without training. Chong Wu 0007, Maolin Che, Zhuoheng Ran, Hong Yan 0001 |
ACM Multimedia | 2 |
| 2025 | DuSA: Fast and Accurate Dual-Stage Sparse Attention Mechanism Accelerating Both Training and InferenceabstractThis paper proposes the Dual-Stage Sparse Attention (DuSA) mechanism for attention acceleration of transformers. In the first stage, DuSA performs intrablock sparse attention to aggregate local inductive biases. In the second stage, DuSA performs interblock sparse attention to obtain long-range dependencies. Both stages have low computational complexity and can be further accelerated by memory acceleration attention mechanisms directly, which makes DuSA faster than some extremely fast attention mechanisms. The dual-stage sparse attention design provides a lower error in approximating vanilla scaled-dot product attention than the basic single-stage sparse attention mechanisms and further advances the basic sparse attention mechanisms to match or even outperform vanilla scaled-dot product attention. Even in some plug and play situations, DuSA can still maintain low performance loss. DuSA can be used in both training and inference acceleration. DuSA achieves leading performance in different benchmarks: long range arena, image classification, semantic segmentation, object detection, text to video generation, and long context understanding, and accelerates models of different sizes. Chong Wu 0007, Jiawang Cao, Zhuoheng Ran, Maolin Che, Hong Yan 0001 |
NeurIPS | 5 |
| 2025 | Gradient neural network models for approximate Tucker decomposition of time-dependent tensors
Maolin Che, Yimin Wei 0001, Hong Yan 0001 |
Neurocomputing | 1 |
| 2024 | Fixed-precision randomized quaternion singular value decomposition algorithm for low-rank quaternion matrix approximations
Yonghe Liu, Fengsheng Wu, Maolin Che, Chaoqian Li |
Neurocomputing | 3 |
| 2024 | Sketch-based multiplicative updating algorithms for symmetric nonnegative tensor factorizations with applications to face image clustering
Maolin Che, Yimin Wei 0001, Hong Yan 0001 |
J. Glob. Optim. | 1 |
| 2023 | Randomized algorithms for the computation of multilinear rank-(μ 1,μ 2,μ 3) approximations
Maolin Che, Yimin Wei 0001, Yanwei Xu 0004 |
J. Glob. Optim. | 1 |
| 2020 | A Unified Self-Stabilizing Neural Network Algorithm for Principal Takagi Component Extraction
Maolin Che, Xuezhong Wang, Yimin Wei 0001 |
Neural Process. Lett. | 1 |
| 2019 | Neural networks based approach solving multi-linear systems with M-tensors
Xuezhong Wang, Maolin Che, Yimin Wei 0001 |
Neurocomputing | 2 |
| 2018 | Adaptive algorithms for computing the principal Takagi vector of a complex symmetric matrix
Maolin Che, Sanzheng Qiao, Yimin Wei 0001 |
Neurocomputing | 1 |
| 2018 | Geometric measures of entanglement in multipartite pure states via complex-valued neural networks
Maolin Che, Liqun Qi 0001, Yimin Wei 0001, Guofeng Zhang 0003 |
Neurocomputing | 1 |
| 2017 | Neural networks for computing best rank-one approximations of tensors and its applications
Maolin Che, Andrzej Cichocki, Yimin Wei 0001 |
Neurocomputing | 1 |
| 2017 | Complex-valued neural networks for the Takagi vector of complex symmetric matrices
Xuezhong Wang, Maolin Che, Yimin Wei 0001 |
Neurocomputing | 2 |
| 2016 | Recurrent neural network for computation of generalized eigenvalue problem with real diagonalizable matrix pair and its applications
Xuezhong Wang, Maolin Che, Yimin Wei 0001 |
Neurocomputing | 2 |