Yuanzhe Cai

dblp:61/2822 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
4since 2021 · last 2025
0000-0002-1449-625XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Hard Sample Aware Robust Contrastive Learning for Multi-View Clustering
abstract
Multi-view clustering aims to divide samples into several clusters, by mining and utilizing the consistency and complementarity of multi-view data. Recent years, numerous deep contrastive multi-view clustering methods have been proposed to address the false negative issue by using self-supervised information. However, the quality of these self-supervised information was rarely taken into consideration, and using these information without discrimination can compromise training, leading to sub optimal performance. To tackle this issue, we propose Hard Sample Aware Robust Contrastive Learning for Multi-View Clustering(HearMVC). Concretely, we use self-supervised information and similarity to determine hard samples. The model focuses on these hard samples by assigning higher weights to enhance discriminative capability. Moreover, we utilize the confidence of self-supervised cluster assignment as weights, to strengthen the learning to confident samples and weaken the influence of unconfident samples. By simultaneously considering the weighting of hardness and confidence, our method can achieve best robustness and strongest discriminative capability. Extensive experiments on public datesets verify the effectiveness of our method.
Yuanzhe Cai, Zhikui Chen, Jing Gao 0007, Peng Li 0027, Jianing Zhang 0001
ICASSP1
2025 Effective and Efficient Similarity Search for DNA Sequences Through de Bruijn Sum Graph Embedding
abstract
Similarity search of DNA sequences is widely used in many genomic analyses, such as pathogen detection, gene function annotation, and evolutionary relationship discovery. Today, sequencing technologies are generating more and more DNA sequences. This requires more accurate and efficient sequence search methods that scale well to large sequence databases. Here, we present a new accurate and efficient DNA search algorithm that scales well to large data as shown in experiments. This algorithm involves three innovative techniques: (i) de Bruijn sum graph, which is a natural representation of multiple DNA sequences, (ii) sampling from equilibrium distribution instead of traditional uniform distribution, and (iii) a technique to solve the sink difficulty in random walk sampling on the directed graph. A sequence corresponds to a path on de Bruijn sum graph, which generates a vector pooled from the path node embedding vectors. A query sequence similarly corresponds to a vector embedding. Thus, the similarity search becomes a vector data search, which can be implemented very efficiently in the vector database. We compare our implementation with MMseqs2 (the most accurate), Bowtie2 (the fastest), DNA2Vec, etc (see more details in experiments). Extensive experiments show that our implementation achieves (i) superior search accuracy (up to 3.5% Top-1 accuracy improvement) to MMseqs2, (ii) comparable search speed to Bowtie2, and (iii) 26.5 times faster than DNA2Vec (146 min vs. 3,869 min) on a 32GB data for learning k-mer embedding. Code is provided in https://github.com/caiyuanzhe/SeqGraph2Vec/.
Zhaochong Yu, Zihang Yang, Chris Ding, Feijuan Huang, Yuanzhe Cai
ICDM6
2023 Music-Graph2Vec: An Efficient Method for Embedding Pitch Segment
abstract
Learning low-dimensional continuous vector representation for short pitch segment extracted from songs is has been confirmed to contain tonal features of music, which is key to melody modeling that can be utilized in many music investigations, such as genre classification, emotion classification, and music retrieval, and so on. The skip-gram version of Word2Vec is ubiquitous, and widely used approach for music pitch segment embedding, but it poorly scales to large data sets due to its extremely long training time. In this paper, we propose a novel efficient graph-based embedding method, named Music-Graph2Vec, to tackle this concern. This approach converts music files into graphs, extracts the rhythmic sequence through random walking, and trains the rhythmic embedding model using skip-gram. Experimental results demonstrate that Music-Graph2Vec outperforms Word2Vec in training rhythmic embedding, with the advantage of being 55 times faster on the top-MAGD dataset (2,134.7s for Word2Vec and 38.9s for Music-Graph2Vec), with the same accuracy for Word2Vec in terms of music genre classification.
Taiwei Wu, Yuanzhe Cai
MMAsia4
2023 Word-Graph2vec: An Efficient Word Embedding Approach on Word Co-occurrence Graph Using Random Walk Technique
Jiahong Xue, Huacan Chen, Feijuan Huang, Yuanzhe Cai
WISE7
2020 AnalyticDB-V: A Hybrid Analytical Engine Towards Query Fusion for Structured and Unstructured Data
abstract
With the explosive growth of unstructured data (such as images, videos, and audios), unstructured data analytics is widespread in a rich vein of real-world applications. Many database systems start to incorporate unstructured data analysis to meet such demands. However, queries over unstructured and structured data are often treated as disjoint tasks in most systems, where hybrid queries ( i.e. , involving both data types) are not yet fully supported. In this paper, we present a hybrid analytic engine developed at Alibaba, named AnalyticDB-V (ADBV), to fulfill such emerging demands. ADBV offers an interface that enables users to express hybrid queries using SQL semantics by converting unstructured data to high dimensional vectors. ADBV adopts the lambda framework and leverages the merits of approximate nearest neighbor search (ANNS) techniques to support hybrid data analytics. Moreover, a novel ANNS algorithm is proposed to improve the accuracy on large-scale vectors representing massive unstructured data. All ANNS algorithms are implemented as physical operators in ADBV, meanwhile, accuracy-aware cost-based optimization techniques are proposed to identify effective execution plans. Experimental results on both public and in-house datasets show the superior performance achieved by ADBV and its effectiveness. ADBV has been successfully deployed on Alibaba Cloud to provide hybrid query processing services for various real-world applications.
Chuangxian Wei, Bin Wu 0003, Sheng Wang 0011, Renjie Lou, Chaoqun Zhan, Feifei Li 0001, Yuanzhe Cai
Proc. VLDB Endow.7
2013 Expertise Ranking of Users in QA Community
Yuanzhe Cai, Sharma Chakravarthy
DASFAA (1)1
2011 Pairwise Similarity Calculation of Information Networks
Yuanzhe Cai, Sharma Chakravarthy
DaWaK1
2011 Low-order tensor decompositions for social tagging recommendation
abstract
Social tagging recommendation is an urgent and useful enabling technology for Web 2.0. In this paper, we present a systematic study of low-order tensor decomposition approach that are specifically targeted at the very sparse data problem in tagging recommendation problem. Low-order polynomials have low functional complexity, are uniquely capable of enhancing statistics and also avoids over-fitting than traditional tensor decompositions such as Tucker and Parafac decompositions. We perform extensive experiments on several datasets and compared with 6 existing methods. Experimental results demonstrate that our approach outperforms existing approaches.
Yuanzhe Cai, Dijun Luo, Chris Ding, Sharma Chakravarthy
WSDM1
2010 Local Methods for Estimating SimRank Score
abstract
SimRank is a well known algorithm which conducts link analysis to measure similarity between each pair of nodes (nodepair). But it suffers from high computational cost, limiting its usage in large-scale datasets. Moreover, Links between nodes are changing over time. It may be desirable to quickly approximate the similarity score between certain nodepair without performing a large-scale computation on the entire graph. In our approach we propose a method to efficiently estimate the similarity score using only a small subgraph of the entire graph. We call this novel algorithm “Local-SimRank”. The experimental results conducted on real datasets and synthetic dataset show that our algorithm efficiently produces good approximations to the global SimRank scores. Meanwhile, we prove that the Local-SimRank score LS(a, b) is always less than original SimRank score S(a, b) mathematically.
Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001, Yuanzhe Cai
APWeb6
2010 Closed form solution of similarity algorithms
abstract
Algorithms defining similarities between objects of an information network are important of many IR tasks. SimRank algorithm and its variations are popularly used in many applications. Many fast algorithms are also developed. In this note, we first reformulate them as random walks on the network and express them using forward and backward transition probably in a matrix form. Second, we show that P-Rank (SimRank is only the special case of P-Rank) has a unique solution of eeT when decay factor c is equal to 1. We also show that SimFusion algorithm is a special case of P-Rank algorithm and prove that the similarity matrix of SimFusion is the product of PageRank vector. Our experiments on the web datasets show that for P-Rank the decay factor c doesn't seriously affect the similarity accuracy and accuracy of P-Rank is also higher than SimFusion and SimRank.
Yuanzhe Cai, Chris Ding, Sharma Chakravarthy
SIGIR1
2009 Calculating Similarity Efficiently in a Small World
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001
ADMA2
2009 An Adaptive Method for the Efficient Similarity Calculation
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001
DASFAA1
2009 Efficient Algorithm for Computing Link-Based Similarity in Real World Networks
abstract
Similarity calculation has many applications, such as information retrieval, and collaborative filtering, among many others. It has been shown that link-based similarity measure, such as SimRank, is very effective in characterizing the object similarities in networks, such as the Web, by exploiting the object-to-object relationship. Unfortunately, it is prohibitively expensive to compute the link-based similarity in a relatively large graph. In this paper, based on the observation that link-based similarity scores of real world graphs follow the power-law distribution, we propose a new approximate algorithm, namely Power-SimRank, with guaranteed error bound to efficiently compute link-based similarity measure. We also prove the convergence of the proposed algorithm. Extensive experiments conducted on real world datasets and synthetic datasets show that the proposed algorithm outperforms SimRank by four-five times in terms of efficiency while the error generated by the approximation is small.
Yuanzhe Cai, Gao Cong, Hongyan Liu 0002, Jun He 0008, Jiaheng Lu, Xiaoyong Du 0001
ICDM1
2009 Exploiting the Block Structure of Link Graph for Efficient Similarity Computation
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001
PAKDD2
2008 S-SimRank: Combining Content and Link Information to Cluster Papers Effectively and Efficiently
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001
ADMA1