EDBT 2026 Demo / reviewers in the wild / expert
Yuanzhe Cai
dblp:61/2822
· DBLP profile ↗
15ranked-venue papers
8as first author
4since 2021 · last 2025
0000-0002-1449-625XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 14 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hard Sample Aware Robust Contrastive Learning for Multi-View ClusteringabstractMulti-view clustering aims to divide samples into several clusters, by mining and utilizing the consistency and complementarity of multi-view data. Recent years, numerous deep contrastive multi-view clustering methods have been proposed to address the false negative issue by using self-supervised information. However, the quality of these self-supervised information was rarely taken into consideration, and using these information without discrimination can compromise training, leading to sub optimal performance. To tackle this issue, we propose Hard Sample Aware Robust Contrastive Learning for Multi-View Clustering(HearMVC). Concretely, we use self-supervised information and similarity to determine hard samples. The model focuses on these hard samples by assigning higher weights to enhance discriminative capability. Moreover, we utilize the confidence of self-supervised cluster assignment as weights, to strengthen the learning to confident samples and weaken the influence of unconfident samples. By simultaneously considering the weighting of hardness and confidence, our method can achieve best robustness and strongest discriminative capability. Extensive experiments on public datesets verify the effectiveness of our method. Yuanzhe Cai, Zhikui Chen, Jing Gao 0007, Peng Li 0027, Jianing Zhang 0001 |
ICASSP | 1 |
| 2025 | Effective and Efficient Similarity Search for DNA Sequences Through de Bruijn Sum Graph EmbeddingabstractSimilarity search of DNA sequences is widely used in many genomic analyses, such as pathogen detection, gene function annotation, and evolutionary relationship discovery. Today, sequencing technologies are generating more and more DNA sequences. This requires more accurate and efficient sequence search methods that scale well to large sequence databases. Here, we present a new accurate and efficient DNA search algorithm that scales well to large data as shown in experiments. This algorithm involves three innovative techniques: (i) de Bruijn sum graph, which is a natural representation of multiple DNA sequences, (ii) sampling from equilibrium distribution instead of traditional uniform distribution, and (iii) a technique to solve the sink difficulty in random walk sampling on the directed graph. A sequence corresponds to a path on de Bruijn sum graph, which generates a vector pooled from the path node embedding vectors. A query sequence similarly corresponds to a vector embedding. Thus, the similarity search becomes a vector data search, which can be implemented very efficiently in the vector database. We compare our implementation with MMseqs2 (the most accurate), Bowtie2 (the fastest), DNA2Vec, etc (see more details in experiments). Extensive experiments show that our implementation achieves (i) superior search accuracy (up to 3.5% Top-1 accuracy improvement) to MMseqs2, (ii) comparable search speed to Bowtie2, and (iii) 26.5 times faster than DNA2Vec (146 min vs. 3,869 min) on a 32GB data for learning k-mer embedding. Code is provided in https://github.com/caiyuanzhe/SeqGraph2Vec/. Zhaochong Yu, Zihang Yang, Chris Ding, Feijuan Huang, Yuanzhe Cai |
ICDM | 6 |
| 2023 | Music-Graph2Vec: An Efficient Method for Embedding Pitch SegmentabstractLearning low-dimensional continuous vector representation for short pitch segment extracted from songs is has been confirmed to contain tonal features of music, which is key to melody modeling that can be utilized in many music investigations, such as genre classification, emotion classification, and music retrieval, and so on. The skip-gram version of Word2Vec is ubiquitous, and widely used approach for music pitch segment embedding, but it poorly scales to large data sets due to its extremely long training time. In this paper, we propose a novel efficient graph-based embedding method, named Music-Graph2Vec, to tackle this concern. This approach converts music files into graphs, extracts the rhythmic sequence through random walking, and trains the rhythmic embedding model using skip-gram. Experimental results demonstrate that Music-Graph2Vec outperforms Word2Vec in training rhythmic embedding, with the advantage of being 55 times faster on the top-MAGD dataset (2,134.7s for Word2Vec and 38.9s for Music-Graph2Vec), with the same accuracy for Word2Vec in terms of music genre classification. Taiwei Wu, Yuanzhe Cai |
MMAsia | 4 |
| 2023 | Word-Graph2vec: An Efficient Word Embedding Approach on Word Co-occurrence Graph Using Random Walk Technique
Jiahong Xue, Huacan Chen, Feijuan Huang, Yuanzhe Cai |
WISE | 7 |
| 2020 | AnalyticDB-V: A Hybrid Analytical Engine Towards Query Fusion for Structured and Unstructured DataabstractWith the explosive growth of unstructured data (such as images, videos, and audios), unstructured data analytics is widespread in a rich vein of real-world applications. Many database systems start to incorporate unstructured data analysis to meet such demands. However, queries over unstructured and structured data are often treated as disjoint tasks in most systems, where hybrid queries ( i.e. , involving both data types) are not yet fully supported. In this paper, we present a hybrid analytic engine developed at Alibaba, named AnalyticDB-V (ADBV), to fulfill such emerging demands. ADBV offers an interface that enables users to express hybrid queries using SQL semantics by converting unstructured data to high dimensional vectors. ADBV adopts the lambda framework and leverages the merits of approximate nearest neighbor search (ANNS) techniques to support hybrid data analytics. Moreover, a novel ANNS algorithm is proposed to improve the accuracy on large-scale vectors representing massive unstructured data. All ANNS algorithms are implemented as physical operators in ADBV, meanwhile, accuracy-aware cost-based optimization techniques are proposed to identify effective execution plans. Experimental results on both public and in-house datasets show the superior performance achieved by ADBV and its effectiveness. ADBV has been successfully deployed on Alibaba Cloud to provide hybrid query processing services for various real-world applications. Chuangxian Wei, Bin Wu 0003, Sheng Wang 0011, Renjie Lou, Chaoqun Zhan, Feifei Li 0001, Yuanzhe Cai |
Proc. VLDB Endow. | 7 |
| 2013 | Expertise Ranking of Users in QA Community
Yuanzhe Cai, Sharma Chakravarthy |
DASFAA (1) | 1 |
| 2011 | Pairwise Similarity Calculation of Information Networks
Yuanzhe Cai, Sharma Chakravarthy |
DaWaK | 1 |
| 2011 | Low-order tensor decompositions for social tagging recommendationabstractSocial tagging recommendation is an urgent and useful enabling technology for Web 2.0. In this paper, we present a systematic study of low-order tensor decomposition approach that are specifically targeted at the very sparse data problem in tagging recommendation problem. Low-order polynomials have low functional complexity, are uniquely capable of enhancing statistics and also avoids over-fitting than traditional tensor decompositions such as Tucker and Parafac decompositions. We perform extensive experiments on several datasets and compared with 6 existing methods. Experimental results demonstrate that our approach outperforms existing approaches. Yuanzhe Cai, Dijun Luo, Chris Ding, Sharma Chakravarthy |
WSDM | 1 |
| 2010 | Local Methods for Estimating SimRank ScoreabstractSimRank is a well known algorithm which conducts link analysis to measure similarity between each pair of nodes (nodepair). But it suffers from high computational cost, limiting its usage in large-scale datasets. Moreover, Links between nodes are changing over time. It may be desirable to quickly approximate the similarity score between certain nodepair without performing a large-scale computation on the entire graph. In our approach we propose a method to efficiently estimate the similarity score using only a small subgraph of the entire graph. We call this novel algorithm “Local-SimRank”. The experimental results conducted on real datasets and synthetic dataset show that our algorithm efficiently produces good approximations to the global SimRank scores. Meanwhile, we prove that the Local-SimRank score LS(a, b) is always less than original SimRank score S(a, b) mathematically. Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001, Yuanzhe Cai |
APWeb | 6 |
| 2010 | Closed form solution of similarity algorithmsabstractAlgorithms defining similarities between objects of an information network are important of many IR tasks. SimRank algorithm and its variations are popularly used in many applications. Many fast algorithms are also developed. In this note, we first reformulate them as random walks on the network and express them using forward and backward transition probably in a matrix form. Second, we show that P-Rank (SimRank is only the special case of P-Rank) has a unique solution of eeT when decay factor c is equal to 1. We also show that SimFusion algorithm is a special case of P-Rank algorithm and prove that the similarity matrix of SimFusion is the product of PageRank vector. Our experiments on the web datasets show that for P-Rank the decay factor c doesn't seriously affect the similarity accuracy and accuracy of P-Rank is also higher than SimFusion and SimRank. Yuanzhe Cai, Chris Ding, Sharma Chakravarthy |
SIGIR | 1 |
| 2009 | Calculating Similarity Efficiently in a Small World
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
ADMA | 2 |
| 2009 | An Adaptive Method for the Efficient Similarity Calculation
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
DASFAA | 1 |
| 2009 | Efficient Algorithm for Computing Link-Based Similarity in Real World NetworksabstractSimilarity calculation has many applications, such as information retrieval, and collaborative filtering, among many others. It has been shown that link-based similarity measure, such as SimRank, is very effective in characterizing the object similarities in networks, such as the Web, by exploiting the object-to-object relationship. Unfortunately, it is prohibitively expensive to compute the link-based similarity in a relatively large graph. In this paper, based on the observation that link-based similarity scores of real world graphs follow the power-law distribution, we propose a new approximate algorithm, namely Power-SimRank, with guaranteed error bound to efficiently compute link-based similarity measure. We also prove the convergence of the proposed algorithm. Extensive experiments conducted on real world datasets and synthetic datasets show that the proposed algorithm outperforms SimRank by four-five times in terms of efficiency while the error generated by the approximation is small. Yuanzhe Cai, Gao Cong, Hongyan Liu 0002, Jun He 0008, Jiaheng Lu, Xiaoyong Du 0001 |
ICDM | 1 |
| 2009 | Exploiting the Block Structure of Link Graph for Efficient Similarity Computation
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
PAKDD | 2 |
| 2008 | S-SimRank: Combining Content and Link Information to Cluster Papers Effectively and Efficiently
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
ADMA | 1 |