VLDB 2026 Research / reviewers in the wild / expert
Zihang Yang
dblp:260/6351
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Index Benefit Estimation via Hierarchical and Two-Dimensional Feature Representation
Feng Liang 0004, Jinqi Quan, Zihang Yang, Runhuai Huang, Xiping Hu, Haipeng Dai |
ICDE | 4 |
| 2025 | Effective and Efficient Similarity Search for DNA Sequences Through de Bruijn Sum Graph EmbeddingabstractSimilarity search of DNA sequences is widely used in many genomic analyses, such as pathogen detection, gene function annotation, and evolutionary relationship discovery. Today, sequencing technologies are generating more and more DNA sequences. This requires more accurate and efficient sequence search methods that scale well to large sequence databases. Here, we present a new accurate and efficient DNA search algorithm that scales well to large data as shown in experiments. This algorithm involves three innovative techniques: (i) de Bruijn sum graph, which is a natural representation of multiple DNA sequences, (ii) sampling from equilibrium distribution instead of traditional uniform distribution, and (iii) a technique to solve the sink difficulty in random walk sampling on the directed graph. A sequence corresponds to a path on de Bruijn sum graph, which generates a vector pooled from the path node embedding vectors. A query sequence similarly corresponds to a vector embedding. Thus, the similarity search becomes a vector data search, which can be implemented very efficiently in the vector database. We compare our implementation with MMseqs2 (the most accurate), Bowtie2 (the fastest), DNA2Vec, etc (see more details in experiments). Extensive experiments show that our implementation achieves (i) superior search accuracy (up to 3.5% Top-1 accuracy improvement) to MMseqs2, (ii) comparable search speed to Bowtie2, and (iii) 26.5 times faster than DNA2Vec (146 min vs. 3,869 min) on a 32GB data for learning k-mer embedding. Code is provided in https://github.com/caiyuanzhe/SeqGraph2Vec/. Zhaochong Yu, Zihang Yang, Chris Ding, Feijuan Huang, Yuanzhe Cai |
ICDM | 2 |
| 2022 | Machine Reading Comprehension Based on Hybrid Attention and Controlled Generation
Feng Gao 0003, Zihang Yang, Jinguang Gu, Junjun Cheng |
WISA | 2 |