VLDB 2026 Research / reviewers in the wild / expert
Gaifan Zhang
dblp:372/4043
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 67% Representation and self-supervised learning · 26% Language models and text generation · 7% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Theoretical computer science
1 paper |
Quantum computing and quantum information · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
1.0 | 1 | 2026 | Map of Encoders - Mapping Sentence Encoders using Quantum Relative Entropy · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis
data annotation |
0.9 | 1 | 2025 | Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › data annotation
LLM-based annotation |
0.9 | 1 | 2025 | Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity |
0.9 | 1 | 2025 | Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models · EMNLP 2025 |
Quantum computing and quantum information › quantum information theory
quantum relative entropy |
0.3 | 1 | 2026 | Map of Encoders - Mapping Sentence Encoders using Quantum Relative Entropy · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
quantum relative entropy · 3.0PIP matrix · 3.0large language model · 0.9data re-annotation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Map of Encoders - Mapping Sentence Encoders using Quantum Relative EntropyabstractWe propose a method to compare and visualise sentence encoders at scale by creating a map of encoders where each sentence encoder is represented in relation to the other sentence encoders. Specifically, we first represent a sentence encoder using an embedding matrix of a sentence set, where each row corresponds to the embedding of a sentence. Next, we compute the PIP matrix for a sentence encoder using its embedding matrix. Finally, we create a feature vector for each sentence encoder that reflects its QRE with respect to a unit base encoder. We construct a map of encoders covering 1101 publicly available sentence encoders, providing a new perspective of the landscape of the pre-trained sentence encoders. Our map accurately reflects various relationships between encoders, where encoders with similar attributes are proximally located on the map. Moreover, our encoder feature vectors can be used to accurately infer downstream task performance of the encoders, such as in retrieval and clustering tasks, demonstrating the correctness of our map. Gaifan Zhang, Danushka Bollegala |
ACL (1) | 1 |
| 2025 | Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language ModelsabstractSemantic similarity between two sentences depends on the aspects considered between those sentences.To study this phenomenon, Deshpande et al. (2023) proposed the Conditional Semantic Textual Similarity (C-STS) task and annotated a human-rated similarity dataset containing pairs of sentences compared under two different conditions.However, Tu et al. (2024)found various annotation issues in this dataset and showed that manually re-annotating a small portion of it leads to more accurate C-STS models.Despite these pioneering efforts, the lack of large and accurately annotated C-STS datasets remains a blocker for making progress on this task as evidenced by the subpar performance of the C-STS models.To address this training data need, we resort to Large Language Models (LLMs) to correct the condition statements and similarity ratings in the original dataset proposed by Deshpande et al. (2023).Our proposed method is able to reannotate a large training dataset for the C-STS task with minimal manual effort.Importantly, by training a supervised C-STS model on our cleaned and re-annotated dataset, we achieve a 5.4% statistically significant improvement in Spearman correlation.The re-annotated dataset is available at https://LivNLP.github. io/CSTS-reannotation. Gaifan Zhang, Yi Zhou 0019, Danushka Bollegala |
EMNLP | 1 |
| 2024 | Evaluating Unsupervised Dimensionality Reduction Methods for Pretrained Sentence EmbeddingsabstractSentence embeddings produced by Pretrained Language Models (PLMs) have received wide attention from the NLP community due to their superior performance when representing texts in numerous downstream applications. However, the high dimensionality of the sentence embeddings produced by PLMs is problematic when representing large numbers of sentences in memory- or compute-constrained devices. As a solution, we evaluate unsupervised dimensionality reduction methods to reduce the dimensionality of sentence embeddings produced by PLMs. Our experimental results show that simple methods such as Principal Component Analysis (PCA) can reduce the dimensionality of sentence embeddings by almost 50%, without incurring a significant loss in performance in multiple downstream tasks. Surprisingly, reducing the dimensionality further improves performance over the original high dimensional versions for the sentence embeddings produced by some PLMs in some tasks. Gaifan Zhang, Yi Zhou 0019, Danushka Bollegala |
LREC/COLING | 1 |