VLDB 2026 Research / reviewers in the wild / expert
Artyom Kozhevnikov
dblp:355/2092
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Representation and self-supervised learning · 100% |
Topics — the 1 heaviest of 1, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
0.9 | 1 | 2025 | Less Mature is More Adaptable for Sentence-level Language Modeling · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
probing · 0.9mean pooling · 0.9fine-tuning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Less Mature is More Adaptable for Sentence-level Language ModelingabstractThis work investigates sentence-level models (i.e., models that operate at the sentence-level) to study how sentence representations from various encoders influence downstream task performance, and which syntactic, semantic, and discourse-level properties are essential for strong performance.Our experiments encompass encoders with diverse training regimes and pretraining domains, as well as various pooling strategies applied to multi-sentence input tasks (including sentence ordering, sentiment classification, and natural language inference) requiring coarse-to-fine-grained reasoning.We find that "less mature" representations (e.g., mean-pooled representations from BERT's first or last layer, or representations from encoders with limited fine-tuning) exhibit greater generalizability and adaptability to downstream tasks compared to representations from extensively fine-tuned models (e.g.,, SBERT or Sim-CSE).These findings are consistent across different pretraining seed initializations for BERT.Our probing analysis reveals that syntactic and discourse-level properties are stronger indicators of downstream performance than MTEB scores or decodability.Furthermore, the data and time efficiency of sentence-level models, often outperforming token-level models, underscores their potential for future research. Abhilasha Sancheti, David Dale, Artyom Kozhevnikov, Maha Elbayad |
ACL (1) | 3 |