Artyom Kozhevnikov

dblp:355/2092 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Representation and self-supervised learning · 100%

Topics — the 1 heaviest of 1, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding
0.912025
Less Mature is More Adaptable for Sentence-level Language Modeling · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

probing · 0.9mean pooling · 0.9fine-tuning · 0.9
YearPublicationVenuePosition
2025 Less Mature is More Adaptable for Sentence-level Language Modeling
abstract
This work investigates sentence-level models (i.e., models that operate at the sentence-level) to study how sentence representations from various encoders influence downstream task performance, and which syntactic, semantic, and discourse-level properties are essential for strong performance.Our experiments encompass encoders with diverse training regimes and pretraining domains, as well as various pooling strategies applied to multi-sentence input tasks (including sentence ordering, sentiment classification, and natural language inference) requiring coarse-to-fine-grained reasoning.We find that "less mature" representations (e.g., mean-pooled representations from BERT's first or last layer, or representations from encoders with limited fine-tuning) exhibit greater generalizability and adaptability to downstream tasks compared to representations from extensively fine-tuned models (e.g.,, SBERT or Sim-CSE).These findings are consistent across different pretraining seed initializations for BERT.Our probing analysis reveals that syntactic and discourse-level properties are stronger indicators of downstream performance than MTEB scores or decodability.Furthermore, the data and time efficiency of sentence-level models, often outperforming token-level models, underscores their potential for future research.
Abhilasha Sancheti, David Dale, Artyom Kozhevnikov, Maha Elbayad
ACL (1)3