Giulia Berardinelli

dblp:362/5233 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 67% Language models and text generation · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › efficient language model
efficient language model architectures
0.812024
ShareBERT: Embeddings Are Capable of Learning Hidden Layers · AAAI 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
ShareBERT: Embeddings Are Capable of Learning Hidden Layers · AAAI 2024
Machine learning › Efficient and distributed learning
parameter sharing
0.812024
ShareBERT: Embeddings Are Capable of Learning Hidden Layers · AAAI 2024

Methods — techniques the papers use, named apart from their topics

parameter sharing · 0.8knowledge distillation · 0.8
YearPublicationVenuePosition
2024 ShareBERT: Embeddings Are Capable of Learning Hidden Layers
abstract
The deployment of Pre-trained Language Models in memory-limited devices is hindered by their massive number of parameters, which motivated the interest in developing smaller architectures. Established works in the model compression literature showcased that small models often present a noticeable performance degradation and need to be paired with transfer learning methods, such as Knowledge Distillation. In this work, we propose a parameter-sharing method that consists of sharing parameters between embeddings and the hidden layers, enabling the design of near-zero parameter encoders. To demonstrate its effectiveness, we present an architecture design called ShareBERT, which can preserve up to 95.5% of BERT Base performances, using only 5M parameters (21.9× fewer parameters) without the help of Knowledge Distillation. We demonstrate empirically that our proposal does not negatively affect the model learning capabilities and that it is even beneficial for representation learning. Code will be available at https://github.com/jchenghu/sharebert.
Jia-Cheng Hu, Roberto Cavicchioli, Giulia Berardinelli, Alessandro Capotondi
AAAI3
2024 Learning from Wrong Predictions in Low-Resource Neural Machine Translation
abstract
Resource scarcity in Neural Machine Translation is a challenging problem in both industry applications and in the support of less-spoken languages represented, in the worst case, by endangered and low-resource languages. Many Data Augmentation methods rely on additional linguistic sources and software tools but these are often not available in less favoured language. For this reason, we present USKI (Unaligned Sentences Keytokens pre-traIning), a pre-training strategy that leverages the relationships and similarities that exist between unaligned sentences. By doing so, we increase the dataset size of endangered and low-resource languages by the square of the initial quantity, matching the typical size of high-resource language datasets such as WMT14 En-Fr. Results showcase the effectiveness of our approach with an increase on average of 0.9 BLEU across the benchmarks using a small fraction of the entire unaligned corpus, suggesting the importance of the research topic and the potential of a currently under-utilized resource and under-explored approach.
Jia-Cheng Hu, Roberto Cavicchioli, Giulia Berardinelli, Alessandro Capotondi
LREC/COLING3