Martin Flechl

dblp:324/2686 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Representation and self-supervised learning · 56% Efficient and distributed learning · 28% Deep learning architectures and training · 17%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference efficiency
1.012026
Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models · ACL (1) 2026
Machine learning › Representation and self-supervised learning
text embedding
1.012026
Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models · ACL (1) 2026
Machine learning › Representation and self-supervised learning › text embedding
text representation learning
1.012026
Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models · ACL (1) 2026
Machine learning › Deep learning architectures and training
recurrent neural network
0.312026
Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models · ACL (1) 2026
Machine learning › Deep learning architectures and training
state space model
0.312026
Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

xLSTM · 1.0mamba2 · 1.0chunked inference · 1.0RWKV · 1.0
YearPublicationVenuePosition
2026 Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models
abstract
Transformer-based embedding models suffer from quadratic computational and linear memory complexity, limiting their utility for long sequences.We propose recurrent architectures as an efficient alternative, introducing a vertically chunked inference strategy that enables fast embedding generation with memory usage that becomes constant in the input length once it exceeds the vertical chunk size.By fine-tuning Mamba2 models, we demonstrate their viability as general-purpose text embedders, achieving competitive performance across a range of benchmarks while maintaining a substantially smaller memory footprint compared to transformer-based counterparts.We empirically validate the applicability of our inference strategy to Mamba2, RWKV, and xLSTM models, confirming consistent runtime-memory trade-offs across architectures and establishing recurrent models as a compelling alternative to transformers for efficient embedding generation.
Tobias Grantner, Emanuel Sallinger, Martin Flechl
ACL (1)3
2022 End-to-end speech recognition modeling from de-identified data
abstract
De-identification of data used for automatic speech recognition modeling is a critical component in protecting privacy, especially in the medical domain.However, simply removing all personally identifiable information (PII) from end-to-end model training data leads to a significant performance degradation in particular for the recognition of names, dates, locations, and words from similar categories.We propose and evaluate a two-step method for partially recovering this loss.First, PII is identified, and each occurrence is replaced with a random word sequence of the same category.Then, corresponding audio is produced via text-to-speech or by splicing together matching audio fragments extracted from the corpus.These artificial audio/label pairs, together with speaker turns from the original data without PII, are used to train models.We evaluate the performance of this method on in-house data of medical conversations and observe a recovery of almost the entire performance degradation in the general word error rate while still maintaining a strong diarization performance.Our main focus is the improvement of recall and precision in the recognition of PII-related words.Depending on the PII category, between 50% -90% of the performance degradation can be recovered using our proposed method.
Martin Flechl, Shou-Chun Yin, Peter Skala
INTERSPEECH1