Georgios Milis

dblp:364/1966 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Digital forensics and information hiding · 100%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
audio generation
0.912025
Robust Distortion-Free Watermark for Autoregressive Audio Generation Models · NeurIPS 2025
Digital forensics and information hiding › watermarking
audio watermarking
0.912025
Robust Distortion-Free Watermark for Autoregressive Audio Generation Models · NeurIPS 2025
Digital forensics and information hiding › watermarking
generative model watermarking
0.912025
Robust Distortion-Free Watermark for Autoregressive Audio Generation Models · NeurIPS 2025
Digital forensics and information hiding
watermarking
0.912025
Robust Distortion-Free Watermark for Autoregressive Audio Generation Models · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

distortion-free watermark · 1.7clustering-based watermark · 1.7
YearPublicationVenuePosition
2025 A Watermark for Auto-Regressive Speech Generation Models
Georgios Milis, Heng Huang 0001
INTERSPEECH3
2025 Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
abstract
The rapid advancement of next-token-prediction models has led to widespread adoption across modalities, enabling the creation of realistic synthetic media. In the audio domain, while autoregressive speech models have propelled conversational interactions forward, the potential for misuse, such as impersonation in phishing schemes or crafting misleading speech recordings, has also increased. Security measures such as watermarking have thus become essential to ensuring the authenticity of digital media. Traditional statistical watermarking methods used for autoregressive language models face challenges when applied to autoregressive audio models, due to the inevitable ``retokenization mismatch'' - the discrepancy between original and retokenized discrete audio token sequences. To address this, we introduce Aligned-IS, a novel, distortion-free watermark, specifically crafted for audio generation models. This technique utilizes a clustering approach that treats tokens within the same cluster equivalently, effectively countering the retokenization mismatch issue. Our comprehensive testing on prevalent audio generation platforms demonstrates that Aligned-IS not only preserves the quality of generated audio but also significantly improves the watermark detectability compared to the state-of-the-art distortion-free watermarking adaptations, establishing a new benchmark in secure audio technology applications.
Georgios Milis, Heng Huang 0001
NeurIPS2