Pol van Rijn

dblp:271/8093 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-4044-9123ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Are Expressions for Music Emotions the Same Across Cultures?
Elif Çelen, Pol van Rijn, Harin Lee, Nori Jacoby
CogSci2
2025 Visual and Musical Aesthetic Preferences Across Cultures
Harin Lee, Eline Van Geert, Elif Çelen, Raja Marjieh, Pol van Rijn, Minsu Park 0002, Nori Jacoby
CogSci5
2024 Characterizing Similarities and Divergences in Conversational Tones in Humans and LLMs by Sampling with People
abstract
Conversational tones -the manners and attitudes in which speakers communicate -are essential to effective communication.Amidst the increasing popularization of Large Language Models (LLMs) over recent years, it becomes necessary to characterize the divergences in their conversational tones relative to humans.However, existing investigations of conversational modalities rely on pre-existing taxonomies or text corpora, which suffer from experimenter bias and may not be representative of real-world distributions for the studies' psycholinguistic domains.Inspired by methods from cognitive science, we propose an iterative method for simultaneously eliciting conversational tones and sentences, where participants alternate between two tasks: (1) one participant identifies the tone of a given sentence and (2) a different participant generates a sentence based on that tone.We run 100 iterations of this process with human participants and GPT-4, then obtain a dataset of sentences and frequent conversational tones.In an additional experiment, humans and GPT-4 annotated all sentences with all tones.With data from 1,339 human participants, 33,370 human judgments, and 29,900 GPT-4 queries, we show how our approach can be used to create an interpretable geometric representation of relations between conversational tones in humans and GPT-4.This work demonstrates how combining ideas from machine learning and cognitive science can address challenges in human-computer interactions.B Sampling paradigm A Problem statement C Quality-of-fit rating D Shared space E Benchmark What are similarities and divergences in conversational tones in humans and LLMs?I am pretty certain, Tom ate all the cookies from the jar.Could it be possible that Tom ate all the cookies from the jar? polite?Humans
Dun-Ming Huang, Pol van Rijn, Ilia Sucholutsky, Raja Marjieh, Nori Jacoby
ACL (1)2
2024 Giving Robots a Voice: Human-in-the-Loop Voice Creation and open-ended Labeling
abstract
Speech is a natural interface for humans to interact with robots. Yet, aligning a robot’s voice to its appearance is challenging due to the rich vocabulary of both modalities. Previous research has explored a few labels to describe robots and tested them on a limited number of robots and existing voices. Here, we develop a robot-voice creation tool followed by large-scale behavioral human experiments (N=2,505). First, participants collectively tune robotic voices to match 175 robot images using an adaptive human-in-the-loop pipeline. Then, participants describe their impression of the robot or their matched voice using another human-in-the-loop paradigm for open-ended labeling. The elicited taxonomy is then used to rate robot attributes and to predict the best voice for an unseen robot. We offer a web interface to aid engineers in customizing robot voices, demonstrating the synergy between cognitive science and machine learning for engineering tools.
Pol van Rijn, Silvan Mertes, Kathrin Janowski, Katharina Weitz, Nori Jacoby, Elisabeth André
CHI1
2024 A Rational Analysis of the Speech-to-Song Illusion
Raja Marjieh, Pol van Rijn, Ilia Sucholutsky, Harin Lee, Thomas L. Griffiths 0001, Nori Jacoby
CogSci2
2024 Studying the Effect of Globalization on Color Perception using Multilingual Online Recruitment and Large Language Models
Jakob Pete Niedermann, Ilia Sucholutsky, Raja Marjieh, Elif Çelen, Thomas L. Griffiths 0001, Nori Jacoby, Pol van Rijn
CogSci7
2023 What Language Reveals about Perception: Distilling Psychophysical Knowledge from Large Language Models
Raja Marjieh, Ilia Sucholutsky, Pol van Rijn, Nori Jacoby, Thomas L. Griffiths 0001
CogSci3
2023 Around the world in 60 words: A generative vocabulary test for online research
Pol van Rijn, Harin Lee, Raja Marjieh, Ilia Sucholutsky, Francesca Lanzarini, Elisabeth André, Nori Jacoby
CogSci1
2023 Words are all you need? Language as an approximation for human similarity judgments
Raja Marjieh, Pol van Rijn, Ilia Sucholutsky, Theodore R. Sumers, Harin Lee, Thomas L. Griffiths 0001, Nori Jacoby
ICLR2
2022 Bridging the prosody GAP: Genetic Algorithm with People to efficiently sample emotional prosody
Pol van Rijn, Harin Lee, Nori Jacoby
CogSci1
2022 VoiceMe: Personalized voice generation in TTS
Pol van Rijn, Silvan Mertes, Dominik Schiller, Piotr Dura, Hubert Siuzdak, Peter M. C. Harrison, Elisabeth André, Nori Jacoby
INTERSPEECH1
2022 WavThruVec: Latent speech representation as intermediate features for neural speech synthesis
abstract
Recent advances in neural text-to-speech research have been dominated by two-stage pipelines utilizing low-level intermediate speech representation such as mel-spectrograms. However, such predetermined features are fundamentally limited, because they do not allow to exploit the full potential of a data-driven approach through learning hidden representations. For this reason, several end-to-end methods have been proposed. However, such models are harder to train and require a large number of high-quality recordings with transcriptions. Here, we propose WavThruVec - a two-stage architecture that resolves the bottleneck by using high-dimensional Wav2Vec 2.0 embeddings as intermediate speech representation. Since these hidden activations provide high-level linguistic features, they are more robust to noise. That allows us to utilize annotated speech datasets of a lower quality to train the first-stage module. At the same time, the second-stage component can be trained on large-scale untranscribed audio corpora, as Wav2Vec 2.0 embeddings are already time-aligned. This results in an increased generalization capability to out-of-vocabulary words, as well as to a better generalization to unseen speakers. We show that the proposed model not only matches the quality of state-of-the-art neural models, but also presents useful properties enabling tasks like voice conversion or zero-shot synthesis.
Hubert Siuzdak, Piotr Dura, Pol van Rijn, Nori Jacoby
INTERSPEECH3
2021 Exploring Emotional Prototypes in a High Dimensional TTS Latent Space
abstract
Recent TTS systems are able to generate prosodically varied and realistic speech. However, it is unclear how this prosodic variation contributes to the perception of speakers' emotional states. Here we use the recent psychological paradigm 'Gibbs Sampling with People' to search the prosodic latent space in a trained GST Tacotron model to explore prototypes of emotional prosody. Participants are recruited online and collectively manipulate the latent space of the generative speech model in a sequentially adaptive way so that the stimulus presented to one group of participants is determined by the response of the previous groups. We demonstrate that (1) particular regions of the model's latent space are reliably associated with particular emotions, (2) the resulting emotional prototypes are well-recognized by a separate group of human raters, and (3) these emotional prototypes can be effectively transferred to new sentences. Collectively, these experiments demonstrate a novel approach to the understanding of emotional speech by providing a tool to explore the relation between the latent space of generative models and human semantics.
Pol van Rijn, Silvan Mertes, Dominik Schiller, Peter M. C. Harrison, Pauline Larrouy-Maestri, Elisabeth André, Nori Jacoby
Interspeech1
2021 Analysis by Synthesis: Using an Expressive TTS Model as Feature Extractor for Paralinguistic Speech Classification
abstract
Modeling adequate features of speech prosody is one key factor to good performance in affective speech classification.However, the distinction between the prosody that is induced by 'how' something is said (i.e., affective prosody) and the prosody that is induced by 'what' is being said (i.e., linguistic prosody) is neglected in state-of-the-art feature extraction systems.This results in high variability of the calculated feature values for different sentences that are spoken with the same affective intent, which might negatively impact the performance of the classification.While this distinction between different prosody types is mostly neglected in affective speech recognition, it is explicitly modeled in expressive speech synthesis to create controlled prosodic variation.In this work, we use the expressive Text-To-Speech model Global Style Token Tacotron to extract features for a speech analysis task.We show that the learned prosodic representations outperform state-of-the-art feature extraction systems in the exemplary use case of Escalation Level Classification.
Dominik Schiller, Silvan Mertes, Pol van Rijn, Elisabeth André
Interspeech3
2020 Gibbs Sampling with People
abstract
A core problem in cognitive science and machine learning is to understand how humans derive semantic representations from perceptual objects, such as color from an apple, pleasantness from a musical chord, or seriousness from a face. Markov Chain Monte Carlo with People (MCMCP) is a prominent method for studying such representations, in which participants are presented with binary choice trials constructed such that the decisions follow a Markov Chain Monte Carlo acceptance rule. However, while MCMCP has strong asymptotic properties, its binary choice paradigm generates relatively little information per trial, and its local proposal function makes it slow to explore the parameter space and find the modes of the distribution. Here we therefore generalize MCMCP to a continuous-sampling paradigm, where in each iteration the participant uses a slider to continuously manipulate a single stimulus dimension to optimize a given criterion such as ‘pleasantness’. We formulate both methods from a utility-theory perspective, and show that the new method can be interpreted as ‘Gibbs Sampling with People’ (GSP). Further, we introduce an aggregation parameter to the transition step, and show that this parameter can be manipulated to flexibly shift between Gibbs sampling and deterministic optimization. In an initial study, we show GSP clearly outperforming MCMCP; we then show that GSP provides novel and interpretable results in three other domains, namely musical chords, vocal emotions, and faces. We validate these results through large-scale perceptual rating experiments. The final experiments use GSP to navigate the latent space of a state-of-the-art image synthesis network (StyleGAN), a promising approach for applying GSP to high-dimensional perceptual spaces. We conclude by discussing future cognitive applications and ethical implications.
Peter M. C. Harrison, Raja Marjieh, Federico Adolfi, Pol van Rijn, Manuel Anglada-Tort, Ofer Tchernichovski, Pauline Larrouy-Maestri, Nori Jacoby
NeurIPS4