EDBT 2026 Demo / reviewers in the wild / expert
Marc-André Carbonneau
dblp:166/3300 · also Marc-Andre Carbonneau
· DBLP profile ↗
21ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-0677-415XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SEREP: Semantic Facial Expression Representation for Robust in-the-Wild Capture and RetargetingabstractMonocular facial performance capture in-the-wild is challenging due to varied capture conditions, face shapes, and expressions. Most current methods rely on linear 3D Morphable Models, which represent facial expressions independently of identity at the vertex displacement level. We propose SEREP (Semantic Expression Representation), a model that disentangles expression from identity at the semantic level. We start by learning an expression representation from high-quality 3D data of unpaired facial expressions. Then, we train a model to predict expression from monocular images relying on a novel semi-supervised scheme using low quality synthetic data. In addition, we introduce MultiREX, a benchmark addressing the lack of evaluation resources for the expression capture task. Our experiments show that SEREP outperforms state-of-the-art methods, capturing challenging expressions and transferring them to new identities. Arthur Josi, Luiz G. Hafemann, Abdallah Dib, Emeline Got, Rafael M. O. Cruz, Marc-André Carbonneau |
ICCV | 6 |
| 2025 | LinearVC: Linear Transformations of Self-Supervised Features Through the Lens of Voice Conversion
Herman Kamper, Benjamin van Niekerk, Julian Zaidi, Marc-André Carbonneau |
INTERSPEECH | 4 |
| 2024 | BinaryAlign: Word Alignment as Binary Sequence LabelingabstractReal world deployments of word alignment are almost certain to cover both high and low resource languages.However, the state-ofthe-art for this task recommends a different model class depending on the availability of gold alignment training data for a particular language pair.We propose BinaryAlign, a novel word alignment technique based on binary sequence labeling that outperforms existing approaches in both scenarios, offering a unifying approach to the task.Additionally, we vary the specific choice of multilingual foundation model, perform stratified error analysis over alignment error type, and explore the performance of BinaryAlign on non-English language pairs.We make our source code publicly available.1 Gaetan Lopez Latouche, Marc-André Carbonneau, Benjamin Swanson |
ACL (1) | 2 |
| 2024 | MoSAR: Monocular Semi-Supervised Model for Avatar Reconstruction using Differentiable ShadingabstractReconstructing an avatar from a portrait image has many applications in multimedia, but remains a challenging research problem. Extracting reflectance maps and geom- etry from one image is ill-posed: recovering geometry is a one-to-many mapping problem and reflectance and light are difficult to disentangle. Accurate geometry and reflectance can be captured under the controlled conditions of a light stage, but it is costly to acquire large datasets in this fash- ion. Moreover, training solely with this type of data leads to poor generalization with in-the-wild images. This moti- vates the introduction of MoSAR, a method for 3D avatar generation from monocular images. We propose a semi- supervised training scheme that improves generalization by learning from both light stage and in-the-wild datasets. This is achieved using a novel differentiable shading formulation. We show that our approach effectively disentangles the intrinsic face parameters, producing relightable avatars. As a result, MoSAR11Project page: https://ubisoft-laforge.github.io/character/mosar estimates a richer set of skin reflectance maps and generates more realistic avatars than existing state-of-the-art methods. We also release a new dataset, that provides intrinsic face attributes (diffuse, specular, am- bient occlusion and translucency maps) for 10k subjects. Abdallah Dib, Luiz G. Hafemann, Emeline Got, Trevor Anderson, Amin Fadaeinejad, Rafael M. O. Cruz, Marc-André Carbonneau |
CVPR | 7 |
| 2024 | UPose3D: Uncertainty-Aware 3D Human Pose Estimation with Cross-view and Temporal Cues
Vandad Davoodnia, Saeed Ghorbani, Marc-André Carbonneau, Alexandre Messier, Ali Etemad |
ECCV (16) | 3 |
| 2024 | Zero-shot Cross-Lingual Transfer for Synthetic Data Generation in Grammatical Error DetectionabstractGrammatical Error Detection (GED) methods rely heavily on human annotated error corpora.However, these annotations are unavailable in many low-resource languages.In this paper, we investigate GED in this context.Leveraging the zero-shot cross-lingual transfer capabilities of multilingual pre-trained language models, we train a model using data from a diverse set of languages to generate synthetic errors in other languages.These synthetic error corpora are then used to train a GED model.Specifically we propose a two-stage fine-tuning pipeline where the GED model is first fine-tuned on multilingual synthetic data from target languages followed by fine-tuning on human-annotated GED corpora from source languages.This approach outperforms current state-of-the-art annotation-free GED methods.We also analyse the errors produced by our method and other strong baselines, finding that our approach produces errors that are more diverse and more similar to human errors. Gaetan Lopez Latouche, Marc-André Carbonneau, Benjamin Swanson |
EMNLP | 2 |
| 2024 | Spoken-Term Discovery using Discrete Speech Units
Benjamin van Niekerk, Julian Zaidi, Marc-André Carbonneau, Herman Kamper |
INTERSPEECH | 3 |
| 2024 | Measuring Disentanglement: A Review of MetricsabstractLearning to disentangle and represent factors of variation in data is an important problem in artificial intelligence. While many advances have been made to learn these representations, it is still unclear how to quantify disentanglement. While several metrics exist, little is known on their implicit assumptions, what they truly measure, and their limits. In consequence, it is difficult to interpret results when comparing different representations. In this work, we survey supervised disentanglement metrics and thoroughly analyze them. We propose a new taxonomy in which all metrics fall into one of the three families: intervention-based, predictor-based, and information-based. We conduct extensive experiments in which we isolate properties of disentangled representations, allowing stratified comparison along several axes. From our experiment results and analysis, we provide insights on relations between disentangled representation properties. Finally, we share guidelines on how to measure disentanglement. Marc-André Carbonneau, Julian Zaidi, Jonathan Boilard, Ghyslain Gagnon |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | ZeroEGGS: Zero-shot Example-based Gesture Generation from SpeechabstractAbstract We present ZeroEGGS, a neural network framework for speech‐driven gesture generation with zero‐shot style control by example. This means style can be controlled via only a short example motion clip, even for motion styles unseen during training. Our model uses a Variational framework to learn a style embedding, making it easy to modify style through latent space manipulation or blending and scaling of style embeddings. The probabilistic nature of our framework further enables the generation of a variety of outputs given the input, addressing the stochastic nature of gesture motion. In a series of experiments, we first demonstrate the flexibility and generalizability of our model to new speakers and styles. In a user study, we then show that our model outperforms previous state‐of‐the‐art techniques in naturalness of motion, appropriateness for speech, and style portrayal. Finally, we release a high‐quality dataset of full‐body gesture motion including fingers, with speech, spanning across 19 different styles. Our code and data are publicly available at https://github.com/ubisoft/ubisoft‐laforge‐ZeroEGGS . Saeed Ghorbani, Ylva Ferstl, Daniel Holden, Nikolaus F. Troje, Marc-André Carbonneau |
Comput. Graph. Forum | 5 |
| 2023 | Rhythm Modeling for Voice ConversionabstractVoice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in the perception of speaker identity. To bridge this gap, we introduce Urhythmic—an unsupervised method for rhythm conversion that does not require parallel data or text transcriptions. Using self-supervised representations, we first divide source audio into segments approximating sonorants, obstruents, and silences. Then we model rhythm by estimating speaking rate or the duration distribution of each segment type. Finally, we match the target speaking rate or rhythm by time-stretching the speech segments. Experiments show that Urhythmic outperforms existing unsupervised methods in terms of quality and prosody. Benjamin van Niekerk, Marc-André Carbonneau, Herman Kamper |
IEEE Signal Process. Lett. | 2 |
| 2022 | A Comparison of Discrete and Soft Speech Units for Improved Voice ConversionabstractThe goal of voice conversion is to transform source speech into a target voice, keeping the content unchanged. In this paper, we focus on self-supervised representation learning for voice conversion. Specifically, we compare discrete and soft speech units as input features. We find that discrete representations effectively remove speaker information but discard some linguistic content – leading to mispronunciations. As a solution, we propose soft speech units learned by predicting a distribution over the discrete units. By modeling uncertainty, soft units capture more content information, improving the intelligibility and naturalness of converted speech.12 Benjamin van Niekerk, Marc-André Carbonneau, Julian Zaidi, Matthew Baas, Hugo Seuté, Herman Kamper |
ICASSP | 2 |
| 2022 | Exemplar-based Stylized Gesture Generation from Speech: An Entry to the GENEA Challenge 2022abstractWe present our entry to the GENEA Challenge of 2022 on data-driven co-speech gesture generation. Our system is a neural network that generates gesture animation from an input audio file. The motion style generated by the model is extracted from an exemplar motion clip. Style is embedded in a latent space using a variational framework. This architecture allows for generating in styles unseen during training. Moreover, the probabilistic nature of our variational framework furthermore enables the generation of a variety of outputs given the same input, addressing the stochastic nature of gesture motion. The GENEA challenge evaluation showed that our model produces full-body motion with highly competitive levels of human-likeness. Saeed Ghorbani, Ylva Ferstl, Marc-André Carbonneau |
ICMI | 3 |
| 2022 | Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech SynthesisabstractThis paper presents Daft-Exprt, a multi-speaker acoustic model advancing the state-of-the-art for cross-speaker prosody transfer on any text.This is one of the most challenging, and rarely directly addressed, task in speech synthesis, especially for highly expressive data.Daft-Exprt uses FiLM conditioning layers to strategically inject different prosodic information in all parts of the architecture.The model explicitly encodes traditional low-level prosody features such as pitch, loudness and duration, but also higher level prosodic information that helps generating convincing voices in highly expressive styles.Speaker identity and prosodic information are disentangled through an adversarial training strategy that enables accurate prosody transfer across speakers.Experimental results show that Daft-Exprt significantly outperforms strong baselines on inter-text crossspeaker prosody transfer tasks, while yielding naturalness comparable to state-of-the-art expressive models.Moreover, results indicate that the model discards speaker identity information from the prosody representation, and consistently generate speech with the desired voice.We publicly release our code 1 and provide speech samples from our experiments 2 . Julian Zaidi, Hugo Seuté, Benjamin van Niekerk, Marc-André Carbonneau |
INTERSPEECH | 4 |
| 2021 | Artist guided generation of video game production quality face texturesabstractWe develop a high resolution face texture generation system which uses artist provided appearance controls as the conditions for a generative network. Artists are able to control various elements in the generated textures, such as the skin, eye, lip, and hair color. This is made possible by reparameterizing our dataset to the same UV mapping, allowing us to utilize image-to-image translation networks. Although our dataset is limited in size, only 126 samples in total, our system is still able to generate realistic face textures which strongly adhere to the input appearance attribute conditions because of our training augmentation methods. Once our system has generated the face texture, it is ready to be used in a modern game production environment. Thanks to our novel SuperResolution and material property recovery methods, our generated face textures are 4K resolution and have the associated material property maps required for raytraced rendering. Christian Murphy, Sudhir P. Mudur, Daniel Holden, Marc-André Carbonneau, Donya Ghafourzadeh, Andre Beauchamp |
Comput. Graph. | 4 |
| 2020 | Appearance Controlled Face Texture Generation for Video Game CharactersabstractManually creating realistic, digital human heads is a difficult and time-consuming task for artists. While 3D scanners and photogrammetry allow for quick and automatic reconstruction of heads, finding an actor who fits specific character appearance descriptions can be difficult. Moreover, modern open-world videogames feature several thousands of characters that cannot realistically all be cast and scanned. Therefore, researchers are investigating generative models to create heads fitting a specific character appearance description. While current methods are able to generate believable head shapes quite well, generating a corresponding high-resolution and high-quality texture which respects the character’s appearance description is not possible using current state of the art methods. Christian Murphy, Sudhir P. Mudur, Daniel Holden, Marc-André Carbonneau, Donya Ghafourzadeh, Andre Beauchamp |
MIG | 4 |
| 2020 | Feature Learning from Spectrograms for Assessment of Personality TraitsabstractSeveral methods have recently been proposed to analyze speech and automatically infer the personality of the speaker. These methods often rely on prosodic and other hand crafted speech processing features extracted with off-the-shelf toolboxes. To achieve high accuracy, numerous features are typically extracted using complex and highly parameterized algorithms. In this paper, a new method based on feature learning and spectrogram analysis is proposed to simplify the feature extraction process while maintaining a high level of accuracy. The proposed method learns a dictionary of discriminant features from patches extracted in the spectrogram representations of training speech segments. Each speech segment is then encoded using the dictionary, and the resulting feature set is used to perform classification of personality traits. Experiments indicate that the proposed method achieves state-of-the-art results with an important reduction in complexity when compared to the most recent reference methods. The number of features, and difficulties linked to the feature extraction process are greatly reduced as only one type of descriptors is used, for which the 7 parameters can be tuned automatically. In contrast, the simplest reference method uses 4 types of descriptors to which 6 functionals are applied, resulting in over 20 parameters to be tuned. Marc-André Carbonneau, Eric Granger, Yazid Attabi, Ghyslain Gagnon |
IEEE Trans. Affect. Comput. | 1 |
| 2019 | Bag-Level Aggregation for Multiple-Instance Active Learning in Instance Classification ProblemsabstractA growing number of applications, e.g., video surveillance and medical image analysis, require training recognition systems from large amounts of weakly annotated data, while some targeted interactions with a domain expert are allowed to improve the training process. In such cases, active learning (AL) can reduce labeling costs for training a classifier by querying the expert to provide the labels of most informative instances. This paper focuses on AL methods for instance classification problems in multiple instance learning (MIL), where data are arranged into sets, called bags, which are weakly labeled. Most AL methods focus on single-instance learning problems. These methods are not suitable for MIL problems because they cannot account for the bag structure of data. In this paper, new methods for bag-level aggregation of instance informativeness are proposed for multiple instance AL (MIAL). The aggregated informativeness method identifies the most informative instances based on classifier uncertainty and queries bags incorporating the most information. The other proposed method, called cluster-based aggregative sampling, clusters data hierarchically in the instance space. The informativeness of instances is assessed by considering bag labels, inferred instance labels, and the proportion of labels that remain to be discovered in clusters. Both proposed methods significantly outperform reference methods in extensive experiments using benchmark data from several application domains. Results indicate that using an appropriate strategy to address MIAL problems yields a significant reduction in the number of queries needed to achieve the same level of performance as single-instance AL methods. Marc-André Carbonneau, Eric Granger, Ghyslain Gagnon |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Multiple instance learning: A survey of problem characteristics and applications
Marc-André Carbonneau, Veronika Cheplygina, Eric Granger, Ghyslain Gagnon |
Pattern Recognit. | 1 |
| 2016 | Witness identification in multiple instance learning using random subspacesabstractMultiple instance learning (MIL) is a form of weakly-supervised learning where instances are organized in bags. A label is provided for bags, but not for instances. MIL literature typically focuses on the classification of bags seen as one object, or as a combination of their instances. In both cases, performance is generally measured using labels assigned to entire bags. In this paper, the MIL problem is formulated as a knowledge discovery task for which algorithms seek to discover the witnesses (i.e. identifying positive instances), using the weak supervision provided by bag labels. Some MIL methods are suitable for instance classification, but perform poorly in application where the witness rate is low, or when the positive class distribution is multimodal. A new method that clusters data projected in random subspaces is proposed to perform witness identification in these adverse settings. The proposed method is assessed on MIL data sets from three application domains, and compared to 7 reference MIL algorithms for the witness identification task. The proposed algorithm constantly ranks among the best methods in all experiments, while all other methods perform unevenly across data sets. Marc-André Carbonneau, Eric Granger, Ghyslain Gagnon |
ICPR | 1 |
| 2016 | Robust multiple-instance learning ensembles using random subspace instance selection
Marc-André Carbonneau, Eric Granger, Alexandre J. Raymond, Ghyslain Gagnon |
Pattern Recognit. | 1 |
| 2015 | Real-time visual play-break detection in sport events using a context descriptorabstractThe detection of play and break segments in team sports is an essential step towards the automation of live game capture and broadcast. This paper presents a two-stage hierarchical method for play-break detection in non-edited video feeds of sport events. Unlike most existing methods, this algorithm performs action and event recognition on content, and thus does not rely on production cues of broadcast feeds. Moreover, the method does not require player tracking, can be used in real-time, and can be easily adapted to different sports. In the first stage, bag-of-words event detectors are trained to recognize key events such as line changes, face-offs and preliminary play-breaks. In the second stage, the output of the detectors along with a novel feature based on spatio-temporal interest points are used to create a context descriptor for the final decision. Experiments demonstrate the efficiency of the proposed method on real hockey game footage, achieving 90% accuracy. Marc-André Carbonneau, Alexandre J. Raymond, Eric Granger, Ghyslain Gagnon |
ISCAS | 1 |