VLDB 2026 Research / reviewers in the wild / expert
Claudiu Cristian Musat
dblp:205/9188 · also Claudiu Musat
· DBLP profile ↗
25ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0003-2156-7738ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Inkeraction: An Interaction Modality Powered by Ink Recognition and SynthesisabstractInk is a powerful medium for note-taking and creativity tasks. Multi-touch devices and stylus input have enabled digital ink to be editable and searchable. To extend the capabilities of digital ink, we introduce Inkeraction, an interaction modality powered by ink recognition and synthesis. Inkeraction segments and classifies digital ink objects (e.g., handwriting and sketches), identifies relationships between them, and generates strokes in different writing styles. Inkeraction reshapes the design space for digital ink by enabling features that include: (1) assisting users to manipulate ink objects, (2) providing word-processor features such as spell checking, (3) automating repetitive writing tasks such as transcribing, and (4) bridging with generative models’ features such as brainstorming. Feedback from two user studies with a total of 22 participants demonstrated that Inkeraction supported writing activities by enabling participants to write faster with fewer steps and achieve better writing quality. Rachel Campbell, Peggy Chi, Maria Cirimele, Mike Cleron, Kirsten Climer, Chelsey Fleming, Ashwin Ganti, Philippe Gervais, Pedro Gonnet, Tayeb A Karim, Andrii Maksai, Chris Melancon, Rob Mickle, Claudiu Cristian Musat, Palash Nandy, Xiaoyu Iris Qu, David Robishaw, Angad Singh, Mathangi Venkatesan |
CHI | 15 |
| 2023 | Sampling and Ranking for Digital Ink Generation on a Tight Computational Budget
Andrei Afonin, Andrii Maksai, Aleksandr Timofeev, Claudiu Cristian Musat |
ICDAR (4) | 4 |
| 2023 | Character Queries: A Transformer-Based Approach to On-line Handwritten Character Segmentation
Michael Jungo, Beat Wolf, Andrii Maksai, Claudiu Cristian Musat, Andreas Fischer 0002 |
ICDAR (1) | 4 |
| 2023 | DSS: Synthesizing Long Digital Ink Using Data Augmentation, Style Encoding and Split Generation
Aleksandr Timofeev, Anastasiia Fadeeva, Andrei Afonin, Claudiu Cristian Musat, Andrii Maksai |
ICDAR (4) | 4 |
| 2021 | Multi-Dimensional Explanation of Target Variables from DocumentsabstractAutomated predictions require explanations to be interpretable by humans. Past work used attention and rationale mechanisms to find words that predict the target variable of a document. Often though, they result in a tradeoff between noisy explanations or a drop in accuracy. Furthermore, rationale methods cannot capture the multi-faceted nature of justifications for multiple targets, because of the non-probabilistic nature of the mask. In this paper, we propose the Multi-Target Masker (MTM) to address these shortcomings. The novelty lies in the soft multi-dimensional mask that models a relevance probability distribution over the set of target variables to handle ambiguities. Additionally, two regularizers guide MTM to induce long, meaningful explanations. We evaluate MTM on two datasets and show, using standard metrics and human annotations, that the resulting masks are more accurate and coherent than those generated by the state-of-the-art methods. Moreover, MTM is the first to also achieve the highest F1 scores for all the target variables simultaneously. Diego Antognini, Claudiu Cristian Musat, Boi Faltings |
AAAI | 2 |
| 2021 | Interacting with Explanations through CritiquingabstractUsing personalized explanations to support recommendations has been shown to increase trust and perceived quality. However, to actually obtain better recommendations, there needs to be a means for users to modify the recommendation criteria by interacting with the explanation. We present a novel technique using aspect markers that learns to generate personalized explanations of recommendations from review texts, and we show that human users significantly prefer these explanations over those produced by state-of-the-art techniques. Our work's most important innovation is that it allows users to react to a recommendation by critiquing the textual explanation: removing (symmetrically adding) certain aspects they dislike or that are no longer relevant (symmetrically that are of interest). The system updates its user model and the resulting recommendations according to the critique. This is based on a novel unsupervised critiquing method for single- and multi-step critiquing with textual explanations. Empirical results show that our system achieves good performance in adapting to the preferences expressed in multi-step critiquing and generates consistent explanations. Diego Antognini, Claudiu Cristian Musat, Boi Faltings |
IJCAI | 2 |
| 2021 | Addressing fairness in classification with a model-agnostic multi-objective algorithmabstractThe goal of fairness in classification is to learn a classifier that does not discriminate against groups of individuals based on sensitive attributes, such as race and gender. One approach to designing fair algorithms is to use relaxations of fairness notions as regularization terms or in a constrained optimization problem. We observe that the hyperbolic tangent function can approximate the indicator function. We leverage this property to define a differentiable relaxation that approximates fairness notions provably better than existing relaxations. In addition, we propose a model-agnostic multi-objective architecture that can simultaneously optimize for multiple fairness notions and multiple sensitive attributes and supports all statistical parity-based notions of fairness. We use our relaxation with the multi-objective architecture to learn fair classifiers. Experiments on public datasets show that our method suffers a significantly lower loss of accuracy than current debiasing algorithms relative to the unconstrained model. Kirtan Padh, Diego Antognini, Emma Lejal Glaude, Boi Faltings, Claudiu Cristian Musat |
UAI | 5 |
| 2021 | Benefiting from Bicubically Down-Sampled Images for Learning Real-World Image Super-ResolutionabstractSuper-resolution (SR) has traditionally been based on pairs of high-resolution images (HR) and their low-resolution (LR) counterparts obtained artificially with bicubic downsampling. However, in real-world SR, there is a large variety of realistic image degradations and analytically modeling these realistic degradations can prove quite difficult. In this work, we propose to handle real-world SR by splitting this ill-posed problem into two comparatively more well-posed steps. First, we train a network to transform real LR images to the space of bicubically down-sampled images in a supervised manner, by using both real LR/HR pairs and synthetic pairs. Second, we take a generic SR network trained on bicubically downsampled images to super-resolve the transformed LR image. The first step of the pipeline addresses the problem by registering the large variety of degraded images to a common, well understood space of images. The second step then leverages the already impressive performance of SR on bicubically downsampled images, sidestepping the issues of end-to-end training on datasets with many different image degradations. We demonstrate the effectiveness of our proposed method by comparing it to recent methods in real-world SR and show that our proposed approach outperforms the state-of-the-art works in terms of both qualitative and quantitative results, as well as results of an extensive user study conducted on several real image datasets. Mohammad Saeed Rad, Thomas Yu, Claudiu Cristian Musat, Hazim Kemal Ekenel, Behzad Bozorgtabar, Jean-Philippe Thiran |
WACV | 3 |
| 2020 | Evaluating The Search Phase of Neural Architecture Search
Kaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Cristian Musat, Mathieu Salzmann |
ICLR | 4 |
| 2020 | Automatic Creation of Text Corpora for Low-Resource Languages from the Internet: The Case of Swiss GermanabstractThis paper presents SwissCrawl, the largest Swiss German text corpus to date. Composed of more than half a million sentences, it was generated using a customized web scraping tool that could be applied to other low-resource languages as well. The approach demonstrates how freely available web pages can be used to construct comprehensive text corpora, which are of fundamental importance for natural language processing. In an experimental evaluation, we show that using the new corpus leads to significant improvements for the task of language modeling. Lucy Linder, Michael Jungo, Jean Hennebert, Claudiu Cristian Musat, Andreas Fischer 0002 |
LREC | 4 |
| 2020 | A Swiss German Dictionary: Variation in Speech and WritingabstractWe introduce a dictionary containing normalized forms of common words in various Swiss German dialects into High German. As Swiss German is, for now, a predominantly spoken language, there is a significant variation in the written forms, even between speakers of the same dialect. To alleviate the uncertainty associated with this diversity, we complement the pairs of Swiss German - High German words with the Swiss German phonetic transcriptions (SAMPA). This dictionary becomes thus the first resource to combine large-scale spontaneous translation with phonetic transcriptions. Moreover, we control for the regional distribution and insure the equal representation of the major Swiss dialects. The coupling of the phonetic and written Swiss German forms is powerful. We show that they are sufficient to train a Transformer-based phoneme to grapheme model that generates credible novel Swiss German writings. In addition, we show that the inverse mapping - from graphemes to phonemes - can be modeled with a transformer trained with the novel dictionary. This generation of pronunciations for previously unknown words is key in training extensible automated speech recognition (ASR) systems, which are key beneficiaries of this dictionary. Larissa Schmidt, Lucy Linder, Sandra Djambazovska, Alexandros Lazaridis, Tanja Samardzic, Claudiu Cristian Musat |
LREC | 6 |
| 2020 | Benefiting from multitask learning to improve single image super-resolution
Mohammad Saeed Rad, Behzad Bozorgtabar, Claudiu Cristian Musat, Urs-Viktor Marti, Max Basler, Hazim Kemal Ekenel, Jean-Philippe Thiran |
Neurocomputing | 3 |
| 2019 | Alleviating Sequence Information Loss with Data Overlapping and Prime Batch SizesabstractIn sequence modeling tasks the token order matters, but this information can be partially lost due to the discretization of the sequence into data points. In this paper, we study the imbalance between the way certain token pairs are included in data points and others are not. We denote this a token order imbalance (TOI) and we link the partial sequence information loss to a diminished performance of the system as a whole, both in text and speech processing tasks. We then provide a mechanism to leverage the full token order information—Alleviated TOI—by iteratively overlapping the token composition of data points. For recurrent networks, we use prime numbers for the batch size to avoid redundancies when building batches from overlapped data points. The proposed method achieved state of the art performance in both text and speech related tasks. Noémien Kocher, Christian Scuito, Lorenzo Tarantino, Alexandros Lazaridis, Andreas Fischer 0002, Claudiu Cristian Musat |
CoNLL | 6 |
| 2019 | Overcoming Multi-model ForgettingabstractWe identify a phenomenon, which we refer to as multi-model forgetting, that occurs when sequentially training multiple deep networks with partially-shared parameters; the performance of previously-trained models degrades as one optimizes a subsequent one, due to the overwriting of shared parameters. To overcome this, we introduce a statistically-justified weight plasticity loss that regularizes the learning of a model’s shared parameters according to their importance for the previous models, and demonstrate its effectiveness when training two models sequentially and for neural architecture search. Adding weight plasticity in neural architecture search preserves the best models to the end of the search and yields improved results in both natural language processing and computer vision tasks. Yassine Benyahia, Kaicheng Yu, Kamil Bennani-Smires, Martin Jaggi, Anthony C. Davison, Mathieu Salzmann, Claudiu Cristian Musat |
ICML | 7 |
| 2018 | Churn Intent Detection in Multilingual Chatbot Conversations and Social MediaabstractChristian Abbet, Meryem M’hamdi, Athanasios Giannakopoulos, Robert West, Andreea Hossmann, Michael Baeriswyl, Claudiu Musat. Proceedings of the 22nd Conference on Computational Natural Language Learning. 2018. Christian Abbet, Meryem M'hamdi, Athanasios Giannakopoulos, Andreea Hossmann, Michael Baeriswyl, Claudiu Cristian Musat |
CoNLL | 7 |
| 2018 | Simple Unsupervised Keyphrase Extraction using Sentence EmbeddingsabstractKeyphrase extraction is the task of automatically selecting a small set of phrases that best describe a given free text document.Supervised keyphrase extraction requires large amounts of labeled training data and generalizes very poorly outside the domain of the training data.At the same time, unsupervised systems have poor accuracy, and often do not generalize well, as they require the input document to belong to a larger corpus also given as input.Addressing these drawbacks, in this paper, we tackle keyphrase extraction from single documents with EmbedRank: a novel unsupervised method, that leverages sentence embeddings.EmbedRank achieves higher F-scores than graph-based state of the art systems on standard datasets and is suitable for real-time processing of large amounts of Web data.With EmbedRank, we also explicitly increase coverage and diversity among the selected keyphrases by introducing an embedding-based maximal marginal relevance (MMR) for new phrases.A user study including over 200 votes showed that, although reducing the phrases' semantic overlap leads to no gains in F-score, our high diversity selection is preferred by humans. Kamil Bennani-Smires, Claudiu Cristian Musat, Andreea Hossmann, Michael Baeriswyl, Martin Jaggi |
CoNLL | 2 |
| 2018 | Submodularity-Inspired Data Selection for Goal-Oriented Chatbot Training Based on Sentence EmbeddingsabstractSpoken language understanding (SLU) systems, such as goal-oriented chatbots or personal assistants, rely on an initial natural language understanding (NLU) module to determine the intent and to extract the relevant information from the user queries they take as input. SLU systems usually help users to solve problems in relatively narrow domains and require a large amount of in-domain training data. This leads to significant data availability issues that inhibit the development of successful systems. To alleviate this problem, we propose a technique of data selection in the low-data regime that enables us to train with fewer labeled sentences, thus smaller labelling costs. We propose a submodularity-inspired data ranking function, the ratio-penalty marginal gain, for selecting data points to label based only on the information extracted from the textual embedding space. We show that the distances in the embedding space are a viable source of information that can be used for data selection. Our method outperforms two known active learning techniques and enables cost-efficient training of the NLU unit. Moreover, our proposed selection technique does not need the model to be retrained in between the selection steps, making it time efficient as well. Mladen Dimovski, Claudiu Cristian Musat, Vladimir Ilievski, Andreea Hossmann, Michael Baeriswyl |
IJCAI | 2 |
| 2018 | Goal-Oriented Chatbot Dialog Management Bootstrapping with Transfer LearningabstractGoal-Oriented (GO) Dialogue Systems, colloquially known as goal oriented chatbots, help users achieve a predefined goal (e.g. book a movie ticket) within a closed domain. A first step is to understand the user's goal by using natural language understanding techniques. Once the goal is known, the bot must manage a dialogue to achieve that goal, which is conducted with respect to a learnt policy. The success of the dialogue system depends on the quality of the policy, which is in turn reliant on the availability of high-quality training data for the policy learning method, for instance Deep Reinforcement Learning. Due to the domain specificity, the amount of available data is typically too low to allow the training of good dialogue policies. In this paper we introduce a transfer learning method to mitigate the effects of the low in-domain data availability. Our transfer learning based approach improves the bot's success rate by 20% in relative terms for distant domains and we more than double it for close domains, compared to the model without transfer learning. Moreover, the transfer learning chatbots learn the policy up to 5 to 10 times faster. Finally, as the transfer learning approach is complementary to additional processing such as warm-starting, we show that their joint application gives the best outcomes. Vladimir Ilievski, Claudiu Cristian Musat, Andreea Hossmann, Michael Baeriswyl |
IJCAI | 2 |
| 2018 | Machine Translation of Low-Resource Spoken Dialects: Strategies for Normalizing Swiss German
Pierre-Edouard Honnet, Andrei Popescu-Belis, Claudiu Cristian Musat, Michael Baeriswyl |
LREC | 3 |
| 2015 | Personalizing Product Rankings Using Collaborative Filtering on Opinion-Derived Topic Profiles
Claudiu Cristian Musat, Boi Faltings |
IJCAI | 1 |
| 2014 | Acquiring Commonsense Knowledge for Sentiment Analysis through Human ComputationabstractMany Artificial Intelligence tasks need large amounts of commonsense knowledge. Because obtaining this knowledge through machine learning would require a huge amount of data, a better alternative is to elicit it from people through human computation. We consider the sentiment classification task, where knowledge about the contexts that impact word polarities is crucial, but hard to acquire from data. We describe a novel task design that allows us to crowdsource this knowledge through Amazon Mechanical Turk with high quality. We show that the commonsense knowledge acquired in this way dramatically improves the performance of established sentiment classification methods. Marina Boia, Claudiu Cristian Musat, Boi Faltings |
AAAI | 2 |
| 2014 | Constructing Context-Aware Sentiment Lexicons with an Asynchronous Game with a Purpose
Marina Boia, Claudiu Cristian Musat, Boi Faltings |
CICLing (2) | 2 |
| 2014 | EmotionWatch: Visualizing Fine-Grained Emotions in Event-Related Tweets
Renato Kempter, Valentina Sintsova, Claudiu Cristian Musat, Pearl Pu |
ICWSM | 3 |
| 2013 | Recommendation Using Textual Opinions
Claudiu Cristian Musat, Yizhong Liang, Boi Faltings |
IJCAI | 1 |
| 2011 | Improving Topic Evaluation Using Conceptual KnowledgeabstractInternational audience Claudiu Cristian Musat, Julien Velcin, Stefan Trausan-Matu, Marian-Andrei Rizoiu |
IJCAI | 1 |