EDBT 2026 Demo / reviewers in the wild / expert
Mihai Dascalu
dblp:67/2120
· DBLP profile ↗
86ranked-venue papers
18as first author
27since 2021 · last 2026
0000-0002-4815-9227ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 63 · 14 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 55 · 13 first-author · 12 since 2021Artificial intelligence and machine learning · 23 · 4 first-author · 13 since 2021Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RO-ABSA: A Romanian Dataset and Baselines for Aspect-Based Sentiment Analysis
Andreea Alina Gheorghe, Claudia Andrei, Elena Ionescu, Stefan Ruseti, Mihai Dascalu |
LREC | 5 |
| 2026 | Automated Extraction of Answer Candidates for Question Generation
Claudia Preda, Mihai Dascalu, Stefan Ruseti, Danielle S. McNamara |
LREC | 2 |
| 2025 | YMCQ: Reasoning-Enhanced MCQ Generation
Andreea-Nicoleta Dutulescu, Stefan Ruseti, Denis Iorga, Mihai Dascalu, Danielle S. McNamara |
AIED (6) | 4 |
| 2025 | One Model to Score Them All: Unified Scoring of Learning Strategies with LLMs
Andreea-Nicoleta Dutulescu, Stefan Ruseti, Mihai Dascalu, Danielle S. McNamara |
EDM | 3 |
| 2025 | The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language ModelsabstractDespite their remarkable progress across diverse domains, Large Language Models (LLMs) consistently fail at simple characterlevel tasks, such as counting letters in words, due to a fundamental limitation: tokenization.In this work, we frame this limitation as a problem of low mutual information and analyze it in terms of concept emergence.Using a suite of 19 synthetic tasks that isolate character-level reasoning in a controlled setting, we show that such capabilities emerge suddenly and only late in training.We find that percolation-based models of concept emergence explain these patterns, suggesting that learning character composition is not fundamentally different from learning commonsense knowledge.To address this bottleneck, we propose a lightweight architectural modification that significantly improves character-level reasoning while preserving the inductive advantages of subword models.Together, our results bridge low-level perceptual gaps in tokenized LMs and provide a principled framework for understanding and mitigating their structural blind spots.We make our code publicly available. Adrian Cosma, Stefan Ruseti, Emilian Radoi, Mihai Dascalu |
EMNLP | 4 |
| 2025 | Code Similarity Detection Using Complexity-Based Birthmarks
Rares Folea, Mihai Dascalu, Emil Slusanschi |
ICCCI (1) | 2 |
| 2025 | Are LLMs Really Underperforming in Stance Detection? Identifying Patterns of Challenging Instances in Zero-Shot Stance DetectionabstractStance Detection (SD) is gaining broad adoption in social media analytics, as it is a refined form of sentiment analysis that focuses on determining whether a text expresses a position in favor of, against, or neutral with a given target. Despite significant progress, current SD approaches still suffer from instructions and labeling inconsistencies, especially for the neutral class, and the influence of annotator bias. In this paper, we conduct a comprehensive analysis of these challenges by combining training dynamics with instruction-based reflective reasoning using Large Language Models. This hybrid approach allows us to characterize a set of difficulties that pose challenges during the annotation process, and also to identify a subset of highly probable mislabeled instances. Our results highlight frequent disagreements between model predictions and crowdsourced labels, particularly on examples involving semantic ambiguity, target misinterpretation and mixed sentiment, along with flawed annotation instructions. To address these issues, we propose strategies that include confidence-guided data filtering and reasoning-based clustering, and generate natural language explanations that clarify the sources of confusion related to either the text or the target. This study presents a novel perspective on diagnosing annotation errors and semantic difficulty in zero-shot stance detection and opens promising directions for integrating LLM reasoning into dataset curation and evaluation workflows. We release the results and the corresponding code as open-source: https://anonymous.4open.science/r/stance-under-the-scope-2643. Alina Gheorghe, Stefan Ruseti, Mihai Dascalu, Cornelia Caragea |
ICTAI | 3 |
| 2024 | Beyond the Obvious Multi-choice Options: Introducing a Toolkit for Distractor Generation Enhanced with NLI Filtering
Andreea-Nicoleta Dutulescu, Stefan Ruseti, Denis Iorga, Mihai Dascalu, Danielle S. McNamara |
AIED (2) | 4 |
| 2024 | Towards Building the LEMI Readability Platform for Children's Literature in the Romanian LanguageabstractReadability is a crucial characteristic of texts, greatly influencing comprehension and reading efficacy. Unfortunately, limited research is available for less-resourced languages, especially for young populations where its impact is even higher. This paper introduces a new readability tool for children’s literature in the Romanian language, explicitly targeting primary school students aged 7-11. The tool consists of a digital repository of school reading texts (self-compiled corpus) and a text analysis interface that generates automatic readability reports for uploaded short texts. The methodology involves extracting, testing, and calibrating a readability formula for Romanian using the children’s literature corpus. Related work on readability and readability tools is discussed, followed by a description of the children’s literature corpus and the platform functionalities. The first steps are presented towards validating the readability formula for children’s literature in Romanian using the ReaderBench framework, while calibration variables relevant to the Romanian language and children’s literature are examined. Currently, no existing platform integrates a research-based readability formula for the Romanian language, making this tool unique. Overall, this research contributes to applied corpus linguistics and Digital Humanities studies and offers a valuable resource for educators, parents, and children in accessing age-appropriate and readable texts. Madalina Chitez, Mihai Dascalu, Aura Cristina Udrea, Cosmin Striletchi, Karla Csürös, Roxana Rogobete, Alexandru Oravitan |
LREC/COLING | 2 |
| 2024 | How Hard can this Question be? An Exploratory Analysis of Features Assessing Question Difficulty using LLMs
Andreea-Nicoleta Dutulescu, Stefan Ruseti, Mihai Dascalu, Danielle S. McNamara |
EDM | 3 |
| 2024 | How Hard is this Test Set? NLI Characterization by Exploiting Training DynamicsabstractNatural Language Inference (NLI) evaluation is crucial for assessing language understanding models; however, popular datasets suffer from systematic spurious correlations that artificially inflate actual model performance.To address this, we propose a method for the automated creation of a challenging test set without relying on the manual construction of artificial and unrealistic examples.We categorize the test set of popular NLI datasets into three difficulty levels by leveraging methods that exploit training dynamics.This categorization significantly reduces spurious correlation measures, with examples labeled as having the highest difficulty showing markedly decreased performance and encompassing more realistic and diverse linguistic phenomena.When our characterization method is applied to the training set, models trained with only a fraction of the data achieve comparable performance to those trained on the full dataset, surpassing other dataset characterization techniques.Our research addresses limitations in NLI dataset construction, providing a more authentic evaluation of model performance with implications for diverse NLU applications. Adrian Cosma, Stefan Ruseti, Mihai Dascalu, Cornelia Caragea |
EMNLP | 3 |
| 2023 | TA-DA: Topic-Aware Domain Adaptation for Scientific Keyphrase Identification and Classification (Student Abstract)abstractKeyphrase identification and classification is a Natural Language Processing and Information Retrieval task that involves extracting relevant groups of words from a given text related to the main topic. In this work, we focus on extracting keyphrases from scientific documents. We introduce TA-DA, a Topic-Aware Domain Adaptation framework for keyphrase extraction that integrates Multi-Task Learning with Adversarial Training and Domain Adaptation. Our approach improves performance over baseline models by up to 5% in the exact match of the F1-score. Razvan-Alexandru Smadu, George-Eduard Zaharia, Andrei-Marius Avram, Dumitru-Clementin Cercel, Mihai Dascalu, Florin Pop |
AAAI | 5 |
| 2023 | The Automated Model of Comprehension Version 3.0: Paying Attention to Context
Dragos Corlatescu, Micah Watanabe, Stefan Ruseti, Mihai Dascalu, Danielle S. McNamara |
AIED | 4 |
| 2022 | Domain Adaptation in Multilingual and Multi-Domain Monolingual Settings for Complex Word IdentificationabstractComplex word identification (CWI) is a cornerstone process towards proper text simplification.CWI is highly dependent on context, whereas its difficulty is augmented by the scarcity of available datasets which vary greatly in terms of domains and languages.As such, it becomes increasingly more difficult to develop a robust model that generalizes across a wide array of input examples.In this paper, we propose a novel training technique for the CWI task based on domain adaptation to improve the target character and context representations.This technique addresses the problem of working with multiple domains, inasmuch as it creates a way of smoothing the differences between the explored datasets.Moreover, we also propose a similar auxiliary task, namely text simplification, that can be used to complement lexical complexity prediction.Our model obtains a boost of up to 2.42% in terms of Pearson Correlation Coefficients in contrast to vanilla training techniques, when considering the CompLex from the Lexical Complexity Prediction 2021 dataset.At the same time, we obtain an increase of 3% in Pearson scores, while considering a cross-lingual setup relying on the Complex Word Identification 2018 dataset.In addition, our model yields state-ofthe-art results in terms of Mean Absolute Error. George-Eduard Zaharia, Razvan-Alexandru Smadu, Dumitru-Clementin Cercel, Mihai Dascalu |
ACL (1) | 4 |
| 2022 | Multitask Summary Scoring with Longformers
Robert-Mihai Botarleanu, Mihai Dascalu, Laura K. Allen, Scott A. Crossley, Danielle S. McNamara |
AIED (1) | 2 |
| 2022 | "A fork is a food stabber": Linguistic creativity in English L1 and L2 speakers
Stephen Cameron Skalicky, Nancy Bell, Mihai Dascalu, Scott A. Crossley |
CogSci | 3 |
| 2022 | Selfit v2 - Challenges Encountered in Building a Psychomotor Intelligent Tutoring System
Laurentiu-Marian Neagu, Eric Rigaud, Vincent Guarnieri, Mihai Dascalu, Sébastien Travadel |
ITS | 4 |
| 2022 | Where are the Large N Studies in Education?: Introducing a Dataset of Scientific Articles and NLP TechniquesabstractResearch, especially in Education, is hampered by scale and our aim is to help shift this dynamic by tracking relevant studies from major scientific venues. The current version of our dataset considers top conferences and journals from the domain (N = 33,531). Several heuristics using advanced text searches and regular expressions, together with NLP techniques ranging from part of speech tagging, syntactic dependency parsing, to semantic models, are employed to extract relevant studies. Our method achieves an F1 score of .633 when employed on a manually annotated subset of 1,000 articles. When applied to the entire dataset, the total number of articles with large N was around 10%, with a positive trend in the last years. [email protected] was by far the venue with the highest density of identified articles, thus arguing for its emphasis on scale. Further filtering of the articles is required when focusing only on learning, as the articles span across multiple domains and have multiple interests. Nonetheless, the order of magnitude raises current problems, namely that these studies are scarce and future endeavors should emphasize the importance of scale. Our dataset, including the validation subset, and the corresponding code for data crawling, PDF processing, and large N extraction mechanisms, have been open-sourced to further support the initiative and stimulate new analyses. Dragos Corlatescu, Stefan Ruseti, Irina Toma, Mihai Dascalu |
L@S | 4 |
| 2022 | Distilling the Knowledge of Romanian BERTs Using Multiple TeachersabstractRunning large-scale pre-trained language models in computationally constrained environments remains a challenging problem yet to be addressed, while transfer learning from these models has become prevalent in Natural Language Processing tasks. Several solutions, including knowledge distillation, network quantization, or network pruning have been previously proposed; however, these approaches focus mostly on the English language, thus widening the gap when considering low-resource languages. In this work, we introduce three light and fast versions of distilled BERT models for the Romanian language: Distil-BERT-base-ro, Distil-RoBERT-base, and DistilMulti-BERT-base-ro. The first two models resulted from the individual distillation of knowledge from two base versions of Romanian BERTs available in literature, while the last one was obtained by distilling their ensemble. To our knowledge, this is the first attempt to create publicly available Romanian distilled BERT models, which were thoroughly evaluated on five tasks: part-of-speech tagging, named entity recognition, sentiment analysis, semantic textual similarity, and dialect identification. Our experimental results argue that the three distilled models offer performance comparable to their teachers, while being twice as fast on a GPU and ~35% smaller. In addition, we further test the similarity between the predictions of our students versus their teachers by measuring their label and probability loyalty, together with regression loyalty - a new metric introduced in this work. Andrei-Marius Avram, Darius Catrina, Dumitru-Clementin Cercel, Mihai Dascalu, Traian Rebedea, Vasile Florian Pais, Dan Tufis |
LREC | 4 |
| 2021 | Multilingual Age of Exposure
Robert-Mihai Botarleanu, Mihai Dascalu, Micah Watanabe, Danielle S. McNamara, Scott A. Crossley |
AIED (1) | 2 |
| 2021 | Automated Model of Comprehension V2.0
Dragos Corlatescu, Mihai Dascalu, Danielle S. McNamara |
AIED (2) | 2 |
| 2021 | Extracting and Clustering Main Ideas from Student Feedback Using Language Models
Mihai Masala, Stefan Ruseti, Mihai Dascalu, Ciprian Dobre |
AIED (1) | 3 |
| 2021 | Exploring Dialogism Using Language Models
Stefan Ruseti, Maria-Dorinela Dascalu, Dragos Corlatescu, Mihai Dascalu, Stefan Trausan-Matu, Danielle S. McNamara |
AIED (2) | 4 |
| 2021 | RoGPT2: Romanian GPT2 for Text GenerationabstractText generation is one of the most important and challenging tasks in NLP, where models have shown a significant performance increase in recent years. However, most generative models are available only for English, whereas low-resource languages like Romanian have no available alternatives. As such, we introduce RoGPT2, a Romanian version of the GPT2 model, trained on the largest corpus available for the Romanian language. Three versions of the model were trained, namely base (124M parameters), medium (354M parameters), and large (774M parameters). Six tasks from the LiRo benchmark were selected to test the performance and limitations of our encoder versus BERT-Base models for Romanian (RoBERT, BERT-ro-base, and RoDiBERT). RoGPT2 manages to achieve similar or even better performance, except for the task of zero-shot learning cross-lingual question answering. RoGPT2 also obtains state-of-the-art results for grammar error correction (RoGEC) using the RONACC corpus, thus arguing for the model’s capability to generate grammatically correct text (F0.5= 69.01). In addition, we introduce two use cases in which we showcase the different versions and explore the extent to which RoGPT2 is able to continue Romanian news articles. After fine-tuning, the model generated rather long text which accounts for the context of the news. Mihai Alexandru Niculescu, Stefan Ruseti, Mihai Dascalu |
ICTAI | 3 |
| 2021 | Automated Summary Scoring with ReaderBench
Robert-Mihai Botarleanu, Mihai Dascalu, Laura K. Allen, Scott A. Crossley, Danielle S. McNamara |
ITS | 2 |
| 2021 | Selfit - An Intelligent Tutoring System for Psychomotor Development
Laurentiu-Marian Neagu, Eric Rigaud, Vincent Guarnieri, Sébastien Travadel, Mihai Dascalu |
ITS | 5 |
| 2021 | Automated Paraphrase Quality Assessment Using Recurrent Neural Networks and Language Models
Bogdan Nicula, Mihai Dascalu, Natalie Newton, Ellen Orcutt, Danielle S. McNamara |
ITS | 2 |
| 2020 | Sequence-to-Sequence Models for Automated Text Simplification
Robert-Mihai Botarleanu, Mihai Dascalu, Scott A. Crossley, Danielle S. McNamara |
AIED (2) | 2 |
| 2020 | Multi-document Cohesion Network Analysis: Visualizing Intratextual and Intertextual Links
Maria-Dorinela Dascalu, Stefan Ruseti, Mihai Dascalu, Danielle S. McNamara, Stefan Trausan-Matu |
AIED (2) | 3 |
| 2020 | Extended Multi-document Cohesion Network Analysis Centered on Comprehension Prediction
Bogdan Nicula, Cecile A. Perret, Mihai Dascalu, Danielle S. McNamara |
AIED (2) | 3 |
| 2020 | RoBERT - A Romanian BERT ModelabstractDeep pre-trained language models tend to become ubiquitous in the field of Natural Language Processing (NLP).These models learn contextualized representations by using a huge amount of unlabeled text data and obtain state of the art results on a multitude of NLP tasks, by enabling efficient transfer learning.For other languages besides English, there are limited options of such models, most of which are trained only on multi-lingual corpora.In this paper we introduce a Romanian-only pre-trained BERT model -RoBERT -and compare it with different multilingual models on seven Romanian specific NLP tasks grouped into three categories, namely: sentiment analysis, dialect and cross-dialect topic identification, and diacritics restoration.Our model surpasses the multi-lingual models, as well as a another mono-lingual implementation of BERT, on all tasks. Mihai Masala, Stefan Ruseti, Mihai Dascalu |
COLING | 3 |
| 2020 | Neural Grammatical Error Correction for RomanianabstractResources for Grammatical Error Correction (GEC) in non-English languages are scarce, while available spellcheckers in these languages are mostly limited to simple corrections and rules. In this paper we introduce a first GEC corpus for Romanian consisting of 10k pairs of sentences. In addition, the German version of ERRANT (ERRor ANnotation Toolkit) scorer was adapted for Romanian to analyze this corpus and extract edits needed for evaluation. Multiple neural models were experimented, together with pretraining strategies, which proved effective for GEC in low-resource settings. Our baseline consists of a small Transformer model trained only on the GEC dataset ( F0.5=44.38), whereas the best performing model is produced by pretraining a larger Transformer model on artificially generated data, followed by finetuning on the actual corpus ( F0.5=53.76). The proposed method for generating additional training examples is easily extensible and can be applied to any language, as it requires only a POS tagger. Teodor-Mihai Cotet, Stefan Ruseti, Mihai Dascalu |
ICTAI | 3 |
| 2020 | Multi-document Cohesion Network Analysis: Automated Prediction of Inferencing across Multiple DocumentsabstractOpen-ended comprehension questions are a common type of assessment used to evaluate how well students understand one of multiple documents. Our aim is to use natural language processing (NLP) to infer the level and type of inferencing within readers' answers to comprehension questions using the linguistic and semantic features within their responses. Our taxonomy considers three types of responses to comprehension questions from students ( N=146) who read four documents: a) textbase responses (i.e., information required for the answer is present in a contiguous short sequence of text); b) single-document inference responses (i.e., requiring information from multiple text segments in a single document); and c) multi-document inference responses (i.e., information spanning multiple documents is required). The classification task was approached in two ways. First, we extracted features from students' answers to the comprehension questions using linguistic and semantic indices related to textual complexity and an extended Cohesion Network Analysis (CNA) graph to assess semantic links between the answers and the reference documents. Second, we compared different Recurrent Neural Networks (RNNs) architectures that rely on word embeddings to encode both answers and reference documents. Our best model based on RNNs predicts the answer type with an accuracy of 81%. Bogdan Nicula, Cecile A. Perret, Mihai Dascalu, Danielle S. McNamara |
ICTAI | 3 |
| 2020 | Cross-Lingual Transfer Learning for Complex Word IdentificationabstractComplex Word Identification (CWI) is a task centered on detecting hard-to-understand words, or groups of words, in texts from different areas of expertise. The purpose of CWI is to highlight problematic structures that non-native speakers would usually find difficult to understand. Our approach uses zero-shot, one-shot, and few-shot learning techniques, alongside state-of-the-art solutions for Natural Language Processing (NLP) tasks (i.e., Transformers). Our aim is to provide evidence that the proposed models can learn the characteristics of complex words in a multilingual environment by relying on the CWI shared task 2018 dataset available for four different languages (i.e., English, German, Spanish, and also French). Our approach surpasses state-of-the-art cross-lingual results in terms of macro F1-score on English (0.774), German (0.782), and Spanish (0.734) languages, for the zero-shot learning scenario. At the same time, our model also outperforms the state-of-the-art monolingual result for German (0.795 macro F1-score). George-Eduard Zaharia, Dumitru-Clementin Cercel, Mihai Dascalu |
ICTAI | 3 |
| 2020 | Cohesion Network Analysis: Predicting Course Grades and Generating Sociograms for a Romanian Moodle Course
Maria-Dorinela Dascalu, Mihai Dascalu, Stefan Ruseti, Mihai Carabas, Stefan Trausan-Matu, Danielle S. McNamara |
ITS | 2 |
| 2020 | Intelligent Tutoring Systems for Psychomotor Training - A Systematic Literature Review
Laurentiu-Marian Neagu, Eric Rigaud, Sébastien Travadel, Mihai Dascalu, Razvan Rughinis |
ITS | 4 |
| 2019 | Predicting Multi-document Comprehension: Cohesion Network Analysis
Bogdan Nicula, Cecile A. Perret, Mihai Dascalu, Danielle S. McNamara |
AIED (1) | 3 |
| 2019 | Semantic Matching of Open Texts to Pre-scripted Answers in Dialogue-Based Learning
Stefan Ruseti, Raja Lala, Gabriel Gutu, Mihai Dascalu, Johan Jeuring, Marcell van Geest |
AIED (2) | 4 |
| 2019 | Modeling Collaboration in Online Conversations Using Time Series Analysis and Dialogism
Robert-Florian Samoilescu, Mihai Dascalu, Maria-Dorinela Sirbu, Stefan Trausan-Matu, Scott A. Crossley |
AIED (1) | 2 |
| 2019 | Automated Scoring of Self-explanations Using Recurrent Neural Networks
Marilena Panaite, Stefan Ruseti, Mihai Dascalu, Renu Balyan, Danielle S. McNamara, Stefan Trausan-Matu |
EC-TEL | 3 |
| 2019 | ReadME - Your Personal Writing Assistant
Irina Toma, Teodor-Mihai Cotet, Mihai Dascalu, Stefan Trausan-Matu |
EC-TEL | 3 |
| 2018 | Modeling Math Success Using Cohesion Network Analysis
Scott A. Crossley, Maria-Dorinela Sirbu, Mihai Dascalu, Tiffany Barnes, Collin F. Lynch, Danielle S. McNamara |
AIED (2) | 3 |
| 2018 | Identifying Implicit Links in CSCL Chats Using String Kernels and Neural Networks
Mihai Masala, Stefan Ruseti, Gabriel Gutu, Traian Rebedea, Mihai Dascalu, Stefan Trausan-Matu |
AIED (2) | 5 |
| 2018 | Bring It on! Challenges Encountered While Building a Comprehensive Tutoring System Using ReaderBench
Marilena Panaite, Mihai Dascalu, Amy M. Johnson, Renu Balyan, Jianmin Dai, Danielle S. McNamara, Stefan Trausan-Matu |
AIED (1) | 2 |
| 2018 | Predicting Question Quality Using Recurrent Neural Networks
Stefan Ruseti, Mihai Dascalu, Amy M. Johnson, Renu Balyan, Kristopher J. Kopp, Danielle S. McNamara, Scott A. Crossley, Stefan Trausan-Matu |
AIED (1) | 2 |
| 2018 | Exploring Online Course Sociograms Using Cohesion Network Analysis
Maria-Dorinela Sirbu, Mihai Dascalu, Scott A. Crossley, Danielle S. McNamara, Tiffany Barnes, Collin F. Lynch, Stefan Trausan-Matu |
AIED (2) | 2 |
| 2018 | Towards an Automated Model of Comprehension (AMoC)
Mihai Dascalu, Ionut Cristian Paraschiv, Danielle S. McNamara, Stefan Trausan-Matu |
EC-TEL | 1 |
| 2018 | Cohesion-Centered Analysis of Sociograms for Online Communities and Courses Using ReaderBench
Mihai Dascalu, Maria-Dorinela Sirbu, Gabriel Gutu, Stefan Ruseti, Scott A. Crossley, Stefan Trausan-Matu |
EC-TEL | 1 |
| 2018 | Help Me Understand This Conversation: Methods of Identifying Implicit Links Between CSCL Contributions
Mihai Masala, Stefan Ruseti, Gabriel Gutu, Traian Rebedea, Mihai Dascalu, Stefan Trausan-Matu |
EC-TEL | 5 |
| 2018 | Modeling Math Identity and Math Success through Sentiment Analysis and Linguistic Features
Scott A. Crossley, Jaclyn Ocumpaugh, Matthew J. Labrum, Franklin Bradfield, Mihai Dascalu, Ryan Baker 0001 |
EDM | 5 |
| 2018 | Scoring Summaries Using Recurrent Neural Networks
Stefan Ruseti, Mihai Dascalu, Amy M. Johnson, Danielle S. McNamara, Renu Balyan, Kathryn S. McCarthy, Stefan Trausan-Matu |
ITS | 2 |
| 2017 | Teaching iSTART to Understand Spanish
Mihai Dascalu, Matthew E. Jacovina, Christian M. Soto, Laura K. Allen, Jianmin Dai, Tricia A. Guerrero, Danielle S. McNamara |
AIED | 1 |
| 2017 | ReaderBench Learns Dutch: Building a Comprehensive Automated Essay Scoring System for Dutch Language
Mihai Dascalu, Wim Westera, Stefan Ruseti, Stefan Trausan-Matu, Hub Kurvers |
AIED | 1 |
| 2017 | Modeling Comprehension Processes via Automated Analyses of Dialogism
Mihai Dascalu, Laura K. Allen, Danielle S. McNamara, Stefan Trausan-Matu, Scott A. Crossley |
CogSci | 1 |
| 2017 | How Well Do Student Nurses Write Case Studies? A Cohesion-Centered Textual Complexity Analysis
Mihai Dascalu, Philippe Dessus, Laurent Thuez, Stefan Trausan-Matu |
EC-TEL | 1 |
| 2017 | ReaderBench: A Multi-lingual Framework for Analyzing Text Complexity
Mihai Dascalu, Gabriel Gutu, Stefan Ruseti, Ionut Cristian Paraschiv, Philippe Dessus, Danielle S. McNamara, Scott A. Crossley, Stefan Trausan-Matu |
EC-TEL | 1 |
| 2017 | Mass Customization in Continuing Medical Education: Automated Extraction of E-Learning Topics
Nicolae Nistor, Mihai Dascalu, Gabriel Gutu, Stefan Trausan-Matu, Sunhea Choi, Ashley Haberman-Lawson, Brigitte Angela Brands, Christian Körner, Berthold Koletzko |
EC-TEL | 2 |
| 2017 | Semantic Boggle: A Game for Vocabulary Acquisition
Irina Toma, Cristina-Elena Alexandru, Mihai Dascalu, Philippe Dessus, Stefan Trausan-Matu |
EC-TEL | 3 |
| 2017 | Unlocking the Power of Word2Vec for Identifying Implicit LinksabstractThis paper presents a research on using Word2Vec for determining implicit links in multi-participant Computer-Supported Collaborative Learning chat conversations. Word2Vec is a powerful and one of the newest Natural Language Processing semantic models used for computing text cohesion and similarity between documents. This research considers cohesion scores in terms of the strength of the semantic relations established between two utterances, the higher the score, the stronger the similarity between two utterances. An implicit link is established based on cohesion to the most similar previous utterance, within an imposed window. Three similarity formulas were used to compute the cohesion score: an unnormalized score, a normalized score with distance and Mihalcea's formula. Our corpus of conversations incorporated explicit references provided by authors, which were used for validation. A window of 5 utterances and a 1-minute time frame provided the highest detection rate both for exact matching and matching of a block of continuous utterances belonging to the same speaker. Moreover, the unnormalized score correctly identified the largest number of implicit links. Gabriel Gutu, Mihai Dascalu, Stefan Ruseti, Traian Rebedea, Stefan Trausan-Matu |
ICALT | 2 |
| 2016 | Age of Exposure: A Model of Word LearningabstractTextual complexity is widely used to assess the difficulty of reading materials and writing quality in student essays. At a lexical level, word complexity can represent a building block for creating a comprehensive model of lexical networks that adequately estimates learners’ understanding. In order to best capture how lexical associations are created between related concepts, we propose automated indices of word complexity based on Age of Exposure (AoE). AOE indices computationally model the lexical learning process as a function of a learner's experience with language. This study describes a proof of concept based on the on a large-scale learning corpus (i.e., TASA). The results indicate that AoE indices yield strong associations with human ratings of age of acquisition, word frequency, entropy, and human lexical response latencies providing evidence of convergent validity. Mihai Dascalu, Danielle S. McNamara, Scott A. Crossley, Stefan Trausan-Matu |
AAAI | 1 |
| 2016 | Document Cohesion Flow: Striving towards Coherence
Scott A. Crossley, Mihai Dascalu, Stefan Trausan-Matu, Laura K. Allen, Danielle S. McNamara |
CogSci | 2 |
| 2016 | Combining Taxonomies using Word2vecabstractTaxonomies have gained a broad usage in a variety of fields due to their extensibility, as well as their use for classification and knowledge organization. Of particular interest is the digital document management domain in which their hierarchical structure can be effectively employed in order to organize documents into content-specific categories. Common or standard taxonomies (e.g., the ACM Computing Classification System) contain concepts that are too general for conceptualizing specific knowledge domains. In this paper we introduce a novel automated approach that combines sub-trees from general taxonomies with specialized seed taxonomies by using specific Natural Language Processing techniques. We provide an extensible and generalizable model for combining taxonomies in the practical context of two very large European research projects. Because the manual combination of taxonomies by domain experts is a highly time consuming task, our model measures the semantic relatedness between concept labels in CBOW or skip-gram Word2vec vector spaces. A preliminary quantitative evaluation of the resulting taxonomies is performed after applying a greedy algorithm with incremental thresholds used for matching and combining topic labels. Tobias Eljasik-Swoboda, Matthias L. Hemmje, Mihai Dascalu, Stefan Trausan-Matu |
DocEng | 3 |
| 2016 | Predicting Academic Performance Based on Students' Blog and Microblog Posts
Mihai Dascalu, Elvira Popescu, Alex Becheru, Scott A. Crossley, Stefan Trausan-Matu |
EC-TEL | 1 |
| 2016 | Finding the Needle in a Haystack: Who are the Most Central Authors Within a Domain?
Ionut Cristian Paraschiv, Mihai Dascalu, Danielle S. McNamara, Stefan Trausan-Matu |
EC-TEL | 2 |
| 2016 | {ENTER}ing the Time Series {SPACE}: Uncovering the Writing Process through Keystroke Analyses
Laura K. Allen, Matthew E. Jacovina, Mihai Dascalu, Rod D. Roscoe, Kevin Kent, Aaron D. Likens, Danielle S. McNamara |
EDM | 3 |
| 2016 | Predicting Student Performance and Differences in Learning Styles Based on Textual Complexity Indices Applied on Blog and Microblog Posts: A Preliminary StudyabstractSocial media tools are increasingly popular in Computer Supported Collaborative Learning and the analysis of students' contributions on these tools is an emerging research direction. Previous studies have mainly focused on examining quantitative behavior indicators on social media tools. In contrast, the approach proposed in this paper relies on the actual content analysis of each student's contributions in a learning environment. More specifically, in this study, textual complexity analysis is applied to investigate how student's writing style on social media tools can be used to predict their academic performance and their learning style. Multiple textual complexity indices are used for analyzing the blog and microblog posts of 27 students engaged in a project-based learning activity. The preliminary results of this pilot study are encouraging, with several indexes predictive of student grades and/or learning styles. Elvira Popescu, Mihai Dascalu, Alex Becheru, Scott A. Crossley, Stefan Trausan-Matu |
ICALT | 2 |
| 2016 | What Makes Your Writing Style Unique? Significant Differences Between Two Famous Romanian Orators
Mihai Dascalu, Daniela Gîfu, Stefan Trausan-Matu |
ICCCI (1) | 1 |
| 2016 | Time Evolution of Writing Styles in Romanian LanguageabstractThis paper presents a diachronic analysis centered on the exploration of differences between the writing styles of journalistic texts in Romanian language. This analysis is focused on the time evolution of this language across two adjacent regions, Bessarabia and Romania in two major periods that were marked by important historical differences. Our aim is to examine these language differences based on corpora of historical and contemporary texts. To this end, we employ the ReaderBench framework to calculate a number of textual complexity indices that can be reliably used to characterize writing style. These analyses are conducted on two independent corpora for each of the two language styles, covering the following time periods: 1941-1991, when Bessarabia was separated from Romania and became a state in the Soviet Union (and there were few connections and language influences with Romania), and after July 1991, when Bessarabia became an independent state, Republic of Moldavia (and many language interactions with Romania occurred). The results of our analyses highlight the lexical and cohesive textual complexity indices that best reflect the differences in writing style, ranging from sentence and paragraph structure to word entropy and cohesion, measured in terms of Latent Semantic Analysis (LSA) and Latent Dirichlet Allocation (LDA). Daniela Gîfu, Mihai Dascalu, Stefan Trausan-Matu, Laura K. Allen |
ICTAI | 2 |
| 2016 | Combining click-stream data with NLP tools to better understand MOOC completionabstractCompletion rates for massive open online classes (MOOCs) are notoriously low. Identifying student patterns related to course completion may help to develop interventions that can improve retention and learning outcomes in MOOCs. Previous research predicting MOOC completion has focused on click-stream data, student demographics, and natural language processing (NLP) analyses. However, most of these analyses have not taken full advantage of the multiple types of data available. This study combines click-stream data and NLP approaches to examine if students' on-line activity and the language they produce in the online discussion forum is predictive of successful class completion. We study this analysis in the context of a subsample of 320 students who completed at least one graded assignment and produced at least 50 words in discussion forums, in a MOOC on educational data mining. The findings indicate that a mix of click-stream data and NLP indices can predict with substantial accuracy (78%) whether students complete the MOOC. This predictive power suggests that student interaction data and language data within a MOOC can help us both to understand student retention in MOOCs and to develop automated signals of student success. Scott A. Crossley, Luc Paquette, Mihai Dascalu, Danielle S. McNamara, Ryan Baker 0001 |
LAK | 3 |
| 2015 | Predicting Comprehension from Students' Summaries
Mihai Dascalu, Larise Lucia Stavarache, Philippe Dessus, Stefan Trausan-Matu, Danielle S. McNamara, Maryse Bianco |
AIED | 1 |
| 2015 | ReaderBench: An Integrated Cohesion-Centered FrameworkabstractReaderBench is an automated software framework designed to support both students and tutors by making use of text mining techniques, advanced natural language processing, and social network analysis tools. ReaderBench is centered on comprehension prediction and assessment based on a cohesion-based representation of the discourse applied on different sources (e.g., textual materials, behavior tracks, metacognitive explanations, Computer Supported Collaborative Learning – CSCL – conversations). Therefore, ReaderBench can act as a Personal Learning Environment (PLE) which incorporates both individual and collaborative assessments. Besides the a priori evaluation of textual materials’ complexity presented to learners, our system supports the identification of reading strategies evident within the learners’ self-explanations or summaries. Moreover, ReaderBench integrates a dedicated cohesion-based module to assess participation and collaboration in CSCL conversations. Mihai Dascalu, Larise Lucia Stavarache, Philippe Dessus, Stefan Trausan-Matu, Danielle S. McNamara, Maryse Bianco |
EC-TEL | 1 |
| 2015 | Informal Learning in Online Knowledge Communities: Predicting Community Response to Visitor InquiriesabstractInformal learning in online knowledge communities (OKCs) comprises visitor inquiries on specific topics. Learning can occur only if the OKC adequately respond. This study aims to predict OKC response, using a social learning analytics approach based on computational linguistics and Bakhtin’s theory of dialogism. Observing the blog topic (cooking vs. politics & economics) and the visitor inquiry format (off-topic vs. on-topic), a field experiment with a 2 × 2 factorial design was conducted on a sample of N = 68 blogger communities with a total of 25,303 members. For the entire sample, the community response was influenced only by the inquiry format. In a separate examination of experimental groups, only for one examined topic (cooking) this remained true, while for the other (politics & economics) the community response only depended on the previously established dialog quality. The findings suggest identification criteria for responsive communities, which can support OKC integration in learning environments. Nicolae Nistor, Mihai Dascalu, Larise Lucia Stavarache, Yvonne Serafin, Stefan Trausan-Matu |
EC-TEL | 2 |
| 2015 | Discourse cohesion: a signature of collaborationabstractAs Computer Supported Collaborative Learning (CSCL) becomes increasingly adopted as an alternative to classic educational scenarios, we face an increasing need for automatic tools designed to support tutors in the time consuming process of analyzing conversations and interactions among students. Therefore, building upon a cohesion-based model of the discourse, we have validated ReaderBench, a system capable of evaluating collaboration based on a social knowledge-building perspective. Through the inter-twining of different participants' points of view, collaboration emerges and this process is reflected in the identified cohesive links between different speakers. Overall, the current experiments indicate that textual cohesion successfully detects collaboration between participants as ideas are shared and exchanged within an ongoing conversation. Mihai Dascalu, Stefan Trausan-Matu, Philippe Dessus, Danielle S. McNamara |
LAK | 1 |
| 2014 | Reflecting Comprehension through French Textual Complexity FactorsabstractResearch efforts in terms of automatic textual complexity analysis are mainly focused on English vocabulary and few adaptations exist for other languages. Starting from a solid base in terms of discourse analysis and existing textual complexity assessment model for English, we introduce a French model trained on 200 documents extracted from school manuals pre-classified into five complexity classes. The underlying textual complexity metrics include surface, syntactic, morphological, semantic and discourse specific factors that are afterwards combined through the use of Support Vector Machines. In the end, each factor is correlated to pupil comprehension metrics scores, spanning throughout multiple classes, therefore creating a clearer perspective in terms of measurements impacting the perceived difficulty of a given text. In addition to purely quantitative surface factors, specific parts of speech and cohesion have proven to be reliable predictors of learners' comprehension level, creating nevertheless a strong background for building dependable French textual complexity models. Mihai Dascalu, Larise Lucia Stavarache, Stefan Trausan-Matu, Philippe Dessus, Maryse Bianco |
ICTAI | 1 |
| 2014 | Are Automatically Identified Reading Strategies Reliable Predictors of Comprehension?
Mihai Dascalu, Philippe Dessus, Maryse Bianco, Stefan Trausan-Matu |
Intelligent Tutoring Systems | 1 |
| 2014 | Validating the Automated Assessment of Participation and of Collaboration in Chat Conversations
Mihai Dascalu, Stefan Trausan-Matu, Philippe Dessus |
Intelligent Tutoring Systems | 1 |
| 2014 | SENSE: A collaborative selfish node detection and incentive mechanism for opportunistic networks
Radu-Ioan Ciobanu, Ciprian Dobre, Mihai Dascalu, Stefan Trausan-Matu, Valentin Cristea |
J. Netw. Comput. Appl. | 3 |
| 2013 | ReaderBench, an Environment for Analyzing Text Complexity and Reading Strategies
Mihai Dascalu, Philippe Dessus, Stefan Trausan-Matu, Maryse Bianco, Aurélie Nardy |
AIED | 1 |
| 2013 | Towards an Integrated Model of Teacher Inquiry into Student Learning, Learning Design and Learning Analytics
Cecilie Hansen, Valérie Emin, Barbara Wasson, Yishay Mor, María Jesús Rodríguez-Triana, Mihai Dascalu, Rebecca Ferguson, Jean-Philippe Pernin |
EC-TEL | 6 |
| 2013 | Virtual Communities of Practice in Academia: Automated Analysis of Collaboration Based on the Social Knowledge-Building Model
Nicolae Nistor, Mihai Dascalu, Stefan Trausan-Matu, Dan Mihaila, Beate Baltes, George Smeaton |
EC-TEL | 2 |
| 2013 | Collaborative selfish node detection with an incentive mechanism for opportunistic networks
Radu-Ioan Ciobanu, Ciprian Dobre, Mihai Dascalu, Stefan Trausan-Matu, Valentin Cristea |
IM | 3 |
| 2012 | A System for the Automatic Analysis of Computer-Supported Collaborative Learning ChatsabstractThe paper presents a system for helping the analysis of Computer-Supported Collaborative Learning chat (instant messenger) sessions, starting from a polyphonic model of the discourse inspired by the dialogistics of Bakhtin. Some theoretical basics of the model are presented, followed by implementation details and validation results. Stefan Trausan-Matu, Mihai Dascalu, Traian Rebedea |
ICALT | 2 |
| 2012 | Textual Complexity and Discourse Structure in Computer-Supported Collaborative Learning
Stefan Trausan-Matu, Mihai Dascalu, Philippe Dessus |
ITS | 2 |
| 2011 | Automatic Assessment of Collaborative Chat Conversations with PolyCAFe
Traian Rebedea, Mihai Dascalu, Stefan Trausan-Matu, Gillian Armitt, Costin-Gabriel Chiru |
EC-TEL | 2 |
| 2011 | Beyond Traditional NLP: A Distributed Solution for Optimizing Chat Processing - Automatic Chat Assessment Using Tagged Latent Semantic AnalysisabstractWith the increasing popularity and evolution of Computer Supported Collaborative Learning systems, the need for developing a tool that automatically assesses instant messaging conversations has become imperative. The main reasons are the high volume of data and the increased amount of time spent for manually assessing conversations. We propose an automated analysis system based on Natural Language Processing (centered on Latent Semantic Analysis and Social Network Analysis) and optimize its runtime performance by means of distributed computing. Moreover, we provide a unique grading mechanism based on a multilayered architecture and induce an increase of speedup by deploying a Replicated Worker architecture. Load balancing and fault tolerance represent key aspects of this approach, besides the actual increase in performance. Mihai Dascalu, Ciprian Dobre, Stefan Trausan-Matu, Valentin Cristea |
ISPDC | 1 |
| 2010 | Overview and Preliminary Results of Using PolyCAFe for Collaboration Analysis and Feedback Generation
Traian Rebedea, Mihai Dascalu, Stefan Trausan-Matu, Dan Banica, Alexandru Gartner, Costin-Gabriel Chiru, Dan Mihaila |
EC-TEL | 2 |