VLDB 2026 Research / reviewers in the wild / expert
Jorge Baptista
dblp:98/6884
· DBLP profile ↗
13ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-4603-4364ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Toward Consistency in Writing Proficiency Assessment: Mitigating Classification Variability in Developmental Education
Miguel Da Corte, Jorge Baptista |
CSEDU (2) | 2 |
| 2025 | Refining English Writing Proficiency Assessment and Placement in Developmental Education Using NLP Tools and Machine Learning
Miguel Da Corte, Jorge Baptista |
CSEDU (2) | 2 |
| 2024 | Charting the Linguistic Landscape of Developing Writers: An Annotation Scheme for Enhancing Native Language ProficiencyabstractThis study describes a pilot annotation task designed to capture orthographic, grammatical, lexical, semantic, and discursive patterns exhibited by college native English speakers participating in developmental education (DevEd) courses. The paper introduces an annotation scheme developed by two linguists aiming at pinpointing linguistic challenges that hinder effective written communication. The scheme builds upon patterns supported by the literature, which are known as predictors of student placement in DevEd courses and English proficiency levels. Other novel, multilayered, linguistic aspects that the literature has not yet explored are also presented. The scheme and its primary categories are succinctly presented and justified. Two trained annotators used this scheme to annotate a sample of 103 text units (3 during the training phase and 100 during the annotation task proper). Texts were randomly selected from a population of 290 community college intending students. An in-depth quality assurance inspection was conducted to assess tagging consistency between annotators and to discern (and address) annotation inaccuracies. Krippendorff’s Alpha (K-alpha) interrater reliability coefficients were calculated, revealing a K-alpha score of k=0.40, which corresponds to a moderate level of agreement, deemed adequate for the complexity and length of the annotation task. Miguel Da Corte, Jorge Baptista |
LREC/COLING | 2 |
| 2024 | Enhancing Writing Proficiency Classification in Developmental Education: The Quest for AccuracyabstractDevelopmental Education (DevEd) courses align students’ college-readiness skills with higher education literacy demands. These courses often use automated assessment tools like Accuplacer for student placement. Existing literature raises concerns about these exams’ accuracy and placement precision due to their narrow representation of the writing process. These concerns warrant further attention within the domain of automatic placement systems, particularly in the establishment of a reference corpus of annotated essays for these systems’ machine/deep learning. This study aims at an enhanced annotation procedure to assess college students’ writing patterns more accurately. It examines the efficacy of machine-learning-based DevEd placement, contrasting Accuplacer’s classification of 100 college-intending students’ essays into two levels (Level 1 and 2) against that of 6 human raters. The classification task encompassed the assessment of the 6 textual criteria currently used by Accuplacer: mechanical conventions, sentence variety & style, idea development & support, organization & structure, purpose & focus, and critical thinking. Results revealed low inter-rater agreement, both on the individual criteria and the overall classification, suggesting human assessment of writing proficiency can be inconsistent in this context. To achieve a more accurate determination of writing proficiency and improve DevEd placement, more robust classification methods are thus required. Miguel Da Corte, Jorge Baptista |
LREC/COLING | 2 |
| 2024 | Leveraging NLP and Machine Learning for English (L1) Writing Assessment in Developmental Education
Miguel Da Corte, Jorge Baptista |
CSEDU (2) | 2 |
| 2022 | Early Experiments on Automatic Annotation of Portuguese Medieval Texts
Maria Inês Bico, Jorge Baptista, Fernando Batista, Esperança Cardeira |
TPDL | 2 |
| 2016 | Automatic generation of exercises on passive transformation in PortugueseabstractTechnology plays a very important role in education and Intelligent Computer-Assisted Language Learning (iCALL) has emerged as a complementary or even alternative method to the conventional language teaching practices. The automatic generation (and correction) of language exercises based on real texts extracted from corpora constitutes a non-trivial challenge to iCALL tutorial systems, and may involve the use of sophisticated Natural Language Processing tools and large-scale linguistic resources. This paper presents the main issues related to the automatic generation of exercises on the Passive transformation, a commonly occurring type of exercises in language textbooks, but also a very complex topic of Portuguese grammar. The paper describes the methods used to produce a large batch of passive-active sentence pairs, where the active sentence was automatically generated from naturally occurring passive sentences, taken from a large-sized, publicly available, corpus. Sentence pairs are ranked by difficulty level. A sample of randomly selected sentence pairs (40 from difficult level, 100 from medium, and 100 from easy level) was manually evaluated by an expert. Results are presented and error analysis is performed. The sentence pairs can be used as prime and correct answer for iCALL systems. Jorge Baptista, Sandra Lourenco, Nuno J. Mamede |
CEC | 1 |
| 2016 | Automated anonymization of text documentsabstractSharing data in the form of text is important for a wide range of activities but it also raises a concern about privacy when sharing data that could be sensitive. Automated text anonymization is a solution for removing all the sensitive information from documents. However, this is a challenging task due to the unstructured form of textual data and the ambiguity of natural language. In this work, we present our implementation of an automated anonymization system, built in a modular structure, for documents written in Portuguese. Four different methods of anonymization are evaluated and compared. Two methods replace the sensitive information by artificial labels: suppression and tagging. The other two methods replace the information by textual expressions: random substitution and generalization. Evaluation showed that the use of the tagging and the generalization methods facilitates the reading of an anonymized text while preventing some semantic drifts caused by the remotion of the original information. Nuno J. Mamede, Jorge Baptista, Francisco Dias |
CEC | 2 |
| 2016 | metaTED: a Corpus of Metadiscourse for Spoken Language
Rui Correia, Nuno J. Mamede, Jorge Baptista, Maxine Eskénazi |
LREC | 3 |
| 2015 | Automatic Text Difficulty Classifier - Assisting the Selection Of Adequate Reading Materials For European Portuguese Teaching
Pedro Curto, Nuno J. Mamede, Jorge Baptista |
CSEDU (1) | 3 |
| 2013 | ASR-based exercises for listening comprehension practice in European Portuguese
Thomas Pellegrini, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede, Maxine Eskénazi |
Comput. Speech Lang. | 4 |
| 2012 | Overview of Computer-assisted Language Learning for European Portuguese at L2f
Thomas Pellegrini, Wang Ling, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede |
CSEDU (2) | 6 |
| 2011 | Automatic Generation of Listening Comprehension Learning Material in European PortugueseabstractThe goal of this work is the automatic selection of materials for a listening comprehension game. We would like to select automatically transcribed sentences from recent broadcast news corpora, in order to gather material for the games with little human effort. The recognized words are used as the ground solution of the exercises, thus sentences with misrecognitions need to be filtered out. Our experiments confirmed the feasibility of the filter chain that automatically selects sentences, although harder confidence thresholds may be needed. Together with the correct words, wrong candidates, namely distractors, are also needed to build the exercises. Two techniques of distractor generation are presented, either based on the confusion networks produced by the recognizer, or on phonetic distances. The experiments confirmed the complementarity of both approaches. Index Terms: CALL, Listening Comprehension, European Portuguese, ASR, distractors Thomas Pellegrini, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede |
INTERSPEECH | 4 |