Maria J. P. Dantas

dblp:414/6204 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Resource-Aware Climate Modeling for Regional Adaptation under Data Scarcity
abstract
Climate modeling is increasingly constrained not only by predictive accuracy, but by how climate intelligence can be computed and deployed under heterogeneous infrastructural conditions. This challenge is particularly critical in data-scarce regions, where local forecasting is limited by sparse observations, hardware constraints, and reliance on centralized global models. This paper presents an architectural contribution for regional climate adaptation, framing climate modeling as a distributed digital infrastructure problem. The approach leverages pretrained global representations and adapts them through a modular pipeline that combines reconstruction of incomplete local observations with physics-consistent inference under limited computational budgets. The primary novelty lies not in individual components, but in the regime-aware integration logic: adaptive routing across reconstruction methods based on data availability, inter-stage confidence propagation, and joint architectural codesign with hardware constraints. Preliminary evidence from the reconstruction layer, using INMET time series and ERA5-Land priors, supports viability under realistic conditions of missing data. By jointly addressing data incompleteness, computational efficiency, and model adaptation, this work introduces a systemsoriented design framework for more accessible and robust climate intelligence.
Maria J. P. Dantas, João V. P. Fernandes, João P. L. Queiroz, Felipe C. V. dos Santos, Norton P. Ricardo, Davi A. S. Dias, João H. S. Miranda, Carine S. Santos, Rafaela M. Silva, Salatiel A. A. Jordão, Reinaldo R. Rosa
COMPSAC1
2026 Style and Stance Alignment in Reddit Discourse
abstract
This work investigates whether writing style is associated with stance alignment in large scale online discourse. We construct a Reddit based pipeline that transforms r/AskReddit questions and comments interactions into stance units using large language models (LLMs), stylistic embeddings, and semantic embeddings. The study begins with 169,053 original AskReddit questions, categorizes them into 48 semantic classes, and filters them for stance suitability. The full question corpus contains 176,932 rows; in the target stance evaluation stage, 22,238 original questions from 10 selected semantic classes are paired with up to 200 top level comments each and annotated by an LLM, yielding 277,474 comment level model outputs and 352,407 stance instances. Style is represented with answer text embeddings trained to capture writing style, while semantic context is represented with embeddings of question summaries, stance targets, and canonical opinions. We analyze 8,000,000 randomly sampled stance instance pairs by measuring style similarity, topic similarity, and stance label agreement. Across topic similarity thresholds from 0.50 to 0.95, pairs with high stylistic similarity consistently show higher model annotated stance agreement than the topic similar baseline, with lifts of approximately 6% to 12% points. A parallel augmentation analysis shows that synthetic questions produced by the tested models under generalize real class diversity, motivating the use of original questions for the main stance analysis. These findings suggest that writing style carries a weak but measurable signal related to stance alignment under semantically controlled conditions.
Salatiel A. A. Jordão, Carine S. Santos, Felipe Santos, Maria J. P. Dantas
COMPSAC4
2025 Multilingual and Informal Web Datasets for Robust Language Modeling
abstract
Large Language Models (LLMs) tend to lack textual, sociolinguistic, and cultural diversity due to being trained on homogeneous data sources. This study promotes the development of innovative datasets from marginal internet sites, such as anonymous imageboards (4chan, 2channel), niche web forums, YouTube comments, X posts, and Reddit sub-reddits, to overcome this shortfall. These datasets prioritize multilingual and multicultural features, reflecting spontaneous language use and distinctive sociocultural communication. Our solution involves web scraping and utilizes existing archives, with NLLB for multilingual translation support, and domain-specific BERT-based model training for improved text embeddings. We also describe building a custom language identification classifier and AI-generated text detector, both trained on a paired human-AI dataset augmented using the Qwen3 model, to enhance data quality and analytical potential for research in linguistic diversity and model robustness.
Salatiel A. A. Jordão, Carine S. Santos, Maria J. P. Dantas
COMPSAC3