Juan Manuel Pérez

dblp:67/67 · also Juan Manuel Pérez-Martínez · DBLP profile ↗
← Back
16ranked-venue papers
10as first author
2since 2021 · last 2025
0000-0002-0782-6716ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 9 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 84% Data integration and cleaning · 11% Query processing and optimization · 5%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
evaluation
0.912025
MessIRve: A Large-Scale Spanish Information Retrieval Dataset · EMNLP 2025
Information retrieval
retrieval models
0.312025
MessIRve: A Large-Scale Spanish Information Retrieval Dataset · EMNLP 2025
Data integration and cleaning
semi-structured data integration
0.112008
Integrating Data Warehouses with Web Data: A Survey · IEEE Trans. Knowl. Data Eng. 2008
Information retrieval
context-based retrieval
0.112007
R-Cubes: OLAP Cubes Contextualized with Documents · ICDE 2007
Query processing and optimization
OLAP
0.112007
R-Cubes: OLAP Cubes Contextualized with Documents · ICDE 2007

Methods — techniques the papers use, named apart from their topics

baseline evaluation · 0.9XML · 0.1OLAP · 0.1information retrieval · 0.1XML document warehousing · 0.1
YearPublicationVenuePosition
2025 MessIRve: A Large-Scale Spanish Information Retrieval Dataset
abstract
Information retrieval (IR) is the task of finding relevant documents in response to a user query.Although Spanish is the second most spoken native language, there are few Spanish IR datasets, which limits the development of information access tools for Spanish speakers.We introduce MessIRve, a large-scale Spanish IR dataset with almost 700,000 queries from Google's autocomplete API and relevant documents sourced from Wikipedia.MessIRve's queries reflect diverse Spanish-speaking regions, unlike other datasets that are translated from English or do not consider dialectal variations.The large size of the dataset allows it to cover a wide variety of topics, unlike smaller datasets.We provide a comprehensive description of the dataset, comparisons with existing datasets, and baseline evaluations of prominent IR models.Our contributions aim to advance Spanish IR research and improve information access for Spanish speakers.
Francisco Valentini, Viviana Cotik, Damián Ariel Furman, Ivan Bercovich, Edgar Altszyler, Juan Manuel Pérez
EMNLP6
2022 RoBERTuito: a pre-trained language model for social media text in Spanish
abstract
Since BERT appeared, Transformer language models and transfer learning have become state-of-the-art for natural language processing tasks. Recently, some works geared towards pre-training specially-crafted models for particular domains, such as scientific papers, medical documents, user-generated texts, among others. These domain-specific models have been shown to improve performance significantly in most tasks; however, for languages other than English, such models are not widely available. In this work, we present RoBERTuito, a pre-trained language model for user-generated text in Spanish, trained on over 500 million tweets. Experiments on a benchmark of tasks involving user-generated text showed that RoBERTuito outperformed other pre-trained language models in Spanish. In addition to this, our model has some cross-lingual abilities, achieving top results for English-Spanish tasks of the Linguistic Code-Switching Evaluation benchmark (LinCE) and also competitive performance against monolingual models in English Twitter tasks. To facilitate further research, we make RoBERTuito publicly available at the HuggingFace model hub together with the dataset used to pre-train it.
Juan Manuel Pérez, Damián Ariel Furman, Laura Alonso Alemany, Franco M. Luque
LREC1
2017 Cross-Linguistic Study of the Production of Turn-Taking Cues in American English and Argentine Spanish
Pablo Brusco, Juan Manuel Pérez, Agustín Gravano
INTERSPEECH2
2017 Using Prosody to Classify Discourse Relations
abstract
Comunicació presentada a: The 18th Annual Conference of the International Speech Communication Association (INTERSPEECH 2017), celebrada a Estocolm, Suència, del 20 al 24 d'agost de 2017.
Janine Kleinhans, Mireia Farrús, Agustín Gravano, Juan Manuel Pérez, Catherine Lai, Leo Wanner
INTERSPEECH4
2016 Disentrainment may be a Positive Thing: A Novel Measure of Unsigned Acoustic-Prosodic Synchrony, and its Relation to Speaker Engagement
Juan Manuel Pérez, Ramiro H. Gálvez, Agustín Gravano
INTERSPEECH1
2009 A relevance model for a data warehouse contextualized with documents
Juan Manuel Pérez, Rafael Berlanga Llavori, María José Aramburu Cabo
Inf. Process. Manag.1
2008 Contextualizing data warehouses with documents
Juan Manuel Pérez, Rafael Berlanga Llavori, María José Aramburu Cabo, Torben Bach Pedersen
Decis. Support Syst.1
2008 Integrating Data Warehouses with Web Data: A Survey
abstract
This paper surveys the most relevant research on combining Data Warehouse (DW) and Web data. It studies the XML technologies that are currently being used to integrate, store, query and retrieve web data, and their application to DWs. The paper reviews different DW distributed architectures and the use of XML languages as an integration tool in these systems. It also introduces the problem of dealing with semi-structured data in a DW. It studies Web data repositories, the design of multidimensional databases for XML data sources and the XML extensions of On-Line Analytical Processing techniques. The paper addresses the application of information retrieval technology in a DW to exploit text-rich documents collections. The authors hope that the paper will help to discover the main limitations and opportunities that offer the combination of the DW and the Web fields, as well as, to identify open research lines.
Juan Manuel Pérez, Rafael Berlanga Llavori, María José Aramburu Cabo, Torben Bach Pedersen
IEEE Trans. Knowl. Data Eng.1
2007 R-Cubes: OLAP Cubes Contextualized with Documents
abstract
Current data warehouse and OLAP (Kimball and Ross, 2002) technologies can be efficiently applied to analyze the huge amounts of structured data that companies produce. These organizations also produce many text documents and use the Web as their largest source of external information. Although these documents include highly valuable information that should also be exploited by companies, they cannot be analyzed by current OLAP technologies because they are unstructured and mainly contain text. The current trend is to find these documents available in XML-like formats. Our proposal is to build XML document warehouses that can be used by companies to store unstructured information coming from their internal and external sources. In (Perez et al., 2005) we proposed an architecture for the integration of a corporate warehouse of structured data with a warehouse of text-rich XML documents. We call the resulting warehouse a contextualized warehouse. Since the XML document warehouse may contain documents about many different topics, we apply well-known information retrieval (IR) (Baeza-Yates and Ribeiro-Neto, 1999) techniques to select the context of analysis from the document warehouse. First, the user specifies an analysis context by supplying a sequence of keywords (e.g., an IR condition like "financial crisis"). Then, the analysis is performed on a so-called R-cube (Relevance cube), which is materialized by retrieving the documents and facts related to the selected context. Each fact in the R-cube will be linked to the set of documents that describe its context, and will have assigned a numerical value representing its relevance with respect to the specified context (e.g., how important the fact is for a "financial crisis"). In (Perez et al., 2005) we provided R-cubes with a data model and an algebra. This paper presents a prototype R-cube system, and explains how to use it.
Juan Manuel Pérez, Rafael Berlanga Llavori, María José Aramburu Cabo, Torben Bach Pedersen
ICDE1
2005 A relevance-extended multi-dimensional model for a data warehouse contextualized with documents
abstract
Current data warehouse and OLAP technologies can be applied to analyze the structured data that companies store in their databases. The circumstances that describe the context associated with these data can be found in other internal and external sources of documents. In this paper we propose to combine the traditional corporate data warehouse with a document warehouse, resulting in a contextualized warehouse. Thus, contextualized warehouses keep a historical record of the fact and their contexts as described by the documents. In this framework, the user selects an analysis context which is represented as a novel type of OLAP cube, here called R-cube. R-cubes are characterized by two special dimensions, namely: the relevance and the context dimensions. The first dimension measures the relevance of each fact in the selected analysis context, whereas the second one relates each fact with the documents that explain their circumstances. In this work we extend an existing multi-dimensional data model and algebra for representing the R-cubes.
Juan Manuel Pérez, Rafael Berlanga Llavori, María José Aramburu Cabo, Torben Bach Pedersen
DOLAP1
2005 IR and OLAP in XML Document Warehouses
Juan Manuel Pérez, Torben Bach Pedersen, Rafael Berlanga Llavori, María José Aramburu Cabo
ECIR1
2004 JERARTOP: A New Topic Detection System
Aurora Pons-Porrata, Rafael Berlanga Llavori, José Ruiz-Shulcloper, Juan Manuel Pérez
CIARP4
2004 A Document Model Based on Relevance Modeling Techniques for Semi-structured Information
Juan Manuel Pérez, Rafael Berlanga Llavori, María José Aramburu Cabo
DEXA1
2003 XML Schemata Inference and Evolution
Ismael Sanz, Juan Manuel Pérez, Rafael Berlanga Llavori, María José Aramburu Cabo
DEXA2
2002 XRL: A XML-Based Query Language for Advanced Services in Digital Libraries
Juan Manuel Pérez, María José Aramburu Cabo, Rafael Berlanga Llavori
DEXA1
2001 Techniques and Tools for the Temporal Analysis of Retrieved Information
Rafael Berlanga Llavori, Juan Manuel Pérez, María José Aramburu Cabo, Dolores María Llidó
DEXA2