Peter Z. Revesz

dblp:r/PZRevesz · also Peter Zsolt Revesz · DBLP profile ↗
← Back
40ranked-venue papers in the field
21as first author
9since 2021 · last 2025
0000-0002-1145-1283ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 35 (21 first)Other / Interdisciplinary · 5
YearPublicationVenuePosition
2025 Data Mining for Language Superfamilies Using Congruent Sound Groups
abstract
There have been several attempts in recent years to prove that some well-known language families can be grouped together into a language superfamily. This paper presents a data mining method to search for a language superfamily. The data mining method is based on the consideration of regular sound changes in various languages. Congruent Sound Groups are derived from the commonly observed regular sound changes. To demonstrate the feasibility of this method, we collected a set of words related to the four basic elements of air, earth, fire and water from seven languages: Hindi, Japanese, Korean, Russian, Sanskrit, Tamil and Telugu. These seven languages are classified into four different language families: Dravidian, Indo-European, Japonic, and Koreanic. The congruent sound group-based analysis enabled the identification of seven cognate groups of words that involve different language families. This suggests that these four different language families originate from a single protolanguage that was likely spoken in Asia more than 10,000 years ago.
Peter Z. Revesz, Mohanendra Siddha
IDEAS1
2025 Generative Adversarial Networks Reveal Carian, Elder Futhark, Old Hungarian and Old Turkic Script Relationships
Shohaib Shaffiey, Peter Z. Revesz
IDEAS2
2025 Automated Identification of Allographs Among the Indus Valley Script Signs
Harsh Tamkiya, Gunjit Agrawal, Chiradeep Debnath, Peter Z. Revesz
IDEAS4
2024 Comparing Related Languages with a Fuzzy Morphism Matching Algorithm
abstract
Abstract This paper proposes a fuzzy morphism matching algorithm for discovering similarities within related languages. The fuzzy morphism matching algorithm takes as input a novel representation of the linguistic structures of the two languages that are compared. This representation is a type of Markov model that is built from an abstract representation of the basic set of words in the languages where the abstraction is based on combinations of six phoneme categories and three positions of those phonemes within the basic sets of words. The limited number of nodes in these Markov models allows efficient calculations of partial subgraph isomorphism matchings between them, and the degree of matching leads to a natural similarity measure that depends not on the number of cognate words but only on the phonetic structure of the languages, which have greater stability. This allows the detection of a strong similarity between closely related languages such as English and German as well as a weaker similarity between more distantly related languages like English and Hungarian.
Daniel Schaefer, Peter Z. Revesz
IDEAS2
2023 A Generalization of the Chomsky-Halle Phonetic Representation using Real Numbers for Robust Speech Recognition in Noisy Environments
abstract
Speech recognition is difficult when the speech signal is weak or occurs in a noisy environment. This paper presents an efficient and robust method that can reconstruct the standard pronunciation of English phonemes and words given a weak or noisy signal. The reconstruction is based on a novel representation of the reconstruction task as a problem of data retrieval from a database in two different cases: (1) when the phonemes are represented in the database as binary tuples and the input is also a binary tuple from which deletion errors occur, and (2) when the phonemes are represented in the database and in the input as tuples of real values ranging between 0 and 1. In the latter case, the input phoneme could contain both a higher or lower value than the standard phoneme in the database that is intended by the speaker. For case (2) a theorem is proven regarding when the data retrieval can be expected to be reliable.
Peter Z. Revesz
IDEAS1
2022 Feature Analysis of Indus Valley and Dravidian Language Scripts with Similarity Matrices
abstract
This paper investigates the similarity between the Indus Valley script and the Kannada, Malayalam, Tamil, and Telugu scripts that are used to write Dravidian languages. The closeness of these scripts is determined by applying a feature analysis of each sign of these scripts and creating similarity matrices that describe the similarity of any pair of signs from two different scripts. The feature list that we use for the analysis of these Dravidian language-related scripts includes six new features beyond the thirteen features that were used for the study of Minoan Linear A and related scripts by Revesz. These new features are the check mark, short vertical line, dot, upper curve, parallel curves, and horizontal line features.
Sarat Sasank Barla, Sai Surya Sanjay Alamuru, Peter Z. Revesz
IDEAS3
2022 Data Science Applied to Discover Ancient Minoan-Indus Valley Trade Routes Implied by Common Weight Measures
abstract
This paper applies data mining of weight measures to discover possible long-distance trade routes among Bronze Age civilizations from the Mediterranean area to India. As a result, a new northern route via the Black Sea is discovered between the Minoan and the Indus Valley civilizations. This discovery enhances the growing set of evidence for a strong and vibrant connection among Bronze Age civilizations.
Peter Z. Revesz
IDEAS1
2022 Measuring Vowel Harmony within Hungarian, the Indus Valley Script Language, Spanish and Turkish Using ERGM
abstract
Front-back vowel harmony is an important characteristic of many languages. Testing whether an untranslated script has vowel harmony may aid its decipherment. This paper tests vowel harmony for three different modern languages (Hungarian, Spanish and Turkish) as well as the extinct underlying language of the undeciphered Indus Valley script. We also introduce a novel vowel harmony index based on the Exponential Random Graph Model for graphs. To achieve this, we first select words from each of the modern languages (Hungarian, Turkish, and Spanish) from their Swadesh list. Then we divide each word into syllables, isolating the vowels. We then analyze the three modern languages using Exponential Random Graph Model methods. The results indicate that this procedure and the vowel harmony index are feasible to define the degree of vowel harmony in a language. The procedure is then extended to the undeciphered Indus Valley Script. Our results indicate that the underlying language of the Indus Valley Script also had vowel harmony. We found that on average the odds of the IVS having vowel harmony were 6.61 times higher than would be found in a random graph.
Josey VanOrsdale, Jigyasa Chauhan, Sai Vivek Potlapally, Srikar Chanamolu, Sai Pratyush Reddy Kasara, Peter Z. Revesz
IDEAS6
2021 Data Mining Autosomal Archaeogenetic Data to Determine Minoan Origins
abstract
This paper presents a method for data mining archaeogenetic autosomal data. The method is applied to the widely debated topic of the origin of the Bronze Age Minoan culture that existed on the island of Crete from 5000 to 3500 years ago. The data is compared with some Neolithic and early Bronze Age samples from the nearby Cycladic islands, mainland Greece and other Neolithic sites. The method shows that a large component of the Minoan autosomal genomes has sources from the Neolithic areas of northern Greece and the rest of the Balkans and a minor component comes directly from Neolithic Anatolia and the Caucasus.
Peter Z. Revesz
IDEAS1
2020 A novel spatio-temporal interpolation algorithm and its application to the COVID-19 pandemic
abstract
This paper describes several interpolation methods for predicting the number of cases of the COVID-19 pandemic. The interpolation methods include some well-known temporal interpolation algorithms including Lagrange interpolation, cubic spline interpolation, and exponential decay interpolation. These temporal interpolation algorithms enable the interpolation of the COVID-19 cases at locations where measures on prior days are available. However, pandemics are not purely temporal but spatio-temporal phenomena. Therefore, the neighboring locations need to be considered too in order to derive accurate interpolation values for future days. This paper introduces a novel spatio-temporal interpolation algorithm that is shown to be better than any purely temporal interpolation algorithm in predicting the COVID-19 cases in the continental United States. In particular, the novel spatio-temporal interpolation method achieves a mean absolute error of 8.44 cases over a million people when predicting two days ahead the number of cases of the COVID-19 pandemic.
Junzhe Cai, Peter Z. Revesz
IDEAS2
2020 Affine-invariant querying of spatial data using a triangle-based logic
Sofie Haesevoets, Bart Kuijpers, Peter Z. Revesz
GeoInformatica3
2019 Data mining ancient scripts to investigate their relationships and origins
abstract
This paper describes a data mining study of a set of ancient scripts in order to discover their relationships, including their possible common origin from a single root script. The data mining uses convolutional neural networks and support vector machines to find the degree of visual similarity between pairs of symbols in eight different ancient scripts. Among the surprising results of the data mining are the following: (1) the Indus Valley Script is visually closest to Sumerian pictographs, and (2) the Linear B script is visually closest to the Cretan Hieroglyphic script.
Shruti Daggumati, Peter Z. Revesz
IDEAS2
2019 The design and implementation of AIDA: ancient inscription database and analytics system
abstract
This paper describes the development of AIDA, the Ancient Inscription Database and Analytics system. The AIDA system currently stores three types of ancient Minoan inscriptions: Linear A, Cretan Hieroglyph and Phaistos Disk inscriptions. In addition, AIDA provides candidate syllabic values and translations of Minoan words and inscriptions into English. The AIDA system allows the users to change these candidate phonetic assignments to the Linear A, Cretan Hieroglyph and Phaistos symbols. Hence the AIDA system provides for various scholars not only a convenient online resource to browse Minoan inscriptions but also provides an analysis tool to explore various options of phonetic assignments and their implications. Such explorations can aid in the decipherment of Minoan inscriptions.
Peter Z. Revesz, M. Parvez Rashid, Yves Tuyishime
IDEAS1
2018 Spatio-Temporal Data Mining of Major European River and Mountain Names Reveals Their Near Eastern and African Origins
Peter Z. Revesz
ADBIS1
2018 Data Mining Ancient Script Image Data Using Convolutional Neural Networks
abstract
The recent surge in ancient scripts has resulted in huge image libraries of ancient texts. Data mining of the collected images enables the study of the evolution of these ancient scripts. In particular, the origin of the Indus Valley script is highly debated. We use convolutional neural networks to test which Phoenician alphabet letters and Brahmi symbols are closest to the Indus Valley script symbols. Surprisingly, our analysis shows that overall the Phoenician alphabet is much closer than the Brahmi script to the Indus Valley script symbols.
Shruti Daggumati, Peter Z. Revesz
IDEAS2
2016 Spatio-temporal traffic video data archiving and retrieval system
Hang Yue, Laurence R. Rilett, Peter Z. Revesz
GeoInformatica3
2015 Data Mining Citation Databases: A New Index Measure that Predicts Nobel Prizewinners
abstract
A new citations index measure that combines total citations and h-index values is proposed. Based on the new citation index, a new algorithm that identifies emerging scientific leaders is also proposed. Experimental results show that the new method can predict Physics Nobel prizewinners from among other highly cited physics researchers many years ahead of their winning of a Nobel Prize. Hence while neither total citations nor h-index alone were good indicators, their combinations can be significant predictors of excellence in science.
Peter Z. Revesz
IDEAS1
2014 Applications of spatio-temporal data mining to north platter river reservoirs
abstract
We propose a spatio-temporal data mining method based on support vector machines regression and spatio-temporal feature reduction by principal component analysis. We apply the spatio-temporal data mining method to derive an automated controller for the reservoirs of the North Platte River. The automated controller opens and closes dams to efficiently and accurately control the reservoirs' water levels.
Abhinaya Mohan, Peter Z. Revesz
IDEAS2
2014 A method for predicting citations to the scientific publications of individual researchers
abstract
Any researcher's publications at any time can be ordered from the highest cited to the lowest cited, yielding a citation curve. We describe a novel method for predicting citation curves of researchers in the future. The method depends on treating the citation curves of researchers for various years as one single spatio-temporal function from rank and time to citations. For each researcher, we derive an estimate of this spatio-temporal function that can be used to predict the total citations of individual publications at any given rank at any time. Experiments show that this method can accurately predict entire citation curves and derived measures, such as, the total citations to all publications and the h-index of the researchers.
Peter Z. Revesz
IDEAS1
2011 Temporal data classification using linear classifiers
Peter Z. Revesz, Thomas Triplet
Inf. Syst.1
2009 Temporal Data Classification Using Linear Classifiers
Peter Z. Revesz, Thomas Triplet
ADBIS1
2009 A comparison of abstract data type and constraint database approaches to GIS query languages
abstract
Designing query languages for geographic information systems using the traditional approach of constantly adding new data types and operations on the new data types has reached a limit beyond which the query language is no longer easy to understand or convenient to use. We advocate instead the design of query languages based on constraint databases. We show that many formerly difficult-looking queries, such as the shortest path query, can be expressed using simple SQL-like queries of constraint databases.
Peter Z. Revesz
GIS1
2009 Efficient MaxCount and threshold operators of moving objects
abstract
Calculating operators of continuously moving objects presents some unique challenges, especially when the operators involve aggregation or the concept of congestion, which happens when the number of moving objects in a changing or dynamic query space exceeds some threshold value. This paper presents the following six d -dimensional moving object operators: (1) M ax C ount (or MinCount ), which finds the Maximum (or Minimum) number of moving objects simultaneously present in the dynamic query space at any time during the query time interval. (2) CountRange , which finds a count of point objects whose trajectories intersect the dynamic query space during the query time interval. (3) ThresholdRange , which finds the set of time intervals during which the dynamic query space is congested. (4) ThresholdSum , which finds the total length of all the time intervals during which the dynamic query space is congested. (5) ThresholdCount , which finds the number of disjoint time intervals during which the dynamic query space is congested. And (6) ThresholdAverage , which finds the average length of time of all the time intervals when the dynamic query space is congested. For these operators separate algorithms are given to find only estimate or only precise values. Experimental results from more than 7,500 queries indicate that the estimation algorithms produce fast, efficient results with error under 5%.
Scot Anderson, Peter Z. Revesz
GeoInformatica2
2008 Reclassification of Linearly Classified Data Using Constraint Databases
Peter Z. Revesz, Thomas Triplet
ADBIS1
2006 On-line maintenance of simplified weighted graphs for efficient distance queries
abstract
We give two efficient on-line algorithms to simplify weighted graphs by eliminating degree-two vertices. Our algorithms are on-line---they react to updates on the data, keeping the simplification up-to-date. We provide both analytical and empirical evaluations of the efficiency of our algorithms. We prove an O(log n) upper bound on the amortized time complexity of our maintenance algorithms, with n the number of insertions. One of our algorithms can handle in logarithmic time the deletions of vertices and edges as well.
Floris Geerts, Peter Z. Revesz, Jan Van den Bussche
GIS2
2005 The Expressivity of Constraint Query Languages with Boolean Algebra Linear Cardinality Constraints
Peter Z. Revesz
ADBIS1
2004 Quantifier-Elimination for the First-Order Theory of Boolean Algebras with Linear Cardinality Constraints
Peter Z. Revesz
ADBIS1
2003 Querying Spatiotemporal XML Using DataFoX
abstract
We describe DataFoX, which is a new query language for XML documents and extends Datalog with support for trees as the domain of the variables. We also introduce for DataFoX a layer algebra, which supports data heterogeneity at the language level, and several algebra-based evaluation techniques.
Peter Z. Revesz
Web Intelligence2
2000 Parametric Rectangles: A Model for Querying and Animation of Spatiotemporal Databases
Mengchu Cai, Dinesh Keshwani, Peter Z. Revesz
EDBT3
2000 Algorithms for Cartogram Animation
abstract
We describe several value-by-area cartogram animation algorithms that can be used to visualize geographically distributed continuous spatiotemporal data that often occur in GIS systems. We implemented the algorithms as part of the graphical user interface of the MLPQ/GIS database system.
Peter Z. Revesz
IDEAS2
2000 The MLPQ/GIS Constraint Database System
abstract
MLPQ/GIS [4,6] is a constraint database [5] system like CCUBE [1] and DEDALE [3] but with a special emphases on spatio-temporal data. Features include data entry tools (first four icons in Fig. 1), icon-based queries such as @@@@ Intersection, @@@@ Union, @@@@ Area, @@@@ Buffer, @@@@ Max and @@@@ Min, which optimize linear objective functions, and @@@@ for Datalog queries. For example, in Fig. 1 we loaded and displayed a constraint database that represents the midwest United States and loaded two contraint relations describing the movements of two persons. The query icon opened a dialog box into which we entered the query which finds (t, i) pairs such that the two people are in the same state i at the same time t.
Peter Z. Revesz, Pradip Kanjamala, Yuguo Liu
SIGMOD Conference1
1999 Constraint-based Interoperability of Spatiotemporal Databases
Jan Chomicki, Peter Z. Revesz
GeoInformatica2
1998 Safe Query Languages for Constraint Databases
abstract
In the database framework of Kanellakis et al. [1990] it was argued that constraint query languages should take constraint databases as input and give other constraint databases that use the same type of atomic constraints as output. This closed-form requirement has been difficult to realize in constraint query languages that contain the negation symbol. This paper describes a general approach to restricting constraint query languages with negation to safe subsets that contain only programs that are evaluable in closed-form on any valid constraint database input.
Peter Z. Revesz
ACM Trans. Database Syst.1
1997 Model-Theoretic Minimal Chenge Operators for Constraint Databases
Peter Z. Revesz
ICDT1
1997 MLPQ: A Linear Constraint Database System with Aggregate Operators
abstract
The paper describes the MLPQ constraint database system. The query language of MLPQ is SQL extended with linear arithmetic constraints. The input and output databases are linear constraint databases (LCDBs). An important feature of the MLPQ system is that it can handle aggregate operators, Min, Max, Sum, Avg, etc. In MLPQ, these operators are evaluated for a series of linear programming (LP) problems. This approach provides an efficient way of evaluation of SQL queries with aggregate operators on linear constraint databases.
Peter Z. Revesz
IDEAS1
1995 Datalog Queries of Set Constraint Databases
Peter Z. Revesz
ICDT1
1993 On the Semantics of Theory Change: Arbitration between Old and New Information
abstract
Katsuno and Mendelzon divide theory change, the problem of adding new information to a logical theory, into two types: revision and update. We propose a third type of theory change: arbitration. The key idea is the following: the new information is considered neither better nor worse than the old information represented by the logical theory. The new information is simply one voice against a set of others already incorporated into the logical theory. From this follows that arbitration should be commutative. First we define arbitration by a set of postulates and then describe a model-theoretic characterization of arbitration for the case of propositional logical theories. We also study weighted arbitration where different models of a theory can have different weights.
Peter Z. Revesz
PODS1
1992 Knowledgebase Transformations
abstract
We propose a language that expresses uniformly queries and updates on knowledgebases consisting of finite sets of relational structures. The language contains an operator that “inserts” arbitrary first-order sentences into knowledgebase. The semantics of the insertion is based on the notion of update formalized by Katsuno and Mendelzon in the context of belief revision theory. Our language can express, among other things, hypothetical queries and queries on recursively indefinite databases. The expressive power of our language lies between existential second-order and general second-order queries. The data complexity is in general within exponential time, although it can be lowered to co-NP and to polynomial time by restricting the form of queries and updates.
Gösta Grahne, Alberto O. Mendelzon, Peter Z. Revesz
PODS3
1990 A Closed Form for Datalog Queries with Integer Order
Peter Z. Revesz
ICDT1
1990 Constraint Query Languages
abstract
We discuss the relationship between constraint programming and database query languages. We show that bottom-up, efficient, declarative database programming can be combined with efficient constraint solving. The key intuition is that the generalization of a ground fact, or tuple, is a conjunction of constraints. We describe the basic Constraint Query Language design principles, and illustrate them with four different classes of constraints: Polynomial, rational order, equality, and Boolean constraints.
Paris C. Kanellakis, Gabriel M. Kuper, Peter Z. Revesz
PODS3