VLDB 2026 Research / reviewers in the wild / expert
Norases Vesdapunt
dblp:85/9733
· DBLP profile ↗
5ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 4 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data integration and cleaning · 54% Web and social media mining · 29% Data mining · 16% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning
entity resolution |
0.4 | 2 | 2015 | Identifying users in social networks with limited information · ICDE 2015 Crowdsourcing Algorithms for Entity Resolution · Proc. VLDB Endow. 2014 |
Web and social media mining › social media analysis › user identification
cross-platform user identification |
0.2 | 1 | 2015 | Identifying users in social networks with limited information · ICDE 2015 |
Data integration and cleaning › entity resolution
low-resource entity resolution |
0.2 | 1 | 2015 | Identifying users in social networks with limited information · ICDE 2015 |
Web and social media mining
social network analysis |
0.2 | 1 | 2015 | Identifying users in social networks with limited information · ICDE 2015 |
Data mining
crowdsourcing |
0.2 | 1 | 2014 | Crowdsourcing Algorithms for Entity Resolution · Proc. VLDB Endow. 2014 |
Data integration and cleaning › crowdsourced data processing
question selection |
0.2 | 1 | 2014 | Crowdsourcing Algorithms for Entity Resolution · Proc. VLDB Endow. 2014 |
Methods — techniques the papers use, named apart from their topics
graph probing · 0.2API-limited querying · 0.2transitivity inference · 0.2theoretical analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Updating an Existing Social Graph Snapshot via a Limited APIabstractWe study the problem of graph tracking with limited information. In this paper, we focus on updating a social graph snapshot. Say we have an existing partial snapshot, G1, of the social graph stored at some system. Over time G1 becomes out of date. We want to update G1 through a public API to the actual graph, restricted by the number of API calls allowed. Periodically recrawling every node in the snapshot is prohibitively expensive. We propose a scheme where we exploit indegrees and outdegrees to discover changes to the actual graph. When there is ambiguity, we probe the graph and verify edges. We propose a novel strategy designed for limited information that can be adapted to different levels of staleness. We evaluate our strategy against recrawling on real datasets and show that it saves an order of magnitude of API calls while introducing minimal errors. Norases Vesdapunt, Hector Garcia-Molina |
CIKM | 1 |
| 2015 | Identifying users in social networks with limited informationabstractWe study the problem of Entity Resolution (ER) with limited information. ER is the problem of identifying and merging records that represent the same real-world entity. In this paper, we focus on the resolution of a single node g from one social graph (Google+ in our case) against a second social graph (Twitter in our case). We want to find the best match for g in Twitter, by dynamically probing the Twitter graph (using a public API), limited by the number of API calls that social systems allow. We propose two strategies that are designed for limited information and can be adapted to different limits. We evaluate our strategies against a naive one on a real dataset and show that our strategies can provide improved accuracy with significantly fewer API calls. Norases Vesdapunt, Hector Garcia-Molina |
ICDE | 1 |
| 2015 | Errata for "Crowdsourcing Algorithms for Entity Resolution" (PVLDB 7(12): 1071-1082)abstractWe discovered that there was a duplicate figure in our paper. We accidentally put Figure 13(b) for Figure 12(b). We have provided the correct Figure 12(b) above (See Figure 1). Figure 1 plots the recall of various strategies as a function of the number of questions asked for Places dataset. There was no error in the discussion in our paper (See Section 6.2.1 in our paper for more details). Norases Vesdapunt, Kedar Bellare, Nilesh N. Dalvi |
Proc. VLDB Endow. | 1 |
| 2014 | Crowdsourcing Algorithms for Entity ResolutionabstractIn this paper, we study a hybrid human-machine approach for solving the problem of Entity Resolution (ER). The goal of ER is to identify all records in a database that refer to the same underlying entity, and are therefore duplicates of each other. Our input is a graph over all the records in a database, where each edge has a probability denoting our prior belief (based on Machine Learning models) that the pair of records represented by the given edge are duplicates. Our objective is to resolve all the duplicates by asking humans to verify the equality of a subset of edges, leveraging the transitivity of the equality relation to infer the remaining edges (e.g. a = c can be inferred given a = b and b = c ). We consider the problem of designing optimal strategies for asking questions to humans that minimize the expected number of questions asked. Using our theoretical framework, we analyze several strategies, and show that a strategy, claimed as " optimal " for this problem in a recent work, can perform arbitrarily bad in theory. We propose alternate strategies with theoretical guarantees. Using both public datasets as well as the production system at Facebook, we show that our techniques are effective in practice. Norases Vesdapunt, Kedar Bellare, Nilesh N. Dalvi |
Proc. VLDB Endow. | 1 |
| 2011 | Scalable aggregation on multicore processorsabstractIn data-intensive and multi-threaded programming, the performance bottleneck has shifted from I/O bandwidth to main memory bandwidth. The availability, size, and other properties of on-chip cache strongly influence performance. A key question is whether to allow different threads to work independently, or whether to coordinate the shared workload among the threads. The independent approach avoids synchronization overhead, but requires resources proportional to the number of threads and thus is not scalable. On the other hand, the shared method suffers from coordination overhead and potential contention. Kenneth A. Ross, Norases Vesdapunt |
DaMoN | 3 |