Harry Halpin

dblp:23/1741 · DBLP profile ↗
← Back
33ranked-venue papers
22as first author
7since 2021 · last 2026
0000-0003-2143-6965ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 17 · 12 first-author · 2 since 2021Security and privacy · 8 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Decentralized Reliability Estimation for Low-Latency Mixnets
Claudia Díaz, Harry Halpin, Aggelos Kiayias
EuroS&P2
2026 Website fingerprinting on Nym: Attacks and Defenses
abstract
Website fingerprinting (WF) enables a passive eavesdropper to infer which web page a client is visiting, even when communications are encrypted or anonymized. In this paper, we study the vulnerability to website fingerprinting of Nym, a mix network based on the Loopix design that enables users to browse the Web. We show that although Nym adds delays and cover traffic to change packet patterns compared to Tor, it still leaks features that website fingerprinting attacks can exploit in both closed‑ and open‑world settings. We carry out an in-depth analysis of the effectiveness of Nym's obfuscation mechanisms, originally designed to provide anonymity in messaging, in thwarting website fingerprinting. We show that mix delays, counterintuitively, not only fail to protect against website fingerprinting but actually make the attack more effective as the mix delays make it easier to distinguish incoming from outgoing packets. We also demonstrate that the current cover traffic strategy of Nym is not effective in thwarting website fingerprinting attacks unless it imposes a large overhead. To address these limitations, we design two new WF defenses based on Nym's existing obfuscation mechanisms that significantly reduce WF effectiveness. The first defense introduces cover traffic to match the bursty nature of real-world web traffic, reducing the F1 score to 0.39 (compared to 0.65 obtained by similar defenses applied on Tor) at moderate overhead increase. The second defense plummets the F1 score to 0.06 by channeling web traffic via Nym's constant traffic capabilities, at the cost of bandwidth.
Eric Jollès, Simon Wicky, Ania M. Piotrowska, Harry Halpin, Carmela Troncoso
Proc. Priv. Enhancing Technol.4
2024 Decentralizing Democracy with Semantic Information Technology: The D-CENT Retrospective
Harry Halpin
WEBIST1
2022 A Critique of EU Digital COVID-19 Certificates: Do Vaccine Passports Endanger Privacy?
abstract
Do COVID-19 vaccine passports come at a fundamental cost for personal privacy? Reviewing proposed COVID-19 credentials from a security and privacy standpoint raises concerns that make deploying COVID-19 digital certificates difficult at best. A closer look into the privacy of the EU Digital COVID-19 certificate presents a fundamental contradiction between two essential security properties: unforgeability and privacy. A substantial reconsideration of the very concept of vaccine passports may be needed to preserve fundamental privacy rights.
Harry Halpin
ARES1
2022 All that is Solid Melts into Air: Towards Decentralized Cryptographic Access Control
abstract
Access control languages are traditionally based on centralized trust models when achieving their security goals. One important reason for a lack of decentralized trust models for access control has been difficulties in referring to and accessing cryptographic key material in a decentralized fashion. This lack of a working and secure decentralized public key infrastructure could be fatal for projects like Berners-Lee’s Solid. To solve this problem, we propose repurposing the access control scheme SDSI (Simple Distributed Security Infrastructure). SDSI is a long-standing and well-studied alternative to the hierarchical PKI infrastructure, but it failed to be adopted due to problems with key discovery and revocation, both of which can be solved with blockchain technology. In this concept note, we outline the next steps for how SDSI could be used as a decentralized trust model using blockchain technology.
Harry Halpin
ARES1
2022 The Knowledge Trust: A Proposal for a Blockchain Consortium for Digital Archives
Harry Halpin
TPDL1
2022 From Indymedia to Tahrir Square: The Revolutionary Origins of Status Updates on Twitter
abstract
One of the most important developments in the history of the Web was the development of the status update. Although social media has been approached by a number of critical theorists as an instrument of control and surveillance, it should be remembered that social media began as liberatory technology harnessed by social movements. In this essay, we trace the origin of the status update for spreading news from protest-driven community networks like Indymedia and text messages for protest coordination via TxtMob. In fact, the use of status updates by Indymedia and the wider anti-globalization movement prefigured their usage in Tahrir Square and the Black Lives Matter movement in the USA. This historical link goes through Twitter itself, as the early Twitter engineers were veterans of Indymedia. There is still much to learn from Indymedia: Framing social media as invented and then harnessed by social movements may even provide innovative solutions to issues of content moderation and censorship. Exploring the origin of social media in social movements provides a perspective on the history of the Web from the tradition of the oppressed.
Harry Halpin, Evan Henshaw-Plath
WWW1
2020 SoK: why Johnny can't fix PGP standardization
abstract
Pretty Good Privacy (PGP) has long been the primary IETF standard for encrypting email, but suffers from widespread usability and security problems that have limited its adoption. As time has marched on, the underlying cryptographic protocol has fallen out of date insofar as PGP is unauthenticated on a per message basis and compresses before encryption. There have been an increasing number of attacks on the increasingly outdated primitives and complex clients used by the PGP eco-system. However, attempts to update the OpenPGP standard have failed at the IETF except for adding modern cryptographic primitives. Outside of official standardization, Autocrypt is a "bottom-up" community attempt to fix PGP, but still falls victim to attacks on PGP involving authentication. The core reason for the inability to "fix" PGP is the lack of a simple AEAD interface which in turn requires a decentralized public key infrastructure to work with email. Yet even if standards like MLS replace PGP, the deployment of a decentralized PKI remains an open issue.
Harry Halpin
ARES1
2017 NEXTLEAP: Decentralizing Identity with Privacy for Secure Messaging
abstract
Identity systems today link users to all of their actions and serve as centralized points of control and data collection. NEXTLEAP proposes an alternative decentralized and privacy-enhanced architecture. First, NEXTLEAP is building privacy-enhanced federated identity systems, using blind signatures based on Algebraic MACs to improve OpenID Connect. Second, secure messaging applications ranging from Signal to WhatsApp may deliver the content in an encrypted form, but they do not protect the metadata of the message and they rely on centralized servers. The EC Project NEXTLEAP is focussed on fixing these two problems by decentralizing traditional identities onto a privacy-enhanced based blockchain that can then be used to build access control lists in a decentralized manner, similar to SDSI. Furthermore, we improve on secure messaging by then using this notion of decentralized identity to build in group messaging, allowing messaging between different servers. NEXTLEAP is also working with the PANORAMIX EC project to use a generic mix networking infrastructure to hide the metadata of the messages themselves and plans to add privacy-enhanced data analytics that work in a decentralized manner.
Harry Halpin
ARES1
2017 Privacy-Preserving Data Dissemination in Untrusted Cloud
abstract
B2B (business-to-business) systems often use service-oriented architecture (SOA) with decomposed business services. These services can interact and share data among each other. Service might use a cloud – hosted database, such as a non - relational encrypted key – value store. However, the cloud platform hosting the database can be untrusted. Data owner needs to be sure that each service can access only those segments of a shared database for which the service is authorized. Furthermore, data requests can come from a service also hosted by untrusted cloud. Hence, there is a need for designing a cloud enterprise framework that can ensure privacy-preserving data dissemination in SOA and accurately detect data leakages. We design and prototype a solution that ensures privacy – preserving dissemination of data. The solution is based on (a) role-based access control, (b) cryptographic capabilities of client's browser, (c) authentication method, (d) subject's trust level. The prototype enables privacy – preserving dissemination of Electronic Health Records (EHRs) hosted in an untrusted cloud.
Denis A. Ulybyshev, Bharat K. Bhargava, Miguel Villarreal-Vasquez, Aala Oqab Alsalem, Donald Steiner, Leon Li, Jason Kobes, Harry Halpin, Rohit Ranchal
CLOUD8
2017 Systematizing Decentralization and Privacy: Lessons from 15 Years of Research and Deployments
abstract
Decentralized systems are a subset of distributed systems where multiple authorities control different components and no authority is fully trusted by all. This implies that any component in a decentralized system is potentially adversarial. We revise fifteen years of research on decentralization and privacy, and provide an overview of key systems, as well as key insights for designers of future systems. We show that decentralized designs can enhance privacy, integrity, and availability but also require careful trade-offs in terms of system complexity, properties provided, and degree of decentralization. These trade-offs need to be understood and navigated by designers. We argue that a combination of insights from cryptography, distributed systems, and mechanism design, aligned with the development of adequate incentives, are necessary to build scalable and successful privacy-preserving decentralized systems.
Carmela Troncoso, Marios Isaakidis, George Danezis, Harry Halpin
Proc. Priv. Enhancing Technol.4
2016 LEAP: A Next-Generation Client VPN and Encrypted Email Provider
Elijah Sparrow, Harry Halpin, Kali Kaneko, Ruben Pollan
CANS2
2015 Crowdmapping Digital Social Innovation with Linked Data
abstract
The European Commission recently became interested in mapping digital social innovation in Europe. In order to understand this rapidly developing if little known area, a visual and interactive survey was made in order to crowd-source a map of digital social innovation, available at http://digitalsocial.eu . Over 900 organizations participated, and Linked Data was used as the backend with a number of valuable advantages. The data was processed using SPARQL and network analysis, and a number of concrete policy recommendations resulted from the analysis.
Harry Halpin, Francesca Bria
ESWC1
2014 Dynamic Provenance for SPARQL Updates
Harry Halpin, James Cheney
ISWC (1)1
2013 An experimental analysis of suggestions in collaborative tagging
abstract
Most tagging systems support the user in the tag selection process by providing tag suggestions, or recommendations, based on a popularity measurement of tags other users provided when tagging the same resource, such as web-page. In this paper we inv
Dirk G. F. M. Bollen, Harry Halpin
Web Intell. Agent Syst.2
2013 Repeatable and reliable semantic search evaluation
Roi Blanco, Harry Halpin, Daniel M. Herzig, Peter Mika, Jeffrey Pound, Henry S. Thompson, Thanh Tran 0001
J. Web Semant.2
2011 Relevance Feedback between Web Search and the Semantic Web
abstract
We investigate the possibility of using structured data to improve search over unstructured documents. In particular, we use relevance feedback to create a 'virtuous cycle' between structured data from the Semantic Web and web-pages from the hypertext Web. Previous approaches have generally considered searching over the Semantic Web and hypertext Web to be entirely disparate, indexing and searching over different domains. Our novel approach is to use relevance feedback from hypertext Web results to improve Semantic Web search, and results from the Semantic Web to improve the retrieval of hypertext Web data. In both cases, our evaluation is based on certain kinds of informational queries (abstract concepts, people, and places) selected from a real-life query log and checked by human judges. We show our relevance model-based system is better than the performance of real-world search engines for both hypertext and Semantic Web search, and we also investigate Semantic Web inference and pseudo-relevance feedback.
Harry Halpin, Victor Lavrenko
IJCAI1
2011 Repeatable and reliable search system evaluation using crowdsourcing
abstract
The primary problem confronting any new kind of search task is how to boot-strap a reliable and repeatable evaluation campaign, and a crowd-sourcing approach provides many advantages. However, can these crowd-sourced evaluations be repeated over long periods of time in a reliable manner? To demonstrate, we investigate creating an evaluation campaign for the semantic search task of keyword-based ad-hoc object retrieval. In contrast to traditional search over web-pages, object search aims at the retrieval of information from factual assertions about real-world objects rather than searching over web-pages with textual descriptions. Using the first large-scale evaluation campaign that specifically targets the task of ad-hoc Web object retrieval over a number of deployed systems, we demonstrate that crowd-sourced evaluation campaigns can be repeated over time and still maintain reliable results. Furthermore, we show how these results are comparable to expert judges when ranking systems and that the results hold over different evaluation and relevance metrics. This work provides empirical support for scalable, reliable, and repeatable search system evaluation using crowdsourcing.
Roi Blanco, Harry Halpin, Daniel M. Herzig, Peter Mika, Jeffrey Pound, Henry S. Thompson, Thanh Tran 0001
SIGIR2
2011 Relevance feedback between hypertext and Semantic Web search: Frameworks and evaluation
Harry Halpin, Victor Lavrenko
J. Web Semant.1
2010 When owl: sameAs Isn't the Same: An Analysis of Identity in Linked Data
Harry Halpin, Patrick J. Hayes, Jamie P. McCusker, Deborah L. McGuinness, Henry S. Thompson
ISWC (1)1
2009 An Ontology of Resources: Solving the Identity Crisis
Harry Halpin, Valentina Presutti
ESWC1
2009 An Experimental Analysis of Suggestions in Collaborative Tagging
abstract
Most tagging systems support the user in the tag selection process by providing tag suggestions, or recommendations, based on a popularity measurement of tags other users provided when tagging the same resource, like a web-page. In this paper we investigate the influence of tag suggestions on the emergence of power-law distributions as a result of collaborative tag behavior. Although previous research has already shown that power-laws emerge in tagging systems, the cause of why power-law distributions emerge is not understood empirically. The majority of theories and mathematical models of tagging found in the literature assume that the emergence of power-laws in tagging systems is mainly driven by the imitation behavior of users when observing tag suggestions provided by the user interface of the tagging system. This imitation behavior leads to a feedback loop in which some tags are reinforced and get more popular which is also known as the `rich get richer' or a preferential attachment model. We present experimental results that show that the power-law distribution forms when tag suggestions are not presented to the users, and the power-law distribution does not hold when there are tag suggestions presented to the user. Furthermore, we show that the real effect of tag suggestions is rather subtle; the power-law distribution that would naturally occur without tag suggestions is `compressed' if tag suggestions are given to the user, resulting in a shorter long tail and a `compressed' top of the power-law distribution. The consequences of this experiment show that tag suggestions by themselves do not account for the formation of power-law distributions in tagging systems.
Dirk G. F. M. Bollen, Harry Halpin
Web Intelligence2
2009 Is there anything worth finding on the semantic web?
abstract
There has recently been an upsurge of interest in the possibilities of combining structured data and ad-hoc information retrieval from traditional hypertext. In this experiment, we run queries extracted from a query log of a major search engine against the Semantic Web to discover if the Semantic Web has anything of interest to the average user. We show that there is indeed much information on the Semantic Web that could be relevant for many queries for people, places and even abstract concepts, although they are overwhelmingly clustered around a Semantic Web-enabled export of Wikipedia known as DBPedia.
Harry Halpin
WWW1
2009 Emergence of consensus and shared vocabularies in collaborative tagging systems
abstract
This article uses data from the social bookmarking site del.icio.us to empirically examine the dynamics of collaborative tagging systems and to study how coherent categorization schemes emerge from unsupervised tagging by individual users. First, we study the formation of stable distributions in tagging systems, seen as an implicit form of “consensus” reached by the users of the system around the tags that best describe a resource. We show that final tag frequencies for most resources converge to power law distributions and we propose an empirical method to examine the dynamics of the convergence process, based on the Kullback-Leibler divergence measure. The convergence analysis is performed for both the most utilized tags at the top of tag distributions and the so-called long tail. Second, we study the information structures that emerge from collaborative tagging, namely tag correlation (or folksonomy) graphs. We show how community-based network techniques can be used to extract simple tag vocabularies from the tag correlation graphs by partitioning them into subsets of related tags. Furthermore, we also show, for a specialized domain, that shared vocabularies produced by collaborative tagging are richer than the vocabularies which can be extracted from large-scale query logs provided by a major search engine. Although the empirical analysis presented in this article is based on a set of tagging data obtained from del.icio.us, the methods developed are general, and the conclusions should be applicable across other websites that employ tagging.
Valentin Robu, Harry Halpin, Hana Shepherd
ACM Trans. Web2
2008 Exploring Semantic Social Networks Using Virtual Reality
Harry Halpin, David J. Zielinski, Rachael Brady, Glenda Kelly
ISWC1
2008 Redgraph: Navigating Semantic Web Networks using Virtual Reality
abstract
We present Redgraph, a generic virtual reality (VR) visualization program for network data, based on resource description framework (RDF), the primary standard of data underlying the semantic Web. Redgraph bypasses a number of problems in 3D graph visualization by relying on users to interactively "extrude" a 2D network into the third dimension. This pilot study applies Redgraph to data from the U.S. Patent and Trademark Office to explore innovations in the history of computer science. Comparison of subjects' response times utilizing 3D pull-out vs. 2D strategies on tasks involving fine grained connectivity or broad network observation found that subjects were faster correctly answering questions involving fine-grained connectivity using 3D strategies, particularly when data was densely clustered. Subjects' qualitative feedback suggests that the most valuable application of this 3D technique lies in untimed exploration to discover relationships in the underlying data structure.
Harry Halpin, David J. Zielinski, Rachael Brady, Glenda Kelly
VR1
2008 In Defense of Ambiguity
abstract
URIs, a universal identification scheme, are different from human names insofar as they can provide the ability to reliably access the thing identified. URIs also can function to reference a non-accessible thing in a similar manner to how names function in natural language. There are two distinctly different relationships between names and things: access and reference. To confuse the two relations leads to underlying problems with Web architecture. Reference is by nature ambiguous in any language. So any attempts by Web architecture to make reference completely unambiguous will fail on the Web. Despite popular belief otherwise, making further ontological distinctions often leads to more ambiguity, not less. Contrary to appeals to Kripke for some sort of eternal and unique identification, reference on the Web uses descriptions and therefore there is no unambiguous resolution of reference. On the Web, what is needed is not just a simple redirection, but a uniform and logically consistent manner of associating descriptions with URIs that can be done in a number of practical ways that should be made consistent.
Patrick J. Hayes, Harry Halpin
Int. J. Semantic Web Inf. Syst.2
2007 The complex dynamics of collaborative tagging
abstract
The debate within the Web community over the optimal means by which to organize information often pits formalized classifications against distributed collaborative tagging systems. A number of questions remain unanswered, however, regarding the nature of collaborative tagging systems including whether coherent categorization schemes can emerge from unsupervised tagging by users. This paper uses data from the social bookmarking site delicio. us to examine the dynamics of collaborative tagging systems. In particular, we examine whether the distribution of the frequency of use of tags for "popular" sites with a long history (many tags and many users) can be described by a power law distribution, often characteristic of what are considered complex systems. We produce a generative model of collaborative tagging in order to understand the basic dynamics behind tagging, including how a power law distribution of tags could arise. We empirically examine the tagging history of sites in order to determine how this distribution arises over time and to determine the patterns prior to a stable distribution. Lastly, by focusing on the high-frequency tags of a site where the distribution of tags is a stabilized power law, we show how tag co-occurrence networks for a sample domain of tags can be used to analyze the meaning of particular tags given their relationship to other tags.
Harry Halpin, Valentin Robu, Hana Shepherd
WWW1
2006 Event Extraction in a Plot Advice Agent
abstract
In this paper we present how the automatic extraction of events from text can be used to both classify narrative texts according to plot quality and produce advice in an interactive learning environment intended to help students with story writing. We focus on the story rewriting task, in which an exemplar story is read to the students and the students rewrite the story in their own words. The system automatically extracts events from the raw text, formalized as a sequence of temporally ordered predicate-arguments. These events are given to a machine-learner that produces a coarse-grained rating of the story. The results of the machine-learner and the extracted events are then used to generate fine-grained advice for the students.
Harry Halpin, Johanna D. Moore
ACL1
2006 Automatic Evaluation and Composition of NLP Pipelines with Web Services
Harry Halpin
LREC1
2006 From Typed-Functional Semantic Web Services to Proofs
Harry Halpin
ISWC1
2006 One document to bind them: combining XML, web services, and the semantic web
abstract
We present a paradigm for uniting the diverse strands of XML-based Web technologies by allowing them to be incorporated within a single document. This overcomes the distinction between programs and data to make XML truly "self-describing." A proposal for a lightweight yet powerful functional XML vocabulary called "Semantic fXML" is detailed, based on the well-understood functional programming paradigm and resembling the embedding of Lisp directly in XML. Infosets are made "dynamic," since documents can now directly embed local processes or Web Services into their Infoset. An optional typing regime for info-sets is provided by Semantic Web ontologies. By regarding Web Services as functions and the Semantic Web as providing types, and tying it all together within a single XML vocabulary, the Web can compute. In this light, the real Web 2.0 can be considered the transformation of the Web from a universal information space to a universal computation space.
Harry Halpin, Henry S. Thompson
WWW1
2004 Automatic Analysis of Plot for Story Rewriting
Harry Halpin, Johanna D. Moore, Judy Robertson
EMNLP1