EDBT 2026 Demo / reviewers in the wild / expert
Neil O'Hare
dblp:63/134
· DBLP profile ↗
24ranked-venue papers
6as first author
3since 2021 · last 2024
0009-0001-7499-7814ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 11 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-authorArtificial intelligence and machine learning · 7 · 2 since 2021Human-computer interaction and ubiquitous computing · 5Applied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multilingual Taxonomic Web Page Categorization Through Ensemble Knowledge DistillationabstractWeb page categorization has been extensively studied in the literature and has been successfully used to improve information retrieval, recommendation, personalization and ad targeting. With the new industry trend of not tracking users' online behavior without their explicit permission, using contextual targeting to accurately understand web pages in order to display ads that are topically relevant to the pages becomes more important. This is challenging, however, because an ad request only contains the URL of a web page. As a result, there is very limited available text for making accurate classifications. In this paper, we propose a unified multilingual model that can seamlessly classify web pages in 5 high-impact languages using either their full content or just their URLs with limited text. We adopt multiple data sampling techniques to increase coverage for rare categories in our training corpus, and modify the loss using class-based re-weighting to smooth the influence of frequent versus rare categories. We also propose using an ensemble of teacher models for knowledge distillation and explore different ways to create a teacher ensemble. Offline evaluation shows at least 2.6% improvement in mean average precision across 5 languages compared to a URL classification model trained with single-teacher knowledge distillation. The unified model for both full-content and URL-only input further improves the mean average precision of the dedicated URL classification model by 0.6%. We launched the proposed models, which achieve at least 37% better mean average precision than the legacy tree-based models, for contextual targeting in the Yahoo Demand Side Platform, leading to a significant ad delivery and revenue increase. Eric Ye, Xiao Bai 0002, Neil O'Hare, Eliyar Asgarieh, Kapil Thadani, Francisco Perez-Sorrosal, Sujyothi Adiga |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Content-Based Email Classification at ScaleabstractUnderstanding the content of email messages can enable new features that highlight what matters to users, making email a more useful tool for people to manage their lives. We present work from a consumer email platform to build multilabel models to classify messages according to a mail-specific, content-based taxonomy that represents the topic, type, and objective of an email. While state-of-the-art Transformer-based language models can achieve impressive results for text classification, these models are too costly to deploy at the scale of email. Using a knowledge distillation framework, we first build a complex, accurate teacher model from limited human-labeled training data and then use a large amount of teacher-labeled data to train lightweight student models that are suitable for deployment. The student models retain up to 91% of the predictive performance of the teacher model while reducing inference cost by three orders of magnitude. Deployed to production in Yahoo Mail, these models classify billions of emails every day and power features that help people tackle their inboxes. Kirstin Early, Neil O'Hare, Christopher LuVogt |
CIKM | 2 |
| 2022 | Multilingual Taxonomic Web Page Classification for Contextual Targeting at YahooabstractAs we move toward a cookie-less world, the ability to track users' online activities for behavior targeting will be drastically reduced, making contextual targeting an appealing alternative for advertising platforms. Category-based contextual targeting displays ads on web pages that are relevant to advertiser-targeted categories, according to a pre-defined taxonomy. Accurate web page classification is key to the success of this approach. In this paper, we use multilingual Transformer-based transfer learning models to classify web pages in five high-impact languages. We adopt multiple data sampling techniques to increase coverage for rare categories, and modify the loss using class-based re-weighting to smooth the influence of frequent versus rare categories. Offline evaluation shows that these are crucial for improving our classifiers. We leverage knowledge distillation to train accurate models that are lightweight in terms of (i) model size, and (ii) the input text used. Classifying web pages using only text from the URL addresses a unique challenge for contextual targeting in that bid requests come to ad systems as URLs without content, while crawling is time consuming and costly. We launched the proposed models for contextual targeting in the Yahoo DSP, significantly increasing its revenue. Eric Ye, Xiao Bai 0002, Neil O'Hare, Eliyar Asgarieh, Kapil Thadani, Francisco Perez-Sorrosal, Sujyothi Adiga |
KDD | 3 |
| 2020 | Effective Few-Shot Classification with Transfer LearningabstractFew-shot learning addresses the the problem of learning based on a small amount of training data. Although more well-studied in the domain of computer vision, recent work has adapted the Amazon Review Sentiment Classification (ARSC) text dataset for use in the few-shot setting. In this work, we use the ARSC dataset to study a simple application of transfer learning approaches to few-shot classification. We train a single binary classifier to learn all few-shot classes jointly by prefixing class identifiers to the input text. Given the text and class, the model then makes a binary prediction for that text/class pair. Our results show that this simple approach can outperform most published results on this dataset. Surprisingly, we also show that including domain information as part of the task definition only leads to a modest improvement in model accuracy, and zero-shot classification, without further fine-tuning on few-shot domains, performs equivalently to few-shot classification. These results suggest that the classes in the ARSC few-shot task, which are defined by the intersection of domain and rating, are actually very similar to each other, and that a more suitable dataset is needed for the study of few-shot text classification. Aakriti Gupta, Kapil Thadani, Neil O'Hare |
COLING | 3 |
| 2017 | Bridging the Aesthetic Gap: The Wild Beauty of Web ImageryabstractTo provide good results, image search engines need to rank not just the most relevant images, but also the highest quality images. To surface beautiful pictures, existing computational aesthetic models are trained with datasets from photo contest websites, dominated by professional photos. Such models fail completely in real web scenarios, where images are extremely diverse in terms of quality and type (e.g. drawings, clip-art, etc). This work aims at bridging and understanding this "aesthetic gap". We collect a dataset of around 100K web images with `quality' and `type' (photo vs non-photo) annotations. We design a set of visual features to describe image pictorial characteristics, and deeply analyse the peculiar beauty of web images as opposed to appealing professional images. Finally, we build a set of computational aesthetic frameworks based on deep learning and hand-crafted features that take into account the diverse quality of web images, and show that they significantly outperform traditional computational aesthetics methods on our dataset. Miriam Redi, Frank Z. Liu, Neil O'Hare |
ICMR | 3 |
| 2017 | Understanding and Discovering Deliberate Self-harm Content in Social MediaabstractStudies suggest that self-harm users found it easier to discuss self-harm-related thoughts and behaviors using social media than in the physical world. Given the enormous and increasing volume of social media data, on-line self-harm content is likely to be buried rapidly by other normal content. To enable voices of self-harm users to be heard, it is important to distinguish self-harm content from other types of content. In this paper, we aim to understand self-harm content and provide automatic approaches to its detection. We first perform a comprehensive analysis on self-harm social media using different input cues. Our analysis, the first of its kind in large scale, reveals a number of important findings. Then we propose frameworks that incorporate the findings to discover self-harm content under both supervised and unsupervised settings. Our experimental results on a large social media dataset from Flickr demonstrate the effectiveness of the proposed frameworks and the importance of our findings in discovering self-harm content. Yilin Wang 0002, Jiliang Tang, Jundong Li, Baoxin Li, Yali Wan, Clayton Mellina, Neil O'Hare, Yi Chang 0001 |
WWW | 7 |
| 2016 | Application of Natural User Interface Devices for Touch-Free Control of Radiological Images During SurgeryabstractNatural User Interface (NUI) systems can enable the scrubbed clinician to assume direct control of medical image interaction while maintaining sterility in the Operating Room. Surgeons and radiologists trialed a touch-free image control system based on the Leap Motion and Microsoft Kinect v2 controllers. Feedback was reported on the perceived utility and usability of both devices. The speed and accuracy of the two controllers was measured. Results showed marginal to average acceptability of both controllers. Surgeons and Interventional Radiologists found Microsoft Kinect to have better utility and to be potentially useful for the majority (54%) of them. The accuracy of the Leap Motion sensor was superior and comparable with that of a computer mouse. A link was established between the system usability and the perception of utility with better usability translating into better utility. Advantages and limitations of each device are highlighted. Design improvements and deployment considerations are discussed. Nikola Nestorov, Peter Hughes, Nuala Healy, Niall Sheehy, Neil O'Hare |
CBMS | 5 |
| 2016 | Leveraging User Interaction Signals for Web Image SearchabstractUser interfaces for web image search engine results differ significantly from interfaces for traditional (text) web search results, supporting a richer interaction. In particular, users can see an enlarged image preview by hovering over a result image, and an `image preview' page allows users to browse further enlarged versions of the results, and to click-through to the referral page where the image is embedded. No existing work investigates the utility of these interactions as implicit relevance feedback for improving search ranking, beyond using clicks on images displayed in the search results page. In this paper we propose a number of implicit relevance feedback features based on these additional interactions: hover-through rate, 'converted-hover' rate, referral page click through, and a number of dwell time features. Also, since images are never self-contained, but always embedded in a referral page, we posit that clicks on other images that are embedded on the same referral webpage as a given image can carry useful relevance information about that image. We also posit that query-independent versions of implicit feedback features, while not expected to capture topical relevance, will carry feedback about the quality or attractiveness of images, an important dimension of relevance for web image search. In an extensive set of ranking experiments in a learning to rank framework, using a large annotated corpus, the proposed features give statistically significant gains of over 2% compared to a state of the art baseline that uses standard click features. Neil O'Hare, Paloma de Juan, Rossano Schifanella, Dawei Yin 0001, Yi Chang 0001 |
SIGIR | 1 |
| 2016 | Predicting celebrity attendees at public events using stock photo metadata
Xin Shuai, Neil O'Hare, Luca Maria Aiello, Alejandro Jaimes |
Multim. Tools Appl. | 2 |
| 2015 | A Large-Scale Study of User Image Search Behavior on the WebabstractIn this study, we analyze user image search behavior from a large-scale Yahoo! Image Search query log, based on the hypothesis that behavior is dependent on query type. We categorize queries using two orthogonal taxonomies (subject-based and facet-based) and identify important query types at the intersection of these taxonomies. We study user search behavior on a large-scale set of search sessions for each query type, examining characteristics of sessions, query reformulation patterns, click patterns, and page view patterns. We identify important behavioral differences across query types, in particular showing that some query types are more exploratory, while others correspond to focused search. We also supplement our study with a survey to link the behavioral differences to users' intent. Our findings shed light on the importance of considering query categories to better understand user behavior on image search platforms. Jaimie Yejean Park, Neil O'Hare, Rossano Schifanella, Alejandro Jaimes, Chin-Wan Chung |
CHI | 2 |
| 2014 | 6 Seconds of Sound and Vision: Creativity in Micro-videosabstractThe notion of creativity, as opposed to related concepts such as beauty or interestingness, has not been studied from the perspective of automatic analysis of multimedia content. Meanwhile, short online videos shared on social media platforms, or micro-videos, have arisen as a new medium for creative expression. In this paper we study creative micro-videos in an effort to understand the features that make a video creative, and to address the problem of automatic detection of creative content. Defining creative videos as those that are novel and have aesthetic value, we conduct a crowdsourcing experiment to create a dataset of over 3, 800 micro-videos labelled as creative and non-creative. We propose a set of computational features that we map to the components of our definition of creativity, and conduct an analysis to determine which of these features correlate most with creative video. Finally, we evaluate a supervised approach to automatically detect creative video, with promising results, showing that it is necessary to model both aesthetic value and novelty to achieve optimal classification accuracy. Miriam Redi, Neil O'Hare, Rossano Schifanella, Michele Trevisiol, Alejandro Jaimes |
CVPR | 2 |
| 2013 | Search behaviour on photo sharing platformsabstractThe behaviour, goals, and intentions of users while searching for images in large scale online collections are not well understood, with image search log analysis providing limited insights, in part because they tend only to have access to user search and result click information. In this paper we study user search behaviour in a large photo-sharing platform, analyzing all user actions during search sessions (i.e. including post result-click pageviews). Search accounts for a significant part of user interactions with such platforms, and we show differences between the queries issued on such platforms and those on general image search. We show that search behaviour is influenced by the query type, and also depends on the user. Finally, we analyse how users behave when they reformulate their queries, and develop URL class prediction models for image search, showing that query-specific models significantly outperform query-agnostic models. The insights provided in this paper are intended as a launching point for the design of better interfaces and ranking models for image search. Silviu Maniu, Neil O'Hare, Luca Maria Aiello, Luca Chiarandini, Alejandro Jaimes |
ICME | 2 |
| 2013 | Competition-based networks for expert findingabstractFinding experts in question answering platforms has important applications, such as question routing or identification of best answers. Addressing the problem of ranking users with respect to their expertise, we propose Competition-Based Expertise Networks (CBEN), a novel community expertise network structure based on the principle of competition among the answerers of a question. We evaluate our approach on a very large dataset from Yahoo! Answers using a variety of centrality measures. We show that it outperforms state-of-the-art network structures and, unlike previous methods, is able to consistly outperform simple metrics like best answer count. We also analyse question answering forums in Yahoo! Answers, and show that they can be characterised by factual or subjective information seeking behavior, social discussions and the conducting of polls or surveys. We find that the ability to identify experts greatly depends on the type of forum, which is directly reflected in the structural properties of the expertise networks. Çigdem Aslay, Neil O'Hare, Luca Maria Aiello, Alejandro Jaimes |
SIGIR | 2 |
| 2013 | Modeling locations with social media
Neil O'Hare, Vanessa Murdock 0001 |
Inf. Retr. | 1 |
| 2012 | An Investigation of Term Weighting Approaches for Microblog Retrieval
Paul Ferguson, Neil O'Hare, James Lanagan, Owen Phelan, Kevin McCarthy |
ECIR | 2 |
| 2012 | Predicting participants in public events using stock photosabstractPictures taken by journalists for distribution and for inclusion in stock photo collections are often enriched with metadata. One key aspect of such photos is that they focus largely on events and feature celebrities and other public figures. They may provide interesting insights into how such public figures are related to each other in terms of the events they attend, and in their social proximity in terms of how often they are photographed together. In this paper, we study a corpus of approximately 9 million stock photographs taken over a 10 year period and, using their metadata, we extract a social network from co-appearance of public figures in events depicted in the photographs. We exploit this latent social information and combine it with the rich image metadata to explore the possibility of predicting attendees at future events, showing promising performance for this task. Neil O'Hare, Luca Maria Aiello, Alejandro Jaimes |
ACM Multimedia | 1 |
| 2012 | Efficient Storage and Decoding of SURF Feature Points
Kevin McGuinness, Kealan McCusker, Neil O'Hare, Noel E. O'Connor |
MMM | 3 |
| 2010 | Coping With Noise in a Real-World Weblog Crawler and Retrieval System
James Lanagan, Paul Ferguson, Neil O'Hare, Alan F. Smeaton |
ICWSM | 3 |
| 2009 | Combining Social Network Analysis and Sentiment Analysis to Explore the Potential for Online RadicalisationabstractThe increased online presence of jihadists has raised the possibility of individuals being radicalised via the Internet. To date, the study of violent radicalisation has focused on dedicated jihadist websites and forums. This may not be the ideal starting point for such research, as participants in these venues may be described as "already made-up minds". Crawling a global social networking platform, such as YouTube, on the other hand, has the potential to unearth content and interaction aimed at radicalisation of those with little or no apparent prior interest in violent jihadism. This research explores whether such an approach is indeed fruitful. We collected a large dataset from a group within YouTube that we identified as potentially having a radicalising agenda. We analysed this data using social network analysis and sentiment analysis tools, examining the topics discussed and what the sentiment polarity (positive or negative) is towards these topics. In particular, we focus on gender differences in this group of users, suggesting most extreme and less tolerant views among female users. Adam Bermingham, Maura Conway, Lisa McInerney, Neil O'Hare, Alan F. Smeaton |
ASONAM | 4 |
| 2009 | Context-Aware Person Identification in Personal Photo CollectionsabstractIdentifying the people in photos is an important need for users of photo management systems. We present MediAssist, one such system which facilitates browsing, searching and semi-automatic annotation of personal photos, using analysis of both image content and the context in which the photo is captured. This semi-automatic annotation includes annotation of the identity of people in photos. In this paper, we focus on such person annotation, and propose person identification techniques based on a combination of context and content. We propose language modelling and nearest neighbor approaches to context-based person identification, in addition to novel face color and image color content-based features (used alongside face recognition and body patch features). We conduct a comprehensive empirical study of these techniques using the real private photo collections of a number of users, and show that combining context- and content-based analysis improves performance over content or context alone. Neil O'Hare, Alan F. Smeaton |
IEEE Trans. Multim. | 1 |
| 2007 | The use of artificial neural networks to stratify the length of stay of cardiac patients based on preoperative and initial postoperative factors
Michael Rowan, Thomas Ryan, Francis Hegarty, Neil O'Hare |
Artif. Intell. Medicine | 4 |
| 2005 | Mobile access to personal digital photograph archivesabstractHandheld computing devices are becoming highly connected devices with high capacity storage. This has resulted in their being able to support storage of, and access to, personal photo archives. However the only means for mobile device users to browse such archives is typically a simple one-by-one scroll through image thumbnails in the order that they were taken, or by manually organising them based on folders. In this paper we describe a system for context-based browsing of personal digital photo archives. Photos are labeled with the GPS location and time they are taken and this is used to derive other context-based metadata such as weather conditions and daylight conditions. We present our prototype system for mobile digital photo retrieval, and an experimental evaluation illustrating the utility of location information for effective personal photo retrieval. Cathal Gurrin, Gareth J. F. Jones, Hyowon Lee 0001, Neil O'Hare, Alan F. Smeaton, Noel Murphy |
Mobile HCI | 4 |
| 2005 | My digital photos: where and when?abstractIn recent years digital cameras have seen an enormous rise in popularity, leading to a huge increase in the quantity of digital photos being taken. This brings with it the challenge of organising these large collections. We preset work which organises personal digital photo collections based on date/time and GPS location, which we believe will become a key organisational methodology over the next few years as consumer digital cameras evolve to incorporate GPS and as cameras in mobile phones spread further. The accompanying video illustrates the results of our research into digital photo management tools which contains a series of screen and user interactions highlighting how a user utilises the tools we are developing to manage a personal archive of digital photos. Neil O'Hare, Cathal Gurrin, Hyowon Lee 0001, Noel Murphy, Alan F. Smeaton, Gareth J. F. Jones |
ACM Multimedia | 1 |
| 2004 | A generic news story segmentation system and its evaluationabstractThe paper presents an approach to segmenting broadcast TV news programmes automatically into individual news stories. We first segment the programme into individual shots, and then a number of analysis tools are run on the programme to extract features to represent each shot. The results of these feature extraction tools are then combined using a support vector machine trained to detect anchorperson shots. A news broadcast can then be segmented into individual stories based on the location of the anchorperson shots within the programme. We use one generic system to segment programmes from two different broadcasters, illustrating the robustness of our feature extraction process to the production styles of different broadcasters. Neil O'Hare, Alan F. Smeaton, Csaba Czirjek, Noel E. O'Connor, Noel Murphy |
ICASSP (3) | 1 |