Nir Grinberg

dblp:133/1786 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-1277-894XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CUPCase: Clinically Uncommon Patient Cases and Diagnoses Dataset
abstract
Medical benchmark datasets significantly contribute to developing Large Language Models (LLMs) for medical knowledge extraction, diagnosis, summarization, and other uses. Yet, current benchmarks are mainly derived from exam questions given to medical students or cases described in the medical literature, lacking the complexity of real-world patient cases that deviate from classic textbook abstractions. These include rare diseases, uncommon presentations of common diseases, and unexpected treatment responses. Here, we construct Clinically Uncommon Patient Cases and Diagnosis Dataset (CUPCase) based on 3,563 real-world case reports from BMC, which we formulate into diagnoses in open-ended textual format and as multiple-choice options with distractors. Using this dataset, we evaluate the ability of state-of-the-art LLMs, including both general-purpose and Clinical LLMs, to identify and correctly diagnose a patient case, and test models' performance when only partial information about cases is available. Our findings show that general-purpose GPT-4o attains the best performance in both the multiple-choice task (average accuracy of 87.9%) and the open-ended task (BERTScore F1 of 0.764), outperforming several LLMs with a focus on the medical domain such as Meditron-70B and MedLM-Large. Moreover, GPT-4o was able to maintain 87% and 88% of its performance with only the first 20% of tokens of the case presentation in multiple-choice and free text, respectively, highlighting the potential of LLMs to aid in early diagnosis in real-world cases. An error analysis demonstrates the complexity of the task, and attempts to hypothesise about the models' reasoning. CUPCase expands our ability to evaluate LLMs for clinical decision support in an open and reproducible manner.
Oriel Perets, Ofir Ben Shoham, Nir Grinberg, Nadav Rappoport
AAAI3
2024 EnronSR: A Benchmark for Evaluating AI-Generated Email Replies
abstract
Human-to-human communication is no longer just mediated by computers, it is increasingly generated by them, including on popular communication platforms such as Gmail, Facebook Messenger, Linkedin, and others. Yet, little is known about the differences between human- and machine-generated responses in complex social settings. Here, we present EnronSR, a novel benchmark dataset that is based on the Enron email corpus and contains both naturally occurring human- and AI-generated email replies for the same set of messages. This resource enables the benchmarking of novel language-generation models in a public and reproducible manner, and facilitates a comparison against the strong, production-level baseline of Google Smart Reply used by millions of people. Moreover, we show that when language models produce responses they could align more closely with human replies in terms of when responses should be offered, their length, sentiment, and semantic meaning. We further demonstrate the utility of this benchmark in a case study of GPT-3, showing significantly better alignment with human responses than Smart Reply, albeit providing no guarantees for quality or safety.
Shay Moran, Roei Davidson, Nir Grinberg
ICWSM3
2024 280 Characters to Employment: Using Twitter to Quantify Job Vacancies
abstract
Accurate assessment of workforce needs is critical for designing well-informed economic policy and improving market efficiency. While surveys are the gold standard for estimating when and where workers are needed, they also have important limitations, most notably their substantial costs, dependence on existing and extensive surveying infrastructure, and limited temporal, geographical, and sectorial resolution. Here, we investigate the potential of social media to provide a complementary signal for estimating labor market demand. We introduce a novel statistical approach for extracting information about the location and occupation advertised in job vacancies posted on Twitter. We then construct an aggregate index of labor market demand by occupational class in every major U.S. city from 2015 to 2022, which we evaluate against two sources of official statistics and an index from a large aggregator of online job postings. We find that the newly constructed index is strongly correlated with official statistics and, in some cases, advantageous compared to statistics from job aggregators. Moreover, we demonstrate that our index can robustly improve the prediction of official statistics across occupations and states.
Boris Sobol, Manuel Tonneau, Samuel P. Fraiberger, Do Lee, Nir Grinberg
ICWSM5
2024 Leveraging Exposure Networks for Detecting Fake News Sources
abstract
The scale and dynamic nature of the Web makes real-time detection of misinformation an extremely difficult task. Prior research mostly focused on offline (retrospective) detection of stories or claims using linguistic features of the content, flagging by users, and crowdsourced labels. Here, we develop a novel machine-learning methodology for detecting fake news sources using active learning, and examine the contribution of network, audience, and text features to the model accuracy. Importantly, we evaluate performance in both offline and online settings, mimicking the strategic choices fact-checkers have to make in practice as news sources emerge over time. We find that exposure networks provide information on considerably more sources than sharing networks (+49.6%), and that the inclusion of exposure features greatly improves classification PR-AUC in both offline (+33%) and online (+69.2%) settings. Textual features perform best in offline settings, but their performance deteriorates by 12.0-18.7% in online settings. Finally, the results show that a few iterations of active learning are sufficient for our model to attain predictive performance to comparable exhaustive labeling while incurring only 24.7% of the labeling costs. These results stress the importance of exposure networks as a source of valuable information for the investigation of information dissemination in social networks and question the robustness of textual features.
Maor Reuben, Lisa Friedland, Rami Puzis, Nir Grinberg
KDD4
2022 Multilingual Detection of Personal Employment Status on Twitter
abstract
Detecting disclosures of individuals' employment status on social media can provide valuable information to match job seekers with suitable vacancies, offer social protection, or measure labor market flows.However, identifying such personal disclosures is a challenging task due to their rarity in a sea of social media content and the variety of linguistic forms used to describe them.Here, we examine three Active Learning (AL) strategies in real-world settings of extreme class imbalance, and identify five types of disclosures about individuals' employment status (e.g.job loss) in three languages using BERT-based classification models.Our findings show that, even under extreme imbalance settings, a small number of AL iterations is sufficient to obtain large and significant gains in precision, recall, and diversity of results compared to a supervised baseline with the same number of labels.We also find that no AL strategy consistently outperforms the rest.Qualitative analysis suggests that AL helps focus the attention mechanism of BERT on core terms and adjust the boundaries of semantic expansion, highlighting the importance of interpretable models to provide greater control and visibility into this dynamic learning process.
Manuel Tonneau, Dhaval Adjodah, João Palotti, Nir Grinberg, Samuel P. Fraiberger
ACL (1)4
2018 Identifying Modes of User Engagement with Online News and Their Relationship to Information Gain in Text
abstract
Prior work established the benefits of server-recorded user engagement measures (e.g. clickthrough rates) for improving the results of search engines and recommendation systems. Client-side measures of post-click behavior received relatively little attention despite the fact that publishers have now the ability to measure how millions of people interact with their content at a fine resolution using client-side logging. In this study, we examine patterns of user engagement in a large, client-side log dataset of over 7.7 million page views (including both mobile and non-mobile devices) of 66,821 news articles from seven popular news publishers. For each page view we use three summary statistics: dwell time, the furthest position the user reached on the page, and the amount of interaction with the page through any form of input (touch, mouse move, etc.). We show that simple transformations on these summary statistics reveal six prototypical modes of reading that range from scanning to extensive reading and persist across sites. Furthermore, we develop a novel measure of information gain in text to capture the development of ideas within the body of articles and investigate how information gain relates to the engagement with articles. Finally, we show that our new measure of information gain is particularly useful for predicting reading of news articles before publication, and that the measure captures unique information not available otherwise.
Nir Grinberg
WWW1
2017 Understanding Feedback Expectations on Facebook
abstract
When people share updates with their friends on Facebook they have varying expectations for the feedback they will receive. In this study, we quantitatively examine the factors contributing to feedback expectations and the potential outcomes of expectation fulfillment. We conducted two sets of surveys: one asking people about their feedback expectations immediately after posting on Facebook and the other asking how the amount of feedback received on a post matched the participant's expectations. Participants were more likely to expect feedback on content they evaluated as more important, and to a lesser extent more personal. Expectations also depended on participants' age, gender, and level of activity on Facebook. When asked about feedback expectations from specific friends, participants were more likely to expect feedback from closer friends, but expectations varied considerably based on recency of communication, geographical proximity, and the type of relationship (e.g. family, co-worker). Finally, receiving more feedback relative to expectations correlated with a greater feeling of connectedness to one's Facebook friends. The findings suggest implications for the theory and the design of social network sites.
Nir Grinberg, Shankar Kalyanaraman, Lada A. Adamic, Mor Naaman
CSCW1
2016 Changes in Engagement Before and After Posting to Facebook
abstract
The asynchronous nature of communications on social network sites creates a unique opportunity for studying how posting content interacts with individuals' engagement. This study focuses on the behavioral changes occurring hours before and after contribution to better understand the changing needs and preferences of contributors. Using observational data analysis of individuals' activity on Facebook, we test hypotheses regarding the motivations for site visits, changes in the distribution of attention to content, and shifts in decisions to interact with others. We find that after posting content people are intrinsically motivated to visit the site more often, are more attentive to content from friends (but not others), and choose to interact more with friends (in large part due to reciprocity). In addition, contributors are more active on the site hours before posting and remain more active for less than a day afterwards. Our study identifies a unique pattern of engagement that accompanies contribution and can inform the design of social network sites to better support contributors.
Nir Grinberg, P. Alex Dow, Lada A. Adamic, Mor Naaman
CHI1
2013 Extracting Diurnal Patterns of Real World Activity from Social Media
Nir Grinberg, Mor Naaman, Blake Shaw, Gilad Lotan
ICWSM1