VLDB 2026 Research / reviewers in the wild / expert
Nikhil Jha
dblp:289/2565
· DBLP profile ↗
9ranked-venue papers
6as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 4 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Effective Predictive Modelling for Emergency Department Visits and Evaluating Exogenous Variables Impact: Using Explainable Meta-Learning Gradient BoostingabstractAccurately predicting Emergency Department (ED) visits is essential for optimising resource allocation, including staffing adjustments and Operating Room scheduling. Despite the proliferation of AI-driven models, effective ED visit prediction remains challenging due to limited generalisability, susceptibility to overfitting and underfitting, scalability and the complexity of fine-tuning hyper-parameters. To address these challenges, we propose a novel Meta-Learning Gradient Booster (Meta-ED) approach to forecast daily ED visits. Meta-ED leverages a comprehensive dataset spanning 23 years from Canberra Hospital, incorporating exogenous variables such as socio-demographic characteristics, healthcare usage, chronic diseases, diagnoses and climate parameters. Meta-ED combines four foundational learners—CatBoost, Random Forest, Extra Trees and LightGBM—with a Multi-Layer Perceptron (MLP) as the master-level learner, thereby enhancing predictive precision by integrating the strengths of diverse base models. Our comparative analysis, which involved testing 23 models against a set of predefined criteria, demonstrates the superior performance of Meta-ED, achieving an accuracy of 85.7% (95% CI [85.4%, 86.0%]) and outperforming prominent models like XGBoost, Random Forest, AdaBoost, LightGBM and Extra Trees by up to 106.3%. Furthermore, incorporating climate features resulted in a 3.25% improvement in prediction accuracy, effectively capturing seasonal variations that influence patient volumes. These results underscore the potential of Meta-ED to advance predictive analytics in complex healthcare environments. Mehdi Neshat, Michael Phipps, Nikhil Jha, Danial Khojasteh, Michael Tong, Amir Hossein Gandomi |
ACM Trans. Comput. Heal. | 3 |
| 2025 | Privacy Policies and Consent Management Platforms: Growth and Users' Interactions over TimeabstractIn response to growing concerns about user privacy, legislators have introduced new regulations and laws, such as the General Data Protection Regulation (GDPR) in the European Union and the California Consumer Privacy Act (CCPA), which force websites to obtain user consent before activating any personal data collection. The cornerstone of this consent-seeking process involves the use of Privacy Banners, the technical tools to collect users’ approval for data collection practices. Consent management platforms (CMPs) have emerged as practical solutions to simplify the configuration and management of such privacy banners for website administrators, allowing them to outsource the complexities of managing user consent and activating advertising features. This article presents a detailed and longitudinal analysis of the evolution of CMPs spanning 9 years. We take a twofold perspective: firstly, thanks to the HTTP Archive dataset, we provide insights into the growth, market share, and geographical spread of CMPs. Noteworthy observations include the substantial impact of the GDPR on the proliferation of CMPs in Europe, where more than 40% of websites currently adopt a CMP. Secondly, we analyse millions of user interactions with a medium-sized CMP present in thousands of websites worldwide. We observe how even small changes in the design of Privacy Banners have a critical impact on the user’s giving or denying one’s consent to data collection. For instance, over 60% of users do not consent when offered a simple “one-click reject-all” option. Conversely, when opting out requires more than one click, about 90% of users prefer to simply give their consent. This hints that their main objective is to eliminate the annoying privacy banner rather than make an informed decision. Curiously, we observe that iOS users exhibit a higher tendency to accept cookies compared with Android users, possibly indicating greater confidence in the privacy offered by Apple devices. We believe that the findings of this article contribute to a deeper understanding of the multifaceted interactions between privacy regulations, technological solutions and user choices in the evolving Web ecosystem. We also show that the availability of large open datasets, although not explicitly designed and collected for our goals, is fundamental to exploring different angles of the internet evolution over time. For this, we make the data and code used in this work available to the community. 1 Nikhil Jha, Martino Trevisan, Marco Mellia, Daniel Fernandez, Rodrigo Irarrazaval |
ACM Trans. Web | 1 |
| 2024 | NeCTAr and RASoC: Tale of Two Class SoCs for Language Model Interference and Robotics in Intel 16abstractThis paper introduces NeCTAr (Near-Cache Transformer Accelerator), a 16nm heterogeneous multicore RISC-V SoC for sparse and dense machine learning kernels with both near-core and near-memory accelerators. A prototype chip runs at 400MHz at 0.85V and performs matrix-vector multiplications with 109 GOPs/W. The effectiveness of the design is demonstrated by running inference on a sparse language model, ReLU-Llama. Viansa Schmulbach, Ethan Gao, Nikhil Jha, Ethan Wu, Oliver Yu, Ben Oliveau, Brendan Roberts, Connor McMahon, Lixiang Yin, Vamber Yang, Brendan Brenner, George Moujaes, Boyu Hao, Lucy Revina, Bryan Ngo, Yufeng Chi, Hongyi Huang, Reza Sajadiany, Raghav Gupta 0001, Ella Schwarz, Jennifer Zhou, Ken Ho, Jerry Zhao, Anita Flynn, Borivoje Nikolic |
HCS | 4 |
| 2024 | Re-Identification Attacks against the Topics APIabstractRecently, Google proposed the Topics API framework as a privacy-friendly alternative for behavioural advertising as a possible solution to balance user’s privacy and advertisement effectiveness. Using the Topics API, the browser builds a user profile based on navigation history, which advertisers can access. The Topics API aim at becoming the new standard for behavioural advertising, thus it is necessary to fully understand its operation and find possible limitations. In this article, we evaluate the robustness of the Topics API to a re-identification attack. To build a user profile, we suppose an attacker accumulates over time the topics a user exposes to different websites. The attacker later re-identifies the same user matching the profiles of their audience. We leverage real traffic traces and realistic population models, and we present increasingly powerful attack threats. We find that the Topics API mitigates but cannot prevent re-identification from taking place, as there is a sizeable chance that a user’s profile remains unique within a website’s audience and the attacker successfully matches it with the profile of the same user on a second website. Depending on environmental factors, the probability of correct re-identification can reach 50%, considering a pool of 1,000 users. We offer the code and data we use in this work to stimulate further studies and the tuning of the Topic API parameters. 1 Nikhil Jha, Martino Trevisan, Emilio Leonardi, Marco Mellia |
ACM Trans. Web | 1 |
| 2023 | FogROS2: An Adaptive Platform for Cloud and Fog Robotics Using ROS 2abstractMobility, power, and price points often dictate that robots do not have sufficient computing power on board to run contemporary robot algorithms at desired rates. Cloud computing providers such as AWS, GCP, and Azure offer immense computing power and increasingly low latency on demand, but tapping into that power from a robot is non-trivial. We present FogROS2, an open-source platform to facilitate cloud and fog robotics that is included in the Robot Operating System 2 (ROS 2) distribution. FogROS2 is distinct from its predecessor FogROS1 in 9 ways, including lower latency, overhead, and startup times; improved usability, and additional automation, such as region and computer type selection. Additionally, FogROS2 gains performance, timing, and additional improvements associated with ROS 2. In common robot applications, FogROS2 reduces SLAM latency by 50 %, reduces grasp planning time from 14 s to 1.2 s, and speeds up motion planning 45x. When compared to FogROS1, FogROS2 reduces network utilization by up to 3.8x, improves startup time by 63 %, and network round-trip latency by 97 % for images using video compression. The source code, examples, and documentation for FogROS2 are available at https://github.com/BerkeleyAutomation/FogROS2, and is available through the official ROS 2 repository at https://index.ros.org/p/FogROS2/. Jeffrey Ichnowski, Kaiyuan Chen 0001, Karthik Dharmarajan, Simeon Adebola, Michael Danielczuk, Victor Mayoral Vilches, Nikhil Jha, Hugo Zhan, Edith LLontop, Derek Xu, Camilo Buscaron, John Kubiatowicz, Ion Stoica, Joseph Gonzalez 0001, Kenneth Y. Goldberg |
ICRA | 7 |
| 2023 | Practical anonymization for data streams: z-anonymity and relation with k-anonymity
Nikhil Jha, Luca Vassio, Martino Trevisan, Emilio Leonardi, Marco Mellia |
Perform. Evaluation | 1 |
| 2023 | On the Robustness of Topics API to a Re-Identification AttackabstractWeb tracking through third-party cookies is considered a threat to users' privacy and is supposed to be abandoned in the near future. Recently, Google proposed the Topics API framework as a privacy-friendly alternative for behavioural advertising. Using this approach, the browser builds a user profile based on navigation history, which advertisers can access. The Topics API has the possibility of becoming the new standard for behavioural advertising, thus it is necessary to fully understand its operation and find possible limitations. This paper evaluates the robustness of the Topics API to a re-identification attack where an attacker reconstructs the user profile by accumulating user's exposed topics over time to later re-identify the same user on a different website. Using real traffic traces and realistic population models, we find that the Topics API mitigates but cannot prevent re-identification to take place, as there is a sizeable chance that a user's profile is unique within a website's audience. Consequently, the probability of correct re-identification can reach 15-17%, considering a pool of 1,000 users. We offer the code and data we use in this work to stimulate further studies and the tuning of the Topic API parameters. Nikhil Jha, Martino Trevisan, Emilio Leonardi, Marco Mellia |
Proc. Priv. Enhancing Technol. | 1 |
| 2022 | The Internet with Privacy Policies: Measuring The Web Upon ConsentabstractTo protect user privacy, legislators have regulated the use of tracking technologies, mandating the acquisition of users’ consent before collecting data. As a result, websites started showing more and more consent management modules–i.e., Consent Banners–the visitors have to interact with to access the website content. Since these banners change the content the browser loads, they challenge web measurement collection, primarily to monitor the extent of tracking technologies, but also to measure web performance. If not correctly handled, Consent Banners prevent crawlers from observing the actual content of the websites. In this paper, we present a comprehensive measurement campaign focusing on popular websites in Europe and the US, visiting both landing and internal pages from different countries around the world. We engineer Priv-Accept , a Web crawler able to accept the Consent Banners, as most users would do in practice. It lets us compare how webpages change before and after accepting such policies, if present. Our results show that all measurements performed ignoring the Consent Banners offer a biased and partial view of the Web. After accepting the privacy policies, web tracking is far more pervasive, and webpages are larger and slower to load. Nikhil Jha, Martino Trevisan, Luca Vassio, Marco Mellia |
ACM Trans. Web | 1 |
| 2020 | z-anonymity: Zero-Delay Anonymization for Data StreamsabstractWith the advent of big data and the birth of the data markets that sell personal information, individuals' privacy is of utmost importance. The classical response is anonymization, i.e., sanitizing the information that can directly or indirectly allow users' re-identification. The most popular solution in the literature is the k-anonymity. However, it is hard to achieve k-anonymity on a continuous stream of data, as well as when the number of dimensions becomes high.In this paper, we propose a novel anonymization property called z-anonymity. Differently from k-anonymity, it can be achieved with zero-delay on data streams and it is well suited for high dimensional data. The idea at the base of z-anonymity is to release an attribute (an atomic information) about a user only if at least z - 1 other users have presented the same attribute in a past time window. z-anonymity is weaker than k-anonymity since it does not work on the combinations of attributes, but treats them individually. In this paper, we present a probabilistic framework to map the z-anonymity into the k-anonymity property. Our results show that a proper choice of the z-anonymity parameters allows the data curator to likely obtain a k-anonymized dataset, with a precisely measurable probability. We also evaluate a real use case, in which we consider the website visits of a population of users and show that z-anonymity can work in practice for obtaining the k-anonymity too. Nikhil Jha, Thomas Favale, Luca Vassio, Martino Trevisan, Marco Mellia |
IEEE BigData | 1 |