VLDB 2026 Research / reviewers in the wild / expert
Carlos Cotrini Jiménez
dblp:150/0652 · also Carlos Cotrini
· DBLP profile ↗
14ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0001-8167-2284ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 8 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Theory of computation · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Steering AI Tutors Through System Prompts: A Crossover Study on Self-Regulated Learning and Cognitive Engagement Scaffolds in CS1abstractBackground. Large language models are increasingly deployed as tutors in introductory programming courses, yet evidence that they actually improve learning remains thin, and their tendency to shortcut productive struggle raises concerns about pedagogical harm. Self-regulated learning (SRL) and cognitive engagement (CE) frameworks offer a principled way to address this, but whether embedding them in system prompts actually changes how students learn is an open question. Maximilian Georg Barth, Sverrir Thorgeirsson, Khashayar Etemadi, Juho Leinonen 0001, Carlos Cotrini Jiménez, Zhendong Su 0001 |
ICER (1) | 5 |
| 2026 | PATHOS: A Pedagogical Method for Sequencing Instruction in Multi-Foundational Machine Learning
Diego Rivera Garrido, Sverrir Thorgeirsson, Damiano Meier, Luigi Pizza, Lahari Goswami, Jesus Solano, Carlos Cotrini Jiménez, Zhendong Su 0001 |
ICER (1) | 7 |
| 2026 | Transforming Confusion into Diffusion: Advancing Machine Learning Education via Bottom-Up InstructionabstractBalancing conceptual depth with practical skill development is a persistent challenge in advanced machine learning (ML) education, where powerful frameworks can obscure underlying mathematical and computational principles. To address this, we define a new principled approach that we call full-stack machine learning (FSML), which emphasizes the construction of large language models and diffusion models from scratch. To evaluate the effectiveness of FSML, we conducted a classroom-based randomized controlled trial (N=208) in which FSML-based instruction was compared against a popular library-based instructional approach. We measured students' conceptual understanding through a specialized assessment and administered a survey capturing knowledge-gap awareness, curiosity, and cognitive load. We found that students who received FSML instruction performed approximately 10% better than control participants in a quiz on transformers and stable diffusion (p=0.006). They also showed increased curiosity and more positive affective responses, suggesting deeper engagement with ML fundamentals. Our findings indicate that our full-stack approach to ML education can improve student learning outcomes, potentially reshaping curricula for ML and other advanced computing topics. Carlos Cotrini Jiménez, Sverrir Thorgeirsson, Jesus Solano, Zhendong Su 0001 |
SIGCSE (1) | 1 |
| 2024 | S-BDT: Distributed Differentially Private Boosted Decision TreesabstractWe introduce S-BDT: a novel (𝜀, 𝛿)-differentially private distributed gradient boosted decision tree (GBDT) learner that improves the protection of single training data points (privacy) while achieving meaningful learning goals, such as accuracy or regression error (utility).S-BDT uses less noise by relying on non-spherical multivariate Gaussian noise, for which we show tight subsampling bounds for privacy amplification and incorporate that into a Rényi filter for individual privacy accounting.We experimentally reach the same utility while saving 50% in terms of epsilon for 𝜀 ≤ 0.5 on the Abalone regression dataset (dataset size ≈ 4𝐾), saving 30% in terms of epsilon for 𝜀 ≤ 0.08 for the Adult classification dataset (dataset size ≈ 50𝐾), and saving 30% in terms of epsilon for 𝜀 ≤ 0.03 for the Spambase classification dataset (dataset size ≈ 5𝐾).Moreover, we show that for situations where a GBDT is learning a stream of data that originates from different subpopulations (non-IID), S-BDT improves the saving of epsilon even further. CCS Concepts• Theory of computation → Theory of database privacy and security; • Computing methodologies → Boosting; Online learning settings. Thorsten Peinemann, Moritz Kirschte, Joshua Stock, Carlos Cotrini Jiménez, Esfandiar Mohammadi |
CCS | 4 |
| 2024 | Automated Large-Scale Analysis of Cookie Notice Compliance
Ahmed Bouhoula, Karel Kubicek 0001, Amit Zac, Carlos Cotrini Jiménez, David A. Basin |
USENIX Security Symposium | 4 |
| 2023 | Invariant Anomaly Detection under Distribution Shifts: A Causal PerspectiveabstractAnomaly detection (AD) is the machine learning task of identifying highly discrepant abnormal samples by solely relying on the consistency of the normal training samples. Under the constraints of a distribution shift, the assumption that training samples and test samples are drawn from the same distribution breaks down. In this work, by leveraging tools from causal inference we attempt to increase the resilience of anomaly detection models to different kinds of distribution shifts. We begin by elucidating a simple yet necessary statistical property that ensures invariant representations, which is critical for robust AD under both domain and covariate shifts. From this property, we derive a regularization term which, when minimized, leads to partial distribution invariance across environments.
Through extensive experimental evaluation on both synthetic and real-world tasks, covering a range of six different AD methods, we demonstrated significant improvements in out-of-distribution performance. Under both covariate and domain shift, models regularized with our proposed term showed marked increased robustness. Code is available at: https://github.com/JoaoCarv/invariant-anomaly-detection João B. S. Carvalho, Mengtao Zhang, Robin C. Geyer, Carlos Cotrini Jiménez, Joachim M. Buhmann |
NeurIPS | 4 |
| 2023 | Locality-Sensitive Hashing Does Not Guarantee Privacy! Attacks on Google's FLoC and the MinHash Hierarchy SystemabstractRecently proposed systems aim at achieving privacy using locality-sensitive hashing. We show how these approaches fail by presenting attacks against two such systems: Google's FLoC proposal for privacy-preserving targeted advertising and the MinHash Hierarchy, a system for processing location trajectories in a privacy-preserving way. Our attacks refute the pre-image resistance, anonymity, and privacy guarantees claimed for these systems. In the case of FLoC, we show how to deanonymize users using Sybil attacks and to reconstruct 10% or more of the browsing history for 30% of its users using Generative Adversarial Networks. We achieve this only analyzing the hashes used by FLoC. For MinHash, we precisely identify the location trajectory of a subset of individuals and, on average, we can limit users' trajectory to just 10% of the possible geographic area, again using just the hashes. In addition, we refute their differential privacy claims. Florian Turati, Karel Kubicek 0001, Carlos Cotrini Jiménez, David A. Basin |
Proc. Priv. Enhancing Technol. | 3 |
| 2022 | Automating Cookie Consent and GDPR Violation Detection
Dino Bollinger, Karel Kubicek 0001, Carlos Cotrini Jiménez, David A. Basin |
USENIX Security Symposium | 3 |
| 2022 | Checking Websites' GDPR Consent Compliance for Marketing EmailsabstractAbstract The sending of marketing emails is regulated to protect users from unsolicited emails. For instance, the European Union’s ePrivacy Directive states that marketers must obtain users’ prior consent, and the General Data Protection Regulation (GDPR) specifies further that such consent must be freely given, specific, informed, and unambiguous. Based on these requirements, we design a labeling of legal characteristics for websites and emails. This leads to a simple decision procedure that detects potential legal violations. Using our procedure, we evaluated 1000 websites and the 5000 emails resulting from registering to these websites. Both datasets and evaluations are available upon request. We find that 21.9% of the websites contain potential violations of privacy and unfair competition rules, either in the registration process (17.3%) or email communication (17.7%). We demonstrate with a statistical analysis the possibility of automatically detecting such potential violations. Karel Kubicek 0001, Jakob Merane, Carlos Cotrini Jiménez, Alexander Stremitzer, Stefan Bechtold, David A. Basin |
Proc. Priv. Enhancing Technol. | 3 |
| 2019 | The Next 700 Policy Miners: A Universal Method for Building Policy MinersabstractA myriad of access control policy languages have been and continue to be proposed. The design of policy miners for each such language is a challenging task that has required specialized machine learning and combinatorial algorithms. We present an alternative method, universal access control policy mining (Unicorn). We show how this method streamlines the design of policy miners for a wide variety of policy languages including ABAC, RBAC, RBAC with user-attribute constraints, RBAC with spatio-temporal constraints, and an expressive fragment of XACML. For the latter two, there were no known policy miners until now. To design a policy miner using Unicorn, one needs a policy language and a metric quantifying how well a policy fits an assignment of permissions to users. From these, one builds the policy miner as a search algorithm that computes a policy that best fits the given permission assignment. We experimentally evaluate the policy miners built with Unicorn on logs from Amazon and access control matrices from other companies. Despite the genericity of our method, our policy miners are competitive with and sometimes even better than specialized state-of-the-art policy miners. The true positive rates of policies we mined differ by only 5% from the policies mined by the state of the art and the false positive rates are always below 5%. In the case of ABAC, it even outperforms the state of the art. Carlos Cotrini Jiménez, Luca Corinzia, Thilo Weghorn, David A. Basin |
CCS | 1 |
| 2018 | Mining ABAC Rules from Sparse LogsabstractDifferent methods have been proposed to mine attribute-based access control (ABAC) rules from logs. In practice, these logs are sparse in that they contain only a fraction of all possible requests. However, for sparse logs, existing methods mine and validate overly permissive rules, enabling privilege abuse. We define a novel measure, reliability, that quantifies how overly permissive a rule is and we show why other standard measures like confidence and entropy fail in quantifying overpermissiveness. We build upon state-of-the-art subgroup discovery algorithms and our new reliability measure to design Rhapsody, the first ABAC mining algorithm with correctness guarantees: Rhapsody mines a rule if and only if the rule covers a significant number of requests, its reliability is above a given threshold, and there is no equivalent shorter rule. We evaluate Rhapsody on different real-world scenarios using logs from Amazon and a computer lab at ETH Zurich. Our results show that Rhapsody generalizes better and produces substantially smaller rules than competing approaches. Carlos Cotrini Jiménez, Thilo Weghorn, David A. Basin |
EuroS&P | 1 |
| 2016 | Basic primal infon logicabstractPrimal infon logic (PIL) was introduced in 2009 in the framework of policy and trust management. In the meantime, some generalizations appeared, and there have been some changes in the syntax of the basic PIL. This article is on the basic PIL, and one of our purposes is to ‘institutionalize’ the changes. We prove a small-model theorem for the propositional fragment of basic primal infon logic (PPIL), give a simple proof of the PPIL locality theorem and present a linear-time decision algorithm (announced earlier) for PPIL in a form convenient for generalizations. For the sake of completeness, we cover the universal fragment of basic PIL. We wish that this article becomes a standard reference on basic PIL. Carlos Cotrini Jiménez, Yuri Gurevich |
J. Log. Comput. | 1 |
| 2015 | Analyzing First-Order Role Based Access ControlabstractWe propose FORBAC, an extension of Role-Based Access Control (RBAC) based on first-order logic. FORBAC is expressive enough to formalize a wide range of access control policies. However, it is simple enough so that relevant policy analysis queries can be analyzed in NP, which we argue is a natural complexity class for this problem. To analyze queries efficiently, we reduce them to the problem of satisfiability modulo appropriate theories, and use off-the-shelf SMT solvers. We evaluate FORBAC's expressiveness and our approach to policy analysis in a case study, analyzing access control in a European bank. Carlos Cotrini Jiménez, Thilo Weghorn, David A. Basin, Manuel Clavel |
CSF | 1 |
| 2014 | Deciding safety and liveness in TPTL
David A. Basin, Carlos Cotrini Jiménez, Felix Klaedtke, Eugen Zalinescu |
Inf. Process. Lett. | 2 |