Shubham Jain 0006

dblp:132/6759-6 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-7978-9504ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2024 Did the Neurons Read your Book? Document-level Membership Inference for Large Language Models
Matthieu Meeus, Shubham Jain 0006, Marek Rei, Yves-Alexandre de Montjoye
USENIX Security Symposium2
2023 Deep perceptual hashing algorithms with hidden dual purpose: when client-side scanning does facial recognition
abstract
End-to-end encryption (E2EE) provides strong technical protections to individuals from interferences. Governments and law enforcement agencies around the world have however raised concerns that E2EE also allows illegal content to be shared undetected. Client-side scanning (CSS), using perceptual hashing (PH) to detect known illegal content before it is shared, is seen as a promising solution to prevent the diffusion of illegal content while preserving encryption. While these proposals raise strong privacy concerns, proponents of the solutions have argued that the risk is limited as the technology has a limited scope: detecting known illegal content. In this paper, we show that modern perceptual hashing algorithms are actually fairly flexible pieces of technology and that this flexibility could be used by an adversary to add a secondary hidden feature to a client-side scanning system. More specifically, we show that an adversary providing the PH algorithm can "hide" a secondary purpose of face recognition of a target individual alongside its primary purpose of image copy detection. We first propose a procedure to train a dual-purpose deep perceptual hashing model by jointly optimizing for both the image copy detection and the targeted facial recognition task. Second, we extensively evaluate our dual-purpose model and show it to be able to reliably identify a target individual 67% of the time while not impacting its performance at detecting illegal content. We also show that our model is neither a general face detection nor a facial recognition model, allowing its secondary purpose to be hidden. Finally, we show that the secondary purpose can be enabled by adding a single illegal looking image to the database. Taken together, our results raise concerns that a deep perceptual hashing-based CSS system could turn billions of user devices into tools to locate targeted individuals.
Shubham Jain 0006, Ana-Maria Cretu 0002, Antoine Cully, Yves-Alexandre de Montjoye
SP1
2022 Adversarial Detection Avoidance Attacks: Evaluating the robustness of perceptual hashing-based client-side scanning
Shubham Jain 0006, Ana-Maria Cretu 0002, Yves-Alexandre de Montjoye
USENIX Security Symposium1
2019 OPAL: High performance platform for large-scale privacy-preserving location data analytics
abstract
Mobile phones and other ubiquitous technologies are generating vast amounts of high-resolution location data. This data has been shown to have a great potential for the public good, e.g. to monitor human migration during crises or to predict the spread of epidemic diseases. Location data is, however, considered one of the most sensitive types of data, and a large body of research has shown the limits of traditional data anonymization methods for big data. Privacy concerns have so far strongly limited the use of location data collected by telcos, especially in developing countries.In this paper, we introduce OPAL (for OPen ALgorithms), an open-source, scalable, and privacy-preserving platform for location data. At its core, OPAL relies on an open algorithm to extract key aggregated statistics from location data for a wide range of potential use cases. We first discuss how we designed the OPAL platform, building a modular and resilient framework for efficient location analytics. We then describe the layered mechanisms we have put in place to protect privacy and discuss the example of a population density algorithm. We finally evaluate the scalability and extensibility of the platform and discuss related work.The code will be open-sourced on GitHub upon publication.
Axel Oehmichen, Shubham Jain 0006, Andrea Gadotti, Yves-Alexandre de Montjoye
IEEE BigData2
2019 UNVEIL: Capture and Visualise WiFi Data Leakages
abstract
In the past few years, numerous privacy vulnerabilities have been discovered in the WiFi standards and their implementations for mobile devices. These vulnerabilities allow an attacker to collect large amounts of data on the device user, which could be used to infer sensitive information such as religion, gender, and sexual orientation. Solutions for these vulnerabilities are often hard to design and typically require many years to be widely adopted, leaving many devices at risk.
Shubham Jain 0006, Eden Bensaid, Yves-Alexandre de Montjoye
WWW1
2018 Deep Learning Techniques for Automatic MRI Cardiac Multi-Structures Segmentation and Diagnosis: Is the Problem Solved?
abstract
Delineation of the left ventricular cavity, myocardium, and right ventricle from cardiac magnetic resonance images (multi-slice 2-D cine MRI) is a common clinical task to establish diagnosis. The automation of the corresponding tasks has thus been the subject of intense research over the past decades. In this paper, we introduce the "Automatic Cardiac Diagnosis Challenge" dataset (ACDC), the largest publicly available and fully annotated dataset for the purpose of cardiac MRI (CMR) assessment. The dataset contains data from 150 multi-equipments CMRI recordings with reference measurements and classification from two medical experts. The overarching objective of this paper is to measure how far state-of-the-art deep learning methods can go at assessing CMRI, i.e., segmenting the myocardium and the two ventricles as well as classifying pathologies. In the wake of the 2017 MICCAI-ACDC challenge, we report results from deep learning methods provided by nine research groups for the segmentation task and four groups for the classification task. Results show that the best methods faithfully reproduce the expert analysis, leading to a mean value of 0.97 correlation score for the automatic extraction of clinical indices and an accuracy of 0.96 for automatic diagnosis. These results clearly open the door to highly accurate and fully automatic analysis of cardiac CMRI. We also identify scenarios for which deep learning methods are still failing. Both the dataset and detailed results are publicly available online, while the platform will remain open for new submissions.
Olivier Bernard 0001, Alain Lalande, Clément Zotti, Frederic Cervenansky, Xin Yang 0009, Pheng-Ann Heng, Irem Cetin, Karim Lekadir, Oscar Camara 0001, Miguel Ángel González Ballester, Gerard Sanroma, Sandy Napel, Steffen E. Petersen, Georgios Tziritas, Ilias Grinias, Mahendra Khened, Alex Varghese, Ganapathy Krishnamurthi, Marc-Michel Rohé, Xavier Pennec, Maxime Sermesant, Fabian Isensee, Paul F. Jaeger, Klaus H. Maier-Hein, Peter M. Full, Ivo Wolf, Sandy Engelhardt, Christian F. Baumgartner, Lisa M. Koch, Jelmer M. Wolterink, Ivana Isgum, Yeonggul Jang, Yoonmi Hong, Jay Patravali, Shubham Jain 0006, Olivier Humbert, Pierre-Marc Jodoin
IEEE Trans. Medical Imaging35