Harshal Shah

dblp:200/4680 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0003-0047-4890ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Block-Accumulate Codes: Accelerated Linear Codes for PCGs and ZK
Vladimir Kolesnikov, Stanislav Peceny, Rahul Rachuri, Srinivasan Raghuraman, Peter Rindal, Harshal Shah
CRYPTO (8)6
2025 Efficient Permutation Correlations and Batched Random Access for Two-Party Computation
abstract
In this work we formalize the notion of a two-party permutation correlation $$(A, B), (C, \pi )$$ s.t. $$\pi (A)=B+C$$ for a random permutation $$\pi $$ of n elements and vectors $$A,B,C\in \mathbb {F}^n$$ . This correlation can be viewed as an abstraction and generalization of the Chase et al. (Asiacrypt 2020) share translation protocol. We give a systematization of knowledge for how such a permutation correlation can be derandomized to allow the parties to perform a wide range of oblivious permutations of secret-shared data. This systematization immediately enables the translation of various popular honest-majority protocols to be efficiently instantiated in the two-party setting, e.g. collaborative filtering, sorting, database joins, graph algorithms, and many more. We give two novel protocols for efficiently generating a random permutation correlation. The first uses MPC-friendly PRFs to generate a correlation of n elements, each of size $$\ell =\log |\mathbb {F}|$$ bits, with $$O(n\ell )$$ bit-OTs, time, communication, and only 3 rounds including setup. Similar asymptotics previously required relatively expensive public-key cryptography, e.g. Paillier or LWE. Our protocol implementation for $$n=2^{20},\ell =128$$ requires just 7 s & $$\sim 2\ell n$$ bits of communication, a respective 40 & $$1.1\times $$ improvement on the LWE solution of Juvekar at al. (CCS 2018). The second protocol is based on pseudo-random correlation generators and achieves an overhead that is sublinear in the string length $$\ell $$ , i.e. the communication and number of OTs is $$O(n\log \ell )$$ . The overhead of the latter protocol has larger hidden constants, and therefore is more efficient only when long strings are permuted, e.g. in graph algorithms. Finally, we present a suite of highly efficient protocols based on permutations for performing various batched random access operations. These include the ability to extract a hidden subset of a secret-shared list. More generally, we give ORAM-like protocols for obliviously reading and writing from a list in a batched manner. We argue that this suite of batched random access protocols should be a first class primitive in the MPC practitioner’s toolbox.(The authors grant IACR a non-exclusive and irrevocable license to distribute the article under the https://creativecommons.org/licenses/by-nc/3.0/ .)
Stanislav Peceny, Srinivasan Raghuraman, Peter Rindal, Harshal Shah
PKC (4)4
2024 Contextual classification of clinical records with bidirectional long short-term memory (Bi-LSTM) and bidirectional encoder representations from transformers (BERT) model
abstract
Abstract Deep learning models have overcome traditional machine learning techniques for text classification domains in the field of natural language processing (NLP). Since, NLP is a branch of machine learning, used for interpreting language, classifying text of interest, and the same can be applied to analyse the medical clinical electronic health records. Medical text consists of lot of rich data which can altogether provide a good insight, by determining patterns from the clinical text data. In this paper, bidirectional‐long short‐term memory (Bi‐LSTM), bi‐LSTM attention and bidirectional encoder representations from transformers (BERT) base models are used to classify the text which are of privacy concern to a person and which should be extracted and can be tagged as sensitive. This text data which we might think not of privacy concern would majorly reveal a lot about the patient's integrity and personal life. Clinical data not only have patient demographic data but lot of hidden data which might go unseen and thus could arise privacy issues. Bi‐LSTM with attention layer is also added on top to realize the importance of critical words which will be of great importance in terms of classification, we are able to achieve accuracy of about 92%. About 206,926 sentences are used out of which 80% are used for training and rest for testing we get accuracy of 90% approx. with Bi‐LSTM alone. The same set of datasets is used for BERT model with accuracy of 93% approx.
Jaya Zalte, Harshal Shah
Comput. Intell.2
2016 Cajun Codefest 4.0 on SMART-on-FHIR apps for Diabetes
Kavishwar B. Wagholikar, Eliel Oliveira, Henry Chu, Harshal Shah, Joshua C. Mandel, Jeffrey G. Klann, Sohail Rao, Kenneth D. Mandl, Shawn N. Murphy, Thomas Carton
AMIA5