Krithika Ramesh

dblp:255/2136 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 36% Efficient and distributed learning · 36% Language models and text generation · 16%
Human-computer interaction and pervasive computing
1 paper
Learning and educational technologies · 100%

Topics — the 6 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › fairness
bias in language models
0.712023
A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language Models · ACL (1) 2023
Machine learning › Trustworthy machine learning
fairness
0.712023
A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language Models · ACL (1) 2023
Machine learning › Efficient and distributed learning
model compression
0.712023
A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language Models · ACL (1) 2023
Machine learning › Efficient and distributed learning › model compression
pruning and quantization
0.712023
A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language Models · ACL (1) 2023
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.612022
'Beach' to 'Bitch': Inadvertent Unsafe Transcription of Kids' Content on YouTube · AAAI 2022
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.212023
A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language Models · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

language model correction · 1.1quantization · 0.7pruning · 0.7human evaluation · 0.7distillation · 0.7benchmarking · 0.7
YearPublicationVenuePosition
2023 A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language Models
abstract
Compression techniques for deep learning have become increasingly popular, particularly in settings where latency and memory constraints are imposed.Several methods, such as pruning, distillation, and quantization, have been adopted for compressing models, each providing distinct advantages.However, existing literature demonstrates that compressing deep learning models could affect their fairness.Our analysis involves a comprehensive evaluation of pruned, distilled, and quantized language models, which we benchmark across a range of intrinsic and extrinsic metrics for measuring bias in text classification.We also investigate the impact of using multilingual models and evaluation measures.Our findings highlight the significance of considering both the pre-trained model and the chosen compression strategy in developing equitable language technologies.The results also indicate that compression strategies can have an adverse effect on fairness measures.
Krithika Ramesh, Arnav Chavan, Shrey Pandit, Sunayana Sitaram
ACL (1)1
2023 MEGA: Multilingual Evaluation of Generative AI
abstract
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, Sunayana Sitaram. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Akshay Uttama Nambi, Tanuja Ganu, Sameer Segal, Kalika Bali, Sunayana Sitaram
EMNLP5
2023 End-to-end Privacy Preserving Training and Inference for Air Pollution Forecasting with Data from Rival Fleets
abstract
Privacy-preserving machine learning (PPML) promises to train machine learning (ML) models by combining data spread across multiple data silos. Theoretically, secure multiparty computation (MPC) allows multiple data owners to train models on their joint data without revealing the data to each other. However, the prior implementations of this secure training using MPC have three limitations: they have only been evaluated on CNNs, and LSTMs have been ignored; fixed point approximations have affected training accuracies compared to training in floating point; and due to significant latency overheads of secure training via MPC, its relevance for practical tasks with streaming data remains unclear. The motivation of this work is to report our experience of addressing the practical problem of secure training and inference of models for urban sensing problems, e.g., traffic congestion estimation, or air pollution monitoring in large cities, where data can be contributed by rival fleet companies while balancing the privacy-accuracy trade-offs using MPC-based techniques.Our first contribution is to design a custom ML model for this task that can be efficiently trained with MPC within a desirable latency. In particular, we design a GCN-LSTM and securely train it on time-series sensor data for accurate forecasting, within 7 minutes per epoch. As our second contribution, we build an end-to-end system of private training and inference that provably matches the training accuracy of cleartext ML training. This work is the first to securely train a model with LSTM cells. Third, this trained model is kept secret-shared between the fleet companies and allows clients to make sensitive queries to this model while carefully handling potentially invalid queries. Our custom protocols allow clients to query predictions from privately trained models in milliseconds, all the while maintaining accuracy and cryptographic security.
Gauri Gupta, Krithika Ramesh, Anwesh Bhattacharya, Divya Gupta 0001, Rahul Sharma 0001, Nishanth Chandran, Rijurekha Sen
Proc. Priv. Enhancing Technol.2
2022 'Beach' to 'Bitch': Inadvertent Unsafe Transcription of Kids' Content on YouTube
abstract
Over the last few years, YouTube Kids has emerged as one of the highly competitive alternatives to television for children's entertainment. Consequently, YouTube Kids' content should receive an additional level of scrutiny to ensure children's safety. While research on detecting offensive or inappropriate content for kids is gaining momentum, little or no current work exists that investigates to what extent AI applications can (accidentally) introduce content that is inappropriate for kids. In this paper, we present a novel (and troubling) finding that well-known automatic speech recognition (ASR) systems may produce text content highly inappropriate for kids while transcribing YouTube Kids' videos. We dub this phenomenon as inappropriate content hallucination. Our analyses suggest that such hallucinations are far from occasional, and the ASR systems often produce them with high confidence. We release a first-of-its-kind data set of audios for which the existing state-of-the-art ASR systems hallucinate inappropriate content for kids. In addition, we demonstrate that some of these errors can be fixed using language models.
Krithika Ramesh, Ashiqur R. KhudaBukhsh
AAAI1
2019 Out of the Closet: Lexicon Based Sentiment Analysis on Tweets about Homosexuality
abstract
Social media sites such as Twitter have given a platform to people to express their thoughts even on the most sensitive of topics. In this paper, we analyze tweets which are posted by people who acknowledged their sexuality and shared their friends and families' response publicly. In our experiment, we have used a set of Tweets data with specific hashtags to analyze the response. Context-based topic modeling was used to gather tweets surrounding this topic to get the relevant tweets. The unwanted tweets discovered by the topic model was discarded. The sentiment of these tweets was then extracted using novel methods. The motivation for this process was to find out how their declaration of sexuality was perceived. Unfortunately, homophobia is one of the battles humanity continues to fight.
Tanvi Anand, Krithika Ramesh, Sanjay Singh 0002
TENCON2