EDBT 2026 Demo / reviewers in the wild / expert
Rishita Anubhai
dblp:119/7727
· DBLP profile ↗
7ranked-venue papers
0as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data integration and cleaning · 50% Knowledge graphs · 25% Information retrieval · 25% | |
| Artificial intelligence
4 papers |
Information extraction and text analysis · 35% Language models and text generation · 25% Probabilistic and Bayesian machine learning · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% | |
| Network and information security
1 paper |
Network security · 50% Cryptographic protocols and secure computation · 50% |
Topics — the 14 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
cross-modal retrieval |
0.8 | 1 | 2024 | BioBridge: Bridging Biomedical Foundation Models via Knowledge Graphs · ICLR 2024 |
Knowledge graphs
knowledge graph embedding |
0.8 | 1 | 2024 | BioBridge: Bridging Biomedical Foundation Models via Knowledge Graphs · ICLR 2024 |
Data integration and cleaning › data transformation
transformation discovery |
0.8 | 1 | 2024 | DATALORE: Can a Large Language Model Find All Lost Scrolls in a Data Repository? · ICDE 2024 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model |
0.5 | 1 | 2021 | Structured Prediction as Translation between Augmented Natural Languages · ICLR 2021 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.5 | 1 | 2021 | Structured Prediction as Translation between Augmented Natural Languages · ICLR 2021 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.4 | 1 | 2020 | To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence Tagging · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis › temporal information extraction
temporal ordering |
0.4 | 1 | 2020 | Severing the Edge Between Before and After: Neural Architectures for Temporal Ordering of Events · EMNLP (1) 2020 |
Machine learning › Deep learning architectures and training › neural network training
end-to-end deep learning |
0.2 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition |
0.2 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Bioinformatics and computational biology
drug discovery |
0.2 | 1 | 2024 | BioBridge: Bridging Biomedical Foundation Models via Knowledge Graphs · ICLR 2024 |
Cryptographic protocols and secure computation › key management › public key infrastructure
certificate validation |
0.1 | 1 | 2012 | The most dangerous code in the world: validating SSL certificates in non-browser software · CCS 2012 |
Network security › secure communication › secure communication protocol
TLS |
0.1 | 1 | 2012 | The most dangerous code in the world: validating SSL certificates in non-browser software · CCS 2012 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.1 | 1 | 2020 | To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence Tagging · EMNLP (1) 2020 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Methods — techniques the papers use, named apart from their topics
parameter-efficient learning · 1.5knowledge graph embedding · 1.5large language model · 0.8batch dispatch · 0.5augmented natural language · 0.5GPU-based inference · 0.5semi-supervised learning · 0.4neural architecture · 0.4BERT · 0.4empirical analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DATALORE: Can a Large Language Model Find All Lost Scrolls in a Data Repository?abstractHow can we effectively generate missing data transformations among tables in a data repository? Multiple versions of the same tables are generated from the iterative process when data scientists and machine learning engineers fine-tune their ML pipelines, making incremental improvements. This process often involves data transformation and augmentation that produces an augmented table based on its base version and related tables. However, data transformations are often not well-documented or completely missing, resulting in poor traceability, reproducibility and explainability of ML pipelines. In this paper, we propose DATALoRE, a framework that explains data changes between an initial dataset and its augmented version to improves traceability. Given a base table, DATALoRE first discovers its potentially related tables from the data repository using a variety of data discovery techniques. DATALoRE then effectively leverages a large language model (LLM) to generate a variety of data transformations that lead to the augmented table. DATALoRE validates these transformations and selects the minimum number of related tables to ensure traceability and reproducibility of the ML pipelines. A preliminary experiment shows that DATALoRE is able to effectively recovery data transformations on two benchmark datasets. Yuze Lou, Chuan Lei, Xiao Qin 0003, Zichen Wang 0002, Christos Faloutsos, Rishita Anubhai, Huzefa Rangwala |
ICDE | 6 |
| 2024 | BioBridge: Bridging Biomedical Foundation Models via Knowledge GraphsabstractFoundation models (FMs) learn from large volumes of unlabeled data to demonstrate superior performance across a wide range of tasks. However, FMs developed for biomedical domains have largely remained unimodal, i.e., independently trained and used for tasks on protein sequences alone, small molecule structures alone, or clinical data alone.
To overcome this limitation, we present BioBridge, a parameter-efficient learning framework, to bridge independently trained unimodal FMs to establish multimodal behavior. BioBridge achieves it by utilizing Knowledge Graphs (KG) to learn transformations between one unimodal FM and another without fine-tuning any underlying unimodal FMs.
Our results demonstrate that BioBridge can
beat the best baseline KG embedding methods (on average by ~ 76.3%) in cross-modal retrieval tasks. We also identify BioBridge demonstrates out-of-domain generalization ability by extrapolating to unseen modalities or relations. Additionally, we also show that BioBridge presents itself as a general-purpose retriever that can aid biomedical multimodal question answering as well as enhance the guided generation of novel drugs. Code is at https://github.com/RyanWangZf/BioBridge. Zifeng Wang 0008, Zichen Wang 0002, Vassilis N. Ioannidis, Huzefa Rangwala, Rishita Anubhai |
ICLR | 6 |
| 2021 | Structured Prediction as Translation between Augmented Natural Languages
Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Alessandro Achille, Rishita Anubhai, Cícero Nogueira dos Santos, Bing Xiang, Stefano Soatto |
ICLR | 6 |
| 2020 | Severing the Edge Between Before and After: Neural Architectures for Temporal Ordering of EventsabstractMiguel Ballesteros, Rishita Anubhai, Shuai Wang, Nima Pourdamghani, Yogarshi Vyas, Jie Ma, Parminder Bhatia, Kathleen McKeown, Yaser Al-Onaizan. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Miguel Ballesteros, Rishita Anubhai, Nima Pourdamghani, Yogarshi Vyas, Jie Ma 0005, Parminder Bhatia, Kathy McKeown, Yaser Al-Onaizan |
EMNLP (1) | 2 |
| 2020 | To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence TaggingabstractKasturi Bhattacharjee, Miguel Ballesteros, Rishita Anubhai, Smaranda Muresan, Jie Ma, Faisal Ladhak, Yaser Al-Onaizan. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Kasturi Bhattacharjee, Miguel Ballesteros, Rishita Anubhai, Smaranda Muresan, Jie Ma 0005, Faisal Ladhak, Yaser Al-Onaizan |
EMNLP (1) | 3 |
| 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and MandarinabstractWe show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale. Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu |
ICML | 3 |
| 2012 | The most dangerous code in the world: validating SSL certificates in non-browser softwareabstractSSL (Secure Sockets Layer) is the de facto standard for secure Internet communications. Security of SSL connections against an active network attacker depends on correctly validating public-key certificates presented when the connection is established. Martin Georgiev, Subodh Iyengar, Suman Jana, Rishita Anubhai, Dan Boneh, Vitaly Shmatikov |
CCS | 4 |