Shishir G. Patil

dblp:251/1906 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-6835-388XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 62% Efficient and distributed learning · 25% Reinforcement learning · 10%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Embedded and real-time systems · 70% Cloud and datacenter computing · 30%
Computer networks
3 papers
Internet of things and sensor networks · 49% Content delivery and video streaming · 31% Network optimization and economics · 11%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Human-computer interaction and pervasive computing
1 paper
Accessibility and assistive technology · 50% Interaction techniques and input · 50%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 19 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › LLM agents
tool use
1.622025
The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models · ICML 2025
Gorilla: Large Language Model Connected with Massive APIs · NeurIPS 2024
Natural language and speech › Language models and text generation
instruction following
1.012026
AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following · ACL (1) 2026
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.012026
AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following · ACL (1) 2026
Embedded and real-time systems › resource-constrained computing › resource-constrained embedded system
deeply embedded systems
1.012026
Efficient ML Model Updates for Deeply Embedded Microcontrollers · EuroSys 2026
Natural language and speech › Language models and text generation › agentic language model › tool-augmented language models
function calling
0.912025
The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models · ICML 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models · ICML 2025
Natural language and speech › Language models and text generation › code generation
API call generation
0.812024
Gorilla: Large Language Model Connected with Massive APIs · NeurIPS 2024
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling
0.812024
LLoCO: Learning Long Contexts Offline · EMNLP 2024
Information retrieval
retrieval-augmented generation
0.812024
Gorilla: Large Language Model Connected with Massive APIs · NeurIPS 2024
Cloud and datacenter computing › cloud networking
cloud data transmission
0.712023
Skyplane: Optimizing Transfer Cost and Throughput Using Cloud-Aware Overlays · NSDI 2023
Machine learning › Efficient and distributed learning
memory-efficient training
0.612022
POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging · ICML 2022
Machine learning › Efficient and distributed learning › edge computing › on-device machine learning
on-device learning
0.612022
POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging · ICML 2022
Accessibility and assistive technology
assistive technology for visual impairment
0.412019
GesturePod: Enabling On-device Gesture-based Interaction for White Cane Users · UIST 2019
Interaction techniques and input › input sensing
gesture recognition
0.412019
GesturePod: Enabling On-device Gesture-based Interaction for White Cane Users · UIST 2019
Internet of things and sensor networks
iot devices
0.312026
Efficient ML Model Updates for Deeply Embedded Microcontrollers · EuroSys 2026
Machine learning › Efficient and distributed learning › inference efficiency
context compression
0.212024
LLoCO: Learning Long Contexts Offline · EMNLP 2024
Network optimization and economics › mechanism design
incentive mechanism
0.212024
Nebula: A Privacy-First Platform for Data Backhaul · SP 2024
Program synthesis and code generation
code generation with language models
0.212024
Gorilla: Large Language Model Connected with Massive APIs · NeurIPS 2024
Internet architecture and protocols
overlay networks
0.212023
Skyplane: Optimizing Transfer Cost and Throughput Using Cloud-Aware Overlays · NSDI 2023

Methods — techniques the papers use, named apart from their topics

retrieval-aware training · 2.3instruction fine-tuning · 2.3decentralized architecture · 1.5overlay routing · 1.3rematerialization · 1.1paging · 1.1reinforcement learning · 1.0abstract syntax tree evaluation · 0.9retrieval · 0.8micropayments · 0.8micropayment · 0.8context compression · 0.8mixed-integer linear programming · 0.6mixed integer linear programming · 0.6machine learning pipeline · 0.4gesture recognition model · 0.4
YearPublicationVenuePosition
2026 AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
abstract
Yun He, Wenzhe Li, Hejia Zhang, Songlin Li, Karishma Mandyam, Sopan Khosla, Yuanhao Xiong, Nanshu Wang, Xiaoliang Peng, Beibin Li, Shengjie Bi, Shishir G Patil, Qi Qi, Shengyu Feng, Julian Katz-Samuels, Richard Yuanzhe Pang, Sujan Kumar Gonugondla, Hunter Lang, Yue Yu, Yundi Qian, Maryam Fazel-Zarandi, Licheng Yu, Amine Benhalloum, Hany Hassan Awadalla, Manaal Faruqui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Karishma Mandyam, Sopan Khosla, Yuanhao Xiong, Nanshu Wang, Xiaoliang Peng, Beibin Li, Shengjie Bi, Shishir G. Patil, Shengyu Feng, Julian Katz-Samuels, Richard Yuanzhe Pang, Sujan K. Gonugondla, Hunter Lang, Yue Yu 0009, Yundi Qian, Maryam Fazel-Zarandi, Licheng Yu, Amine Benhalloum, Hany Hassan, Manaal Faruqui
ACL (1)12
2026 Efficient ML Model Updates for Deeply Embedded Microcontrollers
abstract
Frequent updates to machine learning (ML) models are essential for maintaining accuracy in dynamic environments. However, updating models deployed on IoT devices with low-power microcontrollers remains challenging. Because the model, application, and system are all linked together into a monolithic program, model updates are typically accomplished through full device firmware updates (DFUs). Unfortunately, DFUs are disruptive, inflexible, energy-intensive, and inefficient. We present Minerva, an end-to-end ML model update system for microcontroller-based IoT devices. At the heart of Minerva is a new abstraction and mechanism called capsules. Capsules encapsulate ML models as pure, independent functions and decouple the model from the rest of the system. These properties enable updates that are lightweight, avoiding the time and energy costs of a DFU both in communication and flash reprogramming, and non-disruptive, eliminating the need for a full reboot. Because capsules work at the application level, they require no modifications to the underlying OS or firmware, and thus apply broadly to existing IoT systems. Minerva reduces model update time by up to 89× compared to a full DFU, and up to 73× compared to a state-of-the-art embedded OS, enabling the growing number of deeply embedded devices to receive regular model updates. Minerva is an open-source project available at https://github.com/ShishirPatil/minerva.
Shishir G. Patil, Sam Kumar, Prabal Dutta, Joseph Gonzalez 0001
EuroSys1
2025 The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models
abstract
Function calling, also called tool use, refers to an LLM’s ability to invoke external functions, APIs, or user-defined tools in response to user queries—an essential capability for agentic LLM applications. Despite its prominence, there did not exist a standard benchmark to evaluate function calling abilities, due to two reasons – the challenging nature of evaluating when a function call is valid, and the challenge of acquiring diverse, real-world functions. We present the Berkeley Function Calling Leaderboard (BFCL), a comprehensive benchmark designed to evaluate function calling capabilities in a wide range of real-world settings. The BFCL benchmark evaluates serial and parallel function calls, across various programming languages using a novel Abstract Syntax Tree (AST) evaluation method that can easily scale to thousands of functions. We construct the benchmark using a combination of expert curated, and user-contributed functions and associated prompts. Finally, BFCL benchmark evaluates the ability of models to abstain and reason in stateful multi-step agentic setting. Evaluating a wide range of models, we observe that while state-of-the-art LLMs excel at singleturn calls, memory, dynamic decision-making, and long-horizon reasoning remain open challenges. Since its preview, BFCL has become the defacto standard for evaluating function-calls, and can be accessed at gorilla.cs.berkeley.edu/leaderboard.html.
Shishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, Joseph Gonzalez 0001
ICML1
2024 LLoCO: Learning Long Contexts Offline
abstract
Sijun Tan, Xiuyu Li, Shishir G Patil, Ziyang Wu, Tianjun Zhang, Kurt Keutzer, Joseph E. Gonzalez, Raluca Ada Popa. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Sijun Tan, Xiuyu Li, Shishir G. Patil, Ziyang Wu, Tianjun Zhang, Kurt Keutzer, Joseph Gonzalez 0001, Raluca A. Popa
EMNLP3
2024 Gorilla: Large Language Model Connected with Massive APIs
abstract
Large Language Models (LLMs) have seen an impressive wave of advances, with models now excelling in a variety of tasks, such as mathematical reasoning and program synthesis. However, their potential to effectively use tools via API calls remains unfulfilled. This is a challenging task even for today’s state-of-the-art LLMs such as GPT-4 largely due to their unawareness of what APIs are available and how to use them in a frequently updated tool set. We develop Gorilla, a finetuned LLaMA model that surpasses the performance of GPT-4 on writing API calls. Trained with the novel Retriever Aware Training (RAT), when combined with a document retriever, Gorilla demonstrates a strong capability to adapt to test-time document changes, allowing flexible user updates or version changes. It also substantially mitigates the issue of hallucination, commonly encountered when prompting LLMs directly. To evaluate the model’s ability, we introduce APIBench, a comprehensive dataset consisting of HuggingFace, TorchHub, and TensorHub APIs. The successful integration of the retrieval system with Gorilla demonstrates the potential for LLMs to use tools more accurately, keep up with frequently updated documentation, and consequently increase the reliability and applicability of their outputs. Gorilla’s code, model, data, and demo are available at: https://gorilla.cs.berkeley.edu
Shishir G. Patil, Tianjun Zhang, Xin Wang 0066, Joseph Gonzalez 0001
NeurIPS1
2024 Nebula: A Privacy-First Platform for Data Backhaul
abstract
Imagine being able to deploy a small, battery- powered device nearly anywhere on earth that humans frequent and having it be able to send data to the cloud without needing to provision a network—without buying a physical gateway, setting up WiFi credentials, or acquiring a cellular SIM. Such a capability would address one of the greatest bottlenecks to deploying the long-tail of small, embedded, and power-constrained IoT devices in nearly any setting. Unfortunately, decoupling the device deployment from the network configuration needed to transmit, or backhaul, sensor data to the cloud remains a tricky challenge, but the success of Tile and AirTag offers hope. They have shown that mobile phones can crowd-source worldwide local network coverage to find lost items, yet expanding these systems to enable general-purpose backhaul raises privacy concerns for network participants. In this work, we present Nebula, a privacy-focused architecture for global, intermittent, and low-rate data backhaul to enable nearly any thing to eventually connect to the cloud while (i) preserving the privacy of the mobile network participants from the platform provider by decentralizing data flow through the system, (ii) incentivizing participation through micropayments, and (iii) preventing system abuse.
Jean-Luc Watson, Tess Despres, Alvin Tan, Shishir G. Patil, Prabal Dutta, Raluca A. Popa
SP4
2023 Skyplane: Optimizing Transfer Cost and Throughput Using Cloud-Aware Overlays
Paras Jain 0001, Sam Kumar, Sarah Wooders, Shishir G. Patil, Joseph Gonzalez 0001, Ion Stoica
NSDI4
2022 POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging
abstract
Fine-tuning models on edge devices like mobile phones would enable privacy-preserving personalization over sensitive data. However, edge training has historically been limited to relatively small models with simple architectures because training is both memory and energy intensive. We present POET, an algorithm to enable training large neural networks on memory-scarce battery-operated edge devices. POET jointly optimizes the integrated search search spaces of rematerialization and paging, two algorithms to reduce the memory consumption of backpropagation. Given a memory budget and a run-time constraint, we formulate a mixed-integer linear program (MILP) for energy-optimal training. Our approach enables training significantly larger models on embedded devices while reducing energy consumption while not modifying mathematical correctness of backpropagation. We demonstrate that it is possible to fine-tune both ResNet-18 and BERT within the memory constraints of a Cortex-M class embedded device while outperforming current edge training methods in energy efficiency. POET is an open-source project available at https://github.com/ShishirPatil/poet
Shishir G. Patil, Paras Jain 0001, Prabal Dutta, Ion Stoica, Joseph Gonzalez 0001
ICML1
2019 GesturePod: Enabling On-device Gesture-based Interaction for White Cane Users
abstract
People using white canes for navigation find it challenging to concurrently access devices such as smartphones. Building on prior research on abandonment of specialized devices, we explore a new touch free mode of interaction wherein a person with visual impairment can perform gestures on their existing white cane to trigger tasks on their smartphone. We present GesturePod, an easy-to-integrate device that clips on to any white cane, and detects gestures performed with the cane. With GesturePod, a user can perform common tasks on their smartphone without touch or even removing the phone from their pocket or bag. We discuss the challenges in building the device and our design choices. We propose a novel, efficient machine learning pipeline to train and deploy the gesture recognition model. Our in-lab study shows that GesturePod achieves 92% gesture recognition accuracy and can help perform common smartphone tasks faster. Our in-wild study suggests that GesturePod is a promising tool to improve smartphone access for people with VI, especially in constrained outdoor scenarios.
Shishir G. Patil, Don Kurian Dennis, Chirag Pabbaraju, Nadeem Shaheer, Harsha Vardhan Simhadri, Vivek Seshadri, Manik Varma, Prateek Jain 0002
UIST1