EDBT 2026 Demo / reviewers in the wild / expert
Sandeep Soni
dblp:130/2538
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Parallelism and Energy-Efficiency in SOT-MRAM based CIM Architecture for On-Chip Learning
Anubha Sehgal, Alok Kumar Shukla, Sumit Diware, Sandeep Soni, Seema Dhull, Sonal Shreya, Sourajeet Roy, Rajendra Bishnoi |
DAC | 4 |
| 2025 | Continuous On-Chip Learning in Neural Networks using SOT-MRAM based CIM ArchitecturesabstractComputational-In-Memory (CIM) is an energy-efficient paradigm that integrates computation directly within memory arrays, reducing the bottleneck associated with data transfer. This approach is beneficial for Artificial Intelligence (AI) applications that require on-chip learning for real-time processing. However, implementing on-chip learning in CIM architectures remains challenging due to limited throughput and energy-efficiency during both online training and inference. In conventional architectures, weight updates necessitate the inference process to halt to avoid unintended computation outcomes. To overcome this limitation, this paper presents a novel Spin-Orbit Torque (SOT)-based CIM architecture tailored for continuous on-chip learning applications, which enable weight updates without interrupting the inference. The proposed SOT bit-cell utilizes two read ports and one write port (2R1W) configuration, where one read port (1R) is dedicated to inference and one read and one write (1R1W) for on-chip learning that enables concurrent read and write operations. Our proposed architecture is evaluated at the system-level using the Generic-PDK 45 nm technology node, demonstrating 2.4× improvement in energy-efficiency and 5.4× improvement in throughput compared to state-of-the-art solutions, with minimal overhead. Anubha Sehgal, Sandeep Soni, Sumit Diware, Alok Kumar Shukla, Sourajeet Roy, Rajendra Bishnoi |
ICCAD | 2 |
| 2025 | Words and Action: Modeling Linguistic Leadership in #BlackLivesMatter CommunitiesabstractIn the wake of the 2024 US presidential election, pundits on both the left and the right pointed to a conservative backlash against "woke politics" to explain the election's outcome. These politics, rooted in substantive beliefs about equity and justice--and particularly racial justice--owe their most recent rise to prominence to the Black Lives Matter (BLM) movement. A significant body of work, both qualitative and quantitative, has documented how BLM was able to move these beliefs from the margin to the mainstream. In this paper, we focus on the words that index these beliefs, devising a novel method of modeling semantic leadership across a set of communities associated with the BLM movement that is informed by domain-specific theory about Black Twitter. We describe our bespoke approaches to time-binning, community clustering, and connecting communities over time, as well as our adaptation of state-of-the-art approaches to semantic change detection and semantic leadership induction. We find evidence at scale of the leadership role of BLM activists and progressives, as well as of Black celebrities. We also find evidence of sustained conservative engagement with this discourse, suggesting an alternative explanation for how we have arrived at the present political moment. Dani Roytburg, Deborah Olorunisola, Sandeep Soni, Lauren F. Klein |
ICWSM | 3 |
| 2023 | Grounding Characters and Places in Narrative TextabstractTracking characters and locations throughout a story can help improve the understanding of its plot structure.Prior research has analyzed characters and locations from text independently without grounding characters to their locations in narrative time.Here, we address this gap by proposing a new spatial relationship categorization task.The objective of the task is to assign a spatial relationship category for every character and location co-mention within a window of text, taking into consideration linguistic context, narrative tense, and temporal scope.To this end, we annotate spatial relationships in approximately 2500 book excerpts and train a model using contextual embeddings as features to predict these relationships.When applied to a set of books, this model allows us to test several hypotheses on mobility and domestic space, revealing that protagonists are more mobile than non-central characters and that women as characters tend to occupy more interior space than men.Overall, our work is the first step towards joint modeling and analysis of characters and places in narrative text. Sandeep Soni, Amanpreet Sihra, Elizabeth F. Evans, Matthew Wilkens, David Bamman |
ACL (1) | 1 |
| 2023 | Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4abstractIn this work, we carry out a data archaeology to infer books that are known to ChatGPT and GPT-4 using a name cloze membership inference query.We find that OpenAI models have memorized a wide collection of copyrighted materials, and that the degree of memorization is tied to the frequency with which passages of those books appear on the web.The ability of these models to memorize an unknown set of books complicates assessments of measurement validity for cultural analytics by contaminating test data; we show that models perform much better on memorized books than on nonmemorized books for downstream tasks.We argue that this supports a case for open models whose training data is known. Kent K. Chang, Mackenzie Cramer, Sandeep Soni, David Bamman |
EMNLP | 3 |
| 2022 | Linguistic Characterization of Divisive Topics Online: Case Studies on Contentiousness in Abortion, Climate Change, and Gun Control
Jacob Beel, Tong Xiang, Sandeep Soni, Diyi Yang |
ICWSM | 3 |
| 2021 | Racism is a virus: anti-asian hate and counterspeech in social media during the COVID-19 crisisabstractThe spread of COVID-19 has sparked racism and hate on social media targeted towards Asian communities. However, little is known about how racial hate spreads during a pandemic and the role of counterspeech in mitigating this spread. In this work, we study the evolution and spread of anti-Asian hate speech through the lens of Twitter. We create COVID-HATE, the largest dataset of anti-Asian hate and counterspeech spanning 14 months, containing over 206 million tweets, and a social network with over 127 million nodes. By creating a novel hand-labeled dataset of 3,355 tweets, we train a text classifier to identify hateful and counterspeech tweets that achieves an average macro-F1 score of 0.832. Using this dataset, we conduct longitudinal analysis of tweets and users. Analysis of the social network reveals that hateful and counterspeech users interact and engage extensively with one another, instead of living in isolated polarized communities. We find that nodes were highly likely to become hateful after being exposed to hateful content in the year 2020. Notably, counterspeech messages discourage users from turning hateful, potentially suggesting a solution to curb hate on web and social media platforms. Data and code is available at http://claws.cc.gatech.edu/covid. Bing He 0002, Caleb Ziems, Sandeep Soni, Naren Ramakrishnan, Diyi Yang, Srijan Kumar |
ASONAM | 3 |
| 2021 | Follow the leader: Documents on the leading edge of semantic change get more citationsabstractAbstract Diachronic word embeddings—vector representations of words over time—offer remarkable insights into the evolution of language and provide a tool for quantifying sociocultural change from text documents. Prior work has used such embeddings to identify shifts in the meaning of individual words. However, simply knowing that a word has changed in meaning is insufficient to identify the instances of word usage that convey the historical meaning or the newer meaning. In this study, we link diachronic word embeddings to documents, by situating those documents as leaders or laggards with respect to ongoing semantic changes. Specifically, we propose a novel method to quantify the degree of semantic progressiveness in each word usage, and then show how these usages can be aggregated to obtain scores for each document. We analyze two large collections of documents, representing legal opinions and scientific articles. Documents that are scored as semantically progressive receive a larger number of citations, indicating that they are especially influential. Our work thus provides a new technique for identifying lexical semantic leaders and demonstrates a new link between progressive use of language and influence in a citation network. Sandeep Soni, Kristina Lerman, Jacob Eisenstein |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2017 | Calendar.help: Designing a Workflow-Based Scheduling Agent with Humans in the LoopabstractAlthough we may complain about meetings, they are an essential part of an information worker's work life. Consequently, busy people spend a significant amount of time scheduling meetings. We present Calendar.help, a system that provides fast, efficient scheduling through structured workflows. Users interact with the system via email, delegating their scheduling needs to the system as if it were a human personal assistant. Common scheduling scenarios are broken down using well-defined workflows and completed as a series of microtasks that are automated when possible and executed by a human otherwise. Unusual scenarios fall back to a trained human assistant executing an unstructured macrotask. We describe the iterative approach we used to develop Calendar.help, and share the lessons learned from scheduling thousands of meetings during a year of real-world deployments. Our findings provide insight into how complex information tasks can be broken down into repeatable components that can be executed efficiently to improve productivity. Justin Cranshaw, Emad Elwany, Todd Newman, Rafal Kocielnik, Bowen Yu 0001, Sandeep Soni, Jaime Teevan, Andrés Monroy-Hernández |
CHI | 6 |