Somin Wadhwa

dblp:265/5925 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 43% Efficient and distributed learning · 23% Information extraction and text analysis · 20%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › chain-of-thought reasoning
chain-of-thought distillation
0.812024
Investigating Mysteries of CoT-Augmented Distillation · EMNLP 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.812024
Investigating Mysteries of CoT-Augmented Distillation · EMNLP 2024
Data integration and cleaning
entity matching
0.812024
Learning from Natural Language Explanations for Generalizable Entity Matching · EMNLP 2024
Data integration and cleaning › entity matching
graph entity matching
0.812024
Learning from Natural Language Explanations for Generalizable Entity Matching · EMNLP 2024
Natural language and speech › Language models and text generation › in-context learning
few-shot prompting
0.712023
Revisiting Relation Extraction in the era of Large Language Models · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
relation extraction
0.712023
Revisiting Relation Extraction in the era of Large Language Models · ACL (1) 2023
Machine learning › Trustworthy machine learning › interpretability
natural language explanation
0.212024
Learning from Natural Language Explanations for Generalizable Entity Matching · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 1.5conditional generation · 1.5chain-of-thought prompting · 0.8ablation · 0.8fine-tuning · 0.7few-shot prompting · 0.7chain-of-thought · 0.7
YearPublicationVenuePosition
2024 Investigating Mysteries of CoT-Augmented Distillation
abstract
Eliciting chain of thought (CoT) rationales - sequences of token that convey a “reasoning” process has been shown to consistently improve LLM performance on tasks like question answering. More recent efforts have shown that such rationales can also be used for model distillation: Including CoT sequences (elicited from a large “teacher” model) in addition to target labels when fine-tuning a small student model yields (often substantial) improvements. In this work we ask: Why and how does this additional training signal help in model distillation? We perform ablations to interrogate this, and report some potentially surprising results. Specifically: (1) Placing CoT sequences after labels (rather than before) realizes consistently better downstream performance – this means that no student “reasoning” is necessary at test time to realize gains. (2) When rationales are appended in this way, they need not be coherent reasoning sequences to yield improvements; performance increases are robust to permutations of CoT tokens, for example. In fact, (3) a small number of key tokens are sufficient to achieve improvements equivalent to those observed when full rationales are used in model distillation.
Somin Wadhwa, Silvio Amir, Byron C. Wallace
EMNLP1
2024 Learning from Natural Language Explanations for Generalizable Entity Matching
abstract
Entity matching is the task of linking records from different sources that refer to the same real-world entity.Past work has primarily treated entity linking as a standard supervised learning problem.However, supervised entity matching models often do not generalize well to new data, and collecting exhaustive labeled training data is often cost prohibitive.Further, recent efforts have adopted LLMs for this task in few/zero-shot settings, exploiting their general knowledge.But LLMs are prohibitively expensive for performing inference at scale for real-world entity matching tasks.As an efficient alternative, we re-cast entity matching as a conditional generation task as opposed to binary classification.This enables us to "distill" LLM reasoning into smaller entity matching models via natural language explanations.This approach achieves strong performance, especially on out-of-domain generalization tests (↑10.85%F-1) where standalone generative methods struggle.We perform ablations that highlight the importance of explanations, both for performance and model robustness.Explain matching label class given the entity descriptions: Label: Match E_a: Nike Sportswear AF-1 488298-436 MN Navy.E_b: Air Force 1 [BRAND]
Somin Wadhwa, Adit Krishnan, Runhui Wang, Byron C. Wallace, Luyang Kong
EMNLP1
2024 Distilling Event Sequence Knowledge From Large Language Models
Somin Wadhwa, Oktie Hassanzadeh, Debarun Bhattacharjya, Ken Barker 0002, Jian Ni
ISWC (1)1
2023 Revisiting Relation Extraction in the era of Large Language Models
abstract
Relation extraction (RE) is the core NLP task of inferring semantic relationships between entities from text.Standard supervised RE techniques entail training modules to tag tokens comprising entity spans and then predict the relationship between them.Recent work has instead treated the problem as a sequence-tosequence task, linearizing relations between entities as target strings to be generated conditioned on the input.Here we push the limits of this approach, using larger language models (GPT-3 and Flan-T5 large) than considered in prior work and evaluating their performance on standard RE tasks under varying levels of supervision.We address issues inherent to evaluating generative approaches to RE by doing human evaluations, in lieu of relying on exact matching.Under this refined evaluation, we find that: (1) Few-shot prompting with GPT-3 achieves near SOTA performance, i.e., roughly equivalent to existing fully supervised models; (2) Flan-T5 is not as capable in the fewshot setting, but supervising and fine-tuning it with Chain-of-Thought (CoT) style explanations (generated via GPT-3) yields SOTA results.We release this model as a new baseline for RE tasks 1 .
Somin Wadhwa, Silvio Amir, Byron C. Wallace
ACL (1)1
2020 CommunityClick: Capturing and Reporting Community Feedback from Town Halls to Improve Inclusivity Share on
abstract
Local governments still depend on traditional town halls for community consultation, despite problems such as a lack of inclusive participation for attendees and difficulty for civic organizers to capture attendees' feedback in reports. Building on a formative study with 66 town hall attendees and 20 organizers, we designed and developed CommunityClick, a communitysourcing system that captures attendees' feedback in an inclusive manner and enables organizers to author more comprehensive reports. During the meeting, in addition to recording meeting audio to capture vocal attendees' feedback, we modify iClickers to give voice to reticent attendees by allowing them to provide real-time feedback beyond a binary signal. This information then automatically feeds into a meeting transcript augmented with attendees' feedback and organizers' tags. The augmented transcript along with a feedback-weighted summary of the transcript generated from text analysis methods is incorporated into an interactive authoring tool for organizers to write reports. From a field experiment at a town hall meeting, we demonstrate how CommunityClick can improve inclusivity by providing multiple avenues for attendees to share opinions. Additionally, interviews with eight expert organizers demonstrate CommunityClick's utility in creating more comprehensive and accurate reports to inform critical civic decision-making. We discuss the possibility of integrating CommunityClick with town hall meetings in the future as well as expanding to other domains.
Mahmood Jasim, Pooya Khaloo, Somin Wadhwa, Amy X. Zhang, Ali Sarvghad, Narges Mahyar
Proc. ACM Hum. Comput. Interact.3