Linrong Cai

dblp:356/4050 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 50% Language models and text generation · 17% Learning paradigms · 17%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › data-efficient learning
label-efficient learning
0.812024
Stronger Than You Think: Benchmarking Weak Supervision on Realistic Tasks · NeurIPS 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
Zero-Shot Robustification of Zero-Shot Models · ICLR 2024
Machine learning › Learning paradigms
weakly supervised learning
0.812024
Stronger Than You Think: Benchmarking Weak Supervision on Realistic Tasks · NeurIPS 2024
Machine learning › Trustworthy machine learning › fairness › bias mitigation
word embedding debiasing
0.812024
Zero-Shot Robustification of Zero-Shot Models · ICLR 2024
Machine learning › Trustworthy machine learning › robustness › adversarial robustness › adversarially robust generalization
zero-shot adversarial robustness
0.812024
Zero-Shot Robustification of Zero-Shot Models · ICLR 2024
Natural language and speech › Language models and text generation › large language model inference
zero-shot inference
0.812024
Zero-Shot Robustification of Zero-Shot Models · ICLR 2024

Methods — techniques the papers use, named apart from their topics

language model insights · 0.8labeling functions · 0.8label model · 0.8embedding projection · 0.8
YearPublicationVenuePosition
2024 Zero-Shot Robustification of Zero-Shot Models
abstract
Zero-shot inference is a powerful paradigm that enables the use of large pretrained models for downstream classification tasks without further training. However, these models are vulnerable to inherited biases that can impact their performance. The traditional solution is fine-tuning, but this undermines the key advantage of pretrained models, which is their ability to be used out-of-the-box. We propose RoboShot, a method that improves the robustness of pretrained model embeddings in a fully zero-shot fashion. First, we use language models (LMs) to obtain useful insights from task descriptions. These insights are embedded and used to remove harmful and boost useful components in embeddings---without any supervision. Theoretically, we provide a simple and tractable model for biases in zero-shot embeddings and give a result characterizing under what conditions our approach can boost performance. Empirically, we evaluate RoboShot on nine image and NLP classification tasks and show an average improvement of 15.98% over several zero-shot baselines. Additionally, we demonstrate that RoboShot is compatible with a variety of pretrained and language models and propose a way to further boost performance with a zero-shot adaptation variant.
Dyah Adila, Changho Shin, Linrong Cai, Frederic Sala
ICLR3
2024 Stronger Than You Think: Benchmarking Weak Supervision on Realistic Tasks
abstract
Weak supervision (WS) is a popular approach for label-efficient learning, leveraging diverse sources of noisy but inexpensive weak labels to automatically annotate training data. Despite its wide usage, WS and its practical value are challenging to benchmark due to the many knobs in its setup, including: data sources, labeling functions (LFs), aggregation techniques (called label models), and end model pipelines. Existing evaluation suites tend to be limited, focusing on particular components or specialized use cases. Moreover, they often involve simplistic benchmark tasks or de-facto LF sets that are suboptimally written, producing insights that may not generalize to real-world settings. We address these limitations by introducing a new benchmark, BOXWRENCH, designed to more accurately reflect real-world usages of WS. This benchmark features tasks with (1) higher class cardinality and imbalance, (2) notable domain expertise requirements, and (3) opportunities to re-use LFs across parallel multilingual corpora. For all tasks, LFs are written using a careful procedure aimed at mimicking real-world settings. In contrast to existing WS benchmarks, we show that supervised learning requires substantial amounts (1000+) of labeled examples to match WS in many settings.
Tianyi Zhang 0015, Linrong Cai, Jeffrey Li, Nicholas Roberts, Neel Guha, Frederic Sala
NeurIPS2