VLDB 2026 Research / reviewers in the wild / expert
Linjun Zhang
dblp:117/0963
· DBLP profile ↗
38ranked-venue papers
4as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 2 first-author · 31 since 2021Theory of computation · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PEANuT: Parameter-Efficient Adaptation with Weight-aware Neural TweakersabstractFine-tuning large pre-trained foundation models often yields excellent downstream performance but is prohibitively expensive when updating all parameters. Parameter-efficient fine-tuning (PEFT) methods such as LoRA alleviate this by introducing lightweight update modules, yet they commonly rely on weight-agnostic linear approximations, limiting their expressiveness. In this work, we propose PEANuT, a novel PEFT framework that introduces weight-aware neural tweakers, compact neural modules that generate task-adaptive updates conditioned on frozen pre-trained weights. PEANuT provides a flexible yet efficient way to capture complex update patterns without full model tuning. We theoretically show that PEANuT achieves equivalent or greater expressivity than existing linear PEFT methods with comparable or fewer parameters. Extensive experiments across four benchmarks with over twenty datasets demonstrate that PEANuT consistently outperforms strong baselines in both NLP and vision tasks, while maintaining low computational overhead. Yibo Zhong, Haoxiang Jiang, Lincan Li, Ryumei Nakada, Tianci Liu 0003, Linjun Zhang, Huaxiu Yao, Haoyu Wang 0004 |
KDD (1) | 6 |
| 2025 | MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language ModelsabstractArtificial Intelligence (AI) has demonstrated significant potential in healthcare, particularly in disease diagnosis and treatment planning. Recent progress in Medical Large Vision-Language Models (Med-LVLMs) has opened up new possibilities for interactive diagnostic tools. However, these models often suffer from factual hallucination, which can lead to incorrect diagnoses. Fine-tuning and retrieval-augmented generation (RAG) have emerged as methods to address these issues. However, the amount of high-quality data and distribution shifts between training data and deployment data limit the application of fine-tuning methods. Although RAG is lightweight and effective, existing RAG-based approaches are not sufficiently general to different medical domains and can potentially cause misalignment issues, both between modalities and between the model and the ground truth. In this paper, we propose a versatile multimodal RAG system, MMed-RAG, designed to enhance the factuality of Med-LVLMs. Our approach introduces a domain-aware retrieval mechanism, an adaptive retrieved contexts selection, and a provable RAG-based preference fine-tuning strategy. These innovations make the RAG process sufficiently general and reliable, significantly improving alignment when introducing retrieved contexts. Experimental results across five medical datasets (involving radiology, ophthalmology, pathology) on medical VQA and report generation demonstrate that MMed-RAG can achieve an average improvement of 43.8% in factual accuracy in the factual accuracy of Med-LVLMs. Peng Xia 0005, Kangyu Zhu, Haoran Li 0011, Tianze Wang, Sheng Wang 0014, Linjun Zhang, James Zou 0001, Huaxiu Yao |
ICLR | 7 |
| 2025 | Mitigating Heterogeneous Token Overfitting in LLM Knowledge EditingabstractLarge language models (LLMs) have achieved remarkable performance on various natural language tasks. However, they are trained on static corpora and their knowledge can become outdated quickly in the fast-changing world. This motivates the development of knowledge editing (KE) to update specific knowledge in LLMs without changing unrelated others or compromising their pre-trained capabilities. Previous efforts sought to update a small amount of parameters of a LLM and proved effective for making selective updates. Nonetheless, the edited LLM often exhibits degraded ability to reason about the new knowledge. In this work, we identify a key issue: heterogeneous token overfitting (HTO), where the LLM overfits different tokens in the provided knowledge at varying rates. To tackle this, we propose OVERTONE, a token-level smoothing method that mitigates HTO by adaptively refining the target distribution. Theoretically, OVERTONE offers better parameter updates with negligible computation overhead. It also induces an implicit DPO but does not require preference data pairs. Extensive experiments across four editing methods, two LLMs, and diverse scenarios demonstrate the effectiveness and versatility of our method. Tianci Liu 0003, Ruirui Li 0002, Zihan Dong, Hui Liu 0033, Xianfeng Tang, Qingyu Yin, Linjun Zhang, Haoyu Wang 0004, Jing Gao 0004 |
ICML | 7 |
| 2025 | FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free GuaranteesabstractThe propensity of large language models (LLMs) to generate hallucinations and non-factual content undermines their reliability in high-stakes domains, where rigorous control over Type I errors (the conditional probability of incorrectly classifying hallucinations as truthful content) is essential. Despite its importance, formal verification of LLM factuality with such guarantees remains largely unexplored. In this paper, we introduce FactTest, a novel framework that statistically assesses whether an LLM can provide correct answers to given questions with high-probability correctness guarantees. We formulate hallucination detection as a hypothesis testing problem to enforce an upper bound of Type I errors at user-specified significance levels. Notably, we prove that FactTest also ensures strong Type II error control under mild conditions and can be extended to maintain its effectiveness when covariate shifts exist. Our approach is distribution-free and works for any number of human-annotated samples. It is model-agnostic and applies to any black-box or white-box LM. Extensive experiments on question-answering (QA) benchmarks demonstrate that FactTest effectively detects hallucinations and enable LLMs to abstain from answering unknown questions, leading to an over 40% accuracy improvement. Fan Nie, Xiaotian Hou, Shuhang Lin, James Zou 0001, Huaxiu Yao, Linjun Zhang |
ICML | 6 |
| 2025 | MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference AlignmentabstractReinforcement Learning from Human Feedback (RLHF) has shown promise in aligning large language models (LLMs). Yet its reliance on a singular reward model often overlooks the diversity of human preferences. Recent approaches address this limitation by leveraging multi-dimensional feedback to fine-tune corresponding reward models and train LLMs using reinforcement learning. However, the process is costly and unstable, especially given the competing and heterogeneous nature of human preferences. In this paper, we propose Mixing Preference Optimization (MPO), a post-processing framework for aggregating single-objective policies as an alternative to both multi-objective RLHF (MORLHF) and MaxMin-RLHF. MPO avoids alignment from scratch. Instead, it log-linearly combines existing policies into a unified one with the weight of each policy computed via a batch stochastic mirror descent. Empirical results demonstrate that MPO achieves balanced performance across diverse preferences, outperforming or matching existing models with significantly reduced computational costs. Tianze Wang, Dongnan Gui, Shuhang Lin, Linjun Zhang |
ICML | 5 |
| 2025 | AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-PlayabstractSearch-augmented LLMs often struggle with complex reasoning tasks due to ineffective multi-hop retrieval and limited reasoning ability. We propose AceSearcher, a cooperative self-play framework that trains a single large language model (LLM) to alternate between two roles: a decomposer that breaks down complex queries and a solver that integrates retrieved contexts for answer generation. AceSearcher couples supervised fine-tuning on a diverse mixture of search, reasoning, and decomposition tasks with reinforcement fine-tuning optimized for final answer accuracy, eliminating the need for intermediate annotations. Extensive experiments on three reasoning-intensive tasks across 10 datasets show that AceSearcher outperforms state-of-the-art baselines, achieving an average exact match improvement of 7.6%. Remarkably, on document-level finance reasoning tasks, AceSearcher-32B matches the performance of the giant DeepSeek-V3 model using less than 5% of iits parameters. Even at smaller scales (1.5B and 8B), AceSearcher often surpasses existing search-augmented LLMs with up to 9× more parameters, highlighting its exceptional efficiency and effectiveness in tackling complex reasoning tasks. Ran Xu 0002, Yuchen Zhuang, Zihan Dong, Yue Yu 0001, Joyce C. Ho, Linjun Zhang, Haoyu Wang 0003, Wenqi Shi 0002, Carl Yang 0001 |
NeurIPS | 7 |
| 2025 | Differentially Private Learning Beyond the Classical Dimensionality Regime
Cynthia Dwork, Pranay Tankala, Linjun Zhang |
TCC (4) | 3 |
| 2024 | RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language ModelsabstractThe recent emergence of Medical Large Vision Language Models (Med-LVLMs) has enhanced medical diagnosis.However, current Med-LVLMs frequently encounter factual issues, often generating responses that do not align with established medical facts.Retrieval-Augmented Generation (RAG), which utilizes external knowledge, can improve the factual accuracy of these models but introduces two major challenges.First, limited retrieved contexts might not cover all necessary information, while excessive retrieval can introduce irrelevant and inaccurate references, interfering with the model's generation.Second, in cases where the model originally responds correctly, applying RAG can lead to an over-reliance on retrieved contexts, resulting in incorrect answers.To address these issues, we propose RULE, which consists of two components.First, we introduce a provably effective strategy for controlling factuality risk through the calibrated selection of the number of retrieved contexts.Second, based on samples where over-reliance on retrieved contexts led to errors, we curate a preference dataset to fine-tune the model, balancing its dependence on inherent knowledge and retrieved contexts for generation.We demonstrate the effectiveness of RULE on medical VQA and report generation tasks across three datasets, achieving an average improvement of 47.4% in factual accuracy.We publicly release our benchmark and code in https: //github.com/richard-peng-xia/RULE. Peng Xia 0005, Kangyu Zhu, Haoran Li 0011, Hongtu Zhu, Yun Li 0010, Gang Li 0001, Linjun Zhang, Huaxiu Yao |
EMNLP | 7 |
| 2024 | Analyzing and Mitigating Object Hallucination in Large Vision-Language ModelsabstractLarge vision-language models (LVLMs) have shown remarkable abilities in understanding visual information with human languages. However, LVLMs still suffer from object hallucination, which is the problem of generating descriptions that include objects that do not actually exist in the images. This can negatively impact many vision-language tasks, such as visual summarization and reasoning. To address this issue, we propose a simple yet powerful algorithm, LVLM Hallucination Revisor (LURE), to post-hoc rectify object hallucination in LVLMs by reconstructing less hallucinatory descriptions. LURE is grounded in a rigorous statistical analysis of the key factors underlying object hallucination, including co-occurrence (the frequent appearance of certain objects alongside others in images), uncertainty (objects with higher uncertainty during LVLM decoding), and object position (hallucination often appears in the later part of the generated text). LURE can also be seamlessly integrated with any LVLMs. We evaluate LURE on six open-source LVLMs and found it outperforms the previous best approach in both general object hallucination evaluation metrics, GPT, and human evaluations. Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, Huaxiu Yao |
ICLR | 4 |
| 2024 | Conformal Prediction for Deep Classifier via Label RankingabstractConformal prediction is a statistical framework that generates prediction sets containing ground-truth labels with a desired coverage guarantee. The predicted probabilities produced by machine learning models are generally miscalibrated, leading to large prediction sets in conformal prediction. To address this issue, we propose a novel algorithm named $\textit{Sorted Adaptive Prediction Sets}$ (SAPS), which discards all the probability values except for the maximum softmax probability. The key idea behind SAPS is to minimize the dependence of the non-conformity score on the probability values while retaining the uncertainty information. In this manner, SAPS can produce compact prediction sets and communicate instance-wise uncertainty. Extensive experiments validate that SAPS not only lessens the prediction sets but also broadly enhances the conditional coverage rate of prediction sets. Jianguo Huang, Huajun Xi, Linjun Zhang, Huaxiu Yao, Hongxin Wei |
ICML | 3 |
| 2024 | Fair Risk Control: A Generalized Framework for Calibrating Multi-group Fairness RisksabstractThis paper introduces a framework for post-processing machine learning models so that their predictions satisfy multi-group fairness guarantees. Based on the celebrated notion of multicalibration, we introduce $(s,g,\alpha)-$GMC (Generalized Multi-Dimensional Multicalibration) for multi-dimensional mappings $s$, constraints $g$, and a pre-specified threshold level $\alpha$. We propose associated algorithms to achieve this notion in general settings. This framework is then applied to diverse scenarios encompassing different fairness concerns, including false negative rate control in image segmentation, prediction set conditional uncertainty quantification in hierarchical classification, and de-biased text generation in language models. We conduct numerical studies on several datasets and tasks. Lujing Zhang, Aaron Roth 0001, Linjun Zhang |
ICML | 3 |
| 2024 | Order-Independence Without Fine TuningabstractThe development of generative language models that can create long and coherent textual outputs via autoregression has lead to a proliferation of uses and a corresponding sweep of analyses as researches work to determine the limitations of this new paradigm. Unlike humans, these '*Large Language Models*' (LLMs) are highly sensitive to small changes in their inputs, leading to unwanted inconsistency in their behavior. One problematic inconsistency when LLMs are used to answer multiple-choice questions or analyze multiple inputs is *order dependency*: the output of an LLM can (and often does) change significantly when sub-sequences are swapped, despite both orderings being semantically identical. In this paper we present , a technique that *guarantees* the output of an LLM will not have order dependence on a specified set of sub-sequences. We show that this method *provably* eliminates order dependency, and that it can be applied to *any* transformer-based LLM to enable text generation that is unaffected by re-orderings. Delving into the implications of our method, we show that, despite our inputs being out of distribution, the impact on expected accuracy is small, where the expectation is over the order of uniformly chosen shuffling of the candidate responses, and usually significantly less in practice. Thus, can be used as a '*dropped-in*' method on fully trained models. Finally, we discuss how our method's success suggests that other strong guarantees can be obtained on LLM performance via modifying the input representations.
Code is available at [github.com/reidmcy/set-based-prompting](https://github.com/reidmcy/set-based-prompting.). Reid McIlroy-Young, Katrina Brown, Conlan Olson, Linjun Zhang, Cynthia Dwork |
NeurIPS | 4 |
| 2024 | S2FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured SparsityabstractCurrent PEFT methods for LLMs can achieve high quality, efficient training, or scalable serving, but not all three simultaneously.
To address this limitation, we investigate sparse fine-tuning and observe a remarkable improvement in generalization ability.
Utilizing this key insight, we propose a family of Structured Sparse Fine-Tuning (S${^2}$FT) methods for LLMs, which concurrently achieve state-of-the-art fine-tuning performance, training efficiency, and inference scalability. S${^2}$FT accomplishes this by "selecting sparsely and computing densely". Based on the coupled structures in LLMs, \model selects a few attention heads and channels in the MHA and FFN modules for each Transformer block, respectively. Next, it co-permutes the weight matrices on both sides of all coupled structures to connect the selected subsets in each layer into a dense submatrix. Finally, S${^2}$FT performs in-place gradient updates on all selected submatrices.
Through theoretical analyses and empirical results, our method prevents forgetting while simplifying optimization, delivers SOTA performance on both commonsense and arithmetic reasoning with 4.6% and 1.3% average improvements compared to LoRA, and surpasses full FT by 11.5% when generalizing to various domains after instruction tuning.
Using our partial back-propagation algorithm, S${^2}$FT saves training memory up to 3$\times$ and improves latency by 1.5-2.7$\times$ compared to full FT, while achieving an average 10\% improvement over LoRA on both metrics. We further demonstrate that the weight updates in S${^2}$FT can be decoupled into adapters, enabling effective fusion, fast switch, and efficient parallelism when serving multiple fine-tuned models. Jixuan Leng, Geyang Guo, Ryumei Nakada, Linjun Zhang, Huaxiu Yao, Beidi Chen |
NeurIPS | 6 |
| 2024 | Calibrated Self-Rewarding Vision Language ModelsabstractLarge Vision-Language Models (LVLMs) have made substantial progress by integrating pre-trained large language models (LLMs) and vision models through instruction tuning. Despite these advancements, LVLMs often exhibit the hallucination phenomenon, where generated text responses appear linguistically plausible but contradict the input image, indicating a misalignment between image and text pairs. This misalignment arises because the model tends to prioritize textual information over visual input, even when both the language model and visual representations are of high quality. Existing methods leverage additional models or human annotations to curate preference data and enhance modality alignment through preference optimization. These approaches are resource-intensive and may not effectively reflect the target LVLM's preferences, making the curated preferences easily distinguishable. Our work addresses these challenges by proposing the Calibrated Self-Rewarding (CSR) approach, which enables the model to self-improve by iteratively generating candidate responses, evaluating the reward for each response, and curating preference data for fine-tuning. In the reward modeling, we employ a step-wise strategy and incorporate visual constraints into the self-rewarding process to place greater emphasis on visual input. Empirical results demonstrate that CSR significantly enhances performance and reduces hallucinations across twelve benchmarks and tasks, achieving substantial improvements over existing methods by 7.62\%. Our empirical results are further supported by rigorous theoretical analysis, under mild assumptions, verifying the effectiveness of introducing visual constraints into the self-rewarding paradigm. Additionally, CSR shows compatibility with different vision-language models and the ability to incrementally improve performance through iterative fine-tuning. Yiyang Zhou, Zhiyuan Fan, Dongjie Cheng, Sihan Yang 0001, Zhaorun Chen, Chenhang Cui, Linjun Zhang, Huaxiu Yao |
NeurIPS | 9 |
| 2023 | Reinforcement Learning with Stepwise Fairness ConstraintsabstractAI methods are used in societally important settings, ranging from credit to employment to housing, and it is crucial to provide fairness in regard to automated decision making. Moreover, many settings are dynamic, with populations responding to sequential decision policies. We introduce the study of reinforcement learning (RL) with stepwise fairness constraints, which require group fairness at each time step. In the case of tabular episodic RL, we provide learning algorithms with strong theoretical guarantees in regard to policy optimality and fairness violations. Our framework provides tools to study the impact of fairness constraints in sequential settings and brings up new challenges in RL. Zhun Deng, Steven Z. Wu, Linjun Zhang, David C. Parkes |
AISTATS | 4 |
| 2023 | Understanding Multimodal Contrastive Learning and Incorporating Unpaired DataabstractLanguage-supervised vision models have recently attracted great attention in computer vision. A common approach to build such models is to use contrastive learning on paired data across the two modalities, as exemplified by Contrastive Language-Image Pre-Training (CLIP). In this paper, (i) we initiate the investigation of a general class of nonlinear loss functions for multimodal contrastive learning (MMCL) including CLIP loss and show its connection to singular value decomposition (SVD). Namely, we show that each step of loss minimization by gradient descent can be seen as performing SVD on a contrastive cross-covariance matrix. Based on this insight, (ii) we analyze the performance of MMCL under linear representation settings. We quantitatively show that the feature learning ability of MMCL can be better than that of unimodal contrastive learning applied to each modality even under the presence of wrongly matched pairs. This characterizes the robustness of MMCL to noisy data. Furthermore, when we have access to additional unpaired data, (iii) we propose a new MMCL loss that incorporates additional unpaired datasets. We show that the algorithm can detect the ground-truth pairs and improve performance by fully exploiting unpaired datasets. The performance of the proposed algorithm was verified by numerical experiments. Ryumei Nakada, Halil Ibrahim Gulluk, Zhun Deng, Wenlong Ji, James Zou 0001, Linjun Zhang |
AISTATS | 6 |
| 2023 | Freeze then Train: Towards Provable Representation Learning under Spurious Correlations and Feature NoiseabstractThe existence of spurious correlations such as image backgrounds in the training environment can make empirical risk minimization (ERM) perform badly in the test environment. To address this problem, Kirichenko et al. (2022) empirically found that the core features that are related to the outcome can still be learned well even with the presence of spurious correlations. This opens a promising strategy to first train a feature learner rather than a classifier, and then perform linear probing (last layer retraining) in the test environment. However, a theoretical understanding of when and why this approach works is lacking. In this paper, we find that core features are only learned well when their associated non-realizable noise is smaller than that of spurious features, which is not necessarily true in practice. We provide both theories and experiments to support this finding and to illustrate the importance of non-realizable noise. Moreover, we propose an algorithm called Freeze then Train (FTT), that first freezes certain salient features and then trains the rest of the features using ERM. We theoretically show that FTT preserves features that are more beneficial to test time probing. Across two commonly used spurious correlation datasets, FTT outperforms ERM, IRM, JTT and CVaR-DRO, with substantial improvement in accuracy (by 4.5$%$) when the feature noise is large. FTT also performs better on general distribution shift benchmarks. Haotian Ye, James Zou 0001, Linjun Zhang |
AISTATS | 3 |
| 2023 | FIFA: Making Fairness More Generalizable in Classifiers Trained on Imbalanced Data
Zhun Deng, Jiayao Zhang 0001, Linjun Zhang, Ting Ye, Yates Coley, Weijie J. Su, James Zou 0001 |
ICLR | 3 |
| 2023 | FaiREE: fair classification with finite-sample and distribution-free guarantee
Puheng Li, James Zou 0001, Linjun Zhang |
ICLR | 3 |
| 2023 | Discover and Cure: Concept-aware Mitigation of Spurious CorrelationabstractDeep neural networks often rely on spurious correlations to make predictions, which hinders generalization beyond training environments. For instance, models that associate cats with bed backgrounds can fail to predict the existence of cats in other environments without beds. Mitigating spurious correlations is crucial in building trustworthy models. However, the existing works lack transparency to offer insights into the mitigation process. In this work, we propose an interpretable framework, Discover and Cure (DISC), to tackle the issue. With human-interpretable concepts, DISC iteratively 1) discovers unstable concepts across different environments as spurious attributes, then 2) intervenes on the training data using the discovered concepts to reduce spurious correlation. Across systematic experiments, DISC provides superior generalization ability and interpretability than the existing approaches. Specifically, it outperforms the state-of-the-art methods on an object recognition task and a skin-lesion classification task by 7.5% and 9.6%, respectively. Additionally, we offer theoretical analysis and guarantees to understand the benefits of models trained by DISC. Code and data are available at https://github.com/Wuyxin/DISC. Shirley Wu, Mert Yüksekgönül, Linjun Zhang, James Zou 0001 |
ICML | 3 |
| 2023 | HappyMap : A Generalized Multicalibration MethodabstractModern complex systems, such as radiotherapy machines, require robust strategies for fault detection, diagnosis, and prognosis to ensure operational continuity and patient safety. While data-driven methods have gained traction, few studies address diagnostic and prognostic tasks using multimodal operational data under unsupervised or semi-supervised learning settings. This gap is particularly critical given the scarcity of labeled failure data in real-world environments. This work aims to design a unified approach for fault detection, diagnosis, and prognosis using multimodal data in the absence of complete labeling. To this end, autoencoders (AEs) are employed due to their suitability for unsupervised and self-supervised learning, flexibility in handling heterogeneous data, and ability to construct latent representations optimized for various downstream tasks. A specific implementation based on a Long Short-Term Memory β-Variational Autoencoder (LSTM-β-VAE) was developed to detect anomalies in machine logs. This framework is applied to TomoTherapy® systems - a highly complex and under-explored use case within the radiotherapy domain. Initial results demonstrate strong anomaly detection performance on both a public benchmark dataset (HDFS) and a proprietary dataset derived from real-world TomoTherapy® machine faults. Beyond methodology, the paper includes a concise literature review of multimodal learning and data-driven diagnosis and prognosis with a focus on AEs. Based on this review, key research directions are identified for the continuation of the thesis, especially the integration of explainable AI as a means to enhance diagnosis capabilities in the absence of labeled faults. Zhun Deng, Cynthia Dwork, Linjun Zhang |
ITCS | 3 |
| 2023 | Beyond Confidence: Reliable Models Should Also Consider AtypicalityabstractWhile most machine learning models can provide confidence in their predictions, confidence is insufficient to understand a prediction's reliability. For instance, the model may have a low confidence prediction if the input is not well-represented in the training dataset or if the input is inherently ambiguous. In this work, we investigate the relationship between how atypical~(rare) a sample or a class is and the reliability of a model's predictions. We first demonstrate that atypicality is strongly related to miscalibration and accuracy. In particular, we empirically show that predictions for atypical inputs or atypical classes are more overconfident and have lower accuracy. Using these insights, we show incorporating atypicality improves uncertainty quantification and model performance for discriminative neural networks and large language models. In a case study, we show that using atypicality improves the performance of a skin lesion classifier across different skin tone groups without having access to the group attributes. Overall, we propose that models should use not only confidence but also atypicality to improve uncertainty quantification and performance. Our results demonstrate that simple post-hoc atypicality estimators can provide significant value. Mert Yüksekgönül, Linjun Zhang, James Zou 0001, Carlos Guestrin |
NeurIPS | 2 |
| 2023 | The Power of Contrast for Feature Learning: A Theoretical AnalysisabstractContrastive learning has achieved state-of-the-art performance in various self-supervised learning tasks and even outperforms its supervised counterpart. Despite its empirical success, theoretical understanding of the superiority of contrastive learning is still limited. In this paper, under linear representation settings, (i) we provably show that contrastive learning outperforms the standard autoencoders and generative adversarial networks, two classical generative unsupervised learning methods, for both feature recovery and in-domain downstream tasks; (ii) we also illustrate the impact of labeled data in supervised contrastive learning. This provides theoretical support for recent findings that contrastive learning with labels improves the performance of learned representations in the in-domain downstream task, but it can harm the performance in transfer learning. We verify our theory with numerical experiments. Wenlong Ji, Zhun Deng, Ryumei Nakada, James Zou 0001, Linjun Zhang |
J. Mach. Learn. Res. | 5 |
| 2022 | Meta-Learning with Fewer Tasks through Task Interpolation
Huaxiu Yao, Linjun Zhang, Chelsea Finn |
ICLR | 2 |
| 2022 | Improving Out-of-Distribution Robustness via Selective AugmentationabstractMachine learning algorithms typically assume that training and test examples are drawn from the same distribution. However, distribution shift is a common problem in real-world applications and can cause models to perform dramatically worse at test time. In this paper, we specifically consider the problems of subpopulation shifts (e.g., imbalanced data) and domain shifts. While prior works often seek to explicitly regularize internal representations or predictors of the model to be domain invariant, we instead aim to learn invariant predictors without restricting the model’s internal representations or predictors. This leads to a simple mixup-based technique which learns invariant predictors via selective augmentation called LISA. LISA selectively interpolates samples either with the same labels but different domains or with the same domain but different labels. Empirically, we study the effectiveness of LISA on nine benchmarks ranging from subpopulation shifts to domain shifts, and we find that LISA consistently outperforms other state-of-the-art methods and leads to more invariant predictors. We further analyze a linear setting and theoretically show how LISA leads to a smaller worst-group error. Huaxiu Yao, Yu Wang 0170, Sai Li 0005, Linjun Zhang, Weixin Liang, James Zou 0001, Chelsea Finn |
ICML | 4 |
| 2022 | When and How Mixup Improves CalibrationabstractIn many machine learning applications, it is important for the model to provide confidence scores that accurately capture its prediction uncertainty. Although modern learning methods have achieved great success in predictive accuracy, generating calibrated confidence scores remains a major challenge. Mixup, a popular yet simple data augmentation technique based on taking convex combinations of pairs of training examples, has been empirically found to significantly improve confidence calibration across diverse applications. However, when and how Mixup helps calibration is still a mystery. In this paper, we theoretically prove that Mixup improves calibration in high-dimensional settings by investigating natural statistical models. Interestingly, the calibration benefit of Mixup increases as the model capacity increases. We support our theories with experiments on common architectures and datasets. In addition, we study how Mixup improves calibration in semi-supervised learning. While incorporating unlabeled data can sometimes make the model less calibrated, adding Mixup training mitigates this issue and provably improves calibration. Our analysis provides new insights and a framework to understand Mixup and calibration. Linjun Zhang, Zhun Deng, Kenji Kawaguchi, James Zou 0001 |
ICML | 1 |
| 2022 | C-Mixup: Improving Generalization in RegressionabstractImproving the generalization of deep networks is an important open challenge, particularly in domains without plentiful data. The mixup algorithm improves generalization by linearly interpolating a pair of examples and their corresponding labels. These interpolated examples augment the original training set. Mixup has shown promising results in various classification tasks, but systematic analysis of mixup in regression remains underexplored. Using mixup directly on regression labels can result in arbitrarily incorrect labels. In this paper, we propose a simple yet powerful algorithm, C-Mixup, to improve generalization on regression tasks. In contrast with vanilla mixup, which picks training examples for mixing with uniform probability, C-Mixup adjusts the sampling probability based on the similarity of the labels. Our theoretical analysis confirms that C-Mixup with label similarity obtains a smaller mean square error in supervised regression and meta-regression than vanilla mixup and using feature similarity. Another benefit of C-Mixup is that it can improve out-of-distribution robustness, where the test distribution is different from the training distribution. By selectively interpolating examples with similar labels, it mitigates the effects of domain-associated information and yields domain-invariant representations. We evaluate C-Mixup on eleven datasets, ranging from tabular to video data. Compared to the best prior approach, C-Mixup achieves 6.56%, 4.76%, 5.82% improvements in in-distribution generalization, task generalization, and out-of-distribution robustness, respectively. Code is released at https://github.com/huaxiuyao/C-Mixup. Huaxiu Yao, Yiping Wang 0003, Linjun Zhang, James Zou 0001, Chelsea Finn |
NeurIPS | 3 |
| 2022 | Understanding Dynamics of Nonlinear Representation Learning and Its ApplicationabstractRepresentations of the world environment play a crucial role in artificial intelligence. It is often inefficient to conduct reasoning and inference directly in the space of raw sensory representations, such as pixel values of images. Representation learning allows us to automatically discover suitable representations from raw sensory data. For example, given raw sensory data, a deep neural network learns nonlinear representations at its hidden layers, which are subsequently used for classification (or regression) at its output layer. This happens implicitly during training through minimizing a supervised or unsupervised loss. In this letter, we study the dynamics of such implicit nonlinear representation learning. We identify a pair of a new assumption and a novel condition, called the on-model structure assumption and the data architecture alignment condition. Under the on-model structure assumption, the data architecture alignment condition is shown to be sufficient for the global convergence and necessary for global optimality. Moreover, our theory explains how and when increasing network size does and does not improve the training behaviors in the practical regime. Our results provide practical guidance for designing a model structure; for example, the on-model structure assumption can be used as a justification for using a particular model structure instead of others. As an application, we then derive a new training framework, which satisfies the data architecture alignment condition without assuming it by automatically modifying any given training algorithm dependent on data and architecture. Given a standard training algorithm, the framework running its modified version is empirically shown to maintain competitive (practical) test performances while providing global convergence guarantees for deep residual neural networks with convolutions, skip connections, and batch normalization with standard benchmark data sets, including MNIST, CIFAR-10, CIFAR-100, Semeion, KMNIST, and SVHN. Kenji Kawaguchi, Linjun Zhang, Zhun Deng |
Neural Comput. | 2 |
| 2021 | Improving Adversarial Robustness via Unlabeled Out-of-Domain DataabstractData augmentation by incorporating cheap unlabeled data from multiple domains is a powerful way to improve prediction especially when there is limited labeled data. In this work, we investigate how adversarial robustness can be enhanced by leveraging out-of-domain unlabeled data. We demonstrate that for broad classes of distributions and classifiers, there exists a sample complexity gap between standard and robust classification. We quantify the extent to which this gap can be bridged by leveraging unlabeled samples from a shifted domain by providing both upper and lower bounds. Moreover, we show settings where we achieve better adversarial robustness when the unlabeled data come from a shifted domain rather than the same domain as the labeled data. We also investigate how to leverage out-of-domain data when some structural information, such as sparsity, is shared between labeled and unlabeled domains. Experimentally, we augment object recognition datasets (CIFAR-10, CINIC-10, and SVHN) with easy-to-obtain and unlabeled out-of-domain data and demonstrate substantial improvement in the model’s robustness against $\ell_\infty$ adversarial attacks on the original domain. Zhun Deng, Linjun Zhang, Amirata Ghorbani, James Zou 0001 |
AISTATS | 2 |
| 2021 | How Does Mixup Help With Robustness and Generalization?
Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani, James Zou 0001 |
ICLR | 1 |
| 2021 | Improving Generalization in Meta-learning via Task AugmentationabstractMeta-learning has proven to be a powerful paradigm for transferring the knowledge from previous tasks to facilitate the learning of a novel task. Current dominant algorithms train a well-generalized model initialization which is adapted to each task via the support set. The crux lies in optimizing the generalization capability of the initialization, which is measured by the performance of the adapted model on the query set of each task. Unfortunately, this generalization measure, evidenced by empirical results, pushes the initialization to overfit the meta-training tasks, which significantly impairs the generalization and adaptation to novel tasks. To address this issue, we actively augment a meta-training task with “more data” when evaluating the generalization. Concretely, we propose two task augmentation methods, including MetaMix and Channel Shuffle. MetaMix linearly combines features and labels of samples from both the support and query sets. For each class of samples, Channel Shuffle randomly replaces a subset of their channels with the corresponding ones from a different class. Theoretical studies show how task augmentation improves the generalization of meta-learning. Moreover, both MetaMix and Channel Shuffle outperform state-of-the-art results by a large margin across many datasets and are compatible with existing meta-learning algorithms. Huaxiu Yao, Long-Kai Huang, Linjun Zhang, Ying Wei 0001, James Zou 0001, Junzhou Huang, Zhenhui Li |
ICML | 3 |
| 2021 | Adversarial Training Helps Transfer Learning via Better RepresentationsabstractTransfer learning aims to leverage models pre-trained on source data to efficiently adapt to target setting, where only limited data are available for model fine-tuning. Recent works empirically demonstrate that adversarial training in the source data can improve the ability of models to transfer to new domains. However, why this happens is not known. In this paper, we provide a theoretical model to rigorously analyze how adversarial training helps transfer learning. We show that adversarial training in the source data generates provably better representations, so fine-tuning on top of this representation leads to a more accurate predictor of the target data. We further demonstrate both theoretically and empirically that semi-supervised learning in the source data can also improve transfer learning by similarly improving the representation. Moreover, performing adversarial training on top of semi-supervised learning can further improve transferability, suggesting that the two approaches have complementary benefits on representations. We support our theories with experiments on popular data sets and deep learning architectures. Zhun Deng, Linjun Zhang, Kailas Vodrahalli, Kenji Kawaguchi, James Zou 0001 |
NeurIPS | 2 |
| 2021 | A Central Limit Theorem for Differentially Private Query AnsweringabstractPerhaps the single most important use case for differential privacy is to privately answer numerical queries, which is usually achieved by adding noise to the answer vector. The central question is, therefore, to understand which noise distribution optimizes the privacy-accuracy trade-off, especially when the dimension of the answer vector is high. Accordingly, an extensive literature has been dedicated to the question and the upper and lower bounds have been successfully matched up to constant factors (Bun et al.,2018; Steinke & Ullman, 2017). In this paper, we take a novel approach to address this important optimality question. We first demonstrate an intriguing central limit theorem phenomenon in the high-dimensional regime. More precisely, we prove that a mechanism is approximately Gaussian Differentially Private (Dong et al., 2021) if the added noise satisfies certain conditions. In particular, densities proportional to $\mathrm{e}^{-\|x\|_p^\alpha}$, where $\|x\|_p$ is the standard $\ell_p$-norm, satisfies the conditions. Taking this perspective, we make use of the Cramer--Rao inequality and show an "uncertainty principle"-style result: the product of privacy parameter and the $\ell_2$-loss of the mechanism is lower bounded by the dimension. Furthermore, the Gaussian mechanism achieves the constant-sharp optimal privacy-accuracy trade-off among all such noises. Our findings are corroborated by numerical experiments. Jinshuo Dong, Weijie J. Su, Linjun Zhang |
NeurIPS | 3 |
| 2020 | Interpreting Robust Optimization via Adversarial Influence FunctionsabstractRobust optimization has been widely used in nowadays data science, especially in adversarial training. However, little research has been done to quantify how robust optimization changes the optimizers and the prediction losses comparing to standard training. In this paper, inspired by the influence function in robust statistics, we introduce the Adversarial Influence Function (AIF) as a tool to investigate the solution produced by robust optimization. The proposed AIF enjoys a closed-form and can be calculated efficiently. To illustrate the usage of AIF, we apply it to study model sensitivity — a quantity defined to capture the change of prediction losses on the natural data after implementing robust optimization. We use AIF to analyze how model complexity and randomized smoothing affect the model sensitivity with respect to specific models. We further derive AIF for kernel regressions, with a particular application to neural tangent kernels, and experimentally demonstrate the effectiveness of the proposed AIF. Lastly, the theories of AIF will be extended to distributional robust optimization. Zhun Deng, Cynthia Dwork, Jialiang Wang 0001, Linjun Zhang |
ICML | 4 |
| 2020 | Classifying functional nuclear images with convolutional neural networks: a surveyabstractFunctional imaging has successfully been applied to capture functional changes in the pathological tissues of a body in recent years. Nuclear medicine functional imaging has been used to acquire information about areas of concerns (e.g. lesions and organs) in a non‐invasive manner, enabling semi‐automated or automated decision‐making for disease diagnosis, treatment, evaluation, and prediction. Focusing on functional nuclear medicine images, in this study, the authors review existing work on the classification of single‐photon emission computed tomography, positron emission tomography, and their hybrid modalities with computed tomography and magnetic resonance imaging images by using convolutional neural network (CNN) techniques. Specifically, they first present an overview of nuclear imaging and the CNN technique, such as nuclear imaging modalities, nuclear image data format, CNN architecture, and the main CNN classification models. According to the diseases of concern, they then classify the existing CNN‐based work on the classification of functional nuclear images into three different categories. For the typical work in each of these categories, they present details about their research objectives, adopted CNN models, and achieved main results. Finally, they discuss research challenges and directions for developing technological solutions to classify nuclear medicine images based on the CNN technique. Qiang Lin 0001, Zhengxing Man, Yongchun Cao, Chengcheng Han 0003, Chuangui Cao, Linjun Zhang, Sitao Zeng, Ruiting Gao, Weilan Wang, Jinshui Ji, Xiaodi Huang 0001 |
IET Image Process. | 7 |
| 2018 | Beyond-Line-of-Sight Identification by Using Vehicle-to-Vehicle CommunicationabstractIn this paper, we investigate the identification of the configuration and the dynamics of connected vehicle systems, where wireless vehicle-to-vehicle communication is used to access the motion data of vehicles that are beyond the line of sight. In particular, we first construct a causality detector to determine whether the information received from distant vehicles is relevant to the motion of the receiving vehicle. Then, we design a link-length estimator to identify the number of vehicles between the broadcasting vehicle and the receiving vehicle, which is required for appropriately incorporating the received data into the vehicle control system. Finally, a dynamics identifier is proposed to approximate the nonlinear time-delayed dynamics of vehicle chains, which is needed for the controller design to achieve desired system-level performance. The presented analytical results are validated through numerical simulations using synthetic data and on-road experiments. Linjun Zhang, Gábor Orosz |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Motif-Based Design for Connected Vehicle Systems in Presence of Heterogeneous Connectivity Structures and Time DelaysabstractIn this paper, we investigate the effects of heterogeneous connectivity structures and information delays on the dynamics of connected vehicle systems (CVSs), which are composed of vehicles equipped with connected cruise control (CCC) as well as conventional vehicles. First, a general framework is presented for CCC design that incorporates information delays and allows a large variety of connectivity structures. Then, we present delay-dependent criteria for plant stability and head-to-tail string stability of CVSs. The stability conditions are visualized by using stability diagrams, which allow one to evaluate the robustness of vehicle networks against information delays. To achieve modular and scalable design of large networks, we also propose a motif-based approach. Our results demonstrate the advantages of CCC vehicles in improving traffic efficiency, but also show that increasing the penetration of CCC vehicles does not necessarily improve the robustness if the connectivity structure or the control gains are not appropriately designed. Linjun Zhang, Gábor Orosz |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2014 | How is that complex network complex?abstractEvidence of complex networks in real world settings abounds. Many data sets for physical and social systems display characteristics consistent with various models of complex networks - the most typical examples being scale-free and small-world networks. However, theory does not always match reality. While we see a wide range of real complex networks, simulated data most usually comes from a limited range of generative models (the Barabási-Albert model for scale-free networks, Watt-Strogatz's model for small world networks, and Erdos-Renyi's model of a random graph are the three usual archetypes). We argue that there is much to be learnt by examining what real world data does that these algorithms do not. To do this we propose a variety of new network generation algorithms. These algorithms allow us to sample, in a statistically unbiased manner, from the family of all networks (of a given size N) consistent with a given degree distribution. Using this technique we are able to determine which distributions really are likely origins for various observed data and (equally importantly) observe when particular real world networks are atypical. Examples include the observation that many collaboration networks are not consistent with the Barabási-Albert (BA) model but are typical of the family of graphs that exhibit a power-law degree distribution, Biological networks (protein-protein interaction and cellular metabolic processes) are scale-free (but not BA) networks with atypically large diameter. Michael Small, Kevin Judd, Linjun Zhang |
ISCAS | 3 |