Yikai Zhang 0003

dblp:215/3962-3 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-4924-7824ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PivotAlign: Improve Semi-Supervised Learning by Learning Intra-Class Heterogeneity and Aligning with Pivots
abstract
Self-supervised learning plays an important role in current state-of-the-art semi-supervised learning (SSL) methods. These methods learn inter-class heterogeneity among data and generate pseudo-labels based on class level representations. However, they often neglect intra-class heterogeneity, resulting in the under-exploitation of finer-grained semantic relationships within classes. To address this limitation, we introduce PivotAlign, a novel SSL approach that aims to 1) learn hierarchical representations to detect both interclass and intra-class semantic relationships, and 2) refine pseudo-labels based on learned representations with a class-debiasing strategy. Specifically, we first learn a set of pivots as sub-prototypes of classes. We then train representations so that features align with the assigned pivot and are hierarchically grouped based on both inter-class and intra-class heterogeneity. This allows us to capture both inter-class and intra-class semantic relationships among data and leverage them to better assign and refine pseudo-labels. Additionally, since SSL methods are prone to bias toward classes that are easier to learn, we further re-balance class predictions to alleviate this class bias. We demonstrate the effectiveness of PivotAlign on various SSL benchmarks, where PivotAlign achieves state-of-the-art performances. The source code will be released upon publication of the work.
Lingjie Yi, Tao Sun 0009, Yikai Zhang 0003, Songzhu Zheng, Weimin Lyu, Haibin Ling, Chao Chen 0012
WACV3
2024 Dissecting Human and LLM Preferences
abstract
As a relative quality comparison of model responses, human and Large Language Model (LLM) preferences serve as common alignment goals in model fine-tuning and criteria in evaluation.Yet, these preferences merely reflect broad tendencies, resulting in less explainable and controllable models with potential safety risks.In this work, we dissect the preferences of human and 32 different LLMs to understand their quantitative composition, using annotations from real-world user-model conversations for a fine-grained, scenario-wise analysis.We find that humans are less sensitive to errors, favor responses that support their stances, and show clear dislike when models admit their limits.On the contrary, advanced LLMs like GPT-4-Turbo emphasize correctness, clarity, and harmlessness more.Additionally, LLMs of similar sizes tend to exhibit similar preferences, regardless of their training methods, and fine-tuning for alignment does not significantly alter the preferences of pretrained-only LLMs.Finally, we show that preference-based evaluation can be intentionally manipulated.In both training-free and training-based settings, aligning a model with the preferences of judges boosts scores, while injecting the least preferred properties lowers them.This results in notable score shifts: up to 0.59 on MT-Bench (1-10 scale) and 31.94 on AlpacaEval 2.0 (0-100 scale), highlighting the significant impact of this strategic adaptation.We have made all resources of this project publicly available.
Shichao Sun, Yikai Zhang 0003, Hai Zhao 0001, Pengfei Liu 0003
ACL (1)4
2024 OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
abstract
The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abilities in problem-solving and scientific discovery (i.e., AI4Science) once exclusive to human intellect. To comprehensively evaluate current models' performance in cognitive reasoning abilities, we introduce OlympicArena, which includes 11,163 bilingual problems across both text-only and interleaved text-image modalities. These challenges encompass a wide range of disciplines spanning seven fields and 62 international Olympic competitions, rigorously examined for data leakage. We argue that the challenges in Olympic competition problems are ideal for evaluating AI's cognitive reasoning due to their complexity and interdisciplinary nature, which are essential for tackling complex scientific challenges and facilitating discoveries. Beyond evaluating performance across various disciplines using answer-only criteria, we conduct detailed experiments and analyses from multiple perspectives. We delve into the models' cognitive reasoning abilities, their performance across different modalities, and their outcomes in process-level evaluations, which are vital for tasks requiring complex reasoning with lengthy solutions. Our extensive evaluations reveal that even advanced models like GPT-4o only achieve a 39.97\% overall accuracy (28.67\% for mathematics and 29.71\% for physics), illustrating current AI limitations in complex reasoning and multimodal integration. Through the OlympicArena, we aim to advance AI towards superintelligence, equipping it to address more complex challenges in science and beyond. We also provide a comprehensive set of resources to support AI research, including a benchmark dataset, an open-source annotation platform, a detailed evaluation tool, and a leaderboard with automatic submission features.
Zengzhi Wang, Shijie Xia, Xuefeng Li 0003, Haoyang Zou, Ruijie Xu 0005, Run-Ze Fan, Lyumanshan Ye, Ethan Chern, Yixin Ye, Yikai Zhang 0003, Yuqing Yang 0004, Binjie Wang, Shichao Sun, Yiyuan Li, Steffi Chern, Yiwei Qin, Jiadi Su, Yixiu Liu, Shaoting Zhang 0001, Dahua Lin, Yu Qiao 0001, Pengfei Liu 0003
NeurIPS11
2023 Risk Bounds on Aleatoric Uncertainty Recovery
abstract
Quantifying aleatoric uncertainty is a challenging task in machine learning. It is important for decision making associated with data-dependent uncertainty in model outcomes. Recently, many empirical studies in modeling aleatoric uncertainty under regression settings primarily rely on either a Gaussian likelihood or moment matching. However, the performance of these methods varies for different datasets whereas discussions on their theoretical guarantees are lacking. In this work, we investigate theoretical aspects of these approaches and establish risk bounds for their estimates. We provide conditions that are sufficient to guarantee the PAC-learnablility of the aleatoric uncertainty. The study suggests that the likelihood and moment matching-based methods enjoy different types of guarantee in their risk bounds, i.e., they calibrate different aspects of the uncertainty and thus exhibit distinct properties in different regimes of the parameter space. Finally, we conduct empirical study which shows promising results and supports our theorems.
Yikai Zhang 0003, Jiahe Lin, Fengpei Li, Yeshaya Adler, Kashif Rasul, Anderson Schneider, Yuriy Nevmyvaka
AISTATS1
2023 Learning to Segment from Noisy Annotations: A Spatial Correction Approach
Jiachen Yao, Yikai Zhang 0003, Songzhu Zheng, Mayank Goswami 0001, Prateek Prasanna, Chao Chen 0012
ICLR2
2023 Provably Convergent Schrödinger Bridge with Applications to Probabilistic Time Series Imputation
abstract
The Schrödinger bridge problem (SBP) is gaining increasing attention in generative modeling and showing promising potential even in comparison with the score-based generative models (SGMs). SBP can be interpreted as an entropy-regularized optimal transport problem, which conducts projections onto every other marginal alternatingly. However, in practice, only approximated projections are accessible and their convergence is not well understood. To fill this gap, we present a first convergence analysis of the Schrödinger bridge algorithm based on approximated projections. As for its practical applications, we apply SBP to probabilistic time series imputation by generating missing values conditioned on observed data. We show that optimizing the transport cost improves the performance and the proposed algorithm achieves the state-of-the-art result in healthcare and environmental data while exhibiting the advantage of exploring both temporal and feature patterns in probabilistic time series imputation.
Wei Deng 0002, Shikai Fang, Fengpei Li, Nicole Tianjiao Yang, Yikai Zhang 0003, Kashif Rasul, Shandian Zhe, Anderson Schneider, Yuriy Nevmyvaka
ICML6
2023 Topology-Aware Uncertainty for Image Segmentation
abstract
Segmentation of curvilinear structures such as vasculature and road networks is challenging due to relatively weak signals and complex geometry/topology. To facilitate and accelerate large scale annotation, one has to adopt semi-automatic approaches such as proofreading by experts. In this work, we focus on uncertainty estimation for such tasks, so that highly uncertain, and thus error-prone structures can be identified for human annotators to verify. Unlike most existing works, which provide pixel-wise uncertainty maps, we stipulate it is crucial to estimate uncertainty in the units of topological structures, e.g., small pieces of connections and branches. To achieve this, we leverage tools from topological data analysis, specifically discrete Morse theory (DMT), to first capture the structures, and then reason about their uncertainties. To model the uncertainty, we (1) propose a joint prediction model that estimates the uncertainty of a structure while taking the neighboring structures into consideration (inter-structural uncertainty); (2) propose a novel Probabilistic DMT to model the inherent uncertainty within each structure (intra-structural uncertainty) by sampling its representations via a perturb-and-walk scheme. On various 2D and 3D datasets, our method produces better structure-wise uncertainty maps compared to existing works. Code available at: https://github.com/Saumya-Gupta-26/struct-uncertainty
Saumya Gupta, Yikai Zhang 0003, Xiaoling Hu 0002, Prateek Prasanna, Chao Chen 0012
NeurIPS2
2022 A Manifold View of Adversarial Risk
abstract
The adversarial risk of a machine learning model has been widely studied. Most previous works assume that the data lies in the whole ambient space. We propose to take a new angle and take the manifold assumption into consideration. Assuming data lies in a manifold, we investigate two new types of adversarial risk, the normal adversarial risk due to perturbation along normal direction, and the in-manifold adversarial risk due to perturbation within the manifold. We prove that the classic adversarial risk can be bounded from both sides using the normal and in-manifold adversarial risks. We also show with a surprisingly pessimistic case that the standard adversarial risk can be nonzero even when both normal and in-manifold risks are zero. We finalize the paper with empirical studies supporting our theoretical results. Our results suggest the possibility of improving the robustness of a classifier by only focusing on the normal adversarial risk.
Yikai Zhang 0003, Xiaoling Hu 0002, Mayank Goswami 0001, Chao Chen 0012, Dimitris N. Metaxas
AISTATS2
2022 Stability of SGD: Tightness analysis and improved bounds
abstract
Stochastic Gradient Descent (SGD) based methods have been widely used for training large-scale machine learning models that also generalize well in practice. Several explanations have been offered for this generalization performance, a prominent one being algorithmic stability Hardt et al [2016]. However, there are no known examples of smooth loss functions for which the analysis can be shown to be tight. Furthermore, apart from properties of the loss function, data distribution has also been shown to be an important factor in generalization performance. This raises the question: is the stability analysis of Hardt et al [2016] tight for smooth functions, and if not, for what kind of loss functions and data distributions can the stability analysis be improved? In this paper we first settle open questions regarding tightness of bounds in the data-independent setting: we show that for general datasets, the existing analysis for convex and strongly-convex loss functions is tight, but it can be improved for non-convex loss functions. Next, we give novel and improved data-dependent bounds: we show stability upper bounds for a large class of convex regularized loss functions, with negligible regularization parameters, and improve existing data-dependent bounds in the non-convex setting. We hope that our results will initiate further efforts to better understand the data-dependent setting under non-convex loss functions, leading to an improved understanding of the generalization abilities of deep networks.
Yikai Zhang 0003, Sammy Bald, Vamsi Pingali, Chao Chen 0012, Mayank Goswami 0001
UAI1
2022 Algorithm 1024: Spherical Triangle Algorithm: A Fast Oracle for Convex Hull Membership Queries
abstract
The Convex Hull Membership (CHM) tests whether \( p \in conv(S) \) , where p and the n points of S lie in \( \mathbb { R}^m \) . CHM finds applications in Linear Programming, Computational Geometry, and Machine Learning. The Triangle Algorithm (TA), previously developed, in \( O(1/\varepsilon ^2) \) iterations computes \( p^{\prime } \in conv(S) \) , either an \( \varepsilon \) - approximate solution , or a witness certifying \( p \not\in conv(S) \) . We first prove the equivalence of exact and approximate versions of CHM and Spherical -CHM, where \( p=0 \) and \( \Vert v\Vert =1 \) for each v in S . If for some \( M \ge 1 \) every non-witness with \( \Vert p^{\prime }\Vert \gt \varepsilon \) admits \( v \in S \) satisfying \( \Vert p^{\prime } - v\Vert \ge \sqrt {1+\varepsilon /M} \) , we prove the number of iterations improves to \( O(M/\varepsilon) \) and \( M \le 1/\varepsilon \) always holds. Equivalence of CHM and Spherical-CHM implies Minimum Enclosing Ball (MEB) algorithms can be modified to solve CHM. However, we prove \( (1+ \varepsilon) \) -approximation in MEB is \( \Omega (\sqrt {\varepsilon }) \) -approximation in Spherical-CHM. Thus, even \( O(1/\varepsilon) \) iteration MEB algorithms are not superior to Spherical-TA. Similar weakness is proved for MEB core sets. Spherical-TA also results a variant of the All Vertex Triangle Algorithm (AVTA) for computing all vertices of \( conv(S) \) . Substantial computations on distinct problems demonstrate that TA and Spherical-TA generally achieve superior efficiency over algorithms such as Frank–Wolfe, MEB, and LP-Solver.
Bahman Kalantari, Yikai Zhang 0003
ACM Trans. Math. Softw.2
2021 Learning with Feature-Dependent Label Noise: A Progressive Approach
Yikai Zhang 0003, Songzhu Zheng, Pengxiang Wu, Mayank Goswami 0001, Chao Chen 0012
ICLR1
2021 Topological Detection of Trojaned Neural Networks
abstract
Deep neural networks are known to have security issues. One particular threat is the Trojan attack. It occurs when the attackers stealthily manipulate the model's behavior through Trojaned training samples, which can later be exploited. Guided by basic neuroscientific principles, we discover subtle -- yet critical -- structural deviation characterizing Trojaned models. In our analysis we use topological tools. They allow us to model high-order dependencies in the networks, robustly compare different networks, and localize structural abnormalities. One interesting observation is that Trojaned models develop short-cuts from shallow to deep layers. Inspired by these observations, we devise a strategy for robust detection of Trojaned models. Compared to standard baselines it displays better performance on multiple benchmarks.
Songzhu Zheng, Yikai Zhang 0003, Hubert Wagner, Mayank Goswami 0001, Chao Chen 0012
NeurIPS2
2020 Local Regularizer Improves Generalization
abstract
Regularization plays an important role in generalization of deep learning. In this paper, we study the generalization power of an unbiased regularizor for training algorithms in deep learning. We focus on training methods called Locally Regularized Stochastic Gradient Descent (LRSGD). An LRSGD leverages a proximal type penalty in gradient descent steps to regularize SGD in training. We show that by carefully choosing relevant parameters, LRSGD generalizes better than SGD. Our thorough theoretical analysis is supported by experimental evidence. It advances our theoretical understanding of deep learning and provides new perspectives on designing training algorithms. The code is available at https://github.com/huiqu18/LRSGD.
Yikai Zhang 0003, Dimitris N. Metaxas, Chao Chen 0012
AAAI1
2020 Synthetic Learning: Learn From Distributed Asynchronized Discriminator GAN Without Sharing Medical Image Data
abstract
In this paper, we propose a data privacy-preserving and communication efficient distributed GAN learning framework named Distributed Asynchronized Discriminator GAN (AsynDGAN). Our proposed framework aims to train a central generator learns from distributed discriminator, and use the generated synthetic image solely to train the segmentation model. We validate the proposed framework on the application of health entities learning problem which is known to be privacy sensitive. Our experiments show that our approach: 1) could learn the real image’s distribution from multiple datasets without sharing the patient’s raw data. 2) is more efficient and requires lower bandwidth than other distributed deep learning methods. 3) achieves higher performance compared to the model trained by one real dataset, and almost the same performance compared to the model trained by all real datasets. 4) has provable guarantees that the generator could learn the distributed distribution in an all important fashion thus is unbiased.We release our AsynDGAN source code at: https://github.com/tommy-qichang/AsynDGAN
Yikai Zhang 0003, Mert R. Sabuncu, Chao Chen 0012, Tong Zhang 0001, Dimitris N. Metaxas
CVPR3
2020 Learn Distributed GAN with Temporary Discriminators
Yikai Zhang 0003, Zhennan Yan, Chao Chen 0012, Dimitris N. Metaxas
ECCV (27)2
2019 Taming the Noisy Gradient: Train Deep Neural Networks with Small Batch Sizes
abstract
Deep learning architectures are usually proposed with millions of parameters, resulting in a memory issue when training deep neural networks with stochastic gradient descent type methods using large batch sizes. However, training with small batch sizes tends to produce low quality solution due to the large variance of stochastic gradients. In this paper, we tackle this problem by proposing a new framework for training deep neural network with small batches/noisy gradient. During optimization, our method iteratively applies a proximal type regularizer to make loss function strongly convex. Such regularizer stablizes the gradient, leading to better training performance. We prove that our algorithm achieves comparable convergence rate as vanilla SGD even with small batch size. Our framework is simple to implement and can be potentially combined with many existing optimization algorithms. Empirical results show that our method outperforms SGD and Adam when batch size is small. Our implementation is available at https://github.com/huiqu18/TRAlgorithm.
Yikai Zhang 0003, Chao Chen 0012, Dimitris N. Metaxas
IJCAI1
2018 Robust Vertex Enumeration for Convex Hulls in High Dimensions
abstract
We design a fast and robust algorithm named {All Vertex Traingle Algorithm (AVTA)} for detecting the vertices of the convex hull of a set of points in high dimensions. Our proposed algorithm is very general and works for arbitrary convex hulls. In addition to being a fundamental problem in computational geometry and linear programming, vertex enumeration in high dimensions has numerous applications in machine learning. In particular, we apply AVTA to design new practical algorithms for topic models and non-negative matrix factorization. For topic models, our new algorithm leads to significantly better reconstruction of the topic-word matrix than state of the art approaches. Additionally, we provide a robust analysis of AVTA and empirically demonstrate that it can handle larger amounts of noise than existing methods. For non-negative matrix we show that AVTA is competitive with existing methods that are specialized for this task.
Pranjal Awasthi, Bahman Kalantari, Yikai Zhang 0003
AISTATS3