Ke Bai 0001

dblp:33/8570-1 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
5since 2021 · last 2023
0000-0001-6025-6903ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Representation and self-supervised learning · 24% Deep learning architectures and training · 16% Probabilistic and Bayesian machine learning · 14%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › text embedding › sentence embedding
contrastive sentence embedding
0.712023
OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence Embedding · EMNLP 2023
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding
0.712023
OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence Embedding · EMNLP 2023
Machine learning › Trustworthy machine learning › open-world recognition
open-set recognition
0.612022
Open World Classification with Adaptive Negative Samples · EMNLP 2022
Natural language and speech › Information extraction and text analysis
text classification
0.612022
Open World Classification with Adaptive Negative Samples · EMNLP 2022
Machine learning › Deep learning architectures and training › regularization
optimal transport regularization
0.412020
Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020
Natural language and speech › Language models and text generation › text generation
reinforcement learning for text generation
0.412020
Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation
0.412020
Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020
Machine learning › Generative modeling
generative adversarial network
0.412019
Variational Annealing of GANs: A Langevin Perspective · ICML 2019
Machine learning › Optimization for machine learning
gradient estimation
0.412019
GO Gradient for Expectation-Based Objectives · ICLR (Poster) 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.412019
Variational Annealing of GANs: A Langevin Perspective · ICML 2019
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.112020
Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020

Methods — techniques the papers use, named apart from their topics

contrastive learning · 0.7one-versus-rest classifiers · 0.6adaptive negative sampling · 0.6reinforcement learning · 0.4optimal transport · 0.4maximum likelihood estimation · 0.4likelihood regularization · 0.4langevin dynamics · 0.4fenchel duality · 0.4diffusion process · 0.4
YearPublicationVenuePosition
2023 Estimating Total Correlation with Mutual Information Estimators
abstract
Total correlation (TC) is a fundamental concept in information theory that measures statistical dependency among multiple random variables. Recently, TC has shown noticeable effectiveness as a regularizer in many learning tasks, where the correlation among multiple latent embeddings requires to be jointly minimized or maximized. However, calculating precise TC values is challenging, especially when the closed-form distributions of embedding variables are unknown. In this paper, we introduce a unified framework to estimate total correlation values with sample-based mutual information (MI) estimators. More specifically, we discover a relation between TC and MI and propose two types of calculation paths (tree-like and line-like) to decompose TC into MI terms. With each MI term being bounded, the TC values can be successfully estimated. Further, we provide theoretical analyses concerning the statistical consistency of the proposed TC estimators. Experiments are presented on both synthetic and real-world scenarios, where our estimators demonstrate effectiveness in all TC estimation, minimization, and maximization tasks.
Ke Bai 0001, Pengyu Cheng, Weituo Hao, Ricardo Henao, Larry Carin
AISTATS1
2023 OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence Embedding
abstract
Zhan Shi, Guoyin Wang, Ke Bai, Jiwei Li, Xiang Li, Qingjun Cui, Belinda Zeng, Trishul Chilimbi, Xiaodan Zhu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Guoyin Wang 0002, Ke Bai 0001, Jiwei Li 0001, Qingjun Cui, Belinda Zeng, Trishul Chilimbi, Xiaodan Zhu 0001
EMNLP3
2023 Multiscale Visual-Attribute Co-Attention for Zero-Shot Image Recognition
abstract
Zero-shot image recognition aims to classify data from unseen classes, by exploring the association between visual features and the semantic representations of each class. Most existing approaches focus on learning a shared single-scale embedding space (often at the output layer of the network) for both visual and semantic features, ignoring a fact that different-scale visual features exhibit different semantics. In this article, we propose a multi-scale visual-attribute co-attention (mVACA) model, considering both visual-semantic alignment and visual discrimination at multiple scales. At each scale, a hybrid visual attention is realized by attribute-related attention and visual self-attention. The attribute-related attention is guided by a pseudo attribute vector inferred via a mutual information regularization (MIR). The visual self-attentive features further influence the attribute attention to emphasize visual-associated attributes. Leveraging multiscale visual discrimination, mVACA unifies standard zero-shot learning (ZSL) and generalized ZSL tasks in one framework, achieving state-of-the-art or competitive performance on several commonly used benchmarks of both setups. To better understand the interaction between images and attributes in mVACA, we also provide visualized analysis.
Hao Zhang 0050, Zhengjue Wang, Yishi Xu, Pengyu Cheng, Ke Bai 0001, Bo Chen 0001
IEEE Trans. Neural Networks Learn. Syst.6
2022 Open World Classification with Adaptive Negative Samples
abstract
Open world classification is a task in natural language processing with key practical relevance and impact.Since the open or unknown category data only manifests in the inference phase, finding a model with a suitable decision boundary accommodating for the identification of known classes and discrimination of the open category is challenging.The performance of existing models is limited by the lack of effective open category data during the training stage or the lack of a good mechanism to learn appropriate decision boundaries.We propose an approach based on adaptive negative samples (ANS) designed to generate effective synthetic open category samples in the training stage and without requiring any prior knowledge or external datasets.Empirically, we find a significant advantage in using auxiliary one-versus-rest binary classifiers, which effectively utilize the generated negative samples and avoid the complex threshold-seeking stage in previous works.Extensive experiments on three benchmark datasets show that ANS achieves significant improvements over stateof-the-art methods.
Ke Bai 0001, Guoyin Wang 0002, Jiwei Li 0001, Puyang Xu, Ricardo Henao, Lawrence Carin
EMNLP1
2022 Learning to Weight Filter Groups for Robust Classification
abstract
In many real-world tasks, a canonical “big data” problem is created by combining data from several individual groups or domains. Because test data will likely come from a new group of data, we want to utilize the grouped structure of our training data to enforce generalization between groups of data, not just individual samples. This can be viewed as a multiple-domain generalization problem. Specifically, the goal is to encourage generalization between previously seen labeled source data from multiple domains and unlabeled target domain data. To address this challenge, we introduce Domain-Specific Filter Group (DSFG), where each training domain has a unique filter group and each test data point is predicted by a weighted sum over the outputs of different domain filters. A separate neural network learns to estimate the appropriate filter group weights through a meta-learning strategy. Empirically, experiments on three benchmark datasets demonstrate improved performance compared to current state-of-the-art approaches.
Siyang Yuan, Yitong Li 0001, Dong Wang 0037, Ke Bai 0001, Lawrence Carin, David E. Carlson
WACV4
2020 Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning
abstract
Reinforcement learning (RL) has been widely used to aid training in language generation. This is achieved by enhancing standard maximum likelihood objectives with user-specified reward functions that encourage global semantic consistency. We propose a principled approach to address the difficulties associated with RL-based solutions, namely, high-variance gradients, uninformative rewards and brittle training. By leveraging the optimal transport distance, we introduce a regularizer that significantly alleviates the above issues. Our formulation emphasizes the preservation of semantic features, enabling end-to-end training instead of ad-hoc fine-tuning, and when combined with RL, it controls the exploration space for more efficient model updates. To validate the effectiveness of the proposed solution, we perform a comprehensive evaluation covering a wide variety of NLP tasks: machine translation, abstractive text summarization and image caption, with consistent improvements over competing solutions.
Liqun Chen 0001, Ke Bai 0001, Chenyang Tao, Yizhe Zhang 0002, Guoyin Wang 0002, Wenlin Wang, Ricardo Henao, Lawrence Carin
AAAI2
2020 Advancing weakly supervised cross-domain alignment with optimal transport
Siyang Yuan, Ke Bai 0001, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Chunyuan Li, Guoyin Wang 0002, Ricardo Henao, Lawrence Carin
BMVC2
2019 Adversarial Learning of a Sampler Based on an Unnormalized Distribution
abstract
Fundamental aspects of adversarial learning are investigated, with learning based on samples from the target distribution (conventional GAN setup). With insights so garnered, adversarial learning is extended to the case for which one has access to an unnormalized form $u(x)$ of the target density function, but no samples. Further, new concepts in GAN regularization are developed, based on learning from samples or from $u(x)$. The proposed method is compared to alternative approaches, with encouraging results demonstrated across a range of applications, including deep soft Q-learning.
Chunyuan Li, Ke Bai 0001, Jianqiao Li, Guoyin Wang 0002, Changyou Chen, Lawrence Carin
AISTATS2
2019 GO Gradient for Expectation-Based Objectives
Yulai Cong, Miaoyun Zhao, Ke Bai 0001, Lawrence Carin
ICLR (Poster)3
2019 Variational Annealing of GANs: A Langevin Perspective
abstract
The generative adversarial network (GAN) has received considerable attention recently as a model for data synthesis, without an explicit specification of a likelihood function. There has been commensurate interest in leveraging likelihood estimates to improve GAN training. To enrich the understanding of this fast-growing yet almost exclusively heuristic-driven subject, we elucidate the theoretical roots of some of the empirical attempts to stabilize and improve GAN training with the introduction of likelihoods. We highlight new insights from variational theory of diffusion processes to derive a likelihood-based regularizing scheme for GAN training, and present a novel approach to train GANs with an unnormalized distribution instead of empirical samples. To substantiate our claims, we provide experimental evidence on how our theoretically-inspired new algorithms improve upon current practice.
Chenyang Tao, Shuyang Dai, Liqun Chen 0001, Ke Bai 0001, Junya Chen, Chang Liu 0030, Ruiyi Zhang 0002, Georgiy V. Bobashev, Lawrence Carin
ICML4
2019 On Fenchel Mini-Max Learning
abstract
Inference, estimation, sampling and likelihood evaluation are four primary goals of probabilistic modeling. Practical considerations often force modeling approaches to make compromises between these objectives. We present a novel probabilistic learning framework, called Fenchel Mini-Max Learning (FML), that accommodates all four desiderata in a flexible and scalable manner. Our derivation is rooted in classical maximum likelihood estimation, and it overcomes a longstanding challenge that prevents unbiased estimation of unnormalized statistical models. By reformulating MLE as a mini-max game, FML enjoys an unbiased training objective that (i) does not explicitly involve the intractable normalizing constant and (ii) is directly amendable to stochastic gradient descent optimization. To demonstrate the utility of the proposed approach, we consider learning unnormalized statistical models, nonparametric density estimation and training generative models, with encouraging empirical results presented.
Chenyang Tao, Liqun Chen 0001, Shuyang Dai, Junya Chen, Ke Bai 0001, Dong Wang 0037, Jianfeng Feng, Wenlian Lu, Georgiy V. Bobashev, Lawrence Carin
NeurIPS5