EDBT 2026 Demo / reviewers in the wild / expert
Ke Bai 0001
dblp:33/8570-1
· DBLP profile ↗
11ranked-venue papers
2as first author
5since 2021 · last 2023
0000-0001-6025-6903ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Representation and self-supervised learning · 24% Deep learning architectures and training · 16% Probabilistic and Bayesian machine learning · 14% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › text embedding › sentence embedding
contrastive sentence embedding |
0.7 | 1 | 2023 | OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence Embedding · EMNLP 2023 |
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
0.7 | 1 | 2023 | OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence Embedding · EMNLP 2023 |
Machine learning › Trustworthy machine learning › open-world recognition
open-set recognition |
0.6 | 1 | 2022 | Open World Classification with Adaptive Negative Samples · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis
text classification |
0.6 | 1 | 2022 | Open World Classification with Adaptive Negative Samples · EMNLP 2022 |
Machine learning › Deep learning architectures and training › regularization
optimal transport regularization |
0.4 | 1 | 2020 | Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020 |
Natural language and speech › Language models and text generation › text generation
reinforcement learning for text generation |
0.4 | 1 | 2020 | Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation |
0.4 | 1 | 2020 | Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2019 | Variational Annealing of GANs: A Langevin Perspective · ICML 2019 |
Machine learning › Optimization for machine learning
gradient estimation |
0.4 | 1 | 2019 | GO Gradient for Expectation-Based Objectives · ICLR (Poster) 2019 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.4 | 1 | 2019 | Variational Annealing of GANs: A Langevin Perspective · ICML 2019 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.1 | 1 | 2020 | Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 0.7one-versus-rest classifiers · 0.6adaptive negative sampling · 0.6reinforcement learning · 0.4optimal transport · 0.4maximum likelihood estimation · 0.4likelihood regularization · 0.4langevin dynamics · 0.4fenchel duality · 0.4diffusion process · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Estimating Total Correlation with Mutual Information EstimatorsabstractTotal correlation (TC) is a fundamental concept in information theory that measures statistical dependency among multiple random variables. Recently, TC has shown noticeable effectiveness as a regularizer in many learning tasks, where the correlation among multiple latent embeddings requires to be jointly minimized or maximized. However, calculating precise TC values is challenging, especially when the closed-form distributions of embedding variables are unknown. In this paper, we introduce a unified framework to estimate total correlation values with sample-based mutual information (MI) estimators. More specifically, we discover a relation between TC and MI and propose two types of calculation paths (tree-like and line-like) to decompose TC into MI terms. With each MI term being bounded, the TC values can be successfully estimated. Further, we provide theoretical analyses concerning the statistical consistency of the proposed TC estimators. Experiments are presented on both synthetic and real-world scenarios, where our estimators demonstrate effectiveness in all TC estimation, minimization, and maximization tasks. Ke Bai 0001, Pengyu Cheng, Weituo Hao, Ricardo Henao, Larry Carin |
AISTATS | 1 |
| 2023 | OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence EmbeddingabstractZhan Shi, Guoyin Wang, Ke Bai, Jiwei Li, Xiang Li, Qingjun Cui, Belinda Zeng, Trishul Chilimbi, Xiaodan Zhu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Guoyin Wang 0002, Ke Bai 0001, Jiwei Li 0001, Qingjun Cui, Belinda Zeng, Trishul Chilimbi, Xiaodan Zhu 0001 |
EMNLP | 3 |
| 2023 | Multiscale Visual-Attribute Co-Attention for Zero-Shot Image RecognitionabstractZero-shot image recognition aims to classify data from unseen classes, by exploring the association between visual features and the semantic representations of each class. Most existing approaches focus on learning a shared single-scale embedding space (often at the output layer of the network) for both visual and semantic features, ignoring a fact that different-scale visual features exhibit different semantics. In this article, we propose a multi-scale visual-attribute co-attention (mVACA) model, considering both visual-semantic alignment and visual discrimination at multiple scales. At each scale, a hybrid visual attention is realized by attribute-related attention and visual self-attention. The attribute-related attention is guided by a pseudo attribute vector inferred via a mutual information regularization (MIR). The visual self-attentive features further influence the attribute attention to emphasize visual-associated attributes. Leveraging multiscale visual discrimination, mVACA unifies standard zero-shot learning (ZSL) and generalized ZSL tasks in one framework, achieving state-of-the-art or competitive performance on several commonly used benchmarks of both setups. To better understand the interaction between images and attributes in mVACA, we also provide visualized analysis. Hao Zhang 0050, Zhengjue Wang, Yishi Xu, Pengyu Cheng, Ke Bai 0001, Bo Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Open World Classification with Adaptive Negative SamplesabstractOpen world classification is a task in natural language processing with key practical relevance and impact.Since the open or unknown category data only manifests in the inference phase, finding a model with a suitable decision boundary accommodating for the identification of known classes and discrimination of the open category is challenging.The performance of existing models is limited by the lack of effective open category data during the training stage or the lack of a good mechanism to learn appropriate decision boundaries.We propose an approach based on adaptive negative samples (ANS) designed to generate effective synthetic open category samples in the training stage and without requiring any prior knowledge or external datasets.Empirically, we find a significant advantage in using auxiliary one-versus-rest binary classifiers, which effectively utilize the generated negative samples and avoid the complex threshold-seeking stage in previous works.Extensive experiments on three benchmark datasets show that ANS achieves significant improvements over stateof-the-art methods. Ke Bai 0001, Guoyin Wang 0002, Jiwei Li 0001, Puyang Xu, Ricardo Henao, Lawrence Carin |
EMNLP | 1 |
| 2022 | Learning to Weight Filter Groups for Robust ClassificationabstractIn many real-world tasks, a canonical “big data” problem is created by combining data from several individual groups or domains. Because test data will likely come from a new group of data, we want to utilize the grouped structure of our training data to enforce generalization between groups of data, not just individual samples. This can be viewed as a multiple-domain generalization problem. Specifically, the goal is to encourage generalization between previously seen labeled source data from multiple domains and unlabeled target domain data. To address this challenge, we introduce Domain-Specific Filter Group (DSFG), where each training domain has a unique filter group and each test data point is predicted by a weighted sum over the outputs of different domain filters. A separate neural network learns to estimate the appropriate filter group weights through a meta-learning strategy. Empirically, experiments on three benchmark datasets demonstrate improved performance compared to current state-of-the-art approaches. Siyang Yuan, Yitong Li 0001, Dong Wang 0037, Ke Bai 0001, Lawrence Carin, David E. Carlson |
WACV | 4 |
| 2020 | Sequence Generation with Optimal-Transport-Enhanced Reinforcement LearningabstractReinforcement learning (RL) has been widely used to aid training in language generation. This is achieved by enhancing standard maximum likelihood objectives with user-specified reward functions that encourage global semantic consistency. We propose a principled approach to address the difficulties associated with RL-based solutions, namely, high-variance gradients, uninformative rewards and brittle training. By leveraging the optimal transport distance, we introduce a regularizer that significantly alleviates the above issues. Our formulation emphasizes the preservation of semantic features, enabling end-to-end training instead of ad-hoc fine-tuning, and when combined with RL, it controls the exploration space for more efficient model updates. To validate the effectiveness of the proposed solution, we perform a comprehensive evaluation covering a wide variety of NLP tasks: machine translation, abstractive text summarization and image caption, with consistent improvements over competing solutions. Liqun Chen 0001, Ke Bai 0001, Chenyang Tao, Yizhe Zhang 0002, Guoyin Wang 0002, Wenlin Wang, Ricardo Henao, Lawrence Carin |
AAAI | 2 |
| 2020 | Advancing weakly supervised cross-domain alignment with optimal transport
Siyang Yuan, Ke Bai 0001, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Chunyuan Li, Guoyin Wang 0002, Ricardo Henao, Lawrence Carin |
BMVC | 2 |
| 2019 | Adversarial Learning of a Sampler Based on an Unnormalized DistributionabstractFundamental aspects of adversarial learning are investigated, with learning based on samples from the target distribution (conventional GAN setup). With insights so garnered, adversarial learning is extended to the case for which one has access to an unnormalized form $u(x)$ of the target density function, but no samples. Further, new concepts in GAN regularization are developed, based on learning from samples or from $u(x)$. The proposed method is compared to alternative approaches, with encouraging results demonstrated across a range of applications, including deep soft Q-learning. Chunyuan Li, Ke Bai 0001, Jianqiao Li, Guoyin Wang 0002, Changyou Chen, Lawrence Carin |
AISTATS | 2 |
| 2019 | GO Gradient for Expectation-Based Objectives
Yulai Cong, Miaoyun Zhao, Ke Bai 0001, Lawrence Carin |
ICLR (Poster) | 3 |
| 2019 | Variational Annealing of GANs: A Langevin PerspectiveabstractThe generative adversarial network (GAN) has received considerable attention recently as a model for data synthesis, without an explicit specification of a likelihood function. There has been commensurate interest in leveraging likelihood estimates to improve GAN training. To enrich the understanding of this fast-growing yet almost exclusively heuristic-driven subject, we elucidate the theoretical roots of some of the empirical attempts to stabilize and improve GAN training with the introduction of likelihoods. We highlight new insights from variational theory of diffusion processes to derive a likelihood-based regularizing scheme for GAN training, and present a novel approach to train GANs with an unnormalized distribution instead of empirical samples. To substantiate our claims, we provide experimental evidence on how our theoretically-inspired new algorithms improve upon current practice. Chenyang Tao, Shuyang Dai, Liqun Chen 0001, Ke Bai 0001, Junya Chen, Chang Liu 0030, Ruiyi Zhang 0002, Georgiy V. Bobashev, Lawrence Carin |
ICML | 4 |
| 2019 | On Fenchel Mini-Max LearningabstractInference, estimation, sampling and likelihood evaluation are four primary goals of probabilistic modeling. Practical considerations often force modeling approaches to make compromises between these objectives. We present a novel probabilistic learning framework, called Fenchel Mini-Max Learning (FML), that accommodates all four desiderata in a flexible and scalable manner. Our derivation is rooted in classical maximum likelihood estimation, and it overcomes a longstanding challenge that prevents unbiased estimation of unnormalized statistical models. By reformulating MLE as a mini-max game, FML enjoys an unbiased training objective that (i) does not explicitly involve the intractable normalizing constant and (ii) is directly amendable to stochastic gradient descent optimization. To demonstrate the utility of the proposed approach, we consider learning unnormalized statistical models, nonparametric density estimation and training generative models, with encouraging empirical results presented. Chenyang Tao, Liqun Chen 0001, Shuyang Dai, Junya Chen, Ke Bai 0001, Dong Wang 0037, Jianfeng Feng, Wenlian Lu, Georgiy V. Bobashev, Lawrence Carin |
NeurIPS | 5 |