Yitong Li 0001

dblp:127/0252-1 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
2since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Generative modeling · 27% Graph learning · 19% Probabilistic and Bayesian machine learning · 12%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 28 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
generative adversarial network
1.132020
Sequential Attention GAN for Interactive Image Editing · ACM Multimedia 2020
StoryGAN: A Sequential Conditional GAN for Story Visualization · CVPR 2019
Video Generation From Text · AAAI 2018
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.822020
Sequential Attention GAN for Interactive Image Editing · ACM Multimedia 2020
StoryGAN: A Sequential Conditional GAN for Story Visualization · CVPR 2019
Machine learning › Graph learning
network embedding
0.822020
Dynamic Embedding on Textual Networks via a Gaussian Process · AAAI 2020
Diffusion Maps for Textual Network Embedding · NeurIPS 2018
Machine learning › Deep learning architectures and training
convolutional neural network
0.622018
Learning Context-Aware Convolutional Filters for Text Processing · EMNLP 2018
Targeting EEG/LFP Synchrony with Neural Nets · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
conditional density estimation
0.512021
Estimating Uncertainty Intervals from Collaborating Networks · J. Mach. Learn. Res. 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian prediction
predictive distribution
0.512021
Estimating Uncertainty Intervals from Collaborating Networks · J. Mach. Learn. Res. 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression
0.512021
Estimating Uncertainty Intervals from Collaborating Networks · J. Mach. Learn. Res. 2021
Machine learning › Trustworthy machine learning
uncertainty estimation
0.512021
Estimating Uncertainty Intervals from Collaborating Networks · J. Mach. Learn. Res. 2021
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.412020
Improving Disentangled Text Representation Learning with Information-Theoretic Guidance · ACL 2020
Machine learning › Graph learning
dynamic graph learning
0.412020
Dynamic Embedding on Textual Networks via a Gaussian Process · AAAI 2020
Machine learning › Graph learning
graph representation learning
0.412020
Dynamic Embedding on Textual Networks via a Gaussian Process · AAAI 2020
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.412019
Kernel-Based Approaches for Sequence Modeling: Connections to Neural Methods · NeurIPS 2019
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM
0.412019
Kernel-Based Approaches for Sequence Modeling: Connections to Neural Methods · NeurIPS 2019
Machine learning › Deep learning architectures and training
recurrent neural network
0.412019
Kernel-Based Approaches for Sequence Modeling: Connections to Neural Methods · NeurIPS 2019
Machine learning › Generative modeling › cross-modal generation
story visualization
0.412019
StoryGAN: A Sequential Conditional GAN for Story Visualization · CVPR 2019
Machine learning › Transfer learning and domain adaptation
domain generalization
0.312018
Extracting Relationships by Multi-Domain Matching · NeurIPS 2018
Machine learning › Graph learning › graph neural network › graph convolution
graph diffusion convolution
0.312018
Diffusion Maps for Textual Network Embedding · NeurIPS 2018
Machine learning › Graph learning
graph neural network
0.312018
Diffusion Maps for Textual Network Embedding · NeurIPS 2018
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation
0.312018
Extracting Relationships by Multi-Domain Matching · NeurIPS 2018
Natural language and speech › Information extraction and text analysis › text similarity
paraphrase identification
0.312018
Learning Context-Aware Convolutional Filters for Text Processing · EMNLP 2018
Natural language and speech › Language models and text generation › natural language understanding
sentence pair modeling
0.312018
Learning Context-Aware Convolutional Filters for Text Processing · EMNLP 2018
Natural language and speech › Information extraction and text analysis
text classification
0.312018
Learning Context-Aware Convolutional Filters for Text Processing · EMNLP 2018
Machine learning › Generative modeling › video generation
text-to-video generation
0.312018
Video Generation From Text · AAAI 2018
Machine learning › Generative modeling
variational autoencoder
0.312018
Video Generation From Text · AAAI 2018
Bioinformatics and computational biology › neuroscience › neuroinformatics › neural data analysis
neural signal analysis
0.312017
Targeting EEG/LFP Synchrony with Neural Nets · NIPS 2017
Machine learning › Generative modeling › cross-modal generation
text-conditioned generation
0.222020
Sequential Attention GAN for Interactive Image Editing · ACM Multimedia 2020
Video Generation From Text · AAAI 2018
Machine learning › Representation and self-supervised learning › text embedding
text representation learning
0.112020
Improving Disentangled Text Representation Learning with Information-Theoretic Guidance · ACL 2020
Information retrieval
text analysis
0.112018
Diffusion Maps for Textual Network Embedding · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

neural network · 0.5cumulative distribution function estimation · 0.5asymptotic consistency analysis · 0.5non-parametric model · 0.4neural state tracker · 0.4mutual information · 0.4information-theoretic guidance · 0.4gaussian process · 0.4diffusion · 0.4adversarial training · 0.4diffusion map · 0.3diffusion convolution · 0.3parameterized convolution · 0.3gaussian process adapter · 0.3
YearPublicationVenuePosition
2022 Learning to Weight Filter Groups for Robust Classification
abstract
In many real-world tasks, a canonical “big data” problem is created by combining data from several individual groups or domains. Because test data will likely come from a new group of data, we want to utilize the grouped structure of our training data to enforce generalization between groups of data, not just individual samples. This can be viewed as a multiple-domain generalization problem. Specifically, the goal is to encourage generalization between previously seen labeled source data from multiple domains and unlabeled target domain data. To address this challenge, we introduce Domain-Specific Filter Group (DSFG), where each training domain has a unique filter group and each test data point is predicted by a weighted sum over the outputs of different domain filters. A separate neural network learns to estimate the appropriate filter group weights through a meta-learning strategy. Empirically, experiments on three benchmark datasets demonstrate improved performance compared to current state-of-the-art approaches.
Siyang Yuan, Yitong Li 0001, Dong Wang 0037, Ke Bai 0001, Lawrence Carin, David E. Carlson
WACV2
2021 Estimating Uncertainty Intervals from Collaborating Networks
abstract
Effective decision making requires understanding the uncertainty inherent in a prediction. In regression, this uncertainty can be estimated by a variety of methods; however, many of these methods are laborious to tune, generate overconfident uncertainty intervals, or lack sharpness (give imprecise intervals). We address these challenges by proposing a novel method to capture predictive distributions in regression by defining two neural networks with two distinct loss functions. Specifically, one network approximates the cumulative distribution function, and the second network approximates its inverse. We refer to this method as Collaborating Networks (CN). Theoretical analysis demonstrates that a fixed point of the optimization is at the idealized solution, and that the method is asymptotically consistent to the ground truth distribution. Empirically, learning is straightforward and robust. We benchmark CN against several common approaches on two synthetic and six real-world datasets, including forecasting A1c values in diabetic patients from electronic health records, where uncertainty is critical. In the synthetic data, the proposed approach essentially matches ground truth. In the real-world datasets, CN improves results on many performance metrics, including log-likelihood estimates, mean absolute errors, coverage estimates, and prediction interval widths.
Tianhui Zhou, Yitong Li 0001, David E. Carlson
J. Mach. Learn. Res.2
2020 Dynamic Embedding on Textual Networks via a Gaussian Process
abstract
Textual network embedding aims to learn low-dimensional representations of text-annotated nodes in a graph. Prior work in this area has typically focused on fixed graph structures; however, real-world networks are often dynamic. We address this challenge with a novel end-to-end node-embedding model, called Dynamic Embedding for Textual Networks with a Gaussian Process (DetGP). After training, DetGP can be applied efficiently to dynamic graphs without re-training or backpropagation. The learned representation of each node is a combination of textual and structural embeddings. Because the structure is allowed to be dynamic, our method uses the Gaussian process to take advantage of its non-parametric properties. To use both local and global graph structures, diffusion is used to model multiple hops between neighbors. The relative importance of global versus local structure for the embeddings is learned automatically. With the non-parametric nature of the Gaussian process, updating the embeddings for a changed graph structure requires only a forward pass through the learned model. Considering link prediction and node classification, experiments demonstrate the empirical effectiveness of our method compared to baseline approaches. We further show that DetGP can be straightforwardly and efficiently applied to dynamic textual networks.
Pengyu Cheng, Yitong Li 0001, Xinyuan Zhang 0001, Liqun Chen 0001, David E. Carlson, Lawrence Carin
AAAI2
2020 Improving Disentangled Text Representation Learning with Information-Theoretic Guidance
abstract
Pengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon, Yizhe Zhang, Yitong Li, Lawrence Carin. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Pengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon, Yizhe Zhang 0002, Yitong Li 0001, Lawrence Carin
ACL6
2020 Sequential Attention GAN for Interactive Image Editing
abstract
Most existing text-to-image synthesis tasks are static single-turn generation, based on pre-defined textual descriptions of images. To explore more practical and interactive real-life applications, we introduce a new task - Interactive Image Editing, where users can guide an agent to edit images via multi-turn textual commands on-the-fly. In each session, the agent takes a natural language description from the user as the input, and modifies the image generated in previous turn to a new design, following the user description. The main challenges in this sequential and interactive image generation task are two-fold: 1) contextual consistency between a generated image and the provided textual description; 2) step-by-step region-level modification to maintain visual consistency across the generated image sequence in each session. To address these challenges, we propose a novel Sequential Attention Generative Adversarial Network (SeqAttnGAN), which applies a neural state tracker to encode the previous image and the textual description in each turn of the sequence, and uses a GAN framework to generate a modified version of the image that is consistent with the preceding images and coherent with the description. To achieve better region-specific refinement, we also introduce a sequential attention mechanism into the model. To benchmark on the new task, we introduce two new datasets, Zap-Seq and DeepFashion-Seq, which contain multi-turn sessions with image-description sequences in the fashion domain. Experiments on both datasets show that the proposed SeqAttnGAN model outperforms state-of-the-art approaches on the interactive image editing task across all evaluation metrics including visual quality, image sequence coherence and text-image consistency.
Yu Cheng 0001, Zhe Gan, Yitong Li 0001, Jingjing Liu 0001, Jianfeng Gao 0001
ACM Multimedia3
2019 On Target Shift in Adversarial Domain Adaptation
abstract
Discrepancy between training and testing domains is a fundamental problem in the generalization of machine learning techniques. Recently, several approaches have been proposed to learn domain invariant feature representations through adversarial deep learning. However, label shift, where the percentage of data in each class is different between domains, has received less attention. Label shift naturally arises in many contexts, especially in behavioral studies where the behaviors are freely chosen. In this work, we propose a method called Domain Adversarial nets for Target Shift (DATS) to address label shift while learning a domain invariant representation. This is accomplished by using distribution matching to estimate label proportions in a blind test set. We extend this framework to handle multiple domains by developing a scheme to upweight source domains most similar to the target domain. Empirical results show that this framework performs well under large label shift in synthetic and real experiments, demonstrating the practical importance.
Yitong Li 0001, Michael Murias, Samantha Major, Geraldine Dawson, David E. Carlson
AISTATS1
2019 StoryGAN: A Sequential Conditional GAN for Story Visualization
abstract
In this work, we propose a new task called Story Visualization. Given a multi-sentence paragraph, the story is visualized by generating a sequence of images, one for each sentence. In contrast to video generation, story visualization focuses less on the continuity in generated images (frames), but more on the global consistency across dynamic scenes and characters -- a challenge that has not been addressed by any single-image or video generation methods. Therefore, we propose a new story-to-image-sequence generation model, StoryGAN, based on the sequential conditional GAN framework. Our model is unique in that it consists of a deep Context Encoder that dynamically tracks the story flow, and two discriminators at the story and image levels, to enhance the image quality and the consistency of the generated sequences. To evaluate the model, we modified existing datasets to create the CLEVR-SV and Pororo-SV datasets. Empirically, StoryGAN outperformed state-of-the-art models in image quality, contextual consistency metrics, and human evaluation.
Yitong Li 0001, Zhe Gan, Yelong Shen, Jingjing Liu 0001, Yu Cheng 0001, Yuexin Wu, Lawrence Carin, David E. Carlson, Jianfeng Gao 0001
CVPR1
2019 Kernel-Based Approaches for Sequence Modeling: Connections to Neural Methods
abstract
We investigate time-dependent data analysis from the perspective of recurrent kernel machines, from which models with hidden units and gated memory cells arise naturally. By considering dynamic gating of the memory cell, a model closely related to the long short-term memory (LSTM) recurrent neural network is derived. Extending this setup to $n$-gram filters, the convolutional neural network (CNN), Gated CNN, and recurrent additive network (RAN) are also recovered as special cases. Our analysis provides a new perspective on the LSTM, while also extending it to $n$-gram convolutional filters. Experiments are performed on natural language processing tasks and on analysis of local field potentials (neuroscience). We demonstrate that the variants we derive from kernels perform on par or even better than traditional neural methods. For the neuroscience application, the new models demonstrate significant improvements relative to the prior state of the art.
Kevin J. Liang, Guoyin Wang 0002, Yitong Li 0001, Ricardo Henao, Lawrence Carin
NeurIPS3
2018 Video Generation From Text
abstract
Generating videos from text has proven to be a significant challenge for existing generative models. We tackle this problem by training a conditional generative model to extract both static and dynamic information from text. This is manifested in a hybrid framework, employing a Variational Autoencoder (VAE) and a Generative Adversarial Network (GAN). The static features, called "gist," are used to sketch text-conditioned background color and object layout structure. Dynamic features are considered by transforming input text into an image filter. To obtain a large amount of data for training the deep-learning model, we develop a method to automatically create a matched text-video corpus from publicly available online videos. Experimental results show that the proposed framework generates plausible and diverse short-duration smooth videos, while accurately reflecting the input text information. It significantly outperforms baseline models that directly adapt text-to-image generation procedures to produce videos. Performance is evaluated both visually and by adapting the inception score used to evaluate image generation in GANs.
Yitong Li 0001, Martin Renqiang Min, Dinghan Shen, David E. Carlson, Lawrence Carin
AAAI1
2018 Learning Context-Aware Convolutional Filters for Text Processing
abstract
Convolutional neural networks (CNNs) have recently emerged as a popular building block for natural language processing (NLP).Despite their success, most existing CNN models employed in NLP share the same learned (and static) set of filters for all input sentences.In this paper, we consider an approach of using a small meta network to learn contextaware convolutional filters for text processing.The role of meta network is to abstract the contextual information of a sentence or document into a set of input-aware filters.We further generalize this framework to model sentence pairs, where a bidirectional filter generation mechanism is introduced to encapsulate co-dependent sentence representations.In our benchmarks on four different tasks, including ontology classification, sentiment analysis, answer sentence selection, and paraphrase identification, our proposed model, a modified CNN with context-aware filters, consistently outperforms the standard CNN and attentionbased CNN baselines.By visualizing the learned context-aware filters, we further validate and rationalize the effectiveness of proposed framework.
Dinghan Shen, Martin Renqiang Min, Yitong Li 0001, Lawrence Carin
EMNLP3
2018 Extracting Relationships by Multi-Domain Matching
abstract
In many biological and medical contexts, we construct a large labeled corpus by aggregating many sources to use in target prediction tasks. Unfortunately, many of the sources may be irrelevant to our target task, so ignoring the structure of the dataset is detrimental. This work proposes a novel approach, the Multiple Domain Matching Network (MDMN), to exploit this structure. MDMN embeds all data into a shared feature space while learning which domains share strong statistical relationships. These relationships are often insightful in their own right, and they allow domains to share strength without interference from irrelevant data. This methodology builds on existing distribution-matching approaches by assuming that source domains are varied and outcomes multi-factorial. Therefore, each domain should only match a relevant subset. Theoretical analysis shows that the proposed approach can have a tighter generalization bound than existing multiple-domain adaptation approaches. Empirically, we show that the proposed methodology handles higher numbers of source domains (up to 21 empirically), and provides state-of-the-art performance on image, text, and multi-channel time series classification, including clinically relevant data of a novel treatment of Autism Spectrum Disorder.
Yitong Li 0001, Michael Murias, Geraldine Dawson, David E. Carlson
NeurIPS1
2018 Diffusion Maps for Textual Network Embedding
abstract
Textual network embedding leverages rich text information associated with the network to learn low-dimensional vectorial representations of vertices. Rather than using typical natural language processing (NLP) approaches, recent research exploits the relationship of texts on the same edge to graphically embed text. However, these models neglect to measure the complete level of connectivity between any two texts in the graph. We present diffusion maps for textual network embedding (DMTE), integrating global structural information of the graph to capture the semantic relatedness between texts, with a diffusion-convolution operation applied on the text inputs. In addition, a new objective function is designed to efficiently preserve the high-order proximity using the graph diffusion. Experimental results show that the proposed approach outperforms state-of-the-art methods on the vertex-classification and link-prediction tasks.
Xinyuan Zhang 0001, Yitong Li 0001, Dinghan Shen, Lawrence Carin
NeurIPS2
2017 Targeting EEG/LFP Synchrony with Neural Nets
abstract
We consider the analysis of Electroencephalography (EEG) and Local Field Potential (LFP) datasets, which are “big” in terms of the size of recorded data but rarely have sufficient labels required to train complex models (e.g., conventional deep learning methods). Furthermore, in many scientific applications, the goal is to be able to understand the underlying features related to the classification, which prohibits the blind application of deep networks. This motivates the development of a new model based on {\em parameterized} convolutional filters guided by previous neuroscience research; the filters learn relevant frequency bands while targeting synchrony, which are frequency-specific power and phase correlations between electrodes. This results in a highly expressive convolutional neural network with only a few hundred parameters, applicable to smaller datasets. The proposed approach is demonstrated to yield competitive (often state-of-the-art) predictive performance during our empirical tests while yielding interpretable features. Furthermore, a Gaussian process adapter is developed to combine analysis over distinct electrode layouts, allowing the joint processing of multiple datasets to address overfitting and improve generalizability. Finally, it is demonstrated that the proposed framework effectively tracks neural dynamics on children in a clinical trial on Autism Spectrum Disorder.
Yitong Li 0001, Michael Murias, Samantha Major, Geraldine Dawson, Kafui Dzirasa, Lawrence Carin, David E. Carlson
NIPS1