Shijing Si

dblp:254/1233 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
28since 2021 · last 2026
0000-0003-4346-2574ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Theory of computation · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Exploring the Correlation Between Patient Portal Messages and Inpatient/Outpatient Encounters: A Preliminary Investigation
Shijing Si, Jedrek Wosik, Xiuquan He
ICIC (28)1
2025 GPPT: Gaussian Process-infused Prompt Tuning for Vision-language Models
abstract
Pre-trained vision-language models (VLMs) have achieved remarkable success in image classification tasks, leveraging efficient prompt tuning methods. However, the reliability of fine-tuned VLMs in safety-critical scenarios remains a concern due to the under-explored issue of confidence calibration. To address this limitation, we introduce Gaussian Process-infused Prompt Tuning (GPPT), a novel framework that integrates a Gaussian process into the hidden representations of images. By utilizing random Fourier features (RFF) and Laplace approximation, GPPT enables end-to-end training and seamless integration with existing prompt tuning methods. Our extensive experiments on 11 diverse downstream datasets demonstrate that GPPT achieves competitive performance with deep ensembles in both prediction accuracy and calibration, while requiring only a fraction of the inference time.
Shijing Si, Haixia Sun 0001, Jiawen Gu
ICASSP1
2024 Leveraging BERT and TFIDF Features for Short Text Clustering via Alignment-Promoting Co-Training
abstract
BERT and TFIDF features excel in capturing rich semantics and important words, respectively.Since most existing clustering methods are solely based on the BERT model, they often fall short in utilizing keyword information, which, however, is very useful in clustering short texts.In this paper, we propose a CO-Training Clustering (COTC) framework to make use of the collective strengths of BERT and TFIDF features.Specifically, we develop two modules responsible for the clustering of BERT and TFIDF features, respectively.We use the deep representations and cluster assignments from the TFIDF module outputs to guide the learning of the BERT module, seeking to align them at both the representation and cluster levels.Reversely, we also use the BERT module outputs to train the TFIDF module, thus leading to the mutual promotion.We then show that the alternating co-training framework can be placed under a unified joint training objective, which allows the two modules to be connected tightly and the training signals to be propagated efficiently.Experiments on eight benchmark datasets show that our method outperforms current SOTA methods significantly.
Zetong Li, Qinliang Su, Shijing Si, Jianxing Yu
EMNLP3
2024 Text Feature Adversarial Learning for Text Generation With Knowledge Transfer From GPT2
abstract
Text generation is a key component of many natural language tasks. Motivated by the success of generative adversarial networks (GANs) for image generation, many text-specific GANs have been proposed. However, due to the discrete nature of text, these text GANs often use reinforcement learning (RL) or continuous relaxations to calculate gradients during learning, leading to high-variance or biased estimation. Furthermore, the existing text GANs often suffer from mode collapse (i.e., they have limited generative diversity). To tackle these problems, we propose a new text GAN model named text feature GAN (TFGAN), where adversarial learning is performed in a continuous text feature space. In the adversarial game, GPT2 provides the "true" features, while the generator of TFGAN learns from them. TFGAN is trained by maximum likelihood estimation on text space and adversarial learning on text feature space, effectively combining them into a single objective, while alleviating mode collapse. TFGAN achieves appealing performance in text generation tasks, and it can also be used as a flexible framework for learning text representations.
Hao Zhang 0050, Yulai Cong, Zhengjue Wang, Miaoyun Zhao, Liqun Chen 0001, Shijing Si, Ricardo Henao, Lawrence Carin
IEEE Trans. Neural Networks Learn. Syst.7
2023 On the Calibration and Uncertainty with Pólya-Gamma Augmentation for Dialog Retrieval Models
abstract
Deep neural retrieval models have amply demonstrated their power but estimating the reliability of their predictions remains challenging. Most dialog response retrieval models output a single score for a response on how relevant it is to a given question. However, the bad calibration of deep neural network results in various uncertainty for the single score such that the unreliable predictions always misinform user decisions. To investigate these issues, we present an efficient calibration and uncertainty estimation framework PG-DRR for dialog response retrieval models which adds a Gaussian Process layer to a deterministic deep neural network and recovers conjugacy for tractable posterior inference by Pólya-Gamma augmentation. Finally, PG-DRR achieves the lowest empirical calibration error (ECE) in the in-domain datasets and the distributional shift task while keeping R10@1 and MAP performance.
Shijing Si, Jianzong Wang, Ning Cheng 0001, Zhitao Li 0002, Jing Xiao 0006
AAAI2
2023 Pushing the Efficiency Limit Using Structured Sparse Convolutions
abstract
Weight pruning is among the most popular approaches for compressing deep convolutional neural networks. Recent work suggests that in a randomly initialized deep neural network, there exist sparse subnetworks that achieve performance comparable to the original network. Unfortunately, finding these subnetworks involves iterative stages of training and pruning, which can be computationally expensive. We propose Structured Sparse Convolution (SSC), that leverages the inherent structure in images to reduce the parameters in the convolutional filter. This leads to improved efficiency of convolutional architectures compared to existing methods that perform pruning at initialization. We show that SSC is a generalization of commonly used layers (depthwise, groupwise and pointwise convolution) in "efficient architectures." Extensive experiments on well-known CNN models and datasets show the effectiveness of the proposed method. Architectures based on SSC achieve state-of-the-art performance compared to baselines on CIFAR10, CIFAR-100, Tiny-ImageNet, and ImageNet classification benchmarks. Our source code is publicly available at https://github.com/vkvermaa/SSC.
Vinay Kumar Verma, Nikhil Mehta 0002, Shijing Si, Ricardo Henao, Lawrence Carin
WACV3
2022 Pose Guided Human Image Synthesis with Partially Decoupled GAN
Jianhan Wu 0001, Shijing Si, Jianzong Wang, Xiaoyang Qu, Jing Xiao 0006
ACML2
2022 Machine Unlearning Method Based On Projection Residual
abstract
Machine learning models (mainly neural networks) are used more and more in real life. Users feed their data to the model for training. But these processes are often one-way. Once trained, the model remembers the data. Even when data is removed from the dataset, the effects of these data persist in the model. With more and more laws and regulations around the world protecting data privacy, it becomes even more important to make models forget this data completely through machine unlearning.This paper adopts the projection residual method based on Newton iteration method. The main purpose is to implement machine unlearning tasks in the context of linear regression models and neural network models. This method mainly uses the iterative weighting method to completely forget the data and its corresponding influence, and its computational cost is linear in the feature dimension of the data. This method can improve the current machine learning method. At the same time, it is independent of the size of the training set. Results were evaluated by feature injection testing (FIT). Experiments show that this method is more thorough in deleting data, which is close to model retraining.
Zihao Cao, Jianzong Wang, Shijing Si, Zhangcheng Huang 0002, Jing Xiao 0006
DSAA3
2022 RL-MD: A Novel Reinforcement Learning Approach for DNA Motif Discovery
abstract
The extraction of sequence patterns from a collection of functionally linked unlabeled DNA sequences is known as DNA motif discovery, and it is a key task in computational biology. Several deep learning-based techniques have recently been introduced to address this issue. However, these algorithms can not be used in real-world situations because of the need for labeled data. Here, we presented RL-MD, a novel reinforcement learning based approach for DNA motif discovery task. RL-MD takes unlabelled data as input, employs a relative information-based method to evaluate each proposed motif, and utilizes these continuous evaluation results as the reward. The experiments show that RL-MD can identify high-quality motifs in real-world data.
Wen Wang 0025, Jianzong Wang, Shijing Si, Zhangcheng Huang 0002, Jing Xiao 0006
DSAA3
2022 Efficient Document Retrieval by End-to-End Refining and Quantizing BERT Embedding with Contrastive Product Quantization
abstract
Efficient document retrieval heavily relies on the technique of semantic hashing, which learns a binary code for every document and employs Hamming distance to evaluate document distances.However, existing semantic hashing methods are mostly established on outdated TFIDF features, which obviously do not contain lots of important semantic information about documents.Furthermore, the Hamming distance can only be equal to one of several integer values, significantly limiting its representational ability for document distances.To address these issues, in this paper, we propose to leverage BERT embeddings to perform efficient retrieval based on the product quantization technique, which will assign for every document a real-valued codeword from the codebook, instead of a binary code as in semantic hashing.Specifically, we first transform the original BERT embeddings via a learnable mapping and feed the transformed embedding into a probabilistic product quantization module to output the assigned codeword.The refining and quantizing modules can be optimized in an end-to-end manner by minimizing the probabilistic contrastive loss.A mutual information maximization based method is further proposed to improve the representativeness of codewords, so that documents can be quantized more accurately.Extensive experiments conducted on three benchmarks demonstrate that our proposed method significantly outperforms current state-of-the-art baselines 1 .
Zexuan Qiu, Qinliang Su, Jianxing Yu, Shijing Si
EMNLP4
2022 zkMLaaS: a Verifiable Scheme for Machine Learning as a Service
abstract
Machine Learning as a Service is a promising service for individuals and companies who would like to delegate model training to third parties. The customers desire proof of the integrity of the model training to prevent potential backdoor attacks launched by the server, while the server desires to prove the integrity without revealing their intellectual assets, hyper-parameters of the training scheme. Zero-knowledge proof, a cryptographic tool can theoretically satisfy the above demand, but is still practically infeasible due to the inefficiency of proving. Thus, we propose zkMLaaS, a privacy-preserving and verifiable scheme for efficient training proof generation in the MLaaS scenario. zkMLaaS features a two-round challenge-response pro-tocol equipped with the random sampling. This greatly reduces the time cost of proof generation and ensures the integrity of training procedure simultaneously. We analyze the security of zkMLaaS and conduct comprehensive evaluation which shows it saves around$273\times$times compared with naive scheme.
Jianzong Wang, Huangxun Chen, Shijing Si, Zhangcheng Huang 0002, Jing Xiao 0006
GLOBECOM4
2022 Towards Speaker Age Estimation With Label Distribution Learning
abstract
Existing methods for speaker age estimation usually treat it as a multi-class classification or a regression problem. However, precise age identification remains a challenge due to label ambiguity, i.e., utterances from adjacent age of the same person are often indistinguishable. To address this, we utilize the ambiguous information among the age labels, convert each age label into a discrete label distribution and leverage the label distribution learning (LDL) method to fit the data. For each audio data sample, our method produces a age distribution of its speaker, and on top of the distribution we also perform two other tasks: age prediction and age uncertainty minimization. Therefore, our method naturally combines the age classification and regression approaches, which enhances the robustness of our method. We conduct experiments on the public NIST SRE08-10 dataset and a real-world dataset, which exhibit that our method outperforms baseline methods by a relatively large margin, yielding a 10% reduction in terms of mean absolute error (MAE) on a real-world dataset.
Shijing Si, Jianzong Wang, Junqing Peng, Jing Xiao 0006
ICASSP1
2022 VU-BERT: A Unified Framework for Visual Dialog
abstract
The visual dialog task attempts to train an agent to answer multi-turn questions given an image, which requires the deep understanding of interactions between the image and dialog history. Existing researches tend to employ the modality-specific modules to model the interactions, which might be troublesome to use. To fill in this gap, we propose a unified framework for image-text joint embedding, named VU-BERT, and apply patch projection to obtain vision embedding firstly in visual dialog tasks to simplify the model. The model is trained over two tasks: masked language modeling and next utterance retrieval. These tasks help in learning visual concepts, utterances dependence, and the relationships between these two modalities. Finally, our VU-BERT achieves competitive performance (0.7287 NDCG scores) on VisDial v1.0 Datasets.
Shijing Si, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006
ICASSP2
2022 Boosting StarGANs for Voice Conversion with Contrastive Discriminator
Shijing Si, Jianzong Wang, Xulong Zhang 0001, Xiaoyang Qu, Ning Cheng 0001, Jing Xiao 0006
ICONIP (2)1
2022 A Nearest Neighbor Under-sampling Strategy for Vertical Federated Learning in Financial Domain
abstract
Machine learning techniques have been widely applied in modern financial activities. Participants in the field are aware of the importance of data privacy. Vertical federated learning (VFL) was proposed as a solution to multi-party secure computation for machine learning to obtain the huge data required by the models as well as keep the privacy of the data holders. However, previous research majorly analyzed the algorithms under ideal conditions. Data imbalance in VFL is still an open problem. In this paper, we propose a privacy-preserving sampling strategy for imbalanced VFL based on federated graph embedding of the samples, without leaking any distribution information. The participants of the federation provide partial neighbor information for each sample during the intersection stage and the controversial negative sample will be filtered out. Experiments were conducted on commonly used financial datasets and one real-world dataset. Our proposed approach obtained the leading F1 score on all tested datasets on comparing with the baseline under sampling strategies for VFL.
Denghao Li, Jianzong Wang, Lingwei Kong, Shijing Si, Zhangcheng Huang 0002, Jing Xiao 0006
IH&MMSec4
2022 Federated Split BERT for Heterogeneous Text Classification
abstract
Pre-trained BERT models have achieved impressive performance in many natural language processing (NLP) tasks. However, in many real-world situations, textual data are usually decentralized over many clients and unable to be uploaded to a central server due to privacy protection and regulations. Federated learning (FL) enables multiple clients collaboratively to train a global model while keeping the local data privacy. A few researches have investigated BERT in federated learning setting, but the problem of performance loss caused by heterogeneous (e.g., non-IID) data over clients remain under-explored. To address this issue, we propose a framework, FedSplitBERT, which handles heterogeneous data and decreases the communication cost by splitting the BERT encoder layers into local part and global part. The local part parameters are trained by the local client only while the global part parameters are trained by aggregating gradients of multiple clients. Due to the sheer size of BERT, we explore a quantization method to further reduce the communication cost with minimal performance loss. Our framework is ready-to-use and compatible to many existing federated learning algorithms, including FedAvg, FedProx and FedAdam. Our experiments verify the effectiveness of the proposed framework, which outperforms baseline methods by a significant margin, while FedSplitBERT with quantization can reduce the communication cost by 11.9×.
Shijing Si, Jianzong Wang, Jing Xiao 0006
IJCNN2
2022 Federated Non-negative Matrix Factorization for Short Texts Topic Modeling with Mutual Information
abstract
Non-negative matrix factorization (NMF) based topic modeling is widely used in natural language processing (NLP) to uncover hidden topics of short text documents. Usually, training a high-quality topic model requires large amount of textual data. In many real-world scenarios, customer textual data should be private and sensitive, precluding uploading to data centers. This paper proposes a Federated NMF (FedNMF) framework, which allows multiple clients to collaboratively train a high-quality NMF based topic model with locally stored data. However, standard federated learning will significantly undermine the performance of topic models in downstream tasks (e.g., text classification) when the data distribution over clients is heterogeneous. To alleviate this issue, we further propose FedNMF+MI, which simultaneously maximizes the mutual information (MI) between the count features of local texts and their topic weight vectors to mitigate the performance degradation. Experimental results show that our FedNMF+MI methods outperform Federated Latent Dirichlet Allocation (FedLDA) and the FedNMF without MI methods for short texts by a significant margin on both coherence score and classification F1 score.
Shijing Si, Jianzong Wang, Ruiyi Zhang 0002, Qinliang Su, Jing Xiao 0006
IJCNN1
2022 A Fair Federated Learning Framework With Reinforcement Learning
abstract
Federated learning (FL) is a paradigm where many clients collaboratively train a model under the coordination of a central server, while keeping the training data locally stored. However, heterogeneous data distributions over different clients remain a challenge to mainstream FL algorithms, which may cause slow convergence, overall performance degradation and unfairness of performance across clients. To address these problems, in this study we propose a reinforcement learning framework, called PG-FFL, which automatically learns a policy to assign aggregation weights to clients. Additionally, we propose to utilize Gini coefficient as the measure of fairness for FL. More importantly, we apply the Gini coefficient and validation accuracy of clients in each communication round to construct a reward function for the reinforcement learning. Our PG-FFL is also compatible to many existing FL algorithms. We conduct extensive experiments over diverse datasets to verify the effectiveness of our framework. The experimental results show that our framework can outperform baseline methods in terms of overall performance, fairness and convergence speed.
Yaqi Sun, Shijing Si, Jianzong Wang, Yuhan Dong, Zhitao Zhu, Jing Xiao 0006
IJCNN2
2022 Leveraging Causal Inference for Explainable Automatic Program Repair
abstract
Deep learning models have made significant progress in automatic program repair. However, the black-box nature of these methods has restricted their practical applications. To address this challenge, this paper presents an interpretable approach for program repair based on sequence-to-sequence models with causal inference and our method is called CPR, short for causal program repair. Our CPR can generate explanations in the process of decision making, which consists of groups of causally related input-output tokens. Firstly, our method infers these relations by querying the model with inputs disturbed by data augmentation. Secondly, it generates a graph over tokens from the responses and solves a partitioning problem to select the most relevant components. The experiments on four programming languages (Java, C, Python, and JavaScript) show that CPR can generate causal graphs for reasonable interpretations and boost the performance of bug fixing in automatic program repair.
Jianzong Wang, Shijing Si, Zhitao Zhu, Xiaoyang Qu, Zhenhou Hong, Jing Xiao 0006
IJCNN2
2022 Augmentation-induced Consistency Regularization for Classification
abstract
Deep neural networks have become popular in many supervised learning tasks, but they may suffer from overfitting when the training dataset is limited. To mitigate this, many researchers use data augmentation, which is a widely used and effective method for increasing the variety of datasets. However, the randomness introduced by data augmentation causes inevitable inconsistency between training and inference, which leads to poor improvement. In this paper, we propose a consistency regularization framework based on data augmentation, called CR-Aug, which forces the output distributions of different sub models generated by data augmentation to be consistent with each other. Specifically, CR-Aug evaluates the discrepancy between the output distributions of two augmented versions of each sample, and it utilizes a stop-gradient operation to minimize the consistency loss. We implement CR-Aug to image and audio classification tasks and conduct extensive experiments to verify its effectiveness in improving the generalization ability of classifiers. Our CR-Aug framework is ready-to-use, it can be easily adapted to many state-of-the-art network architectures. Our empirical results show that CR-Aug outperforms baseline methods by a significant margin.
Jianhan Wu 0001, Shijing Si, Jianzong Wang, Jing Xiao 0006
IJCNN2
2022 Improving Human Image Synthesis with Residual Fast Fourier Transformation and Wasserstein Distance
abstract
With the rapid development of the Metaverse, virtual humans have emerged, and human image synthesis and editing techniques, such as pose transfer, have recently become popular. Most of the existing techniques rely on GANs, which can generate good human images even with large variants and occlusions. But from our best knowledge, the existing state-of-the-art method still has the following problems: the first is that the rendering effect of the synthetic image is not realistic, such as poor rendering of some regions. And the second is that the training of GAN is unstable and slow to converge, such as model collapse. Based on the above two problems, we propose several methods to solve them. To improve the rendering effect, we use the Residual Fast Fourier Transform Block to replace the traditional Residual Block. Then, spectral normalization and Wasserstein distance are used to improve the speed and stability of GAN training. Experiments demonstrate that the methods we offer are effective at solving the problems listed above, and we get state-of-the-art scores in LPIPS and PSNR.
Jianhan Wu 0001, Shijing Si, Jianzong Wang, Jing Xiao 0006
IJCNN2
2022 Cali3F: Calibrated Fast Fair Federated Recommendation System
abstract
The increasingly stringent regulations on privacy protection have sparked interest in federated learning. As a distributed machine learning framework, it bridges isolated data islands by training a global model over devices while keeping data localized. Specific to recommendation systems, many federated recommendation algorithms have been proposed to realize the privacy-preserving collaborative recommendation. However, several constraints remain largely unexplored. One big concern is how to ensure fairness between participants of federated learning, that is, to maintain the uniformity of recommendation performance across devices. On the other hand, due to data heterogeneity and limited networks, additional challenges occur in the convergence speed. To address these problems, in this paper, we first propose a personalized federated recommendation system training algorithm to improve the recommendation performance fairness. Then we adopt a clustering-based aggregation method to accelerate the training process. Combining the two components, we proposed Cali3F, a calibrated fast and fair federated recommendation framework. Cali3F not only addresses the convergence problem by a within-cluster parameter sharing approach but also significantly boosts fairness by calibrating local models with the global model. We demonstrate the performance of Cali3F across standard benchmark datasets and explore the efficacy in comparison to traditional aggregation approaches.
Zhitao Zhu, Shijing Si, Jianzong Wang, Jing Xiao 0006
IJCNN2
2022 Uncertainty Calibration for Deep Audio Classifiers
abstract
Although deep Neural Networks (DNNs) have achieved tremendous success in audio classification tasks, their uncertainty calibration are still under-explored.A well-calibrated model should be accurate when it is certain about its prediction and indicate high uncertainty when it is likely to be inaccurate.In this work, we investigate the uncertainty calibration for deep audio classifiers.In particular, we empirically study the performance of popular calibration methods: (i) Monte Carlo Dropout, (ii) ensemble, (iii) focal loss, and (iv) spectralnormalized Gaussian process (SNGP), on audio classification datasets.To this end, we evaluate (i-iv) for the tasks of environment sound and music genre classification.Results indicate that uncalibrated deep audio classifiers may be over-confident, and SNGP performs the best and is very efficient on the two datasets of this paper.
Shijing Si, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006
INTERSPEECH2
2022 Debias the Black-Box: A Fair Ranking Framework via Knowledge Distillation
Zhitao Zhu, Shijing Si, Jianzong Wang, Yaodong Yang 0001, Jing Xiao 0006
WISE2
2021 FairFil: Contrastive Neural Debiasing Method for Pretrained Text Encoders
Pengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si, Lawrence Carin
ICLR4
2021 Speech2Video: Cross-Modal Distillation for Speech to Video Generation
abstract
This paper investigates a novel task of talking face video generation solely from speeches.The speech-to-video generation technique can spark interesting applications in entertainment, customer service, and human-computer-interaction industries.Indeed, the timbre, accent and speed in speeches could contain rich information relevant to speakers' appearance.The challenge mainly lies in disentangling the distinct visual attributes from audio signals.In this article, we propose a light-weight, cross-modal distillation method to extract disentangled emotional and identity information from unlabelled video inputs.The extracted features are then integrated by a generative adversarial network into talking face video clips.With carefully crafted discriminators, the proposed framework achieves realistic generation results.Experiments with observed individuals demonstrated that the proposed framework captures the emotional expressions solely from speeches, and produces spontaneous facial motion in the video output.Compared to the baseline method where speeches are combined with a static image of the speaker, the results of the proposed framework is almost indistinguishable.User studies also show that the proposed method outperforms the existing algorithms in terms of emotion expression in the generated videos.
Shijing Si, Jianzong Wang, Xiaoyang Qu, Ning Cheng 0001, Xinghua Zhu, Jing Xiao 0006
Interspeech1
2021 Variational Information Bottleneck for Effective Low-Resource Audio Classification
abstract
Large-scale deep neural networks (DNNs) such as convolutional neural networks (CNNs) have achieved impressive performance in audio classification for their powerful capacity and strong generalization ability. However, when training a DNN model on low-resource tasks, it is usually prone to overfitting the small data and learning too much redundant information. To address this issue, we propose to use variational information bottleneck (VIB) to mitigate overfitting and suppress irrelevant information. In this work, we conduct experiments on a 4-layer CNN. However, the VIB framework is ready-to-use and could be easily utilized with many other state-of-the-art network architectures. Evaluation on a few audio datasets shows that our approach significantly outperforms baseline methods, yielding _ 5:0% improvement in terms of classification accuracy in some low-source settings. Copyright © 2021 ISCA.
Shijing Si, Jianzong Wang, Huiming Sun, Jianhan Wu 0001, Chuanyao Zhang, Xiaoyang Qu, Ning Cheng 0001, Jing Xiao 0006
Interspeech1
2021 Case Study of Few-Shot Learning in Text Recognition Models
Jianzong Wang, Shijing Si, Zhenhou Hong, Xiaoyang Qu, Xinghua Zhu, Jing Xiao 0006
WISE (2)2
2020 Methods for Numeracy-Preserving Word Embeddings
abstract
Dhanasekar Sundararaman, Shijing Si, Vivek Subramanian, Guoyin Wang, Devamanyu Hazarika, Lawrence Carin. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Dhanasekar Sundararaman, Shijing Si, Vivek Subramanian, Guoyin Wang 0002, Devamanyu Hazarika, Lawrence Carin
EMNLP (1)2