Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yu Cao 0014

dblp:68/6563-14 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Question answering and dialogue systems · 32% Representation and self-supervised learning · 22% Language models and text generation · 21%
Network and information security
1 paper
Security and privacy of machine learning · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
text embedding
0.912025
Enhancing Lexicon-Based Text Embeddings with Large Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model
0.812024
Meta-Task Prompting Elicits Embeddings from Large Language Models · ACL (1) 2024
Information retrieval › document processing › document analysis › document representation
text embedding
0.812024
Meta-Task Prompting Elicits Embeddings from Large Language Models · ACL (1) 2024
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.612022
A Model-agnostic Data Manipulation Method for Persona-based Dialogue Generation · ACL (1) 2022
Natural language and speech › Question answering and dialogue systems › dialogue generation
persona-based dialogue generation
0.612022
A Model-agnostic Data Manipulation Method for Persona-based Dialogue Generation · ACL (1) 2022
Security and privacy of machine learning
adversarial attack
0.612022
TASA: Deceiving Question Answering Models by Twin Answer Sentences Attack · EMNLP 2022
Security and privacy of machine learning › adversarial attack
textual adversarial attack
0.612022
TASA: Deceiving Question Answering Models by Twin Answer Sentences Attack · EMNLP 2022
Machine learning › Transfer learning and domain adaptation
domain generalization
0.412020
Unsupervised Domain Adaptation on Reading Comprehension · AAAI 2020
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.412020
Unsupervised Domain Adaptation on Reading Comprehension · AAAI 2020
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.412020
Unsupervised Domain Adaptation on Reading Comprehension · AAAI 2020
Natural language and speech › Language models and text generation › large language model
large language model representation
0.312025
Enhancing Lexicon-Based Text Embeddings with Large Language Models · ACL (1) 2025
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding
0.212024
Meta-Task Prompting Elicits Embeddings from Large Language Models · ACL (1) 2024
Machine learning › Deep learning architectures and training
data augmentation
0.212022
A Model-agnostic Data Manipulation Method for Persona-based Dialogue Generation · ACL (1) 2022
Machine learning › Trustworthy machine learning › robustness › data poisoning
training data manipulation
0.212022
A Model-agnostic Data Manipulation Method for Persona-based Dialogue Generation · ACL (1) 2022

Methods — techniques the papers use, named apart from their topics

prompt engineering · 1.5meta-task prompting · 1.5filtering · 1.1beam search · 1.1token embedding clustering · 0.9pooling strategies · 0.9bidirectional attention · 0.9perturbation · 0.6data distillation · 0.6data curriculum · 0.6data augmentation · 0.6
YearPublicationVenuePosition
2025 Enhancing Lexicon-Based Text Embeddings with Large Language Models
abstract
Recent large language models (LLMs) have demonstrated exceptional performance on general-purpose text embedding tasks.While dense embeddings have dominated related research, we introduce the first lexicon-based embeddings (LENS) leveraging LLMs that achieve competitive performance on these tasks.LENS consolidates the vocabulary space through token embedding clustering to handle the issue of token redundancy in LLM vocabularies.To further improve performance, we investigate bidirectional attention and various pooling strategies.Specifically, LENS simplifies lexical matching with redundant vocabularies by assigning each dimension to a specific token cluster, where semantically similar tokens are grouped together.Extensive experiments demonstrate that LENS outperforms dense embeddings on the Massive Text Embedding Benchmark (MTEB), delivering compact representations with dimensionality comparable to dense counterparts.Furthermore, LENS inherently supports efficient embedding dimension pruning without any specialized objectives like Matryoshka Representation Learning.Notably, combining LENS with dense embeddings achieves state-of-the-art performance on the retrieval subset of MTEB (i.e., BEIR). 1
Yibin Lei, Tao Shen 0001, Yu Cao 0014, Andrew Yates
ACL (1)3
2025 Code-switching finetuning: Bridging multilingual pretrained language models for enhanced cross-lingual performance
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yu Cao 0014, Weifeng Liu 0001
Eng. Appl. Artif. Intell.4
2024 Meta-Task Prompting Elicits Embeddings from Large Language Models
abstract
We introduce a new unsupervised text embedding method, Meta-Task Prompting with Explicit One-Word Limitation (MetaEOL), for generating high-quality sentence embeddings from Large Language Models (LLMs) without the need for model fine-tuning.Leveraging meta-task prompting, MetaEOL guides LLMs to produce embeddings through a series of carefully designed prompts that address multiple representational aspects.Our comprehensive experiments demonstrate that embeddings averaged from various meta-tasks are versatile embeddings that yield competitive performance on Semantic Textual Similarity (STS) benchmarks and excel in downstream tasks, surpassing contrastive-trained models.Our findings suggest a new scaling law, offering a versatile and resource-efficient approach for embedding generation across diverse scenarios.1
Yibin Lei, Tianyi Zhou 0001, Tao Shen 0001, Yu Cao 0014, Chongyang Tao, Andrew Yates
ACL (1)5
2024 Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
abstract
The Mixture of Experts (MoE) has emerged as a highly successful technique in deep learning, based on the principle of divide-and-conquer to maximize model capacity without significant additional computational cost. Even in the era of large-scale language models (LLMs), MoE continues to play a crucial role, as some researchers have indicated that GPT-4 adopts the MoE structure to ensure diverse inference results. However, MoE is susceptible to performance degeneracy, particularly evident in the issues of imbalance and homogeneous representation among experts. While previous studies have extensively addressed the problem of imbalance, the challenge of homogeneous representation remains unresolved. In this study, we shed light on the homogeneous representation problem, wherein experts in the MoE fail to specialize and lack diversity, leading to frustratingly high similarities in their representations (up to 99% in a well-performed MoE model). This problem restricts the expressive power of the MoE and, we argue, contradicts its original intention. To tackle this issue, we propose a straightforward yet highly effective solution: OMoE, an orthogonal expert optimizer. Additionally, we introduce an alternating training strategy that encourages each expert to update in a direction orthogonal to the subspace spanned by other experts. Our algorithm facilitates MoE training in two key ways: firstly, it explicitly enhances representation diversity, and secondly, it implicitly fosters interaction between experts during orthogonal weights computation. Through extensive experiments, we demonstrate that our proposed optimization algorithm significantly improves the performance of fine-tuning the MoE model on the GLUE benchmark, SuperGLUE benchmark, question-answering task, and name entity recognition tasks.
Boan Liu, Liang Ding 0006, Li Shen 0008, Keqin Peng, Yu Cao 0014, Dazhao Cheng, Dacheng Tao
ECAI5
2022 A Model-agnostic Data Manipulation Method for Persona-based Dialogue Generation
abstract
Towards building intelligent dialogue agents, there has been a growing interest in introducing explicit personas in generation models.However, with limited persona-based dialogue data at hand, it may be difficult to train a dialogue generation model well.We point out that the data challenges of this generation task lie in two aspects: first, it is expensive to scale up current persona-based dialogue datasets; second, each data sample in this task is more complex to learn with than conventional dialogue data.To alleviate the above data issues, we propose a data manipulation method, which is model-agnostic to be packed with any personabased dialogue generation model to improve its performance.The original training samples will first be distilled and thus expected to be fitted more easily.Next, we show various effective ways that can diversify such easier distilled data.A given base model will then be trained via the constructed data curricula, i.e. first on augmented distilled samples and then on original ones.Experiments illustrate the superiority of our method with two strong base dialogue models (Transformer encoderdecoder and GPT2).
Yu Cao 0014, Wei Bi, Shuming Shi 0001, Dacheng Tao
ACL (1)1
2022 On the Complementarity between Pre-Training and Random-Initialization for Resource-Rich Machine Translation
abstract
Pre-Training (PT) of text representations has been successfully applied to low-resource Neural Machine Translation (NMT). However, it usually fails to achieve notable gains (some- times, even worse) on resource-rich NMT on par with its Random-Initialization (RI) counterpart. We take the first step to investigate the complementarity between PT and RI in resource-rich scenarios via two probing analyses, and find that: 1) PT improves NOT the accuracy, but the generalization by achieving flatter loss landscapes than that of RI; 2) PT improves NOT the confidence of lexical choice, but the negative diversity by assigning smoother lexical probability distributions than that of RI. Based on these insights, we propose to combine their complementarities with a model fusion algorithm that utilizes optimal transport to align neurons between PT and RI. Experiments on two resource-rich translation benchmarks, WMT’17 English-Chinese (20M) and WMT’19 English-German (36M), show that PT and RI could be nicely complementary to each other, achieving substantial improvements considering both translation accuracy, generalization, and negative diversity. Probing tools and code are released at: https://github.com/zanchangtong/PTvsRI.
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yu Cao 0014, Weifeng Liu 0001, Dacheng Tao
COLING4
2022 TASA: Deceiving Question Answering Models by Twin Answer Sentences Attack
abstract
We present Twin Answer Sentences Attack (TASA), an adversarial attack method for question answering (QA) models that produces fluent and grammatical adversarial contexts while maintaining gold answers.Despite phenomenal progress on general adversarial attacks, few works have investigated the vulnerability and attack specifically for QA models.In this work, we first explore the biases in the existing models and discover that they mainly rely on keyword matching between the question and context, and ignore the relevant contextual relations for answer prediction.Based on two biases above, TASA attacks the target model in two folds: (1) lowering the model's confidence on the gold answer with a perturbed answer sentence; (2) misguiding the model towards a wrong answer with a distracting answer sentence.Equipped with designed beam search and filtering methods, TASA can generate more effective attacks than existing textual attack methods while sustaining the quality of contexts, in extensive experiments on five QA datasets and human evaluations.
Yu Cao 0014, Dianqi Li, Tianyi Zhou 0001, Yibing Zhan, Dacheng Tao
EMNLP1
2021 Towards Efficiently Diversifying Dialogue Generation Via Embedding Augmentation
abstract
Dialogue generation models face the challenge of producing generic and repetitive responses. Unlike previous augmentation methods that mostly focus on token manipulation and ignore the essential variety within a single sample using hard labels, we propose to promote the generation diversity of the neural dialogue models via soft embedding augmentation along with soft labels in this paper. Particularly, we select some key input tokens and fuse their embeddings together with embeddings from their semantic-neighbor tokens. The new embeddings serve as the input of the model to replace the original one. Besides, soft labels are used in loss calculation, resulting in multi-target supervision for a given input. Our experimental results on two datasets illustrate that our proposed method is capable of generating more diverse responses than raw models while remains a similar n-gram accuracy that ensures the quality of generated responses.
Yu Cao 0014, Liang Ding 0006, Zhiliang Tian
ICASSP1
2021 DAGN: Discourse-Aware Graph Network for Logical Reasoning
abstract
Yinya Huang, Meng Fang, Yu Cao, Liwei Wang, Xiaodan Liang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yinya Huang, Yu Cao 0014, Xiaodan Liang
NAACL-HLT3
2020 Unsupervised Domain Adaptation on Reading Comprehension
abstract
Reading comprehension (RC) has been studied in a variety of datasets with the boosted performance brought by deep neural networks. However, the generalization capability of these models across different domains remains unclear. To alleviate the problem, we investigate unsupervised domain adaptation on RC, wherein a model is trained on the labeled source domain and to be applied to the target domain with only unlabeled samples. We first show that even with the powerful BERT contextual representation, a model can not generalize well from one domain to another. To solve this, we provide a novel conditional adversarial self-training method (CASe). Specifically, our approach leverages a BERT model fine-tuned on the source dataset along with the confidence filtering to generate reliable pseudo-labeled samples in the target domain for self-training. On the other hand, it further reduces domain distribution discrepancy through conditional adversarial learning across domains. Extensive experiments show our approach achieves comparable performance to supervised models on multiple large-scale benchmark datasets.
Yu Cao 0014, Baosheng Yu, Joey Tianyi Zhou
AAAI1