Yang Chen 0065

dblp:48/4792-65 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2025
0009-0003-5667-7803ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Information extraction and text analysis · 50% Vision and language · 28% Transfer learning and domain adaptation · 8%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 25 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
named entity recognition
1.532025
Translation and Fusion Improves Cross-lingual Information Extraction · ACL (1) 2025
An Empirical Study of Pre-trained Transformers for Arabic Information Extraction · EMNLP (1) 2020
Constrained Decoding for Cross-lingual Label Projection · ICLR 2024
Computer vision › Vision and language
vision-language pretraining
1.322023
Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities · ICCV 2023
Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · EMNLP 2023
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual information extraction
0.912025
Translation and Fusion Improves Cross-lingual Information Extraction · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › named entity recognition
low-resource named entity recognition
0.912025
Translation and Fusion Improves Cross-lingual Information Extraction · ACL (1) 2025
Natural language and speech › Language models and text generation › decoding
constrained decoding
0.812024
Constrained Decoding for Cross-lingual Label Projection · ICLR 2024
Natural language and speech › Information extraction and text analysis
geolocation
0.812024
Granular Privacy Control for Geolocation with Vision Language Models · EMNLP 2024
Computer vision › Vision and language
vision-language model
0.812024
Granular Privacy Control for Geolocation with Vision Language Models · EMNLP 2024
Information retrieval
multimodal retrieval
0.812024
UniIR: Training and Benchmarking Universal Multimodal Information Retrievers · ECCV (87) 2024
Information retrieval › multimodal retrieval
universal multimodal retrieval
0.812024
UniIR: Training and Benchmarking Universal Multimodal Information Retrievers · ECCV (87) 2024
Computer vision › Vision and language › visual question answering
knowledge-based visual question answering
0.712023
Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · EMNLP 2023
Natural language and speech › Information extraction and text analysis
misinformation detection
0.712023
Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 Treatments · ACL (1) 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning
multimodal knowledge
0.712023
Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · EMNLP 2023
Natural language and speech › Information extraction and text analysis
stance detection
0.712023
Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 Treatments · ACL (1) 2023
Computer vision › Vision and language
visual entity recognition
0.712023
Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities · ICCV 2023
Computer vision › Vision and language
visual question answering
0.712023
Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · EMNLP 2023
Information retrieval › fact-checking
claim detection
0.712023
Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 Treatments · ACL (1) 2023
Information retrieval
fact-checking
0.712023
Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 Treatments · ACL (1) 2023
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.512021
Model Selection for Cross-lingual Transfer · EMNLP (1) 2021
Machine learning › Learning theory
model selection
0.512021
Model Selection for Cross-lingual Transfer · EMNLP (1) 2021
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.512021
Model Selection for Cross-lingual Transfer · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging
0.412020
An Empirical Study of Pre-trained Transformers for Arabic Information Extraction · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis
relation extraction
0.412020
An Empirical Study of Pre-trained Transformers for Arabic Information Extraction · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis › event extraction
event argument extraction
0.212024
Constrained Decoding for Cross-lingual Label Projection · ICLR 2024
Information retrieval
retrieval models
0.212023
Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · EMNLP 2023
Machine learning › Transfer learning and domain adaptation › cross-lingual transfer
zero-shot cross-lingual transfer
0.112020
An Empirical Study of Pre-trained Transformers for Arabic Information Extraction · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

fine-tuning · 3.3instruction tuning · 2.4prompting · 1.5contrastive learning · 1.5stance classifiers · 1.3human-in-the-loop evaluation · 1.3translation · 0.9annotation fusion · 0.9multilingual LLMs · 0.8constrained decoding · 0.8entity recognition · 0.7
YearPublicationVenuePosition
2025 Translation and Fusion Improves Cross-lingual Information Extraction
abstract
Large language models (LLMs) combined with instruction tuning have shown significant progress in information extraction (IE) tasks, exhibiting strong generalization capabilities to unseen datasets by following annotation guidelines.However, their applicability to lowresource languages remains limited due to lack of both labeled data for fine-tuning, and unlabeled text for pre-training.In this paper, we propose TransFusion, a framework in which models are fine-tuned to use English translations of low-resource language data, enabling more precise predictions through annotation fusion.Based on TransFusion, we introduce GoLLIE-TF, a cross-lingual instruction-tuned LLM for IE tasks, designed to close the performance gap between high and low-resource languages.Our experiments across twelve multilingual IE datasets spanning 50 languages demonstrate that GoLLIE-TF achieves better cross-lingual transfer over the base model.In addition, we show that TransFusion significantly improves low-resource language named entity recognition when applied to proprietary models such as GPT-4 (+5 F1) with a prompting approach, or fine-tuning different language models including decoder-only (+14 F1) and encoder-only (+13 F1) architectures.
Yang Chen 0065, Vedaant Shah, Alan Ritter
ACL (1)1
2024 UniIR: Training and Benchmarking Universal Multimodal Information Retrievers
Cong Wei 0001, Yang Chen 0065, Hexiang Hu, Ge Zhang 0009, Jie Fu 0001, Alan Ritter, Wenhu Chen
ECCV (87)2
2024 Granular Privacy Control for Geolocation with Vision Language Models
abstract
Vision Language Models (VLMs) are rapidly advancing in their capability to answer information-seeking questions. As these models are widely deployed in consumer applications, they could lead to new privacy risks due to emergent abilities to identify people in photos, geolocate images, etc. As we demonstrate, somewhat surprisingly, current open-source and proprietary VLMs are very capable image geolocators, making widespread geolocation with VLMs an immediate privacy risk, rather than merely a theoretical future concern. As a first step to address this challenge, we develop a new benchmark, GPTGeoChat, to test the capability of VLMs to moderate geolocation dialogues with users. We collect a set of 1,000 image geolocation conversations between in-house annotators and GPT-4v, which are annotated with the granularity of location information revealed at each turn. Using this new dataset we evaluate the ability of various VLMs to moderate GPT-4v geolocation conversations by determining when too much location information has been revealed. We find that custom fine-tuned models perform on par with prompted API-based models when identifying leaked location information at the country or city level, however fine-tuning on supervised data appears to be needed to accurately moderate finer granularities, such as the name of a restaurant or building.
Ethan Mendes, Yang Chen 0065, James Hays, Sauvik Das, Wei Xu 0004, Alan Ritter
EMNLP2
2024 Constrained Decoding for Cross-lingual Label Projection
abstract
Zero-shot cross-lingual transfer utilizing multilingual LLMs has become a popular learning paradigm for low-resource languages with no labeled training data. However, for NLP tasks that involve fine-grained predictions on words and phrases, the performance of zero-shot cross-lingual transfer learning lags far behind supervised fine-tuning methods. Therefore, it is common to exploit translation and label projection to further improve the performance by (1) translating training data that is available in a high-resource language (e.g., English) together with the gold labels into low-resource languages, and/or (2) translating test data in low-resource languages to a high-source language to run inference on, then projecting the predicted span-level labels back onto the original test data. However, state-of-the-art marker-based label projection methods suffer from translation quality degradation due to the extra label markers injected in the input to the translation model. In this work, we explore a new direction that leverages constrained decoding for label projection to overcome the aforementioned issues. Our new method not only can preserve the quality of translated texts but also has the versatility of being applicable to both translating training and translating test data strategies. This versatility is crucial as our experiments reveal that translating test data can lead to a considerable boost in performance compared to translating only training data. We evaluate on two cross-lingual transfer tasks, namely Named Entity Recognition and Event Argument Extraction, spanning 20 languages. The results demonstrate that our approach outperforms the state-of-the-art marker-based method by a large margin and also shows better performance than other label projection methods that rely on external word alignment.
Duong Minh Le, Yang Chen 0065, Alan Ritter, Wei Xu 0004
ICLR2
2023 Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 Treatments
abstract
We present a human-in-the-loop evaluation framework for fact-checking novel misinformation claims and identifying social media messages that support them.Our approach extracts check-worthy claims, which are aggregated and ranked for review.Stance classifiers are then used to identify tweets supporting novel misinformation claims, which are further reviewed to determine whether they violate relevant policies.To demonstrate the feasibility of our approach, we develop a baseline system based on modern NLP methods for human-in-the-loop fact-checking in the domain of COVID-19 treatments.We make our data 1 and detailed annotation guidelines available to support the evaluation of human-in-the-loop systems that identify novel misinformation directly from raw usergenerated content.
Ethan Mendes, Yang Chen 0065, Wei Xu 0004, Alan Ritter
ACL (1)2
2023 Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
abstract
Pre-trained vision and language models (Chen et al., 2023b,a;Dai et al., 2023; Li et al., 2023b) have demonstrated state-of-the-art capabilities over existing tasks involving images and texts, including visual question answering.However, it remains unclear whether these models possess the capability to answer questions that are not only querying visual content but knowledge-intensive and informationseeking.In this study, we introduce INFOS-EEK 1 , a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.Using INFOSEEK, we analyze various pre-trained visual question answering models and gain insights into their characteristics.Our findings reveal that state-of-the-art pre-trained multi-modal models (e.g., PaLI-X, BLIP2, etc.) face challenges in answering visual information-seeking questions, but finetuning on the INFOSEEK dataset elicits models to use fine-grained knowledge that was learned during their pre-training.Furthermore, we show that accurate visual entity recognition can be used to improve performance on INFOSEEK by retrieving relevant documents, showing a significant space for improvement.* Work done when interned at Google 1 Our dataset is available at https:// open-vision-language.github.io/infoseek/.Dataset OK-VQA ViQuAE INFOSEEK PaLM (Q-only) 23.8 31.5 5.6 Current SotA 66.1 22.1 18.2 Require Knowledge † 29.2% 95.2% 95.6% † :% of questions that require knowledge to answer.PaLM (Q-only): a question-only baseline using PaLM.
Yang Chen 0065, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, Ming-Wei Chang
EMNLP1
2023 Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities
abstract
Large-scale multi-modal pre-training models such as CLIP [30] and PaLI [8] exhibit strong generalization on various visual domains and tasks. However, existing image classification benchmarks often evaluate recognition on a specific domain (e.g., outdoor images) or a specific task (e.g., classifying plant species), which falls short of evaluating whether pre-trained foundational models are universal visual recognizers. To address this, we formally present the task of Open-domain Visual Entity recognitioN (Oven), where a model need to link an image onto a Wikipedia entity with respect to a text query. We construct Oven-Wiki‡by repurposing 14 existing datasets with all labels grounded onto one single label space: Wikipedia entities. Oven-Wiki challenges models to select among six million possible Wikipedia entities, making it a general visual recognition benchmark with the largest number of labels. Our study on state-ofthe-art pre-trained models reveals large headroom in generalizing to the massive-scale label space. We show that a PaLI-based auto-regressive visual recognition model performs surprisingly well, even on Wikipedia entities that have never been seen during fine-tuning. We also find existing pretrained models yield different strengths: while PaLI-based models obtain higher overall performance, CLIP-based models are better at recognizing tail entities.
Hexiang Hu, Yi Luan, Yang Chen 0065, Urvashi Khandelwal, Mandar Joshi, Kenton Lee, Kristina Toutanova, Ming-Wei Chang
ICCV3
2021 Model Selection for Cross-lingual Transfer
abstract
Transformers that are pre-trained on multilingual corpora, such as, mBERT and XLM-RoBERTa, have achieved impressive crosslingual transfer capabilities.In the zero-shot transfer setting, only English training data is used, and the fine-tuned model is evaluated on another target language.While this works surprisingly well, substantial variance has been observed in target language performance between different fine-tuning runs, and in the zero-shot setup, no target-language development data is available to select among multiple fine-tuned models.Prior work has relied on English dev data to select among models that are fine-tuned with different learning rates, number of steps and other hyperparameters, often resulting in suboptimal choices.In this paper, we show that it is possible to select consistently better models when small amounts of annotated data are available in auxiliary pivot languages.We propose a machine learning approach to model selection that uses the finetuned model's own internal representations to predict its cross-lingual capabilities.In extensive experiments we find that this method consistently selects better models than English validation data across twenty five languages (including eight low-resource languages), and often achieves results that are comparable to model selection using target language development data. 1
Yang Chen 0065, Alan Ritter
EMNLP (1)1
2020 An Empirical Study of Pre-trained Transformers for Arabic Information Extraction
abstract
Multilingual pre-trained Transformers, such as mBERT (Devlin et al., 2019) and XLM-RoBERTa (Conneau et al., 2020a), have been shown to enable the effective cross-lingual zero-shot transfer.However, their performance on Arabic information extraction (IE) tasks is not very well studied.In this paper, we pre-train a customized bilingual BERT, dubbed GigaBERT, that is designed specifically for Arabic NLP and English-to-Arabic zero-shot transfer learning.We study Giga-BERT's effectiveness on zero-short transfer across four IE tasks: named entity recognition, part-of-speech tagging, argument role labeling, and relation extraction.Our best model significantly outperforms mBERT, XLM-RoBERTa, and AraBERT (Antoun et al., 2020) in both the supervised and zero-shot transfer settings.We have made our pre-trained models publicly available at https://github.com
Wuwei Lan, Yang Chen 0065, Wei Xu 0004, Alan Ritter
EMNLP (1)2