EDBT 2026 Demo / reviewers in the wild / expert
Fangyu Liu 0001
dblp:84/11483-1
· DBLP profile ↗
32ranked-venue papers
11as first author
27since 2021 · last 2025
0000-0001-7038-3623ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 10 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text GenerationabstractThis paper introduces the counter-intuitive generalization results of overfitting pre-trained large language models (LLMs) on very small datasets. In the setting of open-ended text generation, it is well-documented that LLMs tend to generate repetitive and dull sequences, a phenomenon that is especially apparent when generating using greedy decoding. This issue persists even with state-of-the-art LLMs containing billions of parameters, trained via next-token prediction on large datasets. We find that by further fine-tuning these models to achieve a near-zero training loss on a small set of samples -- a process we refer to as hyperfitting -- the long-sequence generative capabilities are greatly enhanced.
Greedy decoding with these Hyperfitted models even outperform Top-P sampling over long-sequences, both in terms of diversity and human preferences. This phenomenon extends to LLMs of various sizes, different domains, and even autoregressive image generation. We further find this phenomena to be distinctly different from that of Grokking and double descent. Surprisingly, our experiments indicate that hyperfitted models rarely fall into repeating sequences they were trained on, and even explicitly blocking these sequences results in high-quality output. All hyperfitted models produce extremely low-entropy predictions, often allocating nearly all probability to a single token. Fredrik Carlsson, Fangyu Liu 0001, Daniel Ward, Murathan Kurfali, Joakim Nivre |
ICLR | 2 |
| 2024 | Faithful Chart Summarization with ChaTS-PiabstractChart-to-summary generation can help explore data, communicate insights, and help the visually impaired people.Multi-modal generative models have been used to produce fluent summaries, but they can suffer from factual and perceptual errors.In this work we present CHATS-CRITIC , a reference-free chart summarization metric for scoring faithfulness.CHATS-CRITIC is composed of an image-to-text model to recover the table from a chart, and a tabular entailment model applied to score the summary sentence by sentence.We find that CHATS-CRITIC evaluates the summary quality according to human ratings better than reference-based metrics, either learned or n-gram based, and can be further used to fix candidate summaries by removing not supported sentences.We then introduce CHATS-PI , a chart-to-summary pipeline that leverages CHATS-CRITIC during inference to fix and rank sampled candidates from any chart-summarization model.We evaluate CHATS-PI and CHATS-CRITIC using human raters, establishing state-of-the-art results on two popular chart-to-summary datasets.1 Syrine Krichene, Francesco Piccinno, Fangyu Liu 0001, Julian Martin Eisenschlos |
ACL (1) | 3 |
| 2024 | Reranking Overgenerated Responses for End-to-End Task-Oriented Dialogue SystemsabstractEnd-to-end task-oriented dialogue systems are prone to fall into the so-called ‘likelihood trap’, resulting in generated responses which are dull, repetitive, and often inconsistent with dialogue history. Comparing ranked lists of multiple generated responses against the ‘gold response’ reveals a wide diversity in quality, with many good responses placed lower in the ranked list. The main challenge addressed in this work is how to reach beyond greedily generated system responses, that is, how to obtain and select high-quality responses from the list of overgenerated responses at inference without the availability of the gold response. To this end, we propose a simple yet effective reranking method to select high-quality items from the lists of initially overgenerated responses. The idea is to use any sequence-level scoring function to divide the semantic space of responses into high-scoring versus low-scoring partitions. At training, the high-scoring partition comprises all generated responses whose similarity to the gold response is higher than the similarity of the greedy response to the gold response. At inference, the aim is to estimate the probability that each overgenerated response belongs to the high-scoring partition. We evaluate our proposed method on the standard MultiWOZ dataset, the BiTOD dataset, and with human evaluation. Songbo Hu, Ivan Vulic, Fangyu Liu 0001, Anna Korhonen |
LREC/COLING | 3 |
| 2024 | LUQ: Long-text Uncertainty Quantification for LLMsabstractLarge Language Models (LLMs) have demonstrated remarkable capability in a variety of NLP tasks.However, LLMs are also prone to generate nonfactual content.Uncertainty Quantification (UQ) is pivotal in enhancing our understanding of a model's confidence on its generation, thereby aiding in the mitigation of nonfactual outputs.Existing research on UQ predominantly targets short text generation, typically yielding brief, word-limited responses.However, real-world applications frequently necessitate much longer responses.Our study first highlights the limitations of current UQ methods in handling long text generation.We then introduce LUQ with its two variations: LUQ-ATOMIC and LUQ-PAIR, a series of novel sampling-based UQ approaches specifically designed for long text.Our findings reveal that LUQ outperforms existing baseline methods in correlating with the model's factuality scores (negative coefficient of -0.85 observed for Gemini Pro).To further improve the factuality of LLM responses, we propose LUQ-ENSEMBLE, a method that ensembles responses from multiple models and selects the response with the lowest uncertainty.The ensembling method greatly improves the response factuality upon the best standalone LLM. 1 * Now at Google DeepMind. Caiqi Zhang, Fangyu Liu 0001, Marco Basaldella, Nigel Collier |
EMNLP | 2 |
| 2023 | MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart DerenderingabstractFangyu Liu, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, Julian Eisenschlos. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Fangyu Liu 0001, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, Julian Martin Eisenschlos |
ACL (1) | 1 |
| 2023 | On Reality and the Limits of Language Data: Aligning LLMs with Human Norms
Nigel Collier, Fangyu Liu 0001, Ehsan Shareghi |
CogSci | 2 |
| 2023 | WinoDict: Probing language models for in-context word acquisitionabstractWe introduce a new in-context learning paradigm to measure Large Language Models' (LLMs) ability to learn novel words during inference.In particular, we rewrite Winogradstyle co-reference resolution problems by replacing the key concept word with a synthetic but plausible word that the model must understand to complete the task.Solving this task requires the model to make use of the dictionary definition of the new word given in the prompt.This benchmark addresses word acquisition, one important aspect of the diachronic degradation known to afflict LLMs.As LLMs are frozen in time at the moment they are trained, they are normally unable to reflect the way language changes over time.We show that the accuracy of LLMs compared to the original Winograd tasks decreases radically in our benchmark, thus identifying a limitation of current models and providing a benchmark to measure future improvements in LLMs ability to do in-context learning. Julian Martin Eisenschlos, Jeremy R. Cole, Fangyu Liu 0001, William W. Cohen |
EACL | 3 |
| 2023 | Probing Cross-Lingual Lexical Knowledge from Multilingual Sentence EncodersabstractIvan Vulić, Goran Glavaš, Fangyu Liu, Nigel Collier, Edoardo Maria Ponti, Anna Korhonen. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Ivan Vulic, Goran Glavas, Fangyu Liu 0001, Nigel Collier, Edoardo Maria Ponti, Anna Korhonen |
EACL | 3 |
| 2023 | Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingabstractVisually-situated language is ubiquitous---sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domain-specific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, and image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images. Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu 0001, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw 0004, Ming-Wei Chang, Kristina Toutanova |
ICML | 5 |
| 2023 | Visual Spatial ReasoningabstractAbstract Spatial relations are a basic part of human cognition. However, they are expressed in natural language in a variety of ways, and previous work has suggested that current vision-and-language models (VLMs) struggle to capture relational information. In this paper, we present Visual Spatial Reasoning (VSR), a dataset containing more than 10k natural text-image pairs with 66 types of spatial relations in English (e.g., under, in front of, facing). While using a seemingly simple annotation format, we show how the dataset includes challenging linguistic phenomena, such as varying reference frames. We demonstrate a large gap between human and model performance: The human ceiling is above 95%, while state-of-the-art models only achieve around 70%. We observe that VLMs’ by-relation performances have little correlation with the number of training examples and the tested models are in general incapable of recognising relations concerning the orientations of objects.1 Fangyu Liu 0001, Guy Emerson, Nigel Collier |
Trans. Assoc. Comput. Linguistics | 1 |
| 2023 | Compositional Zero-Shot Domain Transfer with Text-to-Text ModelsabstractAbstract Label scarcity is a bottleneck for improving task performance in specialized domains. We propose a novel compositional transfer learning framework (DoT51) for zero-shot domain transfer. Without access to in-domain labels, DoT5 jointly learns domain knowledge (from masked language modelling of unlabelled in-domain free text) and task knowledge (from task training on more readily available general-domain data) in a multi-task manner. To improve the transferability of task training, we design a strategy named NLGU: We simultaneously train natural language generation (NLG) for in-domain label-to-data generation, which enables data augmentation for self-finetuning and natural language understanding (NLU) for label prediction. We evaluate DoT5 on the biomedical domain and the resource-lean subdomain of radiology, focusing on natural language inference, text summarization, and embedding learning. DoT5 demonstrates the effectiveness of compositional transfer learning through multi-task learning. In particular, DoT5 outperforms the current state-of-the-art in zero-shot transfer by over 7 absolute points in accuracy on RadNLI. We validate DoT5 with ablations and a case study demonstrating its ability to solve challenging NLI examples requiring in-domain expertise. Fangyu Liu 0001, Qianchu Liu, Shruthi Bannur, Fernando Pérez-García, Naoto Usuyama, Sheng Zhang 0012, Tristan Naumann, Aditya V. Nori, Hoifung Poon, Javier Alvarez-Valle, Ozan Oktay, Stephanie L. Hyland |
Trans. Assoc. Comput. Linguistics | 1 |
| 2022 | Fine-Grained Controllable Text Generation Using Non-Residual PromptingabstractFredrik Carlsson, Joey Öhman, Fangyu Liu, Severine Verlinden, Joakim Nivre, Magnus Sahlgren. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Fredrik Carlsson, Joey Öhman, Fangyu Liu 0001, Severine Verlinden, Joakim Nivre, Magnus Sahlgren |
ACL (1) | 3 |
| 2022 | Improving Word Translation via Two-Stage Contrastive LearningabstractWord translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages.In this work, we propose a robust and effective two-stage contrastive learning framework for the BLI task.At Stage C1, we propose to refine standard cross-lingual linear maps between static word embeddings (WEs) via a contrastive learning objective; we also show how to integrate it into the self-learning procedure for even more refined cross-lingual maps.In Stage C2, we conduct BLI-oriented contrastive fine-tuning of mBERT, unlocking its word translation capability.We also show that static WEs induced from the 'C2-tuned' mBERT complement static WEs from Stage C1.Comprehensive experiments on standard BLI datasets for diverse languages and different experimental setups demonstrate substantial gains achieved by our framework.While the BLI method from Stage C1 already yields substantial gains over all state-of-the-art BLI methods in our comparison, even stronger improvements are met with the full two-stage framework: e.g., we report gains for 112/112 BLI setups, spanning 28 language pairs. Yaoyiran Li, Fangyu Liu 0001, Nigel Collier, Anna Korhonen, Ivan Vulic |
ACL (1) | 2 |
| 2022 | Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language ModelsabstractZaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su, Charlotte Collins, Nigel Collier. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Zaiqiao Meng, Fangyu Liu 0001, Ehsan Shareghi, Yixuan Su, Charlotte Collins, Nigel Collier |
ACL (1) | 2 |
| 2022 | Prix-LM: Pretraining for Multilingual Knowledge Base ConstructionabstractKnowledge bases (KBs) contain plenty of structured world and commonsense knowledge.As such, they often complement distributional text-based information and facilitate various downstream tasks.Since their manual construction is resource-and timeintensive, recent efforts have tried leveraging large pretrained language models (PLMs) to generate additional monolingual knowledge facts for KBs.However, such methods have not been attempted for building and enriching multilingual KBs.Besides wider application, such multilingual KBs can provide richer combined knowledge than monolingual (e.g., English) KBs.Knowledge expressed in different languages may be complementary and unequally distributed: this implies that the knowledge available in high-resource languages can be transferred to low-resource ones.To achieve this, it is crucial to represent multilingual knowledge in a shared/unified space.To this end, we propose a unified representation model, Prix-LM , for multilingual KB construction and completion.We leverage two types of knowledge, monolingual triples and cross-lingual links, extracted from existing multilingual KBs, and tune a multilingual language encoder XLM-R via a causal language modeling objective.Prix-LM integrates useful multilingual and KB-based factual knowledge into a single model.Experiments on standard entity-related tasks, such as link prediction in multiple languages, cross-lingual entity linking and bilingual lexicon induction, demonstrate its effectiveness, with gains reported over strong task-specialised baselines. Wenxuan Zhou 0002, Fangyu Liu 0001, Ivan Vulic, Nigel Collier, Muhao Chen 0001 |
ACL (1) | 2 |
| 2022 | Revisiting Parameter-Efficient Tuning: Are We Really There Yet?abstractParameter-Efficient Tuning (PETuning) methods have been deemed by many as the new paradigm for using pretrained language models (PLMs).By tuning just a fraction amount of parameters comparing to full model finetuning, PETuning methods claim to have achieved performance on par with or even better than finetuning.In this work, we take a step back and re-examine these PETuning methods by conducting the first comprehensive investigation into the training and evaluation of them.We found the problematic validation and testing practice in current studies, when accompanied by the instability nature of PETuning methods, has led to unreliable conclusions.When being compared under a truly fair evaluation protocol, PETuning cannot yield consistently competitive performance while finetuning remains to be the best-performing method in medium-and high-resource settings.We delve deeper into the cause of the instability and observed that the number of trainable parameters and training iterations are two main factors: reducing trainable parameters and prolonging training iterations may lead to higher stability in PETuning methods. 1 Guanzheng Chen, Fangyu Liu 0001, Zaiqiao Meng, Shangsong Liang |
EMNLP | 2 |
| 2022 | Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations
Fangyu Liu 0001, Yunlong Jiao, Jordan Massiah, Emine Yilmaz, Serhii Havrylov |
ICLR | 1 |
| 2022 | IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesabstractReliable evaluation benchmarks designed for replicability and comprehensiveness have driven progress in machine learning. Due to the lack of a multilingual benchmark, however, vision-and-language research has mostly focused on English language tasks. To fill this gap, we introduce the Image-Grounded Language Understanding Evaluation benchmark. IGLUE brings together{—}by both aggregating pre-existing datasets and creating new ones{—}visual question answering, cross-modal retrieval, grounded reasoning, and grounded entailment tasks across 20 diverse languages. Our benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups. Based on the evaluation of the available state-of-the-art models, we find that translate-test transfer is superior to zero-shot transfer and that few-shot learning is hard to harness for many tasks. Moreover, downstream performance is partially explained by the amount of available unlabelled textual data for pretraining, and only weakly by the typological distance of target{–}source languages. We hope to encourage future research efforts in this area by releasing the benchmark to the community. Emanuele Bugliarello, Fangyu Liu 0001, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, Ivan Vulic |
ICML | 2 |
| 2022 | Modality-Balanced Embedding for Video RetrievalabstractVideo search has become the main routine for users to discover videos relevant to a text query on large short-video sharing platforms. During training a query-video bi-encoder model using online search logs,\textit we identify a modality bias phenomenon that the video encoder almost entirely relies on text matching, neglecting other modalities of the videos such as vision, audio, \etc This modality imbalance results from a) modality gap: the relevance between a query and a video text is much easier to learn as the query is also a piece of text, with the same modality as the video text; b) data bias: most training samples can be solved solely by text matching. Here we share our practices to improve the first retrieval stage including our solution for the modality imbalance issue. We propose \modelname (short for Modality Balanced Video Retrieval) with two key components: manually generated modality-shuffled (MS) samples and a dynamic margin (DM) based on visual relevance. They can encourage the video encoder to pay balanced attentions to each modality. Through extensive experiments on a real world dataset, we show empirically that our method is both effective and efficient in solving modality bias problem. We have also deployed our ~\modelname~ in a large video platform and observed statistically significant boost over a highly optimized baseline in an A/B test and manual GSB evaluations. Bingqing Ke, Fangyu Liu 0001, Qiushi Xiao |
SIGIR | 4 |
| 2021 | Visual Pivoting for (Unsupervised) Entity AlignmentabstractThis work studies the use of visual semantic representations to align entities in heterogeneous knowledge graphs (KGs). Images are natural components of many existing KGs. By combining visual knowledge with other auxiliary information, we show that the proposed new approach, EVA, creates a holistic entity representation that provides strong signals for cross-graph entity alignment. Besides, previous entity alignment methods require human labelled seed alignment, restricting availability. EVA provides a completely unsupervised solution by leveraging the visual similarity of entities to create an initial seed dictionary (visual pivots). Experiments on benchmark data sets DBP15k and DWY15k show that EVA offers state-of-the-art performance on both monolingual and cross-lingual entity alignment tasks. Furthermore, we discover that images are particularly useful to align long-tail KG entities, which inherently lack the structural contexts necessary for capturing the correspondences. Code release: https://github.com/cambridgeltl/eva; project page: http://cogcomp.org/page/publication view/927. Fangyu Liu 0001, Muhao Chen 0001, Dan Roth 0001, Nigel Collier |
AAAI | 1 |
| 2021 | MirrorWiC: On Eliciting Word-in-Context Representations from Pretrained Language ModelsabstractRecent work indicated that pretrained language models (PLMs) such as BERT and RoBERTa can be transformed into effective sentence and word encoders even via simple self-supervised techniques.Inspired by this line of work, in this paper we propose a fully unsupervised approach to improving word-in-context (WiC) representations in PLMs, achieved via a simple and efficient WiC-targeted fine-tuning procedure: MIRROR-WIC.The proposed method leverages only raw texts sampled from Wikipedia, assuming no sense-annotated data, and learns contextaware word representations within a standard contrastive learning setup.We experiment with a series of standard and comprehensive WiC benchmarks across multiple languages.Our proposed fully unsupervised MIRROR-WIC models obtain substantial gains over offthe-shelf PLMs across all monolingual, multilingual and cross-lingual setups.Moreover, on some standard WiC benchmarks, MIRROR-WIC is even on-par with supervised models fine-tuned with in-task data and sense labels. Qianchu Liu, Fangyu Liu 0001, Nigel Collier, Anna Korhonen, Ivan Vulic |
CoNLL | 2 |
| 2021 | Visually Grounded Reasoning across Languages and CulturesabstractThe design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly overestimate how much this benchmark contributed to progress in computer vision, it is mostly derived from lexical databases and image queries in English, resulting in source material with a North American or Western European bias. Therefore, we devise a new protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. In particular, we let the selection of both concepts and images be entirely driven by native speakers, rather than scraping them automatically. Specifically, we focus on a typologically diverse set of languages, namely, Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish. On top of the concepts and images obtained through this new protocol, we create a multilingual dataset for Multicultural Reasoning over Vision and Language (MaRVL) by eliciting statements from native speaker annotators about pairs of images. The task consists of discriminating whether each grounded statement is true or false. We establish a series of baselines using state-of-the-art models and find that their cross-lingual transfer performance lags dramatically behind supervised performance in English. These results invite us to reassess the robustness and accuracy of current state-of-the-art models beyond a narrow domain, but also open up new exciting challenges for the development of truly multilingual and multicultural systems. Fangyu Liu 0001, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, Desmond Elliott |
EMNLP (1) | 1 |
| 2021 | Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence EncodersabstractPrevious work has indicated that pretrained Masked Language Models (MLMs) are not effective as universal lexical and sentence encoders off-the-shelf, i.e., without further taskspecific fine-tuning on NLI, sentence similarity, or paraphrasing tasks using annotated task data.In this work, we demonstrate that it is possible to turn MLMs into effective lexical and sentence encoders even without any additional data, relying simply on self-supervision.We propose an extremely simple, fast, and effective contrastive learning technique, termed Mirror-BERT, which converts MLMs (e.g., BERT and RoBERTa) into such encoders in 20-30 seconds with no access to additional external knowledge.Mirror-BERT relies on identical and slightly modified string pairs as positive (i.e., synonymous) fine-tuning examples, and aims to maximise their similarity during "identity fine-tuning".We report huge gains over off-the-shelf MLMs with Mirror-BERT both in lexical-level and in sentencelevel tasks, across different domains and different languages.Notably, in sentence similarity (STS) and question-answer entailment (QNLI) tasks, our self-supervised Mirror-BERT model even matches the performance of the Sentence-BERT models from prior work which rely on annotated task data.Finally, we delve deeper into the inner workings of MLMs, and suggest some evidence on why this simple Mirror-BERT fine-tuning approach can yield effective universal lexical and sentence encoders. Fangyu Liu 0001, Ivan Vulic, Anna Korhonen, Nigel Collier |
EMNLP (1) | 1 |
| 2021 | Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERTabstractInfusing factual knowledge into pretrained models is fundamental for many knowledgeintensive tasks.In this paper, we propose Mixture-of-Partitions (MoP), an infusion approach that can handle a very large knowledge graph (KG) by partitioning it into smaller subgraphs and infusing their specific knowledge into various BERT models using lightweight adapters.To leverage the overall factual knowledge for a target task, these sub-graph adapters are further fine-tuned along with the underlying BERT through a mixture layer.We evaluate our MoP with three biomedical BERTs (SciBERT, BioBERT, PubmedBERT) on six downstream tasks (inc.NLI, QA, Classification), and the results show that our MoP consistently enhances the underlying BERTs in task performance, and achieves new SOTA performances on five evaluated datasets.1 Zaiqiao Meng, Fangyu Liu 0001, Thomas Hikaru Clark, Ehsan Shareghi, Nigel Collier |
EMNLP (1) | 2 |
| 2021 | Contrastive Out-of-Distribution Detection for Pretrained TransformersabstractPretrained Transformers achieve remarkable performance when training and test data are from the same distribution.However, in realworld scenarios, the model often faces out-ofdistribution (OOD) instances that can cause severe semantic shift problems at inference time.Therefore, in practice, a reliable model should identify such instances, and then either reject them during inference or pass them over to models that handle another distribution.In this paper, we develop an unsupervised OOD detection method, in which only the indistribution (ID) data are used in training.We propose to fine-tune the Transformers with a contrastive loss, which improves the compactness of representations, such that OOD instances can be better differentiated from ID ones.These OOD instances can then be accurately detected using the Mahalanobis distance in the model's penultimate layer.We experiment with comprehensive settings and achieve near-perfect OOD detection performance, outperforming baselines drastically.We further investigate the rationales behind the improvement, finding that more compact representations through margin-based contrastive learning bring the improvement.We release our code to the community for future research 1 . Wenxuan Zhou 0002, Fangyu Liu 0001, Muhao Chen 0001 |
EMNLP (1) | 2 |
| 2021 | Self-Alignment Pretraining for Biomedical Entity RepresentationsabstractFangyu Liu, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, Nigel Collier. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Fangyu Liu 0001, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, Nigel Collier |
NAACL-HLT | 1 |
| 2021 | DeepOpht: Medical Report Generation for Retinal Images via Deep Models and Visual ExplanationabstractIn this work, we propose an AI-based method that intends to improve the conventional retinal disease treatment procedure and help ophthalmologists increase diagnosis efficiency and accuracy. The proposed method is composed of a deep neural networks-based (DNN-based) module, including a retinal disease identifier and clinical description generator, and a DNN visual explanation module. To train and validate the effectiveness of our DNN-based module, we propose a large-scale retinal disease image dataset. Also, as ground truth, we provide a retinal image dataset manually labeled by ophthalmologists to qualitatively show the proposed AI-based method is effective. With our experimental results, we show that the proposed method is quantitatively and qualitatively effective. Our method is capable of creating meaningful retinal image descriptions and visual explanations that are clinically relevant.https://github.com/Jhhuangkay/DeepOpht-Medical-Report-Generation-for-Retinal-Images-via-Deep-Models-and-Visual-Explanation. Jia-Hong Huang, Chao-Han Huck Yang, Fangyu Liu 0001, Meng Tian 0002, Yi-Chieh Liu, Ting-Wei Wu, I-Hung Lin, Hiromasa Morikawa, Hernghua Chang, Jesper Tegnér, Marcel Worring |
WACV | 3 |
| 2020 | HAL: Improved Text-Image Matching by Mitigating Visual Semantic HubsabstractThe hubness problem widely exists in high-dimensional embedding space and is a fundamental source of error for cross-modal matching tasks. In this work, we study the emergence of hubs in Visual Semantic Embeddings (VSE) with application to text-image matching. We analyze the pros and cons of two widely adopted optimization objectives for training VSE and propose a novel hubness-aware loss function (Hal) that addresses previous methods' defects. Unlike (Faghri et al. 2018) which simply takes the hardest sample within a mini-batch, Hal takes all samples into account, using both local and global statistics to scale up the weights of “hubs”. We experiment our method with various configurations of model architectures and datasets. The method exhibits exceptionally good robustness and brings consistent improvement on the task of text-image matching across all settings. Specifically, under the same model architectures as (Faghri et al. 2018) and (Lee et al. 2018), by switching only the learning objective, we report a maximum R@1 improvement of 7.4% on MS-COCO and 8.3% on Flickr30k.1 Fangyu Liu 0001, Rongtian Ye, Shuaipeng Li |
AAAI | 1 |
| 2020 | COMETA: A Corpus for Medical Entity Linking in the Social MediaabstractWhilst there has been growing progress in Entity Linking (EL) for general language, existing datasets fail to address the complex nature of health terminology in layman's language.Meanwhile, there is a growing need for applications that can understand the public's voice in the health domain.To address this we introduce a new corpus called COMETA, consisting of 20k English biomedical entity mentions from Reddit expert-annotated with links to SNOMED CT, a widely-used medical knowledge graph.Our corpus satisfies a combination of desirable properties, from scale and coverage to diversity and quality, that to the best of our knowledge has not been met by any of the existing resources in the field.Through benchmark experiments on 20 EL baselines from string-to neural-based models we shed light on the ability of these systems to perform complex inference on entities and concepts under 2 challenging evaluation scenarios.Our experimental results on COMETA illustrate that no golden bullet exists and even the best mainstream techniques still have a significant performance gap to fill, while the best solution relies on combining different views of data. Marco Basaldella, Fangyu Liu 0001, Ehsan Shareghi, Nigel Collier |
EMNLP (1) | 2 |
| 2020 | Upgrading the Newsroom: An Automated Image Selection System for News ArticlesabstractWe propose an automated image selection system to assist photo editors in selecting suitable images for news articles. The system fuses multiple textual sources extracted from news articles and accepts multilingual inputs. It is equipped with char-level word embeddings to help both modeling morphologically rich languages, e.g., German, and transferring knowledge across nearby languages. The text encoder adopts a hierarchical self-attention mechanism to attend more to both key words within a piece of text and informative components of a news article. We extensively experiment our system on a large-scale text-image database containing multimodal multilingual news articles collected from Swiss local news media websites. The system is compared with multiple baselines with ablation studies and is shown to beat existing text-image retrieval methods in a weakly supervised learning setting. Besides, we also offer insights on the advantage of using multiple textual sources and multilingual data. Fangyu Liu 0001, Rémi Lebret, Didier Orel, Philippe Sordet, Karl Aberer |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Joint Discriminative Dictionary and Classifier Learning for ALS Point Cloud ClassificationabstractTo efficiently recognize on-ground objects in airborne laser scanning (ALS) point clouds, we design a method that jointly learns a discriminative dictionary and a classifier. In the method, the point cloud is segmented into hierarchical point clusters, which are organized by a tree structure. Then, the feature of each point cluster is extracted. The feature of a leaf node is obtained by aggregating the features of all its parent nodes. The feature of the leaf node is called the hierarchical aggregation feature. The hierarchical aggregation features are encoded by sparse coding. We introduce a new label consistency constraint called “discriminative sparse-code error,” and combine it with the reconstruction error, the classification error, and L1-norm sparsity constraint to form a unified objective function. The objective function is efficiently solved by using the proposed label consistency feature sign method. We obtain an overcomplete discriminative dictionary and an optimal linear classifier. Experiments performed on different ALS point cloud scenes have shown that the hierarchical aggregation features combined with the learned classifier can significantly enhance the classification results, and also demonstrated the superior performance of our method over other techniques in point cloud classification. Zhenxin Zhang, Liqiang Zhang 0001, Yumin Tan, Liang Zhang 0023, Fangyu Liu 0001, Ruofei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | 3DCNN-DQN-RNN: A Deep Reinforcement Learning Framework for Semantic Parsing of Large-Scale 3D Point CloudsabstractSemantic parsing of large-scale 3D point clouds is an important research topic in computer vision and remote sensing fields. Most existing approaches utilize hand-crafted features for each modality independently and combine them in a heuristic manner. They often fail to consider the consistency and complementary information among features adequately, which makes them difficult to capture high-level semantic structures. The features learned by most of the current deep learning methods can obtain high-quality image classification results. However, these methods are hard to be applied to recognize 3D point clouds due to unorganized distribution and various point density of data. In this paper, we propose a 3DCNN-DQN-RNN method which fuses the 3D convolutional neural network (CNN), Deep Q-Network (DQN) and Residual recurrent neural network (RNN)for an efficient semantic parsing of large-scale 3D point clouds. In our method, an eye window under control of the 3D CNN and DQN can localize and segment the points of the object's class efficiently. The 3D CNN and Residual RNN further extract robust and discriminative features of the points in the eye window, and thus greatly enhance the parsing accuracy of large-scale point clouds. Our method provides an automatic process that maps the raw data to the classification results. It also integrates object localization, segmentation and classification into one framework. Experimental results demonstrate that the proposed method outperforms the state-of-the-art point cloud classification methods. Fangyu Liu 0001, Shuaipeng Li, Liqiang Zhang 0001, Chenghu Zhou, Rongtian Ye, Yuebin Wang, Jiwen Lu |
ICCV | 1 |