Anjie Fang

dblp:149/1354 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
7since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Data mining · 64% Web and social media mining · 22% Information retrieval · 14%
Artificial intelligence
3 papers
Information extraction and text analysis · 40% Language models and text generation · 26% Speech recognition and synthesis · 19%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Web and social media mining
social media analysis
1.042017
Examining Information on Social Media: Topic Modelling, Trend Prediction and Community Classification · SIGIR 2017
Using Word Embedding to Evaluate the Coherence of Topics from Twitter Data · SIGIR 2016
Examining the Coherence of the Top Ranked Tweet Topics · SIGIR 2016
Data mining › text mining
topic modeling
1.042017
Examining Information on Social Media: Topic Modelling, Trend Prediction and Community Classification · SIGIR 2017
Using Word Embedding to Evaluate the Coherence of Topics from Twitter Data · SIGIR 2016
Examining the Coherence of the Top Ranked Tweet Topics · SIGIR 2016
Information retrieval
query understanding
0.722022
Gazetteer Enhanced Named Entity Recognition for Code-Mixed Web Queries · SIGIR 2021
CycleKQR: Unsupervised Bidirectional Keyword-Question Rewriting · EMNLP 2022
Data mining › text mining › topic model › topic model evaluation
topic coherence
0.632017
Using Word Embedding to Evaluate the Coherence of Topics from Twitter Data · SIGIR 2016
Examining the Coherence of the Top Ranked Tweet Topics · SIGIR 2016
Examining Information on Social Media: Topic Modelling, Trend Prediction and Community Classification · SIGIR 2017
Machine learning › Learning paradigms › unsupervised learning
cycle-consistency learning
0.612022
CycleNER: An Unsupervised Training Approach for Named Entity Recognition · WWW 2022
Natural language and speech › Information extraction and text analysis
low-resource NLP
0.612022
CycleNER: An Unsupervised Training Approach for Named Entity Recognition · WWW 2022
Natural language and speech › Information extraction and text analysis
named entity recognition
0.612022
CycleNER: An Unsupervised Training Approach for Named Entity Recognition · WWW 2022
Natural language and speech › Information extraction and text analysis
query understanding
0.612022
CycleKQR: Unsupervised Bidirectional Keyword-Question Rewriting · EMNLP 2022
Natural language and speech › Language models and text generation › text generation
text rewriting
0.612022
CycleKQR: Unsupervised Bidirectional Keyword-Question Rewriting · EMNLP 2022
Data mining › text mining › information extraction
named entity recognition
0.512021
Gazetteer Enhanced Named Entity Recognition for Code-Mixed Web Queries · SIGIR 2021
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.412020
Using Phoneme Representations to Build Predictive Models Robust to ASR Errors · SIGIR 2020
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › robust speech recognition
noise-robust speech recognition
0.412020
Using Phoneme Representations to Build Predictive Models Robust to ASR Errors · SIGIR 2020
Natural language and speech › Language models and text generation
text representation
0.412020
Using Phoneme Representations to Build Predictive Models Robust to ASR Errors · SIGIR 2020
Data mining › time series analysis › time series forecasting
trend prediction
0.312017
Examining Information on Social Media: Topic Modelling, Trend Prediction and Community Classification · SIGIR 2017
Data mining › pattern mining
association rule mining
0.212014
Accurate Household Occupant Behavior Modeling Based on Data Mining Techniques · AAAI 2014
Data mining
clustering
0.212014
Accurate Household Occupant Behavior Modeling Based on Data Mining Techniques · AAAI 2014
Data mining
pattern mining
0.212014
Accurate Household Occupant Behavior Modeling Based on Data Mining Techniques · AAAI 2014
Natural language and speech › Language models and text generation
pre-trained language model
0.212022
CycleNER: An Unsupervised Training Approach for Named Entity Recognition · WWW 2022
Natural language and speech › Question answering and dialogue systems
intent detection
0.112020
Using Phoneme Representations to Build Predictive Models Robust to ASR Errors · SIGIR 2020
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.112020
Using Phoneme Representations to Build Predictive Models Robust to ASR Errors · SIGIR 2020
Energy systems and smart grids › building energy
building energy simulation
0.112014
Accurate Household Occupant Behavior Modeling Based on Data Mining Techniques · AAAI 2014

Methods — techniques the papers use, named apart from their topics

unsupervised learning · 1.1non-parallel data · 1.1cycle consistency · 1.1cycle-consistency training · 0.6multilingual transformer · 0.5mixture of experts · 0.5gated architecture · 0.5phoneme representation · 0.4neural network · 0.4topic modeling · 0.3classification · 0.3tweet pooling · 0.2LDA · 0.2nearest neighbor · 0.2markov chain · 0.2association rule learning · 0.2
YearPublicationVenuePosition
2023 Composing Spoken Hints for Follow-on Question Suggestion in Voice Assistants
Pedro Faustini, Besnik Fetahu, Giuseppe Castellucci, Anjie Fang, Oleg Rokhlenko, Shervin Malmasi
INTERSPEECH4
2022 MultiCoNER: A Large-scale Multilingual Dataset for Complex Named Entity Recognition
abstract
We present AnonData, a large multilingual dataset for Named Entity Recognition that covers 3 domains (Wiki sentences, questions, and search queries) across 11 languages, as well as multilingual and code-mixing subsets. This dataset is designed to represent contemporary challenges in NER, including low-context scenarios (short and uncased text), syntactically complex entities like movie titles, and long-tail entity distributions. The 26M token dataset is compiled from public resources using techniques such as heuristic-based sentence sampling, template extraction and slotting, and machine translation. We tested the performance of two NER models on our dataset: a baseline XLM-RoBERTa model, and a state-of-the-art NER GEMNET model that leverages gazetteers. The baseline achieves moderate performance (macro-F1=54%). GEMNET, which uses gazetteers, improvement significantly (average improvement of macro-F1=+30%) and demonstrates the difficulty of our dataset. AnonData poses challenges even for large pre-trained language models, and we believe that it can help further research in building robust NER systems.
Shervin Malmasi, Anjie Fang, Besnik Fetahu, Sudipta Kar, Oleg Rokhlenko
COLING2
2022 CycleKQR: Unsupervised Bidirectional Keyword-Question Rewriting
abstract
Users expect their queries to be answered by search systems, regardless of the query's surface form, which include keyword queries and natural questions.Natural Language Understanding (NLU) components of Search and QA systems may fail to correctly interpret semantically equivalent inputs if this deviates from how the system was trained, leading to suboptimal understanding capabilities.We propose the keyword-question rewriting task to improve query understanding capabilities of NLU systems for all surface forms.To achieve this, we present CycleKQR, an unsupervised approach, enabling effective rewriting between keyword and question queries using non-parallel data.Empirically we show the impact on QA performance of unfamiliar query forms for open domain and Knowledge Base QA systems (trained on either keywords or natural language questions).We demonstrate how CycleKQR significantly improves QA performance by rewriting queries into the appropriate form, while at the same time retaining the original semantic meaning of input queries, allowing CycleKQR to improve performance by up to 3% over supervised baselines.Finally, we release a dataset of 66k keyword-question pairs. 1 1 https://github.com/amzn/kqrHow much is iPhone 13?What is the price of iPhone 13? How much does iPhone 13 cost?iPhone 13 cost, cost of iPhone 13 iPhone 13 price, price of iPhone 13 iPhone 13 offer, iPhone 13dealQuestions
Andrea Iovine, Anjie Fang, Besnik Fetahu, Oleg Rokhlenko, Shervin Malmasi
EMNLP2
2022 Dynamic Gazetteer Integration in Multilingual Models for Cross-Lingual and Cross-Domain Named Entity Recognition
abstract
Besnik Fetahu, Anjie Fang, Oleg Rokhlenko, Shervin Malmasi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Besnik Fetahu, Anjie Fang, Oleg Rokhlenko, Shervin Malmasi
NAACL-HLT2
2022 CycleNER: An Unsupervised Training Approach for Named Entity Recognition
abstract
Named Entity Recognition (NER) is a crucial natural language understanding task for many down-stream tasks such as question answering and retrieval. Despite significant progress in developing NER models for multiple languages and domains, scaling to emerging and/or low-resource domains still remains challenging, due to the costly nature of acquiring training data. We propose CycleNER, an unsupervised approach based on cycle-consistency training that uses two functions: (i) sentence-to-entity – S2E and (ii) entity-to-sentence – E2S, to carry out the NER task. CycleNER does not require annotations but a set of sentences with no entity labels and another independent set of entity examples. Through cycle-consistency training, the output from one function is used as input for the other (e.g. S2E → E2S) to align the representation spaces of both functions and therefore enable unsupervised training. Evaluation on several domains comparing CycleNER against supervised and unsupervised competitors shows that CycleNER achieves highly competitive performance with only a few thousand input sentences. We demonstrate competitive performance against supervised models, achieving 73% of supervised performance without any annotations on CoNLL03, while significantly outperforming unsupervised approaches.
Andrea Iovine, Anjie Fang, Besnik Fetahu, Oleg Rokhlenko, Shervin Malmasi
WWW2
2021 GEMNET: Effective Gated Gazetteer Representations for Recognizing Complex Entities in Low-context Input
abstract
Tao Meng, Anjie Fang, Oleg Rokhlenko, Shervin Malmasi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Anjie Fang, Oleg Rokhlenko, Shervin Malmasi
NAACL-HLT2
2021 Gazetteer Enhanced Named Entity Recognition for Code-Mixed Web Queries
abstract
Named entity recognition (NER) for Web queries is very challenging. Queries often do not consist of well-formed sentences, and contain very little context, with highly ambiguous queried entities. Code-mixed queries, with entities in a different language than the rest of the query, pose a particular challenge in domains like e-commerce (e.g. queries containing movie or product names). This work tackles NER for code-mixed queries, where entities and non-entity query terms co-exist simultaneously in different languages. Our contributions are twofold. First, to address the lack of code-mixed NER data we create EMBER, a large-scale dataset in six languages with four different scripts. Based on Bing query data, we include numerous language combinations that showcase real-world search scenarios. Secondly, we propose a novel gated architecture that enhances existing multi-lingual Transformers with a Mixture-of-Experts model to dynamically infuse multi-lingual gazetteers, allowing it to simultaneously differentiate and handle entities and non-entity query terms in multiple languages. Experimental evaluation on code-mixed queries in several languages shows that our approach efficiently utilizes gazetteers to recognize entities in code-mixed queries with an F1=68%, an absolute improvement of +31% over a non-gazetteer baseline.
Besnik Fetahu, Anjie Fang, Oleg Rokhlenko, Shervin Malmasi
SIGIR2
2020 Using Phoneme Representations to Build Predictive Models Robust to ASR Errors
abstract
Even though Automatic Speech Recognition (ASR) systems significantly improved over the last decade, they still introduce a lot of errors when they transcribe voice to text. One of the most common reasons for these errors is phonetic confusion between similar-sounding expressions. As a result, ASR transcriptions often contain "quasi-oronyms", i.e., words or phrases that sound similar to the source ones, but that have completely different semantics (e.g., "win" instead of "when" or "accessible on defecting" instead of "accessible and affecting"). These errors significantly affect the performance of downstream Natural Language Understanding (NLU) models (e.g., intent classification, slot filling, etc.) and impair user experience. To make NLU models more robust to such errors, we propose novel phonetic-aware text representations. Specifically, we represent ASR transcriptions at the phoneme level, aiming to capture pronunciation similarities, which are typically neglected in word-level representations (e.g., word embeddings). To train and evaluate our phoneme representations, we generate noisy ASR transcriptions of four existing datasets - Stanford Sentiment Treebank, SQuAD, TREC Question Classification and Subjectivity Analysis - and show that common neural network architectures exploiting the proposed phoneme representations can effectively handle noisy transcriptions and significantly outperform state-of-the-art baselines. Finally, we confirm these results by testing our models on real utterances spoken to the Alexa virtual assistant.
Anjie Fang, Simone Filice, Nut Limsopatham, Oleg Rokhlenko
SIGIR1
2019 Evaluating Similarity Metrics for Latent Twitter Topics
Xi Wang 0012, Anjie Fang, Iadh Ounis, Craig Macdonald
ECIR (1)2
2018 An Effective Approach for Modelling Time Features for Classifying Bursty Topics on Twitter
abstract
Several previous approaches attempted to predict bursty topics on Twitter. Such approaches have usually reported that the time information (e.g. the topic popularity over time) of hashtag topics contribute the most to the prediction of bursty topics. In this paper, we propose a novel approach to use time features to predict bursty topics on Twitter. We model the popularity of topics as density curves described by the density function of a beta distribution with different parameters. We then propose various approaches to predict/classify the bursty topics by estimating the parameters of topics, using estimators such as Gradient Decent or Likelihood Maximization. In our experiments, we show that the estimated parameters of topics have a positive effect on classifying bursty topics. In particular, our estimators when combined together improve the bursty topic classification by 6.9 in terms of micro F1 compared to a baseline classifier using hashtag content features.
Anjie Fang, Iadh Ounis, Craig Macdonald, Philip Habel, Xiaoyu Xiong, Hai-Tao Yu 0003
CIKM1
2018 On Refining Twitter Lists as Ground Truth Data for Multi-community User Classification
Ting Su 0003, Anjie Fang, Richard McCreadie, Craig Macdonald, Iadh Ounis
ECIR2
2018 On the Reproducibility and Generalisation of the Linear Transformation of Word Embeddings
Xiao Yang 0004, Iadh Ounis, Richard McCreadie, Craig Macdonald, Anjie Fang
ECIR5
2017 Exploring Time-Sensitive Variational Bayesian Inference LDA for Social Media Data
Anjie Fang, Craig Macdonald, Iadh Ounis, Philip Habel, Xiao Yang 0004
ECIR1
2017 Examining Information on Social Media: Topic Modelling, Trend Prediction and Community Classification
abstract
In the past decade, the use of social media networks (e.g. Twitter) increased dramatically becoming the main channels for the mass public to express their opinions, ideas and preferences, especially during an election or a referendum. Both researchers and the public are interested in understanding what topics are discussed during a real social event, what are the trends of the discussed topics and what is the future topical trend. Indeed, modelling such topics as well as trends offer opportunities for social scientists to continue a long-standing research, i.e. examine the information exchange between people in different communities. We argue that computing science approaches can adequately assist social scientists to extract topics from social media data, to predict their topical trends, or to classify a social media user (e.g. a Twitter user) into a community. However, while topic modelling approaches and classification techniques have been widely used, challenges still exist, such as 1) existing topic modelling approaches can generate topics lacking of coherence for social media data; 2) it is not easy to evaluate the coherence of topics; 3) it can be challenging to generate a large training dataset for developing a social media user classifier. Hence, we identify four tasks to solve these problems and assist social scientists. Initially, we aim to propose topic coherence metrics that effectively evaluate the coherence of topics generated by topic modelling approaches. Such metrics are required to align with human judgements. Since topic modelling approaches cannot always generate useful topics, it is necessary to present users with the most coherent topics using the coherence metrics. Moreover, an effective coherence metric helps us evaluate the performance of our proposed topic modelling approaches. The second task is to propose a topic modelling approach that generates more coherent topics for social media data. We argue that the use of time dimension of social media posts helps a topic modelling approach to distinguish the word usage differences over time, and thus allows to generate topics with higher coherence as well as their trends. A more coherent topic with its trend allows social scientists to quickly identify the topic subject and to focus on analysing the connections between the extracted topics with the social events, e.g., an election. Third, we aim to model and predict the topical trend. Given the timestamps of social media posts within topics, a topical trend can be modelled as a continuous distribution over time. Therefore, we argue that the future trends of topics can be predicted by estimating the density function of their continuous time distribution. By examining the future topical trend, social scientists can ensure the timeliness of their focused events. Politicians and policymakers can keep abreast of the topics that remain salient over time. Finally, we aim to offer a general method that can quickly obtain a large training dataset for constructing a social media user classifier. A social media post contains hashtags and entities. These hashtags (e.g. "#YesScot" in Scottish Independence Referendum) and entities (e.g., job title or parties' name) can reflect the community affiliation of a social media user. We argue that a large and reliable training dataset can be obtained by distinguishing the usage of these hashtags and entities. Using the obtained training dataset, a social media user community classifier can be quickly achieved, and then used as input to assist in examining the different topics discussed in communities. In conclusion, we have identified four aspects for assisting social scientists to better understand the discussed topics on social media networks. We believe that the proposed tools and approaches can help to examine the exchanges of topics among communities on social media networks.
Anjie Fang
SIGIR1
2016 Topics in Tweets: A User Study of Topic Coherence Metrics for Twitter Data
Anjie Fang, Craig Macdonald, Iadh Ounis, Philip Habel
ECIR1
2016 Examining the Coherence of the Top Ranked Tweet Topics
abstract
Topic modelling approaches help scholars to examine the topics discussed in a corpus. Due to the popularity of Twitter, two distinct methods have been proposed to accommodate the brevity of tweets: the tweet pooling method and Twitter LDA. Both of these methods demonstrate a higher performance in producing more interpretable topics than the standard Latent Dirichlet Allocation (LDA) when applied on tweets. However, while various metrics have been proposed to estimate the coherence of the generated topics from tweets, the coherence of the top ranked topics, those that are most likely to be examined by users, has not been investigated. In addition, the effect of the number of generated topics K on the topic coherence scores has not been studied. In this paper, we conduct large-scale experiments using three topic modelling approaches over two Twitter datasets, and apply a state-of-the-art coherence metric to study the coherence of the top ranked topics and how K affects such coherence. Inspired by ranking metrics such as precision at n, we use coherence at n to assess the coherence of a topic model. To verify our results, we conduct a pairwise user study to obtain human preferences over topics. Our findings are threefold: we find evidence that Twitter LDA outperforms both LDA and the tweet pooling method because the top ranked topics it generates have more coherence; we demonstrate that a larger number of topics (K) helps to generate topics with more coherence; and finally, we show that coherence at n is more effective when evaluating the coherence of a topic model than the average coherence score.
Anjie Fang, Craig Macdonald, Iadh Ounis, Philip Habel
SIGIR1
2016 Using Word Embedding to Evaluate the Coherence of Topics from Twitter Data
abstract
Scholars often seek to understand topics discussed on Twitter using topic modelling approaches. Several coherence metrics have been proposed for evaluating the coherence of the topics generated by these approaches, including the pre-calculated Pointwise Mutual Information (PMI) of word pairs and the Latent Semantic Analysis (LSA) word representation vectors. As Twitter data contains abbreviations and a number of peculiarities (e.g. hashtags), it can be challenging to train effective PMI data or LSA word representation. Recently, Word Embedding (WE) has emerged as a particularly effective approach for capturing the similarity among words. Hence, in this paper, we propose new Word Embedding-based topic coherence metrics. To determine the usefulness of these new metrics, we compare them with the previous PMI/LSA-based metrics. We also conduct a large-scale crowdsourced user study to determine whether the new Word Embedding-based metrics better align with human preferences. Using two Twitter datasets, our results show that the WE-based metrics can capture the coherence of topics in tweets more robustly and efficiently than the PMI/LSA-based ones.
Anjie Fang, Craig Macdonald, Iadh Ounis, Philip Habel
SIGIR1
2015 Topic-centric Classification of Twitter User's Political Orientation
abstract
In the recent Scottish Independence Referendum (hereafter, IndyRef), Twitter offered a broad platform for people to express their opinions, with millions of IndyRef tweets posted over the campaign period. In this paper, we aim to classify people's voting intentions by the content of their tweets---their short messages communicated on Twitter. By observing tweets related to the IndyRef, we find that people not only discussed the vote, but raised topics related to an independent Scotland including oil reserves, currency, nuclear weapons, and national debt. We show that the views communicated on these topics can inform us of the individuals' voting intentions ("Yes"--in favour of Independence vs. "No"--Opposed). In particular, we argue that an accurate classifier can be designed by leveraging the differences in the features' usage across different topics related to voting intentions. We demonstrate improvements upon a Naive Bayesian classifier using the topics enrichment method. Our new classifier identifies the closest topic for each unseen tweet, based on those topics identified in the training data. Our experiments show that our Topics-Based Naive Bayesian classifier improves accuracy by 7.8% over the classical Naive Bayesian baseline.
Anjie Fang, Iadh Ounis, Philip Habel, Craig Macdonald, Nut Limsopatham
SIGIR1
2014 Accurate Household Occupant Behavior Modeling Based on Data Mining Techniques
abstract
An important requirement of household energy simulation models is their accuracy in estimating energy demand and its fluctuations. Occupant behavior has a major impact upon energy demand. However, Markov chains, the traditional approach to model occupant behavior, (1) has limitations in accurately capturing the coordinated behavior of occupants and (2) is prone to over-fitting. To address these issues, we propose a novel approach that relies on a combination of data mining techniques. The core idea of our model is to determine the behavior of occupants based on nearest neighbor comparison over a database of sample data. Importantly, the model takes into account features related to the coordination of occupants' activities. We use a customized distance function suited for mixed categorical and numerical data. Further, association rule learning allows us to capture the coordination between occupants. Using real data from four households in Japan we are able to show that our model outperforms the traditional Markov chain model with respect to occupant coordination and generalization of behavior patterns.
Márcia Baptista, Anjie Fang, Helmut Prendinger, Rui Prada, Yohei Yamaguchi
AAAI2