Fangtao Li

dblp:00/8159 · DBLP profile ↗
← Back
24ranked-venue papers
11as first author
6since 2021 · last 2023
0000-0002-8640-3235ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 9 first-authorGraphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2023 Overall-Distinctive GCN for Social Relation Recognition on Videos
Yibo Hu 0005, Chenyu Cao, Fangtao Li, Chenghao Yan, Jinsheng Qi, Bin Wu 0001
MMM (1)3
2023 ReGR: Relation-aware graph reasoning framework for video question answering
abstract
As one of the challenging cross-modal tasks, video question answering (VideoQA) aims to fully understand video content and answer relevant questions. The mainstream approach in current work involves extracting appearance and motion features to characterize videos separately, ignoring the interactions between them and with the question. Furthermore, some crucial semantic interaction details between visual objects are overlooked. In this paper, we propose a novel Relation-aware Graph Reasoning (ReGR) framework for video question answering, which first combines appearance–motion and location–semantic multiple interaction relations between visual objects. For the interaction between appearance and motion, we design the Appearance–Motion Block, which is question-guided to capture the interdependence between appearance and motion. For the interaction between location and semantics, we design the Location–Semantic Block, which utilizes the constructed Multi-Relation Graph Attention Network to capture the geometric position and semantic interaction between objects. Finally, the question-driven Multi-Visual Fusion captures more accurate multimodal representations. Extensive experiments on three benchmark datasets, TGIF-QA, MSVD-QA, and MSRVTT-QA, demonstrate the superiority of our proposed ReGR compared to the state-of-the-art methods.
Fangtao Li, Kaoru Ota, Mianxiong Dong, Bin Wu 0001
Inf. Process. Manag.2
2022 FHGN: Frame-Level Heterogeneous Graph Networks for Video Question Answering
abstract
Video Question Answering (VideoQA), aims to answer the given question based on video grounding, reasoning, and multimodal interacting. Most existing studies usually extract video-level visual and linguistic embeddings separately, then perform spatial-temporal or graph-based reasoning in a single modality. These methods lack the understanding of the interactions among different modalities and ignore the rele-vance of different frames, objects, and modalities to the question. This paper proposes Frame-level Heterogeneous Graph Networks (FHGN), to learn the semantic correlations among different modalities in each frame. Specifically, we construct a uniform frame-level heterogeneous graph to perform both inter- and intra-modality reasoning. Then a two-stage attention module is applied to measure the relevance between the question and different periods, space regions, and modalities. We conduct extensive experiments on a large-scale VideoQA dataset, and the experimental results show that our FHGN achieves the state-of-the-art performance.
Jinsheng Qi, Fangtao Li, Ting Bai 0004, Chenyu Cao, Yibo Hu 0005, Bin Wu 0001
ICME2
2021 Relation-aware Hierarchical Attention Framework for Video Question Answering
abstract
Video Question Answering (VideoQA) is a challenging video understanding task since it requires a deep understanding of both question and video. Previous studies mainly focus on extracting sophisticated visual and language embeddings, fusing them by delicate hand-crafted networks. However, the relevance of different frames, objects, and modalities to the question are varied along with the time, which is ignored in most of existing methods. Lacking understanding of the the dynamic relationships and interactions among objects brings a great challenge to VideoQA task. To address this problem, we propose a novel Relation-aware Hierarchical Attention (RHA) framework to learn both the static and dynamic relations of the objects in videos. In particular, videos and questions are embedded by pre-trained models firstly to obtain the visual and textual features. Then a graph-based relation encoder is utilized to extract the static relationship between visual objects. To capture the dynamic changes of multimodal objects in different video frames, we consider the temporal, spatial, and semantic relations, and fuse the multimodal features by hierarchical attention mechanism to predict the answer. We conduct extensive experiments on a large scale VideoQA dataset, and the experimental results demonstrate that our RHA outperforms the state-of-the-art methods.
Fangtao Li, Ting Bai 0004, Chenyu Cao, Chenghao Yan, Bin Wu 0001
ICMR1
2021 Social Relation Analysis from Videos via Multi-entity Reasoning
abstract
Videos contain rich semantic information. Analyzing social relations in video semantics can help machines interpret the behavior of human beings. However, most of the work related to social relationship recognition is based on still images, while video-based social relationship analysis tasks are less concerned. Here we propose a Multi-entity Relation Reasoning (MRR) framework that can be used for recognizing or predicting social relations in videos. To capture temporal features and contextual cues in videos, and use richer information to represent the person in the video, we track each person's appearance timeline and design a multi-entity representation method to build a social relationship knowledge graph. Then we use graph attention networks to gather information from the entity's neighborhood. Besides, situation information is helpful to identify relationships, we design a situation information extraction module to generate situation embedding from the video clip. Finally, a decoder is adopted to predict relationships between character entities. We evaluate the model on the MovieGraphs dataset and verify the effectiveness of the proposed framework.
Chenghao Yan, Fangtao Li, Chenyu Cao, Bin Wu 0001
ICMR3
2021 Frame Aggregation and Multi-modal Fusion Framework for Video-Based Person Recognition
Fangtao Li, Wenzhe Wang, Chenghao Yan, Bin Wu 0001
MMM (1)1
2020 Multi-Cue and Temporal Attention for Person Recognition in Videos
Wenzhe Wang, Bin Wu 0001, Fangtao Li
PRCV (2)3
2016 Unsupervised Head-Modifier Detection in Search Queries
abstract
Interpreting the user intent in search queries is a key task in query understanding. Query intent classification has been widely studied. In this article, we go one step further to understand the query from the view of head--modifier analysis. For example, given the query “popular iphone 5 smart cover,” instead of using coarse-grained semantic classes (e.g.,find electronic product), we interpret that “smart cover” is the head or the intent of the query and “iphone 5” is its modifier. Query head--modifier detection can help search engines to obtain particularly relevant content, which is also important for applications such as ads matching and query recommendation. We introduce an unsupervised semantic approach for query head--modifier detection. First, we mine a large number of instance level head--modifier pairs from search log. Then, we develop a conceptualization mechanism to generalize the instance level pairs to concept level. Finally, we derive weighted concept patterns that are concise, accurate, and have strong generalization power in head--modifier detection. The developed mechanism has been used in production for search relevance and ads matching. We use extensive experiment results to demonstrate the effectiveness of our approach.
Zhongyuan Wang 0006, Fang Wang 0019, Haixun Wang, Zhirui Hu, Jun Yan 0001, Fangtao Li, Ji-Rong Wen, Zhoujun Li 0001
ACM Trans. Knowl. Discov. Data6
2015 Identifying and constructing elemental parts of shafts based on conditional random fields model
Yamei Wen, Hui Zhang 0013, Fangtao Li, Jia-Guang Sun 0001
Comput. Aided Des.3
2015 TASC: Topic-Adaptive Sentiment Classification on Dynamic Tweets
abstract
Sentiment classification is a topic-sensitive task, i.e., a classifier trained from one topic will perform worse on another. This is especially a problem for the tweets sentiment analysis. Since the topics in Twitter are very diverse, it is impossible to train a universal classifier for all topics. Moreover, compared to product review, Twitter lacks data labeling and a rating mechanism to acquire sentiment labels. The extremely sparse text of tweets also brings down the performance of a sentiment classifier. In this paper, we propose a semi-supervised topic-adaptive sentiment classification (TASC) model, which starts with a classifier built on common features and mixed labeled data from various topics. It minimizes the hinge loss to adapt to unlabeled data and features including topic-related sentiment words, authors' sentiments and sentiment connections derived from“@” mentions of tweets, named as topic-adaptive features. Text and non-text features are extracted and naturally split into two views for co-training. The TASC learning algorithm updates topic-adaptive features based on the collaborative selection of unlabeled data, which in turn helps to select more reliable tweets to boost the performance. We also design the adapting model along a timeline (TASC-t) for dynamic tweets. An experiment on 6 topics from published tweet corpuses demonstrates that TASC outperforms other well-known supervised and ensemble classifiers. It also beats those semi-supervised learning methods without feature adaption. Meanwhile, TASC-t can also achieve impressive accuracy and F-score. Finally, with timeline visualization of “river” graph, people can intuitively grasp the ups and downs of sentiments' evolvement, and the intensity by color gradation.
Shenghua Liu, Xueqi Cheng 0001, Fuxin Li, Fangtao Li
IEEE Trans. Knowl. Data Eng.4
2014 SUIT: A Supervised User-Item Based Topic Model for Sentiment Analysis
abstract
Probabilistic topic models have been widely used for sentiment analysis. However, most of existing topic methods only model the sentiment text, but do not consider the user, who expresses the sentiment, and the item, which the sentiment is expressed on. Since different users may use different sentiment expressions for different items, we argue that it is better to incorporate the user and item information into the topic model for sentiment analysis. In this paper, we propose a new Supervised User-Item based Topic model, called SUIT model, for sentiment analysis. It can simultaneously utilize the textual topic and latent user-item factors. Our proposed method uses the tensor outer product of text topic proportion vector, user latent factor and item latent factor to model the sentiment label generalization. Extensive experiments are conducted on two datasets: review dataset and microblog dataset. The results demonstrate the advantages of our model. It shows significant improvement compared with supervised topic models and collaborative filtering methods.
Fangtao Li, Sheng Wang 0012, Shenghua Liu, Ming Zhang 0004
AAAI1
2014 Ranking Tweets by Labeled and Collaboratively Selected Pairs with Transitive Closure
abstract
Tweets ranking is important for information acquisition in Microblog. Due to the content sparsity and lackof labeled data, it is better to employ semi-supervisedlearning methods to utilize the unlabeled data. However,most of previous semi-supervised learning methods donot consider the pair conflict problem, which means thatthe new selected unlabeled data may conflict with the labeled and previously selected data. It will hurt the learning performance a lot, if the training data contains manyconflict pairs. In this paper, we propose a new collaborative semi-supervised SVM ranking model (CSR-TC)with consideration of the order conflict. The unlabeleddata is selected based on a dynamically maintained transitive closure graph to avoid pair conflict. We also investigate the two views of features, intrinsic and contentrelevant features, for the proposed model. Extensive experiments are conducted on TREC Microblogging corpus. The results demonstrate that our proposed methodachieves significant improvement, compared to severalstate-of-the-art models.
Shenghua Liu, Xueqi Cheng 0001, Fangtao Li
AAAI3
2014 A data-driven study of image feature extraction and fusion
Peng Cui 0001, Fangtao Li, Edward Y. Chang, Shiqiang Yang
Inf. Sci.3
2013 Deceptive Answer Prediction with User Preference Graph
Fangtao Li, Shuchang Zhou 0001, Xiance Si, Decheng Dai
ACL (1)1
2013 Adaptive co-training SVM for sentiment classification on tweets
abstract
Sentiment classification is an important problem in tweets mining. There lack labeled data and rating mechanism for generating them in Twitter service. And topics in Twitter are more diverse while sentiment classifiers always dedicate themselves to a specific domain or topic. Thus it is a challenge to make sentiment classification adaptive to diverse topics without sufficient labeled data. Therefore we formally propose an adaptive multiclass SVM model which transfers an initial common sentiment classifier to a topic-adaptive one. To tackle the tweet sparsity, non-text features are explored besides the conventional text features, which are intuitively split into two views. An iterative algorithm is proposed for solving this model by alternating among three steps: optimization, unlabeled data selection and adaptive feature expansion steps. The algorithm alternatively minimizes the margins of two independent objectives on different views to learn coefficient matrices, which are collaboratively used for unlabeled tweets selection from the topic that the algorithm is adapting to. And then topic-adaptive sentiment words are expended based on the above selection, in turn to help the first two steps find more confident and unlabeled tweets and boost the final performance. Comparing with the well-known supervised sentiment classifiers and semi-supervised approaches, our algorithm achieves promising increases in accuracy averagely on the 6 topics from public tweet corpus.
Shenghua Liu, Fuxin Li, Fangtao Li, Xueqi Cheng 0001, Huawei Shen
CIKM3
2012 Cross-Domain Co-Extraction of Sentiment and Topic Lexicons
Fangtao Li, Sinno Jialin Pan, Ou Jin, Qiang Yang 0001
ACL (1)1
2012 Entity Disambiguation with Freebase
abstract
Entity disambiguation with a knowledge base becomes increasingly popular in the NLP community. In this paper, we employ Freebase as the knowledge base, which contains significantly more entities than Wikipedia and others. While huge in size, Freebase lacks context for most entities, such as the descriptive text and hyperlinks in Wikipedia, which are useful for disambiguation. Instead, we leverage two features of Freebase, namely the naturally disambiguated mention phrases (aka aliases) and the rich taxonomy, to perform disambiguation in an iterative manner. Specifically, we explore both generative and discriminative models for each iteration. Experiments on 2, 430, 707 English sentences and 33, 743 Freebase entities show the effectiveness of the two features, where 90% accuracy can be reached without any labeled data. We also show that discriminative models with proposed split training strategy is robust against over fitting problem, and constantly outperforms the generative ones.
Zhicheng Zheng, Xiance Si, Fangtao Li, Edward Y. Chang, Xiaoyan Zhu 0001
Web Intelligence3
2011 Learning to Identify Review Spam
abstract
In the past few years, sentiment analysis and opinion mining becomes a popular and important task. These studies all assume that their opinion resources are real and trustful. However, they may encounter the faked opinion or opinion spam problem. In this paper, we study this issue in the context of our product review mining system. On product review site, people may write faked reviews, called review spam, to promote their products, or defame their competitors’ products. It is important to identify and filter out the review spam. Previous work only focuses on some heuristic rules, such as helpfulness voting, or rating deviation, which limits the performance of this task. In this paper, we exploit machine learning methods to identify review spam. Toward the end, we manually build a spam collection from our crawled reviews. We first analyze the effect of various features in spam identification. We also observe that the review spammer consistently writes spam. This provides us another view to identify review spam: we can identify if the author of the review is spammer. Based on this observation, we provide a twoview semi-supervised method, co-training, to exploit the large amount of unlabeled data. The experiment results show that our proposed method is effective. Our designed machine learning methods achieve significant improvements in comparison to the heuristic baselines.
Fangtao Li, Minlie Huang, Yi Yang 0038, Xiaoyan Zhu 0001
IJCAI1
2011 Incorporating Reviewer and Product Information for Review Rating Prediction
abstract
Traditional sentiment analysis mainly considers binary classifications of reviews, but in many real-world sentiment classification problems, non-binary review ratings are more useful. This is especially true when consumers wish to compare two products, both of which are not negative. Previous work has addressed this problem by extracting various features from the review text for learning a predictor. Since the same word may have different sentiment effects when used by different reviewers on different products, we argue that it is necessary to model such reviewer and product dependent effects in order to predict review ratings more accurately. In this paper, we propose a novel learning framework to incorporate reviewer and product information into the text based learner for rating prediction. The reviewer, product and text features are modeled as a three-dimension tensor. Tensor factorization techniques can then be employed to reduce the data sparsity problems. We perform extensive experiments to demonstrate the effectiveness of our model, which has a significant improvement compared to state of the art methods, especially for reviews with unpopular products and inactive reviewers.
Fangtao Li, Nathan Nan Liu, Qiang Yang 0001
IJCAI1
2010 Sentiment Analysis with Global Topics and Local Dependency
abstract
With the development of Web 2.0, sentiment analysis has now become a popular research problem to tackle. Recently, topic models have been introduced for the simultaneous analysis for topics and the sentiment in a document. These studies, which jointly model topic and sentiment, take the advantage of the relationship between topics and sentiment, and are shown to be superior to traditional sentiment analysis tools. However, most of them make the assumption that, given the parameters, the sentiments of the words in the document are all independent. In our observation, in contrast, sentiments are expressed in a coherent way. The local conjunctive words, such as “and” or “but”, are often indicative of sentiment transitions. In this paper, we propose a major departure from the previous approaches by making two linked contributions. First, we assume that the sentiments are related to the topic in the document, and put forward a joint sentiment and topic model, i.e. Sentiment-LDA. Second, we observe that sentiments are dependent on local context. Thus, we further extend the Sentiment-LDA model to Dependency-Sentiment-LDA model by relaxing the sentiment independent assumption in Sentiment-LDA. The sentiments of words are viewed as a Markov chain in Dependency-Sentiment-LDA. Through experiments, we show that exploiting the sentiment dependency is clearly advantageous, and that the Dependency-Sentiment-LDA is an effective approach for sentiment analysis.
Fangtao Li, Minlie Huang, Xiaoyan Zhu 0001
AAAI1
2010 Structure-Aware Review Mining and Summarization
Fangtao Li, Minlie Huang, Xiaoyan Zhu 0001, Yingju Xia, Shu Zhang 0004, Hao Yu 0005
COLING1
2010 Learning to Link Entities with Knowledge Base
Zhicheng Zheng, Fangtao Li, Minlie Huang, Xiaoyan Zhu 0001
HLT-NAACL2
2009 Answering Opinion Questions with Random Walks on Graphs
Fangtao Li, Minlie Huang, Xiaoyan Zhu 0001
ACL/IJCNLP1
2008 Classifying What-Type Questions by Head Noun Tagging
Fangtao Li, Xian Zhang 0006, Jinhui Yuan, Xiaoyan Zhu 0001
COLING1