Lei Fang 0004

dblp:66/5168-4 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0003-2510-1281ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 35% Query processing and optimization · 28% Data models and query languages · 22%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Artificial intelligence
4 papers
Representation and self-supervised learning · 38% Question answering and dialogue systems · 38% Language models and text generation · 13%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
analytical query
0.512021
AnaSearch: Extract, Retrieve and Visualize Structured Results from Unstructured Text for Analytical Queries · WSDM 2021
Information retrieval
question answering
0.512021
AnaSearch: Extract, Retrieve and Visualize Structured Results from Unstructured Text for Analytical Queries · WSDM 2021
Visualization and visual analytics › data storytelling
infographic generation
0.412020
Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements · IEEE Trans. Vis. Comput. Graph. 2020
Visualization and visual analytics › visualization generation › automated visualization generation
natural language to visualization
0.412020
Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements · IEEE Trans. Vis. Comput. Graph. 2020
Visualization and visual analytics
visualization authoring
0.412020
Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements · IEEE Trans. Vis. Comput. Graph. 2020
Visualization and visual analytics
visualization generation
0.412020
Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements · IEEE Trans. Vis. Comput. Graph. 2020
Machine learning › Representation and self-supervised learning › word representation
word representation learning
0.412019
Leveraging Web Semantic Knowledge in Word Representation Learning · AAAI 2019
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL
0.412019
Leveraging Adjective-Noun Phrasing Knowledge for Comparison Relation Prediction in Text-to-SQL · EMNLP/IJCNLP (1) 2019
Data integration and cleaning › data extraction
web data extraction
0.112021
AnaSearch: Extract, Retrieve and Visualize Structured Results from Unstructured Text for Analytical Queries · WSDM 2021
Natural language and speech › Language models and text generation
natural language understanding
0.112020
Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements · IEEE Trans. Vis. Comput. Graph. 2020
Natural language and speech › Information extraction and text analysis
relation extraction
0.112019
Leveraging Adjective-Noun Phrasing Knowledge for Comparison Relation Prediction in Text-to-SQL · EMNLP/IJCNLP (1) 2019
Information retrieval › query understanding
natural language query understanding
0.112019
A Split-and-Recombine Approach for Follow-up Query Analysis · EMNLP/IJCNLP (1) 2019

Methods — techniques the papers use, named apart from their topics

natural language processing · 0.9design space study · 0.9word embeddings · 0.8sequence labeling · 0.8semantic parsing · 0.8semantic knowledge · 0.8adjective-noun phrasing knowledge · 0.8natural language query · 0.5keyword matching · 0.5
YearPublicationVenuePosition
2021 AnaSearch: Extract, Retrieve and Visualize Structured Results from Unstructured Text for Analytical Queries
abstract
Modern search engines retrieve results mainly based on the keyword matching techniques, and thus fail to answer analytical queries like "apps with more than 1 billion monthly active users" or "population growth of the US from 2015 to 2019", which requires numerical reasoning or aggregating results from multiple web pages. Such analytical queries are very common in the data analysis area, the expected results would be structured tables or charts. In most cases, these structured results are not available or accessible, they scatter in various text sources. In this work, we build AnaSearch, a search system to support analytical queries, and return structured results that can be visualized in the form of tables or charts. We collect and build structured quantitative data from the unstructured text on the web automatically. With AnaSearch, data analysts could easily derive insights for decision making with keyword or natural language queries. Specifically, we build AnaSearch under the COVID-19 news data, which makes it easy to compare with manually collected structured data.
Tongliang Li, Lei Fang 0004, Jian-Guang Lou, Zhoujun Li 0001, Dongmei Zhang 0001
WSDM2
2020 Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements
abstract
Combining data content with visual embellishments, infographics can effectively deliver messages in an engaging and memorable manner. Various authoring tools have been proposed to facilitate the creation of infographics. However, creating a professional infographic with these authoring tools is still not an easy task, requiring much time and design expertise. Therefore, these tools are generally not attractive to casual users, who are either unwilling to take time to learn the tools or lacking in proper design expertise to create a professional infographic. In this paper, we explore an alternative approach: to automatically generate infographics from natural language statements. We first conducted a preliminary study to explore the design space of infographics. Based on the preliminary study, we built a proof-of-concept system that automatically converts statements about simple proportion-related statistics to a set of infographics with pre-designed styles. Finally, we demonstrated the usability and usefulness of the system through sample results, exhibits, and expert reviews.
Weiwei Cui 0001, Xiaoyu Zhang 0014, Yun Wang 0012, Bei Chen 0008, Lei Fang 0004, Jian-Guang Lou, Dongmei Zhang 0001
IEEE Trans. Vis. Comput. Graph.6
2019 Leveraging Web Semantic Knowledge in Word Representation Learning
Haoyan Liu 0001, Lei Fang 0004, Jian-Guang Lou, Zhoujun Li 0001
AAAI2
2019 A Split-and-Recombine Approach for Follow-up Query Analysis
abstract
Qian Liu, Bei Chen, Haoyan Liu, Jian-Guang Lou, Lei Fang, Bin Zhou, Dongmei Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Qian Liu 0033, Bei Chen 0008, Haoyan Liu 0001, Jian-Guang Lou, Lei Fang 0004, Dongmei Zhang 0001
EMNLP/IJCNLP (1)5
2019 Leveraging Adjective-Noun Phrasing Knowledge for Comparison Relation Prediction in Text-to-SQL
abstract
Haoyan Liu, Lei Fang, Qian Liu, Bei Chen, Jian-Guang Lou, Zhoujun Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Haoyan Liu 0001, Lei Fang 0004, Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Zhoujun Li 0001
EMNLP/IJCNLP (1)2
2015 Leveraging Large Data with Weak Supervision for Joint Feature and Opinion Word Extraction
Lei Fang 0004, Minlie Huang
J. Comput. Sci. Technol.1
2014 Ranking Sentiment Explanations for Review Summarization Using Dual Decomposition
abstract
For online reviews, sentiment explanations refer to the sentences that may suggest detailed reasons of sentiment, which are very important for applications in review mining like opinion summarization. In this paper, we address the problem of ranking sentiment explanations by formulating the process as two subproblems: sentence informativeness ranking and structural sentiment analysis. Tractable inference in joint prediction is performed through dual decomposition. Preliminary experiments on publicly available data demonstrate that our approach obtains promising performance.
Lei Fang 0004, Qiao Qian, Minlie Huang, Xiaoyan Zhu 0001
CIKM1
2013 Exploring weakly supervised latent sentiment explanations for aspect-level review analysis
abstract
In sentiment analysis, aspect-level review analysis has been an important task because it can catalogue, aggregate, or summarize various opinions according to a product's properties. In this paper, we explore a new concept for aspect-level review analysis, latent sentiment explanations, which are defined as a set of informative aspect-specific sentences whose polarities are consistent with that of the review. In other words, sentiment explanations best represent a review in terms of both aspect and polarity. We formulate the problem as a structure learning problem, and sentiment explanations are modeled with latent variables. Training samples are automatically identified through a set of pre-defined aspect signature terms (i.e., without manual annotation on samples), which we term the way weakly supervised.
Lei Fang 0004, Minlie Huang, Xiaoyan Zhu 0001
CIKM1
2010 An efficient location extraction algorithm by leveraging web contextual information
abstract
A typical location extraction approach consists of two steps, location name detection and location entity disambiguation. Promising results have been obtained in the last decade based on natural language processing technologies. However, there are still two challenges which requires further investigation: 1)How to leverage the prior and contextual evidence to improve the location extraction performance, and 2) How to utilize the interdependence information between the named entity recognition step and disambiguation step. In this paper, we propose an iterative detection-ranking framework to address these problems as well as a set of novel features to mine contextual information from web resources. Experimental results show that our solution outperforms the state-of-the-art approaches, including Metacarta GeoTagger and Yahoo Placemaker.
Teng Qin, Rong Xiao 0003, Lei Fang 0004, Xing Xie 0001, Lei Zhang 0001
GIS3