VLDB 2026 Research / reviewers in the wild / expert
Lei Fang 0004
dblp:66/5168-4
· DBLP profile ↗
9ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0003-2510-1281ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 35% Query processing and optimization · 28% Data models and query languages · 22% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% | |
| Artificial intelligence
4 papers |
Representation and self-supervised learning · 38% Question answering and dialogue systems · 38% Language models and text generation · 13% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
analytical query |
0.5 | 1 | 2021 | AnaSearch: Extract, Retrieve and Visualize Structured Results from Unstructured Text for Analytical Queries · WSDM 2021 |
Information retrieval
question answering |
0.5 | 1 | 2021 | AnaSearch: Extract, Retrieve and Visualize Structured Results from Unstructured Text for Analytical Queries · WSDM 2021 |
Visualization and visual analytics › data storytelling
infographic generation |
0.4 | 1 | 2020 | Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements · IEEE Trans. Vis. Comput. Graph. 2020 |
Visualization and visual analytics › visualization generation › automated visualization generation
natural language to visualization |
0.4 | 1 | 2020 | Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements · IEEE Trans. Vis. Comput. Graph. 2020 |
Visualization and visual analytics
visualization authoring |
0.4 | 1 | 2020 | Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements · IEEE Trans. Vis. Comput. Graph. 2020 |
Visualization and visual analytics
visualization generation |
0.4 | 1 | 2020 | Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements · IEEE Trans. Vis. Comput. Graph. 2020 |
Machine learning › Representation and self-supervised learning › word representation
word representation learning |
0.4 | 1 | 2019 | Leveraging Web Semantic Knowledge in Word Representation Learning · AAAI 2019 |
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL |
0.4 | 1 | 2019 | Leveraging Adjective-Noun Phrasing Knowledge for Comparison Relation Prediction in Text-to-SQL · EMNLP/IJCNLP (1) 2019 |
Data integration and cleaning › data extraction
web data extraction |
0.1 | 1 | 2021 | AnaSearch: Extract, Retrieve and Visualize Structured Results from Unstructured Text for Analytical Queries · WSDM 2021 |
Natural language and speech › Language models and text generation
natural language understanding |
0.1 | 1 | 2020 | Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements · IEEE Trans. Vis. Comput. Graph. 2020 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.1 | 1 | 2019 | Leveraging Adjective-Noun Phrasing Knowledge for Comparison Relation Prediction in Text-to-SQL · EMNLP/IJCNLP (1) 2019 |
Information retrieval › query understanding
natural language query understanding |
0.1 | 1 | 2019 | A Split-and-Recombine Approach for Follow-up Query Analysis · EMNLP/IJCNLP (1) 2019 |
Methods — techniques the papers use, named apart from their topics
natural language processing · 0.9design space study · 0.9word embeddings · 0.8sequence labeling · 0.8semantic parsing · 0.8semantic knowledge · 0.8adjective-noun phrasing knowledge · 0.8natural language query · 0.5keyword matching · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | AnaSearch: Extract, Retrieve and Visualize Structured Results from Unstructured Text for Analytical QueriesabstractModern search engines retrieve results mainly based on the keyword matching techniques, and thus fail to answer analytical queries like "apps with more than 1 billion monthly active users" or "population growth of the US from 2015 to 2019", which requires numerical reasoning or aggregating results from multiple web pages. Such analytical queries are very common in the data analysis area, the expected results would be structured tables or charts. In most cases, these structured results are not available or accessible, they scatter in various text sources. In this work, we build AnaSearch, a search system to support analytical queries, and return structured results that can be visualized in the form of tables or charts. We collect and build structured quantitative data from the unstructured text on the web automatically. With AnaSearch, data analysts could easily derive insights for decision making with keyword or natural language queries. Specifically, we build AnaSearch under the COVID-19 news data, which makes it easy to compare with manually collected structured data. Tongliang Li, Lei Fang 0004, Jian-Guang Lou, Zhoujun Li 0001, Dongmei Zhang 0001 |
WSDM | 2 |
| 2020 | Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language StatementsabstractCombining data content with visual embellishments, infographics can effectively deliver messages in an engaging and memorable manner. Various authoring tools have been proposed to facilitate the creation of infographics. However, creating a professional infographic with these authoring tools is still not an easy task, requiring much time and design expertise. Therefore, these tools are generally not attractive to casual users, who are either unwilling to take time to learn the tools or lacking in proper design expertise to create a professional infographic. In this paper, we explore an alternative approach: to automatically generate infographics from natural language statements. We first conducted a preliminary study to explore the design space of infographics. Based on the preliminary study, we built a proof-of-concept system that automatically converts statements about simple proportion-related statistics to a set of infographics with pre-designed styles. Finally, we demonstrated the usability and usefulness of the system through sample results, exhibits, and expert reviews. Weiwei Cui 0001, Xiaoyu Zhang 0014, Yun Wang 0012, Bei Chen 0008, Lei Fang 0004, Jian-Guang Lou, Dongmei Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2019 | Leveraging Web Semantic Knowledge in Word Representation Learning
Haoyan Liu 0001, Lei Fang 0004, Jian-Guang Lou, Zhoujun Li 0001 |
AAAI | 2 |
| 2019 | A Split-and-Recombine Approach for Follow-up Query AnalysisabstractQian Liu, Bei Chen, Haoyan Liu, Jian-Guang Lou, Lei Fang, Bin Zhou, Dongmei Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Qian Liu 0033, Bei Chen 0008, Haoyan Liu 0001, Jian-Guang Lou, Lei Fang 0004, Dongmei Zhang 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Leveraging Adjective-Noun Phrasing Knowledge for Comparison Relation Prediction in Text-to-SQLabstractHaoyan Liu, Lei Fang, Qian Liu, Bei Chen, Jian-Guang Lou, Zhoujun Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Haoyan Liu 0001, Lei Fang 0004, Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Zhoujun Li 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2015 | Leveraging Large Data with Weak Supervision for Joint Feature and Opinion Word Extraction
Lei Fang 0004, Minlie Huang |
J. Comput. Sci. Technol. | 1 |
| 2014 | Ranking Sentiment Explanations for Review Summarization Using Dual DecompositionabstractFor online reviews, sentiment explanations refer to the sentences that may suggest detailed reasons of sentiment, which are very important for applications in review mining like opinion summarization. In this paper, we address the problem of ranking sentiment explanations by formulating the process as two subproblems: sentence informativeness ranking and structural sentiment analysis. Tractable inference in joint prediction is performed through dual decomposition. Preliminary experiments on publicly available data demonstrate that our approach obtains promising performance. Lei Fang 0004, Qiao Qian, Minlie Huang, Xiaoyan Zhu 0001 |
CIKM | 1 |
| 2013 | Exploring weakly supervised latent sentiment explanations for aspect-level review analysisabstractIn sentiment analysis, aspect-level review analysis has been an important task because it can catalogue, aggregate, or summarize various opinions according to a product's properties. In this paper, we explore a new concept for aspect-level review analysis, latent sentiment explanations, which are defined as a set of informative aspect-specific sentences whose polarities are consistent with that of the review. In other words, sentiment explanations best represent a review in terms of both aspect and polarity. We formulate the problem as a structure learning problem, and sentiment explanations are modeled with latent variables. Training samples are automatically identified through a set of pre-defined aspect signature terms (i.e., without manual annotation on samples), which we term the way weakly supervised. Lei Fang 0004, Minlie Huang, Xiaoyan Zhu 0001 |
CIKM | 1 |
| 2010 | An efficient location extraction algorithm by leveraging web contextual informationabstractA typical location extraction approach consists of two steps, location name detection and location entity disambiguation. Promising results have been obtained in the last decade based on natural language processing technologies. However, there are still two challenges which requires further investigation: 1)How to leverage the prior and contextual evidence to improve the location extraction performance, and 2) How to utilize the interdependence information between the named entity recognition step and disambiguation step. In this paper, we propose an iterative detection-ranking framework to address these problems as well as a set of novel features to mine contextual information from web resources. Experimental results show that our solution outperforms the state-of-the-art approaches, including Metacarta GeoTagger and Yahoo Placemaker. Teng Qin, Rong Xiao 0003, Lei Fang 0004, Xing Xie 0001, Lei Zhang 0001 |
GIS | 3 |