Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Binxuan Huang

dblp:195/5963 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0003-3571-5738ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 54% Efficient and distributed learning · 20% Transfer learning and domain adaptation · 20%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 75% Knowledge graphs · 25%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
aspect-level sentiment classification
0.722019
Syntax-Aware Aspect Level Sentiment Classification with Graph Attention Networks · EMNLP/IJCNLP (1) 2019
Parameterized Convolutional Neural Networks for Aspect Level Sentiment Classification · EMNLP 2018
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.722019
Syntax-Aware Aspect Level Sentiment Classification with Graph Attention Networks · EMNLP/IJCNLP (1) 2019
Parameterized Convolutional Neural Networks for Aspect Level Sentiment Classification · EMNLP 2018
Machine learning › Efficient and distributed learning › data-efficient learning
label-efficient learning
0.712023
Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-Pages · Proc. VLDB Endow. 2023
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training
0.712023
Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-Pages · Proc. VLDB Endow. 2023
Data integration and cleaning › schema inference
column relation prediction
0.512021
TCN: Table Convolutional Network for Web Table Interpretation · WWW 2021
Data integration and cleaning › table understanding › table annotation
column type annotation
0.512021
TCN: Table Convolutional Network for Web Table Interpretation · WWW 2021
Knowledge graphs
knowledge graph construction
0.512021
TCN: Table Convolutional Network for Web Table Interpretation · WWW 2021
Data integration and cleaning › table understanding
web table understanding
0.512021
TCN: Table Convolutional Network for Web Table Interpretation · WWW 2021
Natural language and speech › Information extraction and text analysis
geolocation
0.412019
A Hierarchical Location Prediction Neural Network for Twitter User Geolocation · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation
large language model fine-tuning
0.212023
Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-Pages · Proc. VLDB Endow. 2023

Methods — techniques the papers use, named apart from their topics

uncertainty-aware training · 0.7self-training · 0.7generative model · 0.7pre-training · 0.5multi-task learning · 0.5convolutional neural network · 0.5attention mechanism · 0.5neural network · 0.4graph attention network · 0.4dependency parsing · 0.4parameterized gates · 0.3parameterized filters · 0.3parameterized convolutional neural network · 0.3
YearPublicationVenuePosition
2025 Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training
abstract
Yuchen Zhuang, Jingfeng Yang, Haoming Jiang, Xin Liu, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao, Qing Ping, Tianyi Liu, Binxuan Huang, Zheng Li, Zhengyang Wang, Pei Chen, Ruijie Wang, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Bing Yin, Chao Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yuchen Zhuang, Jingfeng Yang 0001, Haoming Jiang, Xin Liu 0039, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao 0001, Qing Ping, Binxuan Huang, Zheng Li 0018, Ruijie Wang 0004, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Chao Zhang 0014
NAACL (Long Papers)10
2025 DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
abstract
Large Language Models (LLMs) are increasingly employed in multi-turn conversational tasks, yet their pre-training data predominantly consists of continuous prose, creating a potential mismatch between required capabilities and training paradigms. We introduce a novel approach to address this discrepancy by synthesizing conversational data from existing text corpora. We present a pipeline that transforms a cluster of multiple related documents into an extended multi-turn, multi-topic information-seeking dialogue. Applying our pipeline to Wikipedia articles, we curate DocTalk, a multi-turn pre-training dialogue corpus consisting of over 730k long conversations. We hypothesize that exposure to such synthesized conversational structures during pre-training can enhance the fundamental multi-turn capabilities of LLMs, such as context memory and understanding. Empirically, we show that incorporating DocTalk during pre-training results in up to 40% gain in context memory and understanding, without compromising base performance. DocTalk is available at https://huggingface.co/datasets/AmazonScience/DocTalk.
Jing Yang Lee, Hamed Bonab, Nasser Zalmout, Ming Zeng 0001, Sanket Lokegaonkar, Colin Lockard, Binxuan Huang, Ritesh Sarkhel
SIGDIAL7
2023 Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-Pages
abstract
Information Extraction (IE) from semi-structured web-pages is a long studied problem. Training a model for this extraction task requires a large number of human-labeled samples. Prior works have proposed transferable models to improve the label-efficiency of this training process. Extraction performance of transferable models however, depends on the size of their fine-tuning corpus. This holds true for large language models (LLM) such as GPT-3 as well. Generalist models like LLMs need to be fine-tuned on in-domain, human-labeled samples for competitive performance on this extraction task. Constructing a large-scale fine-tuning corpus with human-labeled samples, however, requires significant effort. In this paper, we develop aLabel-Efficient Self-Training Algorithm(LEAST) to improve the label-efficiency of this fine-tuning process. Our contributions are two-fold.First, we develop a generative model that facilitates the construction of a large-scale fine-tuning corpus with minimal human-effort.Second, to ensure that the extraction performance does not suffer due to noisy training samples in our fine-tuning corpus, we develop an uncertainty-aware training strategy. Experiments on two publicly available datasets show that LEAST generalizes to multiple verticals and backbone models. Using LEAST, we can train models with less than ten human-labeled pages from each website, outperforming strong baselines while reducing the number of human-labeled training samples needed for comparable performance by up to 11x.
Ritesh Sarkhel, Binxuan Huang, Colin Lockard, Prashant Shiralkar
Proc. VLDB Endow.2
2021 TCN: Table Convolutional Network for Web Table Interpretation
abstract
Information extraction from semi-structured webpages provides valuable long-tailed facts for augmenting knowledge graph. Relational Web tables are a critical component containing additional entities and attributes of rich and diverse knowledge. However, extracting knowledge from relational tables is challenging because of sparse contextual information. Existing work linearize table cells and heavily rely on modifying deep language models such as BERT which only captures related cells information in the same table. In this work, we propose a novel relational table representation learning approach considering both the intra- and inter-table contextual information. On one hand, the proposed Table Convolutional Network model employs the attention mechanism to adaptively focus on the most informative intra-table cells of the same row or column; and, on the other hand, it aggregates inter-table contextual information from various types of implicit connections between cells across different tables. Specifically, we propose three novel aggregation modules for (i) cells of the same value, (ii) cells of the same schema position, and (iii) cells linked to the same page topic. We further devise a supervised multi-task training objective for jointly predicting column type and pairwise column relation, as well as a table cell recovery objective for pre-training. Experiments on real Web table datasets demonstrate our method can outperform competitive baselines by of F1 for column type prediction and by of F1 for pairwise column relation prediction.
Daheng Wang, Prashant Shiralkar, Colin Lockard, Binxuan Huang, Xin Dong 0001, Meng Jiang 0001
WWW4
2020 Entity Linking for Short Text Using Structured Knowledge Graph via Multi-Grained Text Matching
Binxuan Huang
INTERSPEECH1
2019 A large-scale empirical study of geotagging behavior on Twitter
abstract
Geotagging on social media has become an important proxy for understanding people's mobility and social events. Research that uses geotags to infer public opinions relies on several key assumptions about the behavior of geotagged and non-geotagged users. However, these assumptions have not been fully validated. Lack of understanding the geotagging behavior prohibits people further utilizing it. In this paper, we present an empirical study of geotagging behavior on Twitter based on more than 40 billion tweets collected from 20 million users. There are three main findings that may challenge these common assumptions. Firstly, different groups of users have different geotagging preferences. For example, less than 3% of users speaking in Korean are geotagged, while more than 40% of users speaking in Indonesian use geotags. Secondly, users who report their locations in profiles are more likely to use geotags, which may affects the generability of those location prediction systems on non-geotagged users. Thirdly, strong homophily effect exists in users' geotagging behavior, that users tend to connect to friends with similar geotagging preferences.
Binxuan Huang, Kathleen M. Carley
ASONAM1
2019 A Hierarchical Location Prediction Neural Network for Twitter User Geolocation
abstract
Binxuan Huang, Kathleen Carley. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Binxuan Huang, Kathleen M. Carley
EMNLP/IJCNLP (1)1
2019 Syntax-Aware Aspect Level Sentiment Classification with Graph Attention Networks
abstract
Binxuan Huang, Kathleen Carley. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Binxuan Huang, Kathleen M. Carley
EMNLP/IJCNLP (1)1
2018 Parameterized Convolutional Neural Networks for Aspect Level Sentiment Classification
abstract
We introduce a novel parameterized convolutional neural network for aspect level sentiment classification.Using parameterized filters and parameterized gates, we incorporate aspect information into convolutional neural networks (CNN).Experiments demonstrate that our parameterized filters and parameterized gates effectively capture the aspectspecific features, and our CNN-based models achieve excellent results on SemEval 2014 datasets.
Binxuan Huang, Kathleen M. Carley
EMNLP1
2017 The Role of Different Tie Strength in Disseminating Different Topics on a Microblog
abstract
The study of information flow typically does not distinguish the choices of tie strength on which the information flows. All receivers of the information are assumed to have the same potential to pass on the information. Modifying the SEIZ (susceptible, exposed, infected, skeptic) model, we discover that people choose to retweet strong or weak ties based on the topic. We made two modifications in the model. In the first modification (Model I), we assume that the contact rates of agents in different compartment and the probability of an agent transitioning from one compartment to another are different for strong ties and weak ties. In the second modification (Model II), we assume that only the probability of transitioning is different for strong ties and weak ties. We discover that people do not discriminate strong ties and weak ties when retweeting controversial topic, perhaps because this topic can both be personal and breaking news. On the other hand, people discriminate strong ties and weak ties when retweeting non-controversial topic. They prefer to retweet strong ties when the topic is donation, and kids, and weak ties when the topic is news on hurricane and music. Meanwhile, SEIZ model and its modifications are found to be inadequate to model tweets on event promotion.
Felicia Natali, Kathleen M. Carley, Feida Zhu 0001, Binxuan Huang
ASONAM4
2017 RATE: Overcoming Noise and Sparsity of Textual Features in Real-Time Location Estimation
abstract
Real-time location inference of social media users is the fundamental of some spatial applications such as localized search and event detection. While tweet text is the most commonly used feature in location estimation, most of the prior works suffer from either the noise or the sparsity of textual features. In this paper, we aim to tackle these two problems. We use topic modeling as a building block to characterize the geographic topic variation and lexical variation so that "one-hot" encoding vectors will no longer be directly used. We also incorporate other features which can be extracted through the Twitter streaming API to overcome the noise problem. Experimental results show that our RATE algorithm outperforms several benchmark methods, both in the precision of region classification and the mean distance error of latitude and longitude regression.
Yu Zhang 0044, Wei Wei 0019, Binxuan Huang, Kathleen M. Carley, Yan Zhang 0004
CIKM3