Guanghui Qin

dblp:228/5587 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-3009-8614ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Probabilistic and Bayesian machine learning · 23% Knowledge representation and reasoning · 18% Question answering and dialogue systems · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
query log analysis
0.912025
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for Deep Research · SIGIR 2025
Machine learning › Efficient and distributed learning › inference efficiency
context compression
0.812024
Dodo: Dynamic Contextual Compression for Decoder-only LMs · ACL (1) 2024
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context window extension
0.712023
Nugget: Neural Agglomerative Embeddings of Text · ICML 2023
Machine learning › Representation and self-supervised learning › text embedding
text representation learning
0.712023
Nugget: Neural Agglomerative Embeddings of Text · ICML 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic programming
datalog
0.412020
Neural Datalog Through Time: Informed Temporal Modeling via Logical Specification · ICML 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic programming
0.412020
Neural Datalog Through Time: Informed Temporal Modeling via Logical Specification · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process › temporal point process
neural hawkes process
0.412019
Imputing Missing Events in Continuous-Time Event Streams · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › sequential monte carlo
particle smoother
0.412019
Imputing Missing Events in Continuous-Time Event Streams · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
sequential monte carlo
0.412019
Imputing Missing Events in Continuous-Time Event Streams · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
temporal point process
0.412019
Imputing Missing Events in Continuous-Time Event Streams · ICML 2019
Natural language and speech › Information extraction and text analysis › data annotation
semantic annotation
0.312018
Learning Latent Semantic Annotations for Grounding Natural Language to Structured Data · EMNLP 2018
Natural language and speech › Question answering and dialogue systems › question understanding
question decomposition
0.312025
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for Deep Research · SIGIR 2025
Machine learning › Efficient and distributed learning
model compression
0.212024
Dodo: Dynamic Contextual Compression for Decoder-only LMs · ACL (1) 2024
Natural language and speech › Information extraction and text analysis
template-based extraction
0.112018
Learning Latent Semantic Annotations for Grounding Natural Language to Structured Data · EMNLP 2018

Methods — techniques the papers use, named apart from their topics

slow thinking · 1.7large language model · 1.7decomposition · 1.7parameter tuning · 0.8LoRA · 0.8machine translation · 0.7autoencoding · 0.7neural network · 0.4datalog · 0.4bidirectional LSTM · 0.4
YearPublicationVenuePosition
2025 Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for Deep Research
abstract
Existing question answering (QA) datasets are no longer challenging to most powerful Large Language Models (LLMs). Traditional QA benchmarks like TriviaQA, NaturalQuestions, ELI5 and HotpotQA mainly study ''known unknowns'' with clear indications of both what information is missing, and how to find it to answer the question. A yet unmet need of the NLP community is a bank of non-factoid, multi-perspective questions involving a great deal of unclear information needs, i.e. ''unknown unknowns''. We claim we can find such questions in search engine logs, which is surprising because most question-intent queries are indeed factoid. Furthermore, recent products like Google's DeepResearch (announced a year after this resource was released publicly) specifically address such queries, retrieving hundreds of documents to synthesize report-style responses. We present Researchy Questions, the world's first, only and largest public dataset of ''Deep Research'' questions filtered from real search engine logs to be non-factoid, ''decompositional'' and multi-perspective. We show that users spend substantial ''effort'' on these questions in terms of signals like clicks and session length. We also show that ''slow thinking'' answering techniques, like decomposition into sub-questions shows benefit over answering directly. We release (at https://huggingface.co/datasets/corbyrosset/researchy_questions) about 100k Researchy Questions with a permissive CDLA-2.0 license, along with click histograms on over 350k Clueweb22 URLs that were clicked for each question.
Corbin Rosset, Ho-Lam Chung, Guanghui Qin, Ethan C. Chau, Ahmed Awadallah 0001, Jennifer Neville, Nikhil Rao 0001
SIGIR3
2024 Dodo: Dynamic Contextual Compression for Decoder-only LMs
abstract
Transformer-based language models (LMs) are inefficient in long contexts.We propose DODO , a solution for context compression.Instead of one vector per token in a standard transformer model, DODO represents text with a dynamic number of hidden states at each layer, reducing the cost of self-attention to a fraction of typical time and space.Moreover, off-the-shelf models such as LLAMA can be adapted to DODO by efficient parameter tuning methods such as LoRA.In use, DODO can act as either an autoregressive LM or a context compressor for downstream tasks.We demonstrate through experiments in language modeling, question answering, and summarization that DODO retains capabilities in these tasks, while drastically reducing the overhead during decoding.For example, in the autoencoding task, DODO shrinks context at a 20x compression ratio with a BLEU score of 98% for reconstruction, achieving nearly lossless encoding.
Guanghui Qin, Corby Rosset, Ethan C. Chau, Benjamin Van Durme
ACL (1)1
2023 The NLP Task Effectiveness of Long-Range Transformers
abstract
Transformer models cannot easily scale to long sequences due to their O(N 2 ) time and space complexity.This has led to Transformer variants seeking to lower computational complexity, such as Longformer and Performer.While such models have theoretically greater efficiency, their effectiveness on real NLP tasks has not been well studied.We benchmark 7 variants of Transformer models on 5 difficult NLP tasks and 7 datasets.We design experiments to isolate the effect of pretraining and hyperparameter settings, to focus on their capacity for long-range attention.Moreover, we present various methods to investigate attention behaviors to illuminate model details beyond metric scores.We find that the modified attention in long-range transformers has advantages on content selection and query-guided decoding, but they come with previously unrecognized drawbacks such as insufficient attention to distant tokens and accumulated approximation error.
Guanghui Qin, Yukun Feng, Benjamin Van Durme
EACL1
2023 Nugget: Neural Agglomerative Embeddings of Text
abstract
Embedding text sequences is a widespread requirement in modern language understanding. Existing approaches focus largely on constant-size representations. This is problematic, as the amount of information contained in text often varies with the length of the input. We propose a solution called Nugget, which encodes language into a representation based on a dynamically selected subset of input tokens. These nuggets are learned through tasks like autoencoding and machine translation, and intuitively segment language into meaningful units. We demonstrate Nugget outperforms related approaches in tasks involving semantic comparison. Finally, we illustrate these compact units allow for expanding the contextual window of a language model (LM), suggesting new future LMs that can condition on significantly larger amounts of content.
Guanghui Qin, Benjamin Van Durme
ICML1
2021 Everything Is All It Takes: A Multipronged Strategy for Zero-Shot Cross-Lingual Information Extraction
abstract
Mahsa Yarmohammadi, Shijie Wu, Marc Marone, Haoran Xu, Seth Ebner, Guanghui Qin, Yunmo Chen, Jialiang Guo, Craig Harman, Kenton Murray, Aaron Steven White, Mark Dredze, Benjamin Van Durme. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Mahsa Yarmohammadi, Marc Marone, Seth Ebner, Guanghui Qin, Yunmo Chen, Jialiang Guo, Craig Harman, Kenton Murray, Aaron Steven White, Mark Dredze, Benjamin Van Durme
EMNLP (1)6
2021 Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
abstract
Natural-language prompts have recently been used to coax pretrained language models into performing other AI tasks, using a fill-in-theblank paradigm (Petroni et al., 2019) or a few-shot extrapolation paradigm (Brown et al., 2020).For example, language models retain factual knowledge from their training corpora that can be extracted by asking them to "fill in the blank" in a sentential prompt.However, where does this prompt come from?We explore the idea of learning prompts by gradient descent-either fine-tuning prompts taken from previous work, or starting from random initialization.Our prompts consist of "soft words," i.e., continuous vectors that are not necessarily word type embeddings from the language model.Furthermore, for each task, we optimize a mixture of prompts, learning which prompts are most effective and how to ensemble them.Across multiple English LMs and tasks, our approach hugely outperforms previous methods, showing that the implicit factual knowledge in language models was previously underestimated.Moreover, this knowledge is cheap to elicit: random initialization is nearly as good as informed initialization.
Guanghui Qin, Jason Eisner
NAACL-HLT1
2021 Iterative Paraphrastic Augmentation with Discriminative Span Alignment
abstract
Abstract We introduce a novel paraphrastic augmentation strategy based on sentence-level lexically constrained paraphrasing and discriminative span alignment. Our approach allows for the large-scale expansion of existing datasets or the rapid creation of new datasets using a small, manually produced seed corpus. We demonstrate our approach with experiments on the Berkeley FrameNet Project, a large-scale language understanding effort spanning more than two decades of human labor. With four days of training data collection for a span alignment model and one day of parallel compute, we automatically generate and release to the community 495,300 unique (Frame,Trigger) pairs in diverse sentential contexts, a roughly 50-fold expansion atop FrameNet v1.7. The resulting dataset is intrinsically and extrinsically evaluated in detail, showing positive results on a downstream task.
Ryan Culkin, Edward J. Hu, Elias Stengel-Eskin, Guanghui Qin, Benjamin Van Durme
Trans. Assoc. Comput. Linguistics4
2020 Neural Datalog Through Time: Informed Temporal Modeling via Logical Specification
abstract
Learning how to predict future events from patterns of past events is difficult when the set of possible event types is large. Training an unrestricted neural model might overfit to spurious patterns. To exploit domain-specific knowledge of how past events might affect an event’s present probability, we propose using a temporal deductive database to track structured facts over time. Rules serve to prove facts from other facts and from past events. Each fact has a time-varying state—a vector computed by a neural net whose topology is determined by the fact’s provenance, including its experience of past events. The possible event types at any time are given by special facts, whose probabilities are neurally modeled alongside their states. In both synthetic and real-world domains, we show that neural probabilistic models derived from concise Datalog programs improve prediction by encoding appropriate domain knowledge in their architecture.
Hongyuan Mei, Guanghui Qin, Minjie Xu, Jason Eisner
ICML2
2019 Imputing Missing Events in Continuous-Time Event Streams
abstract
Events in the world may be caused by other, unobserved events. We consider sequences of events in continuous time. Given a probability model of complete sequences, we propose particle smoothing—a form of sequential importance sampling—to impute the missing events in an incomplete sequence. We develop a trainable family of proposal distributions based on a type of bidirectional continuous-time LSTM: Bidirectionality lets the proposals condition on future observations, not just on the past as in particle filtering. Our method can sample an ensemble of possible complete sequences (particles), from which we form a single consensus prediction that has low Bayes risk under our chosen loss metric. We experiment in multiple synthetic and real domains, using different missingness mechanisms, and modeling the complete sequences in each domain with a neural Hawkes process (Mei & Eisner 2017). On held-out incomplete sequences, our method is effective at inferring the ground-truth unobserved events, with particle smoothing consistently improving upon particle filtering.
Hongyuan Mei, Guanghui Qin, Jason Eisner
ICML2
2018 Learning Latent Semantic Annotations for Grounding Natural Language to Structured Data
abstract
Previous work on grounded language learning did not fully capture the semantics underlying the correspondences between structured world state representations and texts, especially those between numerical values and lexical terms.In this paper, we attempt at learning explicit latent semantic annotations from paired structured tables and texts, establishing correspondences between various types of values and texts.We model the joint probability of data fields, texts, phrasal spans, and latent annotations with an adapted semi-hidden Markov model, and impose a soft statistical constraint to further improve the performance.As a by-product, we leverage the induced annotations to extract templates for language generation.Experimental results suggest the feasibility of the setting in this study, as well as the effectiveness of our proposed framework.1
Guanghui Qin, Jin-Ge Yao, Jinpeng Wang 0001, Chin-Yew Lin
EMNLP1