VLDB 2026 Research / reviewers in the wild / expert
Giwon Hong
dblp:243/1632
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-5239-9189ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 47% Trustworthy machine learning · 12% Question answering and dialogue systems · 11% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% | |
| Theoretical computer science
1 paper |
Automated reasoning and model checking · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › model steering
language model steering |
1.0 | 1 | 2026 | Compositional Steering of Large Language Models with Steering Tokens · ACL (1) 2026 |
Natural language and speech › Language models and text generation › alignment
LLM behavior control |
1.0 | 1 | 2026 | Compositional Steering of Large Language Models with Steering Tokens · ACL (1) 2026 |
Natural language and speech › Language models and text generation › in-context learning
demonstration selection |
0.9 | 1 | 2025 | Mixtures of In-Context Learners · ACL (1) 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | Mixtures of In-Context Learners · ACL (1) 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | Mixtures of In-Context Learners · ACL (1) 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | Theorem Prover as a Judge for Synthetic Data Generation · ACL (1) 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | GRADA: Graph-based Reranking against Adversarial Documents Attack · EMNLP 2025 |
Machine learning › Generative modeling
synthetic data generation |
0.9 | 1 | 2025 | Theorem Prover as a Judge for Synthetic Data Generation · ACL (1) 2025 |
Information retrieval
retrieval-augmented generation |
0.9 | 1 | 2025 | GRADA: Graph-based Reranking against Adversarial Documents Attack · EMNLP 2025 |
Automated reasoning and model checking › automated reasoning › mathematical reasoning
autoformalization |
0.9 | 1 | 2025 | Theorem Prover as a Judge for Synthetic Data Generation · ACL (1) 2025 |
Automated reasoning and model checking
theorem proving |
0.9 | 1 | 2025 | Theorem Prover as a Judge for Synthetic Data Generation · ACL (1) 2025 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.6 | 1 | 2022 | Graph-Induced Transformers for Efficient Multi-Hop Question Answering · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.5 | 1 | 2021 | Have You Seen That Number? Investigating Extrapolation in Question Answering Models · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › mathematical reasoning
numerical reasoning |
0.5 | 1 | 2021 | Have You Seen That Number? Investigating Extrapolation in Question Answering Models · EMNLP (1) 2021 |
Information retrieval › retrieval models
neural retrieval |
0.5 | 1 | 2021 | Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval · EMNLP (1) 2021 |
Information retrieval
retrieval models |
0.5 | 1 | 2021 | Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval · EMNLP (1) 2021 |
Information retrieval
sparse feature learning |
0.5 | 1 | 2021 | Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval · EMNLP (1) 2021 |
Information retrieval › retrieval models
sparse retrieval |
0.5 | 1 | 2021 | Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.3 | 1 | 2025 | GRADA: Graph-based Reranking against Adversarial Documents Attack · EMNLP 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2022 | Graph-Induced Transformers for Efficient Multi-Hop Question Answering · EMNLP 2022 |
Machine learning › Learning theory › generalization
extrapolation |
0.1 | 1 | 2021 | Have You Seen That Number? Investigating Extrapolation in Question Answering Models · EMNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
theorem prover feedback · 1.7iterative autoformalisation · 1.7graph-based reranking · 1.7self-distillation · 1.0activation steering · 1.0LoRA merging · 1.0gradient-based optimization · 0.9ensemble weighting · 0.9graph-induced attention · 0.6digit-by-digit number representation · 0.5bucketing · 0.5binarization · 0.5BERT · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compositional Steering of Large Language Models with Steering TokensabstractDeploying LLMs in real-world applications requires controllable output that satisfies multiple desiderata at the same time.While existing work extensively addresses LLM steering for a single behavior, compositional steering-i.e., steering LLMs simultaneously towards multiple behaviors-remains an underexplored problem.In this work, we propose compositional steering tokens for multi-behavior steering.We first embed individual behaviors, expressed as natural language instructions, into dedicated tokens via self-distillation.Contrary to most prior work, which operates in the activation space, our behavior steers live in the space of input tokens, enabling more effective zero-shot composition.We then train a dedicated composition token on pairs of behaviors and show that it successfully captures the notion of composition: it generalizes well to unseen compositions, including those with unseen behaviors as well as those with an unseen number of behaviors.Our experiments across different LLM architectures show that steering tokens lead to superior multi-behavior steering of verifiable constraints (e.g., length, format, structure, language) compared to competing approaches (instructions, activation steering, and LoRA merging).Moreover, we show that steering tokens complement natural language instructions, with their combination resulting in further gains. Gorjan Radevski, Kiril Gashteovski, Giwon Hong, Carolin Lawrence, Goran Glavas |
ACL (1) | 3 |
| 2025 | Mixtures of In-Context LearnersabstractIn-context learning (ICL) adapts LLMs by providing demonstrations without fine-tuning the model parameters; however, it is very sensitive to the choice of in-context demonstrations, and processing many demonstrations can be computationally demanding.We propose Mixtures of In-Context Learners (MOICL), a novel approach that uses subsets of demonstrations to train a set of experts via ICL and learns a weighting function to merge their output distributions via gradient-based optimisation.In our experiments, we show performance improvements on 5 out of 7 classification datasets compared to a set of strong baselines (e.g., up to +13% compared to ICL and LENS).Moreover, we improve the Pareto frontier of ICL by reducing the inference time needed to achieve the same performance with fewer demonstrations.Finally, MOICL is more robust to out-ofdomain (up to +11%), imbalanced (up to +49%) and perturbed demonstrations (up to +38%). 1 Giwon Hong, Emile van Krieken, Edoardo Maria Ponti, Nikolay Malkin, Pasquale Minervini |
ACL (1) | 1 |
| 2025 | Theorem Prover as a Judge for Synthetic Data GenerationabstractThe demand for synthetic data in mathematical reasoning has increased due to its potential to enhance the mathematical capabilities of large language models (LLMs).However, ensuring the validity of intermediate reasoning steps remains a significant challenge, affecting data quality.While formal verification via theorem provers effectively validates LLM reasoning, the autoformalisation of mathematical proofs remains error-prone.We introduce iterative autoformalisation, an approach that iteratively refines theorem prover formalisation to mitigate errors, thereby increasing the execution rate on the Lean prover from 60% to 87%.Building upon that, we introduce Theorem Prover as a Judge (TP-as-a-Judge), a method that makes use of theorem prover formalisation to rigorously assess LLM intermediate reasoning, effectively integrating autoformalisation with synthetic data generation.Finally, we present Reinforcement Learning from Theorem Prover Feedback (RLTPF), a framework that replaces human annotation with theorem prover feedback in Reinforcement Learning from Human Feedback (RLHF).Across multiple LLMs, applying TP-as-a-Judge and RLTPF improves benchmarks with only 3,508 samples, achieving 5.56% accuracy gain on Mistral-7B for MultiArith, 6.00% on Llama-2-7B for SVAMP, and 3.55% on Llama-3.1-8B for AQUA. 1 Joshua Ong Jun Leang, Giwon Hong, Shay B. Cohen |
ACL (1) | 2 |
| 2025 | GRADA: Graph-based Reranking against Adversarial Documents AttackabstractRetrieval Augmented Generation (RAG) frameworks can improve the factual accuracy of large language models (LLMs) by integrating external knowledge from retrieved documents, which is useful for overcoming the limitations of models' static intrinsic knowledge.However, these systems are susceptible to adversarial attacks that manipulate the retrieval process by introducing documents that are adversarial yet semantically similar to the query.Notably, while these adversarial documents resemble the query, they exhibit weak similarity to benign documents in the retrieval set.Thus, we propose a simple yet effective Graph-based Reranking against Adversarial Document Attacks (GRADA) framework aimed at preserving retrieval quality while significantly reducing the success of adversaries.Our study evaluates the effectiveness of our approach through experiments conducted on six LLMs: GPT-3.5-Turbo,GPT-4o, Llama3.1-8b-Instruct,Llama3.1-70b-Instruct,Qwen2.5-7b-Instruct, and Qwen2.5-14b-Instruct.We use three datasets to assess performance, with results from the Natural Questions dataset showing up to an 80% reduction in attack success rates while maintaining minimal loss in accuracy. Jingjie Zheng, Aryo Pradipta Gema, Giwon Hong, Xuanli He, Pasquale Minervini, Youcheng Sun, Qiongkai Xu |
EMNLP | 3 |
| 2025 | Are We Done with MMLU?abstractAryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile Van Krieken, Pasquale Minervini. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao 0043, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, Pasquale Minervini |
NAACL (Long Papers) | 3 |
| 2025 | Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation EngineeringabstractYu Zhao, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang, Xuanli He, Kam-Fai Wong, Pasquale Minervini. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yu Zhao 0043, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang 0003, Xuanli He, Kam-Fai Wong, Pasquale Minervini |
NAACL (Long Papers) | 3 |
| 2022 | Graph-Induced Transformers for Efficient Multi-Hop Question AnsweringabstractA graph is a suitable data structure to represent the structural information of text.Recently, multi-hop question answering (MHQA) tasks, which require inter-paragraph/sentence linkages, have come to exploit such properties of a graph.Previous approaches to MHQA relied on leveraging the graph information along with the pre-trained language model (PLM) encoders.However, this trend exhibits the following drawbacks: (i) sample inefficiency while training in a low-resource setting; (ii) lack of reusability due to changes in the model structure or input.Our work proposes the Graph-Induced Transformer (GIT) that applies graphderived attention patterns directly into a PLM, without the need to employ external graph modules.GIT can leverage the useful inductive bias of graphs while retaining the unperturbed Transformer structure and parameters.Our experiments on HotpotQA successfully demonstrate both the sample efficient characteristic of GIT and its capacity to replace the graph modules while preserving model performance. Giwon Hong, Junmo Kang, Sung-Hyon Myaeng |
EMNLP | 1 |
| 2021 | Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text RetrievalabstractThe semantic matching capabilities of neural information retrieval can ameliorate synonymy and polysemy problems of symbolic approaches.However, neural models' dense representations are more suitable for re-ranking, due to their inefficiency.Sparse representations, either in symbolic or latent form, are more efficient with an inverted index.Taking the merits of the sparse and dense representations, we propose an ultra-high dimensional (UHD) representation scheme equipped with directly controllable sparsity.UHD's large capacity and minimal noise and interference among the dimensions allow for binarized representations, which are highly efficient for storage and search.Also proposed is a bucketing method, where the embeddings from multiple layers of BERT are selected/merged to represent diverse linguistic aspects.We test our models with MS MARCO and TREC CAR, showing that our models outperforms other sparse models. Kyoungrok Jang, Junmo Kang, Giwon Hong, Sung-Hyon Myaeng, Joohee Park, Taewon Yoon, Hee-Cheol Seo |
EMNLP (1) | 3 |
| 2021 | Have You Seen That Number? Investigating Extrapolation in Question Answering ModelsabstractNumerical reasoning in machine reading comprehension (MRC) has shown drastic improvements over the past few years.While the previous models for numerical MRC are able to interpolate the learned numerical reasoning capabilities, it is not clear whether they can perform just as well on numbers unseen in the training dataset.Our work rigorously tests state-of-the-art models on DROP, a numerical MRC dataset, to see if they can handle passages that contain out-of-range numbers.One of the key findings is that the models fail to extrapolate to unseen numbers.Presenting numbers as digit-by-digit input to the model, we also propose the E-digit number form that alleviates the lack of extrapolation in models and reveals the need to treat numbers differently from regular words in the text.Our work provides a valuable insight into the numerical MRC models and the way to represent number forms in MRC. Giwon Hong, Junmo Kang, Sung-Hyon Myaeng |
EMNLP (1) | 2 |
| 2020 | Handling Anomalies of Synthetic Questions in Unsupervised Question AnsweringabstractAdvances in Question Answering (QA) research require additional datasets for new domains, languages, and types of questions, as well as for performance increases.Human creation of a QA dataset like SQuAD, however, is expensive.As an alternative, an unsupervised QA approach has been proposed so that QA training data can be generated automatically.However, the performance of unsupervised QA is much lower than that of supervised QA models.We identify two anomalies in the automatically generated questions and propose how they can be mitigated.We show our approach helps improve unsupervised QA significantly across a number of QA tasks. Context Generated Question.. Giwon Hong, Junmo Kang, Doyeon Lim, Sung-Hyon Myaeng |
COLING | 1 |