EDBT 2026 Demo / reviewers in the wild / expert
Bin Wu 0025
dblp:98/4432-25
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-8677-2321ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Context Interference for Reliable and Efficient Search AgentsabstractBoyang Xue, Bin Wu, Shuofei Qiao, Sheng Wang, Rui Wang, Yiming Du, Hongru Wang, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Boyang Xue, Bin Wu 0025, Shuofei Qiao, Rui Wang 0092, Yiming Du, Hongru Wang 0003, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani |
ACL (1) | 2 |
| 2026 | AgentSearch: Indexing, Retrieval, and Ranking of AI Agents
Bin Wu 0025, To Eun Kim, Yue Feng 0002, Fernando Diaz 0001, Zhaochun Ren, Emine Yilmaz |
SIGIR | 1 |
| 2025 | Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search ReasoningabstractXiang Zhuang, Bin Wu, Jiyu Cui, Kehua Feng, Xiaotong Li, Huabin Xing, Keyan Ding, Qiang Zhang, Huajun Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiang Zhuang, Bin Wu 0025, Jiyu Cui, Kehua Feng, Huabin Xing, Keyan Ding, Qiang Zhang 0026, Huajun Chen |
ACL (1) | 2 |
| 2025 | Rethinking the Potential of Multimodality in Collaborative Problem Solving Diagnosis with Large Language Models
Kester Wong, Bin Wu 0025, Sahan Bulathwela, Mutlu Cukurova |
AIED (2) | 2 |
| 2025 | Empirical Analysis on User Profile in Personalized LLMsabstractUtilizing user profiles to personalize Large Language Models (LLMs) has been shown to enhance performance on a wide range of tasks. However, the precise role of user profiles and their effect mechanism on LLMs is unclear. This study first confirms that the effectiveness of user profiles stems primarily from their personalization information, with input-relevant information contributing meaningfully only when built upon personalization. Furthermore, we investigate how user profiles affect the personalization of LLMs. Within the user profile, we reveal that it is the historical personalized response produced or approved by users that plays a pivotal role in personalizing LLMs. This discovery unlocks the potential of LLMs to incorporate more user profiles within the constraints of limited input length. As for the position of user profiles, we observe that user profiles integrated into different positions of the input context do not contribute equally to personalization. Instead, user profiles closer to the beginning have more impact on the personalization of LLMs. Our findings reveal the role of user profiles for the personalization of LLMs, and showcase how incorporating user profiles impacts performance to leverage user profiles effectively. Bin Wu 0025, Zhengyan Shi, Hossein A. Rahmani, Varsha Ramineni, Emine Yilmaz |
CIKM | 1 |
| 2025 | Entropy-Based Decoding for Retrieval-Augmented Large Language ModelsabstractZexuan Qiu, Zijing Ou, Bin Wu, Jingjing Li, Aiwei Liu, Irwin King. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zexuan Qiu, Zijing Ou, Bin Wu 0025, Jingjing Li 0007, Aiwei Liu, Irwin King |
NAACL (Long Papers) | 3 |
| 2024 | Instruction Tuning With Loss Over InstructionsabstractInstruction tuning plays a crucial role in shaping the outputs of language models (LMs) to desired styles. In this work, we propose a simple yet effective method, Instruction Modelling (IM), which trains LMs by applying a loss function to the instruction and prompt part rather than solely to the output part. Through experiments across 21 diverse benchmarks, we show that, in many scenarios, IM can effectively improve the LM performance on both NLP tasks (*e.g.,* MMLU, TruthfulQA, and HumanEval) and open-ended generation benchmarks (*e.g.,* MT-Bench and AlpacaEval). Remarkably, in the most advantageous case, IM boosts model performance on AlpacaEval 1.0 by over 100%. We identify two key factors influencing the effectiveness of IM: (1) The ratio between instruction length and output length in the training data; and (2) The number of training examples. We observe that IM is especially beneficial when trained on datasets with lengthy instructions paired with brief outputs, or under the Superficial Alignment Hypothesis (SAH) where a small amount of training examples are used for instruction tuning. Further analysis substantiates our hypothesis that our improvement can be attributed to reduced overfitting to instruction tuning datasets. It is worth noting that we are not proposing \ours as a replacement for the current instruction tuning process.
Instead, our work aims to provide practical guidance for instruction tuning LMs, especially in low-resource scenarios.
Our code is available at https://github.com/ZhengxiangShi/InstructionModelling. Zhengxiang Shi, Adam X. Yang, Bin Wu 0025, Laurence Aitchison, Emine Yilmaz, Aldo Lipani |
NeurIPS | 3 |
| 2023 | Active Finetuning Protein Language Model: A Budget-Friendly Method for Directed EvolutionabstractDirected evolution is a widely-used strategy of protein engineering to improve protein function via mimicking natural mutation and selection. Machine learning-assisted directed evolution (MLDE) approaches aim to learn a fitness predictor, thereby efficiently searching for optimal mutants within the vast combinatorial mutation space. Since annotating mutants is both costly and labor-intensive, how to efficiently sample and utilize informative protein mutants to train the predictor is a critical problem in MLDE. Previous MLDE works just simply utilized pre-trained protein language models (PPLMs) for sampling without tailoring to the specific target protein of interest, which has not fully exploited the potential of PPLMs. In this work, we propose a novel method, the Actively-Finetuned Protein language model for Directed Evolution(AFP-DE), which leverages PPLMs to actively sample and fine-tune themselves, continuously improving the model’s sampling and overall performance through iterations, to achieve efficient directed protein evolution. Extensive experiments have shown the effectiveness of our method in generating optimal mutants with minimal annotation effort, outperforming previous works even with fewer annotated mutants, making it budget-friendly for biological experiments. Ming Qin, Keyan Ding, Bin Wu 0025, Haihong Yang, Hongbin Ye, Huajun Chen, Qiang Zhang 0026 |
ECAI | 3 |
| 2023 | Adaptive Compositional Continual Meta-LearningabstractThis paper focuses on continual meta-learning, where few-shot tasks are heterogeneous and sequentially available. Recent works use a mixture model for meta-knowledge to deal with the heterogeneity. However, these methods suffer from parameter inefficiency caused by two reasons: (1) the underlying assumption of mutual exclusiveness among mixture components hinders sharing meta-knowledge across heterogeneous tasks. (2) they only allow increasing mixture components and cannot adaptively filter out redundant components. In this paper, we propose an Adaptive Compositional Continual Meta-Learning (ACML) algorithm, which employs a compositional premise to associate a task with a subset of mixture components, allowing meta-knowledge sharing among heterogeneous tasks. Moreover, to adaptively adjust the number of mixture components, we propose a component sparsification method based on evidential theory to filter out redundant components. Experimental results show ACML outperforms strong baselines, showing the effectiveness of our compositional meta-knowledge, and confirming that ACML can adaptively learn meta-knowledge. Bin Wu 0025, Jinyuan Fang, Xiangxiang Zeng, Shangsong Liang, Qiang Zhang 0026 |
ICML | 1 |
| 2023 | Graph Sampling-based Meta-Learning for Molecular Property PredictionabstractMolecular property is usually observed with a limited number of samples, and researchers have considered property prediction as a few-shot problem. One important fact that has been ignored by prior works is that each molecule can be recorded with several different properties simultaneously. To effectively utilize many-to-many correlations of molecules and properties, we propose a Graph Sampling-based Meta-learning (GS-Meta) framework for few-shot molecular property prediction. First, we construct a Molecule-Property relation Graph (MPG): molecule and properties are nodes, while property labels decide edges. Then, to utilize the topological information of MPG, we reformulate an episode in meta-learning as a subgraph of the MPG, containing a target property node, molecule nodes, and auxiliary property nodes. Third, as episodes in the form of subgraphs are no longer independent of each other, we propose to schedule the subgraph sampling process with a contrastive loss function, which considers the consistency and discrimination of subgraphs. Extensive experiments on 5 commonly-used benchmarks show GS-Meta consistently outperforms state-of-the-art methods by 5.71%-6.93% in ROC-AUC and verify the effectiveness of each proposed module. Our code is available at https://github.com/HICAI-ZJU/GS-Meta. Xiang Zhuang, Qiang Zhang 0026, Bin Wu 0025, Keyan Ding, Yin Fang, Huajun Chen |
IJCAI | 3 |
| 2023 | Dynamic Bayesian Contrastive Predictive Coding Model for Personalized Product SearchabstractIn this article, we study the problem of dynamic personalized product search. Due to the data-sparsity problem in the real world, existing methods suffer from the challenge of data inefficiency. We address the challenge by proposing a Dynamic Bayesian Contrastive Predictive Coding model (DBCPC), which aims to capture the rich structured information behind search records to improve data efficiency. Our proposed DBCPC utilizes contrastive predictive learning to jointly learn dynamic embeddings with structure information of entities (i.e., users, products, and words). Specifically, our DBCPC employs structured prediction to tackle the intractability caused by non-linear output space and utilizes the time embedding technique to avoid designing different encoders each time in the Dynamic Bayesian models. In this way, our model jointly learns the underlying embeddings of entities (i.e., users, products, and words) via prediction tasks, which enables the embeddings to focus more on their general attributes and capture the general information during the preference evolution with time. For inferring the dynamic embeddings, we propose an inference algorithm combining the variational objective and the contrastive objectives. Experiments were conducted on an Amazon dataset and the experimental results show that our proposed DBCPC can learn the higher-quality embeddings and outperforms the state-of-the-art non-dynamic and dynamic models for product search. Bin Wu 0025, Zaiqiao Meng, Shangsong Liang |
ACM Trans. Web | 1 |
| 2022 | Meta-Learning Helps Personalized Product SearchabstractPersonalized product search that provides users with customized search services is an important task for e-commerce platforms. This task remains a challenge when inferring users’ preferences from few records or even no records, which is also known as the few-shot or zero-shot learning problem. In this paper, we propose a Bayesian Online Meta-Learning Model (BOML), which transfers meta-knowledge, from the inference for other users’ preferences, to help to infer the current user’s interest behind her/his few or even no historical records. To extract meta-knowledge from various inference patterns, our model constructs a mixture of meta-knowledge and transfers the corresponding meta-knowledge to the specific user according to her/his records. Based on the meta-knowledge learned from other similar inferences, our proposed model searches a ranked list of products to meet users’ personalized query intents for those with few search records (i.e., few-shot learning problem) or even no search records (i.e., zero-shot learning problem). Under the records arriving sequentially setting, we propose an online variational inference algorithm to update meta-knowledge over time. Experimental results demonstrate that our proposed BOML outperforms state-of-the-art algorithms. Bin Wu 0025, Zaiqiao Meng, Qiang Zhang 0026, Shangsong Liang |
WWW | 1 |