EDBT 2026 Demo / reviewers in the wild / expert
Bin Dong 0003
dblp:11/6024-3
· DBLP profile ↗
24ranked-venue papers
1as first author
16since 2021 · last 2026
0009-0000-8431-0644ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RetroLM: Retrieval-Augmented KVs for Long-Context ProcessingabstractLong-context processing remains a significant challenge for large language models (LLMs). Retrieval-augmented generation (RAG) has recently emerged as a promising approach, enabling LLMs to selectively access relevant information from extended contexts to improve efficiency. However, existing RAG approaches often lag behind other efficient long-context processing methods primarily due to inherent limitations on inaccurate retrieval and fragmented contexts. To address these limitations, we propose RetroLM, a novel RAG framework designed for effective long-context processing. Unlike traditional approaches, RetroLM introduces KV-level retrieval augmentation, which partitions the LLM's KV cache into contiguous pages and performs encoding and decoding operations based on the retrieved KV pages. Built upon this framework, we further develop a specialized retriever for precise retrieval of critical pages and conduct unsupervised post-training to optimize the model’s ability to leverage retrieved information. Compared with traditional RAG, the new approach enhances robustness to retrieval inaccuracy, facilitates effective utilization of fragmented contexts, and saves the cost from repeated context-encoding operations. We conduct extensive evaluations across several popular benchmarks, including LongBench, InfiniteBench, and RULER. RetroLM consistently outperforms existing long-LLMs and RAG-based methods, especially in tasks requiring deep reasoning or extreme context lengths. Zheng Liu 0011, Shitao Xiao, Jiabei Chen, Hongjin Qian, Peitian Zhang, Shanshan Jiang 0001, Bin Dong 0003, Jun Zhao 0001, Kang Liu 0001 |
AAAI | 8 |
| 2026 | PaperRegister: Boosting Flexible-grained Paper Search via Hierarchical Register IndexingabstractZhuoqun Li, Xuanang Chen, Hongyu Lin, Yaojie Lu, Xianpei Han, Shanshan Jiang, Bin Dong, Le Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xuanang Chen, Yaojie Lu 0001, Xianpei Han, Shanshan Jiang 0001, Bin Dong 0003, Le Sun 0001 |
ACL (1) | 7 |
| 2026 | GATE: Graph-based Adaptive Tool Evolution Across Diverse TasksabstractJianwen Luo, Yiming Huang, Jinxiang Meng, Fangyu Lei, Shizhu He, Xiao Liu, Shanshan Jiang, Bin Dong, Jun Zhao, Kang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jinxiang Meng, Fangyu Lei, Shizhu He, Shanshan Jiang 0001, Bin Dong 0003, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 8 |
| 2026 | Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait
Chaolong Yang, Yuyao Yan, Chenru Jiang, Weiguang Zhao, Jie Sun 0024, Bin Dong 0003, Kaizhu Huang |
Int. J. Comput. Vis. | 9 |
| 2025 | Towards Robust Comparisons of NLP Models: A Case StudyabstractComparing the test scores of different NLP models across downstream datasets to determine which model leads to the most accurate results is the ultimate step in any experimental work. Doing so via a single mean score may not accurately quantify the real capabilities of the models. Previous works have proposed diverse statistical tests to improve the comparison of NLP models; however, a key statistical phenomenon remains understudied: variability in test scores. We propose a type of regression analysis which better explains this phenomenon by isolating the effect of both nuisance factors (such as random seeds) and datasets from the effects of the models’ capabilities. We showcase our approach via a case study of some of the most popular biomedical NLP models: after isolating nuisance factors and datasets, our results show that the difference between BioLinkBERT and MSR BiomedBERT is, actually, 7 times smaller than previously reported. Vicente Iván Sánchez Carmona, Shanshan Jiang 0001, Bin Dong 0003 |
COLING | 3 |
| 2025 | Why and How LLMs Benefit from Knowledge Introspection in Commonsense ReasoningabstractLarge Language Models (LLMs) can improve commonsense reasoning through generating intermediate knowledge.However, the effectiveness of this knowledge introspection is not always guaranteed.This paper first systematically investigates and reveals an introspection paradox: while simple introspection tends to benefit weaker models, it often degrades the performance of stronger ones, particularly on simpler tasks.Our deep analysis indicates that this paradox arises from a complex interplay among model capability, task difficulty and the quality of generated knowledge.Further interpretability analysis reveals the origins of low-quality knowledge generation.To better employ introspected knowledge in LLM, this paper proposes a training-free Adaptive Introspection Strategy that operates in two stages using only the model's internal states: Knowledge Detection, which dynamically identifies and discards potentially low-quality knowledge, and Knowledge Regeneration, which employs attention smoothing to guide the model away from harmful failure modes during knowledge generation.Extensive experiments on five Llama models with different sizes and eight commonsense reasoning benchmarks demonstrate that our approach effectively mitigates the limitations of standard introspection and has consistent performance gains across almost all settings. Chengfeng Zhao, Shizhu He, Shanshan Jiang 0001, Bin Dong 0003, Jun Zhao 0001, Kang Liu 0001 |
EMNLP | 4 |
| 2025 | From 2D Images to 3D Model: Weakly Supervised Multi-View Face Reconstruction with Deep FusionabstractWhile weakly supervised multi-view face reconstruction (MVR) is garnering increased attention, one critical issue still remains open: how to effectively interact and fuse multiple image information to reconstruct high-precision 3D models. In this regard, we propose a novel pipeline called Deep Fusion MVR (DF-MVR) to explore the feature correspondences between multi-view images and reconstruct high-precision 3D faces. Specifically, we present a novel multi-view feature fusion backbone that utilizes face masks to align features from multiple encoders and integrates one multi-layer attention mechanism to enhance feature interaction and fusion, resulting in one unified facial representation. Additionally, we develop one concise face mask mechanism that facilitates multi-view feature fusion and facial reconstruction by identifying common areas and guiding the network’s focus on critical facial features (e.g., eyes, brows, nose, and mouth). Experiments on Pixel-Face and Bosphorus datasets indicate the superiority of the proposed method. Without the 3D annotation, DF-MVR achieves relative 5.2% and 3.0% RMSE improvement over the existing weakly supervised MVRs, respectively, on Pixel-Face and Bosphorus datasets. Our code is available at https://github.com/weiguangzhao/DF_MVR. Weiguang Zhao, Chaolong Yang, Jianan Ye, Rui Zhang 0012, Yuyao Yan, Xi Yang 0008, Bin Dong 0003, Amir Hussain 0001, Kaizhu Huang |
ICME | 7 |
| 2024 | Few-shot Named Entity Recognition via Superposition Concept DiscriminationabstractFew-shot NER aims to identify entities of target types with only limited number of illustrative instances. Unfortunately, few-shot NER is severely challenged by the intrinsic precise generalization problem, i.e., it is hard to accurately determine the desired target type due to the ambiguity stemming from information deficiency. In this paper, we propose Superposition Concept Discriminator (SuperCD), which resolves the above challenge via an active learning paradigm. Specifically, a concept extractor is first introduced to identify superposition concepts from illustrative instances, with each concept corresponding to a possible generalization boundary. Then a superposition instance retriever is applied to retrieve corresponding instances of these superposition concepts from large-scale text corpus. Finally, annotators are asked to annotate the retrieved instances and these annotated instances together with original illustrative instances are used to learn FS-NER models. To this end, we learn a universal concept extractor and superposition instance retriever using a large-scale openly available knowledge bases. Experiments show that SuperCD can effectively identify superposition concepts from illustrative instances, retrieve superposition instances from large-scale corpus, and significantly improve the few-shot NER performance with minimal additional efforts. Jiawei Chen 0011, Xianpei Han, Yaojie Lu 0001, Shanshan Jiang 0001, Bin Dong 0003, Le Sun 0001 |
LREC/COLING | 6 |
| 2024 | ChatGPT Is a Knowledgeable but Inexperienced Solver: An Investigation of Commonsense Problem in Large Language ModelsabstractLarge language models (LLMs) have made significant progress in NLP. However, their ability to memorize, represent, and leverage commonsense knowledge has been a well-known pain point. In this paper, we specifically focus on ChatGPT, a widely used and easily accessible LLM, and ask the following questions: (1) Can ChatGPT effectively answer commonsense questions? (2) Is ChatGPT aware of the underlying commonsense knowledge for answering a specific question? (3) Is ChatGPT knowledgeable in commonsense? (4) Can ChatGPT effectively leverage commonsense for answering questions? We conduct a series of experiments on 11 datasets to evaluate ChatGPT’s commonsense abilities, including answering commonsense questions, identifying necessary knowledge, generating knowledge descriptions, and using knowledge descriptions to answer questions again. Experimental results show that: (1) ChatGPT can achieve good QA accuracies in commonsense tasks, while still struggling with certain domains of datasets. (2) ChatGPT is knowledgeable, and can accurately generate most of the commonsense knowledge using knowledge prompts. (3) Despite its knowledge, ChatGPT is an inexperienced commonsense problem solver, which cannot precisely identify the needed commonsense for answering a specific question. These findings raise the need to explore improved mechanisms for effectively incorporating commonsense into LLMs like ChatGPT, such as better instruction following and commonsense guidance. Ning Bian, Xianpei Han, Le Sun 0001, Yaojie Lu 0001, Ben He 0001, Shanshan Jiang 0001, Bin Dong 0003 |
LREC/COLING | 8 |
| 2024 | Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language ModelsabstractMemory is one of the most essential cognitive functions serving as a repository of world knowledge and episodes of activities. In recent years, large-scale pre-trained language models have shown remarkable memorizing ability. On the contrary, vanilla neural networks without pre-training have been long observed suffering from the catastrophic forgetting problem. To investigate such a retentive-forgetful contradiction and understand the memorizing dynamic mechanism of language models, we conduct thorough experiments by controlling the target knowledge types, the learning strategies and the learning schedules. We find that: 1) Vanilla language models without pre-training are forgetful; 2) Pre-training leads to retentive language models; 3) Knowledge relevance and diversification significantly influence the memory formation. These conclusions are useful for understanding the abilities of pre-trained language models and shed light on designing and evaluating new learning and inference algorithms of language models. Boxi Cao, Qiaoyu Tang, Shanshan Jiang 0001, Bin Dong 0003, Xianpei Han, Jiawei Chen 0011, Tianshu Wang 0002, Le Sun 0001 |
LREC/COLING | 5 |
| 2024 | Scene Text Recognition via Dual-path Network with Shape-driven Attention AlignmentabstractScene text recognition (STR), one typical sequence-to-sequence problem, has drawn much attention recently in multimedia applications. To guarantee good performance, it is essential for STR to obtain aligned character-wise features from the whole-image feature maps. While most present works adopt fully data-driven attention-based alignment, such practice ignores specific character geometric information. In this article, built upon a group of learnable geometric points, we propose a novel shape-driven attention alignment method that is able to obtain character-wise features. Concretely, we first design a corner detector to generate a shape map to guide the attention alignments explicitly, where a series of points can be learned to represent character-wise features flexibly. We then propose a dual-path network with a mutual learning and cooperating strategy that successfully combines CNN with a ViT-based model, leading to further accuracy improvement. We conduct extensive experiments to evaluate the proposed method on various scene text benchmarks, including six popular regular and irregular datasets, two more challenging datasets (i.e., WordArt and OST), and three Chinese datasets. Experimental results indicate that our method can achieve superior performance with a comparable model size against many state-of-the-art models. Yijie Hu, Bin Dong 0003, Kaizhu Huang, Lei Ding 0012, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Decoupled Learning for Long-Tailed Oracle Character Recognition
Jing Li 0049, Bin Dong 0003, Qiufeng Wang 0001, Lei Ding 0012, Rui Zhang 0012, Kaizhu Huang |
ICDAR (4) | 2 |
| 2023 | Structure First Detail Next: Image Inpainting with Pyramid GeneratorabstractRecent deep generative models have achieved promising performance in image inpainting. However, it is still challenging for a neural network to generate realistic image details and textures due to its inherent spectral bias. We suggest adopting a ‘structure first detail next’ workflow for image inpainting by knowing how artists work. Thus, we propose to build a Pyramid Generator by stacking several sub-generators, where lower-layer sub-generators focus on restoring image structures. In contrast, the higher-layer sub-generators emphasize image details. Our model progressively restores the input through the entire pyramid in a bottom-up fashion. Notably, our approach has a learning scheme of progressively increasing hole size, which allows it to restore large-hole images. In addition, our method could fully exploit the benefits of learning with high-resolution images and hence is suitable for high-resolution image inpainting. Extensive experimental results on benchmark datasets have validated the effectiveness of our approach compared with state-of-the-arts. Shuyi Qu, Zhenxing Niu, Jianke Zhu, Bin Dong 0003, Kaizhu Huang |
ICME | 4 |
| 2023 | Towards Deeper and Better Multi-view Feature Fusion for 3D Semantic Segmentation
Chaolong Yang, Yuyao Yan, Weiguang Zhao, Jianan Ye, Xi Yang 0008, Amir Hussain 0001, Bin Dong 0003, Kaizhu Huang |
ICONIP (15) | 7 |
| 2022 | Towards Accurate Alignment and Sufficient Context in Scene Text Recognition
Yijie Hu, Bin Dong 0003, Qiufeng Wang 0001, Lei Ding 0012, Xiao-Bo Jin, Kaizhu Huang |
ICONIP (3) | 2 |
| 2022 | Rethinking Image Inpainting with Attention Feature Fusion
Shuyi Qu, Kaizhu Huang, Qiufeng Wang 0001, Bin Dong 0003 |
ICONIP (3) | 4 |
| 2019 | Gazetteer-Enhanced Attentive Neural Networks for Named Entity RecognitionabstractHongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Bin Dong, Shanshan Jiang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yaojie Lu 0001, Xianpei Han, Le Sun 0001, Bin Dong 0003, Shanshan Jiang 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2017 | Building a comprehensive syntactic and semantic corpus of Chinese clinical texts
Bin He 0005, Bin Dong 0003, Yi Guan, Jinfeng Yang, Qiubin Yu, Jianyi Cheng, Chunyan Qu |
J. Biomed. Informatics | 2 |
| 2013 | A novel discriminative method for pronunciation quality assessmentabstractThis paper presents a novel method for automatic pronunciation quality assessment. Unlike the traditional “Goodness of Pronunciation” (GOP) method, we judged utterance's pronunciation quality directly by a discriminative method. Under this novel framework, we also designed an algorithm to calculate the assessment confidence. We decoded the student's utterance for two passes. The first-pass decoding was just for getting the phone time points, and the second-pass decoding was for differentiating the pronunciation quality for each triphone. In the second-pass decoding, we used a specially trained acoustic model (AM), where the triphones in different pronunciation qualities were trained as different units. The confidence of the phone-level scoring was also calculated, and the low confidence phone-level scores were excluded in calculating the word-level score. The experimental results shows that the scoring performance was increased significantly compared to the traditional GOP method. Fuping Pan, Bin Dong 0003, Yonghong Yan 0002 |
ICASSP | 3 |
| 2009 | A one-step tone recognition approach using MSD-HMM for continuous speech
Changliang Liu, Fengpei Ge, Fuping Pan, Bin Dong 0003, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2009 | An SVM-Based Mandarin Pronunciation Quality Assessment System
Fengpei Ge, Fuping Pan, Changliang Liu, Bin Dong 0003, Shui-duen Chan, Yonghong Yan 0002 |
ISNN (4) | 4 |
| 2009 | Dynamic Multiple Pronunciation Incorporation in a Refined Search Space for Reading Miscue Detection
Changliang Liu, Fuping Pan, Fengpei Ge, Bin Dong 0003, Shuiduen Chen, Yonghong Yan 0002 |
ISNN (4) | 4 |
| 2008 | Forward optimal modeling of acoustic confusions in Mandarin CALL system
Fengpei Ge, Fuping Pan, Changliang Liu, Bin Dong 0003, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2005 | Fast confidence measure algorithm for continuous speech recognition
Bin Dong 0003, Qingwei Zhao, Yonghong Yan 0002 |
INTERSPEECH | 1 |