Hanbin Wang

dblp:235/8973 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 56% Reinforcement learning · 44%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
Advancing LLM Reasoning Generalists with Preference Trees · ICLR 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Advancing LLM Reasoning Generalists with Preference Trees · ICLR 2025
Information retrieval › document retrieval › domain-specific retrieval
code search
0.912025
Building a Coding Assistant via the Retrieval-Augmented Language Model · ACM Trans. Inf. Syst. 2025
Information retrieval
retrieval models
0.912025
Building a Coding Assistant via the Retrieval-Augmented Language Model · ACM Trans. Inf. Syst. 2025
Program synthesis and code generation › code generation with language models
retrieval-augmented code generation
0.912025
Building a Coding Assistant via the Retrieval-Augmented Language Model · ACM Trans. Inf. Syst. 2025
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.312025
Advancing LLM Reasoning Generalists with Preference Trees · ICLR 2025
Program synthesis and code generation
code completion
0.312025
Building a Coding Assistant via the Retrieval-Augmented Language Model · ACM Trans. Inf. Syst. 2025

Methods — techniques the papers use, named apart from their topics

retrieval-augmented language model · 1.7pre-training · 1.7contrastive alignment · 1.7reward modeling objective · 0.9preference trees · 0.9
YearPublicationVenuePosition
2025 Advancing LLM Reasoning Generalists with Preference Trees
abstract
We introduce EURUS, a suite of large language models (LLMs) optimized for reasoning. Finetuned from Mistral-7B, Llama-3-8B, and Mixtral-8x22B, EURUS models achieve state-of-the-art results among open-source models on a diverse set of benchmarks covering mathematics, code generation, and logical reasoning problems. Notably, EURUX-8X22B outperforms GPT-3.5 Turbo in reasoning through a comprehensive benchmarking across 12 test sets covering five tasks. The strong performance of EURUS can be primarily attributed to ULTRAINTERACT, our newly-curated large-scale, high-quality training data dataset specifically designed for complex reasoning tasks. ULTRAINTERACT can be used in both supervised fine-tuning, preference learning, and reward modeling. It pairs each instruction with a preference tree consisting of (1) reasoning chains with diverse planning strategies in a unified format, (2) multi-turn interaction trajectories with the environment and the critique, and (3) pairwise positive and negative responses to facilitate preference learning. ULTRAINTERACT allows us to conduct an in-depth exploration of preference learning for reasoning tasks. Our investigation reveals that some well-established preference learning algorithms may be less suitable for reasoning tasks compared to their effectiveness in general conversations. The hypothesis is that in reasoning tasks, the space of correct answers is much smaller than that of incorrect ones, so it is necessary to explicitly increase the reward of chosen data. Therefore, in addition to increasing the reward margin as many preference learning algorithms do, the absolute values of positive responses’ rewards should be positive and may serve as a proxy for performance. Inspired by this, we derive a novel reward modeling objective and empirically that it leads to a stable reward modeling curve and better performance. Together with ULTRAINTERACT, we obtain a strong reward model.
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding 0002, Xingyao Wang 0002, Boji Shan, Zeyuan Liu, Ruobing Xie, Yankai Lin 0001, Zhenghao Liu 0001, Bowen Zhou 0002, Hao Peng 0015, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR3
2025 Building a Coding Assistant via the Retrieval-Augmented Language Model
abstract
Pretrained language models have shown strong effectiveness in code-related tasks, such as code retrieval, code generation, code summarization, and code completion tasks. In this article, we propose COde assistaNt viA retrieval-augmeNted language model (CONAN), which aims to build a code assistant by mimicking the knowledge-seeking behaviors of humans during coding. Specifically, it consists of a code structure-aware retriever (CONAN-R) and a dual-view code representation-based retrieval-augmented generation model (CONAN-G). CONAN-R pretrains CodeT5 using Code-Documentation Alignment and Masked Entity Prediction tasks to make language models code structure-aware and learn effective representations for code snippets and documentation. Then CONAN-G designs a dual-view code representation mechanism for implementing a retrieval-augmented code generation model. CONAN-G regards the code documentation descriptions as prompts, which help language models better understand the code semantics. Our experiments show that CONAN achieves convincing performance on different code generation tasks and significantly outperforms previous retrieval augmented code generation models. Our further analyses show that CONAN learns tailored representations for both code snippets and documentation by aligning code-documentation data pairs and capturing structural semantics by masking and predicting entities in the code data. Additionally, the retrieved code snippets and documentation provide necessary information from both program language and natural language to assist the code generation process. CONAN can also be used as an assistant for Large Language Models (LLMs), providing LLMs with external knowledge in shorter code document lengths to improve their effectiveness on various code tasks. It shows the ability of CONAN to extract necessary information and help filter out the noise from retrieved code documents.
Hanbin Wang, Zhenghao Liu 0001, Shi Yu 0001, Shuo Wang 0013, Yukun Yan, Yu Gu 0002, Ge Yu 0001
ACM Trans. Inf. Syst.2
2022 Robust Calibration-Marker and Laser-Line Detection For Underwater 3d Shape Reconstruction By Deep Neural Network
abstract
There are various demands for underwater 3D reconstruction, however, since most active stereo 3D reconstruction methods focus on the air environment, it is difficult to directly apply them to underwater due to the several critical reasons, such as refraction, water flow and severe attenuation. Typically, calibration-markers or laser-lines are strongly blurred and saturated by attenuation, which makes difficult to recover shape in the water. Another problem is that it is difficult to keep cameras, projectors and objects static in the water because of strong water flow, which prevents accurate calibration. In this paper, we propose a method to solve those problems by novel algorithm using deep neural network (DNN), epipolar constraint and specially designed devices. We also built a real system and tested it in the water, e.g., pool and sea. Experimental results confirmed the effectiveness of the proposed method. We also demonstrated real 3D scan in the sea.
Hanbin Wang, Takafumi Iwaguchi, Hiroshi Kawasaki
ICIP1