EDBT 2026 Demo / reviewers in the wild / expert
Fupeng Sun
dblp:286/0449
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0003-4574-1817ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 35% Vision and language · 35% Efficient and distributed learning · 10% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.7 | 2 | 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025 Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
data-efficient learning |
0.9 | 1 | 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Machine learning › Representation and self-supervised learning › pre-training › data-centric pre-training
data selection for pre-training |
0.9 | 1 | 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Machine learning › Trustworthy machine learning
hallucination |
0.9 | 1 | 2025 | Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
0.9 | 1 | 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model training
pretraining data selection |
0.9 | 1 | 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Natural language and speech › Language models and text generation
test-time scaling |
0.9 | 1 | 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025 |
Computer vision › Vision and language
visual reasoning |
0.9 | 1 | 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025 |
Algorithmic game theory and mechanism design › auction theory
all-pay contest |
0.8 | 1 | 2024 | Restricting Entries to All-Pay Contests · EC 2024 |
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts
bayes-nash equilibrium |
0.8 | 1 | 2024 | Restricting Entries to All-Pay Contests · EC 2024 |
Algorithmic game theory and mechanism design › mechanism design
contest design |
0.8 | 1 | 2024 | Restricting Entries to All-Pay Contests · EC 2024 |
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts
symmetric equilibrium |
0.8 | 1 | 2024 | Restricting Entries to All-Pay Contests · EC 2024 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.3 | 1 | 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025 |
Computer vision › Vision and language
image captioning |
0.3 | 1 | 2025 | Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.3 | 1 | 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025 |
Computer vision › Vision and language
visual question answering |
0.3 | 1 | 2025 | Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
verifier-guided reasoning · 0.9supervised fine-tuning · 0.9multi-actor collaboration · 0.9markov decision process · 0.9feature-level consistency loss · 0.9direct preference optimization · 0.9curriculum learning · 0.9bayesian game theory · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor CollaborationabstractEfficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to 10.5% across multiple language model benchmarks compared to the state-of-the-art methods. Code and checkpoints are publicly released at https://github.com/Relaxed-System-Lab/multi-actor-data-selection. Tianyi Bai, Ling Yang 0006, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Lijun Wu 0003, Jiantao Qiu, Wentao Zhang 0001, Binhang Yuan, Conghui He |
ACL (1) | 4 |
| 2025 | Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal LearningabstractMultimodal large language models (MLLMs) have achieved strong performance on vision-language tasks but still struggle with fine-grained visual differences, leading to hallucinations or missed semantic shifts. We attribute this to limitations in both training data and learning objectives. To address these issues, we propose a controlled data generation pipeline that produces minimally edited image pairs with semantically aligned captions. Using this pipeline, we construct the Micro Edit Dataset (MED), containing over 50K image-text pairs spanning 11 fine-grained edit categories, including attribute, count, position, and object presence changes.
Building on MED, we introduce a supervised fine-tuning (SFT) framework with a feature-level consistency loss that promotes stable visual embeddings under small edits. We evaluate our approach on the Micro Edit Detection benchmark, which includes carefully balanced evaluation pairs designed to test sensitivity to subtle visual variations across the same edit categories.
Our method improves difference detection accuracy and reduces hallucinations compared to strong baselines, including GPT-4o. Moreover, it yields consistent gains on standard vision-language tasks such as image captioning and visual question answering. These results demonstrate the effectiveness of combining targeted data and alignment objectives for enhancing fine-grained visual reasoning in MLLMs. Code and datasets are publicly released at https://github.com/Relaxed-System-Lab/hallu_med. Tianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun, Junlin Han, Conghui He, Wentao Zhang 0001, Binhang Yuan |
NeurIPS | 4 |
| 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and VerificationabstractMulti-modal large language models (MLLMs) have achieved remarkable capabilities by integrating visual perception with language understanding, enabling applications such as image-grounded dialogue, visual question answering, and scientific analysis. However, most MLLMs adopt a static inference paradigm, encoding the entire image into fixed visual tokens upfront, which limits their ability to iteratively refine understanding or adapt to context during inference. This contrasts sharply with human perception, which is dynamic, selective, and feedback-driven.
In this work, we introduce a novel framework for inference-time visual token scaling that enables MLLMs to perform iterative, verifier-guided reasoning over visual content. We formulate the problem as a Markov Decision Process, involving a reasoner that proposes visual actions and a verifier—trained via multi-step Direct Preference Optimization (DPO)—that evaluates these actions and determines when reasoning should terminate. To support this, we present a new dataset, VTS, comprising supervised reasoning trajectories (VTS-SFT) and preference-labeled reasoning comparisons (VTS-DPO).
Our method significantly outperforms existing approaches across diverse visual reasoning benchmarks, offering not only improved accuracy but also more interpretable and grounded reasoning processes. These results demonstrate the promise of dynamic inference mechanisms for enabling fine-grained, context-aware visual reasoning in next-generation MLLMs. Code and datasets are publicly released at https://vts-v.github.io/. Tianyi Bai, Zengjie Hu, Fupeng Sun, Jiantao Qiu, Yizhen Jiang, Guangxin He, Bohan Zeng, Conghui He, Binhang Yuan, Wentao Zhang 0001 |
NeurIPS | 3 |
| 2024 | Restricting Entries to All-Pay ContestsabstractWe study an all-pay contest where players with low abilities are filtered prior to the round of competing for prizes. These are often practiced due to limited resources or to enhance the competitiveness of the contest. We consider a setting where the designer admits a certain number of top players into the contest. The players admitted into the contest update their beliefs about their opponents based on the signal that their abilities are among the top. We find that their posterior beliefs, even with IID priors, are correlated and depend on players' private abilities, representing a unique feature of this game. We explicitly characterize the symmetric and unique Bayesian equilibrium strategy and compare it with a contest that admits all players. We also discuss a two-stage extension where players with top first-stage efforts can proceed to the second stage competing for prizes. Fupeng Sun, Chiwei Yan, Li Jin 0004 |
EC | 1 |