VLDB 2026 Research / reviewers in the wild / expert
Tianyi Bai
dblp:322/0707
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Efficient and distributed learning · 24% Vision and language · 23% Language models and text generation · 23% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 26 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
data-efficient learning |
1.7 | 2 | 2025 | Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025 Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Machine learning › Efficient and distributed learning
data selection |
1.7 | 2 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
1.7 | 2 | 2025 | Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025 Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.7 | 2 | 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025 Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation |
1.7 | 2 | 2025 | LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models · ICLR 2025 UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios · AAAI 2025 |
Program synthesis and code generation › code completion
fill-in-the-middle |
1.0 | 1 | 2026 | From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning · ACL (1) 2026 |
Computer vision › 3D vision › 3d scene understanding › multi-view understanding
cross-view reasoning |
0.9 | 1 | 2025 | UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios · AAAI 2025 |
Machine learning › Efficient and distributed learning › data-efficient learning
data-efficient pretraining |
0.9 | 1 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 |
Machine learning › Trustworthy machine learning › Data-centric AI
data influence |
0.9 | 1 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 |
Machine learning › Representation and self-supervised learning › pre-training › data-centric pre-training
data selection for pre-training |
0.9 | 1 | 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Computer vision › 3D vision › visual localization
geo-localization |
0.9 | 1 | 2025 | UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios · AAAI 2025 |
Machine learning › Trustworthy machine learning
hallucination |
0.9 | 1 | 2025 | Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining |
0.9 | 1 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model training
pretraining data selection |
0.9 | 1 | 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Machine learning › Trustworthy machine learning
synthetic data detection |
0.9 | 1 | 2025 | LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models · ICLR 2025 |
Natural language and speech › Language models and text generation
test-time scaling |
0.9 | 1 | 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025 |
Computer vision › Vision and language
visual reasoning |
0.9 | 1 | 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.6 | 1 | 2022 | Transfer Learning based Search Space Design for Hyperparameter Tuning · KDD 2022 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.6 | 1 | 2022 | Transfer Learning based Search Space Design for Hyperparameter Tuning · KDD 2022 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
search space design |
0.6 | 1 | 2022 | Transfer Learning based Search Space Design for Hyperparameter Tuning · KDD 2022 |
Machine learning › Trustworthy machine learning
AI-generated content detection |
0.3 | 1 | 2025 | LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models · ICLR 2025 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.3 | 1 | 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025 |
Computer vision › Vision and language
image captioning |
0.3 | 1 | 2025 | Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.3 | 1 | 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025 |
Computer vision › Vision and language
visual question answering |
0.3 | 1 | 2025 | Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.2 | 1 | 2022 | Transfer Learning based Search Space Design for Hyperparameter Tuning · KDD 2022 |
Methods — techniques the papers use, named apart from their topics
large multimodal model · 1.7search-and-replace instruction tuning · 1.0supervised fine-tuning · 0.9multi-dimensional data selection · 0.9multi-actor collaboration · 0.9kronecker product · 0.9influence functions · 0.9feature-level consistency loss · 0.9curriculum learning · 0.9cross-view detection-matching · 0.9clustering · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction TuningabstractJiajun Zhang, Zeyu Cui, Jiaxi Yang, Lei Zhang, Yuheng Jing, Zeyao Ma, Tianyi Bai, Zilei Wang, Qiang Liu, Liang Wang, Binyuan Hui, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiajun Zhang 0012, Zeyu Cui, Jiaxi Yang 0004, Lei Zhang 0201, Yuheng Jing, Zeyao Ma, Tianyi Bai, Zilei Wang, Qiang Liu 0006, Liang Wang 0001, Binyuan Hui, Junyang Lin |
ACL (1) | 7 |
| 2025 | UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban ScenariosabstractRecent evaluations of Large Multimodal Models (LMMs) have explored their capabilities in various domains, with only few benchmarks specifically focusing on urban environments. Moreover, existing urban benchmarks have been limited to evaluating LMMs with basic region-level urban tasks under singular views, leading to incomplete evaluations of LMMs' abilities in urban environments. To address these issues, we present UrBench, a comprehensive benchmark designed for evaluating LMMs in complex multi-view urban scenarios. UrBench contains 11.6K meticulously curated questions at both region-level and role-level that cover 4 task dimensions: Geo-Localization, Scene Reasoning, Scene Understanding, and Object Understanding, totaling 14 task types. In constructing UrBench, we utilize data from existing datasets and additionally collect data from 11 cities, creating new annotations using a cross-view detection-matching method. With these images and annotations, we then integrate LMM-based, rule-based, and human-based methods to construct large-scale high-quality questions. Our evaluations on 21 LMMs show that current LMMs struggle in the urban environments in several aspects. Even the best performing GPT-4o lags behind humans in most tasks, ranging from simple tasks such as counting to complex tasks such as orientation, localization and object attribute recognition, with an average performance gap of 17.4%. Our benchmark also reveals that LMMs exhibit inconsistent behaviors with different urban views, especially with respect to understanding cross-view relations. Baichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye, Tianyi Bai, Songyang Zhang 0001, Dahua Lin, Conghui He |
AAAI | 5 |
| 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor CollaborationabstractEfficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to 10.5% across multiple language model benchmarks compared to the state-of-the-art methods. Code and checkpoints are publicly released at https://github.com/Relaxed-System-Lab/multi-actor-data-selection. Tianyi Bai, Ling Yang 0006, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Lijun Wu 0003, Jiantao Qiu, Wentao Zhang 0001, Binhang Yuan, Conghui He |
ACL (1) | 1 |
| 2025 | Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language ModelsabstractXinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang, Tianyi Bai, Xingjian Wei, Qiu Jiantao, Chi Zhang, Ying Qian, Conghui He. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang, Tianyi Bai, Xingjian Wei, Jiantao Qiu, Conghui He |
ACL (1) | 5 |
| 2025 | LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal ModelsabstractWith the rapid development of AI-generated content, the future internet may be inundated with synthetic data, making the discrimination of authentic and credible multimodal data increasingly challenging. Synthetic data detection has thus garnered widespread attention, and the performance of large multimodal models (LMMs) in this task has attracted significant interest. LMMs can provide natural language explanations for their authenticity judgments, enhancing the explainability of synthetic content detection. Simultaneously, the task of distinguishing between real and synthetic data effectively tests the perception, knowledge, and reasoning capabilities of LMMs. In response, we introduce LOKI, a novel benchmark designed to evaluate the ability of LMMs to detect synthetic data across multiple modalities. LOKI encompasses video, image, 3D, text, and audio modalities, comprising 18K carefully curated questions across 26 subcategories with clear difficulty levels. The benchmark includes coarse-grained judgment and multiple-choice questions, as well as fine-grained anomaly selection and explanation tasks, allowing for a comprehensive analysis of LMMs. We evaluated 22 open-source LMMs and 6 closed-source models on LOKI, highlighting their potential as synthetic data detectors and also revealing some limitations in the development of LMM capabilities. More information about LOKI can be found at https://opendatalab.github.io/LOKI/. Junyan Ye, Baichuan Zhou, Junan Zhang, Tianyi Bai, Hengrui Kang, Honglin Lin, Zhizheng Wu 0001, Dahua Lin, Conghui He |
ICLR | 5 |
| 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language ModelsabstractData selection is of great significance in pretraining large language models, given the variation in quality within the large-scale available training corpora.
To achieve this, researchers are currently investigating the use of data influence to measure the importance of data instances, $i.e.,$ a high influence score indicates that incorporating this instance to the training set is likely to enhance the model performance. Consequently, they select the top-$k$ instances with the highest scores. However, this approach has several limitations.
(1) Calculating the accurate influence of all available data is time-consuming.
(2) The selected data instances are not diverse enough, which may hinder the pretrained model's ability to generalize effectively to various downstream tasks.
In this paper, we introduce $\texttt{Quad}$, a data selection approach that considers both quality and diversity by using data influence to achieve state-of-the-art pretraining results.
To compute the influence ($i.e.,$ the quality) more accurately and efficiently, we incorporate the attention layers to capture more semantic details, which can be accelerated through the Kronecker product.
For the diversity, $\texttt{Quad}$ clusters the dataset into similar data instances within each cluster and diverse instances across different clusters. For each cluster, if we opt to select data from it, we take some samples to evaluate the influence to prevent processing all instances. Overall, we favor clusters with highly influential instances (ensuring high quality) or clusters that have been selected less frequently (ensuring diversity), thereby well balancing between quality and diversity. Experiments on Slimpajama and FineWeb over 7B large language models demonstrate that $\texttt{Quad}$ significantly outperforms other data selection methods with a low FLOPs consumption. Further analysis also validates the effectiveness of our influence calculation. Chi Zhang 0102, Huaping Zhong, Chengliang Chai, Rui Wang 0119, Xinlin Zhuang, Tianyi Bai, Jiantao Qiu, Lei Cao 0004, Ju Fan, Ye Yuan 0001, Guoren Wang, Conghui He |
ICLR | 7 |
| 2025 | Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal LearningabstractMultimodal large language models (MLLMs) have achieved strong performance on vision-language tasks but still struggle with fine-grained visual differences, leading to hallucinations or missed semantic shifts. We attribute this to limitations in both training data and learning objectives. To address these issues, we propose a controlled data generation pipeline that produces minimally edited image pairs with semantically aligned captions. Using this pipeline, we construct the Micro Edit Dataset (MED), containing over 50K image-text pairs spanning 11 fine-grained edit categories, including attribute, count, position, and object presence changes.
Building on MED, we introduce a supervised fine-tuning (SFT) framework with a feature-level consistency loss that promotes stable visual embeddings under small edits. We evaluate our approach on the Micro Edit Detection benchmark, which includes carefully balanced evaluation pairs designed to test sensitivity to subtle visual variations across the same edit categories.
Our method improves difference detection accuracy and reduces hallucinations compared to strong baselines, including GPT-4o. Moreover, it yields consistent gains on standard vision-language tasks such as image captioning and visual question answering. These results demonstrate the effectiveness of combining targeted data and alignment objectives for enhancing fine-grained visual reasoning in MLLMs. Code and datasets are publicly released at https://github.com/Relaxed-System-Lab/hallu_med. Tianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun, Junlin Han, Conghui He, Wentao Zhang 0001, Binhang Yuan |
NeurIPS | 1 |
| 2025 | Multi-step Visual Reasoning with Visual Tokens Scaling and VerificationabstractMulti-modal large language models (MLLMs) have achieved remarkable capabilities by integrating visual perception with language understanding, enabling applications such as image-grounded dialogue, visual question answering, and scientific analysis. However, most MLLMs adopt a static inference paradigm, encoding the entire image into fixed visual tokens upfront, which limits their ability to iteratively refine understanding or adapt to context during inference. This contrasts sharply with human perception, which is dynamic, selective, and feedback-driven.
In this work, we introduce a novel framework for inference-time visual token scaling that enables MLLMs to perform iterative, verifier-guided reasoning over visual content. We formulate the problem as a Markov Decision Process, involving a reasoner that proposes visual actions and a verifier—trained via multi-step Direct Preference Optimization (DPO)—that evaluates these actions and determines when reasoning should terminate. To support this, we present a new dataset, VTS, comprising supervised reasoning trajectories (VTS-SFT) and preference-labeled reasoning comparisons (VTS-DPO).
Our method significantly outperforms existing approaches across diverse visual reasoning benchmarks, offering not only improved accuracy but also more interpretable and grounded reasoning processes. These results demonstrate the promise of dynamic inference mechanisms for enabling fine-grained, context-aware visual reasoning in next-generation MLLMs. Code and datasets are publicly released at https://vts-v.github.io/. Tianyi Bai, Zengjie Hu, Fupeng Sun, Jiantao Qiu, Yizhen Jiang, Guangxin He, Bohan Zeng, Conghui He, Binhang Yuan, Wentao Zhang 0001 |
NeurIPS | 1 |
| 2022 | Transfer Learning based Search Space Design for Hyperparameter TuningabstractThe tuning of hyperparameters becomes increasingly important as machine learning (ML) models have been extensively applied in data mining applications. Among various approaches, Bayesian optimization (BO) is a successful methodology to tune hyperparameters automatically. While traditional methods optimize each tuning task in isolation, there has been recent interest in speeding up BO by transferring knowledge across previous tasks. In this work, we introduce an automatic method to design the BO search space with the aid of tuning history from past tasks. This simple yet effective approach can be used to endow many existing BO methods with transfer learning capabilities. In addition, it enjoys the three advantages: universality, generality, and safeness. The extensive experiments show that our approach considerably boosts BO by designing a promising and compact search space instead of using the entire space, and outperforms the state-of-the-arts on a wide range of benchmarks, including machine learning and deep learning tuning tasks, and neural architecture search. Yang Li 0106, Yu Shen 0003, Huaijun Jiang, Tianyi Bai, Wentao Zhang 0001, Ce Zhang 0001, Bin Cui 0001 |
KDD | 4 |