Tianyi Bai

dblp:322/0707 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Efficient and distributed learning · 24% Vision and language · 23% Language models and text generation · 23%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 26 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
data-efficient learning
1.722025
Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025
Machine learning › Efficient and distributed learning
data selection
1.722025
Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025
Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model training
language model pretraining
1.722025
Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
1.722025
Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
1.722025
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models · ICLR 2025
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios · AAAI 2025
Program synthesis and code generation › code completion
fill-in-the-middle
1.012026
From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning · ACL (1) 2026
Computer vision › 3D vision › 3d scene understanding › multi-view understanding
cross-view reasoning
0.912025
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios · AAAI 2025
Machine learning › Efficient and distributed learning › data-efficient learning
data-efficient pretraining
0.912025
Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025
Machine learning › Trustworthy machine learning › Data-centric AI
data influence
0.912025
Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025
Machine learning › Representation and self-supervised learning › pre-training › data-centric pre-training
data selection for pre-training
0.912025
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025
Computer vision › 3D vision › visual localization
geo-localization
0.912025
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios · AAAI 2025
Machine learning › Trustworthy machine learning
hallucination
0.912025
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining
0.912025
Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025
Natural language and speech › Language models and text generation › large language model training
pretraining data selection
0.912025
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025
Machine learning › Trustworthy machine learning
synthetic data detection
0.912025
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models · ICLR 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025
Computer vision › Vision and language
visual reasoning
0.912025
Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.612022
Transfer Learning based Search Space Design for Hyperparameter Tuning · KDD 2022
Machine learning › Optimization for machine learning
hyperparameter optimization
0.612022
Transfer Learning based Search Space Design for Hyperparameter Tuning · KDD 2022
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
search space design
0.612022
Transfer Learning based Search Space Design for Hyperparameter Tuning · KDD 2022
Machine learning › Trustworthy machine learning
AI-generated content detection
0.312025
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models · ICLR 2025
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.312025
Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025
Computer vision › Vision and language
image captioning
0.312025
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025
Natural language and speech › Language models and text generation
preference optimization
0.312025
Multi-step Visual Reasoning with Visual Tokens Scaling and Verification · NeurIPS 2025
Computer vision › Vision and language
visual question answering
0.312025
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning · NeurIPS 2025
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.212022
Transfer Learning based Search Space Design for Hyperparameter Tuning · KDD 2022

Methods — techniques the papers use, named apart from their topics

large multimodal model · 1.7search-and-replace instruction tuning · 1.0supervised fine-tuning · 0.9multi-dimensional data selection · 0.9multi-actor collaboration · 0.9kronecker product · 0.9influence functions · 0.9feature-level consistency loss · 0.9curriculum learning · 0.9cross-view detection-matching · 0.9clustering · 0.9
YearPublicationVenuePosition
2026 From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning
abstract
Jiajun Zhang, Zeyu Cui, Jiaxi Yang, Lei Zhang, Yuheng Jing, Zeyao Ma, Tianyi Bai, Zilei Wang, Qiang Liu, Liang Wang, Binyuan Hui, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiajun Zhang 0012, Zeyu Cui, Jiaxi Yang 0004, Lei Zhang 0201, Yuheng Jing, Zeyao Ma, Tianyi Bai, Zilei Wang, Qiang Liu 0006, Liang Wang 0001, Binyuan Hui, Junyang Lin
ACL (1)7
2025 UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
abstract
Recent evaluations of Large Multimodal Models (LMMs) have explored their capabilities in various domains, with only few benchmarks specifically focusing on urban environments. Moreover, existing urban benchmarks have been limited to evaluating LMMs with basic region-level urban tasks under singular views, leading to incomplete evaluations of LMMs' abilities in urban environments. To address these issues, we present UrBench, a comprehensive benchmark designed for evaluating LMMs in complex multi-view urban scenarios. UrBench contains 11.6K meticulously curated questions at both region-level and role-level that cover 4 task dimensions: Geo-Localization, Scene Reasoning, Scene Understanding, and Object Understanding, totaling 14 task types. In constructing UrBench, we utilize data from existing datasets and additionally collect data from 11 cities, creating new annotations using a cross-view detection-matching method. With these images and annotations, we then integrate LMM-based, rule-based, and human-based methods to construct large-scale high-quality questions. Our evaluations on 21 LMMs show that current LMMs struggle in the urban environments in several aspects. Even the best performing GPT-4o lags behind humans in most tasks, ranging from simple tasks such as counting to complex tasks such as orientation, localization and object attribute recognition, with an average performance gap of 17.4%. Our benchmark also reveals that LMMs exhibit inconsistent behaviors with different urban views, especially with respect to understanding cross-view relations.
Baichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye, Tianyi Bai, Songyang Zhang 0001, Dahua Lin, Conghui He
AAAI5
2025 Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
abstract
Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to 10.5% across multiple language model benchmarks compared to the state-of-the-art methods. Code and checkpoints are publicly released at https://github.com/Relaxed-System-Lab/multi-actor-data-selection.
Tianyi Bai, Ling Yang 0006, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Lijun Wu 0003, Jiantao Qiu, Wentao Zhang 0001, Binhang Yuan, Conghui He
ACL (1)1
2025 Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models
abstract
Xinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang, Tianyi Bai, Xingjian Wei, Qiu Jiantao, Chi Zhang, Ying Qian, Conghui He. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang, Tianyi Bai, Xingjian Wei, Jiantao Qiu, Conghui He
ACL (1)5
2025 LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
abstract
With the rapid development of AI-generated content, the future internet may be inundated with synthetic data, making the discrimination of authentic and credible multimodal data increasingly challenging. Synthetic data detection has thus garnered widespread attention, and the performance of large multimodal models (LMMs) in this task has attracted significant interest. LMMs can provide natural language explanations for their authenticity judgments, enhancing the explainability of synthetic content detection. Simultaneously, the task of distinguishing between real and synthetic data effectively tests the perception, knowledge, and reasoning capabilities of LMMs. In response, we introduce LOKI, a novel benchmark designed to evaluate the ability of LMMs to detect synthetic data across multiple modalities. LOKI encompasses video, image, 3D, text, and audio modalities, comprising 18K carefully curated questions across 26 subcategories with clear difficulty levels. The benchmark includes coarse-grained judgment and multiple-choice questions, as well as fine-grained anomaly selection and explanation tasks, allowing for a comprehensive analysis of LMMs. We evaluated 22 open-source LMMs and 6 closed-source models on LOKI, highlighting their potential as synthetic data detectors and also revealing some limitations in the development of LMM capabilities. More information about LOKI can be found at https://opendatalab.github.io/LOKI/.
Junyan Ye, Baichuan Zhou, Junan Zhang, Tianyi Bai, Hengrui Kang, Honglin Lin, Zhizheng Wu 0001, Dahua Lin, Conghui He
ICLR5
2025 Harnessing Diversity for Important Data Selection in Pretraining Large Language Models
abstract
Data selection is of great significance in pretraining large language models, given the variation in quality within the large-scale available training corpora. To achieve this, researchers are currently investigating the use of data influence to measure the importance of data instances, $i.e.,$ a high influence score indicates that incorporating this instance to the training set is likely to enhance the model performance. Consequently, they select the top-$k$ instances with the highest scores. However, this approach has several limitations. (1) Calculating the accurate influence of all available data is time-consuming. (2) The selected data instances are not diverse enough, which may hinder the pretrained model's ability to generalize effectively to various downstream tasks. In this paper, we introduce $\texttt{Quad}$, a data selection approach that considers both quality and diversity by using data influence to achieve state-of-the-art pretraining results. To compute the influence ($i.e.,$ the quality) more accurately and efficiently, we incorporate the attention layers to capture more semantic details, which can be accelerated through the Kronecker product. For the diversity, $\texttt{Quad}$ clusters the dataset into similar data instances within each cluster and diverse instances across different clusters. For each cluster, if we opt to select data from it, we take some samples to evaluate the influence to prevent processing all instances. Overall, we favor clusters with highly influential instances (ensuring high quality) or clusters that have been selected less frequently (ensuring diversity), thereby well balancing between quality and diversity. Experiments on Slimpajama and FineWeb over 7B large language models demonstrate that $\texttt{Quad}$ significantly outperforms other data selection methods with a low FLOPs consumption. Further analysis also validates the effectiveness of our influence calculation.
Chi Zhang 0102, Huaping Zhong, Chengliang Chai, Rui Wang 0119, Xinlin Zhuang, Tianyi Bai, Jiantao Qiu, Lei Cao 0004, Ju Fan, Ye Yuan 0001, Guoren Wang, Conghui He
ICLR7
2025 Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
abstract
Multimodal large language models (MLLMs) have achieved strong performance on vision-language tasks but still struggle with fine-grained visual differences, leading to hallucinations or missed semantic shifts. We attribute this to limitations in both training data and learning objectives. To address these issues, we propose a controlled data generation pipeline that produces minimally edited image pairs with semantically aligned captions. Using this pipeline, we construct the Micro Edit Dataset (MED), containing over 50K image-text pairs spanning 11 fine-grained edit categories, including attribute, count, position, and object presence changes. Building on MED, we introduce a supervised fine-tuning (SFT) framework with a feature-level consistency loss that promotes stable visual embeddings under small edits. We evaluate our approach on the Micro Edit Detection benchmark, which includes carefully balanced evaluation pairs designed to test sensitivity to subtle visual variations across the same edit categories. Our method improves difference detection accuracy and reduces hallucinations compared to strong baselines, including GPT-4o. Moreover, it yields consistent gains on standard vision-language tasks such as image captioning and visual question answering. These results demonstrate the effectiveness of combining targeted data and alignment objectives for enhancing fine-grained visual reasoning in MLLMs. Code and datasets are publicly released at https://github.com/Relaxed-System-Lab/hallu_med.
Tianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun, Junlin Han, Conghui He, Wentao Zhang 0001, Binhang Yuan
NeurIPS1
2025 Multi-step Visual Reasoning with Visual Tokens Scaling and Verification
abstract
Multi-modal large language models (MLLMs) have achieved remarkable capabilities by integrating visual perception with language understanding, enabling applications such as image-grounded dialogue, visual question answering, and scientific analysis. However, most MLLMs adopt a static inference paradigm, encoding the entire image into fixed visual tokens upfront, which limits their ability to iteratively refine understanding or adapt to context during inference. This contrasts sharply with human perception, which is dynamic, selective, and feedback-driven. In this work, we introduce a novel framework for inference-time visual token scaling that enables MLLMs to perform iterative, verifier-guided reasoning over visual content. We formulate the problem as a Markov Decision Process, involving a reasoner that proposes visual actions and a verifier—trained via multi-step Direct Preference Optimization (DPO)—that evaluates these actions and determines when reasoning should terminate. To support this, we present a new dataset, VTS, comprising supervised reasoning trajectories (VTS-SFT) and preference-labeled reasoning comparisons (VTS-DPO). Our method significantly outperforms existing approaches across diverse visual reasoning benchmarks, offering not only improved accuracy but also more interpretable and grounded reasoning processes. These results demonstrate the promise of dynamic inference mechanisms for enabling fine-grained, context-aware visual reasoning in next-generation MLLMs. Code and datasets are publicly released at https://vts-v.github.io/.
Tianyi Bai, Zengjie Hu, Fupeng Sun, Jiantao Qiu, Yizhen Jiang, Guangxin He, Bohan Zeng, Conghui He, Binhang Yuan, Wentao Zhang 0001
NeurIPS1
2022 Transfer Learning based Search Space Design for Hyperparameter Tuning
abstract
The tuning of hyperparameters becomes increasingly important as machine learning (ML) models have been extensively applied in data mining applications. Among various approaches, Bayesian optimization (BO) is a successful methodology to tune hyperparameters automatically. While traditional methods optimize each tuning task in isolation, there has been recent interest in speeding up BO by transferring knowledge across previous tasks. In this work, we introduce an automatic method to design the BO search space with the aid of tuning history from past tasks. This simple yet effective approach can be used to endow many existing BO methods with transfer learning capabilities. In addition, it enjoys the three advantages: universality, generality, and safeness. The extensive experiments show that our approach considerably boosts BO by designing a promising and compact search space instead of using the entire space, and outperforms the state-of-the-arts on a wide range of benchmarks, including machine learning and deep learning tuning tasks, and neural architecture search.
Yang Li 0106, Yu Shen 0003, Huaijun Jiang, Tianyi Bai, Wentao Zhang 0001, Ce Zhang 0001, Bin Cui 0001
KDD4