Yingtian Tang

dblp:295/0111 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0004-9870-5574ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 47% Trustworthy machine learning · 21% Knowledge representation and reasoning · 21%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › language model interpretability
brain alignment
0.912025
From Language to Cognition: How LLMs Outgrow the Human Language Network · EMNLP 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
cognitive modeling
0.912025
From Language to Cognition: How LLMs Outgrow the Human Language Network · EMNLP 2025
Computer vision › Vision and language › vision-language model
vision-language model analysis
0.712023
When are Lemons Purple? The Concept Association Bias of Vision-Language Models · EMNLP 2023
Computer vision › Vision and language
visual question answering
0.712023
When are Lemons Purple? The Concept Association Bias of Vision-Language Models · EMNLP 2023
Computer vision › Vision and language › visual question answering
zero-shot visual question answering
0.712023
When are Lemons Purple? The Concept Association Bias of Vision-Language Models · EMNLP 2023
Machine learning › Reinforcement learning
deep reinforcement learning
0.512021
Learning-Aided Heuristics Design for Storage System · SIGMOD Conference 2021
Storage systems › storage management
storage allocation
0.512021
Learning-Aided Heuristics Design for Storage System · SIGMOD Conference 2021

Methods — techniques the papers use, named apart from their topics

learning-aided heuristic design · 1.0deep reinforcement learning · 1.0representational similarity analysis · 0.9benchmarking · 0.9contrastive learning analysis · 0.7autoregressive loss analysis · 0.7
YearPublicationVenuePosition
2025 From Language to Cognition: How LLMs Outgrow the Human Language Network
abstract
Large language models (LLMs) exhibit remarkable similarity to neural activity in the human language network. However, the key properties of language underlying this alignment—and how brain-like representations emerge and change across training—remain unclear. We here benchmark 34 training checkpoints spanning 300B tokens across 8 different model sizes to analyze how brain alignment relates to linguistic competence. Specifically, we find that brain alignment tracks the development of formal linguistic competence—i.e., knowledge of linguistic rules—more closely than functional linguistic competence. While functional competence, which involves world knowledge and reasoning, continues to develop throughout training, its relationship with brain alignment is weaker, suggesting that the human language network primarily encodes formal linguistic structure rather than broader cognitive functions. Notably, we find that the correlation between next-word prediction, behavioral alignment, and brain alignment fades once models surpass human language proficiency. We further show that model size is not a reliable predictor of brain alignment when controlling for the number of features. Finally, using the largest set of rigorous neural language benchmarks to date, we show that language brain alignment benchmarks remain unsaturated, highlighting opportunities for improving future models. Taken together, our findings suggest that the human language network is best modeled by formal, rather than functional, aspects of language.
Badr AlKhamissi, Greta Tuckute, Yingtian Tang, Taha Binhuraib, Antoine Bosselut, Martin Schrimpf
EMNLP3
2023 When are Lemons Purple? The Concept Association Bias of Vision-Language Models
abstract
Large-scale vision-language models such as CLIP have shown impressive performance on zero-shot image classification and image-totext retrieval.However, such performance does not realize in tasks that require a finergrained correspondence between vision and language, such as Visual Question Answering (VQA).As a potential cause of the difficulty of applying these models to VQA and similar tasks, we report an interesting phenomenon of vision-language models, which we call the Concept Association Bias (CAB).We find that models with CAB tend to treat input as a bag of concepts and attempt to fill in the other missing concept crossmodally, leading to an unexpected zero-shot prediction.We demonstrate CAB by showing that CLIP's zeroshot classification performance greatly suffers when there is a strong concept association between an object (e.g.eggplant) and an attribute (e.g.color purple).We also show that the strength of CAB predicts the performance on VQA.We observe that CAB is prevalent in vision-language models trained with contrastive losses, even when autoregressive losses are jointly employed.However, a model that solely relies on autoregressive loss seems to exhibit minimal or no signs of CAB. * Equal contribution.CLIP: "In this picture, the color of the lemon is purple."
Yingtian Tang, Yutaro Yamada, Yoyo Zhang, Ilker Yildirim
EMNLP1
2022 Accurate Probabilistic Miss Ratio Curve Approximation for Adaptive Cache Allocation in Block Storage Systems
abstract
Cache plays an important role in storage systems. With better allocation of cache space to each storage device, total I/O latency can be reduced remarkably. To achieve this goal, we propose an Accurate Probabilistic miss ratio curve approximation for Adaptive Cache allocation (APAC) system. APAC can obtain near-optimal performance for allocating cache space with low overhead. Specifically, with a linear-time probabilistic approximation of reuse distance of all blocks inside each device, APAC can accurately estimate the miss ratio curve (MRC). Furthermore, APAC utilizes the MRCs to obtain the near-optimal configuration of cache allocation by dynamic programming. Experimental results show that APAC achieves higher accuracy in MRC approximation compared to the state-of-the-art methods, leading to higher hit ratio and lower latency of the block storage systems.
Rongshang Li, Yingtian Tang, Qiquan Shi, Lei Chen 0031, Jikun Jin
DATE2
2021 Learning-Aided Heuristics Design for Storage System
abstract
Computer systems such as storage systems normally require transparent white-box algorithms that are interpretable for human experts. In this work, we propose a learning-aided heuristic design method, which automatically generates human-readable strategies from Deep Reinforcement Learning (DRL) agents. This method benefits from the power of deep learning but avoids the shortcoming of its black-box property. Besides the white-box advantage, experiments in our storage production's resource allocation scenario also show that this solution outperforms the system's default settings and the elaborately handcrafted strategy by human experts.
Yingtian Tang, Han Lu 0004, Xijun Li, Lei Chen 0031, Mingxuan Yuan
SIGMOD Conference1