VLDB 2026 Research / reviewers in the wild / expert
Yingshan Chang
dblp:301/8296
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 33% Language models and text generation · 14% Trustworthy machine learning · 13% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
inductive biases |
0.9 | 1 | 2025 | Language Models Need Inductive Biases to Count Inductively · ICLR 2025 |
Natural language and speech › Language models and text generation › compositional generalization
length generalization |
0.9 | 1 | 2025 | Language Models Need Inductive Biases to Count Inductively · ICLR 2025 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.9 | 1 | 2025 | Language Models Need Inductive Biases to Count Inductively · ICLR 2025 |
Machine learning › Trustworthy machine learning › fairness
bias evaluation |
0.8 | 1 | 2024 | Diffusion PID: Interpreting Diffusion via Partial Information Decomposition · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Diffusion PID: Interpreting Diffusion via Partial Information Decomposition · NeurIPS 2024 |
Machine learning › Generative modeling › normalizing flow
flow-based prior |
0.8 | 1 | 2024 | Flow Priors for Linear Inverse Problems via Iterative Corrupted Trajectory Matching · NeurIPS 2024 |
Machine learning › Generative modeling
flow matching |
0.8 | 1 | 2024 | Flow Priors for Linear Inverse Problems via Iterative Corrupted Trajectory Matching · NeurIPS 2024 |
Machine learning › Learning theory
generalization |
0.8 | 1 | 2024 | Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation · ECCV (87) 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Diffusion PID: Interpreting Diffusion via Partial Information Decomposition · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.8 | 1 | 2024 | Diffusion PID: Interpreting Diffusion via Partial Information Decomposition · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation · ECCV (87) 2024 |
Natural language and speech › Language models and text generation › LLM agents
tool use |
0.8 | 1 | 2024 | Tools Fail: Detecting Silent Errors in Faulty Tools · EMNLP 2024 |
Image and video processing › image restoration
inverse problem |
0.8 | 1 | 2024 | Flow Priors for Linear Inverse Problems via Iterative Corrupted Trajectory Matching · NeurIPS 2024 |
Image and video processing › image restoration › inverse problem
linear inverse problem |
0.8 | 1 | 2024 | Flow Priors for Linear Inverse Problems via Iterative Corrupted Trajectory Matching · NeurIPS 2024 |
Natural language and speech › Question answering and dialogue systems
multimodal question answering |
0.6 | 1 | 2022 | WebQA: Multihop and Multimodal QA · CVPR 2022 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › agent planning
embodied planning |
0.2 | 1 | 2024 | Tools Fail: Detecting Silent Errors in Faulty Tools · EMNLP 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
MAP inference |
0.2 | 1 | 2024 | Flow Priors for Linear Inverse Problems via Iterative Corrupted Trajectory Matching · NeurIPS 2024 |
Information theory › information measures › information decomposition
partial information decomposition |
0.2 | 1 | 2024 | Diffusion PID: Interpreting Diffusion via Partial Information Decomposition · NeurIPS 2024 |
Computer vision › Vision and language
visual question answering |
0.2 | 1 | 2022 | WebQA: Multihop and Multimodal QA · CVPR 2022 |
Information retrieval
multimodal retrieval |
0.2 | 1 | 2022 | WebQA: Multihop and Multimodal QA · CVPR 2022 |
Information retrieval
web search |
0.2 | 1 | 2022 | WebQA: Multihop and Multimodal QA · CVPR 2022 |
Methods — techniques the papers use, named apart from their topics
tweedie's formula · 1.5trajectory matching · 1.5partial information decomposition · 1.5information-theoretic analysis · 1.5flow matching · 1.5knowledge aggregation · 1.1state space model · 0.9positional embedding · 0.9RWKV · 0.9large language model · 0.8multimodal reasoning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Language Models Need Inductive Biases to Count InductivelyabstractCounting constitutes a core skill underlying a wide range of tasks, such as formal language recognition, multi-hop reasoning and simulating algorithms. Generaliz- ing counting inductively is central to task success on out-of-distribution (OOD) instances where testing inputs are longer than those seen in training. While there is a large body of literature reporting poor length generalization in language models, few papers have tried to distill the “reasoning” failure to the simplest case of count- ing failure. We aim to provide a broader picture on whether various language model architectures can a) learn to count, and b) generalize counting inductively. This work provides extensive empirical results on architectures ranging from RNNs, Transformers, State-Space Models and RWKV. We present carefully-designed task formats, auxiliary tasks and positional embeddings to avoid limitations in general- ization with OOD-position and OOD-vocabulary. We find that while traditional RNNs trivially achieve inductive counting, Transformers have to rely on positional embeddings (PEs) to count OOD. Further analyses on interpreting the learned solution reveal that different PEs encode different inductive biases that facilitate counting in different task formats. As counting is the basis for many arguments concerning the expressivity of Transformers, our finding calls for the community to reexamine the application scope of primitive functions defined in formal charac- terizations. Finally, modern RNNs also largely underperform traditional RNNs in generalizing counting inductively, hinting at the tradeoff modern RNNs struggle to balance between parallelized training and maintaining their recurrent nature. Yingshan Chang, Yonatan Bisk |
ICLR | 1 |
| 2024 | Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation
Yingshan Chang, Yasi Zhang, Zhiyuan Fang, Ying Nian Wu, Yonatan Bisk, Feng Gao 0013 |
ECCV (87) | 1 |
| 2024 | Tools Fail: Detecting Silent Errors in Faulty ToolsabstractTools have become a mainstay of LLMs, allowing them to retrieve knowledge not in their weights, to perform tasks on the web, and even to control robots.However, most ontologies and surveys of tool-use have assumed the core challenge for LLMs is choosing the tool.Instead, we introduce a framework for tools more broadly which guides us to explore a model's ability to detect "silent" tool errors, and reflect on how to plan.This more directly aligns with the increasingly popular use of models as tools.We provide an initial approach to failure recovery with promising results both on a controlled calculator setting and embodied agent planning.1 Jimin Sun, So Yeon Min, Yingshan Chang, Yonatan Bisk |
EMNLP | 3 |
| 2024 | Diffusion PID: Interpreting Diffusion via Partial Information DecompositionabstractText-to-image diffusion models have made significant progress in generating naturalistic images from textual inputs, and demonstrate the capacity to learn and represent complex visual-semantic relationships. While these diffusion models have achieved remarkable success, the underlying mechanisms driving their performance are not yet fully accounted for, with many unanswered questions surrounding what they learn, how they represent visual-semantic relationships, and why they sometimes fail to generalize. Our work presents Diffusion Partial Information Decomposition (DiffusionPID), a novel technique that applies information-theoretic principles to decompose the input text prompt into its elementary components, enabling a detailed examination of how individual tokens and their interactions shape the generated image. We introduce a formal approach to analyze the uniqueness, redundancy, and synergy terms by applying PID to the denoising model at both the image and pixel level. This approach enables us to characterize how individual tokens and their interactions affect the model output. We first present a fine-grained analysis of characteristics utilized by the model to uniquely localize specific concepts, we then apply our approach in bias analysis and show it can recover gender and ethnicity biases. Finally, we use our method to visually characterize word ambiguity and similarity from the model’s perspective and illustrate the efficacy of our method for prompt intervention. Our results show that PID is a potent tool for evaluating and diagnosing text-to-image diffusion models. Link to project page: https://rbz-99.github.io/Diffusion-PID/. Shaurya Dewan, Rushikesh Zawar, Prakanshul Saxena, Yingshan Chang, Andrew Luo 0001, Yonatan Bisk |
NeurIPS | 4 |
| 2024 | Flow Priors for Linear Inverse Problems via Iterative Corrupted Trajectory MatchingabstractGenerative models based on flow matching have attracted significant attention for their simplicity and superior performance in high-resolution image synthesis. By leveraging the instantaneous change-of-variables formula, one can directly compute image likelihoods from a learned flow, making them enticing candidates as priors for downstream tasks such as inverse problems. In particular, a natural approach would be to incorporate such image probabilities in a maximum-a-posteriori (MAP) estimation problem. A major obstacle, however, lies in the slow computation of the log-likelihood, as it requires backpropagating through an ODE solver, which can be prohibitively slow for high-dimensional problems. In this work, we propose an iterative algorithm to approximate the MAP estimator efficiently to solve a variety of linear inverse problems. Our algorithm is mathematically justified by the observation that the MAP objective can be approximated by a sum of $N$ ``local MAP'' objectives, where $N$ is the number of function evaluations. By leveraging Tweedie's formula, we show that we can perform gradient steps to sequentially optimize these objectives. We validate our approach for various linear inverse problems, such as super-resolution, deblurring, inpainting, and compressed sensing, and demonstrate that we can outperform other methods based on flow matching. Code is available at \url{https://github.com/YasminZhang/ICTM}. Yasi Zhang, Peiyu Yu, Yaxuan Zhu, Yingshan Chang, Feng Gao 0013, Ying Nian Wu, Oscar Leong |
NeurIPS | 4 |
| 2022 | WebQA: Multihop and Multimodal QAabstractScaling Visual Question Answering (VQA) to the open-domain and multi-hop nature of web searches, requires fundamental advances in visual representation learning, knowledge aggregation, and language generation. In this work, we introduce WEBQA, a challenging new benchmark that proves difficult for large-scale state-of-the-art models which lack language groundable visual representations for novel objects and the ability to reason, yet trivial for humans. WebQA mirrors the way humans use the web: 1) Ask a question, 2) Choose sources to aggregate, and 3) Produce a fluent language response. This is the behavior we should be expecting from IoT devices and digital assistants. Existing work prefers to assume that a model can either reason about knowledge in images or in text. WebQA includes a secondary text-only QA task to ensure improved visual performance does not come at the cost of language understanding. Our challenge for the community is to create unified multimodal reasoning models that answer questions regardless of the source modality, moving us closer to digital assistants that not only query language knowledge, but also the richer visual online world. Yingshan Chang, Guihong Cao, Mridu Narang, Jianfeng Gao 0001, Hisami Suzuki, Yonatan Bisk |
CVPR | 1 |