Jiarui Jiang

dblp:327/4558 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0001-3950-4746ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
1 paper
Design research and methods · 67% User interface design and tools · 33%
Artificial intelligence
2 papers
Deep learning architectures and training · 54% Learning theory · 35% Language models and text generation · 11%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
User interface design and tools
design tools
1.012026
All Futures at Once: Supporting Speculative Design for Placemaking with Multi-Agent Social Simulation · CHI 2026
Design research and methods › design practice
placemaking
1.012026
All Futures at Once: Supporting Speculative Design for Placemaking with Multi-Agent Social Simulation · CHI 2026
Design research and methods
speculative design
1.012026
All Futures at Once: Supporting Speculative Design for Placemaking with Multi-Agent Social Simulation · CHI 2026
Machine learning › Learning theory › overfitting
benign overfitting
0.812024
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization · NeurIPS 2024
Machine learning › Deep learning architectures and training
training dynamics
0.812024
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization · NeurIPS 2024
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.812024
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization · NeurIPS 2024
Natural language and speech › Language models and text generation › LLM agents
LLM-based simulation
0.312026
All Futures at Once: Supporting Speculative Design for Placemaking with Multi-Agent Social Simulation · CHI 2026
Machine learning › Learning theory
generalization bounds
0.212024
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

user study · 2.0multi-agent simulation · 2.0signal-to-noise ratio analysis · 0.8gradient descent · 0.8
YearPublicationVenuePosition
2026 All Futures at Once: Supporting Speculative Design for Placemaking with Multi-Agent Social Simulation
abstract
Placemaking transforms physical spaces into socially meaningful places, with long-term impacts depending on how future communities inhabit and interact with them. Speculative design helps envision such futures, yet existing approaches often produce static representations that emphasize spatial form over evolving activity. We present ParaScape, a design support system that facilitates speculative design for placemaking by generating dynamic speculative objects through an underlying LLM-based multi-agent social simulation framework. The framework models heterogeneous agents with group-specific preferences and sensitivities, simulating context-sensitive behaviors and interactions that produce evolving scenarios. These scenarios are visualized as image sequences, where each scenario depicts multiple activities unfolding within a place at a given moment. ParaScape builds on this framework to allow designers to explore scenarios, analyze activity diversity and evolvability, and reflect on trade-offs among stakeholder needs. Evaluations through two experiments, a user study, and two case studies show that ParaScape supports critical reasoning and inclusive placemaking.
Jiarui Jiang, Shuqing Tang, Mutao Yu, Caoyang Xue, Yunsheng Su, Xinyang Tan, Yang Shi 0007
CHI2
2025 Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
abstract
State-space models (SSMs), particularly Mamba, emerge as an efficient Transformer alternative with linear complexity for long-sequence modeling. Recent empirical works demonstrate Mamba's in-context learning (ICL) capabilities competitive with Transformers, a critical capacity for large foundation models. However, theoretical understanding of Mamba’s ICL remains limited, restricting deeper insights into its underlying mechanisms. Even fundamental tasks such as linear regression ICL, widely studied as a standard theoretical benchmark for Transformers, have not been thoroughly analyzed in the context of Mamba. To address this gap, we study the training dynamics of Mamba on the linear regression ICL task. By developing novel techniques tackling non-convex optimization with gradient descent related to Mamba's structure, we establish an exponential convergence rate to ICL solution, and derive a loss bound that is comparable to Transformer's. Importantly, our results reveal that Mamba can perform a variant of \textit{online gradient descent} to learn the latent function in context. This mechanism is different from that of Transformer, which is typically understood to achieve ICL through gradient descent emulation. The theoretical results are verified by experimental simulation.
Jiarui Jiang, Wei Huang 0034, Miao Zhang 0022, Taiji Suzuki, Liqiang Nie
NeurIPS1
2024 Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
abstract
Transformers have demonstrated great power in the recent development of large foundational models. In particular, the Vision Transformer (ViT) has brought revolutionary changes to the field of vision, achieving significant accomplishments on the experimental side. However, their theoretical capabilities, particularly in terms of generalization when trained to overfit training data, are still not fully understood. To address this gap, this work delves deeply into the \textit{benign overfitting} perspective of transformers in vision. To this end, we study the optimization of a Transformer composed of a self-attention layer with softmax followed by a fully connected layer under gradient descent on a certain data distribution model. By developing techniques that address the challenges posed by softmax and the interdependent nature of multiple weights in transformer optimization, we successfully characterized the training dynamics and achieved generalization in post-training. Our results establish a sharp condition that can distinguish between the small test error phase and the large test error regime, based on the signal-to-noise ratio in the data model. The theoretical results are further verified by experimental simulation. To the best of our knowledge, this is the first work to characterize benign overfitting for Transformers.
Jiarui Jiang, Wei Huang 0034, Miao Zhang 0022, Taiji Suzuki, Liqiang Nie
NeurIPS1