Shuaichen Chang

dblp:230/4596 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Deep learning architectures and training · 26% Information extraction and text analysis · 24% Multi-agent systems · 19%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 78% Data models and query languages · 22%
Theoretical computer science
1 paper
Mathematical optimization · 50% Approximation and online algorithms · 50%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
1.122023
Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness · ICLR 2023
Zero-Shot Text-to-SQL Learning with Auxiliary Task · AAAI 2020
Mathematical optimization
knapsack problem
0.912025
Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection · NeurIPS 2025
Approximation and online algorithms › online algorithms › online packing and covering
online knapsack
0.912025
Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection · NeurIPS 2025
Information retrieval
retrieval-augmented generation
0.812024
RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation · NeurIPS 2024
Information retrieval › evaluation › text generation evaluation
retrieval-augmented generation evaluation
0.812024
RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation · NeurIPS 2024
Information retrieval
retrieval evaluation
0.812024
RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation · NeurIPS 2024
Machine learning › Trustworthy machine learning
robustness evaluation
0.712023
Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness · ICLR 2023
Data models and query languages › natural language interface
natural language interface to database
0.712023
Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness · ICLR 2023
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
sequence-to-sequence generation
0.412020
Zero-Shot Text-to-SQL Learning with Auxiliary Task · AAAI 2020
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.412020
Zero-Shot Text-to-SQL Learning with Auxiliary Task · AAAI 2020
Machine learning › Graph learning
graph neural network
0.412019
Contextualized Non-Local Neural Networks for Sequence Learning · AAAI 2019
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.412019
Contextualized Non-Local Neural Networks for Sequence Learning · AAAI 2019
Machine learning › Deep learning architectures and training
sequence modeling
0.412019
Contextualized Non-Local Neural Networks for Sequence Learning · AAAI 2019

Methods — techniques the papers use, named apart from their topics

utility modeling · 1.7retrieval baseline · 1.7online knapsack · 1.7diagnostic benchmark construction · 1.3meta-evaluation · 0.8seq2seq · 0.4multi-task learning · 0.4self-attention · 0.4graph neural network · 0.4
YearPublicationVenuePosition
2025 PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries
abstract
Mingwen Dong, Nischal Ashok Kumar, Yiqun Hu, Anuj Chauhan, Chung-Wei Hang, Shuaichen Chang, Lin Pan, Wuwei Lan, Henghui Zhu, Jiarong Jiang, Patrick Ng, Zhiguo Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Mingwen Dong, Nischal Ashok Kumar, Yiqun Hu, Anuj Chauhan, Chung-Wei Hang, Shuaichen Chang, Lin Pan 0003, Wuwei Lan, Henghui Zhu, Jiarong Jiang, Patrick Ng, Zhiguo Wang 0006
NAACL (Long Papers)6
2025 You Only Read Once (YORO): Learning to Internalize Database Knowledge for Text-to-SQL
abstract
Hideo Kobayashi, Wuwei Lan, Peng Shi, Shuaichen Chang, Jiang Guo, Henghui Zhu, Zhiguo Wang, Patrick Ng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hideo Kobayashi, Wuwei Lan, Peng Shi 0010, Shuaichen Chang, Henghui Zhu, Zhiguo Wang 0006, Patrick Ng
NAACL (Long Papers)4
2025 Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection
abstract
Designing effective agentic systems requires the seamless composition and integration of agents, tools, and models within dynamic and uncertain environments. Most existing methods rely on static, semantic retrieval approaches for tool or agent discovery. However, effective reuse and composition of existing components remain challenging due to incomplete capability descriptions and the limitations of retrieval methods. Component selection suffers because the decisions are not based on capability, cost, and real-time utility. To address these challenges, we introduce a structured, automated framework for agentic system composition that is inspired by the knapsack problem. Our framework enables a composer agent to systematically identify, select, and assemble an optimal set of agentic components by jointly considering performance, budget constraints, and compatibility. By dynamically testing candidate components and modeling their utility in real-time, our approach streamlines the assembly of agentic systems and facilitates scalable reuse of resources. Empirical evaluation with Claude 3.5 Sonnet across five benchmarking datasets shows that our online-knapsack-based composer consistently lies on the Pareto frontier, achieving higher success rates at significantly lower component costs compared to our baselines. In the single-agent setup, the online knapsack composer shows a success rate improvement of up to 31.6\% in comparison to the retrieval baselines. In multi-agent systems, the online knapsack composer increases success rate from 37\% to 87\% when agents are selected from an agent inventory of 100+ agents. The substantial performance gap confirms the robust adaptability of our method across diverse domains and budget constraints.
Michelle Yuan, Khushbu Pahwa, Shuaichen Chang, Mustafa Kaba, Jiarong Jiang, Monica Sunkara
NeurIPS3
2024 RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
abstract
Despite Retrieval-Augmented Generation (RAG) has shown promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements. In this paper, we propose a fine-grained evaluation framework, RAGChecker, that incorporates a suite of diagnostic metrics for both the retrieval and generation modules. Meta evaluation verifies that RAGChecker has significantly better correlations with human judgments than other evaluation metrics. Using RAGChecker, we evaluate 8 RAG systems and conduct an in-depth analysis of their performance, revealing insightful patterns and trade-offs in the design choices of RAG architectures. The metrics of RAGChecker can guide researchers and practitioners in developing more effective RAG systems.
Dongyu Ru, Xiangkun Hu, Tianhang Zhang, Peng Shi 0010, Shuaichen Chang, Cheng Jiayang, Cunxiang Wang, Shichao Sun, Huanyu Li 0010, Binjie Wang, Jiarong Jiang, Tong He 0002, Zhiguo Wang 0006, Pengfei Liu 0003, Yue Zhang 0004, Zheng Zhang 0001
NeurIPS6
2023 Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness
Shuaichen Chang, Jun Wang 0122, Mingwen Dong, Lin Pan 0003, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang 0029, Jiarong Jiang, Joe Lilien, Steve Ash, William Yang Wang, Zhiguo Wang 0006, Vittorio Castelli, Patrick Ng, Bing Xiang
ICLR1
2020 Zero-Shot Text-to-SQL Learning with Auxiliary Task
abstract
Recent years have seen great success in the use of neural seq2seq models on the text-to-SQL task. However, little work has paid attention to how these models generalize to realistic unseen data, which naturally raises a question: does this impressive performance signify a perfect generalization model, or are there still some limitations?In this paper, we first diagnose the bottleneck of the text-to-SQL task by providing a new testbed, in which we observe that existing models present poor generalization ability on rarely-seen data. The above analysis encourages us to design a simple but effective auxiliary task, which serves as a supportive model as well as a regularization term to the generation task to increase the models' generalization. Experimentally, We evaluate our models on a large text-to-SQL dataset WikiSQL. Compared to a strong baseline coarse-to-fine model, our models improve over the baseline by more than 3% absolute in accuracy on the whole dataset. More interestingly, on a zero-shot subset test of WikiSQL, our models achieve 5% absolute accuracy gain over the baseline, clearly demonstrating its superior generalizability.
Shuaichen Chang, Yun Tang 0002, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001
AAAI1
2019 Contextualized Non-Local Neural Networks for Sequence Learning
abstract
Recently, a large number of neural mechanisms and models have been proposed for sequence learning, of which selfattention, as exemplified by the Transformer model, and graph neural networks (GNNs) have attracted much attention. In this paper, we propose an approach that combines and draws on the complementary strengths of these two methods. Specifically, we propose contextualized non-local neural networks (CN3), which can both dynamically construct a task-specific structure of a sentence and leverage rich local dependencies within a particular neighbourhood.Experimental results on ten NLP tasks in text classification, semantic matching, and sequence labelling show that our proposed model outperforms competitive baselines and discovers task-specific dependency structures, thus providing better interpretability to users.
Pengfei Liu 0003, Shuaichen Chang, Xuanjing Huang 0001, Jackie Chi Kit Cheung
AAAI2