Zhi Rui Tam

dblp:279/1685 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-9968-2416ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 63% Information extraction and text analysis · 16% Generative modeling · 11%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
LLM agents
0.812024
StreamBench: Towards Benchmarking Continuous Improvement of Language Agents · NeurIPS 2024
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.812024
I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL Generation · EMNLP 2024
Machine learning and data management
online learning
0.812024
StreamBench: Towards Benchmarking Continuous Improvement of Language Agents · NeurIPS 2024
Machine learning and data management › model evaluation
streaming evaluation
0.812024
StreamBench: Towards Benchmarking Continuous Improvement of Language Agents · NeurIPS 2024
Natural language and speech › Language models and text generation
alignment
0.712023
OpenAssistant Conversations - Democratizing Large Language Model Alignment · NeurIPS 2023
Natural language and speech › Language models and text generation › alignment › preference alignment
human feedback alignment
0.712023
OpenAssistant Conversations - Democratizing Large Language Model Alignment · NeurIPS 2023
Natural language and speech › Language models and text generation
instruction tuning
0.712023
OpenAssistant Conversations - Democratizing Large Language Model Alignment · NeurIPS 2023
Machine learning › Generative modeling
generative adversarial network
0.512021
Gradient Normalization for Generative Adversarial Networks · ICCV 2021
Machine learning › Deep learning architectures and training › normalization
gradient normalization
0.512021
Gradient Normalization for Generative Adversarial Networks · ICCV 2021
Visual content generation and editing › multimodal content generation
story visualization
0.412020
Character-Preserving Coherent Story Visualization · ECCV (17) 2020
Natural language and speech › Language models and text generation
large language model
0.212024
I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL Generation · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

online learning · 2.3feedback stream · 2.3evaluation metrics · 1.5supervised fine-tuning · 0.7reinforcement learning from human feedback · 0.7spectral normalization · 0.5lipschitz constraint · 0.5gradient penalty · 0.5generative model · 0.4coherence modeling · 0.4
YearPublicationVenuePosition
2025 Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token Learning
abstract
Maintaining consistent model performance across domains is a fundamental challenge in machine learning. While recent work has explored using LLM-generated data for fine-tuning, its impact on cross-domain generalization remains poorly understood. This paper presents a systematic analysis revealing that fine-tuning with LLM-generated data not only improves target task performance but also reduces non-target task degradation compared to fine-tuning with ground truth data. Through analyzing the data sequence in tasks of various domains, we demonstrate that this enhancement of non-target task robustness stems from the reduction of high perplexity tokens found in LLM-generated sequences. Following our findings, we showed that masking high perplexity tokens in ground truth training data achieves similar non-target task performance preservation, comparable to using LLM-generated data. Extensive experiments across different model families and scales, including Gemma 2 IT 2B, Llama 3 8B Instruct, and three additional models, agree with our findings. To the best of our knowledge, this is the first work to provide an empirical explanation based on token perplexity reduction to mitigate catastrophic forgetting in LLMs after fine-tuning, offering valuable insights for developing more robust fine-tuning strategies.
Chao-Chung Wu, Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Shao-Hua Sun, Hung-yi Lee
NeurIPS2
2024 I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL Generation
abstract
This study explores the proactive ability of LLMs to seek user support.We propose metrics to evaluate the trade-off between performance improvements and user burden, and investigate whether LLMs can determine when to request help under varying information availability.Our experiments show that without external feedback, many LLMs struggle to recognize their need for user support.The findings highlight the importance of external signals and provide insights for future research on improving support-seeking strategies.
Cheng-Kuang Wu, Zhi Rui Tam, Chao-Chung Wu, Chieh-Yen Lin, Hung-yi Lee, Yun-Nung Chen
EMNLP2
2024 StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
abstract
Recent works have shown that large language model (LLM) agents are able to improve themselves from experience, which is an important ability for continuous enhancement post-deployment. However, existing benchmarks primarily evaluate their innate capabilities and do not assess their ability to improve over time. To address this gap, we introduce StreamBench, a pioneering benchmark designed to evaluate the continuous improvement of LLM agents over an input-feedback sequence. StreamBench simulates an online learning environment where LLMs receive a continuous flow of feedback stream and iteratively enhance their performance. In addition, we propose several simple yet effective baselines for improving LLMs on StreamBench, and provide a comprehensive analysis to identify critical components that contribute to successful streaming strategies. Our work serves as a stepping stone towards developing effective online learning strategies for LLMs, paving the way for more adaptive AI systems in streaming scenarios.
Cheng-Kuang Wu, Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Hung-yi Lee
NeurIPS2
2024 Personalized EDM Subject Generation via Co-factored User-Subject Embedding
Yu-Hsiu Chen, Zhi Rui Tam, Hong-Han Shuai
PAKDD (2)2
2023 OpenAssistant Conversations - Democratizing Large Language Model Alignment
abstract
Aligning large language models (LLMs) with human preferences has proven to drastically improve usability and has driven rapid adoption as demonstrated by ChatGPT.Alignment techniques such as supervised fine-tuning (\textit{SFT}) and reinforcement learning from human feedback (\textit{RLHF}) greatly reduce the required skill and domain knowledge to effectively harness the capabilities of LLMs, increasing their accessibility and utility across various domains.However, state-of-the-art alignment techniques like \textit{RLHF} rely on high-quality human feedback data, which is expensive to create and often remains proprietary.In an effort to democratize research on large-scale alignment, we release OpenAssistant Conversations, a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292 quality ratings, resulting in over 10,000 complete and fully annotated conversation trees.The corpus is a product of a worldwide crowd-sourcing effort involving over 13,500 volunteers.Models trained on OpenAssistant Conversations show consistent improvements on standard benchmarks over respective base models.We release our code\footnote{\git} and data\footnote{\data} under a fully permissive licence.
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, Alexander Mattick
NeurIPS5
2022 Improving Entity Disambiguation Using Knowledge Graph Regularization
Zhi Rui Tam, Yi-Lun Wu, Hong-Han Shuai
PAKDD (1)1
2021 Gradient Normalization for Generative Adversarial Networks
abstract
In this paper, we propose a novel normalization method called gradient normalization (GN) to tackle the training instability of Generative Adversarial Networks (GANs) caused by the sharp gradient space. Unlike existing work such as gradient penalty and spectral normalization, the proposed GN only imposes a hard 1-Lipschitz constraint on the discriminator function, which increases the capacity of the discriminator. Moreover, the proposed gradient normalization can be applied to different GAN architectures with little modification. Extensive experiments on four datasets show that GANs trained with gradient normalization outperform existing methods in terms of both Frechet Inception Distance and Inception Score.
Yi-Lun Wu, Hong-Han Shuai, Zhi Rui Tam, Hong-Yu Chiu
ICCV3
2020 Character-Preserving Coherent Story Visualization
Yun-Zhu Song, Zhi Rui Tam, Hung-Jen Chen 0001, Huiao-Han Lu, Hong-Han Shuai
ECCV (17)2