Kuang-Da Wang

dblp:334/0632 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0004-0846-8254ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Adapting to Evolving Data: Test-Time Expert Aggregation for Imbalanced Tabular Regression
abstract
Many critical web applications, from e-commerce price prediction to user engagement forecasting, rely on regression models trained on tabular data. These models often face a dual challenge: the inherent imbalance in continuous target values and, more critically, the unpredictable distribution shifts that occur when the model is deployed online. While data imbalance in classification is well-studied, its intersection with regression tasks in dynamic, real-world settings is underexplored. Existing methods for imbalanced regression often assume that the test data distribution is known and stable, an assumption that rarely holds true for live web systems and can lead to significant performance degradation. To address this gap, we propose a novel framework featuring two key innovations: (i) a Region-Aware Mixture of Experts that leverages a Gaussian Mixture Model to identify distinct data sub-populations. This allows us to synthesize targeted training data and train specialized experts, each tailored to a specific data region. (ii) a Test-Time Self-Supervised Expert Aggregation mechanism. This is the core of our adaptation strategy, dynamically adjusting the weights of each expert based on the features of incoming test instances. This enables our model to adapt on-the-fly to varying test distributions without costly retraining. We evaluated our method on four real-world tabular regression datasets: house pricing, bike sharing, and age prediction. These tasks are representative of real-world scenarios that inherently involve both target imbalance and dynamic distribution shifts (e.g., temporal or market-driven changes). The results demonstrate that our approach significantly outperforms existing imbalanced regression methods, especially under these shifts, achieving an average MAE improvement of 7.1%.
Yung-Chien Wang, Kuang-Da Wang, Wei-Yao Wang, Wen-Chih Peng
WSDM2
2025 APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning
abstract
Tabular data are fundamental in common machine learning applications, ranging from finance to genomics and healthcare. This paper focuses on tabular regression tasks, a field where deep learning (DL) methods are not consistently superior to machine learning (ML) models due to the challenges posed by irregular target functions inherent in tabular data, causing sensitive label changes with minor variations from features. To address these issues, we propose a novel Arithmetic-Aware Pre-training and Adaptive-Regularized Fine-tuning framework (APAR), which enables the model to fit irregular target function in tabular data while reducing the negative impact of overfitting. In the pre-training phase, APAR introduces an arithmetic-aware pretext objective to capture intricate sample-wise relationships from the perspective of continuous labels. In the fine-tuning phase, a consistency-based adaptive regularization technique is proposed to self-learn appropriate data augmentation. Extensive experiments across 10 datasets demonstrated that APAR outperforms existing GBDT-, supervised NN-, and pretrain-finetune NN-based methods in RMSE (+9.43% ~ 20.37%), and empirically validated the effects of pre-training tasks, including the study of arithmetic operations.
Hong-Wei Wu, Wei-Yao Wang, Kuang-Da Wang, Wen-Chih Peng
AAAI3
2025 Extending Automatic Machine Translation Evaluation to Book-Length Documents
abstract
Kuang-Da Wang, Shuoyang Ding, Chao-Han Huck Yang, Ping-Chun Hsieh, Wen-Chih Peng, Vitaly Lavrukhin, Boris Ginsburg. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Kuang-Da Wang, Shuoyang Ding, Chao-Han Huck Yang, Ping-Chun Hsieh, Wen-Chih Peng, Vitaly Lavrukhin, Boris Ginsburg
EMNLP1
2025 RallyDiffuser: A Representation-Guided Diffusion Model Framework for Strategic Planning in Badminton
Bing-Zhi Ke, Kuang-Da Wang, Wen-Chih Peng
AAMAS2
2025 DDOT: A Derivative-Directed Dual-Decoder Ordinary Differential Equation Transformer for Dynamic System Modeling
Yang Chang, Kuang-Da Wang, Ping-Chun Hsieh, Cheng-Kuan Lin, Wen-Chih Peng
PAKDD (3)2
2025 Template-Based Financial Report Generation in Agentic and Decomposed Information Retrieval
abstract
Tailoring structured financial reports from companies' earnings releases is crucial for understanding financial performance and has been widely adopted in real-world analytics. However, existing summarization methods often generate broad, high-level summaries, which may lack the precision and detail required for financial reports that typically focus on specific, structured sections. While Large Language Models (LLMs) hold promise, generating reports adhering to predefined multi-section templates remains challenging. This paper investigates two LLM-based approaches popular in industry for generating templated financial reports: an agentic information retrieval (IR) framework and a decomposed IR approach, namely AgenticIR and DecomposedIR. The AgenticIR utilizes collaborative agents prompted with the full template. In contrast, the DecomposedIR approach applies a prompt chaining workflow to break down the template and reframe each section as a query answered by the LLM using the earnings release. To quantitatively assess the generated reports, we evaluated both methods in two scenarios: one using a financial dataset without direct human references, and another with a weather-domain dataset featuring expert-written reports. Experimental results show that while AgenticIR may excel in orchestrating tasks and generating concise reports through agent collaboration, DecomposedIR statistically significantly outperforms AgenticIR approach in providing broader and more detailed coverage in both scenarios, offering reflection on the utilization of the agentic framework in real-world applications.
Yong-En Tian, Yu-Chien Tang, Kuang-Da Wang, An-Zi Yen, Wen-Chih Peng
SIGIR3
2024 Root Cause Analysis in Microservice Using Neural Granger Causal Discovery
abstract
In recent years, microservices have gained widespread adoption in IT operations due to their scalability, maintenance, and flexibility. However, it becomes challenging for site reliability engineers (SREs) to pinpoint the root cause due to the complex relationship in microservices when facing system malfunctions. Previous research employed structure learning methods (e.g., PC-algorithm) to establish causal relationships and derive root causes from causal graphs. Nevertheless, they ignored the temporal order of time series data and failed to leverage the rich information inherent in the temporal relationships. For instance, in cases where there is a sudden spike in CPU utilization, it can lead to an increase in latency for other microservices. However, in this scenario, the anomaly in CPU utilization occurs before the latency increases, rather than simultaneously. As a result, the PC-algorithm fails to capture such characteristics. To address these challenges, we propose RUN, a novel approach for root cause analysis using neural Granger causal discovery with contrastive learning. RUN enhances the backbone encoder by integrating contextual information from time series and leverages a time series forecasting model to conduct neural Granger causal discovery. In addition, RUN incorporates Pagerank with a personalization vector to efficiently recommend the top-k root causes. Extensive experiments conducted on the synthetic and real-world microservice-based datasets demonstrate that RUN noticeably outperforms the state-of-the-art root cause analysis methods. Moreover, we provide an analysis scenario for the sock-shop case to showcase the practicality and efficacy of RUN in microservice-based applications. Our code is publicly available at https://github.com/zmlin1998/RUN.
Cheng-Ming Lin, Ching Chang 0001, Wei-Yao Wang, Kuang-Da Wang, Wen-Chih Peng
AAAI4
2024 The CoachAI Badminton Environment: Bridging the Gap between a Reinforcement Learning Environment and Real-World Badminton Games
abstract
We present the CoachAI Badminton Environment, a reinforcement learning (RL) environment tailored for AI-driven sports analytics. In contrast to traditional environments using rule-based opponents or simplistic physics-based randomness, our environment integrates authentic opponent AIs and realistic randomness derived from real-world matches data to bridge the performance gap encountered in real-game deployments. This novel feature enables RL agents to seamlessly adapt to genuine scenarios. The CoachAI Badminton Environment empowers researchers to validate strategies in intricate real-world settings, offering: i) Realistic opponent simulation for RL training; ii) Visualizations for evaluation; and iii) Performance benchmarks for assessing agent capabilities. By bridging the RL environment with actual badminton games, our environment is able to advance the discovery of winning strategies for players. Our code is available at https://github.com/wywyWang/CoachAI-Projects/tree/main/Strategic%20Environment.
Kuang-Da Wang, Yu-Tse Chen, Yu-Heng Lin, Wei-Yao Wang, Wen-Chih Peng
AAAI1
2024 The CoachAI Badminton Environment: A Novel Reinforcement Learning Environment with Realistic Opponents (Student Abstract)
abstract
The growing demand for precise sports analysis has been explored to improve athlete performance in various sports (e.g., basketball, soccer). However, existing methods for different sports face challenges in validating strategies in environments due to simple rule-based opponents leading to performance gaps when deployed in real-world matches. In this paper, we propose the CoachAI Badminton Environment, a novel reinforcement learning (RL) environment with realistic opponents for badminton, which serves as a compelling example of a turn-based game. It supports researchers in exploring various RL algorithms with the badminton context by integrating state-of-the-art tactical-forecasting models and real badminton game records. The Badminton Benchmarks are proposed with multiple widely adopted RL algorithms to benchmark the performance of simulating matches against real players. To advance novel algorithms and developments in badminton analytics, we make our environment open-source, enabling researchers to simulate more complex badminton sports scenarios based on this foundation. Our code is available at https://github.com/wywyWang/CoachAI-Projects/tree/main/CoachAI%20Badminton%20Environment.
Kuang-Da Wang, Wei-Yao Wang, Yu-Tse Chen, Yu-Heng Lin, Wen-Chih Peng
AAAI1
2024 Offline Imitation of Badminton Player Behavior via Experiential Contexts and Brownian Motion
Kuang-Da Wang, Wei-Yao Wang, Ping-Chun Hsieh, Wen-Chih Peng
ECML/PKDD (10)1
2023 A Reinforcement Learning Badminton Environment for Simulating Player Tactics (Student Abstract)
abstract
Recent techniques for analyzing sports precisely has stimulated various approaches to improve player performance and fan engagement. However, existing approaches are only able to evaluate offline performance since testing in real-time matches requires exhaustive costs and cannot be replicated. To test in a safe and reproducible simulator, we focus on turn-based sports and introduce a badminton environment by simulating rallies with different angles of view and designing the states, actions, and training procedures. This benefits not only coaches and players by simulating past matches for tactic investigation, but also researchers from rapidly evaluating their novel algorithms. Our code is available at https://github.com/wywyWang/CoachAI-Projects/tree/main/Strategic%20Environment.
Li-Chun Huang, Nai-Zen Hseuh, Yen-Che Chien, Wei-Yao Wang, Kuang-Da Wang, Wen-Chih Peng
AAAI5
2023 Enhancing Badminton Player Performance via a Closed-Loop AI Approach: Imitation, Simulation, Optimization, and Execution
abstract
In recent years, the sports industry has witnessed a significant rise in interest in leveraging artificial intelligence to enhance players' performance. However, the application of deep learning to improve badminton athletes' performance faces challenges related to identifying weaknesses, generating winning suggestions, and validating strategy effectiveness. These challenges arise due to the limited availability of realistic environments and agents. This paper aims to address these research gaps and make contributions to the badminton community. To achieve this goal, we propose a closed-loop approach consisting of six key components: Badminton Data Acquisition, Imitating Players' Styles, Simulating Matches, Optimizing Strategies, Training Execution, and Real-World Competitions. Specifically, we developed a novel model called RallyNet, which excels at imitating players' styles, allowing agents to accurately replicate real players' behavior. Secondly, we created a sophisticated badminton simulation environment that incorporates real-world physics, faithfully recreating game situations. Thirdly, we employed reinforcement learning techniques to improve players' strategies, enhancing their chances of winning while preserving their unique playing styles. By comparing strategy differences before and after improvement, we provide winning suggestions to players, which can be validated against diverse opponents within our carefully designed environment. Lastly, through collaborations with badminton venues and players, we apply the generated suggestions to the players' training and competitions, ensuring the effectiveness of our approach. Moreover, we continuously gather data from training and competitions, incorporating it into the closed-loop cycle to refine strategies and suggestions. This research presents an innovative approach for continuously improving players' performance, contributing to the field of AI-driven sports performance enhancement. This dissertation is supervised by Wen-Chih Peng (wcp[email protected]) and Ping-Chun Hsieh ([email protected]).
Kuang-Da Wang
CIKM1