VLDB 2026 Research / reviewers in the wild / expert
Yongchao Chen
dblp:88/9557
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 10 since 2021Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FIGNet: A Robust and Interpretable Fuzzy-Irreversible Gated Network for Auditory Brainstem Response ClassificationabstractAuditory brainstem response (ABR) is an important tool for newborn hearing screening and neurological assessment. However, its signals are often difficult to be accurately resolved due to noise interference and weak waveforms, and the need for repeated measurements under multiple sound intensity conditions results in time-consuming data acquisition. Therefore, there is an urgent need to develop an automatic classification model with high accuracy, robustness and good interpretability to achieve stable and effective recognition performance with minimal ABR data. This study presents FIGNet, a new deep learning model that combines type-2 fuzzy logic with a time-irreversible attention mechanism to address uncertainty and temporal direction in ABR signals. Fuzzy attention helps reduce the impact of noise, while the irreversible attention models the one-way nature of neural responses. Experiments on real ABR datasets show that FIGNet outperforms existing models in both binary and five-class classification tasks. It achieves 93.72% accuracy in binary classification and 84.42% accuracy in five-class classification. Visualization results-including confusion matrices, and accuracy curves under different noise levels-further confirm that FIGNet can focus on key waveform areas and stay reliable even in noisy conditions. These findings demonstrate that FIGNet offers fast, interpretable, and robust performance for clinical ABR analysis, achieving high classification accuracy under both clean and noisy conditions. Ke Zhang 0040, Chunrui Zhao, Zenan Li, Caiwei Li, Desheng Jia, Yongchao Chen, Shang Yan, Xin Wang 0088, Yishu Teng, Hongguang Pan, Shixiong Chen |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Steering Large Language Models between Code Execution and Textual ReasoningabstractWhile a lot of recent research focuses on enhancing the textual reasoning capabilities of Large Language Models (LLMs) by optimizing the multi-agent framework or reasoning chains, several benchmark tasks can be solved with 100\% success through direct coding, which is more scalable and avoids the computational overhead associated with textual iterating and searching. Textual reasoning has inherent limitations in solving tasks with challenges in math, logics, optimization, and searching, which is unlikely to be solved by simply scaling up the model and data size. The recently released OpenAI GPT Code Interpreter and multi-agent frameworks such as AutoGen have demonstrated remarkable proficiency of integrating code generation and execution to solve complex tasks using LLMs. However, based on our experiments on 7 existing popular methods for steering code/text generation in both single- and multi-turn settings with 14 tasks and 6 types of LLMs (including the new O1-preview), currently there is no optimal method to correctly steer LLMs to write code when needed. We discover some interesting patterns on when models use code vs. textual reasoning with the evolution to task complexity and model sizes, which even result in an astonishingly inverse scaling behavior. We also discover that results from LLM written code are not always better than using textual reasoning, even if the task could be solved through code. To mitigate the above issues, we propose three methods to better steer LLM code/text generation and achieve a notable improvement. The costs of token lengths and runtime are thoroughly discussed for all the methods. We believe the problem of steering LLM code/text generation is critical for future research and has much space for further improvement. Project Page, Datasets, and Codes are available at https://yongchao98.github.io/CodeSteer/. Yongchao Chen, Harsh Jhamtani, Srinagesh Sharma, Chuchu Fan |
ICLR | 1 |
| 2025 | CodeSteer: Symbolic-Augmented Language Models via Code/Text GuidanceabstractExisting methods fail to effectively steer Large Language Models (LLMs) between textual reasoning and code generation, leaving symbolic computing capabilities underutilized. We introduce CodeSteer, an effective method for guiding LLM code/text generation. We construct a comprehensive benchmark SymBench comprising 37 symbolic tasks with adjustable complexity and also synthesize datasets of 12k multi-turn guidance/generation trajectories and 5.5k guidance comparison pairs. We fine-tune the Llama-3-8B model with a newly designed multi-turn supervised fine-tuning (SFT) and direct preference optimization (DPO). The resulting model, CodeSteerLLM, augmented with the proposed symbolic and self-answer checkers, effectively guides the code/text generation of larger models. Augmenting GPT-4o with CodeSteer raises its average performance score from 53.3 to 86.4, even outperforming the existing best LLM OpenAI o1 (82.7), o1-preview (74.8), and DeepSeek R1 (76.8) across all 37 tasks (28 seen, 9 unseen). Trained for GPT-4o, CodeSteer demonstrates superior generalizability, providing an average 41.8 performance boost on Claude, Mistral, and GPT-3.5. CodeSteer-guided LLMs fully harness symbolic computing to maintain strong performance on highly complex tasks. Models, Datasets, and Codes are available at https://github.com/yongchao98/CodeSteer-v1.0 and https://huggingface.co/yongchao98. Yongchao Chen, Yilun Hao, Yang Zhang 0001, Chuchu Fan |
ICML | 1 |
| 2025 | Code-as-Symbolic-Planner: Foundation Model-Based Robot Planning via Symbolic Code GenerationabstractRecent works have shown great potential of Large Language Models (LLMs) in robot task and motion planning (TAMP). Current LLM approaches generate text- or code-based reasoning chains with sub-goals and action plans. However, they do not fully leverage LLMs’ symbolic computing and code generation capabilities. Many robot TAMP tasks involve complex optimization under multiple constraints, where pure textual reasoning is insufficient. While augmenting LLMs with predefined solvers and planners improves performance, it lacks generalization across tasks. Given LLMs’ growing coding proficiency, we enhance their TAMP capabilities by steering them to generate code as symbolic planners for optimization and constraint verification. Unlike prior work that uses code to interface with robot action modules or pre-designed planners, we steer LLMs to generate code as solvers, planners, and checkers for TAMP tasks requiring symbolic computing, while still leveraging textual reasoning to incorporate common sense. With a multi-round guidance and answer evolution framework, the proposed Code-as-Symbolic-Planner improves success rates by average 24.1% over best baseline methods across seven typical TAMP tasks and three popular LLMs. Code-as-Symbolic-Planner shows strong effectiveness and generalizability across discrete and continuous environments, 2D/3D simulations and real-world settings, as well as single- and multi-robot tasks with diverse requirements. See our project website†for prompts, videos, and code. Yongchao Chen, Yilun Hao, Yang Zhang 0001, Chuchu Fan |
IROS | 1 |
| 2025 | Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification ToolsabstractYilun Hao, Yongchao Chen, Yang Zhang, Chuchu Fan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yilun Hao, Yongchao Chen, Yang Zhang 0001, Chuchu Fan |
NAACL (Long Papers) | 2 |
| 2024 | Enhancing Imbalanced Classification with Support Vector Machines via Evolutionary Oversampling AlgorithmsabstractSupport Vector Machines (SVMs), as well-known algorithms, have been successfully applied to classification problems. However, when dealing with imbalanced data, the classification performance of SVMs could be significantly compromised. One approach to tackle the class imbalance is oversampling the minority class, exemplified by methods like SMOTE and its variants. These methods generate new samples by interpolation between existing ones and determine the weights based on the ratio of samples from different classes, leading to inaccurate weight assignment, limited generation scope, and indiscriminate sample generation. To address these limitations, we propose novel evolutionary oversampling algorithms based on Support Vector Machine (SVM) and Evolutionary Algorithms (EAs) called SEOA. SEOA leverages the inherent capability of SVM to identify the samples that have a critical influence on the decision boundary and assign them appropriate weights, thereby eliminating the reliance on human experience. Furthermore, SEOA utilizes a novel approach for sample generation and emphasizes the significance of margin for classification, introducing a mechanism that employs margin as the metric to evaluate the quality of generated samples. To assess the performance of SEOA, we conducted a comprehensive comparison against various oversampling methods across 19 real-world datasets. The results underscore SEOA's superiority, showcasing its distinct strengths in addressing the challenges posed by imbalanced classification. Yongchao Chen, Yaqing Hou, Xiangrong Tong, Qiang Zhang 0008 |
CEC | 2 |
| 2024 | PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based SamplingabstractPrompt optimization aims to find the best prompt to a large language model (LLM) for a given task. LLMs have been successfully used to help find and improve prompt candidates for single-step tasks. However, realistic tasks for agents are multi-step and introduce new challenges: (1) Prompt content is likely to be more extensive and complex, making it more difficult for LLMs to analyze errors, (2) the impact of an individual step is difficult to evaluate, and (3) different people may have varied preferences about task execution. While humans struggle to optimize prompts, they are good at providing feedback about LLM outputs; we therefore introduce a new LLM-driven discrete prompt optimization framework PROMST that incorporates human-designed feedback rules to automatically offer direct suggestions for improvement. We also use an extra learned heuristic model that predicts prompt performance to efficiently sample from prompt candidates. This approach significantly outperforms both human-engineered prompts and several other prompt optimization methods across 11 representative multi-step tasks (an average 10.6%-29.3% improvement to current best methods on five LLMs respectively). We believe our work can serve as a benchmark for automatic prompt optimization for LLM-driven multi-step tasks. Yongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang 0001, Nicholas Roy, Chuchu Fan |
EMNLP | 1 |
| 2024 | AutoTAMP: Autoregressive Task and Motion Planning with LLMs as Translators and CheckersabstractFor effective human-robot interaction, robots need to understand, plan, and execute complex, long-horizon tasks described by natural language. Recent advances in large language models (LLMs) have shown promise for translating natural language into robot action sequences for complex tasks. However, existing approaches either translate the natural language directly into robot trajectories or factor the inference process by decomposing language into task sub-goals and relying on a motion planner to execute each sub-goal. When complex environmental and temporal constraints are involved, inference over planning tasks must be performed jointly with motion plans using traditional task-and-motion planning (TAMP) algorithms, making factorization into subgoals untenable. Rather than using LLMs to directly plan task sub-goals, we instead perform few-shot translation from natural language task descriptions to an intermediate task representation that can then be consumed by a TAMP algorithm to jointly solve the task and motion plan. To improve translation, we automatically detect and correct both syntactic and semantic errors via autoregressive re-prompting, resulting in significant improvements in task completion. We show that our approach outperforms several methods using LLMs as planners in complex task domains. See our project website§for prompts, videos, and code. Yongchao Chen, Jacob Arkin, Charles Dawson 0001, Yang Zhang 0001, Nicholas Roy, Chuchu Fan |
ICRA | 1 |
| 2024 | Scalable Multi-Robot Collaboration with Large Language Models: Centralized or Decentralized Systems?abstractA flurry of recent work has demonstrated that pre-trained large language models (LLMs) can be effective task planners for a variety of single-robot tasks. The planning performance of LLMs is significantly improved via prompting techniques, such as in-context learning or re-prompting with state feedback, placing new importance on the token budget for the context window. An under-explored but natural next direction is to investigate LLMs as multi-robot task planners. However, long-horizon, heterogeneous multi-robot planning introduces new challenges of coordination while also pushing up against the limits of context window length. It is therefore critical to find token-efficient LLM planning frameworks that are also able to reason about the complexities of multi-robot coordination. In this work, we compare the task success rate and token efficiency of four multi-agent communication frameworks (centralized, decentralized, and two hybrid) as applied to four coordination-dependent multi-agent 2D task scenarios for increasing numbers of agents. We find that a hybrid framework achieves better task success rates across all four tasks and scales better to more agents. We further demonstrate the hybrid frameworks in 3D simulations where the vision-to-text problem and dynamical errors are considered. See our project website4for prompts, videos, and code. Yongchao Chen, Jacob Arkin, Yang Zhang 0001, Nicholas Roy, Chuchu Fan |
ICRA | 1 |
| 2023 | Surrogate-Assisted Morphology Optimization by Genetic AlgorithmsabstractDeep reinforcement learning has attracted wide interest because of its extraordinary capabilities in multiple fields. However, morphology optimization by using evolutionary computation techniques has not been intensively investigated. In this paper, we explore the use of genetic algorithms (GA) to automatically design the morphology of an agent. Evaluating the performance of an agent is very time-consuming because it needs to be trained from scratch. Moreover, it is computationally infeasible to train separate controllers for all possible different morphologies of agents to identify the optimal ones and is difficult to obtain the accurate cumulative reward of an agent to estimate the performance of the morphologies. To address these issues, we use a morphology comparator as a surrogate model to estimate the probability of one morphology being better than the other, instead of directly predicting the performance of each morphology. A set of surrogate models based on a radial basis function network are developed before evolution to make full use of the data to guide the search. Experimental results indicate that the proposed method is able to efficiently find out optimal morphologies to achieve better performance than the default morphology. Jinlin Jiang, Yongchao Chen, Wenbin Pei, Junxiang Zhang, Yaqing Hou, Hong-Wei Ge, Liang Feng 0001 |
CEC | 2 |
| 2023 | NL2TL: Transforming Natural Languages to Temporal Logics using Large Language ModelsabstractTemporal Logic (TL) can be used to rigorously specify complex high-level specification for systems in many engineering applications.The translation between natural language (NL) and TL has been under-explored due to the lack of dataset and generalizable model across different application domains.In this paper, we propose an accurate and generalizable transformation framework of English instructions from NL to TL, exploring the use of Large Language Models (LLMs) at multiple stages.Our contributions are twofold.First, we develop a framework to create a dataset of NL-TL pairs combining LLMs and human annotation.We publish a dataset with 28K NL-TL pairs.Then, we finetune T5 models on the lifted versions (i.e., the specific Atomic Propositions (AP) are hidden) of the NL and TL.The enhanced generalizability originates from two aspects: 1) Usage of lifted NL-TL characterizes common logical structures, without constraints of specific domains.2) Application of LLMs in dataset creation largely enhances corpus richness.We test the generalization of trained models on five varied domains.To achieve full NL-TL transformation, we either combine the lifted model with AP recognition task or do the further finetuning on each specific domain.During the further finetuning, our model achieves higher accuracy (>95%) using only <10% training data, compared with the baseline sequence to sequence (Seq2Seq) model.12 Yongchao Chen, Rujul Gandhi, Yang Zhang 0001, Chuchu Fan |
EMNLP | 1 |
| 2021 | Improving Text Summarization Using Feature Extraction Approach Based on Pointer-generator with Coverage
Yongchao Chen, Xin He 0021, Guanghui Wang 0003, Junyang Yu |
WISA | 1 |