Chak Shing Lee

dblp:139/1287 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-9573-3463ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Mathematical optimization · 100%
Artificial intelligence
1 paper
Trustworthy machine learning · 100%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation › inductive program synthesis
symbolic regression
1.422025
DisCo-DSO: Coupling Discrete and Continuous Optimization for Efficient Generative Design in Hybrid Spaces · AAAI 2025
A Unified Framework for Deep Symbolic Regression · NeurIPS 2022
Machine learning › Trustworthy machine learning › interpretability › explainable reinforcement learning
decision tree policy
0.912025
DisCo-DSO: Coupling Discrete and Continuous Optimization for Efficient Generative Design in Hybrid Spaces · AAAI 2025
Machine learning › Trustworthy machine learning › interpretability
explainable reinforcement learning
0.912025
DisCo-DSO: Coupling Discrete and Continuous Optimization for Efficient Generative Design in Hybrid Spaces · AAAI 2025
Mathematical optimization
black-box optimization
0.912025
DisCo-DSO: Coupling Discrete and Continuous Optimization for Efficient Generative Design in Hybrid Spaces · AAAI 2025
Mathematical optimization › design optimization
generative design
0.912025
DisCo-DSO: Coupling Discrete and Continuous Optimization for Efficient Generative Design in Hybrid Spaces · AAAI 2025

Methods — techniques the papers use, named apart from their topics

generative model · 2.6deep symbolic optimization · 2.6recursive problem simplification · 1.7pre-training · 1.7neural-guided search · 1.7genetic programming · 1.7
YearPublicationVenuePosition
2025 DisCo-DSO: Coupling Discrete and Continuous Optimization for Efficient Generative Design in Hybrid Spaces
abstract
We consider the challenge of black-box optimization within hybrid discrete-continuous and variable-length spaces, a problem that arises in various applications, such as decision tree learning and symbolic regression. We propose DisCo-DSO (Discrete-Continuous Deep Symbolic Optimization), a novel approach that uses a generative model to learn a joint distribution over discrete and continuous design variables to sample new hybrid designs. In contrast to standard decoupled approaches, in which the discrete and continuous variables are optimized separately, our joint optimization approach uses fewer objective function evaluations, is robust against non-differentiable objectives, and learns from prior samples to guide the search, leading to significant improvement in performance and sample efficiency. Our experiments on a diverse set of optimization tasks demonstrate that the advantages of DisCo-DSO become increasingly evident as problem complexity grows. In particular, we illustrate DisCo-DSO's superiority over the state-of-the-art methods for interpretable reinforcement learning with decision trees.
Jacob F. Pettit, Chak Shing Lee, Alex Ho, Daniel M. Faissol, Brenden K. Petersen, Mikel Landajuela
AAAI2
2025 SRBench++: Principled Benchmarking of Symbolic Regression With Domain-Expert Interpretation
abstract
Symbolic regression searches for analytic expressions that accurately describe studied phenomena. The main promise of this approach is that it may return an interpretable model that can be insightful to users, while maintaining high accuracy. The current standard for benchmarking these algorithms is SRBench, which evaluates methods on hundreds of datasets that are a mix of real-world and simulated processes spanning multiple domains. At present, the ability of SRBench to evaluate interpretability is limited to measuring the size of expressions on real-world data, and the exactness of model forms on synthetic data. In practice, model size is only one of many factors used by subject experts to determine how interpretable a model truly is. Furthermore, SRBench does not characterize algorithm performance on specific, challenging sub-tasks of regression such as feature selection and evasion of local minima. In this work, we propose and evaluate an approach to benchmarking SR algorithms that addresses these limitations of SRBench by 1) incorporating expert evaluations of interpretability on a domain-specific task, and 2) evaluating algorithms over distinct properties of data science tasks. We evaluate 12 modern symbolic regression algorithms on these benchmarks and present an in-depth analysis of the results, discuss current challenges of symbolic regression algorithms and highlight possible improvements for the benchmark itself.
Fabrício Olivetti de França, Marco Virgolin, Michael Kommenda, Maimuna S. Majumder, Miles D. Cranmer, Guilherme Espada, Leon Ingelse, Alcides Fonseca, Mikel Landajuela, Brenden K. Petersen, Ruben Glatt, T. Nathan Mundhenk, Chak Shing Lee, Jacob D. Hochhalter, David L. Randall, P. Kamienny, Hengzhe Zhang, Grant Dick, Alessandro Simon, Bogdan Burlacu, Jaan Kasak, Meera Vieira Machado, Casper Wilstrup, William G. La Cava
IEEE Trans. Evol. Comput.13
2022 A Unified Framework for Deep Symbolic Regression
abstract
The last few years have witnessed a surge in methods for symbolic regression, from advances in traditional evolutionary approaches to novel deep learning-based systems. Individual works typically focus on advancing the state-of-the-art for one particular class of solution strategies, and there have been few attempts to investigate the benefits of hybridizing or integrating multiple strategies. In this work, we identify five classes of symbolic regression solution strategies---recursive problem simplification, neural-guided search, large-scale pre-training, genetic programming, and linear models---and propose a strategy to hybridize them into a single modular, unified symbolic regression framework. Based on empirical evaluation using SRBench, a new community tool for benchmarking symbolic regression methods, our unified framework achieves state-of-the-art performance in its ability to (1) symbolically recover analytical expressions, (2) fit datasets with high accuracy, and (3) balance accuracy-complexity trade-offs, across 252 ground-truth and black-box benchmark problems, in both noiseless settings and across various noise levels. Finally, we provide practical use case-based guidance for constructing hybrid symbolic regression algorithms, supported by extensive, combinatorial ablation studies.
Mikel Landajuela, Chak Shing Lee, Ruben Glatt, Cláudio P. Santiago, Ignacio Aravena, T. Nathan Mundhenk, Garrett Mulcahy, Brenden K. Petersen
NeurIPS2