Parshin Shojaee

dblp:281/9859 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 55% Knowledge representation and reasoning · 22% Representation and self-supervised learning · 14%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 100%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation › inductive program synthesis
symbolic regression
1.522025
LLM-SR: Scientific Equation Discovery via Programming with Large Language Models · ICLR 2025
Transformer-based Planning for Symbolic Regression · NeurIPS 2023
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model
large reasoning model
0.912025
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
scientific knowledge discovery
0.912025
LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models · ICML 2025
Natural language and speech › Language models and text generation › alignment
sycophancy mitigation
0.912025
Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories · EMNLP 2025
Computational science and engineering › AI for science
AI for scientific discovery
0.912025
Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges · AAAI 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
SNIP: Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-training · ICLR 2024
Machine learning › Representation and self-supervised learning › pre-training › multimodal pretraining
cross-modal pre-training
0.812024
SNIP: Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-training · ICLR 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
symbolic reasoning
0.812024
SNIP: Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-training · ICLR 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
symbolic regression
0.812024
SNIP: Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-training · ICLR 2024
Computational science and engineering
scientific machine learning
0.812024
SNIP: Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-training · ICLR 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.712023
Transformer-based Planning for Symbolic Regression · NeurIPS 2023
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.312025
Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

large language model · 3.5program synthesis · 1.7evolutionary search · 1.7pre-training · 1.5contrastive learning · 1.5uncertainty estimation · 0.9theorem proving · 0.9symbolic regression · 0.9reinforcement learning · 0.9puzzle environment · 0.9data-driven modeling · 0.9complexity scaling analysis · 0.9pre-trained transformer · 0.7genetic programming · 0.7
YearPublicationVenuePosition
2025 Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges
abstract
Scientific discovery is a complex cognitive process that has driven human knowledge and technological progress for centuries. While artificial intelligence (AI) has made significant advances in automating aspects of scientific reasoning, simulation, and experimentation, we still lack integrated AI systems capable of performing autonomous long-term scientific research and discovery. This paper examines the current state of AI for scientific discovery, highlighting recent progress in large language models and other AI techniques applied to scientific tasks. We then outline key challenges and promising research directions toward developing more comprehensive AI systems for scientific discovery, including the need for science-focused AI agents, improved benchmarks and evaluation metrics, multimodal scientific representations, and unified frameworks combining reasoning, theorem proving, and data-driven modeling. Addressing these challenges could lead to transformative AI tools to accelerate progress across disciplines towards scientific discovery.
Chandan K. Reddy, Parshin Shojaee
AAAI2
2025 Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
abstract
Mohammad Beigi, Ying Shen, Parshin Shojaee, Qifan Wang, Zichao Wang, Chandan K. Reddy, Ming Jin, Lifu Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Mohammad Beigi, Ying Shen 0006, Parshin Shojaee, Qifan Wang 0001, Zichao Wang 0001, Chandan K. Reddy, Ming Jin 0002, Lifu Huang
EMNLP3
2025 LLM-SR: Scientific Equation Discovery via Programming with Large Language Models
abstract
Mathematical equations have been unreasonably effective in describing complex natural phenomena across various scientific disciplines. However, discovering such insightful equations from data presents significant challenges due to the necessity of navigating extremely large combinatorial hypothesis spaces. Current methods of equation discovery, commonly known as symbolic regression techniques, largely focus on extracting equations from data alone, often neglecting the domain-specific prior knowledge that scientists typically depend on. They also employ limited representations such as expression trees, constraining the search space and expressiveness of equations. To bridge this gap, we introduce LLM-SR, a novel approach that leverages the extensive scientific knowledge and robust code generation capabilities of Large Language Models (LLMs) to discover scientific equations from data. Specifically, LLM-SR treats equations as programs with mathematical operators and combines LLMs' scientific priors with evolutionary search over equation programs. The LLM iteratively proposes new equation skeleton hypotheses, drawing from its domain knowledge, which are then optimized against data to estimate parameters. We evaluate LLM-SR on four benchmark problems across diverse scientific domains (e.g., physics, biology), which we carefully designed to simulate the discovery process and prevent LLM recitation. Our results demonstrate that LLM-SR discovers physically accurate equations that significantly outperform state-of-the-art symbolic regression baselines, particularly in out-of-domain test settings. We also show that LLM-SR's incorporation of scientific priors enables more efficient equation space exploration than the baselines.
Parshin Shojaee, Kazem Meidani, Amir Barati Farimani, Chandan K. Reddy
ICLR1
2025 LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
abstract
Scientific equation discovery is a fundamental task in the history of scientific progress, enabling the derivation of laws governing natural phenomena. Recently, Large Language Models (LLMs) have gained interest for this task due to their potential to leverage embedded scientific knowledge for hypothesis generation. However, evaluating the true discovery capabilities of these methods remains challenging, as existing benchmarks often rely on common equations that are susceptible to memorization by LLMs, leading to inflated performance metrics that do not reflect actual discovery. In this paper, we introduce LLM-SRBench, a comprehensive benchmark with 239 challenging problems across four scientific domains specifically designed to evaluate LLM-based scientific equation discovery methods while preventing trivial memorization. Our benchmark comprises two main categories: LSR-Transform, which transforms common physical models into less common mathematical representations to test reasoning beyond memorization, and LSR-Synth, which introduces synthetic, discovery-driven problems requiring data-driven reasoning. Through extensive evaluation of several state-of-the-art methods on LLM-SRBench, using both open and closed LLMs, we find that the best-performing system so far achieves only 31.5% symbolic accuracy. These findings highlight the challenges of scientific equation discovery, positioning LLM-SRBench as a valuable resource for future research.
Parshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani, Khoa D. Doan, Chandan K. Reddy
ICML1
2025 The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
abstract
Recent generations of frontier language models have introduced Large Reasoning Models (LRMs) that generate detailed thinking processes before providing answers. While these models demonstrate improved performance on reasoning benchmarks, their fundamental capabilities, scaling properties, and limitations remain insufficiently understood. Current evaluations primarily focus on established mathematical and coding benchmarks, emphasizing final answer accuracy. However, this evaluation paradigm often suffers from data contamination and does not provide insights into the reasoning traces' structure and quality. In this work, we systematically investigate these gaps with the help of controllable puzzle environments that allow precise manipulation of compositional complexity while maintaining consistent logical structures. This setup enables the analysis of not only final answers but also the internal reasoning traces, offering insights into how LRMs ``think''. Through extensive experimentation across diverse puzzles, we show that frontier LRMs face a complete accuracy collapse beyond certain complexities. Moreover, they exhibit a counterintuitive scaling limit: their reasoning effort increases with problem complexity up to a point, then declines despite having an adequate token budget. By comparing LRMs with their standard LLM counterparts under equivalent inference compute, we identify three performance regimes: (1) low-complexity tasks where standard models surprisingly outperform LRMs, (2) medium-complexity tasks where additional thinking in LRMs demonstrates advantage, and (3) high-complexity tasks where both models experience complete collapse. We found that LRMs have limitations in exact computation: they fail to use explicit algorithms and reason inconsistently across scales and problems. We also investigate the reasoning traces in more depth, studying the patterns of explored solutions and analyzing the models' computational behavior, shedding light on their strengths, limitations, and ultimately raising questions about the nature for their reasoning capabilities.
Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton, Samy Bengio, Mehrdad Farajtabar
NeurIPS1
2024 SNIP: Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-training
abstract
In an era where symbolic mathematical equations are indispensable for modeling complex natural phenomena, scientific inquiry often involves collecting observations and translating them into mathematical expressions. Recently, deep learning has emerged as a powerful tool for extracting insights from data. However, existing models typically specialize in either numeric or symbolic domains, and are usually trained in a supervised manner tailored to specific tasks. This approach neglects the substantial benefits that could arise from a task-agnostic multi-modal understanding between symbolic equations and their numeric counterparts. To bridge the gap, we introduce SNIP, a Symbolic-Numeric Integrated Pre-training model, which employs contrastive learning between symbolic and numeric domains, enhancing their mutual similarities in the embeddings. By performing latent space analysis, we observe that SNIP provides cross-domain insights into the representations, revealing that symbolic supervision enhances the embeddings of numeric data and vice versa. We evaluate SNIP across diverse tasks, including symbolic-to-numeric mathematical property prediction and numeric-to-symbolic equation discovery, commonly known as symbolic regression. Results show that SNIP effectively transfers to various tasks, consistently outperforming fully supervised baselines and competing strongly with established task-specific methods, especially in the low data regime scenarios where available data is limited.
Kazem Meidani, Parshin Shojaee, Chandan K. Reddy, Amir Barati Farimani
ICLR2
2023 Transformer-based Planning for Symbolic Regression
abstract
Symbolic regression (SR) is a challenging task in machine learning that involves finding a mathematical expression for a function based on its values. Recent advancements in SR have demonstrated the effectiveness of pre-trained transformer models in generating equations as sequences, leveraging large-scale pre-training on synthetic datasets and offering notable advantages in terms of inference time over classical Genetic Programming (GP) methods. However, these models primarily rely on supervised pre-training objectives borrowed from text generation and overlook equation discovery goals like accuracy and complexity. To address this, we propose TPSR, a Transformer-based Planning strategy for Symbolic Regression that incorporates Monte Carlo Tree Search planning algorithm into the transformer decoding process. Unlike conventional decoding strategies, TPSR enables the integration of non-differentiable equation verification feedback, such as fitting accuracy and complexity, as external sources of knowledge into the transformer equation generation process. Extensive experiments on various datasets show that our approach outperforms state-of-the-art methods, enhancing the model's fitting-complexity trade-off, extrapolation abilities, and robustness to noise.
Parshin Shojaee, Kazem Meidani, Amir Barati Farimani, Chandan K. Reddy
NeurIPS1
2022 Task-Driven Privacy-Preserving Data-Sharing Framework for the Industrial Internet
abstract
Industrial Internet provides a collaborative computational platform for participating enterprises, allowing the collection of big data for machine learning tasks. Despite the promise of training and deployment acceleration, and the potential to optimize decision-making processes through data-sharing, the adoption of such technologies is impacted by the increasing concerns about information privacy. As enterprises prefer to keep data private, this limits interoperability. While prior work has largely explored privacy-preserving mechanisms, the proposed methods naively average or randomly sample data shared from all participants instead of selecting the most well-suited subsets for a particular downstream learning task. Motivated by the lack of effective data-sharing mechanisms for heterogeneous machine learning tasks in Industrial Internet, we propose PriED, a task-driven data-sharing framework that selectively fuses shared data and local data from participants to improve supervised learning performance. PriED utilizes privacy-preserving data distillation to facilitate data exchange, and dynamic data selection to optimize downstream machine learning tasks. We demonstrate performance improvements on a real semiconductor manufacturing case study.
Parshin Shojaee, Yingyan Zeng, Muntasir Wahed, Avi Seth, Ismini Lourentzou
IEEE Big Data1
2022 Adaptively Weighted Top-N Recommendation for Organ Matching
abstract
Reducing the shortage of organ donations to meet the demands of patients on the waiting list has being a major challenge in organ transplantation. Because of the shortage, organ matching decision is the most critical decision to assign the limited viable organs to the most “suitable” patients. Currently, organ matching decisions are only made by matching scores calculated via scoring models, which are built by the first principles. However, these models may disagree with the actual post-transplantation matching performance (e.g., patient's post-transplant quality of life (QoL) or graft failure measurements). In this paper, we formulate the organ matching decision-making as a top-N recommendation problem and propose an Adaptively Weighted Top-N Recommendation (AWTR) method. AWTR improves performance of the current scoring models by using limited actual matching performance in historical datasets as well as the collected covariates from organ donors and patients. AWTR sacrifices the overall recommendation accuracy by emphasizing the recommendation and ranking accuracy for top-N matched patients. The proposed method is validated in a simulation study, where KAS [ 60 ] is used to simulate the organ-patient recommendation response. The results show that our proposed method outperforms seven state-of-the-art top-N recommendation benchmark methods.
Parshin Shojaee, Xiaoyu Chen 0005
ACM Trans. Comput. Heal.1