Junyan Cheng

dblp:305/5474 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Deep learning architectures and training · 27% Representation and self-supervised learning · 24% Planning, search and constraint satisfaction · 14%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
automated design
0.912025
Language Modeling by Language Models · NeurIPS 2025
Compilers and program optimization
code generation
0.912025
Language Modeling by Language Models · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems
autonomous agents
0.812024
SocioDojo: Building Lifelong Analytical Agents with Real-world Text and Time Series · ICLR 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
0.812024
Bridging Neural and Symbolic Representations with Transitional Dictionary Learning · ICLR 2024
Machine learning › Representation and self-supervised learning › structured representation
symbolic representation learning
0.812024
Bridging Neural and Symbolic Representations with Transitional Dictionary Learning · ICLR 2024
Machine learning › Deep learning architectures and training › transformer
efficient transformer
0.512021
Multimodal Phased Transformer for Sentiment Analysis · EMNLP (1) 2021
Machine learning › Efficient and distributed learning
parameter sharing
0.512021
Multimodal Phased Transformer for Sentiment Analysis · EMNLP (1) 2021
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
0.512021
Multimodal Phased Transformer for Sentiment Analysis · EMNLP (1) 2021
Machine learning › Deep learning architectures and training
transformer
0.512021
Multimodal Phased Transformer for Sentiment Analysis · EMNLP (1) 2021
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training
0.312025
Language Modeling by Language Models · NeurIPS 2025
Machine learning › Deep learning architectures and training
scaling laws
0.312025
Language Modeling by Language Models · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.212024
Bridging Neural and Symbolic Representations with Transitional Dictionary Learning · ICLR 2024
Natural language and speech › Language models and text generation
prompting
0.212024
SocioDojo: Building Lifelong Analytical Agents with Real-world Text and Time Series · ICLR 2024
Natural language and speech › Information extraction and text analysis › sentiment analysis
multimodal sentiment analysis
0.112021
Multimodal Phased Transformer for Sentiment Analysis · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.112021
Multimodal Phased Transformer for Sentiment Analysis · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

multi-agent LLM · 1.7ladder of scales · 1.7genetic programming · 1.7adversarial review · 1.7time series analysis · 0.8knowledge graph search · 0.8hypothesis-and-proof prompting · 0.8game-theoretic diffusion · 0.8expectation-maximization · 0.8dictionary learning · 0.8
YearPublicationVenuePosition
2025 Language Modeling by Language Models
abstract
*Can we leverage LLMs to model the process of discovering novel language model (LM) architectures?* Inspired by real research, we propose a multi-agent LLM approach that simulates the conventional stages of research, from ideation and literature search (proposal stage) to design implementation (code generation), generative pre-training, and downstream evaluation (verification). Using ideas from scaling laws, our system *Genesys* employs a *Ladder of Scales* approach; new designs are proposed, adversarially reviewed, implemented, and selectively verified at increasingly larger model scales (14M$\sim$350M parameters) with a narrowing budget (the number of models we can train at each scale). To help make discovery efficient and factorizable, Genesys uses a novel genetic programming backbone, which we show has empirical advantages over commonly used direct prompt generation workflows (e.g., $\sim$86\% percentage point improvement in successful design generation, a key bottleneck). We report experiments involving 1,162 newly discovered designs (1,062 fully verified) and find the best designs to be competitive with known architectures (e.g., outperform GPT2, Mamba2, etc., on 6/9 common benchmarks). We couple these results with comprehensive system-level ablations and formal results, which give broader insights into the design of effective autonomous discovery systems.
Junyan Cheng, Peter Clark, Kyle Richardson 0001
NeurIPS1
2024 SocioDojo: Building Lifelong Analytical Agents with Real-world Text and Time Series
abstract
We introduce SocioDojo, an open-ended lifelong learning environment for developing ready-to-deploy autonomous agents capable of performing human-like analysis and decision-making on societal topics such as economics, finance, politics, and culture. It consists of (1) information sources from news, social media, reports, etc., (2) a knowledge base built from books, journals, and encyclopedias, plus a toolbox of Internet and knowledge graph search interfaces, (3) 30K high-quality time series in finance, economy, society, and polls, which support a novel task called "hyperportfolio", that can reliably and scalably evaluate societal analysis and decision-making power of agents, inspired by portfolio optimization with time series as assets to "invest". We also propose a novel Analyst-Assistant-Actuator architecture for the hyperportfolio task, and a Hypothesis & Proof prompting for producing in-depth analyses on input news, articles, etc. to assist decision-making. We perform experiments and ablation studies to explore the factors that impact performance. The results show that our proposed method achieves improvements of 32.4% and 30.4% compared to the state-of-the-art method in the two experimental settings.
Junyan Cheng, Sang (Peter) Chin
ICLR1
2024 Bridging Neural and Symbolic Representations with Transitional Dictionary Learning
abstract
This paper introduces a novel Transitional Dictionary Learning (TDL) framework that can implicitly learn symbolic knowledge, such as visual parts and relations, by reconstructing the input as a combination of parts with implicit relations. We propose a game-theoretic diffusion model to decompose the input into visual parts using the dictionaries learned by the Expectation Maximization (EM) algorithm, implemented as the online prototype clustering, based on the decomposition results. Additionally, two metrics, clustering information gain, and heuristic shape score are proposed to evaluate the model. Experiments are conducted on three abstract compositional visual object datasets, which require the model to utilize the compositionality of data instead of simply exploiting visual features. Then, three tasks on symbol grounding to predefined classes of parts and relations, as well as transfer learning to unseen classes, followed by a human evaluation, were carried out on these datasets. The results show that the proposed method discovers compositional patterns, which significantly outperforms the state-of-the-art unsupervised part segmentation methods that rely on visual features from pre-trained backbones. Furthermore, the proposed metrics are consistent with human evaluations.
Junyan Cheng, Sang (Peter) Chin
ICLR1
2021 Multimodal Phased Transformer for Sentiment Analysis
abstract
Multimodal Transformers achieve superior performance in multimodal learning tasks.However, the quadratic complexity of the selfattention mechanism in Transformers limits their deployment in low-resource devices and makes their inference and training computationally expensive.We propose multimodal Sparse Phased Transformer (SPT) to alleviate the problem of self-attention complexity and memory footprint.SPT uses a sampling function to generate a sparse attention matrix and compress a long sequence to a shorter sequence of hidden states.SPT concurrently captures interactions between the hidden states of different modalities at every layer.To further improve the efficiency of our method, we use Layer-wise parameter sharing and Factorized Co-Attention that share parameters between Cross Attention Blocks, with minimal impact on task performance.We evaluate our model with three sentiment analysis datasets and achieve comparable or superior performance compared with the existing methods, with a 90% reduction in the number of parameters.We conclude that (SPT) along with parameter sharing can capture multimodal interactions with reduced model size and improved sample efficiency.
Junyan Cheng, Iordanis Fostiropoulos, Barry W. Boehm, Mohammad Soleymani 0001
EMNLP (1)1