Emily Jin

dblp:346/1033 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Graph learning · 41% 3D vision · 17% Knowledge representation and reasoning · 10%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 21 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning › graph neural network
expressive power
1.622025
Homomorphism Counts as Structural Encodings for Graph Learning · ICLR 2025
Homomorphism Counts for Graph Neural Networks: All About That Basis · ICML 2024
Computer vision › 3D vision › 3d generation
3d scene generation
0.912025
From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries · NeurIPS 2025
Computer vision › 3D vision
3d scene understanding
0.912025
From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries · NeurIPS 2025
Machine learning › Graph learning
graph neural network
0.912025
Homomorphism Counts as Structural Encodings for Graph Learning · ICLR 2025
Machine learning › Graph learning › graph neural network
graph transformer
0.912025
Homomorphism Counts as Structural Encodings for Graph Learning · ICLR 2025
Computer vision › 3D vision › 3d scene modeling › scene representation
object-centric scene representation
0.912025
Predicate Hierarchies Improve Few-Shot State Classification · ICLR 2025
Machine learning › Graph learning › graph representation learning
structural encoding
0.912025
Homomorphism Counts as Structural Encodings for Graph Learning · ICLR 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.812024
MARPLE: A Benchmark for Long-Horizon Inference · NeurIPS 2024
Machine learning › Graph learning › graph neural network
graph neural network architecture
0.812024
Homomorphism Counts for Graph Neural Networks: All About That Basis · ICML 2024
Machine learning › Graph learning
link prediction
0.712023
Modeling Dynamic Environments with Scene Graph Memory · ICML 2023
Robotics › Robot navigation and mapping
object search
0.712023
Modeling Dynamic Environments with Scene Graph Memory · ICML 2023
Machine learning › Graph learning › link prediction
temporal link prediction
0.712023
Modeling Dynamic Environments with Scene Graph Memory · ICML 2023
Computer vision › Video understanding and tracking › activity recognition
activity parsing
0.612022
MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing · NeurIPS 2022
Computer vision › Video understanding and tracking
activity recognition
0.612022
MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing · NeurIPS 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge engineering › knowledge integration
structured knowledge integration
0.612022
MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing · NeurIPS 2022
Computer vision › Vision and language
video-language model
0.612022
MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing · NeurIPS 2022
Natural language and speech › Language models and text generation
code generation
0.312025
From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries · NeurIPS 2025
Computer vision › Vision and language
multimodal reasoning
0.212024
MARPLE: A Benchmark for Long-Horizon Inference · NeurIPS 2024
Graph algorithms and graph theory
graph homomorphism
0.212024
Homomorphism Counts for Graph Neural Networks: All About That Basis · ICML 2024
Computer vision › Segmentation and scene understanding
scene understanding
0.212023
Modeling Dynamic Environments with Scene Graph Memory · ICML 2023
Computer vision › Video understanding and tracking › activity recognition
hierarchical activity representation
0.212022
MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

graph neural network · 2.2large language model · 1.6homomorphism counting · 1.5self-supervised learning · 0.9program-conditioned pose prediction · 0.9object-centric encoding · 0.9message passing neural network · 0.9hyperbolic embedding · 0.9graph homomorphism counting · 0.9monte carlo simulation · 0.8
YearPublicationVenuePosition
2025 Homomorphism Counts as Structural Encodings for Graph Learning
abstract
Graph Transformers are popular neural networks that extend the well-known Transformer architecture to the graph domain. These architectures operate by applying self-attention on graph nodes and incorporating graph structure through the use of positional encodings (e.g., Laplacian positional encoding) or structural encodings (e.g., random-walk structural encoding). The quality of such encodings is critical, since they provide the necessary \emph{graph inductive biases} to condition the model on graph structure. In this work, we propose \emph{motif structural encoding} (MoSE) as a flexible and powerful structural encoding framework based on counting graph homomorphisms. Theoretically, we compare the expressive power of MoSE to random-walk structural encoding and relate both encodings to the expressive power of standard message passing neural networks. Empirically, we observe that MoSE outperforms other well-known positional and structural encodings across a range of architectures, and it achieves state-of-the-art performance on a widely studied molecular property prediction dataset.
Linus Bao, Emily Jin, Michael M. Bronstein, Ismail Ilkan Ceylan, Matthias Lanzinger
ICLR2
2025 Predicate Hierarchies Improve Few-Shot State Classification
abstract
State classification of objects and their relations is core to many long-horizon tasks, particularly in robot planning and manipulation. However, the combinatorial explosion of possible object-predicate combinations, coupled with the need to adapt to novel real-world environments, makes it a desideratum for state classification models to generalize to novel queries with few examples. To this end, we propose PHIER, which leverages predicate hierarchies to generalize effectively in few-shot scenarios. PHIER uses an object-centric scene encoder, self-supervised losses that infer semantic relations between predicates, and a hyperbolic distance metric that captures hierarchical structure; it learns a structured latent space of image-predicate pairs that guides reasoning over state classification queries. We evaluate PHIER in the CALVIN and BEHAVIOR robotic environments and show that PHIER significantly outperforms existing methods in few-shot, out-of-distribution state classification, and demonstrates strong zero- and few-shot generalization from simulated to real-world tasks. Our results demonstrate that leveraging predicate hierarchies improves performance on state classification tasks with limited data.
Emily Jin, Joy Hsu, Jiajun Wu 0001
ICLR1
2025 From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
abstract
Real-world scenes, such as those in ScanNet, are difficult to capture, with highly limited data available. Generating realistic scenes with varied object poses remains an open and challenging task. In this work, we propose FactoredScenes, a framework that synthesizes realistic 3D scenes by leveraging the underlying structure of rooms while learning the variation of object poses from lived-in scenes. We introduce a factored representation that decomposes scenes into hierarchically organized concepts of room programs and object poses. To encode structure, FactoredScenes learns a library of functions capturing reusable layout patterns from which scenes are drawn, then uses large language models to generate high-level programs, regularized by the learned library. To represent scene variations, FactoredScenes learns a program-conditioned model to hierarchically predict object poses, and retrieves and places 3D objects in a scene. We show that FactoredScenes generates realistic, real-world rooms that are difficult to distinguish from real ScanNet scenes.
Joy Hsu, Emily Jin, Jiajun Wu 0001, Niloy J. Mitra
NeurIPS2
2024 Whodunnit? Inferring what happened from multimodal evidence
Sarah A. Wu, Erik Brockbank, Hannah Cha, Jan-Philipp Fränken, Emily Jin, Zhuoyi Huang, Jiajun Wu 0001, Tobias Gerstenberg
CogSci5
2024 Homomorphism Counts for Graph Neural Networks: All About That Basis
abstract
A large body of work has investigated the properties of graph neural networks and identified several limitations, particularly pertaining to their expressive power. Their inability to count certain patterns (e.g., cycles) in a graph lies at the heart of such limitations, since many functions to be learned rely on the ability of counting such patterns. Two prominent paradigms aim to address this limitation by enriching the graph features with subgraph or homomorphism pattern counts. In this work, we show that both of these approaches are sub-optimal in a certain sense and argue for a more fine-grained approach, which incorporates the homomorphism counts of all structures in the “basis” of the target pattern. This yields strictly more expressive architectures without incurring any additional overhead in terms of computational complexity compared to existing approaches. We prove a series of theoretical results on node-level and graph-level motif parameters and empirically validate them on standard benchmark datasets.
Emily Jin, Michael M. Bronstein, Ismail Ilkan Ceylan, Matthias Lanzinger
ICML1
2024 MARPLE: A Benchmark for Long-Horizon Inference
abstract
Reconstructing past events requires reasoning across long time horizons. To figure out what happened, humans draw on prior knowledge about the world and human behavior and integrate insights from various sources of evidence including visual, language, and auditory cues. We introduce MARPLE, a benchmark for evaluating long-horizon inference capabilities using multi-modal evidence. Our benchmark features agents interacting with simulated households, supporting vision, language, and auditory stimuli, as well as procedurally generated environments and agent behaviors. Inspired by classic ``whodunit'' stories, we ask AI models and human participants to infer which agent caused a change in the environment based on a step-by-step replay of what actually happened. The goal is to correctly identify the culprit as early as possible. Our findings show that human participants outperform both traditional Monte Carlo simulation methods and an LLM baseline (GPT-4) on this task. Compared to humans, traditional inference models are less robust and performant, while GPT-4 has difficulty comprehending environmental changes. We analyze factors influencing inference performance and ablate different modes of evidence, finding that all modes are valuable for performance. Overall, our experiments demonstrate that the long-horizon, multimodal inference tasks in our benchmark present a challenge to current models. Project website: https://marple-benchmark.github.io/.
Emily Jin, Zhuoyi Huang, Jan-Philipp Fränken, Hannah Cha, Erik Brockbank, Sarah A. Wu, Jiajun Wu 0001, Tobias Gerstenberg
NeurIPS1
2023 Modeling Dynamic Environments with Scene Graph Memory
abstract
Embodied AI agents that search for objects in large environments such as households often need to make efficient decisions by predicting object locations based on partial information. We pose this as a new type of link prediction problem: link prediction on partially observable dynamic graphs Our graph is a representation of a scene in which rooms and objects are nodes, and their relationships are encoded in the edges; only parts of the changing graph are known to the agent at each timestep. This partial observability poses a challenge to existing link prediction approaches, which we address. We propose a novel state representation – Scene Graph Memory (SGM) – with captures the agent’s accumulated set of observations, as well as a neural net architecture called a Node Edge Predictor (NEP) that extracts information from the SGM to search efficiently. We evaluate our method in the Dynamic House Simulator, a new benchmark that creates diverse dynamic graphs following the semantic patterns typically seen at homes, and show that NEP can be trained to predict the locations of objects in a variety of environments with diverse object movement dynamics, outperforming baselines both in terms of new scene adaptability and overall accuracy. The codebase and more can be found www.scenegraphmemory.com.
Andrey Kurenkov, Michael Lingelbach, Tanmay Agarwal, Emily Jin, Chengshu Li 0002, Li Fei-Fei 0001, Jiajun Wu 0001, Silvio Savarese, Roberto Martin Martin
ICML4
2022 MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing
abstract
Video-language models (VLMs), large models pre-trained on numerous but noisy video-text pairs from the internet, have revolutionized activity recognition through their remarkable generalization and open-vocabulary capabilities. While complex human activities are often hierarchical and compositional, most existing tasks for evaluating VLMs focus only on high-level video understanding, making it difficult to accurately assess and interpret the ability of VLMs to understand complex and fine-grained human activities. Inspired by the recently proposed MOMA framework, we define activity graphs as a single universal representation of human activities that encompasses video understanding at the activity, sub-activity, and atomic action level. We redefine activity parsing as the overarching task of activity graph generation, requiring understanding human activities across all three levels. To facilitate the evaluation of models on activity parsing, we introduce MOMA-LRG (Multi-Object Multi-Actor Language-Refined Graphs), a large dataset of complex human activities with activity graph annotations that can be readily transformed into natural language sentences. Lastly, we present a model-agnostic and lightweight approach to adapting and evaluating VLMs by incorporating structured knowledge from activity graphs into VLMs, addressing the individual limitations of language and graphical models. We demonstrate strong performance on few-shot activity parsing, and our framework is intended to foster future research in the joint modeling of videos, graphs, and language.
Zelun Luo, Zane Durante, Linden Li, Wanze Xie, Emily Jin, Zhuoyi Huang, Lun Yu Li, Jiajun Wu 0001, Juan Carlos Niebles, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001
NeurIPS6