Jize Cao

dblp:265/5785 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Representation and self-supervised learning · 28% Vision and language · 22% Trustworthy machine learning · 14%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
sequence modeling
0.612022
Symbolic Brittleness in Sequence Models: On Systematic Generalization in Symbolic Mathematics · AAAI 2022
Machine learning › Representation and self-supervised learning
systematic generalization
0.612022
Symbolic Brittleness in Sequence Models: On Systematic Generalization in Symbolic Mathematics · AAAI 2022
Algorithms and data structures
symbolic computation
0.612022
Symbolic Brittleness in Sequence Models: On Systematic Generalization in Symbolic Mathematics · AAAI 2022
Algorithms and data structures › symbolic computation
symbolic integration
0.612022
Symbolic Brittleness in Sequence Models: On Systematic Generalization in Symbolic Mathematics · AAAI 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.512021
MERLOT: Multimodal Neural Script Knowledge Models · NeurIPS 2021
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining
0.512021
MERLOT: Multimodal Neural Script Knowledge Models · NeurIPS 2021
Computer vision › Vision and language › vision-language pretraining
video-language pre-training
0.512021
MERLOT: Multimodal Neural Script Knowledge Models · NeurIPS 2021
Computer vision › Video understanding and tracking
video question answering
0.512021
MERLOT: Multimodal Neural Script Knowledge Models · NeurIPS 2021
Machine learning › Trustworthy machine learning
interpretability
0.412020
Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models · ECCV (6) 2020
Computer vision › Vision and language
vision-language model
0.412020
Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models · ECCV (6) 2020
Machine learning › Trustworthy machine learning
robustness
0.212022
Symbolic Brittleness in Sequence Models: On Systematic Generalization in Symbolic Mathematics · AAAI 2022
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
multimodal self-supervised learning
0.112021
MERLOT: Multimodal Neural Script Knowledge Models · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

sequence-to-sequence model · 1.1genetic algorithm · 1.1temporal objectives · 0.5self-supervised learning · 0.5contrastive learning · 0.5probing · 0.4
YearPublicationVenuePosition
2022 Symbolic Brittleness in Sequence Models: On Systematic Generalization in Symbolic Mathematics
abstract
Neural sequence models trained with maximum likelihood estimation have led to breakthroughs in many tasks, where success is defined by the gap between training and test performance. However, their ability to achieve stronger forms of generalization remains unclear. We consider the problem of symbolic mathematical integration, as it requires generalizing systematically beyond the training set. We develop a methodology for evaluating generalization that takes advantage of the problem domain's structure and access to a verifier. Despite promising in-distribution performance of sequence-to-sequence models in this domain, we demonstrate challenges in achieving robustness, compositionality, and out-of-distribution generalization, through both carefully constructed manual test suites and a genetic algorithm that automatically finds large collections of failures in a controllable manner. Our investigation highlights the difficulty of generalizing well with the predominant modeling and learning approach, and the importance of evaluating beyond the test set, across different aspects of generalization.
Sean Welleck, Peter West, Jize Cao, Yejin Choi 0001
AAAI3
2021 MERLOT: Multimodal Neural Script Knowledge Models
abstract
As humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future. We introduce MERLOT, a model that learns multimodal script knowledge by watching millions of YouTube videos with transcribed speech -- in an entirely label-free, self-supervised manner. By pretraining with a mix of both frame-level (spatial) and video-level (temporal) objectives, our model not only learns to match images to temporally corresponding words, but also to contextualize what is happening globally over time. As a result, MERLOT exhibits strong out-of-the-box representations of temporal commonsense, and achieves state-of-the-art performance on 12 different video QA datasets when finetuned. It also transfers well to the world of static images, allowing models to reason about the dynamic context behind visual scenes. On Visual Commonsense Reasoning, MERLOT~answers questions correctly with 80.6\% accuracy, outperforming state-of-the-art models of similar size by over 3\%, even those that make heavy use of auxiliary supervised data (like object bounding boxes).Ablation analyses demonstrate the complementary importance of: 1) training on videos versus static images; 2) scaling the magnitude and diversity of the pretraining video corpus; and 3) using diverse objectives that encourage full-stack multimodal reasoning, from the recognition to cognition level.
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jize Cao, Ali Farhadi, Yejin Choi 0001
NeurIPS6
2020 Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models
Jize Cao, Zhe Gan, Yu Cheng 0001, Licheng Yu, Yen-Chun Chen 0001, Jingjing Liu 0001
ECCV (6)1