Jiale Zhi

dblp:246/2865 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 20% Deep learning architectures and training · 13% Vision and language · 13%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language pretraining
contrastive vision-language pretraining
0.912025
Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025
Computer vision › 3D vision
depth estimation
0.912025
Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025
Natural language and speech › Language models and text generation › language modeling
multimodal language modeling
0.912025
Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025
Computer vision › Image recognition and object detection
spatial alignment
0.912025
Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025
Machine learning › Deep learning architectures and training
vision encoder
0.912025
Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025
Machine learning › Reinforcement learning › exploration › novelty-based exploration
novelty search
0.412020
Enhanced POET: Open-ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions · ICML 2020
Machine learning › Optimization for machine learning › evolutionary computation
quality-diversity
0.412020
Enhanced POET: Open-ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions · ICML 2020
Machine learning › Reinforcement learning
deep reinforcement learning
0.412019
An Atari Model Zoo for Analyzing, Visualizing, and Comparing Deep Reinforcement Learning Agents · IJCAI 2019
Machine learning › Representation and self-supervised learning
representation analysis
0.412019
An Atari Model Zoo for Analyzing, Visualizing, and Comparing Deep Reinforcement Learning Agents · IJCAI 2019
Computer vision › Video understanding and tracking
video classification
0.312025
Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.112020
Enhanced POET: Open-ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions · ICML 2020

Methods — techniques the papers use, named apart from their topics

contrastive learning · 0.9alignment method · 0.9neural network visualization · 0.8reinforcement learning · 0.4goal-switching · 0.4evolutionary computation · 0.4
YearPublicationVenuePosition
2025 Perception Encoder: The best visual embeddings are not at the output of the network
abstract
We introduce Perception Encoder (PE), a family of state-of-the-art vision encoders for image and video understanding. Traditionally, vision encoders have relied on a variety of pretraining objectives, each excelling at different downstream tasks. Surprisingly, after scaling a carefully tuned image pretraining recipe and refining with a robust video data engine, we find that contrastive vision-language training alone can produce strong, general embeddings for all of these downstream tasks. There is only one caveat: these embeddings are hidden within the intermediate layers of the network. To draw them out, we introduce two alignment methods: language alignment for multimodal language modeling, and spatial alignment for dense prediction. Together, our PE family of models achieves state-of-the-art results on a wide variety of tasks, including zero-shot image and video classification and retrieval; document, image, and video Q&A; and spatial tasks such as detection, tracking, and depth estimation. We release our models, code, and novel dataset of synthetically and human-annotated videos: https://github.com/facebookresearch/perception_models
Daniel Bolya, Po-Yao Huang 0001, Peize Sun, Jang Hyun Cho, Andrea Madotto, Chen Wei 0005, Tengyu Ma 0005, Jiale Zhi, Jathushan Rajasegaran, Hanoona Rasheed, Marco Monteiro, Hu Xu 0001, Shiyu Dong, Nikhila Ravi, Shang-Wen Li 0001, Piotr Dollár, Christoph Feichtenhofer
NeurIPS8
2024 A Biomimetic Robot Crawling Upstream using Adhesive Suckers Inspired by Net-winged Midge Larvae
abstract
Net-winged midge larvae (genus Liponeura) can achieve robust attachment and crawl on the slippery surface in the fast stream with their powerful abdominal suckers. The rigid spine-like structures distributed in the sucker cavity called microtrichia, have been proven to be crucial in the adhesion process. In this work, we carry out a design of the biomimetic sucker with the spine-like structures and then implement various tests using biomimetic suckers to verify the adhesion capacity enhancement brought by spine-like structures. Finally, we assemble the suckers in a quadruped crawling robot capable of locomotion in both aerial and aquatic environments with a speed of 56.4 mm/s (0.225 BL/s) and can crawl upstream against a turbulent flow with a speed of 38.2 mm/s (0.152 BL/s). This study will inspire biomimetic design in future robotics, and pave the way for future robots to realize long-term observation and monitoring in complex environments.
Shuyong Zhao, Jiale Zhi, Chongze Bi
IROS3
2020 Enhanced POET: Open-ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions
abstract
Creating open-ended algorithms, which generate their own never-ending stream of novel and appropriately challenging learning opportunities, could help to automate and accelerate progress in machine learning. A recent step in this direction is the Paired Open-Ended Trailblazer (POET), an algorithm that generates and solves its own challenges, and allows solutions to goal-switch between challenges to avoid local optima. However, the original POET was unable to demonstrate its full creative potential because of limitations of the algorithm itself and because of external issues including a limited problem space and lack of a universal progress measure. Importantly, both limitations pose impediments not only for POET, but for the pursuit of open-endedness in general. Here we introduce and empirically validate two new innovations to the original algorithm, as well as two external innovations designed to help elucidate its full potential. Together, these four advances enable the most open-ended algorithmic demonstration to date. The algorithmic innovations are (1) a domain-general measure of how meaningfully novel new challenges are, enabling the system to potentially create and solve interesting challenges endlessly, and (2) an efficient heuristic for determining when agents should goal-switch from one problem to another (helping open-ended search better scale). Outside the algorithm itself, to enable a more definitive demonstration of open-endedness, we introduce (3) a novel, more flexible way to encode environmental challenges, and (4) a generic measure of the extent to which a system continues to exhibit open-ended innovation. Enhanced POET produces a diverse range of sophisticated behaviors that solve a wide range of environmental challenges, many of which cannot be solved through other means.
Rui Wang 0052, Joel Lehman, Aditya Rawal, Jiale Zhi, Yulun Li, Jeff Clune, Kenneth O. Stanley
ICML4
2019 An Atari Model Zoo for Analyzing, Visualizing, and Comparing Deep Reinforcement Learning Agents
abstract
Much human and computational effort has aimed to improve how deep reinforcement learning (DRL) algorithms perform on benchmarks such as the Atari Learning Environment. Comparatively less effort has focused on understanding what has been learned by such methods, and investigating and comparing the representations learned by different families of DRL algorithms. Sources of friction include the onerous computational requirements, and general logistical and architectural complications for running DRL algorithms at scale. We lessen this friction, by (1) training several algorithms at scale and releasing trained models, (2) integrating with a previous DRL model release, and (3) releasing code that makes it easy for anyone to load, visualize, and analyze such models. This paper introduces the Atari Zoo framework, which contains models trained across benchmark Atari games, in an easy-to-use format, as well as code that implements common modes of analysis and connects such models to a popular neural network visualization library. Further, to demonstrate the potential of this dataset and software package, we show initial quantitative and qualitative comparisons between the performance and representations of several DRL algorithms, highlighting interesting and previously unknown distinctions between them.
Felipe Petroski Such, Vashisht Madhavan, Rosanne Liu, Rui Wang 0052, Pablo Samuel Castro, Yulun Li, Jiale Zhi, Ludwig Schubert, Marc G. Bellemare, Jeff Clune, Joel Lehman
IJCAI7