Zhenyang Feng

dblp:389/4630 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Segmentation and scene understanding · 27% Generative modeling · 27% 3D vision · 21%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 50% Bioinformatics and computational biology · 50%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
collaborative perception
1.122025
Learning 3D Perception from Others' Predictions · ICLR 2025
Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene · CVPR 2025
Computer vision › 3D vision
3d object detection
0.912025
Learning 3D Perception from Others' Predictions · ICLR 2025
Machine learning › Generative modeling › diffusion model
conditional diffusion model
0.912025
Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene · CVPR 2025
Machine learning › Generative modeling › diffusion model › controllable generation
controllable 3d generation
0.912025
Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene · CVPR 2025
Computer vision › Segmentation and scene understanding
pseudo-label learning
0.912025
Learning 3D Perception from Others' Predictions · ICLR 2025
Environmental and earth informatics
biodiversity informatics
0.912025
Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025
Bioinformatics and computational biology
species classification
0.912025
Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025
Machine learning › Trustworthy machine learning › interpretability
explainable AI
0.312025
Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025
Machine learning › Learning paradigms
long-tailed recognition
0.312025
Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025
Computer vision › 3D vision
novel view synthesis
0.312025
Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene · CVPR 2025
Computer vision › 3D vision › point cloud analysis
point cloud perception
0.312025
Learning 3D Perception from Others' Predictions · ICLR 2025

Methods — techniques the papers use, named apart from their topics

machine learning · 1.7computer vision · 1.7transfer learning · 0.9self-training · 0.9pseudo-label refinement · 0.9distance-based curriculum · 0.9diffusion model · 0.9
YearPublicationVenuePosition
2025 Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images
abstract
We introduce Fish-Visual Trait Analysis (Fish-Vista), the first organismal image dataset designed for the analysis of visual traits of aquatic species directly from images using machine learning and computer vision methods. Fish-Vista contains 69,269 annotated images spanning 4,316 fish species, curated and organized to serve three downstream tasks: species classification, trait identification, and trait segmentation. Our work makes two key contributions. First, we provide a fully reproducible data processing pipeline to process fish images sourced from various museum collections, contributing to the advancement of AI in biodiversity science. We annotate the images with carefully curated labels from biological databases and manual annotations to create an AI-ready dataset of visual traits. Second, our work offers fertile grounds for researchers to develop novel methods for a variety of problems in computer vision such as handling long-tailed distributions, out-of-distribution generalization, learning with weak labels, explainable AI, and segmenting small objects. Dataset and code for Fish-Vista are available at https://github.com/Imageomics/Fish-Vista
Kazi Sajeed Mehrab, M. Maruf, Arka Daw, Abhilash Neog, Harish Babu Manogaran, Mridul Khurana, Zhenyang Feng, Bahadir Altintas, Yasin Bakis, Elizabeth G. Campolongo, Matthew J. Thompson, Hilmar Lapp, Tanya Y. Berger-Wolf, Paula M. Mabee, Henry L. Bart Jr., Wei-Lun Chao, Wasila M. Dahdul, Anuj Karpatne
CVPR7
2025 Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene
abstract
Self-driving cars relying solely on ego-centric perception face limitations in sensing, often failing to detect occluded, faraway objects. Collaborative autonomous driving (CAV) seems like a promising direction, but collecting data for development is non-trivial. It requires placing multiple sensor-equipped agents in a real-world driving scene, simultaneously! As such, existing datasets are limited in locations and agents. We introduce a novel surrogate to the rescue, which is to generate realistic perception from different viewpoints in a driving scene, conditioned on a real-world sample—the ego-car’s sensory data. This surrogate has huge potential: it could potentially turn any ego-car dataset into a collaborative driving one to scale up the development of CAV. We present the very first solution, using a combination of simulated collaborative data and real ego-car data. Our method Transfer Your Perspective (TYP) learns a conditioned diffusion model whose output samples are not only realistic but also consistent in both semantics and layouts with the given ego-car data. Empirical results demonstrate TYP’s effectiveness in aiding in a CAV setting. In particular, TYP enables us to (pre-)train collaborative perception algorithms like early and late fusion with little or no real-world collaborative data, greatly facilitating downstream CAV applications.
Tai-Yu Pan, Sooyoung Jeon, Mengdi Fan, Jinsu Yoo, Zhenyang Feng, Mark E. Campbell, Kilian Q. Weinberger, Bharath Hariharan, Wei-Lun Chao
CVPR5
2025 Learning 3D Perception from Others' Predictions
abstract
Accurate 3D object detection in real-world environments requires a huge amount of annotated data with high quality. Acquiring such data is tedious and expensive, and often needs repeated effort when a new sensor is adopted or when the detector is deployed in a new environment. We investigate a new scenario to construct 3D object detectors: *learning from the predictions of a nearby unit that is equipped with an accurate detector.* For example, when a self-driving car enters a new area, it may learn from other traffic participants whose detectors have been optimized for that area. This setting is label-efficient, sensor-agnostic, and communication-efficient: nearby units only need to share the predictions with the ego agent (e.g., car). Naively using the received predictions as ground-truths to train the detector for the ego car, however, leads to inferior performance. We systematically study the problem and identify viewpoint mismatches and mislocalization (due to synchronization and GPS errors) as the main causes, which unavoidably result in false positives, false negatives, and inaccurate pseudo labels. We propose a distance-based curriculum, first learning from closer units with similar viewpoints and subsequently improving the quality of other units' predictions via self-training. We further demonstrate that an effective pseudo label refinement module can be trained with a handful of annotated data, largely reducing the data quantity necessary to train an object detector. We validate our approach on the recently released real-world collaborative driving dataset, using reference cars' predictions as pseudo labels for the ego car. Extensive experiments including several scenarios (e.g., different sensors, detectors, and domains) demonstrate the effectiveness of our approach toward label-efficient learning of 3D perception from other units' predictions.
Jinsu Yoo, Zhenyang Feng, Tai-Yu Pan, Yihong Sun, Cheng Perng Phoo, Xiangyu Chen 0007, Mark E. Campbell, Kilian Q. Weinberger, Bharath Hariharan, Wei-Lun Chao
ICLR2