Zengyi Qin

dblp:230/7736 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
6since 2021 · last 2026
0000-0002-5477-9764ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 39% Reinforcement learning · 15% Generative modeling · 13%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
1.012026
FlowTurbo: Accelerating Flow-Based Image Generation Models via Multi-Stage Refinement · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Generative modeling
normalizing flow
1.012026
FlowTurbo: Accelerating Flow-Based Image Generation Models via Multi-Stage Refinement · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Reinforcement learning
safe reinforcement learning
1.022021
Density Constrained Reinforcement Learning · ICML 2021
Learning Safe Multi-agent Control with Decentralized Neural Barrier Certificates · ICLR 2021
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
monocular 3d object detection
1.022022
MonoGRNet: A General Framework for Monocular 3D Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Triangulation Learning Network: From Monocular to Stereo 3D Object Detection · CVPR 2019
Computer vision › 3D vision
3d object detection
0.822020
Weakly Supervised 3D Object Detection from Point Clouds · ACM Multimedia 2020
Triangulation Learning Network: From Monocular to Stereo 3D Object Detection · CVPR 2019
Computer vision › Image recognition and object detection › object detection
2d object detection
0.722022
MonoGRNet: A General Framework for Monocular 3D Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2022
MonoGRNet: A Geometric Reasoning Network for Monocular 3D Object Localization · AAAI 2019
Machine learning › Reinforcement learning
constrained reinforcement learning
0.512021
Density Constrained Reinforcement Learning · ICML 2021
Knowledge, reasoning and agents › Multi-agent systems
multi-agent control
0.512021
Learning Safe Multi-agent Control with Decentralized Neural Barrier Certificates · ICLR 2021
Robotics › Robot manipulation
grasping
0.412020
KETO: Learning Keypoint Representations for Tool Manipulation · ICRA 2020
Computer vision › 3D vision › 3d object detection
point cloud object detection
0.412020
Weakly Supervised 3D Object Detection from Point Clouds · ACM Multimedia 2020
Robotics › Robot manipulation › object manipulation
tool manipulation
0.412020
KETO: Learning Keypoint Representations for Tool Manipulation · ICRA 2020
Computer vision › 3D vision › 3d object detection › label-efficient 3d object detection
weakly supervised 3d object detection
0.412020
Weakly Supervised 3D Object Detection from Point Clouds · ACM Multimedia 2020
Computer vision › 3D vision
depth estimation
0.412019
MonoGRNet: A Geometric Reasoning Network for Monocular 3D Object Localization · AAAI 2019
Computer vision › 3D vision › 3d object detection › 3d object localization
monocular 3d object localization
0.412019
MonoGRNet: A Geometric Reasoning Network for Monocular 3D Object Localization · AAAI 2019
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
stereo-based 3d object detection
0.412019
Triangulation Learning Network: From Monocular to Stereo 3D Object Detection · CVPR 2019
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.312026
FlowTurbo: Accelerating Flow-Based Image Generation Models via Multi-Stage Refinement · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Robotics › Motion planning and robot control › multi-robot control
decentralized control
0.112021
Learning Safe Multi-agent Control with Decentralized Neural Barrier Certificates · ICLR 2021

Methods — techniques the papers use, named apart from their topics

stage-aware deployment · 1.0sample-aware compilation · 1.0pseudo corrector · 1.0multi-stage refinement · 1.0geometric reasoning · 1.0weakly supervised learning · 0.6neural barrier certificates · 0.5fixed-point iteration · 0.5duality · 0.5control barrier functions · 0.5
YearPublicationVenuePosition
2026 FlowTurbo: Accelerating Flow-Based Image Generation Models via Multi-Stage Refinement
abstract
Building on the. success of diffusion models in visual generation, flow-based models reemerge as another prominent family of generative models that have achieved competitive or better performance in terms of both visual quality and inference speed. By learning the velocity field through flow-matching, flow-based models tend to produce a straighter sampling trajectory, which is advantageous during the sampling process. However, unlike diffusion models for which fast samplers are well-developed, efficient sampling of flow-based generative models has been rarely explored. In this paper, we propose a framework called FlowTurbo to accelerate the sampling of flow-based models while still enhancing the sampling quality. Our primary observation is that the velocity predictor's outputs in the flow-based models will become stable during the sampling, enabling the estimation of velocity via a lightweight velocity refiner. Additionally, we introduce several techniques including a pseudo corrector and sample-aware compilation to further reduce inference time. Since FlowTurbo does not change the multi-step sampling paradigm, it can be effectively applied for various tasks such as image editing, inpainting, etc. Besides, we propose a new multi-stage refinement technique that is designed to reduce the inference costs with large flow-based image generation models. Specifically, the multi-stage refinement split the whole generation procedure on different resolutions, forming a coarse-to-fine text-to-image pipeline. We further adopt a stage-aware deployment strategy that can maximize the inference speed in terms of both latency and throughput. By integrating FlowTurbo into different flow-based models, we obtain an acceleration ratio of 53.1%$\sim$∼58.3% on class-conditional generation and 29.8%$\sim$∼38.5% on text-to-image generation. Notably, FlowTurbo reaches an FID of 2.12 on ImageNet with 100 (ms/img) and FID of 3.93 with 38 (ms/img), achieving the real-time image generation and establishing the new state-of-the-art. Equipped with the recent SD 3.5 Large, we achieved FID of 28.05 with a speed improvement of around 50% on NVIDIA 3090 GPU.
Wenliang Zhao, Minglei Shi, Xumin Yu, Zengyi Qin, Jie Zhou 0001, Jiwen Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 DreamVoice: Text-Guided Voice Conversion
Jiarui Hai, Karan Thakkar, Helin Wang, Zengyi Qin, Mounya Elhilali
INTERSPEECH4
2022 MonoGRNet: A General Framework for Monocular 3D Object Detection
abstract
Detecting and localizing objects in the real 3D space, which plays a crucial role in scene understanding, is particularly challenging given only a monocular image due to the geometric information loss during imagery projection. We propose MonoGRNet for the amodal 3D object detection from a monocular image via geometric reasoning in both the observed 2D projection and the unobserved depth dimension. MonoGRNet decomposes the monocular 3D object detection task into four sub-tasks including 2D object detection, instance-level depth estimation, projected 3D center estimation and local corner regression. The task decomposition significantly facilitates the monocular 3D object detection, allowing the target 3D bounding boxes to be efficiently predicted in a single forward pass, without using object proposals, post-processing or the computationally expensive pixel-level depth estimation utilized by previous methods. In addition, MonoGRNet flexibly adapts to both fully and weakly supervised learning, which improves the feasibility of our framework in diverse settings. Experiments are conducted on KITTI, Cityscapes and MS COCO datasets. Results demonstrate the promising performance of our framework in various scenarios.
Zengyi Qin, Jinglu Wang, Yan Lu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Learning Safe Multi-agent Control with Decentralized Neural Barrier Certificates
Zengyi Qin, Kaiqing Zhang, Yuxiao Chen 0001, Jingkai Chen, Chuchu Fan
ICLR1
2021 Density Constrained Reinforcement Learning
abstract
We study constrained reinforcement learning (CRL) from a novel perspective by setting constraints directly on state density functions, rather than the value functions considered by previous works. State density has a clear physical and mathematical interpretation, and is able to express a wide variety of constraints such as resource limits and safety requirements. Density constraints can also avoid the time-consuming process of designing and tuning cost functions required by value function-based constraints to encode system specifications. We leverage the duality between density functions and Q functions to develop an effective algorithm to solve the density constrained RL problem optimally and the constrains are guaranteed to be satisfied. We prove that the proposed algorithm converges to a near-optimal solution with a bounded error even when the policy update is imperfect. We use a set of comprehensive experiments to demonstrate the advantages of our approach over state-of-the-art CRL methods, with a wide range of density constrained tasks as well as standard CRL benchmarks such as Safety-Gym.
Zengyi Qin, Yuxiao Chen 0001, Chuchu Fan
ICML1
2021 Reactive and Safe Road User Simulations using Neural Barrier Certificates
abstract
Reactive and safe agent modellings are important for nowadays traffic simulator designs and safe planning applications. In this work, we proposed a reactive agent model which can ensure safety without comprising the original purposes, by learning only high-level decisions from expert data and a low level decentralized controller guided by the jointly learned decentralized barrier certificates. Empirical results show that our learned road user simulation models can achieve a significant improvement in safety comparing to state-of-the-art imitation learning and pure control-based methods, while being similar to human agents by having smaller error to the expert data. Moreover, our learned reactive agents are shown to generalize better to unseen traffic conditions, and react better to other road users and therefore can help understand challenging planning problems pragmatically.
Zengyi Qin, Chuchu Fan
IROS2
2020 KETO: Learning Keypoint Representations for Tool Manipulation
abstract
We aim to develop an algorithm for robots to manipulate novel objects as tools for completing different task goals. An efficient and informative representation would facilitate the effectiveness and generalization of such algorithms. For this purpose, we present KETO, a framework of learning keypoint representations of tool-based manipulation. For each task, a set of task-specific keypoints is jointly predicted from 3D point clouds of the tool object by a deep neural network. These keypoints offer a concise and informative description of the object to determine grasps and subsequent manipulation actions. The model is learned from self-supervised robot interactions in the task environment without the need for explicit human annotations. We evaluate our framework in three manipulation tasks with tool use. Our model consistently outperforms state-of-the-art methods in terms of task success rates. Qualitative results of keypoint prediction and tool generation are shown to visualize the learned representations.
Zengyi Qin, Kuan Fang, Yuke Zhu, Li Fei-Fei 0001, Silvio Savarese
ICRA1
2020 Weakly Supervised 3D Object Detection from Point Clouds
abstract
A crucial task in scene understanding is 3D object detection, which aims to detect and localize the 3D bounding boxes of objects belonging to specific classes. Existing 3D object detectors heavily rely on annotated 3D bounding boxes during training, while these annotations could be expensive to obtain and only accessible in limited scenarios. Weakly supervised learning is a promising approach to reducing the annotation requirement, but existing weakly supervised object detectors are mostly for 2D detection rather than 3D. In this work, we propose VS3D, a framework for weakly supervised 3D object detection from point clouds without using any ground truth 3D bounding box for training. First, we introduce an unsupervised 3D proposal module that generates object proposals by leveraging normalized point cloud densities. Second, we present a cross-modal knowledge distillation strategy, where a convolutional neural network learns to predict the final results from the 3D object proposals by querying a teacher network pretrained on image datasets. Comprehensive experiments on the challenging KITTI dataset demonstrate the superior performance of our VS3D in diverse evaluation settings. The source code and pretrained models are publicly available at https://github.com/Zengyi-Qin/Weakly-Supervised-3D-Object-Detection.
Zengyi Qin, Jinglu Wang, Yan Lu 0001
ACM Multimedia1
2019 MonoGRNet: A Geometric Reasoning Network for Monocular 3D Object Localization
abstract
Localizing objects in the real 3D space, which plays a crucial role in scene understanding, is particularly challenging given only a single RGB image due to the geometric information loss during imagery projection. We propose MonoGRNet for the amodal 3D object localization from a monocular RGB image via geometric reasoning in both the observed 2D projection and the unobserved depth dimension. MonoGRNet is a single, unified network composed of four task-specific subnetworks, responsible for 2D object detection, instance depth estimation (IDE), 3D localization and local corner regression. Unlike the pixel-level depth estimation that needs per-pixel annotations, we propose a novel IDE method that directly predicts the depth of the targeting 3D bounding box’s center using sparse supervision. The 3D localization is further achieved by estimating the position in the horizontal and vertical dimensions. Finally, MonoGRNet is jointly learned by optimizing the locations and poses of the 3D bounding boxes in the global context. We demonstrate that MonoGRNet achieves state-of-the-art performance on challenging datasets.
Zengyi Qin, Jinglu Wang, Yan Lu 0001
AAAI1
2019 Triangulation Learning Network: From Monocular to Stereo 3D Object Detection
abstract
In this paper, we study the problem of 3D object detection from stereo images, in which the key challenge is how to effectively utilize stereo information. Different from previous methods using pixel-level depth maps, we propose to employ 3D anchors to explicitly construct object-level correspondences between the regions of interest in stereo images, from which the deep neural network learns to detect and triangulate the targeted object in 3D space. We also introduce a cost-efficient channel reweighting strategy that enhances representational features and weakens noisy signals to facilitate the learning process. All of these are flexibly integrated into a solid baseline detector that inputs monocular images. We demonstrate that both the monocular baseline and the stereo triangulation learning network outperform the prior state-of-the-arts in 3D object detection and localization on the challenging KITTI dataset.
Zengyi Qin, Jinglu Wang, Yan Lu 0001
CVPR1
2019 sEMG-Based Tremor Severity Evaluation for Parkinson's Disease Using a Light-Weight CNN
abstract
We propose a deep learning based approach for quantifying the tremor severity of Parkinson's disease (PD) based on surface electromyography (sEMG). We design the S-Net, a light weight and computational efficient convolutional neural network that learns the similarity between sEMG signals in terms of the tremor severity. Labeled sEMG samples are used for jointly voting for the final results. Experiments on 147 PD patients demonstrate that our approach outperforms traditional methods by a significant margin. In addition, our approach is simple and has potentials in real applications.
Zengyi Qin, Zhenyu Jiang 0002, Jiansheng Chen 0001
IEEE Signal Process. Lett.1