Qi-Zhi Cai

dblp:220/3290 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorSystems, architecture and hardware · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Video understanding and tracking · 33% Trustworthy machine learning · 13% Generative modeling · 10%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking › motion analysis
human motion analysis
0.412020
Long-Term Human Motion Prediction with Scene Context · ECCV (1) 2020
Computer vision › Video understanding and tracking
human motion prediction
0.412020
Long-Term Human Motion Prediction with Scene Context · ECCV (1) 2020
Computer vision › 3D vision
3d object detection
0.412019
Joint Monocular 3D Vehicle Detection and Tracking · ICCV 2019
Robotics › Autonomous driving
driving policy learning
0.412019
Semantic Predictive Control for Explainable and Efficient Policy Learning · ICRA 2019
Machine learning › Reinforcement learning
model-based reinforcement learning
0.412019
Semantic Predictive Control for Explainable and Efficient Policy Learning · ICRA 2019
Computer vision › Video understanding and tracking
multi-object tracking
0.412019
Joint Monocular 3D Vehicle Detection and Tracking · ICCV 2019
Computer vision › Video understanding and tracking
video prediction
0.412019
Disentangling Propagation and Generation for Video Prediction · ICCV 2019
Robotics › Motion planning and robot control › robot learning
visuomotor learning
0.412019
Deep Object-Centric Policies for Autonomous Driving · ICRA 2019
Security and privacy of machine learning
training-time attack
0.412019
Learning to Confuse: Generating Training Time Adversarial Data with Auto-Encoder · NeurIPS 2019
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.312018
Curriculum Adversarial Training · IJCAI 2018
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.312018
Curriculum Adversarial Training · IJCAI 2018
Machine learning › Learning paradigms
curriculum learning
0.312018
Curriculum Adversarial Training · IJCAI 2018
Computer vision › Segmentation and scene understanding › video segmentation
future semantic segmentation prediction
0.112019
Semantic Predictive Control for Explainable and Efficient Policy Learning · ICRA 2019
Computer vision › Image recognition and object detection › object detection
object detection for autonomous driving
0.112019
Deep Object-Centric Policies for Autonomous Driving · ICRA 2019
Computer vision › Segmentation and scene understanding
semantic segmentation
0.112019
Semantic Predictive Control for Explainable and Efficient Policy Learning · ICRA 2019
Machine learning › Generative modeling
video generation
0.112019
Disentangling Propagation and Generation for Video Prediction · ICCV 2019

Methods — techniques the papers use, named apart from their topics

deep learning · 0.4warping · 0.4trajectory prediction · 0.4sampling-based optimization · 0.4optical flow · 0.4model-based reinforcement learning · 0.4inpainting · 0.4depth-ordering matching · 0.4deep neural network · 0.4constrained optimization · 0.4autoencoder · 0.4alternating updates · 0.4LSTM · 0.4
YearPublicationVenuePosition
2020 Long-Term Human Motion Prediction with Scene Context
Zhe Cao 0003, Karttikeya Mangalam, Qi-Zhi Cai, Minh Vo, Jitendra Malik
ECCV (1)4
2019 Disentangling Propagation and Generation for Video Prediction
abstract
A dynamic scene has two types of elements: those that move fluidly and can be predicted from previous frames, and those which are disoccluded (exposed) and cannot be extrapolated. Prior approaches to video prediction typically learn either to warp or to hallucinate future pixels, but not both. In this paper, we describe a computational model for high-fidelity video prediction which disentangles motion-specific propagation from motion-agnostic generation. We introduce a confidence-aware warping operator which gates the output of pixel predictions from a flow predictor for non-occluded regions and from a context encoder for occluded regions. Moreover, in contrast to prior works where confidence is jointly learned with flow and appearance using a single network, we compute confidence after a warping step, and employ a separate network to inpaint exposed regions. Empirical results on both synthetic and real datasets show that our disentangling approach provides better occlusion maps and produces both sharper and more realistic predictions compared to strong baselines.
Huazhe Xu, Qi-Zhi Cai, Ruth Wang, Fisher Yu 0001, Trevor Darrell
ICCV3
2019 Joint Monocular 3D Vehicle Detection and Tracking
abstract
Vehicle 3D extents and trajectories are critical cues for predicting the future location of vehicles and planning future agent ego-motion based on those predictions. In this paper, we propose a novel online framework for 3D vehicle detection and tracking from monocular videos. The framework can not only associate detections of vehicles in motion over time, but also estimate their complete 3D bounding box information from a sequence of 2D images captured on a moving platform. Our method leverages 3D box depth-ordering matching for robust instance association and utilizes 3D trajectory prediction for re-identification of occluded vehicles. We also design a motion learning module based on an LSTM for more accurate long-term motion extrapolation. Our experiments on simulation, KITTI, and Argoverse datasets show that our 3D tracking pipeline offers robust data association and tracking. On Argoverse, our image-based method is significantly better for tracking 3D vehicles within 30 meters than the LiDAR-centric baseline methods.
Hou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin 0002, Min Sun 0001, Philipp Krähenbühl, Trevor Darrell, Fisher Yu 0001
ICCV2
2019 Semantic Predictive Control for Explainable and Efficient Policy Learning
abstract
Visual anticipation of ego and object motion over a short time horizons is a key feature of human-level performance in complex environments. We propose a driving policy learning framework that predicts feature representations of future visual inputs; our predictive model infers not only future events but also semantics, which provide a visual explanation of policy decisions. Our Semantic Predictive Control (SPC) framework predicts future semantic segmentation and events by aggregating multi-scale feature maps. A guidance model assists action selection and enables efficient sampling-based optimization. Experiments on multiple simulation environments show that networks which implement SPC can outperform existing model-based reinforcement learning algorithms in terms of data efficiency and total rewards while providing clear explanations for the policy's behavior.
Xinlei Pan, Qi-Zhi Cai, John F. Canny, Fisher Yu 0001
ICRA3
2019 Deep Object-Centric Policies for Autonomous Driving
abstract
While learning visuomotor skills in an end-to-end manner is appealing, deep neural networks are often uninterpretable and fail in surprising ways. For robotics tasks, such as autonomous driving, models that explicitly represent objects may be more robust to new scenes and provide intuitive visualizations. We describe a taxonomy of “object-centric” models which leverage both object instances and end-to-end learning. In the Grand Theft Auto V simulator, we show that object-centric models outperform object-agnostic methods in scenes with other vehicles and pedestrians, even with an imperfect detector. We also demonstrate that our architectures perform well on real-world environments by evaluating on the Berkeley DeepDrive Video dataset, where an object-centric model outperforms object-agnostic models in the low-data regimes.
Dequan Wang, Coline Devin, Qi-Zhi Cai, Fisher Yu 0001, Trevor Darrell
ICRA3
2019 Monocular Plan View Networks for Autonomous Driving
abstract
Convolutions on monocular dash cam videos capture spatial invariances in the image plane but do not explicitly reason about distances and depth. We propose a simple transformation of observations into a bird's eye view, also known as plan view, for end-to-end control. We detect vehicles and pedestrians in the first person view and project them into an overhead plan view. This representation provides an abstraction of the environment from which a deep network can easily deduce the positions and directions of entities. Additionally, the plan view enables us to leverage advances in 3D object detection in conjunction with deep policy learning. We evaluate our monocular plan view network on the photo-realistic Grand Theft Auto V simulator. A network using both a plan view and front view causes less than half as many collisions as previous detection-based methods and an order of magnitude fewer collisions than pure pixel-based policies.
Dequan Wang, Coline Devin, Qi-Zhi Cai, Philipp Krähenbühl, Trevor Darrell
IROS3
2019 Learning to Confuse: Generating Training Time Adversarial Data with Auto-Encoder
abstract
In this work, we consider one challenging training time attack by modifying training data with bounded perturbation, hoping to manipulate the behavior (both targeted or non-targeted) of any corresponding trained classifier during test time when facing clean samples. To achieve this, we proposed to use an auto-encoder-like network to generate such adversarial perturbations on the training data together with one imaginary victim differentiable classifier. The perturbation generator will learn to update its weights so as to produce the most harmful noise, aiming to cause the lowest performance for the victim classifier during test time. This can be formulated into a non-linear equality constrained optimization problem. Unlike GANs, solving such problem is computationally challenging, we then proposed a simple yet effective procedure to decouple the alternating updates for the two networks for stability. By teaching the perturbation generator to hijacking the training trajectory of the victim classifier, the generator can thus learn to move against the victim classifier step by step. The method proposed in this paper can be easily extended to the label specific setting where the attacker can manipulate the predictions of the victim classifier according to some predefined rules rather than only making wrong predictions. Experiments on various datasets including CIFAR-10 and a reduced version of ImageNet confirmed the effectiveness of the proposed method and empirical results showed that, such bounded perturbations have good transferability across different types of victim classifiers.
Ji Feng, Qi-Zhi Cai, Zhi-Hua Zhou
NeurIPS2
2018 Curriculum Adversarial Training
abstract
Recently, deep learning has been applied to many security-sensitive applications, such as facial authentication. The existence of adversarial examples hinders such applications. The state-of-the-art result on defense shows that adversarial training can be applied to train a robust model on MNIST against adversarial examples; but it fails to achieve a high empirical worst-case accuracy on a more complex task, such as CIFAR-10 and SVHN. In our work, we propose curriculum adversarial training (CAT) to resolve this issue. The basic idea is to develop a curriculum of adversarial examples generated by attacks with a wide range of strengths. With two techniques to mitigate the catastrophic forgetting and the generalization issues, we demonstrate that CAT can improve the prior art's empirical worst-case accuracy by a large margin of 25% on CIFAR-10 and 35% on SVHN. At the same, the model's performance on non-adversarial inputs is comparable to the state-of-the-art models.
Qi-Zhi Cai, Chang Liu 0021, Dawn Song
IJCAI1