EDBT 2026 Demo / reviewers in the wild / expert
Qi-Zhi Cai
dblp:220/3290
· DBLP profile ↗
8ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorSystems, architecture and hardware · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Video understanding and tracking · 33% Trustworthy machine learning · 13% Generative modeling · 10% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking › motion analysis
human motion analysis |
0.4 | 1 | 2020 | Long-Term Human Motion Prediction with Scene Context · ECCV (1) 2020 |
Computer vision › Video understanding and tracking
human motion prediction |
0.4 | 1 | 2020 | Long-Term Human Motion Prediction with Scene Context · ECCV (1) 2020 |
Computer vision › 3D vision
3d object detection |
0.4 | 1 | 2019 | Joint Monocular 3D Vehicle Detection and Tracking · ICCV 2019 |
Robotics › Autonomous driving
driving policy learning |
0.4 | 1 | 2019 | Semantic Predictive Control for Explainable and Efficient Policy Learning · ICRA 2019 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.4 | 1 | 2019 | Semantic Predictive Control for Explainable and Efficient Policy Learning · ICRA 2019 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.4 | 1 | 2019 | Joint Monocular 3D Vehicle Detection and Tracking · ICCV 2019 |
Computer vision › Video understanding and tracking
video prediction |
0.4 | 1 | 2019 | Disentangling Propagation and Generation for Video Prediction · ICCV 2019 |
Robotics › Motion planning and robot control › robot learning
visuomotor learning |
0.4 | 1 | 2019 | Deep Object-Centric Policies for Autonomous Driving · ICRA 2019 |
Security and privacy of machine learning
training-time attack |
0.4 | 1 | 2019 | Learning to Confuse: Generating Training Time Adversarial Data with Auto-Encoder · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.3 | 1 | 2018 | Curriculum Adversarial Training · IJCAI 2018 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.3 | 1 | 2018 | Curriculum Adversarial Training · IJCAI 2018 |
Machine learning › Learning paradigms
curriculum learning |
0.3 | 1 | 2018 | Curriculum Adversarial Training · IJCAI 2018 |
Computer vision › Segmentation and scene understanding › video segmentation
future semantic segmentation prediction |
0.1 | 1 | 2019 | Semantic Predictive Control for Explainable and Efficient Policy Learning · ICRA 2019 |
Computer vision › Image recognition and object detection › object detection
object detection for autonomous driving |
0.1 | 1 | 2019 | Deep Object-Centric Policies for Autonomous Driving · ICRA 2019 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.1 | 1 | 2019 | Semantic Predictive Control for Explainable and Efficient Policy Learning · ICRA 2019 |
Machine learning › Generative modeling
video generation |
0.1 | 1 | 2019 | Disentangling Propagation and Generation for Video Prediction · ICCV 2019 |
Methods — techniques the papers use, named apart from their topics
deep learning · 0.4warping · 0.4trajectory prediction · 0.4sampling-based optimization · 0.4optical flow · 0.4model-based reinforcement learning · 0.4inpainting · 0.4depth-ordering matching · 0.4deep neural network · 0.4constrained optimization · 0.4autoencoder · 0.4alternating updates · 0.4LSTM · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Long-Term Human Motion Prediction with Scene Context
Zhe Cao 0003, Karttikeya Mangalam, Qi-Zhi Cai, Minh Vo, Jitendra Malik |
ECCV (1) | 4 |
| 2019 | Disentangling Propagation and Generation for Video PredictionabstractA dynamic scene has two types of elements: those that move fluidly and can be predicted from previous frames, and those which are disoccluded (exposed) and cannot be extrapolated. Prior approaches to video prediction typically learn either to warp or to hallucinate future pixels, but not both. In this paper, we describe a computational model for high-fidelity video prediction which disentangles motion-specific propagation from motion-agnostic generation. We introduce a confidence-aware warping operator which gates the output of pixel predictions from a flow predictor for non-occluded regions and from a context encoder for occluded regions. Moreover, in contrast to prior works where confidence is jointly learned with flow and appearance using a single network, we compute confidence after a warping step, and employ a separate network to inpaint exposed regions. Empirical results on both synthetic and real datasets show that our disentangling approach provides better occlusion maps and produces both sharper and more realistic predictions compared to strong baselines. Huazhe Xu, Qi-Zhi Cai, Ruth Wang, Fisher Yu 0001, Trevor Darrell |
ICCV | 3 |
| 2019 | Joint Monocular 3D Vehicle Detection and TrackingabstractVehicle 3D extents and trajectories are critical cues for predicting the future location of vehicles and planning future agent ego-motion based on those predictions. In this paper, we propose a novel online framework for 3D vehicle detection and tracking from monocular videos. The framework can not only associate detections of vehicles in motion over time, but also estimate their complete 3D bounding box information from a sequence of 2D images captured on a moving platform. Our method leverages 3D box depth-ordering matching for robust instance association and utilizes 3D trajectory prediction for re-identification of occluded vehicles. We also design a motion learning module based on an LSTM for more accurate long-term motion extrapolation. Our experiments on simulation, KITTI, and Argoverse datasets show that our 3D tracking pipeline offers robust data association and tracking. On Argoverse, our image-based method is significantly better for tracking 3D vehicles within 30 meters than the LiDAR-centric baseline methods. Hou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin 0002, Min Sun 0001, Philipp Krähenbühl, Trevor Darrell, Fisher Yu 0001 |
ICCV | 2 |
| 2019 | Semantic Predictive Control for Explainable and Efficient Policy LearningabstractVisual anticipation of ego and object motion over a short time horizons is a key feature of human-level performance in complex environments. We propose a driving policy learning framework that predicts feature representations of future visual inputs; our predictive model infers not only future events but also semantics, which provide a visual explanation of policy decisions. Our Semantic Predictive Control (SPC) framework predicts future semantic segmentation and events by aggregating multi-scale feature maps. A guidance model assists action selection and enables efficient sampling-based optimization. Experiments on multiple simulation environments show that networks which implement SPC can outperform existing model-based reinforcement learning algorithms in terms of data efficiency and total rewards while providing clear explanations for the policy's behavior. Xinlei Pan, Qi-Zhi Cai, John F. Canny, Fisher Yu 0001 |
ICRA | 3 |
| 2019 | Deep Object-Centric Policies for Autonomous DrivingabstractWhile learning visuomotor skills in an end-to-end manner is appealing, deep neural networks are often uninterpretable and fail in surprising ways. For robotics tasks, such as autonomous driving, models that explicitly represent objects may be more robust to new scenes and provide intuitive visualizations. We describe a taxonomy of “object-centric” models which leverage both object instances and end-to-end learning. In the Grand Theft Auto V simulator, we show that object-centric models outperform object-agnostic methods in scenes with other vehicles and pedestrians, even with an imperfect detector. We also demonstrate that our architectures perform well on real-world environments by evaluating on the Berkeley DeepDrive Video dataset, where an object-centric model outperforms object-agnostic models in the low-data regimes. Dequan Wang, Coline Devin, Qi-Zhi Cai, Fisher Yu 0001, Trevor Darrell |
ICRA | 3 |
| 2019 | Monocular Plan View Networks for Autonomous DrivingabstractConvolutions on monocular dash cam videos capture spatial invariances in the image plane but do not explicitly reason about distances and depth. We propose a simple transformation of observations into a bird's eye view, also known as plan view, for end-to-end control. We detect vehicles and pedestrians in the first person view and project them into an overhead plan view. This representation provides an abstraction of the environment from which a deep network can easily deduce the positions and directions of entities. Additionally, the plan view enables us to leverage advances in 3D object detection in conjunction with deep policy learning. We evaluate our monocular plan view network on the photo-realistic Grand Theft Auto V simulator. A network using both a plan view and front view causes less than half as many collisions as previous detection-based methods and an order of magnitude fewer collisions than pure pixel-based policies. Dequan Wang, Coline Devin, Qi-Zhi Cai, Philipp Krähenbühl, Trevor Darrell |
IROS | 3 |
| 2019 | Learning to Confuse: Generating Training Time Adversarial Data with Auto-EncoderabstractIn this work, we consider one challenging training time attack by modifying training data with bounded perturbation, hoping to manipulate the behavior (both targeted or non-targeted) of any corresponding trained classifier during test time when facing clean samples. To achieve this, we proposed to use an auto-encoder-like network to generate such adversarial perturbations on the training data together with one imaginary victim differentiable classifier. The perturbation generator will learn to update its weights so as to produce the most harmful noise, aiming to cause the lowest performance for the victim classifier during test time. This can be formulated into a non-linear equality constrained optimization problem. Unlike GANs, solving such problem is computationally challenging, we then proposed a simple yet effective procedure to decouple the alternating updates for the two networks for stability. By teaching the perturbation generator to hijacking the training trajectory of the victim classifier, the generator can thus learn to move against the victim classifier step by step. The method proposed in this paper can be easily extended to the label specific setting where the attacker can manipulate the predictions of the victim classifier according to some predefined rules rather than only making wrong predictions. Experiments on various datasets including CIFAR-10 and a reduced version of ImageNet confirmed the effectiveness of the proposed method and empirical results showed that, such bounded perturbations have good transferability across different types of victim classifiers. Ji Feng, Qi-Zhi Cai, Zhi-Hua Zhou |
NeurIPS | 2 |
| 2018 | Curriculum Adversarial TrainingabstractRecently, deep learning has been applied to many security-sensitive applications, such as facial authentication. The existence of adversarial examples hinders such applications. The state-of-the-art result on defense shows that adversarial training can be applied to train a robust model on MNIST against adversarial examples; but it fails to achieve a high empirical worst-case accuracy on a more complex task, such as CIFAR-10 and SVHN. In our work, we propose curriculum adversarial training (CAT) to resolve this issue. The basic idea is to develop a curriculum of adversarial examples generated by attacks with a wide range of strengths. With two techniques to mitigate the catastrophic forgetting and the generalization issues, we demonstrate that CAT can improve the prior art's empirical worst-case accuracy by a large margin of 25% on CIFAR-10 and 35% on SVHN. At the same, the model's performance on non-adversarial inputs is comparable to the state-of-the-art models. Qi-Zhi Cai, Chang Liu 0021, Dawn Song |
IJCAI | 1 |