EDBT 2026 Demo / reviewers in the wild / expert
Jiquan Ngiam
dblp:72/8781
· DBLP profile ↗
19ranked-venue papers
4as first author
5since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
18 papers |
3D vision · 29% Autonomous driving · 17% Efficient and distributed learning · 15% |
Topics — the 30 heaviest of 43, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d object detection |
2.4 | 5 | 2022 | DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection · CVPR 2022 3D-MAN: 3D Multi-Frame Attention Network for Object Detection · CVPR 2021 To the Point: Efficient 3D Object Detection in the Range Image With Graph Convolution Kernels · CVPR 2021 |
Robotics › Autonomous driving
trajectory prediction |
1.1 | 2 | 2022 | Scene Transformer: A unified architecture for predicting future trajectories of multiple agents · ICLR 2022 Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset · ICCV 2021 |
Computer vision › Vision and language › cross-modal alignment
cross-modal feature alignment |
0.6 | 1 | 2022 | DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection · CVPR 2022 |
Machine learning › Representation and self-supervised learning › representation matching
feature alignment |
0.6 | 1 | 2022 | DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection · CVPR 2022 |
Computer vision › 3D vision › multimodal perception
LiDAR-camera fusion |
0.6 | 1 | 2022 | DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection · CVPR 2022 |
Robotics › Autonomous driving › trajectory prediction
multi-agent trajectory prediction |
0.6 | 1 | 2022 | Scene Transformer: A unified architecture for predicting future trajectories of multiple agents · ICLR 2022 |
Computer vision › 3D vision › 3d object detection
multimodal 3d object detection |
0.6 | 1 | 2022 | DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection · CVPR 2022 |
Machine learning › Deep learning architectures and training
training optimization |
0.6 | 2 | 2020 | Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout · NeurIPS 2020 On optimization methods for deep learning · ICML 2011 |
Robotics › Autonomous driving › perception
3d perception |
0.5 | 1 | 2021 | 3D-MAN: 3D Multi-Frame Attention Network for Object Detection · CVPR 2021 |
Robotics › Autonomous driving › trajectory prediction
interactive trajectory prediction |
0.5 | 1 | 2021 | Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset · ICCV 2021 |
Computer vision › 3D vision › 3d scene understanding › multi-view understanding › multi-view fusion
multi-view feature aggregation |
0.5 | 1 | 2021 | 3D-MAN: 3D Multi-Frame Attention Network for Object Detection · CVPR 2021 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.5 | 2 | 2019 | CondConv: Conditionally Parameterized Convolutions for Efficient Inference · NeurIPS 2019 Tiled convolutional neural networks · NIPS 2010 |
Computer vision › Image recognition and object detection › object detection
2d object detection |
0.4 | 1 | 2020 | Scalability in Perception for Autonomous Driving: Waymo Open Dataset · CVPR 2020 |
Machine learning › Learning paradigms
curriculum learning |
0.4 | 1 | 2020 | Learning a Multi-Domain Curriculum for Neural Machine Translation · ACL 2020 |
Machine learning › Deep learning architectures and training
data augmentation |
0.4 | 1 | 2020 | Improving 3D Object Detection Through Progressive Population Based Augmentation · ECCV (21) 2020 |
Machine learning › Efficient and distributed learning
data selection |
0.4 | 1 | 2020 | Learning a Multi-Domain Curriculum for Neural Machine Translation · ACL 2020 |
Machine learning › Deep learning architectures and training › training optimization
gradient balancing |
0.4 | 1 | 2020 | Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout · NeurIPS 2020 |
Machine learning › Optimization for machine learning
gradient conflict resolution |
0.4 | 1 | 2020 | Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout · NeurIPS 2020 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.4 | 1 | 2020 | Scalability in Perception for Autonomous Driving: Waymo Open Dataset · CVPR 2020 |
Machine learning › Learning paradigms
multi-task learning |
0.4 | 1 | 2020 | Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout · NeurIPS 2020 |
Robotics › Autonomous driving
perception |
0.4 | 1 | 2020 | Scalability in Perception for Autonomous Driving: Waymo Open Dataset · CVPR 2020 |
Computer vision › 3D vision › 3d object detection
point cloud object detection |
0.4 | 1 | 2020 | Streaming Object Detection for 3-D Point Clouds · ECCV (18) 2020 |
Machine learning › Efficient and distributed learning › adaptive computation
conditional computation |
0.4 | 1 | 2019 | CondConv: Conditionally Parameterized Convolutions for Efficient Inference · NeurIPS 2019 |
Machine learning › Efficient and distributed learning
distributed training |
0.4 | 1 | 2019 | GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019 |
Machine learning › Efficient and distributed learning › dynamic neural network
dynamic convolution |
0.4 | 1 | 2019 | CondConv: Conditionally Parameterized Convolutions for Efficient Inference · NeurIPS 2019 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.4 | 1 | 2019 | CondConv: Conditionally Parameterized Convolutions for Efficient Inference · NeurIPS 2019 |
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training |
0.4 | 1 | 2019 | GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019 |
Machine learning › Efficient and distributed learning › distributed training › model parallelism
pipeline parallelism |
0.4 | 1 | 2019 | GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019 |
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning |
0.2 | 2 | 2011 | Sparse Filtering · NIPS 2011 ICA with Reconstruction Cost for Efficient Overcomplete Feature Learning · NIPS 2011 |
Computer vision › 3D vision › 3d object annotation
3d bounding box annotation |
0.1 | 1 | 2021 | Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset · ICCV 2021 |
Methods — techniques the papers use, named apart from their topics
transformer architecture · 0.6learnable alignment · 0.6inverse augmentation · 0.6cross-attention · 0.6transformer · 0.5pointnet · 0.5graph convolution kernels · 0.5edge convolution · 0.5cross-modality fusion · 0.5attention network · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object DetectionabstractLidars and cameras are critical sensors that provide complementary information for 3D detection in autonomous driving. While prevalent multi-modal methods [34], [36] simply decorate raw lidar point clouds with camera features and feed them directly to existing 3D detection models, our study shows that fusing camera features with deep lidar features instead of raw points, can lead to better performance. However, as those features are often augmented and aggregated, a key challenge in fusion is how to effectively align the transformed features from two modalities. In this paper, we propose two novel techniques: InverseAug that inverses geometric-related augmentations, e.g., rotation, to enable accurate geometric alignment between lidar points and image pixels, and LearnableAlign that leverages cross-attention to dynamically capture the correlations between image and lidar features during fusion. Based on InverseAug and LearnableAlign, we develop a family of generic multi-modal 3D detection models named DeepFusion, which is more accurate than previous methods. For example, DeepFusion improves Point-Pillars, CenterPoint, and 3D-MAN baselines on Pedestrian detection for 6.7,8.9, and 6.2 LEVEL_2 APH, respectively. Notably, our models achieve state-of-the-art performance on Waymo Open Dataset, and show strong model robustness against input corruptions and out-of-distribution data. Code will be publicly available at https://github.com/tensorflow/lingvo. Yingwei Li 0002, Adams Wei Yu, Tianjian Meng, Benjamin Caine, Jiquan Ngiam, Daiyi Peng, Junyang Shen, Yifeng Lu, Denny Zhou, Quoc V. Le, Alan L. Yuille, Mingxing Tan |
CVPR | 5 |
| 2022 | Scene Transformer: A unified architecture for predicting future trajectories of multiple agents
Jiquan Ngiam, Vijay Vasudevan, Benjamin Caine, Hao-Tien Chiang, Jeffrey Ling, Rebecca Roelofs, Alex Bewley, Chenxi Liu 0001, Ashish Venugopal, David J. Weiss, Benjamin Sapp, Jonathon Shlens |
ICLR | 1 |
| 2021 | To the Point: Efficient 3D Object Detection in the Range Image With Graph Convolution Kernelsabstract3D object detection is vital for many robotics applications. For tasks where a 2D perspective range image exists, we propose to learn a 3D representation directly from this range image view. To this end, we designed a 2D convolutional network architecture that carries the 3D spherical coordinates of each pixel throughout the network. Its layers can consume any arbitrary convolution kernel in place of the default inner product kernel and exploit the underlying local geometry around each pixel. We outline four such kernels: a dense kernel according to the bag-of-words paradigm, and three graph kernels inspired by recent graph neural network advances: the Transformer, the PointNet, and the Edge Convolution. We also explore cross-modality fusion with the camera image, facilitated by operating in the perspective range image view. Our method performs competitively on the Waymo Open Dataset and improves the state-of-the-art AP for pedestrian detection from 69.7% to 75.5%. It is also efficient in that our smallest model, which still outperforms the popular PointPillars in quality, requires 180 times fewer FLOPS and model parameters. Yuning Chai, Jiquan Ngiam, Weiyue Wang 0002, Benjamin Caine, Vijay Vasudevan, Dragomir Anguelov |
CVPR | 3 |
| 2021 | 3D-MAN: 3D Multi-Frame Attention Network for Object Detectionabstract3D object detection is an important module in autonomous driving and robotics. However, many existing methods focus on using single frames to perform 3D detection, and do not fully utilize information from multiple frames. In this paper, we present 3D-MAN: a 3D multi-frame attention network that effectively aggregates features from multiple perspectives and achieves state-of-the-art performance on Waymo Open Dataset. 3D-MAN first uses a novel fast single-frame detector to produce box proposals. The box proposals and their corresponding feature maps are then stored in a memory bank. We design a multi-view alignment and aggregation module, using attention networks, to extract and aggregate the temporal features stored in the memory bank. This effectively combines the features coming from different perspectives of the scene. We demonstrate the effectiveness of our approach on the large-scale complex Waymo Open Dataset, achieving state-of-the-art results compared to published single-frame and multi-frame methods. Zetong Yang, Jiquan Ngiam |
CVPR | 4 |
| 2021 | Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetabstractAs autonomous driving systems mature, motion forecasting has received increasing attention as a critical requirement for planning. Of particular importance are interactive situations such as merges, unprotected turns, etc., where predicting individual object motion is not sufficient. Joint predictions of multiple objects are required for effective route planning. There has been a critical need for high-quality motion data that is rich in both interactions and annotation to develop motion planning models. In this work, we introduce the most diverse interactive motion dataset to our knowledge, and provide specific labels for interacting objects suitable for developing joint prediction models. With over 100,000 scenes, each 20 seconds long at 10 Hz, our new dataset contains more than 570 hours of unique data over 1750 km of roadways. It was collected by mining for interesting interactions between vehicles, pedestrians, and cyclists across six cities within the United States. We use a high-accuracy 3D auto-labeling system to generate high quality 3D bounding boxes for each road agent, and provide corresponding high definition 3D maps for each scene. Furthermore, we introduce a new set of metrics that provides a comprehensive evaluation of both single agent and joint agent interaction motion forecasting models. Finally, we provide strong baseline models for individual-agent prediction and joint-prediction. We hope that this new large-scale interactive motion dataset will provide new opportunities for advancing motion forecasting models. Scott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu 0001, Hang Zhao 0021, Sabeek Pradhan, Yuning Chai, Benjamin Sapp, Charles R. Qi, Zoey Yang, Aurelien Chouard, Jiquan Ngiam, Vijay Vasudevan, Alexander McCauley, Jonathon Shlens, Dragomir Anguelov |
ICCV | 14 |
| 2020 | Learning a Multi-Domain Curriculum for Neural Machine TranslationabstractMost data selection research in machine translation focuses on improving a single domain.We perform data selection for multiple domains at once.This is achieved by carefully introducing instance-level domain-relevance features and automatically constructing a training curriculum to gradually concentrate on multi-domain relevant and noise-reduced data batches.Both the choice of features and the use of curriculum are crucial for balancing and improving all domains, including out-ofdomain.In large-scale experiments, the multidomain curriculum simultaneously reaches or outperforms the individual performance and brings solid gains over no-curriculum training. Wei Wang 0236, Ye Tian 0022, Jiquan Ngiam, Yinfei Yang, Isaac Caswell, Zarana Parekh |
ACL | 3 |
| 2020 | Scalability in Perception for Autonomous Driving: Waymo Open DatasetabstractThe research community has increasing interest in autonomous driving research, despite the resource intensity of obtaining representative real world data. Existing self-driving datasets are limited in the scale and variation of the environments they capture, even though generalization within and between operating regions is crucial to the over-all viability of the technology. In an effort to help align the research community’s contributions with real-world self-driving problems, we introduce a new large scale, high quality, diverse dataset. Our new dataset consists of 1150 scenes that each span 20 seconds, consisting of well synchronized and calibrated high quality LiDAR and camera data captured across a range of urban and suburban geographies. It is 15x more diverse than the largest camera+LiDAR dataset available based on our proposed diversity metric. We exhaustively annotated this data with 2D (camera image) and 3D (LiDAR) bounding boxes, with consistent identifiers across frames. Finally, we provide strong baselines for 2D as well as 3D detection and tracking tasks. We further study the effects of dataset size and generalization across geographies on 3D detection methods. Find data, code and more up-to-date information at http://www.waymo.com/open. Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han 0002, Jiquan Ngiam, Hang Zhao 0021, Aleksei Timofeev, Scott Ettinger, Maxim Krivokon, Amy Gao, Yu Zhang 0033, Jonathon Shlens, Dragomir Anguelov |
CVPR | 13 |
| 2020 | Improving 3D Object Detection Through Progressive Population Based Augmentation
Shuyang Cheng, Zhaoqi Leng, Ekin Dogus Cubuk, Barret Zoph, Chunyan Bai, Jiquan Ngiam, Benjamin Caine, Vijay Vasudevan, Quoc V. Le, Jonathon Shlens, Dragomir Anguelov |
ECCV (21) | 6 |
| 2020 | Streaming Object Detection for 3-D Point Clouds
Wei Han 0002, Benjamin Caine, Brandon Yang, Christoph Sprunk, Ouais Alsharif, Jiquan Ngiam, Vijay Vasudevan, Jonathon Shlens |
ECCV (18) | 7 |
| 2020 | Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutabstractThe vast majority of deep models use multiple gradient signals, typically corresponding to a sum of multiple loss terms, to update a shared set of trainable weights. However, these multiple updates can impede optimal training by pulling the model in conflicting directions. We present Gradient Sign Dropout (GradDrop), a probabilistic masking procedure which samples gradients at an activation layer based on their level of consistency. GradDrop is implemented as a simple deep layer that can be used in any deep net and synergizes with other gradient balancing approaches. We show that GradDrop outperforms the state-of-the-art multiloss methods within traditional multitask and transfer learning settings, and we discuss how GradDrop reveals links between optimal multiloss training and gradient stochasticity. Jiquan Ngiam, Yanping Huang, Thang Luong, Henrik Kretzschmar, Yuning Chai, Dragomir Anguelov |
NeurIPS | 2 |
| 2019 | GPipe: Efficient Training of Giant Neural Networks using Pipeline ParallelismabstractScaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks. In many cases, increasing model capacity beyond the memory limit of a single accelerator has required developing special algorithms or infrastructure. These solutions are often architecture-specific and do not transfer to other machine learning tasks. To address the need for efficient and task-independent model parallelism, we introduce TensorPipe, a pipeline parallelism library that allows scaling any network that can be expressed as a sequence of layers. By pipelining different sub-sequences of layers on separate accelerators, TensorPipe provides the flexibility of scaling a variety of different networks to gigantic sizes efficiently. Moreover, TensorPipe utilizes a novel batch-splitting pipelining algorithm, resulting in almost linear speedup when a model is partitioned across multiple accelerators. We demonstrate the advantages of TensorPipe by training large-scale neural networks on two different tasks with distinct network architectures: (i)Image Classification: We train a 557-million-parameter AmoebaNet model and attain a top-1 accuracy of 84.4% on ImageNet-2012, (ii)Multilingual Neural Machine Translation: We train a single 6-billion-parameter, 128-layer Transformer model on a corpus spanning over 100 languages and achieve better quality than all bilingual models. Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le |
NeurIPS | 8 |
| 2019 | CondConv: Conditionally Parameterized Convolutions for Efficient InferenceabstractConvolutional layers are one of the basic building blocks of modern deep neural networks. One fundamental assumption is that convolutional kernels should be shared for all examples in a dataset. We propose conditionally parameterized convolutions (CondConv), which learn specialized convolutional kernels for each example. Replacing normal convolutions with CondConv enables us to increase the size and capacity of a network, while maintaining efficient inference. We demonstrate that scaling networks with CondConv improves the performance and inference cost trade-off of several existing convolutional neural network architectures on both classification and detection tasks. On ImageNet classification, our CondConv approach applied to EfficientNet-B0 achieves state-ofthe-art performance of 78.3% accuracy with only 413M multiply-adds. Code and checkpoints for the CondConv Tensorflow layer and CondConv-EfficientNet models are available at: https://github.com/tensorflow/tpu/tree/master/ models/official/efficientnet/condconv. Brandon Yang, Gabriel Bender, Quoc V. Le, Jiquan Ngiam |
NeurIPS | 4 |
| 2017 | Real-time programming exercise feedback in MOOCs
Andy Nguyen, Amory Schlender, Jiquan Ngiam |
EDM | 4 |
| 2011 | On optimization methods for deep learning
Quoc V. Le, Jiquan Ngiam, Adam Coates 0002, Ahbik Lahiri, Bobby Prochnow, Andrew Y. Ng |
ICML | 2 |
| 2011 | Learning Deep Energy Models
Jiquan Ngiam, Pang Wei Koh, Andrew Y. Ng |
ICML | 1 |
| 2011 | Multimodal Deep Learning
Jiquan Ngiam, Aditya Khosla, Juhan Nam, Honglak Lee, Andrew Y. Ng |
ICML | 1 |
| 2011 | ICA with Reconstruction Cost for Efficient Overcomplete Feature LearningabstractIndependent Components Analysis (ICA) and its variants have been successfully used for unsupervised feature learning. However, standard ICA requires an orthonoramlity constraint to be enforced, which makes it difficult to learn overcomplete features. In addition, ICA is sensitive to whitening. These properties make it challenging to scale ICA to high dimensional data. In this paper, we propose a robust soft reconstruction cost for ICA that allows us to learn highly overcomplete sparse features even on unwhitened data. Our formulation reveals formal connections between ICA and sparse autoencoders, which have previously been observed only empirically. Our algorithm can be used in conjunction with off-the-shelf fast unconstrained optimizers. We show that the soft reconstruction cost can also be used to prevent replicated features in tiled convolutional neural networks. Using our method to learn highly overcomplete sparse features and tiled convolutional neural networks, we obtain competitive performances on a wide variety of object recognition tasks. We achieve state-of-the-art test accuracies on the STL-10 and Hollywood2 datasets. Quoc V. Le, Alexandre Karpenko, Jiquan Ngiam, Andrew Y. Ng |
NIPS | 3 |
| 2011 | Sparse FilteringabstractUnsupervised feature learning has been shown to be effective at learning representations that perform well on image, video and audio classification. However, many existing feature learning algorithms are hard to use and require extensive hyperparameter tuning. In this work, we present sparse filtering, a simple new algorithm which is efficient and only has one hyperparameter, the number of features to learn. In contrast to most other feature learning methods, sparse filtering does not explicitly attempt to construct a model of the data distribution. Instead, it optimizes a simple cost function -- the sparsity of L2-normalized features -- which can easily be implemented in a few lines of MATLAB code. Sparse filtering scales gracefully to handle high-dimensional inputs, and can also be used to learn meaningful features in additional layers with greedy layer-wise stacking. We evaluate sparse filtering on natural images, object classification (STL-10), and phone classification (TIMIT), and show that our method works well on a range of different modalities. Jiquan Ngiam, Pang Wei Koh, Sonia A. Bhaskar, Andrew Y. Ng |
NIPS | 1 |
| 2010 | Tiled convolutional neural networksabstractConvolutional neural networks (CNNs) have been successfully applied to many tasks such as digit and object recognition. Using convolutional (tied) weights significantly reduces the number of parameters that have to be learned, and also allows translational invariance to be hard-coded into the architecture. In this paper, we consider the problem of learning invariances, rather than relying on hard-coding. We propose tiled convolution neural networks (Tiled CNNs), which use a regular “tiled” pattern of tied weights that does not require that adjacent hidden units share identical weights, but instead requires only that hidden units k steps away from each other to have tied weights. By pooling over neighboring units, this architecture is able to learn complex invariances (such as scale and rotational invariance) beyond translational invariance. Further, it also enjoys much of CNNs’ advantage of having a relatively small number of learned parameters (such as ease of learning and greater scalability). We provide an efficient learning algorithm for Tiled CNNs based on Topographic ICA, and show that learning complex invariant features allows us to achieve highly competitive results for both the NORB and CIFAR-10 datasets. Quoc V. Le, Jiquan Ngiam, Daniel Jin hao Chia, Pang Wei Koh, Andrew Y. Ng |
NIPS | 2 |