Jiquan Ngiam

dblp:72/8781 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
5since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
18 papers
3D vision · 29% Autonomous driving · 17% Efficient and distributed learning · 15%

Topics — the 30 heaviest of 43, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
2.452022
DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection · CVPR 2022
3D-MAN: 3D Multi-Frame Attention Network for Object Detection · CVPR 2021
To the Point: Efficient 3D Object Detection in the Range Image With Graph Convolution Kernels · CVPR 2021
Robotics › Autonomous driving
trajectory prediction
1.122022
Scene Transformer: A unified architecture for predicting future trajectories of multiple agents · ICLR 2022
Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset · ICCV 2021
Computer vision › Vision and language › cross-modal alignment
cross-modal feature alignment
0.612022
DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection · CVPR 2022
Machine learning › Representation and self-supervised learning › representation matching
feature alignment
0.612022
DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection · CVPR 2022
Computer vision › 3D vision › multimodal perception
LiDAR-camera fusion
0.612022
DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection · CVPR 2022
Robotics › Autonomous driving › trajectory prediction
multi-agent trajectory prediction
0.612022
Scene Transformer: A unified architecture for predicting future trajectories of multiple agents · ICLR 2022
Computer vision › 3D vision › 3d object detection
multimodal 3d object detection
0.612022
DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection · CVPR 2022
Machine learning › Deep learning architectures and training
training optimization
0.622020
Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout · NeurIPS 2020
On optimization methods for deep learning · ICML 2011
Robotics › Autonomous driving › perception
3d perception
0.512021
3D-MAN: 3D Multi-Frame Attention Network for Object Detection · CVPR 2021
Robotics › Autonomous driving › trajectory prediction
interactive trajectory prediction
0.512021
Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset · ICCV 2021
Computer vision › 3D vision › 3d scene understanding › multi-view understanding › multi-view fusion
multi-view feature aggregation
0.512021
3D-MAN: 3D Multi-Frame Attention Network for Object Detection · CVPR 2021
Machine learning › Deep learning architectures and training
convolutional neural network
0.522019
CondConv: Conditionally Parameterized Convolutions for Efficient Inference · NeurIPS 2019
Tiled convolutional neural networks · NIPS 2010
Computer vision › Image recognition and object detection › object detection
2d object detection
0.412020
Scalability in Perception for Autonomous Driving: Waymo Open Dataset · CVPR 2020
Machine learning › Learning paradigms
curriculum learning
0.412020
Learning a Multi-Domain Curriculum for Neural Machine Translation · ACL 2020
Machine learning › Deep learning architectures and training
data augmentation
0.412020
Improving 3D Object Detection Through Progressive Population Based Augmentation · ECCV (21) 2020
Machine learning › Efficient and distributed learning
data selection
0.412020
Learning a Multi-Domain Curriculum for Neural Machine Translation · ACL 2020
Machine learning › Deep learning architectures and training › training optimization
gradient balancing
0.412020
Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout · NeurIPS 2020
Machine learning › Optimization for machine learning
gradient conflict resolution
0.412020
Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout · NeurIPS 2020
Computer vision › Video understanding and tracking
multi-object tracking
0.412020
Scalability in Perception for Autonomous Driving: Waymo Open Dataset · CVPR 2020
Machine learning › Learning paradigms
multi-task learning
0.412020
Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout · NeurIPS 2020
Robotics › Autonomous driving
perception
0.412020
Scalability in Perception for Autonomous Driving: Waymo Open Dataset · CVPR 2020
Computer vision › 3D vision › 3d object detection
point cloud object detection
0.412020
Streaming Object Detection for 3-D Point Clouds · ECCV (18) 2020
Machine learning › Efficient and distributed learning › adaptive computation
conditional computation
0.412019
CondConv: Conditionally Parameterized Convolutions for Efficient Inference · NeurIPS 2019
Machine learning › Efficient and distributed learning
distributed training
0.412019
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019
Machine learning › Efficient and distributed learning › dynamic neural network
dynamic convolution
0.412019
CondConv: Conditionally Parameterized Convolutions for Efficient Inference · NeurIPS 2019
Machine learning › Efficient and distributed learning
inference efficiency
0.412019
CondConv: Conditionally Parameterized Convolutions for Efficient Inference · NeurIPS 2019
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training
0.412019
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019
Machine learning › Efficient and distributed learning › distributed training › model parallelism
pipeline parallelism
0.412019
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.222011
Sparse Filtering · NIPS 2011
ICA with Reconstruction Cost for Efficient Overcomplete Feature Learning · NIPS 2011
Computer vision › 3D vision › 3d object annotation
3d bounding box annotation
0.112021
Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset · ICCV 2021

Methods — techniques the papers use, named apart from their topics

transformer architecture · 0.6learnable alignment · 0.6inverse augmentation · 0.6cross-attention · 0.6transformer · 0.5pointnet · 0.5graph convolution kernels · 0.5edge convolution · 0.5cross-modality fusion · 0.5attention network · 0.5
YearPublicationVenuePosition
2022 DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection
abstract
Lidars and cameras are critical sensors that provide complementary information for 3D detection in autonomous driving. While prevalent multi-modal methods [34], [36] simply decorate raw lidar point clouds with camera features and feed them directly to existing 3D detection models, our study shows that fusing camera features with deep lidar features instead of raw points, can lead to better performance. However, as those features are often augmented and aggregated, a key challenge in fusion is how to effectively align the transformed features from two modalities. In this paper, we propose two novel techniques: InverseAug that inverses geometric-related augmentations, e.g., rotation, to enable accurate geometric alignment between lidar points and image pixels, and LearnableAlign that leverages cross-attention to dynamically capture the correlations between image and lidar features during fusion. Based on InverseAug and LearnableAlign, we develop a family of generic multi-modal 3D detection models named DeepFusion, which is more accurate than previous methods. For example, DeepFusion improves Point-Pillars, CenterPoint, and 3D-MAN baselines on Pedestrian detection for 6.7,8.9, and 6.2 LEVEL_2 APH, respectively. Notably, our models achieve state-of-the-art performance on Waymo Open Dataset, and show strong model robustness against input corruptions and out-of-distribution data. Code will be publicly available at https://github.com/tensorflow/lingvo.
Yingwei Li 0002, Adams Wei Yu, Tianjian Meng, Benjamin Caine, Jiquan Ngiam, Daiyi Peng, Junyang Shen, Yifeng Lu, Denny Zhou, Quoc V. Le, Alan L. Yuille, Mingxing Tan
CVPR5
2022 Scene Transformer: A unified architecture for predicting future trajectories of multiple agents
Jiquan Ngiam, Vijay Vasudevan, Benjamin Caine, Hao-Tien Chiang, Jeffrey Ling, Rebecca Roelofs, Alex Bewley, Chenxi Liu 0001, Ashish Venugopal, David J. Weiss, Benjamin Sapp, Jonathon Shlens
ICLR1
2021 To the Point: Efficient 3D Object Detection in the Range Image With Graph Convolution Kernels
abstract
3D object detection is vital for many robotics applications. For tasks where a 2D perspective range image exists, we propose to learn a 3D representation directly from this range image view. To this end, we designed a 2D convolutional network architecture that carries the 3D spherical coordinates of each pixel throughout the network. Its layers can consume any arbitrary convolution kernel in place of the default inner product kernel and exploit the underlying local geometry around each pixel. We outline four such kernels: a dense kernel according to the bag-of-words paradigm, and three graph kernels inspired by recent graph neural network advances: the Transformer, the PointNet, and the Edge Convolution. We also explore cross-modality fusion with the camera image, facilitated by operating in the perspective range image view. Our method performs competitively on the Waymo Open Dataset and improves the state-of-the-art AP for pedestrian detection from 69.7% to 75.5%. It is also efficient in that our smallest model, which still outperforms the popular PointPillars in quality, requires 180 times fewer FLOPS and model parameters.
Yuning Chai, Jiquan Ngiam, Weiyue Wang 0002, Benjamin Caine, Vijay Vasudevan, Dragomir Anguelov
CVPR3
2021 3D-MAN: 3D Multi-Frame Attention Network for Object Detection
abstract
3D object detection is an important module in autonomous driving and robotics. However, many existing methods focus on using single frames to perform 3D detection, and do not fully utilize information from multiple frames. In this paper, we present 3D-MAN: a 3D multi-frame attention network that effectively aggregates features from multiple perspectives and achieves state-of-the-art performance on Waymo Open Dataset. 3D-MAN first uses a novel fast single-frame detector to produce box proposals. The box proposals and their corresponding feature maps are then stored in a memory bank. We design a multi-view alignment and aggregation module, using attention networks, to extract and aggregate the temporal features stored in the memory bank. This effectively combines the features coming from different perspectives of the scene. We demonstrate the effectiveness of our approach on the large-scale complex Waymo Open Dataset, achieving state-of-the-art results compared to published single-frame and multi-frame methods.
Zetong Yang, Jiquan Ngiam
CVPR4
2021 Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset
abstract
As autonomous driving systems mature, motion forecasting has received increasing attention as a critical requirement for planning. Of particular importance are interactive situations such as merges, unprotected turns, etc., where predicting individual object motion is not sufficient. Joint predictions of multiple objects are required for effective route planning. There has been a critical need for high-quality motion data that is rich in both interactions and annotation to develop motion planning models. In this work, we introduce the most diverse interactive motion dataset to our knowledge, and provide specific labels for interacting objects suitable for developing joint prediction models. With over 100,000 scenes, each 20 seconds long at 10 Hz, our new dataset contains more than 570 hours of unique data over 1750 km of roadways. It was collected by mining for interesting interactions between vehicles, pedestrians, and cyclists across six cities within the United States. We use a high-accuracy 3D auto-labeling system to generate high quality 3D bounding boxes for each road agent, and provide corresponding high definition 3D maps for each scene. Furthermore, we introduce a new set of metrics that provides a comprehensive evaluation of both single agent and joint agent interaction motion forecasting models. Finally, we provide strong baseline models for individual-agent prediction and joint-prediction. We hope that this new large-scale interactive motion dataset will provide new opportunities for advancing motion forecasting models.
Scott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu 0001, Hang Zhao 0021, Sabeek Pradhan, Yuning Chai, Benjamin Sapp, Charles R. Qi, Zoey Yang, Aurelien Chouard, Jiquan Ngiam, Vijay Vasudevan, Alexander McCauley, Jonathon Shlens, Dragomir Anguelov
ICCV14
2020 Learning a Multi-Domain Curriculum for Neural Machine Translation
abstract
Most data selection research in machine translation focuses on improving a single domain.We perform data selection for multiple domains at once.This is achieved by carefully introducing instance-level domain-relevance features and automatically constructing a training curriculum to gradually concentrate on multi-domain relevant and noise-reduced data batches.Both the choice of features and the use of curriculum are crucial for balancing and improving all domains, including out-ofdomain.In large-scale experiments, the multidomain curriculum simultaneously reaches or outperforms the individual performance and brings solid gains over no-curriculum training.
Wei Wang 0236, Ye Tian 0022, Jiquan Ngiam, Yinfei Yang, Isaac Caswell, Zarana Parekh
ACL3
2020 Scalability in Perception for Autonomous Driving: Waymo Open Dataset
abstract
The research community has increasing interest in autonomous driving research, despite the resource intensity of obtaining representative real world data. Existing self-driving datasets are limited in the scale and variation of the environments they capture, even though generalization within and between operating regions is crucial to the over-all viability of the technology. In an effort to help align the research community’s contributions with real-world self-driving problems, we introduce a new large scale, high quality, diverse dataset. Our new dataset consists of 1150 scenes that each span 20 seconds, consisting of well synchronized and calibrated high quality LiDAR and camera data captured across a range of urban and suburban geographies. It is 15x more diverse than the largest camera+LiDAR dataset available based on our proposed diversity metric. We exhaustively annotated this data with 2D (camera image) and 3D (LiDAR) bounding boxes, with consistent identifiers across frames. Finally, we provide strong baselines for 2D as well as 3D detection and tracking tasks. We further study the effects of dataset size and generalization across geographies on 3D detection methods. Find data, code and more up-to-date information at http://www.waymo.com/open.
Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han 0002, Jiquan Ngiam, Hang Zhao 0021, Aleksei Timofeev, Scott Ettinger, Maxim Krivokon, Amy Gao, Yu Zhang 0033, Jonathon Shlens, Dragomir Anguelov
CVPR13
2020 Improving 3D Object Detection Through Progressive Population Based Augmentation
Shuyang Cheng, Zhaoqi Leng, Ekin Dogus Cubuk, Barret Zoph, Chunyan Bai, Jiquan Ngiam, Benjamin Caine, Vijay Vasudevan, Quoc V. Le, Jonathon Shlens, Dragomir Anguelov
ECCV (21)6
2020 Streaming Object Detection for 3-D Point Clouds
Wei Han 0002, Benjamin Caine, Brandon Yang, Christoph Sprunk, Ouais Alsharif, Jiquan Ngiam, Vijay Vasudevan, Jonathon Shlens
ECCV (18)7
2020 Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
abstract
The vast majority of deep models use multiple gradient signals, typically corresponding to a sum of multiple loss terms, to update a shared set of trainable weights. However, these multiple updates can impede optimal training by pulling the model in conflicting directions. We present Gradient Sign Dropout (GradDrop), a probabilistic masking procedure which samples gradients at an activation layer based on their level of consistency. GradDrop is implemented as a simple deep layer that can be used in any deep net and synergizes with other gradient balancing approaches. We show that GradDrop outperforms the state-of-the-art multiloss methods within traditional multitask and transfer learning settings, and we discuss how GradDrop reveals links between optimal multiloss training and gradient stochasticity.
Jiquan Ngiam, Yanping Huang, Thang Luong, Henrik Kretzschmar, Yuning Chai, Dragomir Anguelov
NeurIPS2
2019 GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
abstract
Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks. In many cases, increasing model capacity beyond the memory limit of a single accelerator has required developing special algorithms or infrastructure. These solutions are often architecture-specific and do not transfer to other machine learning tasks. To address the need for efficient and task-independent model parallelism, we introduce TensorPipe, a pipeline parallelism library that allows scaling any network that can be expressed as a sequence of layers. By pipelining different sub-sequences of layers on separate accelerators, TensorPipe provides the flexibility of scaling a variety of different networks to gigantic sizes efficiently. Moreover, TensorPipe utilizes a novel batch-splitting pipelining algorithm, resulting in almost linear speedup when a model is partitioned across multiple accelerators. We demonstrate the advantages of TensorPipe by training large-scale neural networks on two different tasks with distinct network architectures: (i)Image Classification: We train a 557-million-parameter AmoebaNet model and attain a top-1 accuracy of 84.4% on ImageNet-2012, (ii)Multilingual Neural Machine Translation: We train a single 6-billion-parameter, 128-layer Transformer model on a corpus spanning over 100 languages and achieve better quality than all bilingual models.
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le
NeurIPS8
2019 CondConv: Conditionally Parameterized Convolutions for Efficient Inference
abstract
Convolutional layers are one of the basic building blocks of modern deep neural networks. One fundamental assumption is that convolutional kernels should be shared for all examples in a dataset. We propose conditionally parameterized convolutions (CondConv), which learn specialized convolutional kernels for each example. Replacing normal convolutions with CondConv enables us to increase the size and capacity of a network, while maintaining efficient inference. We demonstrate that scaling networks with CondConv improves the performance and inference cost trade-off of several existing convolutional neural network architectures on both classification and detection tasks. On ImageNet classification, our CondConv approach applied to EfficientNet-B0 achieves state-ofthe-art performance of 78.3% accuracy with only 413M multiply-adds. Code and checkpoints for the CondConv Tensorflow layer and CondConv-EfficientNet models are available at: https://github.com/tensorflow/tpu/tree/master/ models/official/efficientnet/condconv.
Brandon Yang, Gabriel Bender, Quoc V. Le, Jiquan Ngiam
NeurIPS4
2017 Real-time programming exercise feedback in MOOCs
Andy Nguyen, Amory Schlender, Jiquan Ngiam
EDM4
2011 On optimization methods for deep learning
Quoc V. Le, Jiquan Ngiam, Adam Coates 0002, Ahbik Lahiri, Bobby Prochnow, Andrew Y. Ng
ICML2
2011 Learning Deep Energy Models
Jiquan Ngiam, Pang Wei Koh, Andrew Y. Ng
ICML1
2011 Multimodal Deep Learning
Jiquan Ngiam, Aditya Khosla, Juhan Nam, Honglak Lee, Andrew Y. Ng
ICML1
2011 ICA with Reconstruction Cost for Efficient Overcomplete Feature Learning
abstract
Independent Components Analysis (ICA) and its variants have been successfully used for unsupervised feature learning. However, standard ICA requires an orthonoramlity constraint to be enforced, which makes it difficult to learn overcomplete features. In addition, ICA is sensitive to whitening. These properties make it challenging to scale ICA to high dimensional data. In this paper, we propose a robust soft reconstruction cost for ICA that allows us to learn highly overcomplete sparse features even on unwhitened data. Our formulation reveals formal connections between ICA and sparse autoencoders, which have previously been observed only empirically. Our algorithm can be used in conjunction with off-the-shelf fast unconstrained optimizers. We show that the soft reconstruction cost can also be used to prevent replicated features in tiled convolutional neural networks. Using our method to learn highly overcomplete sparse features and tiled convolutional neural networks, we obtain competitive performances on a wide variety of object recognition tasks. We achieve state-of-the-art test accuracies on the STL-10 and Hollywood2 datasets.
Quoc V. Le, Alexandre Karpenko, Jiquan Ngiam, Andrew Y. Ng
NIPS3
2011 Sparse Filtering
abstract
Unsupervised feature learning has been shown to be effective at learning representations that perform well on image, video and audio classification. However, many existing feature learning algorithms are hard to use and require extensive hyperparameter tuning. In this work, we present sparse filtering, a simple new algorithm which is efficient and only has one hyperparameter, the number of features to learn. In contrast to most other feature learning methods, sparse filtering does not explicitly attempt to construct a model of the data distribution. Instead, it optimizes a simple cost function -- the sparsity of L2-normalized features -- which can easily be implemented in a few lines of MATLAB code. Sparse filtering scales gracefully to handle high-dimensional inputs, and can also be used to learn meaningful features in additional layers with greedy layer-wise stacking. We evaluate sparse filtering on natural images, object classification (STL-10), and phone classification (TIMIT), and show that our method works well on a range of different modalities.
Jiquan Ngiam, Pang Wei Koh, Sonia A. Bhaskar, Andrew Y. Ng
NIPS1
2010 Tiled convolutional neural networks
abstract
Convolutional neural networks (CNNs) have been successfully applied to many tasks such as digit and object recognition. Using convolutional (tied) weights significantly reduces the number of parameters that have to be learned, and also allows translational invariance to be hard-coded into the architecture. In this paper, we consider the problem of learning invariances, rather than relying on hard-coding. We propose tiled convolution neural networks (Tiled CNNs), which use a regular “tiled” pattern of tied weights that does not require that adjacent hidden units share identical weights, but instead requires only that hidden units k steps away from each other to have tied weights. By pooling over neighboring units, this architecture is able to learn complex invariances (such as scale and rotational invariance) beyond translational invariance. Further, it also enjoys much of CNNs’ advantage of having a relatively small number of learned parameters (such as ease of learning and greater scalability). We provide an efficient learning algorithm for Tiled CNNs based on Topographic ICA, and show that learning complex invariant features allows us to achieve highly competitive results for both the NORB and CIFAR-10 datasets.
Quoc V. Le, Jiquan Ngiam, Daniel Jin hao Chia, Pang Wei Koh, Andrew Y. Ng
NIPS2