EDBT 2026 Demo / reviewers in the wild / expert
Jonathan Huang
dblp:55/2421
· DBLP profile ↗
56ranked-venue papers
11as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 11 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Principles of Visual Tokens for Efficient Video UnderstandingabstractVideo understanding has made huge strides in recent years, relying largely on the power of transformers. As this architecture is notoriously expensive and video data is highly redundant, research into improving efficiency has become particularly relevant. Some creative solutions include token selection and merging. While most methods succeed in reducing the cost of the model and maintaining accuracy, an interesting pattern arises: most methods do not outperform the baseline of randomly discarding tokens. In this paper we take a closer look at this phenomenon and observe 5 principles of the nature of visual tokens. For example, we observe that the value of tokens follows a clear Pareto-distribution where most tokens have remarkably low value, and just a few carry most of the perceptual information. We build on these and further insights to propose a lightweight video model, LITE, that can select a small number of tokens effectively, outperforming state-of-the-art and existing baselines across datasets (Kinetics-400 and Something-Something-V2) in the challenging trade-off of computation (GFLOPs) vs accuracy. Experiments also show that LITE generalizes across datasets and even other tasks without the need for retraining. Xinyue Hao 0001, Gen Li 0008, Shreyank N. Gowda, Robert B. Fisher, Jonathan Huang, Anurag Arnab, Laura Sevilla-Lara |
ICCV | 5 |
| 2025 | Enabling Controllable, Identity Preserving, Non-Rigid Edits in Human-Centric ImagesabstractWe approach the problem of inserting a person into a novel scene and controlling their pose via text guidance. Given an image of a person, a masked image of a scene, and a text description of the target pose, our model generates realistic, highly controllable images. We validate the robustness of our model’s true-to-text accuracy and identity preservation via a user study on in-the-wild images. In addition, we present a novel dataset containing pairs of frames from human-centric and action-rich videos, with text captions of the difference in human pose between frames. We also explore the challenges of controllable identity preservation for in-the-wild scenes and the failure modes of similar models. Our methods achieve a 10% increase in pose adherence ([email protected]) over comparable methods without compromising visual fidelity, and show a clear qualitative improvement. Nikolai Warner, Jack Kolb, Meera Hahn, Jonathan Huang, Vighnesh Birodkar, Irfan A. Essa |
ICIP | 4 |
| 2025 | Visually Consistent Hierarchical Image ClassificationabstractHierarchical classification predicts labels across multiple levels of a taxonomy, e.g., from coarse-level \textit{Bird} to mid-level \textit{Hummingbird} to fine-level \textit{Green hermit}, allowing flexible recognition under varying visual conditions. It is commonly framed as multiple single-level tasks, but each level may rely on different visual cues. Distinguishing \textit{Bird} from \textit{Plant} relies on {\it global features} like {\it feathers} or {\it leaves}, while separating \textit{Anna's hummingbird} from \textit{Green hermit} requires {\it local details} such as {\it head coloration}.
Prior methods improve accuracy using external semantic supervision, but such statistical learning criteria fail to ensure consistent visual grounding at test time, resulting in incorrect hierarchical classification. We propose, for the first time, to enforce \textit{internal visual consistency} by aligning fine-to-coarse predictions through intra-image segmentation. Our method outperforms zero-shot CLIP and state-of-the-art baselines on hierarchical classification benchmarks, achieving both higher accuracy and more consistent predictions. It also improves internal image segmentation without requiring pixel-level annotations. Seulki Park, Youren Zhang, Stella X. Yu, Sara Beery, Jonathan Huang |
ICLR | 5 |
| 2025 | Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You ThinkabstractRecent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that one main bottleneck in training large-scale diffusion models for generation lies in effectively learning these representations. Moreover, training can be made easier by incorporating high-quality external visual representations, rather than relying solely on the diffusion models to learn them independently. We study this by introducing a straightforward regularization called REPresentation Alignment (REPA), which aligns the projections of noisy input hidden states in denoising networks with clean image representations obtained from external, pretrained visual encoders. The results are striking: our simple strategy yields significant improvements in both training efficiency and generation quality when applied to popular diffusion and flow-based transformers, such as DiTs and SiTs. For instance, our method can speed up SiT training by over 17.5$\times$, matching the performance (without classifier-free guidance) of a SiT-XL model trained for 7M steps in less than 400K steps. In terms of final generation quality, our approach achieves state-of-the-art results of FID=1.42 using classifier-free guidance with the guidance interval. Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, Saining Xie |
ICLR | 5 |
| 2024 | Optimizing Factorized Encoder Models: Time and Memory Reduction for Scalable and Efficient Action Recognition
Shreyank N. Gowda, Anurag Arnab, Jonathan Huang |
ECCV (10) | 3 |
| 2024 | Tree-D Fusion: Simulation-Ready Tree Dataset from Single Images with Diffusion Priors
Jae Joong Lee, Bosheng Li, Sara Beery, Jonathan Huang, Songlin Fei, Raymond A. Yeh, Bedrich Benes |
ECCV (41) | 4 |
| 2024 | VideoPoet: A Large Language Model for Zero-Shot Video GenerationabstractWe present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that can be adapted for a range of video generation tasks. We present empirical results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the ability to generate high-fidelity motions. Project page: http://sites.research.google/videopoet/ Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng 0003, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Hartwig Adam, Ming-Hsuan Yang 0001, Irfan A. Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang 0004 |
ICML | 5 |
| 2023 | Text and Click inputs for unambiguous open vocabulary instance segmentation
Vighnesh Birodkar, Jonathan Huang, Meera Hahn, Irfan A. Essa, Nikolai Warner |
BMVC | 2 |
| 2023 | Learning to Detect Novel and Fine-Grained Acoustic Sequences Using Pretrained Audio RepresentationsabstractThis work investigates pretrained audio representations for few shot Sound Event Detection. We specifically address the task of few shot detection of novel acoustic sequences, or sound events with semantically meaningful temporal structure, without assuming access to non-target audio. We develop procedures for pretraining suitable representations, and methods which transfer them to our few shot learning scenario. Our experiments evaluate the general purpose utility of our pretrained representations on AudioSet, and the utility of proposed few shot methods via tasks constructed from real-world acoustic sequences. Our pretrained embeddings are suitable to the proposed task, and enable multiple aspects of our few shot framework. Vasudha Kowtha, Miquel Espi Marques, Jonathan Huang, Carlos Avendaño |
ICASSP | 3 |
| 2023 | Partisan US News Media Representations of Syrian RefugeesabstractWe investigate how representations of Syrian refugees (2011-2021) differ across US partisan news outlets. We analyze 47,388 articles from the online US media about Syrian refugees to detail differences in reporting between left- and right-leaning media. We use various NLP techniques to understand these differences. Our polarization and question answering results indicated that left-leaning media tended to represent refugees as child victims, welcome in the US, and right-leaning media cast refugees as Islamic terrorists. We noted similar results with our sentiment and offensive speech scores over time, which detail possibly unfavorable representations of refugees in right-leaning media. A strength of our work is how the different techniques we have applied validate each other. Based on our results, we provide several recommendations. Stakeholders may utilize our findings to intervene around refugee representations, and design communications campaigns that improve the way society sees refugees and possibly aid refugee outcomes. Marzieh Babaeianjelodar, Yiwen Shi, Kamila Janmohamed, Rupak Sarkar, Ingmar Weber, Thomas Davidson, Munmun De Choudhury, Jonathan Huang, Shweta Yadav 0001, Ashiqur R. KhudaBukhsh, Chris T. Bauch, Preslav Nakov, Orestis Papakyriakopoulos, Koustuv Saha, Kaveh Khoshnood, Navin Kumar 0004 |
ICWSM | 9 |
| 2023 | DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation ModelabstractObserving the close relationship among panoptic, semantic and instance segmentation tasks, we propose to train a universal multi-dataset multi-task segmentation model: DaTaSeg. We use a shared representation (mask proposals with class predictions) for all tasks. To tackle task discrepancy, we adopt different merge operations and post-processing for different tasks. We also leverage weak-supervision, allowing our segmentation model to benefit from cheaper bounding box annotations. To share knowledge across datasets, we use text embeddings from the same semantic embedding space as classifiers and share all network parameters among datasets. We train DaTaSeg on ADE semantic, COCO panoptic, and Objects365 detection datasets. DaTaSeg improves performance on all datasets, especially small-scale datasets, achieving 54.0 mIoU on ADE semantic and 53.5 PQ on COCO panoptic. DaTaSeg also enables weakly-supervised knowledge transfer on ADE panoptic and Objects365 instance segmentation. Experiments show DaTaSeg scales with the number of training datasets and enables open-vocabulary segmentation through direct transfer. In addition, we annotate an Objects365 instance segmentation set of 1,000 images and release it as a public evaluation benchmark on https://laoreja.github.io/dataseg. Xiuye Gu, Yin Cui, Jonathan Huang, Abdullah Rashwan, Xingyi Zhou, Golnaz Ghiasi, Weicheng Kuo, Huizhong Chen, Liang-Chieh Chen, David A. Ross |
NeurIPS | 3 |
| 2022 | The Auto Arborist Dataset: A Large-Scale Benchmark for Multiview Urban Forest Monitoring Under Domain ShiftabstractGeneralization to novel domains is a fundamental chal-lenge for computer vision. Near-perfect accuracy on bench-marks is common, but these models do not work as expected when deployed outside of the training distribution. To build computer vision systems that truly solve real-world prob-lems at global scale, we need benchmarks that fully capture real-world complexity, including geographic domain shift, long-tailed distributions, and data noise. We propose urban forest monitoring as an ideal testbed for studying and improving upon these computer vision challenges, while working towards filling a crucial environ-mental and societal need. Urban forests provide significant benefits to urban societies. However, planning and main-taining these forests is expensive. One particularly costly aspect of urban forest management is monitoring the ex-isting trees in a city: e.g., tracking tree locations, species, and health. Monitoring efforts are currently based on tree censuses built by human experts, costing cities millions of dollars per census and thus collected infrequently. Previous investigations into automating urban forest monitoring focused on small datasets from single cities, covering only common categories. To address these short-comings, we introduce a new large-scale dataset that joins public tree censuses from 23 cities with a large collection of street level and aerial imagery. Our Auto Arborist dataset contains over 2.5M trees and 344 genera and is >2 or-ders of magnitude larger than the closest dataset in the literature. We introduce baseline results on our dataset across modalities as well as metrics for the detailed analy-sis of generalization with respect to geographic distribution shifts, vital for such a system to be deployed at-scale. Sara Beery, Guanhang Wu, Trevor Edwards, Filip Pavetic, Bo Majewski, Shreyasee Mukherjee, Stanley Chan, John Morgan, Vivek Rathod, Jonathan Huang |
CVPR | 10 |
| 2022 | PERF-Net: Pose Empowered RGB-Flow NetabstractIn recent years, many works in the video action recognition literature have shown that two stream models (combining spatial and temporal input streams) are necessary for achieving state-of-the-art performance. In this paper we show the benefits of including yet another stream based on human pose estimated from each frame — specifically by rendering pose on input RGB frames. At first blush, this additional stream may seem redundant given that human pose is fully determined by RGB pixel values — however we show (perhaps surprisingly) that this simple and flexible addition can provide complementary gains. Using this insight, we propose a new model, which we dub PERF-Net (short for Pose Empowered RGB-Flow Net), which combines this new pose stream with the standard RGB and flow based input streams via distillation techniques and show that our model outperforms the state-of-the-art by a large margin in a number of human action recognition datasets while not requiring flow or pose to be explicitly computed at inference time. The proposed pose stream is also part of the winner solution of the ActivityNet Kinetics Challenge 2020 [1]. Yinxiao Li, Zhichao Lu, Xuehan Xiong, Jonathan Huang |
WACV | 4 |
| 2021 | The surprising impact of mask-head architecture on novel class segmentationabstractInstance segmentation models today are very accurate when trained on large annotated datasets, but collecting mask annotations at scale is prohibitively expensive. We address the partially supervised instance segmentation problem in which one can train on (significantly cheaper) bounding boxes for all categories but use masks only for a subset of categories. In this work, we focus on a popular family of models which apply differentiable cropping to a feature map and predict a mask based on the resulting crop. Under this family, we study Mask R-CNN and discover that instead of its default strategy of training the mask-head with a combination of proposals and groundtruth boxes, training the mask-head with only groundtruth boxes dramatically improves its performance on novel classes. This training strategy also allows us to take advantage of alternative mask-head architectures, which we exploit by replacing the typical mask-head of 2-4 layers with significantly deeper off-the-shelf architectures (e.g. ResNet, Hourglass models). While many of these architectures perform similarly when trained in fully supervised mode, our main finding is that they can generalize to novel classes in dramatically different ways. We call this ability of mask-heads to generalize to unseen classes the strong mask generalization effect and show that without any specialty modules or losses, we can achieve state-of-the-art results in the partially supervised COCO instance segmentation benchmark. Finally, we demonstrate that our effect is general, holding across underlying detection methodologies (including anchor-based, anchor-free or no detector at all) and across different backbone networks. Code and pre-trained models are available at https://git.io/deepmac. Vighnesh Birodkar, Zhichao Lu, Siyang Li 0002, Vivek Rathod, Jonathan Huang |
ICCV | 5 |
| 2020 | Context R-CNN: Long Term Temporal Context for Per-Camera Object DetectionabstractIn static monitoring cameras, useful contextual information can stretch far beyond the few seconds typical video understanding models might see: subjects may exhibit similar behavior over multiple days, and background objects remain static. Due to power and storage constraints, sampling frequencies are low, often no faster than one frame per second, and sometimes are irregular due to the use of a motion trigger. In order to perform well in this setting, models must be robust to irregular sampling rates. In this paper we propose a method that leverages temporal context from the unlabeled frames of a novel camera to improve performance at that camera. Specifically, we propose an attention-based approach that allows our model, Context R-CNN, to index into a long term memory bank constructed on a per-camera basis and aggregate contextual features from other frames to boost object detection performance on the current frame. We apply Context R-CNN to two settings: (1) species detection using camera traps, and (2) vehicle detection in traffic cameras, showing in both settings that Context R-CNN leads to performance gains over strong baselines. Moreover, we show that increasing the contextual time horizon leads to improved results. When applied to camera trap data from the Snapshot Serengeti dataset, Context R-CNN with context from up to a month of images outperforms a single-frame baseline by 17.9% mAP, and outperforms S3D (a 3d convolution based baseline) by 11.2% mAP. Sara Beery, Guanhang Wu, Vivek Rathod, Ronny Votel, Jonathan Huang |
CVPR | 5 |
| 2020 | RetinaTrack: Online Single Stage Joint Detection and TrackingabstractTraditionally multi-object tracking and object detection are performed using separate systems with most prior works focusing exclusively on one of these aspects over the other. Tracking systems clearly benefit from having access to accurate detections, however and there is ample evidence in literature that detectors can benefit from tracking which, for example, can help to smooth predictions over time. In this paper we focus on the tracking-by-detection paradigm for autonomous driving where both tasks are mission critical. We propose a conceptually simple and efficient joint model of detection and tracking, called RetinaTrack, which modifies the popular single stage RetinaNet approach such that it is amenable to instance-level embedding training. We show, via evaluations on the Waymo Open Dataset, that we outperform a recent state of the art tracking algorithm while requiring significantly less computation. We believe that our simple yet effective approach can serve as a strong baseline for future work in this area. Zhichao Lu, Vivek Rathod, Ronny Votel, Jonathan Huang |
CVPR | 4 |
| 2020 | Structural Sparsification for Far-Field Speaker Recognition with Intel® GnaabstractRecently, deep neural networks (DNN) have been widely used in speaker recognition area. In order to achieve fast response time and high accuracy, the requirements for hardware resources increase rapidly. However, as the speaker recognition application is often implemented on mobile devices, it is necessary to maintain a low computational cost while keeping high accuracy in far-field condition. In this paper, we apply structural sparsification on time-delay neural networks (TDNN) to remove redundant structures and accelerate the execution. On our targeted hardware, our model can remove 60% of parameters and only slightly increasing equal error rate (EER) by 0.18% while our structural sparse model can achieve more than 1.5× speedup. Jingchi Zhang, Jonathan Huang, Michael Deisher, Hai Li 0001, Yiran Chen 0001 |
ICASSP | 2 |
| 2020 | Length- and Noise-Aware Training Techniques for Short-Utterance Speaker RecognitionabstractSpeaker recognition performance has been greatly improved with the emergence of deep learning. Deep neural networks show the capacity to effectively deal with impacts of noise and reverberation, making them attractive to far-field speaker recognition systems. The x-vector framework is a popular choice for generating speaker embeddings in recent literature due to its robust training mechanism and excellent performance in various test sets. In this paper, we start with early work on including invariant representation learning (IRL) to the loss function and modify the approach with centroid alignment (CA) and length variability cost (LVC) techniques to further improve robustness in noisy, far-field applications. This work mainly focuses on improvements for short-duration test utterances (1-8s). We also present improved results on long-duration tasks. In addition, this work discusses a novel self-attention mechanism. On the VOiCES far-field corpus, the combination of the proposed techniques achieves relative improvements of 7.0% for extremely short and 8.2% for full-duration test utterances on equal error rate (EER) over our baseline system. Wenda Chen, Jonathan Huang, Tobias Bocklet |
INTERSPEECH | 2 |
| 2020 | Compact Speaker Embedding: lrx-VectorabstractDeep neural networks (DNN) have recently been widely used in speaker recognition systems, achieving state-of-the-art performance on various benchmarks. The x-vector architecture is especially popular in this research community, due to its excellent performance and manageable computational complexity. In this paper, we present the lrx-vector system, which is the low-rank factorized version of the x-vector embedding network. The primary objective of this topology is to further reduce the memory requirement of the speaker recognition system. We discuss the deployment of knowledge distillation for training the lrx-vector system and compare against low-rank factorization with SVD. On the VOiCES 2019 far-field corpus we were able to reduce the weights by 28% compared to the full-rank x-vector system while keeping the recognition rate constant (1.83% EER). Munir Georges, Jonathan Huang, Tobias Bocklet |
INTERSPEECH | 2 |
| 2020 | Investigating topics, audio representations and attention for multimodal scene-aware dialog
Shachi H. Kumar, Eda Okur, Saurav Sahay, Jonathan Huang, Lama Nachman |
Comput. Speech Lang. | 4 |
| 2019 | Diverse Generation for Multi-Agent Sports GamesabstractIn this paper, we propose a new generative model for multi-agent trajectory data, focusing on the case of multi-player sports games. Our model leverages graph neural networks (GNNs) and variational recurrent neural networks (VRNNs) to achieve a permutation equivariant model suitable for sports. On two challenging datasets (basketball and soccer), we show that we are able to produce more accurate forecasts than previous methods. We assess accuracy using various metrics, such as log-likelihood and "best of N" loss, based on N different samples of the future. We also measure the distribution of statistics of interest, such as player location or velocity, and show that the distribution induced by our generative model better matches the empirical distribution of the test set. Finally, we show that our model can perform conditional prediction, which lets us answer counterfactual questions such as “how will the players move differently if A passes the ball to B instead of C?” Raymond A. Yeh, Alexander G. Schwing, Jonathan Huang, Kevin Murphy 0002 |
CVPR | 3 |
| 2019 | Uncertainty-Aware Audiovisual Activity Recognition Using Deep Bayesian Variational InferenceabstractDeep neural networks (DNNs) provide state-of-the-art results for a multitude of applications, but the approaches using DNNs for multimodal audiovisual applications do not consider predictive uncertainty associated with individual modalities. Bayesian deep learning methods provide principled confidence and quantify predictive uncertainty. Our contribution in this work is to propose an uncertainty aware multimodal Bayesian fusion framework for activity recognition. We demonstrate a novel approach that combines deterministic and variational layers to scale Bayesian DNNs to deeper architectures. Our experiments using in- and out-of-distribution samples selected from a subset of Moments-in-Time (MiT) dataset show a more reliable confidence measure as compared to the non-Bayesian baseline and the Monte Carlo dropout (MC dropout) approximate Bayesian inference. We also demonstrate the uncertainty estimates obtained from the proposed framework can identify out-of-distribution data on the UCF101 and MiT datasets. In the multimodal setting, the proposed framework improved precision-recall AUC by 10.2% on the subset of MiT dataset as compared to non-Bayesian baseline. Mahesh Subedar, Ranganath Krishnan, Paulo Lopez-Meyer, Omesh Tickoo, Jonathan Huang |
ICCV | 5 |
| 2019 | Intel Far-Field Speaker Recognition System for VOiCES Challenge 2019
Jonathan Huang, Tobias Bocklet |
INTERSPEECH | 1 |
| 2018 | Progressive Neural Architecture Search
Chenxi Liu 0001, Barret Zoph, Maxim Neumann, Jonathon Shlens, Li-Jia Li 0001, Li Fei-Fei 0001, Alan L. Yuille, Jonathan Huang, Kevin Murphy 0002 |
ECCV (1) | 9 |
| 2018 | Learning to Segment via Cut-and-Paste
Tal Remez, Jonathan Huang, Matthew Brown 0001 |
ECCV (7) | 2 |
| 2018 | Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification
Saining Xie, Chen Sun 0002, Jonathan Huang, Zhuowen Tu, Kevin Murphy 0002 |
ECCV (15) | 3 |
| 2018 | Sufficiency Quantification for Seamless Text-Independent Speaker EnrollmentabstractText-independent speaker recognition (TI-SR) requires a lengthy enrollment process that involves asking dedicated time from the user to create a reliable model of their voice. Seamless enrollment is a highly attractive feature which refers to the enrollment process that happens in the background and asks for no dedicated time from the user. One of the key problems in a fully automated seamless enrollment process is to determine the sufficiency of a given utterance collection for the purpose of TI-SR. No known metric exists in the literature to quantify sufficiency. This paper introduces a novel metric called phoneme-richness score. Quality of a sufficiency metric can be assessed via its correlation with the TI-SR performance. Our assessment shows that phoneme-richness score achieves -0.96 correlation with TI-SR performance (measured in equal error rate), which is highly significant, whereas a naive sufficiency metric like speech duration achieves only -0.68 correlation. Gokcen Cilingir, Jonathan Huang, Mandar S. Joshi, Narayan Biswal |
ICASSP | 2 |
| 2018 | Generative Models of Visually Grounded Imagination
Ramakrishna Vedantam, Ian Fischer, Jonathan Huang, Kevin Murphy 0002 |
ICLR (Poster) | 3 |
| 2017 | Spatially Adaptive Computation Time for Residual NetworksabstractThis paper proposes a deep learning architecture based on Residual Network that dynamically adjusts the number of executed layers for the regions of the image. This architecture is end-to-end trainable, deterministic and problem-agnostic. It is therefore applicable without any modifications to a wide range of computer vision problems such as image classification, object detection and image segmentation. We present experimental results showing that this model improves the computational efficiency of Residual Networks on the challenging ImageNet classification and COCO object detection datasets. Additionally, we evaluate the computation time maps on the visual saliency dataset cat2000 and find that they correlate surprisingly well with human eye fixation positions. Michael Figurnov, Maxwell D. Collins, Yukun Zhu, Li Zhang 0003, Jonathan Huang, Dmitry P. Vetrov, Ruslan Salakhutdinov |
CVPR | 5 |
| 2017 | Speed/Accuracy Trade-Offs for Modern Convolutional Object DetectorsabstractThe goal of this paper is to serve as a guide for selecting a detection architecture that achieves the right speed/memory/accuracy balance for a given application and platform. To this end, we investigate various ways to trade accuracy for speed and memory usage in modern convolutional object detection systems. A number of successful systems have been proposed in recent years, but apples-toapples comparisons are difficult due to different base feature extractors (e.g., VGG, Residual Networks), different default image resolutions, as well as different hardware and software platforms. We present a unified implementation of the Faster R-CNN [30], R-FCN [6] and SSD [25] systems, which we view as meta-architectures and trace out the speed/accuracy trade-off curve created by using alternative feature extractors and varying other critical parameters such as image size within each of these meta-architectures. On one extreme end of this spectrum where speed and memory are critical, we present a detector that achieves real time speeds and can be deployed on a mobile device. On the opposite end in which accuracy is critical, we present a detector that achieves state-of-the-art performance measured on the COCO detection task. Jonathan Huang, Vivek Rathod, Chen Sun 0002, Menglong Zhu, Anoop Korattikara Balan, Alireza Fathi, Ian Fischer, Zbigniew Wojna, Yang Song 0009, Sergio Guadarrama, Kevin Murphy 0002 |
CVPR | 1 |
| 2016 | Generation and Comprehension of Unambiguous Object DescriptionsabstractWe propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to infer which object is being described. We show that our method outperforms previous methods that generate descriptions of objects without taking into account other potentially ambiguous objects in the scene. Our model is inspired by recent successes of deep learning methods for image captioning, but while image captioning is difficult to evaluate, our task allows for easy objective evaluation. We also present a new large-scale dataset for referring expressions, based on MSCOCO. We have released the dataset and a toolbox for visualization and evaluation, see https://github.com/ mjhucla/Google_Refexp_toolbox. Junhua Mao, Jonathan Huang, Alexander Toshev, Oana-Maria Camburu, Alan L. Yuille, Kevin Murphy 0002 |
CVPR | 2 |
| 2016 | Detecting Events and Key Actors in Multi-person VideosabstractMulti-person event recognition is a challenging task, often with many people active in the scene but only a small subset contributing to an actual event. In this paper, we propose a model which learns to detect events in such videos while automatically "attending" to the people responsible for the event. Our model does not use explicit annotations regarding who or where those people are during training and testing. In particular, we track people in videos and use a recurrent neural network (RNN) to represent the track features. We learn time-varying attention weights to combine these features at each time-instant. The attended features are then processed using another RNN for event detection/ classification. Since most video datasets with multiple people are restricted to a small number of videos, we also collected a new basketball dataset comprising 257 basketball games with 14K event annotations corresponding to 11 event classes. Our model outperforms state-of-the-art methods for both event classification and detection on this new dataset. Additionally, we show that the attention mechanism is able to consistently localize the relevant players. Vignesh Ramanathan, Jonathan Huang, Sami Abu-El-Haija, Alexander N. Gorban, Kevin Murphy 0002, Li Fei-Fei 0001 |
CVPR | 2 |
| 2015 | Im2Calories: Towards an Automated Mobile Vision Food DiaryabstractWe present a system which can recognize the contents of your meal from a single image, and then predict its nutritional contents, such as calories. The simplest version assumes that the user is eating at a restaurant for which we know the menu. In this case, we can collect images offline to train a multi-label classifier. At run time, we apply the classifier (running on your phone) to predict which foods are present in your meal, and we lookup the corresponding nutritional facts. We apply this method to a new dataset of images from 23 different restaurants, using a CNN-based classifier, significantly outperforming previous work. The more challenging setting works outside of restaurants. In this case, we need to estimate the size of the foods, as well as their labels. This requires solving segmentation and depth / volume estimation from a single image. We present CNN-based approaches to these problems, with promising preliminary results. Austin Myers, Nicholas Johnston, Vivek Rathod, Anoop Korattikara Balan, Alexander N. Gorban, Nathan Silberman, Sergio Guadarrama, George Papandreou, Jonathan Huang, Kevin Murphy 0002 |
ICCV | 9 |
| 2015 | Learning Program Embeddings to Propagate Feedback on Student CodeabstractProviding feedback, both assessing final work and giving hints to stuck students, is difficult for open-ended assignments in massive online classes which can range from thousands to millions of students. We introduce a neural network method to encode programs as a linear mapping from an embedded precondition space to an embedded postcondition space and propose an algorithm for feedback at scale using these linear maps as features. We apply our algorithm to assessments from the Code.org Hour of Code and Stanford University’s CS1 course, where we propagate human comments on student assignments to orders of magnitude more submissions. Chris Piech, Jonathan Huang, Andy Nguyen, Mike Phulsuksombati, Mehran Sahami, Leonidas J. Guibas |
ICML | 2 |
| 2015 | Meeting assistant application
Michel Assayag, Jonathan Huang, Jonathan Mamou, Oren Pereg, Saurav Sahay, Oren Shamir, Georg Stemmer, Moshe Wasserblat |
INTERSPEECH | 2 |
| 2015 | Autonomously Generating Hints by Inferring Problem Solving PoliciesabstractExploring the whole sequence of steps a student takes to produce work, and the patterns that emerge from thousands of such sequences is fertile ground for a richer understanding of learning. In this paper we autonomously generate hints for the Code.org `Hour of Code,' (which is to the best of our knowledge the largest online course to date) using historical student data. We first develop a family of algorithms that can predict the way an expert teacher would encourage a student to make forward progress. Such predictions can form the basis for effective hint generation systems. The algorithms are more accurate than current state-of-the-art methods at recreating expert suggestions, are easy to implement and scale well. We then show that the same framework which motivated the hint generating algorithms suggests a sequence-based statistic that can be measured for each learner. We discover that this statistic is highly predictive of a student's future success. Chris Piech, Mehran Sahami, Jonathan Huang, Leonidas J. Guibas |
L@S | 3 |
| 2015 | What's Cookin'? Interpreting Cooking Videos using Text, Speech and VisionabstractJonathan Malmaud, Jonathan Huang, Vivek Rathod, Nicholas Johnston, Andrew Rabinovich, Kevin Murphy. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nicholas Johnston, Andrew Rabinovich, Kevin Murphy 0002 |
HLT-NAACL | 2 |
| 2015 | Deep Knowledge TracingabstractKnowledge tracing, where a machine models the knowledge of a student as they interact with coursework, is an established and significantly unsolved problem in computer supported education.In this paper we explore the benefit of using recurrent neural networks to model student learning.This family of models have important advantages over current state of the art methods in that they do not require the explicit encoding of human domain knowledge,and have a far more flexible functional form which can capture substantially more complex student interactions.We show that these neural networks outperform the current state of the art in prediction on real student data,while allowing straightforward interpretation and discovery of structure in the curriculum.These results suggest a promising new line of research for knowledge tracing. Chris Piech, Jonathan Bassen, Jonathan Huang, Surya Ganguli, Mehran Sahami, Leonidas J. Guibas, Jascha Sohl-Dickstein |
NIPS | 3 |
| 2014 | Superposter behavior in MOOC forumsabstractDiscussion forums, employed by MOOC providers as the primary mode of interaction among instructors and students, have emerged as one of the important components of online courses. We empirically study contribution behavior in these online collaborative learning forums using data from 44 MOOCs hosted on Coursera, focusing primarily on the highest-volume contributors---"superposters"---in a forum. We explore who these superposters are and study their engagement patterns across the MOOC platform, with a focus on the following question---to what extent is superposting a positive phenomenon for the forum? Specifically, while superposters clearly contribute heavily to the forum in terms of quantity, how do these contributions rate in terms of quality, and does this prolific posting behavior negatively impact contribution from the large remainder of students in the class? Jonathan Huang, Anirban Dasgupta 0001, Arpita Ghosh, Jane Manning, Marc Sanders |
L@S | 1 |
| 2014 | Codewebs: scalable homework search for massive open online programming coursesabstractMassive open online courses (MOOCs), one of the latest internet revolutions have engendered hope that constant iterative improvement and economies of scale may cure the ``cost disease" of higher education. While scalable in many ways, providing feedback for homework submissions (particularly open-ended ones) remains a challenge in the online classroom. In courses where the student-teacher ratio can be ten thousand to one or worse, it is impossible for instructors to personally give feedback to students or to understand the multitude of student approaches and pitfalls. Organizing and making sense of massive collections of homework solutions is thus a critical web problem. Despite the challenges, the dense solution space sampling in highly structured homeworks for some MOOCs suggests an elegant solution to providing quality feedback to students on a massive scale. Andy Nguyen, Chris Piech, Jonathan Huang, Leonidas J. Guibas |
WWW | 3 |
| 2013 | Tuned Models of Peer Assessment in MOOCs
Chris Piech, Jonathan Huang, Chuong B. Do, Andrew Y. Ng, Daphne Koller |
EDM | 2 |
| 2012 | Probabilistic Event Cascades for Alzheimer's diseaseabstractAccurate and detailed models of the progression of neurodegenerative diseases such as Alzheimer's (AD) are crucially important for reliable early diagnosis and the determination and deployment of effective treatments. In this paper, we introduce the ALPACA (Alzheimer's disease Probabilistic Cascades) model, a generative model linking latent Alzheimer's progression dynamics to observable biomarker data. In contrast with previous works which model disease progression as a fixed ordering of events, we explicitly model the variability over such orderings among patients which is more realistic, particularly for highly detailed disease progression models. We describe efficient learning algorithms for ALPACA and discuss promising experimental results on a real cohort of Alzheimer's patients from the Alzheimer's Disease Neuroimaging Initiative. Jonathan Huang, Daniel C. Alexander |
NIPS | 1 |
| 2012 | Riffled Independence for Efficient Inference with Partial RankingsabstractDistributions over rankings are used to model data in a multitude of real world settings such as preference analysis and political elections. Modeling such distributions presents several computational challenges, however, due to the factorial size of the set of rankings over an item set. Some of these challenges are quite familiar to the artificial intelligence community, such as how to compactly represent a distribution over a combinatorially large space, and how to efficiently perform probabilistic inference with these representations. With respect to ranking, however, there is the additional challenge of what we refer to as human task complexity users are rarely willing to provide a full ranking over a long list of candidates, instead often preferring to provide partial ranking information. Simultaneously addressing all of these challenges i.e., designing a compactly representable model which is amenable to efficient inference and can be learned using partial ranking data is a difficult task, but is necessary if we would like to scale to problems with nontrivial size. In this paper, we show that the recently proposed riffled independence assumptions cleanly and efficiently address each of the above challenges. In particular, we establish a tight mathematical connection between the concepts of riffled independence and of partial rankings. This correspondence not only allows us to then develop efficient and exact algorithms for performing inference tasks using riffled independence based represen- tations with partial rankings, but somewhat surprisingly, also shows that efficient inference is not possible for riffle independent models (in a certain sense) with observations which do not take the form of partial rankings. Finally, using our inference algorithm, we introduce the first method for learning riffled independence based models from partially ranked data. Jonathan Huang, Ashish Kapoor, Carlos Guestrin |
J. Artif. Intell. Res. | 1 |
| 2011 | Fourier-Information Duality in the Identity Management Problem
Xiaoye Jiang, Jonathan Huang, Leonidas J. Guibas |
ECML/PKDD (2) | 2 |
| 2011 | Efficient Probabilistic Inference with Partial Ranking Queries
Jonathan Huang, Ashish Kapoor, Carlos Guestrin |
UAI | 1 |
| 2010 | Learning Hierarchical Riffle Independent Groupings from Rankings
Jonathan Huang, Carlos Guestrin |
ICML | 1 |
| 2009 | Hilbert space embeddings of conditional distributions with applications to dynamical systemsabstractIn this paper, we extend the Hilbert space embedding approach to handle conditional distributions. We derive a kernel estimate for the conditional embedding, and show its connection to ordinary embeddings. Condi-tional embeddings largely extend our ability to manipulate distributions in Hilbert spaces, and as an example, we derive a nonpara-metric method for modeling dynamical sys-tems where the belief state of the system is maintained as a conditional embedding. Our method is very general in terms of both the domains and the types of distributions that it can handle, and we demonstrate the ef-fectiveness of our method in various dynami-cal systems. We expect that conditional em-beddings will have wider applications beyond modeling dynamical systems. 1. Jonathan Huang, Alexander J. Smola, Kenji Fukumizu |
ICML | 2 |
| 2009 | Riffled Independence for Ranked DataabstractRepresenting distributions over permutations can be a daunting task due to the fact that the number of permutations of n objects scales factorially in n. One recent way that has been used to reduce storage complexity has been to exploit probabilistic independence, but as we argue, full independence assumptions impose strong sparsity constraints on distributions and are unsuitable for modeling rankings. We identify a novel class of independence structures, called riffled independence, which encompasses a more expressive family of distributions while retaining many of the properties necessary for performing efficient inference and reducing sample complexity. In riffled independence, one draws two permutations independently, then performs the riffle shuffle, common in card games, to combine the two permutations to form a single permutation. In ranking, riffled independence corresponds to ranking disjoint sets of objects independently, then interleaving those rankings. We provide a formal introduction and present algorithms for using riffled independence within Fourier-theoretic frameworks which have been explored by a number of recent papers. Jonathan Huang, Carlos Guestrin |
NIPS | 1 |
| 2009 | Fourier Theoretic Probabilistic Inference over Permutations
Jonathan Huang, Carlos Guestrin, Leonidas J. Guibas |
J. Mach. Learn. Res. | 1 |
| 2007 | Efficient Inference for Distributions on PermutationsabstractPermutations are ubiquitous in many real world problems, such as voting, rankings and data association. Representing uncertainty over permutations is challenging, since there are n! possibilities, and typical compact representations such as graphical models cannot efficiently capture the mutual exclusivity con- straints associated with permutations. In this paper, we use the “low-frequency” terms of a Fourier decomposition to represent such distributions compactly. We present Kronecker conditioning, a general and efficient approach for maintaining these distributions directly in the Fourier domain. Low order Fourier-based approximations can lead to functions that do not correspond to valid distributions. To address this problem, we present an efficient quadratic program defined directly in the Fourier domain to project the approximation onto a relaxed form of the marginal polytope. We demonstrate the effectiveness of our approach on a real camera-based multi-people tracking setting. Jonathan Huang, Carlos Guestrin, Leonidas J. Guibas |
NIPS | 1 |
| 2005 | The Intel Mote platform: a bluetooth-based sensor network for industrial monitoringabstractThe Intel mote is a new sensor node platform motivated by several design goals: increased CPU performance, improved radio bandwidth and reliability, and the usage of commercial off-the-shelf components in order to maintain cost-effectiveness. This new platform is built around an integrated wireless microcontroller consisting of an ARM*7 core, a Bluetooth radio, SRAM and FLASH memory, as well as various I/O options. The Intel Mote software architecture is based on an ARM port of TinyOS. Networking and routing layers have been created on top of the TinyOS base to provide Bluetooth-based multi-hop functionality. The network is self-organizing on startup and has mechanisms to repair failed links and circumvent failed nodes. A reliable high bandwidth streaming transport layer has also been created. The Intel Mote was deployed in an equipment monitoring application using industrial vibration sensors. This application was chosen since it benefits from the increased platform capabilities and network bandwidth of the Intel Mote platform. The paper presents a detailed analysis of the observed network operation, packet transfer rates, and power consumption. Lama Nachman, Ralph Kling, Robert Adler, Jonathan Huang, Vincent Hummel |
IPSN | 4 |
| 2005 | Intel mote 2: an advanced platform for demanding sensor network applicationsabstractNo abstract available. Robert Adler, Mick Flanigan, Jonathan Huang, Ralph Kling, Nandakishore Kushalnagar, Lama Nachman, Chieh-Yih Wan, Mark D. Yarvis |
SenSys | 3 |
| 2004 | Intel Mote: using bluetooth in sensor networksabstractThe Intel Mote is a new sensor node platform motivated by several design goals: increased CPU performance, improved radio bandwidth and reliability and the usage of commercial off-the-shelf components in order to maintain cost-effectiveness. This new platform is built around an integrated wireless microcontroller consisting of an ARM*7 core, a Bluetooth* radio, RAM and FLASH memory as well as various I/O options. Due to the connection-oriented nature of Bluetooth, a new network formation and maintenance algorithms that are optimized for this protocol have been created. In particular, the "scatternet" mode of Bluetooth has been successfully adapted to form networks comprised of multiple piconet. Ralph Kling, Robert Adler, Jonathan Huang, Vincent Hummel, Lama Nachman |
SenSys | 3 |
| 2000 | Subband-based adaptive decorrelation filtering for co-channel speech separationabstractA subband-based adaptive decorrelation filtering algorithm (SBADF) is proposed for co-channel speech separation. The SBADF decomposes the input signals into several frequency subbands, and uses the adaptive decorrelation filtering algorithm (ADF) to process the signals in each subband independently. The processed subband signals are then combined for each channel to form the separated speech. Experimental results show that while the fullband ADF can achieve better separation performance after reaching convergence, the SBADF has the advantage of improved convergence rate and reduced computational complexity by a factor of approximately two. Jonathan Huang, Kuan-Chieh Yen, Yunxin Zhao |
IEEE Trans. Speech Audio Process. | 1 |
| 1999 | NotePals: Light Weight Note Sharing by the Group, for the GroupabstractNotePals is a lightweight note sharing system that gives group members easy access to each others experiences through their personal notes. The system allows notes taken by group members in any context to be uploaded to a shared repository. Group members view these notes with browsers that allow them to retrieve all notes taken in a given context or to access notes from other related notes or documents. This is possible because NotePals records the context in which each note is created (e.g., its author, subject, and creation time). The system is lightweight because it fits easily into group members regular note- taking practices, and uses informal, ink-based user interfaces that run on portable, inexpensive hardware. In this paper we describe NotePals, show how we have used it to share our notes, and present our evaluations of the system. Richard C. Davis, James A. Landay, Jonathan Huang, Rebecca B. Lee |
CHI | 4 |
| 1992 | Inductive techniques for formal verification of systolic array designs in DSP applicationsabstractIt is shown how several inductive techniques can be utilized to provide fast and efficient proofs to the correctness of systolic designs in digital signal processing (DSP) applications. These techniques exploit the repeatability, regularity, and locality nature of systolic arrays and algorithms in DSP to produce fast proofs independent of the array size. How inductive techniques can be applied to different array topologies suitable for DSP is also shown, and the structure of the verifier developed to automate induction using logic programming is illustrated.> Nam Ling, Timothy K. Shih, Jonathan Huang |
ICASSP | 3 |