VLDB 2026 Research / reviewers in the wild / expert
Jing Huang 0020
dblp:14/4834-20
· DBLP profile ↗
19ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0003-3488-3641ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions—including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. https://ego-exo4d-data.org/ Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, David Crandall, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
Int. J. Comput. Vis. | 23 |
| 2024 | Real-Time Simulated Avatar from Head-Mounted SensorsabstractWe present SimXR, a methodfor controlling a simulated avatar from information (headset pose and cameras) ob-tained from AR / VR headsets. Due to the challenging view-point of head-mounted cameras, the human body is often clipped out of view, making traditional image-based ego-centric pose estimation challenging. On the other hand, headset poses provide valuable information about overall body motion, but lack fine-grained details about the hands and feet. To synergize headset poses with cameras, we control a humanoid to track headset movement while analyzing input images to decide body movement. When body parts are seen, the movements of hands and feet will be guided by the images; when unseen, the laws of physics guide the controller to generate plausible motion. We design an end-to-end method that does not rely on any intermediate representations and learns to directly map from images and headset poses to humanoid control signals. To train our method, we also propose a large-scale synthetic dataset created using camera configurations compatible with a commercially available VR headset (Quest 2) and show promising results on real-world captures. To demonstrate the applicability of our framework, we also test it on an AR headset with a forward-facing camera. Zhengyi Luo 0002, Jinkun Cao, Rawal Khirodkar, Alexander Winkler, Jing Huang 0020, Kris Makoto Kitani, Weipeng Xu |
CVPR | 5 |
| 2024 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multi-modal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured ego-centric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is un-precedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions-including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
CVPR | 23 |
| 2024 | Universal Humanoid Motion Representations for Physics-Based ControlabstractWe present a universal motion representation that encompasses a comprehensive range of motor skills for physics-based humanoid control. Due to the high dimensionality of humanoids and the inherent difficulties in reinforcement learning, prior methods have focused on learning skill embeddings for a narrow range of movement styles (e.g. locomotion, game characters) from specialized motion datasets. This limited scope hampers their applicability in complex tasks. We close this gap by significantly increasing the coverage of our motion representation space. To achieve this, we first learn a motion imitator that can imitate all of human motion from a large, unstructured motion dataset. We then create our motion representation by distilling skills directly from the imitator. This is achieved by using an encoder-decoder structure with a variational information bottleneck. Additionally, we jointly learn a prior conditioned on proprioception (humanoid's own pose and velocities) to improve model expressiveness and sampling efficiency for downstream tasks. By sampling from the prior, we can generate long, stable, and diverse human motions. Using this latent space for hierarchical RL, we show that our policies solve tasks using human-like behavior. We demonstrate the effectiveness of our motion representation by solving generative tasks (e.g. strike, terrain traversal) and motion tracking using VR controllers. Zhengyi Luo 0002, Jinkun Cao, Josh Merel, Alexander Winkler, Jing Huang 0020, Kris Makoto Kitani, Weipeng Xu |
ICLR | 5 |
| 2024 | HyperMix: Out-of-Distribution Detection and Classification in Few-Shot SettingsabstractOut-of-distribution (OOD) detection is an important topic for real-world machine learning systems, but settings with limited in-distribution samples have been underexplored. Such few-shot OOD settings are challenging, as models have scarce opportunities to learn the data distribution before being tasked with identifying OOD samples. Indeed, we demonstrate that recent state-of-the-art OOD methods fail to outperform simple baselines in the few-shot setting. We thus propose a hypernetwork framework called HyperMix, using Mixup on the generated classifier parameters, as well as a natural out-of-episode outlier exposure technique that does not require an additional outlier dataset. We conduct experiments on CIFAR-FS and MiniImageNet, significantly outperforming other OOD methods in the few-shot regime. Nikhil Mehta 0002, Kevin J. Liang, Jing Huang 0020, Fu-Jen Chu, Tal Hassner |
WACV | 3 |
| 2023 | Unsupervised Melody-to-Lyrics GenerationabstractYufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Gunnar Sigurdsson, Chenyang Tao, Wenbo Zhao, Tagyoung Chung, Jing Huang, Nanyun Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Gunnar A. Sigurdsson, Chenyang Tao, Wenbo Zhao 0006, Tagyoung Chung, Jing Huang 0020, Nanyun Peng 0001 |
ACL (1) | 9 |
| 2023 | Self-Supervised Object Detection from Egocentric VideosabstractUnderstanding the visual world from human perspectives has been a long-standing challenge in computer vision. Egocentric videos exhibit high scene complexity and irregular motion flows compared to typical video understanding tasks. With the egocentric domain in mind, we address the problem of self-supervised, class-agnostic object detection, aiming to locate all objects in a given view, without any annotations or pre-trained weights. Our method, self-supervised object detection from egocentric videos (DEVI), generalizes appearance-based methods to learn features end-to-end that are category-specific and invariant to viewing angle and illumination. Our approach leverages natural human behavior in egocentric perception to sample diverse views of objects for our multi-view and scale-regression losses, and our cluster residual module learns multi-category patches for complex scene understanding. DEVI results in gains up to 4.11% AP50, 0.11% AR1, 1.32% AR10, and 5.03% AR100on recent egocentric datasets, while significantly reducing model complexity. We also demonstrate competitive performance on out-of-domain datasets without additional training or fine-tuning. Peri Akiva, Jing Huang 0020, Kevin J. Liang, Rama Kovvuri, Matt Feiszli, Kristin J. Dana, Tal Hassner |
ICCV | 2 |
| 2022 | ExPUNations: Augmenting Puns with Keywords and ExplanationsabstractJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Tagyoung Chung, Jing Huang, Yang Liu, Nanyun Peng. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Tagyoung Chung, Jing Huang 0020, Yang Liu 0004, Nanyun Peng 0001 |
EMNLP | 6 |
| 2022 | Context-Situated Pun GenerationabstractJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Shuyang Gao, Tagyoung Chung, Jing Huang, Yang Liu, Nanyun Peng. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jiao Sun, Anjali Narayan-Chen, Shereen Oraby, Shuyang Gao, Tagyoung Chung, Jing Huang 0020, Yang Liu 0004, Nanyun Peng 0001 |
EMNLP | 6 |
| 2021 | A Multiplexed Network for End-to-End, Multilingual OCRabstractRecent advances in OCR have shown that an end-to-end (E2E) training pipeline that includes both detection and recognition leads to the best results. However, many existing methods focus primarily on Latin-alphabet languages, often even only case-insensitive English characters. In this paper, we propose an E2E approach, Multiplexed Multilingual Mask TextSpotter, that performs script identification at the word level and handles different scripts with different recognition heads, all while maintaining a unified loss that simultaneously optimizes script identification and multiple recognition heads. Experiments show that our method outperforms the single-head model with similar number of parameters in end-to-end recognition tasks, and achieves state-of-the-art results on MLT17 and MLT19 joint text detection and script identification benchmarks. We believe that our work is a step towards the end-to-end trainable and scalable multilingual multi-purpose OCR system. Our code and model will be released. Jing Huang 0020, Guan Pang, Rama Kovvuri, Mandy Toh, Kevin J. Liang, Praveen Krishnan, Xi Yin 0001, Tal Hassner |
CVPR | 1 |
| 2021 | TextOCR: Towards Large-Scale End-to-End Reasoning for Arbitrary-Shaped Scene TextabstractA crucial component for the scene text based reasoning required for TextVQA and TextCaps datasets involve detecting and recognizing text present in the images using an optical character recognition (OCR) system. The current systems are crippled by the unavailability of ground truth text annotations for these datasets as well as lack of scene text detection and recognition datasets on real images disallowing the progress in the field of OCR and evaluation of scene text based reasoning in isolation from OCR systems. In this work, we propose TextOCR, an arbitrary-shaped scene text detection and recognition with 900k annotated words collected on real images from TextVQA dataset. We show that current state-of-the-art text-recognition (OCR) models fail to perform well on TextOCR and that training on TextOCR helps achieve state-of-the-art performance on multiple other OCR datasets as well. We use a TextOCR trained OCR model to create PixelM4C model which can do scene text based reasoning on an image in an end-to-end fashion, allowing us to revisit several design choices to achieve new state-of-the-art performance on TextVQA dataset. Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang 0020, Wojciech Galuba, Tal Hassner |
CVPR | 4 |
| 2020 | Mask TextSpotter v3: Segmentation Proposal Network for Robust Scene Text Spotting
Minghui Liao, Guan Pang, Jing Huang 0020, Tal Hassner, Xiang Bai |
ECCV (11) | 3 |
| 2019 | A Computer Vision Perspective on Analyzing and Synthesizing Geospatial DataabstractThe AI era sustains its foundations from the availability of large datasets. Especially geospatial datasets are very interesting from a computer vision perspective, as they enable us to understand the world we live in. Although many application domains arise from analyzing such big data, analysis itself is not enough for impacting lives. As its counter part, synthesis approaches are recently being developed for mimicking real-world data for completing and creating new worlds. In this paper, we will explore not only example analysis methods developed using large public datasets, but also some generative models to propose realistic and impactful solutions for going beyond observations. Ilke Demir, Guan Pang, Jing Huang 0020 |
IGARSS | 3 |
| 2016 | Vehicle detection in urban point clouds with orthogonal-view convolutional neural networkabstractIn this paper, we aim at detecting vehicles from the point clouds scanned from the urban area. Our detection method consists of a segmentation stage and a classification stage. Prior knowledge for vehicles and urban environment is utilized to help the detection process. Specifically, we incorporate curb detection and removal in the segmentation stage. Moreover, our approach is able to estimate the orientation of the candidates and use it to handle the difficult cases such as the vehicles in the parking lot. In order to distinguish the vehicles from other segments among the 3D point cloud candidates, we develop three architectures of the orthogonal-view CNN, which are based on the orthogonal view projections of the candidates. Detailed evaluations and comparisons are performed on a challenging point cloud dataset of urban area. Jing Huang 0020, Suya You |
ICIP | 1 |
| 2016 | Point cloud labeling using 3D Convolutional Neural NetworkabstractIn this paper, we tackle the labeling problem for 3D point clouds. We introduce a 3D point cloud labeling scheme based on 3D Convolutional Neural Network. Our approach minimizes the prior knowledge of the labeling problem and does not require a segmentation step or hand-crafted features as most previous approaches did. Particularly, we present solutions for large data handling during the training and testing process. Experiments performed on the urban point cloud dataset containing 7 categories of objects show the robustness of our approach. Jing Huang 0020, Suya You |
ICPR | 1 |
| 2015 | Pole-like object detection and classification from urban point cloudsabstractThis paper focuses on detecting and classifying pole-like objects from point clouds obtained in urban areas. To achieve our goal, we propose a system consisting of three stages: localization, segmentation and classification. The localization algorithm based on slicing, clustering, pole seed generation and bucket augmentation takes advantage of the unique characteristics of pole-like objects and avoids heavy computation on the feature of every point in traditional methods. Then, the bucket-shaped neighborhood of the segments is integrated and trimmed with region growing algorithms, reducing the noises within candidate's neighborhood. Finally, we introduce a representation of six attributes based on the height and five point classes closely related to the pole categories and apply SVM to classify the candidate objects into 4 categories, including 3 pole categories light, utility pole and sign, and the non-pole category. The performance of our method is demonstrated through comparison with previous works on a large-scale urban dataset. Jing Huang 0020, Suya You |
ICRA | 1 |
| 2015 | Change Detection in Laser-Scanned Data of Industrial SitesabstractAs laser scanners become widely used in 3D data acquisition of industrial sites, one challenging problem emerges: given two data of the same site scanned/modeled at different times, how can we tell the difference between the two? In this paper, we formulate this problem as the 3D change detection problem, and propose a novel method for detecting object-level changes. In general, we notice that the changes can be viewed as the inconsistency between the global alignment and the local alignment. Therefore, we propose a change detection framework that comprises global alignment, local object detection and a novel change detection method. Specifically, we propose a series of change evaluation functions for pair wise change inference, based on which we formulate the many-to-many object change correlation problem as the weighted bipartite matching problem which could be solved efficiently. Finally, we demonstrate the feasibility of our approach through experiments on both synthetic and real industrial datasets. Jing Huang 0020, Suya You |
WACV | 1 |
| 2014 | Segmentation and matching: Towards a robust object detection systemabstractThis paper focuses on detecting parts in laser-scanned data of a cluttered industrial scene. To achieve the goal, we propose a robust object detection system based on segmentation and matching, as well as an adaptive segmentation algorithm and an efficient pose extraction algorithm based on correspondence filtering. We also propose an overlapping-based criterion that exploits more information of the original point cloud than the number-of-matching criterion that only considers key-points. Experiments show how each component works and the results demonstrate the performance of our system compared to the state of the art. Jing Huang 0020, Suya You |
WACV | 1 |
| 2013 | Detecting Objects in Scene Point Cloud: A Combinational ApproachabstractObject detection is a fundamental task in computer vision. As the 3D scanning techniques become popular, directly detecting objects through 3D point cloud of a scene becomes an immediate need. We propose an object detection framework combining learning-Based classification, local descriptor, a new variance of RANSAC imposing rigid-body constraint and an iterative process for multi-object detection in continuous point clouds. The framework not only takes global and local information into account, but also benefits from both learning and empirical methods. The experiments performed on the challenging ground Lidar dataset show the effectiveness of our method. Jing Huang 0020, Suya You |
3DV | 1 |