VLDB 2026 Research / reviewers in the wild / expert
Huaizu Jiang
dblp:128/7890
· DBLP profile ↗
41ranked-venue papers
12as first author
24since 2021 · last 2026
0000-0002-2300-4237ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 8 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 11 first-author · 12 since 2021Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SNAP: Towards Segmenting Anything in Any Point CloudabstractInteractive 3D point cloud segmentation enables efficient annotation of complex 3D scenes through user-guided prompts. However, current approaches are typically restricted in scope to a single domain (indoor or outdoor), and to a single form of user interaction (either spatial clicks or textual prompts). Moreover, training on multiple datasets often leads to negative transfer, resulting in domain-specific tools that lack generalizability. To address these limitations, we present SNAP (Segment aNything in Any Point cloud), a unified model for interactive 3D segmentation that supports both point-based and text-based prompts across diverse domains. Our approach achieves cross-domain generalizability by training on 7 datasets spanning indoor, outdoor, and aerial environments, while employing domain-adaptive normalization to prevent negative transfer. For text-prompted segmentation, we automatically generate mask proposals without human intervention and match them against CLIP embeddings of textual queries, enabling both panoptic and open-vocabulary segmentation. Extensive experiments demonstrate that SNAP consistently delivers high-quality segmentation results. We achieve state-of-the-art performance on 8 out of 9 zero-shot benchmarks for spatial-prompted segmentation and demonstrate competitive results on all 5 text-prompted benchmarks. These results show that a unified model can match or exceed specialized domain-specific approaches, providing a practical tool for scalable 3D annotation. Project page is at https://neu-vi.github.io/SNAP/ Aniket Gupta, Hanhui Wang, Charles Saunders, Aruni Roy Chowdhury, Hanumant Singh, Huaizu Jiang |
3DV | 6 |
| 2026 | Anatomically-guided masked autoencoder pre-training for aneurysm detectionabstractIntracranial aneurysms are a major cause of morbidity and mortality worldwide, and detecting them manually is a complex, time-consuming task. Albeit automated solutions are desirable, the limited availability of training data makes it difficult to develop such solutions using typical supervised learning frameworks. In this work, we propose a novel pre-training strategy using more widely available unannotated head CT scan data to pre-train a 3D Vision Transformer model prior to fine-tuning for the aneurysm detection task. Specifically, we modify masked auto-encoder (MAE) pre-training in the following ways: we use a factorized self-attention mechanism to make 3D attention computationally viable, we restrict the masked patches to areas near arteries to focus on areas where aneurysms are likely to occur, and we reconstruct not only CT scan intensity values but also artery distance maps, which describe the distance between each voxel and the closest artery, thereby enhancing the backbone's learned representations. Compared with SOTA aneurysm detection models, our approach gains +2-7% absolute Sensitivity at a false positive rate of 0.5. Code and weights are available at: https://github.com/alceballosa/aneurysm-mae. Alberto Mario Ceballos-Arroyo, Chu-Hsuan Lin, Geoffrey S. Young, Huaizu Jiang |
WACV | 6 |
| 2025 | Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked AutoregressionabstractSince 2023, Vector Quantization (VQ)-based discrete generation methods have rapidly dominated human motion generation, primarily surpassing diffusion-based continuous generation methods in standard performance metrics. However, VQ-based methods have inherent limitations. Representing continuous motion data as limited discrete tokens leads to inevitable information loss, reduces the diversity of generated motions, and restricts their ability to function effectively as motion priors or generation guidance. In contrast, the continuous space generation nature of diffusion-based methods makes them well-suited to address these limitations and with even potential for model scalability. In this work, we systematically investigate why current VQ-based methods perform well and explore the limitations of existing diffusion-based methods from the perspective of motion data representation and distribution. Drawing on these insights, we preserve the inherent strengths of a diffusion-based human motion generation model and gradually optimize it with inspiration from VQ-based approaches. Our approach introduces a human motion diffusion model enabled to perform masked autoregression, optimized with a reformed data representation and distribution. Additionally, we propose a more robust evaluation method to assess different approaches. Extensive experiments on various datasets demonstrate our method outperforms previous methods and achieves state-of-the-art performances. Zichong Meng, Xiaogang Peng, Zeyu Han, Huaizu Jiang |
CVPR | 5 |
| 2025 | HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Models
Vikram Voleti, Varun Jampani, Huaizu Jiang |
ICCV | 5 |
| 2025 | SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation
Chun-Han Yao, Vikram Voleti, Huaizu Jiang, Varun Jampani |
ICCV | 4 |
| 2025 | SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View ConsistencyabstractWe present Stable Video 4D (SV4D) — a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and novel view synthesis, we design a unified diffusion model to generate novel view videos of dynamic 3D objects. Specifically, given a monocular reference video, SV4D generates novel views for each video frame that are temporally consistent. We then use the generated novel view videos to optimize an implicit 4D representation (dynamic NeRF) efficiently, without the need for cumbersome SDS-based optimization used in most prior works. To train our unified novel view video generation model, we curate a dynamic 3D object dataset from the existing Objaverse dataset. Extensive experimental results on multiple datasets and user studies demonstrate SV4D's state-of-the-art performance on novel-view video synthesis as well as 4D generation compared to prior works. Project page: https://sv4d.github.io. Chun-Han Yao, Vikram Voleti, Huaizu Jiang, Varun Jampani |
ICLR | 4 |
| 2025 | A System for Multi-View Mapping of Dynamic Scenes Using Time-Synchronized UAVsabstractRecent advances in 3D scene reconstruction, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting, have demonstrated remarkable results in novel view synthesis and dynamic scene representation. Despite these successes, existing approaches rely on time-synchronized multi-view imagery captured using specialized camera rigs in controlled environments. This reliance limits their applicability in uncontrolled, unbounded dynamic scenes. In this work, we propose a novel Unmanned Aerial Vehicle (UAV) based multi-view capture system that leverages GNSS Pulse Per Second (PPS) signals for precise frame synchronization across multiple cameras. Our system eliminates the need for fixed infrastructure, enabling flexible and scalable data collection for dynamic scene reconstruction in diverse environments. In addition to the system architecture, we also introduce a dataset of synchronized multi-view images captured in unbounded outdoor scenes from four synchronized UAVs, each carrying a stereo camera rig. We benchmark several 3D and 4D representation methods on our dataset and highlight the challenges associated with data collection in unstructured outdoor settings such as sparse views, varied lighting conditions, visual degradation etc. Our hardware configuration details, software details and dataset is available at https://github.com/neufieldrobotics/Dynamic_Mapping. Aniket Gupta, Dennis Giaya, Vishnu Rohit Annadanam, Mithun Diddi, Huaizu Jiang, Hanumant Singh |
IROS | 5 |
| 2025 | NeuFlow-V2: Push High-Efficiency Optical Flow To the LimitabstractReal-time high-accuracy optical flow estimation is critical for a variety of real-world robotic applications. However, current learning-based methods often struggle to balance accuracy and computational efficiency: methods that achieve high accuracy typically demand substantial processing power, while faster approaches tend to sacrifice precision. These fast approaches specifically falter in their generalization capabilities and do not perform well across diverse real-world scenarios. In this work, we revisit the limitations of the SOTA methods and present NeuFlow-V2, a novel method that offers both — high accuracy in real-world datasets coupled with low computational overhead. In particular, we introduce a novel light-weight backbone and a fast refinement module to keep computational demands tractable while delivering accurate optical flow. Experimental results on synthetic and real-world datasets demonstrate that NeuFlow-V2 provides similar accuracy to SOTA methods while achieving 10x-70x speedups. It is capable of running at over 20 FPS on 512x384 resolution images on a Jetson Orin Nano. The full training and evaluation code is available at https://github.com/ neufieldrobotics/NeuFlow_v2. Aniket Gupta, Huaizu Jiang, Hanumant Singh |
IROS | 3 |
| 2025 | Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMsabstractUnlocking spatial reasoning in Multimodal Large Language Models (MLLMs) is crucial for enabling intelligent interaction with 3D environments. While prior efforts often rely on explicit 3D inputs or specialized model architectures, we ask: can MLLMs reason about 3D space using only structured 2D representations derived from perception?
In this work, we introduce Struct2D, a perception-guided prompting framework that combines bird’s-eye-view (BEV) images with object marks and object-centric metadata, optionally incorporating egocentric keyframes when needed. Using Struct2D, we conduct an in-depth zero-shot analysis of closed-source MLLMs (e.g., GPT-o3) and find that they exhibit surprisingly strong spatial reasoning abilities when provided with projected 2D inputs, effectively handling tasks such as relative direction estimation and route planning.
Motivated by these findings, we construct a large-scale instructional tuning dataset, \textbf{Struct2D-Set}, using an automated pipeline that generates fine-grained QA pairs grounded in 3D indoor scenes. We then fine-tune an open-source MLLM (Qwen2.5VL) using Struct2D-Set, relying on noisy 3D perception rather than ground-truth annotations. Despite this, the tuned model achieves strong performance across multiple spatial reasoning benchmarks, including 3D question answering, captioning, and object grounding, spanning eight diverse reasoning categories.
Our approach demonstrates that structured 2D inputs can effectively bridge perception and language reasoning in MLLMs—without requiring explicit 3D representations as input. We will release both our code and dataset to support future research. Fangrui Zhu, Hanhui Wang, Tianye Ding, Huaizu Jiang |
NeurIPS | 7 |
| 2025 | Diagnosing Human-Object Interaction DetectorsabstractAbstract We have witnessed significant progress in human-object interaction (HOI) detection. However, relying solely on mAP (mean Average Precision) scores as a summary metric does not provide sufficient insight into the nuances of model performance ( e.g. , why one model outperforms another), which can hinder further innovation in this field. To address this issue, we introduce a diagnosis toolbox in this paper to offer a detailed quantitative breakdown of HOI detection models, inspired by the success of object detection diagnosis tools. We first conduct a holistic investigation into the HOI detection pipeline. By defining a set of errors and using oracles to fix each one, we quantitatively analyze the significance of different errors based on the mAP improvement gained from fixing them. Next, we explore the two key sub-tasks of HOI detection: human-object pair localization and interaction classification. For the pair localization task, we compute the coverage of ground-truth human-object pairs and assess the noisiness of the localization results. For the classification task, we measure a model’s ability to distinguish between positive and negative detection results and to classify actual interactions when human-object pairs are correctly localized. We analyze eight state-of-the-art HOI detection models, providing valuable diagnostic insights to guide future research. For instance, our diagnosis reveals that the state-of-the-art model RLIPv2 outperforms others primarily due to its significant improvement in multi-label interaction classification accuracy. Our toolbox is applicable across various methods and datasets and is available at https://neu-vi.github.io/Diag-HOI/ . Fangrui Zhu, Weidi Xie, Huaizu Jiang |
Int. J. Comput. Vis. | 4 |
| 2024 | Zero-Shot Referring Expression Comprehension via Structural Similarity Between Images and CaptionsabstractZero-shot referring expression comprehension aims at localizing bounding boxes in an image corresponding to provided textual prompts, which requires: (i) afine-grained disentanglement of complex visual scene and textual context, and (ii) a capacity to understand relationships among disentangled entities. Unfortunately, existing large vision-language alignment (VLA) models, e.g., CLIP, struggle with both aspects so cannot be directly used for this task. To mitigate this gap, we leverage large foundation models to disentangle both images and texts into triplets in the for-mat of (subject, predicate, object). After that, grounding is accomplished by calculating the structural similarity matrix between visual and textual triplets with a VLA model, and subsequently propagate it to an instance-level similarity matrix. Furthermore, to equip VLA models with the ability of relationship understanding, we design a triplet-matching objective to fine-tune the VLA models on a collection of curated dataset containing abundant entity relationships. Experiments demonstrate that our visual grounding performance increase of up to 19.5% over the SOTA zero-shot model on RefCOCO/+/g. On the more challenging Who's Waldo dataset, our zero-shot approach achieves comparable accuracy to the fully super-vised model. Code is available at https://github.com/Show-han/Zeroshot_REC. Zeyu Han, Fangrui Zhu, Qianru Lao, Huaizu Jiang |
CVPR | 4 |
| 2024 | SMooDi: Stylized Motion Diffusion Model
Varun Jampani, Deqing Sun, Huaizu Jiang |
ECCV (1) | 5 |
| 2024 | OmniControl: Control Any Joint at Any Time for Human Motion GenerationabstractWe present a novel approach named OmniControl for incorporating flexible spatial control signals into a text-conditioned human motion generation model based on the diffusion process. Unlike previous methods that can only control the pelvis trajectory, OmniControl can incorporate flexible spatial control signals over different joints at different times with only one model. Specifically, we propose analytic spatial guidance that ensures the generated motion can tightly conform to the input control signals. At the same time, realism guidance is introduced to refine all the joints to generate more coherent motion. Both the spatial and realism guidance are essential and they are highly complementary for balancing control accuracy and motion realism. By combining them, OmniControl generates motions that are realistic, coherent, and consistent with the spatial constraints. Experiments on HumanML3D and KIT-ML datasets show that OmniControl not only achieves significant improvement over state-of-the-art methods on pelvis control but also shows promising results when incorporating the constraints over other joints. Project page: https://neu-vi.github.io/omnicontrol/. Varun Jampani, Deqing Sun, Huaizu Jiang |
ICLR | 5 |
| 2024 | StereoNavNet: Learning to Navigate using Stereo Cameras with Auxiliary Occupancy VoxelsabstractVisual navigation has received significant attention recently. Most of the prior works focus on predicting navigation actions based on semantic features extracted from visual encoders. However, these approaches often rely on large datasets and exhibit limited generalizability. In contrast, our approach draws inspiration from traditional navigation planners that operate on geometric representations, such as occupancy maps. We propose StereoNavNet (SNN), a novel visual navigation approach employing a modular learning framework comprising perception and policy modules. Within the perception module, we estimate an auxiliary 3D voxel occupancy grid from stereo RGB images and extract geometric features from it. These features, along with user-defined goals, are utilized by the policy module to predict navigation actions. Through extensive empirical evaluation, we demonstrate that SNN outperforms baseline approaches in terms of success rates, success weighted by path length, and navigation error. Furthermore, SNN exhibits better generalizability, characterized by maintaining leading performance when navigating across previously unseen environments. Hongyu Li 0003, Taskin Padir, Huaizu Jiang |
IROS | 3 |
| 2024 | ODTFormer: Efficient Obstacle Detection and Tracking with Stereo Cameras Based on TransformerabstractObstacle detection and tracking represent a critical component in robot autonomous navigation. In this paper, we propose ODTFormer, a Transformer-based model that addresses both obstacle detection and tracking problems. For the detection task, our approach leverages deformable attention to construct a 3D cost volume, which is decoded progressively in the form of voxel occupancy grids. We further track the obstacles by matching the voxels between consecutive frames. The entire model can be optimized in an end-to-end manner. Through extensive experiments on DrivingStereo and KITTI benchmarks, our model achieves state-of-the-art performance in the obstacle detection task. We also report comparable accuracy to state-of-the-art obstacle tracking models while requiring only a fraction of their computation cost, typically ten-fold to twenty-fold less. Our code is available on https://github.com/neu-vi/ODTFormer. Tianye Ding, Hongyu Li 0003, Huaizu Jiang |
IROS | 3 |
| 2024 | NeuFlow: Real-time, High-accuracy Optical Flow Estimation on Robots Using Edge DevicesabstractReal-time high-accuracy optical flow estimation is a crucial component in various applications, including localization and mapping in robotics, object tracking, and activity recognition in computer vision. While recent learning-based optical flow methods have achieved high accuracy, they often come with heavy computation costs. In this paper, we propose a highly efficient optical flow architecture, called NeuFlow, that addresses both high accuracy and computational cost concerns. The architecture follows a global-to-local scheme. Given the features of the input images extracted at different spatial resolutions, global matching is employed to estimate an initial optical flow on the 1/16 resolution, capturing large displacement, which is then refined on the 1/8 resolution with lightweight CNN layers for better accuracy. We evaluate our approach on Jetson Orin Nano and RTX 2080 to demonstrate efficiency improvements across different computing platforms. We achieve a notable 10-80 speedup compared to several state-of-the-art methods, while maintaining comparable accuracy. Our approach achieves around 30 FPS on edge computing platforms, which represents a significant breakthrough in deploying complex computer vision tasks such as SLAM on small robots like drones. The full training and evaluation code is available at https://github.com/neufieldrobotics/NeuFlow. Huaizu Jiang, Hanumant Singh |
IROS | 2 |
| 2024 | Vessel-Aware Aneurysm Detection Using Multi-scale Deformable 3D Attention
Alberto Mario Ceballos-Arroyo, Fangrui Zhu, Shrikanth M. Yadav, Geoffrey S. Young, Huaizu Jiang |
MICCAI (5) | 8 |
| 2024 | Towards Flexible Visual Relationship SegmentationabstractVisual relationship understanding has been studied separately in human-object interaction(HOI) detection, scene graph generation(SGG), and referring relationships(RR) tasks.
Given the complexity and interconnectedness of these tasks, it is crucial to have a flexible framework that can effectively address these tasks in a cohesive manner.
In this work, we propose FleVRS, a single model that seamlessly integrates the above three aspects in standard and promptable visual relationship segmentation, and further possesses the capability for open-vocabulary segmentation to adapt to novel scenarios.
FleVRS leverages the synergy between text and image modalities,
to ground various types of relationships from images and use textual features from vision-language models to visual conceptual understanding.
Empirical validation across various datasets demonstrates that our framework outperforms existing models in standard, promptable, and open-vocabulary tasks, e.g., +1.9 $mAP$ on HICO-DET, +11.4 $Acc$ on VRD, +4.7 $mAP$ on unseen HICO-DET.
Our FleVRS represents a significant step towards a more intuitive, comprehensive, and scalable understanding of visual relationships. Fangrui Zhu, Huaizu Jiang |
NeurIPS | 3 |
| 2023 | Pixel-Aligned Recurrent Queries for Multi-View 3D Object DetectionabstractWe present PARQ - a multi-view 3D object detector with transformer and pixel-aligned recurrent queries. Unlike previous works that use learnable features or only encode 3D point positions as queries in the decoder, PARQ leverages appearance-enhanced queries initialized from reference points in 3D space and updates their 3D location with recurrent cross-attention operations. Incorporating pixel-aligned features and cross attention enables the model to encode the necessary 3D-to-2D correspondences and capture global contextual information of the input images. PARQ outperforms prior best methods on the ScanNet and ARKitScenes datasets, learns and detects faster, is more robust to distribution shifts in reference points, can leverage additional input views without retraining, and can adapt inference compute by changing the number of recurrent iterations. Code is available at https://ymingxie.github.io/parq. Huaizu Jiang, Georgia Gkioxari, Julian Straub |
ICCV | 2 |
| 2023 | StereoVoxelNet: Real-Time Obstacle Detection Based on Occupancy Voxels from a Stereo Camera Using Deep Neural NetworksabstractObstacle detection is a safety-critical problem in robot navigation, where stereo matching is a popular vision-based approach. While deep neural networks have shown impressive results in computer vision, most of the previous obstacle detection works only leverage traditional stereo matching techniques to meet the computational constraints for real-time feedback. This paper proposes a computationally efficient method that employs a deep neural network to detect occupancy from stereo images directly. Instead of learning the point cloud correspondence from the stereo data, our approach extracts the compact obstacle distribution based on volumetric representations. In addition, we prune the computation of safety irrelevant spaces in a coarse-to-fine manner based on octrees generated by the decoder. As a result, we achieve real-time performance on the onboard computer (NVIDIA Jetson TX2). Our approach detects obstacles accurately in the range of 32 meters and achieves better IoU (Intersection over Union) and CD (Chamfer Distance) scores with only 2% of the computation cost of the state-of-the-art stereo model. Furthermore, we validate our method's robustness and real-world feasibility through autonomous navigation experiments with a real robot. Hence, our work contributes toward closing the gap between the stereo-based system in robot perception and state-of-the-art stereo models in computer vision. To counter the scarcity of high-quality real-world indoor stereo datasets, we collect a 1.36 hours stereo dataset with a mobile robot which is used to fine-tune our model. The dataset, the code, and further details including additional visualizations are available at https://lhy.xyz/stereovoxelnet/. Hongyu Li 0003, Zhengang Li 0001, Neset Ünver Akmandor, Huaizu Jiang, Yanzhi Wang 0001, Taskin Padir |
ICRA | 4 |
| 2023 | DCVNet: Dilated Cost Volume Networks for Fast Optical FlowabstractThe cost volume, capturing the similarity of possible correspondences across two input images, is a key ingredient in state-of-the-art optical flow approaches. When sampling correspondences to build the cost volume, a large neighborhood radius is required to deal with large displacements, introducing a significant computational burden. To address this, coarse-to-fine or recurrent processing of the cost volume is usually adopted, where correspondence sampling in a local neighborhood with a small radius suffices. In this paper, we propose an alternative by constructing cost volumes with different dilation factors to capture small and large displacements simultaneously. A U-Net with sikp connections is employed to convert the dilated cost volumes into interpolation weights between all possible captured displacements to get the optical flow. Our proposed model DCVNet only needs to process the cost volume once in a simple feedforward manner and does not rely on the sequential processing strategy. DCVNet obtains comparable accuracy to existing approaches and achieves real-time inference (30 fps on a mid-end 1080ti GPU). Huaizu Jiang, Erik G. Learned-Miller |
WACV | 1 |
| 2022 | Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object InteractionsabstractA significant gap remains between today's visual pattern recognition models and humanlevel visual cognition especially when it comes to fewshot learning and compositional reasoning of novel concepts. We introduce Bongard-HOI, a new visual reasoning benchmark that focuses on compositional learning of humanobject interactions (HOIs) from natural images. It is inspired by two desirable characteristics from the classical Bongard problems (BPs): 1) fewshot concept learning, and 2) contextdependent reasoning. We carefully curate the fewshot instances with hard negatives, where positive and negative images only disagree on action labels, making mere recognition of object categories insufficient to complete our benchmarks. We also design multiple test sets to systematically study the generalization of visual learning models, where we vary the overlap of the HOI concepts between the training and test sets of fewshot instances, from partial to no overlaps. Bongard-HOI presents a substantial challenge to today's visual recognition models. The state-of-the-art HOI detection model achieves only 62% accuracy on fewshot binary prediction while even amateur human testers on MTurk have 91% accuracy. With the Bongard-HOI benchmark, we hope to further advance research efforts in visual reasoning, especially in holistic perception-reasoning systems and better representation learning. Huaizu Jiang, Xiaojian Ma 0001, Weili Nie, Zhiding Yu, Yuke Zhu, Anima Anandkumar |
CVPR | 1 |
| 2022 | PlanarRecon: Realtime 3D Plane Detection and Reconstruction from Posed Monocular VideosabstractWe present PlanarRecon - a novel framework for globally coherent detection and reconstruction of 3D planes from a posed monocular video. Unlike previous works that detect planes in 2D from a single image, PlanarRecon incrementally detects planes in 3D for each video fragment, which consists of a set of key frames, from a volumetric representation of the scene using neural networks. A learning-based tracking and fusion module is designed to merge planes from previous fragments to form a coherent global plane reconstruction. Such design allows Planar-Recon to integrate observations from multiple views within each fragment and temporal information across different ones, resulting in an accurate and coherent reconstruction of the scene abstraction with low-polygonal geometry. Experiments show that the proposed approach achieves state-of-the-art performances on the ScanNet dataset while being real-time. Code is available at the project page: https://neu-vi.github.io/planarrecon/. Matheus Gadelha, Fengting Yang, Xiaowei Zhou 0001, Huaizu Jiang |
CVPR | 5 |
| 2022 | RelViT: Concept-guided Vision Transformer for Visual Relational Reasoning
Xiaojian Ma 0001, Weili Nie, Zhiding Yu, Huaizu Jiang, Chaowei Xiao, Yuke Zhu, Song-Chun Zhu, Anima Anandkumar |
ICLR | 4 |
| 2020 | In Defense of Grid Features for Visual Question AnsweringabstractPopularized as `bottom-up' attention, bounding box (or region) based visual features have recently surpassed vanilla grid-based convolutional features as the de facto standard for vision and language tasks like visual question answering (VQA). However, it is not clear whether the advantages of regions (e.g. better localization) are the key reasons for the success of bottom-up attention. In this paper, we revisit grid features for VQA, and find they can work surprisingly well -- running more than an order of magnitude faster with the same accuracy (e.g. if pre-trained in a similar fashion). Through extensive experiments, we verify that this observation holds true across different VQA models (reporting a state-of-the-art accuracy on VQA 2.0 test-std, 72.71), datasets, and generalizes well to other tasks like image captioning. As grid features make the model design and training process much simpler, this enables us to train them end-to-end and also use a more flexible network design. We learn VQA models end-to-end, from pixels directly to answers, and show that strong performance is achievable without using any region annotations in pre-training. We hope our findings help further improve the scientific understanding and the practical application of VQA. Code and features will be made available. Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik G. Learned-Miller, Xinlei Chen |
CVPR | 1 |
| 2019 | Automatic Adaptation of Object Detectors to New Domains Using Self-TrainingabstractThis work addresses the unsupervised adaptation of an existing object detector to a new target domain. We assume that a large number of unlabeled videos from this domain are readily available. We automatically obtain labels on the target data by using high-confidence detections from the existing detector, augmented with hard (misclassified) examples acquired by exploiting temporal cues using a tracker. These automatically-obtained labels are then used for re-training the original model. A modified knowledge distillation loss is proposed, and we investigate several ways of assigning soft-labels to the training examples from the target domain. Our approach is empirically evaluated on challenging face and pedestrian detection tasks: a face detector trained on WIDER-Face, which consists of high-quality images crawled from the web, is adapted to a large-scale surveillance data set; a pedestrian detector trained on clear, daytime images from the BDD-100K driving data set is adapted to all other scenarios such as rainy, foggy, night-time. Our results demonstrate the usefulness of incorporating hard examples obtained from tracking, the advantage of using soft-labels via distillation loss versus hard-labels, and show promising performance as a simple method for unsupervised domain adaptation of object detectors, with minimal dependence on hyper-parameters. Aruni Roy Chowdhury, Prithvijit Chakrabarty, SouYoung Jin, Huaizu Jiang, Liangliang Cao, Erik G. Learned-Miller |
CVPR | 5 |
| 2019 | SENSE: A Shared Encoder Network for Scene-Flow EstimationabstractWe introduce a compact network for holistic scene flow estimation, called SENSE, which shares common encoder features among four closely-related tasks: optical flow estimation, disparity estimation from stereo, occlusion estimation, and semantic segmentation. Our key insight is that sharing features makes the network more compact, induces better feature representations, and can better exploit interactions among these tasks to handle partially labeled data. With a shared encoder, we can flexibly add decoders for different tasks during training. This modular design leads to a compact and efficient model at inference time. Exploiting the interactions among these tasks allows us to introduce distillation and self-supervised losses in addition to supervised losses, which can better handle partially labeled real-world data. SENSE achieves state-of-the-art results on several optical flow benchmarks and runs as fast as networks specifically designed for optical flow. It also compares favorably against the state of the art on stereo and scene flow, while consuming much less memory. Huaizu Jiang, Deqing Sun, Varun Jampani, Zhaoyang Lv, Erik G. Learned-Miller, Jan Kautz |
ICCV | 1 |
| 2019 | Salient object detection: A surveyabstractDetecting and segmenting salient objects from natural scenes, often referred to as salient object detection, has attracted great interest in computer vision. While many models have been proposed and several applications have emerged, a deep understanding of achievements and issues remains lacking. We aim to provide a comprehensive review of recent progress in salient object detection and situate this field among other closely related areas such as generic scene segmentation, object proposal generation, and saliency for fixation prediction. Covering 228 publications, we survey i) roots, key concepts, and tasks, ii) core techniques and main modeling trends, and iii) datasets and evaluation metrics for salient object detection. We also discuss open problems such as evaluation metrics and dataset bias in model performance, and suggest future research directions. Ali Borji, Ming-Ming Cheng, Qibin Hou, Huaizu Jiang, Jia Li 0003 |
Comput. Vis. Media | 4 |
| 2019 | Joint salient object detection and existence prediction
Huaizu Jiang, Ming-Ming Cheng, Shijie Li 0006, Ali Borji, Jingdong Wang 0001 |
Frontiers Comput. Sci. | 1 |
| 2018 | Super SloMo: High Quality Estimation of Multiple Intermediate Frames for Video InterpolationabstractGiven two consecutive frames, video interpolation aims at generating intermediate frame(s) to form both spatially and temporally coherent video sequences. While most existing methods focus on single-frame interpolation, we propose an end-to-end convolutional neural network for variable-length multi-frame video interpolation, where the motion interpretation and occlusion reasoning are jointly modeled. We start by computing bi-directional optical flow between the input images using a U-Net architecture. These flows are then linearly combined at each time step to approximate the intermediate bi-directional optical flows. These approximate flows, however, only work well in locally smooth regions and produce artifacts around motion boundaries. To address this shortcoming, we employ another U-Net to refine the approximated flow and also predict soft visibility maps. Finally, the two input images are warped and linearly fused to form each intermediate frame. By applying the visibility maps to the warped images before fusion, we exclude the contribution of occluded pixels to the interpolated intermediate frame to avoid artifacts. Since none of our learned network parameters are time-dependent, our approach is able to produce as many intermediate frames as needed. To train our network, we use 1,132 240-fps video clips, containing 300K individual video frames. Experimental results on several datasets, predicting different numbers of interpolated frames, demonstrate that our approach performs consistently better than existing methods. Huaizu Jiang, Deqing Sun, Varun Jampani, Ming-Hsuan Yang 0001, Erik G. Learned-Miller, Jan Kautz |
CVPR | 1 |
| 2018 | Self-Supervised Relative Depth Learning for Urban Scene Understanding
Huaizu Jiang, Gustav Larsson, Michael Maire, Gregory Shakhnarovich, Erik G. Learned-Miller |
ECCV (11) | 1 |
| 2018 | Unsupervised Hard Example Mining from Videos for Improved Object Detection
SouYoung Jin, Aruni Roy Chowdhury, Huaizu Jiang, Aditya Prasad, Deep Chakraborty, Erik G. Learned-Miller |
ECCV (13) | 3 |
| 2017 | Face Detection with the Faster R-CNNabstractWhile deep learning based methods for generic object detection have improved rapidly in the last two years, most approaches to face detection are still based on the R-CNN framework [11], leading to limited accuracy and processing speed. In this paper, we investigate applying the Faster RCNN [26], which has recently demonstrated impressive results on various object detection benchmarks, to face detection. By training a Faster R-CNN model on the large scale WIDER face dataset [34], we report state-of-the-art results on the WIDER test set as well as two other widely used face detection benchmarks, FDDB and the recently released IJB-A. Huaizu Jiang, Erik G. Learned-Miller |
FG | 1 |
| 2017 | Reasoning About Fine-Grained Attribute Phrases Using Reference GamesabstractWe present a framework for learning to describe finegrained visual differences between instances using attribute phrases. Attribute phrases capture distinguishing aspects of an object (e.g., “propeller on the nose” or “door near the wing” for airplanes) in a compositional manner. Instances within a category can be described by a set of these phrases and collectively they span the space of semantic attributes for a category. We collect a large dataset of such phrases by asking annotators to describe several visual differences between a pair of instances within a category. We then learn to describe and ground these phrases to images in the context of a reference game between a speaker and a listener. The goal of a speaker is to describe attributes of an image that allows the listener to correctly identify it within a pair. Data collected in a pairwise manner improves the ability of the speaker to generate, and the ability of the listener to interpret visual descriptions. Moreover, due to the compositionality of attribute phrases, the trained listeners can interpret descriptions not seen during training for image retrieval, and the speakers can generate attribute-based explanations for differences between previously unseen categories. We also show that embedding an image into the semantic space of attribute phrases derived from listeners offers 20% improvement in accuracy over existing attributebased representations on the FGVC-aircraft dataset. Jong-Chyi Su, Chenyun Wu, Huaizu Jiang, Subhransu Maji |
ICCV | 3 |
| 2017 | Salient Object Detection: A Discriminative Regional Feature Integration Approach
Jingdong Wang 0001, Huaizu Jiang, Zejian Yuan, Ming-Ming Cheng, Xiaowei Hu 0003, Nanning Zheng 0001 |
Int. J. Comput. Vis. | 2 |
| 2015 | Salient Object Detection: A BenchmarkabstractWe extensively compare, qualitatively and quantitatively, 41 state-of-the-art models (29 salient object detection, 10 fixation prediction, 1 objectness, and 1 baseline) over seven challenging data sets for the purpose of benchmarking salient object detection and segmentation methods. From the results obtained so far, our evaluation shows a consistent rapid progress over the last few years in terms of both accuracy and running time. The top contenders in this benchmark significantly outperform the models identified as the best in the previous benchmark conducted three years ago. We find that the models designed specifically for salient object detection generally work better than models in closely related areas, which in turn provides a precise definition and suggests an appropriate treatment of this problem that distinguishes it from other problems. In particular, we analyze the influences of center bias and scene complexity in model performance, which, along with the hard cases for the state-of-the-art models, provide useful hints toward constructing more challenging large-scale data sets and better saliency models. Finally, we propose probable solutions for tackling several open problems, such as evaluation scores and data set bias, which also suggest future research directions in the rapidly growing field of salient object detection. Ali Borji, Ming-Ming Cheng, Huaizu Jiang, Jia Li 0003 |
IEEE Trans. Image Process. | 3 |
| 2015 | Online Multi-Target Tracking With Unified Handling of Complex ScenariosabstractComplex scenarios, including miss detections, occlusions, false detections, and trajectory terminations, make the data association challenging. In this paper, we propose an online tracking-by-detection method to track multiple targets with unified handling of aforementioned complex scenarios, where current detection responses are linked to the previous trajectories. We introduce a dummy node to each trajectory to allow it to temporally disappear. If a trajectory fails to find its matching detection, it will be linked to its corresponding dummy node until the emergence of its matching detection. Source nodes are also incorporated to account for the entrance of new targets. The standard Hungarian algorithm, extended by the dummy nodes, can be exploited to solve the online data association implicitly in a global manner, although it is formulated between two consecutive frames. Moreover, as dummy nodes tend to accumulate in a fake or disappeared trajectory while they only occasionally appear in a real trajectory, we can deal with false detections and trajectory terminations by simply checking the number of consecutive dummy nodes. Our approach works on a single, uncalibrated camera, and requires neither scene prior knowledge nor explicit occlusion reasoning, running at 132 frames/s on the PETS09-S2L1 benchmark sequence. The experimental results validate the effectiveness of the dummy nodes in complex scenarios and show that our proposed approach is robust against false detections and miss detections. Quantitative comparisons with other methods on five benchmark sequences demonstrate that we can achieve comparable results with the most existing offline methods and better results than other online algorithms. Huaizu Jiang, Jinjun Wang, Yihong Gong, Na Rong, Zhenhua Chai, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Regularized Tree Partitioning and Its Application to Unsupervised Image SegmentationabstractIn this paper, we propose regularized tree partitioning approaches. We study normalized cut (NCut) and average cut (ACut) criteria over a tree, forming two approaches: 1) normalized tree partitioning (NTP) and 2) average tree partitioning (ATP). We give the properties that result in an efficient algorithm for NTP and ATP. In addition, we present the relations between the solutions of NTP and ATP over the maximum weight spanning tree of a graph and NCut and ACut over this graph. To demonstrate the effectiveness of the proposed approaches, we show its application to image segmentation over the Berkeley image segmentation data set and present qualitative and quantitative comparisons with state-of-the-art methods. Jingdong Wang 0001, Huaizu Jiang, Yangqing Jia, Xian-Sheng Hua 0001, Changshui Zhang, Long Quan |
IEEE Trans. Image Process. | 2 |
| 2013 | Salient Object Detection: A Discriminative Regional Feature Integration ApproachabstractSalient object detection has been attracting a lot of interest, and recently various heuristic computational models have been designed. In this paper, we regard saliency map computation as a regression problem. Our method, which is based on multi-level image segmentation, uses the supervised learning approach to map the regional feature vector to a saliency score, and finally fuses the saliency scores across multiple levels, yielding the saliency map. The contributions lie in two-fold. One is that we show our approach, which integrates the regional contrast, regional property and regional background ness descriptors together to form the master saliency map, is able to produce superior saliency maps to existing algorithms most of which combine saliency maps heuristically computed from different types of features. The other is that we introduce a new regional feature vector, background ness, to characterize the background, which can be regarded as a counterpart of the objectness descriptor [2]. The performance evaluation on several popular benchmark data sets validates that our approach outperforms existing state-of-the-arts. Huaizu Jiang, Jingdong Wang 0001, Zejian Yuan, Yang Wu 0001, Nanning Zheng 0001, Shipeng Li 0001 |
CVPR | 1 |
| 2013 | Probabilistic salient object contour detection based on superpixelsabstractIn this paper, we propose a data-driven approach to detect the probabilistic salient object contour, which is formulated as predicting the probability of superpixel boundaries being on the object contour based on the learned regressor. Each superpixel boundary is jointly described by the superpixel saliency, superpixel contrast, and boundary geometry features. Experimental results on the benchmark data set validate the effectiveness of our approach. Furthermore, we demonstrate that the predicted probabilistic salient object contour is useful for improving the multiple segmentations for salient object detection. Huaizu Jiang, Yang Wu 0001, Zejian Yuan |
ICIP | 1 |
| 2011 | Automatic salient object segmentation based on context and shape priorabstractWe propose a novel automatic salient object segmentation algorithm which integrates both bottom-up salient stimuli and object-level shape prior, i.e., a salient object has a well-defined closed boundary. Our approach is formalized as an iterative energy minimization framework, leading to binary segmentation of the salient object. Such energy minimization is initialized with a saliency map which is computed through context analysis based on multi-scale superpixels. Object-level shape prior is then extracted combining saliency with object boundary information. Both saliency map and shape prior update after each iteration. Experimental results on two public benchmark datasets show that our proposed approach outperforms state-of-the-art methods. 1 Huaizu Jiang, Jingdong Wang 0001, Zejian Yuan, Nanning Zheng 0001 |
BMVC | 1 |