Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hou-Ning Hu

dblp:199/3014 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-7564-1473ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Video understanding and tracking · 47% Vision and language · 17% 3D vision · 16%
Computer graphics and multimedia
4 papers
Computational photography and imaging · 47% Virtual and augmented reality · 25% Image and video processing · 24%
Human-computer interaction and pervasive computing
1 paper
Immersive interaction · 50% Interaction techniques and input · 50%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
1.022023
Monocular Quasi-Dense 3D Object Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Joint Monocular 3D Vehicle Detection and Tracking · ICCV 2019
Computer vision › Video understanding and tracking
multi-object tracking
1.022023
Monocular Quasi-Dense 3D Object Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Joint Monocular 3D Vehicle Detection and Tracking · ICCV 2019
Computer vision › Vision and language › cross-modal alignment
image-text alignment
0.812024
Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation · CVPR 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
0.812024
Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation · CVPR 2024
Computer vision › Video understanding and tracking › object tracking
3d object tracking
0.712023
Monocular Quasi-Dense 3D Object Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Video understanding and tracking › object tracking › 3d object tracking
monocular 3d tracking
0.712023
Monocular Quasi-Dense 3D Object Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Video understanding and tracking › multi-object tracking
trajectory association
0.712023
Monocular Quasi-Dense 3D Object Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computational photography and imaging
high dynamic range imaging
0.712023
Learning Continuous Exposure Value Representations for Single-Image HDR Reconstruction · ICCV 2023
Image and video processing
image restoration
0.712023
Learning Continuous Exposure Value Representations for Single-Image HDR Reconstruction · ICCV 2023
Computational photography and imaging › high dynamic range imaging
single-shot HDR
0.712023
Learning Continuous Exposure Value Representations for Single-Image HDR Reconstruction · ICCV 2023
Computer vision › Vision and language
visual grounding
0.312018
Self-View Grounding Given a Narrated 360° Video · AAAI 2018
Virtual and augmented reality › immersive video
360-degree video
0.312018
Self-View Grounding Given a Narrated 360° Video · AAAI 2018
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.312017
Deep 360 Pilot: Learning a Deep Agent for Piloting through 360° Sports Videos · CVPR 2017
Virtual and augmented reality › immersive video
360° video viewing
0.312017
Deep 360 Pilot: Learning a Deep Agent for Piloting through 360° Sports Videos · CVPR 2017
Immersive interaction › virtual reality experience
360° video viewing
0.312017
Tell Me Where to Look: Investigating Ways for Assisting Focus in 360° Video · CHI 2017
Robotics › Autonomous driving
perception
0.212023
Monocular Quasi-Dense 3D Object Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Multimedia analysis and retrieval
video understanding
0.112018
Self-View Grounding Given a Narrated 360° Video · AAAI 2018
Computer vision › Image recognition and object detection
object detection
0.112017
Deep 360 Pilot: Learning a Deep Agent for Piloting through 360° Sports Videos · CVPR 2017
Virtual and augmented reality
immersive experience
0.112017
Tell Me Where to Look: Investigating Ways for Assisting Focus in 360° Video · CHI 2017

Methods — techniques the papers use, named apart from their topics

trajectory prediction · 1.0LSTM · 1.0prompt learning · 0.8contrastive learning · 0.8sentence decoder · 0.7quasi-dense similarity learning · 0.7implicit function · 0.7depth-ordering heuristics · 0.7cycle training · 0.7convolutional neural network · 0.7GRU · 0.7experiment · 0.6depth-ordering matching · 0.4soft attention · 0.3recurrent neural network · 0.3qualitative feedback · 0.3policy gradient · 0.3object detector · 0.3
YearPublicationVenuePosition
2026 HDR Reconstruction Boosting with Training-Free and Exposure-Consistent Diffusion
Yo-Tin Lin, Su-Kai Chen, Hou-Ning Hu, Yen-Yu Lin
WACV3
2025 ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark Detection
abstract
Although facial landmark detection (FLD) has gained significant progress, existing FLD methods still suffer from performance drops on partially non-visible faces, such as faces with occlusions or under extreme lighting conditions or poses. To address this issue, we introduce ORFormer, a novel transformer-based method that can detect non-visible regions and recover their missing features from vis-ible parts. Specifically, ORFormer associates each image patch token with one additional learnable token called the messenger token. The messenger token aggregates features from all but its patch. This way, the consensus between a patch and other patches can be assessed by referring to the similarity between its regular and messenger embeddings, enabling non-visible region identification. Our method then recovers occluded patches with features aggregated by the messenger tokens. Leveraging the recovered features, OR-Former compiles high-quality heatmaps for the downstream FLD task. Extensive experiments show that our method generates heatmaps resilient to partial occlusions. By inte-grating the resultant heatmaps into existing FLD methods, our method performs favorably against the state of the arts on challenging datasets such as WFLWand COFW.
Jui-Che Chiang, Hou-Ning Hu, Bo-Syuan Hou, Chia-Yu Tseng, Yu-Lun Liu 0001, Min-Hung Chen, Yen-Yu Lin
WACV2
2024 Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation
abstract
This paper addresses text-supervised semantic segmentation, aiming to learn a model capable of segmenting arbitrary visual concepts within images by using only image-text pairs without dense annotations. Existing methods have demonstrated that contrastive learning on image-text pairs effectively aligns visual segments with the meanings of texts. We notice that there is a discrepancy between text alignment and semantic segmentation: A text often consists of multiple semantic concepts, whereas semantic segmentation strives to create semantically homogeneous segments. To address this issue, we propose a novel framework, Image-Text Co-Decomposition (CoDe), where the paired image and text are jointly decomposed into a set of image regions and a set of word segments, respectively, and contrastive learning is developed to enforce region-word alignment. To work with a vision-language model, we present a prompt learning mechanism that derives an extra representation to highlight an image segment or a word segment of interest, with which more effective features can be extracted from that segment. Comprehensive experimental results demonstrate that our method performs favorably against existing text-supervised semantic segmentation methods on six benchmark datasets. The code is available at https://github.com/072jiajia/image-text-co-decomposition.
Ji-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang, Chun-Pei Chen, Yu-Lun Liu 0001, Min-Hung Chen, Hou-Ning Hu, Yung-Yu Chuang, Yen-Yu Lin
CVPR7
2023 Learning Continuous Exposure Value Representations for Single-Image HDR Reconstruction
abstract
Deep learning is commonly used to reconstruct HDR images from LDR images. LDR stack-based methods are used for single-image HDR reconstruction, generating an HDR image from a deep learning-generated LDR stack. However, current methods generate the stack with predetermined exposure values (EVs), which may limit the quality of HDR reconstruction. To address this, we propose the continuous exposure value representation (CEVR), which uses an implicit function to generate LDR images with arbitrary EVs, including those unseen during training. Our approach generates a continuous stack with more images containing diverse EVs, significantly improving HDR reconstruction. We use a cycle training strategy to supervise the model in generating continuous EV LDR images without corresponding ground truths. Our CEVR model outperforms existing methods, as demonstrated by experimental results.
Su-Kai Chen, Hung-Lin Yen, Yu-Lun Liu 0001, Min-Hung Chen, Hou-Ning Hu, Wen-Hsiao Peng, Yen-Yu Lin
ICCV5
2023 Monocular Quasi-Dense 3D Object Tracking
abstract
A reliable and accurate 3D tracking framework is essential for predicting future locations of surrounding objects and planning the observer's actions in numerous applications such as autonomous driving. We propose a framework that can effectively associate moving objects over time and estimate their full 3D bounding box information from a sequence of 2D images captured on a moving platform. The object association leverages quasi-dense similarity learning to identify objects in various poses and viewpoints with appearance cues only. After initial 2D association, we further utilize 3D bounding boxes depth-ordering heuristics for robust instance association and motion-based 3D trajectory prediction for re-identification of occluded vehicles. In the end, an LSTM-based object velocity learning module aggregates the long-term trajectory information for more accurate motion extrapolation. Experiments on our proposed simulation data and real-world benchmarks, including KITTI, nuScenes, and Waymo datasets, show that our tracking framework offers robust object association and tracking on urban-driving scenarios. On the Waymo Open benchmark, we establish the first camera-only baseline in the 3D tracking and 3D detection challenges. Our quasi-dense 3D tracking pipeline achieves impressive improvements on the nuScenes 3D tracking benchmark with near five times tracking accuracy of the best vision-only submission among all published methods.
Hou-Ning Hu, Yung-Hsu Yang, Tobias Fischer 0004, Trevor Darrell, Fisher Yu 0001, Min Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Joint Monocular 3D Vehicle Detection and Tracking
abstract
Vehicle 3D extents and trajectories are critical cues for predicting the future location of vehicles and planning future agent ego-motion based on those predictions. In this paper, we propose a novel online framework for 3D vehicle detection and tracking from monocular videos. The framework can not only associate detections of vehicles in motion over time, but also estimate their complete 3D bounding box information from a sequence of 2D images captured on a moving platform. Our method leverages 3D box depth-ordering matching for robust instance association and utilizes 3D trajectory prediction for re-identification of occluded vehicles. We also design a motion learning module based on an LSTM for more accurate long-term motion extrapolation. Our experiments on simulation, KITTI, and Argoverse datasets show that our 3D tracking pipeline offers robust data association and tracking. On Argoverse, our image-based method is significantly better for tracking 3D vehicles within 30 meters than the LiDAR-centric baseline methods.
Hou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin 0002, Min Sun 0001, Philipp Krähenbühl, Trevor Darrell, Fisher Yu 0001
ICCV1
2019 3D LiDAR and Stereo Fusion using Stereo Matching Network with Conditional Cost Volume Normalization
abstract
The complementary characteristics of active and passive depth sensing techniques motivate the fusion of the LiDAR sensor and stereo camera for improved depth perception. Instead of directly fusing estimated depths across LiDAR and stereo modalities, we take advantages of the stereo matching network with two enhanced techniques: Input Fusion and Conditional Cost Volume Normalization (CCVNorm) on the LiDAR information. The proposed framework is generic and closely integrated with the cost volume component that is commonly utilized in stereo matching neural networks. We experimentally verify the efficacy and robustness of our method on the KITTI Stereo and Depth Completion datasets, obtaining favorable performance against various fusion strategies. Moreover, we demonstrate that, with a hierarchical extension of CCVNorm, the proposed method brings only slight overhead to the stereo matching network in terms of computation time and model size.
Tsun-Hsuan Wang, Hou-Ning Hu, Chieh Hubert Lin, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001
IROS2
2018 Self-View Grounding Given a Narrated 360° Video
abstract
Narrated 360° videos are typically provided in many touring scenarios to mimic real-world experience. However, previous work has shown that smart assistance (i.e., providing visual guidance) can significantly help users to follow the Normal Field of View (NFoV) corresponding to the narrative.In this project, we aim at automatically grounding the NFoVs of a 360° video given subtitles of the narrative (referred to as ''NFoV-grounding"). We propose a novel Visual Grounding Model (VGM) to implicitly and efficiently predict the NFoVs given the video content and subtitles. Specifically, at each frame, we efficiently encode the panorama into feature map of candidate NFoVs using a Convolutional Neural Network (CNN) and the subtitles to the same hidden space using an RNN with Gated Recurrent Units (GRU). Then, we apply soft-attention on candidate NFoVs to trigger sentence decoder aiming to minimize the reconstruct loss between the generated and given sentence. Finally, we obtain the NFoV as the candidate NFoV with the maximum attention without any human supervision.To train VGM more robustly, we also generate a reverse sentence conditioning on one minus the soft-attention such that the attention focuses on candidate NFoVs less relevant to the given sentence. The negative log reconstruction loss of the reverse sentence (referred to as ''irrelevant loss") is jointly minimized to encourage the reverse sentence to be different from the given sentence. To evaluate our method, we collect the first narrated 360° videos dataset and achieve state-of-the-art NFoV-grounding performance.
Shih-Han Chou, Kuo-Hao Zeng, Hou-Ning Hu, Jianlong Fu, Min Sun 0001
AAAI4
2018 Self-supervised Learning of Depth and Camera Motion from 360 ^\circ Videos
Fu-En Wang, Hou-Ning Hu, Hsien-Tzu Cheng, Juan-Ting Lin, Shang-Ta Yang, Meng-Li Shih, Hung-Kuo Chu, Min Sun 0001
ACCV (5)2
2017 Tell Me Where to Look: Investigating Ways for Assisting Focus in 360° Video
abstract
360° videos give viewers a spherical view and immersive experience of surroundings. However, one challenge of watching 360° videos is continuously focusing and re-focusing intended targets. To address this challenge, we developed two Focus Assistance techniques: Auto Pilot (directly bringing viewers to the target), and Visual Guidance (indicating the direction of the target). We conducted an experiment to measure viewers' video-watching experience and discomfort using these techniques and obtained their qualitative feedback. We showed that: 1) Focus Assistance improved ease of focus. 2) Focus Assistance techniques have specificity to video content. 3) Participants' preference of and experience with Focus Assistance depended not only on individual difference but also on their goal of watching the video. 4) Factors such as view-moving-distance, salience of the intended target and guidance, and language comprehension affected participants' video-watching experience. Based on these findings, we provide design implications for better 360° video focus assistance.
Yen-Chen Lin, Yung-Ju Chang, Hou-Ning Hu, Hsien-Tzu Cheng, Chi-Wen Huang, Min Sun 0001
CHI3
2017 Deep 360 Pilot: Learning a Deep Agent for Piloting through 360° Sports Videos
abstract
Watching a 360° sports video requires a viewer to continuously select a viewing angle, either through a sequence of mouse clicks or head movements. To relieve the viewer from this “360 piloting” task, we propose “deep 360 pilot” - a deep learning-based agent for piloting through 360° sports videos automatically. At each frame, the agent observes a panoramic image and has the knowledge of previously selected viewing angles. The task of the agent is to shift the current viewing angle (i.e. action) to the next preferred one (i.e., goal). We propose to directly learn an online policy of the agent from data. Specifically, we leverage a state-of-the-art object detector to propose a few candidate objects of interest (yellow boxes in Fig. 1). Then, a recurrent neural network is used to select the main object (green dash boxes in Fig. 1). Given the main object and previously selected viewing angles, our method regresses a shift in viewing angle to move to the next one. We use the policy gradient technique to jointly train our pipeline, by minimizing: (1) a regression loss measuring the distance between the selected and ground truth viewing angles, (2) a smoothness loss encouraging smooth transition in viewing angle, and (3) maximizing an expected reward offocusing on a foreground object. To evaluate our method, we built a new 360-Sports video dataset consisting offive sports domains. We trained domain-specific agents and achieved the best performance on viewing angle selection accuracy and users' preference compared to [53] and other baselines.
Hou-Ning Hu, Yen-Chen Lin, Ming-Yu Liu 0001, Hsien-Tzu Cheng, Yung-Ju Chang, Min Sun 0001
CVPR1