Wei-Chih Hung

dblp:70/2879 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
10since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 2
YearPublicationVenuePosition
2024 3D Open-Vocabulary Panoptic Segmentation with 2D-3D Vision-Language Distillation
Zihao Xiao 0001, Longlong Jing, Shangxuan Wu, Alex Zihao Zhu, Jingwei Ji, Chiyu Max Jiang, Wei-Chih Hung, Thomas A. Funkhouser, Weicheng Kuo, Anelia Angelova, Shiwei Sheng
ECCV (40)7
2024 LET-3D-AP: Longitudinal Error Tolerant 3D Average Precision for Camera-Only 3D Detection
abstract
The 3D Average Precision (3DAP) relies on the intersection over union between predictions and ground truth objects. However, camera-only detectors have limited depth accuracy, which may cause otherwise reasonable predictions that suffer from such longitudinal localization errors to be treated as false positives. We therefore propose variants of the 3DAP metric to be more permissive with respect to depth estimation errors. Specifically, our novel longitudinal error tolerant metrics, LET-3D-AP and LET-3D-APL, allow longitudinal localization errors of the prediction boxes up to a given tolerance. To evaluate the proposed metrics, we also construct a new test set for the Waymo Open Dataset, tailored to camera-only 3D detection methods. Surprisingly, we find that state-of-the-art camera-based detectors can outperform popular LiDAR-based detectors with our new metrics past at 10% depth error tolerance, suggesting that existing camera-based detectors already have the potential to surpass LiDAR-based detectors in downstream applications. We believe the proposed metrics and the new benchmark dataset will facilitate advances in the field of camera-only 3D detection by providing more informative signals that can better indicate the system-level performance.
Wei-Chih Hung, Vincent Casser, Henrik Kretzschmar, Jyh-Jing Hwang, Dragomir Anguelov
ICRA1
2024 STT: Stateful Tracking with Transformers for Autonomous Driving
abstract
Tracking objects in three-dimensional space is critical for autonomous driving. To ensure safety while driving, the tracker must be able to reliably track objects across frames and accurately estimate their states such as velocity and acceleration in the present. Existing works frequently focus on the association task while either neglecting the model’s performance on state estimation or deploying complex heuristics to predict the states. In this paper, we propose STT, a Stateful Tracking model built with Transformers, that can consistently track objects in the scenes while also predicting their states accurately. STT consumes rich appearance, geometry, and motion signals through long term history of detections and is jointly optimized for both data association and state estimation tasks. Since the standard tracking metrics like MOTA and MOTP do not capture the combined performance of the two tasks in the wider spectrum of object states, we extend them with new metrics called S-MOTA and MOTPSthat address this limitation. STT achieves competitive real-time performance on the Waymo Open Dataset.
Longlong Jing, Ruichi Yu, Zhengli Zhao, Shiwei Sheng, Colin Graber, Qinru Li, Shangxuan Wu, Chris Sweeney, Wei-Chih Hung, Xingyi Zhou, Farshid Moussavi, James Guo, Mingxing Tan, Weilong Yang
ICRA14
2023 Hi-LASSIE: High-Fidelity Articulated Shape and Skeleton Discovery from Sparse Image Ensemble
abstract
Automatically estimating 3D skeleton, shape, camera viewpoints, and part articulation from sparse in-the-wild image ensembles is a severely under-constrained and challenging problem. Most prior methods rely on large-scale image datasets, dense temporal correspondence, or human annotations like camera pose, 2D keypoints, and shape templates. We propose Hi-LASSIE, which performs 3D articulated reconstruction from only 20–30 online images in the wild without any user-defined shape or skeleton templates. We follow the recent work of LASSIE that tackles a similar problem setting and make two significant advances. First, instead of relying on a manually annotated 3D skeleton, we automatically estimate a class-specific skeleton from the selected reference image. Second, we improve the shape reconstructions with novel instance-specific optimization strategies that allow reconstructions to faithful fit on each instance while preserving the class-specific priors learned across all images. Experiments on in-the-wild image ensembles show that Hi-LASSIE obtains higher fidelity state-of-the-art 3D reconstructions despite requiring minimum user input. Project page: chhankyao.github.io/hi-lassie/
Chun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein, Ming-Hsuan Yang 0001, Varun Jampani
CVPR2
2023 ARTIC3D: Learning Robust Articulated 3D Shapes from Noisy Web Image Collections
abstract
Estimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture, lighting, etc. We propose ARTIC3D, a self-supervised framework to reconstruct per-instance 3D shapes from a sparse image collection in-the-wild. Specifically, ARTIC3D is built upon a skeleton-based surface representation and is further guided by 2D diffusion priors from Stable Diffusion. First, we enhance the input images with occlusions/truncation via 2D diffusion to obtain cleaner mask estimates and semantic features. Second, we perform diffusion-guided 3D optimization to estimate shape and texture that are of high-fidelity and faithful to input images. We also propose a novel technique to calculate more stable image-level gradients via diffusion models compared to existing alternatives. Finally, we produce realistic animations by fine-tuning the rendered shape and texture under rigid part transformations. Extensive evaluations on multiple existing datasets as well as newly introduced noisy web image collections with occlusions and truncation demonstrate that ARTIC3D outputs are more robust to noisy images, higher quality in terms of shape and texture details, and more realistic when animated.
Chun-Han Yao, Amit Raj, Wei-Chih Hung, Michael Rubinstein, Yuanzhen Li, Ming-Hsuan Yang 0001, Varun Jampani
NeurIPS3
2022 Incremental False Negative Detection for Contrastive Learning
Tsai-Shien Chen, Wei-Chih Hung, Hung-Yu Tseng, Shao-Yi Chien, Ming-Hsuan Yang 0001
ICLR2
2022 LASSIE: Learning Articulated Shapes from Sparse Image Ensemble via 3D Part Discovery
abstract
Creating high-quality articulated 3D models of animals is challenging either via manual creation or using 3D scanning tools. Therefore, techniques to reconstruct articulated 3D objects from 2D images are crucial and highly useful. In this work, we propose a practical problem setting to estimate 3D pose and shape of animals given only a few (10-30) in-the-wild images of a particular animal species (say, horse). Contrary to existing works that rely on pre-defined template shapes, we do not assume any form of 2D or 3D ground-truth annotations, nor do we leverage any multi-view or temporal information. Moreover, each input image ensemble can contain animal instances with varying poses, backgrounds, illuminations, and textures. Our key insight is that 3D parts have much simpler shape compared to the overall animal and that they are robust w.r.t. animal pose articulations. Following these insights, we propose LASSIE, a novel optimization framework which discovers 3D parts in a self-supervised manner with minimal user intervention. A key driving force behind LASSIE is the enforcing of 2D-3D part consistency using self-supervisory deep features. Experiments on Pascal-Part and self-collected in-the-wild animal datasets demonstrate considerably better 3D reconstructions as well as both 2D and 3D part discovery compared to prior arts. Project page: https://chhankyao.github.io/lassie/
Chun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein, Ming-Hsuan Yang 0001, Varun Jampani
NeurIPS2
2022 Semi-supervised Multi-task Learning for Semantics and Depth
abstract
Multi-Task Learning (MTL) aims to enhance the model generalization by sharing representations between related tasks for better performance. Typical MTL methods are jointly trained with the complete multitude of ground-truths for all tasks simultaneously. However, one single dataset may not contain the annotations for each task of interest. To address this issue, we propose the Semi-supervised Multi-Task Learning (SemiMTL) method to leverage the available supervisory signals from different datasets, particularly for semantic segmentation and depth estimation tasks. To this end, we design an adversarial learning scheme in our semi-supervised training by leveraging unlabeled data to optimize all the task branches simultaneously and accomplish all tasks across datasets with partial annotations. We further present a domain-aware discriminator structure with various alignment formulations to mitigate the domain discrepancy issue among datasets. Finally, we demonstrate the effectiveness of the proposed method to learn across different datasets on challenging street view and remote sensing benchmarks.
Yufeng Wang 0004, Yi-Hsuan Tsai, Wei-Chih Hung, Wenrui Ding, Ming-Hsuan Yang 0001
WACV3
2021 Discovering 3D Parts from Image Collections
abstract
Reasoning 3D shapes from 2D images is an essential yet challenging task, especially when only single-view images are at our disposal. While an object can have a complicated shape, individual parts are usually close to geometric primitives and thus are easier to model. Furthermore, parts provide a mid-level representation that is robust to appearance variations across objects in a particular category. In this work, we tackle the problem of 3D part discovery from only 2D image collections. Instead of relying on manually annotated parts for supervision, we propose a self-supervised approach, latent part discovery (LPD). Our key insight is to learn a novel part shape prior that allows each part to fit an object shape faithfully while constrained to have simple geometry. Extensive experiments on the synthetic ShapeNet, PartNet, and real-world Pascal 3D+ datasets show that our method discovers consistent object parts and achieves favorable reconstruction accuracy compared to the existing methods with the same level of supervision. Our project page with code is at https://chhankyao.github.io/lpd/.
Chun-Han Yao, Wei-Chih Hung, Varun Jampani, Ming-Hsuan Yang 0001
ICCV2
2021 Learning to Caricature via Semantic Shape Transform
abstract
Abstract Caricature is an artistic drawing created to abstract or exaggerate facial features of a person. Rendering visually pleasing caricatures is a difficult task that requires professional skills, and thus it is of great interest to design a method to automatically generate such drawings. To deal with large shape changes, we propose an algorithm based on a semantic shape transform to produce diverse and plausible shape exaggerations. Specifically, we predict pixel-wise semantic correspondences and perform image warping on the input photo to achieve dense shape transformation. We show that the proposed framework is able to render visually pleasing shape exaggerations while maintaining their facial structures. In addition, our model allows users to manipulate the shape via the semantic map. We demonstrate the effectiveness of our approach on a large photograph-caricature benchmark dataset with comparisons to the state-of-the-art methods.
Wenqing Chu, Wei-Chih Hung, Yi-Hsuan Tsai, Yu-Ting Chang, Yijun Li 0001, Deng Cai 0001, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.2
2020 Mixup-CAM: Weakly-supervised Semantic Segmentation via Uncertainty Regularization
Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001
BMVC3
2020 Weakly-Supervised Semantic Segmentation via Sub-Category Exploration
abstract
Existing weakly-supervised semantic segmentation methods using image-level annotations typically rely on initial responses to locate object regions. However, such response maps generated by the classification network usually focus on discriminative object parts, due to the fact that the network does not need the entire object for optimizing the objective function. To enforce the network to pay attention to other parts of an object, we propose a simple yet effective approach that introduces a self-supervised task by exploiting the sub-category information. Specifically, we perform clustering on image features to generate pseudo sub-categories labels within each annotated parent class, and construct a sub-category objective to assign the network to a more challenging task. By iteratively clustering image features, the training process does not limit itself to the most discriminative object parts, hence improving the quality of the response maps. We conduct extensive analysis to validate the proposed method and show that our approach performs favorably against the state-of-the-art approaches.
Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001
CVPR3
2020 From Image Collections to Point Clouds With Self-Supervised Shape and Pose Networks
abstract
Reconstructing 3D models from 2D images is one of the fundamental problems in computer vision. In this work, we propose a deep learning technique for 3D object reconstruction from a single image. Contrary to recent works that either use 3D supervision or multi-view supervision, we use only single view images with no pose information during training as well. This makes our approach more practical requiring only an image collection of an object category and the corresponding silhouettes. We learn both 3D point cloud reconstruction and pose estimation networks in a self-supervised manner, making use of differentiable point cloud renderer to train with 2D supervision. A key novelty of the proposed technique is to impose 3D geometric reasoning into predicted 3D point clouds by rotating them with randomly sampled poses and then enforcing cycle consistency on both 3D reconstructions and poses. In addition, using single-view supervision allows us to do test-time optimization on a given test image. Experiments on the synthetic ShapeNet and real-world Pix3D datasets demonstrate that our approach, despite using less supervision, can achieve competitive performance compared to pose-supervised and multi-view supervised approaches.
Navaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung, Varun Jampani, Venkatesh Babu Radhakrishnan
CVPR4
2020 Image Hashing via Linear Discriminant Learning
abstract
Hashing has attracted attention in recent years due to the rapid growth of image and video data on the web. Benefiting from recent advances in deep learning, deep supervised hashing has achieved promising results for image retrieval. However, existing methods are either less efficient in data usage or incapable of learning linearly discriminative binary codes. In this paper, we revisit linear discriminative analysis and propose a linear discriminative hashing (LDH) objective that is efficient in training and achieves better accuracy in retrieval. With the joint supervision of a classification loss, we design a robust deep network to obtain binary codes that are inter-class separable and intra-class compact, which provides better representations for image retrieval. We conduct extensive experiments on three benchmark datasets, and our LDH algorithm performs favorably against existing state-of-the-art deep supervised hashing methods.
Weixiang Hong 0001, Yu-Ting Chang, Haifang Qin, Wei-Chih Hung, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001
WACV4
2020 Progressive Domain Adaptation for Object Detection
abstract
Recent deep learning methods for object detection rely on a large amount of bounding box annotations. Collecting these annotations is laborious and costly, yet supervised models do not generalize well when testing on images from a different distribution. Domain adaptation provides a solution by adapting existing labels to the target testing data. However, a large gap between domains could make adaptation a challenging task, which leads to unstable training processes and sub-optimal results. In this paper, we propose to bridge the domain gap with an intermediate domain and progressively solve easier adaptation subtasks. This intermediate domain is constructed by translating the source images to mimic the ones in the target domain. To tackle the domain-shift problem, we adopt adversarial learning to align distributions at the feature level. In addition, a weighted task loss is applied to deal with unbalanced image quality in the intermediate domain. Experimental results show that our method performs favorably against the state-of-the-art method in terms of the performance on the target domain.
Han-Kai Hsu, Chun-Han Yao, Yi-Hsuan Tsai, Wei-Chih Hung, Hung-Yu Tseng, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
WACV4
2019 A Top-Down Unified Framework for Instance-level Human Parsing
Haifang Qin, Weixiang Hong 0001, Wei-Chih Hung, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001
BMVC3
2019 SCOPS: Self-Supervised Co-Part Segmentation
abstract
Parts provide a good intermediate representation of objects that is robust with respect to camera, pose and appearance variations. Existing work on part segmentation is dominated by supervised approaches that rely on large amounts of manual annotations and also can not generalize to unseen object categories. We propose a self-supervised deep learning approach for part segmentation, where we devise several loss functions that aids in predicting part segments that are geometrically concentrated, robust to object variations and are also semantically consistent across different object instances. Extensive experiments on different types of image collections demonstrate that our approach can produce part segments that adhere to object boundaries and also more semantically consistent across object instances compared to existing self-supervised techniques.
Wei-Chih Hung, Varun Jampani, Sifei Liu, Pavlo Molchanov 0001, Ming-Hsuan Yang 0001, Jan Kautz
CVPR1
2019 Weakly-Supervised Caricature Face Parsing Through Domain Adaptation
abstract
A caricature is an artistic form of a person's picture in which certain striking characteristics are abstracted or exaggerated in order to create a humor or sarcasm effect. For numerous caricature related applications such as attribute recognition and caricature editing, face parsing is an essential pre-processing step that provides a complete facial structure understanding. However, current state-of-the-art face parsing methods require large amounts of labeled data on the pixel-level and such process for caricature is tedious and labor-intensive. For real photos, there are numerous labeled datasets for face parsing. Thus, we formulate caricature face parsing as a domain adaptation problem, where real photos play the role of the source domain, adapting to the target caricatures. Specifically, we first leverage a spatial transformer based network to enable shape domain shifts. A feed-forward style transfer network is then utilized to capture texture-level domain gaps. With these two steps, we synthesize face caricatures from real photos, and thus we can use parsing ground truths of the original photos to learn the parsing model. Experimental results on the synthetic and real caricatures demonstrate the effectiveness of the proposed domain adaptation algorithm. Code is available at: https://github.com/ZJULearning/CariFaceParsing.
Wenqing Chu, Wei-Chih Hung, Yi-Hsuan Tsai, Deng Cai 0001, Ming-Hsuan Yang 0001
ICIP2
2018 Adversarial Learning for Semi-supervised Semantic Segmentation
Wei-Chih Hung, Yi-Hsuan Tsai, Yan-Ting Liou, Yen-Yu Lin, Ming-Hsuan Yang 0001
BMVC1
2018 Fast and Accurate Online Video Object Segmentation via Tracking Parts
abstract
Online video object segmentation is a challenging task as it entails to process the image sequence timely and accurately. To segment a target object through the video, numerous CNN-based methods have been developed by heavily finetuning on the object mask in the first frame, which is time-consuming for online applications. In this paper, we propose a fast and accurate video object segmentation algorithm that can immediately start the segmentation process once receiving the images. We first utilize a part-based tracking method to deal with challenging factors such as large deformation, occlusion, and cluttered background. Based on the tracked bounding boxes of parts, we construct a region-of-interest segmentation network to generate part masks. Finally, a similarity-based scoring function is adopted to refine these object parts by comparing them to the visual information in the first frame. Our method performs favorably against state-of-the-art algorithms in accuracy on the DAVIS benchmark dataset, while achieving much faster runtime performance.
Jingchun Cheng, Yi-Hsuan Tsai, Wei-Chih Hung, Shengjin Wang, Ming-Hsuan Yang 0001
CVPR3
2018 Learning to Adapt Structured Output Space for Semantic Segmentation
abstract
Convolutional neural network-based approaches for semantic segmentation rely on supervision with pixel-level ground truth, but may not generalize well to unseen image domains. As the labeling process is tedious and labor intensive, developing algorithms that can adapt source ground truth labels to the target domain is of great interest. In this paper, we propose an adversarial learning method for domain adaptation in the context of semantic segmentation. Considering semantic segmentations as structured outputs that contain spatial similarities between the source and target domains, we adopt adversarial learning in the output space. To further enhance the adapted model, we construct a multi-level adversarial network to effectively perform output space domain adaptation at different feature levels. Extensive experiments and ablation study are conducted under various domain adaptation settings, including synthetic-to-real and cross-city scenarios. We show that the proposed method performs favorably against the state-of-the-art methods in terms of accuracy and visual quality.
Yi-Hsuan Tsai, Wei-Chih Hung, Samuel Schulter, Kihyuk Sohn, Ming-Hsuan Yang 0001, Manmohan Krishna Chandraker
CVPR2
2018 Learning to Blend Photos
Wei-Chih Hung, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Joon-Young Lee, Ming-Hsuan Yang 0001
ECCV (7)1
2017 Scene Parsing with Global Context Embedding
abstract
We present a scene parsing method that utilizes global context information based on both the parametric and nonparametric models. Compared to previous methods that only exploit the local relationship between objects, we train a context network based on scene similarities to generate feature representations for global contexts. In addition, these learned features are utilized to generate global and spatial priors for explicit classes inference. We then design modules to embed the feature representations and the priors into the segmentation network as additional global context cues. We show that the proposed method can eliminate false positives that are not compatible with the global context representations. Experiments on both the MIT ADE20K and PASCAL Context datasets show that the proposed method performs favorably against existing methods.
Wei-Chih Hung, Yi-Hsuan Tsai, Xiaohui Shen, Zhe Lin 0001, Kalyan Sunkavalli, Xin Lu 0006, Ming-Hsuan Yang 0001
ICCV1
2016 Unsupervised Visual Representation Learning by Graph-Based Consistent Constraints
Dong Li 0025, Wei-Chih Hung, Jia-Bin Huang 0001, Shengjin Wang, Narendra Ahuja, Ming-Hsuan Yang 0001
ECCV (4)2
2014 Gaze direction estimation using support vector machine with active appearance model
Yi-Leh Wu, Chun-Tsai Yeh, Wei-Chih Hung, Cheng-Yuan Tang
Multim. Tools Appl.3
2013 Iterative decoding for uncompressed wireless video transmission
abstract
Due to the emergence of mmWave systems that provide multi-Gbps data rate over 60GHz band, uncompressed video transmission has been considered as a commonly used feature for wireless multimedia transmission over the wireless personal area networks (WPANs). However, mmWaves signals usually have higher attenuation than the conventional low-frequency wireless signals, and therefore supporting transmission in a low SNR condition becomes a challenging problem for such systems. In this paper, we propose a iterative decoding method for the wireless transmission of uncompressed video. With our proposed iterative decoding structure, the video decoder utilizes the temporal redundancy among the video frames to jointly decode the received coded bits from different video frames with the aid of correlation noise model. The simulation results show that the proposed method can significantly enhanced the video quality in terms of PSNR, even while operating in an extremely low SNR condition.
Wei-Chih Hung, Ping-Cheng Yeh
GLOBECOM2
2012 Dynamic source-channel rate-distortion control under time-varying complexity constraint for wireless video transmission
abstract
In recent years, a number of rate-distortion (R-D) optimization schemes have been proposed. These works discussed the joint source-channel rate allocation problem for achieving best end-to-end video transmission quality under limited transmission bit rate. However, encoding complexity, a critical concern for video transmission from mobile devices, was hardly considered. Especially, the time-varying nature of computational resources for mobile devices was largely ignored. In this paper, an on-line adaptive video encoder parameters control scheme is proposed. Simulation results show that the proposed scheme, by dynamically adjusting the encoder parameters in H.264 using the precomputed rate-complexity-distortion (R-C-D) tables, achieves low end-to-end distortion under time-varying channel condition and dynamic computational complexity constraints.
Tsu-Hao Kuo, Po-Hsuan Chen, Wei-Chih Hung, Chih-Yu Huang, Chia-han Lee, Ping-Cheng Yeh
WCNC3
2010 Point-of-Regard Measurement via Iris Contour with One Eye from Single Image
abstract
Eye gaze tracking is a technique which is commonly used in human-computer interaction. By determining the eye gaze, the point-of-regard can be estimated by intersecting the gaze line and the target plane. A particular assumption is that irises are regarded as ellipses rather than circles in our approach. Besides, only one eye is required in the processes of the approach. The unique eye gaze can be computed via some image processing techniques such as edge detection and ellipse fitting and exploiting geometric properties of circles and ellipses in the 3D space. One of the key phases in our approach is the ellipse fitting process due to the noise of the image. Conventional algorithms of ellipse fitting are sensitive to noise. In this paper we have improved the randomized Hough transform to reduce the effect of noise. It is also confirmed to be robust by experimenting with real images. The estimation results of point-of-regard are illustrated and four distinct clusters can be separated robustly with both the influences of resolution and eyelids.
Shang-Che Huang, Yi-Leh Wu, Wei-Chih Hung, Cheng-Yuan Tang
ISM3