VLDB 2026 Research / reviewers in the wild / expert
Tzu-Yi Hung
dblp:05/6184
· DBLP profile ↗
22ranked-venue papers
8as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Dynamic Weight Adjustment for Spatial-Temporal Trajectory Planning in Crowd NavigationabstractRobot navigation in dense human crowds poses a significant challenge due to the complexity of human behavior in dynamic and obstacle-rich environments. In this work, we propose a dynamic weight adjustment scheme using a neural network to predict the optimal weights of objectives in an optimization-based motion planner. We adopt a spatial-temporal trajectory planner and incorporate diverse objectives to achieve a balance among safety, efficiency, and goal achievement in complex and dynamic environments. We design the network structure, observation encoding, and reward function to effectively train the policy network using reinforcement learning, allowing the robot to adapt its behavior in real time based on environmental and pedestrian information. Simulation results show improved safety compared to the fixed-weight planner and the state-of-the-art learning-based methods, and verify the ability of the learned policy to adaptively adjust the weights based on the observed situations. The feasibility of the approach is demonstrated in a navigation task using an autonomous delivery robot across a crowded corridor over a 300 m distance. Video: https://youtu.be/nSCbNaaF_VM Muqing Cao, Xinhang Xu, Yizhuo Yang 0001, Jianping Li 0004, Tongxing Jin, Tzu-Yi Hung, Guosheng Lin, Lihua Xie 0001 |
ICRA | 7 |
| 2025 | Graph Optimality-Aware Stochastic LiDAR Bundle Adjustment With Progressive Spatial SmoothingabstractLarge-scale LiDAR Bundle Adjustment (LBA) to refine sensor orientation and point cloud accuracy simultaneously for building navigation maps is a fundamental task in logistics, intelligent transportation, and robotics. In the context of autonomous delivery and smart mobility, the 3D map obtained by accurate and robust LBA plays a pivotal role in enabling reliable localization and navigation across complex, large-scale urban environments. Unlike pose-graph-based methods that rely solely on pairwise relationships between LiDAR frames, LBA leverages raw LiDAR correspondences to achieve more precise results, especially when initial pose estimates are unreliable for low-cost sensors. However, existing LBA methods face challenges such as simplistic planar correspondences, extensive observations, and dense normal matrices in the least-squares problem, which limit robustness, efficiency, and scalability. To address these issues, we propose a Graph Optimality-aware Stochastic Optimization scheme with Progressive Spatial Smoothing, namely PSS-GOSO, to achieverobust,efficient, andscalableLBA. The Progressive Spatial Smoothing (PSS) module extractsrobustLiDAR feature association exploiting the prior structure information obtained by the polynomial smooth kernel. The Graph Optimality-aware Stochastic Optimization (GOSO) module first sparsifies the graph according to optimality for anefficientoptimization. GOSO then utilizes stochastic clustering and graph marginalization to solve the large-scale state estimation problem for ascalableLBA. We validate PSS-GOSO across diverse scenes captured by various platforms, demonstrating its superior performance compared to existing methods. Moreover, the resulting point cloud maps are used for automatic last-mile delivery in large-scale complex scenes, showcasing the practical benefits of our method in modern intelligent transportation systems. The project page can be found at:https://kafeiyin00.github.io/PSS-GOSO/ Jianping Li 0004, Thien-Minh Nguyen, Muqing Cao, Shenghai Yuan 0001, Tzu-Yi Hung, Lihua Xie 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Dense Supervision Propagation for Weakly Supervised Semantic Segmentation on 3D Point CloudsabstractSemantic segmentation on 3D point clouds is an important task for 3D scene understanding. While dense labeling on 3D data is expensive and time-consuming, only a few works address weakly supervised semantic point cloud segmentation methods to relieve the labeling cost by learning from simpler and cheaper labels. Meanwhile, there are still huge performance gaps between existing weakly supervised methods and state-of-the-art fully supervised methods. In this paper, we propose Dense Supervision Propagation (DSP) to train a semantic point cloud segmentation network with only a small portion of points being labeled. We argue that we can better utilize the limited supervision information as we densely propagate the supervision signal from the labeled points to other points within and across the input samples. Specifically, we propose a cross-sample feature reallocating module to transfer similar features and therefore re-route the gradients across two samples with common classes and an intra-sample feature redistribution module to propagate supervision signals on unlabeled points across and within point cloud samples. We conduct extensive experiments on public datasets S3DIS and ScanNet. Our weakly supervised method with only 10% and 1% of labels can produce competitive results with the fully supervised counterpart. Jiacheng Wei, Guosheng Lin, Kim-Hui Yap, Fayao Liu, Tzu-Yi Hung |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Cross-Image Region Mining With Region Prototypical Network for Weakly Supervised SegmentationabstractWeakly supervised image segmentation trained with image-level labels usually suffers from inaccurate coverage of object areas during the generation of the pseudo groundtruth. This is because the object activation maps are trained with the classification objective and lack the ability to generalize. To improve the generality of the object activation maps, we propose a region prototypical network (RPNet) to explore the cross-image object diversity of the training set. Similar object parts across images are identified via region feature comparison. Object confidence is propagated between regions to discover new object areas while background regions are suppressed. Experiments show that the proposed method generates more complete and accurate pseudo object masks while achieving state-of-the-art performance on PASCAL VOC 2012 and MS COCO. In addition, we investigate the robustness of the proposed method on reduced training sets. The code is available athttps://github.com/liuweide01/RPNet-Weakly-Supervised-Segmentation. Weide Liu, Xiangfei Kong, Tzu-Yi Hung, Guosheng Lin |
IEEE Trans. Multim. | 3 |
| 2023 | Few-Shot Segmentation With Optimal Transport Matching and Message FlowabstractWe tackle the challenging task of few-shot segmentation in this work. It is essential for few-shot semantic segmentation to fully utilize the support information. Previous methods typically adopt masked average pooling over the support feature to extract the support clues as a global vector, usually dominated by the salient part and lost certain essential clues. In this work, we argue that every support pixel’s information is desired to be transferred to all query pixels and propose a Correspondence Matching Network (CMNet) with an Optimal Transport Matching module to mine out the correspondence between the query and support images. Besides, it is critical to fully utilize both local and global information from the annotated support images. To this end, we propose a Message Flow module to propagate the message along the inner-flow inside the same image and cross-flow between support and query images, which greatly helps enhance the local feature representations. Experiments on PASCAL VOC 2012, MS COCO, and FSS-1000 datasets show that our network achieves new state-of-the-art few-shot segmentation performance. Weide Liu, Chi Zhang 0007, Henghui Ding, Tzu-Yi Hung, Guosheng Lin |
IEEE Trans. Multim. | 4 |
| 2022 | Learning Regional Purity for Instance Segmentation on 3D Point Clouds
Guosheng Lin, Tzu-Yi Hung |
ECCV (30) | 3 |
| 2021 | Defect-GAN: High-Fidelity Defect Synthesis for Automated Defect InspectionabstractAutomated defect inspection is critical for effective and efficient maintenance, repair, and operations in advanced manufacturing. On the other hand, automated defect inspection is often constrained by the lack of defect samples, especially when we adopt deep neural networks for this task. This paper presents Defect-GAN, an automated defect synthesis network that generates realistic and diverse defect samples for training accurate and robust defect inspection networks. Defect-GAN learns through defacement and restoration processes, where the defacement generates defects on normal surface images while the restoration removes defects to generate normal images. It employs a novel compositional layer-based architecture for generating realistic defects within various image backgrounds with different textures and appearances. It can also mimic the stochastic variations of defects and offer flexible control over the locations and categories of the generated defects within the image background. Extensive experiments show that Defect-GAN is capable of synthesizing various defects with superior diversity and fidelity. In addition, the synthesized defect samples demonstrate their effectiveness in training better defect inspection networks. Gongjie Zhang, Kaiwen Cui, Tzu-Yi Hung, Shijian Lu |
WACV | 3 |
| 2020 | SpSequenceNet: Semantic Segmentation Network on 4D Point CloudsabstractPoint clouds are useful in many applications like autonomous driving and robotics as they provide natural 3D information of the surrounding environments. While there are extensive research on 3D point clouds, scene understanding on 4D point clouds, a series of consecutive 3D point clouds frames, is an emerging topic and yet under-investigated. With 4D point clouds (3D point cloud videos), robotic systems could enhance their robustness by leveraging the temporal information from previous frames. However, the existing semantic segmentation methods on 4D point clouds suffer from low precision due to the spatial and temporal information loss in their network structures. In this paper, we propose SpSequenceNet to address this problem. The network is designed based on 3D sparse convolution. And we introduce two novel modules, a cross-frame global attention module and a cross-frame local interpolation module, to capture spatial and temporal information in 4D point clouds. We conduct extensive experiments on SemanticKITTI, and achieve the state-of-the-art result of 43.1% on mIoU, which is 1.5% higher than the previous best approach. Hanyu Shi 0002, Guosheng Lin, Hao Wang 0094, Tzu-Yi Hung, Zhenhua Wang 0003 |
CVPR | 4 |
| 2020 | Multi-Path Region Mining for Weakly Supervised 3D Semantic Segmentation on Point CloudsabstractPoint clouds provide intrinsic geometric information and surface context for scene understanding. Existing methods for point cloud segmentation require a large amount of fully labeled data. Using advanced depth sensors, collection of large scale 3D dataset is no longer a cumbersome process. However, manually producing point-level label on the large scale dataset is time and labor-intensive. In this paper, we propose a weakly supervised approach to predict point-level results using weak labels on 3D point clouds. We introduce our multi-path region mining module to generate pseudo point-level labels from a classification network trained with weak labels. It mines the localization cues for each class from various aspects of the network feature using different attention modules. Then, we use the point-level pseudo label to train a point cloud segmentation network in a fully supervised manner. To the best of our knowledge, this is the first method that uses cloud-level weak labels on raw 3D space to train a point cloud semantic segmentation network. In our setting, the 3D weak labels only indicate the classes that appeared in our input sample. We discuss both scene- and subcloud-level weakly labels on raw 3D point cloud data and perform in-depth experiments on them. On ScanNet dataset, our result trained with subcloud-level labels is compatible with some fully supervised methods. Jiacheng Wei, Guosheng Lin, Kim-Hui Yap, Tzu-Yi Hung, Lihua Xie 0001 |
CVPR | 4 |
| 2020 | Weakly Supervised Segmentation with Maximum Bipartite Graph MatchingabstractIn the weakly supervised segmentation task with only image-level labels, a common step in many existing algorithms is first to locate the image regions corresponding to each existing class with the Class Activation Maps (CAMs), and then generate the pseudo ground truth masks based on the CAMs to train a segmentation network in the fully supervised manner. The quality of the CAMs has a crucial impact on the performance of the segmentation model. We propose to improve the CAMs from a novel graph perspective. We model paired images containing common classes with a bipartite graph and use the maximum matching algorithm to locate corresponding areas in two images. The matching areas are then used to refine the predicted object regions in the CAMs. The experiments on Pascal VOC 2012 dataset show that our network can effectively boost the performance of the baseline model and achieves new state-of-the-art performance. Weide Liu, Chi Zhang 0007, Guosheng Lin, Tzu-Yi Hung, Chunyan Miao |
ACM Multimedia | 4 |
| 2020 | Deep historical long short-term memory network for action recognition
Junlin Hu 0001, Tzu-Yi Hung, Yap-Peng Tan |
Neurocomputing | 4 |
| 2020 | RGBD Salient Object Detection via Disentangled Cross-Modal FusionabstractDepth is beneficial for salient object detection (SOD) for its additional saliency cues. Existing RGBD SOD methods focus on tailoring complicated cross-modal fusion topologies, which although achieve encouraging performance, are with a high risk of over-fitting and ambiguous in studying cross-modal complementarity. Different from these conventional approaches combining cross-modal features entirely without differentiating, we concentrate our attention on decoupling the diverse cross-modal complements to simplify the fusion process and enhance the fusion sufficiency. We argue that if cross-modal heterogeneous representations can be disentangled explicitly, the cross-modal fusion process can hold less uncertainty, while enjoying better adaptability. To this end, we design a disentangled cross-modal fusion network to expose structural and content representations from both modalities by cross-modal reconstruction. For different scenes, the disentangled representations allow the fusion module to easily identify, and incorporate desired complements for informative multi-modal fusion. Extensive experiments show the effectiveness of our designs and a large outperformance over state-of-the-art methods. Hao Chen 0034, Yongjian Deng, Youfu Li 0001, Tzu-Yi Hung, Guosheng Lin |
IEEE Trans. Image Process. | 4 |
| 2017 | Phase Fourier Reconstruction for Anomaly Detection on Metal Surface Using Salient Irregularity
Tzu-Yi Hung, Sriram Vaikundam, Vidhya Natarajan, Liang-Tien Chia |
MMM (1) | 1 |
| 2016 | Anomaly region detection and localization in metal surface inspectionabstractVisual inspection and identification of anomalies are important steps in the manufacturing process. A data set containing images of metallic components at different orientations and lighting conditions have been used for this experiment. Our aim is to find the anomaly in an image by subtracting it from an equivalent reference image with no anomalies. This is challenging due to the availability of multiple and slightly changing viewpoints for the reference image. The anomalies can only be detected in the residue obtained by using an identical reference image with no anomalies in the subtraction process. The proposed system is based on finding the best reference from a pool of similar images. A method for ranking the images based on similarity is also introduced. The proposed method scales well for high contrast images with dark colored anomalies. Sriram Vaikundam, Tzu-Yi Hung, Liang-Tien Chia |
ICIP | 2 |
| 2014 | Efficient Sparsity Estimation via Marginal-Lasso Coding
Tzu-Yi Hung, Jiwen Lu, Yap-Peng Tan, Shenghua Gao |
ECCV (4) | 1 |
| 2013 | Activity-based human identificationabstractWe investigate in this paper the problem of activity-based human identification. Different from most existing gait recognition methods where only human walking activity is considered and utilized for person identification, we aim to identify people from various activities such as eating, jumping, and weaving. For each video clip, we first extract binary human body masks by using background substraction, followed by computing the average energy image (AEI) features to represent each video clip. Then, a mapping is learned by applying an adaptive discriminant analysis (ADA) method to project AEI features into a low-dimensional subspace, such that the intra-class (activities performed by the same person) variations are minimized and the interclass (activities performed by different persons) are maximized, simultaneously. Moreover, interclass samples with large similarity difference are deemphasized and those with small difference are emphasized, such that more discriminative information can be used for recognition. Experimental results on three publicly available databases show the efficacy of our proposed approach. Tzu-Yi Hung, Jiwen Lu, Junlin Hu 0001, Yap-Peng Tan, Yongxin Ge |
ICASSP | 1 |
| 2013 | Graph-based sparse coding and embedding for activity-based human identificationabstractIn this paper, we propose a new graph-based sparse coding and embedding (GSCE) method for activity-based human identification. Different from human activity recognition which recognizes different types of human activities such as walking, running, eating, and drinking, in this study, we aim to identify persons from his/her activities. To our best knowledge, this problem has been seldom investigated in the literature. Given a training set of video clips, we first extract human body mask in each frame and learn a codebook to quantize these masks into a histogram feature by using a graphbased sparse coding technique to better preserve the similarity information of different frames within a same video clip. Moreover, we also learn a mapping to project each frame into a low-dimensional subspace to speed up the quantization procedure, such that more discriminative information can be further exploited for classification. Experimental results on three databases are presented to show the efficacy of the proposed method. Tzu-Yi Hung, Jiwen Lu, Yap-Peng Tan |
ICME | 1 |
| 2013 | Cross-scene abnormal event detectionabstractThis paper presents an cross-scene abnormal event detection method by adopting Bag of Words (BoW) model with Spatial Pyramid Matching Kernel (SPM) cooperating with SIFT features and a SVM classifier. Different from existing abnormal event detection methods where abnormal events happened in a well-learned scene are considered and detected, we aim to detect concerned events in public where scenes can be unlearned before. Our method is motivated by the fact that the pattern of the notable events are similar and the learned models should be transferable to examine the events in other unlearned public scenes. To learn the patterns for an abnormal event, we divide the proposed method into two steps: feature coding and spatial pooling. For the feature coding step, the codebook is generated and the feature is quantized based on small patches. For the spatial pooling step, the patches are concatenating to exploit the spatial information of local regions. The intersection kernel is used to integrate with a SVM classifier. Experimental results on two benchmark databases demonstrate the efficacy of our proposed approach. Tzu-Yi Hung, Jiwen Lu, Yap-Peng Tan |
ISCAS | 1 |
| 2012 | Video organization: Near-Duplicate Video clusteringabstractIt is not uncommon to see several videos of almost identical content on the internet. These near duplicates, coupled with the sheer number of videos, pose a big challenge to the effective organization of video clips online. We propose an adaptive classification approach to detect near-duplicate versions, and an integrated voting strategy to group clusters and to elect a representative for each cluster. Our proposed methods are based on our observation that near-duplicate videos usually span a small, albeit variable area in the feature space, while videos of different contents are scattered far apart. The classification method aims to select a suitable threshold by maximizing the margin for each video sequence in the similarity space, and the voting scheme focuses on merging subsets with mutual information based on neighbor information and inverted indices. Experimental results on an unconstrained web dataset including over 10000 videos demonstrate the efficacy of the proposed methods. Tzu-Yi Hung, Ce Zhu, Gao Yang 0001, Yap-Peng Tan |
ISCAS | 1 |
| 2011 | Packet scheduling with playout adaptation for scalable video delivery over wireless networks
Tzu-Yi Hung, Yap-Peng Tan |
J. Vis. Commun. Image Represent. | 1 |
| 2010 | Playout adaptation based packet scheduling for scalable video delivery over wireless linksabstractIn this paper, we propose an efficient packet scheduling algorithm based on a novel adaptive playout for scalable video delivery over wireless links. The proposed playout adaptation algorithm monitors the playout buffer status to adjust the playout speed by active and passive ways to prevent playout buffer from underflow and avoid annoying playout interruptions. The active playout adaptation proportionally adjusts the playout speed based on the buffer fullness. The passive playout adaptation is enabled when the network is seriously congested and adjusts the playout speed to the lowest one directly. In addition to the adaptive playout deadline, the packet priority and channel conditions are incorporated in the proposed packet scheduling algorithm. Packets are selected for transmission by maximizing the quality of the received video. The playout-deadline aware packet retransmission is also performed. The simulation results show that the proposed approach can efficiently reduce the playout latency and minimize the video quality degradation. Tzu-Yi Hung, Yap-Peng Tan |
ICME | 1 |
| 2008 | Video forgery detection using correlation of noise residueabstractWe propose a new approach for locating forged regions in a video using correlation of noise residue. In our method, block-level correlation values of noise residual are extracted as a feature for classification. We model the distribution of correlation of temporal noise residue in a forged video as a Gaussian mixture model (GMM). We propose a two-step scheme to estimate the model parameters. Consequently, a Bayesian classifier is used to find the optimal threshold value based on the estimated parameters. Two video inpainting schemes are used to simulate two different types of forgery processes for performance evaluation. Simulation results show that our method achieves promising accuracy in video forgery detection. Chih-Chung Hsu, Tzu-Yi Hung, Chia-Wen Lin, Chiou-Ting Hsu |
MMSP | 2 |