Huei-Fang Yang

dblp:56/7379 · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
7since 2021 · last 2026
0000-0001-8261-6965ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 2 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Syntax Element Encryption for H.265/HEVC Using Chaotic Map-Based Coefficient Scrambling Scheme
abstract
In today’s digital landscape, high-efficiency video coding (H.265/HEVC) has emerged as the most widely used video coding standard, employing selective encryption schemes to protect the privacy of video content while maintaining efficient compression performance. However, existing coefficient scrambling methods impose a significant computational load, leading to increased bit rate overhead due to encryption, longer execution times, and insufficient safety measures. To address these issues, a new coefficient scrambling scheme based onchaotic mapsis proposed. This approach leverages the pseudorandomness, ergodicity, and sensitivity to initial conditions inherent in chaotic maps to generate highly unpredictable coefficient distributions, thereby strengthening security while preserving low complexity. Unlike conventional scrambling, chaotic maps ensure minimal correlation between encrypted coefficients, enhancing resistance against statistical and differential attacks. Additionally, the scrambling conditions are specifically designed to minimize the impact on the bit rate overhead. Furthermore, when combined with syntax element encryption (SEC), which includes motion vector difference (MVD), quantized transform coefficients (QTC), and luma intraprediction mode (Luma IPM), this method effectively distorts video content. The proposed scheme operates synchronously with slices, ensuring that the decryption of video content remains intact even if some slices are lost. Additionally, a random sequence generated by AES-CTR is incorporated with the H.265 encoded stream to protect against chosen-plaintext attacks. The experimental results indicate that this scheme features high security, compliance with format standards, fast execution times, synchronous updates with slices, and resilience against common attacks, all while achieving a reduced bit rate overhead of 45.13% with a lowered average execution time overhead of 1.91%.
Liang-Wei Li, Chung-Nan Lee, Kishu Gupta, Huei-Fang Yang, Ashutosh Kumar Singh 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Continual Learning for Weakly-Supervised Histopathology Tissue Segmentation
abstract
Weakly supervised histopathology segmentation is a widely studied field that aims to achieve pixel-level semantic segmentation using image-level annotations, reducing the need for labor-intensive labeling. Despite significant advances in this task, existing methods assume the availability of all training data at once during training. Since medical image collections typically expand over time in practice, such methods become impractical. Meanwhile, research on continual semantic segmentation has also made significant progress. However, most existing works still rely on pixel-level annotations to train models. As a result, integrating continual learning into weakly supervised segmentation models has emerged as a promising direction. To address this challenge, we propose CL4WSeg, a novel end-to-end transformer-based framework that employs temporal distillation to leverage features from previous models for continual weakly supervised segmentation. Furthermore, we utilize a controllable diffusion model to enable generative replay and integrate an image quality filter to collect high-quality images, alleviating catastrophic forgetting. Experiments on the LUAD-HistoSeg, BCSS-WSSS, and WSSS4LUAD datasets demonstrate that our approach outperforms state-of-the-art methods.
Huei-Fang Yang, Chu-Song Chen
CIBCB2
2025 PDSeg: Patch-Wise Distillation and Controllable Image Generation for Weakly-Supervised Histopathology Tissue Segmentation
abstract
Weakly-supervised semantic segmentation, which achieves pixel-wise segmentation using image-level labels, has emerged as an alternative to fully supervised methods by reducing the need for detailed annotations. Inspired by the recent success of the teacher-student strategy in various vision tasks, we present a transformer-based weakly supervised framework that distills knowledge from a CNN teacher. Specifically, we incorporate a sequence of patch-wise distillation tokens into the transformer student, with each token focused on learning a specific patch under the teacher’s guidance. This design enables the teacher to provide more reliable supervision to the student. On the other hand, in pathology images, it is often observed that certain tissue types are less represented than others. This class imbalance poses a significant challenge for many WSSS algorithms. To address this issue, we further introduce a data synthesis pipeline using a diffusion model conditioned on semantic label maps to mitigate the effects of class imbalance in histopathology images. Unlike previous methods that rely on full annotations to construct semantic label maps, our approach leverages the intrinsic characteristics of histopathology images. This leads to an approach that does not require full annotations and is well-suited for weakly-supervised scenarios. Through extensive experiments on the LUAD-HistoSeg and BCSS-WSSS datasets, we demonstrate that our approach outperforms state-of-the-art methods.
Yu-Hsing Hsieh, Huei-Fang Yang, Chu-Song Chen
ICASSP3
2023 Continual Cell Instance Segmentation of Microscopy Images
abstract
A continual cell instance segmenter aims to continually learn to segment new objects while preserving the ability to localize and distinguish old objects without access to previous data. Besides catastrophic forgetting, background shift, where the background class could contain objects in the old and unseen future classes, could occur. In addition, as acquiring annotations is label-intensive, cell images can be partially labeled. In this paper, we present iMRCNN, which extends Mask R-CNN with knowledge distillation and pseudo labeling, to address these challenges. To preserve the learned skills, the current student distills knowledge from the former teacher at output and feature levels. Furthermore, we employ a pseudo labeling scheme, where the teacher is utilized to identify objects with no labels provided, to deal with background shift and partially labeled data. Experiments on two microscopy image sets demonstrate the effectiveness of iMRCNN over other alternatives in various incremental learning scenarios.
Tzu-Ting Chuang, Ting-Yun Wei, Yu-Hsing Hsieh, Chu-Song Chen, Huei-Fang Yang
ICASSP5
2023 Class-incremental Continual Learning for Instance Segmentation with Image-level Weak Supervision
abstract
Instance segmentation requires labor-intensive manual labeling of the contours of complex objects in images for training. The labels can also be provided incrementally in practice to balance the human labor in different time steps. However, research on incremental learning for instance segmentation with only weak labels is still lacking. In this paper, we propose a continual-learning method to segment object instances from image-level labels. Unlike most weakly-supervised instance segmentation (WSIS) which relies on traditional object proposals, we transfer the semantic knowledge from weakly-supervised semantic segmentation (WSSS) to WSIS to generate instance cues. To address the background shift problem in continual learning, we employ the old class segmentation results generated by the previous model to provide more reliable semantic and peak hypotheses. To our knowledge, this is the first work on weakly-supervised continual learning for instance segmentation of images. Experimental results show that our method can achieve better performance on Pascal VOC and COCO datasets under various incremental settings1.
Yu-Hsing Hsieh, Guan-Sheng Chen, Shun-Xian Cai, Ting-Yun Wei, Huei-Fang Yang, Chu-Song Chen
ICCV5
2022 Contrastive Self-Supervised Learning as a Strong Baseline for Unsupervised Hashing
abstract
Contrastive self-supervised learning has shown to learn representations transferable to a variety of downstream applications, e.g., object detection and classification. While utilizing a contrastive self-supervised objective to learn generalizable features has been much explored, employing it to directly learn binary representations for image search is yet to be studied. This paper presents Contrastive Self-supervised deep Hashing (CSHash), a simple yet effective unsupervised hashing framework aimed at producing compact binary hash codes that preserve semantic similarity between data without relying on human annotations. CSHash is trained on the basis of a contrastive learning objective, pulling together the augmentations of the same sample and keep apart those of different samples in the hash code space. Evaluation on three datasets shows that simply trained by a contrastive self-supervised loss, CSHash is a strong baseline for unsupervised hashing. It yields discriminative, high-quality binary codes and performs comparably to other unsupervised hashing methods. Additionally, we perform thorough analyses on the main components of CSHash to provide a better insight into the framework.
Huei-Fang Yang
MMSP1
2022 Learning Binary Hash Codes Based on Adaptable Label Representations
abstract
The goal of supervised hashing is to construct hash mappings from collections of images and semantic annotations such that semantically relevant images are embedded nearby in the learned binary hash representations. Existing deep supervised hashing approaches that employ classification frameworks with a classification training objective for learning hash codes often encode class labels as one-hot or multi-hot vectors. We argue that such label encodings do not well reflect semantic relations among classes and instead, effective class label representations ought to be learned from data, which could provide more discriminative signals for hashing. In this article, we introduce Adaptive Labeling Deep Hashing (AdaLabelHash) that learns binary hash codes based on learnable class label representations. We treat the class labels as the vertices of a K -dimensional hypercube, which are trainable variables and adapted together with network weights during the backward network training procedure. The label representations, referred to as codewords, are the target outputs of hash mapping learning. In the label space, semantically relevant images are then expressed by the codewords that are nearby regarding Hamming distances, yielding compact and discriminative binary hash representations. Furthermore, we find that the learned label representations well reflect semantic relations. Our approach is easy to realize and can simultaneously construct both the label representations and the compact binary embeddings. Quantitative and qualitative evaluations on several popular benchmarks validate the superiority of AdaLabelHash in learning effective binary codes for image search.
Huei-Fang Yang, Cheng-Hao Tu 0001, Chu-Song Chen
IEEE Trans. Neural Networks Learn. Syst.1
2020 Cross-Batch Reference Learning for Deep Retrieval
abstract
Learning effective representations that exhibit semantic content is crucial to image retrieval applications. Recent advances in deep learning have made significant improvements in performance on a number of visual recognition tasks. Studies have also revealed that visual features extracted from a deep network learned on a large-scale image data set (e.g., ImageNet) for classification are generic and perform well on new recognition tasks in different domains. Nevertheless, when applied to image retrieval, such deep representations do not attain performance as impressive as used for classification. This is mainly because the deep features are optimized for classification rather than for the desired retrieval task. We introduce the cross-batch reference (CBR), a novel training mechanism that enables the optimization of deep networks with a retrieval criterion. With the CBR, the networks leverage both the samples in a single minibatch and the samples in the others for weight updates, enhancing the stochastic gradient descent (SGD) training by enabling interbatch information passing. This interbatch communication is implemented as a cross-batch retrieval process in which the networks are trained to maximize the mean average precision (mAP) that is a popular performance measure in retrieval. Maximizing the cross-batch mAP is equivalent to centralizing the samples relevant to each other in the feature space and separating the samples irrelevant to each other. The learned features can discriminate between relevant and irrelevant samples and thus are suitable for retrieval. To circumvent the discrete, nondifferentiable mAP maximization, we derive an approximate, differentiable lower bound that can be easily optimized in deep networks. Furthermore, the mAP loss can be used alone or with a classification loss. Experiments on several data sets demonstrate that our CBR learning provides favorable performance, validating its effectiveness.
Huei-Fang Yang, Ting-Yen Chen, Chu-Song Chen
IEEE Trans. Neural Networks Learn. Syst.1
2019 Adaptive Labeling For Hash Code Learning Via Neural Networks
abstract
Learning-based hash has been widely used for large-scale similarity retrieval due to the efficient computation and condensed storage of binary representations. In this paper, we propose AdaLabelHash, a hash function learning approach via neural networks. In AdaLabelHash, class label representations are adaptable during the network training. We express the labels as hypercube vertices in a K-dimensional space, and both the network weights and class label representations are updated in the learning process. As the label representations are explored from data, semantically similar categories will be assigned with the label representations that are close to each other in terms of Hamming distance in the label space. The label representations then serve as the desired output of the hash function learning so as to yield compact and discriminating binary hash codes via the network. AdaLabelHash is simple but effective, which can jointly learn label representations and infer compact binary codes from data. It is applicable to both supervised and semi-supervised learning of hash codes. Experimental results on standard benchmarks show the effectiveness of AdaLabelHash.
Huei-Fang Yang, Cheng-Hao Tu 0001, Chu-Song Chen
ICIP1
2018 Supervised Learning of Semantics-Preserving Hash via Deep Convolutional Neural Networks
abstract
This paper presents a simple yet effective supervised deep hash approach that constructs binary hash codes from labeled data for large-scale image search. We assume that the semantic labels are governed by several latent attributes with each attribute on or off, and classification relies on these attributes. Based on this assumption, our approach, dubbed supervised semantics-preserving deep hashing (SSDH), constructs hash functions as a latent layer in a deep network and the binary codes are learned by minimizing an objective function defined over classification error and other desirable hash codes properties. With this design, SSDH has a nice characteristic that classification and retrieval are unified in a single learning model. Moreover, SSDH performs joint learning of image representations, hash codes, and classification in a point-wised manner, and thus is scalable to large-scale datasets. SSDH is simple and can be realized by a slight enhancement of an existing deep architecture for classification; yet it is effective and outperforms other hashing approaches on several benchmarks and large datasets. Compared with state-of-the-art approaches, SSDH achieves higher retrieval accuracy, while the classification performance is not sacrificed.
Huei-Fang Yang, Chu-Song Chen
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Equivalent Scanning Network of Unpadded CNNs
abstract
This letter presents a theory of scanning a signal with a sliding window, where the window's mapping function is built upon a convolutional neural network (CNN). When using a CNN as the sliding window, we show that the resultant feature maps are equivalent to the maps obtained by applying another CNN (called EQ-ScanNet) to the whole signal. The EQ-ScanNet can be established by reconfiguring the original CNN with dilated (i.e., sparse kernel) convolutions. We clarify that, this property is originated from the noble identity (i.e., the swapping equivalence of downsample and FIR filter), and extend the property to the generalized convolution that subsumes CNN's window-sliding operations. We further show that an unpadded CNN is a necessary condition for formulating the EQ-ScanNet.
Huei-Fang Yang, Ting-Yen Chen, Cheng-Hao Tu 0001, Chu-Song Chen
IEEE Signal Process. Lett.1
2018 Joint Estimation of Age and Expression by Combining Scattering and Convolutional Networks
abstract
This article tackles the problem of joint estimation of human age and facial expression. This is an important yet challenging problem because expressions can alter face appearances in a similar manner to human aging. Different from previous approaches that deal with the two tasks independently, our approach trains a convolutional neural network (CNN) model that unifies ordinal regression and multi-class classification in a single framework. We demonstrate experimentally that our method performs more favorably against state-of-the-art approaches.
Huei-Fang Yang, Bo-Yao Lin, Kuang-Yu Chang, Chu-Song Chen
ACM Trans. Multim. Comput. Commun. Appl.1
2016 Cross-batch Reference Learning for Deep Classification and Retrieval
abstract
Learning feature representations for image retrieval is essential to multimedia search and mining applications. Recently, deep convolutional networks (CNNs) have gained much attention due to their impressive performance on object detection and image classification, and the feature representations learned from a large-scale generic dataset (e.g., ImageNet) can be transferred to or fine-tuned on the datasets of other domains. However, when the feature representations learned with a deep CNN are applied to image retrieval, the performance is still not as good as they are used for classification, which restricts their applicability to relevant image search. To ensure the retrieval capability of the learned feature space, we introduce a new idea called cross-batch reference (CBR) to enhance the stochastic-gradient-descent (SGD) training of CNNs. In each iteration of our training process, the network adjustment relies not only on the training samples in a single batch, but also on the information passed by the samples in the other batches. This inter-batches communication mechanism is formulated as a cross-batch retrieval process based on the mean average precision (MAP) criterion, where the relevant and irrelevant samples are encouraged to be placed on top and rear of the retrieval list, respectively. The learned feature space is not only discriminative to different classes, but the samples that are relevant to each other or of the same class are also enforced to be centralized. To maximize the cross-batch MAP, we design a loss function that is an approximated lower bound of the MAP on the feature layer of the network, which is differentiable and easier for optimization. By combining the intra-batch classification and inter-batch cross-reference losses, the learned features are effective for both classification and retrieval tasks. Experimental results on various benchmarks demonstrate the effectiveness of our approach.
Huei-Fang Yang, Chu-Song Chen
ACM Multimedia1
2015 Automatic Age Estimation from Face Images via Deep Ranking
abstract
This paper focuses on automatic age estimation (AAE) from face images, which amounts to determining the exact age or age group of a face image according to features from faces. Although great effort has been devoted to AAE [1, 4, 6], it remains a challenging problem. The difficulties are due to large facial appearance variations resulting from a number of factors, e.g., aging and facial expressions. AAE algorithms need to overcome heterogeneity in facial appearance changes to provide accurate age estimates. To this end, we propose a generic, deep network model for AAE (see Figure 1). Given a face image, our network first extracts features from the face through a 3-layer scattering network (ScatNet) [2], then reduces the feature dimension by principal component analysis (PCA), and finally predicts the age via category-wise rankers constructed as a 3-layer fullyconnected network. The contributions are: (1) Our ranking method is point-wised and thus is easily scaled up to large-scale datasets; (2) our deep ranking model is general and can be applied to age estimation from faces with large facial appearance variations as a result of aging or facial expression changes; and (3) we show that the high-level concepts learned from large-scale neutral faces can be transferred to estimating ages from faces under expression changes, leading to improved performance. Our model is with the following characteristics: (1) The scattering features are invariant to translation and small deformations. ScatNet is a deep convolutional network of specific characteristics. It uses predefined wavelets and computes scattering representations via a cascade of wavelet transforms and modulus pooling operators from shallow to deep layers. With the nonlinear modulus and averaging operators, ScatNet can produce representations that are discriminative as well as invariant to translation and small deformations. As ScatNet provides fundamentally invariant representations for discriminating feature extraction, only the weights of the fully-connected layers are learned in our network model, which considerably reduces the training time. (2) The rank labels encoded in the network exploit the ordering relation among labels. Each category-wise ranker is an ordinal regression ranker. We encode the age rank based on the reduction framework [5]. Given a set of training samples X = {(xi,yi), i = 1 · · ·N}, let xi ∈ RD be the input image and yi be a rank label (yi ∈ {1, . . . ,K}), respectively, where K is the number of age ranks. For rank k, we separate X into two subsets, X k and X − k , as follows: X k = {(xi,+1)|yi > k} X− k = {(xi,−1)|yi ≤ k}. (1)
Huei-Fang Yang, Bo-Yao Lin, Kuang-Yu Chang, Chu-Song Chen
BMVC1
2015 Rapid Clothing Retrieval via Deep Learning of Binary Codes and Hierarchical Search
abstract
This paper deals with the problem of clothing retrieval in a recommendation system. We develop a hierarchical deep search framework to tackle this problem. We use a pre-trained network model that has learned rich mid-level visual representations in module 1. Then, in module 2, we add a latent layer to the network and have neurons in this layer to learn hashes-like representations while fine-tuning it on the clothing dataset. Finally, module 3 achieves fast clothing retrieval using the learned hash codes and representations via a coarse-to-fine strategy. We use a large clothing dataset where 161,234 clothes images are collected and labeled. Experiments demonstrate the potential of our proposed framework for clothing retrieval in a large corpus.
Huei-Fang Yang, Kuan-Hsien Liu, Jen-Hao Hsiao, Chu-Song Chen
ICMR2
2015 Quick browsing and retrieval for surveillance videos
Cheng-Chieh Chiang, Huei-Fang Yang
Multim. Tools Appl.2
2012 Tracking Growing Axons by Particle Filtering in 3D + t Fluorescent Two-Photon Microscopy Images
Huei-Fang Yang, Xavier Descombes, Charles Kervrann, Caroline Medioni, Florence Besse
ACCV (3)1
2009 3D volume extraction of densely packed cells in EM data stack by forward and backward graph cuts
abstract
3D reconstruction on dense nanoscale medical images is a very challenging research topic. The challenge comes from the fact that boundaries of objects on such images are not always very clear due to imperfect staining. This makes the segmentation of dense nanoscale medical images very difficult and thus increases the difficulty in 3D reconstruction. In this paper, we proposed a method based on watershed and an interactive segmentation technique, graph cuts, to extract 3D volumes from dense nanoscale medical images. In our method, images are first segmented by a marker-controlled watershed algorithm. Markers for watershed segmentation algorithm are seed points generated by using distance transform, followed by a new grouping method that clusters seed points that are too close. Regions obtained by watershed transform segmentation algorithms are considered as nodes in a graph. Edges are to connect between the nodes in adjacent image slices. The weight on each edge is defined based on the overlapped area between nodes. User-selected nodes (regions) in an initial image slice serve as hard constraints in the minimization process. A globally optimal 3D volume is obtained by minimizing MAP-MRF energy function via graph cuts. In our application, in order to obtain a complete 3D volume structures including branching, the final 3D volume is the union of two 3D volumes obtained by performing the minimization of MAP-MRF energy function using graph cuts forwards and backwards through the image stack. Experiments are conducted both on synthetic data and on nanoscale image sequences from the Serial Block Face Scanning Electron Microscope (SBF-SEM). The results show that our method can successfully extract 3D volumes.
Huei-Fang Yang, Yoonsuck Choe
CIMSIVP1
2009 Indexing and Teaching Focus Mining of Lecture Videos
abstract
This paper proposes an indexing and teaching focus mining system for lecture videos recorded in an unconstraint environment. The slide structure can be reconstructed by an edge-based shot change detection algorithm. Besides, the teaching focus can be extracted according to instructor’s behavior, including the gesture, the lecture time for each slide, and the speech speed. Experiment results show the feasibility of the proposed method, that is, the slide shots can be correctly detected even if the illumination conditions is variant or the slides are obstructed by the instructor or students, and the teaching focus can be well extracted to provide learners an efficient way to study.
Yu-Tzu Lin, Bai-Jang Yen, Chia-Hu Chang, Huei-Fang Yang, Greg C. Lee
ISM4