Daniel Stanley Tan

dblp:216/6502 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
12since 2021 · last 2024
0000-0002-8071-9060ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Unilateral Facial Action Unit Detection: Revealing Nuanced Facial Expressions
abstract
Facial expressions are essential for non-verbal human communication as they convey behavioral intentions and emotional states. While facial action units (AUs) can occur bilaterally or unilaterally, the existing research in affective computing predominantly concentrates on bilateral expressions, largely due to the lack of datasets with unilateral AU labels. In this study, we present a method for generating unilateral AU labels and assess its efficacy against expert-labeled facial images. Furthermore, we introduce a dedicated model trained on the generated data and evaluate its performance across multiple datasets. Our findings offer insights into feature extraction for unilateral facial expression recognition. This research contributes to advancing the understanding and recognition of nuanced facial expressions, with potential applications in various domains such as healthcare and human-computer interaction.
Deniz Iren, Daniel Stanley Tan
ACII2
2023 Language Modeling in Logistics: Customer Calling Prediction
abstract
Customer centers in logistics companies deal with many customer calls and requests daily.One of the most common calls is related to requesting an update on the shipment status.Proactively sending message updates to customers can reduce the number of calls.However, naively sending updates to everyone can cause unnecessary anxiety to people who do not want it, thus leading to customer dissatisfaction or even more calls.If a machine learning model could predict shipments leading to a customer call based on its journey, it could be possible to proactively send message updates only to customers likely to make a call.Therefore, reducing the workload in the customer center while increasing customer satisfaction.In large logistic companies where the volume of calls can reach a million calls per month, even 10% of the reduction of calls could already significantly reduce the additional expenses and workload associated with tracing a shipment.In this paper, we formulate the shipment journey as a variant of a language model.Specifically, we treat checkpoints (station, facility, time, event code) as tokens and predict the next checkpoint(station, facility, time delta, event code).Our core insight is that shipment checkpoints follow a set of rules that dictate the possible sequence of checkpoints.This is similar to how grammar rules dictate which words can follow another.Despite remaining a difficult problem, our experiments show that features learned by modeling shipment checkpoints as a language model can improve customer calling prediction.
Giacomo Anerdi, Daniel Stanley Tan, Stefano Bromuri
ESANN3
2023 MCGAN: mask controlled generative adversarial network for image retargeting
Jilyan Bianca Dy, John Jethro Virtusio, Daniel Stanley Tan, Yong-Xiang Lin, Joel P. Ilao, Yung-Yao Chen, Kai-Lung Hua
Neural Comput. Appl.3
2023 MACnet: Mask augmented counting network for class-agnostic counting
Tadhg McCarthy, John Jethro Virtusio, Jose Jaena Mari Ople, Daniel Stanley Tan, Divina Amalin, Kai-Lung Hua
Pattern Recognit. Lett.4
2023 ConCoNet: Class-agnostic counting with positive and negative exemplars
abstract
Class-agnostic counting is usually phrased as a matching problem between a user-defined exemplar patch and a query image. The count is derived based on the number of objects similar to the exemplar patch. However, defining a target class using only positive exemplar patches inevitably miscounts unintended objects that are visually alike to the exemplar. In this paper, we propose to include negative exemplars that define what not to count. This allows the model to calibrate its notion of what is similar based on both positive and negative exemplars. It effectively disentangles visually similar negatives, leading to a more discriminative definition of the target object. We designed our method such that it can be incorporated with other class-agnostic counting models. Moreover, application-wise, our model can be used into a semi-automatic labeling tool to simplify the job of the annotator
Adrienne Francesca O. Soliven, John Jethro Virtusio, Jose Jaena Mari Ople, Daniel Stanley Tan, Divina Amalin, Kai-Lung Hua
Pattern Recognit. Lett.4
2022 CLIPCAM: A Simple Baseline For Zero-Shot Text-Guided Object And Action Localization
abstract
The key for the contemporary deep learning-based object and action localization algorithms to work is the large-scale annotated data. However, in real-world scenarios, since there are infinite amounts of unlabeled data beyond the categories of publicly available datasets, it is not only time- and manpower-consuming to annotate all the data but also requires a lot of computational resources to train the detectors. To address these issues, we show a simple and reliable baseline that can be easily obtained and work directly for the zero-shot text-guided object and action localization tasks without introducing additional training costs by using Grad-CAM, the widely used class visual saliency map generator, with the help of the recently released Contrastive Language-Image Pre-Training (CLIP) model by OpenAI, which is trained contrastively using the dataset of 400 million image-sentence pairs with rich cross-modal information between text semantics and image appearances. With extensive experiments on the Open Images and HICO-DET datasets, the results demonstrate the effectiveness of the proposed approach for the text-guided unseen object and action localization tasks for images.
Hsuan-An Hsia, Che-Hsien Lin, Bo-Han Kung, Jhao-Ting Chen, Daniel Stanley Tan, Jun-Cheng Chen, Kai-Lung Hua
ICASSP5
2022 Controllable and Identity-Aware Facial Attribute Transformation
abstract
Modifying facial attributes without the paired dataset proves to be a challenging task. Previously, approaches either required supervision from a ground-truth transformed image or required training a separate model for mapping every pair of attributes. These limit the scalability of the models to accommodate a larger set of attributes since the number of models that we need to train grows exponentially large. Another major drawback of the previous approaches is the unintentional gain of the identity of the person as they transform the facial attributes. We propose a method that allows for controllable and identity-aware transformations across multiple facial attributes using only a single model. Our approach is to train a generative adversarial network (GAN) with a multitask conditional discriminator that recognizes the identity of the face, distinguishes real images from fake, as well as identifies facial attributes present in an image. This guides the generator into producing an output that is realistic while preserving the person's identity and facial attributes. Through this framework, our model also learns meaningful image representations in a lower dimensional latent space and semantically associate separate parts of the encoded vector with both the person's identity and facial attributes. This opens up the possibility of generating new faces and other transformations such as making the face thinner or chubbier. Furthermore, our model only encodes the image once and allows for multiple transformations using the encoded vector. This allows for faster transformations since it does not need to reprocess the entire image for every transformation. We show the effectiveness of our proposed method through both qualitative and quantitative evaluations, such as ablative studies, visual inspection, and face verification. Competitive results are achieved compared to the main competition (CycleGAN), however, at great space and extensibility gain by using a single model.
Daniel Stanley Tan, Jonathan Hans Soeseno, Kai-Lung Hua
IEEE Trans. Cybern.1
2021 Naturalistic Physical Adversarial Patch for Object Detectors
abstract
Most prior works on physical adversarial attacks mainly focus on the attack performance but seldom enforce any restrictions over the appearance of the generated adversarial patches. This leads to conspicuous and attention-grabbing patterns for the generated patches which can be easily identified by humans. To address this issue, we pro-pose a method to craft physical adversarial patches for object detectors by leveraging the learned image manifold of a pretrained generative adversarial network (GAN) (e.g., BigGAN and StyleGAN) upon real-world images. Through sampling the optimal image from the GAN, our method can generate natural looking adversarial patches while maintaining high attack performance. With extensive experiments on both digital and physical domains and several independent subjective surveys, the results show that our proposed method produces significantly more realistic and natural looking patches than several state-of-the-art base-lines while achieving competitive attack performance.1
Yu-Chih-Tuan Hu, Jun-Cheng Chen, Bo-Han Kung, Kai-Lung Hua, Daniel Stanley Tan
ICCV5
2021 TrustMAE: A Noise-Resilient Defect Classification Framework using Memory-Augmented Auto-Encoders with Trust Regions
abstract
In this paper, we propose a framework called Trust-MAE to address the problem of product defect classification. Instead of relying on defective images that are difficult to collect and laborious to label, our framework can accept datasets with unlabeled images. Moreover, unlike most anomaly detection methods, our approach is robust against noises, or defective images, in the training dataset. Our framework uses a memory-augmented auto-encoder with a sparse memory addressing scheme to avoid over-generalizing the auto-encoder, and a novel trust-region memory updating scheme to keep the noises away from the memory slots. The result is a framework that can reconstruct defect-free images and identify the defective regions using a perceptual distance network. When compared against various state-of-the-art baselines, our approach performs competitively under noise-free MVTec datasets. More importantly, it remains effective at a noise level up to 40% while significantly outperforming other baselines.
Daniel Stanley Tan, Trista Pei-Chun Chen, Wei-Chao Chen
WACV1
2021 Incremental Learning of Multi-Domain Image-to-Image Translations
abstract
Current multi-domain image-to-image translation models assume a fixed set of domains and that all the data are always available during training. However, over time, we may want to include additional domains to our model. Existing methods either require re-training the whole model with data from all domains or require training several additional modules to accommodate new domains. To address these limitations, we present IncrementalGAN, a multi-domain image-to-image translation model that can incrementally learn new domains using only a single generator. Our approach first decouples the domain label representation from the generator to allow it to be re-used for new domains without any architectural modification. Next, we introduce a distillation loss that prevents the model from forgetting previously learned domains. Our model compares favorably against several state-of-the-art baselines.
Daniel Stanley Tan, Yong-Xiang Lin, Kai-Lung Hua
IEEE Trans. Circuits Syst. Video Technol.1
2021 Neural Style Palette: A Multimodal and Interactive Style Transfer From a Single Style Image
abstract
Despite the myriad of attributes found in a single style image, existing neural style transfer methods produce outputs with limited variety–typically only a single realization of the style image. They also do not provide an easy way to control the stylization process, limiting the creative freedom of users. In this paper, we propose Neural Style Palette (NSP), a method for interactively generating a variety of stylized images from only a single style input. Our approach allows human influence in the stylization process, a design inspired by Hybrid Human-Artificial Intelligence. Like a color palette,NSPenables a meaningful interaction by presenting a collection of sub-textures, which we also refer to as anchor styles, that act as a visual guide for the users. These anchor styles capture different attributes in the single style image that the users can creatively blend to create their desired realizations. To offer a diversified selection in theNSP, we constrain the anchor styles to be distant from one another while maintaining faithfulness to the original style image. This is possible through our two proposed novel losses: a style-separation loss that encourages the sub-textures to be distinct and a unification loss to ensure that the sub-textures center around the original style while encouraging additional diversity. We perform several experiments to prove the effectiveness of our method and generalize to improve existing methods.
John Jethro Virtusio, Jose Jaena Mari Ople, Daniel Stanley Tan, Muhammad Tanveer 0001, Neeraj Kumar 0001, Kai-Lung Hua
IEEE Trans. Multim.3
2021 Enabling Artistic Control Over Pattern Density and Stroke Strength
abstract
Despite the remarkable results and numerous advancements in neural style transfer, achieving artistic control is still a challenging feat, primarily since existing methodologies treat the style representation as a black-box model. This oversight significantly limits the range of possible artistic manipulations. In this paper, we propose a method to enable artistic control on any correlation-based style transfer models along with guiding intuitions. Our focus is on controlling two perceptual factors: Pattern Density and Stroke Strength. To achieve this, we introduce the centered Gram style representation and manipulate it with our variance-aware adaptive weighting and correlation-based selective masking. Through several experiments and comparisons with the state-of-the-art, we show that we can achieve artistic control with competitive stylization quality. Additionally, since our method involves manipulating style representation, it can easily be adapted to popular style transfer models. We analyze different style representation properties to propose rules that govern the style transfer process, which is critical towards achieving artistic control over pattern density and stroke strength.
John Jethro Virtusio, Daniel Stanley Tan, Wen-Huang Cheng, Muhammad Tanveer 0001, Kai-Lung Hua
IEEE Trans. Multim.2
2020 Super-Resolution by Image Enhancement Using Texture Transfer
abstract
Recent deep learning approaches in single image super-resolution (SISR) can generate high-definition textures for super-resolved (SR) images. However, they tend to hallucinate fake textures and even produce artifacts. An alternative to SISR, reference-based SR (RefSR) approaches use high-resolution (HR) reference (Ref) images to provide HR details that are missing in the low-resolution (LR) input image. We propose a novel framework that leverages existing SISR approaches and enhances them with RefSR. Specifically, we refine the output of SISR methods using neural texture transfer, where HR features are queried from the Ref images. The query is conducted by computing the similarity of textural and semantic features between the input image and the Ref images. The most similar HR features, patch-wise, to the LR image is used to augment the SR image through an augmentation network. In the case of dissimilar Ref images from the LR input image, we prevent performance degradation by including the similarity scores in the input features of the network. Furthermore, we use random texture patches during the training to condition our augmentation network to not always trust the queried texture features. Different from past RefSR approaches, our method can use arbitrary Ref images and its lower-bound performance is based on the SR image. We showcase that our method drastically improves the performance of the base SISR approach.
Jose Jaena Mari Ople, Daniel Stanley Tan, Arnulfo P. Azcarraga, Chao-Lung Yang, Kai-Lung Hua
ICIP2
2019 Single-Fusion Detector: Towards Faster Multi-Scale Object Detection
abstract
Despite recent improvements, the arbitrary sizes of objects still impede the predictive ability of object detectors. Recent solutions combine feature maps of different receptive fields to detect multi-scale objects. However, these methods have large computational costs resulting to slower inference time, which is not practical for real-time applications. Contrarily, fusion methods depending on large networks with many skip connections demand larger memory requirement, prohibiting usage in devices with limited memory. In this paper, we propose a more computationally efficient fusion method which integrates higher-order information to low-level feature maps using a single operation. Our method can flexibly adapt to any base network, allowing tailored performance for different computational requirements. Our approach achieves 81.7% mAP at 41 FPS on the PASCAL VOC dataset using ResNet-50 as the base network, which is superior in terms of both speed and mAP as compared to several state-of-the-art baselines, even those which use larger base networks.
Arren Matthew C. Antioquia, Daniel Stanley Tan, Arnulfo P. Azcarraga, Kai-Lung Hua
ICIP2
2019 Spatially-Aware Domain Adaptation for Semantic Segmentation of Urban Scenes
abstract
It is very expensive and time consuming to collect a large enough dataset with pixel-level annotations to train a semantic segmentation model. Synthetic datasets are common alternatives for training segmentation models, however models trained on synthetic data do not necessarily perform well on real world images due to the domain shift problem. Domain adaptation techniques address this problem by leveraging on adversarial training to align features. Prior works have mostly performed global feature alignment. They do not consider the positions of objects. However, objects in urban scenes are highly correlated with their spatial locations. For example, the sky will always appear on top while cars will usually appear in the middle of the image. Based on this insight, we propose a spatial-aware discriminator that accounts for the spatial prior on the objects in order to improve the feature alignment. We demonstrate in our experiments that our model outperforms several state-of-the-art baselines in terms of mean intersection over union (mIoU).
Yong-Xiang Lin, Daniel Stanley Tan, Wen-Huang Cheng, Yung-Yao Chen, Kai-Lung Hua
ICIP2
2019 Segmenting Hepatic Lesions Using Residual Attention U-Net with an Adaptive Weighted Dice Loss
abstract
We propose a novel network architecture called Residual Attention U-Net (ResAttU-Net) for segmenting hepatic lesions. Our model incorporates residual blocks that can extract more complex features as compared with traditional convolutional layers combined with a skip-connection attention module that learns to focus on the relevant features for the task of hepatic lesions segmentation. Moreover, we train our model using an adaptive weighted dice loss that prioritizes the pixels of the tumor class over the pixels of the background class. We evaluate our model on the MICCAI Liver Tumor Segmentation (LiTS) benchmark dataset. Our experimental results show that our method significantly improves upon several state-of-the-art baselines for hepatic lesion or liver tumor segmentation.
Daniel Stanley Tan, Jyh-Cheng Chen, Wen-Huang Cheng, Kai-Lung Hua
ICIP2
2019 Layout and Context Understanding for Image Synthesis with Scene Graphs
abstract
Advancements on text-to-image synthesis generate remarkable images from textual descriptions. However, these methods are designed to generate only one object with varying attributes. They face difficulties with complex descriptions having multiple arbitrary objects since it would require information on the placement and sizes of each object in the image. Recently, a method that infers object layouts from scene graphs has been proposed as a solution to this problem. However, their method uses only object labels in describing the layout, which fail to capture the appearance of some objects. Moreover, their model is biased towards generating rectangular shaped objects in the absence of ground-truth masks. In this paper, we propose an object encoding module to capture object features and use it as additional information to the image generation network. We also introduce a graph-cuts based segmentation method that can infer the masks of objects from bounding boxes to better model object shapes. Our method produces more discernible images with more realistic shapes as compared to the images generated by the current state-of-the-art method.
Arces Talavera, Daniel Stanley Tan, Arnulfo P. Azcarraga, Kai-Lung Hua
ICIP2
2019 Adapting Semantic Segmentation of Urban Scenes via Mask-Aware Gated Discriminator
abstract
Training a deep neural network for semantic segmentation relies on pixel-level ground truth labels for supervision. However, collecting large datasets with pixel-level annotations is very expensive and time consuming. One workaround is to utilize synthetic data where we can generate potentially unlimited data with their corresponding ground truth labels. Unfortunately, networks trained on synthetic data perform poorly on real images due to the domain shift problem. Domain adaptation techniques have shown potential in transferring the knowledge learned from synthetic data to real world data. Prior works have mostly leveraged on adversarial training to perform a global aligning of features. However, we observed that background objects have lesser variations across different domains as opposed to foreground objects. Using this insight, we propose a method for domain adaptation that models and adapts foreground objects and background objects separately. Our approach starts with a fast style transfer to match the appearance of the inputs. This is followed by a foreground adaptation module that learns a foreground mask that is used by our gated discriminator in order to adapt the foreground and background objects separately. We demonstrate in our experiments that our model outperforms several state-of-the-art baselines in terms of mean intersection over union (mIoU).
Yong-Xiang Lin, Daniel Stanley Tan, Wen-Huang Cheng, Kai-Lung Hua
ICME2
2019 3D Object Completion via Class-Conditional Generative Adversarial Network
Yu-Chieh Chen, Daniel Stanley Tan, Wen-Huang Cheng, Kai-Lung Hua
MMM (2)2
2018 Pedestrian Detection from Lidar Data via Cooperative Deep and Hand-Crafted Features
abstract
Autopilot systems need to be able to detect pedestrians with high precision and recall regardless of whether it is during the day or night. This means that we cannot rely on normal cameras to sense the surroundings due to its sensitivity to lighting conditions. An alternative for images is to use light detection and ranging sensors (LiDAR) that produces three-dimensional point clouds where each point represents the distance to an object. However, most pedestrian detection systems are designed for image inputs and not on distance point clouds. In this paper, we propose a method for detecting pedestrians using only the three-dimensional point clouds generated by the LiDAR. Our approach first projects the three-dimensional point cloud into a two-dimensional plane. We then extract both hand-crafted features and learned features from a convolutional neural network in order to train a support vector machine (SVM) to detect pedestrians. Our proposed method achieved significant improvements in terms of F1-measurement over prior state-of-the-art methods.
Tzu-Chieh Lin, Daniel Stanley Tan, Hsueh-Ling Tang, Shih-Che Chien, Feng-Chia Chang, Yung-Yao Chen, Wen-Huang Cheng, Kai-Lung Hua
ICIP2
2018 ZipNet: ZFNet-level Accuracy with 48× Fewer Parameters
abstract
With the introduction of Convolutional Neural Networks, models for image classification achieve higher classification accuracy. Based on the pattern of the design of CNN architectures, increasing the number of layers equates to a higher classification accuracy, but also increases the number of parameters and model size. This negatively affects the model training time, processing time, and memory requirement. We develop ZipNet, a CNN architecture with a higher classification accuracy than ZFNet, the winner of ILSVRC 2013, but with 48.5× smaller model size and 48.7× fewer parameters. The classification accuracy of ZipNet is higher than the performance of ZFNet and SqueezeNet on all configurations of the Caltech-256 dataset with varying number of training examples.
Arren Matthew C. Antioquia, Daniel Stanley Tan, Arnulfo P. Azcarraga, Wen-Huang Cheng, Kai-Lung Hua
VCIP2
2018 Interactive Style Transfer: Towards Styling User-Specified Object
abstract
Researches dealing with the task of style transfer have focused and produced results that only transfers the style of a reference image to the entirety of another image. As this domain would be beneficial in digital art, it would be preferable for such algorithms to support a more flexible style transfer, such that the style would be applied only on a specific portion of an image. We propose a framework that can selectively apply a given style onto an object using only 4 user-defined points. Our approach combines a style transfer module and an object segmentation module to synthesize the stylized image. As the ultimate goal is to develop an artistic application tool, we also introduce a method that makes use of specific filters to preserve certain characteristics of an image, such as its high frequency components. Experiments show that our proposed method is able to produce pleasing results with minimal effort while also being able to handle more complicated tasks such as the application of multiple reference style images onto different objects in an image.
John Jethro Virtusio, Arces Talavera, Daniel Stanley Tan, Kai-Lung Hua, Arnulfo P. Azcarraga
VCIP3
2018 DeepDemosaicking: Adaptive Image Demosaicking via Multiple Deep Fully Convolutional Networks
abstract
Convolutional neural networks are currently the state-of-the-art solution for a wide range of image processing tasks. Their deep architecture extracts low and high-level features from images, thus, improving the model's performance. In this paper, we propose a method for image demosaicking based on deep convolutional neural networks. Demosaicking is the task of reproducing full color images from incomplete images formed from overlaid color filter arrays on image sensors found in digital cameras. Instead of producing the output image directly, the proposed method divides the demosaicking task into an initial demosaicking step and a refinement step. The initial step produces a rough demosaicked image containing unwanted color artifacts. The refinement step then reduces these color artifacts using deep residual estimation and multi-model fusion producing a higher quality image. Experimental results show that the proposed method outperforms several existing and state-of-the-art methods in terms of both subjective and objective evaluations.
Daniel Stanley Tan, Wei-Yang Chen, Kai-Lung Hua
IEEE Trans. Image Process.1