Ding-Jie Chen

dblp:123/2959 · DBLP profile ↗
← Back
24ranked-venue papers
12as first author
8since 2021 · last 2023
0000-0001-7649-7824ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 10 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2023 One-Shot Action Detection via Attention Zooming In
abstract
Hinted by a modest support set, few-shot action detection (FSAD) aims at localizing the action instances of unseen classes within an untrimmed query video. Existing FSAD techniques mostly rely on generating a set of class-agnostic action proposals from the query video and then finding the most plausible ones by assessing their correlation to the support set. Such two-stage approaches are feasible but not efficient, largely due to neglecting the support information in generating the proposals. This work focuses on the one-shot image scenario and introduces the attention zooming in strategy to effectively and progressively carry out support-query cross-attention while generating proposals. The resulting one-stage model yields high-quality action proposals for boosting one-shot action detection (OSAD) performance. Our extensive experiments on the ActivityNet-1.3 and THUMOS-14 datasets demonstrate that the proposed framework can achieve state-of-the-art performance in tack-ling challenging image-based OSAD tasks.
He-Yen Hsieh, Ding-Jie Chen, Cheng-Wei Chang, Tyng-Luh Liu
ICASSP2
2023 Contrastive Feature Decoupling for Weakly-Supervised Disease Detection
Jhih-Ciang Wu, Ding-Jie Chen, Chiou-Shann Fuh
MICCAI (5)2
2023 Aggregating Bilateral Attention for Few-Shot Instance Localization
abstract
Attention filtering under various learning scenarios has proven advantageous in enhancing the performance of many neural network architectures. The mainstream attention mechanism is established upon the non-local block, also known as an essential component of the prominent Transformer networks, to catch long-range correlations. However, such unilateral attention is often hampered by sparse and obscure responses, revealing insufficient dependencies across images/patches, and high computational cost, especially for those employing the multi-head design. To overcome these issues, we introduce a novel mechanism of aggregating bilateral attention (ABA) and validate its usefulness in tackling the task of few-shot instance localization, reflecting the underlying query-support dependency. Specifically, our method facilitates uncovering informative features via assessing: i) an embedding norm for exploring the semantically-related cues; ii) context awareness for correlating the query data and support regions. ABA is then carried out by integrating the affinity relations derived from the two measurements to serve as a lightweight but effective query-support attention mechanism with high localization recall. We evaluate ABA on two localization tasks, namely, few-shot action localization and one-shot object detection. Extensive experiments demonstrate that the proposed ABA achieves superior performances over existing methods.
He-Yen Hsieh, Ding-Jie Chen, Cheng-Wei Chang, Tyng-Luh Liu
WACV2
2022 Self-supervised Sparse Representation for Video Anomaly Detection
Jhih-Ciang Wu, He-Yen Hsieh, Ding-Jie Chen, Chiou-Shann Fuh, Tyng-Luh Liu
ECCV (13)3
2022 Contextual Proposal Network for Action Localization
abstract
This paper investigates the problem of Temporal Action Proposal (TAP) generation, which aims to provide a set of high-quality video segments that potentially contain actions events locating in long untrimmed videos. Based on the goal to distill available contextual information, we introduce a Contextual Proposal Network (CPN) composing of two context-aware mechanisms. The first mechanism, i.e., feature enhancing, integrates the inception-like module with long-range attention to capture the multi-scale temporal contexts for yielding a robust video segment representation. The second mechanism, i.e., boundary scoring, employs the bi-directional recurrent neural networks (RNN) to capture bi-directional temporal contexts that explicitly model actionness, background, and confidence of proposals. While generating and scoring proposals, such bi-directional temporal contexts are helpful to retrieve high-quality proposals of low false positives for covering the video action instances. We conduct experiments on two challenging datasets of ActivityNet-1.3 and THUMOS-14 to demonstrate the effectiveness of the proposed Contextual Proposal Network (CPN). In particular, our method respectively surpasses state-of-the-art TAP methods by 1.54% AUC on ActivityNet-1.3 test split and by 0.61% AR@200 on THUMOS-14 dataset.
He-Yen Hsieh, Ding-Jie Chen, Tyng-Luh Liu
WACV2
2021 Adaptive Image Transformer for One-Shot Object Detection
abstract
One-shot object detection tackles a challenging task that aims at identifying within a target image all object instances of the same class, implied by a query image patch. The main difficulty lies in the situation that the class label of the query patch and its respective examples are not available in the training data. Our main idea leverages the concept of language translation to boost metric-learning-based detection methods. Specifically, we emulate the language translation process to adaptively translate the feature of each object proposal to better correlate the given query feature for discriminating the class-similarity among the proposal-query pairs. To this end, we propose the Adaptive Image Transformer (AIT) module that deploys an attention-based encoder-decoder architecture to simultaneously explore intra-coder and inter-coder (i.e., each proposal-query pair) attention. The adaptive nature of our design turns out to be flexible and effective in addressing the one-shot learning scenario. With the informative attention cues, the proposed model excels in predicting the class-similarity between the target image proposals and the query image patch. Though conceptually simple, our model significantly outperforms a state-of-the-art technique, improving the unseen-class object classification from 63.8 mAP and 22.0 AP50 to 72.2 mAP and 24.3 AP50 on the PASCAL-VOC and MS-COCO benchmark datasets, respectively.
Ding-Jie Chen, He-Yen Hsieh, Tyng-Luh Liu
CVPR1
2021 Learning Unsupervised Metaformer for Anomaly Detection
abstract
Anomaly detection (AD) aims to address the task of classification or localization of image anomalies. This paper addresses two pivotal issues of reconstruction-based approaches to AD in images, namely, model adaptation and reconstruction gap. The former generalizes an AD model to tackling a broad range of object categories, while the latter provides useful clues for localizing abnormal regions. At the core of our method is an unsupervised universal model, termed as Metaformer, which leverages both meta-learned model parameters to achieve high model adaptation capability and instance-aware attention to emphasize the focal regions for localizing abnormal regions, i.e., to explore the reconstruction gap at those regions of interest. We justify the effectiveness of our method with SOTA results on the MVTec AD dataset of industrial images and highlight the adaptation flexibility of the universal Metaformer with multi-class and few-shot scenarios.
Jhih-Ciang Wu, Ding-Jie Chen, Chiou-Shann Fuh, Tyng-Luh Liu
ICCV2
2021 Referring Image Segmentation via Language-Driven Attention
abstract
This paper aims to tackle the problem of referring image segmentation, which is targeted at reasoning the region of interest referred by a query natural language sentence. One key issue to address the referring image segmentation is how to establish the cross-modal representation for encoding the two modalities, namely, the query sentence and the input image. Most existing methods are designed to concatenate the features from each modality or to gradually encode the cross-modal representation concerning each word’s effect. In contrast, our approach leverages the correlation between the two modalities for constructing the cross-modal representation. To make the resulting cross-modal representation more discriminative for the segmentation task, we propose a novel mechanism of language-driven attention to encode the cross-modal representation for reflecting the attention between every single visual element and the entire query sentence. The proposed mechanism, named as Language-Driven Attention (LDA), first decouples the cross-modal correlation to channel-attention and spatial-attention and then integrates the two attentions for obtaining the cross-modal representation. The channel attention and the spatial attention respectively reveal how sensitive each channel or each pixel of a particular feature map is with respect to the query sentence. With a proper fusion of the two kinds of feature attention, the proposed LDA model can effectively guide the generation of the final cross-modal representation. The resulting representation is further strengthened for capturing the multi-receptive-field and multi-level-semantic for the intended segmentation. We assess our referring image segmentation model on four public benchmark datasets, and the experimental results show that our model achieves state-of-the-art performance
Ding-Jie Chen, He-Yen Hsieh, Tyng-Luh Liu
ICRA1
2020 Temporal Action Proposal Generation Via Deep Feature Enhancement
abstract
Temporal action proposal generation (TAPG) is a challenging problem for analyzing video content. It aims to localize the video segments which are likely to contain actions or events. Intuitively, making a satisfying prediction of these video segments is directly relies on their representation quality. A typical representation of a video segment is applying a two-stream feature, which comprises appearance and motion information. Rather than directly concatenating the two-stream features as the previous methods, we illustrate a feature-aggregation network (FA-Net) concerning the feature-relation among neighboring video segments for obtaining the high-quality representation that better characterizing the actions or events. Further, we design a feature-expansion network (FE-Net) to extract multi-granularity features for retrieving the proposals of high action-instance covering confidence. We evaluate our approach on two challenging datasets: ActivityNet-1.3 and THUMOS-14. The experiments showed that the proposed approach consistently outperforms the existing state-of-the-art TAPG methods.
He-Yen Hsieh, Ding-Jie Chen, Tyng-Luh Liu
ICIP2
2020 SwipeCut: Interactive Segmentation via Seed Grouping
abstract
Interactive image segmentation algorithms rely on the user to provide annotations as the guidance. When the task of interactive segmentation is performed on a small touchscreen device, the requirement of providing precise annotations could be cumbersome to the user. We design a new interaction mechanism that actively queries seeds for guiding the user to label. Our method enforces sparsity and diversity criteria on the selection of query seeds, and at each round of interaction, the user is only presented with a small number of informative query seeds that are far apart from each other. Therefore, the user merely has to swipe through on the ROI-relevant query seeds for checking which ones of the query seeds are inside the region of interest (ROI). This kind of interaction should be easy since those gestures are commonly used on a touchscreen. As a result, we are able to derive a user-friendly interaction mechanism for annotation on small touchscreen devices. The performance of our algorithm is evaluated on six publicly available datasets. The evaluation results show that our algorithm achieves high segmentation accuracy, with short computational time and less user feedback.
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
IEEE Trans. Circuits Syst. Video Technol.1
2019 Unsupervised Meta-Learning of Figure-Ground Segmentation via Imitating Visual Effects
Ding-Jie Chen, Jui-Ting Chien, Hwann-Tzong Chen, Tyng-Luh Liu
AAAI1
2019 Instance-Level Meta Normalization
abstract
This paper presents a normalization mechanism called Instance-Level Meta Normalization (ILM Norm) to address a learning-to-normalize problem. ILM~Norm learns to predict the normalization parameters via both the feature feed-forward and the gradient back-propagation paths. ILM Norm provides a meta normalization mechanism and has several good properties. It can be easily plugged into existing instance-level normalization schemes such as Instance Normalization, Layer Normalization, or Group Normalization. ILM~Norm normalizes each instance individually and therefore maintains high performance even when small mini-batch is used. The experimental results show that ILM~Norm well adapts to different network architectures and tasks, and it consistently improves the performance of the original models.
Songhao Jia, Ding-Jie Chen, Hwann-Tzong Chen
CVPR2
2019 See-Through-Text Grouping for Referring Image Segmentation
abstract
Motivated by the conventional grouping techniques to image segmentation, we develop their DNN counterpart to tackle the referring variant. The proposed method is driven by a convolutional-recurrent neural network (ConvRNN) that iteratively carries out top-down processing of bottom-up segmentation cues. Given a natural language referring expression, our method learns to predict its relevance to each pixel and derives a See-through-Text Embedding Pixelwise (STEP) heatmap, which reveals segmentation cues of pixel level via the learned visual-textual co-embedding. The ConvRNN performs a top-down approximation by converting the STEP heatmap into a refined one, whereas the improvement is expected from training the network with a classification loss from the ground truth. With the refined heatmap, we update the textual representation of the referring expression by re-evaluating its attention distribution and then compute a new STEP heatmap as the next input to the ConvRNN. Boosting by such collaborative learning, the framework can progressively and simultaneously yield the desired referring segmentation and reasonable attention distribution over the referring sentence. Our method is general and does not rely on, say, the outcomes of object detection from other DNN models, while achieving state-of-the-art performance in all of the four datasets in the experiments.
Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen, Tyng-Luh Liu
ICCV1
2018 Tap and Shoot Segmentation
abstract
We present a new segmentation method that leverages latent photographic information available at the moment of taking pictures. Photography on a portable device is often done by tapping to focus before shooting the picture. This tap-and-shoot interaction for photography not only specifies the region of interest but also yields useful focus/defocus cues for image segmentation. However, most of the previous interactive segmentation methods address the problem of image segmentation in a post-processing scenario without considering the action of taking pictures. We propose a learning-based approach to this new tap-and-shoot scenario of interactive segmentation. The experimental results on various datasets show that, by training a deep convolutional network to integrate the selection and focus/defocus cues, our method can achieve higher segmentation accuracy in comparison with existing interactive segmentation methods.
Ding-Jie Chen, Jui-Ting Chien, Hwann-Tzong Chen, Long-Wen Chang
AAAI1
2018 A2A: Attention to Attention Reasoning for Movie Question Answering
Chao-Ning Liu, Ding-Jie Chen, Hwann-Tzong Chen, Tyng-Luh Liu
ACCV (6)2
2018 Toward a unified scheme for fast interactive segmentation
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
J. Vis. Commun. Image Represent.1
2018 Interactive 1-bit feedback segmentation using transductive inference
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
Mach. Vis. Appl.1
2017 Video segmentation via boundary-aware flow
abstract
We present a new algorithm for unsupervised video segmentation based on boundary-aware optical flow. Existing video segmentation methods usually tweak their segmentation model to tolerate the inaccuracy in the estimation of optical flow around object boundaries. In contrast, we directly manipulate the optical flow for better quality. We smooth the optical flow via transductive inference to make the flow consistent within the object and fit to the object boundaries. We then use the boundary-aware optical flow to estimate the initial foreground object region from each frame for learning the appearance model. The learned appearance model is consequently used to refine the segmentation result. Experiments on the DAVIS dataset show that our method performs favorably against the existing ones.
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
ICIP1
2016 Interactive Segmentation from 1-Bit Feedback
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
ACCV (1)1
2016 Fast defocus map estimation
abstract
This paper presents a fast algorithm for deriving the defocus map from a single image. Existing methods of defocus map estimation often include a pixel-level propagation step to spread the measured sparse defocus cues over the whole image. Since the pixel-level propagation step is time-consuming, we develop an effective method to obtain the whole-image defocus blur using oversegmentation and transductive inference. Oversegmentation produces the superpixels and hence greatly reduces the computation costs for subsequent procedures. Transductive inference provides a way to calculate the similarity between superpixels, and thus helps to infer the defocus blur of each superpixel from all other superpixels. The experimental results show that our method is efficient and able to estimate a plausible superpixel-level defocus map from a given single image.
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
ICIP1
2014 Unsupervised Image Co-segmentation Based on Cooperative Game
Bo-Chen Lin, Ding-Jie Chen, Long-Wen Chang
ACCV (3)2
2014 Modified soft-decision adaptive interpolation by an evolutionary game
abstract
Soft-decision adaptive interpolation (SAI) provides a powerful result in preserving edge structures for interpolation from low-resolution images to obtain high-resolution images. However, the SAI algorithm may produce artifacts in the smooth regions and texture patterns of the interpolated images. To improve the SAI algorithm, we propose an evolutionary game, where every pixel is regarded as a player. The players are divided into two roles, one for unknown high-resolution pixels and the other for known low-resolution pixels. Each role has a different strategy set. The evolutionarily stable strategy in each local image region is a mixed strategy adopted by every player. By considering the mixed strategy as weights for interpolation, we can adaptively estimate the high-resolution image. Experimental results show that the proposed algorithm improves the SAI algorithm by alleviating the artifacts both in PSNR and visual quality.
Pei-Chi Hsiao, Ding-Jie Chen, Long-Wen Chang
ICIP2
2012 Video object cosegmentation
abstract
We introduce and address the problem of video object cosegmentation, which concerns the task of segmenting the common object in a pair of video sequences. We present a new algorithm that works on super-voxels in videos to solve this task. The algorithm computes i the intra-video relative motion derived from dense optical flow and ii) the inter-video co-features based on Gaussian mixture models. The experimental results show that, by integrating the intra-video and inter-video information, our algorithm is able to obtain better results of segmenting video objects.
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
ACM Multimedia1
2008 An RST-resilient image copyright protection scheme based on the invariant domain and image secret sharing
abstract
This paper proposes a novel scheme that can protect the copyrights of images based on the RST-invariant domain and Image Secret Sharing (ISS) scheme. The proposed scheme aims at resisting the rotation, scaling, and translation (RST) of the geometric distortions. The scheme first extracts the features from the RST-invariant domain of the host image. It then utilizes the ISS scheme to encode a binary logo according to the extracted features. Finally, it generates an image for authentication, which can later be used to reconstruct the logo for copyright verification when disputes over the image ownership occur. The experimental results show that the proposed scheme can achieve a satisfactory accuracy rate for images suffering RST distortions.
Shang-Lin Hsieh, Ding-Jie Chen, Bin-Yuan Huang, I-Ju Tsai
SMC2