EDBT 2026 Demo / reviewers in the wild / expert
Ding-Jie Chen
dblp:123/2959
· DBLP profile ↗
24ranked-venue papers
12as first author
8since 2021 · last 2023
0000-0001-7649-7824ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 10 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | One-Shot Action Detection via Attention Zooming InabstractHinted by a modest support set, few-shot action detection (FSAD) aims at localizing the action instances of unseen classes within an untrimmed query video. Existing FSAD techniques mostly rely on generating a set of class-agnostic action proposals from the query video and then finding the most plausible ones by assessing their correlation to the support set. Such two-stage approaches are feasible but not efficient, largely due to neglecting the support information in generating the proposals. This work focuses on the one-shot image scenario and introduces the attention zooming in strategy to effectively and progressively carry out support-query cross-attention while generating proposals. The resulting one-stage model yields high-quality action proposals for boosting one-shot action detection (OSAD) performance. Our extensive experiments on the ActivityNet-1.3 and THUMOS-14 datasets demonstrate that the proposed framework can achieve state-of-the-art performance in tack-ling challenging image-based OSAD tasks. He-Yen Hsieh, Ding-Jie Chen, Cheng-Wei Chang, Tyng-Luh Liu |
ICASSP | 2 |
| 2023 | Contrastive Feature Decoupling for Weakly-Supervised Disease Detection
Jhih-Ciang Wu, Ding-Jie Chen, Chiou-Shann Fuh |
MICCAI (5) | 2 |
| 2023 | Aggregating Bilateral Attention for Few-Shot Instance LocalizationabstractAttention filtering under various learning scenarios has proven advantageous in enhancing the performance of many neural network architectures. The mainstream attention mechanism is established upon the non-local block, also known as an essential component of the prominent Transformer networks, to catch long-range correlations. However, such unilateral attention is often hampered by sparse and obscure responses, revealing insufficient dependencies across images/patches, and high computational cost, especially for those employing the multi-head design. To overcome these issues, we introduce a novel mechanism of aggregating bilateral attention (ABA) and validate its usefulness in tackling the task of few-shot instance localization, reflecting the underlying query-support dependency. Specifically, our method facilitates uncovering informative features via assessing: i) an embedding norm for exploring the semantically-related cues; ii) context awareness for correlating the query data and support regions. ABA is then carried out by integrating the affinity relations derived from the two measurements to serve as a lightweight but effective query-support attention mechanism with high localization recall. We evaluate ABA on two localization tasks, namely, few-shot action localization and one-shot object detection. Extensive experiments demonstrate that the proposed ABA achieves superior performances over existing methods. He-Yen Hsieh, Ding-Jie Chen, Cheng-Wei Chang, Tyng-Luh Liu |
WACV | 2 |
| 2022 | Self-supervised Sparse Representation for Video Anomaly Detection
Jhih-Ciang Wu, He-Yen Hsieh, Ding-Jie Chen, Chiou-Shann Fuh, Tyng-Luh Liu |
ECCV (13) | 3 |
| 2022 | Contextual Proposal Network for Action LocalizationabstractThis paper investigates the problem of Temporal Action Proposal (TAP) generation, which aims to provide a set of high-quality video segments that potentially contain actions events locating in long untrimmed videos. Based on the goal to distill available contextual information, we introduce a Contextual Proposal Network (CPN) composing of two context-aware mechanisms. The first mechanism, i.e., feature enhancing, integrates the inception-like module with long-range attention to capture the multi-scale temporal contexts for yielding a robust video segment representation. The second mechanism, i.e., boundary scoring, employs the bi-directional recurrent neural networks (RNN) to capture bi-directional temporal contexts that explicitly model actionness, background, and confidence of proposals. While generating and scoring proposals, such bi-directional temporal contexts are helpful to retrieve high-quality proposals of low false positives for covering the video action instances. We conduct experiments on two challenging datasets of ActivityNet-1.3 and THUMOS-14 to demonstrate the effectiveness of the proposed Contextual Proposal Network (CPN). In particular, our method respectively surpasses state-of-the-art TAP methods by 1.54% AUC on ActivityNet-1.3 test split and by 0.61% AR@200 on THUMOS-14 dataset. He-Yen Hsieh, Ding-Jie Chen, Tyng-Luh Liu |
WACV | 2 |
| 2021 | Adaptive Image Transformer for One-Shot Object DetectionabstractOne-shot object detection tackles a challenging task that aims at identifying within a target image all object instances of the same class, implied by a query image patch. The main difficulty lies in the situation that the class label of the query patch and its respective examples are not available in the training data. Our main idea leverages the concept of language translation to boost metric-learning-based detection methods. Specifically, we emulate the language translation process to adaptively translate the feature of each object proposal to better correlate the given query feature for discriminating the class-similarity among the proposal-query pairs. To this end, we propose the Adaptive Image Transformer (AIT) module that deploys an attention-based encoder-decoder architecture to simultaneously explore intra-coder and inter-coder (i.e., each proposal-query pair) attention. The adaptive nature of our design turns out to be flexible and effective in addressing the one-shot learning scenario. With the informative attention cues, the proposed model excels in predicting the class-similarity between the target image proposals and the query image patch. Though conceptually simple, our model significantly outperforms a state-of-the-art technique, improving the unseen-class object classification from 63.8 mAP and 22.0 AP50 to 72.2 mAP and 24.3 AP50 on the PASCAL-VOC and MS-COCO benchmark datasets, respectively. Ding-Jie Chen, He-Yen Hsieh, Tyng-Luh Liu |
CVPR | 1 |
| 2021 | Learning Unsupervised Metaformer for Anomaly DetectionabstractAnomaly detection (AD) aims to address the task of classification or localization of image anomalies. This paper addresses two pivotal issues of reconstruction-based approaches to AD in images, namely, model adaptation and reconstruction gap. The former generalizes an AD model to tackling a broad range of object categories, while the latter provides useful clues for localizing abnormal regions. At the core of our method is an unsupervised universal model, termed as Metaformer, which leverages both meta-learned model parameters to achieve high model adaptation capability and instance-aware attention to emphasize the focal regions for localizing abnormal regions, i.e., to explore the reconstruction gap at those regions of interest. We justify the effectiveness of our method with SOTA results on the MVTec AD dataset of industrial images and highlight the adaptation flexibility of the universal Metaformer with multi-class and few-shot scenarios. Jhih-Ciang Wu, Ding-Jie Chen, Chiou-Shann Fuh, Tyng-Luh Liu |
ICCV | 2 |
| 2021 | Referring Image Segmentation via Language-Driven AttentionabstractThis paper aims to tackle the problem of referring image segmentation, which is targeted at reasoning the region of interest referred by a query natural language sentence. One key issue to address the referring image segmentation is how to establish the cross-modal representation for encoding the two modalities, namely, the query sentence and the input image. Most existing methods are designed to concatenate the features from each modality or to gradually encode the cross-modal representation concerning each word’s effect. In contrast, our approach leverages the correlation between the two modalities for constructing the cross-modal representation. To make the resulting cross-modal representation more discriminative for the segmentation task, we propose a novel mechanism of language-driven attention to encode the cross-modal representation for reflecting the attention between every single visual element and the entire query sentence. The proposed mechanism, named as Language-Driven Attention (LDA), first decouples the cross-modal correlation to channel-attention and spatial-attention and then integrates the two attentions for obtaining the cross-modal representation. The channel attention and the spatial attention respectively reveal how sensitive each channel or each pixel of a particular feature map is with respect to the query sentence. With a proper fusion of the two kinds of feature attention, the proposed LDA model can effectively guide the generation of the final cross-modal representation. The resulting representation is further strengthened for capturing the multi-receptive-field and multi-level-semantic for the intended segmentation. We assess our referring image segmentation model on four public benchmark datasets, and the experimental results show that our model achieves state-of-the-art performance Ding-Jie Chen, He-Yen Hsieh, Tyng-Luh Liu |
ICRA | 1 |
| 2020 | Temporal Action Proposal Generation Via Deep Feature EnhancementabstractTemporal action proposal generation (TAPG) is a challenging problem for analyzing video content. It aims to localize the video segments which are likely to contain actions or events. Intuitively, making a satisfying prediction of these video segments is directly relies on their representation quality. A typical representation of a video segment is applying a two-stream feature, which comprises appearance and motion information. Rather than directly concatenating the two-stream features as the previous methods, we illustrate a feature-aggregation network (FA-Net) concerning the feature-relation among neighboring video segments for obtaining the high-quality representation that better characterizing the actions or events. Further, we design a feature-expansion network (FE-Net) to extract multi-granularity features for retrieving the proposals of high action-instance covering confidence. We evaluate our approach on two challenging datasets: ActivityNet-1.3 and THUMOS-14. The experiments showed that the proposed approach consistently outperforms the existing state-of-the-art TAPG methods. He-Yen Hsieh, Ding-Jie Chen, Tyng-Luh Liu |
ICIP | 2 |
| 2020 | SwipeCut: Interactive Segmentation via Seed GroupingabstractInteractive image segmentation algorithms rely on the user to provide annotations as the guidance. When the task of interactive segmentation is performed on a small touchscreen device, the requirement of providing precise annotations could be cumbersome to the user. We design a new interaction mechanism that actively queries seeds for guiding the user to label. Our method enforces sparsity and diversity criteria on the selection of query seeds, and at each round of interaction, the user is only presented with a small number of informative query seeds that are far apart from each other. Therefore, the user merely has to swipe through on the ROI-relevant query seeds for checking which ones of the query seeds are inside the region of interest (ROI). This kind of interaction should be easy since those gestures are commonly used on a touchscreen. As a result, we are able to derive a user-friendly interaction mechanism for annotation on small touchscreen devices. The performance of our algorithm is evaluated on six publicly available datasets. The evaluation results show that our algorithm achieves high segmentation accuracy, with short computational time and less user feedback. Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Unsupervised Meta-Learning of Figure-Ground Segmentation via Imitating Visual Effects
Ding-Jie Chen, Jui-Ting Chien, Hwann-Tzong Chen, Tyng-Luh Liu |
AAAI | 1 |
| 2019 | Instance-Level Meta NormalizationabstractThis paper presents a normalization mechanism called Instance-Level Meta Normalization (ILM Norm) to address a learning-to-normalize problem. ILM~Norm learns to predict the normalization parameters via both the feature feed-forward and the gradient back-propagation paths. ILM Norm provides a meta normalization mechanism and has several good properties. It can be easily plugged into existing instance-level normalization schemes such as Instance Normalization, Layer Normalization, or Group Normalization. ILM~Norm normalizes each instance individually and therefore maintains high performance even when small mini-batch is used. The experimental results show that ILM~Norm well adapts to different network architectures and tasks, and it consistently improves the performance of the original models. Songhao Jia, Ding-Jie Chen, Hwann-Tzong Chen |
CVPR | 2 |
| 2019 | See-Through-Text Grouping for Referring Image SegmentationabstractMotivated by the conventional grouping techniques to image segmentation, we develop their DNN counterpart to tackle the referring variant. The proposed method is driven by a convolutional-recurrent neural network (ConvRNN) that iteratively carries out top-down processing of bottom-up segmentation cues. Given a natural language referring expression, our method learns to predict its relevance to each pixel and derives a See-through-Text Embedding Pixelwise (STEP) heatmap, which reveals segmentation cues of pixel level via the learned visual-textual co-embedding. The ConvRNN performs a top-down approximation by converting the STEP heatmap into a refined one, whereas the improvement is expected from training the network with a classification loss from the ground truth. With the refined heatmap, we update the textual representation of the referring expression by re-evaluating its attention distribution and then compute a new STEP heatmap as the next input to the ConvRNN. Boosting by such collaborative learning, the framework can progressively and simultaneously yield the desired referring segmentation and reasonable attention distribution over the referring sentence. Our method is general and does not rely on, say, the outcomes of object detection from other DNN models, while achieving state-of-the-art performance in all of the four datasets in the experiments. Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen, Tyng-Luh Liu |
ICCV | 1 |
| 2018 | Tap and Shoot SegmentationabstractWe present a new segmentation method that leverages latent photographic information available at the moment of taking pictures. Photography on a portable device is often done by tapping to focus before shooting the picture. This tap-and-shoot interaction for photography not only specifies the region of interest but also yields useful focus/defocus cues for image segmentation. However, most of the previous interactive segmentation methods address the problem of image segmentation in a post-processing scenario without considering the action of taking pictures. We propose a learning-based approach to this new tap-and-shoot scenario of interactive segmentation. The experimental results on various datasets show that, by training a deep convolutional network to integrate the selection and focus/defocus cues, our method can achieve higher segmentation accuracy in comparison with existing interactive segmentation methods. Ding-Jie Chen, Jui-Ting Chien, Hwann-Tzong Chen, Long-Wen Chang |
AAAI | 1 |
| 2018 | A2A: Attention to Attention Reasoning for Movie Question Answering
Chao-Ning Liu, Ding-Jie Chen, Hwann-Tzong Chen, Tyng-Luh Liu |
ACCV (6) | 2 |
| 2018 | Toward a unified scheme for fast interactive segmentation
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Interactive 1-bit feedback segmentation using transductive inference
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang |
Mach. Vis. Appl. | 1 |
| 2017 | Video segmentation via boundary-aware flowabstractWe present a new algorithm for unsupervised video segmentation based on boundary-aware optical flow. Existing video segmentation methods usually tweak their segmentation model to tolerate the inaccuracy in the estimation of optical flow around object boundaries. In contrast, we directly manipulate the optical flow for better quality. We smooth the optical flow via transductive inference to make the flow consistent within the object and fit to the object boundaries. We then use the boundary-aware optical flow to estimate the initial foreground object region from each frame for learning the appearance model. The learned appearance model is consequently used to refine the segmentation result. Experiments on the DAVIS dataset show that our method performs favorably against the existing ones. Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang |
ICIP | 1 |
| 2016 | Interactive Segmentation from 1-Bit Feedback
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang |
ACCV (1) | 1 |
| 2016 | Fast defocus map estimationabstractThis paper presents a fast algorithm for deriving the defocus map from a single image. Existing methods of defocus map estimation often include a pixel-level propagation step to spread the measured sparse defocus cues over the whole image. Since the pixel-level propagation step is time-consuming, we develop an effective method to obtain the whole-image defocus blur using oversegmentation and transductive inference. Oversegmentation produces the superpixels and hence greatly reduces the computation costs for subsequent procedures. Transductive inference provides a way to calculate the similarity between superpixels, and thus helps to infer the defocus blur of each superpixel from all other superpixels. The experimental results show that our method is efficient and able to estimate a plausible superpixel-level defocus map from a given single image. Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang |
ICIP | 1 |
| 2014 | Unsupervised Image Co-segmentation Based on Cooperative Game
Bo-Chen Lin, Ding-Jie Chen, Long-Wen Chang |
ACCV (3) | 2 |
| 2014 | Modified soft-decision adaptive interpolation by an evolutionary gameabstractSoft-decision adaptive interpolation (SAI) provides a powerful result in preserving edge structures for interpolation from low-resolution images to obtain high-resolution images. However, the SAI algorithm may produce artifacts in the smooth regions and texture patterns of the interpolated images. To improve the SAI algorithm, we propose an evolutionary game, where every pixel is regarded as a player. The players are divided into two roles, one for unknown high-resolution pixels and the other for known low-resolution pixels. Each role has a different strategy set. The evolutionarily stable strategy in each local image region is a mixed strategy adopted by every player. By considering the mixed strategy as weights for interpolation, we can adaptively estimate the high-resolution image. Experimental results show that the proposed algorithm improves the SAI algorithm by alleviating the artifacts both in PSNR and visual quality. Pei-Chi Hsiao, Ding-Jie Chen, Long-Wen Chang |
ICIP | 2 |
| 2012 | Video object cosegmentationabstractWe introduce and address the problem of video object cosegmentation, which concerns the task of segmenting the common object in a pair of video sequences. We present a new algorithm that works on super-voxels in videos to solve this task. The algorithm computes i the intra-video relative motion derived from dense optical flow and ii) the inter-video co-features based on Gaussian mixture models. The experimental results show that, by integrating the intra-video and inter-video information, our algorithm is able to obtain better results of segmenting video objects. Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang |
ACM Multimedia | 1 |
| 2008 | An RST-resilient image copyright protection scheme based on the invariant domain and image secret sharingabstractThis paper proposes a novel scheme that can protect the copyrights of images based on the RST-invariant domain and Image Secret Sharing (ISS) scheme. The proposed scheme aims at resisting the rotation, scaling, and translation (RST) of the geometric distortions. The scheme first extracts the features from the RST-invariant domain of the host image. It then utilizes the ISS scheme to encode a binary logo according to the extracted features. Finally, it generates an image for authentication, which can later be used to reconstruct the logo for copyright verification when disputes over the image ownership occur. The experimental results show that the proposed scheme can achieve a satisfactory accuracy rate for images suffering RST distortions. Shang-Lin Hsieh, Ding-Jie Chen, Bin-Yuan Huang, I-Ju Tsai |
SMC | 2 |