Tsu-Jui Fu

dblp:218/5366 · DBLP profile ↗
← Back
31ranked-venue papers
13as first author
21since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 12 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 first-author · 8 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
Tsu-Jui Fu, Yusu Qian, Wenze Hu, Zhe Gan, Yinfei Yang
ICCV1
2025 STIV: Scalable Text and Image Conditioned Video Generation
abstract
The field of video generation has made remarkable advancements, yet there remains a pressing need for a clear, systematic recipe that can guide the development of robust and scalable models. In this work, we present a comprehensive study that systematically explores the interplay of model architectures, training recipes, and data curation strategies, culminating in a simple and scalable text-image-conditioned video generation method, named STIV. Our framework integrates image condition into a Diffusion Transformer (DiT) through frame replacement, while incorporating text conditioning via a joint image-text conditional classifier-free guidance. This design enables STIV to perform both text-to-video (T2V) and text-image-to-video (TI2V) tasks simultaneously. Additionally, STIV can be easily extended to various applications, such as video prediction, frame interpolation, multi-view generation, and long video generation, etc. With comprehensive ablation studies on T2I, T2V, and TI2V, STIV demonstrate strong performance, despite its simple design. An 8.7B model with 512 resolution achieves 83.1 on VBench T2V, surpassing both leading open and closed-source models like CogVideoX-5B, Pika, Kling, and Gen-3. The same-sized model also achieves a state-of-the-art result of 90.1 on VBench I2V task at 512 resolution. By providing a transparent and extensible recipe for building cutting-edge video generation models, we aim to empower future research and accelerate progress toward more versatile and reliable video generation solutions.
Zongyu Lin, Chen Chen 0005, Jiasen Lu, Wenze Hu, Tsu-Jui Fu, Jesse Allardice, Zhengfeng Lai, Liangchen Song, Bowen Zhang 0002, Cha Chen, Yiran Fei, Lezhi Li, Yinfei Yang, Yizhou Sun, Kai-Wei Chang 0001
ICCV6
2025 CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching
abstract
Conditional generative modeling aims to learn a conditional data distribution from samples containing data-condition pairs. For this, diffusion and flow-based methods have attained compelling results. These methods use a learned (flow) model to transport an initial standard Gaussian noise that ignores the condition to the conditional data distribution. The model is hence required to learn both mass transport \emph{and} conditional injection. To ease the demand on the model, we propose \emph{Condition-Aware Reparameterization for Flow Matching} (CAR-Flow) -- a lightweight, learned \emph{shift} that conditions the source, the target, or both distributions. By relocating these distributions, CAR-Flow shortens the probability path the model must learn, leading to faster training in practice. On low-dimensional synthetic data, we visualize and quantify the effects of CAR-Flow. On higher-dimensional natural image data (ImageNet-256), equipping SiT-XL/2 with CAR-Flow reduces FID from 2.07 to 1.68, while introducing less than \(0.6\%\) additional parameters.
Chen Chen 0005, Pengsheng Guo, Liangchen Song, Jiasen Lu, Rui Qian 0003, Tsu-Jui Fu, Xinze Wang, Yinfei Yang, Alex Schwing 0002
NeurIPS6
2024 VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
abstract
Incremental decision making in real-world environments is one of the most challenging tasks in embodied artificial intelligence. One particularly demanding scenario is Vision and Language Navigation (VLN) which requires visual and natural language understanding as well as spatial and temporal reasoning capabilities. The embodied agent needs to ground its understanding of navigation instructions in observations of a real-world environment like Street View. Despite the impressive results of LLMs in other research areas, it is an ongoing problem of how to best connect them with an interactive visual environment. In this work, we propose VELMA, an embodied LLM agent that uses a verbalization of the trajectory and of visual environment observations as contextual prompt for the next action. Visual information is verbalized by a pipeline that extracts landmarks from the human written navigation instructions and uses CLIP to determine their visibility in the current panorama view. We show that VELMA is able to successfully follow navigation instructions in Street View with only two in-context examples. We further finetune the LLM agent on a few thousand examples and achieve around 25% relative improvement in task completion over the previous state-of-the-art for two datasets.
Raphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu, Stefan Riezler, William Yang Wang
AAAI4
2024 Guiding Instruction-based Image Editing via Multimodal Large Language Models
abstract
Instruction-based image editing improves the controllability and flexibility of image manipulation via natural commands without elaborate descriptions or regional masks. However, human instructions are sometimes too brief for current methods to capture and follow. Multimodal large language models (MLLMs) show promising capabilities in cross-modal understanding and visual-aware response generation via LMs. We investigate how MLLMs facilitate edit instructions and present MLLM-Guided Image Editing (MGIE). MGIE learns to derive expressive instructions and provides explicit guidance. The editing model jointly captures this visual imagination and performs manipulation through end-to-end training. We evaluate various aspects of Photoshop-style modification, global photo optimization, and local editing. Extensive experimental results demonstrate that expressive instructions are crucial to instruction-based image editing, and our MGIE can lead to a notable improvement in automatic metrics and human evaluation while maintaining competitive inference efficiency.
Tsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang, Yinfei Yang, Zhe Gan
ICLR1
2024 T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback
abstract
Diffusion-based text-to-video (T2V) models have achieved significant success but continue to be hampered by the slow sampling speed of their iterative sampling processes. To address the challenge, consistency models have been proposed to facilitate fast inference, albeit at the cost of sample quality. In this work, we aim to break the quality bottleneck of a video consistency model (VCM) to achieve **both fast and high-quality video generation**. We introduce T2V-Turbo, which integrates feedback from a mixture of differentiable reward models into the consistency distillation (CD) process of a pre-trained T2V model. Notably, we directly optimize rewards associated with single-step generations that arise naturally from computing the CD loss, effectively bypassing the memory constraints imposed by backpropagating gradients through an iterative sampling process. Remarkably, the 4-step generations from our T2V-Turbo achieve the highest total score on VBench, even surpassing Gen-2 and Pika. We further conduct human evaluations to corroborate the results, validating that the 4-step generations from our T2V-Turbo are preferred over the 50-step DDIM samples from their teacher models, representing more than a tenfold acceleration while improving video generation quality.
Weixi Feng, Tsu-Jui Fu, Xinyi Wang 0003, Sugato Basu, Wenhu Chen, William Yang Wang
NeurIPS3
2023 An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling
abstract
Masked visual modeling (MVM) has been recently proven effective for visual pre-training. While similar reconstructive objectives on video inputs (e.g., masked frame modeling) have been explored in video-language (VidL) pre-training, previous studies fail to find a truly effective MVM strategy that can largely benefit the downstream performance. In this work, we systematically examine the potential of MVM in the context of VidL learning. Specifically, we base our study on a fully end-to-end VIdeO-LanguagE Transformer (VIOLET) [15], where the supervision from MVM training can be backpropogated to the video pixel space. In total, eight different reconstructive targets of MVM are explored, from low-level pixel values and oriented gradients to high-level depth maps, optical flow, discrete visual tokens and latent visual features. We conduct comprehensive experiments and provide insights into the factors leading to effective MVM training, resulting in an enhanced model VIOLETv2. Empirically, we show VIOLETv2 pre-trained with MVM objective achieves notable improvements on 13 VidL benchmarks, ranging from video question answering, video captioning, to text-to-video retrieval.11Code has been released at https://github.com/tsujuifu/pytorch_empirical-mvm
Tsu-Jui Fu, Zhe Gan, William Yang Wang, Zicheng Liu 0001
CVPR1
2023 Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation
abstract
Generating a video given the first several static frames is challenging as it anticipates reasonable future frames with temporal coherence. Besides video prediction, the ability to rewind from the last frame or infilling between the head and tail is also crucial, but they have rarely been explored for video completion. Since there could be different outcomes from the hints of just a few frames, a system that can follow natural language to perform video completion may significantly improve controllability. Inspired by this, we introduce a novel task, text-guided video completion (TVC), which requests the model to generate a video from partial frames guided by an instruction. We then propose Multimodal Masked Video Generation (MMVG) to address this TVC task. During training, MMVG discretizes the video frames into visual tokens and masks most of them to perform video completion from any time point. At inference time, a single MMVG model can address all 3 cases of TVC, including video prediction, rewind, and infilling, by applying corresponding masking conditions. We evaluate MMVG in various video scenarios, including egocentric, animation, and gaming. Extensive experimental results indicate that MMVG is effective in generating high-quality visual appearances with text guidance for TVC.
Tsu-Jui Fu, Licheng Yu, Ning Zhang 0014, Cheng-Yang Fu, Jong-Chyi Su, William Yang Wang, Sean Bell
CVPR1
2023 EDIS: Entity-Driven Image Search over Multimodal Web Content
abstract
Making image retrieval methods practical for real-world search applications requires significant progress in dataset scales, entity comprehension, and multimodal information fusion.In this work, we introduce Entity-Driven Image Search (EDIS), a challenging dataset for cross-modal image search in the news domain.EDIS consists of 1 million web images from actual search engine results and curated datasets, with each image paired with a textual description.Unlike datasets that assume a small set of single-modality candidates, EDIS reflects realworld web image search scenarios by including a million multimodal image-text pairs as candidates.EDIS encourages the development of retrieval models that simultaneously address cross-modal information fusion and matching.To achieve accurate ranking results, a model must: 1) understand named entities and events from text queries, 2) ground entities onto images or text descriptions, and 3) effectively fuse textual and visual representations.Our experimental results show that EDIS challenges stateof-the-art methods with dense entities and the large-scale candidate set.The ablation study also proves that fusing textual features with visual features is critical in improving retrieval results.
Weixi Feng, Tsu-Jui Fu, Wenhu Chen, William Yang Wang
EMNLP3
2023 Collaborative Generative AI: Integrating GPT-k for Efficient Editing in Text-to-Image Generation
abstract
The field of text-to-image (T2I) generation has garnered significant attention both within the research community and among everyday users.Despite the advancements of T2I models, a common issue encountered by users is the need for repetitive editing of input prompts in order to receive a satisfactory image, which is time-consuming and labor-intensive.Given the demonstrated text generation power of largescale language models, such as GPT-k, we investigate the potential of utilizing such models to improve the prompt editing process for T2I generation.We conduct a series of experiments to compare the common edits made by humans and GPT-k, evaluate the performance of GPT-k in prompting T2I, and examine factors that may influence this process.We found that GPT-k models focus more on inserting modifiers while humans tend to replace words and phrases, which includes changes to the subject matter.Experimental results show that GPT-k are more effective in adjusting modifiers rather than predicting spontaneous changes in the primary subject matters.Adopting the edit suggested by GPT-k models may reduce the percentage of remaining edits by 20-30%. 1 Our experiments are conducted upon StableDiffusion since it is a wide-adopted open-source large text-to-image generative model with SoTA performance.
Wanrong Zhu, Xinyi Wang 0003, Tsu-Jui Fu, Xin Wang 0061, Miguel P. Eckstein, William Yang Wang
EMNLP4
2023 Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Arjun R. Akula, Pradyumna Narayana, Sugato Basu, Xin Wang 0061, William Yang Wang
ICLR3
2023 LayoutGPT: Compositional Visual Planning and Generation with Large Language Models
abstract
Attaining a high degree of user controllability in visual generation often requires intricate, fine-grained inputs like layouts. However, such inputs impose a substantial burden on users when compared to simple text inputs. To address the issue, we study how Large Language Models (LLMs) can serve as visual planners by generating layouts from text conditions, and thus collaborate with visual generative models. We propose LayoutGPT, a method to compose in-context visual demonstrations in style sheet language to enhance visual planning skills of LLMs. We show that LayoutGPT can generate plausible layouts in multiple domains, ranging from 2D images to 3D indoor scenes. LayoutGPT also shows superior performance in converting challenging language concepts like numerical and spatial relations to layout arrangements for faithful text-to-image generation. When combined with a downstream image generation model, LayoutGPT outperforms text-to-image models/systems by 20-40\% and achieves comparable performance as human users in designing visual layouts for numerical and spatial correctness. Lastly, LayoutGPT achieves comparable performance to supervised methods in 3D indoor scene synthesis, demonstrating its effectiveness and potential in multiple visual domains.
Weixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani, Arjun R. Akula, Xuehai He, Sugato Basu, Xin Wang 0061, William Yang Wang
NeurIPS3
2023 PHOTOSWAP: Personalized Subject Swapping in Images
abstract
In an era where images and visual content dominate our digital landscape, the ability to manipulate and personalize these images has become a necessity. Envision seamlessly substituting a tabby cat lounging on a sunlit window sill in a photograph with your own playful puppy, all while preserving the original charm and composition of the image. We present \emph{Photoswap}, a novel approach that enables this immersive image editing experience through personalized subject swapping in existing images. \emph{Photoswap} first learns the visual concept of the subject from reference images and then swaps it into the target image using pre-trained diffusion models in a training-free manner. We establish that a well-conceptualized visual subject can be seamlessly transferred to any image with appropriate self-attention and cross-attention manipulation, maintaining the pose of the swapped subject and the overall coherence of the image. Comprehensive experiments underscore the efficacy and controllability of \emph{Photoswap} in personalized subject swapping. Furthermore, \emph{Photoswap} significantly outperforms baseline methods in human ratings across subject swapping, background preservation, and overall quality, revealing its vast application potential, from entertainment to professional editing.
Yilin Wang 0002, Nanxuan Zhao, Tsu-Jui Fu, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Xin Wang 0061
NeurIPS4
2022 DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents
abstract
Creating presentation materials requires complex multimodal reasoning skills to summarize key concepts and arrange them in a logical and visually pleasing manner. Can machines learn to emulate this laborious process? We present a novel task and approach for document-to-slide generation. Solving this involves document summarization, image and text retrieval, slide structure and layout prediction to arrange key elements in a form suitable for presentation. We propose a hierarchical sequence-to-sequence approach to tackle our task in an end-to-end manner. Our approach exploits the inherent structures within documents and slides and incorporates paraphrasing and layout prediction modules to generate slides. To help accelerate research in this domain, we release a dataset about 6K paired documents and slide decks used in our experiments. We show that our approach outperforms strong baselines and produces slides with rich content and aligned imagery.
Tsu-Jui Fu, William Yang Wang, Daniel McDuff, Yale Song
AAAI1
2022 M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers
abstract
Video editing tools are widely used nowadays for digital design. Although the demand for these tools is high, the prior knowledge required makes it difficult for novices to get started. Systems that could follow natural language instructions to perform automatic editing would significantly improve accessibility. This paper introduces the language-based video editing (LBVE) task, which allows the model to edit, guided by text instruction, a source video into a target video. LBVE contains two features: 1) the scenario of the source video is preserved instead of generating a completely different video; 2) the semantic is presented differently in the target video, and all changes are controlled by the given instruction. We propose a Multi-Modal Multi-Level Transformer (M3L) to carry out LBVE. M3L dynamically learns the correspondence between video perception and language semantic at different levels, which benefits both the video understanding and video frame synthesis. We build three new datasets for evaluation, including two diagnostic and one from natural videos with human-labeled text. Extensive experimental results show that M3L is effective for video editing and that LBVE can lead to a new field toward vision-and-language research.
Tsu-Jui Fu, Xin Wang 0061, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
CVPR1
2022 Language-Driven Artistic Style Transfer
Tsu-Jui Fu, Xin Wang 0061, William Yang Wang
ECCV (36)1
2022 ULN: Towards Underspecified Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) is a task to guide an embodied agent moving to a target position using language instructions.Despite the significant performance improvement, the wide use of fine-grained instructions fails to characterize more practical linguistic variations in reality.To fill in this gap, we introduce a new setting, namely Underspecified vision-and-Language Navigation (ULN), and associated evaluation datasets.ULN evaluates agents using multi-level underspecified instructions instead of purely fine-grained or coarsegrained, which is a more realistic and general setting.As a primary step toward ULN, we propose a VLN framework that consists of a classification module, a navigation agent, and an Exploitation-to-Exploration (E2E) module.Specifically, we propose to learn Granularity Specific Sub-networks (GSS) for the agent to ground multi-level instructions with minimal additional parameters.Then, our E2E module estimates grounding uncertainty and conducts multi-step lookahead exploration to improve the success rate further.Experimental results show that existing VLN models are still brittle to multi-level language underspecification.Our framework is more robust and outperforms the baselines on ULN by "10% relative success rate across all levels. 1
Weixi Feng, Tsu-Jui Fu, William Yang Wang
EMNLP2
2022 CPL: Counterfactual Prompt Learning for Vision and Language Models
abstract
Xuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu, Arjun Akula, Varun Jampani, Pradyumna Narayana, Sugato Basu, William Yang Wang, Xin Wang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu, Arjun R. Akula, Varun Jampani, Pradyumna Narayana, Sugato Basu, William Yang Wang, Xin Wang 0061
EMNLP4
2021 L2C: Describing Visual Differences Needs Semantic Understanding of Individuals
abstract
Recent advances in language and vision push forward the research of captioning a single image to describing visual differences between image pairs.Suppose there are two images, I 1 and I 2 , and the task is to generate a description W 1,2 comparing them, existing methods directly model ⟨I 1 , I 2 ⟩ → W 1,2 mapping without the semantic understanding of individuals.In this paper, we introduce a Learningto-Compare (L2C) model, which learns to understand the semantic structures of these two images and compare them while learning to describe each one.We demonstrate that L2C benefits from a comparison between explicit semantic representations and singleimage captions, and generalizes better on the new testing image pairs.It outperforms the baseline on both automatic evaluation and human evaluation for the Birds-to-Words dataset.
An Yan 0003, Xin Wang 0061, Tsu-Jui Fu, William Yang Wang
EACL3
2021 Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation
abstract
Wanrong Zhu, Xin Wang, Tsu-Jui Fu, An Yan, Pradyumna Narayana, Kazoo Sone, Sugato Basu, William Yang Wang. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Wanrong Zhu, Xin Wang 0061, Tsu-Jui Fu, An Yan 0003, Pradyumna Narayana, Kazoo Sone, Sugato Basu, William Yang Wang
EACL3
2021 Semi-Supervised Policy Initialization for Playing Games with Language Hints
abstract
Using natural language as a hint can supply an additional reward for playing sparse-reward games.Achieving a goal should involve several different hints, while the given hints are usually incomplete.Those unmentioned latent hints still rely on the sparse reward signal, and make the learning process difficult.In this paper, we propose semi-supervised initialization (SSI) that allows the agent to learn from various possible hints before training under different tasks.Experiments show that SSI not only helps to learn faster (1.2x) but also has a higher success rate (11% relative improvement) of the final policy.
Tsu-Jui Fu, William Yang Wang
NAACL-HLT1
2020 Why Attention? Analyze BiLSTM Deficiency and Its Remedies in the Case of NER
abstract
BiLSTM has been prevalently used as a core module for NER in a sequence-labeling setup. State-of-the-art approaches use BiLSTM with additional resources such as gazetteers, language-modeling, or multi-task supervision to further improve NER. This paper instead takes a step back and focuses on analyzing problems of BiLSTM itself and how exactly self-attention can bring improvements. We formally show the limitation of (CRF-)BiLSTM in modeling cross-context patterns for each word – the XOR limitation. Then, we show that two types of simple cross-structures – self-attention and Cross-BiLSTM – can effectively remedy the problem. We test the practical impacts of the deficiency on real-world NER datasets, OntoNotes 5.0 and WNUT 2017, with clear and consistent improvements over the baseline, up to 8.7% on some of the multi-token entity mentions. We give in-depth analyses of the improvements across several aspects of NER, especially the identification of multi-token mentions. This study should lay a sound foundation for future improvements on sequence-labeling NER1.
Peng-Hsuan Li, Tsu-Jui Fu, Wei-Yun Ma
AAAI2
2020 Counterfactual Vision-and-Language Navigation via Adversarial Path Sampler
Tsu-Jui Fu, Xin Wang 0061, Matthew F. Peterson, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
ECCV (6)1
2020 SSCR: Iterative Language-Based Image Editing via Self-Supervised Counterfactual Reasoning
abstract
Iterative Language-Based Image Editing (IL-BIE) tasks follow iterative instructions to edit images step by step.Data scarcity is a significant issue for ILBIE as it is challenging to collect large-scale examples of images before and after instruction-based changes.However, humans still accomplish these editing tasks even when presented with an unfamiliar image-instruction pair.Such ability results from counterfactual thinking and the ability to think about alternatives to events that have happened already.In this paper, we introduce a Self-Supervised Counterfactual Reasoning (SSCR) framework that incorporates counterfactual thinking to overcome data scarcity.SSCR allows the model to consider out-ofdistribution instructions paired with previous images.With the help of cross-task consistency (CTC), we train these counterfactual instructions in a self-supervised scenario.Extensive results show that SSCR improves the correctness of ILBIE in terms of both object identity and position, establishing a new state of the art (SOTA) on two IBLIE datasets (i-CLEVR and CoDraw).Even with only 50% of the training data, SSCR achieves a comparable result to using complete data.
Tsu-Jui Fu, Xin Wang 0061, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
EMNLP (1)1
2019 GraphRel: Modeling Text as Relational Graphs for Joint Entity and Relation Extraction
abstract
In this paper, we present GraphRel, an end-to-end relation extraction model which uses graph convolutional networks (GCNs) to jointly learn named entities and relations.In contrast to previous baselines, we consider the interaction between named entities and relations via a relation-weighted GCN to better extract relations.Linear and dependency structures are both used to extract both sequential and regional features of the text, and a complete word graph is further utilized to extract implicit features among all word pairs of the text.With the graph-based approach, the prediction for overlapping relations is substantially improved over previous sequential approaches.We evaluate GraphRel on two public datasets: NYT and WebNLG.Results show that GraphRel maintains high precision while increasing recall substantially.Also, GraphRel outperforms previous work by 3.2% and 5.8% (F1 score), achieving a new state-of-the-art for relation extraction.
Tsu-Jui Fu, Peng-Hsuan Li, Wei-Yun Ma
ACL (1)1
2019 A Distributed Scheme for Accelerating Semantic Video Segmentation on An Embedded Cluster
abstract
We present a methodology for enhancing the throughput of semantic video segmentation tasks on an embedded cluster containing multiple embedded processing elements (ePEs). The methodology embraces a scalable master-slave hierarchy and features a global and local key management scheme for allocating video frames to different ePEs. The master ePE divides each video frame into frame regions, and dynamically distributes different regions to different slave ePEs. Each slave ePE executes either a segmentation path or a flow path: the former is highly accurate but slower, while the latter is faster but less accurate. A lightweight decision network is employed to determine the execution path for each slave ePE. We propose a global and local key management scheme to facilitate the execution of the embedded cluster, such that the average processing latency of each frame is significantly reduced. We evaluate the performance of our methodology on a real embedded cluster in terms of accuracy and frame rate, and validate its effectiveness and efficiency for various ePE configurations. We further provide a detailed latency analysis for different configurations of ePEs.
Hsuan-Kung Yang, Tsu-Jui Fu, Po-Han Chiang, Kuan-Wei Ho, Chun-Yi Lee
ICCD2
2019 Attentive and Adversarial Learning for Video Summarization
abstract
This paper aims to address the video summarization problem via attention-aware and adversarial training. We formulate the problem as a sequence-to-sequence task, where the input sequence is an original video and the output sequence is its summarization. We propose a GAN-based training framework, which combines the merits of unsupervised and supervised video summarization approaches. The generator is an attention-aware Ptr-Net that generates the cutting points of summarization fragments. The discriminator is a 3D CNN classifier to judge whether a fragment is from a ground-truth or a generated summarization. The experiments show that our method achieves state-of-the-art results on SumMe, TVSum, YouTube, and LoL datasets with 1.5% to 5.6% improvements. Our Ptr-Net generator can overcome the unbalanced training-test length in the seq2seq problem, and our discriminator is effective in leveraging unpaired summarizations to achieve better performance.
Tsu-Jui Fu, Shao-Heng Tai, Hwann-Tzong Chen
WACV1
2018 Region-Semantics Preserving Image Synthesis
Kang-Jun Liu, Tsu-Jui Fu, Shan-Hung Wu
ACCV (4)2
2018 Dynamic Video Segmentation Network
abstract
In this paper, we present a detailed design of dynamic video segmentation network (DVSNet) for fast and efficient semantic video segmentation. DVSNet consists of two convolutional neural networks: a segmentation network and a flow network. The former generates highly accurate semantic segmentations, but is deeper and slower. The latter is much faster than the former, but its output requires further processing to generate less accurate semantic segmentations. We explore the use of a decision network to adaptively assign different frame regions to different networks based on a metric called expected confidence score. Frame regions with a higher expected confidence score traverse the flow network. Frame regions with a lower expected confidence score have to pass through the segmentation network. We have extensively performed experiments on various configurations of DVSNet, and investigated a number of variants for the proposed decision network. The experimental results show that our DVSNet is able to achieve up to 70.4% mIoU at 19.8 fps on the Cityscape dataset. A high speed version of DVSNet is able to deliver an fps of 30.4 with 63.2% mIoU on the same dataset. DVSNet is also able to reduce up to 95% of the computational workloads.
Yu-Syuan Xu, Tsu-Jui Fu, Hsuan-Kung Yang, Chun-Yi Lee
CVPR2
2018 Speed Reading: Learning to Read ForBackward via Shuttle
abstract
We present LSTM-Shuttle, which applies human speed reading techniques to natural language processing tasks for accurate and efficient comprehension.In contrast to previous work, LSTM-Shuttle not only reads shuttling forward but also goes back.Shuttling forward enables high efficiency, and going backward gives the model a chance to recover lost information, ensuring better prediction.We evaluate LSTM-Shuttle on sentiment analysis, news classification, and cloze on IMDB, Rotten Tomatoes, AG, and Children's Book Test datasets.We show that LSTM-Shuttle predicts both better and more quickly.To demonstrate how LSTM-Shuttle actually behaves, we also analyze the shuttling operation and present a case study.
Tsu-Jui Fu, Wei-Yun Ma
EMNLP1
2018 Diversity-Driven Exploration Strategy for Deep Reinforcement Learning
abstract
Efficient exploration remains a challenging research problem in reinforcement learning, especially when an environment contains large state spaces, deceptive local optima, or sparse rewards. To tackle this problem, we present a diversity-driven approach for exploration, which can be easily combined with both off- and on-policy reinforcement learning algorithms. We show that by simply adding a distance measure to the loss function, the proposed methodology significantly enhances an agent's exploratory behaviors, and thus preventing the policy from being trapped in local optima. We further propose an adaptive scaling method for stabilizing the learning process. We demonstrate the effectiveness of our method in huge 2D gridworlds and a variety of benchmark environments, including Atari 2600 and MuJoCo. Experimental results show that our method outperforms baseline approaches in most tasks in terms of mean scores and exploration efficiency.
Zhang-Wei Hong, Tzu-Yun Shann, Shih-Yang Su, Yi-Hsiang Chang, Tsu-Jui Fu, Chun-Yi Lee
NeurIPS5