Takashi Isobe

dblp:48/2800 · DBLP profile ↗
← Back
15ranked-venue papers
10as first author
9since 2021 · last 2025
0009-0001-5170-5436ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 7 first-author · 6 since 2021Computer networks · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 ReNeg: Learning Negative Embedding with Reward Guidance
abstract
In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In this paper, we introduce ReNeg, an end-to-end method designed to learn improved Negative embeddings guided by a Reward model. We employ a reward feedback learning framework and integrate classifier-free guidance (CFG) into the training process, which was previously utilized only during inference, thus enabling the effec tive learning of negative embeddings. We also propose two strategies for learning both global and per-sample negative embeddings. Extensive experiments show that the learned negative embedding significantly outperforms null-text and handcrafted counterparts, achieving substantial improvements in human preference alignment. Additionally, the negative embedding learned within the same text embedding space exhibits strong generalization capabilities. For example, using the same CLIP text encoder, the negative embedding learned on SD1.5 can be seamlessly transferred to text-to-image or even text-to-video models such as ControlNet, ZeroScope, and VideoCrafter2, resulting in consistent performance improvements across the board. Code is available at https://github.com/AMD-AIG-AIMA/ReNeg.
Xiaomin Li 0001, Yixuan Liu 0004, Takashi Isobe, Xu Jia 0012, Qinpeng Cui, Dong Zhou 0003, Dong Li 0025, You He 0002, Huchuan Lu, Zhongdao Wang, Emad Barsoum
CVPR3
2024 Customizing Text-to-Image Generation with Inverted Interaction
abstract
Subject-driven image generation, aimed at customizing user-specified subjects, has experienced rapid progress. However, most of them focus on transferring the customized appearance of subjects. In this work, we consider a novel concept customization task, that is, capturing the interaction between subjects in exemplar images and transferring the learned concept of interaction to achieve customized text-to-image generation. Intrinsically, the interaction between subjects is diverse and is difficult to describe in only a few words. In addition, typical exemplar images are about the interaction between humans, which further intensifies the challenge of interaction-driven image generation with various categories of subjects. To address this task, we adopt a divide-and-conquer strategy and propose a two-stage interaction inversion framework. The framework begins by learning a pseudo-word for a single pose of each subject in the interaction. This is then employed to promote the learning of the concept for the interaction. In addition, language prior and cross-attention loss are incorporated into the optimization process to encourage the modeling of interaction. Extensive experiments demonstrate that the proposed methods are able to effectively invert the interactive pose from exemplar images and apply it to the customized generation with user-specified interaction.
Mengmeng Ge 0002, Xu Jia 0012, Takashi Isobe, Xiaomin Li 0001, Dong Zhou 0003, Li Wang 0125, Huchuan Lu, Ashish Sirasao, Emad Barsoum
ACM Multimedia3
2023 Compression-Aware Video Super-Resolution
abstract
Videos stored on mobile devices or delivered on the Internet are usually in compressed format and are of various unknown compression parameters, but most video super-resolution (VSR) methods often assume ideal inputs resulting in large performance gap between experimental settings and real-world applications. In spite of a few pioneering works being proposed recently to super-resolve the compressed videos, they are not specially designed to deal with videos of various levels of compression. In this paper, we propose a novel and practical compression-aware video super-resolution model, which could adapt its video enhancement process to the estimated compression level. A compression encoder is designed to model compression levels of input frames, and a base VSR model is then conditioned on the implicitly computed representation by inserting compression-aware modules. In addition, we propose to further strengthen the VSR model by taking full advantage of meta data that is embedded naturally in compressed video streams in the procedure of information fusion. Extensive experiments are conducted to demonstrate the effectiveness and efficiency of the proposed method on compressed VSR benchmarks. The codes will be available at https://github.com/aprBlue/CAVSR
Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Huchuan Lu, Yu-Wing Tai
CVPR2
2023 List-Mode PET Image Reconstruction Using Deep Image Prior
abstract
List-mode positron emission tomography (PET) image reconstruction is an important tool for PET scanners with many lines-of-response and additional information such as time-of-flight and depth-of-interaction. Deep learning is one possible solution to enhance the quality of PET image reconstruction. However, the application of deep learning techniques to list-mode PET image reconstruction has not been progressed because list data is a sequence of bit codes and unsuitable for processing by convolutional neural networks (CNN). In this study, we propose a novel list-mode PET image reconstruction method using an unsupervised CNN called deep image prior (DIP) which is the first trial to integrate list-mode PET image reconstruction and CNN. The proposed list-mode DIP reconstruction (LM-DIPRecon) method alternatively iterates the regularized list-mode dynamic row action maximum likelihood algorithm (LM-DRAMA) and magnetic resonance imaging conditioned DIP (MR-DIP) using an alternating direction method of multipliers. We evaluated LM-DIPRecon using both simulation and clinical data, and it achieved sharper images and better tradeoff curves between contrast and noise than the LM-DRAMA, MR-DIP and sinogram-based DIPRecon methods. These results indicated that the LM-DIPRecon is useful for quantitative PET imaging with limited events while keeping accurate raw data information. In addition, as list data has finer temporal information than dynamic sinograms, list-mode deep image prior reconstruction is expected to be useful for 4D PET imaging and motion correction.
Kibo Ote, Fumio Hashimoto, Yuya Onishi, Takashi Isobe, Yasuomi Ouchi
IEEE Trans. Medical Imaging4
2022 Look Back and Forth: Video Super-Resolution with Explicit Temporal Difference Modeling
abstract
Temporal modeling is crucial for video super-resolution. Most of the video super-resolution methods adopt the optical flow or deformable convolution for explicitly motion compensation. However, such temporal modeling techniques increase the model complexity and might fail in case of occlusion or complex motion, resulting in serious distortion and artifacts. In this paper, we propose to explore the role of explicit temporal difference modeling in both LR and HR space. Instead of directly feeding consecutive frames into a VSR model, we propose to compute the temporal difference between frames and divide those pixels into two subsets according to the level of difference. They are separately processed with two branches of different receptive fields in order to better extract complementary information. To further enhance the super-resolution result, not only spatial residual features are extracted, but the difference between consecutive frames in high-frequency domain is also computed. It allows the model to exploit intermediate SR results in both future and past to refine the current SR output. The difference at different time steps could be cached such that information from further distance in time could be propagated to the current frame for refinement. Experiments on several video super-resolution benchmark datasets demonstrate the effectiveness of the proposed method and its favorable performance against state-of-the-art methods.
Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Ruihuang Li, Yongjie Shi, Huchuan Lu, Yu-Wing Tai
CVPR1
2021 Chemicals Informatics: Explore Structural Factors and Potential Chemicals based on Public Literatures
abstract
Chemical industry pays much cost and long time to develop new compounds or composites that have aimed properties. The developers need to efficiently discover initial candidates before simulation, actual synthesis, optimization, and evaluation. To meet their needs, we have developed CI (Chemicals Informatics) to efficiently discover potential candidates of compounds or composites based on large number of public literatures. Our system now has the data of 117M existing compounds, and 61 properties extracted by analyzing a public chemical database linked to worldwide 33M papers and 30M patents in addition to 11M new structures generated based on existing compounds. The compounds are shown as vectors including 41 organic and 70 inorganic features. The properties are extracted from literatures using rule-based NLP (Natural Language Processing). Our system predicts 61 properties by each space of neighbor, structural crossover, and compound crossover. Users can extract structural factors and highly probable combinations of structures and compounds that contribute to good properties. Moreover, they can find potential combinations that have no patents of aimed property from 128M to the fourth power. In the biochemical use case, we explored candidates for antiCoronavirus medicine. CI extracted three candidates and we confirmed that the results matched with those of past papers about docking simulator, clinical trial, and in-vitro test. In the other case of kinase inhibitors, we could find new structures derived from existing inhibitors or like the inhibitor of past paper. In other cases, we explored structural factors and new probable combinations for optical resin lens, low dielectric and etching gas for semiconductor. CI could extract structural factors that contributed to good properties and search highly probable combinations that had a new structure replicating the same idea as the past literatures.
Takashi Isobe, Yoshihiro Okada
IEEE BigData1
2021 Multi-Target Domain Adaptation With Collaborative Consistency Learning
abstract
Recently unsupervised domain adaptation for the semantic segmentation task has become more and more popular due to high-cost of pixel-level annotation on real-world images. However, most domain adaptation methods are only restricted to single-source-single-target pair, and can not be directly extended to multiple target domains. In this work, we propose a collaborative learning framework to achieve unsupervised multi-target domain adaptation. An unsupervised domain adaptation expert model is first trained for each source-target pair and is further encouraged to collaborate with each other through a bridge built between different target domains. These expert models are further improved by adding the regularization of making the consistent pixel-wise prediction for each sample with the same structured context. To obtain a single model that works across multiple target domains, we propose to simultaneously learn a student model which is trained to not only imitate the output of each expert on the corresponding target domain, but also to pull different expert close to each other with regularization on their weights. Extensive experiments demonstrate that the proposed method can effectively exploit rich structured information contained in both labeled source domain and multiple unlabeled target domains. Not only does it perform well across multiple target domains but also performs favorably against state-of-the-art unsupervised domain adaptation methods specially trained on a single source-target pair. Code is available at https://github.com/junpan19/MTDA.
Takashi Isobe, Xu Jia 0012, Shuaijun Chen, Yongjie Shi, Jianzhuang Liu, Huchuan Lu, Shengjin Wang
CVPR1
2021 Frame-Rate-Aware Aggregation for Efficient Video Super-Resolution
abstract
Video super-resolution, which aims at producing a high-resolution video from its corresponding low-resolution version, recently draws increasing attention. In contrast to the previous works that perform explicit motion estimation and compensation, we propose a novel deep neural network which performs implicit motion estimation with frame-rate-based temporal aggregation. Specifically, the input frames are first aggregated by a frame-rate-aware 3D convolution layer, where neighboring frames are integrated with the reference frame according to the corresponding frame rate. Then, the aggregated features are fed into several branches for further aggregation. Different branches correspond to a kind of motion rate, which provides complementary information to recover missing details in the reference frame. Extensive experiments demonstrate that our method is able to handle various motion types and achieves state-of-the-art performance on several benchmarks. In addition, our model is light-weight and requires an extremely less computational load than other state-of-the-art methods.
Takashi Isobe, Shengjin Wang
ICASSP1
2021 Towards Discriminative Representation Learning for Unsupervised Person Re-identification
abstract
In this work, we address the problem of unsupervised domain adaptation for person re-ID where annotations are available for the source domain but not for target. Previous methods typically follow a two-stage optimization pipeline, where the network is first pre-trained on source and then fine-tuned on target with pseudo labels created by feature clustering. Such methods sustain two main limitations. (1) The label noise may hinder the learning of discriminative features for recognizing target classes. (2) The domain gap may hinder knowledge transferring from source to target. We propose three types of technical schemes to alleviate these issues. First, we propose a cluster-wise contrastive learning algorithm (CCL) by iterative optimization of feature learning and cluster refinery to learn noise-tolerant representations in the unsupervised manner. Second, we adopt a progressive domain adaptation (PDA) strategy to gradually mitigate the domain gap between source and target data. Third, we propose Fourier augmentation (FA) for further maximizing the class separability of re-ID models by imposing extra constraints in the Fourier space. We observe that these proposed schemes are capable of facilitating the learning of discriminative feature representations. Experiments demonstrate that our method consistently achieves notable improvements over the state-of-the-art unsupervised re-ID methods on multiple benchmarks, e.g., surpassing MMT largely by 8.1%, 9.9%, 11.4% and 11.1% mAP on the Market-to-Duke, Duke-to-Market, Market-to-MSMT and Duke-to-MSMT tasks, respectively.
Takashi Isobe, Dong Li 0025, Shengjin Wang
ICCV1
2020 Revisiting Temporal Modeling for Video Super-resolution
Takashi Isobe, Shengjin Wang
BMVC1
2020 Video Super-Resolution With Temporal Group Attention
abstract
Video super-resolution, which aims at producing a high-resolution video from its corresponding low-resolution version, has recently drawn increasing attention. In this work, we propose a novel method that can effectively incorporate temporal information in a hierarchical way. The input sequence is divided into several groups, with each one corresponding to a kind of frame rate. These groups provide complementary information to recover missing details in the reference frame, which is further integrated with an attention module and a deep intra-group fusion module. In addition, a fast spatial alignment is proposed to handle videos with large motion. Extensive results demonstrate the capability of the proposed model in handling videos with various motion. It achieves favorable performance against state-of-the-art methods on several benchmark datasets.
Takashi Isobe, Songjiang Li, Xu Jia 0012, Shanxin Yuan, Gregory Slabaugh, Chunjing Xu, Yali Li 0001, Shengjin Wang, Qi Tian 0001
CVPR1
2020 Video Super-Resolution with Recurrent Structure-Detail Network
Takashi Isobe, Xu Jia 0012, Shuhang Gu, Songjiang Li, Shengjin Wang, Qi Tian 0001
ECCV (12)1
2020 Intra-Clip Aggregation For Video Person Re-Identification
abstract
Video-based person re-identification has drawn massive attention in recent years due to its extensive applications in video surveillance. While deep learning based methods have led to significant progress, these methods are limited by ineffectively using complementary information, which is blamed on necessary data augmentation in training process. Data augmentation has been widely used to mitigate the overfitting trap and improve the ability of network representation. However, the previous methods adopt image-based data augmentation scheme to individually process the input frames, which corrupts the complementary information between consecutive frames and causes performance degradation. In this paper, we propose a novel video-based data augmentation scheme, termed as Synchronous Data Augmentation, to address the challenge above. In order to represent discriminative clip-level features, we also propose a cascade integration module which hierarchically aggregates the intra-clip features with a linear-nonlinear combining projection. Extensive experiments on three benchmark datasets demonstrate that our framework outperforms the most recent state-of-the-art methods. We also perform cross-dataset validation to prove the generality of our method.
Takashi Isobe, Yali Li 0001, Shengjin Wang
ICIP1
2010 10 Gbps implementation of TLS/SSL accelerator on FPGA
abstract
This paper proposes the one-chip architecture to mount all processes for TLS/SSL ciphered communication into one FPGA or ASIC, and shows the 10 Gbps implementation of low-power (23 W) TLS/SSL accelerator on 65 nm FPGA. The usage of FPGA/ASIC enables high efficient processing and low-power consumption by using parallel, optimized and pipelined processing. One-chip architecture achieves high throughput by using a switch to avoid the congestion in exchanging data between multiple processing-blocks. In this research, to reduce the circuit area in the one-chip architecture, high-efficient processing design (a parallel processing circuit shared with multiple data, and a circuit shared in transmitting and receiving) was used. In addition, to enhance the operating frequency, a switch downsized by sharing a port to exchange data with multiple blocks decreased the number of wires. By means of these designs, circuit area to implement all TLS/SSL processes was reduced to less than that of 65 nm FPGA used in this research, and 166 MHz operating frequency required to realize 10 Gbps throughput at 64-bit pipeline was achieved. In experimental evaluation using prototype, 23 W power consumption and 10 Gbps encryption throughput were achieved.
Takashi Isobe, Satoshi Tsutsumi, Koichiro Seto, Kenji Aoshima, Kazutoshi Kariya
IWQoS1
2004 New Bandwidth-Control Design: Policer for Probable Packet Discard (PPPD)
abstract
A new bandwidth-control design - called policer for packet probable discard (PPPD) - has been developed. This not only limits the bandwidth of each user to a policing bandwidth, but also achieves TCP throughputs equal to the policing bandwidth. The proposed PPPD uses an improved leaky bucket algorithm, by which packets are discarded at a certain probability when the received bandwidth is judged to be more than the policing bandwidth. Simulations show that the proposed PPPD increases TCP throughput from 74% to 95%, or more, of the policing bandwidth.
Takeki Yazaki, Takashi Isobe, Yuichi Ishikawa, Hiroki Yano
LCN2