EDBT 2026 Demo / reviewers in the wild / expert
Yunbin Tu
dblp:207/1931
· DBLP profile ↗
26ranked-venue papers
12as first author
23since 2021 · last 2026
0000-0002-9525-9060ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 12 · 8 first-author · 12 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
Zhuo Tao, Liang Li 0003, Qi Chen 0014, Yunbin Tu, Zhengjun Zha, Amin Beheshti, Qingming Huang, Yuankai Qi, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | Learning to Change by Critique and Correction: A Synergistic Framework for Remote Sensing Change Detection and CaptioningabstractChange Detection (CD) and Change Captioning (CC) are two core tasks for understanding land-cover evolution in remote sensing imagery. Existing approaches have explored CD-CC joint modeling through shared representations, task-specific decoding, feature interaction, and semantic guidance. However, the lack of explicit cross-task feedback mechanisms often leads to mutual interference, making it difficult to achieve both accurate detection and expressive descriptions. To address this issue, we propose the Learning to Change by Critique and Correction (LCCC) framework, which reformulates CD and CC as a critique-correction closed-loop process. In LCCC, CD and CC no longer passively share features but interact through bidirectional critique and correction: the CD task provides explicit spatial constraints for CC, and CC, in turn, supervises CD via a Text-Guided Critique Attention (TGCA) mechanism, establishing a synergistic relationship where both tasks act as critics and correctors. Furthermore, we design a Reciprocal Suppression and Enhancement (RSE) module to purify cross-task representations and propose a Key Complementary Feature Fusion (KCFF) mechanism to bridge the gap between high-level semantics and low-level visual features, ensuring a balance between task specialization and cross-task enhancement. Extensive experiments demonstrate that LCCC significantly outperforms existing methods in both detection accuracy and description quality, validating the effectiveness and generality of the proposed critique-correction paradigm for synergistic multi-task modeling. The code of the proposed method is available at https://github.com/Throb16/Lccc. Huafeng Li 0001, Yamin Zhang, Yunbin Tu, Liang Li 0003 |
IEEE Trans. Image Process. | 3 |
| 2025 | Unsupervised Photometric-Consistent Depth Estimation from Endoscopic Monocular VideoabstractRecent advancements in unsupervised monocular depth estimation typically rely on an assumption that image photometry remains consistent across consecutive frames. However, this assumption often fails in endoscopic scenes due to: 1) local photometric inconsistency caused by specular reflections creating highlights; and 2) global photometric inconsistency resulting from the simultaneous movement of the light source and the camera. Since unsupervised depth estimation methods rely on appearance discrepancies between frames as a supervisory signal, these photometric inconsistencies inevitably deteriorate loss function calculation. In this paper, our goal is to obtain a strong and reliable supervisory signal for achieving photometric-consistent depth estimation. To this end, for local photometric inconsistency, we utilize the specular reflection model to introduce a Highlight Loss for handling the estimation of highlight regions. For global photometric inconsistency, we design a Photometric Match module, which utilizes the spotlight illumination model to derive an analytical expression, achieving photometric alignment across different frames. Unlike previous works that introduce additional optical flow or networks, our method is simpler and more efficient. Extensive experiments demonstrate our method achieves the state-of-the-art results on C3VD, SCARED and SERV-CT datasets. Weijun Lin, Qingyuan Xiang, Yunbin Tu, Shitan Asu |
AAAI | 4 |
| 2025 | Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-CaptioningabstractVideo has emerged as a favored multimedia format on the internet. To better gain video contents, a new topic HIREST is presented, including video retrieval, moment retrieval, moment segmentation, and step-captioning. The pioneering work chooses the pre-trained CLIP-based model for video retrieval, and leverages it as a feature extractor for other three challenging tasks solved in a multi-task learning paradigm. Nevertheless, this work struggles to learn the comprehensive cognition of user-preferred content, due to disregarding the hierarchies and association relations across modalities. In this paper, guided by the shallow-to-deep principle, we propose a query-centric audio-visual cognition (QUAG) network to construct a reliable multi-modal representation for moment retrieval, segmentation and step-captioning. Specifically, we first design the modality-synergistic perception to obtain rich audio-visual content, by modeling global contrastive alignment and local fine-grained interaction between visual and audio modalities. Then, we devise the query-centric cognition that uses the deep-level query to perform the temporal-channel filtration on the shallow-level audio-visual representation. This can cognize user-preferred content and thus attain a query-centric audio-visual representation for three tasks. Extensive experiments show QUAG achieves the SOTA results on HIREST. Further, we test QUAG on the query-based video summarization task and verify its good generalization. Yunbin Tu, Liang Li 0003, Li Su 0003, Qingming Huang |
AAAI | 1 |
| 2025 | Generalizing Single-Frame Supervision to Event-Level Understanding for Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) aims to identify abnormal frames from discrete events within video sequences. Existing VAD methods suffer from heavy annotation burdens in fully-supervised paradigm, insensitivity to subtle anomalies in semi-supervised paradigm, and vulnerability to noise in weakly-supervised paradigm. To address these limitations, we propose a novel paradigm: Single-Frame supervised VAD (SF-VAD), which uses a single annotated abnormal frame per abnormal video. SF-VAD ensures annotation efficiency while offering precise anomaly reference, facilitating robust anomaly modeling, and enhancing the detection of subtle anomalies in complex visual contexts. To validate its effectiveness, we construct three SF-VAD benchmarks by manually re-annotating the ShanghaiTech, UCF-Crime, and XD-Violence datasets in a practical procedure. Further, we devise Frame-guided Progressive Learning (FPL), to generalize sparse frame supervision to event-level anomaly understanding. FPL first leverages evidential learning to estimate anomaly relevance guided by annotated frames. Then it extends anomaly supervision by mining discrete abnormal events based on anomaly relevance and feature similarity. Meanwhile, FPL decouples normal patterns by isolating distinct normal frames outside abnormal events, reducing false alarms. Extensive experiments show SF-VAD achieves state-of-the-art detection results while offering a favorable trade-off between performance and annotation cost. Junxi Chen, Liang Li 0003, Yunbin Tu, Li Su 0003, Zhe Xue, Qingming Huang |
NeurIPS | 3 |
| 2025 | Dynamic Strategy Prompt Reasoning for Emotional Support ConversationabstractAn emotional support conversation (ESC) system aims to reduce users' emotional distress by engaging in conversation using various reply strategies as guidance. To develop instructive reply strategies for an ESC system, it is essential to consider the dynamic transitions of users' emotional states through the conversational turns. However, existing methods for strategy-guided ESC systems struggle to capture these transitions as they overlook the inference of fine-grained user intentions. This oversight poses a significant obstacle, impeding the model's ability to derive pertinent strategy information and, consequently, hindering its capacity to generate emotionally supportive responses. To tackle this limitation, we propose a novel dynamic strategy prompt reasoning model (DSR), which leverages sparse context relation deduction to acquire adaptive representation of reply strategies as prompts for guiding the response generation process. Specifically, we first perform turn-level commonsense reasoning with different approaches to extract auxiliary knowledge, which enhances the comprehension of user intention. Then we design a context relation deduction module to dynamically integrate interdependent dialogue information, capturing granular user intentions and generating effective strategy prompts. Finally, we utilize the strategy prompts to guide the generation of more relevant and supportive responses. DSR model is validated through extensive experiments conducted on a benchmark dataset, demonstrating its superior performance compared to the latest competitive methods in the field. Yiting Liu 0007, Liang Li 0003, Yunbin Tu, Beichen Zhang 0006, Zhengjun Zha, Qingming Huang |
IEEE Trans. Multim. | 3 |
| 2025 | SketchRefiner: Text-Guided Sketch Refinement Through Latent Diffusion ModelsabstractFree-hand sketches serve as efficient tools for creativity and communication, yet expressing ideas clearly through sketches remains challenging for untrained individuals. Optimizing sketches through text guidance can enhance individuals' ability to effectively convey their ideas and improve overall communication efficiency. While recent advancements in Artificial Intelligence Generated Content (AIGC) have been notable, research on optimizing free-hand sketches remains relatively unexplored. In this paper, we introduce SketchRefiner, an innovative method designed to refine rough sketches from various categories into polished versions guided by text prompts. SketchRefiner utilizes a latent diffusion model with ControlNet to guide a differentiable rasterizer in optimizing a set of Bézier curves. We extend the score distillation sampling (SDS) loss and introduce a joint semantic loss to encourage sketches aligned with given text prompts and free-hand sketches. Additionally, we propose a fusion attention-map stroke initialization strategy to improve the quality of refined sketches. Furthermore, SketchRefiner provides users with fine-grained control over text guidance. Through extensive experiments, we demonstrate that our method can generate accurate and aesthetically pleasing refined sketches that closely align with input text prompts and sketches. Yingjie Tian 0001, Minghao Liu 0022, Yunbin Tu, Duo Su |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Context-aware Difference Distilling for Multi-change CaptioningabstractMulti-change captioning aims to describe complex and coupled changes within an image pair in natural language.Compared with singlechange captioning, this task requires the model to have higher-level cognition ability to reason an arbitrary number of changes.In this paper, we propose a novel context-aware difference distilling (CARD) network to capture all genuine changes for yielding sentences.Given an image pair, CARD first decouples context features that aggregate all similar/dissimilar semantics, termed common/difference context features.Then, the consistency and independence constraints are designed to guarantee the alignment/discrepancy of common/difference context features.Further, the common context features guide the model to mine locally unchanged features, which are subtracted from the pair to distill locally difference features.Next, the difference context features augment the locally difference features to ensure that all changes are distilled.In this way, we obtain an omni-representation of all changes, which is translated into linguistic sentences by a transformer decoder.Extensive experiments on three public datasets show CARD performs favourably against state-of-the-art methods. Yunbin Tu, Liang Li 0003, Li Su 0003, Zhengjun Zha, Chenggang Yan 0001, Qingming Huang |
ACL (1) | 1 |
| 2024 | Dual Attention Encoder with Joint Preservation for Medical Image SegmentationabstractTransformers have recently gained considerable popularity for capturing long-range dependencies in the medical image segmentation. However, most transformer-based segmentation methods primarily focus on modeling global dependencies and fail to fully explore the complementary nature of different dimensional dependencies within features. These methods simply treat the aggregation of multi-dimensional dependencies as auxiliary modules for incorporating context into the Transformer architecture, thereby limiting the model’s capability to learn rich feature representations. To address this issue, we introduce the Dual Attention Encoder with Joint Preservation (DANIE) for medical image segmentation, which synergistically aggregates spatial-channel dependencies across both local and global areas through attention learning. Additionally, we design a lightweight aggregation mechanism, termed Joint Preservation, which learns a composite feature representation, allowing different dependencies to complement each other. Without bells and whistles, our DANIE significantly improves the performance of previous state-of-the-art methods on five popular medical image segmentation benchmarks, including Synapse, ACDC, ISIC 2017, ISIC 2018 and GlaS. Yunbin Tu, Bowen Zhong |
ECAI | 2 |
| 2024 | Distractors-Immune Representation Learning with Cross-Modal Contrastive Regularization for Change Captioning
Yunbin Tu, Liang Li 0003, Li Su 0003, Chenggang Yan 0001, Qingming Huang |
ECCV (43) | 1 |
| 2024 | MAGIC: Rethinking Dynamic Convolution Design for Medical Image SegmentationabstractRecently, dynamic convolution shows performance boost for the CNN-related networks in medical image segmentation. The core idea is to replace static convolutional kernel with a linear combination of multiple convolutional kernels, conditioned on input-dependent attention function. However, the existing dynamic convolution design suffers from two limitations: i) The convolutional kernels are weighted by enforcing a single-dimensional attention function upon the input maps, overlooking the synergy in multi-dimensional information. This results in sub-optimal computations of convolution kernels. ii) The linear kernel aggregation is inefficient, restricting the model's capacity to learn more intricate patterns. In this paper, we rethink the dynamic convolution design to address these limitations and propose multi-dimensional aggregation dynamic convolution (MAGIC). Specifically, our MAGIC introduce a dimensional-reciprocal fusion module to capture correlations among input maps across the spatial, channel, and global dimensions simultaneously for computing convolutional kernels. Furthermore, we design kernel recalculation module, which enhances the efficiency of aggregation through learning the interaction between kernels. As a drop-in replacement for regular convolution, our MAGIC can be flexibly integrated into prevalent pure CNN or hybrid CNN-Transformer backbones. The extensive experiments on four benchmarks demonstrate that our MAGIC outperforms regular convolution and existing dynamic convolution. Code is available at: https://github.com/Segment82/MAGIC Yunbin Tu, Qingyuan Xiang |
ACM Multimedia | 2 |
| 2024 | SMART: Syntax-Calibrated Multi-Aspect Relation Transformer for Change CaptioningabstractChange captioning aims to describe the semantic change between two similar images. In this process, as the most typical distractor, viewpoint change leads to the pseudo changes about appearance and position of objects, thereby overwhelming the real change. Besides, since the visual signal of change appears in a local region with weak feature, it is difficult for the model to directly translate the learned change features into the sentence. In this paper, we propose a syntax-calibrated multi-aspect relation transformer to learn effective change features under different scenes, and build reliable cross-modal alignment between the change features and linguistic words during caption generation. Specifically, a multi-aspect relation learning network is designed to 1) explore the fine-grained changes under irrelevant distractors (e.g., viewpoint change) by embedding the relations of semantics and relative position into the features of each image; 2) learn two view-invariant image representations by strengthening their global contrastive alignment relation, so as to help capture a stable difference representation; 3) provide the model with the prior knowledge about whether and where the semantic change happened by measuring the relation between the representations of captured difference and the image pair. Through the above manner, the model can learn effective change features for caption generation. Further, we introduce the syntax knowledge of Part-of-Speech (POS) and devise a POS-based visual switch to calibrate the transformer decoder. The POS-based visual switch dynamically utilizes visual information during different word generation based on the POS of words. This enables the decoder to build reliable cross-modal alignment, so as to generate a high-level linguistic sentence about change. Extensive experiments show that the proposed method achieves the state-of-the-art performance on the three public datasets. Yunbin Tu, Liang Li 0003, Li Su 0003, Zhengjun Zha, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Multi-Grained Representation Aggregating Transformer with Gating Cycle for Change CaptioningabstractChange captioning aims to describe the difference within an image pair in natural language, which combines visual comprehension and language generation. Although significant progress has been achieved, it remains a key challenge of perceiving the object change from different perspectives, especially the severe situation with drastic viewpoint change. In this article, we propose a novel full-attentive network, namely Multi-grained Representation Aggregating Transformer (MURAT), to distinguish the actual change from viewpoint change. Specifically, the Pair Encoder first captures similar semantics between pairwise objects in a multi-level manner, which are regarded as the semantic cues of distinguishing the irrelevant change. Next, a novel Multi-grained Representation Aggregator (MRA) is designed to construct the reliable difference representation by employing both coarse- and fine-grained semantic cues. Finally, the language decoder generates a description of the change based on the output of MRA. Besides, the Gating Cycle Mechanism is introduced to facilitate the semantic consistency between difference representation learning and language generation with a reverse manipulation, so as to bridge the semantic gap between change features and text features. Extensive experiments demonstrate that the proposed MURAT can greatly improve the ability to describe the actual change in the distraction of irrelevant change and achieves state-of-the-art performance on three benchmarks, CLEVR-Change, CLEVR-DC, and Spot-the-Diff. Shengbin Yue, Yunbin Tu, Liang Li 0003, Shengxiang Gao, Zhengtao Yu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Self-supervised Cross-view Representation Reconstruction for Change CaptioningabstractChange captioning aims to describe the difference between a pair of similar images. Its key challenge is how to learn a stable difference representation under pseudo changes caused by viewpoint change. In this paper, we address this by proposing a self-supervised cross-view representation reconstruction (SCORER) network. Concretely, we first design a multi-head token-wise matching to model relationships between cross-view features from similar/dissimilar images. Then, by maximizing cross-view contrastive alignment of two similar images, SCORER learns two view-invariant image representations in a self-supervised way. Based on these, we reconstruct the representations of unchanged objects by cross-attention, thus learning a stable difference representation for caption generation. Further, we devise a cross-modal backward reasoning to improve the quality of caption. This module reversely models a "hallucination" representation with the caption and "before" representation. By pushing it closer to the "after" representation, we enforce the caption to be informative about the difference in a self-supervised manner. Extensive experiments show our method achieves the state-of-the-art results on four datasets. The code is available at https://github.com/tuyunbin/SCORER. Yunbin Tu, Liang Li 0003, Li Su 0003, Zhengjun Zha, Chenggang Yan 0001, Qingming Huang |
ICCV | 1 |
| 2023 | Relation-aware attention for video captioning via graph learning
Yunbin Tu, Huafeng Li 0001, Shengxiang Gao, Zhengtao Yu 0001 |
Pattern Recognit. | 1 |
| 2023 | Viewpoint-Adaptive Representation Disentanglement Network for Change CaptioningabstractChange captioning is to describe the fine-grained change between a pair of images. The pseudo changes caused by viewpoint changes are the most typical distractors in this task, because they lead to the feature perturbation and shift for the same objects and thus overwhelm the real change representation. In this paper, we propose a viewpoint-adaptive representation disentanglement network to distinguish real and pseudo changes, and explicitly capture the features of change to generate accurate captions. Concretely, a position-embedded representation learning is devised to facilitate the model in adapting to viewpoint changes via mining the intrinsic properties of two image representations and modeling their position information. To learn a reliable change representation for decoding into a natural language sentence, an unchanged representation disentanglement is designed to identify and disentangle the unchanged features between the two position-embedded representations. Extensive experiments show that the proposed method achieves the state-of-the-art performance on the four public datasets. The code is available at https://github.com/tuyunbin/VARD. Yunbin Tu, Liang Li 0003, Li Su 0003, Junping Du 0001, Ke Lu 0002, Qingming Huang |
IEEE Trans. Image Process. | 1 |
| 2023 | Neighborhood Contrastive Transformer for Change CaptioningabstractChange captioning is to describe the semantic change between a pair of similar images in natural language. It is more challenging than general image captioning, because it requires capturing fine-grained change information while being immune to irrelevant viewpoint changes, and solving syntax ambiguity in change descriptions. In this paper, we propose a neighborhood contrastive transformer to improve the model's perceiving ability for various changes under different scenes and cognition ability for complex syntax structure. Concretely, we first design a neighboring feature aggregating to integrate neighboring context into each feature, which helps quickly locate the inconspicuous changes under the guidance of conspicuous referents. Then, we devise a common feature distilling to compare two images at neighborhood level and extract common properties from each image, so as to learn effective contrastive information between them. Finally, we introduce the explicit dependencies between words to calibrate the transformer decoder, which helps better understand complex syntax structure during training. Extensive experimental results demonstrate that the proposed method achieves the state-of-the-art performance on three public datasets with different change scenarios. The code is available athttps://github.com/tuyunbin/NCT. Yunbin Tu, Liang Li 0003, Li Su 0003, Ke Lu 0002, Qingming Huang |
IEEE Trans. Multim. | 1 |
| 2023 | I3N: Intra- and Inter-Representation Interaction Network for Change CaptioningabstractChange captioning aims to describe the disagreement of image pairs with a linguistic sentence. Compared with single image captioning, change captioning requires not only understanding the fine-grained information of each image, but also determining whether change occurs and further representing the differences of image pairs. Although much progress has been made, it remains a severe challenge of the precise difference representation in the distraction of viewpoint change, especially that of tiny difference. In this paper, we propose a novel Intra- and Inter-representation Interaction Network (I3N) to learn the fine difference representation and be immune to viewpoint change. In the Intra-representation Interaction stage, we design Geometry-Semantic Interaction Refining (GSIR) to explore the positional and semantic interactions of intra-image, which can be a prior knowledge of enduring viewpoint change and reinforce the cognition of semantic change. In the Inter-representation Interaction stage, to endow the model with the capability of pinpointing the latent difference in viewpoint change, Hierarchical Representation Interaction (HRI) models difference from coarse to fine representations through the Semantic Matcher and Change Amplifier module. The proposed approach outperforms the state-of-the-art methods with an encouraging performance on the existing change captioning benchmarks. Our code is available athttps://github.com/yueshengbin/I3N. Shengbin Yue, Yunbin Tu, Liang Li 0003, Shengxiang Gao, Zhengtao Yu 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency ReasoningabstractIterative language-based image manipulation aims to edit images step by step according to user's linguistic instructions. The existing methods mostly focus on aligning the attributes and appearance of new-added visual elements with current instruction. However, they fail to maintain consistency between instructions and images as iterative rounds increase. To address this issue, we propose a novel Long and Short term consistency reasoning Generative Adversarial Network (LS-GAN), which enhances the awareness of previous objects with current instruction and better maintains the consistency with the user's intent under the continuous iterations. Specifically, we first design a Context-aware Phrase Encoder (CPE) to learn the user's intention by extracting different phrase-level information about the instruction. Further, we introduce a Long and Short term Consistency Reasoning (LSCR) mechanism. The long-term reasoning improves the model on semantic understanding and positional reasoning, while short-term reasoning ensures the ability to construct visual scenes based on linguistic instructions. Extensive results show that LS-GAN improves the generation quality in terms of both object identity and position, and achieves the state-of-the-art performance on two public datasets. Gaoxiang Cong 0001, Liang Li 0003, Zhenhuan Liu, Yunbin Tu, Weijun Qin, Shenyuan Zhang, Chengang Yan, Bin Jiang 0011 |
ACM Multimedia | 4 |
| 2022 | Long Short-Term Relation Transformer With Global Gating for Video CaptioningabstractVideo captioning aims to generate a natural language sentence to describe the main content of a video. Since there are multiple objects in videos, taking full exploration of the spatial and temporal relationships among them is crucial for this task. The previous methods wrap the detected objects as input sequences, and leverage vanilla self-attention or graph neural network to reason about visual relations. This cannot make full use of the spatial and temporal nature of a video, and suffers from the problems of redundant connections, over-smoothing, and relation ambiguity. In order to address the above problems, in this paper we construct a long short-term graph (LSTG) that simultaneously captures short-term spatial semantic relations and long-term transformation dependencies. Further, to perform relational reasoning over the LSTG, we design a global gated graph reasoning module (G3RM), which introduces a global gating based on global context to control information propagation between objects and alleviate relation ambiguity. Finally, by introducing G3RM into Transformer instead of self-attention, we propose the long short-term relation transformer (LSRT) to fully mine objects' relations for caption generation. Experiments on MSVD and MSR-VTT datasets show that the LSRT achieves superior performance compared with state-of-the-art methods. The visualization results indicate that our method alleviates problem of over-smoothing and strengthens the ability of relational reasoning. Liang Li 0003, Xingyu Gao 0001, Jincan Deng, Yunbin Tu, Zhengjun Zha, Qingming Huang |
IEEE Trans. Image Process. | 4 |
| 2022 | I2Transformer: Intra- and Inter-Relation Embedding Transformer for TV Show CaptioningabstractTV show captioning aims to generate a linguistic sentence based on the video and its associated subtitle. Compared to purely video-based captioning, the subtitle can provide the captioning model with useful semantic clues such as actors’ sentiments and intentions. However, the effective use of subtitle is also very challenging, because it is the pieces of scrappy information and has semantic gap with visual modality. To organize the scrappy information together and yield a powerful omni-representation for all the modalities, an efficient captioning model requires understanding video contents, subtitle semantics, and the relations in between. In this paper, we propose an Intra- and Inter-relation Embedding Transformer (I2Transformer), consisting of an Intra-relation Embedding Block (IAE) and an Inter-relation Embedding Block (IEE) under the framework of a Transformer. First, the IAE captures the intra-relation in each modality via constructing the learnable graphs. Then, IEE learns the cross attention gates, and selects useful information from each modality based on their inter-relations, so as to derive the omni-representation as the input to the Transformer. Experimental results on the public dataset show that the I2Transformer achieves the state-of-the-art performance. We also evaluate the effectiveness of the IAE and IEE on two other relevant tasks of video with text inputs,i.e., TV show retrieval and video-guided machine translation. The encouraging performance further validates that the IAE and IEE blocks have a good generalization ability. The code is available athttps://github.com/tuyunbin/I2Transformer. Yunbin Tu, Liang Li 0003, Li Su 0003, Shengxiang Gao, Chenggang Yan 0001, Zhengjun Zha, Zhengtao Yu 0001, Qingming Huang |
IEEE Trans. Image Process. | 1 |
| 2021 | R\^3Net: Relation-embedded Representation Reconstruction Network for Change CaptioningabstractChange captioning is to use a natural language sentence to describe the fine-grained disagreement between two similar images.Viewpoint change is the most typical distractor in this task, because it changes the scale and location of the objects and overwhelms the representation of real change.In this paper, we propose a Relation-embedded Representation Reconstruction Network (R 3 Net) to explicitly distinguish the real change from the large amount of clutter and irrelevant changes.Specifically, a relation-embedded module is first devised to explore potential changed objects in the large amount of clutter.Then, based on the semantic similarities of corresponding locations in the two images, a representation reconstruction module (RRM) is designed to learn the reconstruction representation and further model the difference representation.Besides, we introduce a syntactic skeleton predictor (SSP) to enhance the semantic interaction between change localization and caption generation.Extensive experiments show that the proposed method achieves the state-of-the-art results on two public datasets 1 . Yunbin Tu, Liang Li 0003, Chenggang Yan 0001, Shengxiang Gao, Zhengtao Yu 0001 |
EMNLP (1) | 1 |
| 2021 | Enhancing the alignment between target words and corresponding frames for video captioning
Yunbin Tu, Shengxiang Gao, Zhengtao Yu 0001 |
Pattern Recognit. | 1 |
| 2020 | STAT: Spatial-Temporal Attention Mechanism for Video CaptioningabstractVideo captioning refers to automatic generate natural language sentences, which summarize the video contents. Inspired by the visual attention mechanism of human beings, temporal attention mechanism has been widely used in video description to selectively focus on important frames. However, most existing methods based on temporal attention mechanism suffer from the problems of recognition error and detail missing, because temporal attention mechanism cannot further catch significant regions in frames. In order to address above problems, we propose the use of a novel spatial-temporal attention mechanism (STAT) within an encoder-decoder neural network for video captioning. The proposed STAT successfully takes into account both the spatial and temporal structures in a video, so it makes the decoder to automatically select the significant regions in the most relevant temporal segments for word prediction. We evaluate our STAT on two well-known benchmarks: MSVD and MSR-VTT-10K. Experimental results show that our proposed STAT achieves the state-of-the-art performance with several popular evaluation metrics: BLEU-4, METEOR, and CIDEr. Chenggang Yan 0001, Yunbin Tu, Xingzheng Wang, Yongbing Zhang 0002, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai |
IEEE Trans. Multim. | 2 |
| 2020 | Corrections to "STAT: Spatial-Temporal Attention Mechanism for Video Captioning"abstractPresents corrections to affiliations in the above named paper. Chenggang Yan 0001, Yunbin Tu, Xingzheng Wang, Yongbing Zhang 0002, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai |
IEEE Trans. Multim. | 2 |
| 2017 | Video Description with Spatial-Temporal AttentionabstractTemporal attention has been widely used in video description to adaptively focus on important frames. However, most existing methods based on temporal attention suffer from the problems of recognition error and detail missing, because only coarse frame-level global features are employed. Inspired by recent successful work in image description using spatial attention, we propose a spatial-temporal attention (STAT) method to address such problems. In particular, first, we take advantage of object-level local features to address the problem of detail missing. Second, the STAT method further selects relevant local features by spatial attention and then attend to important frames by temporal attention to recognize related semantics. The proposed two-stage attention mechanism can recognize the salient objects more precisely with high recall and automatically focus on the most relevant spatial-temporal segments given the sentence context. Extensive experiments on two well-known benchmarks suggest that STAT method outperforms the state-of-the-art methods on MSVD with BLEU4 score 0.511, and achieves superior BLEU4 score 0.374 on MSR-VTT-10K. Compared to the method without local features, the relative improvements derived from our STAT method are 10.1% and 0.8% respectively on two benchmarks. Compared to the method using only temporal attention, the relative improvements derived from our STAT method are 18.3% and 9.0% respectively on two benchmarks. Yunbin Tu, Xishan Zhang, Bingtao Liu, Chenggang Yan 0001 |
ACM Multimedia | 1 |