EDBT 2026 Demo / reviewers in the wild / expert
Bing-Kun Bao
dblp:61/8802 · also Bingkun Bao
· DBLP profile ↗
121ranked-venue papers
12as first author
79since 2021 · last 2026
0000-0001-5956-831XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 93 · 10 first-author · 65 since 2021Artificial intelligence and machine learning · 32 · 3 first-author · 19 since 2021Computer networks · 14 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Social Event Prediction via Fourier Graph LearningabstractSocial event prediction has garnered increasing attention in web-centered society. Most existing studies represent web-based event stream as chronological graph sequences, then leverage RNNs and GNNs to model temporal and relational patterns. However, this paradigm is inherently flawed: (1) RNNs struggle to capture long-term temporal dependencies, ignoring those temporally distant but influential events. (2) Spatio-temporal GNNs exhibit high computational complexity on large-scale real-time event streams, which hinders their web applications. To this end, we explore a novel paradigm called Fourier Graph Learning from the perspective of frequency domain. Specifically, we first define a novel data structure called Fourier Graph (FG). In FG, both nodes and edges are complex vectors, with real part encoding semantics and imaginary part representing semantic-specific temporal patterns. These temporal patterns are obtained by semantic-aware frequency filter, which utilizes semantics as guidance to adaptively incorporates both long-term dependency and short-term dynamic. Based on FG, we further propose Fourier Graph Neural Network (FGNN). It replaces time-domain convolution with frequency-domain multiplication for efficient aggregation. FGNN also includes a complex-valued event decoder, which fully leverages semantics and temporal patterns from complex space to predict future event probabilities. Extensive experiments show our superior performance with higher accuracy, less complexity and better interpretability compared with baselines. Mingjie Qiu, Zhiyi Tan 0002, Bing-Kun Bao |
WWW | 3 |
| 2026 | Towards Rare Social Event Prediction via Mediator LearningabstractRare social events are infrequent yet influential incidents. Predicting such events is practically significant yet inherently challenging due to their extreme scarcity in web-based event stream. Existing studies view this task as an imbalanced classification problem and adopt static rebalancing methods to mitigate scarcity. However, they (1) ignore inter-event dependency that represents the interactions between different event streams, failing to capture precursors that lead to rare events and fundamentally limits performance. They (2) overlook intra-event dependency between different time points within single event stream, which prevents the model from adapting to shifting event patterns and degrades its generalization ability. To this end, we propose a novel Mediator Learning (ML) framework, which introduces mediators to explicitly model complex dependencies within web-based event streams. Specifically, we propose (1) Precursor Event Router (PER) that utilizes an information-theoretic routing approach to extract precursor events as mediators from massive event streams. Based on extracted mediators, (2) Conditional Hierarchical Graph Network (CHG) is introduced to model observed events, mediators and rare events into bottom-up graph levels, where its upper-level propagation is conditioned on bottom-level probability distribution. Jointly, these two modules decompose the imbalanced task into two more balanced stages, which not only mitigates the scarcity of rare events but also explicitly model inter-event dependencies, so as to capture precursor events leading to target rare events. Finally, we design (3) Adaptive Information Regularizer (AIR) to optimize the two stages. It dynamically adjusts the information flow between two stages, which models intra-event dependencies and facilitate adaption to drifting event patterns. We theoretically reveal the effectiveness of ML by framing it within Information Bottleneck (IB) principle. Extensive experiments show our superior accuracy and interpretability compared with SOTA. Mingjie Qiu, Zhiyi Tan 0002, Bing-Kun Bao |
WWW | 3 |
| 2026 | Deep Orientational Representation Learning for Ordinal RegressionabstractOrdinal regression aims to predict ordered classes. Existing methods mainly focus on label distribution shapes and feature distance relationships, while the directional characteristics in the representation space remain underexplored. In this paper, we propose deep orientational representation learning (ORL), aiming to ensure the trajectory of features sequentially connected by ordinal categories approximates a geodesic. We treat the output layer weights as ordinal prototypes and introduce two constraints, the co-directional constraint and the counter-directional constraint. They operate by constraining the angles between pairs of vectors. The former minimizes the angle between vectors with matching start and end categories, while the latter maximizes the angle between vectors whose start categories are the same but whose end categories are on opposite sides. The two constraints optimize the representation from different ordinal directions. ORL is extended to a multi-prototype setting (MORL) to mitigate misalignment between features and oriented prototypes caused by large intra-class variations. Theoretical analysis links ORL to distribution unimodality and distance orderliness, highlighting its advantages. The effectiveness of ORL (MORL) is demonstrated on various tasks including facial age estimation, historical image dating, and aesthetic quality assessment. Gengyun Jia, Xin Ma 0031, Bing-Kun Bao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Goal-Guided Prompting With Adaptive Modality Selection for Efficient Assembly Activity Anticipation in Egocentric VideosabstractWith the functions of egocentric observation and multimodal perception equipped in augmented reality (AR) devices, the next generation of smart assistants has the potential to reduce human labor and enhance execution efficiency in assembly tasks. Among diverse assembly activity understanding tasks, anticipating the near future activities is crucial yet challenging, which can assist humans or agents to actively plan and engage in interactions with the environment. However, the existing egocentric activity anticipation methods still struggle to achieve a decent trade-off between accuracy and computational efficiency, hindering them to be deployed in practical applications. To address this dilemma, in this paper, we propose a goal-guided prompting framework with adaptive modality selection (GP-AMS), for assembly activity anticipation in egocentric videos. For bridging the semantic gap between the historical observations and unobserved future activities, we inject the inferred high-level goal clues into the constructed prompts, which are further utilized to guide a pre-trained vision-language (V-L) model to compensate relevant semantics of unseen future. Moreover, a mask-and-predict strategy is adopted with two imposed constraints, i.e., casual masking and probabilistic token-dropping, to mine the intrinsic associations between the assembly activities within a specific procedure. For maintaining the benefits of exploiting multimodal information while avoiding extensively increasing the computational burdens, an adaptive modality selection strategy is designed to train a policy network, which learns to dynamically decide which modalities should be sampled for processing by the anticipation model on a per observation time-step basis. By allocating major computation to the selected indicative modalities on-the-fly, the efficiency of the overall model can be improved, thus paving the way for feasibility on real-world devices. Extensive experimental results on two public data sets validate that the proposed method yields not only consistent improvements in anticipation accuracy, but also significant savings in computation budgets. Tianshan Liu, Bing-Kun Bao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | An Episode Memory-Guided Dual-Stage Framework for Long-Form Video Temporal GroundingabstractVideo temporal grounding (VTG) aims to localize video moments that are semantically related to a given natural language query. In spite of recent progress in short-form videos, research on VTG in long-form videos (e.g., hours long) remains highly demanded yet underexplored. Existing methods predominantly adopt sliding window-based or multi-scale anchor-based strategies to generate temporal proposals, which require time-consuming post-processing or are independent of video content, thereby limiting their performance and efficiency. To address this dilemma, in this paper, we propose an episode memory-prompted (EMP) two-stage framework for temporal grounding in long-form videos. Specifically, the first stage generates a set of dynamic episode memories, which explicitly summarize various activities occurring throughout the lengthy video. An unsupervised memory learning paradigm is formulated by imposing discriminability and diversity constraints, eliminating the reliance on additional activity-instance annotations. Then, in the second stage, based on the supplement of frame-level detailed content and the guidance of a language query, the augmented memory prompts function as anchors for efficiently regressing the refined boundaries of the target video moment. Extensive experimental results on two public long-form video data sets, i.e., MAD and Ego4d, validate that the proposed EMP framework saves more than 8.5% trainable parameters and 13.9% FLOPs, while still achieving comparable performance with existing methods. Tianshan Liu, Bing-Kun Bao, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Question Understanding and Temporality Guiding for Video Question AnsweringabstractVideo Question Answering (VideoQA) aims to answer a question based on the content of a given video. Recent methods adapt image-text pre-trained models to the VideoQA task by designing learnable temporal modules within the image encoder. However, these methods struggle to fully comprehend the questions and effectively extract temporal information due to 1) over-reliance on candidate answers and 2) lack of explicit temporal modeling. Specifically, since the question is fixed in different question-answer pairs, existing models tend to focus on the varying candidate answers. Moreover, existing methods merely utilize the classification loss to constrain the confidence of candidate answers, failing to differentiate the effectiveness of temporal information and to explicitly guide temporal modeling. In this paper, we introduce the Question Understanding and Temporality Guiding (QU-TG) method to address the aforementioned limitations. To reduce over-reliance on candidate answers, we propose providing diverse questions through question selection and enhancing the model's comprehensive understanding of questions through question-video matching. To conduct explicit temporal modeling guiding, we propose negative video prevention and positive video guidance to conduct explicit temporal modeling guiding. Negative video prevention incorporates a prevention loss to discourage the model from making predictions based on erroneous temporal cues, whereas positive video guidance utilizes classification loss to encourage the model to derive correct answers from positive videos. Extensive experiments on the NExT-QA, IntentQA, STAR-QA, and Causal-VidQA datasets demonstrate the effectiveness and generalization of our method. Sisi You, Bing-Kun Bao |
IEEE Trans. Multim. | 3 |
| 2026 | ColView: Consistent Text-Guided Grayscale Scene Colorization From Multi-View ImagesabstractThe colorization of scenes from multi-view grayscale images plays a crucial role in applications such as augmented reality and virtual exhibitions. Existing methods combine NeRF with an automatic colorization model, averaging multiple colorized patches to reduce inconsistency. However, they still face three key limitations: (1) Current methods cannot produce diverse colorization results due to the lack of multimodal conditional inputs, (2) They struggle to maintain multi-view consistency caused by unreliable geometric correspondence and ineffective propagation mechanisms, and (3) Computational inefficiency from NeRF's dense ray sampling and numerical integration. In this paper, we propose ColView, a unified framework for text-guided grayscale scene colorization that achieves both automatic and controllable colorization of grayscale scenes from multi-view grayscale images. First, for flexible color control, we leverage text description as the input to guide the colorization process, which allows users to specify desired colors through natural language descriptions. Second, to ensure multi-view consistency, we introduce a multi-view consistent colorization module that explicitly models dependencies between different views. This module follows three key steps: cross-view attention mechanism for collaborative key-view colorization, feature matching for inter-view correspondence establishment, and correspondence-guided feature propagation. Third, to improve computational efficiency, we adopt 3D Gaussian Splatting as our underlying representation. This explicit point-based representation renders significantly faster than NeRF. Extensive experimental results demonstrate that our method achieves superior visual quality and computational efficiency. Our code and models are publicly available athttps://github.com/ChchNiu/ColView. Chaochao Niu, Ming Tao 0002, Bing-Kun Bao, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2026 | FoodDiff: A Collaborative Relationship Perception Framework for Food Image Synthesis Using Diffusion ModelsabstractFood image generation is a typical application of text-to-image (T2I) models. The core difference between food image synthesis and other T2I tasks is that there exist complex collaborative relationships among ingredients, cooking actions, and food images, which determine the appearance of dishes. However, existing food image generation models generally ignore or fail to sufficiently utilize such collaborative relationships, which hinders the model from precisely perceiving the shapes and details of food. Furthermore, the pre-training distribution of T2I models is usually noisy and differs from the user-preferred food feature distributions, resulting in deviations from human aesthetics. To address the above issues, we proposeFoodDiff, a collaborative relationship-aware diffusion model for food image generation, which consists of three key components: (1) To perceive collaborative relationships, we propose a collaborative relation module to extract these relations and inject them into the image generation process. (2) To sufficiently interact with the relationships between recipe semantics and food representations, we propose a recipe fusion fine-tuning module to precisely fuse recipe semantics with visual features and fine-tune the pre-trained model. (3) To make the pre-training feature distribution conform to human preference, we introduce an image reward feedback mechanism to optimize the aesthetics of food images. In addition, we propose a high-quality food dataset named Food-Aesthetic with exquisite plates and elaborate annotations. Extensive experiments and human evaluations show that FoodDiff has superior image aesthetics and semantic consistency. Mengling Xu, Sisi You, Bing-Kun Bao |
IEEE Trans. Multim. | 3 |
| 2026 | Light Field Reconstruction Using Multi-orientation Epipolar Plane ImagesabstractLight field reconstruction is one of the most important techniques for future glass-free 3D media production. However, current techniques suffer from low view-consistency and fixed patterns of view-trajectory. This article presents the 3D Multi-orientation Epipolar Plane Image (MOEPI) representation for high-quality light field reconstruction with both the inter-view and extra-view settings. Each layer in MOEPI is composed of EPI lines with a fixed orientation. To infer the MOEPI, a new Multi-reference Focal Stack (MRFS) intermediate representation is proposed. The optimization of MOEPI could be regarded as the problem of the most-focused content extraction from the MRFS. This optimization is implemented with a 3D U-shaped network. We also propose the LPIPS-EPI metric for evaluating the view-consistency. Experiments on light fields with both high and low signal-to-noise ratios demonstrate that the proposed MOEPI representation could synthesize high-quality light fields especially in occlusion or un-captured areas. Hao Zhu 0005, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2026 | MyGO: Modality-incomplete Fake News Video Detection via Prompt-assisted Modality Disentangling ModelabstractFake news video detection has become a pressing concern with the growth of short video platforms. However, previous studies have primarily focused on videos with modality-complete data, failing to handle the uncertain missing modality issue in real-world applications. They fall short in two key aspects: (1) Highly coupled feature fusion hinders the model to learn intra- and inter-modality dependencies, making it difficult to form robust multimodal representations when facing uncertain modality missing. (2) Excessive reliance on discriminative modality combinations toward fake news, which leads to inferior performance on other modality combinations. To this end, we propose a novel model for modality-incomplete fake news video detection called MyGO. It contains three modules: (1) Caption-guided Keyframe Attention (CKA) leverages embedded captions to guide feature extraction, which adaptively excludes irrelevant frames to enhance the learning of intra-modality dependencies, resulting in refined modality features. (2) Based on refined modality features from CKA, Modality Disentangling Network (MDN) is designed to decompose them into shared and specific parts, which captures fine-grained inter-modality dependencies effectively. These two kinds of dependencies help avoid coupled multimodal fusion and bridge information gaps caused by missing modalities. (3) Furthermore, missing prompts are newly introduced to explicitly mark modality combinations within each news video. By integrating missing prompts with aforementioned inter-modality dependencies within Prompt-assisted Modality Aligning (PMA) Module, we alleviate over-reliance on discriminative modality combinations and enhancing the representation of less discriminative ones. Extensive experiments showcase that MyGO achieves 3.79–4.85% improvements in accuracy, demonstrating its performance over state-of-the-art approaches under different missing conditions. Mingjie Qiu, Zhiyi Tan 0002, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2026 | AdaEdit: Adaptive Diffusion Model for Invisible Target Oriented Text-Conditioned Image EditingabstractText-conditioned image editing aims to modify a source image into a target image according to a specified text description, tackling two core challenges: locating target editing regions and ensuring consistency in non-target editing areas. Existing approaches utilize manual selection or cross-modal attention to define editing regions and deploy diffusion models to generate edited images. Despite these recent advancements, two problems remain. First, current methods fail to locate editing areas described in the text but invisible in the image. Second, they struggle to ensure spatial consistency in non-targeted regions due to the global noise addition along with excessive denoising during the diffusion process. To overcome these limitations, we propose AdaEdit, which comprises an adaptive mask localization module and an adaptive denoising strategy for text-conditioned image editing. AdaEdit can accurately identify the editing area via the measurement of cross-modal semantic mismatch, even when the visual details are not explicitly described in the text inputs. The adaptive denoising strategy applies varying noise levels to differentiate between targeted and non-targeted regions, enhancing the stability and consistency of the non-edited areas. Extensive experiments demonstrate that our proposed method achieves excellent performance on MS-COCO, MagicBrush, and Laion. We also expand our application to iterative editing tasks, thereby extending its utility for generalized editing scenarios. Yefei Sheng, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2026 | A Reciprocal Interaction Framework for Collaborative Temporal Grounding and Question Answering in Egocentric VideosabstractCollaborative Temporal Grounding and Question Answering (CTGQA) in egocentric videos enables users to inquire about past visual experiences and obtain corresponding temporal segments and answers. Existing CTGQA methods typically treat Video Temporal Grounding (VTG) and Video Question Answering (VQA) as separate tasks, overlooking their inherent semantic and temporal complementarity. As a result, VQA models often generate ambiguous answers due to the lack of precise temporal cues, while VTG models fail to fully exploit the high-level semantic information embedded in the answers. To address these limitations, we propose a Reciprocal Interaction Framework (RIF). RIF employs a two-branch interaction structure to enhance the performance of both VTG and VQA. RIF consists of two modules: Localization-Guided Answering (LGA) and Answer-Enhanced Temporal Grounding (AETG). The LGA module assists the VQA model in generating high-quality answers by highlighting relevant segments while minimizing the influence of irrelevant content. To mitigate model overconfidence, we propose a progressive feature fusion strategy that dynamically adjusts the weights of relevant segments, thus preventing localization errors. The AETG module leverages additional information embedded in the generated answer to improve VTG performance. Moreover, we employ a perplexity-based filtering strategy to ensure the reliability of the answer. Extensive experiments show that our framework performs well on the QAEGO4D and Ego4D-NLQ benchmarks. Tianshan Liu, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | InstantPainting: Expanding GANs for Efficient Text-Conditioned Image Generation PlatformabstractText-conditioned image generation enables cross-modal comprehension. Recent emergence of many platforms have found applications in diverse domains like assisted designing and video gaming. However, there still exist challenges in existing platforms due to their expensive training and time-consuming generation processes. In this paper, we introduce an efficient text-conditioned image generation platform, termed InstantPainting. Unlike existing platforms based on large-scale pre-trained diffusion models, InstantPainting expands generative adversarial networks (GANs) to achieve efficient generation by using only about three percent pre-training data of other platforms. Compared to existing platforms, InstantPainting achieves the following functions at a very low deployment cost and approximately 4 to 5 times faster generation speeds: (1) Multi-category and multi-size image generation (2) Image stylization and controlled generation (3) Creative generation, including the generation of poetry pictures and counterfactual images. The proposed platform provides web application implementations for PC and mobile, users can create high-quality images directly through the user interface. Bing-Kun Bao, Yefei Sheng, Jie Wang 0061, Sisi You |
AAAI | 1 |
| 2025 | BeFA: A General Behavior-driven Feature Adapter for Multimedia RecommendationabstractMultimedia recommender systems focus on utilizing behavioral information and content information to model user preferences. Typically, it employs pre-trained feature encoders to extract content features, then fuses them with behavioral features. However, pre-trained feature encoders often extract features from the entire content simultaneously, including excessive preference-irrelevant details.We speculate that it may result in the extracted features not containing sufficient features to accurately reflect user preferences. To verify our hypothesis, we introduce an attribution analysis method for visually and intuitively analyzing the content features. The results indicate that certain items’ content features exhibit the issues of information drift and information omission, reducing the expressive ability of features. Building upon this finding, we propose an effective and efficient general Behaviordriven Feature Adapter (BeFA) to tackle these issues. This adapter reconstructs the content feature with the guidance of behavioral information, enabling content features accurately reflecting user preferences. Extensive experiments demonstrate the effectiveness of the adapter across all multimedia recommendation methods. Qile Fan, Penghang Yu, Zhiyi Tan 0002, Bing-Kun Bao, Guanming Lu |
AAAI | 4 |
| 2025 | Leveraging Group Classification with Descending Soft Labeling for Deep Imbalanced RegressionabstractDeep imbalanced regression (DIR), where the target values have a highly skewed distribution and are also continuous, is an intriguing yet under-explored problem in machine learning. While recent works have already shown that incorporating various classification-based regularizers can produce enhanced outcomes, the role of classification remains elusive in DIR. Moreover, such regularizers (e.g., contrastive penalties) merely focus on learning discriminative features of data, which inevitably results in ignorance of either continuity or similarity across the data. To address these issues, we first bridge the connection between the objectives of DIR and classification from a Bayesian perspective. Consequently, this motivates us to decompose the objective of DIR into a combination of classification and regression tasks, which naturally guides us toward a divide-and-conquer manner to solve the DIR problem. Specifically, by aggregating the data at nearby labels into the same groups, we introduce an ordinal group-aware contrastive learning loss along with a multi-experts regressor to tackle the different groups of data thereby maintaining the data continuity. Meanwhile, considering the similarity between the groups, we also propose a symmetric descending soft labeling strategy to exploit the intrinsic similarity across the data, which allows classification to facilitate regression more effectively. Extensive experiments on real-world datasets also validate the effectiveness of our method. Ruizhi Pu, Gezheng Xu, Ruiyi Fang, Bing-Kun Bao, Charles Ling 0001, Boyu Wang 0004 |
AAAI | 4 |
| 2025 | Mind Individual Information! Principal Graph Learning for Multimedia RecommendationabstractGraph Neural Network (GNN)-based methods have recently emerged as effective approaches for multimedia recommendation. Typically, these methods employ message passing on the user-item interaction graph, and model user preferences by exploiting co-occurrence patterns. Despite their effectiveness, we argue that they insufficiently exploit the individual information, potentially limiting recommendation performance. To validate our argument, we first analyze existing methods from spectral graph theory. We identify that existing methods focus on capturing global structural features, but underutilize local structural features that convey individual information. Further detailed experiments reveal that such an underutilization leads to overly similar user preferences modeling. Furthermore, we propose a novel Principal Graph Learning (PGL) framework to address this issue. The idea is to enhance user preference modeling by effectively mining and utilizing principal local structural features. PGL first extracts the principal subgraph from the user-item interaction graph using two novel extraction operators: global-aware and local-aware subgraph extraction. It then employs message passing on the principal subgraph to comprehensively model user perference, with the aim of simultaneously capturing co-occurrence patterns and individual information. Compared to existing methods, PGL achieves an average performance improvement of 9%. Penghang Yu, Zhiyi Tan 0002, Guanming Lu, Bing-Kun Bao |
AAAI | 4 |
| 2025 | Causal Debiasing for Visual Commonsense ReasoningabstractVisual Commonsense Reasoning (VCR) refers to answering questions and providing explanations based on images. While existing methods achieve high prediction accuracy, they often overlook bias in datasets and lack debiasing strategies. In this paper, our analysis reveals co-occurrence and statistical biases in both textual and visual data. We introduce the VCR-OOD datasets, comprising VCR-OOD-QA and VCR-OOD-VA subsets, which are designed to evaluate the generalization capabilities of models across two modalities. Furthermore, we analyze the causal graphs and prediction shortcuts in VCR and adopt a backdoor adjustment method to remove bias. Specifically, we create a dictionary based on the set of correct answers to eliminate prediction shortcuts. Experiments demonstrate the effectiveness of our debiasing method across different datasets. Jiayi Zou, Gengyun Jia, Bing-Kun Bao |
ICASSP | 3 |
| 2025 | Spatial-Temporal Prior Knowledge Guidance for Long-term Action AnticipationabstractFor long-term action anticipation (LAA), the primary focus is on understanding the observed video content and anticipating future actions, including both the names of upcoming actions and their corresponding durations. This requires the model to fully grasp the patterns of action transitions. To facilitate this, we employ spatial-temporal prior knowledge to guide the LAA model in capturing these transition patterns, which is referred to as explicit learning. Additionally, we use a transformer structure which incorporates a parallel decoding mechanism in which mitigates error accumulation and can be regarded as implicit learning. Consequently, we propose a novel model that integrates explicit and implicit learning approaches, combining the advantages of both. On the benchmarks for long-term action anticipation, our method achieves state-of-the-art results on the 50Salads and Breakfast. Yiming Li 0008, Miao Ji, Sisi You, Bing-Kun Bao |
ICME | 4 |
| 2025 | CLGC: Continuous Layout Guidance for Consistent Text-to-Video EditingabstractText-to-Video (T2V) editing aims to produce temporally consistent videos aligned with text prompts, simultaneously reconstructing the original spatial structure. Existing methods rely on cross-attention maps generated from fixed text prompts, which lack sufficient spatial information, leading to inaccurate object positioning across frames. Additionally, existing methods only rely on the first or former frame to synthesize the current frame, offering limited viewpoint information and causing flickering artifacts. To address these issues, we propose CLGC, a training-free framework for continuous layout-guided T2V editing. First, we introduce semantic masks for continuous object position layout guidance, refining cross-attention maps and ensuring accurate object positioning across frames. Second, we adaptively integrate extra reference frames into the self-attention for the current frame synthesis, enhancing temporal consistency in the edited video. Finally, we integrate parallel null-text inversion to improve DDIM sampling, achieving accurate reconstruction results. Extensive experiments demonstrate that CLGC excels at attribute editing and shape transformation, confirming its effectiveness in T2V editing. Xuancheng Xu, Ming Tao 0002, Bing-Kun Bao |
ICME | 3 |
| 2025 | Test-Time Selective Adaptation for Uni-Modal Distribution Shift in Multi-Modal DataabstractModern machine learning applications are characterized by the increasing size of deep models and the growing diversity of data modalities. This trend underscores the importance of efficiently adapting pre-trained multi-modal models to the test distribution in real time, i.e., multi-modal test-time adaptation. In practice, the magnitudes of multi-modal shifts vary because multiple data sources interact with the impact factor in diverse manners. In this research, we investigate the the under-explored practical scenario uni-modal distribution shift, where the distribution shift influences only one modality, leaving the others unchanged. Through theoretical and empirical analyses, we demonstrate that the presence of such shift impedes multi-modal fusion and leads to the negative transfer phenomenon in existing test-time adaptation techniques. To flexibly combat this unique shift, we propose a selective adaptation schema that incorporates multiple modality-specific adapters to accommodate potential shifts and a “router” module that determines which modality requires adaptation. Finally, we validate the effectiveness of our proposed method through extensive experimental evaluations. Code available at https://github.com/chenmc1996/Uni-Modal-Distribution-Shift. Mingcai Chen, Baoming Zhang, Zongbo Han, Yanmeng Wang, Yuntao Du 0001, Bing-Kun Bao |
ICML | 8 |
| 2025 | Graph Prompts: Adapting Video Graph for Video Question AnsweringabstractDue to the dynamic nature in videos, it is evident that perceiving and reasoning about temporal information are the key focus of Video Question Answering (VideoQA). In recent years, several methods have explored relationship-level temporal modeling with graph-structured video representation. Unfortunately, these methods heavily rely on the question text, thus making it challenging to perceive and reason about video content that is not explicitly mentioned in the question. To address the above challenge, we propose Graph Prompts-based VideoQA (GP-VQA), which adopts a video-based graph structure for enhanced video understanding. The proposed GP-VQA contains two stages, i.e., pre-training and prompt tuning. In pre-training, we define the pretext task that requires GP-VQA to reason about the randomly masked nodes or edges in the video graph, thus prompting GP-VQA to learn the reasoning ability with video-guided information. In prompt-tuning, we organize the textual question into question graph and implement message passing from video graph to question graph, therefore inheriting the video-based reasoning ability from video graph completion to VideoQA. Extensive experiments on various datasets have demonstrated the promising performance of GP-VQA. Yiming Li 0008, Xiaoshan Yang, Bing-Kun Bao, Changsheng Xu |
IJCAI | 3 |
| 2025 | SCVBench: A Benchmark with Multi-turn Dialogues for Story-Centric Video UnderstandingabstractVideo understanding seeks to enable machines to interpret visual content across three levels: action, event, and story. Existing models are limited in their ability to perform high-level long-term story understanding, due to (1) the oversimplified treatment of temporal information and (2) the training bias introduced by action/event-centric datasets. To address this, we introduce SCVBench, a novel benchmark for story-centric video understanding. SCVBench evaluates LVLMs through an event ordering task decomposed into sub-questions leading to a final question, quantitatively measuring historical dialogue exploration. We collected 1,253 final questions and 6,027 sub-question pairs from 925 videos, constructing continuous multi-turn dialogues. Experimental results show that while closed-source GPT-4o outperforms other models, most open-source LVLMs struggle with story-centric video understanding. Additionally, our StoryCoT model significantly surpasses open-source LVLMs on SCVBench. SCVBench aims to advance research by comprehensively analyzing LVLMs' temporal reasoning and comprehension capabilities. Code can be accessed at https://github.com/yuanrr/SCVBench. Sisi You, Bing-Kun Bao |
IJCAI | 3 |
| 2025 | DToMA: Training-free Dynamic Token MAnipulation for Long Video UnderstandingabstractVideo Large Language Models (VideoLLMs) often require thousands of visual tokens to process long videos, leading to substantial computational costs, further exacerbated by visual token inefficiency. Existing token reduction and alternative video representation methods improve efficiency but often compromise comprehension abilities. In this work, we analyze the reasoning processes of VideoLLMs in multi-choice VideoQA task, identifying three reasoning stages—shallow, intermediate, and deep stages—that closely mimic human cognitive processing. Our analysis reveals specific inefficiencies at each stage: in shallow layers, VideoLLMs attempt to memorize all video details without prioritizing relevant content; in intermediate layers, models fail to re-examine uncertain content dynamically; and in deep layers, they continue processing video even when sufficiently confident. To bridge this gap, we propose DToMA, a training-free Dynamic Token MAnipulation method inspired by human adjustment mechanisms in three aspects: 1) Text-guided keyframe-aware reorganization to prioritize keyframes and reduce redundancy, 2) Uncertainty-based visual injection to revisit content dynamically, and 3) Early-exit pruning to halt visual tokens when confident. Experiments on 6 long video understanding benchmarks show that DToMA enhances both efficiency and comprehension, outperforming state-of-the-art methods and generalizing well across 3 VideoLLM architectures and sizes. Code is available at https://github.com/yuanrr/DToMA. Sisi You, Bing-Kun Bao |
IJCAI | 3 |
| 2025 | Retaining Temporal Semantics and Relation Topologies for Continual Weakly-Supervised Audio-Visual Video ParsingabstractTo achieve audio and visual action detections in a given video accompanying with only video-level labels, a group of weakly-supervised audio-visual video parsing methods have been explored. Throughout their training processes, the action categories are typically assumed to be static, which is not always satisfied. Consequently, these methods can not be employed to handle dynamic scenarios involving continuously growing novel classes. To alleviate the above issue, we introduce a novel Continual Weakly-Supervised Audio-Visual Video Parsing (C-WSAVVP) task, where maintaining the knowledge of historic categories remains the eternal topic. Distinctly, owing to the weakly-supervised and multi-modal characteristics, two core challenges are more obvious in C-WSAVVP: (1) Compared with the continual audio-visual video classification task, where distilling video-level coarse action semantic of trimmed videos is sufficient for mitigating catastrophic forgetting, C-WSAVVP has to retain more fine-grained temporal semantic information of untrimmed videos containing both actions and backgrounds. (2) The semantics of different actions generally exhibit a certain degree of correlation, which is beneficial for understanding related actions, but how to maintain the semantic correlations? To address the specific challenges, the Semantic Prototype-based Action Refinement (SPAR) and Inter-Class Relation Topology Preservation (IRTP) modules are explored, where the former devotes to utilizing various semantic prototypes to refine more reliable temporal action intervals for distillation and the latter focuses on retaining semantic correlations between different actions in both modalities during continual learning. Comprehensive experiments on our reconstructed C-LLP dataset demonstrate the effectiveness and generalization capability of our proposed method. Jie Fu 0004, Bing-Kun Bao |
ACM Multimedia | 2 |
| 2025 | D2Gaussian: Dynamic Control with Discretized 3D View Modeling for Text-Driven 3D Gaussian Splatting EditingabstractCurrent advances in text-driven 3D scene editing tasks typically render the 3D representations into multi-view images and modify the images with the text instructions. Context consistency across multiple views and cross-modal consistency in the single-view are the keys to effective 3D editing. Accordingly, existing methods introduce additional image constraints and apply pre-trained 2D editing models. However, they fix the same text instruction across all views and freeze the pre-trained 2D model for single-view editing, leading to deficient modeling of 3D scene views and results in inconsistent generations with visual artifacts. To address these limitations, we introduce a discretized 3D view modeling method and a diffusion-based multi-view consistent editing pipeline for text-driven 3D gaussian splatting editing, abbreviated as D2Gaussian. Specifically, our approach constructs a codebook that encodes continuous 3D view information into discrete token embeddings to model the spatial feature expressions. Then, the token embeddings are proposed to guide and finetune the diffusion-based image editing model with the dynamic addition of control conditions, yielding a multi-view consistent editing pipeline. Finally, we introduce a 3D editing dataset generation approach along with a 3D-CLIP-SIM metric to form a benchmark, 3D-MagicBrush, to provide more diverse evaluation scenarios for future 3D editing works. Experiments demonstrate that our method achieves better visual results and multi-view consistency than previous state-of-the-art methods. Yefei Sheng, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao |
ACM Multimedia | 4 |
| 2025 | Chain-of-Cooking: Cooking Process Visualization via Bidirectional Chain-of-Thought GuidanceabstractCooking process visualization is a promising task in the intersection of image generation and food analysis, which aims to generate an image for each cooking step of a recipe. However, most existing works focus on generating images of finished foods based on the given recipes, and face two challenges in visualizing the cooking process. First, the appearance of ingredients changes variously across cooking steps, it is difficult to generate the correct appearances of foods that match the textual description, leading to semantic inconsistency. Second, the current step might depend on the operations of previous step, it is crucial to maintain the contextual coherence of images in sequential order. In this work, we present a cooking process visualization model, called Chain-of-Cooking. Specifically, to generate correct appearances of ingredients, we present a Dynamic Patch Selection Module to retrieve previously generated image patches as references, which are most related to current textual contents. Furthermore, to enhance the coherence and keep the rational order of generated images, we propose a Semantic Evolution Module and a Bidirectional Chain-of-Thought (CoT) Guidance. To better utilize the semantics of previous texts, the Semantic Evolution Module establishes the semantical association between latent prompts and current cooking step, and merges it with the latent features. Then the CoT Guidance updates the merged features to guide the current cooking step remain coherent with the previous step. Moreover, we construct a dataset named CookViz, consisting of intermediate image-text pairs for the cooking process. Quantitative and qualitative experiments show that our method outperforms existing methods in generating coherent and semantic consistent cooking process. Mengling Xu, Ming Tao 0002, Bing-Kun Bao |
ACM Multimedia | 3 |
| 2025 | DMC3: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question AnsweringabstractEgocentric Video Question Answering (Egocentric VideoQA) plays an important role in egocentric video understanding, which refers to answering questions based on first-person videos. Although existing methods have made progress through the paradigm of pre-training and fine-tuning, they ignore the unique challenges posed by the first-person perspective, such as understanding multiple events and recognizing hand-object interactions. To deal with these challenges, we propose a Dual-Modal Counterfactual Contrastive Construction (DMC3) framework, which contains an egocentric videoqa baseline, a counterfactual sample construction module and a counterfactual sample-involved contrastive optimization. Specifically, We first develop a counterfactual sample construction module to generate positive and negative samples for textual and visual modalities through event description paraphrasing and core interaction mining, respectively. Then, We feed these samples together with the original samples into the baseline. Finally, in the counterfactual sample-involved contrastive optimization module, we apply contrastive loss to minimize the distance between the original sample features and the positive sample features, while maximizing the distance from the negative samples. Experiments show that our method achieve 52.51% and 46.04% on the normal and indirect splits of EgoTaskQA, and 13.2% on QAEGO4D, both reaching the state-of-the-art performance. Jiayi Zou, Bing-Kun Bao, Changsheng Xu |
ACM Multimedia | 3 |
| 2025 | In-Context Fully Decentralized Cooperative Multi-Agent Reinforcement LearningabstractIn this paper, we consider fully decentralized cooperative multi-agent reinforcement learning, where each agent has access only to the states, its local actions, and the shared rewards. The absence of information about other agents' actions typically leads to the non-stationarity problem during per-agent value function updates, and the relative overgeneralization issue during value function estimation. However, existing works fail to address both issues simultaneously, as they lack the capability to model the agents' joint policy in a fully decentralized setting. To overcome this limitation, we propose a simple yet effective method named Return-Aware Context (RAC). RAC formalizes the dynamically changing task, as locally perceived by each agent, as a contextual Markov Decision Process (MDP), and addresses both non-stationarity and relative overgeneralization through return-aware context modeling. Specifically, the contextual MDP attributes the non-stationary local dynamics of each agent to switches between contexts, each corresponding to a distinct joint policy. Then, based on the assumption that the joint policy changes only between episodes, RAC distinguishes different joint policies by the training episodic return and constructs contexts using discretized episodic return values. Accordingly, RAC learns a context-based value function for each agent to address the non-stationarity issue during value function updates. For value function estimation, an individual optimistic marginal value is constructed to encourage the selection of optimal joint actions, thereby mitigating the relative overgeneralization problem. Experimentally, we evaluate RAC on various cooperative tasks (including matrix game, predator and prey, and SMAC), and its significant performance validates its effectiveness. Bing-Kun Bao, Yang Gao 0001 |
NeurIPS | 2 |
| 2025 | A Memory-Assisted Knowledge Transferring Framework with Curriculum Anticipation for Weakly Supervised Online Activity Detection
Tianshan Liu, Kin-Man Lam 0001, Bing-Kun Bao |
Int. J. Comput. Vis. | 3 |
| 2025 | LD4MRec: simplifying and powering diffusion model for multimedia recommendation
Jiarui Zhu, Penghang Yu, Zhiyi Tan 0002, Bing-Kun Bao |
Multim. Syst. | 5 |
| 2025 | SEMACOL: Semantic-enhanced multi-scale approach for text-guided grayscale image colorization
Chaochao Niu, Ming Tao 0002, Bing-Kun Bao |
Pattern Recognit. | 3 |
| 2025 | DMRFlow: 4D Radar Scene Flow Estimation With Decoupled Matching and RefinementabstractScene flow estimation from 4D radar sensors has become increasingly popular in recent years. In this paper, we propose a matching and refinement decoupling method to estimate scene flow from 4D radar point clouds. Since 4D radar point clouds are much sparser and noisier than LiDAR point clouds, it is challenging to effectively establish correspondences between two frames and properly refine flow fields in the 3D space. To address this issue, we present decoupled correlation fields and decoupled flow fields for scene flow estimation, named DMRFlow. On the one hand, we propose a position-velocity decoupled matching approach that decouples the positional features from the velocity features of two adjacent point clouds and matches them separately. On the other hand, we design a dynamic-static decoupled refinement approach that splits initial flow fields into two groups according to motion segmentation maps and refines them separately. By integrating the matching and refinement decoupling method, our DMRFlow is able to effectively reduce mutual interference between different features during the matching and refinement process. We evaluate the proposed approach on the View-of-Delft (VoD) dataset. Experimental results show that DMRFlow yields competitive performance in autonomous driving scenarios compared to recent 4D radar scene flow estimation methods. Mingliang Zhai, Bing-Kun Bao, Xuezhi Xiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Scene Flow Estimation for Autonomous Driving via Correlation Compensation and Initial Motion CheckabstractScene flow estimation from LiDAR sensors is a crucial task for dynamic environmental perception in autonomous driving scenarios. Recently, self-supervised approaches have gained attention for their ability to reduce the burden of point-wise annotation. Although existing methods have been able to generate initial flow fields by constructing point-to-point correspondences between adjacent frames of point clouds, the reliability of correlation extraction and initial motion measurement has not been adequately considered. To address this problem, we propose a novel deep neural network to estimate scene flow from LiDAR sensors. Unlike previous works, our approach incorporates a Statistical-based Correlation Compensation Module (SCCM) that leverages statistical features to capture more reasonable correspondences. Furthermore, we design a Holistic Correlation Compensation Module (HCCM) to capture the overall correspondence that reflects most of the rigid motion in dynamic environments. In addition, an Initial Motion Check Mechanism (IMCM) is introduced to calibrate the initial flow and provide more reliable motion priors for subsequent flow refinement. Extensive experimental results on public scene flow benchmarks show that the proposed approach achieves competitive performance in autonomous driving scenarios. Mingliang Zhai, Bing-Kun Bao, Xuezhi Xiang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | SVSRD: Spatial Visual and Statistical Relation Distillation for Class-Incremental Semantic SegmentationabstractClass-incremental semantic segmentation (CISS) aims to incrementally learn novel classes while retaining the ability to segment old classes, and suffers catastrophic forgetting since the old-class labels are unavailable. Most existing methods typically impose strict constraints on the consistency between the extracted features or output logits of each pixel from old and current models in an attempt to prevent forgetting through knowledge distillation (KD), which 1) results in a significant transfer of redundant knowledge while limiting the restoration of old classes (rigidity) due to potentially overlooking essential knowledge extraction, and 2) imposes strong constraints at the pixel level making it challenging for the model to learn novel classes (plasticity). To solve the above limitations, we propose a novel Spatial Visual and Statistical Relation Distillation (SVSRD) by applying multi-scale visual and statistical position relation distillation for CISS, which enjoys several merits. First, we introduce a region-based similarity matrix and impose a consistency constraint between current and old models, which preserves the essential visual knowledge to enhance the rigidity. Second, we propose a novel statistical feature calculation algorithm to investigate the distribution of the data and further preserve the rules of statistics through statistical consistency, which also promotes the model on the novel-class learning for improving the plasticity. Finally, the aforementioned constraints are jointly applied in multiple scales to alleviate old-class forgetting and enhance novel-class learning. Extensive experiments on Pascal-VOC 2012 and ADE20 K demonstrate that the proposed approach performs favorably against the state-of-the-art CISS methods. Yuyang Chang, Yifan Jiao, Bing-Kun Bao |
IEEE Trans. Multim. | 3 |
| 2025 | CookGALIP: Recipe Controllable Generative Adversarial CLIPs With Sequential Ingredient Prompts for Food Image GenerationabstractGenerating food images from recipes is a challenging task in food analysis, as recipes contain lengthy texts far beyond the semantic information in food images, making it difficult to align the features of two modalities. Existing studies usually concatenate the representations of ingredients and cooking instructions directly, and use the concatenated representations to generate food images through generative adversarial networks (GANs). However, previous models generally ignore the sequential information contained in complicated procedural instructions, which leads to semantic inconsistency between recipes and generated food images. Furthermore, it is still difficult for current models to distinguish and control fine-grained features, causing the entangled ingredient features in food images. To this end, we propose CookGALIP, which strengthens semantic consistency and controllability for food image generation. Based on the recently proposed text-to-image framework GALIP, two modules are specially designed: 1) To incorporate the sequential relationships into the food image generation process, we propose a Recipe Fusion Module (RFM) to fuse the semantics of cooking instructions, so as to balance the semantic complexity between modalities and improve the semantic consistency of recipes and generated food images. 2) To distinguish and control the fine-grained ingredient features, we introduce the Ingredient Control Module (ICM) to generate sequential ingredient prompts, which enables more refined control over the recipe-to-food synthesis process. Experimental results on Recipe1M and Vireo Food-172 datasets show that the proposed model outperforms the state-of-the-art methods. Mengling Xu, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao, Changsheng Xu |
IEEE Trans. Multim. | 4 |
| 2025 | Relation Inference Enhancement Network for Visual Commonsense ReasoningabstractWhen presented with a question regarding an image, Visual Commonsense Reasoning (VCR) offers not only a correct answer but also a rationale to justify the answer. Existing methods simply combine features from multiple modalities onto a shared dimension space, which doesn't align with human reasoning patterns, resulting in inadequate cross-modal and intra-modal reasoning behaviors. On the one hand, inadequate cross-modal reasoning arises from existing models relying on semantic correlations between answers and rationales in both textual modalities rather than the generative process of human reasoning from visual to textual modality. On the other hand, inadequate intra-modal reasoning arises from the incapacity of existing models to leverage previously acquired object relations beyond current observations like humans. To this end, we propose a novel Relation Inference Enhancement Network (RIE-Net), which enhances reasoning ability based on cross-modal image analysis and introduces intra-modal relational reasoning modules to memorize reasoning knowledge. To enhance the cross-modal association between images and rationales, RIE-Net introduces a cross-modal image analysis module, which eliminates language bias between answers and rationales by generating rationale from images. In addition, to comprehend and retain relational knowledge, RIE-Net introduces intra-modal relational reasoning modules to capture prior knowledge associated with various object categories and enhance the model's understanding of visual-spatial relationships. Quantitative and qualitative evaluations of the public VCR dataset demonstrate that our approach performs favorably against state-of-the-art methods. Gengyun Jia, Bing-Kun Bao |
IEEE Trans. Multim. | 3 |
| 2025 | Replay-Based Incremental Object Detection With Local Response ExplorationabstractIncremental object detection (IOD) aims to train an object detector on non-stationary data streams without forgetting previous knowledge. Prevalent replay-based methods keep a buffer composed of carefully selected instances towards this goal. However, due to the limited storage space and uniform feature distribution, existing methods are prone to overfit on replayed instances, leading to poor generalization on diverse test data. Additionally, the imbalance in data quantity makes the detector fail to distinguish old and new classes that are visually similar, introducing bias toward new classes. To enhance the diversity of stored instances and eliminate bias, we propose a Local Response Exploration (LRE) framework, which comprises three modules. First, Region-Entropy Instance Selector (REIS) introduces a novel metric to assess instance diversity based on the entropy of local responses. Second, Confusion-Guided Instance Replay (CGIR) replaces the previous random replay approach by replaying specific old class instances based on class similarity, ensuring that parameters for similar new and old classes are updated together, thereby mitigating bias and helping mining discriminative patterns. Third, Confusion-Aware Region Segregation (CARS) adaptively differentiates biased regions from other regions based on local responses, reducing bias toward new classes while preserving relationships between new and old classes. Extensive evaluations on Pascal-VOC and MS COCO datasets demonstrate that our approach outperforms State-of-the-Art methods in incremental object detection. Yifan Jiao, Bing-Kun Bao |
IEEE Trans. Multim. | 3 |
| 2025 | Unified Text-Image Space Alignment with Cross-Modal Prompting in CLIP for UDAabstractUnsupervised Domain Adaptation (UDA) aims to transfer models trained on a labeled source domain to an unlabeled target domain. Due to the excellent generalization ability of Vision Language Models (VLMs) such as CLIP in downstream tasks, most recent methods apply CLIP to UDA tasks through learning domain-specific text prompts for source and target domains separately. However, these methods fail to dynamically adjust image features based on the characteristics of their respective domains, thereby limiting their alignment with domain-specific text prompts in CLIP’s joint space, which is a key factor in improving classification performance in the target domain. To bridge this gap, we propose a Unified Text-Image Space Alignment with Cross-Modal Prompting (UTISA) framework for UDA. First, we introduce a Cross-Modal Prompt Learning (CMP) module to generate domain-specific image prompts and layer-specific image prompts for the visual branch to encode domain-specific knowledge globally and locally. Second, under the guidance of image prompts, we introduce a Domain-Aware Multi-Layer Feature Fusion (DMF) module to construct multi-layer domain features for each domain and enhance the image features with these multi-layer domain features, which enables the image features to better reflect the characteristics of their respective domains, thereby promoting their alignment with domain-specific text prompts. Moreover, we introduce a Perturbation-Driven Regularization (PDR) mechanism for the target domain to enhance the robustness and generalization of the model. The experiments demonstrate that UTISA achieves the best performance on three mainstream UDA benchmarks, including 87.9% on Office-Home, 90.9% on VisDA-2017, and 62.4% on DomainNet. Yifan Jiao, Chenglong Cai, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Bool Prompt with Decomposition and Enhancement: Zero-Shot VQA Based on PVLMsabstractZero-Shot Visual Question Answering (ZSVQA) aims to answer questions about images without prior training on explicit image question pairs. Most existing methods usually apply Pre-trained Visual and Language Models (PVLMs) by designing prompts to convert questions into predefined input templates, which (1) ignores the text details and associations of the question when the question is complex, hindering comprehensive understanding, and (2) does not pay attention to local information of image, resulting in overlooking some details of the image that are important to the question when the image content is particularly complex or requires detailed observation. To address these challenges, we propose the Bool Prompt with Decomposition and Enhancement (BPDE) framework for ZSVQA. Specifically, we propose the Bool Sub-Questions Generating module to extract keywords from the original question and generate captions from the image, then use these keywords and captions to guide the transformation of original questions into simpler bool sub-questions, which focus on a specific logical point or piece of information, and guided from the captions can provide the model with local visual information, thereby enhancing the model’s understanding of complex questions and attention to local visual information. Additionally, an Adaptive Sub-Questions Selecting mechanism is designed to ensure non-redundant selection and that the meanings of the sub-questions can cover the original question. Extensive experiments on VQAv2 and AOKVQA demonstrate that the proposed approach performs favorably against the state-of-the-art methods. Liyong Xu, Yifan Jiao, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
Ming Tao 0002, Bing-Kun Bao, Hao Tang 0005, Yaowei Wang 0001, Changsheng Xu |
ECCV (56) | 2 |
| 2024 | Inferring the effectiveness of epidemic prevention measures based on spatial heterogeneity modelingabstractFaced with recent outbreaks of various epidemics, governments worldwide have adopted various countermeasures to slow down the virus spread. Accurately evaluating the effectiveness of these interventions is crucial for their efficiency. However, most existing models assume uniform mixing of infected and susceptible populations throughout the country, which becomes inadequate when strict restrictions significantly reduce interregional population mobility. This bias may lead to unreliable estimates of intervention effects. To this end, we propose a multi-scale regeneration model that considers spatial inhomogeneity, for more accurate evaluation of intervention effects. Specially, by leveraging the idea of renormalization groups to model the spread process between regions at different scales, our model is able to simulate the spatial heterogeneity in epidemic spreading, and achieve a reliable estimation of the interventions’ effect. Experiments on real data show that compared with traditional epidemic models, our model can more accurately describe the epidemic, thereby achieving more reliable evaluation of epidemic intervention effects. Mingyu Wu 0008, Zhiyi Tan 0002, Bing-Kun Bao |
ICME | 3 |
| 2024 | Dynamic Scene Graph Generation with Unified Temporal ModelingabstractDynamic scene graph generation requires understanding the spatial information intra-frame and temporal information between different frames. Existing methods utilize implicit and explicit modeling algorithms to capture temporal information and correlations by designing network architectures or incorporating prior knowledge. However, they exclusively rely on relationship evolution patterns within distinct post-processing modules that are independent of the temporal encoder, leading to solely modifying the results of relationship prediction and difficulty fully harnessing the temporal cues inherent in the video. To address the above challenge, we propose a Unified Temporal Modeling (UTM) that can integrate temporal encoding and temporal correlation modeling. We leverage the relationship evolution patterns to model temporal correlations that can be adapted to capture more relevant temporal cues during temporal encoding. Additionally, our model can be applied to existing image-based scene graph generation methods, extending their capabilities to video tasks. Extensive experiments on the Action Genome dataset demonstrate the robustness of UTM. Sisi You, Bing-Kun Bao |
ICME | 2 |
| 2024 | Label Text-aided Hierarchical Semantics Mining for Panoramic Activity RecognitionabstractPanoramic activity recognition is a comprehensive yet challenging task in crowd scene understanding, which aims to concurrently identify multi-grained human behaviors, including individual actions, social group activities, and global activities. Previous studies tend to capture cross-granularity activity-semantics relations from solely the video input, thus ignoring the intrinsic semantic hierarchy in label-text space. To this end, we propose a label text-aided hierarchical semantics mining (THSM) framework, which explores multi-level cross-modal associations by learning hierarchical semantic alignment between visual content and label texts. Specifically, a hierarchical encoder is first constructed to encode the visual and text inputs into semantics-aligned representations at different granularities. To fully exploit the cross-modal semantic correspondence learned by the encoder, a hierarchical decoder is further developed, which progressively integrates the lower-level representations with the higher-level contextual knowledge for coarse-to-fine action/activity recognition. Extensive experimental results on the public JRDB-PAR benchmark validate the superiority of the proposed THSM framework over state-of-the-art methods. Tianshan Liu, Kin-Man Lam 0001, Bing-Kun Bao |
ACM Multimedia | 3 |
| 2024 | CoIn: A Lightweight and Effective Framework for Story Visualization and ContinuationabstractStory visualization aims to generate realistic and coherent images based on multi-sentence stories. However, current methods face challenges in achieving high-quality image generation while maintaining lightweight models and a fast generation speed. The main issue lies in the two existing frameworks. The independent framework prioritizes speed but sacrifices image quality with the non-collaborative image generation process and basic GAN-based learning. The autoregressive framework modifies the large pretrained text-to-image model in an auto-regressive manner with additional history modules, leading to large model size, resource-intensive requirements, and slow generation speed. To address these issues, we propose a lightweight and effective framework, namely CoIn. Specifically, we introduce a Context-aware Story Generator to predict shared context semantics for each image generator. Additionally, we propose an Intra-Story Interchange module that allows each image generator to exchange visual information with other image generators. Furthermore, we incorporate DINOv2 into the story and image discriminators to assess the story image quality more accurately. Extensive experiments show that our CoIn keeps the model size and generation speed of the independent framework, while achieving promising story image quality. Ming Tao 0002, Bing-Kun Bao, Hao Tang 0005, Yaowei Wang 0001, Changsheng Xu |
ACM Multimedia | 2 |
| 2024 | Two-Stage Reasoning Network with Modality Decomposition for Text VQA
Shengrong Ling, Sisi You, Bing-Kun Bao |
MMM (3) | 3 |
| 2024 | MSGNN: Multi-scale Spatio-temporal Graph Neural Network for epidemic forecasting
Mingjie Qiu, Zhiyi Tan 0002, Bing-Kun Bao |
Data Min. Knowl. Discov. | 3 |
| 2024 | Game and reference: efficient policy making for epidemic prevention and control
Zhiyi Tan 0002, Bing-Kun Bao |
Multim. Syst. | 2 |
| 2024 | Multimodal Imbalance-Aware Gradient Modulation for Weakly-Supervised Audio-Visual Video ParsingabstractWeakly-supervised audio-visual video parsing (WS-AVVP) aims to localize the temporal extents of audio, visual and audio-visual event instances as well as identify the corresponding event categories with only video-level category labels for training. Most previous efforts have been devoted to refining the supervision for each modality or extracting fruitful cross-modality information for more reliable feature learning. None of them have noticed the imbalanced feature learning between different modalities in the task. In this paper, to balance the feature learning processes of different modalities, a dynamic gradient modulation (DGM) mechanism is explored, where a novel and effective metric function is designed to measure the imbalanced feature learning between audio and visual modalities. Furthermore, by going in depth into the principle of traditional WS-AVVP pipelines, two additional challenges are identified: confusing multimodal calculation will hamper the precise measurement of audio-visual imbalanced feature learning, as well as the global supervision provided by video-level labels can not provide explicit guidance for robust semantic feature learning in each action subspace. To cope with the above issues, the modality-separated decision unit (MSDU) and semantic-aware feature extractor (SAFE) are designed for precise measurement of imbalanced feature learning and unambiguous semantic-aware feature extraction separately. Comprehensive experiments are conducted on public benchmarks and the corresponding experimental results demonstrate the effectiveness of our proposed method. Jie Fu 0004, Junyu Gao 0002, Bing-Kun Bao, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Source-Guided Target Feature Reconstruction for Cross-Domain Classification and DetectionabstractExisting cross-domain classification and detection methods usually apply a consistency constraint between the target sample and its self-augmentation for unsupervised learning without considering the essential source knowledge. In this paper, we propose a Source-guided Target Feature Reconstruction (STFR) module for cross-domain visual tasks, which applies source visual words to reconstruct the target features. Since the reconstructed target features contain the source knowledge, they can be treated as a bridge to connect the source and target domains. Therefore, using them for consistency learning can enhance the target representation and reduce the domain bias. Technically, source visual words are selected and updated according to the source feature distribution, and applied to reconstruct the given target feature via a weighted combination strategy. After that, consistency constraints are built between the reconstructed and original target features for domain alignment. Furthermore, STFR is connected with the optimal transportation algorithm theoretically, which explains the rationality of the proposed module. Extensive experiments onnine benchmarksandtwo cross-domain visual tasksprove the effectiveness of the proposed STFR module,e.g., 1)cross-domain image classification: obtaining average accuracy of 91.0%, 73.9%, and 87.4% onOffice-31,Office-Home, andVisDA-2017, respectively; 2)cross-domain object detection: obtaining mAP of 44.50% onCityscapes→Foggy Cityscapes, AP on car of 78.10% onCityscapes→KITTI, MR-2of 8.63%, 12.27%, 22.10%, and 40.58% onCOCOPersons→Caltech,CityPersons→Caltech,COCOPersons→CityPersons, andCaltech→CityPersons, respectively. Yifan Jiao, Hantao Yao, Bing-Kun Bao, Changsheng Xu |
IEEE Trans. Image Process. | 3 |
| 2024 | Injecting Text Clues for Improving Anomalous Event Detection From Weakly Labeled VideosabstractVideo anomaly detection (VAD) aims at localizing the snippets containing anomalous events in long unconstrained videos. The weakly supervised (WS) setting, where solely video-level labels are available during training, has attracted considerable attention, owing to its satisfactory trade-off between the detection performance and annotation cost. However, due to lack of snippet-level dense labels, the existing WS-VAD methods still get easily stuck on the detection errors, caused by false alarms and incomplete localization. To address this dilemma, in this paper, we propose to inject text clues of anomaly-event categories for improving WS-VAD, via a dedicated dual-branch framework. For suppressing the response of confusing normal contexts, we first present a text-guided anomaly discovering (TAG) branch based on a hierarchical matching scheme, which utilizes the label-text queries to search the discriminative anomalous snippets in a global-to-local fashion. To facilitate the completeness of anomaly-instance localization, an anomaly-conditioned text completion (ATC) branch is further designed to perform an auxiliary generative task, which intrinsically forces the model to gather sufficient event semantics from all the relevant anomalous snippets for completely reconstructing the masked description sentence. Furthermore, to encourage the cross-branch knowledge sharing, a mutual learning strategy is introduced by imposing a consistency constraint on the anomaly scores of these two branches. Extensive experimental results on two public benchmarks validate that the proposed method achieves superior performance over the competing methods. Tianshan Liu, Kin-Man Lam 0001, Bing-Kun Bao |
IEEE Trans. Image Process. | 3 |
| 2024 | Improving Graph Collaborative Filtering with Directional Behavior Enhanced Contrastive LearningabstractGraph Collaborative Filtering is a widely adopted approach for recommendation, which captures similar behavior features through Graph Neural Network (GNN). Recently, Contrastive Learning (CL) has been demonstrated as an effective method to enhance the performance of graph collaborative filtering. Typically, CL-based methods first perturb users’ history behavior data (e.g., drop clicked items), then construct a self-discriminating task for behavior representations under different random perturbations. However, for widely existing inactive users, random perturbation makes their sparse behavior information more incomplete, thereby harming the behavior feature extraction. To tackle the above issue, we design a novel directional perturbation-based CL method to improve the graph collaborative filtering performance. The idea is to perturb node representations through directionally enhancing behavior features. To do so, we propose a simple yet effective feedback mechanism, which fuses the representations of nodes based on behavior similarity. Then, to avoid irrelevant behavior preferences introduced by the feedback mechanism, we construct a behavior self-contrast task before and after feedback, to align the node representations between the final output and the first layer of GNN. Different from the widely adopted self-discriminating task, the behavior self-contrast task avoids complex message propagation on different perturbed graphs, which is more efficient than previous methods. Extensive experiments on three public datasets demonstrate that the proposed method has distinct advantages over other CL methods on recommendation accuracy. Penghang Yu, Bing-Kun Bao, Zhiyi Tan 0002, Guanming Lu |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | GPT-Based Knowledge Guiding Network for Commonsense Video CaptioningabstractVideo-based commonsense captioning aims to generate captions for the video content while providing multiple commonsense about the underlying event. Existing methods utilize video features to explore and generate commonsense containing latent semantics. However, this process needs to overcome the complex semantic gap between visible videos and invisible commonsense, which is not supported by the limited knowledge in existing video captioning datasets. To this end, we propose a novel GPT-based Two-stage Knowledge Guiding Network (TKG-Net), which uses GPT to augment datasets knowledge and introduces a cross-attention mechanism to fuse multimodal knowledge. Specifically, to augment knowledge, we set prompts and finetune GPT to imagine and reason based on the video content description at the first stage. At the second stage, to prevent over-reasoning caused by the loss of visual features in GPT, TKG-Net extracts high-level semantic representations of commonsense knowledge and fuses them with video features in a cross-attention mechanism for multimodal semantic interaction. Our experiments on the large-scale Video-to-Commonsense dataset manifest significant improvements over the previous state-of-the-art approach on all metrics. Gengyun Jia, Bing-Kun Bao |
IEEE Trans. Multim. | 3 |
| 2024 | Semantic Distance Adversarial Learning for Text-to-Image SynthesisabstractText-to-Image (T2I) synthesis is a cross-modality task that requires a text description as input to generate a realistic and semantically consistent image. To guarantee semantic consistency, previous studies regenerate text descriptions from synthetic images and align them with the given descriptions. However, the existing redescription modules lack explicit modeling of their training objectives, which is crucial for reliable measurement of semantic distance between redescriptions and given text inputs. Consequently, the aligned text redescriptions suffer from training bias caused by the emergence of adversarial image samples, unseen semantics, and mistaken contents from low-quality synthesized images. To this end, we propose a SEMantic distance Adversarial learning (SEMA) framework for Text-to-Image synthesis which strengthens semantic consistency from two aspects: 1) We introduce adversarial learning between the image generator and the text redescription module to mutually promote or demote the quality of generated image or text instances. This learning model ensures accurate redescription of image contents, thus diminishing the generation of adversarial image samples. 2) We introduce two-fold semantic distance discrimination (SEM distance) to characterize semantic relevance between matching text or image pairs. The unseen semantics and mistaken contents will be penalized with a large SEM distance. The proposed discrimination method also simplifies the model training process with no need to optimize multiple discriminators. Experimental results on CUB Birds 200 and MS-COCO datasets show that the proposed model outperforms the state-of-the-art methods. Yefei Sheng, Bing-Kun Bao, Yi-Ping Phoebe Chen, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2024 | ISF-GAN: Imagine, Select, and Fuse with GPT-Based Text Enrichment for Text-to-Image SynthesisabstractText-to-Image synthesis aims to generate an accurate and semantically consistent image from a given text description. However, it is difficult for existing generative methods to generate semantically complete images from a single piece of text. Some works try to expand the input text to multiple captions via retrieving similar descriptions of the input text from the training set but still fail to fill in missing image semantics. In this article, we propose a GAN-based approach to Imagine, Select, and Fuse for Text-to-image synthesis, named ISF-GAN. The proposed ISF-GAN contains Imagine Stage and Select and Fuse Stage to solve the above problems. First, the Imagine Stage proposes a text completion and enrichment module. This module guides a GPT-based model to enrich the text expression beyond the original dataset. Second, the Select and Fuse Stage selects qualified text descriptions and then introduces a cross-modal attentional mechanism to interact these different sentence embeddings with the image features at different scales. In short, our proposed model enriches the input text information for completing missing semantics and introduces a cross-modal attentional mechanism to maximize the utilization of enriched text information to generate semantically consistent images. Experimental results on CUB, Oxford-102, and CelebA-HQ datasets prove the effectiveness and superiority of the proposed network. Code is available at https://github.com/Feilingg/ISF-GAN Yefei Sheng, Ming Tao 0002, Jie Wang 0061, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Multi-object Tracking with Spatial-Temporal Tracklet AssociationabstractRecently, the tracking-by-detection methods have achieved excellent performance in Multi-Object Tracking (MOT), which focuses on obtaining a robust feature for each object and generating tracklets based on feature similarity. However, they are confronted with two issues: (1) unstable features in short-term occlusion and (2) insufficient matching in long-term occlusion. Specifically, the unstable feature is caused by the appearance variation under occlusion, and the association with the current unstable feature will lead to insufficient matching in long-term occlusion. To address the above issues, we propose a two-stage tracklet-level association method, Spatial-Temporal Tracklet Association (STTA), to effectively combine spatial-temporal context between feature extraction and data association. In the first stage, we propose the Tracklet-guided Spatial-Temporal Attention network (TSTA) to generate robust and stable features. Specifically, TSTA captures spatial-temporal context to obtain the most salient regions between the current and previous clips. In the second stage, we design the Bi-Tracklet Spatial-Temporal association (BTST) module to fully exploit the spatial-temporal context in data association. Specifically, we leverage BTST to merge different tracklets into long-term trajectories by jointly learning visual feature and spatial-temporal context and designing a bidirectional interpolation to recover the missed objects between matched tracklets. Extensive experiments of public and private detections on four benchmarks demonstrate the robustness of STTA. Furthermore, the proposed method is a model-agnostic method, which can be plugged and played with existing methods to boost their performance, e.g., obtain 11.0%, 10.1%, 2.9%, 3.2%, and 7.8% improvement on IDF1 in the MOT16 validation dataset for Tracktor, CenterTrack, Deepsort, JDE, and CTracker, respectively. Sisi You, Hantao Yao, Bing-Kun Bao, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Unbiased Feature Learning with Causal Intervention for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to match individuals across different modalities. Existing methods can learn class-separable features but still struggle with modality gaps within class due to the modality-specific information, which is discriminative in one modality but not present in another (e.g., a black striped shirt). The presence of the interfering information creates a spurious correlation with the class label, which hinders alignment across modalities. To this end, we propose an Unbiased feature learning method based on Causal inTervention for VI-ReID from three aspects. Firstly, through the proposed structural causal graph, we demonstrate that modality-specific information acts as a confounder that restricts the intra-class feature alignment. Secondly, we propose a causal intervention method to remove the confounder using an effective approximation of backdoor adjustment, which involves adjusting the spurious correlation between features and labels. Thirdly, we incorporate the proposed approximation method into the basic VI-ReID model. Specifically, the confounder can be removed by adjusting the extracted features with a set of weighted pre-trained class prototypes from different modalities, where the weight is adapted based on the features. Extensive experiments on the SYSU-MM01 and RegDB datasets demonstrate that our method outperforms state-of-the-art methods. Code is available at https://github.com/NJUPT-MCC/UCT . Sisi You, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | DE-net: Dynamic Text-Guided Image Editing Adversarial NetworksabstractText-guided image editing models have shown remarkable results. However, there remain two problems. First, they employ fixed manipulation modules for various editing requirements (e.g., color changing, texture changing, content adding and removing), which results in over-editing or insufficient editing. Second, they do not clearly distinguish between text-required and text-irrelevant parts, which leads to inaccurate editing. To solve these limitations, we propose: (i) a Dynamic Editing Block (DEBlock) that composes different editing modules dynamically for various editing requirements. (ii) a Composition Predictor (Comp-Pred), which predicts the composition weights for DEBlock according to the inference on target texts and source images. (iii) a Dynamic text-adaptive Convolution Block (DCBlock) that queries source image features to distinguish text-required parts and text-irrelevant parts. Extensive experiments demonstrate that our DE-Net achieves excellent performance and manipulates source images more correctly and accurately. Ming Tao 0002, Bing-Kun Bao, Hao Tang 0005, Fei Wu 0004, Longhui Wei, Qi Tian 0001 |
AAAI | 2 |
| 2023 | GALIP: Generative Adversarial CLIPs for Text-to-Image SynthesisabstractSynthesizing high-fidelity complex images from text is challenging. Based on large pretraining, the autoregressive and diffusion models can synthesize photo-realistic images. Although these large models have shown notable progress, there remain three flaws. 1) These models require tremendous training data and parameters to achieve good performance. 2) The multi-step generation design slows the image synthesis process heavily. 3) The synthesized visual features are challenging to control and require delicately designed prompts. To enable high-quality, efficient, fast, and controllable text-to-image synthesis, we propose Generative Adversarial CLIPs, namely GALIP. GALIP leverages the powerful pretrained CLIP model both in the discriminator and generator. Specifically, we propose a CLIP-based discriminator. The complex scene understanding ability of CLIP enables the discriminator to accurately assess the image quality. Furthermore, we propose a CLIP-empowered generator that induces the visual concepts from CLIP through bridge features and prompts. The CLIP-integrated generator and discriminator boost training efficiency, and as a result, our model only requires about 3% training data and 6% learnable parameters, achieving comparable results to large pretrained autoregressive and diffusion models. Moreover, our model achieves ~120×faster synthesis speed and inherits the smooth latent space from GAN. The extensive experimental results demonstrate the excellent performance of our GALIP. Code is available at https://github.com/tobran/GALIP. Ming Tao 0002, Bing-Kun Bao, Hao Tang 0005, Changsheng Xu |
CVPR | 2 |
| 2023 | UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature EnhancementabstractRecently, Multiple Object Tracking has achieved great success, which consists of object detection, feature embedding, and identity association. Existing methods apply the three-step or two-step paradigm to generate robust trajectories, where identity association is independent of other components. However, the independent identity association results in the identity-aware knowledge contained in the tracklet not be used to boost the detection and embedding modules. To overcome the limitations of existing methods, we introduce a novel Unified Tracking Model (UTM) to bridge those three components for generating a positive feedback loop with mutual benefits. The key insight of UTM is the Identity-Aware Feature Enhancement (IAFE), which is applied to bridge and benefit these three components by utilizing the identity-aware knowledge to boost detection and embedding. Formally, IAFE contains the Identity-Aware Boosting Attention (IABA) and the Identity-Aware Erasing Attention (IAEA), where IABA enhances the consistent regions between the current frame feature and identity-aware knowledge, and IAEA suppresses the distracted regions in the current frame feature. With better detections and embeddings, higher-quality tracklets can also be generated. Extensive experiments of public and private detections on three benchmarks demonstrate the robustness of UTM. Sisi You, Hantao Yao, Bing-Kun Bao, Changsheng Xu |
CVPR | 3 |
| 2023 | TPM: Two-Stage Prediction Mechanism for Universal Adversarial Patch Defense
Huaize Dong, Yifan Jiao, Bing-Kun Bao |
ICIG (5) | 3 |
| 2023 | Multi-level Semantic Extraction Using Graph Pooling Network for Text Representation
Tiankui Fu, Bing-Kun Bao, Xi Shao |
ICIG (4) | 2 |
| 2023 | Multi-modal Context-Aware Network for Scene Graph Generation
Bing-Kun Bao, Zhiyi Tan 0002 |
ICIG (2) | 2 |
| 2023 | Recombination Samples Training for Robust Natural Language Visual ReasoningabstractNatural Language for Visual Reasoning (NLVR) is a challenging binary classification reasoning task, requiring the model to determine whether a textual description is true about a pair of images. However, in this task, if the part of input content (such as only color or quantity) is changed, the recent model is difficult to make the correct judgment. To this end, we propose a novel Recombination Samples Training (RST) scheme for increasing model robustness with stronger reasoning ability. Specifically, the RST synthesizes a lot of unlabeled new samples through the recombination and uses entropy minimization loss to train them. These samples through different combinations of the same content make the model focus on reasoning logic rather than some superficial correlations. Moreover, the RST calculates the consistency loss between the model trained with and without new samples to prevent overfitting. Experiments in the NLVR2 dataset demonstrate the effectiveness of the method. Yuling Jiang, Yingyuan Zhao, Bing-Kun Bao |
ICME | 3 |
| 2023 | Multi-View Graph Convolutional Network for Multimedia RecommendationabstractMultimedia recommendation has received much attention in recent years. It models user preferences based on both behavior information and item multimodal information. Though current GCN-based methods achieve notable success, they suffer from two limitations: (1) Modality noise contamination to the item representations. Existing methods often mix modality features and behavior features in a single view (e.g., user-item view) for propagation, the noise in the modality features may be amplified and coupled with behavior features. In the end, it leads to poor feature discriminability; (2) Incomplete user preference modeling caused by equal treatment of modality features. Users often exhibit distinct modality preferences when purchasing different items. Equally fusing each modality feature ignores the relative importance among different modalities, leading to the suboptimal user preference modeling. Penghang Yu, Zhiyi Tan 0002, Guanming Lu, Bing-Kun Bao |
ACM Multimedia | 4 |
| 2023 | Self-PT: Adaptive Self-Prompt Tuning for Low-Resource Visual Question AnsweringabstractPretraining and finetuning large vision-language models (VLMs) have achieved remarkable success in visual question answering (VQA). However, finetuning VLMs requires heavy computation, expensive storage costs, and is prone to overfitting for VQA in low-resource settings. Existing prompt tuning methods have reduced the number of tunable parameters, but they cannot capture valid context-aware information during prompt encoding, resulting in 1) poor generalization of unseen answers and 2) lower improvements with more parameters. To address these issues, we propose a prompt tuning method for low-resource VQA named Adaptive Self-Prompt Tuning (Self-PT), which utilizes representations of question-image pairs as conditions to obtain context-aware prompts. To enhance the generalization of unseen answers, Self-PT uses dynamic instance-level prompts to avoid overfitting the correlations between static prompts and seen answers observed during training. To reduce parameters, we utilize hyper-networks and low-rank parameter factorization to make Self-PT more flexible and efficient. The hyper-network decouples the number of parameters and prompt length to generate flexible-length prompts by the fixed number of parameters. While the low-rank parameter factorization decomposes and reparameterizes the weights of the prompt encoder into a low-rank subspace for better parameter efficiency. Experiments conducted on VQA v2, GQA, and OK-VQA with different low-resource settings show that our Self-PT outperforms the state-of-the-art parameter-efficient methods, especially in lower-shot settings, e.g., 6% average improvements cross three datasets in 16-shot. Code is available at https://github.com/NJUPT-MCC/Self-PT. Sisi You, Bing-Kun Bao |
ACM Multimedia | 3 |
| 2023 | PRM-KGED: paper recommender model using knowledge graph embedding and deep neural network
Nimbeshaho Thierry, Bing-Kun Bao, Zafar Ali, Zhiyi Tan 0002, Ingabire Batamira Christ Chatelain, Pavlos Kefalas |
Appl. Intell. | 2 |
| 2023 | Centralized sub-critic based hierarchical-structured reinforcement learning for temporal sentence grounding
Yingyuan Zhao, Zhiyi Tan 0002, Bing-Kun Bao, Zhengzheng Tu |
Multim. Syst. | 3 |
| 2023 | Adaptive Text Denoising Network for Image Caption EditingabstractImage caption editing, which aims at editing the inaccurate descriptions of the images, is an interdisciplinary task of computer vision and natural language processing. As the task requires encoding the image and its corresponding inaccurate caption simultaneously and decoding to generate an accurate image caption, the encoder-decoder framework is widely adopted for image caption editing. However, existing methods mostly focus on the decoder, yet ignore a big challenge on the encoder: the semantic inconsistency between image and caption. To this end, we propose a novel A daptive T ext D enoising Net work (ATD-Net) to filter out noises at the word level and improve the model’s robustness at sentence level. Specifically, at the word level, we design a cross-attention mechanism called Textual Attention Mechanism (TAM), to differentiate the misdescriptive words. The TAM is designed to encode the inaccurate caption word by word based on the content of both image and caption. At the sentence level, in order to minimize the influence of misdescriptive words on the semantic of an entire caption, we introduce a Bidirectional Encoder to extract the correct semantic representation from the raw caption. The Bidirectional Encoder is able to model the global semantics of the raw caption, which enhances the robustness of the framework. We extensively evaluate our proposals on the MS-COCO image captioning dataset and prove the effectiveness of our method when compared with the state-of-the-arts. Bing-Kun Bao, Zhiyi Tan 0002, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisabstractSynthesizing high-quality realistic images from text descriptions is a challenging task. Existing text-to-image Generative Adversarial Networks generally employ a stacked architecture as the backbone yet still remain three flaws. First, the stacked architecture introduces the entanglements between generators of different image scales. Second, existing studies prefer to apply and fix extra networks in adversarial learning for text-image semantic consistency, which limits the supervision capability of these networks. Third, the cross-modal attention-based text-image fusion that widely adopted by previous works is limited on several special image scales because of the computational cost. To these ends, we propose a simpler but more effective Deep Fusion Generative Adversarial Networks (DF-GAN). To be specific, we propose: (i) a novel one-stage text-to-image backbone that directly synthesizes high-resolution images without entanglements between different generators, (ii) a novel Target-Aware Discriminator composed of Matching-Aware Gradient Penalty and One-Way Output, which enhances the text-image semantic consistency without introducing extra networks, (iii) a novel deep text-image fusion block, which deepens the fusion process to make a full fusion between text and visual features. Compared with current state-of-the-art methods, our proposed DF-GAN is simpler but more efficient to synthesize realistic and text-matching images and achieves better performance on widely used datasets. Code is available at https://github.com/tobran/DF-GAN. Ming Tao 0002, Hao Tang 0005, Fei Wu 0004, Xiaoyuan Jing, Bing-Kun Bao, Changsheng Xu |
CVPR | 5 |
| 2022 | Unbiased feature enhancement framework for cross-modality person re-identification
Bairu Chen, Zhiyi Tan 0002, Xi Shao, Bing-Kun Bao |
Multim. Syst. | 5 |
| 2022 | Truncated γ norm-based low-rank and sparse decomposition
Zhenzhen Yang, Bing-Kun Bao |
Multim. Tools Appl. | 4 |
| 2022 | River Channel Extraction in SAR Images Using Level Sets Driven by Symmetric Kullback-Leibler DistanceabstractThis article proposes a novel level set method (LSM) that improves river channel extraction accuracy in synthetic aperture radar (SAR) images by developing global median image fitting energies. First, we define a new global median fitting image (GMFI) to approximate the input image and use this GMFI to construct the fitting energy based on the symmetric Kullback–Leibler distance (SKL). Second, to exploit more image grayscale features, a squared global median fitting image (SGMFI) is derived and another fitting energy is similarly constructed using this SGMFI based on SKL. Third, we integrate the above two fitting energies and introduce additional regularized energies. The proposed LSM is verified and compared with several state-of-the-art methods on real SAR images. The river channel extraction results indicate that our proposed LSM has a clear advantage in accuracy and is robust to level set initialization. Bing-Kun Bao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | DualVGR: A Dual-Visual Graph Reasoning Unit for Video Question AnsweringabstractVideo question answering is a challenging task, which requires agents to be able to understand rich video contents and perform spatial-temporal reasoning. However, existing graph-based methods fail to perform multi-step reasoning well, neglecting two properties of VideoQA: (1) Even for the same video, different questions may require different amount of video clips or objects to infer the answer with relational reasoning; (2) During reasoning, appearance and motion features have complicated interdependence which are correlated and complementary to each other. Based on these observations, we propose a Dual-Visual Graph Reasoning Unit (DualVGR) which reasons over videos in an end-to-end fashion. The first contribution of our DualVGR is the design of an explainable Query Punishment Module, which can filter out irrelevant visual features through multiple cycles of reasoning. The second contribution is the proposed Video-based Multi-view Graph Attention Network, which captures the relations between appearance and motion features. Our DualVGR network achieves state-of-the-art performance on the benchmark MSVD-QA and SVQA datasets, and demonstrates competitive results on benchmark MSRVTT-QA datasets. Our code is available athttps://github.com/MM-IR/DualVGR-VideoQA. Jianyu Wang 0010, Bing-Kun Bao, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2021 | Multi-pose Facial Expression Recognition Based on Unpaired Images
Bairu Chen, Yibo Gan, Bing-Kun Bao |
ICIG (2) | 3 |
| 2021 | Meta-Learning Causal Feature Selection for Stable PredictionabstractConventional predictive models in machine learning are based on I.I.D. hypothesis between training and testing data. However, such a hypothesis is fragile in the real world, and the model minimizing empirical errors on training data does not perform well on testing data, which makes the prediction unstable. This instability can be found widely in domain generalization, active learning, and transfer learning, etc. In this paper, we propose a novel Meta-learning Causal Feature Selection (MCFS) model for general Non-I.I.D. image classification. In MCFS, we jointly optimize a convolutional network and a causal parameter for identifying causal variables on meta-training and meta-testing data which simulate the distribution shifts in Non-I.I.D. problems. Extensive experiments conducted on public VLCS and NICO datasets demonstrate the effectiveness of the proposed MCFS, which outperforms the state-of-the-art methods. Zhaoquan Yuan, Xiao Wu 0001, Bing-Kun Bao, Changsheng Xu |
ICME | 4 |
| 2021 | A parallel multi-block alternating direction method of multipliers for tensor completionabstractAbstract This paper proposes an algorithm for the tensor completion problem of estimating multi‐linear data under the limitation of observation rate. Many tensor completion methods are based on nuclear norm minimization, they may fail to achieve the global solution for solving nuclear norm minimization in tensor completion problem with high missing ratio. To tackle this issue, an adaptive tensor completion method based on parallel multi‐block alternating direction method of multipliers (ADMM) algorithm is proposed, it can derive the model from the initial estimate and compute the next estimate from the current solution. The parallel multi‐block ADMM with global convergence is adopted to solve the dual problem, which greatly improves the processing power and reliability of the algorithm. Hu Zhu, Taiyu Yan, Yu-Feng Yu 0001, Lizhen Deng, Bing-Kun Bao |
IET Image Process. | 6 |
| 2021 | RoDeRain: Rotational Video Derain via Nonconvex and Nonsmooth Optimization
Lizhen Deng, Guoxia Xu, Hu Zhu, Bing-Kun Bao |
Mob. Networks Appl. | 4 |
| 2021 | Cross-Domain Object Representation via Robust Low-Rank Correlation AnalysisabstractCross-domain data has become very popular recently since various viewpoints and different sensors tend to facilitate better data representation. In this article, we propose a novel cross-domain object representation algorithm (RLRCA) which not only explores the complexity of multiple relationships of variables by canonical correlation analysis (CCA) but also uses a low rank model to decrease the effect of noisy data. To the best of our knowledge, this is the first try to smoothly integrate CCA and a low-rank model to uncover correlated components across different domains and to suppress the effect of noisy or corrupted data. In order to improve the flexibility of the algorithm to address various cross-domain object representation problems, two instantiation methods of RLRCA are proposed from feature and sample space, respectively. In this way, a better cross-domain object representation can be achieved through effectively learning the intrinsic CCA features and taking full advantage of cross-domain object alignment information while pursuing low rank representations. Extensive experimental results on CMU PIE, Office-Caltech, Pascal VOC 2007, and NUS-WIDE-Object datasets, demonstrate that our designed models have superior performance over several state-of-the-art cross-domain low rank methods in image clustering and classification tasks with various corruption levels. Xiangjun Shen, Jinghui Zhou, Zhongchen Ma, Bing-Kun Bao, Zhengjun Zha |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | A Densely Connected Network Based on U-Net for Medical Image SegmentationabstractThe U-Net has become the most popular structure in medical image segmentation in recent years. Although its performance for medical image segmentation is outstanding, a large number of experiments demonstrate that the classical U-Net network architecture seems to be insufficient when the size of segmentation targets changes and the imbalance happens between target and background in different forms of segmentation. To improve the U-Net network architecture, we develop a new architecture named densely connected U-Net (DenseUNet) network in this article. The proposed DenseUNet network adopts a dense block to improve the feature extraction capability and employs a multi-feature fuse block fusing feature maps of different levels to increase the accuracy of feature extraction. In addition, in view of the advantages of the cross entropy and the dice loss functions, a new loss function for the DenseUNet network is proposed to deal with the imbalance between target and background. Finally, we test the proposed DenseUNet network and compared it with the multi-resolutional U-Net (MultiResUNet) and the classic U-Net networks on three different datasets. The experimental results show that the DenseUNet network has significantly performances compared with the MultiResUNet and the classic U-Net networks. Zhenzhen Yang, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Relative coordinates constraint for face alignment
Fudong Nian, Teng Li 0001, Bing-Kun Bao, Changsheng Xu |
Neurocomputing | 3 |
| 2020 | Gradient-based discriminative modeling for blind image deblurring
Wenze Shao, Yunzhi Lin, Li-Qian Wang, Qi Ge, Bing-Kun Bao, Haibo Li 0001 |
Neurocomputing | 6 |
| 2020 | Robust low rank representation via feature and sample scaling
Xiangjun Shen, Liangjun Wang, Sumet Mehta, Bing-Kun Bao, Jianping Fan 0001 |
Neurocomputing | 5 |
| 2020 | A generalized least-squares approach regularized with graph embedding for dimensionality reduction
Xiangjun Shen, Si-Xing Liu, Bing-Kun Bao, Chunhong Pan, Zhengjun Zha, Jianping Fan 0001 |
Pattern Recognit. | 3 |
| 2020 | DeblurGAN+: Revisiting blind motion deblurring using conditional adversarial networks
Wenze Shao, Lu-Yue Ye, Li-Qian Wang, Qi Ge, Bing-Kun Bao, Haibo Li 0001 |
Signal Process. | 6 |
| 2019 | Adversarial Representation Learning for Dynamic Scene Deblurring: A Simple, Fast and Robust ApproachabstractIn this paper, we investigate a novel learning-based method for dynamic scene deblurring. Since the inference model is formulated as an encoder-decoder, the core task has turned to learning blur-invariant hidden features from the blurred images to a great degree. To achieve state-of-the-art results in terms of both deblurring accuracy and efficiency, a simple, robust and computationally efficient deep auto-encoder is developed tailored specifically for blind deblurring, which is learned in an adversarial fashion based on use of the recent Wasserstein generative adversarial networks. Thanks to the designed framework, the new model has shown comparable or superior performance both qualitatively and quantitatively to existing state-of-the-art methods. What is more important, instead of exploiting the multi-scale strategy as previous methods, our model is just single-scale capable of achieving 6 times more efficiency than the closest competitor by Nah et al. [1], which is a more complex multi-scale deep method. Lu-Yue Ye, Wenze Shao, Qi Ge, Li-Qian Wang, Bing-Kun Bao, Haibo Li 0001 |
ICIP | 6 |
| 2019 | Multimodal Latent Factor Model with Language Constraint for Predicate DetectionabstractNowadays, visual relationship detection has shown an important utility in scene understanding. Predicate detection, which aims to detect the predicate between entities in an image, is an important part of visual relationship detection. In this paper, we propose Multimodal Latent Factor Model with Language Constraint (MMLFM-LC) for predicate detection with the novelty of integrating knowledge learned from multiple modalities, valid relationships and semantical similarities. Representations of visual and textual modalities are firstly input into the constructed model. Secondly, a bilinear structure is introduced to model the relationships using valid relationships, while a language constraint is also built utilizing semantical similarities. Lastly, visual and textual representations are fused in an embedded subspace for predicate detection. Experiments on both Visual Relationship and Visual Genome datasets show that our method outperforms other methods on predicate detection. Bing-Kun Bao, Lingling Yao, Changsheng Xu |
ICIP | 2 |
| 2019 | Training Efficient Saliency Prediction Models with Knowledge DistillationabstractRecently, deep learning-based saliency prediction methods have achieved significant accuracy improvements. However, they are hard to embed in practical multimedia applications due to large memory consumption and running time caused by complicated architectures. In addition, most methods are fine-tuned from pre-trained models for classification tasks, and networks cannot flexibly be transferred for a new task. In this paper, a condensed and randomly initialized student network is employed to achieve higher efficiency by transferring knowledge from complicated and well-trained teacher networks. This is the first use of knowledge distillation for efficient pixel-wise saliency prediction. Instead of directly minimizing Euclidean distance between feature maps, we propose two statistical representations of feature maps (i.e., first-order and second-order statistics) as knowledge. We conduct experiments on three kinds of teacher networks and four benchmark datasets to verify the effectiveness of the proposed method. Compared with the teacher networks, the student networks achieve an acceleration ratio of 4.56-4.73. Compared with state-of-the-art approaches, the proposed model achieves competitive accuracy with faster running speed (up to 4.38 times) and smaller model size (up to 93.27% reduction). We further embedded the proposed saliency prediction model into a video captioning application. The saliency-embedded approaches improve video captioning on all test metrics with a small complexity cost. The student-model embedded approach achieves 25% time saving with similar performance to the teacher embedded one. Peng Zhang 0024, Li Su 0003, Liang Li 0003, Bing-Kun Bao, Pamela C. Cosman, Guorong Li, Qingming Huang |
ACM Multimedia | 4 |
| 2019 | Session details: Human Analysis in MultimediaabstractNo abstract available. Bing-Kun Bao |
MMAsia | 1 |
| 2019 | Sentiment-Aware Multi-modal Recommendation on Tourist Attractions
Bing-Kun Bao, Changsheng Xu |
MMM (1) | 2 |
| 2019 | Dictionary-induced least squares framework for multi-view dimensionality reduction with multi-manifold embeddingsabstractThis study proposes a novel dimensionality reduction (DR) method for multi‐view datasets. The principal component analysis (PCA) idea of minimising least squares reconstruction errors is extended to consider both data distribution and penalty weights called dictionary to recover outliers free global structures from missing and noisy data points. In this way, PCA is viewed as a special instance of the authors’ proposed dictionary induced least squares framework (DLS). Furthermore, to appropriately handle multi‐view DR, we combine the DLS with multiple manifold embeddings (DLSME). Therefore it can obtain lower projections while maintaining a balance between preserving global structures with DLS and local structures with multi‐manifold embeddings. Extensive experiments on object and face recognition datasets verify that the DLS achieves better classification results with lower dimensional projections than PCA. Also, on many multi‐view datasets of visual recognition and web image annotation, the DLSME method demonstrates more effectiveness than Graph‐Laplacian PCA (gLPCA), robust PCA‐optimal mean, canonical correlation analysis (CCA), bilinear models (BLM), neighbourhood preserving embedding, locality preserving projections, and locality sensitive discriminant analysis. Timothy Apasiba Abeo, Xiangjun Shen, Jianping Gou, Qirong Mao, Bing-Kun Bao |
IET Comput. Vis. | 5 |
| 2019 | On potentials of regularized Wasserstein generative adversarial networks for realistic hallucination of tiny faces
Wenze Shao, Qi Ge, Li-Qian Wang, Bing-Kun Bao, Haibo Li 0001 |
Neurocomputing | 6 |
| 2019 | Enhancing Blurred Low-Resolution Images via Exploring the Potentials of Learning-Based Super-ResolutionabstractThis paper aims to propose a candidate solution to the challenging task of single-image blind super-resolution (SR), via extensively exploring the potentials of learning-based SR schemes in the literature. The task is formulated into an energy functional to be minimized with respect to both an intermediate super-resolved image and a nonparametric blur-kernel. The functional includes a so-called convolutional consistency term which incorporates a nonblind learning-based SR result to better guide the kernel estimation process, and a bi-[Formula: see text]-[Formula: see text]-norm regularization imposed on both the super-resolved sharp image and the nonparametric blur-kernel. A numerical algorithm is deduced via coupling the splitting augmented Lagrangian (SAL) and the conjugate gradient (CG) method. With the estimated blur-kernel, the final SR image is reconstructed using a simple TV-based nonblind SR method. The proposed blind SR approach is demonstrated to achieve better performance than [T. Michaeli and M. Irani, Nonparametric Blind Super-resolution, in Proc. IEEE Conf. Comput. Vision (IEEE Press, Washington, 2013), pp. 945–952.] in terms of both blur-kernel estimation accuracy and image ehancement quality. In the meanwhile, the experimental results demonstrate surprisingly that the local linear regression-based SR method, anchored neighbor regression (ANR) serves the proposed functional more appropriately than those harnessing the deep convolutional neural networks. Wenze Shao, Bing-Kun Bao, Haibo Li 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2019 | A generalized multi-dictionary least squares framework regularized with multi-graph embeddings
Timothy Apasiba Abeo, Xiangjun Shen, Bing-Kun Bao, Zhengjun Zha, Jianping Fan 0001 |
Pattern Recognit. | 3 |
| 2018 | Blind Deblurring Using Discriminative Image Smoothing
Wenze Shao, Yunzhi Lin, Bing-Kun Bao, Liqian Wang, Qi Ge, Haibo Li 0001 |
PRCV (1) | 3 |
| 2018 | You Are What You Eat: Exploring Rich Recipe Information for Cross-Region Food AnalysisabstractCuisine is a style of cooking and usually associated with a specific geographic region. Recipes from different cuisines shared on the web are an indicator of culinary cultures in different countries. Therefore, analysis of these recipes can lead to deep understanding of food from the cultural perspective. In this paper, we perform the first cross-region recipe analysis by jointly using the recipe ingredients, food images, and attributes such as the cuisine and course (e.g., main dish and dessert). For that solution, we propose a culinary culture analysis framework to discover the topics of ingredient bases and visualize them to enable various applications. We first propose a probabilistic topic model to discover cuisine-course specific topics. The manifold ranking method is then utilized to incorporate deep visual features to retrieve food images for topic visualization. At last, we applied the topic modeling and visualization method for three applications: 1) multimodal cuisine summarization with both recipe ingredients and images, 2) cuisine-course pattern analysis including topic-specific cuisine distribution and cuisine-specific course distribution of topics, and 3) cuisine recommendation for both cuisine-oriented and ingredient-oriented queries. Through these three applications, we can analyze the culinary cultures at both macro and micro levels. We conduct the experiment on a recipe database Yummly-66K with 66,615 recipes from 10 cuisines in Yummly. Qualitative and quantitative evaluation results have validated the effectiveness of topic modeling and visualization, and demonstrated the advantage of the framework in utilizing rich recipe information to analyze and interpret the culinary cultures from different regions. Weiqing Min, Bing-Kun Bao, Shuhuan Mei, Yong Rui, Shuqiang Jiang |
IEEE Trans. Multim. | 2 |
| 2017 | Multi-Modal Knowledge Representation Learning via Webly-Supervised Relationships MiningabstractKnowledge representation learning (KRL) encodes enormous structured information with entities and relations into a continuous low-dimensional semantic space. Most conventional methods solely focus on learning knowledge representation from single modality, yet neglect the complementary information from others. The more and more rich available multi-modal data on Internet also drive us to explore a novel approach for KRL in multi-modal way, and overcome the limitations of previous single-modal based methods. This paper proposes a novel multi-modal knowledge representation learning (MM-KRL) framework which attempts to handle knowledge from both textual and visual modal web data. It consists of two stages, i.e., webly-supervised multi-modal relationship mining, and bi-enhanced cross-modal knowledge representation learning. Compared with existing knowledge representation methods, our framework has several advantages: (1) It can effectively mine multi-modal knowledge with structured textual and visual relationships from web automatically. (2) It is able to learn a common knowledge space which is independent to both task and modality by the proposed Bi-enhanced Cross-modal Deep Neural Network (BC-DNN). (3) It has the ability to represent unseen multi-modal relationships by transferring the learned knowledge with isolated seen entities and relations into unseen relationships. We build a large-scale multi-modal relationship dataset (MMR-D) and the experimental results show that our framework achieves excellent performance in zero-shot multi-modal retrieval and visual relationship recognition. Fudong Nian, Bing-Kun Bao, Teng Li 0001, Changsheng Xu |
ACM Multimedia | 2 |
| 2017 | A discriminative graph inferring framework towards weakly supervised image parsing
Bing-Kun Bao, Changsheng Xu |
Multim. Syst. | 2 |
| 2016 | Depth-aware layered edge for object proposalabstractObject proposal, typically served as preprocessing of various multimedia applications, aims to detect the bounding boxes of possible objects in an image. In this paper, we propose a novel object proposal method for RGB-D images based on layered edges, which can effectively eliminate the influence of the mixture of edges from objects and background and improve the accuracy of proposals. Firstly, we detect the sparse edges and correct depth on super-pixel representation. Then, we use depth-adaptive sliding windows in sampling of depth distribution and measure the objectness of each candidate box in multiple depth layers. Finally, the candidate boxes are ranked according to the integrated scores of all the depth layers, and the final proposals are generated. The experimental results show that the proposed method can outperform the state-of-the-art methods on the largest RGB-D image dataset for object proposal. Tongwei Ren, Bing-Kun Bao, Jia Bei |
ICME | 3 |
| 2016 | An incremental probabilistic model for temporal theme analysis of landmarks
Weiqing Min, Bing-Kun Bao, Changsheng Xu |
Multim. Syst. | 2 |
| 2016 | Recent advances in social multimedia big data mining and applications
Yue Gao 0002, Bing-Kun Bao, Cees Snoek, Qionghai Dai |
Multim. Syst. | 3 |
| 2016 | Guest Editorial: Learning Multimedia for Real World Applications
Bing-Kun Bao, Congyan Lang, Tao Mei 0001, Alberto Del Bimbo |
Multim. Tools Appl. | 1 |
| 2015 | Joint Local and Global Consistency on Interdocument and Interword Relationships for Co-ClusteringabstractCo-clustering has recently received a lot of attention due to its effectiveness in simultaneously partitioning words and documents by exploiting the relationships between them. However, most of the existing co-clustering methods neglect or only partially reveal the interword and interdocument relationships. To fully utilize those relationships, the local and global consistencies on both word and document spaces need to be considered, respectively. Local consistency indicates that the label of a word/document can be predicted from its neighbors, while global consistency enforces a smoothness constraint on words/documents labels over the whole data manifold. In this paper, we propose a novel co-clustering method, called co-clustering via local and global consistency, to not only make use of the relationship between word and document, but also jointly explore the local and global consistency on both word and document spaces, respectively. The proposed method has the following characteristics: 1) the word-document relationships is modeled by following information-theoretic co-clustering (ITCC); 2) the local consistency on both interword and interdocument relationships is revealed by a local predictor; and 3) the global consistency on both interword and interdocument relationships is explored by a global smoothness regularization. All the fitting errors from these three-folds are finally integrated together to formulate an objective function, which is iteratively optimized by a convergence provable updating procedure. The extensive experiments on two benchmark document datasets validate the effectiveness of the proposed co-clustering method. Bing-Kun Bao, Weiqing Min, Teng Li 0001, Changsheng Xu |
IEEE Trans. Cybern. | 1 |
| 2015 | Cross-Platform Multi-Modal Topic Modeling for Personalized Inter-Platform RecommendationabstractIn this paper, we investigate a novel cross- platform multimedia problem: given two platforms, Flickr and Foursquare, we conduct the recommendation between these two platforms, namely the photo recommendation from Flickr to Foursquare users and the venue recommendation from Foursquare to Flickr users. Such inter-platform recommendations enable users from one single platform to enjoy different recommendation services effectively . To solve the problem, we propose a cross- platform multi-modal topic model ( CM3TM), which is capable of: 1) differentiating between two kinds of topics, i.e., platform- specific topics only relevant to a certain platform and shared topics characterizing the knowledge shared by different platforms and 2) aligning multiple modalities from different platforms. Specifically, CM3TM can not only split the topic space into the shared topic space and platform-specific topic space and learn them simultaneously, but also enable the alignment among different modalities through the learned topic space. Given the location information, we applied the proposed CM3TM into two inter-platform recommendation applications: 1) personalized venue recommendation from Foursquare to Flickr users and 2) personalized image recommendation from Flickr to Foursquare users. We have conducted experiments on the collected large-scale real-world dataset from Flickr and Foursquare. Qualitative and quantitative evaluation results validate the effectiveness of our method and demonstrate the advantage of connecting different platforms with different modalities for the inter-platform recommendation. Weiqing Min, Bing-Kun Bao, Changsheng Xu, M. Shamim Hossain |
IEEE Trans. Multim. | 2 |
| 2015 | Knowing Verb From Object: Retagging With Transfer Learning on Verb-Object Concept ImagesabstractImage retagging is significant and essential for tag-based applications, such as search and browsing. However, most existing image retagging approaches are typically based on enriching-and-removing and/or reranking strategies, which lead to two drawbacks: 1) since the object and/or human appeared in the images are tagged as individuals, the meanings represented by the mutual context of object and human are ignored and not tagged, and 2) some images which are visually dissimilar but semantically similar could be filtered incorrectly, as they are conflict with the content consistency rule. These two defects are distinct especially when images with human-object interactions are retagged. To tackle these defects, in this paper we propose a Bayesian approach to jointly consider the human and object in an image and retag it properly. In our approach, human and objects in images are detected and their interrelationships are taken into account. Tags which represent the mutual context of human and objects are then mapped to those interrelationships by a probabilistic graphical model. For a new image which lacks the tag representing the interaction between human and object, our model can correctly retag it for the interaction. In this paper, those images involving human-object interactions are called verb-object concept images, and experiments on a 60-class dataset demonstrate the capacity of our Bayesian retagging approach of verb-object concept images (BRVOI). Bing-Kun Bao, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2015 | Cross-Platform Emerging Topic Detection and Elaboration from Multimedia StreamsabstractWith the explosive growth of online media platforms in recent years, it becomes more and more attractive to provide users a solution of emerging topic detection and elaboration. And this posts a real challenge to both industrial and academic researchers because of the overwhelming information available in multiple modalities and with large outlier noises. This article provides a method on emerging topic detection and elaboration using multimedia streams cross different online platforms. Specifically, Twitter, New York Times and Flickr are selected for the work to represent the microblog, news portal and imaging sharing platforms. The emerging keywords of Twitter are firstly extracted using aging theory. Then, to overcome the nature of short length message in microblog, Robust Cross-Platform Multimedia Co-Clustering (RCPMM-CC) is proposed to detect emerging topics with three novelties: 1) The data from different media platforms are in multimodalities; 2) The coclustering is processed based on a pairwise correlated structure, in which the involved three media platforms are pairwise dependent; 3) The noninformative samples are automatically pruned away at the same time of coclustering. In the last step of cross-platform elaboration, we enrich each emerging topic with the samples from New York Times and Flickr by computing the implicit links between social topics and samples from selected news and Flickr image clusters, which are obtained by RCPMM-CC. Qualitative and quantitative evaluation results demonstrate the effectiveness of our method. Bing-Kun Bao, Changsheng Xu, Weiqing Min, M. Shamim Hossain |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2014 | Scene and viewpoint based visual summarization for landmarksabstractVisual summarization of landmarks is an important task for applications, such as landmark organization, search and browsing. In this work, we make the first attempt towards landmark summarization by simultaneously considering both the scenes (e.g., sunny view and night view) and viewpoints (e.g., front-side and close-distant viewpoint). In the proposed framework of landmark summarization, we first group images into different clusters by viewpoints, then the distinctive scenes for each viewpoint cluster are discovered by the proposed scene-viewpoint based theme modeling. Compared with the existing topic models, our model is capable of mining scene-viewpoint themes directly from all viewpoint clusters and meanwhile differentiating among these themes by viewpoints. The landmark summary is generated by the discovered scene-viewpoint themes, where each theme is represented by the selected images with one certain scene and viewpoint. The experimental results validate the proposed method and demonstrate its advantage in improving user experience. Weiqing Min, Bing-Kun Bao, Changsheng Xu |
ICIP | 2 |
| 2014 | Robust object removal with an exemplar-based image inpainting approach
Jing Wang 0093, Ke Lu 0002, Daru Pan, Bing-Kun Bao |
Neurocomputing | 5 |
| 2014 | Single-image motion deblurring using an adaptive image prior
Ke Lu 0002, Bing-Kun Bao, Lulu Zhang 0006, Jinbao Wang 0001 |
Inf. Sci. | 3 |
| 2014 | Inductive hierarchical nonnegative graph embedding for "verb-object" image classification
Bing-Kun Bao, Changsheng Xu |
Mach. Vis. Appl. | 2 |
| 2014 | Mobile Landmark Search with 3D ModelsabstractLandmark search is crucial to improve the quality of travel experience. Smart phones make it possible to search landmarks anytime and anywhere. Most of the existing work computes image features on smart phones locally after taking a landmark image. Compared with sending original image to the remote server, sending computed features saves network bandwidth and consequently makes sending process fast. However, this scheme would be restricted by the limitations of phone battery power and computational ability. In this paper, we propose to send compressed (low resolution) images to remote server instead of computing image features locally for landmark recognition and search. To this end, a robust 3D model based method is proposed to recognize query images with corresponding landmarks. Using the proposed method, images with low resolution can be recognized accurately, even though images only contain a small part of the landmark or are taken under various conditions of lighting, zoom, occlusions and different viewpoints. In order to provide an attractive landmark search result, a 3D texture model is generated to respond to a landmark query. The proposed search approach, which opens up a new direction, starts from a 2D compressed image query input and ends with a 3D model search result. Weiqing Min, Changsheng Xu, Min Xu 0001, Xian Xiao, Bing-Kun Bao |
IEEE Trans. Multim. | 5 |
| 2013 | Latent support vector machine for sign language recognition with KinectabstractIn this paper, we propose a novel algorithm to model and recognize sign language with Kinect sensor. We assume that in a sign language video, some frames are expected to be both discriminative and representative. Under this assumption, each frame in training videos is assigned a binary latent variable indicating its discriminative capability. A Latent Support Vector Machine model is then developed to classify the signs, as well as localize the discriminative and representative frames in videos. In addition, we utilize the depth map together with color image captured by Kinect sensor to obtain more effective and accurate feature to enhance the recognition accuracy. To evaluate our approach, we collected an American Sign Language (ASL) dataset which included approximately 2000 phrases, while each phrase was captured by Kinect sensor and hence included color, depth and skeleton information. Experiments on our dataset demonstrate the effectiveness of the proposed method for sign language recognition. Tianzhu Zhang 0001, Bing-Kun Bao, Changsheng Xu |
ICIP | 3 |
| 2013 | Social event detection with robust high-order co-clusteringabstractThis paper is devoted to detecting social, real-world events from the sharing images/videos on social media sites like Flickr and YouTube. The fast growing contents make the social media sites become gold mines for social event detection, but we still need to overcome the challenge of processing the associated heterogeneous metadata, such as time-stamp, location, visual content and textual content. Different from the traditional early or late fusion with different types of metadata, we represent them into a star-structured $K$-partite graph, that is, social media itself is regarded as the central vertices set and different types of metadata are treated as the auxiliary vertices sets which are pairwise independent with each other but correlated with the central one. Based on this graph, Social Event Detection with Robust High-Order Co-Clustering (SED-RHOCC) algorithm is proposed and it includes two steps: 1) coarse event detection, 2) clusters and samples refinement. In the first step, by revealing the inter-relationship on the constructed star-structured $K$-partite graph and the intra-relationship within some metadata sets such as time-stamp, we co-cluster social media and the associated metadata separately and iteratively to avoid information loss in early/late fusion. After that, a post process is utilized to refine the clusters and social media samples in the second step. MediaEval Social Event Detection Dataset [1] and its subset are selected to demonstrate the effectiveness of our proposed approach in handling the datasets with and without non-event samples. Bing-Kun Bao, Weiqing Min, Ke Lu 0002, Changsheng Xu |
ICMR | 1 |
| 2013 | Landmark History Visualization
Weiqing Min, Bing-Kun Bao, Changsheng Xu |
MMM (2) | 2 |
| 2013 | Verb-Object Concepts Image Classification via Hierarchical Nonnegative Graph Embedding
Bing-Kun Bao, Changsheng Xu |
MMM (1) | 2 |
| 2013 | Discriminative Exemplar Coding for Sign Language Recognition With KinectabstractSign language recognition is a growing research area in the field of computer vision. A challenge within it is to model various signs, varying with time resolution, visual manual appearance, and so on. In this paper, we propose a discriminative exemplar coding (DEC) approach, as well as utilizing Kinect sensor, to model various signs. The proposed DEC method can be summarized as three steps. First, a quantity of class-specific candidate exemplars are learned from sign language videos in each sign category by considering their discrimination. Then, every video of all signs is described as a set of similarities between frames within it and the candidate exemplars. Instead of simply using a heuristic distance measure, the similarities are decided by a set of exemplar-based classifiers through the multiple instance learning, in which a positive (or negative) video is treated as a positive (or negative) bag and those frames similar to the given exemplar in Euclidean space as instances. Finally, we formulate the selection of the most discriminative exemplars into a framework and simultaneously produce a sign video classifier to recognize sign. To evaluate our method, we collect an American sign language dataset, which includes approximately 2000 phrases, while each phrase is captured by Kinect sensor with color, depth, and skeleton information. Experimental results on our dataset demonstrate the feasibility and effectiveness of the proposed approach for sign language recognition. Tianzhu Zhang 0001, Bing-Kun Bao, Changsheng Xu, Tao Mei 0001 |
IEEE Trans. Cybern. | 3 |
| 2013 | General Subspace Learning With Corrupted Training Data Via Graph EmbeddingabstractWe address the following subspace learning problem: supposing we are given a set of labeled, corrupted training data points, how to learn the underlying subspace, which contains three components: an intrinsic subspace that captures certain desired properties of a data set, a penalty subspace that fits the undesired properties of the data, and an error container that models the gross corruptions possibly existing in the data. Given a set of data points, these three components can be learned by solving a nuclear norm regularized optimization problem, which is convex and can be efficiently solved in polynomial time. Using the method as a tool, we propose a new discriminant analysis (i.e., supervised subspace learning) algorithm called Corruptions Tolerant Discriminant Analysis (CTDA), in which the intrinsic subspace is used to capture the features with high within-class similarity, the penalty subspace takes the role of modeling the undesired features with high between-class similarity, and the error container takes charge of fitting the possible corruptions in the data. We show that CTDA can well handle the gross corruptions possibly existing in the training data, whereas previous linear discriminant analysis algorithms arguably fail in such a setting. Extensive experiments conducted on two benchmark human face data sets and one object recognition data set show that CTDA outperforms the related algorithms. Bing-Kun Bao, Guangcan Liu, Richang Hong, Shuicheng Yan, Changsheng Xu |
IEEE Trans. Image Process. | 1 |
| 2013 | Robust Image Analysis With Sparse Representation on Quantized Visual FeaturesabstractRecent techniques based on sparse representation (SR) have demonstrated promising performance in high-level visual recognition, exemplified by the highly accurate face recognition under occlusion and other sparse corruptions. Most research in this area has focused on classification algorithms using raw image pixels, and very few have been proposed to utilize the quantized visual features, such as the popular bag-of-words feature abstraction. In such cases, besides the inherent quantization errors, ambiguity associated with visual word assignment and misdetection of feature points, due to factors such as visual occlusions and noises, constitutes the major cause of dense corruptions of the quantized representation. The dense corruptions can jeopardize the decision process by distorting the patterns of the sparse reconstruction coefficients. In this paper, we aim to eliminate the corruptions and achieve robust image analysis with SR. Toward this goal, we introduce two transfer processes (ambiguity transfer and mis-detection transfer) to account for the two major sources of corruption as discussed. By reasonably assuming the rarity of the two kinds of distortion processes, we augment the original SR-based reconstruction objective with l(0) norm regularization on the transfer terms to encourage sparsity and, hence, discourage dense distortion/transfer. Computationally, we relax the nonconvex l(0) norm optimization into a convex l(1) norm optimization problem, and employ the accelerated proximal gradient method to optimize the convergence provable updating procedure. Extensive experiments on four benchmark datasets, Caltech-101, Caltech-256, Corel-5k, and CMU pose, illumination, and expression, manifest the necessity of removing the quantization corruptions and the various advantages of the proposed framework. Bing-Kun Bao, Guangyu Zhu 0002, Jialie Shen 0001, Shuicheng Yan |
IEEE Trans. Image Process. | 1 |
| 2012 | Multimedia news digger on emerging topics from social streamsabstractWith the overwhelming information from social media networks and news portals, it is crucial to provide users a complete package of visual and textual information with popular interests automatically. To this concern, we present a news detection and pushing system, called Me-Digger (Multimedia News Digger), which not only effectively detects emerging topics from social streams but also provides the corresponding information in multiple modalities. Me-digger is the first systematic effort to leverage three sources of data, that is, Twitter, Flickr and Google news, to output with vivid visual and textual contents on emerging topics. Enabled by a novel general-structured high-order co-clustering approach, it has a more accurate detection of emerging topics compared to the existing methods on micro-blog social streams. Bing-Kun Bao, Weiqing Min, Changsheng Xu |
ACM Multimedia | 1 |
| 2012 | Inductive Robust Principal Component AnalysisabstractIn this paper we address the error correction problem that is to uncover the low-dimensional subspace structure from high-dimensional observations, which are possibly corrupted by errors. When the errors are of Gaussian distribution, Principal Component Analysis (PCA) can find the optimal (in terms of least-square-error) low-rank approximation to highdimensional data. However, the canonical PCA method is known to be extremely fragile to the presence of gross corruptions. Recently, Wright et al. established a so-called Robust Principal Component Analysis (RPCA) method, which can well handle grossly corrupted data [14]. However, RPCA is a transductive method and does not handle well the new samples which are not involved in the training procedure. Given a new datum, RPCA essentially needs to recalculate over all the data, resulting in high computational cost. So, RPCA is inappropriate for the applications that require fast online computation. To overcome this limitation, in this paper we propose an Inductive Robust Principal Component Analysis (IRPCA) method. Given a set of training data, unlike RPCA that targets on recovering the original data matrix, IRPCA aims at learning the underlying projection matrix, which can be used to efficiently remove the possible corruptions in any datum. The learning is done by solving a nuclear norm regularized minimization problem, which is convex and can be solved in polynomial time. Extensive experiments on a benchmark human face dataset and two video surveillance datasets show that IRPCA can not only be robust to gross corruptions, but also handle well the new data in an efficient way. Bing-Kun Bao, Guangcan Liu, Changsheng Xu, Shuicheng Yan |
IEEE Trans. Image Process. | 1 |
| 2012 | Hidden-Concept Driven Multilabel Image Annotation and Label RankingabstractConventional semisupervised image annotation algorithms usually propagate labels predominantly via holistic similarities over image representations and do not fully consider the label locality, inter-label similarity, and intra-label diversity among multilabel images. Taking these problems into consideration, we present the hidden-concept driven image annotation and label ranking algorithm (HDIALR), which conducts label propagation based on the similarity over a visually semantically consistent hidden-concepts space. The proposed method has the following characteristics: 1) each holistic image representation is implicitly decomposed into label representations to reveal label locality: the decomposition is guided by the so-called hidden concepts, characterizing image regions and reconstructing both visual and nonvisual labels of the entire image; 2) each label is represented by a linear combination of hidden concepts, while the similar linear coefficients reveal the inter-label similarity; 3) each hidden concept is expressed as a respective subspace, and different expressions of the same label over the subspace then induce the intra-label diversity; and 4) the sparse coding-based graph is proposed to enforce the collective consistency between image labels and image representations, such that it naturally avoids the dilemma of possible inconsistency between the pairwise label similarity and image representation similarity in multilabel scenario. These properties are finally embedded in a regularized nonnegative data factorization formulation, which decomposes images representations into label representations over both labeled and unlabeled data for label propagation and ranking. The objective function is iteratively optimized by a convergence provable updating procedure. Extensive experiments on three benchmark image datasets well validate the effectiveness of our proposed solution to semisupervised multilabel image annotation and label ranking problem. Bing-Kun Bao, Teng Li 0001, Shuicheng Yan |
IEEE Trans. Multim. | 1 |
| 2011 | Efficient region-aware large graph construction towards scalable multi-label propagation
Bing-Kun Bao, Bingbing Ni, Yadong Mu, Shuicheng Yan |
Pattern Recognit. | 1 |