VLDB 2026 Research / reviewers in the wild / expert
Xiaoshan Yang
dblp:74/9989
· DBLP profile ↗
97ranked-venue papers
13as first author
69since 2021 · last 2026
0000-0001-5453-9755ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 80 · 11 first-author · 57 since 2021Artificial intelligence and machine learning · 27 · 2 first-author · 21 since 2021Computer networks · 12 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving localization precision in open-vocabulary object detection through reinforcement learning-based model collaboration
Xiaoshan Yang |
Multim. Syst. | 4 |
| 2026 | Toward Visual Grounding: A SurveyabstractVisual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression text. This task simulates the common referential relationships between visual and linguistic modalities, enabling machines to develop human-like multimodal comprehension capabilities. Consequently, it has extensive applications in various domains. However, since 2021, visual grounding has witnessed significant advancements, with emerging new concepts such as grounded pre-training, grounding multimodal LLMs, generalized visual grounding, and giga-pixel grounding, which have brought numerous new challenges. In this survey, we first examine the developmental history of visual grounding and provide an overview of essential background knowledge, including fundamental concepts and evaluation metrics. We systematically track and summarize the advancements, and then meticulously define and organize the various settings to standardize future research and ensure a fair comparison. In the dataset section, we compile a comprehensive list of current relevant datasets, conduct a fair comparative analysis, and provide ultimate performance prediction to inspire the development of new standard benchmarks. Additionally, we delve into numerous applications and highlight several advanced topics. Finally, we outline the challenges confronting visual grounding and propose valuable directions for future research, which may serve as inspiration for subsequent researchers. By extracting common technical details, this survey encompasses the representative work in each subtopic over the past decade. To the best of our knowledge, this paper represents the most comprehensive overview currently available in the field of visual grounding. This survey is designed to be suitable for both beginners and experienced researchers, serving as an invaluable resource for understanding key concepts and tracking the latest research developments. Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, Yaowei Wang 0001, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Group-Relative Visual Discrimination Enhancement for Unlocking Intrinsic Capability of MLLMsabstractAlthough Multimodal Large Language Models (MLLMs) have shown remarkable generalization across diverse vision-language tasks, recent studies reveal their limitations in visual discrimination. These challenges arise not from insufficient model capacity, but from existing training paradigms that favor linguistic priors over detailed visual analysis. While existing approaches address this limitation through external interventions such as feature integration or knowledge augmentation, we propose a Group-Relative Visual Discrimination Enhancement framework to unlock intrinsic capability of MLLMs and requires no external resources. Our method introduces a Group-Relative Reinforcement Learning paradigm equipped with a lightweight Visual Patch Selection Plugin to dynamically select discriminative visual tokens. The framework establishes a self-feedback loop between visual encoder and language decoder, leveraging the dual reward-penalty signals derived from the model’s internal language feedback to optimize the visual focus, thereby enhancing the model’s visual discrimination capabilities. Extensive experimental results across six visual recognition benchmarks and two VQA benchmarks demonstrate the effectiveness of our method. Code is available at https://github.com/FannierPeng/GROVE. Xiaoshan Yang, Yaowei Wang 0001, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Fed-DiffLoRA: Personalized Federated Style Transfer for T2I Diffusion ModelsabstractLow-Rank Adaptation (LoRA) merging enables efficient customization of T2I diffusion models; nevertheless, centralized aggregation raises serious privacy concerns. While federated LoRA adaptations mitigate these risks, diffusion models present unique challenges: structural heterogeneity and vulnerability to member inference attacks. To overcome these limitations, we propose Fed-DiffLoRA, a privacy-preserving framework that securely aggregates LoRA adapters. At the client level, we disentangle user-specific adaptations into two orthogonal subspaces: content LoRAs preserving semantic fidelity and style LoRAs encoding stylistic features, isolating sensitive attributes from stylistic components. We design a learnable aggregation operator that dynamically optimizes cross-client style LoRA fusion based on the semantic vectors of clients' content LoRAs, achieving high-fidelity style blending while suppressing client-identifiable patterns. We also provide theoretical guarantees for convergence and privacy. Extensive experiments validate the effectiveness of the proposed approach, achieving substantial reductions in attack success rates and consistently high stylization fidelity. Fan Qi, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Image Process. | 3 |
| 2026 | Progressive Learning of Instance-Level Proxy Semantics for Few-Shot Action RecognitionabstractFew-shot action recognition is a crucial task for mitigating the challenges of data scarcity in video understanding. Recent advancements in large-scale pre-trained models have introduced the potential of incorporating semantic knowledge from multi-modal pre-trained models, such as CLIP, to alleviate these challenges. Although some progress have been made, existing methods still rely on class-level text embeddings that are inherently low in diversity, limiting their ability to generalize to unseen actions. To overcome this limitation, we propose a novel framework called Progressive Learning of Instance-Level Proxy Semantics (ProLIPS). ProLIPS integrates Proxy Semantic Diffusion (PSD) to generate rich, instance-level proxy semantic features with diverse semantic contents and temporal dynamics, utilizing a multi-step CLIP-guidance mechanism and a time-conditioned reverse diffusion process. Our approach preserves the diversity of semantic-aligned visual features, significantly improving the generalization and robustness of few-shot action recognition. Extensive experiments on five challenging benchmarks demonstrate the effectiveness of ProLIPS. Xiaoshan Yang, Yaowei Wang 0001, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2026 | Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain AdaptationabstractIn this paper, we study the challenging task of Few-Shot Video Domain Adaptation (FSVDA). The multimodal nature of videos introduces unique challenges, necessitating the simultaneous consideration of both domain alignment and modality collaboration in a few-shot scenario, which is ignored in previous literature. We observe that, under the influence of domain shift, the generalization performance on the target domain of each individual modality, as well as that of fused multimodal features, is constrained. Because each modality is comprised of coupled features with multiple components that exhibit different domain shifts. This variability increases the complexity of domain adaptation, thereby reducing the effectiveness of multimodal feature integration. To address these challenges, we introduce a novel framework of Modality-Collaborative Low Rank Decomposers (MC-LRD) to decompose modality-unique and modality-shared features with different domain shift levels from each modality that are more friendly for domain alignment. The MC-LRD comprises multiple decomposers for each modality and Multimodal Decomposition Routers (MDR). Each decomposer has progressively shared parameters across different modalities. The MDR is leveraged to selectively activate the decomposers to produce modality-unique and modality-shared features. To ensure efficient decomposition, we apply orthogonal decorrelation constraints separately to decomposers and sub routers, enhancing their diversity. Furthermore, we propose a cross-domain activation consistency loss to guarantee that target and source samples of the same category exhibit consistent activation preferences of the decomposers, thereby facilitating domain alignment. Extensive experimental results on three public benchmarks demonstrate that our model achieves significant improvements over existing methods. Yuyang Wanyan, Xiaoshan Yang, Weiming Dong, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2025 | Pseudo Informative Episode Construction for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) studies how to empower the machine learning system to learn novel classes with only a few annotated examples continually. To tackle the FSCIL task, recent state-of-the-art methods propose to employ the meta-learning mechanism, which constructs the pseudo incremental episodes/tasks in the training phase. However, these methods only select part of the base classes to construct the pseudo novel classes in the feature space of the base classes, which cannot mimic the real novel classes of the testing scenario. To deal with this problem, we propose a new Pseudo Informative Episode Construction (PIEC) framework. Specifically, we first perform distribution-level mixing to generate a set of pseudo novel classes in the feature space of the novel class. Then, we propose two diversity criteria to select the informative pseudo novel classes that have large discrepancies with each other and high information gain over the base classes to construct the pseudo incremental session. In this way, we can allow the model to learn rich new concepts beyond the base classes as in the real incremental session during the episodic training procedure, thus improving its generalization ability. Extensive experiments on three popular classification benchmarks (i.e., CUB200, miniImageNet, and CIFAR100) show that the proposed framework can outperform other state-of-the-art methods. Xiaoshan Yang, Changsheng Xu |
AAAI | 2 |
| 2025 | Pilot: Building the Federated Multimodal Instruction Tuning FrameworkabstractIn this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal instruction data on distributed devices. To solve the new task, we propose a federated multimodal instruction tuning framework(Pilot). Our framework integrates two-stage of ``adapter on adapter” into the connector of the vision encoder and the LLM. In stage 1, we extract task-specific features and client-specific features from visual information. In stage 2, we build the cross-task Mixture-of-Adapters(CT-MoA) module to perform cross-task interaction. Each client can not only capture personalized information of local data and learn task-related multimodal information, but also learn general knowledge from other tasks. In addition, we introduce an adaptive parameter aggregation strategy for text training parameters, which optimizes parameter aggregation by calculating weights based on the euclidean distance between parameters, so that parameter aggregation can benefit from positive effects to the greatest extent while effectively reducing negative effects. Our framework can collaboratively exploit distributed data from different local clients to learn cross-task knowledge without being affected by the task heterogeneity during instruction tuning. The effectiveness of our method is verified in two different cross-task scenarios. Baochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang 0001, Changsheng Xu |
AAAI | 2 |
| 2025 | M3E: Mixture of Multi-scale Multi-modal Experts for Time Series Forecasting
Shaobo Xie, Chenlin Zhao, Xiaoshan Yang |
ICIG (1) | 4 |
| 2025 | Towards A Real-World Road Damage Detection DatasetabstractRoad damage represents a serious challenge to the health of road infrastructure and driving safety, making deep learning-based image analysis for road damage detection (RDD) an important research focus. The limited diversity in road damage types, road image collection, size, environment, and imperfect damage definitions within current RDD datasets restrict the real-world applications of RDD. To address this issue, this paper constructs PCL-RDD, a new and extensive RDD dataset. It comprises 24,765 road images, 54,732 instances, and 19 types of road damage. Compared to the existing datasets that mainly include common road damage, the proposed dataset contains a variety of rare and urgent road damages. Besides, we collect road facility-related damages, which also affect traffic safety. We evaluate eight well-established object detection algorithms on the dataset, highlighting the limitations of state-of-the-art detection algorithms under complex conditions. This study contributes a significant dataset to the RDD field and can advance artificial intelligence in both city infrastructure management and environmental perception for autonomous driving. The dataset is available at https://github.com/humh-c/PCL-RDD. Menghao Hu, Zuogan Tang, Xiaoshan Yang, Zhe Wu 0006, Zhouxin Yang, Shaocong Wu, Yaguang Song, Kui Hou, Yaowei Wang 0001 |
ICME | 3 |
| 2025 | Graph Prompts: Adapting Video Graph for Video Question AnsweringabstractDue to the dynamic nature in videos, it is evident that perceiving and reasoning about temporal information are the key focus of Video Question Answering (VideoQA). In recent years, several methods have explored relationship-level temporal modeling with graph-structured video representation. Unfortunately, these methods heavily rely on the question text, thus making it challenging to perceive and reason about video content that is not explicitly mentioned in the question. To address the above challenge, we propose Graph Prompts-based VideoQA (GP-VQA), which adopts a video-based graph structure for enhanced video understanding. The proposed GP-VQA contains two stages, i.e., pre-training and prompt tuning. In pre-training, we define the pretext task that requires GP-VQA to reason about the randomly masked nodes or edges in the video graph, thus prompting GP-VQA to learn the reasoning ability with video-guided information. In prompt-tuning, we organize the textual question into question graph and implement message passing from video graph to question graph, therefore inheriting the video-based reasoning ability from video graph completion to VideoQA. Extensive experiments on various datasets have demonstrated the promising performance of GP-VQA. Yiming Li 0008, Xiaoshan Yang, Bing-Kun Bao, Changsheng Xu |
IJCAI | 2 |
| 2025 | VidCog: Empowering LLM with Long Video Understanding via Human-like Temporal Cognitive LoopabstractComprehending long-form videos, with their extensive temporal contexts and rich semantic complexities, remains a frontier challenge in video understanding. Recently, many existing methods offer promise for long video understanding yet often exhibit operational inefficiencies and suboptimal reasoning. These core challenges typically stem from fragmented multi-step reasoning, unreliable iterative control over information gathering, and visual retrieval strategies that inadequately adapt to query-aware granularities. In this paper, we propose a novel framework named VidCog that mimicks human cognitive processes to achieve robust and efficient long video understanding with Large Language Models (LLMs). VidCog features a Unified Reasoning Engine (URE) that transforms the discrete reasoning tasks into a single, cohesive LLM invocation, and a Contrastive Policy-Optimized Reasoning Gate (CPRG) that learns from relative preferences among contrastive query-exemplars to ensure reliable iterative decision-making. Furthermore, we propose Triadic Optimal Transport Visual Evidence Miner (TOT-VEM) to adaptively capture global-local temporal visual evidence by modeling it as a novel triadic optimal transport problem. Experiments on challenging long-video benchmarks demonstrate that VidCog consistently outperforms the strong baseline in both reasoning accuracy and efficiency, validating the superiority of the human-like cognitive loop. Xiaoshan Yang, Changsheng Xu |
MMAsia | 2 |
| 2025 | Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI AutomationabstractIn recent years, Multimodal Large Language Models (MLLMs) have been extensively utilized for multimodal reasoning tasks, including Graphical User Interface (GUI) automation. Unlike general offline multimodal tasks, GUI automation is executed in online interactive environments, necessitating step-by-step decision-making based on the real-time status of the environment. This task has a lower tolerance for decision-making errors at each step, as any mistakes may cumulatively disrupt the process and potentially lead to irreversible outcomes like deletions or payments. To address these issues, we introduce a pre-operative critic mechanism that provides effective feedback prior to the actual execution, by reasoning about the potential outcome and correctness of actions. Specifically, we propose a Suggestion-aware Group Relative Policy Optimization (S-GRPO) strategy to construct our pre-operative critic model GUI-Critic-R1, incorporating a novel suggestion reward to enhance the reliability of the model's feedback. Furthermore, we develop a reasoning-bootstrapping based data collection pipeline to create a GUI-Critic-Train and a GUI-Critic-Test, filling existing gaps in GUI critic data. Static experiments on the GUI-Critic-Test across both mobile and web domains reveal that our GUI-Critic-R1 offers significant advantages in critic accuracy compared to current MLLMs. Dynamic evaluation on GUI automation benchmark further highlights the effectiveness and superiority of our model, as evidenced by improved success rates and operational efficiency. The code is available at https://github.com/X-PLUG/MobileAgent/tree/main/GUI-Critic-R1. Yuyang Wanyan, Haiyang Xu 0001, Junyang Wang 0001, Jiabo Ye, Yutong Kou, Ming Yan 0008, Fei Huang 0002, Xiaoshan Yang, Weiming Dong, Changsheng Xu |
NeurIPS | 10 |
| 2025 | A Comprehensive Review of Few-Shot Action Recognition
Yuyang Wanyan, Xiaoshan Yang, Weiming Dong, Changsheng Xu |
Int. J. Comput. Vis. | 2 |
| 2025 | Component-Coordinated and Uncertainty-Enhanced LoRA for Few-shot Source-Free Domain Adaptive Object Detection
Xiaoshan Yang |
Neurocomputing | 3 |
| 2025 | FedMRG: federated medical report generation via text-aware learning rate adjustment and multi-level prototype collaboration
Hichem Metmer, Xiaoshan Yang |
Multim. Syst. | 2 |
| 2025 | Compact Latent Primitive Space Learning for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) aims to recognize novel compositions formed by known primitives (attribute and object). The key challenge of CZSL is the visual diversity of the primitive caused by the dependencies of attributes and objects. To solve this problem, most existing methods attempt to mine primitive-invariant features shared in all compositions or learn primitive-variant features specialized for each composition. However, these methods overlook that the primitives have inherent similarities and differences in different compositions, i.e., one primitive may exhibit a common visual appearance under some compositions, but have different expressions in other partial compositions. To sufficiently explore the partial similarity and visual diversity of primitives, we propose a compact latent primitive space learning framework, which explicitly leverages various codewords to encode the primitive features to make a balance between generality and diversity. Specifically, we borrow the idea from discriminative sparse coding to learn these representative codewords to build the latent primitive space. Through the sparse reconstruction loss, contrastive loss and orthogonal constraint, our model can adaptively reconstruct the primitive features according to the similarity weights between the primitive features and codewords. Comprehensive experiments on four benchmarks demonstrate that the proposed method achieves better performance than previous methods. Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2025 | SPDQ: Synergetic Prompts as Disentanglement Queries for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) aims to identify novel compositions formed by known primitives (attributes and objects). Motivated by recent advancements in pre-trained vision-language models such as CLIP, many methods attempt to fine-tune CLIP for CZSL and achieve remarkable performance. However, the existing CLIP-based CZSL methods focus mainly on text prompt tuning, which lacks the flexibility to dynamically adapt both modalities. To solve this issue, an intuitive solution is to additionally introduce visual prompt tuning. This insight is not trivial to achieve because effectively learning prompts for CZSL involves the challenge of entanglement between visual primitives as well as appearance shifts in different compositions. In this paper, we propose a novel Synergetic Prompts as Disentanglement Queries (SPDQ) framework for CZSL. It can disentangle primitive features based on synergetic prompts to jointly alleviate these challenges. Specifically, we first design a low-rank primitive modulator to produce synergetic adaptive attribute and object prompts based on prior knowledge of each instance for model adaptation. Then, we additionally utilize text prefix prompts to construct synergetic prompt queries, which are used to resample corresponding visual features from local visual patches. Comprehensive experiments conducted on three benchmarks demonstrate that our SPDQ approach achieves state-of-the-art results. Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2025 | Health-oriented Multimodal Food Question Answering with Implicit and Explicit KnowledgeabstractHealth-oriented food analysis has become a research hotspot in recent years because it can help people keep away from unhealthy diets. Remarkable advancements have been made in recipe retrieval, food recommendation, nutrition analysis, and calorie estimation. However, existing works still cannot well balance the individual preference and the health. Multimodal food question and answering (MFQA) presents substantial promise for practical applications, yet it remains underexplored. In this article, we introduce a health-oriented MFQA dataset with 9,000 Chinese question−answer pairs based on a multimodal food knowledge graph (MFKG) collected from a food-sharing Web site. Additionally, we propose a novel framework for MFQA in the health domain that leverages implicit general knowledge and explicit domain-specific knowledge. The framework comprises four key components: implicit general knowledge injection module (IGKIM), explicit domain-specific knowledge retrieval module (EDKRM), ranking module, and answer module. The IGKIM facilitates knowledge acquisition at both the feature and text levels. The EDKRM retrieves the most relevant candidate knowledge from the knowledge graph based on the given question. The ranking module sorts the results retrieved by EDKRM and further retrieve candidate knowledge relevant to the problem. Subsequently, the answer module thoroughly analyzes the multimodal information in the query along with the retrieved relevant knowledge to predict accurate answers. Extensive experimental results on the MFQA dataset demonstrate the effectiveness of our proposed method. The code and dataset are available at https://github.com/Wjianghai/HMFQA . Menghao Hu, Yaguang Song, Xiaoshan Yang, Yaowei Wang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Modality-Collaborative Test-Time Adaptation for Action RecognitionabstractVideo-based Unsupervised Domain Adaptation (VUDA) method improves the generalization of the video model, en-abling it to be applied to action recognition tasks in different environments. However, these methods require contin-uous access to source data during the adaptation process, which are impractical in real scenarios where the source videos are not available with concerns in transmission efficiency or privacy issues. To address this problem, in this paper, we focus on the Multimodal Video Test- Time Adaptation (MVTTA) task. Existing image-based TTA methods cannot be directly applied to this task because videos have domain shifts in multimodal and temporal, which brings difficulties to adaptation. To address the above challenges, we propose a Modality-Collaborative Test-Time Adaptation (MC-TTA) Network. MC-TTA contains maintain teacher and student memory banks respectively for generating pseudo-prototypes and target-prototypes. In the teacher model, we propose Self-assembled Source-friendly Feature Reconstruction (SSFR) to encourage the teacher memory bank to store features that are more likely to be consistent with the source distribution. Through multimodal prototype alignment and cross-modal relative consistency, our method can effectively alleviate domain shift in videos. We evaluate the proposed model on four public video datasets. The results show that our model outperforms existing state-of-the-art methods. Baochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang 0001, Changsheng Xu |
CVPR | 2 |
| 2024 | VG-Annotator: Vision-Language Models as Query Annotators for Unsupervised Visual GroundingabstractVisual grounding focuses on localizing objects referred to by natural language queries. Existing fully and weakly supervised methods rely on a mass of language queries for training. However, collecting natural language queries corresponding to specific objects by annotators is expensive. To reduce the reliance on human-written queries, we propose a novel unsupervised visual grounding framework named VG-Annotator. Different from the existing unsupervised methods that rely on manually designed rules to link objects and language queries. The key idea of VG-Annotator lies in that vision-language pre-trained (VLP) generation models can be language query annotators. Thanks to the powerful multi-modal understanding ability implicitly learned from large-scale pre-training, we consider stimulating models to explicitly generate appropriate descriptions for specific objects in natural language. To this end, we explore a series of multi-modal instructions to indicate which object should be described. We also introduce a supervised fine-tuning process to teach the vision-language models to follow the instructions. Extensive experiments show that the proposed method obtains high-quality language queries. The visual grounding model trained with the generated queries outperforms state-of-the-art unsupervised methods on five widely used datasets. Jiabo Ye, Xiaoshan Yang, Zhenru Zhang, Anwen Hu, Ming Yan 0008, Ji Zhang 0011, Liang He 0001, Xin Lin 0001 |
ICME | 3 |
| 2024 | Libra: Building Decoupled Vision System on Large Language ModelsabstractIn this work, we introduce **Libra**, a prototype model with a decoupled vision system on a large language model (LLM). The decoupled vision system decouples inner-modal modeling and cross-modal interaction, yielding unique visual information modeling and effective cross-modal comprehension. Libra is trained through discrete auto-regressive modeling on both vision and language inputs. Specifically, we incorporate a routed visual expert with a cross-modal bridge module into a pretrained LLM to route the vision and language flows during attention computing to enable different attention patterns in inner-modal modeling and cross-modal interaction scenarios. Experimental results demonstrate that the dedicated design of Libra achieves a strong MLLM baseline that rivals existing works in the image-to-text scenario with merely 50 million training data, providing a new perspective for future multimodal foundation models. Code is available at https://github.com/YifanXu74/Libra. Yifan Xu 0008, Xiaoshan Yang, Yaguang Song, Changsheng Xu |
ICML | 2 |
| 2024 | HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual GroundingabstractVisual grounding, which aims to ground a visual region via natural language, is a task that heavily relies on cross-modal alignment. Existing works utilized uni-modal pre-trained models to transfer visual or linguistic knowledge separately while ignoring the multimodal corresponding information. Motivated by recent advancements in contrastive language-image pre-training and low-rank adaptation (LoRA) methods, we aim to solve the grounding task based on multimodal pre-training. However, there exists significant task gaps between pre-training and grounding. Therefore, to address these gaps, we propose a concise and efficient hierarchical multimodal fine-grained modulation framework, namely HiVG. Specifically, HiVG consists of a multi-layer adaptive cross-modal bridge and a hierarchical multimodal low-rank adaptation (HiLoRA) paradigm. The cross-modal bridge can address the inconsistency between visual features and those required for grounding, and establish a connection between multi-level visual and text features. HiLoRA prevents the accumulation of perceptual errors by adapting the cross-modal features from shallow to deep layers in a hierarchical manner. Experimental results on five datasets demonstrate the effectiveness of our approach and showcase the significant grounding capabilities as well as promising energy efficiency advantages. The project page: https://github.com/linhuixiao/HiVG. Linhui Xiao, Xiaoshan Yang, Yaowei Wang 0001, Changsheng Xu |
ACM Multimedia | 2 |
| 2024 | Part-Aware Prompt Tuning for Weakly Supervised Referring Expression Grounding
Chenlin Zhao, Jiabo Ye, Yaguang Song, Ming Yan 0008, Xiaoshan Yang, Changsheng Xu |
MMM (3) | 5 |
| 2024 | OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring ModelingabstractConstrained by the separate encoding of vision and language, existing grounding and referring segmentation works heavily rely on bulky Transformer-based fusion en-/decoders and a variety of early-stage interaction technologies. Simultaneously, the current mask visual language modeling (MVLM) fails to capture the nuanced referential relationship between image-text in referring tasks. In this paper, we propose **OneRef**, a minimalist referring framework built on the modality-shared one-tower transformer that unifies the visual and linguistic feature spaces. To modeling the referential relationship, we introduce a novel MVLM paradigm called Mask Referring Modeling (**MRefM**), which encompasses both referring-aware mask image modeling and referring-aware mask language modeling. Both modules not only reconstruct modality-related content but also cross-modal referring content. Within MRefM, we propose a referring-aware dynamic image masking strategy that is aware of the referred region rather than relying on fixed ratios or generic random masking schemes. By leveraging the unified visual language feature space and incorporating MRefM's ability to model the referential relations, our approach enables direct regression of the referring results without resorting to various complex techniques. Our method consistently surpasses existing approaches and achieves SoTA performance on both grounding and segmentation tasks, providing valuable insights for future research. Our code and models are available at https://github.com/linhuixiao/OneRef. Linhui Xiao, Xiaoshan Yang, Yaowei Wang 0001, Changsheng Xu |
NeurIPS | 2 |
| 2024 | An open chest X-ray dataset with benchmarks for automatic radiology report generation in French
Hichem Metmer, Xiaoshan Yang |
Neurocomputing | 2 |
| 2024 | Self-supervised spatial-temporal feature enhancement for one-shot video object detection
Xiaoshan Yang |
Neurocomputing | 2 |
| 2024 | Cross-Modal Federated Human Activity RecognitionabstractFederated human activity recognition (FHAR) has attracted much attention due to its great potential in privacy protection. Existing FHAR methods can collaboratively learn a global activity recognition model based on unimodal or multimodal data distributed on different local clients. However, it is still questionable whether existing methods can work well in a more common scenario where local data are from different modalities, e.g., some local clients may provide motion signals while others can only provide visual data. In this article, we study a new problem of cross-modal federated human activity recognition (CM-FHAR), which is conducive to promote the large-scale use of the HAR model on more local devices. CM-FHAR has at least three dedicated challenges: 1) distributive common cross-modal feature learning, 2) modality-dependent discriminate feature learning, 3) modality imbalance issue. To address these challenges, we propose a modality-collaborative activity recognition network (MCARN), which can comprehensively learn a global activity classifier shared across all clients and multiple modality-dependent private activity classifiers. To produce modality-agnostic and modality-specific features, we learn an altruistic encoder and an egocentric encoder under the constraint of a separation loss and an adversarial modality discriminator collaboratively learned in hyper-sphere. To address the modality imbalance issue, we propose an angular margin adjustment scheme to improve the modality discriminator on modality-imbalanced data by enhancing the intra-modality compactness of the dominant modality and increase the inter-modality discrepancy. Moreover, we propose a relation-aware global-local calibration mechanism to constrain class-level pairwise relationships for the parameters of the private classifier. Finally, through decentralized optimization with alternative steps of adversarial local updating and modality-aware global aggregation, the proposed MCARN obtains state-of-the-art performance on both modality-balanced and modality-imbalanced data. Xiaoshan Yang, Baochen Xiong, Yi Huang 0037, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | A Versatile Multimodal Learning Framework for Zero-Shot Emotion RecognitionabstractMulti-modal Emotion Recognition (MER) aims to identify various human emotions from heterogeneous modalities. With the development of emotional theories, there are more and more novel and fine-grained concepts to describe human emotional feelings. Real-world recognition systems often encounter unseen emotion labels. To address this challenge, we propose a versatile zero-shot MER framework to refine emotion label embeddings for capturing inter-label relationships and improving discrimination between labels. We integrate prior knowledge into a novel affective graph space that generates tailored label embeddings capturing inter-label relationships. To obtain multimodal representations, we disentangle the features of each modality into egocentric and altruistic components using adversarial learning. These components are then hierarchically fused using a hybrid co-attention mechanism. Furthermore, an emotion-guided decoder exploits label-modal dependencies to generate adaptive multimodal representations guided by emotion embeddings. We conduct extensive experiments with different multimodal combinations, including visual-acoustic and visual-textual inputs, on four datasets in both single-label and multi-label zero-shot settings. Results demonstrate the superiority of our proposed framework over state-of-the-art methods. Fan Qi, Huaiwen Zhang, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Exploring Multi-Modal Contextual Knowledge for Open-Vocabulary Object DetectionabstractWe explore multi-modal contextual knowledge learned through multi-modal masked language modeling to provide explicit localization guidance for novel classes in open-vocabulary object detection (OVD). Intuitively, a well-modeled and correctly predicted masked concept word should effectively capture the textual contexts, visual contexts, and the cross-modal correspondence between texts and regions, thereby automatically activating high attention on corresponding regions. In light of this, we propose a multi-modal contextual knowledge distillation framework, MMC-Det, to explicitly supervise a student detector with the context-aware attention of the masked concept words in a teacher fusion transformer. The teacher fusion transformer is trained with our newly proposed diverse multi-modal masked language modeling (D-MLM) strategy, which significantly enhances the fine-grained region-level visual context modeling in the fusion transformer. The proposed distillation process provides additional contextual guidance to the concept-region matching of the detector, thereby further improving the OVD performance. Extensive experiments performed upon various detection datasets show the effectiveness of our multi-modal context learning strategy. Yifan Xu 0008, Mengdan Zhang, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Image Process. | 3 |
| 2024 | SgVA-CLIP: Semantic-Guided Visual Adapting of Vision-Language Models for Few-Shot Image ClassificationabstractAlthough significant progress has been made in few-shot learning, most of existing few-shot image classification methods require supervised pre-training on a large amount of samples of base classes, which limits their generalization ability in real world application. Recently, large-scale Vision-Language Pre-trained models (VLPs) have been gaining increasing attention in few-shot learning because they can provide a new paradigm for transferable visual representation learning with easily available text on the Web. However, the VLPs may neglect detailed visual information that is difficult to describe by language sentences, but important for learning an effective classifier to distinguish different images. To address the above problem, we propose a new framework, named Semantic-guided Visual Adapting (SgVA), which can effectively extend vision-language pre-trained models to produce discriminative adapted visual features by comprehensively using an implicit knowledge distillation, a vision-specific contrastive loss, and a cross-modal contrastive loss. The implicit knowledge distillation is designed to transfer the fine-grained cross-modal knowledge to guide the updating of the vision adapter. State-of-the-art results on 13 datasets demonstrate that the adapted visual features can well complement the cross-modal features to improve few-shot image classification. Xiaoshan Yang, Linhui Xiao, Yaowei Wang 0001, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2024 | Recovering Generalization via Pre-Training-Like Knowledge Distillation for Out-of-Distribution Visual Question AnsweringabstractWith the emergence of large-scale multi-modal foundation models, significant improvements have been made towards Visual Question Answering (VQA) in recent years via the “Pre-training and Fine-tuning” paradigm. However, the fine-tuned VQA model, which is more specialized for the downstream training data, may fail to generalize well when there is a distribution shift between the training and test data, which is defined as the Out-of-Distribution (OOD) problem. An intuitive way to solve this problem is to transfer the common knowledge from the foundation model to the fine-tuned VQA model via knowledge distillation for better generalization. However, the generality of distilled knowledge based on the task-specific training data is questionable due to the bias between the training and test data. An ideal way is to adopt the pre-training data to distill the common knowledge shared by the training and OOD test samples, which however is impracticable due to the huge size of pre-training data. Based on the above considerations, in this article, we propose a method, named Pre-training-like Knowledge Distillation (PKD), to imitate the pre-training feature distribution and leverage it to distill the common knowledge, which can improve the generalization performance of the fine-tuned model for OOD VQA. Specifically, we first leverage the in-domain VQA data as guidance and adopt two cross-modal feature prediction networks, which are learned under the supervision of image-text matching loss and feature divergence loss, to estimate pre-training-like vision and text features. Next, we conduct feature-level distillation by explicitly integrating the downstream VQA input features with the predicted pre-training-like features through a memory mechanism. In the meantime, we also conduct model-level distillation by constraining the image-text matching output of the downstream VQA model and the output of the foundation model for the pre-training-like image and text features. Extensive experiments on the VQA-CP v2 and VQA v2 datasets demonstrate the effectiveness of our method. Yaguang Song, Xiaoshan Yang, Yaowei Wang 0001, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2024 | CLIP-VG: Self-Paced Curriculum Adapting of CLIP for Visual GroundingabstractVisual Grounding (VG) is a crucial topic in the field of vision and language, which involves locating a specific region described by expressions within an image. To reduce the reliance on manually labeled data, unsupervised methods have been developed to locate regions using pseudo-labels. However, the performance of existing unsupervised methods is highly dependent on the quality of pseudo-labels and these methods always encounter issues with limited diversity. In order to utilize vision and language pre-trained models to address the grounding problem, and reasonably take advantage of pseudo-labels, we propose CLIP-VG, a novel method that can conduct self-paced curriculum adapting of CLIP with pseudo-language labels. We propose a simple yet efficient end-to-end network architecture to realize the transfer of CLIP to the visual grounding. Based on the CLIP-based architecture, we further propose single-source and multi-source curriculum adapting algorithms, which can progressively find more reliable pseudo-labels to learn an optimal model, thereby achieving a balance between reliability and diversity for the pseudo-language labels. Our method outperforms the current state-of-the-art unsupervised method by a significant margin on RefCOCO/+/g datasets in both single-source and multi-source scenarios, with improvements ranging from 6.78% to 10.67% and 11.39% to 14.87%, respectively. Furthermore, our approach even outperforms existing weakly supervised methods. Linhui Xiao, Xiaoshan Yang, Ming Yan 0008, Yaowei Wang 0001, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2024 | UniQRNet: Unifying Referring Expression Grounding and Segmentation with QRNetabstractReferring expression comprehension aims to align natural language queries with visual scenes, which requires establishing fine-grained correspondence between vision and language. This has important applications in multi-modal reasoning systems. Existing methods typically use text-agnostic visual backbones to extract features independently without considering the specific text input. However, we argue that the extracted visual features can be inconsistent with the referring expression, which hurts multi-modal understanding. To address this, we first propose Query-modulated Refinement Network (QRNet) that leverages language guidance to guide visual feature extraction. However, it only focuses on the grounding task that can only provide coarse-grained annotations in the form of bounding box coordinates. The guidance for the visual backbone is indirect, and the inconsistent issue still exists. To this end, we further propose UniQRNet, a multi-task framework over the QRNet to learn referring expression grounding and segmentation jointly. The framework introduces a multi-task head that leverages fine-grained pixel-level supervision from the segmentation task to directly guide the intermediate layers of QRNet to learn text-consistent visual features. Besides, UniQRNet also includes a loss balance strategy that allows two types of supervision signals to cooperate and optimize the model together. We conduct the most comprehensive comparison experiment covering four major datasets, ten evaluation set and three evaluation metrics used in previous work. UniQRNet outperforms previous state-of-the-art methods by a large margin on both referring comprehensive grounding (1.8%~5.09%) and segmentation tasks (0.57%~5.56%). Ablation and analysis reveal that UniQRNet can improve the consistency of visual features with text input and can bring significant performance improvement. Jiabo Ye, Ming Yan 0008, Haiyang Xu 0001, Qinghao Ye, Yaya Shi, Xiaoshan Yang, Xuwu Wang, Ji Zhang 0011, Liang He 0001, Xin Lin 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | Active Exploration of Multimodal Complementarity for Few-Shot Action RecognitionabstractRecently, few-shot action recognition receives increasing attention and achieves remarkable progress. However, previous methods mainly rely on limited unimodal data (e.g., RGB frames) while the multimodal information remains relatively underexplored. In this paper, we propose a novel Active Multimodal Few-shot Action Recognition (AMFAR) framework, which can actively find the reliable modality for each sample based on task-dependent context information to improve few-shot reasoning procedure. In meta-training, we design an Active Sample Selection (ASS) module to organize query samples with large differences in the reliability of modalities into different groups based on modality-specific posterior distributions. In addition, we design an Active Mutual Distillation (AMD) to capture discriminative task-specific knowledge from the reliable modality to improve the representation learning of unreliable modality by bidirectional knowledge distillation. In meta-test, we adopt Adaptive Multimodal Inference (AMI) to adaptively fuse the modality-specific posterior distributions with a larger weight on the reliable modality. Extensive experimental results on four public benchmarks demonstrate that our model achieves significant improvements over existing unimodal and multimodal methods. Yuyang Wanyan, Xiaoshan Yang, Changsheng Xu |
CVPR | 2 |
| 2023 | Fine-grained Primitive Representation Learning for Compositional Zero-shot ClassificationabstractCompositional zero-shot learning (CZSL) aims to recognize attribute-object compositions that are never seen in the training set. Existing methods solve this problem mainly by learning a single primitive representation for each attribute or object, ignoring the natural intra-attribute or intra-object diversity, i.e., the same attribute (or object) presents dramatically different visual appearance in different compositions. In this paper, we treat the same attribute (or object) under different compositions as fine-grained classes and propose a novel fine-grained primitive representation learning framework to learn more discriminative primitive representations by multiple compact feature sub-spaces. We employ attribute- or object-specific fine-grained primitive prototypes and embed them into a cross-modal space to improve the discriminative ability of primitive representations. Moreover, we leverage semantic-guided sample synthesis to estimate primitive representations of unseen compositions. Extensive experiments on three public benchmarks indicate that our approach can achieve state-of-the-art performance. Xiaoshan Yang, Changsheng Xu |
ICME | 2 |
| 2023 | Iterative Learning with Extra and Inner Knowledge for Long-tail Dynamic Scene Graph GenerationabstractDynamic scene graphs have become a powerful tool for higher-level visual understanding tasks, and the interest in dynamic scene graph generation (dynamic SGG) is grown over time. Recently, numbers of existing methods achieve significant progress in dynamic SGG by capturing temporal information with transformer or recurrent network structures. However, most existing methods only focus on predicting the head predicates, which ignore the long-tail phenomenon, thus the tail predicates are hard to be recognized. In this paper, we propose a novel method named Iterative Learning with Extra and Inner Knowledge (I2LEK) to address the long-tail problem in dynamic SGG. The extra knowledge is obtained from commonsense, while inner knowledge is defined as the temporal evolution patterns of visual relationships. Specifically, we introduce extra knowledge to enrich the representations of predicates in the spatial dimension and adopt inner knowledge to implement knowledge sharing in the temporal dimension. With enriched representations and shared knowledge, I2LEK can accurately predict both the tail and head predicates. Moreover, an iterative learning strategy is proposed to fuse the extra knowledge, inner knowledge, and spatial-temporal context contained in videos, which further enhances the model's understanding of visual relationships. Our experimental results on the public Action Genome dataset demonstrate that our model achieves state-of-the-art performance. Yiming Li 0008, Xiaoshan Yang, Changsheng Xu |
ACM Multimedia | 2 |
| 2023 | Client-Adaptive Cross-Model Reconstruction Network for Modality-Incomplete Multimodal Federated LearningabstractMultimodal federated learning (MFL) is an emerging field that allows many distributed clients, each with multimodal data, to work together to train models targeting multimodal tasks without sharing local data. Whereas, existing methods assume that all modalities for each sample are complete, which limits their practicality. In this paper, we propose a Client-Adaptive Cross-Modal Reconstruction Network (CACMRN) to solve the modality-incomplete multimodal federated learning (MI-MFL). Compared to existing centralized methods for reconstructing missing modality, the local client data in federated learning is typically much less, which makes it challenging to train a reliable reconstruction model that can accurately predict missing data. We propose a cross-modal reconstruction transformer, which can prevent the model overfitting on the local client by exploring instance-instance relationships within the local client and utilizing normalized self-attention to conduct data-depended partial updating. Using federated optimization with alternative local updating and global aggregation, our method can not only collaboratively utilize the distributed data on different local clients to learn the cross-modal reconstruction transformer, but also prevent the reconstruction model from overfitting the data on the local client. Extensive experimental results on three datasets demonstrate the effectiveness of our method. Baochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang 0001, Changsheng Xu |
ACM Multimedia | 2 |
| 2023 | mPLUG-Octopus: The Versatile Assistant Empowered by A Modularized End-to-End Multimodal LLMabstractInspired by the recent developments of large language models (LLMs), we propose mPLUG-Octopus, a versatile conversational assistant designed to provide users with coherent, engaging, and helpful interaction experiences in both text-only and multi-modal scenarios. Unlike traditional pipeline chatting systems, mPLUG-Octopus offers a diverse range of creative capabilities including open-domain QA, multi-turn chatting, and multi-modal creation, all built with a unified multimodal LLM without relying on any external API. With the modularized end-to-end multimodal LLM technology, mPLUG-Octopus efficiently facilitates engaging and open-domain conversation experience. It exhibits a wide range of uni/multi-modal elemental capabilities, enabling it to seamlessly communicate with users on open-domain topics and engage in multi-turn conversations. It also assists users in accomplishing various content creation and application tasks. Our conversational assistant can also be deployed on smart hardware to drive advanced AIGC applications. Qinghao Ye, Haiyang Xu 0001, Ming Yan 0008, Chenlin Zhao, Junyang Wang 0001, Xiaoshan Yang, Ji Zhang 0011, Fei Huang 0002, Jitao Sang 0001, Changsheng Xu |
ACM Multimedia | 6 |
| 2023 | Health-Oriented Multimodal Food Question Answering
Jianghai Wang, Menghao Hu, Yaguang Song, Xiaoshan Yang |
MMM (1) | 4 |
| 2023 | Multi-modal Queried Object Detection in the WildabstractWe introduce MQ-Det, an efficient architecture and pre-training strategy design to utilize both textual description with open-set generalization and visual exemplars with rich description granularity as category queries, namely, Multi-modal Queried object Detection, for real-world detection with both open-vocabulary categories and various granularity. MQ-Det incorporates vision queries into existing well-established language-queried-only detectors. A plug-and-play gated class-scalable perceiver module upon the frozen detector is proposed to augment category text with class-wise visual information. To address the learning inertia problem brought by the frozen detector, a vision conditioned masked language prediction strategy is proposed. MQ-Det's simple yet effective architecture and training strategy design is compatible with most language-queried object detectors, thus yielding versatile applications. Experimental results demonstrate that multi-modal queries largely boost open-world detection. For instance, MQ-Det significantly improves the state-of-the-art open-set detector GLIP by +7.8% AP on the LVIS benchmark via multi-modal queries without any downstream finetuning, and averagely +6.3% AP on 13 few-shot downstream tasks, with merely additional 3% modulating time required by GLIP. Code is available at https://github.com/YifanXu74/MQ-Det. Yifan Xu 0008, Mengdan Zhang, Chaoyou Fu, Peixian Chen, Xiaoshan Yang, Ke Li 0015, Changsheng Xu |
NeurIPS | 5 |
| 2023 | Towards a multimodal human activity dataset for healthcare
Menghao Hu, Mingxuan Luo, Menghua Huang, Wenhua Meng, Baochen Xiong, Xiaoshan Yang, Jitao Sang 0001 |
Multim. Syst. | 6 |
| 2023 | Postpartum pelvic organ prolapse assessment via adversarial feature complementation in heterogeneous data
Mingxuan Luo, Xiaoshan Yang |
Neural Comput. Appl. | 2 |
| 2023 | Category Knowledge-Guided Parameter Calibration for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) aims to adapt generic detectors to the novel categories with only a few annotations, which is an important and realistic task. Although the generic object detection has been widely studied over the past years, the FSOD is under explored. In this paper, we propose a novel Category Knowledge-guided Parameter Calibration (CKPC) framework to solve the FSOD task. We first propagate the category relation information to explore the representative category knowledge. Then, we explore the RoI-RoI and RoI-Category relations to capture the local-global context information to enhance the RoI (Region of Interest) features. Next, we project the knowledge representations of foreground categories into a parameter space by a linear transformation to generate the parameters of the category-level classifier. For the background, we learn a proxy category by concluding the global characteristics of all foreground categories to help ensure the discrepancy between the foreground and background, which is then projected into the parameter space by the same linear transformation. Finally, we leverage the parameters of the category-level classifier to explicitly calibrate the instance-level classifier learned on the enhanced RoI features for both the foreground and background categories to improve the detection performance. We conduct extensive experiments on two popular FSOD benchmarks (i.e., Pascal VOC and MS COCO), and the experimental results show that the proposed framework can outperform state-of-the-art methods. Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Image Process. | 2 |
| 2023 | Zero-Shot Predicate Prediction for Scene Graph ParsingabstractThe scene graph is a structured semantic representation of an image, which represents objects and relationships with vertices and edges, respectively. Since it is impossible to manually label all potential relationships in the real world, some previous methods try to apply the zero-shot method for scene graph generation. However, existing methods take triplet (i.e., hsubject-predicate-objecti) as the basic unit of a relationship. Each element (i.e., subject, predicate, or object) of the unseen relationship is actually seen in the training data. Therefore, they ignore the unseen predicate. To predict the unseen predicate, we introduce a novel task named zero-shot predicate prediction, which is crucial to extending existing scene graph generation methods to recognize more relationship classes. The new task is challenging and cannot be simply resolved through conventional zero-shot learning methods because there is a large intra-class variation of each predicate. Firstly, the large intra-class variation leads to the difficulty of computing the discriminative instancelevel feature of the predicate class. Secondly, the large intraclass variation also brings more difficulties when knowledge is transferred from seen classes to unseen classes. For the first challenge, we propose distilling lexical knowledge of different objects and construct multi-modal representations of pairwise objects to reduce the intra-class variation of the predicate. To respond to the second challenge, we build a compact semantic space where the representations of unseen classes are reconstructed based on the seen classes for zero-shot predicate classification. We evaluate the proposed method on the public dataset Visual Genome. The extensive experiment results under the zeroshot/few-shot/supervised settings demonstrate the effectiveness of the proposed method. Yiming Li 0008, Xiaoshan Yang, Xuhui Huang, Zhe Ma 0001, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2023 | Many Hands Make Light Work: Transferring Knowledge From Auxiliary Tasks for Video-Text RetrievalabstractThe problem of video-text retrieval, which searches videos via natural language descriptions or vice versa, has attracted growing attention due to the explosive scale of videos produced every day. The dominant approaches for this problem follow the pipeline that firstly learns compact feature representations of videos and texts, and then jointly embeds them into a common feature space where matched video-text pairs are close and unmatched pairs are far away. However, most of them neither consider the structural similarities among cross-modal samples in a global view, nor leverage useful information from other relevant retrieval processes. We argue that both information has great potential for video-text retrieval. In this paper, we treat the relevant retrieval processes as auxiliary tasks and we extract useful knowledge from them by exploiting structural similarities via Graph Neural Networks (GNNs). We then progressively transfer the knowledge from auxiliary tasks in a general-to-specific manner to assist the main task of the current retrieval process. Specifically, for the retrieval of the given query, we first construct a sequence of query-graphs whose central queries are chosen from distant to close to the given query. Then we conduct knowledge-guided message passing in each query-graph to exploit regional structural similarities and gather knowledge of different levels from the updated query-graphs with a knowledge-based attention mechanism. Finally, we transfer the extracted useful knowledge from general to specific to assist the current retrieval process. Extensive experimental results show that our model outperforms the state-of-the-arts on four benchmarks. Wei Wang 0354, Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2023 | Counterfactual Scenario-relevant Knowledge-enriched Multi-modal Emotion ReasoningabstractMulti-modal video emotion reasoning (MERV) has recently attracted increasing attention due to its potential application in human-computer interaction. This task needs to not only recognize utterance-level emotions for conspicuous speakers, but also perceive the emotions of non-speakers in videos. Existing methods focus on modeling multi-modal multi-level contexts to capture emotion-relevant clues from the complex scenarios in videos. However, the context information is far from enough to infer the emotion labels of non-speakers due to the large gap between the scenario situation and emotions labels. Inspired by the observation that humans can find solutions to complex problems with the leverage of experience and knowledge, we propose SK-MER , a Scenario-relevant Knowledge-enhanced Multi-modal Emotion Reasoning framework for MERV task, which can leverage external knowledge to enhance the video scenario understanding and emotion reasoning. Specifically, we use scenario concepts extracted from videos to build knowledge subgraphs from external knowledge bases. The knowledge subgraphs are then utilized to obtain scenario-relevant knowledge representations through dynamic knowledge graph attention. Next, we incorporate the knowledge representations into context modeling to enhance emotion reasoning with external scenario-relevant knowledge. In addition, we propose a counterfactual knowledge representation learning approach to obtain more effective scenario-relevant knowledge representations. Extensive experimental results on MEmoR dataset show that the proposed SK-MER framework achieves new state-of-the-art results. Xiaoshan Yang, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Multi-Source Knowledge Reasoning Graph Network for Multi-Modal Commonsense InferenceabstractAs a crucial part of natural language processing, event-centered commonsense inference task has attracted increasing attention. With a given observed event, the intention and reaction of the people involved in the event are required to be inferred with artificial intelligent algorithms. To solve this problem, sequence-to-sequence methods are widely studied, where the event is first encoded into a specific representation and then decoded to generate the results. However, all the existing methods learn the event representation only with the textual information, while the visual information is ignored, which is actually helpful for the commonsense reference. In this article, we first define a new task of multi-modal commonsense reference with both textual and visual information. A new event-centered multi-modal dataset is also provided. Then we propose a multi-source knowledge reasoning graph network to solve this task, where three kinds of relational knowledge are considered. Multi-modal correlations are learned to get the event’s multi-modal representation from a global perspective. Intra-event object relations are explored to capture the fine-grained event feature with an object graph. Inter-event semantic relations are also explored through the external knowledge to understand the semantic associations among events with an event graph. We conduct extensive experiments on the new dataset, and the results show the effectiveness of our method. Xiaoshan Yang, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Self-supervised Calorie-aware Heterogeneous Graph Networks for Food RecommendationabstractWith the rapid development of online recipe sharing platforms, food recommendation is emerging as an important application. Although recent studies have made great progress on food recommendation, they have two shortcomings that are likely to affect the recommendation performance. (1) The relations between ingredients are not considered, which may lead to sub-optimal representations of recipes and further result in the neglect of the user’s personalized ingredient combination preference. (2) Existing methods do not consider the impact of users’ preferences on calories in users’ food decision-making process. In this article, we propose a Self-supervised Calorie-aware Heterogeneous Graph Network (SCHGN) to model the relations between ingredients and incorporate calories of food simultaneously. Specifically, we first incorporate users, recipes, ingredients, and calories into a heterogeneous graph and explicitly present the complex relations among them with directed edges. Then, we explore the co-occurrence relation of ingredients in different recipes via self-supervised ingredient prediction. To capture users’ dynamic preferences on calories of food, we learn calorie-aware user representations by hierarchical message passing and compute a comprehensive user-guided recipe representation by attention mechanism. The final food recommendation is accomplished based on the similarity between a user’s calorie-aware representation and the user-guided representation of a recipe. Extensive experiment results on benchmark datasets demonstrate the effectiveness of the proposed method. Yaguang Song, Xiaoshan Yang, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Dual Scene Graph Convolutional Network for Motivation PredictionabstractHumans can easily infer the motivations behind human actions from only visual data by comprehensively analyzing the complex context information and utilizing abundant life experiences. Inspired by humans’ reasoning ability, existing motivation prediction methods have improved image-based deep classification models using the commonsense knowledge learned by pre-trained language models. However, the knowledge learned from public text corpora is probably incompatible with the task-specific data of the motivation prediction, which may impact the model performance. To address this problem, this paper proposes a dual scene graph convolutional network (dual-SGCN) to comprehensively explore the complex visual information and semantic context prior from the image data for motivation prediction. The proposed dual-SGCN has a visual branch and a semantic branch. For the visual branch, we build a visual graph based on scene graph where object nodes and relation edges are represented by visual features. For the semantic branch, we build a semantic graph where nodes and edges are directly represented by the word embeddings of the object and relation labels. In each branch, node-oriented and edge-oriented message passing is adopted to propagate interaction information between different nodes and edges. Besides, a multi-modal interactive attention mechanism is adopted to cooperatively attend and fuse the visual and semantic information. The proposed dual-SGCN is learned in an end-to-end form by a multi-task co-training scheme. In the inference stage, Total Direct Effect is adopted to alleviate the bias caused by the semantic context prior. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance. Yuyang Wanyan, Xiaoshan Yang, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Cross-Modal Federated Human Activity Recognition via Modality-Agnostic and Modality-Specific Representation LearningabstractIn this paper, we propose a new task of cross-modal federated human activity recognition (CMF-HAR), which is conducive to promote the large-scale use of the HAR model on more local devices. To address the new task, we propose a feature-disentangled activity recognition network (FDARN), which has five important modules of altruistic encoder, egocentric encoder, shared activity classifier, private activity classifier and modality discriminator. The altruistic encoder aims to collaboratively embed local instances on different clients into a modality-agnostic feature subspace. The egocentric encoder aims to produce modality-specific features that cannot be shared across clients with different modalities. The modality discriminator is used to adversarially guide the parameter learning of the altruistic and egocentric encoders. Through decentralized optimization with a spherical modality discriminative loss, our model can not only generalize well across different clients by leveraging the modality-agnostic features but also capture the modality-specific discriminative characteristics of each client. Extensive experiment results on four datasets demonstrate the effectiveness of our method. Xiaoshan Yang, Baochen Xiong, Yi Huang 0037, Changsheng Xu |
AAAI | 1 |
| 2022 | Dynamic Scene Graph Generation via Anticipatory Pre-trainingabstractHumans can not only see the collection of objects in visual scenes, but also identify the relationship between objects. The visual relationship in the scene can be abstracted into the semantic representation of a triple (subject, predicate, object) and thus results in a scene graph, which can convey a lot of information for visual understanding. Due to the motion of objects, the visual relationship between two objects in videos may vary, which makes the task of dynamically generating scene graphs from videos more complicated and challenging than the conventional image-based static scene graph generation. Inspired by the ability of humans to infer the visual relationship, we propose a novel anticipatory pre-training paradigm based on Transformer to explicitly model the temporal correlation of visual relationships in different frames to improve dynamic scene graph generation. In pre-training stage, the model predicts the visual relationships of current frame based on the previous frames by extracting intra-frame spatial information with a spatial encoder and inter-frame temporal correlations with a progressive temporal encoder. In the fine-tuning stage, we reuse the spatial encoder and the progressive temporal encoder while the information of the current frame is combined for predicting the visual relationship. Extensive experiments demonstrate that our method achieves state-of-the-art performance on Action Genome dataset. Yiming Li 0008, Xiaoshan Yang, Changsheng Xu |
CVPR | 2 |
| 2022 | Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual GroundingabstractVisual grounding focuses on establishing fine-grained alignment between vision and natural language, which has essential applications in multimodal reasoning systems. Existing methods use pre-trained query-agnostic visual backbones to extract visual feature maps independently without considering the query information. We argue that the visual features extracted from the visual backbones and the features really needed for multimodal reasoning are inconsistent. One reason is that there are differences between pre-training tasks and visual grounding. Moreover, since the backbones are query-agnostic, it is difficult to completely avoid the inconsistency issue by training the visual backbone end-to-end in the visual grounding framework. In this paper, we propose a Query-modulated Refinement Network (QRNet) to address the inconsistent issue by adjusting intermediate features in the visual backbone with a novel Query-aware Dynamic Attention (QD-ATT) mechanism and query-aware multiscale fusion. The QD-ATT can dynamically compute query-dependent visual attention at the spatial and channel levels of the feature maps produced by the visual backbone. We apply the QRNet to an end-to-end visual grounding framework. Extensive experiments show that the proposed method outperforms state-of-the-art methods on five widely used datasets. Our code is available at https://github.com/LukeForeverYoung/QRNet. Jiabo Ye, Ming Yan 0008, Xiaoshan Yang, Xuwu Wang, Ji Zhang 0011, Liang He 0001, Xin Lin 0001 |
CVPR | 4 |
| 2022 | Attribute-guided Dynamic Routing Graph Network for Transductive Few-shot LearningabstractMotivated by the structured form of human cognition, attributes have been introduced in few-shot classification to learn more representative sample features. However, existing attribute-based methods usually treat the importance of different attributes as equals to conclude the sample relations, which cannot distinguish the classes with many similar attributes well. In order to address this problem, we propose an Attribute-guided Dynamic Routing Graph Network (ADRGN) to explicitly learn task-dependent attribute importance scores to help explore the sample relations in a fine-grained manner for adaptive graph-based inference. Specifically, we first leverage a CNN backbone and a transformation network to generate attribute-specific sample representations according to attribute annotations. Next, we treat the attribute-specific sample representations as visual primary capsules and employ an inter-sample routing to explore the visual diversity of each attribute in the current task. Based on the generated diversity capsules, we perform an inter-attribute routing to explore the relations between different attributes to predict the visual attribute importance scores. Meanwhile, we design an attribute semantic routing module to predict the semantic attribute importance from the semantic attribute embeddings to help the learning of the visual attribute importance prediction with a knowledge distillation strategy. Finally, we utilize the visual attribute importance scores to adaptively aggregate sample similarities computed based on the attribute-specific representations to capture the global fine-grained sample relations for message passing and graph-based inference. Experimental results on three few-shot classification benchmarks show that the proposed ADRGN obtains state-of-the-art performance. Xiaoshan Yang, Ming Yan 0008, Changsheng Xu |
ACM Multimedia | 2 |
| 2022 | Relative Alignment Network for Source-Free Multimodal Video Domain AdaptationabstractVideo domain adaptation aims to transfer knowledge from labeled source videos to unlabeled target videos. Existing video domain adaptation methods require full access to the source videos to reduce the domain gap between the source and target videos, which are impractical in real scenarios where the source videos are not available with concerns in transmission efficiency or privacy issues. To address this problem, in this paper, we propose to solve a source-free domain adaptation task for videos where only a pre-trained source model and unlabeled target videos are available for learning a multimodal video classification model. Existing source-free domain adaptation methods cannot be directly applied to this task, since videos always suffer from domain discrepancy along both the multimodal and temporal aspects, which brings difficulties in domain adaptation especially when the source data are unavailable. In this paper, we propose a Multimodal and Temporal Relative Alignment Network (MTRAN) to deal with the above challenges. To explicitly imitate the domain shifts contained in the multimodal information and the temporal dynamics of the source and target videos, we divide the target videos into two splits according to the self-entropy values of the classification results. The low-entropy videos are deemed to be source-like while the high-entropy videos are deemed to be target-like. Then, we adopt a self-entropy-guided MixUp strategy to generate synthetic samples and hypothetical samples as instance-level based on source-like and target-like videos, and push each synthetic sample to be similar with the corresponding hypothetical sample that is slightly closer to the source-like videos than the synthetic sample by multimodal and temporal relative alignment schemes. We evaluate the proposed model on four public video datasets. The results show that our model outperforms existing state-of-the-art methods. Yi Huang 0037, Xiaoshan Yang, Ji Zhang 0011, Changsheng Xu |
ACM Multimedia | 2 |
| 2022 | A unified framework for multi-modal federated learning
Baochen Xiong, Xiaoshan Yang, Fan Qi, Changsheng Xu |
Neurocomputing | 2 |
| 2022 | Holographic Feature Learning of Egocentric-Exocentric Videos for Multi-Domain Action RecognitionabstractThough existing cross-domain action recognition methods successfully improve the performance on videos of one view (e.g., egocentric videos) by transferring the knowledge from videos of another view (e.g., exocentric videos), they have limitations in generality because the source and target domains need to be fixed aforehand. In this paper, we propose to solve a more practical task of multi-domain action recognition on egocentric-exocentric videos, which aims to learn a single model to recognize test videos from either egocentric perspective or exocentric perspective by transferring knowledge between two domains. Though previous cross-domain methods can also transfer knowledge from one domain to another one by learning view-invariant representations of two video domains, they are not suitable for the multi-domain action recognition task because they always suffer from the problem of losing view-specific visual information. As a solution to the multi-domain action recognition task, we propose to map a video from either egocentric perspective or exocentric perspective to a global feature space (we call it holographic feature space) that shares both view-invariant and view-specific visual knowledge of two views. Specially, we decompose the video feature into view-invariant component and view-specific component, where view-specific component is written into memory networks for saving view-specific visual knowledge. The final holographic feature combines view-invariant feature and view-specific features of two views based on the memory networks. We demonstrate the effectiveness of the proposed method with extensive experimental results on two public datasets. Moreover, the good performances under the semi-supervised setting show the generality of our model. Yi Huang 0037, Xiaoshan Yang, Junyun Gao, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2022 | The Model May Fit You: User-Generalized Cross-Modal RetrievalabstractIn real-world applications, a cross-model retrieval model trained on multimodal instances without considering differences in data distributions among users, termed as user domain shift, usually cannot generalize well to unknown user domains. In this paper, we define a new task of user-generalized cross-modal retrieval, and propose a novel Meta-Learning Multimodal User Generalization (MLMUG) method to solve it. MLMUG simulates the user domain shift with meta-optimization, which aims to embed multimodal data effectively and generalize the cross-modal retrieval model to any unknown user domains. We design a cross-modal embedding network with a learnable meta covariant attention module to encode transferable knowledge among different user domains. A user-adaptive metaoptimization scheme is proposed to adaptively aggregate gradients and meta-gradients for fast and stable meta-optimization.We build two benchmarks for user-generalized cross-modal retrieval evaluation. Experiments on the proposed benchmarks validate the generalization of our method compared with several stateof-the-art methods. Xinhong Ma, Xiaoshan Yang, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2022 | Learning Hierarchical Video Graph Networks for One-Stop Video DeliveryabstractThe explosive growth of video data has brought great challenges to video retrieval, which aims to find out related videos from a video collection. Most users are usually not interested in all the content of retrieved videos but have a more fine-grained need. In the meantime, most existing methods can only return a ranked list of retrieved videos lacking a proper way to present the video content. In this paper, we introduce a distinctively new task, namely One-Stop Video Delivery (OSVD) aiming to realize a comprehensive retrieval system with the following merits: it not only retrieves the relevant videos but also filters out irrelevant information and presents compact video content to users, given a natural language query and video collection. To solve this task, we propose an end-to-end Hierarchical Video Graph Reasoning framework (HVGR) , which considers relations of different video levels and jointly accomplishes the one-stop delivery task. Specifically, we decompose the video into three levels, namely the video-level, moment-level, and the clip-level in a coarse-to-fine manner, and apply Graph Neural Networks (GNNs) on the hierarchical graph to model the relations. Furthermore, a pairwise ranking loss named Progressively Refined Loss is proposed based on prior knowledge that there is a relative order of the similarity of query-video, query-moment, and query-clip due to the different granularity of matched information. Extensive experimental results on benchmark datasets demonstrate that the proposed method achieves superior performance compared with baseline methods. Yaguang Song, Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | ECKPN: Explicit Class Knowledge Propagation Network for Transductive Few-Shot LearningabstractRecently, the transductive graph-based methods have achieved great success in the few-shot classification task. However, most existing methods ignore exploring the class-level knowledge that can be easily learned by humans from just a handful of samples. In this paper, we propose an Explicit Class Knowledge Propagation Network (ECKPN), which is composed of the comparison, squeeze and calibration modules, to address this problem. Specifically, we first employ the comparison module to explore the pairwise sample relations to learn rich sample representations in the instance-level graph. Then, we squeeze the instance-level graph to generate the class-level graph, which can help obtain the class-level visual knowledge and facilitate modeling the relations of different classes. Next, the calibration module is adopted to characterize the relations of the classes explicitly to obtain the more discriminative class-level knowledge representations. Finally, we combine the class-level knowledge with the instance-level sample representations to guide the inference of the query samples. We conduct extensive experiments on four few-shot classification benchmarks, and the experimental results show that the proposed ECKPN significantly outperforms the state-of-the art methods. Xiaoshan Yang, Changsheng Xu, Xuhui Huang, Zhe Ma 0001 |
CVPR | 2 |
| 2021 | Few-shot Learning for Multi-Modality TasksabstractRecent deep learning methods rely on a large amount of labeled data to achieve high performance. These methods may be impractical in some scenarios, where manual data annotation is costly or the samples of certain categories are scarce (e.g., tumor lesions, endangered animals and rare individual activities). When only limited annotated samples are available, these methods usually suffer from the overfitting problem severely, which degrades the performance significantly. In contrast, humans can recognize the objects in the images rapidly and correctly with their prior knowledge after exposed to only a few annotated samples. To simulate the learning schema of humans and relieve the reliance on the large-scale annotation benchmarks, researchers start shifting towards the few-shot learning problem: they try to learn a model to correctly recognize novel categories with only a few annotated samples. Jie Chen 0001, Qixiang Ye, Xiaoshan Yang, Shaohua Kevin Zhou, Xiaopeng Hong, Li Zhang 0040 |
ACM Multimedia | 3 |
| 2021 | Multimodal Global Relation Knowledge Distillation for Egocentric Action AnticipationabstractIn this paper, we consider the task of action anticipation on egocentric videos. Previous methods ignore explicit modeling of the global context relation among past and future actions, which is not an easy task due to the vacancy of unobserved videos. To solve this problem, we propose a Multimodal Global Relation Knowledge Distillation (MGRKD) framework to distill the knowledge learned from full videos to improve the action anticipation task on partially observed videos. The proposed MGRKD has a teacher-student learning strategy, where either the teacher or student model has three branches of global relation graph networks (GRGN) to explore the pairwise relations between past and future actions based on three kinds of features (i.e., RGB, motion or object). The teacher model has a similar architecture with the student model, except that the teacher model uses true feature of the future video snippet to build the graph in GRGN while the student model uses a progressive GRU to predict an initialized node feature of future snippet in GRGN. Through the teacher-student learning strategy, the discriminative features and relation knowledge of the past and future actions learned in the teacher model can be distilled to the student model. The experiments on two egocentric video datasets EPIC-Kitchens and EGTEA Gaze+ show that the proposed framework achieves state-of-the-art performances. Yi Huang 0037, Xiaoshan Yang, Changsheng Xu |
ACM Multimedia | 2 |
| 2021 | Zero-shot Video Emotion Recognition via Multimodal Protagonist-aware Transformer NetworkabstractRecognizing human emotions from videos has attracted significant attention in numerous computer vision and multimedia applications, such as human-computer interaction and health care. It aims to understand the emotional response of humans, where candidate emotion categories are generally defined by specific psychological theories. However, with the development of psychological theories, emotion categories become increasingly diverse and fine-grained, samples are also increasingly difficult to collect. In this paper, we investigate a new task of zero-shot video emotion recognition, which aims to recognize rare unseen emotions. Specifically, we propose a novel multimodal protagonist-aware transformer network, which is composed of two branches: one is equipped with a novel dynamic emotional attention mechanism and a visual transformer to learn better visual representations; the other is an acoustic transformer for learning discriminative acoustic representations. We manage to align the visual and acoustic representations with semantic embeddings of fine-grained emotion labels through jointly mapping them into a common space under a noise contrastive estimation objective. Extensive experimental results on three datasets demonstrate the effectiveness of the proposed method. Fan Qi, Xiaoshan Yang, Changsheng Xu |
ACM Multimedia | 2 |
| 2021 | Few-shot Egocentric Multimodal Activity RecognitionabstractActivity recognition based on egocentric multimodal data collected by wearable devices has become increasingly popular recently. However, conventional activity recognition methods face the dilemma of the lack of large-scale labeled egocentric multimodal datasets due to the high cost of data collection. In this paper, we propose a new task of few-shot egocentric multimodal activity recognition, which has at least two significant challenges. On the one hand, it is difficult to extract effective features from the multimodal data sequences of video and sensor signals due to the scarcity of the samples. On the other hand, how to robustly recognize novel activity classes with very few labeled samples becomes another more critical challenge due to the complexity of the multimodal data. To resolve the challenges, we propose a two-stream graph network, which consists of a heterogeneous graph-based multimodal association module and a knowledge-aware activity classifier module. The former uses a heterogeneous graph network to comprehensively capture the dynamic and complementary information contained in the multimodal data stream. The latter learns robust activity classifiers through knowledge propagation among the classifier parameters of different classes. In addition, we adopt episodic training strategy to improve the generalization ability of the proposed few-shot activity recognition model. Experiments on two public datasets show that the proposed model achieves better performances than other baseline models. Jinxing Pan, Xiaoshan Yang, Yi Huang 0037, Changsheng Xu |
MMAsia | 2 |
| 2021 | Unsupervised Video Summarization via Relation-Aware Assignment LearningabstractWe address the problem of unsupervised video summarization that automatically selects key video clips. Most state-of-the-art approaches suffer from two issues: (1) they model video clips without explicitly exploiting their relations, and (2) they learn soft importance scores over all the video clips to generate the summary representation. However, a meaningful video summary should be inferred by taking the relation-aware context of the original video into consideration, and directly selecting a subset of clips with a hard assignment. In this paper, we propose to exploit clip-clip relations to learn relation-aware hard assignments for selecting key clips in an unsupervised manner. First, we consider the clips as graph nodes to construct an assignment-learning graph. Then, we utilize the magnitude of the node features to generate hard assignments as the summary selection. Finally, we optimize the whole framework via a proposed multi-task loss including a reconstruction constraint, and a contrastive constraint. Extensive experimental results on three popular benchmarks demonstrate the favourable performance of our approach. Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2021 | Emotion Knowledge Driven Video Highlight DetectionabstractThis paper addresses video highlight detection which aims to select a small subset of frames according to user's major or special interest. The performances of conventional methods highly depend on large-scale manually labeled training data which are time-consuming and labor-intensive to collect. To deal with this problem, we trace back to the original problem definition and find that whether a user is interested in a specific video segment heavily depends on human's subjective emotions. Leveraging this insight, we introduce an emotion knowledge driven video detection framework for modeling human's general emotion and inferencing highlight strength. Firstly, we obtain the concept-level representation of the video clip with a front-end network. The concepts are used as nodes to build an emotion-related knowledge graph, and their relationships in the graph are modeled via external public knowledge graphs. Then we adopt Siamese GCNs to model the dependencies between nodes in the graph and propagate messages along the edges. Finally, we compute the emotion-aware representation of the video clip based on the GCN layers and further use it to predict the highlight score. Our framework, including the front-end network, graph convolution layers and the highlight mapping network, can be trained in an end-to-end manner with the constraint of a ranking loss. Experiments on two benchmark datasets show that our proposed method performs favorably against the state-of-the-art methods. Fan Qi, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2021 | Learning Coarse-to-Fine Graph Neural Networks for Video-Text RetrievalabstractWe address the problem of video-text retrieval that searches videos via natural language description or vice versa. Most state-of-the-art methods only consider cross-modal learning for two or three data points in isolation, ignoring to get benefit from the structural information of other data points from a global view. In this paper, we propose to exploit the comprehensive relationships among cross-modal samples via Graph Neural Networks (GNN). To improve the discriminative ability for accurately finding the positive sample, a Coarse-to-Fine GNN is constructed, which can progressively optimize the retrieval results via multi-step reasoning. Specifically, we first adopt heuristic edge features to represent relationships. Then we design a scoring module in each layer to rank the edges connected to the query node and drop the edges with lower scores. Finally, to alleviate the class imbalance issue, we propose a random-drop focal loss to optimize the whole framework. Extensive experimental results show that our method consistently outperforms the state-of-the-arts on four benchmarks. Wei Wang 0354, Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2021 | Knowledge-driven Egocentric Multimodal Activity RecognitionabstractRecognizing activities from egocentric multimodal data collected by wearable cameras and sensors, is gaining interest, as multimodal methods always benefit from the complementarity of different modalities. However, since high-dimensional videos contain rich high-level semantic information while low-dimensional sensor signals describe simple motion patterns of the wearer, the large modality gap between the videos and the sensor signals raises a challenge for fusing the raw data. Moreover, the lack of large-scale egocentric multimodal datasets due to the cost of data collection and annotation processes makes another challenge for employing complex deep learning models. To jointly deal with the above two challenges, we propose a knowledge-driven multimodal activity recognition framework that exploits external knowledge to fuse multimodal data and reduce the dependence on large-scale training samples. Specifically, we design a dual-GCLSTM (Graph Convolutional LSTM) and a multi-layer GCN (Graph Convolutional Network) to collectively model the relations among activities and intermediate objects. The dual-GCLSTM is designed to fuse temporal multimodal features with top-down relation-aware guidance. In addition, we apply a co-attention mechanism to adaptively attend to the features of different modalities at different timesteps. The multi-layer GCN aims to learn relation-aware classifiers of activity categories. Experimental results on three publicly available egocentric multimodal datasets show the effectiveness of the proposed model. Yi Huang 0037, Xiaoshan Yang, Junyu Gao 0002, Jitao Sang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Health Status Prediction with Local-Global Heterogeneous Behavior GraphabstractHealth management is getting increasing attention all over the world. However, existing health management mainly relies on hospital examination and treatment, which are complicated and untimely. The emergence of mobile devices provides the possibility to manage people’s health status in a convenient and instant way. Estimation of health status can be achieved with various kinds of data streams continuously collected from wearable sensors. However, these data streams are multi-source and heterogeneous, containing complex temporal structures with local contextual and global temporal aspects, which makes the feature learning and data joint utilization challenging. We propose to model the behavior-related multi-source data streams with a local-global graph, which contains multiple local context sub-graphs to learn short-term local context information with heterogeneous graph neural networks and a global temporal sub-graph to learn long-term dependency with self-attention networks. Then health status is predicted based on the structure-aware representation learned from the local-global behavior graph. We take experiments on the StudentLife dataset, and extensive results demonstrate the effectiveness of our proposed model. Xiaoshan Yang, Junyu Gao 0002, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Find Objects and Focus on Highlights: Mining Object Semantics for Video Highlight Detection via Graph Neural NetworksabstractWith the increasing prevalence of portable computing devices, browsing unedited videos is time-consuming and tedious. Video highlight detection has the potential to significantly ease this situation, which discoveries moments of user's major or special interest in a video. Existing methods suffer from two problems. Firstly, most existing approaches only focus on learning holistic visual representations of videos but ignore object semantics for inferring video highlights. Secondly, current state-of-the-art approaches often adopt the pairwise ranking-based strategy, which cannot enjoy the global information to infer highlights. Therefore, we propose a novel video highlight framework, named VH-GNN, to construct an object-aware graph and model the relationships between objects from a global view. To reduce computational cost, we decompose the whole graph into two types of graphs: a spatial graph to capture the complex interactions of object within each frame, and a temporal graph to obtain object-aware representation of each frame and capture the global information. In addition, we optimize the framework via a proposed multi-stage loss, where the first stage aims to determine the highlight-probability and the second stage leverage the relationships between frames and focus on hard examples from the former stage. Extensive experiments on two standard datasets strongly evidence that VH-GNN obtains significant performance compared with state-of-the-arts. Junyu Gao 0002, Xiaoshan Yang, Yan Li 0068, Changsheng Xu |
AAAI | 3 |
| 2020 | Structured Neural Motifs: Scene Graph Parsing via Enhanced Context
Yiming Li 0008, Xiaoshan Yang, Changsheng Xu |
MMM (2) | 2 |
| 2020 | Multi-hop Interactive Cross-Modal Retrieval
Xuecheng Ning, Xiaoshan Yang, Changsheng Xu |
MMM (2) | 2 |
| 2020 | Discriminative multimodal embedding for event classification
Fan Qi, Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu |
Neurocomputing | 2 |
| 2020 | Asymmetric multi-stage CNNs for small-scale pedestrian detection
Xiaoshan Yang, Changsheng Xu |
Neurocomputing | 2 |
| 2020 | Cross-domain personalized image captioning
Cuirong Long, Xiaoshan Yang, Changsheng Xu |
Multim. Tools Appl. | 2 |
| 2019 | Exploring Feature Representation and Training Strategies in Temporal Action LocalizationabstractTemporal action localization has recently attracted significant interest in the Computer Vision community. However, despite the great progress, it is hard to identify which aspects of the proposed methods contribute most to the increase in localization performance. To address this issue, we conduct ablative experiments on feature extraction methods, fixed-size feature representation methods and training strategies, and report how each influences the overall performance. Based on our findings, we propose a two-stage detector that outperforms the state of the art in THUMOS14, achieving a mAP@tIoU=0.5 equal to 44.20%. Tingting Xie, Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu, Ioannis Patras |
ICIP | 2 |
| 2019 | Biomedia ACM MM Grand Challenge 2019: Using Data Enhancement to Solve Sample UnbalanceabstractThe Biomedia ACM MM Grand Challenge focuses on medical applications with a task to detect and classify abnormalities within gastrointestinal (GI) tract. As a part of the submission for this challenge, several methods we applied are reported in this paper. The data we used is from the KVASIR dataset and the NEETHUS dataset. It contains training and test data in form of image or video. The main challenge of this task is the data's insufficiency and unbalance, which will significantly decrease the performance. To solve this problem and achieve better result, we conduct multiple data enhancement operations on the data. The method is proved to be efficient. Apart from the operations applied on the data, we also test several classification structures include SCNN (Shallow Convolutional Neural Network), SCNN-SVM (Support Vector Machine), ResNet32-SVM, SVM, ResNet16, ResNet32, ResNet50, ResNet101 and Residual Attention Network. We finally chose ResNet50 as the main structure considered with the balance of accuracy and efficiency. We obtain 91.51% precision, 87.45% sensitivity, 99.48% specificity, 87.93% F1-score (the harmonic mean of precision and sensitivity) and 91.40% MCC (Matthews correlation coefficient) on the test dataset. Wenhua Meng, Xiaoshan Yang, Changsheng Xu, Xiaowen Huang 0001 |
ACM Multimedia | 4 |
| 2019 | Multimodal Attribute and Feature Embedding for Activity RecognitionabstractHuman Activity Recognition (HAR) automatically recognizes human activities such as daily life and work based on digital records, which is of great significance to medical and health fields. Egocentric video and human acceleration data comprehensively describe human activity patterns from different aspects, which have laid a foundation for activity recognition based on multimodal behavior data. However, on the one hand, the low-level multimodal signal structures differ greatly and the mapping to high-level activities is complicated. On the other hand, the activity labeling based on multimodal behavior data has high cost and limited data amount, which limits the technical development in this field. In this paper, an activity recognition model MAFE based on multimodal attribute feature embedding is proposed. Before the activity recognition, the middle-level attribute features are extracted from the low-level signals of different modes. On the one hand, the mapping complexity from the low-level signals to the high-level activities is reduced, and on the other hand, a large number of middle-level attribute labeling data can be used to reduce the dependency on the activity labeling data. We conducted experiments on Stanford-ECM datasets to verify the effectiveness of the proposed MAFE method. Yi Huang 0037, Wanting Yu, Xiaoshan Yang, Wei Wang 0354, Jitao Sang 0001 |
MMAsia | 4 |
| 2019 | Time-Guided High-Order Attention Model of Longitudinal Heterogeneous Healthcare Data
Yi Huang 0037, Xiaoshan Yang, Changsheng Xu |
PRICAI (1) | 2 |
| 2019 | Image Captioning by Asking QuestionsabstractImage captioning and visual question answering are typical tasks that connect computer vision and natural language processing. Both of them need to effectively represent the visual content using computer vision methods and smoothly process the text sentence using natural language processing skills. The key problem of these two tasks is to infer the target result based on the interactive understanding of the word sequence and the image. Though they practically use similar algorithms, they are studied independently in the past few years. In this article, we attempt to exploit the mutual correlation between these two tasks. We propose the first VQA-improved image-captioning method that transfers the knowledge learned from the VQA corpora to the image-captioning task. A VQA model is first pretrained on image--question--answer instances. Then, the pretrained VQA model is used to extract VQA-grounded semantic representations according to selected free-form open-ended visual question--answer pairs. The VQA-grounded features are complementary to the visual features, because they interpret images from a different perspective. We incorporate the VQA model into the image-captioning model by adaptively fusing the VQA-grounded feature and the attended visual feature. We show that such simple VQA-improved image-captioning (VQA-IIC) models perform better than conventional image-captioning methods on large-scale public datasets. Xiaoshan Yang, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | A Unified Framework for Multimodal Domain AdaptationabstractDomain adaptation aims to train a model on labeled data from a source domain while minimizing test error on a target domain. Most of existing domain adaptation methods only focus on reducing domain shift of single-modal data. In this paper, we consider a new problem of multimodal domain adaptation and propose a unified framework to solve it. The proposed multimodal domain adaptation neural networks(MDANN) consist of three important modules. (1) A covariant multimodal attention is designed to learn a common feature representation for multiple modalities. (2) A fusion module adaptively fuses attended features of different modalities. (3) Hybrid domain constraints are proposed to comprehensively learn domain-invariant features by constraining single modal features, fused features, and attention scores. Through jointly attending and fusing under an adversarial objective, the most discriminative and domain-adaptive parts of the features are adaptively fused together. Extensive experimental results on two real-world cross-domain applications (emotion recognition and cross-media retrieval) demonstrate the effectiveness of the proposed method. Fan Qi, Xiaoshan Yang, Changsheng Xu |
ACM Multimedia | 2 |
| 2018 | P2T: Part-to-Target Tracking via Deep Regression LearningabstractMost existing part based tracking methods are part-to-part trackers, which usually have two separated steps including part matching and target localization. Different from existing methods, in this paper, we propose a novel part-totarget (P2T) tracker in a unified fashion by inferring target location from parts directly. To achieve this goal, we propose a novel deep regression model for part to target regression in an end-to-end framework via Convolutional Neural Networks. The proposed model is able to not only exploit part context information to preserve object spatial layout structure, but also learn part reliability to emphasize part importance for robust part to target regression. We evaluate the proposed tracker on 4 challenging benchmark sequences, and extensive experimental results demonstrate that our method performs favorably against state-of-the-art trackers because of the powerful capacity of the proposed deep regression model. Junyu Gao 0002, Tianzhu Zhang 0001, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Image Process. | 3 |
| 2018 | Three-Dimensional Attention-Based Deep Ranking Model for Video Highlight DetectionabstractThe video highlight detection task is to localize key elements (moments of user's major or special interest) in a video. Most of existing highlight detection approaches extract features from the video segment as a whole without considering the difference of local features both temporally and spatially. Due to the complexity of video content, this kind of mixed features will impact the final highlight prediction. In temporal extent, not all frames are worth watching because some of them only contain the background of the environment without human or other moving objects. In spatial extent, it is similar that not all regions in each frame are highlights especially when there are lots of clutters in the background. To solve the above problem, we propose a novel three-dimensional (3-D) (spatial+temporal) attention model that can automatically localize the key elements in a video without any extra supervised annotations. Specifically, the proposed attention model produces attention weights of local regions along both the spatial and temporal dimensions of the video segment. The regions of key elements in the video will be strengthened with large weights. Thus, the more effective feature of the video segment is obtained to predict the highlight score. The proposed 3-D attention scheme can be easily integrated into a conventional end-to-end deep ranking model that aims to learn a deep neural network to compute the highlight score of each video segment. Extensive experimental results on the YouTube and SumMe datasets demonstrate that the proposed approach achieves significant improvement over state-of-the-art methods. With the proposed 3-D attention model, video highlights can be accurately retrieved in spatial and temporal dimensions without human supervision in several domains, such as gymnastics, parkour, skating, skiing, surfing, and dog activities, on the public datasets. Yifan Jiao, Zhetao Li, Shucheng Huang, Xiaoshan Yang, Bin Liu 0014, Tianzhu Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Deep-Structured Event Modeling for User-Generated PhotosabstractVision-based event analysis is difficult because of the following challenges. The first challenge is intraclass variation. Photos uploaded by users are sparsely sampled visual appearances of an event over time. Thus, each photo may only capture a single object or scene of a specific complex event. The second challenge is interclass confusion. Photos related to different events may contain similar objects or scenes. Third, unusual events are characterized by scarcity, and only a few samples are available for use in learning event patterns. In this paper, by considering the photo timestamp, we propose a structured event modeling (SEM) framework for event analysis that exploits the temporal information of visual features and event classes in a photo sequence. Specifically, the temporal event patterns of the photo sequence and the relationships of different photos are jointly learned using deep neural networks (convolutional neural networks and recurrent neural networks) and a conditional random field. We evaluate the proposed SEM framework in two applications: multiclass event recognition and unusual event detection in photo sequences. The results of extensive experiments performed on a public event recognition dataset and a collected unusual event dataset demonstrate the effectiveness of the proposed method. Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2018 | Text2Video: An End-to-end Learning Framework for Expressing Text With VideosabstractVideo creation is a challenging and highly profession-al task that generally involves substantial manual efforts. To ease this burden, a better approach is to automatically produce new videos based on clips from the massive amount of existing videos according to arbitrary text. In this paper, we formulate video creation as a problem of retrieving a sequence of videos for a sentence stream. To achieve this goal, we propose a novel multimodal recurrent architecture for automatic video production. Compared with existing methods, the proposed model has three major advantages. First, it is the first completely integrated end-to-end deep learning system for real-world production to the best of our knowledge. We are among the first to address the problem of retrieving a sequence of videos for a sentence stream. Second, it can effectively exploit the correspondence between sentences and video clips through semantic consistency modeling. Third, it can model the visual coherence well by requiring that the produced videos should be organized coherently in terms of visual appearance. We have conducted extensive experiments on two applications, including video retrieval and video composition. The qualitative and quantitative results obtained on two public datasets used in the Large Scale Movie Description Challenge 2016 both demonstrate the effectiveness of the proposed model compared with other state-of-the-art algorithms. Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2017 | Research on endurance evaluation for NAND flash-based solid state driveabstractThe solid state drive based on NAND flash memory has been widely used, but the limited number of Program/Erase operation cycles has lead to its limited lifetime and reliability, reduced the endurance of solid state drive. The lifetime of the NAND flash-based solid state drive depends on its endurance, and the write amplification is an important factor that affects the endurance. It is of great value to test and evaluate the write amplification in real and accurate. The accuracy and credibility of the evaluation model are seldom mentioned in the existing research, and some characteristics of the solid state drive are neglected, therefore, the simulation model based on solid state drive has a credibility problem. In this paper, the actual volume data of programming operation in NAND flash memory chip is obtained by detecting the change signal of the pin on the flash chip, flash memory chip write amplification is tested and quantified, solid state drive Program/Erase wastage evaluation model based on NAND flash memory is studied and established, namely the endurance evaluation model (EEM). The experimental results show that the EEM is feasible and can be used to evaluate the endurance of solid state drive. Xiaoshan Yang, Ligu Zhu, Qicong Zhang |
ICIS | 1 |
| 2017 | Video Highlight Detection via Deep Ranking Modeling
Yifan Jiao, Xiaoshan Yang, Tianzhu Zhang 0001, Shucheng Huang, Changsheng Xu |
PSIVT | 2 |
| 2017 | Deep Relative TrackingabstractMost existing tracking methods are direct trackers, which directly exploit foreground or/and background information for object appearance modeling and decide whether an image patch is target object or not. As a result, these trackers cannot perform well when target appearance changes heavily and becomes different from its model. To deal with this issue, we propose a novel relative tracker, which can effectively exploit the relative relationship among image patches from both foreground and background for object appearance modeling. Different from direct trackers, the proposed relative tracker is robust to localize target object by use of the best image patch with the highest relative score to target appearance model. To model relative relationship among large-scale image patch pairs, we propose a novel and effective deep relative learning algorithm via Convolutional Neural Network. We test the proposed approach on challenging sequences involving heavy occlusion, drastic illumination changes, and large pose variations. Experimental results show that our method consistently outperforms state-of-the-art trackers due to the powerful capacity of the proposed deep relative model. Junyu Gao 0002, Tianzhu Zhang 0001, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Image Process. | 3 |
| 2016 | Abnormal Event Discovery in User Generated PhotosabstractVision based event analysis plays a very critical role in automatically organizing user generated photos. As one of the important tasks in event analysis, abnormal event discovery still does not obtain much attentions. It is difficult because only few samples can be used for event pattern learning. In this paper, by considering the photo taken time, we propose a novel one-class structured event modeling (OSEM) where we explore the temporal event patterns in negative photos of the event using the continuous conditional random field (CRF). With the estimated piecewise training of CRF, the proposed OSEM can be efficiently solved using stochastic gradients descent (SGD) in an end-to-end form. The extensive experimental results on a collected abnormal event dataset demonstrate the effectiveness of the proposed OSEM. Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu |
ACM Multimedia | 1 |
| 2016 | Deep Relative AttributesabstractRelative attribute (RA) learning aims to learn the ranking function describing the relative strength of the attribute. Most of current learning approaches learn a linear ranking function for each attribute by use of the hand-crafted visual features. Different from the existing study, in this paper, we propose a novel deep relative attributes (DRA) algorithm to learn visual features and the effective nonlinear ranking function to describe the RA of image pairs in a unified framework. Here, visual features and the ranking function are learned jointly, and they can benefit each other. The proposed DRA model is comprised of five convolutional neural layers, five fully connected layers, and a relative loss function which contains the contrastive constraint and the similar constraint corresponding to the ordered image pairs and the unordered image pairs, respectively. To train the DRA model effectively, we make use of the transferred knowledge from the large scale visual recognition on ImageNet [1] to the RA learning task. We evaluate the proposed DRA model on three widely used datasets. Extensive experimental results demonstrate that the proposed DRA model consistently and significantly outperforms the state-of-the-art RA learning methods. On the public OSR, PubFig, and Shoes datasets, compared with the previous RA learning results [2], the average ranking accuracies have been significantly improved by about 8%, 9%, and 14%, respectively. Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu, Shuicheng Yan, M. Shamim Hossain, Ahmed Ghoneim |
IEEE Trans. Multim. | 1 |
| 2016 | Semantic Feature Mining for Video Event UnderstandingabstractContent-based video understanding is extremely difficult due to the semantic gap between low-level vision signals and the various semantic concepts (object, action, and scene) in videos. Though feature extraction from videos has achieved significant progress, most of the previous methods rely only on low-level features, such as the appearance and motion features. Recently, visual-feature extraction has been improved significantly with machine-learning algorithms, especially deep learning. However, there is still not enough work focusing on extracting semantic features from videos directly. The goal of this article is to adopt unlabeled videos with the help of text descriptions to learn an embedding function, which can be used to extract more effective semantic features from videos when only a few labeled samples are available for video recognition. To achieve this goal, we propose a novel embedding convolutional neural network (ECNN). We evaluate our algorithm by comparing its performance on three challenging benchmarks with several popular state-of-the-art methods. Extensive experimental results show that the proposed ECNN consistently and significantly outperforms the existing methods. Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2015 | A new discriminative coding method for image classification
Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu |
Multim. Syst. | 1 |
| 2015 | Cross-Domain Feature Learning in MultimediaabstractIn the Web 2.0 era, a huge number of media data, such as text, image/video, and social interaction information, have been generated on the social media sites (e.g., Facebook, Google, Flickr, and YouTube). These media data can be effectively adopted for many applications (e.g., image/video annotation, image/video retrieval, and event classification) in multimedia. However, it is difficult to design an effective feature representation to describe these data because they have multi-modal property (e.g., text, image, video, and audio) and multi-domain property (e.g., Flickr, Google, and YouTube). To deal with these issues, we propose a novel cross-domain feature learning (CDFL) algorithm based on stacked denoising auto-encoders. By introducing the modal correlation constraint and the cross-domain constraint in conventional auto-encoder, our CDFL can maximize the correlations among different modalities and extract domain invariant semantic features simultaneously. To evaluate our CDFL algorithm , we apply it to three important applications: sentiment classification, spam filtering, and event classification. Comprehensive evaluations demonstrate the encouraging performance of the proposed approach. Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2015 | Automatic Visual Concept Learning for Social Event UnderstandingabstractVision-based event analysis is extremely difficult due to the various concepts (object, action, and scene) contained in videos. Though visual concept-based event analysis has achieved significant progress, it has two disadvantages: visual concept is defined manually, and has only one corresponding classifier in traditional methods. To deal with these issues, we propose a novel automatic visual concept learning algorithm for social event understanding in videos. First, instead of defining visual concept manually, we propose an effective automatic concept mining algorithm with the help of Wikipedia, N-gram Web services, and Flickr. Then, based on the learned visual concept, we propose a novel boosting concept learning algorithm to iteratively learn multiple classifiers for each concept to enhance its representative discriminability. The extensive experimental evaluations on the collected dataset well demonstrate the effectiveness of the proposed algorithm for social event understanding. Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu, M. Shamim Hossain |
IEEE Trans. Multim. | 1 |
| 2015 | Boosted Multifeature Learning for Cross-Domain TransferabstractConventional learning algorithm assumes that the training data and test data share a common distribution. However, this assumption will greatly hinder the practical application of the learned model for cross-domain data analysis in multimedia. To deal with this issue, transfer learning based technology should be adopted. As a typical version of transfer learning, domain adaption has been extensively studied recently due to its theoretical value and practical interest. In this article, we propose a boosted multifeature learning (BMFL) approach to iteratively learn multiple representations within a boosting procedure for unsupervised domain adaption. The proposed BMFL method has a number of properties. (1) It reuses all instances with different weights assigned by the previous boosting iteration and avoids discarding labeled instances as in conventional methods. (2) It models the instance weight distribution effectively by considering the classification error and the domain similarity, which facilitates learning new feature representation to correct the previously misclassified instances. (3) It learns multiple different feature representations to effectively bridge the source and target domains. We evaluate the BMFL by comparing its performance on three applications: image classification, sentiment classification and spam filtering. Extensive experimental results demonstrate that the proposed BMFL algorithm performs favorably against state-of-the-art domain adaption methods. Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu, Ming-Hsuan Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2013 | Intrinsic Image Decomposition Using Optimization and User ScribblesabstractIn this paper, we present a novel high-quality intrinsic image recovery approach using optimization and user scribbles. Our approach is based on the assumption of color characteristics in a local window in natural images. Our method adopts a premise that neighboring pixels in a local window having similar intensity values should have similar reflectance values. Thus, the intrinsic image decomposition is formulated by minimizing an energy function with the addition of a weighting constraint to the local image properties. In order to improve the intrinsic image decomposition results, we further specify local constraint cues by integrating the user strokes in our energy formulation, including constant-reflectance, constant-illumination, and fixed-illumination brushes. Our experimental results demonstrate that the proposed approach achieves a better recovery result of intrinsic reflectance and illumination components than the previous approaches. Jianbing Shen, Xiaoshan Yang, Xuelong Li 0001, Yunde Jia |
IEEE Trans. Cybern. | 2 |
| 2011 | Intrinsic images using optimizationabstractIn this paper, we present a novel intrinsic image recovery approach using optimization. Our approach is based on the assumption of in a local window in natural images. Our method adopts a premise that neighboring pixels in a local window of a single image having similar intensity values should have similar reflectance values. Thus the intrinsic image decomposition is formulated by optimizing an energy function with adding a weighting constraint to the local image properties. In order to improve the intrinsic image extraction results, we specify local constrain cues by integrating the user strokes in our energy formulation, including constant-reflectance, constant-illumination and fixed-illumination brushes. Our experimental results demonstrate that our approach achieves a better recovery of intrinsic reflectance and illumination components than by previous approaches. Jianbing Shen, Xiaoshan Yang, Yunde Jia, Xuelong Li 0001 |
CVPR | 2 |