VLDB 2026 Research / reviewers in the wild / expert
Shuhui Wang
dblp:37/2537
· DBLP profile ↗
203ranked-venue papers
23as first author
96since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 137 · 12 first-author · 58 since 2021Artificial intelligence and machine learning · 84 · 10 first-author · 49 since 2021Databases, data management, data science and information retrieval · 16 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Force everything to one: Targeted output redirection against diffusion-based customization
Xingzhi Xu, Lihua Tian, Shuhui Wang |
Neurocomputing | 4 |
| 2026 | D-LIM: A neural network for interpretable gene-gene interactionsabstractRecent advances in gene editing can produce large genotype-fitness maps for targeted genes, yet predicting the effects of mutations between genes remains challenging. Indeed, biochemical models require knowledge of underlying parameters and interactions, whereas machine learning methods typically lack interpretability, as they do not link model parameters to biological quantities. We introduce D-LIM, a neural network that infers low-dimensional fitness landscapes directly from mutation-fitness data. The distinctive feature of D-LIM is that it assumes genes act through independent gene-specific molecular phenotypes whose nonlinear interactions determine fitness. When this assumption holds, the model yields accurate predictions and interpretable effective phenotypes. Conversely, failure reveals that a low-dimensional model is insufficient. Applied to deep mutational scanning of metabolic pathways, protein-protein interactions, and yeast environmental adaptation, D-LIM achieves state-of-the-art predictive accuracy. The inferred phenotype-fitness landscapes reveal whether epistatic interactions can be captured by a low-dimensional continuous model and identify potential trade-offs. Moreover, D-LIM estimates mutational effects on the effective phenotypes, enabling weak extrapolation beyond the training domain. D-LIM demonstrates how simple structure constraints in a neural network can help inference and hypothesis generation in biology. Shuhui Wang, Alexandre Allauzen, Philippe Nghe, Vaitea Opuu |
PLoS Comput. Biol. | 1 |
| 2026 | Graph Continual Learning Network: An Incremental Intelligent Diagnosis Method of Machines for New Fault DetectionabstractStreaming data of machines is continuously collected in practical applications, which produces new fault information with respect to the health change. Therefore, a lifelong-learning intelligent diagnosis model is desired for new fault type recognition based on the streaming data. However, existing research in intelligent fault diagnosis always treats new fault type detection and class incremental learning as two independent problems, which reduces their practicality in industrial applications. To tackle this limitation, a graph continual learning network is constructed for incremental intelligent diagnosis of new faults. The method integrates the advantages of both new fault type detection and class incremental learning. In the method, a graph convolutional network (GCN) based model is formulated for detecting new classes to prejudge whether the DL model needs to be updated. Once any new class is detected, class incremental learning is started automatically to update the DL model without leading to catastrophic forgetting. The proposed method is applied to a pump fault diagnosis case with incremental fault types. Results show that the proposed method offers an effective solution for online intelligent fault diagnosis with satisfactory classification performanceNote to Practitioners—Existing DL-based intelligent diagnosis models often assume the closed-set assumption, i.e., fault types of the monitoring data are the same as those of the training data. In the situations where the assumption is held, DL models receive high-precision recognition results towards instances from the monitoring data stream. However, when a new fault pattern appears in the monitoring data stream, the previous well-trained diagnosis model will inevitably classify the instances into one of the known patterns, resulting in untrustworthy results. This paper proposes a method with the aim of tackling the above issue. The method is able to realize the following functions: Once the instances in the streaming data are detected as a new fault pattern, the class incremental learning will be started automatically. The detected class is used to update the diagnosis model. On the contrary, if the instances are not detected as a new class, they will be sent to the diagnosis model for fault recognition. An incremental diagnosis task is designed under a centrifugal pump application scenario. The results demonstrate the feasibility of the proposed method. The method is anticipated with more application scenarios to verify its superiority. Shuhui Wang, Yaguo Lei, Bin Yang 0014, Xiang Li 0018, Naipeng Li |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Asking Questions to Alleviate Object Hallucination in Large Vision-Language ModelsabstractLarge vision language models (LVLMs) have achieved rapid development. However, just like large language models (LLMs), LVLMs face the critical challenge of hallucination, which refers to the phenomenon that the generation text containing references or descriptions of the input image is incorrect or inconsistent. The causes of hallucinations are complex and therefore difficult to avoid directly during the generation process. To alleviate hallucinations, existing studies mainly employ an instruction-tuning approach that requires model retraining with specific data. Other methods use decoding constraints to penalize specific tokens during the decoding process. These will incur expensive annotation costs and computation burden. In this paper, we propose a framework named AQAH to alleviate hallucinations without relying on manual data and large-scale parameter tuning. AQAH compares multiple generated samples to locate the hallucination factors, and then asks questions about the uncertain information. Finally, the answers to the questions are used to add auxiliary information to the prompt to correct the hallucination of LVLMs during regeneration. To facilitate this process, we constructed an automatic process that involves the training of a small model for question generation, and the agent collaboration framework including the small question generation model and large question answering foundation model. Since AQAH does not directly constrain the decoding, it will not cause a significant degradation in inference efficiency, nor force LVLMs to suffer the notorious problem of shortened text generation length. We experimentally demonstrate the effectiveness of AQAH in hallucination alleviation through the proposed “active questioning & answer verification” paradigm in various multimodal tasks such as captioning and visual question answering. Beyond the promising performance and fewer training/inference time costs against other hallucination reduction methods, our method is highly interpretable and flexible, showing great potential in improving LVLMs by exploiting small-scale models. The code is available at https://github.com/bcxbg/AQAH. Chao Bi, Tiantian Dang, Shuhui Wang, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Divide-and-Conquer: Tree-structured Strategy with Answer Distribution Estimator for Goal-Oriented Visual DialogueabstractGoal-oriented visual dialogue involves multi-round interaction between artificial agents, which has been of remarkable attention due to its wide applications. Given a visual scene, this task occurs when a Questioner asks an action-oriented question and an Answerer responds with the intent of letting the Questioner know the correct action to take. The quality of questions affects the accuracy and efficiency of the target search progress. However, existing methods lack a clear strategy to guide the generation of questions, resulting in the randomness in the search process and inconvergent results. We propose a Tree-Structured Strategy with Answer Distribution Estimator (TSADE) which guides the question generation by excluding half of the current candidate objects in each round. The above process is implemented by maximizing a binary reward inspired by the ``divide-and-conquer'' paradigm. We further design a candidate-minimization reward which encourages the model to narrow down the scope of candidate objects toward the end of the dialogue. We experimentally demonstrate that our method can enable the agents to achieve high task-oriented accuracy with fewer repeating questions and rounds compared to traditional ergodic question generation approaches. Qualitative results further show that TSADE facilitates agents to generate higher-quality questions. Shuo Cai, Xinzhe Han, Shuhui Wang |
AAAI | 3 |
| 2025 | Dis²Booth: Learning Image Distribution with Disentangled Features for Text-to-Image Diffusion ModelsabstractPersonalized image generation enables customized content creation based on the text-to-image diffusion models.However, existing personalization methods focus on fine-tuning generative models to learn to generate specific single individuals or concepts, such as an image of a specific Corgi, but are unable to generate data for multiple individuals or concepts with common characteristics, such as images of multiple different Corgis. In this work, we focus on personalizing a diffusion model to generated varied data usually containing multiple subjects, which has a more diverse and complex data distribution. Our basic assumption is that the varied data distribution is composed of the common features shared among all samples, as well as the reasonable variations within it. Accordingly, we are capable to decompose the learning process of complex data distributions into two simpler sub-tasks, employing a divide-and-conquer approach. To this end we propose Dis2Booth, a framework that can learn complex image Distribution by Disentangling data distribution in an unsupervised manner.Specifically, Dis2Booth contains two modules, Anchor LoRA and Delta LoRA, that are tasked with learning the common features and variational features constrained by Contextual Loss and Delta Loss unsupervisedly. Besides, the Asynchronous Optimization Strategy is proposed to ensure the collaborative training of the two modules. Extensive experiments suggest that Dis2Booth is able to learn the data distribution with higher diversity and complexity while maintaining the same level of flexibility as LoRA. Guanqi Ding, Shuhui Wang, Jinzhe Zhang, Xin Jin 0004, Qingming Huang |
AAAI | 3 |
| 2025 | MSR: A Multifaceted Self-Retrieval Framework for Microscopic Cascade PredictionabstractThe microscopic cascade prediction task has wide applications in downstream areas like ''rumor detection''. Its goal is to forecast the diffusion routines of information cascade within networks. Existing works typically formulate it as a classification task, which fails to well align with the Social Homophily assumption, as it just use the features of ''infected'' users while neglecting those of ''uninfected'' users in representation learning. Moreover, these methods focus primarily on social relationships, thereby dismissing other vital dimensions like users' historical behavior and the underlying preferences behind it. To address these challenges, we introduce the MSR (Multifaceted Self-Retrieval) framework. During encoding, in addition to the existing social graph, we construct a preference graph to represent ''behavioral preferences'' and further propose a modified multi-channel GRAU for multi-view analysis of cascade phenomenon. For decoding, our approach diverges from classification-based methods by reformulating the task as an information retrieval problem that predicts the target user with similarity measures. Empirical evaluations on public datasets demonstrate that this framework significantly outperforms baselines on Hits@κ and MAP@κ, affirming its enhanced ability. Dongsheng Hong, Xujia Li, Shuhui Wang, Wen Lin 0002, Xiangwen Liao |
AAAI | 4 |
| 2025 | Image-to-video Adaptation with Outlier Modeling and Robust Self-learningabstractThe image-to-video adaptation task seeks to effectively harness both labeled images and unlabeled videos for achieving effective video recognition. The modality gap of the image and video modalities and the domain discrepancy across the two domains are the two essential challenges in this task. Existing methods reduce the domain discrepancy via close-set domain adaptation techniques, resulting in inaccurate domain alignment as there exist outlier target frames. To tackle this issue, we extend the vanilla classifier with outlier classes, where each outlier class responsible for capturing outlier frames for a specific class via batch nuclear norm maximization loss. We further propose a new loss by treating the source images apart from class c as instances from outlier class specific for c. As for the modality gap, existing methods usually utilize the pseudo labels obtained from an image-level adapted model to learn a video-level model. Rare efforts are dedicated to handling the noise in pseudo labels. We proposed a new metric based on label propagation consistency to select samples for training a better video-level model. Experiments on 3 benchmarks validating the effectiveness of our method. Junbao Zhuo, Shuhui Wang, Zhenghan Chen, Li Shen 0005, Qingming Huang, Huimin Ma 0001 |
AAAI | 2 |
| 2025 | Video Language Model Pretraining with Spatio-temporal MaskingabstractThe development of self-supervised video-language models based on mask learning has significantly advanced downstream video tasks. These models leverage masked reconstruction to facilitate joint learning of visual and linguistic information. However, recent study reveals that reconstructing image features yields superior downstream performance compared to video feature reconstruction. We hypothesize that this performance gap stems from the way how masking strategies influence the model’s attention to temporal dynamics. To validate this hypothesis, we performed two sets of experiments that demonstrate that alignment between the masked target and the reconstruction target is crucial for self-supervised video-language learning. Based on these findings, we propose a spatio-temporal masking strategy (STM) for video-language model pretraining that operates across adjacent frames, and a decoder leverages semantic information to enhance the spatio-temporal representations of masked tokens. Thanks to the combination of masking strategy and reconstruction decoder, STM enforces the model to learn spatio-temporal feature representation comprehensively. Experiments in three video understanding downstream tasks validate the superiority of our method. Codes are available here. Zhaobo Qi, Junshu Sun, Yaowei Wang 0001, Qingming Huang, Shuhui Wang |
CVPR | 6 |
| 2025 | Enhancing Pre-trained Representation Classifiability can Boost its InterpretabilityabstractThe visual representation of a pre-trained model prioritizes the classifiability on downstream tasks, while the widespread applications for pre-trained visual models have posed new requirements for representation interpretability. However, it remains unclear whether the pre-trained representations can achieve high interpretability and classifiability simultaneously. To answer this question, we quantify the representation interpretability by leveraging its correlation with the ratio of interpretable semantics within the representations. Given the pre-trained representations, only the interpretable semantics can be captured by interpretations, whereas the uninterpretable part leads to information loss. Based on this fact, we propose the Inherent Interpretability Score (IIS) that evaluates the information loss, measures the ratio of interpretable semantics, and quantifies the representation interpretability. In the evaluation of the representation interpretability with different classifiability, we surprisingly discover that the interpretability and classifiability are positively correlated, i.e., representations with higher classifiability provide more interpretable semantics that can be captured in the interpretations. This observation further supports two benefits to the pre-trained representations. First, the classifiability of representations can be further improved by fine-tuning with interpretability maximization. Second, with the classifiability improvement for the representations, we obtain predictions based on their interpretations with less accuracy degradation. The discovered positive correlation and corresponding applications show that practitioners can unify the improvements in interpretability and classifiability for pre-trained vision models. Codes are available at https://github.com/ssfgunner/IIS. Shufan Shen, Zhaobo Qi, Junshu Sun, Qingming Huang, Qi Tian 0001, Shuhui Wang |
ICLR | 6 |
| 2025 | Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video RetrievalabstractWith the explosive growth of video data, finding videos that meet detailed requirements in large datasets has become a challenge. To address this, the composed video retrieval task has been introduced, enabling users to retrieve videos using complex queries that involve both visual and textual information. However, the inherent heterogeneity between the modalities poses significant challenges. Textual data are highly abstract, while video content contains substantial redundancy. The modality gap in information representation makes existing methods struggle with the modality fusion and alignment required for fine-grained composed retrieval. To overcome these challenges, we first introduce FineCVR-1M, a fine-grained composed video retrieval dataset containing 1,010,071 video-text triplets with detailed textual descriptions. This dataset is constructed through an automated process that identifies key concept changes between video pairs to generate textual descriptions for both static and action concepts. For fine-grained retrieval methods, the key challenge lies in understanding the detailed requirements. Text description serves as clear expressions of intent, but it requires models to distinguish subtle differences in the description of video semantics. Therefore, we propose a textual Feature Disentanglement and Cross-modal Alignment framework (FDCA) that disentangles features at both the sentence and token levels. At the sequence level, we separate text features into retained and injected features. At the token level, an Auxiliary Token Disentangling mechanism is proposed to disentangle texts into retained, injected, and excluded tokens. The disentanglement at both levels extracts fine-grained features, which are aligned and fused with the reference video to extract global representations for video retrieval. Experiments on FineCVR-1M dataset demonstrate the superior performance of FDCA. Our code and dataset are available at: https://may2333.github.io/FineCVR/. Zhaobo Qi, Yiling Wu, Junshu Sun, Yaowei Wang 0001, Shuhui Wang |
ICLR | 6 |
| 2025 | Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional VideosabstractIn this paper, we address the challenge of procedure planning in instructional videos, aiming to generate coherent and task-aligned action sequences from start and end visual observations. Previous work has mainly relied on text-level supervision to bridge the gap between observed states and unobserved actions, but it struggles with capturing intricate temporal relationships among actions. Building on these efforts, we propose the Masked Temporal Interpolation Diffusion (MTID) model that introduces a latent space temporal interpolation module within the diffusion model. This module leverages a learnable interpolation matrix to generate intermediate latent features, thereby augmenting visual supervision with richer mid-state details. By integrating this enriched supervision into the model, we enable end-to-end training tailored to task-specific requirements, significantly enhancing the model's capacity to predict temporally coherent action sequences. Additionally, we introduce an action-aware mask projection mechanism to restrict the action generation space, combined with a task-adaptive masked proximity loss to prioritize more accurate reasoning results close to the given start and end states over those in intermediate steps. Simultaneously, it filters out task-irrelevant action predictions, leading to contextually aware action sequences. Experimental results across three widely used benchmark datasets demonstrate that our MTID achieves promising action planning performance on most metrics. Zhaobo Qi, Lingshuai Lin, Junqi Jing, Tingting Chai, Beichen Zhang 0006, Shuhui Wang, Weigang Zhang |
ICLR | 7 |
| 2025 | Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMsabstractLifelong knowledge editing enables continuous, precise updates to outdated knowledge in large language models (LLMs) without computationally expensive full retraining. However, existing methods often accumulate errors throughout the editing process, causing a gradual decline in both editing accuracy and generalization. To tackle this problem, we propose Neuron-Specific Masked Knowledge Editing (NMKE), a novel fine-grained editing framework that combines neuron-level attribution with dynamic sparse masking.
Leveraging neuron functional attribution, we identify two key types of knowledge neurons, with knowledge-general neurons activating consistently across prompts and knowledge-specific neurons activating to specific prompts.
NMKE further introduces an entropy-guided dynamic sparse mask, locating relevant neurons to the target knowledge. This strategy enables precise neuron-level knowledge editing with fewer parameter modifications.
Experimental results from thousands of sequential edits demonstrate that NMKE outperforms existing methods in maintaining high editing success rates and preserving model general capabilities in lifelong editing. Jinzhe Liu, Junshu Sun, Shufan Shen, Chenxue Yang, Shuhui Wang |
NeurIPS | 5 |
| 2025 | VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept SetabstractThe alignment of vision-language representations endows current Vision-Language Models (VLMs) with strong multi-modal reasoning capabilities. However, the interpretability of the alignment component remains uninvestigated due to the difficulty in mapping the semantics of multi-modal representations into a unified concept set. To address this problem, we propose VL-SAE, a sparse autoencoder that encodes vision-language representations into its hidden activations. Each neuron in the hidden layer correlates to a concept represented by semantically similar images and texts, thereby interpreting these representations with a unified concept set. To establish the neuron-concept correlation, we encourage semantically similar representations to exhibit consistent neuron activations during self-supervised training. First, to measure the semantic similarity of multi-modal representations, we perform their alignment in an explicit form based on cosine similarity. Second, we construct the VL-SAE with a distance-based encoder and two modality-specific decoders to ensure the activation consistency of semantically similar representations. Experiments across multiple VLMs (e.g., CLIP, LLaVA) demonstrate the superior capability of VL-SAE in interpreting and enhancing the vision-language alignment. For interpretation, the alignment between vision and language representations can be understood by comparing their semantics with concepts. For enhancement, the alignment can be strengthened by aligning vision-language representations at the concept level, contributing to performance improvements in downstream tasks, including zero-shot image classification and hallucination elimination. Codes are provided in the supplementary and will be released to GitHub. Shufan Shen, Junshu Sun, Qingming Huang, Shuhui Wang |
NeurIPS | 4 |
| 2025 | Relieving the Over-Aggregating Effect in Graph TransformersabstractGraph attention has demonstrated superior performance in graph learning tasks. However, learning from global interactions can be challenging due to the large number of nodes. In this paper, we discover a new phenomenon termed over-aggregating. Over-aggregating arises when a large volume of messages is aggregated into a single node with less discrimination, leading to the dilution of the key messages and potential information loss. To address this, we propose Wideformer, a plug-and-play method for graph attention. Wideformer divides the aggregation of all nodes into parallel processes and guides the model to focus on specific subsets of these processes. The division can limit the input volume per aggregation, avoiding message dilution and reducing information loss. The guiding step sorts and weights the aggregation outputs, prioritizing the informative messages. Evaluations show that Wideformer can effectively mitigate over-aggregating. As a result, the backbone methods can focus on the informative messages, achieving superior performance compared to baseline methods. Junshu Sun, Wanxing Chang, Chenxue Yang, Qingming Huang, Shuhui Wang |
NeurIPS | 5 |
| 2025 | PIC: Domain generalization by path information constraint
Jilong Zhu, Junbao Zhuo, Shuhui Wang |
Pattern Recognit. | 3 |
| 2025 | Adaptive Fault-Tolerant Optimized Platoon Cloud Tracking Control for Heterogeneous Vehicles via Dual Learning MechanismabstractThis work investigates an adaptive fault-tolerant cloud tracking control scheme for heterogeneous vehicular platoon systems (HVPSs) based on a dual learning mechanism (DLM). A human-vehicle-road-cloud traffic scenario is built, in which the cloud control platform operates HVPSs by sending the command signal to a non-autonomous leader vehicle, which generates the required trajectory. Moreover, a DLM consisting of an improved reinforcement learning mechanism and a composite learning mechanism is proposed to optimize the control performance of the entire vehicular platoon. Furthermore, unlike existing fault-tolerant control methods, the proposed DLM can effectively reduce the adverse effects caused by actuator fault. Via the DLM, and cloud control platform, an adaptive fault-tolerant optimized platoon technology is developed to guarantee the bistability. Finally, simulation results certify the effectiveness and superiority of the developed strategy.Note to Practitioners—This paper is inspired by the problem of vehicular platoon control, but the proposed approach is not only applicable to transportation systems, but also to a variety of scenarios such as smart grids, supply chain management, and fire rescue. The intelligent transportation system aims to realize V2P, V2V, and V2I. However, there still exist several problems, such as the lack of obstacle avoidance/collision mechanism, the occurrence of platoon actuator fault, and the excessive consumption of vehicle energy, making it difficult to apply in complex traffic. In this paper, the proposed method can effectively solve the above problems, and simulation experiments have shown that this method is feasible and superior. However, vehicular platoon path programming and external network attacks have not yet been considered. To approach the real traffic environment more closely also prompts us to conduct research in the future. Jiaxin An, Yingxun Wang, Ahmer Khan Jadoon, Shuhui Wang |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Stable Attribute Group Editing for Reliable Few-Shot Image GenerationabstractFew-shot image generation aims to generate data of an unseen category based on only a few samples. Apart from basic content generation, a bunch of downstream applications hopefully benefit from this task, such as low-data detection and few-shot classification. To achieve this goal, the generated images should guarantee category retention for classification beyond the visual quality and diversity. In our preliminary work, we present an “editing-based” framework, Attribute Group Editing (AGE), for reliable few-shot image generation, which largely improves the performance compared with existing methods that require re-training a GAN with limited data. Nevertheless, AGE’s performance on downstream classification is not as satisfactory as expected. Furthermore, existing generative models suffer from similar issues. This paper focuses on addressing the issue of universal class inconsistency in all generative models. It not only improves AGE to enhance its ability to preserve class information but also conducts a comprehensive analysis of the causes of this problem in generative models from multiple perspectives, proposing potential directions for resolution. We first propose Stable Attribute Group Editing (SAGE) for more stable class-relevant image generation. SAGE corrects the inaccurate assumptions in AGE and leverages the distribution information from seen categories to accurately estimate the data distribution of unseen categories, thereby eliminating the class inconsistency issue in the generated data. We apply SAGE to both GANs and diffusion models to verify its flexibility and further achieve promising generation performance. Going one step further, we find that even though the generated images look photo-realistic and require no category-relevant editing, they are usually of limited help for downstream classification. We systematically discuss this issue from both the generation and classification perspectives, and propose to boost the downstream classification performance of SAGE by enhancing the pixel and frequency components. Extensive experiments provide valuable insights into extending image generation to wider downstream applications. Codes are available at https://github.com/UniBester/SAGE. Guanqi Ding, Xinzhe Han, Shuhui Wang, Xin Jin 0004, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Uncertainty-Aware Mixture of Experts for Video Action AnticipationabstractAnticipating future actions in daily life videos is crucial for seamless human-machine collaboration. However, accurately predicting these actions is challenging due to the inherent uncertainty and non-determinism of future events. To address this, we propose the uncertainty-aware mixture-of-experts framework for action anticipation (AntMoE), which employs multiple anticipation experts to model diverse video evolution patterns through learnable expert embeddings. These anticipation experts generate diverse predictions by integrating the top-k semantically similar observed video frames related to the current predicted feature representation, along with their corresponding expert embeddings. An anticipation router then aggregates these predictions based on the relationship between the current feature representation and all expert embeddings. To enhance the effectiveness of AntMoE, we introduce an expert regularization loss with three components: orthogonal loss promotes orthogonality among expert embeddings; expert balance loss ensures equal activation of all experts during training; and stability loss encourages the generation of numerically stable aggregation weights. Additionally, we incorporate an anticipation ranking loss function that aligns the model’s confidence across varying anticipation time durations with the ground-truth ranking order, where a shorter anticipation time length corresponds to a higher confidence level. Experimental results across multiple benchmarks demonstrate that our method achieves remarkable anticipation performance. Zhaobo Qi, Shuhui Wang, Weigang Zhang, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Prompting Video-Language Foundation Models With Domain-Specific Fine-Grained Heuristics for Video Question AnsweringabstractVideo Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and sophisticated cross-modal interaction for accurate inference. Despite advancements in multi-modal pre-trained models and video-language foundation models, these systems often struggle with domain-specific VideoQA due to their generalized pre-training objectives. Addressing this gap necessitates bridging the divide between broad cross-modal knowledge and the specific inference demands of VideoQA tasks. To this end, we introduce HeurVidQA, a framework that leverages domain-specific entity-action heuristics to refine pre-trained video-language foundation models. Our approach treats these models as implicit knowledge engines, employing domain-specific entity-action prompters to direct the model’s focus toward precise cues that enhance reasoning. By delivering fine-grained heuristics, we improve the model’s ability to identify and interpret key entities and actions, thereby enhancing its reasoning capabilities. Extensive evaluations across multiple VideoQA datasets demonstrate that our method significantly outperforms existing models, underscoring the importance of integrating domain-specific knowledge into video-language models for more accurate and context-aware VideoQA. Ting Yu 0002, Kunhao Fu, Shuhui Wang, Qingming Huang, Jun Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Enhancing the Robustness of Vision-Language Foundation Models by Alignment PerturbationabstractWhile Vision-Language Models (VLMs) based on large-scale models have shown revolutionary advancements across various vision-language tasks, research on improving VLM robustness remains underexplored. Existing studies primarily focus on attacking VLM after the pretrained visual or textual encoders, typically requiring obvious noise or long inference time. In this study, we look into VLM structure and highlight alignment module’s role as a protective filter that enhances VLM robustness against various perturbations. Motivated by these insights, we investigate VLM from both user and model developer perspectives and introduce the alignment perturbation strategy, which consists of multimodal, visual, and textual perturbations. Multimodal perturbation aims to achieve targeted textual output generation and is further utilized to enhance VLM robustness. Minimal perturbations to visual or textual inputs can lead to significant changes in the overall output of VLMs, revealing their sensitivity to both visual and textual input variations. Building on the alignment perturbation strategy, we propose alignment robust training, which efficiently improves VLM robustness by finetuning the parameters of alignment module without excessive resource consumption. Experiment results across various tasks and models demonstrate the effectiveness of the proposed alignment perturbation and alignment robust training. These methods deepen the understanding of VLM robustness, allowing for secure and reliable deployment towards diverse real-world scenarios. Codes are available at https://github.com/zhangconghhh/RobustVLMs. Shuhui Wang, Yao Zhu 0003, Honggang Qi, Qingming Huang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Consistency Conditioned Memory Augmented Dynamic Diagnosis Model for Medical Visual Question AnsweringabstractMedical Visual Question Answering (Med-VQA) holds immense promise as an invaluable medical assistance aid, offering timely diagnostic outcomes based on medical images and accompanying questions, thereby supporting medical professionals in making accurate clinical decisions. However, Med-VQA is still in its infancy, with existing solutions falling short in imitating human diagnostic processes and ensuring result consistency. To address these challenges, we propose a Consistency Conditioned Memory augmented Dynamic diagnosis model (CoCoMeD), incorporating two core components: a dynamic memory diagnosis engine and a consistency-conditioned enforcer. The dynamic memory diagnosis engine enables intricate diagnostic interactions by retaining vital visual cues from medical images and iteratively updating pertinent memories. This dynamic reasoning capability mirrors the cognitive processes observed in skilled medical diagnosticians, thus effectively enhancing the model's ability to reason over diverse medical visual facts and patient-specific questions. Moreover, to strengthen diagnostic coherence, the consistency-conditioned enforcer imposes coherence constraints linking interrelated questions with identical medical facts, ensuring the credibility and reliability of its diagnostic outcomes. Additionally, we present C-SLAKE, an extended Med-VQA dataset encompassing diverse medical image types, and categorized diagnostic question-answer pairs for consistent Med-VQA evaluation on rich medical sources. Comprehensive experiments on DME and C-SLAKE showcase CoCoMeD's superior performance and potential to advance trustworthy multi-source medical question answering. Ting Yu 0002, Binhui Ge, Shuhui Wang, Qingming Huang, Jun Yu 0002 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Boost Tracking by Natural Language With Prompt-Guided GroundingabstractTNL (Tracking by Natural Language) aims to locate the target described by a natural language sentence in a video. Most existing TNL methods are typically composed of three modules: object grounding, object tracking, and switching module, and their performance is limited by the poor performance of the grounding and switching modules due to the complex backgrounds and inaccurate information stored in the memory. This paper presents a global-local framework to address these issues, which includes a prompt-guided grounding module, a trained local tracking module, and a memory-based switcher module. The prompt-guided grounding module uses noun prompts to guide the CLIP model in focusing more on target regions and aligning visual features semantically with linguistic features, avoiding being misled by distractors and background. The memory-based switch module stores historical information with higher-quality memory, allowing the model to make more accurate decisions based on reliable data, thus improving the overall performance. Experiments on TNL2K, LaSOT, and OTB-Lang demonstrate the effectiveness and generalizability of the proposed framework. Hengyou Li, Xinyan Liu 0008, Guorong Li, Shuhui Wang, Laiyun Qing, Qingming Huang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Region Uncertainty Estimation for Medical Image Segmentation With Noisy LabelsabstractThe success of deep learning in 3D medical image segmentation hinges on training with a large dataset of fully annotated 3D volumes, which are difficult and time-consuming to acquire. Although recent foundation models (e.g., segment anything model, SAM) can utilize sparse annotations to reduce annotation costs, segmentation tasks involving organs and tissues with blurred boundaries remain challenging. To address this issue, we propose a region uncertainty estimation framework for Computed Tomography (CT) image segmentation using noisy labels. Specifically, we propose a sample-stratified training strategy that stratifies samples according to their varying quality labels, prioritizing confident and fine-grained information at each training stage. This sample-to-voxel level processing enables more reliable supervision information to propagate to noisy label data, thus effectively mitigating the impact of noisy annotations. Moreover, we further design a boundary-guided regional uncertainty estimation module that adapts sample hierarchical training to assist in evaluating sample confidence. Experiments conducted across multiple CT datasets demonstrate the superiority of our proposed method over several competitive approaches under various noise conditions. Our proposed reliable label propagation strategy not only significantly reduces the cost of medical image annotation and robust model training but also improves the segmentation performance in scenarios with imperfect annotations, thus paving the way towards the application of medical segmentation foundation models under low-resource and remote scenarios. Code will be available at https://github.com/KHan-UJS/NoisyLabel. Kai Han 0006, Shuhui Wang, Jun Chen 0030, Chengxuan Qian, Chongwen Lyu, Siqi Ma 0004, Cheng-Jian Qiu, Victor S. Sheng, Qingming Huang, Zhe Liu 0004 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Inferential and Commonsense Visual Question GenerationabstractThe Visual Question Generation (VQG) task generally aims to produce questions based on images in natural language. Existing studies often handle VQG as a reverse Visual Question Answering (VQA), training data-driven generators on VQA datasets. However, this solution pipeline struggles to generate high-quality questions that effectively challenge robots and humans, even by leveraging the most advanced large-scale foundational models. There are also some other VQG methods depending on elaborate and costly manual preprocessing heavily. To address these limitations, we propose a novel method with a two-module framework for automatically generating inferential visual questions that also follow commonsense. The “Scene Graph Generation” module constructs specialized scene graphs by progressively expanding connections from high-confidence nodes. This module ensures semantic consistency by aligning visual, textual, and salient features. Additionally, we incorporate external knowledge to extend abstract semantic concepts and associated facts, enriching the content of generated questions and facilitating the generated question to better follow the commonsense of human. Another module “Question Generation” utilizes the above scene graph as a foundation to search and instantiate for the question. The generated questions will match with the program templates and have diverse inferential paths. Experimental results demonstrate that our method is both effective and highly scalable. The generated questions are controllable in terms of semantic richness and difficulty, exhibiting clear inferential and commonsense properties. Furthermore, we automatically utilize our method to create a large-scale dataset, ICVQA, which includes approximately 160,000 images and 800,000 questionanswer pairs, thereby facilitating further research in VQA and visual dialogue. Chao Bi, Shuhui Wang, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2025 | Boosting Dataset Distillation With the Assistance of Crucial Samples for Visual Learning
Yao Zhu 0003, Yuefeng Chen, Cen Chen 0001, Jianmei Guo, Shuhui Wang |
IEEE Trans. Multim. | 6 |
| 2025 | Introduction to the Special Issue on Deep Learning for Robust Human Body Language UnderstandingabstractThis editorial introduces the Special Issue on Deep Learning for Robust Human Body Language Understanding, hosted by the ACM Transactions on Multimedia Computing, Communications, and Applications in 2024. Human body language understanding has emerged as a critical research area, addressing challenges in analyzing, recognizing, and synthesizing multimodal human behavioral data such as gestures, poses, facial expressions. This Special Issue highlights recent advancements in deep learning techniques that enhance the robustness, scalability, and applicability of human body language understanding in diverse scenarios, including healthcare, education, and barrier-free human-computer interaction systems. The issue features a total of eight research articles focusing on key aspects of human body language understanding. These contributions are categorized into major research areas: gesture and sign language understanding, pose and action recognition, and facial expression and emotion analysis. Each article provides novel insights into challenges such as data scarcity, multimodal integration, adversarial robustness, and cross-modal generative modeling. We summarize the main contributions of the included works and emphasize their role in advancing the field of human body language understanding. Finally, we discuss ongoing challenges and future opportunities in this rapidly evolving domain, particularly in the context of integrating human-centric AI systems into real-world applications. Dan Guo 0001, Troy McDaniel, Shuhui Wang, Meng Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in VideoabstractTemporal Sentence Grounding in Video (TSGV) is troubled by dataset bias issue, which is caused by the uneven temporal distribution of the target moments for samples with similar semantic components in input videos or query texts. Existing methods resort to utilizing prior knowledge about bias to artificially break this uneven distribution, which only removes a limited amount of significant language biases. In this work, we propose the bias-conflict sample synthesis and adversarial removal debias strategy (BSSARD), which dynamically generates bias-conflict samples by explicitly leveraging potentially spurious correlations between single-modality features and the temporal position of the target moments. Through adversarial training, its bias generators continuously introduce biases and generate bias-conflict samples to deceive its grounding model. Meanwhile, the grounding model continuously eliminates the introduced biases, which requires it to model multi-modality alignment information. BSSARD will cover most kinds of coupling relationships and disrupt language and visual biases simultaneously. Extensive experiments on Charades-CD and ActivityNet-CD demonstrate the promising debiasing capability of BSSARD. Source codes are available at https://github.com/qzhb/BSSARD. Zhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang, Weigang Zhang, Qingming Huang |
AAAI | 4 |
| 2024 | Confusing Pair Correction Based on Category Prototype for Domain Adaptation under Noisy EnvironmentsabstractIn this paper, we address unsupervised domain adaptation under noisy environments, which is more challenging and practical than traditional domain adaptation. In this scenario, the model is prone to overfitting noisy labels, resulting in a more pronounced domain shift and a notable decline in the overall model performance. Previous methods employed prototype methods for domain adaptation on robust feature spaces. However, these approaches struggle to effectively classify classes with similar features under noisy environments. To address this issue, we propose a new method to detect and correct confusing class pair. We first divide classes into easy and hard classes based on the small loss criterion. We then leverage the top-2 predictions for each sample after aligning the source and target domain to find the confusing pair in the hard classes. We apply label correction to the noisy samples within the confusing pair. With the proposed label correction method, we can train our model with more accurate labels. Extensive experiments confirm the effectiveness of our method and demonstrate its favorable performance compared with existing state-of-the-art methods. Our codes are publicly available at https://github.com/Hehxcf/CPC/. Churan Zhi, Junbao Zhuo, Shuhui Wang |
AAAI | 3 |
| 2024 | CTSM: Combining Trait and State Emotions for Empathetic Response ModelabstractEmpathetic response generation endeavors to empower dialogue systems to perceive speakers’ emotions and generate empathetic responses accordingly. Psychological research demonstrates that emotion, as an essential factor in empathy, encompasses trait emotions, which are static and context-independent, and state emotions, which are dynamic and context-dependent. However, previous studies treat them in isolation, leading to insufficient emotional perception of the context, and subsequently, less effective empathetic expression. To address this problem, we propose Combining Trait and State emotions for Empathetic Response Model (CTSM). Specifically, to sufficiently perceive emotions in dialogue, we first construct and encode trait and state emotion embeddings, and then we further enhance emotional perception capability through an emotion guidance module that guides emotion representation. In addition, we propose a cross-contrastive learning decoder to enhance the model’s empathetic expression capability by aligning trait and state emotions between generated responses and contexts. Both automatic and manual evaluation results demonstrate that CTSM outperforms state-of-the-art baselines and can generate more empathetic responses. Our code is available at https://github.com/wangyufeng-empty/CTSM Zhou Yang 0012, Shuhui Wang, Xiangwen Liao |
LREC/COLING | 4 |
| 2024 | Learning Invariant Representation with Consistency and Diversity for Semi-Supervised Source Hypothesis TransferabstractSemi-supervised Domain adaptation (SSDA) has shown promising results by leveraging unlabeled data and limited labeled samples in the target domain. However, accessibility to source data is hindered by data privacy concerns, giving rise to Semi-supervised Source Hypothesis Transfer (SSHT). Integrating the SSDA methods directly into SSHT tasks is straightforward but poses two significant challenges: i) The hypothesis (classifier) is no longer supervised by source labels, and relying on only a few labels may result in hypothesis collapse; ii) Trained source models often exhibit bias, making them susceptible to misclassifying samples from minority categories into majority ones. We examined the recent methods in the SSHT setting and observed variations in performance compared to SSDA. To address these challenges, we first mitigate model overfitting to target labeled data by promoting prediction consistency between two types of randomly augmented unlabeled data, thereby preventing training collapse. Additionally, we maintain both the prediction diversity and discriminability by leveraging unlabeled data. Experiments on SSHT tasks show that our method yields more stable and competitive results compared with state-of-the-art methods. Junbao Zhuo, Shuhao Cui, Shuhui Wang, Yuejian Fang |
ICASSP | 4 |
| 2024 | R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image GenerationabstractRecent text-to-image (T2I) diffusion models have achieved remarkable progress in generating high-quality images given text-prompts as input. However, these models fail to convey appropriate spatial composition specified by a layout instruction. In this work, we probe into zero-shot grounded T2I generation with diffusion models, that is, generating images corresponding to the input layout information without training auxiliary modules or finetuning diffusion models. We propose a **R**egion and **B**oundary (R&B) aware cross-attention guidance approach that gradually modulates the attention maps of diffusion model during generative process, and assists the model to synthesize images (1) with high fidelity, (2) highly compatible with textual input, and (3) interpreting layout instructions accurately. Specifically, we leverage the discrete sampling to bridge the gap between consecutive attention maps and discrete layout constraints, and design a region-aware loss to refine the generative layout during diffusion process. We further propose a boundary-aware loss to strengthen object discriminability within the corresponding regions. Experimental results show that our method outperforms existing state-of-the-art zero-shot grounded T2I generation methods by a large margin both qualitatively and quantitatively on several benchmarks.
Project page: https://sagileo.github.io/Region-and-Boundary. Jiayu Xiao, Henglei Lv, Liang Li 0003, Shuhui Wang, Qingming Huang |
ICLR | 4 |
| 2024 | Modeling Language Tokens as Functionals of Semantic FieldsabstractRecent advances in natural language processing have relied heavily on using Transformer-based language models. However, Transformers often require large parameter sizes and model depth. Existing Transformer-free approaches using state-space models demonstrate superiority over Transformers, yet they still lack a neuro-biologically connection to the human brain. This paper proposes ${\it LasF}$, representing ${\bf L}$anguage tokens ${\bf as}$ ${\bf F}$unctionals of semantic fields, to simulate the neuronal behaviors for better language modeling. The ${\it LasF}$ module is equivalent to a nonlinear approximator tailored for sequential data. By replacing the final layers of pre-trained language models with the ${\it LasF}$ module, we obtain ${\it LasF}$-based models. Experiments conducted for standard reading comprehension and question-answering tasks demonstrate that the ${\it LasF}$-based models consistently improve accuracy with fewer parameters. Besides, we use CommonsenseQA's blind test set to evaluate a full-parameter tuned ${\it LasF}$-based model, which outperforms the prior best ensemble and single models by $0.4\%$ and $3.1\%$, respectively. Furthermore, our ${\it LasF}$-only language model trained from scratch outperforms existing parameter-efficient language models on standard datasets such as WikiText103 and PennTreebank. Zhengqi Pei, Shuhui Wang, Qingming Huang |
ICML | 3 |
| 2024 | Data-free Neural Representation Compression with Riemannian Neural DynamicsabstractNeural models are equivalent to dynamic systems from a physics-inspired view, implying that computation on neural networks can be interpreted as the dynamical interactions between neurons. However, existing work models neuronal interaction as a weight-based linear transformation, and the nonlinearity comes from the nonlinear activation functions, which leads to limited nonlinearity and data-fitting ability of the whole neural model. Inspired by Riemannian geometry, we interpret neural structures by projecting neurons onto the Riemannian neuronal state space and model neuronal interaction with Riemannian metric (${\it RieM}$), which provides a more efficient neural representation with higher parameter efficiency. With ${\it RieM}$, we further design a novel data-free neural compression mechanism that does not require additional fine-tuning with real data. Using backbones like ResNet and Vision Transformer, we conduct extensive experiments on datasets such as MNIST, CIFAR-100, ImageNet-1k, and COCO object detection. Empirical results show that, under equal compression rates and computational complexity, models compressed with ${\it RieM}$ achieve superior inference accuracy compared to existing data-free compression methods. Zhengqi Pei, Shuhui Wang, Xiangyang Ji, Qingming Huang |
ICML | 3 |
| 2024 | Unsupervised Image-to-Video Adaptation via Category-aware Flow Memory Bank and Realistic Video Generation
Kenan Huang, Junbao Zhuo, Shuhui Wang, Chi Su, Qingming Huang, Huimin Ma 0001 |
ACM Multimedia | 3 |
| 2024 | Expanding Sparse Tuning for Low Memory UsageabstractParameter-efficient fine-tuning (PEFT) is an effective method for adapting pre-trained vision models to downstream tasks by tuning a small subset of parameters. Among PEFT methods, sparse tuning achieves superior performance by only adjusting the weights most relevant to downstream tasks, rather than densely tuning the whole weight matrix. However, this performance improvement has been accompanied by increases in memory usage, which stems from two factors, i.e., the storage of the whole weight matrix as learnable parameters in the optimizer and the additional storage of tunable weight indexes. In this paper, we propose a method named SNELL (Sparse tuning with kerNELized LoRA) for sparse tuning with low memory usage. To achieve low memory usage, SNELL decomposes the tunable matrix for sparsification into two learnable low-rank matrices, saving from the costly storage of the whole original matrix. A competition-based sparsification mechanism is further proposed to avoid the storage of tunable weight indexes. To maintain the effectiveness of sparse tuning with low-rank matrices, we extend the low-rank decomposition by applying nonlinear kernel functions to the whole-matrix merging. Consequently, we gain an increase in the rank of the merged matrix, enhancing the ability of SNELL in adapting the pre-trained models to downstream tasks. Extensive experiments on multiple downstream tasks show that SNELL achieves state-of-the-art performance with low memory usage, endowing PEFT with sparse tuning to large-scale models. Codes are available at https://github.com/ssfgunner/SNELL. Shufan Shen, Junshu Sun, Xiangyang Ji, Qingming Huang, Shuhui Wang |
NeurIPS | 5 |
| 2024 | Towards Dynamic Message Passing on GraphsabstractMessage passing plays a vital role in graph neural networks (GNNs) for effective feature learning. However, the over-reliance on input topology diminishes the efficacy of message passing and restricts the ability of GNNs. Despite efforts to mitigate the reliance, existing study encounters message-passing bottlenecks or high computational expense problems, which invokes the demands for flexible message passing with low complexity. In this paper, we propose a novel dynamic message-passing mechanism for GNNs. It projects graph nodes and learnable pseudo nodes into a common space with measurable spatial relations between them. With nodes moving in the space, their evolving relations facilitate flexible pathway construction for a dynamic message-passing process. Associating pseudo nodes to input graphs with their measured relations, graph nodes can communicate with each other intermediately through pseudo nodes under linear complexity. We further develop a GNN model named $\mathtt{N^2}$ based on our dynamic message-passing mechanism. $\mathtt{N^2}$ employs a single recurrent layer to recursively generate the displacements of nodes and construct optimal dynamic pathways. Evaluation on eighteen benchmarks demonstrates the superior performance of $\mathtt{N^2}$ over popular GNNs. $\mathtt{N^2}$ successfully scales to large-scale benchmarks and requires significantly fewer parameters for graph classification with the shared recurrent layer. Junshu Sun, Chenxue Yang, Xiangyang Ji, Qingming Huang, Shuhui Wang |
NeurIPS | 5 |
| 2024 | A noise generative network to reduce the gap between simulation and measurement signals in mechanical fault diagnosis
Hui Wang 0140, Shuhui Wang, Ronggang Yang, Jiawei Xiang |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Learning Hierarchical Modular Networks for Video CaptioningabstractVideo captioning aims to generate natural language descriptions for a given video clip. Existing methods mainly focus on end-to-end representation learning via word-by-word comparison between predicted captions and ground-truth texts. Although significant progress has been made, such supervised approaches neglect semantic alignment between visual and linguistic entities, which may negatively affect the generated captions. In this work, we propose a hierarchical modular network to bridge video representations and linguistic semantics at four granularities before generating captions: entity, verb, predicate, and sentence. Each level is implemented by one module to embed corresponding semantics into video representations. Additionally, we present a reinforcement learning module based on the scene graph of captions to better measure sentence similarity. Extensive experimental results show that the proposed method performs favorably against the state-of-the-art models on three widely-used benchmark datasets, including microsoft research video description corpus (MSVD), MSR-video to text (MSR-VTT), and video-and-TEXt (VATEX). Guorong Li, Hanhua Ye, Yuankai Qi, Shuhui Wang, Laiyun Qing, Qingming Huang, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Uncertainty-Boosted Robust Video Activity AnticipationabstractVideo activity anticipation aims to predict what will happen in the future, embracing a broad application prospect ranging from robot vision and autonomous driving. Despite the recent progress, the data uncertainty issue, reflected as the content evolution process and dynamic correlation in event labels, has been somehow ignored. This reduces the model generalization ability and deep understanding on video content, leading to serious error accumulation and degraded performance. In this paper, we address the uncertainty learning problem and propose an uncertainty-boosted robust video activity anticipation framework, which generates uncertainty values to indicate the credibility of the anticipation results. The uncertainty value is used to derive a temperature parameter in the softmax function to modulate the predicted target activity distribution. To guarantee the distribution adjustment, we construct a reasonable target activity label representation by incorporating the activity evolution from the temporal class correlation and the semantic relationship. Moreover, we quantify the uncertainty into relative values by comparing the uncertainty among sample pairs and their temporal-lengths. This relative strategy provides a more accessible way in uncertainty modeling than quantifying the absolute uncertainty values on the whole dataset. Experiments on multiple backbones and benchmarks show our framework achieves promising performance and better robustness/interpretability. Zhaobo Qi, Shuhui Wang, Weigang Zhang, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Inductive State-Relabeling Adversarial Active Learning With Heuristic Clique RescalingabstractActive learning (AL) is to design label-efficient algorithms by labeling the most representative samples. It reduces annotation cost and attracts increasing attention from the community. However, previous AL methods suffer from the inadequacy of annotations and unreliable uncertainty estimation. Moreover, we find that they ignore the intra-diversity of selected samples, which leads to sampling redundancy. In view of these challenges, we propose an inductive state-relabeling adversarial AL model (ISRA) that consists of a unified representation generator, an inductive state-relabeling discriminator, and a heuristic clique rescaling module. The generator introduces contrastive learning to leverage unlabeled samples for self-supervised training, where the mutual information is utilized to improve the representation quality for AL selection. Then, we design an inductive uncertainty indicator to learn the state score from labeled data and relabel unlabeled data with different importance for better discrimination of instructive samples. To solve the problem of sampling redundancy, the heuristic clique rescaling module measures the intra-diversity of candidate samples and recurrently rescales them to select the most informative samples. The experiments conducted on eight datasets and two imbalanced scenarios show that our model outperforms the previous state-of-the-art AL methods. As an extension on the cross-modal AL task, we apply ISRA to the image captioning and it also achieves superior performance. Beichen Zhang 0006, Liang Li 0003, Shuhui Wang, Shaofei Cai, Zhengjun Zha, Qi Tian 0001, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Rethink video retrieval representation for video captioning
Mingkai Tian, Guorong Li, Yuankai Qi, Shuhui Wang, Quan Z. Sheng, Qingming Huang |
Pattern Recognit. | 4 |
| 2024 | Linguistic Hallucination for Text-Based Video RetrievalabstractText-based video retrieval is a crucial technology for video and multimodal applications. Although in traditional Text-Video Retrieval caption-video pairs are supposed to be entirely relevant, there is still information missing in text when compared to the video content. In a specific application scenario of Text-Video Retrieval, where the given caption corresponds to only a segment of the target video, the challenge of aligning two modalities becomes particularly difficult. To address this issue, we introduce context information as an auxiliary to enrich text representation and enhance alignment. In this work, we propose an effective Linguistic Hallucination framework, which incorporates context captions during training and replaces them with hallucinated textual representations predicted from the source sentence at inference. Specific hallucination loss and consistency loss are designed to supervise the learning process. Besides, Curriculum Learning is introduced at both data-level and model-level, which makes the training procedure more stable and improves the retrieval performance simultaneously. Extensive comparison experiments and ablation studies on benchmark datasets demonstrate the effectiveness of our framework. Moreover, we also apply our proposed method to other cross-modal tasks and the promising experimental results prove its generalization ability. Our codes and datasets are available in https://github.com/silenceFS/Linguistic-Hallucination. Tiantian Dang, Shuhui Wang, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Collaborative Debias Strategy for Temporal Sentence Grounding in VideoabstractTemporal sentence grounding in video has witnessed significant advancements, but suffers from substantial dataset bias, which undermines its generalization ability. Existing debias approaches primarily concentrate on well-known distribution and linguistic biases, while overlooking the relationship among different biases, limiting their debias capability. In this work, we delve into the existence of visual bias and combinatorial bias in the widely used datasets, and introduce a collaborative debias structure that can be seamlessly integrated into present methods. It encompasses four low-capacity models, a re-label module, and a main model. Each biased model deliberately leverages bias as shortcut information to accurately perform grounding, achieved by customizing the appropriate model structure and input data format to align with the bias characteristics. During the training phase, the gradient descent direction for optimizing the main model should align with the negative gradient descent direction of the biased model that is optimized by utilizing ground truth labels. Subsequently, the re-label module introduces a gradient aggregation function, consolidating the gradient descent direction from these biased models and constructing new labels to compel the main model to effectively capture multi-modality alignment features instead of relying on shortcut contents for grounding. Finally, we design two debias structures, P-Debias and C-Debias, to exploit the independence and inclusion relationships between different types of biases. Extensive experiments on multiple span-based models over Charades-CD and ActivityNet-CD demonstrate the exceptional debias capability of our strategy (https://github.com/qzhb/CDS). Zhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang, Weigang Zhang, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | SpikeODE: Image Reconstruction for Spike Camera With Neural Ordinary Differential EquationabstractThe recently invented retina-inspired spike camera has shown great potential for capturing dynamic scenes. However, reconstructing high-quality images from the binary spike data remains a challenge due to the existence of noises in the camera. This paper proposes SpikeODE, a novel approach to reconstructing clear images by exploring temporal-spatial correlation to depress noises. The main idea of our method is to restore the continuous dynamic process of real scenes in a latent space and learn the temporal correlations in a fine-grained manner. Furthermore, to model the dynamic process more effectively, we design a conditional ODE where the latent state of each timestamp is conditioned on the observed spike data. Subsequently, forward and backward inferences are conducted through the ODE to investigate the correlations between the representation of the target timestamp and the information from both past and future contexts. Additionally, we incorporate a Unet structure with a pixel-wise attention mechanism at each level to learn spatial correlations. Experimental results demonstrate that our method outperforms state-of-the-art methods across several metrics. Chen Yang 0034, Guorong Li, Shuhui Wang, Li Su 0003, Laiyun Qing, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D ScenesabstractThree-Dimensional (3D) dense captioning is an emerging vision-language bridging task that aims to generate multiple detailed and accurate descriptions for 3D scenes. It presents significant potential and challenges due to its closer representation of the real world compared to 2D visual captioning, as well as complexities in data collection and processing of 3D point cloud sources. Despite the popularity and success of existing methods, there is a lack of comprehensive surveys summarizing the advancements in this field, which hinders its progress. In this paper, we provide a comprehensive review of 3D dense captioning, covering task definition, architecture classification, dataset analysis, evaluation metrics, and in-depth prosperity discussions. Based on a synthesis of previous literature, we refine a standard pipeline that serves as a common paradigm for existing methods. We also introduce a clear taxonomy of existing models, summarize technologies involved in different modules, and conduct detailed experiment analysis. Instead of a chronological order introduction, we categorize the methods into different classes to facilitate exploration and analysis of the differences and connections among existing techniques. We also provide a reading guideline to assist readers with different backgrounds and purposes in reading efficiently. Furthermore, we propose a series of promising future directions for 3D dense captioning by identifying challenges and aligning them with the development of related tasks, offering valuable insights and inspiring future research in this field. Our aim is to provide a comprehensive understanding of 3D dense captioning, foster further investigations, and contribute to the development of novel applications in multimedia and related domains. Ting Yu 0016, Shuhui Wang, Weiguo Sheng 0001, Qingming Huang, Jun Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | COMICS: End-to-End Bi-Grained Contrastive Learning for Multi-Face Forgery DetectionabstractDeepFakes have raised serious societal concerns, leading to a great surge in detection-based forensics methods in recent years. Face forgery recognition is a standard detection method that usually follows a two-phase pipeline,i.e., it extracts the face first and then determines its authenticity by classification. While those methods perform well in ideal experimental environment, they face challenges when dealing with DeepFakes in the wild involving complex background and multiple faces of varying sizes. Moreover, most face forgery recognition methods can only process one face at a time. One straightforward way to address this issue is to simultaneous process multi-face by integrating face extraction and forgery detection in an end-to-end fashion by adapting advanced object detection architectures. However, as these object detection architectures are designed to capture the discriminative features of different object categories rather than the subtle forgery traces among the faces, the direct adaptation suffers from limited representation ability. In this paper, we propose Contrastive Multi-FaceForensics (COMICS), an end-to-end framework for multi-face forgery detection. COMICS integrates face extraction and forgery detection in a seamless manner and adapts to the advanced object detection architectures. The core of the proposed framework is a bi-grained contrastive learning approach that explores face forgery traces at both the coarse- and fine-grained levels. Specifically, coarse-grained level contrastive learning captures the discriminative features among positive and negative proposal pairs at multiple layers produced by the proposal generator, and the fine-grained level contrastive learning captures the pixel-wise discrepancy between the forged and original areas of the same face and the pixel-wise content inconsistency among different faces. Extensive experiments on the OpenForensics and FFIW datasets demonstrate that our method outperforms other counterparts and shows great potential for being integrated into various architectures. Codes are available at https://github.com/zhangconghhh/COMICS. Honggang Qi, Shuhui Wang, Yuezun Li, Siwei Lyu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Dual-View Curricular Optimal Transport for Cross-Lingual Cross-Modal RetrievalabstractCurrent research on cross-modal retrieval is mostly English-oriented, as the availability of a large number of English-oriented human-labeled vision-language corpora. In order to break the limit of non-English labeled data, cross-lingual cross-modal retrieval (CCR) has attracted increasing attention. Most CCR methods construct pseudo-parallel vision-language corpora via Machine Translation (MT) to achieve cross-lingual transfer. However, the translated sentences from MT are generally imperfect in describing the corresponding visual contents. Improperly assuming the pseudo-parallel data are correctly correlated will make the networks overfit to the noisy correspondence. Therefore, we propose Dual-view Curricular Optimal Transport (DCOT) to learn with noisy correspondence in CCR. In particular, we quantify the confidence of the sample pair correlation with optimal transport theory from both the cross-lingual and cross-modal views, and design dual-view curriculum learning to dynamically model the transportation costs according to the learning stage of the two views. Extensive experiments are conducted on two multilingual image-text datasets and one video-text dataset, and the results demonstrate the effectiveness and robustness of the proposed method. Besides, our proposed method also shows a good expansibility to cross-lingual image-text baselines and a decent generalization on out-of-domain data. Shuhui Wang, Hao Luo 0004, Jianfeng Dong, Fan Wang 0019, Xun Wang 0007, Meng Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | ImageNet-E: Benchmarking Neural Network Robustness via Attribute EditingabstractRecent studies have shown that higher accuracy on ImageNet usually leads to better robustness against different corruptions. Therefore, in this paper, instead of following the traditional research paradigm that investigates new out-of-distribution corruptions or perturbations deep models may encounter, we conduct model debugging in in-distribution data to explore which object attributes a model may be sensitive to. To achieve this goal, we create a toolkit for object editing with controls of backgrounds, sizes, positions, and directions, and create a rigorous benchmark named ImageNet-E(diting) for evaluating the image classifier robustness in terms of object attributes. With our ImageNet-E, we evaluate the performance of current deep learning models, including both convolutional neural networks and vision transformers. We find that most models are quite sensitive to attribute changes. A small change in the background can lead to an average of 9.23% drop on top-1 accuracy. We also evaluate some robust models including both adversarially trained models and other robust trained models and find that some models show worse robustness against attribute changes than vanilla models. Based on these findings, we discover ways to enhance attribute robustness with preprocessing, architecture designs, and training strategies. We hope this work can provide some insights to the community and open up a new avenue for research in robust computer vision. The code and dataset are available at https://github.com/alibaba/easyrobust. Yuefeng Chen, Yao Zhu 0003, Shuhui Wang, Rong Zhang 0006, Hui Xue 0001 |
CVPR | 4 |
| 2023 | Exploiting Completeness and Uncertainty of Pseudo Labels for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection aims to identify abnormal events in videos using only video-level labels. Recently, two-stage self-training methods have achieved significant improvements by self-generating pseudo labels and self-refining anomaly scores with these labels. As the pseudo labels play a crucial role, we propose an enhancement framework by exploiting completeness and uncertainty properties for effective self-training. Specifically, we first design a multi-head classification module (each head serves as a classifier) with a diversity loss to maximize the distribution differences of predicted pseudo labels across heads. This encourages the generated pseudo labels to cover as many abnormal events as possible. We then devise an iterative uncertainty pseudo label refinement strategy, which improves not only the initial pseudo labels but also the updated ones obtained by the desired classifier in the second stage. Extensive experimental results demonstrate the proposed method performs favorably against state-of-the-art approaches on the UCF-Crime, TAD, and XD-Violence benchmark datasets. Chen Zhang 0013, Guorong Li, Yuankai Qi, Shuhui Wang, Laiyun Qing, Qingming Huang, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2023 | The Euclidean Space is Evil: Hyperbolic Attribute Editing for Few-shot Image GenerationabstractFew-shot image generation is a challenging task since it aims to generate diverse new images for an unseen category with only a few images. Existing methods suffer from the trade-off between the quality and diversity of generated images. To tackle this problem, we propose Hyperbolic Attribute Editing (HAE), a simple yet effective method. Unlike other methods that work in Euclidean space, HAE captures the hierarchy among images using data from seen categories in hyperbolic space. Given a well-trained HAE, images of unseen categories can be generated by moving the latent code of a given image toward any meaningful directions in the Poincaré disk with a fixing radius. Most importantly, the hyperbolic space allows us to control the semantic diversity of the generated images by setting different radii in the disk. Extensive experiments and visualizations demonstrate that HAE is capable of not only generating images with promising quality and diversity using limited data but achieving a highly controllable and interpretable editing process. Code is available at https://github.com/lingxiao-li/HAE. Yi Zhang 0120, Shuhui Wang |
ICCV | 3 |
| 2023 | Dynamics-inspired Neuromorphic Visual Representation LearningabstractThis paper investigates the dynamics-inspired neuromorphic architecture for visual representation learning following Hamilton’s principle. Our method converts weight-based neural structure to its dynamics-based form that consists of finite sub-models, whose mutual relations measured by computing path integrals amongst their dynamical states are equivalent to the typical neural weights. Based on the entropy reduction process derived from the Euler-Lagrange equations, the feedback signals interpreted as stress forces amongst sub-models push them to move. We first train a dynamics-based neural model from scratch and observe that this model outperforms traditional neural models on MNIST. We then convert several pre-trained neural structures into dynamics-based forms, followed by fine-tuning via entropy reduction to obtain the stabilized dynamical states. We observe consistent improvements in these transformed models over their weight-based counterparts on ImageNet and WebVision in terms of computational complexity, parameter size, testing accuracy, and robustness. Besides, we show the correlation between model performance and structural entropy, providing deeper insight into weight-free neuromorphic learning. Zhengqi Pei, Shuhui Wang |
ICML | 2 |
| 2023 | All in a Row: Compressed Convolution Networks for GraphsabstractCompared to Euclidean convolution, existing graph convolution methods generally fail to learn diverse convolution operators under limited parameter scales and depend on additional treatments of multi-scale feature extraction. The challenges of generalizing Euclidean convolution to graphs arise from the irregular structure of graphs. To bridge the gap between Euclidean space and graph space, we propose a differentiable method for regularization on graphs that applies permutations to the input graphs. The permutations constrain all nodes in a row regardless of their input order and therefore enable the flexible generalization of Euclidean convolution. Based on the regularization of graphs, we propose Compressed Convolution Network (CoCN) for hierarchical graph representation learning. CoCN follows the local feature learning and global parameter sharing mechanisms of Convolution Neural Networks. The whole model can be trained end-to-end and is able to learn both individual node features and the corresponding structure features. We validate CoCN on several node classification and graph classification benchmarks. CoCN achieves superior performance over competitive convolutional GNNs and graph pooling models. Codes are available at https://github.com/sunjss/CoCN. Junshu Sun, Shuhui Wang, Xinzhe Han, Zhe Xue, Qingming Huang |
ICML | 2 |
| 2023 | CoP: Chain-of-Pose for Image Animation in Large Pose ChangesabstractImage animation involves generating a video of a source image imitating the pose of a driving video. Despite recent advancements in the image animation task, most state-of-the-art methods remain vulnerable to large pose changes. In cases of large pose changes, existing methods struggle to model the complex nonlinear motion and yield distorted results, which greatly restricts their application in the real world. To tackle this problem, we present a novel approach called Chain-of-Pose (CoP) that decomposes large pose changes into a sequence of intermediate pose changes. This enables us to handle simplified pose changes and improves the accuracy of pose estimation. Furthermore, to better preserve the appearance of the source object, we introduce the Appearance Refinement Module (ARM) that effectively integrates the appearance texture feature of the source image with the structural pose feature from the pose chain. Our experimental results demonstrate that our method qualitatively and quantitatively outperforms state-of-the-art approaches on four diverse datasets, comprising talking faces, human bodies, and pixel animals. Notably, our approach significantly improves video quality in the case of large object pose changes. Our code is attached to the supplementary material. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Shuhui Wang, Jiao Dai, Jizhong Han |
ACM Multimedia | 4 |
| 2023 | Conversational Composed Retrieval with Iterative Sequence RefinementabstractDue to the progress of large-scale multimodal model pretraining, existing cross-modal retrieval techniques is accurate to align text description to the target image when they show close and clear semantic correspondence. However, in real situations, users only provide ambiguous text queries, making it difficult to retrieve the desired images. To address this issue, we introduce the conversational composed retrieval paradigm, inspired by conversational search which models complex user intent through iterative interaction. This paradigm enhances the model capacity in learning fine-grained correspondences. To train the cross-modal conversational retrieval, we propose the Iterative Refining Retrieval (IRR) framework. It formalizes the reference images and modification texts in each session as a multimodal sequence, which is fed into the generative model to predict the information in the sequence autoregressively, and ultimately predicting the target image feature. In the conversational retrieval paradigm, the model refines the learned correspondences based on the interaction in the later stage of the retrieval session, thus captures fine-grained semantic correspondence to enforce the cross-modal representation. We propose a domain-specific multimodal pretraining method and the full sequence sampling augmentation method to fully utilize the session information. Extensive experiments demonstrate that the iterative refining retrieval method achieves state-of-the-art performance on sessions of varying lengths. Shuhui Wang, Zhe Xue, Shengbo Chen, Qingming Huang |
ACM Multimedia | 2 |
| 2023 | Orthogonal Temporal Interpolation for Zero-Shot Video RecognitionabstractZero-shot video recognition (ZSVR) is a task that aims to recognize video categories that have not been seen during the model training process. Recently, vision-language models (VLMs) pre-trained on large-scale image-text pairs have demonstrated impressive transferability for ZSVR. To make VLMs applicable to the video domain, existing methods often use an additional temporal learning module after the image-level encoder to learn the temporal relationships among video frames. Unfortunately, for video from unseen categories, we observe an abnormal phenomenon where the model that uses spatial-temporal feature performs much worse than the model that removes temporal learning module and uses only spatial feature. We conjecture that improper temporal modeling on video disrupts the spatial feature of the video. To verify our hypothesis, we propose Feature Factorization to retain the orthogonal temporal feature of the video and use interpolation to construct refined spatial-temporal feature. The model using appropriately refined spatial-temporal feature performs better than the one using only spatial feature, which verifies the effectiveness of the orthogonal temporal feature for the ZSVR task. Therefore, an Orthogonal Temporal Interpolation module is designed to learn a better refined spatial-temporal video feature during training. Additionally, a Matching Loss is introduced to improve the quality of the orthogonal temporal feature. We propose a model called OTI for ZSVR by employing orthogonal temporal interpolation and the matching loss based on VLMs. The ZSVR accuracies on popular video datasets (i.e., Kinetics-600, UCF101 and HMDB51) show that OTI outperforms the previous state-of-the-art method by a clear margin.Our codes are publicly available at https://github.com/yanzhu/mm2023_oti. Junbao Zhuo, Bin Ma 0028, Jiajia Geng, Xiaoming Wei, Xiaolin Wei, Shuhui Wang |
ACM Multimedia | 7 |
| 2023 | Adaptive Feature Swapping for Unsupervised Domain AdaptationabstractThe bottleneck of visual domain adaptation always lies in the learning of domain invariant representations. In this paper, we present a simple but effective technique named Adaptive Feature Swapping for learning domain invariant features in Unsupervised Domain Adaptation (UDA). Adaptive Feature Swapping aims to select semantically irrelevant features from labeled source data and unlabeled target data and swap these features with each other. Then the merged representations are also utilized for training with prediction consistency constraints. In this way, the model is encouraged to learn representations that are robust to domain-specific information. We develop two swapping strategies including channel swapping and spatial swapping. The former encourages the model to squeeze redundancy out of features and pay more attention to semantic information. The latter motivates the model to be robust to the background and focus on objects. We conduct experiments on object recognition and semantic segmentation in UDA setting and the results show that Adaptive Feature Swapping can promote various existing UDA methods. Our codes are publicly available at https://github.com/junbaoZHUO/AFS. Junbao Zhuo, Xingyu Zhao 0005, Shuhao Cui, Qingming Huang, Shuhui Wang |
ACM Multimedia | 5 |
| 2023 | Synthesizing Videos from Images for Image-to-Video AdaptationabstractWe address the image-to-video adaptation task that aims to leverage labeled images and unlabeled videos for video recognition. There are two major challenges in this task, including the domain discrepancy between the two domains, and the modality gap between the image and video modalities. Existing methods mainly employ a two-stage paradigm by first adopting frame-level adaptation to reduce the domain discrepancy and then learning a spatio-temporal model to bridge the modality gap. In this paper, we provide a new perspective and propose a single-stage method that synthesizes video from the source static image and converts the image-to-video adaptation problem into a video-to-video adaptation problem. With the synthesized video, we present a simple baseline that a spatio-temporal model is trained with cross entropy loss with source labels and the Batch Nuclear norm Maximization loss to encourage the classification responses of target videos maintain the discriminability and diversity. We further propose a new pseudo label generation method that inherits the robustness of class prototype and the effectiveness of the small loss criterion. Based on the constructed baseline and the proposed pseudo label generation method, we train a model that achieves state-of-the-art performances or gets comparable performances on three standard benchmarks. Our codes are publicly available at https://github.com/junbaoZHUO/ST-I2V. Junbao Zhuo, Xingyu Zhao 0005, Shuhui Wang, Huimin Ma 0001, Qingming Huang |
ACM Multimedia | 3 |
| 2023 | EcoDialTest: Adaptive Mutation Schedule for Automated Dialogue Systems TestingabstractWith the rapid growth of Artificial Intelligence, dialogue systems have become increasingly powerful. Though Recurrent Neural Network power the dialogue systems, it also bring challenges to the systems’ testing. In order to ensure the safety of these systems, which we must pay attention to, DialTest showed up. DialTest broke the traditional test methods, it made innovation at many levels. We have to acknowledge this great contribution. However, DialTest has a smattering of shortcomings. It treats all seeds as equal, implying that it cannot adjust the energy assignment quickly, resulting in energy waste. Moreover, DialTest’s mutant sentences generated by a few original seed sentences in the late stage of variation. This paper presents an improved DialTest with an adaptive mutation schedule, we called it EcoDialTest. EcoDialTest divides all the seed into three states, different states have different energy distribution strategies. We devise a new mutation strategy to improve the effectiveness and dependability of the seeds in the transformed seed set. All of these were implemented based on DialTest, we still adopt DeepGini impurity as the main guidance to guide the test generation process and utilize the three mutation operators as it does. Through ATIS, Snips and Facebook datasets, EcoDialTest was evaluated by two state-of-the-art models in the experiment. According to the result, we found that EcoDialTest attained lower values in both intent accuracy and slot accuracy than DialTest. Xiangchen Shen, Haibo Chen 0005, Jinfu Chen 0001, Shuhui Wang |
SANER | 5 |
| 2023 | A graph neural network-based data cleaning method to prevent intelligent fault diagnosis from data contamination
Shuhui Wang, Yaguo Lei, Bin Yang 0014, Xiang Li 0018, Yue Shu |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | General Greedy De-Bias LearningabstractNeural networks often make predictions relying on the spurious correlations from the datasets rather than the intrinsic properties of the task of interest, facing with sharp degradation on out-of-distribution (OOD) test data. Existing de-bias learning frameworks try to capture specific dataset bias by annotations but they fail to handle complicated OOD scenarios. Others implicitly identify the dataset bias by special design low capability biased models or losses, but they degrade when the training and testing data are from the same distribution. In this paper, we propose a General Greedy De-bias learning framework (GGD), which greedily trains the biased models and base model. The base model is encouraged to focus on examples that are hard to solve with biased models, thus remaining robust against spurious correlations in the test stage. GGD largely improves models' OOD generalization ability on various tasks, but sometimes over-estimates the bias level and degrades on the in-distribution test. We further re-analyze the ensemble process of GGD and introduce the Curriculum Regularization inspired by curriculum learning, which achieves a good trade-off between in-distribution (ID) and out-of-distribution performance. Extensive experiments on image classification, adversarial question answering, and visual question answering demonstrate the effectiveness of our method. GGD can learn a more robust base model under the settings of both task-specific biased models with prior knowledge and self-ensemble biased model without prior knowledge. Codes are available at https://github.com/GeraldHan/GGD. Xinzhe Han, Shuhui Wang, Chi Su, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Entity-Enhanced Adaptive Reconstruction Network for Weakly Supervised Referring Expression GroundingabstractWeakly supervised Referring Expression Grounding (REG) aims to ground a particular target in an image described by a language expression while lacking the correspondence between target and expression. Two main problems exist in weakly supervised REG. First, the lack of region-level annotations introduces ambiguities between proposals and queries. Second, most previous weakly supervised REG methods ignore the discriminative location and context of the referent, causing difficulties in distinguishing the target from other same-category objects. To address the above challenges, we design an entity-enhanced adaptive reconstruction network (EARN). Specifically, EARN includes three modules: entity enhancement, adaptive grounding, and collaborative reconstruction. In entity enhancement, we calculate semantic similarity as supervision to select the candidate proposals. Adaptive grounding calculates the ranking score of candidate proposals upon subject, location and context with hierarchical attention. Collaborative reconstruction measures the ranking result from three perspectives: adaptive reconstruction, language reconstruction and attribute classification. The adaptive mechanism helps to alleviate the variance of different referring expressions. Experiments on five datasets show EARN outperforms existing state-of-the-art methods. Qualitative results demonstrate that the proposed EARN can better handle the situation where multiple objects of a particular category are situated together. Xuejing Liu, Liang Li 0003, Shuhui Wang, Zhengjun Zha, Zechao Li, Qi Tian 0001, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Self-Regulated Learning for Egocentric Video Activity AnticipationabstractFuture activity anticipation is a challenging problem in egocentric vision. As a standard future activity anticipation paradigm, recursive sequence prediction suffers from the accumulation of errors. To address this problem, we propose a simple and effective Self-Regulated Learning framework, which aims to regulate the intermediate representation consecutively to produce representation that (a) emphasizes the novel information in the frame of the current time-stamp in contrast to previously observed content, and (b) reflects its correlation with previously observed frames. The former is achieved by minimizing a contrastive loss, and the latter can be achieved by a dynamic reweighing mechanism to attend to informative frames in the observed content with a similarity comparison between feature of the current frame and observed frames. The learned final video representation can be further enhanced by multi-task learning which performs joint feature learning on the target activity labels and the automatically detected action and object class tokens. SRL sharply outperforms existing state-of-the-art in most cases on two egocentric video datasets and two third-person video datasets. Its effectiveness is also verified by the experimental fact that the action and object concepts that support the activity semantics can be accurately identified. Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Uncertainty Modeling for Robust Domain Adaptation Under Noisy EnvironmentsabstractIn this paper, we tackle the task of domain adaptation under noisy environments; this is a practical and challenging problem in which the source domain is corrupted with noise in its labels, its features, or both. Noise in the source domain leads to inaccurate visual representations and makes it harder to estimate and reduce the domain discrepancy between the source and target domains, resulting in severe performance degradation in the target domain. These challenges can be addressed with offline source sample selection following robust domain discrepancy reduction. To achieve reliable sample selection, we model the uncertainty in the predictions of a convolutional neural network (CNN) classifier and reweight the classification loss by this uncertainty. Such a reweighting mechanism reduces the contribution of noise, leading to improved noise robustness. We further propose UncertaintyRank, a novel regularizer, to encourage the uncertainty to be more sensitive to noisy labels, as label corruption brings more severe degradation. The uncertainty is also aggregated with the classification loss to eliminate the adverse effects of noisy representations while estimating the domain discrepancy. Extensive experiments validate the effectiveness of our method and verify that it performs favorably against existing state-of-the-art methods. Junbao Zhuo, Shuhui Wang, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2023 | Weakly Supervised Text-based Actor-Action Video Segmentation by Clip-level Multi-instance LearningabstractIn real-world scenarios, it is common that a video contains multiple actors and their activities. Selectively localizing one specific actor and its action spatially and temporally via a language query becomes a vital and challenging task. Existing fully supervised methods require extensive elaborately annotated data and are sensitive to the class labels, which cannot satisfy real-world applications’ needs. Thus, we introduce the task of weakly supervised actor-action video segmentation from a sentence query (AAVSS) in this work, where only the video-sentence pairs are provided. To the best of our knowledge, our work is the first to perform AAVSS under weakly supervised situations. However, this task is extremely challenging not only because the task aims to learn the complex interactions between two heterogeneous modalities but also because the task needs to learn fine-grained analysis of video content without pixel-level annotations. To overcome the challenges, we propose a two-stage network. The network first follows the sentence guidance to localize the candidate region and then performs segmentation to achieve selective segmentation. Specifically, a novel tracker-based clip-level multiple instance learning paradigm is proposed in this article to learn the matches between regions and sentences, which makes our two-stage network robust to the region proposal network. Furthermore, two intrinsic characteristics of the video, temporal consistency and motion information, are utilized in companion with the weak supervision to facilitate the region-query matching. Through extensive experiments, the proposed method achieves comparable performance to state-of-the-art fully supervised approaches on two large-scale benchmarks, including A2D Sentences and J-HMDB Sentences. Weidong Chen 0013, Guorong Li, Xinfeng Zhang 0001, Shuhui Wang, Liang Li 0003, Qingming Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Temporal Dynamic Concept Modeling Network for Explainable Video Event RecognitionabstractRecently, with the vigorous development of deep learning and multimedia technology, intelligent urban computing has received more and more extensive attention from academia and industry. Unfortunately, most of the related technologies are black-box paradigms that lack interpretability. Among them, video event recognition is a basic technology. Event contains multiple concepts and their rich interactions, which can assist us to construct explainable event recognition methods. However, the crucial concepts needed to recognize events have various temporal existing patterns, and the relationship between events and the temporal characteristics of concepts has not been fully exploited. This brings great challenges for concept-based event categorization. To address the above issues, we introduce the temporal concept receptive field, which is the length of the temporal window size required to capture key concepts for concept-based event recognition methods. Accordingly, we introduce the temporal dynamic convolution (TDC) to model the temporal concept receptive field dynamically according to different events. Its core idea is to combine the results of multiple convolution layers with the learned coefficients from two complementary perspectives. These convolution layers contain a variety of kernel sizes, which can provide temporal concept receptive fields of different lengths. Similarly, we also propose the cross-domain temporal dynamic convolution (CrTDC) with the help of the rich relationship between different concepts. Different coefficients can help us to capture suitable temporal concept receptive field sizes and highlight crucial concepts to obtain accurate and complete concept representations for event analysis. Based on the TDC and CrTDC, we introduce the temporal dynamic concept modeling network (TDCMN) for explainable video event recognition. We evaluate TDCMN on large-scale and challenging datasets FCVID, ActivityNet, and CCV. Experimental results show that TDCMN significantly improves the event recognition performance of concept-based methods, and the explainability of our method inspires us to construct more explainable models from the perspective of the temporal concept receptive field. Weigang Zhang, Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Unsupervised Coherent Video Cartoonization with Perceptual Motion ConsistencyabstractIn recent years, creative content generations like style transfer and neural photo editing have attracted more and more attention. Among these, cartoonization of real-world scenes has promising applications in entertainment and industry. Different from image translations focusing on improving the style effect of generated images, video cartoonization has additional requirements on the temporal consistency. In this paper, we propose a spatially-adaptive semantic alignment framework with perceptual motion consistency for coherent video cartoonization in an unsupervised manner. The semantic alignment module is designed to restore deformation of semantic structure caused by spatial information lost in the encoder-decoder architecture. Furthermore, we introduce the spatio-temporal correlative map as a style-independent, global-aware regularization on perceptual motion consistency. Deriving from similarity measurement of high-level features in photo and cartoon frames, it captures global semantic information beyond raw pixel-value of optical flow. Besides, the similarity measurement disentangles temporal relationship from domain-specific style properties, which helps regularize the temporal consistency without hurting style effects of cartoon images. Qualitative and quantitative experiments demonstrate our method is able to generate highly stylistic and temporal consistent cartoon videos. Zhenhuan Liu, Liang Li 0003, Huajie Jiang, Xin Jin 0004, Dandan Tu, Shuhui Wang, Zhengjun Zha |
AAAI | 6 |
| 2022 | Revisiting Unsupervised Domain Adaptation Models: A Smoothness Perspective
Junbao Zhuo, Shuhui Wang, Yuejian Fang |
ACCV (6) | 4 |
| 2022 | Attribute Group Editing for Reliable Few-shot Image GenerationabstractFew-shot image generation is a challenging task even using the state-of-the-art Generative Adversarial Networks (GANs). Due to the unstable GAN training process and the limited training data, the generated images are often of low quality and low diversity. In this work, we propose a new “editing-based” method, i.e., Attribute Group Editing (AGE), for few-shot image generation. The basic assumption is that any image is a collection of attributes and the editing direction for a specific attribute is shared across all categories. AGE examines the internal representation learned in GANs and identifies semantically meaningful directions. Specifically, the class embedding, i.e., the mean vector of the latent codes from a specific category, is used to represent the category-relevant attributes, and the category-irrelevant attributes are learned globally by Sparse Dictionary Learning on the difference between the sample embedding and the class embedding. Given a GAN well trained on seen categories, diverse images of unseen categories can be synthesized through editing category-irrelevant attributes while keeping category-relevant attributes unchanged. Without re-training the GAN, AGE is capable of not only producing more realistic and diverse images for downstream visual applications with limited data but achieving controllable image editing with interpretable category-irrelevant directions. Code is available at https://github.com/UniBester/AGE. Guanqi Ding, Xinzhe Han, Shuhui Wang, Shuzhe Wu, Xin Jin 0004, Dandan Tu, Qingming Huang |
CVPR | 3 |
| 2022 | DeeCap: Dynamic Early Exiting for Efficient Image CaptioningabstractBoth accuracy and efficiency are crucial for image captioning in real-world scenarios. Although Transformer-based models have gained significant improved captioning performance, their computational cost is very high. A feasible way to reduce the time complexity is to exit the prediction early in internal decoding layers without passing the entire model. However, it is not straightforward to devise early exiting into image captioning due to the following issues. On one hand, the representation in shallow layers lacks high-level semantic and sufficient cross-modal fusion information for accurate prediction. On the other hand, the exiting decisions made by internal classifiers are unreliable sometimes. To solve these issues, we propose DeeCap framework for efficient image captioning, which dynamically selects proper-sized decoding layers from a global perspective to exit early. The key to successful early exiting lies in the specially designed imitation learning mechanism, which predicts the deep layer activation with shallow layer features. By deliberately merging the imitation learning into the whole image captioning architecture, the imitated deep layer representation can mitigate the loss brought by the missing of actual deep layers when early exiting is undertaken, resulting in significant reduction in calculation cost with small sacrifice of accuracy. Experiments on the MS COCO and Flickr30k datasets demonstrate the DeeCap can achieve competitive performances with 4× speed-up. Code is available at: https://github.com/feizc/DeeCap. Zhengcong Fei, Shuhui Wang, Qi Tian 0001 |
CVPR | 3 |
| 2022 | Hierarchical Modular Network for Video CaptioningabstractVideo captioning aims to generate natural language descriptions according to the content, where representation learning plays a crucial role. Existing methods are mainly developed within the supervised learning framework via word-by-word comparison of the generated caption against the ground-truth text without fully exploiting linguistic semantics. In this work, we propose a hierarchical modular network to bridge video representations and linguistic semantics from three levels before generating captions. In particular, the hierarchy is composed of: (I) Entity level, which highlights objects that are most likely to be mentioned in captions. (II) Predicate level, which learns the actions conditioned on highlighted objects and is supervised by the predicate in captions. (III) Sentence level, which learns the global semantic representation and is supervised by the whole caption. Each level is implemented by one module. Extensive experimental results show that the proposed method performs favorably against the state-of-the-art models on the two widely-used benchmarks: MSVD 104.0% and MSR-VTT 51.5% in CIDEr score. Code will be made available at https://github.com/MarcusNerva/HMN. Hanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang, Qingming Huang, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2022 | Learning Linguistic Association Towards Efficient Text-Video Retrieval
Shuhui Wang, Junbao Zhuo, Xinzhe Han, Qingming Huang |
ECCV (36) | 2 |
| 2022 | Inferential Visual Question GenerationabstractThe task of Visual Question Generation (VQG) aims to generate natural language questions for images. Many methods regard it as a reverse Visual Question Answering (VQA) task. They trained a data-driven generator on VQA datasets, which is hard to obtain questions that can challenge robots and humans. Other methods rely heavily on elaborate but expensive artificial preprocessing to generate. To overcome these limitations, we propose a method to generate inferential questions from the image with noisy captions. Our method first introduces a core scene graph generation module, which can align text features and salient visual features to the initial scene graph. It constructs a special core scene graph with expanded linkage outwards from the high-confidence nodes hop by hop. Next, a question generation module uses the core scene graph as a basis to instantiate the function templates, resulting in questions with varying inferential paths. Experiments show that the visual questions generated by our method are controllable in both content and difficulty, and demonstrate clear inferential properties. In addition, since the salient region, captions, and function templates can be replaced by human-customized ones, our method has strong scalability and potential for more interactive applications. Finally, we use our method to automatically build a new dataset, InVQA, containing about 120k images and 480k question-answer pairs, to facilitate the development of more versatile VQA models. Chao Bi, Shuhui Wang, Zhe Xue, Shengbo Chen, Qingming Huang |
ACM Multimedia | 2 |
| 2022 | Multi-Attention Network for Compressed Video Referring Object SegmentationabstractReferring video object segmentation aims to segment the object referred by a given language expression. Existing works typically require compressed video bitstream to be decoded to RGB frames before being segmented, which increases computation and storage requirements and ultimately slows the inference down. This may hamper its application in real-world computing resource limited scenarios, such as autonomous cars and drones. To alleviate this problem, in this paper, we explore the referring object segmenta- tion task on compressed videos, namely on the original video data flow. Besides the inherent difficulty of the video referring object segmentation task itself, obtaining discriminative representation from compressed video is also rather challenging. To address this problem, we propose a multi-attention network which consists of dual-path dual-attention module and a query-based cross-modal Transformer module. Specifically, the dual-path dual-attention module is designed to extract effective representation from compressed data in three modalities, i.e., I-frame, Motion Vector and Residual. The query-based cross-modal Transformer firstly models the corre- lation between linguistic and visual modalities, and then the fused multi-modality features are used to guide object queries to generate a content-aware dynamic kernel and to predict final segmentation masks. Different from previous works, we propose to learn just one kernel, which thus removes the complicated post mask-matching procedure of existing methods. Extensive promising experimental results on three challenging datasets show the effectiveness of our method compared against several state-of-the-art methods which are proposed for processing RGB data. Source code is available at: https://github.com/DexiangHong/MANet. Weidong Chen 0013, Dexiang Hong, Yuankai Qi, Zhenjun Han, Shuhui Wang, Laiyun Qing, Qingming Huang, Guorong Li |
ACM Multimedia | 5 |
| 2022 | Concept Propagation via Attentional Knowledge Graph Reasoning for Video-Text RetrievalabstractDue to the rapid growth of online video data, video-text retrieval techniques are in urgent need, which aim to search for the most relevant video given a natural language caption and vice versa. The major challenge of this task is how to identify the true fine-grained semantic correspondence between videos and texts, using only the document-level correspondence. To deal with this issue, we propose a simple yet effective two-stream framework which takes the concept information into account and introduces a new branch of semantic-level matching. We further propose a concept propagation mechanism for mining the latent semantics in videos and achieving enriched representations. The concept propagation is achieved by building a commonsense graph distilled from ConceptNet with concepts extracted from videos and captions. The original concepts of videos are detected by pretrained detectors as the initial concept representations. By conducting attentional graph reasoning on the commonsense graph with the guidance of external knowledge, we can extend some new concepts in a detector-free manner for further enriching the video representations. In addition, a propagated BCE loss is designed for supervising the concept propagation procedure. Common space learning is then constructed for cross-modal matching. We conduct extensive experiments on various baseline models and several benchmark datasets. Promising experimental results demonstrate the effectiveness and generalization ability of our method. Shuhui Wang, Junbao Zhuo, Qingming Huang, Bin Ma 0028, Xiaoming Wei, Xiaolin Wei |
ACM Multimedia | 2 |
| 2022 | Synthesizing Counterfactual Samples for Effective Image-Text MatchingabstractImage-text matching is a fundamental research topic bridging vision and language. Recent works use hard negative mining to capture the multiple correspondences between visual and textual domains. Unfortunately, the truly informative negative samples are quite sparse in the training data, which are hard to obtain only in a randomly sampled mini-batch. Motivated by causal inference, we aim to overcome this shortcoming by carefully analyzing the analogy between hard negative mining and causal effects optimizing. Further, we propose Counterfactual Matching (CFM) framework for more effective image-text correspondence mining. CFM contains three major components, \ie, Gradient-Guided Feature Selection for automatic casual factor identification, Self-Exploration for causal factor completeness, and Self-Adjustment for counterfactual sample synthesis. Compared with traditional hard negative mining, our method largely alleviates the over-fitting phenomenon and effectively captures the fine-grained correlations between image and text modality. We evaluate our CFM in combination with three state-of-the-art image-text matching architectures. Quantitative and qualitative experiments conducted on two publicly available datasets demonstrate its strong generality and effectiveness. Code is available at: https://github.com/weihao20/cfm. Shuhui Wang, Xinzhe Han, Zhe Xue, Bin Ma 0028, Xiaoming Wei, Xiaolin Wei |
ACM Multimedia | 2 |
| 2022 | Zero-shot Video Classification with Appropriate Web and Task Knowledge TransferabstractZero-shot video classification (ZSVC) that aims to recognize video classes that have never been seen during model training, has become a thriving research direction. ZSVC is achieved by building mappings between visual and semantic embeddings. Recently, ZSVC has been achieved by automatically mining the underlying objects in videos as attributes and incorporating external commonsense knowledge. However, the object mined from seen categories can not generalized to unseen ones. Besides, the category-object relationships are usually extracted from commonsense knowledge or word embedding, which is not consistent with video modality. To tackle these issues, we propose to mine associated objects and category-object relationships for each category from retrieved web images. The associated objects of all categories are employed as generic attributes and the mined category-object relationships could narrow the modality inconsistency for better knowledge transfer. Another issue of existing ZSVC methods is that the model sufficiently trained with labeled seen categories may not generalize well to distinct unseen categories. To encourage a more reliable transfer, we propose Task Similarity aware Representation Learning (TSRL). In TSRL, the similarity between seen categories and the unseen ones is estimated and used to regularize the model in an appropriate way. We construct a model for ZSVC based on the constructed attributes, the mined category-object relationships and the proposed TSRL. Experimental results on four public datasets, i.e., FCVID, UCF101, HMDB51 and Olympic Sports, show that our model performs favorably against state-of-the-art methods. Our codes are publicly available at https://github.com/junbaoZHUO/TSRL. Junbao Zhuo, Shuhao Cui, Shuhui Wang, Bin Ma 0028, Qingming Huang, Xiaoming Wei, Xiaolin Wei |
ACM Multimedia | 4 |
| 2022 | Improved surrogate-assisted whale optimization algorithm for fractional chaotic systems ' parameters identification
Shuhui Wang, Wei Hu 0013, Ignacio Riego, Yongguang Yu |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | Semantic inpainting on segmentation map via multi-expansion loss
Xuchao Zhang, Shuo Lei, Shuhui Wang, Chang-Tien Lu, Bei Xiao |
Neurocomputing | 4 |
| 2022 | A string matching based ultra-low complexity lossless screen content coding technique
Yufen Yang, Tao Lin 0005, Liping Zhao 0005, Kailun Zhou, Shuhui Wang |
Multim. Tools Appl. | 5 |
| 2022 | Syntax-Guided Hierarchical Attention Network for Video CaptioningabstractVideo captioning is a challenging task that aims to generate linguistic description based on video content. Most methods only incorporate visual features (2D/3D) as input for generating visual and non-visual words in the caption. However, generating non-visual words usually depends more on sentence-context than visual features. The wrong non-visual words can reduce the sentence fluency and even change the meaning of sentence. In this paper, we propose a syntax-guided hierarchical attention network (SHAN), which leverages semantic and syntax cues to integrate visual and sentence-context features for captioning. First, a globally-dependent context encoder is designed to extract the global sentence-context feature that facilitates generating non-visual words. Then, we introduce hierarchical content attention and syntax attention to adaptively integrate features in terms of temporality and feature characteristics respectively. Content attention helps focus on time intervals related to the semantic of current word, while cross-modal syntax attention uses syntax information to model importance of different features for target word’s generation. Moreover, such hierarchical attention can enhance the model interpretability for captioning. Experiments on MSVD and MSR-VTT datasets show the comparable performance of our method compared with current methods. Jincan Deng, Liang Li 0003, Beichen Zhang 0006, Shuhui Wang, Zhengjun Zha, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Transfer Relation Network for Fault Diagnosis of Rotating Machinery With Small DataabstractMany deep-learning methods have been developed for fault diagnosis. However, due to the difficulty of collecting and labeling machine fault data, the datasets in some practical applications are relatively much smaller than the other big data benchmarks. In addition, the fault data come from different machines. Therefore, on some occasions, fault diagnosis is a multidomain problem with small data, where satisfactory transfer performance is difficult to obtain and has been rarely explored from the few-shot learning viewpoint. Different from the existing deep transfer learning solutions, a novel transfer relation network (TRN), combining a few-shot learning mechanism and transfer learning, is developed in this study. Specifically, the fault diagnosis problem has been treated as a similarity metric-learning problem instead of solely feature weighted classification. A feature net and a relation net have been, respectively, constructed for feature extraction and relation computation. The Siamese structure has been borrowed to extract the features of the source and the target domain samples with shared weights. Multikernel maximum mean discrepancy (MK-MMD) is employed on several higher layers with different tradeoff parameters to enable an efficient domain feature transfer considering different feature properties. To implement efficient diagnosis based on small data, an episode-based few-shot training strategy is adopted to train TRN. Average pooling has been adopted to suppress the noise influence from the vibration sequence which turns out to be important for the success of time sequence-based fault diagnosis. Transfer experiments on four datasets have verified the superior performance of TRN. A significant improvement of classification accuracy has been made compared with the state-of-the-art methods on the adopted datasets. Huiyang Hu, Yaguo Lei, Shuhui Wang |
IEEE Trans. Cybern. | 5 |
| 2021 | Composite Adversarial AttacksabstractAdversarial attack is a technique for deceiving Machine Learning (ML) models, which provides a way to evaluate the adversarial robustness. In practice, attack algorithms are artificially selected and tuned by human experts to break a ML system. However, manual selection of attackers tends to be sub-optimal, leading to a mistakenly assessment of model security. In this paper, a new procedure called Composite Adversarial Attack (CAA) is proposed for automatically searching the best combination of attack algorithms and their hyper-parameters from a candidate pool of 32 base attackers. We design a search space where attack policy is represented as an attacking sequence, i.e., the output of the previous attacker is used as the initialization input for successors. Multi-objective NSGA-II genetic algorithm is adopted for finding the strongest attack policy with minimum complexity. The experimental result shows CAA beats 10 top attackers on 11 diverse defenses with less elapsed time (6 × faster than AutoAttack), and achieves the new state-of-the-art on linf, l2 and unrestricted adversarial attacks. Xiaofeng Mao, Yuefeng Chen, Shuhui Wang, Hang Su 0006, Yuan He 0011, Hui Xue 0001 |
AAAI | 3 |
| 2021 | QAIR: Practical Query-Efficient Black-Box Attacks for Image RetrievalabstractWe study the query-based attack against image retrieval to evaluate its robustness against adversarial examples under the black-box setting, where the adversary only has query access to the top-k ranked unlabeled images from the database. Compared with query attacks in image classification, which produce adversaries according to the returned labels or confidence score, the challenge becomes even more prominent due to the difficulty in quantifying the attack effectiveness on the partial retrieved list. In this paper, we make the first attempt in Query-based Attack against Image Retrieval (QAIR), to completely subvert the top-k retrieval results. Specifically, a new relevance-based loss is designed to quantify the attack effects by measuring the set similarity on the top-k retrieval results before and after attacks and guide the gradient optimization. To further boost the attack efficiency, a recursive model stealing method is proposed to acquire transferable priors on the target model and generate the prior-guided gradients. Comprehensive experiments show that the proposed attack achieves a high attack success rate with few queries against the image retrieval systems under the black-box setting. The attack evaluations on the real-world visual search engine show that it successfully deceives a commercial system such as Bing Visual Search with 98% attack success rate by only 33 queries on average. Yuefeng Chen, Shaokai Ye, Yuan He 0011, Shuhui Wang, Hang Su 0006, Hui Xue 0001 |
CVPR | 6 |
| 2021 | Greedy Gradient Ensemble for Robust Visual Question AnsweringabstractLanguage bias is a critical issue in Visual Question Answering (VQA), where models often exploit dataset biases for the final decision without considering the image information. As a result, they suffer from performance drop on out-of-distribution data and inadequate visual explanation. Based on experimental analysis for existing robust VQA methods, we stress the language bias in VQA that comes from two aspects, i.e., distribution bias and shortcut bias. We further propose a new de-bias framework, Greedy Gradient Ensemble (GGE), which combines multiple biased models for unbiased base model learning. With the greedy strategy, GGE forces the biased models to over-fit the biased data distribution in priority, thus makes the base model pay more attention to examples that are hard to solve by biased models. The experiments demonstrate that our method makes better use of visual information and achieves state-of-the-art performance on diagnosing dataset VQACP without using extra annotations. Xinzhe Han, Shuhui Wang, Chi Su, Qingming Huang, Qi Tian 0001 |
ICCV | 2 |
| 2021 | Cascade Cross-modal Attention Network for Video Actor and Action Segmentation from a SentenceabstractIn this paper, we address the problem that selectively segments the actor and its action in the video clip given the sentence description. The main challenge is to match the local semantic features of the video with the heterogeneous textual features. A widely used language processing method in previous works is to leverage bi-LSTM and self-attention, which fixed the attention of the sentence and neglected the personality of the video, leading the attention of the sentence mismatch the most discriminative feature of the video. The proposed algorithm in this paper allows the sentence to learn the most discriminative features of the video, remarkably improving the accuracy of matching and segmentation. Specifically, we propose a cascade cross-modal attention to leverage two perspectives visual features to attend language from coarse to fine to generate the discriminative vision-aware language features. Moreover, equipping our framework with a contrastive learning method and a designed hard negative mining strategy benefits our proposed network from identifying the positive sample from numbers of negatives, and further improving the performance. To demonstrate the effectiveness of our approach, we conduct experiments on two datasets: A2D Sentences and J-HMDB Sentences. Experimental results show that our method significantly improves the performance over recent state-of-the-art methods. Weidong Chen 0013, Guorong Li, Xinfeng Zhang 0001, Hongyang Yu 0001, Shuhui Wang, Qingming Huang |
ACM Multimedia | 5 |
| 2021 | Multimodal Entity Linking: A New Dataset and A BaselineabstractIn this paper, we introduce a new Multimodal Entity Linking (MEL) task on the multimodal data. The MEL task discovers entities in multiple modalities and various forms within large-scale multimodal data and maps multimodal mentions in a document to entities in a structured knowledge base such as Wikipedia. Different from the conventional Neural Entity Linking (NEL) task that focuses on textual information solely, MEL aims at achieving human-level disambiguation among entities in images, texts, and knowledge bases. Due to the lack of sufficient labeled data for the MEL task, we release a large-scale multimodal entity linking dataset M3EL (abbreviated for MultiModal Movie Entity Linking). Specifically, we collect reviews and images of 1,100 movies, extract textual and visual mentions, and label them with entities registered in Wikipedia. In addition, we construct a new baseline method to solve the MEL problem, which models the alignment of textual and visual mentions as a bipartite graph matching problem and solves it with an optimal-transportation-based linking method. Extensive experiments on the M3EL dataset verify the quality of the dataset and the effectiveness of the proposed method. We envision this work to be helpful for soliciting more research effort and applications regarding multimodal computing and inference in the future. We make the dataset and the baseline algorithm publicly available at https://jingrug.github.io/research/M3EL. Jingru Gan, Jinchang Luo, Shuhui Wang, Qingming Huang |
ACM Multimedia | 4 |
| 2021 | Semi-Autoregressive Image CaptioningabstractCurrent state-of-the-art approaches for image captioning typically adopt an autoregressive manner, i.e., generating descriptions word by word, which suffers from slow decoding issue and becomes a bottleneck in real-time applications. Non-autoregressive image captioning with continuous iterative refinement, which eliminates the sequential dependence in a sentence generation, can achieve comparable performance to the autoregressive counterparts with a considerable acceleration. Nevertheless, based on a well-designed experiment, we empirically proved that iteration times can be effectively reduced when providing sufficient prior knowledge for the language decoder. Towards that end, we propose a novel two-stage framework, referred to as Semi-Autoregressive Image Captioning (SAIC), to make a better trade-off between performance and speed. The proposed SAIC model maintains autoregressive property in global but relieves it in local. Specifically, SAIC model first jumpily generates an intermittent sequence in an autoregressive manner, that is, it predicts the first word in every word group in order. Then, with the help of the partially deterministic prior information and image features, SAIC model non-autoregressively fills all the skipped words with one iteration. Experimental results on the MS COCO benchmark demonstrate that our SAIC model outperforms the preceding non-autoregressive image captioning models while obtaining a competitive inference speedup. Zhengcong Fei, Zekang Li, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 4 |
| 2021 | Mining Latent Structures for Multimedia RecommendationabstractMultimedia content is of predominance in the modern Web era. Investigating how users interact with multimodal items is a continuing concern within the rapid development of recommender systems. The majority of previous work focuses on modeling user-item interactions with multimodal features included as side information. However, this scheme is not well-designed for multimedia recommendation. Specifically, only collaborative item-item relationships are implicitly modeled through high-order item-user-item relations. Considering that items are associated with rich contents in multiple modalities, we argue that the latent semantic item-item structures underlying these multimodal contents could be beneficial for learning better item representations and further boosting recommendation. To this end, we propose a LATent sTructure mining method for multImodal reCommEndation, which we term LATTICE for brevity. To be specific, in the proposed LATTICE model, we devise a novel modality-aware structure learning layer, which learns item-item structures for each modality and aggregates multiple modalities to obtain latent item graphs. Based on the learned latent graphs, we perform graph convolutions to explicitly inject high-order item affinities into item representations. These enriched item representations can then be plugged into existing collaborative filtering methods to make more accurate recommendations. Extensive experiments on three real-world datasets demonstrate the superiority of our method over state-of-the-art multimedia recommendation methods and validate the efficacy of mining latent item-item relationships from multimodal features. Yanqiao Zhu 0001, Qiang Liu 0006, Shuhui Wang, Liang Wang 0001 |
ACM Multimedia | 5 |
| 2021 | Local-binarized very deep residual network for visual categorization
Xuejing Liu, Liang Li 0003, Shuhui Wang, Zhengjun Zha, Qingming Huang |
Neurocomputing | 3 |
| 2021 | Harmonized Multimodal Learning with Gaussian Process Latent Variable ModelsabstractMultimodal learning aims to discover the relationship between multiple modalities. It has become an important research topic due to extensive multimodal applications such as cross-modal retrieval. This paper attempts to address the modality heterogeneity problem based on Gaussian process latent variable models (GPLVMs) to represent multimodal data in a common space. Previous multimodal GPLVM extensions generally adopt individual learning schemes on latent representations and kernel hyperparameters, which ignore their intrinsic relationship. To exploit strong complementarity among different modalities and GPLVM components, we develop a novel learning scheme called Harmonization, where latent representations and kernel hyperparameters are jointly learned from each other. Beyond the correlation fitting or intra-modal structure preservation paradigms widely used in existing studies, the harmonization is derived in a model-driven manner to encourage the agreement between modality-specific GP kernels and the similarity of latent representations. We present a range of multimodal learning models by incorporating the harmonization mechanism into several representative GPLVM-based approaches. Experimental results on four benchmark datasets show that the proposed models outperform the strong baselines for cross-modal retrieval tasks, and that the harmonized multimodal learning method is superior in discovering semantically consistent latent representation. Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Pano-SfMLearner: Self-Supervised Multi-Task Learning of Depth and Semantics in Panoramic VideosabstractWith the advent of virtual reality and augment reality applications, omnidirectional imaging and$360^{\circ }$cameras become increasingly popular in many scenarios such as entertainment and autonomous systems. In this paper, we propose a self-supervised framework for multi-task learning on depth, camera motion and semantics from panoramic videos. Specifically, our method is based on differentiable warping of adjacent views to the target. Two improvements are provided. First, we introduce a view synthesis module based on equirectangular projection to enable direct optimization on panoramic images. Second, we introduce a self-supervised segmentation branch to involve the constraint of semantic consistency for further improvement. Extensive experiments on two$360^{\circ }$video and two$360^{\circ }$image datasets demonstrate that our method outperforms the state-of-the-art and achieves favorable cross-modality performance. Shuhui Wang, Yulan Guo, Yuan He 0011, Hui Xue 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Learning Feature Representation and Partial Correlation for Multimodal Multi-Label DataabstractUser-provided annotations in existing multimodal datasets sometimes are inappropriate for model learning and can hinder the task of cross-modal retrieval. To handle this issue, we propose a discriminative and noise-robust cross-modal retrieval method, called FLPCL, which consists of deep feature learning and partial correlation learning. Deep feature learning is implemented by utilizing label supervised information to guide the training of deep neural network for each modality, which aims to find modality-specific deep feature representations that preserve the similarity and discrimination information among multimodal data. Based on deep feature learning, partial correlation learning is proposed to infer direct association between different modalities by removing the effect of common underlying semantics from each modality. It is achieved by maximizing the canonical correlation of the feature representations of different modalities conditioned on the label modality. Different from existing works that build indirect association between modalities via incorporating semantic labels, our FLPCL method can learn more effective and robust multimodal latent representations by explicitly preserving both intra-modal and inter-modal relationship among multimodal data. Extensive experiments on three cross-modal datasets show that our method outperforms state-of-the-art methods on cross-modal retrieval tasks. Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Augmented Adversarial Training for Cross-Modal RetrievalabstractCross-modal retrieval has received considerable attention in recent years. The core of cross-modal retrieval is to find a representation space to align data from different modalities according to their semantics. In this paper, we propose a cross-modal retrieval method that aligns data from different modalities by transferring one source modality to another target modality with augmented adversarial training. To preserve the semantic meaning in the modality transfer process, we employ the idea of conditional GANs and augment it. The key idea is to incorporate semantic information from the label space into the adversarial training process by sampling more semantic relevant and irrelevant source-target sample pairs. The augmented sample pairs improve the alignment from two aspects. First, relevant source-target sample pairs provide more training samples, leading to a better guidance of the alignment of fake targets and true paired targets. Second, relevant and irrelevant source-target sample pairs teach the discriminator to better distinguish true relevant pairs from fake relevant pairs, which guides the generator to better transfer from the source modality to the target modality. Extensive experiments compared with state-of-the-art methods show the promising power of our approach. Yiling Wu, Shuhui Wang, Guoli Song, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2021 | Graph Regularized Encoder-Decoder Networks for Image Representation LearningabstractImage representation learning with encoder-decoder networks plays a fundamental role in multimedia processing. Recent findings show that traditional encoder-decoders can be negatively affected by small visual perturbations. The learned non-smooth feature embedding cannot guarantee to capture semantic-meaningful geometric distance between visually-similar image samples. Inspired by manifold learning, we propose a graph regularized encoder-decoder network, which can preserve local geometric information of the code embedding space. More discriminative feature embedding is learnt to attain both high-level image semantic and neighbor relationship of image clusters. The proposed graph regularizer is formulated upon multi-layer perceptions. It uses the local invariance principle to explicitly reconstruct the geometric similarity graph. Theoretical analysis is provided to show the connection between our deep regularizer and traditional graph Laplacian regularizer. Practically, the network complexity is alleviated by anchor based bipartite graph, and this leverages our method into large scale scenario. Experimental evaluations show the comparable results of the proposed method with state-of-the-art models on different tasks. Liang Li 0003, Shuhui Wang, Weigang Zhang, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | String Prediction for 4: 2: 0 Format Screen Content Coding and Its Implementation in AVS3abstractIn the past, string prediction (also known as string matching) was applied only to RGB and YUV 4:4:4 format screen content coding. This paper proposes a string prediction approach to 4:2:0 format screen content coding implemented in the third generation of Audio Video Standard (AVS3) in China. String prediction is applied to both YUV CU and Y CU. To further improve the coding performance, several improved technicals of string prediction are presented, including a mixed string searching strategy for finding the optimal reference string, a joint picture-level, CU-level, and pixel-level early termination strategy to reduce coding complexity, and two effective coding methods for string prediction parameters. For low-complexity hardware implementation of string prediction decoder, the memory access bandwidth is reduced by introducing string constraints. Meanwhile, string prediction reuses the reference pixel buffer of intra block copy (IBC). Compared with the newest AVS3 reference software HPM7.0 with string prediction disabled, the proposed string prediction approach achieves up to 18.48% Y BD-rate reduction. Using AVS3 Screen Content Coding (SCC) Common Test Condition and YUV test sequences in Text and Graphics with Motion category, the proposed technique achieves an average Y BD-rate reduction of 10.33%, 8.47%, 6.91% for All Intra (AI), Random Access (RA) and Low Delay (LD) configurations, respectively, with low additional encoding and decoding complexity. The proposed string prediction approach has been adopted in the newest AVS3 reference software HPM7.0. Qingyang Zhou, Liping Zhao 0005, Kailun Zhou, Tao Lin 0005, Shuhui Wang, Mengcao Jiao |
IEEE Trans. Multim. | 6 |
| 2020 | F³Net: Fusion, Feedback and Focus for Salient Object DetectionabstractMost of existing salient object detection models have achieved great progress by aggregating multi-level features extracted from convolutional neural networks. However, because of the different receptive fields of different convolutional layers, there exists big differences between features generated by these layers. Common feature fusion strategies (addition or concatenation) ignore these differences and may cause suboptimal solutions. In this paper, we propose the F3Net to solve above problem, which mainly consists of cross feature module (CFM) and cascaded feedback decoder (CFD) trained by minimizing a new pixel position aware loss (PPA). Specifically, CFM aims to selectively aggregate multi-level features. Different from addition and concatenation, CFM adaptively selects complementary components from input features before fusion, which can effectively avoid introducing too much redundant information that may destroy the original features. Besides, CFD adopts a multi-stage feedback mechanism, where features closed to supervision will be introduced to the output of previous layers to supplement them and eliminate the differences between features. These refined features will go through multiple similar iterations before generating the final saliency maps. Furthermore, different from binary cross entropy, the proposed PPA loss doesn't treat pixels equally, which can synthesize the local structure information of a pixel to guide the network to focus more on local details. Hard pixels from boundaries or error-prone parts will be given more attention to emphasize their importance. F3Net is able to segment salient object regions accurately and provide clear local details. Comprehensive experiments on five benchmark datasets demonstrate that F3Net outperforms state-of-the-art approaches on six evaluation metrics. Code will be released at https://github.com/weijun88/F3Net. Jun Wei 0006, Shuhui Wang, Qingming Huang |
AAAI | 2 |
| 2020 | Towards Discriminability and Diversity: Batch Nuclear-Norm Maximization Under Label Insufficient SituationsabstractThe learning of the deep networks largely relies on the data with human-annotated labels. In some label insufficient situations, the performance degrades on the decision boundary with high data density. A common solution is to directly minimize the Shannon Entropy, but the side effect caused by entropy minimization, \it i.e., reduction of the prediction diversity, is mostly ignored. To address this issue, we reinvestigate the structure of classification output matrix of a randomly selected data batch. We find by theoretical analysis that the prediction discriminability and diversity could be separately measured by the Frobenius-norm and rank of the batch output matrix. Besides, the nuclear-norm is an upperbound of the Frobenius-norm, and a convex approximation of the matrix rank. Accordingly, to improve both discriminability and diversity, we propose Batch Nuclear-norm Maximization (BNM) on the output matrix. BNM could boost the learning under typical label insufficient learning scenarios, such as semi-supervised learning, domain adaptation and open domain recognition. On these tasks, extensive experimental results show that BNM outperforms competitors and works well with existing well-known methods. The code is available at https://github.com/cuishuhao/BNM. Shuhao Cui, Shuhui Wang, Junbao Zhuo, Liang Li 0003, Qingming Huang, Qi Tian 0001 |
CVPR | 2 |
| 2020 | Gradually Vanishing Bridge for Adversarial Domain AdaptationabstractIn unsupervised domain adaptation, rich domain-specific characteristics bring great challenge to learn domain-invariant representations. However, domain discrepancy is considered to be directly minimized in existing solutions, which is difficult to achieve in practice. Some methods alleviate the difficulty by explicitly modeling domain-invariant and domain-specific parts in the representations, but the adverse influence of the explicit construction lies in the residual domain-specific characteristics in the constructed domain-invariant representations. In this paper, we equip adversarial domain adaptation with Gradually Vanishing Bridge (GVB) mechanism on both generator and discriminator. On the generator, GVB could not only reduce the overall transfer difficulty, but also reduce the influence of the residual domain-specific characteristics in domain-invariant representations. On the discriminator, GVB contributes to enhance the discriminating ability, and balance the adversarial training process. Experiments on three challenging datasets show that our GVB methods outperform strong competitors, and cooperate well with other adversarial methods. The code is available at https://github.com/cuishuhao/GVB. Shuhao Cui, Shuhui Wang, Junbao Zhuo, Chi Su, Qingming Huang, Qi Tian 0001 |
CVPR | 2 |
| 2020 | Parsing-Based View-Aware Embedding Network for Vehicle Re-IdentificationabstractVehicle Re-Identification is to find images of the same vehicle from various views in the cross-camera scenario. The main challenges of this task are the large intra-instance distance caused by different views and the subtle inter-instance discrepancy caused by similar vehicles. In this paper, we propose a parsing-based view-aware embedding network (PVEN) to achieve the view-aware feature alignment and enhancement for vehicle ReID. First, we introduce a parsing network to parse a vehicle into four different views and then align the features by mask average pooling. Such alignment provides a fine-grained representation of the vehicle. Second, in order to enhance the view-aware features, we design a common-visible attention to focus on the common visible views, which not only shortens the distance among intra-instances, but also enlarges the discrepancy of inter-instances. The PVEN helps capture the stable discriminative information of vehicle under different views. The experiments conducted on three datasets show that our model outperforms state-of-the-art methods by a large margin. Dechao Meng, Liang Li 0003, Xuejing Liu, Zhengjun Zha, Xingyu Gao 0001, Shuhui Wang, Qingming Huang |
CVPR | 8 |
| 2020 | Label Decoupling Framework for Salient Object DetectionabstractTo get more accurate saliency maps, recent methods mainly focus on aggregating multi-level features from fully convolutional network (FCN) and introducing edge information as auxiliary supervision. Though remarkable progress has been achieved, we observe that the closer the pixel is to the edge, the more difficult it is to be predicted, because edge pixels have a very imbalance distribution. To address this problem, we propose a label decoupling framework (LDF) which consists of a label decoupling (LD) procedure and a feature interaction network (FIN). LD explicitly decomposes the original saliency map into body map and detail map, where body map concentrates on center areas of objects and detail map focuses on regions around edges. Detail map works better because it involves much more pixels than traditional edge supervision. Different from saliency map, body map discards edge pixels and only pays attention to center areas. This successfully avoids the distraction from edge pixels during training. Therefore, we employ two branches in FIN to deal with body map and detail map respectively. Feature interaction (FI) is designed to fuse the two complementary branches to predict the saliency map, which is then used to refine the two branches again. This iterative refinement is helpful for learning better representations and more precise saliency maps. Comprehensive experiments on six benchmark datasets demonstrate that LDF outperforms state-of-the-art approaches on different evaluation metrics. Jun Wei 0006, Shuhui Wang, Zhe Wu 0006, Chi Su, Qingming Huang, Qi Tian 0001 |
CVPR | 2 |
| 2020 | State-Relabeling Adversarial Active LearningabstractActive learning is to design label-efficient algorithms by sampling the most representative samples to be labeled by an oracle. In this paper, we propose a state relabeling adversarial active learning model (SRAAL), that leverages both the annotation and the labeled/unlabeled state information for deriving the most informative unlabeled samples. The SRAAL consists of a representation generator and a state discriminator. The generator uses the complementary annotation information with traditional reconstruction information to generate the unified representation of samples, which embeds the semantic into the whole data representation. Then, we design an online uncertainty indicator in the discriminator, which endues unlabeled samples with different importance. As a result, we can select the most informative samples based on the discriminator's predicted state. We also design an algorithm to initialize the labeled pool, which makes subsequent sampling more efficient. The experiments conducted on various datasets show that our model outperforms the previous state-of-art active learning methods and our initially sampling algorithm achieves better performance. Beichen Zhang 0006, Liang Li 0003, Shuhui Wang, Zhengjun Zha, Qingming Huang |
CVPR | 4 |
| 2020 | Interpretable Visual Reasoning via Probabilistic Formulation Under Natural Supervision
Xinzhe Han, Shuhui Wang, Chi Su, Weigang Zhang, Qingming Huang, Qi Tian 0001 |
ECCV (9) | 2 |
| 2020 | A Structured Latent Variable Recurrent Network With Stochastic Attention For Generating Weibo CommentsabstractBuilding intelligent agents to generate realistic Weibo comments is challenging. For such realistic Weibo comments, the key criterion is improving diversity while maintaining coherency. Considering that the variability of linguistic comments arises from multi-level sources, including both discourse-level properties and word-level selections, we improve the comment diversity by leveraging such inherent hierarchy. In this paper, we propose a structured latent variable recurrent network, which exploits the hierarchical-structured latent variables with stochastic attention to model the variations of comments. First, we endow both discourse-level and word-level latent variables with hierarchical and temporal dependencies for constructing multi-level hierarchy. Second, we introduce a stochastic attention to infer the key-words of interest in the input post. As a result, diverse comments can be generated with both discourse-level properties and local-word selections. Experiments on open-domain Weibo data show that our model generates more diverse and realistic comments. Liang Li 0003, Shuhui Wang, Weigang Zhang, Qingming Huang, Qi Tian 0001 |
IJCAI | 3 |
| 2020 | Sharp Multiple Instance Learning for DeepFake Video DetectionabstractWith the rapid development of facial manipulation techniques, face forgery has received considerable attention in multimedia and computer vision community due to security concerns. Existing methods are mostly designed for single-frame detection trained with precise image-level labels or for video-level prediction by only modeling the inter-frame inconsistency, leaving potential high risks for DeepFake attackers. In this paper, we introduce a new problem of partial face attack in DeepFake video, where only video-level labels are provided but not all the faces in the fake videos are manipulated. We address this problem by multiple instance learning framework, treating faces and input video as instances and bag respectively. A sharp MIL (S-MIL) is proposed which builds direct mapping from instance embeddings to bag prediction, rather than from instance embeddings to instance prediction and then to bag prediction in traditional MIL. Theoretical analysis proves that the gradient vanishing in traditional MIL is relieved in S-MIL. To generate instances that can accurately incorporate the partially manipulated faces, spatial-temporal encoded instance is designed to fully model the intra-frame and inter-frame inconsistency, which further helps to promote the detection performance. We also construct a new dataset FFPMS for partially attacked DeepFake video detection, which can benefit the evaluation of different methods at both frame and video levels. Experiments on FFPMS and the widely used DFDC dataset verify that S-MIL is superior to other counterparts for partially attacked DeepFake video detection. In addition, S-MIL can also be adapted to traditional DeepFake image detection tasks and achieve state-of-the-art performance on single-frame datasets. Yining Lang, Yuefeng Chen, Xiaofeng Mao, Yuan He 0011, Shuhui Wang, Hui Xue 0001 |
ACM Multimedia | 6 |
| 2020 | Diverter-Guider Recurrent Network for Diverse Poems Generation from ImageabstractPoem generation from image aims to automatically generate the poetic sentences for presenting the image content or overtone. Previous works focused on 1-to-1 image-poem generation with the demands of poeticness and content relevance. This paper proposes the paradigm of multiple poems generation from one image, which is closer to human poetizing but more challenging. Its key problem is to simultaneously guarantee the diversity of multiple poems with poeticness and relevance. To this end, we propose an end-to-end probabilistic Diverter-Guider Recurrent Network (DG-Net), which is a context-based encoder-decoder generative model with the hierarchical stochastic variables. Specifically, the diverter-variable represents the decoding-context inferred from the input image to diversify the poem themes; the guider-variable is introduced as an attribute decoder to restricts the word-choice with supervised information. Extensive experiments on automatic evaluations and human judgments demonstrate the superior performance of DG-Net than existing poem generation methods. Qualitative study show that our model can generate diverse poems with the poeticness and relevance. Liang Li 0003, Li Su 0003, Shuhui Wang, Chenggang Yan 0001, Zhengjun Zha, Qingming Huang |
ACM Multimedia | 4 |
| 2020 | IR-GAN: Image Manipulation with Linguistic Instruction by Increment ReasoningabstractConditional image generation is an active research topic including text2image and image translation. Recently image manipulation with linguistic instruction brings new challenges of multimodal conditional generation. However, traditional conditional image generation models mainly focus on generating high-quality and visually realistic images, and lack resolving the partial consistency between image and instruction. To address this issue, we propose an Increment Reasoning Generative Adversarial Network (IR-GAN), which aims to reason the consistency between visual increment in images and semantic increment in instructions. First, we introduce the word-level and instruction-level instruction encoders to learn user's intention from history-correlated instructions as semantic increment. Second, we embed the representation of semantic increment into that of source image for generating target image, where source image plays the role of referring auxiliary. Finally, we propose a reasoning discriminator to measure the consistency between visual increment and semantic increment, which purifies user's intention and guarantees the good logic of generated target image. Extensive experiments and visualization conducted on two datasets show the effectiveness of IR-GAN. Zhenhuan Liu, Jincan Deng, Liang Li 0003, Shaofei Cai, Qianqian Xu 0001, Shuhui Wang, Qingming Huang |
ACM Multimedia | 6 |
| 2020 | Transferrable Referring Expression Grounding with Concept Transfer and Context InheritanceabstractReferring Expression Grounding (REG) aims at localizing a particular object in an image according to a language expression. Recent REG methods have achieved promising performance, but most of them are constrained to limited object categories due to the scale of current REG datasets. In this paper, we explore REG in a new scenario, where the REG model can ground novel objects out of REG training data. With this motivation, we propose a Concept-Context Disentangled network (CCD) which transfers concepts from auxiliary classification data with new categories meanwhile inherits context from REG data to ground new objects. Specially, we design a subject encoder to learn a cross-modal common semantic space, which can bridge the semantic and domain gap between auxiliary classification data and REG data. This common space guarantees CCD can transfer and recognize novel categories. Further, we learn the correspondence between image proposal and referring expression upon location and relationship. Benefiting from the disentangled structure, the context is relatively independent of the subject, so it can be better inherited from the REG training data. Finally, a language attention is learned to adaptively assign different importance to subject and context for grounding target objects. Experiments on four REG datasets show our method outperforms the compared approach on the new-category test datasets. Xuejing Liu, Liang Li 0003, Shuhui Wang, Zhengjun Zha, Dechao Meng, Qingming Huang |
ACM Multimedia | 3 |
| 2020 | Fine-grained Feature Alignment with Part Perspective Transformation for Vehicle ReIDabstractGiven a query image, vehicle Re-Identification is to search the same vehicle in multi-camera scenarios, which are attracting much attention in recent years. However, vehicle ReID severely suffers from the perspective variation problem. For different vehicles with similar color and type which are taken from different perspectives, all visual patterns are misaligned and warped, which is hard for the model to find out the exact discriminative regions. In this paper, we propose part perspective transformation module (PPT) to map the different parts of vehicle into a unified perspective respectively. The PPT disentangles the vehicle features of different perspectives and then aligns them in a fine-grained level. Further, we propose a dynamically batch hard triplet loss to select the common visible regions of the compared vehicles. Our approach helps the model to generate the perspective invariant features and find out the exact distinguishable regions for vehicle ReID. Extensive experiments on three standard vehicle ReID datasets show the effectiveness of our method. Dechao Meng, Liang Li 0003, Shuhui Wang, Xingyu Gao 0001, Zhengjun Zha, Qingming Huang |
ACM Multimedia | 3 |
| 2020 | Towards More Explainability: Concept Knowledge Mining Network for Event RecognitionabstractEvent recognition of untrimmed video is a challenging task due to the big gap between low level visual features and event semantics. Beyond feature learning via deep neural networks, some recent works focus on analyzing event videos using concept-based representation. However, these methods simply aggregate the concept representation vectors of frames or segments, which inevitably introduces information loss on video-level concept knowledge. Moreover, the diversified relation between different concept domains (e.g., scene, object and action) has not been fully explored. To address the above issues, we propose a concept knowledge mining network (CKMN) for event recognition. CKMN is composed of an intra-domain concept knowledge mining subnetwork (IaCKM) and an inter-domain concept knowledge mining subnetwork~(IrCKM). IaCKM aims to obtain a complete concept representation by mining the existing pattern of each concept at different time granularities with dilated temporal pyramid convolution and temporal self-attention, while IrCKM explores the interaction between different types of concepts with co-attention style learning. We evaluate our method on FCVID and ActivityNet datasets. Experimental results show the effectiveness and better interpretability of our model on event analytics. Code is available at https://github.com/qzhb/CKMN. Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2020 | Modeling Temporal Concept Receptive Field Dynamically for Untrimmed Video AnalysisabstractEvent analysis in untrimmed videos has attracted increasing attention due to the application of cutting-edge techniques such as CNN. As a well studied property for CNN-based models, the receptive field is a measurement for measuring the spatial range covered by a single feature response, which is crucial in improving the image categorization accuracy. In video domain, video event semantics are actually described by complex interaction among different concepts, while their behaviors vary drastically from one video to another, leading to the difficulty in concept-based analytics for accurate event categorization. To model the concept behavior, we study temporal concept receptive field of concept-based event representation, which encodes the temporal occurrence pattern of different mid-level concepts. Accordingly, we introduce temporal dynamic convolution (TDC) to give stronger flexibility to concept-based event analytics. TDC can adjust the temporal concept receptive field size dynamically according to different inputs. Notably, a set of coefficients are learned to fuse the results of multiple convolutions with different kernel widths that provide various temporal concept receptive field sizes. Different coefficients can generate appropriate and accurate temporal concept receptive field size according to input videos and highlight crucial concepts. Based on TDC, we propose the temporal dynamic concept modeling network~(TDCMN) to learn an accurate and complete concept representation for efficient untrimmed video analysis. Experiment results on FCVID and ActivityNet show that TDCMN demonstrates adaptive event recognition ability conditioned on different inputs, and improve the event recognition performance of Concept-based methods by a large margin. Code is available at https://github.com/qzhb/TDCMN. Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Weigang Zhang, Qingming Huang |
ACM Multimedia | 2 |
| 2020 | Structural Semantic Adversarial Active Learning for Image CaptioningabstractMost image captioning models achieve superior performances with the help of large-scale surprised training data, but it is prohibitively costly to label the image captions. To solve this problem, we propose a structural semantic adversarial active learning (SSAAL) model that leverages both visual and textual information for deriving the most representative samples while maximizing the image captioning performance. SSAAL consists of a semantic constructor, a snapshot& caption (SC) supervisor, and a labeled/unlabeled state discriminator. The constructor is designed to generate a structural semantic representation describing the objects, attributes and object relationships in the image. The SC supervisor is proposed to supervise this representation at the word-level and sentence-level in a multi-task learning manner, which directly relates the representation to ground-truth captions and updates it in the caption generating process. Finally, we introduce a state discriminator to predict the sample state and select images with sufficient semantic and fine-grained diversity. Extensive experiments on standard captioning dataset show that our model outperforms other active learning methods and achieves a competitive performance even though selecting a small amount of samples. Beichen Zhang 0006, Liang Li 0003, Li Su 0003, Shuhui Wang, Jincan Deng, Zhengjun Zha, Qingming Huang |
ACM Multimedia | 4 |
| 2020 | Heuristic Domain AdaptationabstractIn visual domain adaptation (DA), separating the domain-specific characteristics from the domain-invariant representations is an ill-posed problem. Existing methods apply different kinds of priors or directly minimize the domain discrepancy to address this problem, which lack flexibility in handling real-world situations. Another research pipeline expresses the domain-specific information as a gradual transferring process, which tends to be suboptimal in accurately removing the domain-specific properties. In this paper, we address the modeling of domain-invariant and domain-specific information from the heuristic search perspective. We identify the characteristics in the existing representations that lead to larger domain discrepancy as the heuristic representations. With the guidance of heuristic representations, we formulate a principled framework of Heuristic Domain Adaptation (HDA) with well-founded theoretical guarantees. To perform HDA, the cosine similarity scores and independence measurements between domain-invariant and domain-specific representations are cast into the constraints at the initial and final states during the learning procedure. Similar to the final condition of heuristic search, we further derive a constraint enforcing the final range of heuristic network output to be small. Accordingly, we propose Heuristic Domain Adaptation Network (HDAN), which explicitly learns the domain-invariant and domain-specific representations with the above mentioned constraints. Extensive experiments show that HDAN has exceeded state-of-the-art on unsupervised DA, multi-source DA and semi-supervised DA. The code is available at https://github.com/cuishuhao/HDA. Shuhao Cui, Xuan Jin, Shuhui Wang, Yuan He 0011, Qingming Huang |
NeurIPS | 3 |
| 2020 | Two-stream deep sparse network for accurate and efficient image restoration
Shuhui Wang, Liang Li 0003, Weigang Zhang, Qingming Huang |
Comput. Vis. Image Underst. | 1 |
| 2020 | A minimum entropy deconvolution-enhanced convolutional neural networks for fault diagnosis of axial piston pumps
Shuhui Wang, Jiawei Xiang |
Soft Comput. | 1 |
| 2020 | Textual-Visual Reference-Aware Attention Network for Visual DialogabstractVisual dialog is a challenging task in multimedia understanding, which requires the dialog agent to answer a series of questions that are based on an input image. The critical issue to produce an exact answer is how to model the mutual semantic interaction among feature representations of the image, question-answer history, and current question. In this study, we propose a textual-visual Reference-Aware Attention Network (RAA-Net), which aims to effectively fuse Q (question), H (history), Vl (local vision), and Vg (global vision) to infer the exact answer. In the multimodal feature flows, RAA-Net first learns the textual context through multi-head attention between Q and H and then guides the textual reference semantics to the image to capture visual reference semantics by self-and cross-reference-aware attention in and between Vl and Vg. In the proposed RAA-Net, we exploit the two-stage (intraand inter-) visual reasoning mechanism on Vl and Vg. Extensive experiments on the VisDial v0.9 and v1.0 datasets show that RAA-Net achieves state-of-the-art performance. Visualization results on both visual and textual attention maps further validate the remarkable interpretability achieved by our solution. Dan Guo 0001, Hui Wang 0079, Shuhui Wang, Meng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Online Fast Adaptive Low-Rank Similarity Learning for Cross-Modal RetrievalabstractThe semantic similarity among cross-modal data objects, e.g., similarities between images and texts, are recognized as the bottleneck of cross-modal retrieval. However, existing batch-style correlation learning methods suffer from prohibitive time complexity and extra memory consumption in handling large-scale high dimensional cross-modal data. In this paper, we propose a Cross-Modal Online Low-Rank Similarity function learning (CMOLRS) method, which learns a low-rank bilinear similarity measurement for cross-modal retrieval. We model the cross-modal relations by relative similarities on the training data triplets and formulate the relative relations as convex hinge loss. By adapting the margin in hinge loss with pair-wise distances in feature space and label space, CMOLRS effectively captures the multi-level semantic correlation and adapts to the content divergence among cross-modal data. Imposed with a low-rank constraint, the similarity function is trained by online learning in the manifold of low-rank matrices. The low-rank constraint not only endows the model learning process with faster speed and better scalability, but also improves the model generality. We further propose fast-CMOLRS combining multiple triplets for each query instead of standard process using single triplet at each model update step, which further reduces the times of gradient updates and retractions. Extensive experiments are conducted on four public datasets, and comparisons with state-of-the-art methods show the effectiveness and efficiency of our approach. Yiling Wu, Shuhui Wang, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2020 | An Ultra-Low Complexity and High Efficiency Approach for Lossless Alpha Channel CodingabstractAlpha channel is being applied in an increasing number of mobile web applications on mobile devices that require ultra-low power consumption in all cases including compute-intensive video encoding and decoding. Thus, we propose an ultra-low coding complexity and high efficiency alpha channel lossless coding approach. A novel coding framework and four new coding schemes are proposed for alpha channel coding. The framework fuses a string matching technique and a proposed prediction coding scheme named bit-depth preserving prediction (BDPP) together to reduce the correlations within and between repeated identical patterns and neighboring pixels. To achieve a good tradeoff between complexity and efficiency, either the unmatchable bytes are coded directly or the BDPP residuals of unmatchable bytes are coded by a proposed bytewise entropy coding scheme named 0.5-1-2byte-size-code. The other string matching parameters are coded by another proposed bytewise entropy coding scheme named byte-size multi-variable-length-code. To speed up the string-matching search, we apply a fast string search scheme that combines special position search and hash-based search. For the selected typical 236 alpha test images, compared with x265 in the fastest configuration and lossless mode, the proposed lossless approach achieves 14.33% less total compressed bytes with only 2.75% encoding and 1.83% decoding runtime. The proposed approach also outperforms the conventional lossless coding techniques such as LZ4HC, ZLIB, and PNG. Liping Zhao 0005, Tao Lin 0005, Kailun Zhou, Shuhui Wang |
IEEE Trans. Multim. | 5 |
| 2019 | Unsupervised Open Domain Recognition by Semantic Discrepancy MinimizationabstractWe address the unsupervised open domain recognition (UODR) problem, where categories in labeled source domain S is only a subset of those in unlabeled target domain T. The task is to correctly classify all samples in T including known and unknown categories. UODR is challenging due to the domain discrepancy, which becomes even harder to bridge when a large number of unknown categories exist in T. Moreover, the classification rules propagated by graph CNN (GCN) may be distracted by unknown categories and lack generalization capability. To measure the domain discrepancy for asymmetric label space between S and T, we propose Semantic-Guided Matching Discrepancy (SGMD), which first employs instance matching between S and T, and then the discrepancy is measured by a weighted feature distance between matched instances. We further design a limited balance constraint to achieve a more balanced classification output on known and unknown categories. We develop Unsupervised Open Domain Transfer Network (UODTN), which learns both the backbone classification network and GCN jointly by reducing the SGMD, enforcing the limited balance constraint and minimizing the classification loss on S. UODTN better preserves the semantic structure and enforces the consistency between the learned domain invariant visual features and the semantic embeddings. Experimental results show superiority of our method on recognizing images of both known and unknown categories. Junbao Zhuo, Shuhui Wang, Shuhao Cui, Qingming Huang |
CVPR | 2 |
| 2019 | Adaptive Reconstruction Network for Weakly Supervised Referring Expression GroundingabstractWeakly supervised referring expression grounding aims at localizing the referential object in an image according to the linguistic query, where the mapping between the referential object and query is unknown in the training stage. To address this problem, we propose a novel end-to-end adaptive reconstruction network (ARN). It builds the correspondence between image region proposal and query in an adaptive manner: adaptive grounding and collaborative reconstruction. Specifically, we first extract the subject, location and context features to represent the proposals and the query respectively. Then, we design the adaptive grounding module to compute the matching score between each proposal and query by a hierarchical attention model. Finally, based on attention score and proposal features, we reconstruct the input query with a collaborative loss of language reconstruction loss, adaptive reconstruction loss, and attribute classification loss. This adaptive mechanism helps our model to alleviate the variance of different referring expressions. Experiments on four large-scale datasets show ARN outperforms existing state-of-the-art methods by a large margin. Qualitative results demonstrate that the proposed ARN can better handle the situation where multiple objects of a particular category situated together. Xuejing Liu, Liang Li 0003, Shuhui Wang, Zhengjun Zha, Dechao Meng, Qingming Huang |
ICCV | 3 |
| 2019 | Knowledge-guided Pairwise Reconstruction Network for Weakly Supervised Referring Expression GroundingabstractWeakly supervised referring expression grounding (REG) aims at localizing the referential entity in an image according to linguistic query, where the mapping between the image region (proposal) and the query is unknown in the training stage. In referring expressions, people usually describe a target entity in terms of its relationship with other contextual entities as well as visual attributes. However, previous weakly supervised REG methods rarely pay attention to the relationship between the entities. In this paper, we propose a knowledge-guided pairwise reconstruction network (KPRN), which models the relationship between the target entity (subject) and contextual entity (object) as well as grounds these two entities. Specifically, we first design a knowledge extraction module to guide the proposal selection of subject and object. The prior knowledge is obtained in a specific form of semantic similarities between each proposal and the subject/object. Second, guided by such knowledge, we design the subject and object attention module to construct the subject-object proposal pairs. The subject attention excludes the unrelated proposals from the candidate proposals. The object attention selects the most suitable proposal as the contextual proposal. Third, we introduce a pairwise attention and an adaptive weighting scheme to learn the correspondence between these proposal pairs and the query. Finally, a pairwise reconstruction module is used to measure the grounding for weakly supervised learning. Extensive experiments on four large-scale datasets show our method outperforms existing state-of-the-art methods by a large margin. Xuejing Liu, Liang Li 0003, Shuhui Wang, Zhengjun Zha, Li Su 0003, Qingming Huang |
ACM Multimedia | 3 |
| 2019 | Learning Fragment Self-Attention Embeddings for Image-Text MatchingabstractIn image-text matching task, the key to good matching quality is to capture the rich contextual dependencies between fragments of image and text. However, previous works either simply aggregate the similarity of all possible pairs of image regions and words, or take multi-step cross attention to attend to image regions and words with each other as context, which requires exhaustive similarity computation between all image region and word pairs. In this paper, we propose Self-Attention Embeddings (SAEM) to exploit fragment relations in images or texts by self-attention mechanism, and aggregate fragment information into visual and textual embeddings. Specifically, SAEM extracts salient image regions based on bottom-up attention, and takes WordPiece tokens as sentence fragments. The self-attention layers are built to model subtle and fine-grained fragment relation in image and text respectively, which consists of multi-head self-attention sub-layer and position-wise feed-forward network sub-layer. Consequently, the fragment self-attention mechanism can discover the fragment relations and identify the semantically salient regions in images or words in sentences, and capture their interaction more accurately. By simultaneously exploiting the fine-grained fragment relation in both visual and textual modalities, our method produces more semantically consistent embeddings for representing images and texts, and demonstrates promising image-text matching accuracy and high efficiency on Flickr30K and MSCOCO datasets. Yiling Wu, Shuhui Wang, Guoli Song, Qingming Huang |
ACM Multimedia | 2 |
| 2019 | Structured Stochastic Recurrent Network for Linguistic Video PredictionabstractIntelligent machines are expected to have the capability of predicting impending occurrences. Inspired by video frame prediction and video captioning, we introduce a new task of Linguistic Video Prediction (LVP), which aims to predict the forthcoming events based on past video content and generate corresponding linguistic descriptions. Different from traditional video captioning that describes one specifically happened event, LVP is an open task involving one-to-many mappings between past and future. It explores different visual clues and associates them with potential events to generate corresponding descriptions. To address this task, we propose an end-to-end probabilistic approach named structured stochastic recurrent network (SRN) to characterize the one-to-many connections between past visual clues and possible future events. Specially, we first propose hierarchical-structured latent variables to represent the choice of event theme. Second, we introduce a stochastic attention module to capture the variations of the focused visual clues. Given a video, our model is able to generate multiple linguistic predictions by focusing on different event themes and visual clues. Experiments on ActivityNet dataset showed that the proposed model not only yields more informative predictions measured by BLEU, METEOR, ROUGE-L, CIDEr and SPICE scores, but also generates significantly more diverse predictions with higher recall rates to correctly hit the ground-truth. Liang Li 0003, Shuhui Wang, Dechao Meng, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2019 | Active Perception Network for Salient Object DetectionabstractTo get better saliency maps for salient object detection, recent methods fuse features from different levels of convolutional neural networks and have achieved remarkable progress. However, the differences between different feature levels bring difficulties to the fusion process, thus it may lead to unsatisfactory saliency predictions. To address this issue, we propose Active Perception Network (APN) to enhance inter-feature consistency for salient object detection. First, Mutual Projection Module (MPM) is developed to fuse different features, which uses high-level features as guided information to extract complementary components from low-level features, and can suppress background noises and improve semantic consistency. Self Projection Module (SPM) is designed to further refine the fused features, which can be considered as the extended version of residual connection. Features that pass through SPM can produce more accurate saliency maps. Finally, we propose Head Projection Module (HPM) to aggregate global information, which brings strong semantic consistency to the whole network. Comprehensive experiments on five benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches on different evaluation metrics. Jun Wei 0006, Shuhui Wang, Liang Li 0003, Qingming Huang |
MMAsia | 2 |
| 2019 | Regularized topic-aware latent influence propagation in dynamic relational networks
Shuhui Wang, Liang Li 0003, Chenxue Yang, Qingming Huang |
GeoInformatica | 1 |
| 2019 | Multi-modal semantic autoencoder for cross-modal retrieval
Yiling Wu, Shuhui Wang, Qingming Huang |
Neurocomputing | 2 |
| 2019 | Beyond global fusion: A group-aware fusion approach for multi-view image clustering
Zhe Xue, Guorong Li, Shuhui Wang, Jun Huang 0003, Weigang Zhang, Qingming Huang |
Inf. Sci. | 3 |
| 2019 | Online Asymmetric Metric Learning With Multi-Layer Similarity Aggregation for Cross-Modal RetrievalabstractCross-modal retrieval has attracted intensive attention in recent years, where a substantial yet challenging problem is how to measure the similarity between heterogeneous data modalities. Despite using modality-specific representation learning techniques, most existing shallow or deep models treat different modalities equally and neglect the intrinsic modality heterogeneity and information imbalance among images and texts. In this paper, we propose an online similarity function learning framework to learn the metric that can well reflect the cross-modal semantic relation. Considering that multiple CNN feature layers naturally represent visual information from low-level visual patterns to high-level semantic abstraction, we propose a new asymmetric image-text similarity formulation which aggregates the layer-wise visual-textual similarities parameterized by different bilinear parameter matrices. To effectively learn the aggregated similarity function, we develop three different similarity combination strategies, i.e., average kernel, multiple kernel learning, and layer gating. The former two kernel-based strategies assign uniform weights on different layers to all data pairs; the latter works on the original feature representation and assigns instance-aware weights on different layers to different data pairs, and they are all learned by preserving the bi-directional relative similarity expressed by a large number of cross-modal training triplets. The experiments conducted on three public datasets well demonstrate the effectiveness of our methods. Yiling Wu, Shuhui Wang, Guoli Song, Qingming Huang |
IEEE Trans. Image Process. | 2 |
| 2019 | SkeletonNet: A Hybrid Network With a Skeleton-Embedding Process for Multi-View Image Representation LearningabstractMulti-view representation learning plays a fundamental role in multimedia data analysis. Some specific inter-view alignment principles are adopted in conventional models, where there is an assumption that different views share a common latent subspace. However, when dealing views on diverse semantic levels, the view-specific characteristics are neglected, and the divergent inconsistency of similarity measurements hinders sufficient information sharing. This paper proposes a hybrid deep network by introducing tensor factorization into the multi-view deep auto-encoder. The network adopts skeleton-embedding process for unsupervised multi-view subspace learning. It takes full consideration of view-specific characteristics, and leverages the strength of both shallow and deep architectures for modeling low- and high-level views, respectively. We first formulate the high-level-view semantic distribution as the underlying skeleton structure of the learned subspace, and then infer the local tangent structures according to the affinity propagation of low-level-view geometric correlations. As a consequence, more discriminative subspace representation can be learned from global semantic pivots to local geometric details. Experimental comparisons on three benchmark image datasets show the promising performance and flexibility of our model. Liang Li 0003, Shuhui Wang, Weigang Zhang, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Reverse Densely Connected Feature Pyramid Network for Object Detection
Yongjian Xin, Shuhui Wang, Liang Li 0003, Weigang Zhang, Qingming Huang |
ACCV (5) | 2 |
| 2018 | Less Is More: Picking Informative Frames for Video Captioning
Shuhui Wang, Weigang Zhang, Qingming Huang |
ECCV (13) | 2 |
| 2018 | Attentive Recurrent Neural Network for Weak-supervised Multi-label Image ClassificationabstractMulti-label image classification is a fundamental and challenging task in computer vision, and recently achieved significant progress by exploiting semantic relations among labels. However, the spatial positions of labels for multi-labels images are usually not provided in real scenarios, which brings insuperable barrier to conventional models. In this paper, we propose an end-to-end attentive recurrent neural network for multi-label image classification under only image-level supervision, which learns the discriminative feature representations and models the label relations simultaneously. First, inspired by attention mechanism, we propose a recurrent highlight network (RHN) which focuses on the most related regions in the image to learn the discriminative feature representations for different objects in an iterative manner. Second, we develop a gated recurrent relation extractor (GRRE) to model the label relations using multiplicative gates in a recurrent fashion, which learns to decide how multiple labels of the image influence the relation extraction. Extensive experiments on three benchmark datasets show that our model outperforms the state-of-the-arts, and performs better on small-object categories and under the scenario with large number of labels. Liang Li 0003, Shuhui Wang, Shuqiang Jiang, Qingming Huang |
ACM Multimedia | 2 |
| 2018 | Joint Global and Co-Attentive Representation Learning for Image-Sentence RetrievalabstractIn image-sentence retrieval task, correlated images and sentences involve different levels of semantic relevance. However, existing multi-modal representation learning paradigms fail to capture the meaningful component relation on word and phrase level, while the attention-based methods still suffer from component-level mismatching and huge computation burden. We propose a Joint Global and Co-Attentive Representation learning method (JGCAR) for image-sentence retrieval. We formulate a global representation learning task which utilizes both intra-modal and inter-modal relative similarity to optimize the semantic consistency of the visual/textual component representations. We further develop a co-attention learning procedure to fully exploit different levels of visual-linguistic relations. We design a novel softmax-like bi-directional ranking loss to learn the co-attentive representation for image-sentence similarity computation. It is capable of discovering the correlative components and rectifying inappropriate component-level correlation to produce more accurate sentence-level ranking results. By joint global and co-attentive representation learning, the latter benefits from the former by producing more semantically consistent component representation, and the former also benefits from the latter by back-propagating the contextual information. Image-sentence retrieval is performed as a two-step process in the testing stage, inheriting advantages on both effectiveness and efficiency. Experiments show that JGCAR outperforms existing methods on MSCOCO and Flickr30K image-sentence retrieval tasks. Shuhui Wang, Junbao Zhuo, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2018 | Learning Semantic Structure-preserved Embeddings for Cross-modal RetrievalabstractThis paper learns semantic embeddings for multi-label cross-modal retrieval. Our method exploits the structure in semantics represented by label vectors to guide the learning of embeddings. First, we construct a semantic graph based on label vectors which incorporates data from both modalities, and enforce the embeddings to preserve the local structure of this semantic graph. Second, we enforce the embeddings to well reconstruct the labels, i.e., the global semantic structure. In addition, we encourage the embeddings to preserve local geometric structure of each modality. Accordingly, the local and global semantic structure consistencies as well as the local geometric structure consistency are enforced, simultaneously. The mappings between inputs and embeddings are designed to be nonlinear neural network with larger capacity and more flexibility. The overall objective function is optimized by stochastic gradient descent to gain the scalability on large datasets. Experiments conducted on three real world datasets clearly demonstrate the superiority of our proposed approach over the state-of-the-art methods. Yiling Wu, Shuhui Wang, Qingming Huang |
ACM Multimedia | 2 |
| 2018 | Multi-label double-layer learning for cross-modal retrieval
Bingpeng Ma, Shuhui Wang, Yugui Liu, Qingming Huang |
Neurocomputing | 3 |
| 2018 | Semantic invariant cross-domain image generation with generative adversarial networks
Xiaofeng Mao, Shuhui Wang, Liying Zheng, Qingming Huang |
Neurocomputing | 2 |
| 2018 | Heterogeneous anomaly detection in social diffusion with discriminative feature discovery
Siyuan Liu 0001, Qiang Qu 0001, Shuhui Wang |
Inf. Sci. | 3 |
| 2018 | Convolutional neural network-based hidden Markov models for rolling element bearing fault identification
Shuhui Wang, Jiawei Xiang, Yongteng Zhong |
Knowl. Based Syst. | 1 |
| 2018 | Bilevel Multiview Latent Space LearningabstractDifferent kinds of features describe different aspects of image data, and each feature can be treated as a view when we take it as a particular understanding of images. Leveraging multiple views provides a richer and comprehensive description than using only a single view. However, multiview data are often represented by high-dimensional heterogeneous features, so it is meaningful to find a low-dimensional consensus representation from multiple views. In this paper, we propose an unsupervised multiview dimensionality reduction method for images based on bilevel latent space learning. As different views have different physical meanings and statistical properties, they are not directly comparable. Therefore, we learn the comparable representation for each view in the first level. The shared and the private nature of multiview data are exploited to accurately preserve the information of each view. Then, we fuse different views into a low-dimensional representation by conducting joint matrix factorization in the second level. To guarantee the low-dimensional representation to be compact and discriminative, the intrinsic geometric structure of data is utilized. Besides, our method considers resisting the outliers and noise contained in multiview data, which may influence the learned representation and deteriorate its semantic consistency. We design appropriate optimization objectives to learn the latent spaces in different levels. Compared with the existing methods, our method could provide a more flexible multiview learning strategy that not only accurately captures the information of each view but also is robust to outliers and noise, which can obtain a more discriminative and compact low-dimensional representation. Experiments on two real-world image data sets demonstrate the advantages of our method over the existing multiview dimensionality reduction methods. Zhe Xue, Guorong Li, Shuhui Wang, Weigang Zhang, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | A Universal String Matching Approach to Screen Content CodingabstractThis paper proposes a universal string matching (USM) approach to screen content coding (SCC). USM uses a primary reference buffer and a secondary reference buffer for string matching and includes three modes: general string (GS) mode, constrained string 1 (CS1) mode, and constrained string 2 (CS2) mode. The CS1 mode and CS2 mode are constrained cases of the GS mode. Due to the diversity of the screen content, each of the three modes plays an indispensable role in coding some types of screen content.When using USM to code a coding unit (CU), one of the three modes is selected to code the CU. Compared with high-efficiency video coding (HEVC) SCC reference software HM-16.6 + SCM-5.2 of full frame search range for intrablock copy, USM achieves an average Y BD-rate of -28.4% for five text and graphics with motion (TGM) sequences from the audio video coding standard SCC common test condition (CTC) test suite and -5.8% for eight TGM test sequences from the HEVC SCC CTC test suite in all intraconfigurations, with a nearly 10% decrease in encoding runtime and almost the same decoding runtime. Liping Zhao 0005, Kailun Zhou, Shuhui Wang, Tao Lin 0005 |
IEEE Trans. Multim. | 4 |
| 2017 | Online Asymmetric Similarity Learning for Cross-Modal RetrievalabstractCross-modal retrieval has attracted intensive attention in recent years. Measuring the semantic similarity between heterogeneous data objects is an essential yet challenging problem in cross-modal retrieval. In this paper, we propose an online learning method to learn the similarity function between heterogeneous modalities by preserving the relative similarity in the training data, which is modeled as a set of bi-directional hinge loss constraints on the cross-modal training triplets. The overall online similarity function learning problem is optimized by the margin based Passive-Aggressive algorithm. We further extend the approach to learn similarity function in reproducing kernel Hilbert spaces by kernelizing the approach and combining multiple kernels derived from different layers of the CNN features using the Hedging algorithm. Theoretical mistake bounds are given for our methods. Experiments conducted on real world datasets well demonstrate the effectiveness of our methods. Yiling Wu, Shuhui Wang, Qingming Huang |
CVPR | 2 |
| 2017 | A Graph Regularized Deep Neural Network for Unsupervised Image Representation LearningabstractDeep Auto-Encoder (DAE) has shown its promising power in high-level representation learning. From the perspective of manifold learning, we propose a graph regularized deep neural network (GR-DNN) to endue traditional DAEs with the ability of retaining local geometric structure. A deep-structured regularizer is formulated upon multi-layer perceptions to capture this structure. The robust and discriminative embedding space is learned to simultaneously preserve the high-level semantics and the geometric structure within local manifold tangent space. Theoretical analysis presents the close relationship between the proposed graph regularizer and the graph Laplacian regularizer in terms of the optimization objective. We also alleviate the growth of the network complexity by introducing the anchor-based bipartite graph, which guarantees the good scalability for large scale data. The experiments on four datasets show the comparable results of the proposed GR-DNN with the state-of-the-art methods. Liang Li 0003, Shuhui Wang, Weigang Zhang, Qingming Huang |
CVPR | 3 |
| 2017 | Multimodal Gaussian Process Latent Variable Models with HarmonizationabstractIn this work, we address multimodal learning problem with Gaussian process latent variable models (GPLVMs) and their application to cross-modal retrieval. Existing GPLVM based studies generally impose individual priors over the model parameters and ignore the intrinsic relations among these parameters. Considering the strong complementarity between modalities, we propose a novel joint prior over the parameters for multimodal GPLVMs to propagate multimodal information in both kernel hyperparameter spaces and latent space. The joint prior is formulated as a harmonization constraint on the model parameters, which enforces the agreement among the modality-specific GP kernels and the similarity in the latent space. We incorporate the harmonization mechanism into the learning process of multimodal GPLVMs. The proposed methods are evaluated on three widely used multimodal datasets for cross-modal retrieval. Experimental results show that the harmonization mechanism is beneficial to the GPLVM algorithms for learning non-linear correlation among heterogeneous modalities. Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
ICCV | 2 |
| 2017 | Online low-rank similarity function learning with adaptive relative margin for cross-modal retrievalabstractThis paper presents a Cross-Modal Online Low-Rank Similarity function learning method (CMOLRS) for cross-modal retrieval, which learns a low-rank bilinear similarity measure on data from different modalities. CMOLRS models the cross-modal relations by relative similarities on a set of training data triplets and formulates the relative relations as convex hinge loss functions. By adapting the margin of hinge loss using information from feature space and label space for each triplet, CMOLRS effectively captures the multi-level semantic correlation among cross-modal data. The similarity function is learned by online learning in the manifold of low-rank matrices, thus good scalability is gained when processing large scale datasets. Extensive experiments are conducted on three public datasets. Comparisons with the state-of-the-art methods show the effectiveness and efficiency of our approach. Yiling Wu, Shuhui Wang, Weigang Zhang, Qingming Huang |
ICME | 2 |
| 2017 | A Delicious Recipe Analysis Framework for Exploring Multi-Modal Recipes with Various AttributesabstractHuman beings have developed a diverse food culture. Many factors like ingredients, visual appearance, courses (e.g., breakfast and lunch), flavor and geographical regions affect our food perception and choice. In this work, we focus on multi-dimensional food analysis based on these food factors to benefit various applications like summary and recommendation. For that solution, we propose a delicious recipe analysis framework to incorporate various types of continuous and discrete attribute features and multi-modal information from recipes. First, we develop a Multi-Attribute Theme Modeling (MATM) method, which can incorporate arbitrary types of attribute features to jointly model them and the textual content. We then utilize a multi-modal embedding method to build the correlation between the learned textual theme features from MATM and visual features from the deep learning network. By learning attribute-theme relations and multi-modal correlation, we are able to fulfill different applications, including (1) flavor analysis and comparison for better understanding the flavor patterns from different dimensions, such as the region and course, (2) region-oriented multi-dimensional food summary with both multi-modal and multi-attribute information and (3) multi-attribute oriented recipe recommendation. Furthermore, our proposed framework is flexible and enables easy incorporation of arbitrary types of attributes and modalities. Qualitative and quantitative evaluation results have validated the effectiveness of the proposed method and framework on the collected Yummly dataset. Weiqing Min, Shuqiang Jiang, Shuhui Wang, Shuhuan Mei |
ACM Multimedia | 3 |
| 2017 | Deep Unsupervised Convolutional Domain AdaptationabstractIn multimedia analysis, the task of domain adaptation is to adapt the feature representation learned in the source domain with rich label information to the target domain with less or even no label information. Significant research endeavors have been devoted to aligning the feature distributions between the source and the target domains in the top fully connected layers based on unsupervised DNN-based models. However, the domain adaptation has been arbitrarily constrained near the output ends of the DNN models, which thus brings about inadequate knowledge transfer in DNN-based domain adaptation process, especially near the input end. We develop an attention transfer process for convolutional domain adaptation. The domain discrepancy, measured in correlation alignment loss, is minimized on the second-order correlation statistics of the attention maps for both source and target domains. Then we propose Deep Unsupervised Convolutional Domain Adaptation DUCDA method, which jointly minimizes the supervised classification loss of labeled source data and the unsupervised correlation alignment loss measured on both convolutional layers and fully connected layers. The multi-layer domain adaptation process collaborately reinforces each individual domain adaptation component, and significantly enhances the generalization ability of the CNN models. Extensive cross-domain object classification experiments show DUCDA outperforms other state-of-the-art approaches, and validate the promising power of DUCDA towards large scale real world application. Junbao Zhuo, Shuhui Wang, Weigang Zhang, Qingming Huang |
ACM Multimedia | 2 |
| 2017 | Multi-label classification by exploiting local positive and negative pairwise label correlation
Jun Huang 0003, Guorong Li, Shuhui Wang, Zhe Xue, Qingming Huang |
Neurocomputing | 3 |
| 2017 | A survey on context-aware mobile visual recognition
Weiqing Min, Shuqiang Jiang, Shuhui Wang, Ruihan Xu 0001, Yushan Cao, Luis Herranz, Zhiqiang He 0002 |
Multim. Syst. | 3 |
| 2017 | Multimodal Similarity Gaussian Process Latent Variable ModelabstractData from real applications involve multiple modalities representing content with the same semantics from complementary aspects. However, relations among heterogeneous modalities are simply treated as observation-to-fit by existing work, and the parameterized modality specific mapping functions lack flexibility in directly adapting to the content divergence and semantic complicacy in multimodal data. In this paper, we build our work based on the Gaussian process latent variable model (GPLVM) to learn the non-parametric mapping functions and transform heterogeneous modalities into a shared latent space. We propose multimodal Similarity Gaussian Process latent variable model (m-SimGP), which learns the mapping functions between the intra-modal similarities and latent representation. We further propose multimodal distance-preserved similarity GPLVM (m-DSimGP) to preserve the intra-modal global similarity structure, and multimodal regularized similarity GPLVM (m-RSimGP) by encouraging similar/dissimilar points to be similar/dissimilar in the latent space. We propose m-DRSimGP, which combines the distance preservation in m-DSimGP and semantic preservation in m-RSimGP to learn the latent representation. The overall objective functions of the four models are solved by simple and scalable gradient decent techniques. They can be applied to various tasks to discover the nonlinear correlations and to obtain the comparable low-dimensional representation for heterogeneous modalities. On five widely used real-world data sets, our approaches outperform existing models on cross-modal content retrieval and multimodal classification. Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Location-Based Parallel Tag Completion for Geo-Tagged Social Image RetrievalabstractHaving benefited from tremendous growth of user-generated content, social annotated tags get higher importance in the organization and retrieval of large-scale image databases on Online Sharing Websites (OSW). To obtain high-quality tags from existing community contributed tags with missing information and noise, tag-based annotation or recommendation methods have been proposed for performance promotion of tag prediction. While images from OSW contain rich social attributes, they have not taken full advantage of rich social attributes and auxiliary information associated with social images to construct global information completion models. In this article, beyond the image-tag relation, we take full advantage of the ubiquitous GPS locations and image-user relationship to enhance the accuracy of tag prediction and improve the computational efficiency. For GPS locations, we define the popular geo-locations where people tend to take more images as Points of Interests (POI), which are discovered by mean shift approach. For image-user relationship, we integrate a localized prior constraint, expecting the completed tag sub-matrix in each POI to maintain consistency with users’ tagging behaviors. Based on these two key issues, we propose a unified tag matrix completion framework, which learns the image-tag relation within each POI. To solve the optimization problem, an efficient proximal sub-gradient descent algorithm is designed. The model optimization can be easily parallelized and distributed to learn the tag sub-matrix for each POI. Extensive experimental results reveal that the learned tag sub-matrix of each POI reflects the major trend of users’ tagging results with respect to different POIs and users, and the parallel learning process provides strong support for processing large-scale online image databases. To fit the response time requirement and storage limitations of Tag-based Image Retrieval (TBIR) on mobile devices, we introduce Asymmetric Locality Sensitive Hashing (ALSH) to reduce the time cost and meanwhile improve the efficiency of retrieval. Shuhui Wang, Qingming Huang |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2017 | Trajectory Community Discovery and Recommendation by Multi-Source Diffusion ModelingabstractIn this paper, we detect communities from trajectories. Existing algorithms for trajectory clustering usually rely on simplex representation and a single proximity-related metric. Unfortunately, additional information markers (e.g., social interactions or semantics in the spatial layout) are ignored, leading to the inability to fully discover the communities in trajectory database. This is especially true for human-generated trajectories, where additional fine-grained markers (e.g., movement velocity at certain locations, or the sequence of semantic spaces visited) are especially useful in capturing latent relationships among community members. To overcome this limitation, we propose TODMIS, a general framework for Trajectory-based community Detection by diffusion modeling on Multiple Information Sources. TODMIS combines additional information with raw trajectory data and construct the diffusion process on multiple similarity metrics. It also learns the consistent graph Laplacians by constructing the multi-modal diffusion process and optimizing the heat kernel coupling on each pair of similarity matrices from multiple information sources. Then, dense sub-graph detection is used to discover the set of distinct communities (including community size) on the coupled multi-graph representation. At last, based on the community information, we propose a novel model for online recommendation. We evaluate TODMIS and our online recommendation methods using different real-life datasets. Experimental results demonstrate the effectiveness and efficiency of our methods. Siyuan Liu 0001, Shuhui Wang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | Cross-modal Retrieval by Real Label Partial Least SquaresabstractThis paper proposes a novel method named Real Label Partial Least Squares (RL-PLS) for the task of cross-modal retrieval. Pervious works just take the texts and images as two modalities in PLS. But in RL-PLS, considering that the class label is more related to the semantics directly, we take the class label as the assistant modality. Specially, we build two KPLS models and project both images and texts into the label space. Then, the similarity of images and texts can be measured more accurately in the label space. Furthermore, we do not restrict the label indicator values as the binary values as the traditional methods. By contraries, in RL-PLS, the label indicator values are set to the real values. Specially, the label indicator values are comprised by two parts: positive or negative represents the sample class while the absolute value represents the local structure in the class. By this way, the discriminate ability of RL-PLS is improved greatly. To show the effectiveness of RL-PLS, the experiments are conducted on two cross-modal retrieval tasks (Wiki and Pascal Voc2007), on which the competitive results are obtained. Bingpeng Ma, Shuhui Wang, Yugui Liu, Qingming Huang |
ACM Multimedia | 3 |
| 2016 | Effective Multimodality Fusion Framework for Cross-Media Topic DetectionabstractDue to the prevalence of We-Media, information is quickly published and received in various forms anywhere and anytime through the Internet. The rich cross-media information carried by the multimodal data in multiple media has a wide audience, deeply reflects the social realities, and brings about much greater social impact than any single media information. Therefore, automatically detecting topics from cross media is of great benefit for the organizations (i.e., advertising agencies and governments) that care about the social opinions. However, cross-media topic detection is challenging from the following aspects: 1) the multimodal data from different media often involve distinct characteristics and 2) topics are presented in an arbitrary manner among the noisy web data. In this paper, we propose a multimodality fusion framework and a topic recovery (TR) approach to effectively detect topics from cross-media data. The multimodality fusion framework flexibly incorporates the heterogeneous multimodal data into a multimodality graph, which takes full advantage from the rich cross-media information to effectively detect topic candidates (T.C.). The TR approach solidly improves the entirety and purity of detected topics by: 1) merging the T.C. that are highly relevant themes of the same real topic and 2) filtering out the less-relevant noise data in the merged T.C. Extensive experiments on both single-media and cross-media data sets demonstrate the promising flexibility and effectiveness of our method in detecting topics from cross media. Lingyang Chu, Guorong Li, Shuhui Wang, Weigang Zhang, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Cross-Modal Correlation Learning by Adaptive Hierarchical Semantic AggregationabstractWith the explosive growth of web data, effective and efficient technologies are in urgent need for retrieving semantically relevant contents of heterogeneous modalities. Previous studies devote efforts to modeling simple cross-modal statistical dependencies, and globally projecting the heterogeneous modalities into a measurable subspace. However, global projections cannot appropriately adapt to diverse contents, and the naturally existing multilevel semantic relation in web data is ignored. We study the problem of semantic coherent retrieval, where documents from different modalities should be ranked by the semantic relevance to the query. Accordingly, we propose TINA, a correlation learning method by adaptive hierarchical semantic aggregation. First, by joint modeling of content and ontology similarities, we build a semantic hierarchy to measure multilevel semantic relevance. Second, with a set of local linear projections and probabilistic membership functions, we propose two paradigms for local expert aggregation, i.e., local projection aggregation and local distance aggregation. To learn the cross-modal projections, we optimize the structure risk objective function that involves semantic coherence measurement, local projection consistency, and the complexity penalty of local projections. Compared to existing approaches, a better bias-variance tradeoff is achieved by TINA in real-world cross-modal correlation learning tasks. Extensive experiments on widely used NUS-WIDE and ICML-Challenge for image-text retrieval demonstrate that TINA better adapts to the multilevel semantic relation and content divergence, and, thus, outperforms state of the art with better semantic coherence. Yan Hua, Shuhui Wang, Siyuan Liu 0001, Anni Cai, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2016 | Corrections to "Cross-Modal Correlation Learning by Adaptive Hierarchical Semantic Aggregation"abstractPresents corrections to the paper, "Cross-modal correlation learning by adaptive hierarchical semantic aggregation," (Hua, Y., et al) IEEE Trans. Multimedia, vol. 18, no. 6, pp. 1201-1216, Jun. 2016. Yan Hua, Shuhui Wang, Siyuan Liu 0001, Anni Cai, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2016 | Pseudo 2D String Matching Technique for High Efficiency Screen Content CodingabstractThis paper proposes a pseudo 2D string matching (P2SM) technique for high efficiency screen content coding (SCC). The technique uses a primary reference buffer (PRB) and a secondary reference buffer (SRB) for string matching and string copying. In the encoder, optimal reference string searching is performed in both PRB and SRB, and either a PRB or an SRB string is selected as an optimal reference string on a string-by-string basis. If no reference string of at least one pixel is founded for a current pixel, then the current pixel is coded as an unmatched pixel. Compared with HM-16.4${+}$SCM-4.0 reference software, the proposed P2SM technique achieves up to 37.7% Y BD-rate reduction for a screen snapshot of a spreadsheet. On average, using HEVC SCC common test condition and YUV test sequences in text and graphics with motion category, the proposed technique achieves Y BD-rate reduction of 7.7%, 5.0%, 2.6% for all intra (AI), random access (RA) and low-delay B (LB) configurations, respectively in lossy coding with both intra block copy (IBC) and P2SM having the same 4 coding tree units (CTUs) searching range, and bit-rate saving of 6.0%, 3.9%, 3.1% for AI, RA, LB configurations, respectively in lossless coding with IBC having full frame searching range while P2SM having only 2 CTUs searching range, at very low additional encoding and decoding complexity. Liping Zhao 0005, Tao Lin 0005, Kailun Zhou, Shuhui Wang, Xianyi Chen |
IEEE Trans. Multim. | 4 |
| 2015 | Similarity Gaussian Process Latent Variable Model for Multi-modal Data AnalysisabstractData from real applications involve multiple modalities representing content with the same semantics and deliver rich information from complementary aspects. However, relations among heterogeneous modalities are simply treated as observation-to-fit by existing work, and the parameterized cross-modal mapping functions lack flexibility in directly adapting to the content divergence and semantic complicacy of multi-modal data. In this paper, we build our work based on Gaussian process latent variable model (GPLVM) to learn the non-linear non-parametric mapping functions and transform heterogeneous data into a shared latent space. We propose multi-modal Similarity Gaussian Process latent variable model (m-SimGP), which learns the nonlinear mapping functions between the intra-modal similarities and latent representation. We further propose multi-modal regularized similarity GPLVM (m-RSimGP) by encouraging similar/dissimilar points to be similar/dissimilar in the output space. The overall objective functions are solved by simple and scalable gradient decent techniques. The proposed models are robust to content divergence and high-dimensionality in multi-modal representation. They can be applied to various tasks to discover the non-linear correlations and obtain the comparable low-dimensional representation for heterogeneous modalities. On two widely used real-world datasets, we outperform previous approaches for cross-modal content retrieval and cross-modal classification. Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
ICCV | 2 |
| 2015 | Group sensitive Classifier Chains for multi-label classificationabstractIn multi-label classification, labels often have correlations with each other. Exploiting label correlations can improve the performances of classifiers. Current multi-label classification methods mainly consider the global label correlations. However, the label correlations may be different over different data groups. In this paper, we propose a simple and efficient framework for multi-label classification, called Group sensitive Classifier Chains. We assume that similar examples not only share the same label correlations, but also tend to have similar labels. We augment the original feature space with label space and cluster them into groups, then learn the label dependency graph in each group respectively and build the classifier chains on each group specific label dependency graph. The group specific classifier chains which are built on the nearest group of the test example are used for prediction. Comparison results with the state-of-the-art approaches manifest competitive performances of our method. Jun Huang 0003, Guorong Li, Shuhui Wang, Weigang Zhang, Qingming Huang |
ICME | 3 |
| 2015 | Improving cross-modal correlation learning with hyperlinksabstractWe propose a new cross-modal correlation learning framework which boosts the performance of correlation learning models using the hyperlink information. First, we design a neighborhood selection paradigm using the hyperlink structure and content similarities to identify a set of semantically related documents for each multi-modal document in both training and testing stage. Based on the neighborhood structure, we revise two well-established content-based correlation learning models, i.e., canonical correlation analysis (CCA) and kernel canonical correlation analysis (KCCA) with a structure coding matrix. Third, we develop a correlation score aggregation technique to discover more semantically relevant cross-modal documents. To our best knowledge, this is the first to introduce hyperlink information into cross-modal correlation learning. Experimental results demonstrate that our proposed framework can significantly improve the model generality towards real-world cross-modal retrieval. Shuhui Wang, Yiling Wu, Qingming Huang |
ICME | 1 |
| 2015 | GOMES: A group-aware multi-view fusion approach towards real-world image clusteringabstractDifferent features describe different views of visual appearance, multi-view based methods can integrate the information contained in each view and improve the image clustering performance. Most of the existing methods assume that the importance of one type of feature is the same to all the data. However, the visual appearance of images are different, so the description abilities of different features vary with different images. To solve this problem, we propose a group-aware multi-view fusion approach. Images are partitioned into groups which consist of several images sharing similar visual appearance. We assign different weights to evaluate the pairwise similarity between different groups. Then the clustering results and the fusion weights are learned by an iterative optimization procedure. Experimental results indicate that our approach achieves promising clustering performance compared with the existing methods. Zhe Xue, Guorong Li, Shuhui Wang, Chunjie Zhang 0001, Weigang Zhang, Qingming Huang |
ICME | 3 |
| 2015 | Location-Based Parallel Tag Completion for Geo-tagged Social Image RetrievalabstractBenefit from tremendous growth of user-generated content, social annotated tags get higher importance in organization and retrieval of large scale image database on Online Sharing Websites (OSW). To obtain high-quality tags from existing community contributed tags with missing information and noise, tag-based annotation or recommendation methods have been proposed for performance promotion of tag prediction. While images from OSW contain rich social attributes, existing studies only utilize the relations between visual content and tags to construct global information completion models. In this paper, beyond the image-tag relation, we take full advantage of the ubiquitous GPS locations and image-user relationship, to enhance the accuracy of tag prediction and improve the computational efficiency. For GPS locations, we define the popular geo-locations where people tend to take more images as Points of Interests (POI), which are discovered by mean shift approach. For image-user relationship, we integrate a localized prior constraint, expecting the completed tag sub-matrix in each POI to maintain consistency with users' tagging behaviors. Based on these two key issues, we propose a unified tag matrix completion framework which learns the image-tag relation within each POI. To solve the proposed model, an efficient proximal sub-gradient descent algorithm is designed. The model optimization can be easily parallelized and distributed to learn the tag sub-matrix for each POI. Extensive experimental results reveal that the learned tag sub-matrix of each POI reflects the major trend of users' tagging results with respect to different POIs and users, and the parallel learning process provides strong support for processing large scale online image database. Shuhui Wang, Qingming Huang |
ICMR | 2 |
| 2015 | Understanding taxi drivers' routing choices from spatial and social traces
Siyuan Liu 0001, Shuhui Wang, Ramayya Krishnan |
Frontiers Comput. Sci. | 2 |
| 2015 | Cluster-sensitive Structured Correlation Analysis for Web cross-modal retrieval
Shuhui Wang, Fuzhen Zhuang, Shuqiang Jiang, Qingming Huang, Qi Tian 0001 |
Neurocomputing | 1 |
| 2015 | Polysemious visual representation based on feature aggregation for large scale image applications
Xinhang Song, Shuqiang Jiang, Shuhui Wang, Liang Li 0003, Qingming Huang |
Multim. Tools Appl. | 3 |
| 2015 | Pseudo-2D-matching based enhancement to high efficiency video coding for screen contents
Shuhui Wang, Tao Lin 0005, Kailun Zhou, Peijun Zhang, Xianyi Chen |
Multim. Tools Appl. | 1 |
| 2015 | ALID: Scalable Dominant Cluster DetectionabstractDetecting dominant clusters is important in many analytic applications. The state-of-the-art methods find dense subgraphs on the affinity graph as dominant clusters. However, the time and space complexities of those methods are dominated by the construction of affinity graph, which is quadratic with respect to the number of data points, and thus are impractical on large data sets. To tackle the challenge, in this paper, we apply Evolutionary Game Theory (EGT) and develop a scalable algorithm, Approximate Localized Infection Immunization Dynamics (ALID). The major idea is to perform Localized Infection Immunization Dynamics (LID) to find dense subgraphs within local ranges of the affinity graph. LID is further scaled up with guaranteed high efficiency and detection quality by an estimated Region of Interest (ROI) and a Candidate Infective Vertex Search method (CIVS). ALID only constructs small local affinity graphs and has time complexity O ( C ( a * + δ ) n ) and space complexity O ( a * ( a * + δ )), where a * is the size of the largest dominant cluster, and C « n and δ « n are small constants. We demonstrate by extensive experiments on both synthetic data and real world data that ALID achieves the state-of-the-art detection quality with much lower time and space cost on single machine. We also demonstrate the encouraging parallelization performance of ALID by implementing the Parallel ALID (PALID) on Apache Spark. PALID processes 50 million SIFT data points in 2.29 hours, achieving a speedup ratio of 7.51 with 8 executors. Lingyang Chu, Shuhui Wang, Siyuan Liu 0001, Qingming Huang, Jian Pei 0001 |
Proc. VLDB Endow. | 2 |
| 2015 | Multi-Level Discriminative Dictionary Learning With Application to Large Scale Image ClassificationabstractThe sparse coding technique has shown flexibility and capability in image representation and analysis. It is a powerful tool in many visual applications. Some recent work has shown that incorporating the properties of task (such as discrimination for classification task) into dictionary learning is effective for improving the accuracy. However, the traditional supervised dictionary learning methods suffer from high computation complexity when dealing with large number of categories, making them less satisfactory in large scale applications. In this paper, we propose a novel multi-level discriminative dictionary learning method and apply it to large scale image classification. Our method takes advantage of hierarchical category correlation to encode multi-level discriminative information. Each internal node of the category hierarchy is associated with a discriminative dictionary and a classification model. The dictionaries at different layers are learnt to capture the information of different scales. Moreover, each node at lower layers also inherits the dictionary of its parent, so that the categories at lower layers can be described with multi-scale information. The learning of dictionaries and associated classification models is jointly conducted by minimizing an overall tree loss. The experimental results on challenging data sets demonstrate that our approach achieves excellent accuracy and competitive computation cost compared with other sparse coding methods for large scale image classification. Li Shen 0005, Gang Sun 0005, Qingming Huang, Shuhui Wang, Zhouchen Lin, Enhua Wu |
IEEE Trans. Image Process. | 4 |
| 2015 | Rationality Analytics from TrajectoriesabstractThe availability of trajectories tracking the geographical locations of people as a function of time offers an opportunity to study human behaviors. In this article, we study rationality from the perspective of user decision on visiting a point of interest (POI) which is represented as a trajectory. However, the analysis of rationality is challenged by a number of issues, for example, how to model a trajectory in terms of complex user decision processes? and how to detect hidden factors that have significant impact on the rational decision making? In this study, we propose Rationality Analysis Model (RAM) to analyze rationality from trajectories in terms of a set of impact factors. In order to automatically identify hidden factors, we propose a method, Collective Hidden Factor Retrieval (CHFR), which can also be generalized to parse multiple trajectories at the same time or parse individual trajectories of different time periods. Extensive experimental study is conducted on three large-scale real-life datasets (i.e., taxi trajectories, user shopping trajectories, and visiting trajectories in a theme park). The results show that the proposed methods are efficient, effective, and scalable. We also deploy a system in a large theme park to conduct a field study. Interesting findings and user feedback of the field study are provided to support other applications in user behavior mining and analysis, such as business intelligence and user management for marketing purposes. Siyuan Liu 0001, Qiang Qu 0001, Shuhui Wang |
ACM Trans. Knowl. Discov. Data | 3 |
| 2015 | Structured Learning from Heterogeneous Behavior for Social Identity LinkageabstractSocial identity linkage across different social media platforms is of critical importance to business intelligence by gaining from social data a deeper understanding and more accurate profiling of users. In this paper, we propose a solution framework, HYDRA, which consists of three key steps: (I) we model heterogeneous behavior by long-term topical distribution analysis and multi-resolution temporal behavior matching against high noise and information missing, and the behavior similarity are described by multi-dimensional similarity vector for each user pair; (II) we build structure consistency models to maximize the structure and behavior consistency on users' core social structure across different platforms, thus the task of identity linkage can be performed on groups of users, which is beyond the individual level linkage in previous study; and (III) we propose a normalized-margin-based linkage function formulation, and learn the linkage function by multi-objective optimization where both supervised pair-wise linkage function learning and structure consistency maximization are conducted towards a unified Pareto optimal solution. The model is able to deal with drastic information missing, and avoid the curse-of-dimensionality in handling high dimensional sparse representation. Extensive experiments on 10 million users across seven popular social networks platforms demonstrate that HYDRA correctly identifies real user linkage across different platforms from massive noisy user behavior data records, and outperforms existing state-of-the-art approaches by at least 20 percent under different settings, and four times better in most settings. Siyuan Liu 0001, Shuhui Wang, Feida Zhu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | TINA: Cross-Modal Correlation Learning by Adaptive Hierarchical Semantic AggregationabstractWith the explosive growth of web data, effective and efficient technologies are in urgent needs for retrieving semantically relevant contents of heterogeneous modalities. Previous studies construct global transformations to project the heterogeneous data into a measurable subspace. However, global projections cannot appropriately adapt to diverse contents, and the naturally existing multi-level semantic relation in web data is ignored. We study the problem of semantic coherent retrieval, where documents from different modalities should be ranked by the semantic relevance to the queries. Accordingly, we propose TINA, a correlation learning method by Adaptive Hierarchical Semantic Aggregation. First, by joint modeling of content and ontology similarities, we build a semantic hierarchy to measure multi-level semantic relevance. Second, with a set of local linear projections aggregated by gating functions, we optimize the structure risk objective function that involves semantic coherence measurement, local projection consistency and the complexity penalty of local projections. Therefore, semantic coherence and a better bias-variance trade-off can be achieved by TINA. Extensive experiments on widely used NUS-WIDE and ICML-Challenge datasets demonstrate that TINA outperforms state-of-the-art, and achieves better adaptation to the multi-level semantic relation and content divergence. Yan Hua, Shuhui Wang, Siyuan Liu 0001, Qingming Huang, Anni Cai |
ICDM | 2 |
| 2014 | Cross modal metric learning with multi-level semantic relevanceabstractThe Mahalanobis metric learning is an effective tool for constructing semantic consistent distance among data in single modal data analysis. However, distance metric learning is a more challenging issue for cross modal data, where less attention has been paid in previous studies. In this paper, we propose Cross mOdal Large mArgin metric leaRning (COLAR) with multi-level semantic relevance. With large margin principle, we model different levels of the semantic relations across modalities, e.g., the one-to-one correspondence and intra-class relation, while traditional correlation learning approaches (such as CCA and its variants) can only handle the one-to-one correspondence or treat them indiscriminatively. As a result, the distances of multi-level relevance among cross modal data are optimized based on a regularized learning framework. Promising performance is achieved on cross modal retrieval, i.e., image-to-text retrieval and text-to-image retrieval. Yan Hua, Shuhui Wang, Zhicheng Zhao 0001, Qingming Huang, Anni Cai |
ICIP | 2 |
| 2014 | Sharing model with multi-level feature representationsabstractHierarchical classification models have been proposed to achieve high accuracy by transferring effective information across the categories. One important challenge for this paradigm is to design what can be transferred across the categories. In this paper, we propose a novel method to learn a sharing model by taking advantage of multi-level feature representations. Unlike many of the existing methods which learn the sharing model based on identical feature space, multi-level feature detectors enable our model to capture rich visual information in hierarchical category structure. Moreover, hierarchical classifier parameters associated with multi-level feature representations are learned to model the visual correlation in the hierarchy. The experimental results on Caltech-256 dataset and ImageNet subset demonstrate that our method achieves excellent performance compared with some state-of-the-art methods, and shows the advantage of multi-level information transfer. Li Shen 0005, Gang Sun 0005, Shuhui Wang, Enhua Wu, Qingming Huang |
ICIP | 3 |
| 2014 | Graph-Density-based visual word vocabulary for image retrievalabstractDescriptive visual word vocabulary serves as the foundation of large scale image retrieval systems. However, the visual word descriptive power is limited by the construction mechanisms based on either cluster center or partitioned feature space, since such mechanisms may merge the sparsely distributed features and split the densely distributed features. Besides, there are a large number of outlier features that are not similar with any visual word. Quantizing such features into visual words inevitably decreases the visual word descriptive power. In this paper, we propose a novel Graph-Density-based visual word Vocabulary (GDV), which constructs the visual word by dense feature subgraph and directly measures the intra-word similarity by the corresponding graph density. Our method remarkably enhances the visual word descriptive power from the following three aspects: 1) GDV guarantees the high intra-word similarity by constructing visual words under the criterion of large graph density; 2) GDV improves the inter-word dissimilarity by alleviating the unexpected effect of subgraph splitting; 3) GDV suppresses the influence of outlier features by selectively quantizing only the features that are similar enough with the visual words. Extensive experiments demonstrate GDV's advanced descriptive power over traditional visual word vocabularies in enhancing both the retrieval accuracy and efficiency, which provides a higher level starting point for most image retrieval systems. Lingyang Chu, Shuhui Wang, Shuqiang Jiang, Qingming Huang |
ICME | 2 |
| 2014 | Cross media topic analytics based on synergetic content and user behavior modelingabstractHot topic in cross media, defined as a set of Web documents containing similar semantic information, describes the same real world event with significant social impact. However, cross media topic detection still remains a challenging issue since it is unclear how users interact with the cross media topics, leading to difficulties in identifying the influence of different topics on different user communities. In this paper, we propose a solution framework for cross media topic analysis based on synergetic modeling of multi-modal content and user behavior. First, we detect atom topics by multi-modal topic detection. Second, we propose a multi-resolution user behavior modeling method to discover communities on the active Web users by considering the distribution of related atom topics along the temporal axis. We analyze the topic-topic, topic-community and community-community relations by using the proposed method. Consequently, the macro-topic and macro-community structures can be obtained for better understanding of the interaction between topics and communities. Experiments show the capability of our method in knowledge discovery on cross media topics. Shuhui Wang, Zhenjun Wang, Shuqiang Jiang, Qingming Huang |
ICME | 1 |
| 2014 | Persistent Community Detection in Dynamic Social Networks
Siyuan Liu 0001, Shuhui Wang, Ramayya Krishnan |
PAKDD (1) | 2 |
| 2014 | HYDRA: large-scale social identity linkage via heterogeneous behavior modelingabstractWe study the problem of large-scale social identity linkage across different social media platforms, which is of critical importance to business intelligence by gaining from social data a deeper understanding and more accurate profiling of users. This paper proposes HYDRA, a solution framework which consists of three key steps: (I) modeling heterogeneous behavior by long-term behavior distribution analysis and multi-resolution temporal information matching; (II) constructing structural consistency graph to measure the high-order structure consistency on users' core social structures across different platforms; and (III) learning the mapping function by multi-objective optimization composed of both the supervised learning on pair-wise ID linkage information and the cross-platform structure consistency maximization. Extensive experiments on 10 million users across seven popular social network platforms demonstrate that HYDRA correctly identifies real user linkage across different platforms, and outperforms existing state-of-the-art algorithms by at least 20% under different settings, and 4 times better in most settings. Siyuan Liu 0001, Shuhui Wang, Feida Zhu 0001, Ramayya Krishnan |
SIGMOD Conference | 2 |
| 2014 | United coding method for compound image compression
Shuhui Wang, Tao Lin 0005 |
Multim. Tools Appl. | 1 |
| 2013 | Antihypertensive effect of acupuncturing at KI3 in spontaneously hypertensive ratsabstractNowadays, hypertension is a major issue in public health worldwide, because it contributes to vascular and renal morbidity, cardiovascular mortality, and economic burden. Thus, many animal and clinical studies have done to find ways to the management of hypertension. Many researches have shown that acupuncture is an effective alternative treatment in antihypertensive. This study investigated the regulation effect of acupuncture at both Taixi (KI3) acupoints to the blood pressure in spontaneously hypertensive rats (SHRs). Eighteen SHRs were randomly divided into three groups (SHR group, KI group, non-acupoint group), six Wistar Kyoto rats(WKY) were used as normal controls. The original and 7 days' systolic blood pressure were measured, using a computerized rat tail-cuff technique. Results: The systolic blood pressure of rats in SHR group were much higer than WKY group. The blood pressures in both KI group and non-acupoint (NON) group were decreased significantly after acupuncture (P<;0.05). Compared with the non-acupoint group, the antihypertensive effect of acupoint is more obvious(P<;0.05). Conclusion: Acupuncture at Taixi (KI3) had stable antihypertensive effect in spontaneously hypertensive rats (SHRs), and is better than non-acupoint. Shaoyang Cui, Shuhui Wang, Chunzhi Tang, Xinsheng Lai, Zhiqi Fan |
BIBM | 3 |
| 2013 | TODMIS: mining communities from trajectoriesabstractExisting algorithms for trajectory-based clustering usually rely on simplex representation and a single proximity-related distance (or similarity) measure. Consequently, additional information markers (e.g., social interactions or the semantics of the spatial layout) are usually ignored, leading to the inability to fully discover the communities in the trajectory database. This is especially true for human-generated trajectories, where additional fine-grained markers (e.g., movement velocity at certain locations, or the sequence of semantic spaces visited) can help capture latent relationships between cluster members. To address this limitation, we propose TODMIS: a general framework for Trajectory cOmmunity Discovery using Multiple Information Sources. TODMIS combines additional information with raw trajectory data and creates multiple similarity metrics. In our proposed approach, we first develop a novel approach for computing semantic level similarity by constructing a Markov Random Walk model from the semantically-labeled trajectory data, and then measuring similarity at the distribution level. In addition, we also extract and compute pair-wise similarity measures related to three additional markers, namely trajectory level spatial alignment (proximity), temporal patterns and multi-scale velocity statistics. Finally, after creating a single similarity metric from the weighted combination of these multiple measures, we apply dense sub-graph detection to discover the set of distinct communities. We evaluated TODMIS extensively using traces of (i) student movement data in a campus, (ii) customer trajectories in a shopping mall, and (iii) city-scale taxi movement data. Experimental results demonstrate that TODMIS correctly and efficiently discovers the real grouping behaviors in these diverse settings. Siyuan Liu 0001, Shuhui Wang, Kasthuri Jayarajah, Archan Misra, Ramayya Krishnan |
CIKM | 2 |
| 2013 | Multi-level Discriminative Dictionary Learning towards Hierarchical Visual CategorizationabstractFor the task of visual categorization, the learning model is expected to be endowed with discriminative visual feature representation and flexibilities in processing many categories. Many existing approaches are designed based on a flat category structure, or rely on a set of pre-computed visual features, hence may not be appreciated for dealing with large numbers of categories. In this paper, we propose a novel dictionary learning method by taking advantage of hierarchical category correlation. For each internode of the hierarchical category structure, a discriminative dictionary and a set of classification models are learnt for visual categorization, and the dictionaries in different layers are learnt to exploit the discriminative visual properties of different granularity. Moreover, the dictionaries in lower levels also inherit the dictionary of ancestor nodes, so that categories in lower levels are described with multi-scale visual information using our dictionary learning approach. Experiments on Image Net object data subset and SUN397 scene dataset demonstrate that our approach achieves promising performance on data with large numbers of classes compared with some state-of-the-art methods, and is more efficient in processing large numbers of categories. Li Shen 0005, Shuhui Wang, Gang Sun 0005, Shuqiang Jiang, Qingming Huang |
CVPR | 2 |
| 2013 | WIKI-CMR: A web cross modality dataset for studying and evaluation of cross modality retrieval modelsabstractWith the popularity of Web multimedia data, cross-modality retrieval becomes an urgent and challenging problem. Bridging the semantic gap between different modalities and dealing with abundant data are the main challenges for cross-modality retrieval. A well-designed dataset could provide a platform for developing the state-of-the-art cross-modality retrieval algorithms. However, existing Web cross-modality datasets are small in size, or do not contain the full information, for example, the hyperlink structure. In this paper, we introduce a new Web cross-modality dataset called “WIKI-CMR” by selecting Wikipedia as the reliable and information-rich data resource, and collect data with a smart crawling strategy. This dataset is comprised of 74961 documents with textual paragraphs, images and hyperlinks. All documents are categorized into 11 semantic topics. We point out several challenges on this dataset and use this dataset to evaluate some well-known cross-modality retrieval models. Shuhui Wang, Chunjie Zhang 0001, Qingming Huang |
ICME | 2 |
| 2013 | Cross-media topic detection: A multi-modality fusion frameworkabstractDetecting topics from Web data attracts increasing attention in recent years. Most previous works on topic detection mainly focus on the data from single medium, however, the rich and complementary information carried by multiple media can be used to effectively enhance the topic detection performance. In this paper, we propose a flexible data fusion framework to detect topics that simultaneously exist in different mediums. The framework is based on a multi-modality graph (MMG), which is obtained by fusing two single-modality graphs together: a text graph and a visual graph. Each node of MMGrepresents a multi-modal data and the edge weight between two nodes jointly measures their content and upload-time similarities. Since the data about the same topic often have similar content and are usually uploaded in a similar period of time, they would naturally form a dense (namely, strongly connected) subgraph in MMG. Such dense subgraph is robust to noise and can be efficiently detected by pair-wise clustering methods. The experimental results on single-medium and cross-media datasets demonstrate the flexibility and effectiveness of our method. Guorong Li, Lingyang Chu, Shuhui Wang, Weigang Zhang, Qingming Huang |
ICME | 4 |
| 2013 | Beyond bag of words: image representation in sub-semantic spaceabstractDue to the semantic gap, the low-level features are not able to semantically represent images well. Besides, traditional semantic related image representation may not be able to cope with large inter class variations and are not very robust to noise. To solve these problems, in this paper, we propose a novel image representation method in the sub-semantic space. First, examplar classifiers are trained by separating each training image from the others and serve as the weak semantic similarity measurement. Then a graph is constructed by combining the visual similarity and weak semantic similarity of these training images. We partition this graph into visually and semantically similar sub-sets. Each sub-set of images are then used to train classifiers in order to separate this sub-set from the others. The learned sub-set classifiers are then used to construct a sub-semantic space based representation of images. This sub-semantic space is not only more semantically meaningful but also more reliable and resistant to noise. Finally, we make categorization of images using this sub-semantic space based representation on several public datasets to demonstrate the effectiveness of the proposed method. Chunjie Zhang 0001, Shuhui Wang, Chao Liang 0001, Jing Liu 0001, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2013 | Undo the codebook bias by linear transformation for visual applicationsabstractThe bag of visual words model (BoW) and its variants have demonstrate their effectiveness for visual applications and have been widely used by researchers. The BoW model first extracts local features and generates the corresponding codebook, the elements of a codebook are viewed as visual words. The local features within each image are then encoded to get the final histogram representation. However, the codebook is dataset dependent and has to be generated for each image dataset. This costs a lot of computational time and weakens the generalization power of the BoW model. To solve these problems, in this paper, we propose to undo the dataset bias by codebook linear transformation. To represent every points within the local feature space using Euclidean distance, the number of bases should be no less than the space dimensions. Hence, each codebook can be viewed as a linear transformation of these bases. In this way, we can transform the pre-learned codebooks for a new dataset. However, not all of the visual words are equally important for the new dataset, it would be more effective if we can make some selection using sparsity constraints and choose the most discriminative visual words for transformation. We propose an alternative optimization algorithm to jointly search for the optimal linear transformation matrixes and the encoding parameters. Image classification experimental results on several image datasets show the effectiveness of the proposed method. Chunjie Zhang 0001, Yifan Zhang 0001, Shuhui Wang, Junbiao Pang, Chao Liang 0001, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2013 | Cross Concept Local Fisher Discriminant Analysis for Image Classification
Xinhang Song, Shuqiang Jiang, Shuhui Wang, Jinhui Tang 0001, Qingming Huang |
MMM (2) | 3 |
| 2013 | Arbitrary shape matching for screen content codingabstractIn this paper, we present an arbitrary shape matching (ASM) coding technique for screen contents. In ASM coder, a Coding Unit (CU) is broken into multiple pixel sample strings. Each string called matched string in the CU has a matching string in the previously coded and reconstructed pixel buffer. Then the distance (either 1D or 2D) between the matching string and matched string and the length of the string are entropy-coded into the bitstream buffer. If the length is zero (no matching is found), then the original pixel sample is entropy-coded into the bitstream buffer. Experiments show that for some types of screen contents, ASM can significantly improve coding performance. Tao Lin 0005, Kailun Zhou, Xianyi Chen, Shuhui Wang |
PCS | 4 |
| 2013 | Shared Structure Learning for Multiple Tasks with Multiple Views
Xin Jin 0004, Fuzhen Zhuang, Shuhui Wang, Qing He 0003, Zhongzhi Shi |
ECML/PKDD (2) | 3 |
| 2013 | Compound image compression based on unified LZ and hybrid codingabstractThis study proposes a unified LZ and hybrid coding (ULHC) method for compound image and video compression of visually lossless quality and high compression ratio. The method is macroblock‐based for ultra‐low coding latency and compatibility with the conventional and most popular hybrid video coding standards such as H.264 and MPEG‐2. First, each macroblock is coded by two tools: (i) gzip, a popular lossless LZ coding tool, modified to be macroblock‐oriented and to be seamlessly unifiable with a lossy hybrid coding tool; (ii) H.264, an advanced lossy hybrid coding tool. Then rate‐distortion optimisation is used to select either the modified gzip or H.264 as the final coding. To seamlessly unify these two coding tools for maximum quality and high compression ratio in ULHC, the modified gzip uses the most recent reconstructed and specially serialised macroblock data as the dictionary. Experimental results show that for images and videos composed of natural or synthesised picture, text and graphics, the proposed method provides higher peak signal‐to‐noise ratio and better subjective quality than H.264 at the same bitrate, and also achieves much higher compression ratio than gzip without any visual quality loss. In fact, ULHC can achieve partial‐lossless and partial‐near‐lossless coding with a high compression ratio. Shuhui Wang, Tao Lin 0005 |
IET Image Process. | 1 |
| 2013 | Laplacian affine sparse coding with tilt and orientation consistency for image classification
Chunjie Zhang 0001, Shuhui Wang, Qingming Huang, Chao Liang 0001, Jing Liu 0001, Qi Tian 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Image classification using spatial pyramid robust sparse coding
Chunjie Zhang 0001, Shuhui Wang, Qingming Huang, Jing Liu 0001, Chao Liang 0001, Qi Tian 0001 |
Pattern Recognit. Lett. | 2 |
| 2013 | Mixed Chroma Sampling-Rate High Efficiency Video Coding for Full-Chroma Screen ContentabstractComputer screens contain discontinuous-tone content and continuous-tone content. Thus, the most effective way for screen content coding (SCC) is to use two essentially different coders: a dictionary-entropy coder and a traditional hybrid coder. Although screen content is originally in a full-chroma (e.g., YUV444) format, the current method of compression is to first subsample chroma of pictures and then compress pictures using a chroma-subsampled (e.g., YUV420) coder. Using two chroma-subsampled coders cannot achieve high-quality SCC, but using two full-chroma coders is overkill and inefficient for SCC. To solve the dilemma, this paper proposes a mixed chroma sampling-rate approach for SCC. An original full-chroma input macroblock (coding unit) or its prediction residual is chroma-subsampled. One full-chroma base coder and one chroma-subsampled base coder are used simultaneously to code the original and the chroma-subsampled macroblock, respectively. The coder minimizing rate-distortion (R-D) is selected as the final coder for the macroblock. The two base coders are coherently unified and optimized to get the best overall coding performance and share coding components and resources as much as possible. The approach achieves very high visual quality with minimal computing complexity increment for SCC, and has better R-D performance than two full-chroma coders approach, especially in low bitrate. Tao Lin 0005, Peijun Zhang, Shuhui Wang, Kailun Zhou, Xianyi Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Robust Spatial Consistency Graph Model for Partial Duplicate Image RetrievalabstractPartial duplicate images often have large non-duplicate regions and small duplicate regions with random rotation, which lead to the following problems: 1) large number of noisy features from the non-duplicate regions; 2) small number of representative features from the duplicate regions; 3) randomly rotated or deformed duplicate regions. These problems challenge many content based image retrieval (CBIR) approaches, since most of them cannot distinguish the representative features from a large proportion of noisy features in a rotation invariant way. In this paper, we propose a rotation invariant partial duplicate image retrieval (PDIR) approach, which effectively and efficiently retrieves the partial duplicate images by accurately matching the representative SIFT features. Our method is based on the Combined-Orientation-Position (COP) consistency graph model, which consists of the following two parts: 1) The COP consistency, which is a rotation invariant measurement of the relative spatial consistency among the candidate matches of SIFT features; it uses a coarse-to-fine family of evenly sectored polar coordinate systems to softly quantize and combine the orientations and positions of the SIFT features. 2) The consistency graph model, which robustly rejects the spatially inconsistent noisy features by effectively detecting the group of candidate feature matches with the largest average COP consistency. Extensive experiments on five large scale image data sets show promising retrieval performances. Lingyang Chu, Shuqiang Jiang, Shuhui Wang, Qingming Huang |
IEEE Trans. Multim. | 3 |
| 2013 | Accurate and efficient cross-domain visual matching leveraging multiple feature representations
Gang Sun 0005, Shuhui Wang, Xuehui Liu, Qingming Huang, Yanyun Chen, Enhua Wu |
Vis. Comput. | 2 |
| 2012 | Multi-feature metric learning with knowledge transfer among semantics and social taggingabstractPrevious metric learning approaches learn a unified metric for all the classes on single feature representation, thus cannot be directly transplanted to applications involving multiple features, hundreds to thousands of hierarchical structured semantics and abundant social tagging. In this paper, we propose a novel multi-task multi-feature metric learning method which models the information sharing mechanism among different learning tasks. We decompose the real world multi-class problems such as semantic categorization or automatic tagging into a set of tasks where each task corresponds to several classes with strong visual correlation. We conduct metric learning to learn a set of (hyper)category-specific metrics for all the tasks. By encouraging model sharing among tasks, more generalization power is acquired. Another advantage is the capability of simultaneous learning with semantic information and social tagging based on the multi-task learning framework, and thus they both benefit from the information provided by each other. Experiments demonstrate the advantages on applications including semantic categorization and automatic tagging compared with other popular metric learning approaches. Shuhui Wang, Shuqiang Jiang, Qingming Huang, Qi Tian 0001 |
CVPR | 1 |
| 2012 | Nearest-neighbor method using multiple neighborhood similarities for social media data mining
Shuhui Wang, Qingming Huang, Shuqiang Jiang, Qi Tian 0001 |
Neurocomputing | 1 |
| 2012 | S3MKL: Scalable Semi-Supervised Multiple Kernel Learning for Real-World Image ApplicationsabstractWe study the visual learning models that could work efficiently with little ground-truth annotation and a mass of noisy unlabeled data for large scale Web image applications, following the subroutine of semi-supervised learning (SSL) that has been deeply investigated in various visual classification tasks. However, most previous SSL approaches are not able to incorporate multiple descriptions for enhancing the model capacity. Furthermore, sample selection on unlabeled data was not advocated in previous studies, which may lead to unpredictable risk brought by real-world noisy data corpse. We propose a learning strategy for solving these two problems. As a core contribution, we propose a scalable semi-supervised multiple kernel learning method$({\rm S}^{3}{\rm MKL})$to deal with the first problem. The aim is to minimize an overall objective function composed of log-likelihood empirical loss, conditional expectation consensus (CEC) on the unlabeled data and group LASSO regularization on model coefficients. We further adapt CEC into a group-wise formulation so as to better deal with the intrinsic visual property of real-world images. We propose a fast block coordinate gradient descent method with several acceleration techniques for model solution. Compared with previous approaches, our model better makes use of large scale unlabeled images with multiple feature representation with lower time complexity. Moreover, to address the issue of reducing the risk of using unlabeled data, we design a multiple kernel hashing scheme to identify the “informative” and “compact” unlabeled training data subset. Comprehensive experiments are conducted and the results show that the proposed learning framework provides promising power for real-world image applications, such as image categorization and personalized Web image re-ranking with very little user interaction. Shuhui Wang, Qingming Huang, Shuqiang Jiang, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2011 | Efficient lp-norm multiple feature metric learning for image categorizationabstractPrevious metric learning approaches are only able to learn the metric based on single concatenated multivariate feature representation. However, for many real world problems with multiple feature representation such as image categorization, the model trained by previous approaches will degrade because of sparsity brought by significant dimension growth and uncontrolled influence from each feature channel. In this paper, we propose an efficient distance metric learning model which adapts Distance Metric Learning on multiple feature representations. The aim is to learn the Mahalanobis matrices for each independent feature and their non-sparse lp-norm weight coefficients simultaneously by maximizing the margin of the overall learned distance metric among the pairs from the same class and the distance of pairs from different classes. We further extend this method to nonlinear kernel learning and category specific metric learning, which demonstrate the applicability of using many existing kernels for image data and exploring the hierarchical semantic structures for large scale image datasets. Experiments on various datasets demonstrate the promising power of our method. Shuhui Wang, Qingming Huang, Shuqiang Jiang, Qi Tian 0001 |
CIKM | 1 |
| 2011 | Shape and location design of supporting legs for a new Water Strider RobotabstractIn this paper, the problems are discussed for shape design and position arrangement for Water Strider Robot's supporting legs are discussed. A supporting leg is approached as Euler-Bernoulli elastic curved beam and a method for designing its optimal shape is proposed by analysing elastic deformation and stress-strain. The objective of the proposed optimal method is to attain the maximum lift force in leg operation. The effectiveness and validity of design results are verified through simulations and lab experiments. A method for properly locating supporting legs on the robot body is proposed by analysing the influence of leg location to lift force and the relationship of supporting legs' with robot's roll-resistant capability. A layout scheme for the Water Dancer II-a prototype with ten supporting legs is presented with its operation successful designed. Licheng Wu, Shuhui Wang, Marco Ceccarelli, Haiwen Yuan, Guosheng Yang |
IROS | 2 |
| 2010 | Multiple Kernel Learning with High Order KernelsabstractPrevious Multiple Kernel Learning approaches (MKL) employ different kernels by their linear combination. Though some improvements have been achieved over methods using single kernel, the advantages of employing multiple kernels for machine learning are far from being fully developed. In this paper, we propose to use “high order kernels” to enhance the learning of MKL when a set of original kernels are given. High order kernels are generated by the products of real power of the original kernels. We incorporate the original kernels and high order kernels into a unified localized kernel logistic regression model. To avoid over-fitting, we apply group LASSO regularization to the kernel coefficients of each training sample. Experiments on image classification prove that our approach outperforms many of the existing MKL approaches. Shuhui Wang, Shuqiang Jiang, Qingming Huang, Qi Tian 0001 |
ICPR | 1 |
| 2010 | Nearest-neighbor classification using unlabeled data for real world image applicationabstractCurrently, Nearest-Neighbor approaches (NN) have been widely applied to real world image data mining. These approaches have the following three disadvantages: (i) the performance is inferior on small datasets; (ii) the performance of approximated nearest neighbor search will degrade for data with high dimensions; (iii) they are heavily dependent on the chosen feature and distance measure. To overcome these intrinsic weaknesses, we propose a novel Nearest-Neighbor method, which improves the original NN approaches from three aspects. Firstly, we propose a novel neighborhood similarity measure, where the similarity between test images and labeled images in the database is calculated jointly by the original image-to-image similarity and the average similarity of their neighboring unlabeled data. Secondly, we adopt the kernelized locality sensitive hashing to effectively conduct the nearest neighbor search for high dimensional data. Finally, to enhance the robustness of the method on different genres of images, we propose to fuse the discrimination power of different features by considering all the retrieved nearest neighbors via hashing systems using different features/kernels. Experimental result shows the advantage over traditional Nearest-Neighbor methods using the labeled data only. Even when the ratio of labeled data is very small, our method could also achieve remarkable results, thanks to the help of unlabeled data and multiple features. Shuhui Wang, Qingming Huang, Shuqiang Jiang, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2010 | S3MKL: scalable semi-supervised multiple kernel learning for image data miningabstractFor large scale image data mining, a challenging problem is to design a method that could work efficiently under the situation of little ground-truth annotation and a mass of unlabeled or noisy data. As one of the major solutions, semi-supervised learning (SSL) has been deeply investigated and widely used in image classification, ranking and retrieval. However, most SSL approaches are not able to incorporate multiple information sources. Furthermore, no sample selection is done on unlabeled data, leading to the unpredictable risk brought by uncontrolled unlabeled data and heavy computational burden that is not suitable for learning on real world dataset. In this paper, we propose a scalable semi-supervised multiple kernel learning method (S3MKL) to deal with the first problem. Our method imposes group LASSO regularization on the kernel coefficients to avoid over-fitting and conditional expectation consensus for regularizing the behaviors of different kernel on the unlabeled data. To reduce the risk of using unlabeled data, we also design a hashing system where multiple kernel locality sensitive hashing (MKLSH) are constructed with respect to different kernels to identify a set of "informative" and "compact" unlabeled training subset from a large unlabeled data corpus. Combining S3MKL with MKLSH, the method is suitable for real world image classification and personalized web image re-ranking with very little user interaction. Comprehensive experiments are conducted to test the performance of our method, and the results show that our method provides promising powers for large scale real world image classification and retrieval. Shuhui Wang, Shuqiang Jiang, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2009 | Cloudlet-screen computing: A multi-core-based, cloud-computing-oriented, traditional-computing-compatible parallel computing Paradigm for the massesabstractThis paper proposes a computing paradigm where many users share a host platform consisting of one or more multicore CPU(s) and GPU(s). Each user is connected to the host by a link transferring primarily compressed compound video of screen, mouse and keyboard data. All computation and processing tasks including but not limited to 2D/3D graphics and multimedia processing are done on the host using a traditional or existing architecture. Since all computing and storage resources are centralized on the "cloudlet" (host) and each user site has only a high-resolution monitor, a mouse, and a keyboard with some optional I/O devices, it is a highly energy-efficient, secure, reliable, and low total-ownership-cost paradigm. The compound video link between the host and the user is partial-lossless (partially lossless and partially near-lossless), and has ultra-low latency, so the proposed paradigm still has no-compromised 2D/3D graphics and multimedia performance demanded by the mass market. Tao Lin 0005, Shuhui Wang |
ICME | 2 |
| 2008 | Shot classification for action movies based on motion characteristicsabstractIn this paper, we propose a shot classification method for action movies. Considering that motion characteristic is very important for semantic movie analysis, and it contains abundant information in action movies, the structure tensor analysis is used for feature extraction due to its capability of representing both spatial and temporal characteristics of a shot. Firstly, the movie shots with known labels are decomposed into a set of overlapped fixed-length segments and their structure tensor histogram are computed. The labels of segments are identical to the shots they belong to. Then Adaboost is used to train the semantic classifier with these structure tensor histogram sets. In testing procedure, the unknown shot are decomposed in the same way, and feature vector of each segment is extracted and classified by the classifier. Finally, the label of the shot is generalized by the segment label voting scheme. Experimental results show that this scheme could effectively deal with multiple motion patterns within shots and promising results are achieved. Shuhui Wang, Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICIP | 1 |