Zhou Yu 0001

dblp:83/3205-1 · DBLP profile ↗
← Back
59ranked-venue papers
15as first author
38since 2021 · last 2026
0000-0001-8407-1137ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 9 first-author · 23 since 2021Artificial intelligence and machine learning · 28 · 9 first-author · 18 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images
abstract
Feed-forward 3D reconstruction from sparse, low-resolution (LR) images is a crucial capability for real-world applications, such as autonomous driving and embodied AI. However, existing methods often fail to recover fine texture details. This limitation stems from the inherent lack of high-frequency information in LR inputs. To address this, we propose SRSplat, a feed-forward framework that reconstructs high-resolution 3D scenes from only a few LR views. Our main insight is to compensate for the deficiency of texture information by jointly leveraging external high-quality reference images and internal texture cues. We first construct a scene-specific reference gallery, generated for each scene using Multimodal Large Language Models (MLLMs) and diffusion models. To integrate this external information, we introduce the Reference-Guided Feature Enhancement (RGFE) module, which aligns and fuses features from the LR input images and their reference twin image. Subsequently, we train a decoder to predict the Gaussian primitives using the multi-view fused feature obtained from RGFE. To further refine predicted Gaussian primitives, we introduce Texture-Aware Density Control (TADC), which adaptively adjusts Gaussian density based on the internal texture richness of the LR inputs. Extensive experiments demonstrate that our SRSplat outperforms existing methods on various datasets, including RealEstate10K, ACID, and DTU, and exhibits strong cross-dataset and cross-resolution generalization capabilities.
Changyue Shi, Chuxiao Yang, Jiajun Ding, Zhou Yu 0001, Min Tan 0005
AAAI8
2026 Sparse4DGS: 4D Gaussian Splatting for Sparse-Frame Dynamic Scene Reconstruction
abstract
Dynamic Gaussian Splatting approaches have achieved remarkable performance for 4D scene reconstruction. However, these approaches rely on dense-frame video sequences for photorealistic reconstruction. In real-world scenarios, due to equipment constraints, sometimes only sparse frames are accessible. In this paper, we propose Sparse4DGS, the first method for sparse-frame dynamic scene reconstruction. We observe that dynamic reconstruction methods fail in both canonical and deformed spaces under sparse-frame settings, especially in areas with high texture richness. Sparse4DGS tackles this challenge by focusing on texture-rich areas. For the deformation network, we propose Texture-Aware Deformation Regularization, which introduces a texture-based depth alignment loss to regulate Gaussian deformation. For the canonical Gaussian field, we introduce Texture-Aware Canonical Optimization, which incorporates texture-based noise into the gradient descent process of canonical Gaussians. Extensive experiments show that when taking sparse frames as inputs, our method outperforms existing dynamic or few-shot techniques on NeRF-Synthetic, HyperNeRF, NeRF-DS, and our iPhone-4D datasets.
Changyue Shi, Chuxiao Yang, Wenwen Pan 0003, Jiajun Ding, Zhou Yu 0001, Jun Yu 0002
AAAI8
2026 KF-GS: Kalman filter-guided Gaussian splatting for real-time high-quality dynamic scene reconstruction
Qingyuan Tang, Yufei Yin, Yanming Zhu 0001, Zhou Yu 0001, Zhenzhong Kuang, Jiajun Ding, Jifa He
J. Vis. Commun. Image Represent.4
2026 GC-GS: Gradient control Gaussian splatting with various image degradation
Qida Cao, Jiajun Ding, Qingyuan Tang, Tianning Zhao, Xiaoling Gu, Jianping Fan 0001, Zhou Yu 0001
Pattern Recognit.7
2026 Emotional conflict adaptation for multimodal sentiment analysis
Tingting Han 0003, Lingyun Yu 0004, Min Tan 0005, Zhou Yu 0001, Hongxun Yao
Pattern Recognit.4
2026 HEART: Emotionally Grounded Video Captioning via Hierarchical Emotion-Aligned Representation
abstract
Emotional Video Captioning (EVC) seeks to generate video descriptions that are both factually accurate and emotionally expressive. However, existing approaches often lack structured semantic grounding and fine-grained temporal modeling, leading to incomplete or emotionally inconsistent captions. To address these issues, we proposeHEART(HierarchicalEmotion-AlignedRepresentation withTemporal structure), a unified framework that jointly models hierarchical visual semantics and multi-scale temporal context. Specifically, HEART introduces a Hierarchical Semantic Extraction Module that decomposes visual content into entity-, action-, and event-level representations, providing a rich foundation for multi-level emotional alignment. A Temporal Pyramid Module captures short- and long-range temporal dependencies through multi-scale convolution, enabling temporally coherent captioning. Together, these components enable HEART to generate captions that are both emotionally grounded and temporally complete. To support this framework, we construct EmoStruct, a new benchmark dataset with fine-grained emotional annotations at the subject and predicate levels. Experiments on EmoStruct and public datasets demonstrate that HEART significantly outperforms prior methods in both semantic and emotional dimensions.
Tingting Han 0003, Yuxuan Gong, Sicheng Zhao, Min Tan 0005, Zhou Yu 0001, Hongxun Yao
IEEE Trans. Affect. Comput.5
2026 TailorEdit: An Adaptive Framework for Instruction-Guided Fashion Image Editing
abstract
Fashion image editing has garnered significant attention due to its growing demand in e-commerce, social media, and virtual try-on applications. However, existing methods are typically designed for specific editing tasks in isolation, lacking a unified framework capable of handling diverse editing requirements. This work addresses this limitation from two critical perspectives. First, we constructInstructFashion, a large-scale, high-quality dataset specifically curated for instruction-guided fashion image editing. It is generated through carefully designed pipelines that cover four distinct editing tasks. Second, we proposeTailorEdit, an adaptive framework for instruction-guided fashion image editing. It integrates human segmentation map-based denoising guidance, modular LoRA-based editing experts, and a dynamic expert routing mechanism to enable precise and semantically coherent modifications. Extensive quantitative and qualitative evaluations demonstrate that TailorEdit consistently outperforms state-of-the-art methods in terms of realism, coherence, and instruction adherence. Our code is available at https://github.com/EndaJude/TailorEdit.
Xiaoling Gu, Lingda Zhu, Yongkang Wong, Zhou Yu 0001, Huan Li 0003, Zizhao Wu, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.4
2026 Fuzzy Language Gaussian Splatting
abstract
Recent advancements in open-vocabulary 3D querying have achieved remarkable progress. However, existing approaches such as LERF and LangSplat still rely heavily on specific query text for accurate 3D target identification. They fail to accurately comprehend fuzzy query text (which describes a target's properties, functions, or traits rather than naming it directly), causing localization mistakes and poor 3D segmentation results. In this work, we introduce a novel task—open-vocabulary 3D fuzzy query—which aims to locate and generate precise 3D target masks based on fuzzy query text, a capability that is crucial for building a 3D intelligent system. To address this challenge, we proposeFuzzy Language Gaussian Splatting (FL-GS), a framework consisting of three key stages: We first leverage a multimodal large language model (MLLM) to identify and localize potential targets in representative views based on fuzzy query text, thereby generating initial masks using the segment anything model (SAM). Subsequently, we compute the pairwise similarity between all target masks and construct an undirected, unweighted graph, from which the maximum clique is identified, thus obtaining the most probable correct masks, referred to as refined masks. Finally, we propose single-dimensional mask encoding to efficiently achieve precise 3D target masks through supervision with refined masks. Furthermore, we manually annotated two key datasets and established a benchmark for this new task. Experimental results clearly demonstrate that FL-GS outperforms existing methods in open-vocabulary 3D fuzzy querying.
Jiajun Ding, Yaowei Liu, Hongxi Zhu, Min Tan 0005, Zhou Yu 0001
IEEE Trans. Fuzzy Syst.7
2026 Compositional Text-to-Image Synthesis With Training-Free Layout-Guided Diffusion
abstract
Recent text-to-image (T2I) diffusion models have made significant strides in generating high-quality images from diverse textual prompts. Despite this progress, these models often face challenges in accurately understanding and synthesizing complex prompts, primarily due to their limited compositional capabilities. In this study, we propose a novel approach for compositional T2I synthesis using layout-guided diffusion models, which do not require additional training. Specifically, we leverage the chain-of-code prompting technique of large language models to interpret textual prompts and generate object layouts with spatial coherence. To enhance the alignment between generated images and textual descriptions, we introduce two innovative layoutguided loss functions: Patch-oriented Cross-Attention (PCA) loss and Region-oriented Cross-Attention (RCA) loss. The PCA loss emphasizes high activation values for image patches that attend to all tokens in the prompt across the layout. The RCA loss enhances the average attention within the layout, thereby increasing the accuracy of generating objects and their associated attributes within specified regions. These proposed loss functions reassign cross-attention in diffusion models during the denoising process. Our comprehensive experiments consistently demonstrate the effectiveness of our approach in improving semantic alignment between generated images and a diverse range of textual prompts, while ensuring high usability as a ready-to-use plugin. Our code is available athttps://github.com/gxl-groups/Compositional-T2I.
Xiaoling Gu, Lingwei Luo, Shengqi Wu, Zizhao Wu, Zhenzhong Kuang, Zhou Yu 0001
IEEE Trans. Multim.6
2025 Growing a Twig to Accelerate Large Vision-Language Models
Zhenwei Shao, Zhou Yu 0001, Wenwen Pan 0003, Hongyuan Zhang 0001, Wei Chen 0001, Jun Yu 0002
ICCV3
2025 MMCNav: MLLM-empowered Multi-agent Collaboration for Outdoor Visual Language Navigation
Suguo Zhu, Tingting Han 0003, Zhou Yu 0001
ICMR5
2025 DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency Regularization
abstract
The rapid evolution of the online fashion industry has intensified the demand for interactive fashion retrieval systems capable of precise and flexible searches based on user-specified attribute modifications. However, prevailing fashion retrieval methods often overlook the distinctive distributional properties of fashion images and struggle to preserve semantic consistency during attribute manipulation. To address these limitations, we propose DiSCo, a novel disentangled attribute manipulation retrieval framework via semantic reconstruction and consistency regularization. Our approach comprises three key components: (1) An attribute-aware manipulation network that constructs target fashion embeddings through cross-modal attribute modification deltas, leveraging dedicated fashion attribute encoders; (2) A cross-modal semantic reconstruction network that synthesizes target images directly from modified attribute descriptions, supervised by adversarial and attribute classification losses to ensure interpretable edits; (3) An adaptive fusion mechanism that dynamically integrates attribute-modified embeddings with reconstructed image features. Extensive evaluations on two benchmark datasets (DeepFashion and Shopping100K) demonstrate that DiSCo achieves superior retrieval accuracy over state-of-the-arts while maintaining high-fidelity editing. Quantitative and qualitative analyses further confirm that DiSCo generates more realistic fashion representations, underscoring its effectiveness in attribute-aware retrieval tasks.
Min Tan 0005, Guanhao Liu, Huijing Zhan, Yuyu Yin, Zhou Yu 0001, Jiajun Ding, Yinfu Feng
ACM Multimedia5
2025 Modality-aware contrast and fusion for multi-modal summarization
Lixin Dai, Tingting Han 0003, Zhou Yu 0001, Jun Yu 0002, Min Tan 0005
Neurocomputing3
2025 Prophet: Prompting Large Language Models With Complementary Answer Heuristics for Knowledge-Based Visual Question Answering
abstract
Knowledge-based visual question answering (VQA) requires external knowledge beyond the image to answer the question. Early studies retrieve required knowledge from explicit knowledge bases (KBs), which often introduces irrelevant information to the question, hence restricting the performance of their models. Recent works have resorted to using a powerful large language model (LLM) as an implicit knowledge engine to acquire the necessary knowledge for answering. Despite the encouraging results achieved by these methods, we argue that they have not fully activated the capacity of the LLM as the provided textual input is insufficient to depict the required visual information to answer the question. In this paper, we present Prophet-a conceptually simple, flexible, and general framework designed to prompt LLM with answer heuristics for knowledge-based VQA. Specifically, we first train a vanilla VQA model on a specific knowledge-based VQA dataset without external knowledge. After that, we extract two types of complementary answer heuristics from the VQA model: answer candidates and answer-aware examples. The two types of answer heuristics are jointly encoded into a formatted prompt to facilitate the LLM's understanding of both the image and question, thus generating a more accurate answer. By incorporating the state-of-the-art LLM GPT-3 (Brown et al. 2020), Prophet significantly outperforms existing state-of-the-art methods on four challenging knowledge-based VQA datasets. Prophet is general that can be instantiated with the combinations of different VQA models (i.e., both discriminative and generative ones) and different LLMs (i.e., both commercial and open-source ones). Moreover, Prophet can also be integrated with modern large multimodal models in different stages, which is named Prophet++, to further improve the capabilities on knowledge-based VQA tasks.
Zhou Yu 0001, Xuecheng Ouyang, Zhenwei Shao, Meng Wang 0001, Jun Yu 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Action-Driven Semantic Representation and Aggregation for Video Captioning
abstract
Video captioning, a challenging task that entails generating natural language descriptions of visual content, often fails to effectively grasp the essence of action semantics. To harness the power of action detection to facilitate a deeper understanding of the video content, we propose an action-driven method, named Hierarchical Semantic Representation and Aggregation (HSRA) network. This method explicitly exploits action clues with a hierarchical semantic representation module, which models visual semantics in a three-level structure: “object-action-event”. By employing learnable action queries, our approach injects extensive action semantics into the model, thereby enabling more accurate and context-rich captions. To further enhance semantic alignment and understanding, we introduce a semantic aggregation composed of a semantic interaction module and a semantic refinement module. This component facilitates the alignment of semantics across different levels and emphasizes key information, ultimately leading to significant improvements in semantic consistency between the video and generated captions. We performed extensive evaluations on two well-established public datasets, MSVD and MSR-VTT, and the findings consistently demonstrate that our proposed HSRA network outperforms contemporary state-of-the-art methods.
Tingting Han 0003, Yaochen Xu, Jun Yu 0002, Zhou Yu 0001, Sicheng Zhao
IEEE Trans. Circuits Syst. Video Technol.4
2025 Benchmarking and Enhancing Geospatial Visual Reasoning Over Street Maps
abstract
Recent advances in large multimodal models (LMMs) have enabled substantial progress in various visual question answering (VQA) benchmarks, including the challenging text-centric ones that require a simultaneous understanding of both the visual and textual contents in the images. Despite the prominence of existing text-centric VQA benchmarks, they either have limited textual information or have a limited number of questions requiring complex reasoning skills beyond the basic OCR. To this end, we present SMVQA—a novel text-centric VQA benchmark based on street map images. SMVQA contains more than 10K real-world street map images from the open geospatial database OpenStreetMap. Each image in SMVQA is also associated with detailed geospatial annotations, enabling it to automatically generate up to 57.5K distinctive QA pairs of five representative question types. In addition to the standard test split, SMVQA introduces an extra test split to verify the generalization abilities over out-of-domain images and novel reasoning skills. The evaluation of the state-of-the-art open-source and commercial LMMs reflects the great challenge posed by SMVQA. The latest LMMs, such as GPT-4o, only achieve accuracies of 49.9%, showing plenty of room for improvement. To further improve the latest LMMs’ performance on SMVQA, we introduce a LMM-based agentic framework LHR, which consists of the localizing, highlighting, and reasoning stages. Specifically, LHR first prompts the LMM to localize region-of-interest (RoI) to the question and then highlight the RoI and perform chain-of-thought reasoning for answer prediction. By integrating LHR with GPT-4o, we observe a significant improvement over the vanilla counterpart, showing the effectiveness of our framework.
Wenwen Pan 0003, Haiting Zhou, Zhenwei Shao, Shuai Shao 0012, Suguo Zhu, Min Tan 0005, Jun Yu 0002, Zhou Yu 0001
IEEE Trans. Geosci. Remote. Sens.8
2025 ScatDiff: Physical Diffusion Model for Electromagnetic Computational Imaging
abstract
Electromagnetic computational imaging offers a promising solution to electromagnetic inverse scattering problems. Whereas, it is challenged by its ill-posed nature and non-linearity. Traditional iterative methods are often slow and prone to local minima, while recent deep generative models overlook the physical principles that govern the transformation from scattering fields to images of constitute parameters, limiting their interpretability, generalization, and robustness. To address these issues, we propose ScatDiff, a novel Scatter-to-image Diffusion model that integrates electromagnetic data with fundamental physical principles. ScatDiff uses a time-aware, backpropagation-enhanced diffusion to generate noisy images embedded with electromagnetic priors, along with a denoising module that uses cross-attention to adaptively integrate scattering fields. Additionally, a physics-driven reconstruction module incorporates an induced current model into the loss function to enhance interpretability. Experiments on three MNIST variants, Gesture dataset, the “Austria” profile, and the “FoamDielExt” profile show that ScatDiff outperforms traditional iterative methods and deep learning models in both imaging quality and efficiency, with strong generalization and robustness under high noise conditions. Code and datasets are available on https://github.com/Scatdif.
Min Tan 0005, Kuiwen Xu, Zhou Yu 0001, Jun Yu 0002
IEEE Trans. Geosci. Remote. Sens.6
2025 Spatio-Temporal and Retrieval-Augmented Modeling for Chest X-Ray Report Generation
abstract
Chest X-ray report generation has attracted increasing research attention. However, most existing methods neglect the temporal information and typically generate reports conditioned on a fixed number of images. In this paper, we propose STREAM: Spatio-Temporal and REtrieval-Augmented Modelling for automatic chest X-ray report generation. It mimics clinical diagnosis by integrating current and historical studies to interpret the present condition (temporal), with each study containing images from multi-views (spatial). Concretely, our STREAM is built upon an encoder-decoder architecture, utilizing a large language model (LLM) as the decoder. Overall, spatio-temporal visual dynamics are packed as visual prompts and regional semantic entities are retrieved as textual prompts. First, a token packer is proposed to capture condensed spatio-temporal visual dynamics, enabling the flexible fusion of images from current and historical studies. Second, to augment the generation with existing knowledge and regional details, a progressive semantic retriever is proposed to retrieve semantic entities from a preconstructed knowledge bank as heuristic text prompts. The knowledge bank is constructed to encapsulate anatomical chest X-ray knowledge into structured entities, each linked to a specific chest region. Extensive experiments on public datasets have shown the state-of-the-art performance of our method. Related codes and the knowledge bank are available at https://github.com/yangyan22/STREAM.
Xiaoxing You, Ke Zhang 0029, Zhenqi Fu, Xianyun Wang, Jiajun Ding, Jiamei Sun, Zhou Yu 0001, Qingming Huang, Weidong Han 0001, Jun Yu 0002
IEEE Trans. Medical Imaging8
2025 Imp: Highly Capable Large Multimodal Models for Mobile Devices
abstract
By harnessing the capabilities of large language models (LLMs), recent large multimodal models (LMMs) have shown remarkable versatility in open-world multimodal understanding. Nevertheless, they are usually parameter-heavy and computation-intensive, thus hindering their applicability in resource-constrained scenarios. To this end, several lightweight LMMs have been proposed successively to maximize the capabilities under constrained scale (e.g., 3B). Despite the encouraging results achieved by these methods, most of them only focus on one or two aspects of the design space, and the key design choices that influence model capability have not yet been thoroughly investigated. In this paper, we conduct a systematic study for lightweight LMMs from the aspects of model architecture, training strategy, and training data. Based on our findings, we obtain Imp—a family of highly capable LMMs at the 2B$\sim$4B scales. Notably, our Imp-3B model steadily outperforms all the existing lightweight LMMs of similar size, and even surpasses the state-of-the-art LMMs at the 13B scale. With low-bit quantization and resolution reduction techniques, our Imp model can be deployed on a Qualcomm Snapdragon 8Gen3 mobile chip with a high inference speed of about 13 tokens/s.
Zhenwei Shao, Zhou Yu 0001, Jun Yu 0002, Xuecheng Ouyang, Lihao Zheng 0001, Zhenbiao Gai, Zhenzhong Kuang, Jiajun Ding
IEEE Trans. Multim.2
2024 MPOD123: One Image to 3D Content Generation Using Mask-Enhanced Progressive Outline-to-Detail Optimization
abstract
Recent advancements in single image driven 3D content generation have been propelled by leveraging prior knowledge from pretrained 2D diffusion models. However, the 3D content generated by existing methods often exhibits distorted outline shapes and inadequate details. To solve this problem, we propose a novel framework called Mask-enhanced Progressive Outline-to-Detail optimization (aka. MPOD123), which consists of two stages. Specifically, in the first stage, MPOD123 utilizes the pretrained view-conditioned diffusion model to guide the outline shape optimization of the 3D content. Given certain viewpoint, we estimate outline shape priors in the form of 2D mask from the 3D content by leveraging opacity calculation. In the second stage, MPOD123 incorporates Detail Appearance Inpainting (DAI) to guide the refinement on local geometry and texture with the shape priors. The essence of DAI lies in the Mask Rectified Cross-Attention (MRCA), which can be conveniently plugged in the stable diffusion model. The MRCA module utilizes the mask to rectify the attention map from each cross-attention layer. Accompanied with this new module, DAI is capable of guiding the detail refinement of the 3D content, while better preserves the outline shape. To assess the applicability in practical scenarios, we contribute a new dataset modeled on real-world e-commerce environments. Extensive quantitative and qualitative experiments on this dataset and open benchmarks demonstrate the effectiveness of MPOD123 over the state-of-the-arts.
Jimin Xu, Tianbao Wang, Tao Jin 0004, Shengyu Zhang 0001, Jiangjing Lyu, Chengfei Lv, Chaoyue Niu, Zhou Yu 0001, Zhou Zhao 0001, Fei Wu 0001
CVPR10
2024 Benchmarking Geospatial Visual Reasoning over Street Map Images
abstract
In this paper, we present a novel VQA benchmark SM-VQA, which is built upon street map images. Specifically, SM-VQA contains about 9.5K real-world street map images collected from the open geospatial database OpenStreetMap. Each image in SM-VQA is also associated with detailed geospatial annotations, enabling it to automatically generate up to 50K distinctive QA pairs of five types of geospatial reasoning abilities. The evaluation of the state-of-the-art open-source and commercial LMMs reflects the great challenge posed by SM-VQA. The cutting-edge Phi-3V and GPT-4o models merely achieve accuracies of 45.3% and 46.7% respectively.
Haiting Zhou, Zhou Yu 0001
CW2
2024 3D Question Answering with Scene Graph Reasoning
Zizhao Wu, Haohan Li, Gongyi Chen, Zhou Yu 0001, Xiaoling Gu, Yigang Wang
ACM Multimedia4
2024 Confidence correction for trained graph convolutional networks
abstract
Adopting Graph Convolutional Networks (GCNs) for transductive node classification is a hot research direction in artificial intelligence . Vanilla GCNs are primarily under-confident and struggle to clarify the final classification results explicitly due to the lack of supervision. Existing works mainly alleviated this issue by improving annotation deficiency and introducing addition regularization terms. However, these methods need to re-train the model from the beginning, which is computationally expensive for large dataset and model. To deal with this problem, a novel confidence correction mechanism (CCM) for trained GCNs is proposed in this work. Such mechanism aims at calibrating the confidence output of each node in the inference stage by jointly inferring the feature and predicted pseudo label. Specifically, in the inference stage, it uses the predicted pseudo label to select target-related features over all network to obtain a more confident and better result. Such selectivity is formulated as an optimization problem to maximize the category score of each node. In addition, the greedy optimization strategy is utilized to solve this problem and we have mathematically proven that the proposed mechanism can reach the local optimum by mathematical induction . Note that such mechanism is flexible and can be introduced to most GCN-based model. Extensive experimental results on benchmark datasets show that the proposed method can promote the confidence of the final target category and improve the performance of GCNs in the inference stage.
Junqing Yuan, Huanlei Guo, Chenyi Zhou, Jiajun Ding, Zhenzhong Kuang, Zhou Yu 0001
Pattern Recognit.6
2024 Effective Video Summarization by Extracting Parameter-Free Motion Attention
abstract
Video summarization remains a challenging task despite increasing research efforts. Traditional methods focus solely on long-range temporal modeling of video frames, overlooking important local motion information that cannot be captured by frame-level video representations. In this article, we propose the Parameter-free Motion Attention Module (PMAM) to exploit the crucial motion clues potentially contained in adjacent video frames, using a multi-head attention architecture. The PMAM requires no additional training for model parameters, leading to an efficient and effective understanding of video dynamics. Moreover, we introduce the Multi-feature Motion Attention Network (MMAN), integrating the PMAM with local and global multi-head attention based on object-centric and scene-centric video representations. The synergistic combination of local motion information, extracted by the proposed PMAM, with long-range interactions modeled by the local and global multi-head attention mechanism, can significantly enhance the performance of video summarization. Extensive experimental results on the benchmark datasets, SumMe and TVSum, demonstrate that the proposed MMAN outperforms other state-of-the-art methods, resulting in remarkable performance gains.
Tingting Han 0003, Jun Yu 0002, Zhou Yu 0001, Sicheng Zhao
ACM Trans. Multim. Comput. Commun. Appl.4
2023 ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed Videos
abstract
Building benchmarks to systemically analyze different capabilities of video question answering (VideoQA) models is challenging yet crucial. Existing benchmarks often use non-compositional simple questions and suffer from language biases, making it difficult to diagnose model weaknesses incisively. A recent benchmark AGQA [8] poses a promising paradigm to generate QA pairs automatically from pre-annotated scene graphs, enabling it to measure diverse reasoning abilities with granular control. However, its questions have limitations in reasoning about the fine-grained semantics in videos as such information is absent in its scene graphs. To this end, we present ANetQA, a large-scale benchmark that supports fine-grained compositional reasoning over the challenging untrimmed videos from ActivityNet [4]. Similar to AGQA, the QA pairs in ANetQA are automatically generated from annotated video scene graphs. The fine-grained properties of ANetQA are reflected in the following: (i) untrimmed videos with fine-grained semantics; (ii) spatio-temporal scene graphs with fine-grained taxonomies; and (iii) diverse questions generated from fine-grained templates. ANetQA attains 1.4 billion unbalanced and 13.4 million balanced QA pairs, which is an order of magnitude larger than AGQA with a similar number of videos. Comprehensive experiments are performed for state-of-the-art methods. The best model achieves 44.5% accuracy while human performance tops out at 84.5%, leaving sufficient room for improvement.
Zhou Yu 0001, Lixiang Zheng, Zhou Zhao 0001, Fei Wu 0001, Jianping Fan 0001, Kui Ren 0001, Jun Yu 0002
CVPR1
2023 Prompting Large Language Models with Answer Heuristics for Knowledge-Based Visual Question Answering
abstract
Knowledge-based visual question answering (VQA) requires external knowledge beyond the image to answer the question. Early studies retrieve required knowledge from explicit knowledge bases (KBs), which often introduces irrelevant information to the question, hence restricting the performance of their models. Recent works have sought to use a large language model (i.e., GPT-3 [3]) as an implicit knowledge engine to acquire the necessary knowledge for answering. Despite the encouraging results achieved by these methods, we argue that they have not fully activated the capacity of GPT-3 as the provided input information is insufficient. In this paper, we present Prophet-a conceptually simple framework designed to$prompt$GPT-3 with answer heuristics for knowledge-based VQA. Specifically, we first train a vanilla VQA model on a specific knowledge-based VQA dataset without external knowledge. After that, we extract two types of complementary answer heuristics from the model: answer candidates and answer-aware examples. Finally, the two types of answer heuristics are encoded into the prompts to enable GPT-3 to better comprehend the task thus enhancing its capacity. Prophet significantly outperforms all existing state-of-the-art methods on two challenging knowledge-based VQA datasets, OK-VQA and A-OKVQA, delivering 61.1% and 55.7% accuracies on their testing sets, respectively.
Zhenwei Shao, Zhou Yu 0001, Meng Wang 0001, Jun Yu 0002
CVPR2
2023 Parameter-Efficient Transfer Learning for Audio-Visual-Language Tasks
abstract
The pretrain-then-finetune paradigm has been widely used in various unimodal and multimodal tasks. However, finetuning all the parameters of a pre-trained model becomes prohibitive as the model size grows exponentially. To address this issue, the adapter mechanism that freezes the pre-trained model and only finetunes a few extra parameters is introduced and delivers promising results. Most studies on adapter architectures are dedicated to unimodal or bimodal tasks, while the adapter architectures for trimodal tasks have not been investigated yet. This paper introduces a novel Long Short-Term Trimodal Adapter (LSTTA) approach for video understanding tasks involving audio, visual, and language modalities. Based on the pre-trained from the three modalities, the designed adapter module is inserted between the sequential blocks to model the dense interactions across the three modalities. Specifically, LSTTA consists of two types of complementary adapter modules, namely the long-term semantic filtering module and the short-term semantic interaction module. The long-term semantic filtering aims to characterize the temporal importance of the video frames and the short-term semantic interaction module models local interactions within short periods. Compared to previous state-of-the-art trimodal learning methods pre-trained on a large-scale trimodal corpus, LSTTA is more flexible and can inherit any powerful unimodal or bimodal models. Experimental results on four typical trimodal learning tasks show the effectiveness of LSTTA over existing state-of-the-art methods.
Hongye Liu, Xianhai Xie, Zhou Yu 0001
ACM Multimedia4
2023 Contrastive Perturbation Network for Weakly Supervised Temporal Sentence Grounding
Tingting Han 0003, Yuanxin Lv, Zhou Yu 0001, Jun Yu 0002, Jianping Fan 0001
PRCV (1)3
2023 MARN: Multi-level Attentional Reconstruction Networks for Weakly Supervised Video Temporal Grounding
Yijun Song, Jingwen Wang 0003, Lin Ma 0002, Jun Yu 0002, Jinxiu Liang, Zhou Yu 0001
Neurocomputing7
2023 Bilaterally Slimmable Transformer for Elastic and Efficient Visual Question Answering
abstract
Recent advances in Transformer architectures [1] have brought remarkable improvements to visual question answering (VQA). Nevertheless, Transformer-based VQA models are usually deep and wide to guarantee good performance, so they can only run on powerful GPU servers and cannot run on capacity-restricted platforms such as mobile phones. Therefore, it is desirable to learn an elastic VQA model that supports adaptive pruning at runtime to meet the efficiency constraints of different platforms. To this end, we present the bilaterally slimmable Transformer (BST), a general framework that can be seamlessly integrated into arbitrary Transformer-based VQA models to train a single model once and obtain various slimmed submodels of different widths and depths. To verify the effectiveness and generality of this method, we integrate the proposed BST framework with three typical Transformer-based VQA approaches, namely MCAN [2], UNITER [3], and CLIP-ViL [4], and conduct extensive experiments on two commonly-used benchmark datasets. In particular, one slimmed MCAN$_\mathsf {BST}$submodel achieves comparable accuracy on VQA-v2, while being 0.38× smaller in model size and having 0.27× fewer FLOPs than the reference MCAN model. The smallest MCAN$_\mathsf {BST}$submodel only has 9M parameters and 0.16G FLOPs during inference, making it possible to deploy it on a mobile device with less than 60 ms latency.
Zhou Yu 0001, Zitian Jin, Jun Yu 0002, Mingliang Xu 0001, Jianping Fan 0007
IEEE Trans. Multim.1
2022 Delegate-based Utility Preserving Synthesis for Pedestrian Image Anonymization
abstract
The rapidly growing application of pedestrian images has aroused wide concern on visual privacy protection because personal information is under the risk of privacy disclosure. Anonymization is regarded as an effective solution by identity obfuscation. Most recent methods focus on face, but it is not enough when the presence of human body carries lots of identifiable information. This paper presents a new delegate-based utility preserving synthesis (DUPS) approach for pedestrian image anonymization. This is challenging because one may expect that the anonymized image can still be useful in various computer vision tasks. We model DUPS as an adaptive translation process from source to target. To provide a comprehensive identity protection, we first perform anonymous delegate sampling based on image-level differential privacy. To synthesize anonymous images, we then introduce an adaptive translation network and optimize it with a multi-task loss function. Our approach is theoretically sound and can generate diverse results by preserving data utility. The experiments on multiple datasets show that DUPS can not only achieve superior anonymization performance against deep pedestrian recognizers, but also can obtain a better tradeoff between privacy protection and utility preservation compared with state-of-the-art methods.
Zhenzhong Kuang, Longbin Teng, Zhou Yu 0001, Jun Yu 0002, Jianping Fan 0001, Mingliang Xu 0001
ACM Multimedia3
2022 Question-relationship guided graph attention network for visual question answer
Liansheng Zhuang, Zhou Yu 0001, Tian Bai 0005
Multim. Syst.3
2022 Deep relational self-Attention networks for scene graph generation
Zhou Yu 0001, Yibing Zhan
Pattern Recognit. Lett.2
2021 ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge Integration
abstract
Vision-and-language pretraining (VLP) aims to learn generic multimodal representations from massive image-text pairs. While various successful attempts have been proposed, learning fine-grained semantic alignments between image-text pairs plays a key role in their approaches. Nevertheless, most existing VLP approaches have not fully utilized the intrinsic knowledge within the image-text pairs, which limits the effectiveness of the learned alignments and further restricts the performance of their models. To this end, we introduce a new VLP method called ROSITA, which integrates the cross- and intra-modal knowledge in a unified scene graph to enhance the semantic alignments. Specifically, we introduce a novel structural knowledge masking (SKM) strategy to use the scene graph structure as a priori to perform masked language (region) modeling, which enhances the semantic alignments by eliminating the interference information within and across modalities. Extensive ablation studies and comprehensive analysis verifies the effectiveness of ROSITA in semantic alignments. Pretrained with both in-domain and out-of-domain datasets, ROSITA significantly outperforms existing state-of-the-art VLP methods on three typical vision-and-language tasks over six benchmark datasets.
Yuhao Cui, Zhou Yu 0001, Chunqi Wang, Zhongzhou Zhao, Ji Zhang 0011, Meng Wang 0001, Jun Yu 0002
ACM Multimedia2
2021 Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems
abstract
Derek Chen, Howard Chen, Yi Yang, Alexander Lin, Zhou Yu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Derek Chen, Howard Chen 0003, Yi Yang 0001, Alexander Lin, Zhou Yu 0001
NAACL-HLT5
2021 Accelerated masked transformer for dense video captioning
Zhou Yu 0001, Nanjia Han
Neurocomputing1
2021 Long-Term Video Question Answering via Multimodal Hierarchical Memory Attentive Networks
abstract
Long-term Video Question Answering plays an essential role in visual information retrieval, which aims at generating natural language answers to discretionary free-form questions about the referenced long-term video. Rather than remember the video as a sequence of visual content, humans have an innate cognitive ability to identify the critical moments related to the question at first glance, then tie together the specific evidence around these critical moments for further analysis and reasoning. Motivated by this intuition, we propose the multimodal hierarchical memory attentive networks with two heterogeneous memory subnetworks: the top guided memory network and the bottom enhanced multimodal memory attentive network. The top guided memory network serves as a shallow inference engine to pick relevant and informative moments of questions and obtain salient video content at a coarse-grained level. Subsequently, the bottom enhanced multimodal memory attentive network is designed as an in-depth reasoning engine to perform more accurate attention with cues from video bottom evidence in a fine-grained level to enhance question answering quality. We evaluate the proposed method on three publicly available video question answering benchmarks, namely ActivityNet-QA, MSRVTT-QA, and MSVD-QA. Experimental results demonstrate that the proposed approach significantly outperforms other state-of-the-art methods for long-term videos. Extensive ablation studies are carried out to explore the reasons behind the proposed model's effectiveness.
Ting Yu 0016, Jun Yu 0002, Zhou Yu 0001, Qingming Huang, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 SPRNet: Single-Pixel Reconstruction for One-Stage Instance Segmentation
abstract
Object instance segmentation is one of the most fundamental but challenging tasks in computer vision, and it requires the pixel-level image understanding. Most existing approaches address this problem by adding a mask prediction branch to a two-stage object detector with the region proposal network (RPN). Although producing good segmentation results, the efficiency of these two-stage approaches is far from satisfactory, restricting their applicability in practice. In this article, we propose a one-stage framework, single-pixel reconstruction net (SPRNet), which performs efficient instance segmentation by introducing a single-pixel reconstruction (SPR) branch to off-the-shelf one-stage detectors. The added SPR branch reconstructs the pixel-level mask from every single pixel in the convolution feature map directly. Using the same ResNet-50 backbone, SPRNet achieves comparable mask AP with Mask R-CNN at a higher inference speed and gains all-round improvements on box AP at every scale compared with RetinaNet.
Jun Yu 0002, Jinghan Yao, Jian Zhang 0026, Zhou Yu 0001, Dacheng Tao
IEEE Trans. Cybern.4
2020 Deep Multimodal Neural Architecture Search
abstract
Designing effective neural networks is fundamentally important in deep multimodal learning. Most existing works focus on a single task and design neural architectures manually, which are highly task-specific and hard to generalize to different tasks. In this paper, we devise a generalized deep multimodal neural architecture search (MMnas) framework for various multimodal learning tasks. Given multimodal input, we first define a set of primitive operations, and then construct a deep encoder-decoder based unified backbone, where each encoder or decoder block corresponds to an operation searched from a predefined operation pool. On top of the unified backbone, we attach task-specific heads to tackle different multimodal learning tasks. By using a gradient-based NAS algorithm, the optimal architectures for different tasks are learned efficiently. Extensive ablation studies, comprehensive analysis, and comparative experimental results show that the obtained MMnasNet significantly outperforms existing state-of-the-art approaches across three multimodal learning tasks (over five datasets), including visual question answering, image-text matching, and visual grounding.
Zhou Yu 0001, Yuhao Cui, Jun Yu 0002, Meng Wang 0001, Dacheng Tao, Qi Tian 0001
ACM Multimedia1
2020 Intra- and Inter-modal Multilinear Pooling with Multitask Learning for Video Grounding
Zhou Yu 0001, Yijun Song, Jun Yu 0002, Meng Wang 0001, Qingming Huang
Neural Process. Lett.1
2020 Multimodal Transformer With Multi-View Visual Representation for Image Captioning
abstract
Image captioning aims to automatically generate a natural language description of a given image, and most state-of-the-art models have adopted an encoder-decoder framework. The framework consists of a convolution neural network (CNN)-based image encoder that extracts region-based visual features from the input image, and an recurrent neural network (RNN) based caption decoder that generates the output caption words based on the visual features with the attention mechanism. Despite the success of existing studies, current methods only model the co-attention that characterizes the inter-modal interactions while neglecting the self-attention that characterizes the intra-modal interactions. Inspired by the success of the Transformer model in machine translation, here we extend it to a Multimodal Transformer (MT) model for image captioning. Compared to existing image captioning approaches, the MT model simultaneously captures intra- and inter-modal interactions in a unified attention block. Due to the in-depth modular composition of such attention blocks, the MT model can perform complex multimodal reasoning and output accurate captions. Moreover, to further improve the image captioning performance, multi-view visual features are seamlessly introduced into the MT model. We quantitatively and qualitatively evaluate our approach using the benchmark MSCOCO image captioning dataset and conduct extensive ablation studies to investigate the reasons behind its effectiveness. The experimental results show that our method significantly outperforms the previous state-of-the-art methods. With an ensemble of seven models, our solution ranks the 1st place on the real-time leaderboard of the MSCOCO image captioning challenge at the time of the writing of this paper.
Jun Yu 0002, Jing Li 0099, Zhou Yu 0001, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2020 Compositional Attention Networks With Two-Stream Fusion for Video Question Answering
abstract
Given a video, Video Question Answering (VideoQA) aims at answering arbitrary free-form questions about the video content in natural language. A successful VideoQA framework usually has the following two key components: 1) a discriminative video encoder that learns the effective video representation to maintain as much information as possible about the video and 2) a question-guided decoder that learns to select the most related features to perform spatiotemporal reasoning, as well as outputs the correct answer. We propose compositional attention networks (CAN) with two-stream fusion for VideoQA tasks. For the encoder, we sample video snippets using a two-stream mechanism (i.e., a uniform sampling stream and an action pooling stream) and extract a sequence of visual features for each stream to represent the video semantics with implementation. For the decoder, we propose a compositional attention module to integrate the two-stream features with the attention mechanism. The compositional attention module is the core of CAN and can be seen as a modular combination of a unified attention block. With different fusion strategies, we devise five compositional attention module variants. We evaluate our approach on one long-term VideoQA dataset, ActivityNet-QA, and two short-term VideoQA datasets, MSRVTT-QA and MSVD-QA. Our CAN model achieves new state-of-the-art results on all the datasets.
Ting Yu 0016, Jun Yu 0002, Zhou Yu 0001, Dacheng Tao
IEEE Trans. Image Process.3
2019 ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
abstract
Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answering (VideoQA). Compared to the image domain where large scale and fully annotated benchmark datasets exists, VideoQA datasets are limited to small scale and are automatically generated, etc. These limitations restrict their applicability in practice. Here we introduce ActivityNet-QA, a fully annotated and large scale VideoQA dataset. The dataset consists of 58,000 QA pairs on 5,800 complex web videos derived from the popular ActivityNet dataset. We present a statistical analysis of our ActivityNet-QA dataset and conduct extensive experiments on it by comparing existing VideoQA baselines. Moreover, we explore various video representation strategies to improve VideoQA performance, especially for long videos.
Zhou Yu 0001, Dejing Xu, Jun Yu 0002, Ting Yu 0016, Zhou Zhao 0001, Yueting Zhuang, Dacheng Tao
AAAI1
2019 Deep Modular Co-Attention Networks for Visual Question Answering
abstract
Visual Question Answering (VQA) requires a fine-grained and simultaneous understanding of both the visual content of images and the textual content of questions. Therefore, designing an effective `co-attention' model to associate key words in questions with key objects in images is central to VQA performance. So far, most successful attempts at co-attention learning have been achieved by using shallow models, and deep co-attention models show little improvement over their shallow counterparts. In this paper, we propose a deep Modular Co-Attention Network (MCAN) that consists of Modular Co-Attention (MCA) layers cascaded in depth. Each MCA layer models the self-attention of questions and images, as well as the question-guided-attention of images jointly using a modular composition of two basic attention units. We quantitatively and qualitatively evaluate MCAN on the benchmark VQA-v2 dataset and conduct extensive ablation studies to explore the reasons behind MCAN's effectiveness. Experimental results demonstrate that MCAN significantly outperforms the previous state-of-the-art. Our best single model delivers 70.63% overall accuracy on the test-dev set.
Zhou Yu 0001, Jun Yu 0002, Yuhao Cui, Dacheng Tao, Qi Tian 0001
CVPR1
2019 End-to-end visual grounding via region proposal networks and bilinear pooling
abstract
Phrase‐based visual grounding aims to localise the object in the image referred by a textual query phrase. Most existing approaches adopt a two‐stage mechanism to address this problem: first, an off‐the‐shelf proposal generation model is adopted to extract region‐based visual features, and then a deep model is designed to score the proposals based on the query phrase and extracted visual features. In contrast to that, the authors design an end‐to‐end approach to tackle the visual grounding problem in this study. They use a region proposal network to generate object proposals and the corresponding visual features simultaneously, and multi‐modal factorised bilinear pooling model to fuse the multi‐modal features effectively. After that, two novel losses are posed on top of the multi‐modal features to rank and refine the proposals, respectively. To verify the effectiveness of the proposed approach, the authors conduct experiments on three real‐world visual grounding datasets, namely Flickr‐30k Entities, ReferItGame and RefCOCO. The experimental results demonstrate the significant superiority of the proposed method over the existing state‐of‐the‐arts.
Chenchao Xiang, Zhou Yu 0001, Suguo Zhu, Jun Yu 0002, Xiaokang Yang 0001
IET Comput. Vis.2
2018 Rethinking Diversified and Discriminative Proposal Generation for Visual Grounding
abstract
Visual grounding aims to localize an object in an image referred to by a textual query phrase. Various visual grounding approaches have been proposed, and the problem can be modularized into a general framework: proposal generation, multi-modal feature representation, and proposal ranking. Of these three modules, most existing approaches focus on the latter two, with the importance of proposal generation generally neglected. In this paper, we rethink the problem of what properties make a good proposal generator. We introduce the diversity and discrimination simultaneously when generating proposals, and in doing so propose Diversified and Discriminative Proposal Networks model (DDPN). Based on the proposals generated by DDPN, we propose a high performance baseline model for visual grounding and evaluate it on four benchmark datasets. Experimental results demonstrate that our model delivers significant improvements on all the tested data-sets (e.g., 18.8% improvement on ReferItGame and 8.2% improvement on Flickr30k Entities over the existing state-of-the-arts respectively).
Zhou Yu 0001, Jun Yu 0002, Chenchao Xiang, Zhou Zhao 0001, Qi Tian 0001, Dacheng Tao
IJCAI1
2018 Open-Ended Long-form Video Question Answering via Adaptive Hierarchical Reinforced Networks
abstract
Open-ended long-form video question answering is challenging problem in visual information retrieval, which automatically generates the natural language answer from the referenced long-form video content according to the question. However, the existing video question answering works mainly focus on the short-form video question answering, due to the lack of modeling the semantic representation of long-form video contents. In this paper, we consider the problem of long-form video question answering from the viewpoint of adaptive hierarchical reinforced encoder-decoder network learning. We propose the adaptive hierarchical encoder network to learn the joint representation of the long-form video contents according to the question with adaptive video segmentation. we then develop the reinforced decoder network to generate the natural language answer for open-ended video question answering. We construct a large-scale long-form video question answering dataset. The extensive experiments show the effectiveness of our method.
Zhou Zhao 0001, Shuwen Xiao, Zhou Yu 0001, Jun Yu 0002, Deng Cai 0001, Fei Wu 0001, Yueting Zhuang
IJCAI4
2018 Comprehensive Distance-Preserving Autoencoders for Cross-Modal Retrieval
abstract
In this paper, we propose a novel method with comprehensive distance-preserving autoencoders (CDPAE) to address the problem of unsupervised cross-modal retrieval. Previous unsupervised methods rely primarily on pairwise distances of representations extracted from cross media spaces that co-occur and belong to the same objects. However, besides pairwise distances, the CDPAE also considers heterogeneous distances of representations extracted from cross media spaces as well as homogeneous distances of representations extracted from single media spaces that belong to different objects. The CDPAE consists of four components. First, denoising autoencoders are used to retain the information from the representations and to reduce the negative influence of redundant noises. Second, a comprehensive distance-preserving common space is proposed to explore the correlations among different representations. This aims to preserve the respective distances between the representations within the common space so that they are consistent with the distances in their original media spaces. Third, a novel joint loss function is defined to simultaneously calculate the reconstruction loss of the denoising autoencoders and the correlation loss of the comprehensive distance-preserving common space. Finally, an unsupervised cross-modal similarity measurement is proposed to further improve the retrieval performance. This is carried out by calculating the marginal probability of two media objects based on a kNN classifier. The CDPAE is tested on four public datasets with two cross-modal retrieval tasks: "query images by texts" and "query texts by images". Compared with eight state-of-the-art cross-modal retrieval methods, the experimental results demonstrate that the CDPAE outperforms all the unsupervised methods and performs competitively with the supervised methods.
Yibing Zhan, Jun Yu 0002, Zhou Yu 0001, Rong Zhang 0004, Dacheng Tao, Qi Tian 0001
ACM Multimedia3
2018 Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering
abstract
Visual question answering (VQA) is challenging, because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both the image and the question; 2) multimodal feature fusion that is able to capture the complex interactions between multimodal features; and 3) automatic answer prediction that is able to consider the complex correlations between multiple diverse answers for the same question. For fine-grained image and question representations, a "coattention" mechanism is developed using a deep neural network (DNN) architecture to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features effectively and obtain more discriminative features for image and question representations. For multimodal feature fusion, a generalized multimodal factorized high-order pooling approach (MFH) is developed to achieve more effective fusion of multimodal features by exploiting their correlations sufficiently, which can further result in superior VQA performance as compared with the state-of-the-art approaches. For answer prediction, the Kullback-Leibler divergence is used as the loss function to achieve precise characterization of the complex correlations between multiple diverse answers with the same or similar meaning, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction. A DNN architecture is designed to integrate all these aforementioned modules into a unified model for achieving superior VQA performance. With an ensemble of our MFH models, we achieve the state-of-the-art performance on the large-scale VQA data sets and win the runner-up in VQA Challenge 2017.
Zhou Yu 0001, Jun Yu 0002, Chenchao Xiang, Jianping Fan 0001, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2018 User-Click-Data-Based Fine-Grained Image Recognition via Weakly Supervised Metric Learning
abstract
We present a novel fine-grained image recognition framework using user click data, which can bridge the semantic gap in distinguishing categories that are similar in visual. As query set in click data is usually large-scale and redundant, we first propose a click-feature-based query-merging approach to merge queries with similar semantics and construct a compact click feature. Afterward, we utilize this compact click feature and convolutional neural network (CNN)-based deep visual feature to jointly represent an image. Finally, with the combined feature, we employ the metriclearning-based template-matching scheme for efficient recognition. Considering the heavy noise in the training data, we introduce a reliability variable to characterize the image reliability, and propose a weakly-supervised metric and template leaning with smooth assumption and click prior (WMTLSC) method to jointly learn the distance metric, object templates, and image reliability. Extensive experiments are conducted on a public Clickture-Dog dataset and our newly established Clickture-Bird dataset. It is shown that the click-data-based query merging helps generating a highly compact (the dimension is reduced to 0.9%) and dense click feature for images, which greatly improves the computational efficiency. Also, introducing this click feature into CNN feature further boosts the recognition accuracy. The proposed framework performs much better than previous state-of-the-arts in fine-grained recognition tasks.
Min Tan 0005, Jun Yu 0002, Zhou Yu 0001, Fei Gao 0006, Yong Rui, Dacheng Tao
ACM Trans. Multim. Comput. Commun. Appl.3
2017 Multi-modal Factorized Bilinear Pooling with Co-attention Learning for Visual Question Answering
abstract
Visual question answering (VQA) is challenging because it requires a simultaneous understanding of both the visual content of images and the textual content of questions. The approaches used to represent the images and questions in a fine-grained manner and questions and to fuse these multimodal features play key roles in performance. Bilinear pooling based models have been shown to outperform traditional linear models for VQA, but their high-dimensional representations and high computational complexity may seriously limit their applicability in practice. For multimodal feature fusion, here we develop a Multi-modal Factorized Bilinear (MFB) pooling approach to efficiently and effectively combine multi-modal features, which results in superior performance for VQA compared with other bilinear pooling approaches. For fine-grained image and question representation, we develop a `co-attention' mechanism using an end-to-end deep network architecture to jointly learn both the image and question attentions. Combining the proposed MFB approach with co-attention learning in a new network architecture provides a unified model for VQA. Our experimental results demonstrate that the single MFB with co-attention model achieves new state-of-theart performance on the real-world VQA dataset. Code available at https://github.com/yuzcccc/mfb.
Zhou Yu 0001, Jun Yu 0002, Jianping Fan 0001, Dacheng Tao
ICCV1
2017 Deep Mixture of Experts with Diverse Task Spaces
abstract
In this paper, a deep mixture algorithm is developed to support large-scale visual recognition (e.g., recognizing tens of thousands of object classes) by seamlessly combining a set of base deep CNNs (AlexNet) with diverse task spaces, e.g., such base deep CNNs (i.e., diverse experts) are trained to recognize different subsets of tens of thousands of object classes rather than the same set of object classes. Our experimental results have demonstrated that our deep mixture algorithm can achieve very competitive results on large-scale visual recognition.
Jianping Fan 0001, Zhenzhong Kuang, Zhou Yu 0001, Jun Yu 0002
ICMLA4
2017 Privacy Setting Recommendation for Image Sharing
abstract
This paper aims to simultaneously consider two inseparable issues for privacy setting recommendation: (1) sensitiveness of visual content of the images being shared; and (2) trustworthiness of users being granted. First, an object-based approach is developed for image content sensitiveness (privacy) representation. Secondly, the users on a social network are clustered into a set of representative social groups to generate a discriminative dictionary for user trustworthiness characterization. Finally, a tree classifier is trained hierarchically to recommend appropriate privacy settings for image sharing.
Jun Yu 0002, Zhenzhong Kuang, Zhou Yu 0001, Dan Lin 0001, Jianping Fan 0001
ICMLA3
2014 Cross-media hashing with kernel regression
abstract
Cross-media retrieval is a challenging problem in multimedia retrieval area. In the real-world, many applications involve multi-modal data, e.g., web pages containing both images and texts. How to utilize the intrinsic intra-modality and inter-modality similarity to learn the appropriate relationships of the data objects and provide efficient search across different modalities is the core of cross-media retrieval. Inspired by the fact that hashing methods well address the fast retrieval problem in the large-scale data settings, designing a cross-media hashing approach which can perform efficient retrieval over heterogenous high-dimensional feature spaces is highly desirable. In this paper, we propose a cross-media hashing approach based on kernel regression (abbreviated as KRCMH) to obtain the hash codes for the data objects across different modalities. The experiments on two real-world data sets show that KRCMH achieves superior cross-media retrieval performance comparing with the state-of-the-art methods.
Zhou Yu 0001, Yin Zhang 0006, Siliang Tang, Yi Yang 0001, Qi Tian 0001, Jiebo Luo 0001
ICME1
2014 Cross-Media Hashing with Neural Networks
abstract
Cross-media hashing, which conducts cross-media retrieval by embedding data from different modalities into a common low-dimensional hamming space, has attracted intensive attention in recent years. This is motivated by the facts a) the multi-modal data is widespread, e.g., the web images on Flickr are associated with tags, and b) hashing is an effective technique towards large-scale high-dimensional data processing, which is exactly the situation of cross-media retrieval. Inspired by recent advances in deep learning, we propose a cross-media hashing approach based on multi-modal neural networks. By restricting in the learning objective a) the hash codes for relevant cross-media data being similar, and b) the hash codes being discriminative for predicting the class labels, the learned Hamming space is expected to well capture the cross-media semantic relationships and to be semantically discriminative. The experiments on two real-world data sets show that our approach achieves superior cross-media retrieval performance compared with the state-of-the-art methods.
Yueting Zhuang, Zhou Yu 0001, Wei Wang 0059, Fei Wu 0001, Siliang Tang, Jian Shao 0001
ACM Multimedia2
2014 Discriminative coupled dictionary hashing for fast cross-media retrieval
abstract
Cross-media hashing, which conducts cross-media retrieval by embedding data from different modalities into a common low-dimensional Hamming space, has attracted intensive attention in recent years. The existing cross-media hashing approaches only aim at learning hash functions to preserve the intra-modality and inter-modality correlations, but do not directly capture the underlying semantic information of the multi-modal data. We propose a discriminative coupled dictionary hashing (DCDH) method in this paper. In DCDH, the coupled dictionary for each modality is learned with side information (e.g., categories). As a result, the coupled dictionaries not only preserve the intra-similarity and inter-correlation among multi-modal data, but also contain dictionary atoms that are semantically discriminative (i.e., the data from the same category is reconstructed by the similar dictionary atoms). To perform fast cross-media retrieval, we learn hash functions which map data from the dictionary space to a low-dimensional Hamming space. Besides, we conjecture that a balanced representation is crucial in cross-media retrieval. We introduce multi-view features on the relatively ``weak'' modalities into DCDH and extend it to multi-view DCDH (MV-DCDH) in order to enhance their representation capability. The experiments on two real-world data sets show that our DCDH and MV-DCDH outperform the state-of-the-art methods significantly on cross-media retrieval.
Zhou Yu 0001, Fei Wu 0001, Yi Yang 0001, Qi Tian 0001, Jiebo Luo 0001, Yueting Zhuang
SIGIR1
2014 Hashing with List-Wise learning to rank
abstract
Hashing techniques have been extensively investigated to boost similarity search for large-scale high-dimensional data. Most of the existing approaches formulate the their objective as a pair-wise similarity-preserving problem. In this paper, we consider the hashing problem from the perspective of optimizing a list-wise learning to rank problem and propose an approach called List-Wise supervised Hashing (LWH). In LWH, the hash functions are optimized by employing structural SVM in order to explicitly minimize the ranking loss of the whole list-wise permutations instead of merely the point-wise or pair-wise supervision. We evaluate the performance of LWH on two real-world data sets. Experimental results demonstrate that our method obtains a significant improvement over the state-of-the-art hashing approaches due to both structural large margin and list-wise ranking pursuing in a supervised manner.
Zhou Yu 0001, Fei Wu 0001, Yin Zhang 0006, Siliang Tang, Jian Shao 0001, Yueting Zhuang
SIGIR1
2014 Sparse Multi-Modal Hashing
abstract
Learning hash functions across heterogenous high-dimensional features is very desirable for many applications involving multi-modal data objects. In this paper, we propose an approach to obtain the sparse codesets for the data objects across different modalities via joint multi-modal dictionary learning, which we call sparse multi-modal hashing (abbreviated as${\rm SM}^{2}{\rm H}$). In${\rm SM}^{2}{\rm H}$, both intra-modality similarity and inter-modality similarity are first modeled by a hypergraph, then multi-modal dictionaries are jointly learned by Hypergraph Laplacian sparse coding. Based on the learned dictionaries, the sparse codeset of each data object is acquired and conducted for multi-modal approximate nearest neighbor retrieval using a sensitive Jaccard metric. The experimental results show that${\rm SM}^{2}{\rm H}$outperforms other methods in terms of mAP and Percentage on two real-world data sets.
Fei Wu 0001, Zhou Yu 0001, Yi Yang 0001, Siliang Tang, Yin Zhang 0006, Yueting Zhuang
IEEE Trans. Multim.2
2010 Fire Surveillance Method Based on Quaternionic Wavelet Features
Zhou Yu 0001, Yi Xu 0001, Xiaokang Yang 0001
MMM1