VLDB 2026 Research / reviewers in the wild / expert
Jun Yu 0002
dblp:50/5754-2
· DBLP profile ↗
270ranked-venue papers
41as first author
148since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 145 · 16 first-author · 84 since 2021Artificial intelligence and machine learning · 126 · 21 first-author · 72 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 5 since 2021Computer networks · 8 · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorSecurity and privacy · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse4DGS: 4D Gaussian Splatting for Sparse-Frame Dynamic Scene ReconstructionabstractDynamic Gaussian Splatting approaches have achieved remarkable performance for 4D scene reconstruction. However, these approaches rely on dense-frame video sequences for photorealistic reconstruction. In real-world scenarios, due to equipment constraints, sometimes only sparse frames are accessible. In this paper, we propose Sparse4DGS, the first method for sparse-frame dynamic scene reconstruction. We observe that dynamic reconstruction methods fail in both canonical and deformed spaces under sparse-frame settings, especially in areas with high texture richness. Sparse4DGS tackles this challenge by focusing on texture-rich areas. For the deformation network, we propose Texture-Aware Deformation Regularization, which introduces a texture-based depth alignment loss to regulate Gaussian deformation. For the canonical Gaussian field, we introduce Texture-Aware Canonical Optimization, which incorporates texture-based noise into the gradient descent process of canonical Gaussians. Extensive experiments show that when taking sparse frames as inputs, our method outperforms existing dynamic or few-shot techniques on NeRF-Synthetic, HyperNeRF, NeRF-DS, and our iPhone-4D datasets. Changyue Shi, Chuxiao Yang, Wenwen Pan 0003, Jiajun Ding, Zhou Yu 0001, Jun Yu 0002 |
AAAI | 9 |
| 2026 | Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image CaptioningabstractNews image captioning aims to produce journalistically informative descriptions by combining visual content with contextual cues from associated articles. Despite recent advances, existing methods struggle with three key challenges: (1) incomplete information coverage, (2) weak cross-modal alignment, and (3) suboptimal visual-entity grounding. To address these issues, we introduce MERGE, the first Multimodal Entity-aware Retrieval-augmented GEneration framework for news image captioning. MERGE constructs an entity-centric multimodal knowledge base (EMKB) that integrates textual, visual, and structured knowledge, enabling enriched background retrieval. It improves cross-modal alignment through a multistage hypothesis-caption strategy and enhances visual-entity matching via dynamic retrieval guided by image content. Extensive experiments on GoodNews and NYTimes800k show that MERGE significantly outperforms state-of-the-art baselines, with CIDEr gains of +6.84 and +1.16 in caption quality, and F1-score improvements of +4.14 and +2.64 in named entity recognition. Notably, MERGE also generalizes well to the unseen Visual News dataset, achieving +20.17 in CIDEr and +6.22 in F1-score, demonstrating strong robustness and domain adaptability. Xiaoxing You, Chi Zhang 0022, Min Zhang 0005, Jun Yu 0002 |
AAAI | 7 |
| 2026 | Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real PhotosabstractVision-Language Models (VLMs) are increasingly deployed in socially consequential settings, raising concerns about social bias driven by demographic cues. A central challenge in measuring such social bias is attribution under visual confounding: real-world images entangle race and gender with correlated factors such as background and clothing, obscuring attribution. We propose a \textbf{face-only counterfactual evaluation paradigm} that isolates demographic effects while preserving real-image realism. Starting from real photographs, we generate counterfactual variants by editing only facial attributes related to race and gender, keeping all other visual factors fixed. Based on this paradigm, we construct \textbf{FOCUS}, a dataset of 480 scene-matched counterfactual images across six occupations and ten demographic groups, and propose \textbf{REFLECT}, a benchmark comprising three decision-oriented tasks: two-alternative forced choice, multiple-choice socioeconomic inference, and numeric salary recommendation. Experiments on five state-of-the-art VLMs reveal that demographic disparities persist under strict visual control and vary substantially across task formulations. These findings underscore the necessity of controlled, counterfactual audits and highlight task design as a critical factor in evaluating social bias in multimodal models. Qiuping Jiang, Xiaojun Chang, Jun Yu 0002 |
ACL (1) | 6 |
| 2026 | Towards alleviating hallucination in text-to-image retrieval for CLIP in zero-shot learning
Hanyao Wang, Yibing Zhan, Liu Liu 0014, Liang Ding 0006, Jun Yu 0002 |
Neurocomputing | 6 |
| 2026 | Divide-and-conquer towards optimal adaptation of pre-trained model to medical tasks
Zhanghui Huang, Zunlei Feng, Xiaoyan Sun 0006, Shuifa Sun, Zhenming Yuan, Jun Yu 0002, Jian Zhang 0026 |
Pattern Recognit. | 6 |
| 2026 | Morphology semantics-guided vision language alignment for cervical cell image classification
Jiaxin Lei, Zunlei Feng, Jingwen Ye, Zhenming Yuan, Jun Yu 0002 |
Pattern Recognit. | 6 |
| 2026 | TAG: Triple Alignment With Rationale Generation for Knowledge-Based Visual Question AnsweringabstractKnowledge-based Visual Question Answering (VQA) involves answering questions based not only on the given image, but also on external knowledge. Existing methods for knowledge-based VQA can be classified into two main categories: those that rely on external knowledge bases, and those that use Large Language Models (LLMs) as implicit knowledge engines. However, the former approach heavily relies on the quality of information retrieval, introducing additional information bias to the entire system. And the latter approach suffers from the extremely high computational cost and the loss of image information. To address these issues, we propose a novel framework called TAG that reformulates knowledge-based VQA as a contrastive learning problem. We innovatively propose a triple asymmetric paradigm, which aligns a lightweight text encoder to the image space with an extremely low training cost (0.0152B trainable parameters), and enhance its understanding ability on semantic granularity. TAG is both computation-efficient and effective, and we evaluate it on the knowledge-based VQA datasets, A-OKVQA, OK-VQA and VCR. The results show that TAG (0.387B) achieves the state-of-the-art performance when compared to methods using less than 1B parameters. Besides, TAG still shows competitive performance when compared to methods with LLM. Sihang Cai, Xuan Lin, Jingtong Wu, Tao Jin 0004, Zhou Zhao 0001, Fei Wu 0001, Jun Yu 0002 |
IEEE Trans. Big Data | 8 |
| 2026 | DKGZSL: Leveraging Dynamic Visual-Semantic Knowledge for Generative Zero-Shot LearningabstractGenerative Zero-Shot Learning (GZSL) methods address the challenge of recognizing unseen classes by synthesizing visual features, thereby converting ZSL into a supervised learning task. However, existing approaches are predominantly constrained to two multi-stage strategies: pre-generation prior knowledge enhancement and post-generation feature refinement. These paradigms often suffer from error propagation across stages, ultimately limiting generation quality and representational fidelity. To overcome these limitations, we propose DKGZSL, a novel generative framework that injects dynamic visual-semantic knowledge directly into the feature synthesis process, effectively unifying generation and refinement into a single cohesive stage. Specifically, a Knowledge Transfer Network (KTN) is introduced to convert semantic information into hierarchical visual knowledge representations. To ensure accurate semanticvisual alignment, we further design a Semantic-Oriented Visual Refinement (SOVR) module that reshapes real visual features into semantically aligned and noise-suppressed representations, providing precise guidance for the KTN. Moreover, hierarchical knowledge extracted from each KTN layer is progressively transmitted to the generator via Meta-Fusion Units (MFUs), enabling dynamic semantic guidance and improving generation quality. Extensive experiments on three benchmark datasets demonstrate that DKGZSL achieves consistent state-of-the-art performance with both ResNet-101 and ViT-B/16 feature extractors. Comprehensive ablation studies further confirm the effectiveness and complementarity of each proposed component. The code is available at https://github.com/JingHu-gdut/DKGZSL. Min Meng 0001, Jigang Liu, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Exploring Hierarchical Cross-Modal Correlation Consistency for Partial MismatchingabstractCross-modal retrieval facilitates more flexible information access and improves semantic understanding across different modalities. However, traditional cross-modal retrieval models rely on well-aligned datasets, which are often labor-intensive and costly to obtain. In real-world applications, data inevitably includes mismatched pairs, and these semantically inconsistent pairs can significantly degrade retrieval performance. Previous approaches have assumed ideal loss value distributions to optimize models for accurate semantic matching through soft-label estimation. However, the absence of hierarchical semantic correlation learning limits the effectiveness of these models in scenarios involving partial mismatches. To address these challenges, we propose Exploring Hierarchical Cross-Modal Correlation Consistency (EH3C) for cross-modal retrieval under partially mismatched conditions. Specifically, our approach first leverages neighborhood correlation distributions among samples to optimize cross-modal alignment, without assuming ideal distributions. This allows for the measurement of soft matching degrees between cross-modal data pairs and facilitates the effective learning of their positive correlations. Next, we enhance inter-class separability through intra-modal correlation learning by exploiting negative correlations between reliable negative sample pairs, thus enabling a more comprehensive exploration of cross-modal correlations. Finally, to assess the effectiveness and robustness of our approach, we conducted extensive experiments on three benchmark datasets. The results demonstrate that the proposed EH3C significantly improves cross-modal retrieval performance in scenarios involving partial mismatches. Zhiwen Yu 0002, Jun Yu 0002, Huanqiang Zeng, Zhuoyao Wang 0001, C. L. Philip Chen |
IEEE Trans. Image Process. | 3 |
| 2026 | Zero-Shot Video Translation via Token WarpingabstractWith the revolution of generative AI, video-related tasks have been widely studied. However, current state-of-the-art video models still lag behind image models in visual quality and user control over generated content. In this paper, we introduce TokenWarping, a novel framework for temporally coherent video translation. Existing diffusion-based video editing approaches rely solely on key and value patches in self-attention to ensure temporal consistency, often sacrificing the preservation of local and structural regions. Critically, these methods overlook the significance of the query patches in achieving accurate feature aggregation and temporal coherence. In contrast, TokenWarping leverages complementary token priors by constructing temporal correlations across different frames. Our method begins by extracting optical flows from source videos. During the denoising process of the diffusion model, these optical flows are used to warp the previous frame's query, key, and value patches, aligning them with the current frame's patches. By directly warping the query patches, we enhance feature aggregation in self-attention, while warping the key and value patches ensures temporal consistency across frames. This token warping imposes explicit constraints on the self-attention layer outputs, effectively ensuring temporally coherent translation. Our framework does not require any additional training or fine-tuning and can be seamlessly integrated with existing text-to-image editing methods. We conduct extensive experiments on various video translation tasks, demonstrating that TokenWarping surpasses state-of-the-art methods both qualitatively and quantitatively. Video demonstrations are available in supplementary materials. Haiming Zhu, Yangyang Xu 0003, Jun Yu 0002, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | Balancing Relevance and Diversity in k-Maximum Inner Product Search
Yanhao Wang 0001, Yiqun Sun, Anthony K. H. Tung, Jun Yu 0002 |
VLDB J. | 5 |
| 2025 | Fine-grained Adaptive Visual Prompt for Generative Medical Visual Question AnsweringabstractMedical Visual Question Answering (MedVQA) serves as an automated medical assistant, capable of answering patient queries and aiding physician diagnoses based on medical images and questions. Recent advancements have shown that incorporating Large Language Models (LLMs) into MedVQA tasks significantly enhances the capability for answer generation. However, for tasks requiring fine-grained organ-level precise localization, relying solely on language prompts struggles to accurately locate relevant regions within medical images due to substantial background noise. To address this challenge, we explore the use of visual prompts in MedVQA tasks for the first time and propose fine-grained adaptive visual prompts to enhance generative MedVQA. Specifically, we introduce an Adaptive Visual Prompt Creator that adaptively generates region-level visual prompts based on image characteristics of various organs, providing fine-grained references for LLMs during answer retrieval and generation from the medical domain, thereby improving the model's precise cross-modal localization capabilities on original images. Furthermore, we incorporate a Hierarchical Answer Generator with Parameter-Efficient Fine-Tuning (PEFT) techniques, significantly enhancing the model's understanding of spatial and contextual information with minimal parameter increase, promoting the alignment of representation learning with the medical space. Extensive experiments on VQA-RAD, SLAKE, and DME datasets validate the effectiveness of our proposed method, demonstrating its potential in generative MedVQA. Ting Yu 0016, Zixuan Tong, Jun Yu 0002, Ke Zhang 0029 |
AAAI | 3 |
| 2025 | A Similarity Paradigm Through Textual Regularization Without ForgettingabstractPrompt learning has emerged as a promising method for adapting pre-trained visual-language models (VLMs) to a range of downstream tasks. While optimizing the context can be effective for improving performance on specific tasks, it can often lead to poor generalization performance on unseen classes or datasets sampled from different distributions. It may be attributed to the fact that textual prompts tend to overfit downstream data distributions, leading to the forgetting of generalized knowledge derived from hand-crafted prompts. In this paper, we propose a novel method called Similarity Paradigm with Textual Regularization (SPTR) for prompt learning without forgetting. SPTR is a two-pronged design based on hand-crafted prompts that is an inseparable framework. 1) To avoid forgetting general textual knowledge, we introduce the optimal transport as a textual regularization to finely ensure approximation with hand-crafted features and tuning textual features. 2) In order to continuously unleash the general ability of multiple hand-crafted prompts, we propose a similarity paradigm for natural alignment score and adversarial alignment score to improve model robustness for generalization. Both modules share a common objective in addressing generalization issues, aiming to maximize the generalization capability derived from multiple hand-crafted prompts. Four representative tasks (i.e., non-generalization few-shot learning, base-to-novel generalization, cross-dataset generalization, domain generalization) across 11 datasets demonstrate that SPTR outperforms existing prompt learning methods. Fangming Cui, Jan Fong, Rongfei Zeng, Xinmei Tian 0001, Jun Yu 0002 |
AAAI | 5 |
| 2025 | MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teamingabstractThe proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures.However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more prone to produce harmful responses.In this paper, we propose the Multi-Turn Safety Alignment (MTSA) framework, to address the challenge of securing LLMs in multi-round interactions.It consists of two stages: In the thought-guided attack learning stage, the redteam model learns about thought-guided multiround jailbreak attacks to generate adversarial prompts.In the adversarial iterative optimization stage, the red-team model and the target model continuously improve their respective capabilities in interaction.Furthermore, we introduce a multi-turn reinforcement learning algorithm based on future rewards to enhance the robustness of safety alignment.Experimental results show that the red-team model exhibits state-of-the-art attack capabilities, while the target model significantly improves its performance on safety benchmarks. Weiyang Guo, Jing Li 0034, Wenya Wang 0001, Yu Li 0007, Daojing He, Jun Yu 0002, Min Zhang 0005 |
ACL (1) | 6 |
| 2025 | Safety Alignment via Constrained Knowledge UnlearningabstractZesheng Shi, Yucheng Zhou, Jing Li, Yuxin Jin, Yu Li, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zesheng Shi, Yucheng Zhou 0001, Jing Li 0034, Yu Li 0007, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu 0002, Min Zhang 0005 |
ACL (1) | 9 |
| 2025 | PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-EncoderabstractSemantic Text Embedding is a fundamental NLP task that encodes textual content into vector representations, where proximity in the embedding space reflects semantic similarity.While existing embedding models excel at capturing general meaning, they often overlook ideological nuances, limiting their effectiveness in tasks that require an understanding of political bias.To address this gap, we introduce PRISM, the first framework designed to Produce inteRpretable polItical biaS eMbeddings.PRISM operates in two key stages: (1) Controversial Topic Bias Indicator Mining, which systematically extracts fine-grained political topics and their corresponding bias indicators from weakly labeled news data, and (2) Cross-Encoder Political Bias Embedding, which assigns structured bias scores to news articles based on their alignment with these indicators.This approach ensures that embeddings are explicitly tied to bias-revealing dimensions, enhancing both interpretability and predictive power.Through extensive experiments on two large-scale datasets, we demonstrate that PRISM outperforms stateof-the-art text embedding models in political bias classification while offering highly interpretable representations that facilitate diversified retrieval and ideological analysis. Yiqun Sun, Anthony K. H. Tung, Jun Yu 0002 |
ACL (1) | 4 |
| 2025 | Towards Text-Image Interleaved RetrievalabstractXin Zhang, Ziqi Dai, Yongqi Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jun Yu, Wenjie Li, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xin Zhang 0097, Ziqi Dai, Yongqi Li 0001, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jun Yu 0002, Wenjie Li 0002, Min Zhang 0005 |
ACL (1) | 8 |
| 2025 | Speed Up Your Code: Progressive Code Acceleration Through Bidirectional Tree EditingabstractLonghui Zhang, Jiahao Wang, Meishan Zhang, GaoXiong Cao, Ensheng Shi, Mayuchi Mayuchi, Jun Yu, Honghai Liu, Jing Li, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Longhui Zhang, Meishan Zhang, GaoXiong Cao, Ensheng Shi, Mayuchi Mayuchi, Jun Yu 0002, Honghai Liu 0001, Jing Li 0034, Min Zhang 0005 |
ACL (1) | 7 |
| 2025 | Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and ReasoningabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks.Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs.To study the reason behind these limitations, we propose VGCURE, a comprehensive benchmark covering 22 tasks for examining the fundamental graph understanding and reasoning capacities of LVLMs.Extensive evaluations conducted on 14 LVLMs reveal that LVLMs are weak in basic graph understanding and reasoning tasks, particularly those concerning relational or structurally complex information.Based on this observation, we propose a structure-aware fine-tuning framework to enhance LVLMs with structure learning abilities through three self-supervised learning tasks.Experiments validate the effectiveness of our method in improving LVLMs' performance on fundamental and downstream graph learning tasks, as well as enhancing their robustness against complex visual graphs. Yingjie Zhu, Xuefeng Bai 0001, Kehai Chen, Yang Xiang 0003, Jun Yu 0002, Min Zhang 0005 |
ACL (1) | 5 |
| 2025 | Recognition-Synergistic Scene Text EditingabstractScene text editing aims to modify text content within scene images while maintaining style consistency. Traditional methods achieve this by explicitly disentangling style and content from the source image and then fusing the style with the target content, while ensuring content consistency using a pre-trained recognition model. Despite notable progress, these methods suffer from complex pipelines, leading to suboptimal performance in complex scenarios. In this work, we introduce Recognition-Synergistic Scene Text Editing (RS-STE), a novel approach that fully exploits the intrinsic synergy of text recognition for editing. Our model seamlessly integrates text recognition with text editing within a unified framework, and leverages the recognition model’s ability to implicitly disentangle style and content while ensuring content consistency. Specifically, our approach employs a multi-modal parallel decoder based on transformer architecture, which predicts both text content and stylized images in parallel. Additionally, our cyclic self-supervised fine-tuning strategy enables effective training on unpaired real-world data without ground truth, enhancing style and content consistency through a twice-cyclic generation process. Built on a relatively simple architecture, RS-STE achieves state-of-the-art performance on both synthetic and real-world benchmarks, and further demonstrates the effectiveness of leveraging the generated hard cases to boost the performance of downstream recognition tasks. Code is available at https://github.com/ZhengyaoFang/RS-STE. Zhengyao Fang, Pengyuan Lv, Chengquan Zhang, Jun Yu 0002, Guangming Lu 0002, Wenjie Pei |
CVPR | 5 |
| 2025 | Learning Compatible Multi-Prize Subnetworks for Asymmetric RetrievalabstractAsymmetric retrieval is a typical scenario in real-world retrieval systems, where compatible models of varying capacities are deployed on platforms with different resource configurations. Existing methods generally train pre-defined networks or subnetworks with capacities specifically designed for pre-determined platforms, using compatible learning. Nevertheless, these methods suffer from limited flexibility for multi-platform deployment. For example, when introducing a new platform into the retrieval systems, developers have to train an additional model at an appropriate capacity that is compatible with existing models via backward-compatible learning. In this paper, we propose a Prunable Network with self-compatibility, which allows developers to generate compatible subnetworks at any desired capacity through post-training pruning. Thus it allows the creation of a sparse subnetwork matching the resources of the new platform without additional training. Specifically, we optimize both the architecture and weight of subnetworks at different capacities within a dense network in compatible learning. We also design a conflict-aware gradient integration scheme to handle the gradient conflicts between the dense network and subnetworks during compatible learning. Extensive experiments on diverse benchmarks and visual backbones demonstrate the effectiveness of our method. The code will be made publicly available. Yushuai Sun, Zikun Zhou, Dongmei Jiang, Yaowei Wang 0001, Jun Yu 0002, Guangming Lu 0002, Wenjie Pei |
CVPR | 5 |
| 2025 | AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMsabstractDespite the impressive performance of large language models (LLMs) in general domains, they often underperform in specialized domains.Existing approaches typically rely on data synthesis methods and yield promising results by using unlabeled data to capture domain-specific features.However, these methods either incur high computational costs or suffer from performance limitations, while also demonstrating insufficient generalization across different tasks.To address these challenges, we propose AQuilt, a framework for constructing instruction-tuning data for any specialized domains from corresponding unlabeled data, including Answer, Question, Unlabeled data, Inspection, Logic, and Task type.By incorporating logic and inspection, we encourage reasoning processes and self-inspection to enhance model performance.Moreover, customizable task instructions enable high-quality data generation for any task.As a result, we construct a dataset of 703k examples to train a powerful data synthesis model.Experiments show that AQuilt is comparable to DeepSeek-V3 while utilizing just 17% of the production cost.Further analysis demonstrates that our generated data exhibits higher relevance to downstream tasks. Xiaopeng Ke, Hexuan Deng, Xuebo Liu 0002, Jun Rao, Zhenxi Song, Jun Yu 0002, Min Zhang 0005 |
EMNLP | 6 |
| 2025 | D2 ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-Shot Action Recognition
Wenjie Pei, Qizhong Tan, Guangming Lu 0002, Jiandong Tian, Jun Yu 0002 |
ICCV | 5 |
| 2025 | Growing a Twig to Accelerate Large Vision-Language Models
Zhenwei Shao, Zhou Yu 0001, Wenwen Pan 0003, Hongyuan Zhang 0001, Wei Chen 0001, Jun Yu 0002 |
ICCV | 10 |
| 2025 | Stable Score DistillationabstractText-guided image and 3D editing have advanced with diffusion-based models, yet methods like Delta Denoising Score often struggle with stability, spatial control, and editing strength. These limitations stem from reliance on complex auxiliary structures, which introduce conflicting optimization signals and restrict precise, localized edits. We introduce Stable Score Distillation (SSD), a streamlined framework that enhances stability and alignment in the editing process by anchoring a single classifier to the source prompt. Specifically, SSD utilizes Classifier-Free Guidance (CFG) equation to achieves cross-prompt alignment, and introduces a constant term null-text branch to stabilize the optimization process. This approach preserves the original content's structure and ensures that editing trajectories are closely aligned with the source prompt, enabling smooth, prompt-specific modifications while maintaining coherence in surrounding regions. Additionally, SSD incorporates a prompt enhancement branch to boost editing strength, particularly for style transformations. Our method achieves state-of-the-art results in 2D and 3D editing tasks, including NeRF and text-driven style edits, with faster convergence and reduced complexity, providing a robust and efficient solution for text-guided editing. Haiming Zhu, Yangyang Xu 0003, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du 0003, Jun Yu 0002, Shengfeng He |
ICCV | 7 |
| 2025 | OmniKV: Dynamic Context Selection for Efficient Long-Context LLMsabstractDuring the inference phase of Large Language Models (LLMs) with long context, a substantial portion of GPU memory is allocated to the KV cache, with memory usage increasing as the sequence length grows. To mitigate the GPU memory footprint associate with KV cache, some previous studies have discarded less important tokens based on the sparsity identified in attention scores in long context scenarios. However, we argue that attention scores cannot indicate the future importance of tokens in subsequent generation iterations, because attention scores are calculated based on current hidden states. Therefore, we propose OmniKV, a token-dropping-free and training-free inference method, which achieves a 1.68x speedup without any loss in performance. It is well-suited for offloading, significantly reducing KV cache memory usage by up to 75% with it. The core innovative insight of OmniKV is: Within a single generation iteration, there is a high degree of similarity in the important tokens identified across consecutive layers. Extensive experiments demonstrate that OmniKV achieves state-of-the-art performance across multiple benchmarks, with particularly advantages in chain-of-thoughts scenarios. OmniKV extends the maximum context length supported by a single A100 for Llama-3-8B from 128K to 450K. Our code is available at https://github.com/antgroup/OmniKV.git. Jitai Hao, Yuke Zhu, Jun Yu 0002, Xin Xin 0003, Bo Zheng 0007, Zhaochun Ren, Sheng Guo 0005 |
ICLR | 4 |
| 2025 | A General Framework for Producing Interpretable Semantic Text EmbeddingsabstractSemantic text embedding is essential to many tasks in Natural Language Processing (NLP). While black-box models are capable of generating high-quality embeddings, their lack of interpretability limits their use in tasks that demand transparency. Recent approaches have improved interpretability by leveraging domain-expert-crafted or LLM-generated questions, but these methods rely heavily on expert input or well-prompt design, which restricts their generalizability and ability to generate discriminative questions across a wide range of tasks. To address these challenges, we introduce \algo{CQG-MBQA} (Contrastive Question Generation - Multi-task Binary Question Answering), a general framework for producing interpretable semantic text embeddings across diverse tasks. Our framework systematically generates highly discriminative, low cognitive load yes/no questions through the \algo{CQG} method and answers them efficiently with the \algo{MBQA} model, resulting in interpretable embeddings in a cost-effective manner. We validate the effectiveness and interpretability of \algo{CQG-MBQA} through extensive experiments and ablation studies, demonstrating that it delivers embedding quality comparable to many advanced black-box models while maintaining inherently interpretability. Additionally, \algo{CQG-MBQA} outperforms other interpretable text embedding methods across various downstream tasks. The source code is available at \url{https://github.com/dukesun99/CQG-MBQA}. Yiqun Sun, Anthony K. H. Tung, Jun Yu 0002 |
ICLR | 5 |
| 2025 | Enhancing Target-unspecific Tasks through a Features MatrixabstractRecent developments in prompt learning of large Vision-Language Models (VLMs) have significantly improved performance in target-specific tasks. However, these prompting methods often struggle to tackle the target-unspecific or generalizable tasks effectively. It may be attributed to the fact that overfitting training causes the model to forget its general knowledge. The general knowledge has a strong promotion on target-unspecific tasks. To alleviate this issue, we propose a novel Features Matrix (FM) approach designed to enhance these models on target-unspecific tasks. Our method extracts and leverages general knowledge, shaping a Features Matrix (FM). Specifically, the FM captures the semantics of diverse inputs from a deep and fine perspective, preserving essential general knowledge, which mitigates the risk of overfitting. Representative evaluations demonstrate that: 1) the FM is compatible with existing frameworks as a generic and flexible module, and 2) the FM significantly showcases its effectiveness in enhancing target-unspecific tasks (base-to-novel generalization, domain generalization, and cross-dataset generalization), achieving state-of-the-art performance. Fangming Cui, Yonggang Zhang 0003, Xinmei Tian 0001, Jun Yu 0002 |
ICML | 5 |
| 2025 | Function-to-Style Guidance of LLMs for Code TranslationabstractLarge language models (LLMs) have made significant strides in code translation tasks. However, ensuring both the correctness and readability of translated code remains a challenge, limiting their effective adoption in real-world software development. In this work, we propose F2STrans, a function-to-style guiding paradigm designed to progressively improve the performance of LLMs in code translation. Our approach comprises two key stages: (1) Functional learning, which optimizes translation correctness using high-quality source-target code pairs mined from online programming platforms, and (2) Style learning, which improves translation readability by incorporating both positive and negative style examples. Additionally, we introduce a novel code translation benchmark that includes up-to-date source code, extensive test cases, and manually annotated ground-truth translations, enabling comprehensive functional and stylistic evaluations. Experiments on both our new benchmark and existing datasets demonstrate that our approach significantly improves code translation performance. Notably, our approach enables Qwen-1.5B to outperform prompt-enhanced Qwen-32B and GPT-4 on average across 20 diverse code translation scenarios. Longhui Zhang, Bin Wang 0004, Hao Yang 0007, Meishan Zhang, Yu Li 0007, Jing Li 0034, Jun Yu 0002, Min Zhang 0005 |
ICML | 10 |
| 2025 | DiffusionMat: Alpha Matting as Deterministic Sequential Refinement Learning
Yangyang Xu 0003, Shengfeng He, Wenqi Shao, Yong Du 0003, Kwan-Yee Kenneth Wong, Yu Qiao 0001, Jun Yu 0002, Ping Luo 0002 |
ACM Multimedia | 7 |
| 2025 | Towards Generalizable Detector for Generated ImageabstractThe effective detection of generated images is crucial to mitigate potential risks associated with their misuse. Despite significant progress, a fundamental challenge remains: ensuring the generalizability of detectors. To address this, we propose a novel perspective on understanding and improving generated image detection, inspired by the human cognitive process: Humans identify an image as unnatural based on specific patterns because these patterns lie outside the space spanned by those of natural images. This is intrinsically related to out-of-distribution (OOD) detection, which identifies samples whose semantic patterns (i.e., labels) lie outside the semantic pattern space of in-distribution (ID) samples.
By treating patterns of generated images as OOD samples, we demonstrate that models trained merely over natural images bring guaranteed generalization ability under mild assumptions.
This transforms the generalization challenge of generated image detection into the problem of fitting natural image patterns.
Based on this insight, we propose a generalizable detection method through the lens of ID energy. Theoretical results capture the generalization risk of the proposed method. Experimental results across multiple benchmarks demonstrate the effectiveness of our approach. Qianshu Cai, Chao Wu 0001, Yonggang Zhang 0003, Jun Yu 0002, Xinmei Tian 0001 |
NeurIPS | 4 |
| 2025 | A Token is Worth over 1, 000 Tokens: Efficient Knowledge Distillation through Low-Rank CloneabstractTraining high-performing Small Language Models (SLMs) remains computationally expensive, even with knowledge distillation and pruning from larger teacher models.
Existing approaches often face three key challenges: (1) information loss from hard pruning, (2) inefficient alignment of representations, and (3) underutilization of informative activations, particularly from Feed-Forward Networks (FFNs).
To address these challenges, we introduce \textbf{Low-Rank Clone (LRC)}, an efficient pre-training method that constructs SLMs aspiring to behavioral equivalence with strong teacher models.
LRC trains a set of low-rank projection matrices that jointly enable soft pruning by compressing teacher weights, and activation clone by aligning student activations, including FFN signals, with those of the teacher.
This unified design maximizes knowledge transfer while removing the need for explicit alignment modules.
Extensive experiments with open-source teachers such as Llama-3.2-3B-Instruct and Qwen2.5-3B/7B-Instruct show that LRC matches or surpasses the performance of state-of-the-art models trained on trillions of tokens--using only 20B tokens, achieving over \textbf{1,000$\times$} greater training efficiency.
Our codes and model checkpoints are available at https://github.com/CURRENTF/LowRankClone and https://huggingface.co/JitaiHao/LRC-4B-Base. Jitai Hao, Xinyan Xiao, Zhaochun Ren, Jun Yu 0002 |
NeurIPS | 6 |
| 2025 | EditInfinity: Image Editing with Binary-Quantized Generative ModelsabstractAdapting pretrained diffusion-based generative models for text-driven image editing with negligible tuning overhead has demonstrated remarkable potential. A classical adaptation paradigm, as followed by these methods, first infers the generative trajectory inversely for a given source image by image inversion, then performs image editing along the inferred trajectory guided by the target text prompts. However, the performance of image editing is heavily limited by the approximation errors introduced during image inversion by diffusion models, which arise from the absence of exact supervision in the intermediate generative steps. To circumvent this issue, we investigate the parameter-efficient adaptation of binary-quantized generative models for image editing, and leverage their inherent characteristic that the exact intermediate quantized representations of a source image are attainable, enabling more effective supervision for precise image inversion. Specifically, we propose EditInfinity, which adapts Infinity, a binary-quantized generative model, for image editing. We propose an efficient yet effective image inversion mechanism that integrates text prompting rectification and image style preservation, enabling precise image inversion. Furthermore, we devise a holistic smoothing strategy which allows our EditInfinity to perform image editing with high fidelity to source images and precise semantic alignment to the text prompts. Extensive experiments on the PIE-Bench benchmark across add, change, and delete editing operations, demonstrate the superior performance of our model compared to state-of-the-art diffusion-based baselines. Code available at: https://github.com/yx-chen-ust/EditInfinity. Jun Yu 0002, Guangming Lu 0002, Wenjie Pei |
NeurIPS | 3 |
| 2025 | Enhancing LLM Planning for Robotics Manipulation through Hierarchical Procedural Knowledge GraphsabstractLarge Language Models (LLMs) have shown the promising planning capabilities for robotic manipulation, which advances the development of embodied intelligence significantly. However, existing LLM-driven robotic manipulation approaches excel at simple pick-and-place tasks but are insufficient for complex manipulation tasks due to inaccurate procedural knowledge. Besides, for embodied intelligence, equipping a large scale LLM is energy-consuming and inefficient, which affects its real-world application.
To address the above problems, we propose Hierarchical Procedural Knowledge Graphs (\textbf{HP-KG}) to enhance LLMs for complex robotic planning while significantly reducing the demand for LLM scale in robotic manipulation.
Considering that the complex real-world tasks require multiple steps, and each step is composed of robotic-understandable atomic actions, we design a hierarchical knowledge graph structure to model the relationships between tasks, steps, and actions. This design bridges the gap between human instructions and robotic manipulation actions. To construct HP-KG, we develop an automatic knowledge graph construction framework powered by LLM-based multi-agents, which eliminates costly manual efforts while maintaining high-quality graph structures.
The resulting HP-KG encompasses over 40k activity steps across more than 6k household tasks, spanning diverse everyday scenarios. Extensive experiments demonstrate that small scale LLMs (7B) enhanced by our HP-KG significantly improve the planning capabilities, which are stronger than 72B LLMs only. Encouragingly, our approach remains effective on the most powerful GPT-4o model. Jiacong Zhou, Jiaxu Miao, Xianyun Wang, Jun Yu 0002 |
NeurIPS | 4 |
| 2025 | Distribution-Aware Multi-Attention Tri-branch Networks with Feedforward Differential Features for semi-supervised medical image segmentation
Peilian Shi, Shuchang Zhao, Shiqing Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Hongsheng Lu, Jun Yu 0002 |
Expert Syst. Appl. | 10 |
| 2025 | Modality-aware contrast and fusion for multi-modal summarization
Lixin Dai, Tingting Han 0003, Zhou Yu 0001, Jun Yu 0002, Min Tan 0005 |
Neurocomputing | 4 |
| 2025 | Prophet: Prompting Large Language Models With Complementary Answer Heuristics for Knowledge-Based Visual Question AnsweringabstractKnowledge-based visual question answering (VQA) requires external knowledge beyond the image to answer the question. Early studies retrieve required knowledge from explicit knowledge bases (KBs), which often introduces irrelevant information to the question, hence restricting the performance of their models. Recent works have resorted to using a powerful large language model (LLM) as an implicit knowledge engine to acquire the necessary knowledge for answering. Despite the encouraging results achieved by these methods, we argue that they have not fully activated the capacity of the LLM as the provided textual input is insufficient to depict the required visual information to answer the question. In this paper, we present Prophet-a conceptually simple, flexible, and general framework designed to prompt LLM with answer heuristics for knowledge-based VQA. Specifically, we first train a vanilla VQA model on a specific knowledge-based VQA dataset without external knowledge. After that, we extract two types of complementary answer heuristics from the VQA model: answer candidates and answer-aware examples. The two types of answer heuristics are jointly encoded into a formatted prompt to facilitate the LLM's understanding of both the image and question, thus generating a more accurate answer. By incorporating the state-of-the-art LLM GPT-3 (Brown et al. 2020), Prophet significantly outperforms existing state-of-the-art methods on four challenging knowledge-based VQA datasets. Prophet is general that can be instantiated with the combinations of different VQA models (i.e., both discriminative and generative ones) and different LLMs (i.e., both commercial and open-source ones). Moreover, Prophet can also be integrated with modern large multimodal models in different stages, which is named Prophet++, to further improve the capabilities on knowledge-based VQA tasks. Zhou Yu 0001, Xuecheng Ouyang, Zhenwei Shao, Meng Wang 0001, Jun Yu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Adversarial temporal sentence grounding by learning from external data
Tingting Han 0003, Kai Wang 0036, Jun Yu 0002, Sicheng Zhao, Jianping Fan 0001 |
Pattern Recognit. | 3 |
| 2025 | Self-adaptive image-text fusion for medical image classification
Jian Zhang 0026, Kaihao He, Zunlei Feng, Shuifa Sun, Xiaoyan Sun 0006, Zhenming Yuan, Jun Yu 0002 |
Pattern Recognit. | 7 |
| 2025 | Action-Driven Semantic Representation and Aggregation for Video CaptioningabstractVideo captioning, a challenging task that entails generating natural language descriptions of visual content, often fails to effectively grasp the essence of action semantics. To harness the power of action detection to facilitate a deeper understanding of the video content, we propose an action-driven method, named Hierarchical Semantic Representation and Aggregation (HSRA) network. This method explicitly exploits action clues with a hierarchical semantic representation module, which models visual semantics in a three-level structure: “object-action-event”. By employing learnable action queries, our approach injects extensive action semantics into the model, thereby enabling more accurate and context-rich captions. To further enhance semantic alignment and understanding, we introduce a semantic aggregation composed of a semantic interaction module and a semantic refinement module. This component facilitates the alignment of semantics across different levels and emphasizes key information, ultimately leading to significant improvements in semantic consistency between the video and generated captions. We performed extensive evaluations on two well-established public datasets, MSVD and MSR-VTT, and the findings consistently demonstrate that our proposed HSRA network outperforms contemporary state-of-the-art methods. Tingting Han 0003, Yaochen Xu, Jun Yu 0002, Zhou Yu 0001, Sicheng Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | CoDi: Contrastive Disentanglement Generative Adversarial Networks for Zero-Shot Sketch-Based 3D Shape RetrievalabstractSketch-based 3D shape retrieval has attracted increasing attention in recent years. Most existing methods fail to address the zero-shot scenario, and the few dedicated to zero-shot learning encounter the following two issues: 1) the features learned by these methods lack informativeness and generalization, rendering them ineffective in identifying unseen samples; 2) the generation of low-quality samples, aimed at facilitating the recognition of unseen categories, paradoxically diminishes their ability to identify these unseen classes. This paper introduces a novel contrastive disentanglement generative adversarial networks (CoDi) tailored for zero-shot sketch-based 3D shape retrieval. Initially, we introduce a paradoxical feature construction approach designed to assist the networks in capturing certain low-level features. Despite their weak semantic relevance, these features play a crucial role in sample recognition. Subsequently, a SemContrast fusion module is employed to align the semantic space with the prototype embedding space of categories. This alignment facilitates knowledge transfer to unseen classes and promotes the generation of high-quality samples. The networks are jointly trained on real and generated samples to achieve retrieval for unseen categories. Extensive experiments demonstrate a significant improvement in retrieval performance for unseen categories using our method. Min Meng 0001, Wenhang Chen, Jigang Liu, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | MDKAT: Multimodal Decoupling With Knowledge Aggregation and Transfer for Video Emotion RecognitionabstractMultimodal Emotion Recognition (MER) leverages multiple input signals to identify the expressed emotions in user-generated data. Currently, effectively addressing both modality heterogeneity and homogeneity on MER tasks is a challenging issue due to the diversity of multimodal inputs in videos. To address this issue, this work proposes an efficient Multimodal Decoupling Method with Knowledge Aggregation and Transfer (MDKAT) for robust multimodal feature learning in emotional videos. MDKAT is consisted of three key steps: modality-independent feature extraction, modality-specific feature extraction, and multi-loss integration for decoupling. In these three steps, four crucial modules are individually designed to improve different aspects of multimodal learning on MER tasks, including a Cross-modal Feature Fusion (CFF) module for enhancing modality-independent features, an Adaptive Masked Self-Attention (AMSA) module for feature refinement, a Knowledge Aggregation (KA) module for ensuring the semantic similarity of modality-independent features, and a Knowledge Transfer (KT) module for balancing the strengths of different modalities. Experimental results on the typical CMU-MOSI and CMU-MOSEI datasets show that MDKAT obtains superior performance over state-of-the-art methods, demonstrating the effectiveness of MDKAT on MER tasks. Jian Wang 0066, Shuchang Zhao, Shiqing Zhang, Xiaoming Zhao 0002, Jun Yu 0002, Yaowei Wang 0001, Yi Yang 0001, Siwei Ma 0001, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Prompting Video-Language Foundation Models With Domain-Specific Fine-Grained Heuristics for Video Question AnsweringabstractVideo Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and sophisticated cross-modal interaction for accurate inference. Despite advancements in multi-modal pre-trained models and video-language foundation models, these systems often struggle with domain-specific VideoQA due to their generalized pre-training objectives. Addressing this gap necessitates bridging the divide between broad cross-modal knowledge and the specific inference demands of VideoQA tasks. To this end, we introduce HeurVidQA, a framework that leverages domain-specific entity-action heuristics to refine pre-trained video-language foundation models. Our approach treats these models as implicit knowledge engines, employing domain-specific entity-action prompters to direct the model’s focus toward precise cues that enhance reasoning. By delivering fine-grained heuristics, we improve the model’s ability to identify and interpret key entities and actions, thereby enhancing its reasoning capabilities. Extensive evaluations across multiple VideoQA datasets demonstrate that our method significantly outperforms existing models, underscoring the importance of integrating domain-specific knowledge into video-language models for more accurate and context-aware VideoQA. Ting Yu 0002, Kunhao Fu, Shuhui Wang, Qingming Huang, Jun Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Benchmarking and Enhancing Geospatial Visual Reasoning Over Street MapsabstractRecent advances in large multimodal models (LMMs) have enabled substantial progress in various visual question answering (VQA) benchmarks, including the challenging text-centric ones that require a simultaneous understanding of both the visual and textual contents in the images. Despite the prominence of existing text-centric VQA benchmarks, they either have limited textual information or have a limited number of questions requiring complex reasoning skills beyond the basic OCR. To this end, we present SMVQA—a novel text-centric VQA benchmark based on street map images. SMVQA contains more than 10K real-world street map images from the open geospatial database OpenStreetMap. Each image in SMVQA is also associated with detailed geospatial annotations, enabling it to automatically generate up to 57.5K distinctive QA pairs of five representative question types. In addition to the standard test split, SMVQA introduces an extra test split to verify the generalization abilities over out-of-domain images and novel reasoning skills. The evaluation of the state-of-the-art open-source and commercial LMMs reflects the great challenge posed by SMVQA. The latest LMMs, such as GPT-4o, only achieve accuracies of 49.9%, showing plenty of room for improvement. To further improve the latest LMMs’ performance on SMVQA, we introduce a LMM-based agentic framework LHR, which consists of the localizing, highlighting, and reasoning stages. Specifically, LHR first prompts the LMM to localize region-of-interest (RoI) to the question and then highlight the RoI and perform chain-of-thought reasoning for answer prediction. By integrating LHR with GPT-4o, we observe a significant improvement over the vanilla counterpart, showing the effectiveness of our framework. Wenwen Pan 0003, Haiting Zhou, Zhenwei Shao, Shuai Shao 0012, Suguo Zhu, Min Tan 0005, Jun Yu 0002, Zhou Yu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | ScatDiff: Physical Diffusion Model for Electromagnetic Computational ImagingabstractElectromagnetic computational imaging offers a promising solution to electromagnetic inverse scattering problems. Whereas, it is challenged by its ill-posed nature and non-linearity. Traditional iterative methods are often slow and prone to local minima, while recent deep generative models overlook the physical principles that govern the transformation from scattering fields to images of constitute parameters, limiting their interpretability, generalization, and robustness. To address these issues, we propose ScatDiff, a novel Scatter-to-image Diffusion model that integrates electromagnetic data with fundamental physical principles. ScatDiff uses a time-aware, backpropagation-enhanced diffusion to generate noisy images embedded with electromagnetic priors, along with a denoising module that uses cross-attention to adaptively integrate scattering fields. Additionally, a physics-driven reconstruction module incorporates an induced current model into the loss function to enhance interpretability. Experiments on three MNIST variants, Gesture dataset, the “Austria” profile, and the “FoamDielExt” profile show that ScatDiff outperforms traditional iterative methods and deep learning models in both imaging quality and efficiency, with strong generalization and robustness under high noise conditions. Code and datasets are available on https://github.com/Scatdif. Min Tan 0005, Kuiwen Xu, Zhou Yu 0001, Jun Yu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Enhanced Cross-Modal Hashing via Hybrid Distillation and Structural RefinementabstractSince cross-modal hashing requires minimal storage and computation, it is becoming increasingly popular with the exponential growth of multimedia content on the internet. However, the lack of accurate supervisory data has curtailed the effectiveness of unsupervised hashing techniques. Conversely, supervised hashing strategies necessitate considerable human and financial resources for data annotation. To address this limitation, we propose a novel semi-supervised cross-modal hashing method called Enhanced Cross-Modal Hashing via Hybrid Distillation and Structural Refinement (HDSR). Specifically, we first learn the features of inter-modal and inter-instance similarity relationships through pointwise semantic alignment and listwise similarity partial order learning, respectively, to extract refined structural representations from partially labeled data. Secondly, by fusing inter-modal similarity to construct higher-order affinity matrices, we precisely delineate the semantic correlation information across cross-modal data, facilitating stable self-supervised training of unlabeled data through the application of momentum fusion strategies. Finally, the refined structural representation of labeled data is transferred into unlabeled branches through hybrid distillation, enhancing the performance of cross-modal hash learning by generating compact and accurate hash codes. The proposed HDSR is compared with several state-of-the-art deep cross-modal hashing methods on three widely used benchmark databases, and the experimental results verify its efficiency and superiority. Zhiwen Yu 0002, Kaixiang Yang 0001, Jun Yu 0002, Huanqiang Zeng, C. L. Philip Chen |
IEEE Trans. Image Process. | 4 |
| 2025 | Multi-Scale Group Agent Attention-Based Graph Convolutional Decoding Networks for 2D Medical Image SegmentationabstractAutomated medical image segmentation plays a crucial role in assisting doctors in diagnosing diseases. Feature decoding is a critical yet challenging issue for medical image segmentation. To address this issue, this work proposes a novel feature decoding network, called multi-scale group agent attention-based graph convolutional decoding networks (MSGAA-GCDN), to learn local-global features in graph structures for 2D medical image segmentation. The proposed MSGAA-GCDN combines graph convolutional network (GCN) and a lightweight multi-scale group agent attention (MSGAA) mechanism to represent features globally and locally within a graph structure. Moreover, in skip connections a simple yet efficient attention-based upsampling convolution fusion (AUCF) module is designed to enhance encoder-decoder feature fusion in both channel and spatial dimensions. Extensive experiments are conducted on three typical medical image segmentation tasks, namely Synapse abdominal multi-organs, Cardiac organs, and Polyp lesions. Experimental results demonstrate that the proposed MSGAA-GCDN outperforms the state-of-the-art methods, and the designed MSGAA is a lightweight yet effective attention architecture. The proposed MSGAA-GCDN can be easily taken as a plug-and-play decoder cascaded with other encoders for general medical image segmentation tasks. Shuchang Zhao, Shiqing Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Hongsheng Lu, Jun Yu 0002, Qi Tian 0001 |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | Consistency Conditioned Memory Augmented Dynamic Diagnosis Model for Medical Visual Question AnsweringabstractMedical Visual Question Answering (Med-VQA) holds immense promise as an invaluable medical assistance aid, offering timely diagnostic outcomes based on medical images and accompanying questions, thereby supporting medical professionals in making accurate clinical decisions. However, Med-VQA is still in its infancy, with existing solutions falling short in imitating human diagnostic processes and ensuring result consistency. To address these challenges, we propose a Consistency Conditioned Memory augmented Dynamic diagnosis model (CoCoMeD), incorporating two core components: a dynamic memory diagnosis engine and a consistency-conditioned enforcer. The dynamic memory diagnosis engine enables intricate diagnostic interactions by retaining vital visual cues from medical images and iteratively updating pertinent memories. This dynamic reasoning capability mirrors the cognitive processes observed in skilled medical diagnosticians, thus effectively enhancing the model's ability to reason over diverse medical visual facts and patient-specific questions. Moreover, to strengthen diagnostic coherence, the consistency-conditioned enforcer imposes coherence constraints linking interrelated questions with identical medical facts, ensuring the credibility and reliability of its diagnostic outcomes. Additionally, we present C-SLAKE, an extended Med-VQA dataset encompassing diverse medical image types, and categorized diagnostic question-answer pairs for consistent Med-VQA evaluation on rich medical sources. Comprehensive experiments on DME and C-SLAKE showcase CoCoMeD's superior performance and potential to advance trustworthy multi-source medical question answering. Ting Yu 0002, Binhui Ge, Shuhui Wang, Qingming Huang, Jun Yu 0002 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Adapter-Enhanced Hierarchical Cross-Modal Pre-Training for Lightweight Medical Report GenerationabstractAutomatic medical report generation is an emerging field that aims to transform medical images into descriptive, clinically relevant narratives, potentially reducing the workload for radiologists significantly. Despite substantial progress, the increasing model parameter size and corresponding marginal performance gains have limited further development and application. To address this challenge, we introduce an Adapter-enhanced Hierarchical cross-modal Pre-training (AHP) strategy for lightweight medical report generation. This approach significantly reduces the pre-trained model's parameter size while maintaining superior report generation performance through our proposed spatial adapters. To further address the issue of inadequate representation of visual space details, we employ a convolutional stem combined with hierarchical injectors and extractors, fully integrating with traditional Vision Transformers to achieve more comprehensive visual representations. Additionally, our cross-modal pre-training model effectively handles the inherent complex visual-textual relationships in medical imaging. Extensive experiments on multiple datasets, including IU X-Ray, MIMIC-CXR, and bladder pathology, demonstrate our model's exceptional generalization and transfer performance in downstream medical report generation tasks, highlighting AHP's potential in significantly reducing model parameters while enhancing report generation accuracy and efficiency. Ting Yu 0002, Wangwen Lu, Weidong Han 0001, Qingming Huang, Jun Yu 0002, Ke Zhang 0029 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Spatio-Temporal and Retrieval-Augmented Modeling for Chest X-Ray Report GenerationabstractChest X-ray report generation has attracted increasing research attention. However, most existing methods neglect the temporal information and typically generate reports conditioned on a fixed number of images. In this paper, we propose STREAM: Spatio-Temporal and REtrieval-Augmented Modelling for automatic chest X-ray report generation. It mimics clinical diagnosis by integrating current and historical studies to interpret the present condition (temporal), with each study containing images from multi-views (spatial). Concretely, our STREAM is built upon an encoder-decoder architecture, utilizing a large language model (LLM) as the decoder. Overall, spatio-temporal visual dynamics are packed as visual prompts and regional semantic entities are retrieved as textual prompts. First, a token packer is proposed to capture condensed spatio-temporal visual dynamics, enabling the flexible fusion of images from current and historical studies. Second, to augment the generation with existing knowledge and regional details, a progressive semantic retriever is proposed to retrieve semantic entities from a preconstructed knowledge bank as heuristic text prompts. The knowledge bank is constructed to encapsulate anatomical chest X-ray knowledge into structured entities, each linked to a specific chest region. Extensive experiments on public datasets have shown the state-of-the-art performance of our method. Related codes and the knowledge bank are available at https://github.com/yangyan22/STREAM. Xiaoxing You, Ke Zhang 0029, Zhenqi Fu, Xianyun Wang, Jiajun Ding, Jiamei Sun, Zhou Yu 0001, Qingming Huang, Weidong Han 0001, Jun Yu 0002 |
IEEE Trans. Medical Imaging | 11 |
| 2025 | Boundary Discretization and Reliable Classification Network for Temporal Action DetectionabstractTemporal action detection aims to recognize the action category and determine each action instance's starting and ending time in untrimmed videos. The mixed method has demonstrated notable performance by integrating both anchor-based and anchor-free approaches. However, while it leverages the strengths of each method, it also retains their respective limitations. For instance, the anchor-based approach depends on manually crafted anchors tailored to specific datasets, while the anchor-free approach predicts potential action instances at each temporal position, resulting in a significant number of false positives in category prediction. The inclusion of these limitations undermines the potential benefits of the mixed method. In this paper, we propose a novel Boundary Discretization and Reliable Classification Network (BDRC-Net) that addresses the issues above by introducing boundary discretization and reliable classification modules. Specifically, the boundary discretization module (BDM) elegantly merges anchor-based and anchor-free approaches in the form of boundary discretization, eliminating the need for the traditional handcrafted anchor design. Furthermore, the reliable classification module (RCM) predicts reliable global action categories to reduce false positives. Extensive experiments conducted on different benchmarks demonstrate that our proposed method achieves competitive detection performance. Zhenying Fang, Jun Yu 0002, Richang Hong |
IEEE Trans. Multim. | 2 |
| 2025 | Imp: Highly Capable Large Multimodal Models for Mobile DevicesabstractBy harnessing the capabilities of large language models (LLMs), recent large multimodal models (LMMs) have shown remarkable versatility in open-world multimodal understanding. Nevertheless, they are usually parameter-heavy and computation-intensive, thus hindering their applicability in resource-constrained scenarios. To this end, several lightweight LMMs have been proposed successively to maximize the capabilities under constrained scale (e.g., 3B). Despite the encouraging results achieved by these methods, most of them only focus on one or two aspects of the design space, and the key design choices that influence model capability have not yet been thoroughly investigated. In this paper, we conduct a systematic study for lightweight LMMs from the aspects of model architecture, training strategy, and training data. Based on our findings, we obtain Imp—a family of highly capable LMMs at the 2B$\sim$4B scales. Notably, our Imp-3B model steadily outperforms all the existing lightweight LMMs of similar size, and even surpasses the state-of-the-art LMMs at the 13B scale. With low-bit quantization and resolution reduction techniques, our Imp model can be deployed on a Qualcomm Snapdragon 8Gen3 mobile chip with a high inference speed of about 13 tokens/s. Zhenwei Shao, Zhou Yu 0001, Jun Yu 0002, Xuecheng Ouyang, Lihao Zheng 0001, Zhenbiao Gai, Zhenzhong Kuang, Jiajun Ding |
IEEE Trans. Multim. | 3 |
| 2025 | MossVLN: Memory-Observation Synergistic System for Continuous Vision-Language NavigationabstractNavigating in continuous environments with vision-language cues presents critical challenges, particularly in the accuracy of waypoint prediction and the quality of navigation decision-making. Traditional methods, which predominantly rely on spatial data from depth images or straightforward RGB-depth integrations, frequently encounter difficulties in environments where waypoints share similar spatial characteristics, leading to erroneous navigational outcomes. Additionally, the capacity for effective navigation decisions is often hindered by the inadequacies of traditional topological maps and the issue of uneven data sampling. In response, this paper introduces a robust memory-observation synergistic vision-language navigation framework to substantially enhance the navigation capabilities of agents operating in continuous environments. We present an advanced observation-driven waypoint predictor that effectively utilizes spatial data and integrates aligned visual and textual cues to significantly improve the accuracy of waypoint predictions within complex real-world scenarios. Additionally, we develop a strategic memory-observation planning approach that leverages memory panoramic environmental data and detailed current observation information, enabling more informed and precise navigation decisions. Our framework sets new performance benchmarks on the VLN-CE dataset, achieving a 60.25% Success Rate (SR) and a 50.89% Path Length Score (SPL) on the R2R-CE dataset's unseen validation splits. Furthermore, when adapted to a discrete environment, our model also shows exceptional performance on the R2R dataset, achieving a 74% SR and a 64% SPL on the unseen validation split. The code is available athttps://github.com/OpenMICG/MossVLN. Ting Yu 0002, Qiongjie Cui, Qingming Huang, Jun Yu 0002 |
IEEE Trans. Multim. | 5 |
| 2025 | Semi-Supervised RGB-D Hand Gesture Recognition via Mutual Learning of Self-Supervised ModelsabstractHuman hand gesture recognition is important to human–computer interaction. Gesture recognition based on RGB and Depth (RGB-D) data exploits both RGB and depth images to provide comprehensive results. However, the research under scenario with insufficient annotated data is not adequate. In view of the problem, our insight is to perform self-supervised learning with respect to each modality, transfer the learned information to modality-specific classifiers, and then fuse their results for final decision. To this end, we propose a semi-supervised hand gesture recognition method known as Mutual Learning of Rotation-Aware Gesture Predictors (MLRAGP), which exploits unlabeled training RGB and depth images via self-supervised learning and achieves multi-modal decision fusion through deep mutual learning. For each modality, we rotate both labeled and unlabeled images to fixed angles and train an angle predictor to predict the angles, then we use the feature extraction part of the angle predictor to construct the category predictor and train it through labeled data. We subsequently fuse the category predictors about both modalities by impelling each of them to simulate the probability estimation produced by the other, and making the prediction of labeled images to approach the ground truth annotation. During the training of category predictor and mutual learning, the parameters of feature extractors can be slighted fine-tuned to avoid under-fitting. Experimental results on NTU-Microsoft Kinect Hand Gesture dataset and Washington RGB-D dataset demonstrate the superiority of this framework to existing methods. Jian Zhang 0026, Kaihao He, Ting Yu 0016, Jun Yu 0002, Zhenming Yuan |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Graph Context Transformation Learning for Progressive Correspondence PruningabstractMost of existing correspondence pruning methods only concentrate on gathering the context information as much as possible while neglecting effective ways to utilize such information. In order to tackle this dilemma, in this paper we propose Graph Context Transformation Network (GCT-Net) enhancing context information to conduct consensus guidance for progressive correspondence pruning. Specifically, we design the Graph Context Enhance Transformer which first generates the graph network and then transforms it into multi-branch graph contexts. Moreover, it employs self-attention and cross-attention to magnify characteristics of each graph context for emphasizing the unique as well as shared essential information. To further apply the recalibrated graph contexts to the global domain, we propose the Graph Context Guidance Transformer. This module adopts a confident-based sampling strategy to temporarily screen high-confidence vertices for guiding accurate classification by searching global consensus between screened vertices and remaining ones. The extensive experimental results on outlier removal and relative pose estimation clearly demonstrate the superior performance of GCT-Net compared to state-of-the-art methods across outdoor and indoor datasets. Junwen Guo, Guobao Xiao, Shiping Wang, Jun Yu 0002 |
AAAI | 4 |
| 2024 | BCLNet: Bilateral Consensus Learning for Two-View Correspondence PruningabstractCorrespondence pruning aims to establish reliable correspondences between two related images and recover relative camera motion. Existing approaches often employ a progressive strategy to handle the local and global contexts, with a prominent emphasis on transitioning from local to global, resulting in the neglect of interactions between different contexts. To tackle this issue, we propose a parallel context learning strategy that involves acquiring bilateral consensus for the two-view correspondence pruning task. In our approach, we design a distinctive self-attention block to capture global context and parallel process it with the established local context learning module, which enables us to simultaneously capture both local and global consensuses. By combining these local and global consensuses, we derive the required bilateral consensus. We also design a recalibration block, reducing the influence of erroneous consensus information and enhancing the robustness of the model. The culmination of our efforts is the Bilateral Consensus Learning Network (BCLNet), which efficiently estimates camera pose and identifies inliers (true correspondences). Extensive experiments results demonstrate that our network not only surpasses state-of-the-art methods on benchmark datasets but also showcases robust generalization abilities across various feature extraction techniques. Noteworthily, BCLNet obtains significant improvement gains over the second best method on unknown outdoor dataset, and obviously accelerates model training speed. Xiangyang Miao, Guobao Xiao, Shiping Wang, Jun Yu 0002 |
AAAI | 4 |
| 2024 | Multi-Domain Deep Learning from a Multi-View Perspective for Cross-Border E-commerce SearchabstractBuilding click-through rate (CTR) and conversion rate (CVR) prediction models for cross-border e-commerce search requires modeling the correlations among multi-domains. Existing multi-domain methods would suffer severely from poor scalability and low efficiency when number of domains increases. To this end, we propose a Domain-Aware Multi-view mOdel (DAMO), which is domain-number-invariant, to effectively leverage cross-domain relations from a multi-view perspective. Specifically, instead of working in the original feature space defined by different domains, DAMO maps everything to a new low-rank multi-view space. To achieve this, DAMO firstly extracts multi-domain features in an explicit feature-interactive manner. These features are parsed to a multi-view extractor to obtain view-invariant and view-specific features. Then a multi-view predictor inputs these two sets of features and outputs view-based predictions. To enforce view-awareness in the predictor, we further propose a lightweight view-attention estimator to dynamically learn the optimal view-specific weights w.r.t. a view-guided loss. Extensive experiments on public and industrial datasets show that compared with state-of-the-art models, our DAMO achieves better performance with lower storage and computational costs. In addition, deploying DAMO to a large-scale cross-border e-commence platform leads to 1.21%, 1.76%, and 1.66% improvements over the existing CGC-based model in the online AB-testing experiment in terms of CTR, CVR, and Gross Merchandises Value, respectively. Yinfu Feng, Yunan Ye, Min Tan 0005, Rong Xiao 0005, Haihong Tang, Jiajun Ding, Jun Yu 0002 |
AAAI | 9 |
| 2024 | Integrating Representation Subspace Mapping with Unimodal Auxiliary Loss for Attention-based Multimodal Emotion RecognitionabstractMultimodal emotion recognition (MER) aims to identify emotions by utilizing affective information from multiple modalities. Due to the inherent disparities among these heterogeneous modalities, there is a large modality gap in their representations, leading to the challenge of fusing multiple modalities for MER. To address this issue, this work proposes a novel attention-based MER framework by integrating representation subspace mapping with unimodal auxiliary loss for enhancing multimodal fusion capabilities. Initially, a representation subspace mapping module is proposed to map each modality into two distinct subspaces. One is modality-public, enabling the acquisition of common representations and reducing the discrepancies across modalities. The other is modality-unique, retaining the unique characteristics of each modality while eliminating redundant inter-modal attributes. Then, a cross-modality attention is leveraged to bridge the modality gap in unique representations and facilitate modality adaptation. Additionally, our method designs an unimodal auxiliary loss to remove the noise unrelated to emotion classification, resulting in robust and meaningful representations for MER. Comprehensive experiments are conducted on the IEMOCAP and MSP-Improv datasets, and experiment results show that our method achieves superior performance to state-of-the-art MER methods. Keywords: Multimodal emotion recognition, representation subspace mapping, cross-modality attention, unimodal auxiliary loss, fusion Xulong Du, Xingnan Zhang, Shiqing Zhang, Xiaoming Zhao 0002, Jun Yu 0002, Liangliang Lou |
LREC/COLING | 8 |
| 2024 | GLOW: Global Layout Aware Attacks on Object DetectionabstractAdversarial attacks aim to perturb images such that a predictor outputs incorrect results. Due to the limited research in structured attacks, imposing consistency checks on natural multi-object scenes is a practical defense against conventional adversarial attacks. More desired attacks should be able to fool defenses with such consistency checks. Therefore, we present the first approach GLOW that copes with various attack requests by generating global layout-aware adversar-ial attacks, in which both categorical and geometric layout constraints are explicitly established. Specifically, we focus on object detection tasks and given a victim image, GLOW first localizes victim objects according to target labels. And then it generates multiple attack plans, together with their context-consistency scores. GLOW, on the one hand, is ca-pable of handling various types of requests, including single or multiple victim objects, with or without specified victim objects. On the other hand, it produces a consistency score for each attack plan, reflecting the overall contextual consistency that both semantic category and global scene layout are considered. We conduct our experiments on MS COCO and Pascal. Extensive experimental results demonstrate that we can achieve about 30% average relative improvement compared to state-of-the-art methods in conventional single object attack request; Moreover, such superiority is also valid across more generic attack requests, under both white-box and zero-query black-box settings. Finally, we conduct comprehensive human analysis, which not only validates our claim further but also provides strong evidence that our evaluation metrics reflect human reviews well. Jun Bao, Buyu Liu, Kui Ren 0001, Jun Yu 0002 |
CVPR | 4 |
| 2024 | Facial Identity Anonymization via Intrinsic and Extrinsic Attention DistractionabstractThe unprecedented capture and application of face images raise increasing concerns on anonymization to fight against privacy disclosure. Most existing methods may suffer from the problem of excessive change of the identity-independent information or insufficient identity protection. In this paper, we present a new face anonymization approach by distracting the intrinsic and extrinsic identity attentions. On the one hand, we anonymize the identity information in the feature space by distracting the intrinsic identity attention. On the other, we anonymize the visual clues (i.e. appearance and geometry structure) by distracting the extrinsic identity attention. Our approach allows for flexible and intuitive manipulation of face appearance and geometry structure to produce diverse results, and it can also be used to instruct users to perform personalized anonymization. We conduct extensive experiments on multiple datasets and demonstrate that our approach outperforms state-of-the-art methods. Zhenzhong Kuang, Yingjie Shen, Jun Yu 0002 |
CVPR | 5 |
| 2024 | Latent Representation Reorganization for Face Privacy ProtectionabstractThe issue of face privacy protection has aroused wide social concern along with the increasing applications of face images. The latest methods focus on achieving a good privacy-utility tradeoff so that the protected results can still be used to support the downstream computer vision tasks. However, they may suffer from limited flexibility in manipulating this tradeoff because the practical requirements may vary under different scenarios. In this paper, we present a novel recurrent latent representation reorganization (LReOrg) framework to deal with the problem. LReOrg relies on two key modules to deal with the privacy-utility tradeoff, where the first one is responsible for anonymizing the privacy sensitive information and the other is responsible for recovering the destroyed useful insensitive information according to user requirements. LReOrg is advantageous in: (a) enabling users to recurrently process fine-grained attributes; (b) providing flexible control over privacy-utility tradeoff by manipulating which attributes to anonymize or preserve using cross-modal keywords; and (c) eliminating the need of data annotations for network training. The experimental results on benchmark datasets have reported the superior ability of our approach for providing flexible protection on facial information. Zhenzhong Kuang, Jianan Lu, Chenhui Hong, Haobin Huang, Suguo Zhu, Jun Yu 0002, Jianping Fan 0007 |
ACM Multimedia | 7 |
| 2024 | MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability
Buyu Liu, Kai Wang 0036, Jun Bao, Tingting Han 0003, Jun Yu 0002 |
ACM Multimedia | 6 |
| 2024 | Advancing Incremental Few-Shot Semantic Segmentation via Semantic-Guided Relation Alignment and Adaptation
Yuan Zhou 0016, Xin Chen 0033, Yanrong Guo, Jun Yu 0002, Richang Hong, Qi Tian 0001 |
MMM (1) | 4 |
| 2024 | Learnability Matters: Active Learning for Video CaptioningabstractThis work focuses on the active learning in video captioning. In particular, we propose to address the learnability problem in active learning, which has been brought up by collective outliers in video captioning and neglected in the literature. To start with, we conduct a comprehensive study of collective outliers, exploring their hard-to-learn property and concluding that ground truth inconsistency is one of the main causes. Motivated by this, we design a novel active learning algorithm that takes three complementary aspects, namely learnability, diversity, and uncertainty, into account. Ideally, learnability is reflected by ground truth consistency. Under the active learning scenario where ground truths are not available until human involvement, we measure the consistency on estimated ground truths, where predictions from off-the-shelf models are utilized as approximations to ground truths. These predictions are further used to estimate sample frequency and reliability, evincing the diversity and uncertainty respectively. With the help of our novel caption-wise active learning protocol, our algorithm is capable of leveraging knowledge from humans in a more effective yet intellectual manner. Results on publicly available video captioning datasets with diverse video captioning models demonstrate that our algorithm outperforms SOTA active learning methods by a large margin, e.g. we achieve about 103% of full performance on CIDEr with 25% of human annotations on MSR-VTT. Buyu Liu, Jun Bao, Min Zhang 0005, Jun Yu 0002 |
NeurIPS | 6 |
| 2024 | ZS-SRT: An efficient zero-shot super-resolution training method for Neural Radiance Fields
Yongbo He, Chengkai Wang, Zhenzhong Kuang, Jiajun Ding, Fei-wei Qin, Jun Yu 0002, Jianping Fan 0001 |
Neurocomputing | 8 |
| 2024 | Multi2Human: Controllable human image generation with multimodal controls
Xiaoling Gu, Shengwenzhuo Xu, Yongkang Wong, Zizhao Wu, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli |
Neurocomputing | 5 |
| 2024 | 3D human pose estimation with multi-hypotheses gated transformer
Xiena Dong, Jian Zhang 0026, Jun Yu 0002, Ting Yu 0016 |
Multim. Syst. | 3 |
| 2024 | GTADT: Gated tone-sensitive acne grading via augmented domain transfer
Min Tan 0005, Ruirui Wang, Ankur Purwar, Tao Jin 0004, Jun Yu 0002, Alex Chichung Kot |
Multim. Tools Appl. | 5 |
| 2024 | Regularly Truncated M-Estimators for Learning With Noisy LabelsabstractThe sample selection approach is very popular in learning with noisy labels. As deep networks "learn pattern first", prior methods built on sample selection share a similar training procedure: the small-loss examples can be regarded as clean examples and used for helping generalization, while the large-loss examples are treated as mislabeled ones and excluded from network parameter updates. However, such a procedure is arguably debatable from two folds: (a) it does not consider the bad influence of noisy labels in selected small-loss examples; (b) it does not make good use of the discarded large-loss examples, which may be clean or have meaningful information for generalization. In this paper, we propose regularly truncated M-estimators (RTME) to address the above two issues simultaneously. Specifically, RTME can alternately switch modes between truncated M-estimators and original M-estimators. The former can adaptively select small-losses examples without knowing the noise rate and reduce the side-effects of noisy labels in them. The latter makes the possibly clean examples but with large losses involved to help generalization. Theoretically, we demonstrate that our strategies are label-noise-tolerant. Empirically, comprehensive experimental results show that our method can outperform multiple baselines and is robust to broad noise types and levels. Xiaobo Xia, Pengqian Lu, Chen Gong 0002, Bo Han 0003, Jun Yu 0001, Jun Yu 0002, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Latent Semantic Consensus for Deterministic Geometric Model FittingabstractEstimating reliable geometric model parameters from the data with severe outliers is a fundamental and important task in computer vision. This paper attempts to sample high-quality subsets and select model instances to estimate parameters in the multi-structural data. To address this, we propose an effective method called Latent Semantic Consensus (LSC). The principle of LSC is to preserve the latent semantic consensus in both data points and model hypotheses. Specifically, LSC formulates the model fitting problem into two latent semantic spaces based on data points and model hypotheses, respectively. Then, LSC explores the distributions of points in the two latent semantic spaces, to remove outliers, generate high-quality model hypotheses, and effectively estimate model instances. Finally, LSC is able to provide consistent and reliable solutions within only a few milliseconds for general multi-structural model fitting, due to its deterministic fitting nature and efficiency. Compared with several state-of-the-art model fitting methods, our LSC achieves significant superiority for the performance of both accuracy and speed on synthetic data and real images. Guobao Xiao, Jun Yu 0002, Jiayi Ma 0001, Deng-Ping Fan, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Semantic-aware hyper-space deformable neural radiance fields for facial avatar reconstruction
Kaixin Jin, Xiaoling Gu, Zhenzhong Kuang, Zizhao Wu, Min Tan 0005, Jun Yu 0002 |
Pattern Recognit. Lett. | 7 |
| 2024 | MTDAN: A Lightweight Multi-Scale Temporal Difference Attention Networks for Automated Video Depression DetectionabstractDeep learning based video depression analysis has been recently an interesting and challenging topic. Most of existing works focus on learning single-scale facial dynamics of participants for depression detection. Besides, they usually adopt expensive deep learning models with high computational complexity, resulting in difficulty in real-time clinical applications. To address these two issues, this work proposes a lightweight Multi-scale Temporal Difference Attention Networks (MTDAN) integrating the temporal difference and attention mechanism to model both short-term and long-term temporal facial behaviors for automated video depression detection. Initially, two simple yet effective sub-branches, i.e., a Short-term Temporal Difference Attention Network (ST-TDAN), and a Long-term Temporal Difference Attention Network (LT-TDAN), are designed to perform individually short-term and long-term depressive behavior modeling. Then, a simple Interactive Multi-head Attention Fusion (IMHAF) strategy is employed for integrating short-term and long-term spatiotemporal features, followed by a linear fully-collected layer for depression score prediction. Experiments on two public AVEC2013 and AVEC2014 datasets show that our proposed method not only achieves highly competitive performance to state-of-the-art methods, but also has much smaller computational complexity than them on video depression detection tasks. Shiqing Zhang, Xingnan Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Mingyue Niu, Ziping Zhao 0001, Jun Yu 0002, Qi Tian 0001 |
IEEE Trans. Affect. Comput. | 7 |
| 2024 | MSGA-Net: Progressive Feature Matching via Multi-Layer Sparse Graph AttentionabstractFeature matching is an essential computer vision task that requires the establishment of high-quality correspondences between two images. Constructing sparse dynamic graphs and extracting contextual information by searching for neighbors in feature space is a prevalent strategy in numerous previous works. Nonetheless, these works often neglect the potential connections between dynamic graphs from different layers, leading to underutilization of available information. To tackle this issue, we introduce a Sparse Dynamic Graph Interaction block for feature matching. This innovation facilitates the implicit establishment of dependencies by enabling interaction and aggregation among dynamic graphs across various layers. In addition, we design a novel Multiple Sparse Transformer to enhance the capture of the global context from the sparse graph. This block selectively mines significant global contextual information along spatial and channel dimensions, respectively. Ultimately, we present the Multi-layer Sparse Graph Attention Network (MSGA-Net), a framework designed to predict probabilities of correspondences as inliers and to recover camera poses. Experimental results demonstrate that our proposed MSGA-Net surpasses state-of-the-art methods on challenging indoor and outdoor datasets. Code will be available at https://github.com/gongzhepeng/MSGA-Net. Zhepeng Gong, Guobao Xiao, Ziwei Shi, Riqing Chen, Jun Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Semantic Disentanglement Adversarial Hashing for Cross-Modal RetrievalabstractCross-modal hashing has gained considerable attention in cross-modal retrieval due to its low storage cost and prominent computational efficiency. However, preserving more semantic information in the compact hash codes to bridge the modality gap still remains challenging. Most existing methods unconsciously neglect the influence of modality-private information on semantic embedding discrimination, leading to unsatisfactory retrieval performance. In this paper, we propose a novel deep cross-modal hashing method, called Semantic Disentanglement Adversarial Hashing (SDAH), to tackle these challenges for cross-modal retrieval. Specifically, SDAH is designed to decouple the original features of each modality into modality-common features with semantic information and modality-private features with disturbing information. After the preliminary decoupling, the modality-private features are shuffled and treated as positive interactions to enhance the learning of modality-common features, which can significantly boost the discriminative and robustness of semantic embeddings. Moreover, the variational information bottleneck is introduced in the hash feature learning process, which can avoid the loss of a large amount of semantic information caused by the high-dimensional feature compression. Finally, the discriminative and compact hash codes can be computed directly from the hash features. A large number of comparative and ablation experiments show that SDAH achieves superior performance than other state-ofthe- art methods. Min Meng 0001, Jiaxuan Sun, Jigang Liu, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D ScenesabstractThree-Dimensional (3D) dense captioning is an emerging vision-language bridging task that aims to generate multiple detailed and accurate descriptions for 3D scenes. It presents significant potential and challenges due to its closer representation of the real world compared to 2D visual captioning, as well as complexities in data collection and processing of 3D point cloud sources. Despite the popularity and success of existing methods, there is a lack of comprehensive surveys summarizing the advancements in this field, which hinders its progress. In this paper, we provide a comprehensive review of 3D dense captioning, covering task definition, architecture classification, dataset analysis, evaluation metrics, and in-depth prosperity discussions. Based on a synthesis of previous literature, we refine a standard pipeline that serves as a common paradigm for existing methods. We also introduce a clear taxonomy of existing models, summarize technologies involved in different modules, and conduct detailed experiment analysis. Instead of a chronological order introduction, we categorize the methods into different classes to facilitate exploration and analysis of the differences and connections among existing techniques. We also provide a reading guideline to assist readers with different backgrounds and purposes in reading efficiently. Furthermore, we propose a series of promising future directions for 3D dense captioning by identifying challenges and aligning them with the development of related tasks, offering valuable insights and inspiring future research in this field. Our aim is to provide a comprehensive understanding of 3D dense captioning, foster further investigations, and contribute to the development of novel applications in multimedia and related domains. Ting Yu 0016, Shuhui Wang, Weiguo Sheng 0001, Qingming Huang, Jun Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Learning to Discover Knowledge: A Weakly-Supervised Partial Domain Adaptation ApproachabstractDomain adaptation has shown appealing performance by leveraging knowledge from a source domain with rich annotations. However, for a specific target task, it is cumbersome to collect related and high-quality source domains. In real-world scenarios, large-scale datasets corrupted with noisy labels are easy to collect, stimulating a great demand for automatic recognition in a generalized setting, i.e., weakly-supervised partial domain adaptation (WS-PDA), which transfers a classifier from a large source domain with noises in labels to a small unlabeled target domain. As such, the key issues of WS-PDA are: 1) how to sufficiently discover the knowledge from the noisy labeled source domain and the unlabeled target domain, and 2) how to successfully adapt the knowledge across domains. In this paper, we propose a simple yet effective domain adaptation approach, termed as self-paced transfer classifier learning (SP-TCL), to address the above issues, which could be regarded as a well-performing baseline for several generalized domain adaptation tasks. The proposed model is established upon the self-paced learning scheme, seeking a preferable classifier for the target domain. Specifically, SP-TCL learns to discover faithful knowledge via a carefully designed prudent loss function and simultaneously adapts the learned knowledge to the target domain by iteratively excluding source examples from training under the self-paced fashion. Extensive evaluations on several benchmark datasets demonstrate that SP-TCL significantly outperforms state-of-the-art approaches on several generalized domain adaptation tasks. Code is available at https://github.com/mc-lan/SP-TCL. Mengcheng Lan, Min Meng 0001, Jun Yu 0002, Jigang Wu |
IEEE Trans. Image Process. | 3 |
| 2024 | Multi-Granularity Contrastive Cross-Modal Collaborative Generation for End-to-End Long-Term Video Question AnsweringabstractLong-term Video Question Answering (VideoQA) is a challenging vision-and-language bridging task focusing on semantic understanding of untrimmed long-term videos and diverse free-form questions, simultaneously emphasizing comprehensive cross-modal reasoning to yield precise answers. The canonical approaches often rely on off-the-shelf feature extractors to detour the expensive computation overhead, but often result in domain-independent modality-unrelated representations. Furthermore, the inherent gradient blocking between unimodal comprehension and cross-modal interaction hinders reliable answer generation. In contrast, recent emerging successful video-language pre-training models enable cost-effective end-to-end modeling but fall short in domain-specific ratiocination and exhibit disparities in task formulation. Toward this end, we present an entirely end-to-end solution for long-term VideoQA: Multi-granularity Contrastive cross-modal collaborative Generation (MCG) model. To derive discriminative representations possessing high visual concepts, we introduce Joint Unimodal Modeling (JUM) on a clip-bone architecture and leverage Multi-granularity Contrastive Learning (MCL) to harness the intrinsically or explicitly exhibited semantic correspondences. To alleviate the task formulation discrepancy problem, we propose a Cross-modal Collaborative Generation (CCG) module to reformulate VideoQA as a generative task instead of the conventional classification scheme, empowering the model with the capability for cross-modal high-semantic fusion and generation so as to rationalize and answer. Extensive experiments conducted on six publicly available VideoQA datasets underscore the superiority of our proposed method. Ting Yu 0016, Kunhao Fu, Jian Zhang 0026, Qingming Huang, Jun Yu 0002 |
IEEE Trans. Image Process. | 5 |
| 2024 | Token-Mixer: Bind Image and Text in One Embedding Space for Medical Image ReportingabstractMedical image reporting focused on automatically generating the diagnostic reports from medical images has garnered growing research attention. In this task, learning cross-modal alignment between images and reports is crucial. However, the exposure bias problem in autoregressive text generation poses a notable challenge, as the model is optimized by a word-level loss function using the teacher-forcing strategy. To this end, we propose a novel Token-Mixer framework that learns to bind image and text in one embedding space for medical image reporting. Concretely, Token-Mixer enhances the cross-modal alignment by matching image-to-text generation with text-to-text generation that suffers less from exposure bias. The framework contains an image encoder, a text encoder and a text decoder. In training, images and paired reports are first encoded into image tokens and text tokens, and these tokens are randomly mixed to form the mixed tokens. Then, the text decoder accepts image tokens, text tokens or mixed tokens as prompt tokens and conducts text generation for network optimization. Furthermore, we introduce a tailored text decoder and an alternative training strategy that well integrate with our Token-Mixer framework. Extensive experiments across three publicly available datasets demonstrate Token-Mixer successfully enhances the image-text alignment and thereby attains a state-of-the-art performance. Related codes are available at https://github.com/yangyan22/Token-Mixer. Jun Yu 0002, Zhenqi Fu, Ke Zhang 0029, Ting Yu 0016, Xianyun Wang, Hanliang Jiang, Junhui Lv, Qingming Huang, Weidong Han 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Attribute Prototype-Guided Iterative Scene Graph for Explainable Radiology Report GenerationabstractThe potential of automated radiology report generation in alleviating the time-consuming tasks of radiologists is increasingly being recognized in medical practice. Existing report generation methods have evolved from using image-level features to the latest approach of utilizing anatomical regions, significantly enhancing interpretability. However, directly and simplistically using region features for report generation compromises the capability of relation reasoning and overlooks the common attributes potentially shared across regions. To address these limitations, we propose a novel region-based Attribute Prototype-guided Iterative Scene Graph generation framework (AP-ISG) for report generation, utilizing scene graph generation as an auxiliary task to further enhance interpretability and relational reasoning capability. The core components of AP-ISG are the Iterative Scene Graph Generation (ISGG) module and the Attribute Prototype-guided Learning (APL) module. Specifically, ISSG employs an autoregressive scheme for structural edge reasoning and a contextualization mechanism for relational reasoning. APL enhances intra-prototype matching and reduces inter-prototype semantic overlap in the visual space to fully model the potential attribute commonalities among regions. Extensive experiments on the MIMIC-CXR with Chest ImaGenome datasets demonstrate the superiority of AP-ISG across multiple metrics. Ke Zhang 0029, Jun Yu 0002, Jianping Fan 0007, Hanliang Jiang, Qingming Huang, Weidong Han 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | FedSea: Federated Learning via Selective Feature Alignment for Non-IID Multimodal DataabstractThe growing demands for privacy protection challenge the joint training of one model by leveraging multiple datasets. Federated learning (FL) provides a new way to overcome this challenge and has attracted many research interests, which enables multiple parties to collaboratively train a machine learning model without exchanging their local data. Despite some success, the non-independent and identically distributed (non-IID) data distributions in different parties remain challenging and easily damage the performance of FL methods, specifically for the heterogeneous multimodal data. Existing FL studies on non-IID data settings are often dedicated to the label space, neglecting the non-IID issues in feature space, thus limiting their performance when the parties with non-IID multimodal data. This paper proposes a newFederated learning method viaSelective featureAlignment (FedSea) to align representations across multiple parties in the feature space. FedSea uses a domain adversarial learning framework consisting of an affine-transform-based generator and a gradient-reversal-based client discriminator to perform IID transformation and reduce data source distinguishability, respectively. An attention-based mask module and a feature IID confidence quantification method are introduced to effectively address the diverse feature non-IID levels across multimodal data. Comprehensive experiments are conducted on three widely-used public datasets and one large-scale industrial dataset, showing FedSea has: 1) better performance than state-of-the-art FL methods on both multimodal and single-modal datasets; 2) superior feature alignment ability on non-IID datasets, and 3) good model interpretability. Min Tan 0005, Yinfu Feng, Lingqiang Chu, Jingcheng Shi, Rong Xiao 0005, Haihong Tang, Jun Yu 0002 |
IEEE Trans. Multim. | 7 |
| 2024 | DSIS-DPR:Structured Instance Segmentation and Diffusion Prior Refinement for Dental Anatomy LearningabstractInstance segmentation in medical imaging plays a crucial role in clinical diagnostic tasks, and have shown promising performance in practical applications. In this paper, we discuss a more fine-grained instance segmentation task: dental structured instance segmentation based on panoramic radiographs. However, direct segmentation of tooth structures encounters inherent challenges. Traditional instance segmentation networks often fall short in capturing intricate internal features, and exacerbated by the frequent blurring found in medical imaging, which can result in the deficiency of anatomical details. To deal with these problems, we propose a novel framework called DSISDPR, which combines a dental structured instance segmentation (DSIS) network with an enhanced diffusion prior refinement (DPR) method. Specifically, our innovatively designed structureaware network leverages fine-grained feature fusion, acquiring a richer representation of internal anatomical structures. With the integration of adversarial learning, the model is primed to deliver holistic and subtle predictions of tooth structures. Furthermore, taking inspiration from dentists’ inherent ability to utilize prior knowledge, such as understanding dental structures to label invisible anatomical structures, we propose a diffusion inpainting to refine the results of DSIS without additional annotations. Equipped with built-in structure learning, DPR is capable of modifying anomalies within each predicted segmentation, resulting in a more robust and complete structured segmentation result. Meanwhile, we ensure rigorous oversight over the reconstruction of areas affected by abnormalities, ensuring that any introduced adjustments minimally disrupt the wellpredicted structured segmentation results. Extensive experiments have demonstrated that our DSIS-DPR outperforms all existing classical instance segmentation networks. The collected dataset is available:https://github.com/Zzz512/TSD. Xianyun Wang, Linhong Wang, Zhenchen Yang, Jiacong Zhou, Yuchen Zheng 0004, Richang Hong, Jun Yu 0002, Fan Yang 0063 |
IEEE Trans. Multim. | 8 |
| 2024 | Semi-Supervised Medical Report Generation via Graph-Guided Hybrid Feature ConsistencyabstractMedical report generation generates the corresponding report according to the given radiology image, which has been attracting increasing research interest. However, existing methods mainly adopt supervised training which rely on large amount of medical reports that are actually unavailable owing to the labor-intensive labeling process and privacy protection protocol. In the meanwhile, the intrinsic relationships between local pathological changes in the image are often ignored, which actually are important hints to high quality report generation. To this end, we propose a Relation-Aware Mean Teacher (RAMT) framework, which follows a standard mean teacher paradigm for semi-supervised report generation. The key to the encoder of the backbone network is the Graph-guided Hybrid Feature Encoding (GHFE) module, which exploits a prior disease knowledge graph to encode the intrinsic relations between pathological changes into the graph embedding and learns a word dictionary to retrieve the semantic embedding for each potential pathological change. GHFE combines the graph embedding, semantic embedding and visual features to form hybrid features, which are sent to a Transformer-based decoder for report generation. Extensive experiments on the MIMIC-CXR and IU X-Ray datasets demonstrate the effectiveness of our proposed approach. Ke Zhang 0029, Hanliang Jiang, Jian Zhang 0026, Qingming Huang, Jianping Fan 0007, Jun Yu 0002, Weidong Han 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Multi-Task Paired Masking With Alignment Modeling for Medical Vision-Language Pre-TrainingabstractIn recent years, the growing demand for medical imaging diagnosis has placed a significant burden on radiologists. As a solution, Medical Vision-Language Pre-training (Med-VLP) methods have been proposed to learn universal representations from medical images and reports, benefiting downstream tasks without requiring fine-grained annotations. However, existing methods have overlooked the importance of cross-modal alignment in joint image-text reconstruction, resulting in insufficient cross-modal interaction. To address this limitation, we propose a unified Med-VLP framework based on Multi-task Paired Masking with Alignment (MPMA) to integrate the cross-modal alignment task into the joint image-text reconstruction framework to achieve more comprehensive cross-modal interaction, while a Global and Local Alignment (GLA) module is designed to assist self-supervised paradigm in obtaining semantic representations with rich domain knowledge. Furthermore, we introduce a Memory-Augmented Cross-Modal Fusion (MA-CMF) module to fully integrate visual information to assist report reconstruction and fuse the multi-modal representations adequately. Experimental results demonstrate that the proposed unified approach outperforms previous methods in all downstream tasks, including uni-modal, cross-modal, and multi-modal tasks. Ke Zhang 0029, Jun Yu 0002, Hanliang Jiang, Jianping Fan 0007, Qingming Huang, Weidong Han 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Video Moment Retrieval With Noisy LabelsabstractVideo moment retrieval (VMR) aims to localize the target moment in an untrimmed video according to the given nature language query. The existing algorithms typically rely on clean annotations to train their models. However, making annotations by human labors may introduce much noise. Thus, the video moment retrieval models will not be well trained in practice. In this article, we present a simple yet effective video moment retrieval framework via bottom-up schema, which is in end-to-end manners and robust to noisy label training. Specifically, we extract the multimodal features by syntactic graph convolutional networks and multihead attention layers, which are fused by the cross gates and the bilinear approach. Then, the feature pyramid networks are constructed to encode plentiful scene relationships and capture high semantics. Furthermore, to mitigate the effects of noisy annotations, we devise the multilevel losses characterized by two levels: a frame-level loss that improves noise tolerance and an instance-level loss that reduces adverse effects of negative instances. For the frame level, we adopt the Gaussian smoothing to regard noisy labels as soft labels through the partial fitting. For the instance level, we exploit a pair of structurally identical models to let them teach each other during iterations. This leads to our proposed robust video moment retrieval model, which experimentally and significantly outperforms the state-of-the-art approaches on standard public datasets ActivityCaption and textually annotated cooking scene (TACoS). We also evaluate the proposed approach on the different manual annotation noises to further demonstrate the effectiveness of our model. Wenwen Pan 0003, Zhou Zhao 0001, Wencan Huang, Liyong Fu, Jun Yu 0002, Fei Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | PAINT: Photo-realistic Fashion Design SynthesisabstractIn this article, we investigate a new problem of generating a variety of multi-view fashion designs conditioned on a human pose and texture examples of arbitrary sizes, which can replace the repetitive and low-level design work for fashion designers. To solve this challenging multi-modal image translation problem, we propose a novel Photo-reAlistic fashIon desigN synThesis (PAINT) framework, which decomposes the framework into three manageable stages. In the first stage, we employ a Layout Generative Network (LGN) to transform an input human pose into a series of person semantic layouts. In the second stage, we propose a Texture Synthesis Network (TSN) to synthesize textures on all transformed semantic layouts. Specifically, we design a novel attentive texture transfer mechanism for precisely expanding texture patches to the irregular clothing regions of the target fashion designs. In the third stage, we leverage an Appearance Flow Network (AFN) to generate the fashion design images of other viewpoints from a single-view observation by learning 2D multi-scale appearance flow fields. Experimental results demonstrate that our method is capable of generating diverse photo-realistic multi-view fashion design images with fine-grained appearance details conditioned on the provided multiple inputs. The source code and trained models are available at https://github.com/gxl-groups/PAINT . Xiaoling Gu, Jie Huang 0033, Yongkang Wong, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Recurrent Appearance Flow for Occlusion-Free Virtual Try-OnabstractImage-based virtual try-on aims at transferring a target in-shop garment onto a reference person, and has garnered significant attention from the research communities recently. However, previous methods have faced severe challenges in handling occlusion problems. To address this limitation, we classify occlusion problems into three types based on the reference person’s arm postures: single-arm occlusion , two-arm non-crossed occlusion , and two-arm crossed occlusion . Specifically, we propose a novel Occlusion-Free Virtual Try-On Network (OF-VTON) that effectively overcomes these occlusion challenges. The OF-VTON framework consists of two core components: (i) a new Recurrent Appearance Flow based Deformation (RAFD) model that robustly aligns the in-shop garment to the reference person by adopting a multi-task learning strategy . This model jointly produces the dense appearance flow to warp the garment and predicts a human segmentation map to provide semantic guidance for the subsequent image synthesis model. (ii) a powerful Multi-mask Image SynthesiS (MISS) model that generates photo-realistic try-on results by introducing a new mask generation and selection mechanism . Experimental results demonstrate that our proposed OF-VTON significantly outperforms existing state-of-the-art methods by mitigating the impact of occlusion problems. Our code is available at https://github.com/gxl-groups/OF-VTON . Xiaoling Gu, Junkai Zhu, Yongkang Wong, Zizhao Wu, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Effective Video Summarization by Extracting Parameter-Free Motion AttentionabstractVideo summarization remains a challenging task despite increasing research efforts. Traditional methods focus solely on long-range temporal modeling of video frames, overlooking important local motion information that cannot be captured by frame-level video representations. In this article, we propose the Parameter-free Motion Attention Module (PMAM) to exploit the crucial motion clues potentially contained in adjacent video frames, using a multi-head attention architecture. The PMAM requires no additional training for model parameters, leading to an efficient and effective understanding of video dynamics. Moreover, we introduce the Multi-feature Motion Attention Network (MMAN), integrating the PMAM with local and global multi-head attention based on object-centric and scene-centric video representations. The synergistic combination of local motion information, extracted by the proposed PMAM, with long-range interactions modeled by the local and global multi-head attention mechanism, can significantly enhance the performance of video summarization. Extensive experimental results on the benchmark datasets, SumMe and TVSum, demonstrate that the proposed MMAN outperforms other state-of-the-art methods, resulting in remarkable performance gains. Tingting Han 0003, Jun Yu 0002, Zhou Yu 0001, Sicheng Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Knowledge-Constrained Answer Generation for Open-Ended Video Question AnsweringabstractOpen-ended Video question answering (open-ended VideoQA) aims to understand video content and question semantics to generate the correct answers. Most of the best performing models define the problem as a discriminative task of multi-label classification. In real-world scenarios, however, it is difficult to define a candidate set that includes all possible answers. In this paper, we propose a Knowledge-constrained Generative VideoQA Algorithm (KcGA) with an encoder-decoder pipeline, which enables out-of-domain answer generation through an adaptive external knowledge module and a multi-stream information control mechanism. We use ClipBERT to extract the video-question features, extract framewise object-level external knowledge from a commonsense knowledge base and compute the contextual-aware episode memory units via an attention based GRU to form the external knowledge features, and exploit multi-stream information control mechanism to fuse video-question and external knowledge features such that the semantic complementation and alignment are well achieved. We evaluate our model on two open-ended benchmark datasets to demonstrate that we can effectively and robustly generate high-quality answers without restrictions of training data. Guocheng Niu, Xinyan Xiao, Jian Zhang 0026, Xi Peng 0001, Jun Yu 0002 |
AAAI | 6 |
| 2023 | ShiftDDPMs: Exploring Conditional Diffusion Models by Shifting Diffusion TrajectoriesabstractDiffusion models have recently exhibited remarkable abilities to synthesize striking image samples since the introduction of denoising diffusion probabilistic models (DDPMs). Their key idea is to disrupt images into noise through a fixed forward process and learn its reverse process to generate samples from noise in a denoising way. For conditional DDPMs, most existing practices relate conditions only to the reverse process and fit it to the reversal of unconditional forward process. We find this will limit the condition modeling and generation in a small time window. In this paper, we propose a novel and flexible conditional diffusion model by introducing conditions into the forward process. We utilize extra latent space to allocate an exclusive diffusion trajectory for each condition based on some shifting rules, which will disperse condition modeling to all timesteps and improve the learning capacity of model. We formulate our method, which we call ShiftDDPMs, and provide a unified point of view on existing related methods. Extensive qualitative and quantitative experiments on image synthesis demonstrate the feasibility and effectiveness of ShiftDDPMs. Zijian Zhang 0002, Zhou Zhao 0001, Jun Yu 0002, Qi Tian 0001 |
AAAI | 3 |
| 2023 | ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed VideosabstractBuilding benchmarks to systemically analyze different capabilities of video question answering (VideoQA) models is challenging yet crucial. Existing benchmarks often use non-compositional simple questions and suffer from language biases, making it difficult to diagnose model weaknesses incisively. A recent benchmark AGQA [8] poses a promising paradigm to generate QA pairs automatically from pre-annotated scene graphs, enabling it to measure diverse reasoning abilities with granular control. However, its questions have limitations in reasoning about the fine-grained semantics in videos as such information is absent in its scene graphs. To this end, we present ANetQA, a large-scale benchmark that supports fine-grained compositional reasoning over the challenging untrimmed videos from ActivityNet [4]. Similar to AGQA, the QA pairs in ANetQA are automatically generated from annotated video scene graphs. The fine-grained properties of ANetQA are reflected in the following: (i) untrimmed videos with fine-grained semantics; (ii) spatio-temporal scene graphs with fine-grained taxonomies; and (iii) diverse questions generated from fine-grained templates. ANetQA attains 1.4 billion unbalanced and 13.4 million balanced QA pairs, which is an order of magnitude larger than AGQA with a similar number of videos. Comprehensive experiments are performed for state-of-the-art methods. The best model achieves 44.5% accuracy while human performance tops out at 84.5%, leaving sufficient room for improvement. Zhou Yu 0001, Lixiang Zheng, Zhou Zhao 0001, Fei Wu 0001, Jianping Fan 0001, Kui Ren 0001, Jun Yu 0002 |
CVPR | 7 |
| 2023 | Prompting Large Language Models with Answer Heuristics for Knowledge-Based Visual Question AnsweringabstractKnowledge-based visual question answering (VQA) requires external knowledge beyond the image to answer the question. Early studies retrieve required knowledge from explicit knowledge bases (KBs), which often introduces irrelevant information to the question, hence restricting the performance of their models. Recent works have sought to use a large language model (i.e., GPT-3 [3]) as an implicit knowledge engine to acquire the necessary knowledge for answering. Despite the encouraging results achieved by these methods, we argue that they have not fully activated the capacity of GPT-3 as the provided input information is insufficient. In this paper, we present Prophet-a conceptually simple framework designed to$prompt$GPT-3 with answer heuristics for knowledge-based VQA. Specifically, we first train a vanilla VQA model on a specific knowledge-based VQA dataset without external knowledge. After that, we extract two types of complementary answer heuristics from the model: answer candidates and answer-aware examples. Finally, the two types of answer heuristics are encoded into the prompts to enable GPT-3 to better comprehend the task thus enhancing its capacity. Prophet significantly outperforms all existing state-of-the-art methods on two challenging knowledge-based VQA datasets, OK-VQA and A-OKVQA, delivering 61.1% and 55.7% accuracies on their testing sets, respectively. Zhenwei Shao, Zhou Yu 0001, Meng Wang 0001, Jun Yu 0002 |
CVPR | 4 |
| 2023 | Graph Matching with Bi-level Noisy CorrespondenceabstractIn this paper, we study a novel and widely existing problem in graph matching (GM), namely, Bi-level Noisy Correspondence (BNC), which refers to node-level noisy correspondence (NNC) and edge-level noisy correspondence (ENC). In brief, on the one hand, due to the poor recognizability and viewpoint differences between images, it is inevitable to inaccurately annotate some keypoints with offset and confusion, leading to the mismatch between two associated nodes, i.e., NNC. On the other hand, the noisy node-to-node correspondence will further contaminate the edge-to-edge correspondence, thus leading to ENC. For the BNC challenge, we propose a novel method termed Contrastive Matching with Momentum Distillation. Specifically, the proposed method is with a robust quadratic contrastive loss which enjoys the following merits: i) better exploring the node-to-node and edge-to-edge correlations through a GM customized quadratic contrastive learning paradigm; ii) adaptively penalizing the noisy assignments based on the confidence estimated by the momentum teacher. Extensive experiments on three real-world datasets show the robustness of our model compared with 12 competitive baselines. The code is available at https://github.com/XLearning-SCU/2023-ICCV-COMMON. Yijie Lin 0001, Mouxing Yang, Jun Yu 0002, Peng Hu 0002, Changqing Zhang 0002, Xi Peng 0001 |
ICCV | 3 |
| 2023 | Follow-me: Deceiving Trackers with Fabricated PathsabstractConvolutional Neural Networks (CNNs) are vulnerable to adversarial attacks in which visually imperceptible perturbations can deceive CNN-based models. While current research on adversarial attacks in single object tracking exists, it overlooks a critical aspect of manipulating predicted trajectories to follow user-defined paths regardless of the actual location of the targeted object. To address this, we propose the very first white-box attack algorithm that is capable of deceiving victim trackers by compelling them to generate trajectories that adhere to predetermined counterfeit paths. Specifically, we focus on Siamese-based trackers as our victim models. Given an arbitrary counterfeit path, we first decompose it into discrete target locations in each frame, with the assumption of constant velocity. These locations are converted to heatmap anchors, which represent the offset of their location from the target object's location in the previous frame. Later on, we design a novel loss function to minimize the gap between above-mentioned anchors and our predicted ones. Finally, the gradients computed by such loss are used to update the original video, resulting in our adversarial video. To validate our ideas, we design three sets of counterfeit paths as well as novel evaluation metrics to measure the path-following properties. Experiments with two victim models on three publicly available datasets, OTB100, VOT2018, and VOT2016, demonstrate that our algorithm not only outperforms SOTA methods significantly under conventional evaluation metrics, e.g. 90% and 68.4% precision and successful rate drop on OTB100, but also follows the counterfeit paths well, which is beyond any existing attack methods. The source code is available at https://github.com/loushengtao/Follow-me. Shengtao Lou, Buyu Liu, Jun Bao, Jiajun Ding, Jun Yu 0002 |
ACM Multimedia | 5 |
| 2023 | Contrastive Perturbation Network for Weakly Supervised Temporal Sentence Grounding
Tingting Han 0003, Yuanxin Lv, Zhou Yu 0001, Jun Yu 0002, Jianping Fan 0001 |
PRCV (1) | 4 |
| 2023 | MARN: Multi-level Attentional Reconstruction Networks for Weakly Supervised Video Temporal Grounding
Yijun Song, Jingwen Wang 0003, Lin Ma 0002, Jun Yu 0002, Jinxiu Liang, Zhou Yu 0001 |
Neurocomputing | 4 |
| 2023 | Multi-level uncertainty aware learning for semi-supervised dental panoramic caries segmentation
Xianyun Wang, Sizhe Gao, Kaisheng Jiang, Huicong Zhang, Linhong Wang, Jun Yu 0002, Fan Yang 0063 |
Neurocomputing | 7 |
| 2023 | EGRA-NeRF: Edge-Guided Ray Allocation for Neural Radiance Fields
Zhenbiao Gai, Zhenyang Liu, Min Tan 0005, Jiajun Ding, Jun Yu 0002, Mingzhao Tong, Junqing Yuan |
Image Vis. Comput. | 5 |
| 2023 | An efficient multi-path structure with staged connection and multi-scale mechanism for text-to-image synthesis
Jiajun Ding, Beili Liu, Jun Yu 0002, Huanlei Guo, Kenong Shen |
Multim. Syst. | 3 |
| 2023 | Position constrained network for 3D human pose estimation
Xiena Dong, Jun Yu 0002, Jian Zhang 0026 |
Multim. Syst. | 2 |
| 2023 | LPR: learning point-level temporal action localization through re-trainingabstractAbstract Point-level temporal action localization (PTAL) aims to locate action instances in untrimmed videos with only one timestamp annotation for each action instance. Existing methods adopt the localization-by-classification paradigm to locate action boundaries in the temporal class activation map (TCAM) by thresholding, also known as TCAM-based method. However, TCAM-based methods are limited by the gap between classification and localization tasks, since TCAM is generated by a classification network. To address this issue, we propose a re-training framework for the PTAL task, also known as LPR. This framework consists of two stages: pseudo-label generation and re-training. In the pseudo-label generation stage, we propose a feature embedding module based on a transformer encoder to capture global context features and optimize pseudo-labels’ quality by leveraging point-level annotations. In the re-training stage, LPR uses the above pseudo-labels as supervision to locate action instances with a temporal action localization network rather than generating TCAMs. Furthermore, to alleviate the effects of label noise in the pseudo-labels, we propose a joint learning classification module (JLCM) in the re-training stage. This module contains two classification sub-modules that simultaneously predict action categories and are guided by a jointly determined clean set for network training. The proposed framework achieves state-of-the-art localization performance on both the THUMOS’14 and BEOID datasets. Zhenying Fang, Jianping Fan 0001, Jun Yu 0002 |
Multim. Syst. | 3 |
| 2023 | Import vertical characteristic of rain streak for single image deraining
Zhexin Zhang, Jiajun Ding, Jun Yu 0002, Yiming Yuan, Jianping Fan 0001 |
Multim. Syst. | 3 |
| 2023 | Concept Parser With Multimodal Graph Learning for Video CaptioningabstractConventional video captioning methods are either stage-wise or simple end-to-end. While the former might introduce additional noise when exploiting off-the-shelf models to provide extra information, the latter suffers from lacking high-level cues. Therefore, a more desired framework should be able to capture multi-aspects of videos consistently. To this end, we present a concept-aware and task-specific model named CAT that accounts for both low-level visual and high-level concept cues, and incorporates them effectively in an end-to-end manner. Specifically, low-level visual and high-level concept features are obtained from the video transformer and concept parser of CAT. And a concept loss is further introduced to regularize the learning process of concept parser w.r.t. generated pseudo ground truth. To combine multi-level features, a caption transformer is later introduced in CAT, where visual and concept features are the inputs and caption is its output. In particular, we make critical design choices in the caption transformer to learn to exploit these cues with a multi-modal graph. This is achieved by a graph loss that enforces effective learning of intra and inter correlations between multi-level cues. Extensive experiments on three benchmark datasets demonstrate that CAT achieves 2.3 and 0.7 improvements in the CIDEr metric on MSVD and MSR-VTT compared to the state-of-the-art method SwinBERT and also achieves a competitive result on VATEX. Bofeng Wu, Buyu Liu, Jun Bao, Peng Xi, Jun Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Electromagnetic Imaging Boosted Visual Object Recognition Under Difficult Visual ConditionsabstractObject imaging and recognition under difficult visual conditions is extremely challenging due to the captured low-quality images, and traditional optical-based recognition methods always fail in this task. In this paper, we propose to utilize the visual-microwave image pairs captured by both visual cameras and microwave sensors for imaging and recognition. To address the heavy noises in the low-quality optical images, we retrieve the physically quantitative images from associated scattered field data, and enhance visual features by both optical and retrieval images. We develop a cross-modal Enhanced Attentive Visual-Microwave Fusion (EAVMF) object recognition model to jointly learn the cross-modal generator and multimodal recognizer. In addition, an attention module for the visual subnetwork is utilized to highlight the regions of interest. Two multimodal datasets with synthetic visual-microwave image pairs are built to simulate the difficult visual condition. The numerical results on these datasets demonstrate that: 1) both the multimodal fusion, cross-modal enhancement, and visual attention module can enhance the performance; and 2) compared with existing methods, the proposed EAVMF not only performs better in terms of accuracy but also has good scalability and one-shot learning ability. Min Tan 0005, Tao Jin 0004, Danhui Ye, Kuiwen Xu, Xiaoling Gu, Jun Yu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Dual-Level Adaptive and Discriminative Knowledge Transfer for Cross-Domain RecognitionabstractUnsupervised domain adaptation is an appealing technique to learn robust classifiers for unlabeled target domain by borrowing knowledge from well-established source domain. However, previous works mainly suffer from two limitations: 1) the classifier trained on labeled source data may be prone to overfitting the source distribution, lowering its performance on the target domain; 2) the adaptation process will be misled by conditional distribution matching using hard pseudo labels of target samples. This paper presents a Dual-Level Adaptive and Discriminative (DLAD) classifier learning framework, in which transfer classifier and distribution adaptation can be mutually beneficial for effective knowledge transfer. Specifically, we aim to achieve a domain-level adaptive classifier by considering structural risk minimization (SRM) on both domains and performing weighted distribution adaptation, which facilitates joint classifier learning in a semi-supervised manner. To further achieve a class-level discriminative classifier, we explicitly leverage unlabeled target data to promote classifier learning based on class probabilities, which refines the decision boundary to be more discriminative for unlabeled target data. To the best of our knowledge, DLAD is the first attempt to consider the principle of SRM on the target domain, which significantly boosts the discriminative power of transfer classifier and yields a tighter generalization bound. Experimental evaluations on several standard cross-domain datasets show that DLAD significantly outperforms other competitive methods. Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu, Ligang Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Joint Embedding of Deep Visual and Semantic Features for Medical Image Report GenerationabstractMedical image report generation (MeIRG) aims at generating associated diagnosis descriptions with natural language sentences from medical images, which is essential in the computer-aided diagnosis system. Nevertheless, this task remains challenging in that medical images and linguistic expressions should be understood jointly which however show great discrepancies in the modality. To fill this visual-to-semantic gap, we propose a novel framework that follows the encoder-decoder pipeline. Our framework is characterized by encoding both deep visual and semantic embeddings through a triple-branch network (TriNet) during the encoding phase. The visual attention branch captures attended visual embeddings from medical images with the soft-attention mechanism. The medical report (MeRP) embedding branch predicts semantic report embeddings. The embedding branch of medical subject headings (MeSH) obtains semantic embeddings of related medical tags as complementary information. Then, outputs of these branches are fused and fed into a decoder for the report generation. Experimental results on two benchmark datasets have demonstrated the excellent performance of our method. Related codes are available athttps://github.com/yangyan22/Medical-Report-Generation-TriNet. Jun Yu 0002, Jian Zhang 0026, Weidong Han 0001, Hanliang Jiang, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2023 | Bilaterally Slimmable Transformer for Elastic and Efficient Visual Question AnsweringabstractRecent advances in Transformer architectures [1] have brought remarkable improvements to visual question answering (VQA). Nevertheless, Transformer-based VQA models are usually deep and wide to guarantee good performance, so they can only run on powerful GPU servers and cannot run on capacity-restricted platforms such as mobile phones. Therefore, it is desirable to learn an elastic VQA model that supports adaptive pruning at runtime to meet the efficiency constraints of different platforms. To this end, we present the bilaterally slimmable Transformer (BST), a general framework that can be seamlessly integrated into arbitrary Transformer-based VQA models to train a single model once and obtain various slimmed submodels of different widths and depths. To verify the effectiveness and generality of this method, we integrate the proposed BST framework with three typical Transformer-based VQA approaches, namely MCAN [2], UNITER [3], and CLIP-ViL [4], and conduct extensive experiments on two commonly-used benchmark datasets. In particular, one slimmed MCAN$_\mathsf {BST}$submodel achieves comparable accuracy on VQA-v2, while being 0.38× smaller in model size and having 0.27× fewer FLOPs than the reference MCAN model. The smallest MCAN$_\mathsf {BST}$submodel only has 9M parameters and 0.16G FLOPs during inference, making it possible to deploy it on a mobile device with less than 60 ms latency. Zhou Yu 0001, Zitian Jin, Jun Yu 0002, Mingliang Xu 0001, Jianping Fan 0007 |
IEEE Trans. Multim. | 3 |
| 2022 | ESCNet: Gaze Target Detection with the Understanding of 3D ScenesabstractThis paper aims to address the single image gaze target detection problem. Conventional methods either focus on 2D visual cues or exploit additional depth information in a very coarse manner. In this work, we propose to explicitly and effectively model 3D geometry under challenging scenario where only 2D annotations are available. We first obtain 3D point clouds of given scene with estimated depth and reference objects. Then we figure out the front-most points in all possible 3D directions of given person. These points are later leveraged in our ESCNet model. Specifically, ESCNet consists of geometry and scene parsing modules. The former produces an initial heatmap inferring the probability that each front-most point has been looking at according to estimated 3D gaze direction. And the latter further explores scene contextual cues to regulate detection results. We validate our idea on two publicly available dataset, GazeFollow and VideoAttentionTarget, and demon-strate the state-of-the-art performance. Our method also beats the human in terms of AUC on GazeFollow. Our code can be found here https://github.com/bjj9/ESCNet. Jun Bao, Buyu Liu, Jun Yu 0002 |
CVPR | 3 |
| 2022 | Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross- Modal Denoising NetworksabstractAudio-Guided video object segmentation is a challenging problem in visual analysis and editing, which automatically separates foreground objects from the background in a video sequence according to the referring audio expressions. However, existing referring video object segmentation works mainly focus on the guidance of text-based referring expressions, due to the lack of modeling the semantic representations of audio-video interaction contents. In this paper, we consider the problem of audio-guided video semantic segmentation from the viewpoint of end-to-end denoising encoder-decoder network learning. We propose the wavelet-based encoder network to learn the cross-modal representations of the video contents with audio-form queries. Specifically, we adopt the multi-head cross-modal attention layers to explore the potential relations of video and query contents. A 2-dimension discrete wavelet trans-form is merged into the transformer encoder to decompose the audio-video features. Next, we maximize mutual information between the encoded features and multi-modal features after cross-modal attention layers to enhance the au-dio guidance. Then, a self attention-free decoder network is developed to generate the target masks with frequency-domain transforms. In addition, we construct the first large-scale audio-guided video semantic segmentation dataset. The extensive experiments show the effectiveness of our method11Code is available at: https://github.com/asudahkzj/Wnet.git. Wenwen Pan 0003, Zhou Zhao 0001, Jieming Zhu, Xiuqiang He 0001, Lianli Gao, Jun Yu 0002, Fei Wu 0001, Qi Tian 0001 |
CVPR | 8 |
| 2022 | Group Correspondence: A Statistical Perspective for Incomplete Multi-View Clustering AugmentationabstractCross-view consistency is the fundamental property of multiview clustering. However, in incomplete multi-view scenarios, existing methods can only pursue consistency through the paired data while ignoring the information in unpaired data. In this paper, we show a new insight from the data pattern and provide a novel perspective to incorporate unpaired data for consistency maximization by mining group correspondence. We first formulate cross-view consistency in a statistical perspective to by-pass the strict demand of instance correspondence, and then propose a technique to construct corresponding groups across views to enhance the objective of consistency maximization. Our proposal can be used as a universal plug-in to augment existing approaches. We test the efficacy and generality of our proposal by adapting it to two base methods as augmentations and comparing the augmented models against the original ones and other baselines. Experiment results demonstrate the effectiveness of our proposal and validate the value of our insight. Tianyou Liang, Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu |
ICME | 4 |
| 2022 | Triple Disentangling Network for Unsupervised Domain AdaptationabstractMost existing unsupervised domain adaptation methods learn domain-invariant representations with entangled domain in-formation, semantic information, and instance information. Differently, in this paper, we propose a Triple Disentangling Network (TDN), to disentangle these three types of information and then predict the target labels merely using semantic information. Specifically, TDN consists of a reconstruction module and a disentanglement module. In the reconstruction module, TDN utilizes a variational auto-encoder to re-construct the domain, semantic, and instance latent variables behind the data. In the disentanglement module, adversar-ial learning, discriminative clustering, and instance separation are seamlessly integrated to disentangle these three sets of re-constructed latent variables. Significantly, TDN can not only effectively alleviate the negative transfer of outliers through disentangling instance information, but also disentangle se-mantic information more thoroughly by exploring discriminative structure knowledge. Experimental studies on two bench-mark datasets demonstrate the superiority of TDN. Zhuanghui Wu, Tianyou Liang, Min Meng 0001, Jigang Liu, Jun Yu 0002, Jigang Wu |
ICME | 5 |
| 2022 | Delegate-based Utility Preserving Synthesis for Pedestrian Image AnonymizationabstractThe rapidly growing application of pedestrian images has aroused wide concern on visual privacy protection because personal information is under the risk of privacy disclosure. Anonymization is regarded as an effective solution by identity obfuscation. Most recent methods focus on face, but it is not enough when the presence of human body carries lots of identifiable information. This paper presents a new delegate-based utility preserving synthesis (DUPS) approach for pedestrian image anonymization. This is challenging because one may expect that the anonymized image can still be useful in various computer vision tasks. We model DUPS as an adaptive translation process from source to target. To provide a comprehensive identity protection, we first perform anonymous delegate sampling based on image-level differential privacy. To synthesize anonymous images, we then introduce an adaptive translation network and optimize it with a multi-task loss function. Our approach is theoretically sound and can generate diverse results by preserving data utility. The experiments on multiple datasets show that DUPS can not only achieve superior anonymization performance against deep pedestrian recognizers, but also can obtain a better tradeoff between privacy protection and utility preservation compared with state-of-the-art methods. Zhenzhong Kuang, Longbin Teng, Zhou Yu 0001, Jun Yu 0002, Jianping Fan 0001, Mingliang Xu 0001 |
ACM Multimedia | 4 |
| 2022 | Unsupervised Domain Adaptation Integrating Transformer and Mutual Information for Cross-Corpus Speech Emotion RecognitionabstractThis paper focuses on an interesting task, i.e., unsupervised cross-corpus Speech Emotion Recognition (SER), in which the labelled training (source) corpus and the unlabelled testing (target) corpus have different feature distributions, resulting in the discrepancy between the source and target domains. To address this issue, this paper proposes an unsupervised domain adaptation method integrating Transformers and Mutual Information (MI) for cross-corpus SER. Initially, our method employs encoder layers of Transformers to capture long-term temporal dynamics in an utterance from the extracted segment-level log-Mel spectrogram features, thereby producing the corresponding utterance-level features for each utterance in two domains. Then, we propose an unsupervised feature decomposition method with a hybrid Max-Min MI strategy to separately learn domain-invariant features and domain-specific features from the extracted mixed utterance-level features, in which the discrepancy between two domains is eliminated as much as possible and meanwhile their individual characteristic is preserved. Finally, an interactive Multi-Head attention fusion strategy is designed to learn the complementarity between domain-invariant features and domain-specific features so that they can be interactively fused for SER. Extensive experiments on the IEMOCAP and MSP-Improv datasets demonstrate the effectiveness of our proposed method on unsupervised cross-corpus SER tasks, outperforming state-of-the-art unsupervised cross-corpus SER methods. Shiqing Zhang, Ruixin Liu, Yijiao Yang, Xiaoming Zhao 0002, Jun Yu 0002 |
ACM Multimedia | 5 |
| 2022 | Complex-valued Reinforcement Learning Based Dynamic Beamforming Design for IRS Aided Time-Varying Downlink ChannelabstractThe intelligent reflecting surface (IRS) is an artificial metasurface making the communication environment smart and controllable. The IRS on an aerial platform (AIRS) expands the wireless network to the three-dimensional space, thus improving the degree of freedom (DoF) for the signal adjustment. Since the AIRS-enabled wireless channel is generally time-variant in practice, herein, this paper considers the time-varying characteristic of the downlink channels, and proposes a complex-valued ResNet-based deep Q-learning (DQN) algorithm to maximize the sum-rate at user equipment (UE) side, by jointly designing the transmit beamforming at base station (BS) side and the reconfigurable phase shifts at AIRS side. Our results reveal that the proposed complex-valued deep reinforcement learning (DRL) approach shows stronger generalization ability in comparison with the real-valued DRL algorithms, and is validated to be able to mitigate the problem of gradient vanishing and improve the performance over the time-varying downlink channels. Mengfan Liu, Rui Wang 0001, Zhe Xing, Jun Yu 0002 |
VTC Spring | 4 |
| 2022 | Guest Editorial: Intelligent information processing and services in media convergence
Meng Wang 0001, Chi Zhang 0022, Shijie Hao, Jun Yu 0002, Tingting Mu |
Int. J. Intell. Syst. | 4 |
| 2022 | Semisupervised image classification by mutual learning of multiple self-supervised modelsabstractImage classification has been widely adopted by current social media applications. Compared with fully supervised classification, semisupervised classification attracts more attention because it is commonly observed that category labels are only available for a small portion of images while most images on social media platforms do not have labels. To this end, we propose a two-stage semisupervised learning framework. In the first stage, we train two Self-supervised Models (SSMs). One model is initialized by predicting the rotation angles of pretransformed training images and then further trained by the labeled images. The other model is initialized by making consistent predictions for the transformed images in color, shape, and quality from the same sample image, and then further trained by the labeled images. In the second stage, we fuse the two SSMs through deep mutual learning, which enhances each of the two SSMs with the complementary information provided by the other such that the correct prediction could be shared. Experimental results on CIFAR and Caltech-256 data sets demonstrate the effect of the proposed framework. Jian Zhang 0026, Jun Yu 0002, Jianping Fan 0001 |
Int. J. Intell. Syst. | 3 |
| 2022 | Joint usage of global and local attentions in hourglass network for human pose estimation
Xiena Dong, Jun Yu 0002, Jian Zhang 0026 |
Neurocomputing | 2 |
| 2022 | Modeling long-term video semantic distribution for temporal action proposal generation
Tingting Han 0003, Sicheng Zhao, Xiaoshuai Sun, Jun Yu 0002 |
Neurocomputing | 4 |
| 2022 | Interaction augmented transformer with decoupled decoding for video captioning
Tao Jin 0004, Zhou Zhao 0001, Jun Yu 0002, Fei Wu 0001 |
Neurocomputing | 4 |
| 2022 | A contrastive triplet network for automatic chest X-ray reporting
Jun Yu 0002, Hanliang Jiang, Weidong Han 0001, Jian Zhang 0026 |
Neurocomputing | 2 |
| 2022 | Graph and dynamics interpretation in robotic reinforcement learning task
Zonggui Yao, Jun Yu 0002, Jian Zhang 0026, Wei He 0001 |
Inf. Sci. | 2 |
| 2022 | Weakly supervised moment localization with natural language based on semantic reconstruction
Tingting Han 0003, Kai Wang 0036, Jun Yu 0002, Jianping Fan 0001 |
Image Vis. Comput. | 3 |
| 2022 | Hierarchical Deep Click Feature Prediction for Fine-Grained Image RecognitionabstractThe click feature of an image, defined as the user click frequency vector of the image on a predefined word vocabulary, is known to effectively reduce the semantic gap for fine-grained image recognition. Unfortunately, user click frequency data are usually absent in practice. It remains challenging to predict the click feature from the visual feature, because the user click frequency vector of an image is always noisy and sparse. In this paper, we devise a Hierarchical Deep Word Embedding (HDWE) model by integrating sparse constraints and an improved RELU operator to address click feature prediction from visual features. HDWE is a coarse-to-fine click feature predictor that is learned with the help of an auxiliary image dataset containing click information. It can therefore discover the hierarchy of word semantics. We evaluate HDWE on three dog and one bird image datasets, in which Clickture-Dog and Clickture-Bird are utilized as auxiliary datasets to provide click data, respectively. Our empirical studies show that HDWE has 1) higher recognition accuracy, 2) a larger compression ratio, and 3) good one-shot learning ability and scalability to unseen categories. Jun Yu 0002, Min Tan 0005, Hongyuan Zhang 0001, Yong Rui, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Generalized Multi-View Collaborative Subspace ClusteringabstractIn real-world applications, complete or incomplete multi-view data are common, which leads to the problem of generalized multi-view clustering. Recently, researchers attempt to learn the latent representation in the common subspace from heterogeneous data, which usually suffers from feature degeneration. Moreover, there are limited efforts on simultaneously revealing the underlying subspace structure and exploring the complementary information from incomplete multiple views. In this paper, we introduce a novel Generalized Multi-view Collaborative Subspace Clustering (GMCSC) framework to address the above issues, in which consensus subspace structure of all views and embedding subspaces for each view are jointly learned to benefit each other. Specifically, we develop a novel collaborative subspace learning strategy based on self-representation learning, which provides a brand-new way of pursuing the complete subspace structure directly from multi-view data. Furthermore, we explore complementary information by enforcing the consistency across different views and preserving the view-specific information of each view, which can alleviate the problem of feature degeneration and enhance the reasonability of using a consensus representation for multiple views. Experimental results on six benchmark datasets demonstrate that the proposed method can significantly outperform the state-of-the-art algorithms. Mengcheng Lan, Min Meng 0001, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Exploring Fine-Grained Cluster Structure Knowledge for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation aims to leverage knowledge from a labeled source domain to learn an accurate model in an unlabeled target domain. However, many previous approaches propose to learn domain agnostic feature representations using a global distribution alignment objective, which does not consider the fine-grained cluster structures in the source and target domains. As such, the goal of this paper is to address two challenging problems:1) how to thoroughly explore fine-grained cluster structure knowledge in the source and target domains, 2) how to effectively incorporate these structure knowledge for adaptation.Regarding the first point, we are motivated by structural domain similarity assumption and propose structural representation learning, which is achieved by enforcing structural consistency between the source and target domains while retaining their individual discriminative properties. Regarding the second point, we firstly devise a novel structural centroid-based label prediction method, which explicitly models structural representations to form discriminative source and target cluster centroids, and estimates the label distribution of each target sample through the cosine similarity between its corresponding target cluster centroid and all the other source cluster centroids. Then, we adopt clustering learning to incorporate these discriminative structure knowledge for adaptation by minimizing the KL divergence between the predictive target label distribution and an introduced auxiliary one. Comprehensive experiments and analyses on four benchmark datasets demonstrate the superiority of the proposed discriminative clustering framework. Min Meng 0001, Zhuanghui Wu, Tianyou Liang, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Towards Knowledge-Aware Video Captioning via Transitive Visual Relationship DetectionabstractVideo captioning can be enhanced by incorporating the knowledge, which is usually represented as relationships of objects. However, the previous methods construct only superficial or static object relationships, and often introduce noise into the task through irrelevant common sense or fixed syntax templates. These problems mitigate the model interpretability and lead to the undesirable consequence. To overcome these limitations, we propose to enhance video captioning with deep-level object relationships that are adaptively explored during training. Specifically, we present a Transitive Visual Relationship Detection (TVRD) module in which we estimate the actions of the visual objects, and construct an Object-Action Graph (OAG) to describe the shallow relationship between the objects and actions. Then we bridge the gap between the objects via the actions to transitively infer an Object-Object Graph (OOG) which reflects the deep-level relationship. We further feed the OOG to a graph convolutional network to refine the object representation by deep-level relationships. With the refined representation, we capitalize on an LSTM-based decoder for caption generation. Experimental results on two benchmark datasets: MSVD, MSR-VTT demonstrate that the proposed method achieves state-of-the-art performance. Lastly, we present comprehensive ablation studies as well as visualization of visual relationships to demonstrate the effectiveness and interpretability of our model. Bofeng Wu, Guocheng Niu, Jun Yu 0002, Xinyan Xiao, Jian Zhang 0026, Hua Wu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Local-Global Graph Pooling via Mutual Information Maximization for Video-Paragraph RetrievalabstractAs a task of cross-modal retrieval between long videos and paragraphs, video-paragraph retrieval is a non-trivial task. Unlike traditional video-text retrieval, the video in video-paragraph retrieval usually contains multiple clips. Each clip corresponds to a descriptive sentence; all the sentences constitute the corresponding paragraph of the video. Previous methods for video-paragraph retrieval usually encode videos and para-graphs from segment-level (clips and sentences) and overall-level (videos and paragraphs). However, there are also contents about actions and objects that exist in the segment. Hence, we propose a Local-Global Graph Pooling Network (LGGP) via Mutual Information Maximization for video-paragraph retrieval. Our model disentangles videos and paragraphs into four levels: overall-level, segment-level, motion-level, and object-level. We construct the Hierarchical Local Graph (segment-level, motion-level, and object-level) and the Hierarchical Global Graph (overall-level, segment-level, motion-level, and object-level), respectively, for semantic interaction among different levels. Meanwhile, to obtain hierarchical pooling features with fine-grained semantic information, we design hierarchical graph pooling methods to maximize the mutual information between pooling features and corresponding graph nodes. We evaluate our model on two video-paragraph retrieval datasets with three different video features. The experimental results show that our model establishes state-of-the-art results for video-paragraph retrieval. Our code will be released athttps://github.com/PengchengZhang1997/LGGP. Zhou Zhao 0001, Nannan Wang 0001, Jun Yu 0002, Fei Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Multiview Consensus Structure DiscoveryabstractMultiview subspace learning has attracted much attention due to the efficacy of exploring the information on multiview features. Most existing methods perform data reconstruction on the original feature space and thus are vulnerable to noisy data. In this article, we propose a novel multiview subspace learning method, called multiview consensus structure discovery (MvCSD). Specifically, we learn the low-dimensional subspaces corresponding to different views and simultaneously pursue the structure consensus over subspace clustering for multiple views. In such a way, latent subspaces from different views regularize each other toward a common consensus that reveals the underlying cluster structure. Compared to existing methods, MvCSD leverages the consensus structure derived from the subspaces of diverse views to better exploit the intrinsic complementary information that well reflects the essence of data. Accordingly, the proposed MvCSD is capable of producing a more robust and accurate representation structure which is crucial for multiview subspace learning. The proposed method can be optimized effectively, with theoretical convergence guarantee, by alternatively iterating the argument Lagrangian multiplier algorithm and the eigendecomposition. Extensive experiments on diverse datasets demonstrate the advantages of our method over the state-of-the-art methods. Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu |
IEEE Trans. Cybern. | 3 |
| 2022 | An Individual-Difference-Aware Model for Cross-Person Gaze EstimationabstractWe propose a novel method on refining cross-person gaze prediction task with eye/face images only by explicitly modelling the person-specific differences. Specifically, we first assume that we can obtain some initial gaze prediction results with existing method, which we refer to as InitNet, and then introduce three modules, the Validity Module (VM), Self-Calibration (SC) and Person-specific Transform (PT) module. By predicting the reliability of current eye/face images, VM is able to identify invalid samples, e.g. eye blinking images, and reduce their effects in modelling process. SC and PT module then learn to compensate for the differences on valid samples only. The former models the translation offsets by bridging the gap between initial predictions and dataset-wise distribution. And the later learns more general person-specific transformation by incorporating the information from existing initial predictions of the same person. We validate our ideas on three publicly available datasets, EVE, XGaze, and MPIIGaze dataset. We demonstrate that our proposed method outperforms the SOTA methods significantly on all of them, e.g. respectively 21.7%, 36.0%, and 32.9% relative performance improvements. We are the winner of the GAZE 2021 EVE Challenge and our code can be found here https://github.com/bjj9/EVE_SCPT. Jun Bao, Buyu Liu, Jun Yu 0002 |
IEEE Trans. Image Process. | 3 |
| 2022 | TaoHighlight: Commodity-Aware Multi-Modal Video Highlight Detection in E-CommerceabstractIn e-commerce, product related video is important content to introduce product characteristics and attract consumers. Especially in the recommendation system of e-commerce platform, video highlight detection methods are usually adopted to capture the most attractive clips for showing to consumers, so as to improve the click through rate of products. However, the effect of the current research methods applied to the actual scene is not satisfactory. Compared with other video understanding tasks, video highlight detection is relatively abstract and subjective, and it is difficult to make accurate judgment only by using visual information. Consequently, we put forward multi-modal video highlight detection task, which introduces video related linguistic information as supervised information. And we propose a graph-based commodity-aware model to solve multi-modal video highlight detection in e-commerce scene. Our model consists of multi-modal highlight detection stage and graph-based fine-tuning stage, in which we adopt graph aggregation method to fuse multi-source natural language information and introduce effective visual feature composition method for graph convolution network based highlight detection. Besides, we release the largest e-commerce video highlight detection dataset, TaoHighlight, in which the videos and related data are collected from Taobao e-commerce platform. Our model achieves state-of-art in all separate categories and overall dataset of TaoHighlight, which shows the superiority of our model. Zhaoyu Guo, Zhou Zhao 0001, Weike Jin, Dazhou Wang, Ruitao Liu, Jun Yu 0002 |
IEEE Trans. Multim. | 6 |
| 2022 | Domain-invariant Graph for Adaptive Semi-supervised Domain AdaptationabstractDomain adaptation aims to generalize a model from a source domain to tackle tasks in a related but different target domain. Traditional domain adaptation algorithms assume that enough labeled data, which are treated as the prior knowledge are available in the source domain. However, these algorithms will be infeasible when only a few labeled data exist in the source domain, thus the performance decreases significantly. To address this challenge, we propose a Domain-invariant Graph Learning (DGL) approach for domain adaptation with only a few labeled source samples. Firstly, DGL introduces the Nyström method to construct a plastic graph that shares similar geometric property with the target domain. Then, DGL flexibly employs the Nyström approximation error to measure the divergence between the plastic graph and source graph to formalize the distribution mismatch from the geometric perspective. Through minimizing the approximation error, DGL learns a domain-invariant geometric graph to bridge the source and target domains. Finally, we integrate the learned domain-invariant graph with the semi-supervised learning and further propose an adaptive semi-supervised model to handle the cross-domain problems. The results of extensive experiments on popular datasets verify the superiority of DGL, especially when only a few labeled source samples are available. Weifeng Liu 0001, Yicong Zhou, Jun Yu 0002, Dapeng Tao, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Fine-grained Image Classification via Multi-scale Selective Hierarchical Biquadratic PoolingabstractHow to extract distinctive features greatly challenges the fine-grained image classification tasks. In previous models, bilinear pooling has been frequently adopted to address this problem. However, most bilinear pooling models neglect either intra or inter layer feature interaction. This insufficient interaction brings in the loss of discriminative information. In this article, we devise a novel fine-grained image classification approach named M ulti-scale S elective H ierarchical bi Q uadratic P ooling (MSHQP). The proposed biquadratic pooling simultaneously models intra and inter layer feature interactions and enhances part response by integrating multi-layer features. The subsequent coarse-to-fine multi-scale interaction structure captures the complementary information within features. Finally, the active interaction selection module adaptively learns the optimal interaction subset for a specific dataset. Consequently, we obtain a robust image representation with coarse-to-fine semantics. We conduct experiments on five benchmark datasets. The experimental results demonstrate that MSHQP achieves competitive or even match the state-of-the-art methods in terms of both accuracy and computational efficiency, with 89.0%, 94.9%, 93.4%, 90.4%, and 91.5% top-1 classification accuracy on CUB-200-2011, Stanford-Cars, FGVC-Aircraft, Stanford-Dog, and VegFru, respectively. Min Tan 0005, Fu Yuan, Jun Yu 0002, Guijun Wang, Xiaoling Gu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal HashingabstractUnsupervised cross-modal hashing (UCMH) has become a hot topic recently. Current UCMH focuses on exploring data similarities. However, current UCMH methods calculate the similarity between two data, mainly relying on the two data's cross-modal features. These methods suffer from inaccurate similarity problems that result in a suboptimal retrieval Hamming space, because the cross-modal features between the data are not sufficient to describe the complex data relationships, such as situations where two data have different feature representations but share the inherent concepts. In this paper, we devise a deep graph-neighbor coherence preserving network (DGCPN). Specifically, DGCPN stems from graph models and explores graph-neighbor coherence by consolidating the information between data and their neighbors. DGCPN regulates comprehensive similarity preserving losses by exploiting three types of data similarities (i.e., the graph-neighbor coherence, the coexistent similarity, and the intra- and inter-modality consistency) and designs a half-real and half-binary optimization strategy to reduce the quantization errors during hashing. Essentially, DGCPN addresses the inaccurate similarity problem by exploring and exploiting the data's intrinsic relationships in a graph. We conduct extensive experiments on three public UCMH datasets. The experimental results demonstrate the superiority of DGCPN, e.g., by improving the mean average precision from 0.722 to 0.751 on MIRFlickr-25K using 64-bit hashing codes to retrieval texts from images. We will release the source code package and the trained model on https://github.com/Atmegal/DGCPN. Jun Yu 0002, Yibing Zhan, Dacheng Tao |
AAAI | 1 |
| 2021 | Learning Controlled Semantic Embedding for Cross-Modal RetrievalabstractCross-modal retrieval has caught appealing attentions as it supports querying across different modalities. However, most existing methods have emphasized on directly mapping heterogeneous features into the common subspace, which inevitably results in highly entangled representations, thereby preventing them from bridging the modality gap. This paper presents a novel deep framework called Controlled Semantic Embedding (CSE), which is the first attempt to learn disentangled representations with controlled semantic structure for cross-modal retrieval. Specifically, we design two generative networks based on variational autoencoder, which incorporate semantic discriminators for effective prediction of structured semantics. Meanwhile, a self-supervised semantic network is seamlessly integrated into the generative networks to supervise the semantic embedding process, which is further coupled with a quantizer for controlling the quantizability of semantic representations. Extensive experiments show the superiority of CSE over other state-of-the-art methods in cross-modal retrieval. Min Meng 0001, Jun Yu 0002, Jigang Wu |
ICME | 3 |
| 2021 | Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal MatchingabstractThis paper proposes an approach to Dense Video Captioning (DVC) without pairwise event-sentence annotation. First, we adopt the knowledge distilled from relevant and well solved tasks to generate high-quality event proposals. Then we incorporate contrastive loss and cycle-consistency loss typically applied to cross-modal retrieval tasks to build semantic matching between the proposals and sentences, which are eventually used to train the caption generation module. In addition, the parameters of matching module are initialized via pre-training based on annotated images to improve the matching performance. Extensive experiments on ActivityNet-Caption dataset reveal the significance of distillation-based event proposal generation and cross-modal retrieval-based semantic matching to weakly supervised DVC, and demonstrate the superiority of our method to existing state-of-the-art methods. Bofeng Wu, Guocheng Niu, Jun Yu 0002, Xinyan Xiao, Jian Zhang 0026, Hua Wu 0003 |
IJCAI | 3 |
| 2021 | ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge IntegrationabstractVision-and-language pretraining (VLP) aims to learn generic multimodal representations from massive image-text pairs. While various successful attempts have been proposed, learning fine-grained semantic alignments between image-text pairs plays a key role in their approaches. Nevertheless, most existing VLP approaches have not fully utilized the intrinsic knowledge within the image-text pairs, which limits the effectiveness of the learned alignments and further restricts the performance of their models. To this end, we introduce a new VLP method called ROSITA, which integrates the cross- and intra-modal knowledge in a unified scene graph to enhance the semantic alignments. Specifically, we introduce a novel structural knowledge masking (SKM) strategy to use the scene graph structure as a priori to perform masked language (region) modeling, which enhances the semantic alignments by eliminating the interference information within and across modalities. Extensive ablation studies and comprehensive analysis verifies the effectiveness of ROSITA in semantic alignments. Pretrained with both in-domain and out-of-domain datasets, ROSITA significantly outperforms existing state-of-the-art VLP methods on three typical vision-and-language tasks over six benchmark datasets. Yuhao Cui, Zhou Yu 0001, Chunqi Wang, Zhongzhou Zhao, Ji Zhang 0011, Meng Wang 0001, Jun Yu 0002 |
ACM Multimedia | 7 |
| 2021 | Effective De-identification Generative Adversarial Network for Face AnonymizationabstractThe growing application of face images and modern AI technology has raised another important concern in privacy protection. In many real scenarios like scientific research, social sharing and commercial application, lots of images are released without privacy processing to protect people's identity. In this paper, we develop a novel effective de-identification generative adversarial network (DeIdGAN) for face anonymization by seamlessly replacing a given face image with a different synthesized yet realistic one. Our approach consists of two steps. First, we anonymize the input face to obfuscate its original identity. Then, we use our designed de-identification generator to synthesize an anonymized face. During the training process, we leverage a pair of identity-adversarial discriminators to explicitly constrain identity protection by pushing the synthesized face away from the predefined sensitive faces to resist re-identification and identity invasion. Finally, we validate the effectiveness of our approach on public datasets. Compared with existing methods, our approach can not only achieve better identity protection rates but also preserve superior image quality and data reusability, which suggests the state-of-the-art performance. Zhenzhong Kuang, Huigui Liu, Jun Yu 0002, Aikui Tian, Jianping Fan 0001, Noboru Babaguchi |
ACM Multimedia | 3 |
| 2021 | Federated Learning Model Training Method Based on Data Features Perception AggregationabstractThe rapidly expanding number of Internet of Things (IoT) devices is generating huge quantities of data, but public concern over data privacy means users are apprehensive to send data to a central server for machine learning purposes. Federated learning is an emerging concept, which allows edge devices to collaboratively learn and share models, while keeping training data on devices. Federated learning decouples “model training” and “direct access to original training data”. But, in the IoT where the wireless network resource is constrained, the key problem of federated learning is the communication overhead for parameter synchronization, which wastes bandwidth, increases training time, and even impacts the model accuracy. Moreover, the IoT devices collect data from different users, so the distribution over devices can be highly non-independent identically distributed (non-IID), which results in the variation of feature distribution and label distribution. As a result, the test accuracy of the federated model is reduced, and the communication cost of training the federated model is increased. In this paper, we propose the FedCC algorithm, to improve the accuracy of the federated model in the non-IID scenario. The FedCC method constructs client groups by mining data similarity, and selects one model of every client group to upload to the cloud server for model aggregation. Our experiments show that FedCC not only outperforms popular state-of-the-art federated learning algorithms on CNN and MLP architectures trained on MNIST and CIFAR-10 datasets, but also reduces the overall communication cost. Zeng Yan, Yan Zhong Yi, Nailiang Zhao, Jian Wan 0001, Jun Yu 0002 |
VTC Fall | 7 |
| 2021 | Contrastive learning of graph encoder for accelerating pedestrian trajectory prediction trainingabstractAbstract In the area of pedestrian trajectory prediction, the hybrid structures of temporal feature extractor or spatial feature extractor have paved the way for the precise prediction model, and they are in larger and larger scale. Learning of specific feature encoding model not only influenced by the structure of the network, but also by the learning manners such as supervised learning and unsupervised learning. Previous works concentrated on more comprehensive encoders and more delicate designs of feature extractors. However, the mutual influence factors from the neighbour pedestrians associate with the distance to the centre pedestrian seldomly noticed. Most of the existed feature extractors in prediction models trained in the way of supervised learning other than unsupervised manners caused the problem that the extracted features are always handcrafted without the natural distinction of obscure situations. The graph contrastive accelerating encoder is proposed, which accelerates the pedestrian trajectory prediction training process of the state of the art method of spatio‐temporal graph transformer networks. Employing the unsupervised contrastive learning process and the graph of neighbours representing distance affection of nearest and farthest pedestrian to the centre pedestrian, the graph contrastive accelerating encoder significantly shrinked the training time. Holding the final performance on to state of the art level, the proposed method let the lowest pedestrian trajectory prediction error show up in the obviously earlier training steps. Zonggui Yao, Jun Yu 0002, Jiajun Ding |
IET Image Process. | 2 |
| 2021 | Unnoticeable synthetic face replacement for image privacy protection
Zhenzhong Kuang, Zhiqiang Guo, Jinglong Fang, Jun Yu 0002, Noboru Babaguchi, Jianping Fan 0001 |
Neurocomputing | 4 |
| 2021 | Deep embedding of concept ontology for hierarchical fashion recognition
Zhenzhong Kuang, Xin Zhang 0063, Jun Yu 0002, Zongmin Li, Jianping Fan 0001 |
Neurocomputing | 3 |
| 2021 | Distributed feedback network for single-image deraining
Jiajun Ding, Huanlei Guo, Jun Yu 0002, Xiongxiong He, Bo Jiang 0016 |
Inf. Sci. | 4 |
| 2021 | Coupled Knowledge Transfer for Visual Data RecognitionabstractTransfer learning aims to learn an effective classifier for unlabeled target data by borrowing knowledge from well-labeled source data. However, most existing work has emphasized on learning domain invariant features to reduce the distribution discrepancy, which may suffer from the negative transfer problem caused by structure inconsistencies or distribution outliers. To address this challenge, in this paper, we propose a novel transfer learning approach, which seamlessly integrates domain invariant feature learning, discriminative structure preservation and sample reweighting into a unified learning model. Specifically, we attempt to learn domain invariant features by jointly adapting the marginal and conditional distributions. To transfer discriminative knowledge inferred from data, we enforce the structure consistency between the original feature space and the latent feature space. Furthermore, to enhance the robustness of our model, an efficient and more generalized sample reweighting strategy is developed to assign target predictions with different levels of confidence. The key advantage over previous methods is that our model can adaptively select pivot samples in target domain and retain the properties of discriminative structures underlying data domains, which enables coupled knowledge transfer during the learning process. Experimental results on several benchmark datasets have verified the superiority of the proposed method over other state-of-the-art algorithms. Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Long-Term Video Question Answering via Multimodal Hierarchical Memory Attentive NetworksabstractLong-term Video Question Answering plays an essential role in visual information retrieval, which aims at generating natural language answers to discretionary free-form questions about the referenced long-term video. Rather than remember the video as a sequence of visual content, humans have an innate cognitive ability to identify the critical moments related to the question at first glance, then tie together the specific evidence around these critical moments for further analysis and reasoning. Motivated by this intuition, we propose the multimodal hierarchical memory attentive networks with two heterogeneous memory subnetworks: the top guided memory network and the bottom enhanced multimodal memory attentive network. The top guided memory network serves as a shallow inference engine to pick relevant and informative moments of questions and obtain salient video content at a coarse-grained level. Subsequently, the bottom enhanced multimodal memory attentive network is designed as an in-depth reasoning engine to perform more accurate attention with cues from video bottom evidence in a fine-grained level to enhance question answering quality. We evaluate the proposed method on three publicly available video question answering benchmarks, namely ActivityNet-QA, MSRVTT-QA, and MSVD-QA. Experimental results demonstrate that the proposed approach significantly outperforms other state-of-the-art methods for long-term videos. Extensive ablation studies are carried out to explore the reasons behind the proposed model's effectiveness. Ting Yu 0016, Jun Yu 0002, Zhou Yu 0001, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Toward Realistic Face Photo-Sketch Synthesis via Composition-Aided GANsabstractFace photo-sketch synthesis aims at generating a facial sketch/photo conditioned on a given photo/sketch. It covers wide applications including digital entertainment and law enforcement. Precisely depicting face photos/sketches remains challenging due to the restrictions on structural realism and textural consistency. While existing methods achieve compelling results, they mostly yield blurred effects and great deformation over various facial components, leading to the unrealistic feeling of synthesized images. To tackle this challenge, in this article, we propose using facial composition information to help the synthesis of face sketch/photo. Especially, we propose a novel composition-aided generative adversarial network (CA-GAN) for face photo-sketch synthesis. In CA-GAN, we utilize paired inputs, including a face photo/sketch and the corresponding pixelwise face labels for generating a sketch/photo. Next, to focus training on hard-generated components and delicate facial structures, we propose a compositional reconstruction loss. In addition, we employ a perceptual loss function to encourage the synthesized image and real image to be perceptually similar. Finally, we use stacked CA-GANs (SCA-GANs) to further rectify defects and add compelling details. The experimental results show that our method is capable of generating both visually comfortable and identity-preserving face sketches/photos over a wide range of challenging data. In addition, our method significantly decreases the best previous Fréchet inception distance (FID) from 36.2 to 26.2 for sketch synthesis, and from 60.9 to 30.5 for photo synthesis. Besides, we demonstrate that the proposed method is of considerable generalization ability. Jun Yu 0002, Xingxin Xu, Fei Gao 0006, Shengjie Shi, Meng Wang 0001, Dacheng Tao, Qingming Huang |
IEEE Trans. Cybern. | 1 |
| 2021 | SPRNet: Single-Pixel Reconstruction for One-Stage Instance SegmentationabstractObject instance segmentation is one of the most fundamental but challenging tasks in computer vision, and it requires the pixel-level image understanding. Most existing approaches address this problem by adding a mask prediction branch to a two-stage object detector with the region proposal network (RPN). Although producing good segmentation results, the efficiency of these two-stage approaches is far from satisfactory, restricting their applicability in practice. In this article, we propose a one-stage framework, single-pixel reconstruction net (SPRNet), which performs efficient instance segmentation by introducing a single-pixel reconstruction (SPR) branch to off-the-shelf one-stage detectors. The added SPR branch reconstructs the pixel-level mask from every single pixel in the convolution feature map directly. Using the same ResNet-50 backbone, SPRNet achieves comparable mask AP with Mask R-CNN at a higher inference speed and gains all-round improvements on box AP at every scale compared with RetinaNet. Jun Yu 0002, Jinghan Yao, Jian Zhang 0026, Zhou Yu 0001, Dacheng Tao |
IEEE Trans. Cybern. | 1 |
| 2021 | Complementary, Heterogeneous and Adversarial Networks for Image-to-Image TranslationabstractImage-to-image translation is to transfer images from a source domain to a target domain. Conditional Generative Adversarial Networks (GANs) have enabled a variety of applications. Initial GANs typically conclude one single generator for generating a target image. Recently, using multiple generators has shown promising results in various tasks. However, generators in these works are typically of homogeneous architectures. In this paper, we argue that heterogeneous generators are complementary to each other and will benefit the generation of images. By heterogeneous, we mean that generators are of different architectures, focus on diverse positions, and perform over multiple scales. To this end, we build two generators by using a deep U-Net and a shallow residual network, respectively. The former concludes a series of down-sampling and up-sampling layers, which typically have large perception field and great spatial locality. In contrast, the residual network has small perceptual fields and works well in characterizing details, especially textures and local patterns. Afterwards, we use a gated fusion network to combine these two generators for producing a final output. The gated fusion unit automatically induces heterogeneous generators to focus on different positions and complement each other. Finally, we propose a novel approach to integrate multi-level and multi-scale features in the discriminator. This multi-layer integration discriminator encourages generators to produce realistic details from coarse to fine scales. We quantitatively and qualitatively evaluate our model on various benchmark datasets. Experimental results demonstrate that our method significantly improves the quality of transferred images, across a variety of image-to-image translation tasks. We have made our code and results publicly available: http://aiart.live/chan/. Fei Gao 0006, Xingxin Xu, Jun Yu 0002, Meimei Shang, Xiang Li 0205, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2021 | Asymmetric Supervised Consistent and Specific Hashing for Cross-Modal RetrievalabstractHashing-based techniques have provided attractive solutions to cross-modal similarity search when addressing vast quantities of multimedia data. However, existing cross-modal hashing (CMH) methods face two critical limitations: 1) there is no previous work that simultaneously exploits the consistent or modality-specific information of multi-modal data; 2) the discriminative capabilities of pairwise similarity is usually neglected due to the computational cost and storage overhead. Moreover, to tackle the discrete constraints, relaxation-based strategy is typically adopted to relax the discrete problem to the continuous one, which severely suffers from large quantization errors and leads to sub-optimal solutions. To overcome the above limitations, in this article, we present a novel supervised CMH method, namely Asymmetric Supervised Consistent and Specific Hashing (ASCSH). Specifically, we explicitly decompose the mapping matrices into the consistent and modality-specific ones to sufficiently exploit the intrinsic correlation between different modalities. Meanwhile, a novel discrete asymmetric framework is proposed to fully explore the supervised information, in which the pairwise similarity and semantic labels are jointly formulated to guide the hash code learning process. Unlike existing asymmetric methods, the discrete asymmetric structure developed is capable of solving the binary constraint problem discretely and efficiently without any relaxation. To validate the effectiveness of the proposed approach, extensive experiments on three widely used datasets are conducted and encouraging results demonstrate the superiority of ASCSH over other state-of-the-art CMH methods. Min Meng 0001, Haitao Wang 0026, Jun Yu 0002, Jigang Wu |
IEEE Trans. Image Process. | 3 |
| 2021 | Toward Multi-Modal Conditioned Fashion Image TranslationabstractHaving the capability to synthesize photo-realistic fashion product images conditioned on multiple attributes or modalities would bring many new exciting applications. In this work, we propose an end-to-end network architecture that built upon a new generative adversarial network for automatically synthesizing photo-realistic images of fashion products under multiple conditions. Given an input pose image that consists of a 2D skeleton pose and a sentence description of products, our model synthesizes a fashion image preserving the same pose and wearing the fashion products described as the text. Specifically, the generator$G$tries to generate realistic-looking fashion images based on a$\langle \mathsf {pose}, \mathsf {text} \rangle$pair condition to fool the discriminator. An attention network is added for enhancing the generator, which predicts a probability map indicating which part of the image needs to be attended for translation. In contrast, the discriminator$D$distinguishes real images from the translated ones based on the input pose image and text description. The discriminator is divided into two multi-scale sub-discriminators for improving image distinguishing task. Quantitative and qualitative analysis demonstrates that our method is capable of synthesizing realistic images that retain the poses of given images while matching the semantics of provided sentence descriptions. Xiaoling Gu, Jun Yu 0002, Yongkang Wong, Mohan Kankanhalli |
IEEE Trans. Multim. | 2 |
| 2020 | Diversified Bayesian Nonnegative Matrix Factorization
Maoying Qiao, Jun Yu 0002, Tongliang Liu, Xinchao Wang, Dacheng Tao |
AAAI | 2 |
| 2020 | Deep Multimodal Neural Architecture SearchabstractDesigning effective neural networks is fundamentally important in deep multimodal learning. Most existing works focus on a single task and design neural architectures manually, which are highly task-specific and hard to generalize to different tasks. In this paper, we devise a generalized deep multimodal neural architecture search (MMnas) framework for various multimodal learning tasks. Given multimodal input, we first define a set of primitive operations, and then construct a deep encoder-decoder based unified backbone, where each encoder or decoder block corresponds to an operation searched from a predefined operation pool. On top of the unified backbone, we attach task-specific heads to tackle different multimodal learning tasks. By using a gradient-based NAS algorithm, the optimal architectures for different tasks are learned efficiently. Extensive ablation studies, comprehensive analysis, and comparative experimental results show that the obtained MMnasNet significantly outperforms existing state-of-the-art approaches across three multimodal learning tasks (over five datasets), including visual question answering, image-text matching, and visual grounding. Zhou Yu 0001, Yuhao Cui, Jun Yu 0002, Meng Wang 0001, Dacheng Tao, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2020 | Relationship graph learning network for visual relationship detectionabstractVisual relationship detection aims to predict the relationships between detected object pairs. It is well believed that the correlations between image components (i.e., objects and relationships between objects) are significant considerations when predicting objects' relationships. However, most current visual relationship detection methods only exploited the correlations among objects, and the correlations among objects' relationships remained underexplored. This paper proposes a relationship graph learning network (RGLN) to explore the correlations among objects' relationships for visual relationship detection. Specifically, RGLN obtains image objects using an object detector, and then, every pair of objects constitutes a relationship proposal. All relationship proposals construct a relationship graph, in which the proposals are treated as nodes. Accordingly, RGLN designs bi-stream graph attention subnetworks to detect relationship proposals, in which one graph attention subnetwork analyzes correlations among relationships based on visual and spatial information, and the other analyzes correlations based on semantic and spatial information. Besides, RGLN exploits a relationship selection subnetwork to ignore redundant information of object pairs with no relationships. We conduct extensive experiments on two public datasets: the VRD and the VG datasets. The experimental results compared with the state-of-the-art demonstrate the competitiveness of RGLN. Jun Yu 0002, Yibing Zhan, Zhi Chen 0010 |
MMAsia | 2 |
| 2020 | Discriminative Regions Erasing Strategy for Weakly-Supervised Temporal Action Localization
Huanbin Zeng, Suguo Zhu, Jun Yu 0002 |
PRCV (2) | 3 |
| 2020 | Representation learning of image composition for aesthetic prediction
Meimei Shang, Fei Gao 0006, Rongsheng Li, Jun Yu 0002 |
Comput. Vis. Image Underst. | 6 |
| 2020 | Multi-task Compositional Network for Visual Relationship Detection
Yibing Zhan, Jun Yu 0002, Ting Yu 0016, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2020 | Style-adaptive photo aesthetic rating via convolutional neural networks and multi-task learning
Fei Gao 0006, Ziyun Li 0002, Jun Yu 0002, Junze Yu, Qingming Huang, Qi Tian 0001 |
Neurocomputing | 3 |
| 2020 | Fine-grained visual understanding and reasoning
Jun Yu 0002, Yezhou Yang, Fionn Murtagh, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2020 | Incremental focal loss GANs
Fei Gao 0006, Jingjie Zhu, Hanliang Jiang, Zhenxing Niu, Weidong Han 0001, Jun Yu 0002 |
Inf. Process. Manag. | 6 |
| 2020 | Fine-grained image classification with factorized deep user click feature
Min Tan 0005, Zhiyou Peng, Jun Yu 0002, Fang Tang |
Inf. Process. Manag. | 4 |
| 2020 | Intra- and Inter-modal Multilinear Pooling with Multitask Learning for Video Grounding
Zhou Yu 0001, Yijun Song, Jun Yu 0002, Meng Wang 0001, Qingming Huang |
Neural Process. Lett. | 3 |
| 2020 | TVENet: Temporal variance embedding network for fine-grained action representation
Tingting Han 0003, Hongxun Yao, Wenlong Xie, Xiaoshuai Sun, Sicheng Zhao, Jun Yu 0002 |
Pattern Recognit. | 6 |
| 2020 | Guest Editorial Introduction to the Special Section on Representation Learning for Visual Content UnderstandingabstractRepresentation learning methods allow a system to automatically learn robust and discriminative features from raw data for given goals, which play an important role in various visual content understanding applications, such as visual object segmentation, detection, tracking, recognition, and search. The performance of visual content understanding tasks is heavily dependent on the choice of data representation (or features) on which they are applied. Conventional feature representation methods usually employ transformations of data that make it easier to extract useful information, such as scale-invariant feature transform (SIFT), local binary patterns (LBP), and histogram of oriented gradients (HOG). In recent years, deep learning techniques have been widely applied to learn data-driven representations with supervised annotations and achieved great success in different visual content understanding tasks. Representative methods include the ResNet method for image classification, the DeepFace method for face recognition, and the feature pyramid networks (FPNs) method for object detection. Despite recent progresses on deep representation learning with a great amount of annotated data, how to effectively learn visual representation with limited data annotations still requires many efforts. This special section focuses on data-effective representation learning methods for visual content understanding. Jiwen Lu, Yuxin Peng 0001, Guo-Jun Qi, Jun Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Multimodal Transformer With Multi-View Visual Representation for Image CaptioningabstractImage captioning aims to automatically generate a natural language description of a given image, and most state-of-the-art models have adopted an encoder-decoder framework. The framework consists of a convolution neural network (CNN)-based image encoder that extracts region-based visual features from the input image, and an recurrent neural network (RNN) based caption decoder that generates the output caption words based on the visual features with the attention mechanism. Despite the success of existing studies, current methods only model the co-attention that characterizes the inter-modal interactions while neglecting the self-attention that characterizes the intra-modal interactions. Inspired by the success of the Transformer model in machine translation, here we extend it to a Multimodal Transformer (MT) model for image captioning. Compared to existing image captioning approaches, the MT model simultaneously captures intra- and inter-modal interactions in a unified attention block. Due to the in-depth modular composition of such attention blocks, the MT model can perform complex multimodal reasoning and output accurate captions. Moreover, to further improve the image captioning performance, multi-view visual features are seamlessly introduced into the MT model. We quantitatively and qualitatively evaluate our approach using the benchmark MSCOCO image captioning dataset and conduct extensive ablation studies to investigate the reasons behind its effectiveness. The experimental results show that our method significantly outperforms the previous state-of-the-art methods. With an ensemble of seven models, our solution ranks the 1st place on the real-time leaderboard of the MSCOCO image captioning challenge at the time of the writing of this paper. Jun Yu 0002, Jing Li 0099, Zhou Yu 0001, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Constrained Discriminative Projection Learning for Image ClassificationabstractProjection learning is widely used in extracting discriminative features for classification. Although numerous methods have already been proposed for this goal, they barely explore the label information during projection learning and fail to obtain satisfactory performance. Besides, many existing methods can learn only a limited number of projections for feature extraction which may degrade the performance in recognition. To address these problems, we propose a novel constrained discriminative projection learning (CDPL) method for image classification. Specifically, CDPL can be formulated as a joint optimization problem over subspace learning and classification. The proposed method incorporates the low-rank constraint to learn a robust subspace which can be used as a bridge to seamlessly connect the original visual features and objective outputs. A regression function is adopted to explicitly exploit the class label information so as to enhance the discriminability of subspace. Unlike existing methods, we use two matrices to perform feature learning and regression, respectively, such that the proposed approach can obtain more projections and achieve superior performance in classification tasks. The experiments on several datasets show clearly the advantages of our method against other state-of-the-art methods. Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu, Dapeng Tao |
IEEE Trans. Image Process. | 3 |
| 2020 | Compositional Attention Networks With Two-Stream Fusion for Video Question AnsweringabstractGiven a video, Video Question Answering (VideoQA) aims at answering arbitrary free-form questions about the video content in natural language. A successful VideoQA framework usually has the following two key components: 1) a discriminative video encoder that learns the effective video representation to maintain as much information as possible about the video and 2) a question-guided decoder that learns to select the most related features to perform spatiotemporal reasoning, as well as outputs the correct answer. We propose compositional attention networks (CAN) with two-stream fusion for VideoQA tasks. For the encoder, we sample video snippets using a two-stream mechanism (i.e., a uniform sampling stream and an action pooling stream) and extract a sequence of visual features for each stream to represent the video semantics with implementation. For the decoder, we propose a compositional attention module to integrate the two-stream features with the attention mechanism. The compositional attention module is the core of CAN and can be seen as a modular combination of a unified attention block. With different fusion strategies, we devise five compositional attention module variants. We evaluate our approach on one long-term VideoQA dataset, ActivityNet-QA, and two short-term VideoQA datasets, MSRVTT-QA and MSVD-QA. Our CAN model achieves new state-of-the-art results on all the datasets. Ting Yu 0016, Jun Yu 0002, Zhou Yu 0001, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2020 | Spatial Pyramid-Enhanced NetVLAD With Weighted Triplet Loss for Place RecognitionabstractWe propose an end-to-end place recognition model based on a novel deep neural network. First, we propose to exploit the spatial pyramid structure of the images to enhance the vector of locally aggregated descriptors (VLAD) such that the enhanced VLAD features can reflect the structural information of the images. To encode this feature extraction into the deep learning method, we build a spatial pyramid-enhanced VLAD (SPE-VLAD) layer. Next, we impose weight constraints on the terms of the traditional triplet loss (T-loss) function such that the weighted T-loss (WT-loss) function avoids the suboptimal convergence of the learning process. The loss function can work well under weakly supervised scenarios in that it determines the semantically positive and negative samples of each query through not only the GPS tags but also the Euclidean distance between the image representations. The SPE-VLAD layer and the WT-loss layer are integrated with the VGG-16 network or ResNet-18 network to form a novel end-to-end deep neural network that can be easily trained via the standard backpropagation method. We conduct experiments on three benchmark data sets, and the results demonstrate that the proposed model defeats the state-of-the-art deep learning approaches applied to place recognition. Jun Yu 0002, Jian Zhang 0026, Qingming Huang, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Proposal Complementary Action DetectionabstractTemporal action detection not only requires correct classification but also needs to detect the start and end times of each action accurately. However, traditional approaches always employ sliding windows or actionness to predict the actions, and it is different to train to model with sliding windows or actionness by end-to-end means. In this article, we attempt a different idea to detect the actions end-to-end, which can calculate the probabilities of actions directly through one network as one part of the results. We present PCAD, a novel proposal complementary action detector to deal with video streams under continuous, untrimmed conditions. Our approach first uses a simple fully 3D convolutional network to encode the video streams and then generates candidate temporal proposals for activities by using anchor segments. To generate more precise proposals, we also design a boundary proposal network to offer some complementary information for the candidate proposals. Finally, we learn an efficient classifier to classify the generated proposals into different activities and refine their temporal boundaries at the same time. Our model can achieve end-to-end training by jointly optimizing classification loss and regression loss. When evaluating on the THUMOS’14 detection benchmark, PCAD achieves state-of-the-art performance in high-speed models. Suguo Zhu, Xiaoxian Yang, Jun Yu 0002, Zhenying Fang, Meng Wang 0001, Qingming Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question AnsweringabstractRecent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answering (VideoQA). Compared to the image domain where large scale and fully annotated benchmark datasets exists, VideoQA datasets are limited to small scale and are automatically generated, etc. These limitations restrict their applicability in practice. Here we introduce ActivityNet-QA, a fully annotated and large scale VideoQA dataset. The dataset consists of 58,000 QA pairs on 5,800 complex web videos derived from the popular ActivityNet dataset. We present a statistical analysis of our ActivityNet-QA dataset and conduct extensive experiments on it by comparing existing VideoQA baselines. Moreover, we explore various video representation strategies to improve VideoQA performance, especially for long videos. Zhou Yu 0001, Dejing Xu, Jun Yu 0002, Ting Yu 0016, Zhou Zhao 0001, Yueting Zhuang, Dacheng Tao |
AAAI | 3 |
| 2019 | Embedding Complementary Deep Networks for Image ClassificationabstractIn this paper, a deep embedding algorithm is developed to achieve higher accuracy rates on large-scale image classification. By adapting the importance of the object classes to their error rates, our deep embedding algorithm can train multiple complementary deep networks sequentially, where each of them focuses on achieving higher accuracy rates for different subsets of object classes in an easy-to-hard way. By integrating such complementary deep networks to generate an ensemble network, our deep embedding algorithm can improve the accuracy rates for the hard object classes (which initially have higher error rates) at certain degrees while effectively preserving high accuracy rates for the easy object classes. Our deep embedding algorithm has achieved higher overall accuracy rates on large scale image classification. Qiuyu Chen, Wei Zhang 0016, Jun Yu 0002, Jianping Fan 0001 |
CVPR | 3 |
| 2019 | Deep Modular Co-Attention Networks for Visual Question AnsweringabstractVisual Question Answering (VQA) requires a fine-grained and simultaneous understanding of both the visual content of images and the textual content of questions. Therefore, designing an effective `co-attention' model to associate key words in questions with key objects in images is central to VQA performance. So far, most successful attempts at co-attention learning have been achieved by using shallow models, and deep co-attention models show little improvement over their shallow counterparts. In this paper, we propose a deep Modular Co-Attention Network (MCAN) that consists of Modular Co-Attention (MCA) layers cascaded in depth. Each MCA layer models the self-attention of questions and images, as well as the question-guided-attention of images jointly using a modular composition of two basic attention units. We quantitatively and qualitatively evaluate MCAN on the benchmark VQA-v2 dataset and conduct extensive ablation studies to explore the reasons behind MCAN's effectiveness. Experimental results demonstrate that MCAN significantly outperforms the previous state-of-the-art. Our best single model delivers 70.63% overall accuracy on the test-dev set. Zhou Yu 0001, Jun Yu 0002, Yuhao Cui, Dacheng Tao, Qi Tian 0001 |
CVPR | 2 |
| 2019 | On Exploring Undetermined Relationships for Visual Relationship DetectionabstractIn visual relationship detection, human-notated relationships can be regarded as determinate relationships. However, there are still large amount of unlabeled data, such as object pairs with less significant relationships or even with no relationships. We refer to these unlabeled but potentially useful data as undetermined relationships. Although a vast body of literature exists, few methods exploit these undetermined relationships for visual relationship detection. In this paper, we explore the beneficial effect of undetermined relationships on visual relationship detection. We propose a novel multi-modal feature based undetermined relationship learning network (MF-URLN) and achieve great improvements in relationship detection. In detail, our MF-URLN automatically generates undetermined relationships by comparing object pairs with human-notated data according to a designed criterion. Then, the MF-URLN extracts and fuses features of object pairs from three complementary modals: visual, spatial, and linguistic modals. Further, the MF-URLN proposes two correlated subnetworks: one subnetwork decides the determinate confidence, and the other predicts the relationships. We evaluate the MF-URLN on two datasets: the Visual Relationship Detection (VRD) and the Visual Genome (VG) datasets. The experimental results compared with state-of-the-art methods verify the significant improvements made by the undetermined relationships, e.g., the top-50 relation detection recall improves from 19.5% to 23.9% on the VRD dataset. Yibing Zhan, Jun Yu 0002, Ting Yu 0016, Dacheng Tao |
CVPR | 2 |
| 2019 | PCPCAD: Proposal Complementary Action DetectorabstractTemporal action detection is still a challenging task. This task not only requires correct classification, but also needs to accurately detect the start and end times of each action. In this paper, we present a novel proposal complementary action detector (PCAD) to deal with video streams under continuous, untrimmed conditions. Our approach first uses a simple fully 3D convolutional (Conv3D) network to encode the video streams and then generates candidate temporal proposals for activities by using anchor segments. To generate more precise proposals, we also designed a boundary proposal network (BPN) to offer some complementary information for the candidate proposals. Finally, we learn an efficient classifier to classify the generated proposals into different activities and refine their temporal boundaries at the same time. Our model can achieve end-to-end training by jointly optimizing classification loss and regression loss. When evaluating on THUMOS'14 detection benchmark, PCAD achieves the state-of-the-art performance in high-speed models. Zhenying Fang, Suguo Zhu, Jun Yu 0002, Qi Tian 0001 |
ICME | 3 |
| 2019 | Multi-interaction Network with Object Relation for Video Question AnsweringabstractVideo question answering is an important task for testing machine's ability of video understanding. The existing methods normally focus on the combination of recurrent and convolutional neural networks to capture spatial and temporal information of the video. Recently, some work has also shown that using attention mechanism can achieve better performance. In this paper, we propose a new model called Multi-interaction network for video question answering. There are two types of interactions in our model. The first type is the multi-modal interaction between the visual and textual information. The second type is the multi-level interaction inside the multi-modal interaction. Specifically, instead of using original self-attention, we propose a new attention mechanism called multi-interaction, which can capture both element-wise and segment-wise sequence interactions, simultaneously. And in addition to the normal frame-level interaction, we also take the object relations into consideration, in order to obtain more fine-grained information, such as motions and other potential relations among these objects. We evaluate our method on TGIF-QA and other two video QA datasets. The qualitative and quantitative experimental results show the effectiveness of our model, which achieves the new state-of-the-art performance. Weike Jin, Zhou Zhao 0001, Mao Gu, Jun Yu 0002, Jun Xiao 0001, Yueting Zhuang |
ACM Multimedia | 4 |
| 2019 | Video Dialog via Multi-Grained Convolutional Self-Attention Context NetworksabstractVideo dialog is a new and challenging task, which requires an AI agent to maintain a meaningful dialog with humans in natural language about video contents. Specifically, given a video, a dialog history and a new question about the video, the agent has to combine video information with dialog history to infer the answer. And due to the complexity of video information, the methods of image dialog might be ineffectively applied directly to video dialog. In this paper, we propose a novel approach for video dialog called multi-grained convolutional self-attention context network, which combines video information with dialog history. Instead of using RNN to encode the sequence information, we design a multi-grained convolutional self-attention mechanism to capture both element and segment level interactions which contain multi-grained sequence information. Then, we design a hierarchical dialog history encoder to learn the context-aware question representation and a two-stream video encoder to learn the context-aware video representation. We evaluate our method on two large-scale datasets. Due to the flexibility and parallelism of the new attention mechanism, our method can achieve higher time efficiency, and the extensive experiments also show the effectiveness of our method. Weike Jin, Zhou Zhao 0001, Mao Gu, Jun Yu 0002, Jun Xiao 0001, Yueting Zhuang |
SIGIR | 4 |
| 2019 | End-to-end visual grounding via region proposal networks and bilinear poolingabstractPhrase‐based visual grounding aims to localise the object in the image referred by a textual query phrase. Most existing approaches adopt a two‐stage mechanism to address this problem: first, an off‐the‐shelf proposal generation model is adopted to extract region‐based visual features, and then a deep model is designed to score the proposals based on the query phrase and extracted visual features. In contrast to that, the authors design an end‐to‐end approach to tackle the visual grounding problem in this study. They use a region proposal network to generate object proposals and the corresponding visual features simultaneously, and multi‐modal factorised bilinear pooling model to fuse the multi‐modal features effectively. After that, two novel losses are posed on top of the multi‐modal features to rank and refine the proposals, respectively. To verify the effectiveness of the proposed approach, the authors conduct experiments on three real‐world visual grounding datasets, namely Flickr‐30k Entities, ReferItGame and RefCOCO. The experimental results demonstrate the significant superiority of the proposed method over the existing state‐of‐the‐arts. Chenchao Xiang, Zhou Yu 0001, Suguo Zhu, Jun Yu 0002, Xiaokang Yang 0001 |
IET Comput. Vis. | 4 |
| 2019 | Multimodal activity recognition with local block CNN and attention-based spatial weighted CNN
Suguo Zhu, Zhenying Fang, Jun Yu 0002, Junping Du 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Deep Mixture of Diverse Experts for Large-Scale Visual RecognitionabstractIn this paper, a deep mixture of diverse experts algorithm is developed to achieve more efficient learning of a huge (mixture) network for large-scale visual recognition application. First, a two-layer ontology is constructed to assign large numbers of atomic object classes into a set of task groups according to the similarities of their learning complexities, where certain degrees of inter-group task overlapping are allowed to enable sufficient inter-group message passing. Second, one particular base deep CNNs with M+1 outputs is learned for each task group to recognize its M atomic object classes and identify one special class of "not-in-group", where the network structure (numbers of layers and units in each layer) of the well-designed deep CNNs (such as AlexNet, VGG, GoogleNet, ResNet) is directly used to configure such base deep CNNs. For enhancing the separability of the atomic object classes in the same task group, two approaches are developed to learn more discriminative base deep CNNs: (a) our deep multi-task learning algorithm that can effectively exploit the inter-class visual similarities; (b) our two-layer network cascade approach that can improve the accuracy rates for the hard object classes at certain degrees while effectively maintaining the high accuracy rates for the easy ones. Finally, all these complementary base deep CNNs with diverse but overlapped outputs are seamlessly combined to generate a mixture network with larger outputs for recognizing tens of thousands of atomic object classes. Our experimental results have demonstrated that our deep mixture of diverse experts algorithm can achieve very competitive results on large-scale visual recognition. Qiuyu Chen, Zhenzhong Kuang, Jun Yu 0002, Wei Zhang 0016, Jianping Fan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Adapting Stochastic Block Models to Power-Law Degree DistributionsabstractStochastic block models (SBMs) have been playing an important role in modeling clusters or community structures of network data. But, it is incapable of handling several complex features ubiquitously exhibited in real-world networks, one of which is the power-law degree characteristic. To this end, we propose a new variant of SBM, termed power-law degree SBM (PLD-SBM), by introducing degree decay variables to explicitly encode the varying degree distribution over all nodes. With an exponential prior, it is proved that PLD-SBM approximately preserves the scale-free feature in real networks. In addition, from the inference of variational E-Step, PLD-SBM is indeed to correct the bias inherited in SBM with the introduced degree decay factors. Furthermore, experiments conducted on both synthetic networks and two real-world datasets including Adolescent Health Data and the political blogs network verify the effectiveness of the proposed model in terms of cluster prediction accuracies. Maoying Qiao, Jun Yu 0002, Wei Bian 0003, Qiang Li 0024, Dacheng Tao |
IEEE Trans. Cybern. | 2 |
| 2019 | Multimodal Face-Pose Estimation With Multitask Manifold Deep LearningabstractFace-pose estimation aims at estimating the gazing direction with two-dimensional face images. It gives important communicative information and visual saliency. However, it is challenging because of lights, background, face orientations, and appearance visibility. Therefore, a descriptive representation of face images and mapping it to poses are critical. In this paper, we use multimodal data and propose a novel face-pose estimation framework named multitask manifold deep learning ($\text{M}^2\text{DL}$). It is based on feature extraction with improved convolutional neural networks (CNNs) and multimodal mapping relationship with multitask learning. In the proposed CNNs, manifold regularized convolutional layers learn the relationship between outputs of neurons in a low-rank space. Besides, in the proposed mapping relationship learning method, different modals of face representations are naturally combined by applying multitask learning with incoherent sparse and low-rank learning with a least-squares loss. Experimental results on three challenging benchmark datasets demonstrate the performance of$\text{M}^2\text{DL}$. Jun Yu 0002, Jian Zhang 0026, Xiongnan Jin, Kyong-Ho Lee |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | Image Recognition by Predicted User Click Feature With Multidomain Multitask Transfer Deep NetworkabstractThe click feature of an image, defined as a user click count vector based on click data, has been demonstrated to be effective for reducing the semantic gap for image recognition. Unfortunately, most of the traditional image recognition datasets do not contain click data. To address this problem, researchers have begun to develop a click prediction model using assistant datasets containing click information and have adapted this predictor to a common click-free dataset for different tasks. This method can be customized to our problem, but it has two main limitations: 1) the predicted click feature often performs badly in the recognition task since the prediction model is constructed independently of the subsequent recognition problem and 2) transferring the predictor from one dataset to another is challenging due to the large cross-domain diversity. In this paper, we devise a multitask and multidomain deep network with varied modals (MTMDD-VM) to formulate image recognition and click prediction tasks in a unified framework. Datasets with and without click information are integrated in the training. Furthermore, a nonlinear word embedding with a position-sensitive loss function is designed to discover the visual click correlation. We evaluate the proposed method on three public dog breed image datasets, and we utilize the Clickture-Dog dataset as the auxiliary dataset that provides click data. The experimental results show that: 1) the nonlinear word embedding and position-sensitive loss function largely enhance the predicted click feature in the recognition task, realizing a 32% improvement in accuracy; 2) the multitask learning framework improves accuracies in both image recognition and click prediction; and 3) the unified training using the combined dataset with and without click data further improves the performance. Compared with the state-of-the-art methods, the proposed approach not only performs much better in accuracy but also achieves good scalability and one-shot learning ability. Min Tan 0005, Jun Yu 0002, Hongyuan Zhang 0001, Yong Rui, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2019 | Zero-Shot Learning via Robust Latent Representation and Manifold RegularizationabstractZero-shot learning (ZSL) for visual recognition aims to accurately recognize the objects of unseen classes through mapping the visual feature to an embedding space spanned by class semantic information. However, the semantic gap across visual features and their underlying semantics is still a big obstacle in ZSL. Conventional ZSL methods construct that the mapping typically focus on the original visual features that are independent of the ZSL tasks, thus degrading the prediction performance. In this paper, we propose an effective method to uncover an appropriate latent representation of data for the purpose of zero-shot classification. Specifically, we formulate a novel framework to jointly learn the latent subspace and cross-modal embedding to link visual features with their semantic representations. The proposed framework combines feature learning and semantics prediction, such that the learned data representation is more discriminative to predict the semantic vectors, hence improving the overall classification performance. To learn a robust latent subspace, we explicitly avoid the information loss by ensuring the reconstruction ability of the obtained data representation. An efficient algorithm is designed to solve the proposed optimization problem. To fully exploit the intrinsic geometric structure of data, we develop a manifold regularization strategy to refine the learned semantic representations, leading to further improvements of the classification performance. To validate the effectiveness of the proposed approach, extensive experiments are conducted on three ZSL benchmarks and encouraging results are achieved compared with the state-of-the-art ZSL methods. Min Meng 0001, Jun Yu 0002 |
IEEE Trans. Image Process. | 2 |
| 2019 | Scalable Zero-Shot Learning via Binary Visual-Semantic EmbeddingsabstractZero-shot learning aims to classify visual instances from unseen classes in the absence of training examples. This is typically achieved by directly mapping visual features to a semantic embedding space of classes (e.g., attributes or word vectors), where the similarity between the two modalities can be readily measured. However, the semantic space may not be reliable for recognition due to the noisy class embeddings or visual bias problem. In this work, we propose a novel Binary embedding based Zero-Shot Learning (BZSL) method, which recognizes visual instances from unseen classes through an intermediate discriminative Hamming space. Specifically, BZSL jointly learns two binary coding functions to encode both visual instances and class embeddings into the Hamming space, which well alleviates the visual-semantic bias problem. As a desiring property, classifying an unseen instance thereby can be efficiently done by retrieving its nearest-class codes with minimal Hamming distance. During training, by introducing two auxiliary variables for the coding functions, we formulate an equivalent correlation maximization problem, which admits an analytical solution. The resulting algorithm thus enjoys both highly efficient training and scalable novel class inferring. Extensive experiments on four benchmark datasets, including the full ImageNet Fall 2011 dataset with over 20K unseen classes, demonstrate the superiority of our method on the zero-shot learning task. Particularly, we show that increasing the binary embedding dimension can inevitably improve the recognition accuracy. Fumin Shen, Xiang Zhou 0008, Jun Yu 0002, Yang Yang 0002, Li Liu 0004, Heng Tao Shen |
IEEE Trans. Image Process. | 3 |
| 2019 | Long-Form Video Question Answering via Dynamic Hierarchical Reinforced NetworksabstractOpen-ended long-form video question answering is a challenging task in visual information retrieval, which automatically generates a natural language answer from the referenced long-form video contents according to a given question. However, the existing works mainly focus on short-form video question answering, due to the lack of modeling semantic representations from long-form video contents. In this paper, we introduce a dynamic hierarchical reinforced network for open-ended long-form video question answering, which employs an encoder-decoder architecture with a dynamic hierarchical encoder and a reinforced decoder. Concretely, we first propose a frame-level dynamic long-short term memory (LSTM) network with binary segmentation gate to learn frame-level semantic representations according to the given question. We then develop a segment-level highway LSTM network with a question-aware highway gate for segment-level semantic modeling. Furthermore, we devise the reinforced decoder with a hierarchical attention mechanism to generate natural language answers. We construct a large-scale long-form video question answering dataset. The extensive experiments on the long-form dataset and another public short-form dataset show the effectiveness of our method. Zhou Zhao 0001, Shuwen Xiao, Zhenxin Xiao, Jun Yu 0002, Deng Cai 0001, Fei Wu 0001 |
IEEE Trans. Image Process. | 6 |
| 2019 | Effective 3-D Shape Retrieval by Integrating Traditional Descriptors and Pointwise ConvolutionabstractThe applications of isometric 3-D objects have recently received sufficient attention and, thus, it is very attractive to retrieve such isometric 3-D objects from large-scale collections. Although existing approaches have presented some interesting ideas, their performance is limited to their ability on feature representation. To improve the performance of 3-D object (shape) recognition, some recent algorithms prefer using complicated deep neural networks to learn discriminative features, but they consume huge amounts of computing resources. Instead, this paper presents a more effective solution by seamlessly integrating the traditional local descriptor with a deep pointwise convolutional network to extract 1-D features for shape recognition and retrieval. To reduce the costs of designing a complicated deep network, the first step of our algorithm is to describe the shape deformation by sampling a set of intrinsic point descriptors. Then, we introduce a simple yet effective pointwise convolutional network to integrate these descriptors as a global feature and the learning process can be significantly accelerated with the help of downsampling. Furthermore, a knowledge transfer strategy is used to upgrade our feature by compensating for information loss. Finally, we carry out experimental evaluations over popular shape benchmarks, and the results suggest that our approach exhibits superior accuracy rates and robustness on shape recognition and retrieval. Zhenzhong Kuang, Jun Yu 0002, Suguo Zhu, Zongmin Li, Jianping Fan 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | FishEyeRecNet: A Multi-context Collaborative Deep Network for Fisheye Image Rectification
Xiaoqing Yin, Xinchao Wang, Jun Yu 0002, Maojun Zhang, Pascal Fua, Dacheng Tao |
ECCV (10) | 3 |
| 2018 | Deep Point Convolutional Approach for 3D Model RetrievalabstractWith the increasing popularity of 3D models, retrieving deformable 3D objects is becoming a crucial task. The state-of-the-art methods use complex deep neural networks to address this problem, which require lots of computational resources. In this paper, we develop a more effective solution by using point convolution. Our algorithm takes local point descriptors as the input and produces a global vector for shape retrieval. To save the efforts of designing complex deep convolutional neural network (CNN), we first use intrinsic point descriptors to describe the shape deformations. Then, a simple but effective point CNN network is developed to integrate the local shape information by performing subspace compression and fusion, which depends on an end-to-end learning process to link the local and global information for discriminative shape representation. The experimental results on popular benchmarks have verified that our algorithm is able to outperform the state-of-the-art methods. Zhenzhong Kuang, Jun Yu 0002, Jianping Fan 0001, Min Tan 0005 |
ICME | 2 |
| 2018 | Rethinking Diversified and Discriminative Proposal Generation for Visual GroundingabstractVisual grounding aims to localize an object in an image referred to by a textual query phrase. Various visual grounding approaches have been proposed, and the problem can be modularized into a general framework: proposal generation, multi-modal feature representation, and proposal ranking. Of these three modules, most existing approaches focus on the latter two, with the importance of proposal generation generally neglected. In this paper, we rethink the problem of what properties make a good proposal generator. We introduce the diversity and discrimination simultaneously when generating proposals, and in doing so propose Diversified and Discriminative Proposal Networks model (DDPN). Based on the proposals generated by DDPN, we propose a high performance baseline model for visual grounding and evaluate it on four benchmark datasets. Experimental results demonstrate that our model delivers significant improvements on all the tested data-sets (e.g., 18.8% improvement on ReferItGame and 8.2% improvement on Flickr30k Entities over the existing state-of-the-arts respectively). Zhou Yu 0001, Jun Yu 0002, Chenchao Xiang, Zhou Zhao 0001, Qi Tian 0001, Dacheng Tao |
IJCAI | 2 |
| 2018 | Open-Ended Long-form Video Question Answering via Adaptive Hierarchical Reinforced NetworksabstractOpen-ended long-form video question answering is challenging problem in visual information retrieval, which automatically generates the natural language answer from the referenced long-form video content according to the question. However, the existing video question answering works mainly focus on the short-form video question answering, due to the lack of modeling the semantic representation of long-form video contents. In this paper, we consider the problem of long-form video question answering from the viewpoint of adaptive hierarchical reinforced encoder-decoder network learning. We propose the adaptive hierarchical encoder network to learn the joint representation of the long-form video contents according to the question with adaptive video segmentation. we then develop the reinforced decoder network to generate the natural language answer for open-ended video question answering. We construct a large-scale long-form video question answering dataset. The extensive experiments show the effectiveness of our method. Zhou Zhao 0001, Shuwen Xiao, Zhou Yu 0001, Jun Yu 0002, Deng Cai 0001, Fei Wu 0001, Yueting Zhuang |
IJCAI | 5 |
| 2018 | Deep Learning for Multimedia: Science or Technology?abstractDeep learning has been successfully explored in addressing different multimedia topics recent years, ranging from object detection, semantic classification, entity annotation, to multimedia captioning, multimedia question answering and storytelling. Open source libraries and platforms such as Tensorflow, Caffe, MXnet significantly help promote the wide deployment of deep learning in solving real-world applications. On one hand, deep learning practitioners, while not necessary to understand the involved math behind, are able to set up and make use of a complex deep network. One recent deep learning tool based on Keras even provides the graphical interface to enable straightforward 'drag and drop' operation for deep learning programming. On the other hand, however, some general theoretical problems of learning such as the interpretation and generalization, have only achieved limited progress. Most deep learning papers published these days follow the pipeline of designing/modifying network structures - tuning parameters - reporting performance improvement in specific applications. We have even seen many deep learning application papers without one single equation. Theoretical interpretation and the science behind the study are largely ignored. While excited about the successful application of deep learning in classical and novel problems, we multimedia researchers are responsible to think and solve the fundamental topics in deep learning science. Prof. Guanrong Chen recently wrote an editorial note titled 'Science and Technology, not SciTech' [1]. This panel falls into similar discussion and aims to invite prestigious multimedia researchers and active deep learning practitioners to discuss the positioning of deep learning research now and in the future. Specifically, each panelist is asked to present their opinions on the following five questions: 1)How do you think the current phenomenon that deep learning applications are explosively growing, while the general theoretical problems remain slow progress? 2)Do you agree that deployment of deep learning techniques is getting easy (with a low barrier), while deep learning research is difficult (with a high barrier) 3)What do you think are the core problems for deep learning techniques? 4)What do you think are the core problems for deep learning science? 5)What's your suggestion on the multimedia research in the post-deep learning era? Jun Yu 0002, Ramesh Jain 0001, Rainer Lienhart, Peng Cui 0001, Jiashi Feng |
ACM Multimedia | 2 |
| 2018 | Comprehensive Distance-Preserving Autoencoders for Cross-Modal RetrievalabstractIn this paper, we propose a novel method with comprehensive distance-preserving autoencoders (CDPAE) to address the problem of unsupervised cross-modal retrieval. Previous unsupervised methods rely primarily on pairwise distances of representations extracted from cross media spaces that co-occur and belong to the same objects. However, besides pairwise distances, the CDPAE also considers heterogeneous distances of representations extracted from cross media spaces as well as homogeneous distances of representations extracted from single media spaces that belong to different objects. The CDPAE consists of four components. First, denoising autoencoders are used to retain the information from the representations and to reduce the negative influence of redundant noises. Second, a comprehensive distance-preserving common space is proposed to explore the correlations among different representations. This aims to preserve the respective distances between the representations within the common space so that they are consistent with the distances in their original media spaces. Third, a novel joint loss function is defined to simultaneously calculate the reconstruction loss of the denoising autoencoders and the correlation loss of the comprehensive distance-preserving common space. Finally, an unsupervised cross-modal similarity measurement is proposed to further improve the retrieval performance. This is carried out by calculating the marginal probability of two media objects based on a kNN classifier. The CDPAE is tested on four public datasets with two cross-modal retrieval tasks: "query images by texts" and "query texts by images". Compared with eight state-of-the-art cross-modal retrieval methods, the experimental results demonstrate that the CDPAE outperforms all the unsupervised methods and performs competitively with the supervised methods. Yibing Zhan, Jun Yu 0002, Zhou Yu 0001, Rong Zhang 0004, Dacheng Tao, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2018 | Click data guided query modeling with click propagation and sparse coding
Min Tan 0005, Jun Yu 0002, Qingming Huang, Weichen Wu |
Multim. Tools Appl. | 2 |
| 2018 | Machine learning for big visual analysis
Jun Yu 0002, Xue Mei, Fatih Porikli, Jason J. Corso |
Mach. Vis. Appl. | 1 |
| 2018 | Blind image quality prediction by exploiting multi-level deep representations
Fei Gao 0006, Jun Yu 0002, Suguo Zhu, Qingming Huang, Qi Tian 0001 |
Pattern Recognit. | 2 |
| 2018 | Integrating multi-level deep learning and concept ontology for large-scale visual recognition
Zhenzhong Kuang, Jun Yu 0002, Zongmin Li, Baopeng Zhang, Jianping Fan 0001 |
Pattern Recognit. | 2 |
| 2018 | Face biometric quality assessment via light CNN
Jun Yu 0002, Kejia Sun, Fei Gao 0006, Suguo Zhu |
Pattern Recognit. Lett. | 1 |
| 2018 | Unsupervised image segmentation via Stacked Denoising Auto-encoder and hierarchical patch indexing
Jun Yu 0002, Di Huang 0007, Zhongliang Wei |
Signal Process. | 1 |
| 2018 | Leveraging Content Sensitiveness and User Trustworthiness to Recommend Fine-Grained Privacy Settings for Social Image SharingabstractTo configure successful privacy settings for social image sharing, two issues are inseparable: 1) content sensitiveness of the images being shared; and 2) trustworthiness of the users being granted to see the images. This paper aims to consider these two inseparable issues simultaneously to recommend fine-grained privacy settings for social image sharing. For achieving more compact representation of image content sensitiveness (privacy), two approaches are developed: 1) a deep network is adapted to extract 1024-D discriminative deep features; and 2) a deep multiple instance learning algorithm is adopted to identify 280 privacy-sensitive object classes and events. Second, users on the social network are clustered into a set of representative social groups to generate a discriminative dictionary for user trustworthiness characterization. Finally, both the image content sensitiveness and the user trustworthiness are integrated to train a tree classifier to recommend fine-grained privacy settings for social image sharing. Our experimental studies have demonstrated both the efficiency and the effectiveness of our proposed algorithms. Jun Yu 0002, Zhenzhong Kuang, Baopeng Zhang, Wei Zhang 0016, Dan Lin 0001, Jianping Fan 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | Local Deep-Feature Alignment for Unsupervised Dimension ReductionabstractThis paper presents an unsupervised deep-learning framework named Local Deep-Feature Alignment (LDFA) for dimension reduction. We construct neighbourhood for each data sample and learn a local Stacked Contractive Auto-encoder (SCAE) from the neighbourhood to extract the local deep features. Next, we exploit an affine transformation to align the local deep features of each neighbourhood with the global features. Moreover, we derive an approach from LDFA to map explicitly a new data sample into the learned low-dimensional subspace. The advantage of the LDFA method is that it learns both local and global characteristics of the data sample set: the local SCAEs capture local characteristics contained in the data set, while the global alignment procedures encode the interdependencies between neighbourhoods into the final lowdimensional feature representations. Experimental results on data visualization, clustering and classification show that the LDFA method is competitive with several well-known dimension reduction techniques, and exploiting locality in deep learning is a research topic worth further exploring. Jian Zhang 0026, Jun Yu 0002, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2018 | Embedding Visual Hierarchy With Deep Networks for Large-Scale Visual RecognitionabstractIn this paper, a layer-wise mixture model (LMM) is developed to support hierarchical visual recognition, where a Bayesian approach is used to automatically adapt the visual hierarchy to the progressive improvements of the deep network along the time. Our LMM algorithm can provide an end-to-end approach for jointly learning: (a) the deep network for achieving more discriminative deep representations for object classes and their inter-class visual similarities; (b) the tree classifier for recognizing large numbers of object classes hierarchically; and (c) the visual hierarchy adaptation for achieving more accurate assignment and organization of large numbers of object classes. By learning the tree classifier, the deep network and the visual hierarchy adaptation jointly in an end-to-end manner, our LMM algorithm can achieve higher accuracy rates on hierarchical visual recognition. Our experiments are carried on ImageNet1K and ImageNet10K image sets, which have demonstrated that our LMM algorithm can achieve very competitive results on the accuracy rates as compared with the baseline methods. Baopeng Zhang, Wei Zhang 0016, Jun Yu 0002, Jianping Fan 0001 |
IEEE Trans. Image Process. | 6 |
| 2018 | Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question AnsweringabstractVisual question answering (VQA) is challenging, because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both the image and the question; 2) multimodal feature fusion that is able to capture the complex interactions between multimodal features; and 3) automatic answer prediction that is able to consider the complex correlations between multiple diverse answers for the same question. For fine-grained image and question representations, a "coattention" mechanism is developed using a deep neural network (DNN) architecture to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features effectively and obtain more discriminative features for image and question representations. For multimodal feature fusion, a generalized multimodal factorized high-order pooling approach (MFH) is developed to achieve more effective fusion of multimodal features by exploiting their correlations sufficiently, which can further result in superior VQA performance as compared with the state-of-the-art approaches. For answer prediction, the Kullback-Leibler divergence is used as the loss function to achieve precise characterization of the complex correlations between multiple diverse answers with the same or similar meaning, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction. A DNN architecture is designed to integrate all these aforementioned modules into a unified model for achieving superior VQA performance. With an ensemble of our MFH models, we achieve the state-of-the-art performance on the large-scale VQA data sets and win the runner-up in VQA Challenge 2017. Zhou Yu 0001, Jun Yu 0002, Chenchao Xiang, Jianping Fan 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | User-Click-Data-Based Fine-Grained Image Recognition via Weakly Supervised Metric LearningabstractWe present a novel fine-grained image recognition framework using user click data, which can bridge the semantic gap in distinguishing categories that are similar in visual. As query set in click data is usually large-scale and redundant, we first propose a click-feature-based query-merging approach to merge queries with similar semantics and construct a compact click feature. Afterward, we utilize this compact click feature and convolutional neural network (CNN)-based deep visual feature to jointly represent an image. Finally, with the combined feature, we employ the metriclearning-based template-matching scheme for efficient recognition. Considering the heavy noise in the training data, we introduce a reliability variable to characterize the image reliability, and propose a weakly-supervised metric and template leaning with smooth assumption and click prior (WMTLSC) method to jointly learn the distance metric, object templates, and image reliability. Extensive experiments are conducted on a public Clickture-Dog dataset and our newly established Clickture-Bird dataset. It is shown that the click-data-based query merging helps generating a highly compact (the dimension is reduced to 0.9%) and dense click feature for images, which greatly improves the computational efficiency. Also, introducing this click feature into CNN feature further boosts the recognition accuracy. The proposed framework performs much better than previous state-of-the-arts in fine-grained recognition tasks. Min Tan 0005, Jun Yu 0002, Zhou Yu 0001, Fei Gao 0006, Yong Rui, Dacheng Tao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | Multi-modal Factorized Bilinear Pooling with Co-attention Learning for Visual Question AnsweringabstractVisual question answering (VQA) is challenging because it requires a simultaneous understanding of both the visual content of images and the textual content of questions. The approaches used to represent the images and questions in a fine-grained manner and questions and to fuse these multimodal features play key roles in performance. Bilinear pooling based models have been shown to outperform traditional linear models for VQA, but their high-dimensional representations and high computational complexity may seriously limit their applicability in practice. For multimodal feature fusion, here we develop a Multi-modal Factorized Bilinear (MFB) pooling approach to efficiently and effectively combine multi-modal features, which results in superior performance for VQA compared with other bilinear pooling approaches. For fine-grained image and question representation, we develop a `co-attention' mechanism using an end-to-end deep network architecture to jointly learn both the image and question attentions. Combining the proposed MFB approach with co-attention learning in a new network architecture provides a unified model for VQA. Our experimental results demonstrate that the single MFB with co-attention model achieves new state-of-theart performance on the real-world VQA dataset. Code available at https://github.com/yuzcccc/mfb. Zhou Yu 0001, Jun Yu 0002, Jianping Fan 0001, Dacheng Tao |
ICCV | 2 |
| 2017 | Convolutional neural networks for intestinal hemorrhage detection in wireless capsule endoscopy imagesabstractWireless capsule endoscopy (WCE) can painlessly capture a large number of images inside the intestine. However, only a small portion of these WCE images contain hemorrhage. It is thus critical to develop automated hemorrhage detection method to facilitate the diagnosis of intestinal diseases. However, automated hemorrhage detection is complicated by 1) the extreme imbalance between the amount of hemorrhage images and that of normal images; and 2) the variety of the appearance, texture, and luminance inside the intestine. In this paper, we proposed to learn a robust intestinal hemorrhage detection model via Convolutional Neural Networks (CNNs), because of CNNs' extraordinary performance in solving various image understanding tasks. Specially, we explored different CNN architectures and data augmentation methods. Besides, we investigated the correlation between hemorrhage detection accuracy and image quality. Across about 1.3k hemorrhage images and 40k normal images, the learned CNN model achieves an F-measure of 98.87%. Panpeng Li, Ziyun Li 0002, Fei Gao 0006, Jun Yu 0002 |
ICME | 5 |
| 2017 | Fine-grained image recognition via weakly supervised click data guided bilinear CNN modelabstractBilinear convolutional neural networks (BCNN) model, the state-of-the-art in fine-grained image recognition, fails in distinguishing the categories with subtle visual differences. We design a novel BCNN model guided by user click data (C-BCNN) to improve the performance via capturing both the visual and semantical content in images. Specially, to deal with the heavy noise in large-scale click data, we propose a weakly supervised learning approach to learn the C-BCNN, namely W-C-BCNN. It can automatically weight the training images based on their reliability. Extensive experiments are conducted on the public Clickture-Dog dataset. It shows that: (1) integrating CNN with click feature largely improves the performance; (2) both the click data and visual consistency can help to model image reliability. Moreover, the method can be easily customized to medical image recognition. Our model performs much better than conventional BCNN models on both the Clickture-Dog and medical image dataset. Guangjian Zheng, Min Tan 0005, Jun Yu 0002, Jianping Fan 0001 |
ICME | 3 |
| 2017 | Deep Mixture of Experts with Diverse Task SpacesabstractIn this paper, a deep mixture algorithm is developed to support large-scale visual recognition (e.g., recognizing tens of thousands of object classes) by seamlessly combining a set of base deep CNNs (AlexNet) with diverse task spaces, e.g., such base deep CNNs (i.e., diverse experts) are trained to recognize different subsets of tens of thousands of object classes rather than the same set of object classes. Our experimental results have demonstrated that our deep mixture algorithm can achieve very competitive results on large-scale visual recognition. Jianping Fan 0001, Zhenzhong Kuang, Zhou Yu 0001, Jun Yu 0002 |
ICMLA | 5 |
| 2017 | Privacy Setting Recommendation for Image SharingabstractThis paper aims to simultaneously consider two inseparable issues for privacy setting recommendation: (1) sensitiveness of visual content of the images being shared; and (2) trustworthiness of users being granted. First, an object-based approach is developed for image content sensitiveness (privacy) representation. Secondly, the users on a social network are clustered into a set of representative social groups to generate a discriminative dictionary for user trustworthiness characterization. Finally, a tree classifier is trained hierarchically to recommend appropriate privacy settings for image sharing. Jun Yu 0002, Zhenzhong Kuang, Zhou Yu 0001, Dan Lin 0001, Jianping Fan 0001 |
ICMLA | 1 |
| 2017 | Improving Stochastic Block Models by Incorporating Power-Law Degree CharacteristicabstractStochastic block models (SBMs) provide a statistical way modeling network data, especially in representing clusters or community structures. However, most block models do not consider complex characteristics of networks such as scale-free feature, making them incapable of handling degree variation of vertices, which is ubiquitous in real networks. To address this issue, we introduce degree decay variables into SBM, termed power-law degree SBM (PLD-SBM), to model the varying probability of connections between node pairs. The scale-free feature is approximated by a power-law degree characteristic. Such a property allows PLD-SBM to correct the distortion of degree distribution in SBM, and thus improves the performance of cluster prediction. Experiments on both simulated networks and two real-world networks including the Adolescent Health Data and the political blogs network demonstrate the validity of the motivation of PLD-SBM, and its practical superiority. Maoying Qiao, Jun Yu 0002, Wei Bian 0003, Qiang Li 0024, Dacheng Tao |
IJCAI | 2 |
| 2017 | DeepSim: Deep similarity for image quality assessment
Fei Gao 0006, Panpeng Li, Min Tan 0005, Jun Yu 0002, Yani Zhu |
Neurocomputing | 5 |
| 2017 | Exemplar-based 3D human pose estimation with sparse spectral embedding
Jun Yu 0002 |
Neurocomputing | 1 |
| 2017 | Machine learning and signal processing for big multimedia analysis
Jun Yu 0002, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2017 | Three-dimensional image-based human pose recovery with hypergraph regularized autoencoders
Jun Yu 0002, Jane You, Zhiwen Yu 0002 |
Multim. Tools Appl. | 2 |
| 2017 | Diversified dictionaries for multi-instance learning
Maoying Qiao, Liu Liu 0014, Jun Yu 0002, Chang Xu 0002, Dacheng Tao |
Pattern Recognit. | 3 |
| 2017 | Constrained Low-Rank Learning Using Least Squares-Based RegularizationabstractLow-rank learning has attracted much attention recently due to its efficacy in a rich variety of real-world tasks, e.g., subspace segmentation and image categorization. Most low-rank methods are incapable of capturing low-dimensional subspace for supervised learning tasks, e.g., classification and regression. This paper aims to learn both the discriminant low-rank representation (LRR) and the robust projecting subspace in a supervised manner. To achieve this goal, we cast the problem into a constrained rank minimization framework by adopting the least squares regularization. Naturally, the data label structure tends to resemble that of the corresponding low-dimensional representation, which is derived from the robust subspace projection of clean data by low-rank learning. Moreover, the low-dimensional representation of original data can be paired with some informative structure by imposing an appropriate constraint, e.g., Laplacian regularizer. Therefore, we propose a novel constrained LRR method. The objective function is formulated as a constrained nuclear norm minimization problem, which can be solved by the inexact augmented Lagrange multiplier algorithm. Extensive experiments on image classification, human pose estimation, and robust face recovery have confirmed the superiority of our method. Ping Li 0006, Jun Yu 0002, Meng Wang 0001, Deng Cai 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2017 | Deep Multimodal Distance Metric Learning Using Click Constraints for Image RankingabstractHow do we retrieve images accurately? Also, how do we rank a group of images precisely and efficiently for specific queries? These problems are critical for researchers and engineers to generate a novel image searching engine. First, it is important to obtain an appropriate description that effectively represent the images. In this paper, multimodal features are considered for describing images. The images unique properties are reflected by visual features, which are correlated to each other. However, semantic gaps always exist between images visual features and semantics. Therefore, we utilize click feature to reduce the semantic gap. The second key issue is learning an appropriate distance metric to combine these multimodal features. This paper develops a novel deep multimodal distance metric learning (Deep-MDML) method. A structured ranking model is adopted to utilize both visual and click features in distance metric learning (DML). Specifically, images and their related ranking results are first collected to form the training set. Multimodal features, including click and visual features, are collected with these images. Next, a group of autoencoders is applied to obtain initially a distance metric in different visual spaces, and an MDML method is used to assign optimal weights for different modalities. Next, we conduct alternating optimization to train the ranking model, which is used for the ranking of new queries with click features. Compared with existing image ranking methods, the proposed method adopts a new ranking model to use multimodal features, including click features and visual features in DML. We operated experiments to analyze the proposed Deep-MDML in two benchmark data sets, and the results validate the effects of the method. Jun Yu 0002, Xiaokang Yang 0001, Fei Gao 0006, Dacheng Tao |
IEEE Trans. Cybern. | 1 |
| 2017 | Coupled Deep Autoencoder for Single Image Super-ResolutionabstractSparse coding has been widely applied to learning-based single image super-resolution (SR) and has obtained promising performance by jointly learning effective representations for low-resolution (LR) and high-resolution (HR) image patch pairs. However, the resulting HR images often suffer from ringing, jaggy, and blurring artifacts due to the strong yet ad hoc assumptions that the LR image patch representation is equal to, is linear with, lies on a manifold similar to, or has the same support set as the corresponding HR image patch representation. Motivated by the success of deep learning, we develop a data-driven model coupled deep autoencoder (CDA) for single image SR. CDA is based on a new deep architecture and has high representational capability. CDA simultaneously learns the intrinsic representations of LR and HR image patches and a big-data-driven function that precisely maps these LR representations to their corresponding HR representations. Extensive experimentation demonstrates the superior effectiveness and efficiency of CDA for single image SR compared to other state-of-the-art methods on Set5 and Set14 datasets. Jun Yu 0002, Ruxin Wang 0002, Cuihua Li, Dacheng Tao |
IEEE Trans. Cybern. | 2 |
| 2017 | iPrivacy: Image Privacy Protection by Identifying Sensitive Objects via Deep Multi-Task LearningabstractTo achieve automatic recommendation of privacy settings for image sharing, a new tool called iPrivacy (image privacy) is developed for releasing the burden from users on setting the privacy preferences when they share their images for special moments. Specifically, this paper consists of the following contributions: 1) massive social images and their privacy settings are leveraged to learn the object-privacy relatedness effectively and identify a set of privacy-sensitive object classes automatically; 2) a deep multi-task learning algorithm is developed to jointly learn more representative deep convolutional neural networks and more discriminative tree classifier, so that we can achieve fast and accurate detection of large numbers of privacy-sensitive object classes; 3) automatic recommendation of privacy settings for image sharing can be achieved by detecting the underlying privacy-sensitive objects from the images being shared, recognizing their classes, and identifying their privacy settings according to the object-privacy relatedness; and 4) one simple solution for image privacy protection is provided by blurring the privacy-sensitive objects automatically. We have conducted extensive experimental studies on real-world images and the results have demonstrated both the efficiency and effectiveness of our proposed approach. Jun Yu 0002, Baopeng Zhang, Zhenzhong Kuang, Dan Lin 0001, Jianping Fan 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | HD-MTL: Hierarchical Deep Multi-Task Learning for Large-Scale Visual RecognitionabstractIn this paper, a hierarchical deep multi-task learning (HD-MTL) algorithm is developed to support large-scale visual recognition (e.g., recognizing thousands or even tens of thousands of atomic object classes automatically). First, multiple sets of multi-level deep features are extracted from different layers of deep convolutional neural networks (deep CNNs), and they are used to achieve more effective accomplishment of the coarseto- fine tasks for hierarchical visual recognition. A visual tree is then learned by assigning the visually-similar atomic object classes with similar learning complexities into the same group, which can provide a good environment for determining the interrelated learning tasks automatically. By leveraging the inter-task relatedness (inter-class similarities) to learn more discriminative group-specific deep representations, our deep multi-task learning algorithm can train more discriminative node classifiers for distinguishing the visually-similar atomic object classes effectively. Our hierarchical deep multi-task learning (HD-MTL) algorithm can integrate two discriminative regularization terms to control the inter-level error propagation effectively, and it can provide an end-to-end approach for jointly learning more representative deep CNNs (for image representation) and more discriminative tree classifier (for large-scale visual recognition) and updating them simultaneously. Our incremental deep learning algorithms can effectively adapt both the deep CNNs and the tree classifier to the new training images and the new object classes. Our experimental results have demonstrated that our HD-MTL algorithm can achieve very competitive results on improving the accuracy rates for large-scale visual recognition. Jianping Fan 0001, Zhenzhong Kuang, Yu Zheng 0006, Ji Zhang 0005, Jun Yu 0002, Jinye Peng 0001 |
IEEE Trans. Image Process. | 6 |
| 2016 | Multi-modal Image Re-ranking with Autoencoders and Click Semantics
Chaohui Tang, Qingxin Zhu, Jun Yu 0002 |
MMM (1) | 4 |
| 2016 | Data-driven facial animation via hypergraph learningabstractData-driven facial animation has attracted much attention in recent years. Existing facial animation methods may not preserve the topology structure, and cannot achieve a natural face. This paper proposes a new data-driven facial animation method based on hypergraph learning. It drives a neutral face to a certain expression face. This paper assumes that neutral face has similar topology with the expression face, we compute the alignment laplacian matrix using hypergraph learning. To get a natural face, we add a constraint item which is consisted of a set of motion data. Experiment results demonstrate that our method can achieve a natural expression face. And the results show the superiority over the state-of-art. Jun Yu 0002, Fei Gao 0006, Jian Zhang 0026 |
SMC | 2 |
| 2016 | Photo aesthetic quality assessment via label distribution learningabstractAutomatic prediction of photo aesthetic quality is useful for many practical purposes. Current computational approaches typically solved this problem by assigning a categorical label (good or bad) to a photo. However, due to the subjectivity and complexity of humans aesthetic judgments, only a categorical label is insufficient to represent humans perceived aesthetic quality of a photo. This paper focuses on an interesting problem: is it possible to predict the crowed opinions about the aesthetic quality of a photo? The crowed opinion here is expressed by the distribution of scores given by a number of subjects. For each given photo, a deep convolutional neural network (DCNN) is utilized to calculate its feature representation. Afterwards, the crowed opinion prediction problem is formulated as one of label distribution learning (LDL). Experiments show that the proposed method is highly effective and outperforms state-of-the-art algorithms. Fei Gao 0006, Di Huang 0007, Min Tan 0005, Jun Yu 0002 |
SMC | 5 |
| 2016 | Recent developments on deep big vision
Jun Yu 0002, Dapeng Tao, Richang Hong, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2016 | Boosting video popularity through keyword suggestion and recommendation systems
Samamon Khemmarat, Lixin Gao 0001, Jian Wan 0001, Yuyu Yin, Jun Yu 0002 |
Neurocomputing | 7 |
| 2016 | Towards robust subspace recovery via sparsity-constrained latent low-rank representation
Ping Li 0006, Jiajun Bu, Jun Yu 0002, Chun Chen 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Realtime and robust object matching with a large number of templates
Jianke Zhu, Jun Yu 0002, Jun Cheng 0002 |
Multim. Tools Appl. | 3 |
| 2016 | Data-driven facial animation via semi-supervised local patch alignment
Jian Zhang 0026, Jun Yu 0002, Jane You, Dapeng Tao, Jun Cheng 0002 |
Pattern Recognit. | 2 |
| 2016 | Biologically inspired image quality assessment
Fei Gao 0006, Jun Yu 0002 |
Signal Process. | 2 |
| 2015 | Hessian Regularized Sparse Coding for Human Action Recognition
Weifeng Liu 0001, Zhen Wang 0004, Dapeng Tao, Jun Yu 0002 |
MMM (2) | 4 |
| 2015 | Human pose recovery by supervised spectral embedding
Jun Yu 0002, Yukun Guo, Dapeng Tao, Jian Wan 0001 |
Neurocomputing | 1 |
| 2015 | l2, 1 Norm regularized fisher criterion for optimal feature selection
Jian Zhang 0026, Jun Yu 0002, Jian Wan 0001 |
Neurocomputing | 2 |
| 2015 | Multi-view ensemble manifold regularization for 3D object recognition
Jun Yu 0002, Jane You, Dapeng Tao |
Inf. Sci. | 2 |
| 2015 | Low-rank matrix factorization with multiple Hypergraph regularizer
Taisong Jin, Jun Yu 0002, Jane You, Cuihua Li, Zhengtao Yu 0001 |
Pattern Recognit. | 2 |
| 2015 | Semantic embedding for indoor scene recognition by weighted hypergraph learning
Jun Yu 0002, Dapeng Tao, Meng Wang 0001 |
Signal Process. | 1 |
| 2015 | Machine learning and signal processing for human pose recovery and behavior analysis
Jun Yu 0002, Huiyu Zhou 0001, Xinbo Gao 0001 |
Signal Process. | 1 |
| 2015 | Learning to Rank Using User Clicks and Visual Features for Image RetrievalabstractThe inconsistency between textual features and visual contents can cause poor image search results. To solve this problem, click features, which are more reliable than textual information in justifying the relevance between a query and clicked images, are adopted in image ranking model. However, the existing ranking model cannot integrate visual features, which are efficient in refining the click-based search results. In this paper, we propose a novel ranking model based on the learning to rank framework. Visual features and click features are simultaneously utilized to obtain the ranking model. Specifically, the proposed approach is based on large margin structured output learning and the visual consistency is integrated with the click features through a hypergraph regularizer term. In accordance with the fast alternating linearization method, we design a novel algorithm to optimize the objective function. This algorithm alternately minimizes two different approximations of the original objective function by keeping one function unchanged and linearizing the other. We conduct experiments on a large-scale dataset collected from the Microsoft Bing image search engine, and the results demonstrate that the proposed learning to rank models based on visual features and user clicks outperforms state-of-the-art algorithms. Jun Yu 0002, Dacheng Tao, Meng Wang 0001, Yong Rui |
IEEE Trans. Cybern. | 1 |
| 2015 | Semiautomated Extraction of Street Light Poles From Mobile LiDAR Point-CloudsabstractThis paper proposes a novel algorithm for extracting street light poles from vehicleborne mobile light detection and ranging (LiDAR) point-clouds. First, the algorithm rapidly detects curb-lines and segments a point-cloud into road and nonroad surface points based on trajectory data recorded by the integrated position and orientation system onboard the vehicle. Second, the algorithm accurately extracts street light poles from the segmented nonroad surface points using a novel pairwise 3-D shape context. The proposed algorithm is tested on a set of point-clouds acquired by a RIEGL VMX-450 mobile LiDAR system. The results show that road surfaces are correctly segmented, and street light poles are robustly extracted with a completeness exceeding 99%, a correctness exceeding 97%, and a quality exceeding 96%, thereby demonstrating the efficiency and feasibility of the proposed algorithm to segment road surfaces and extract street light poles from huge volumes of mobile LiDAR point-clouds. Yongtao Yu, Jonathan Li 0001, Haiyan Guan, Cheng Wang 0003, Jun Yu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2015 | Multimodal Deep Autoencoder for Human Pose RecoveryabstractVideo-based human pose recovery is usually conducted by retrieving relevant poses using image features. In the retrieving process, the mapping between 2D images and 3D poses is assumed to be linear in most of the traditional methods. However, their relationships are inherently non-linear, which limits recovery performance of these methods. In this paper, we propose a novel pose recovery method using non-linear mapping with multi-layered deep neural network. It is based on feature extraction with multimodal fusion and back-propagation deep learning. In multimodal fusion, we construct hypergraph Laplacian with low-rank representation. In this way, we obtain a unified feature description by standard eigen-decomposition of the hypergraph Laplacian matrix. In back-propagation deep learning, we learn a non-linear mapping from 2D images to 3D poses with parameter fine-tuning. The experimental results on three data sets show that the recovery error has been reduced by 20%-25%, which demonstrates the effectiveness of the proposed method. Jun Yu 0002, Jian Wan 0001, Dacheng Tao, Meng Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Structured action classification with hypergraph regularizationabstractTraditional multi-class classifying methods treat outputs separately. It leads to a multiclass problem with a very large number of classes and downgrades the performance of classifiers. Actually, the outputs of different testing samples are usually interdependent. Therefore, we propose a novel method of structured classification based on SVM and hypergraph regularization (Hyper-SSVM). First, it exploits the structure and dependencies within classifying outputs. Second, we impose local constraints to samples by using Hypergraph regularization. We apply the proposed Hyper-SSVM to action classification. The experimental results demonstrate the effectiveness of the proposed method. Jun Yu 0002 |
SMC | 2 |
| 2014 | Genetic algorithm for spanning tree construction in P2P distributed interactive applications
Yusen Li, Jun Yu 0002, Dapeng Tao |
Neurocomputing | 2 |
| 2014 | Image clustering by hyper-graph regularized non-negative matrix factorization
Jun Yu 0002, Cuihua Li, Jane You, Taisong Jin |
Neurocomputing | 2 |
| 2014 | Semantic preserving distance metric learning and applications
Jun Yu 0002, Dapeng Tao, Jonathan Li 0001, Jun Cheng 0002 |
Inf. Sci. | 1 |
| 2014 | Automated Detection of Road Manhole and Sewer Well Covers From Mobile LiDAR Point CloudsabstractA novel object detection algorithm is developed for automatically detecting road manhole and sewer well covers from mobile light detection and ranging point clouds. This algorithm takes advantage of a marked point process of disks and rectangles to model the locations of manhole and sewer well covers and their geometric dimensions. A reversible jump Markov chain Monte Carlo algorithm is implemented for simulating the posterior distribution obtained using a Bayesian paradigm. The detection results obtained from the road surface point clouds acquired by a RIEGL VMX-450 system show that the manhole and sewer well covers can be detected automatically and accurately. The performance achieved using the proposed algorithm is much more accurate and effective than those of the other three existing algorithms. Yongtao Yu, Jonathan Li 0001, Haiyan Guan, Cheng Wang 0003, Jun Yu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2014 | Pairwise Three-Dimensional Shape Context for Partial Object Matching and Retrieval on Mobile Laser Scanning DataabstractA novel pairwise 3-D shape context for partial object matching and retrieval is developed for extracting 3-D light poles and trees from mobile laser scanning (MLS) point clouds in a typical urban street scene. Unlike the single-point shape context describing only the local topology of a shape, the pairwise 3-D shape context can simultaneously model the local and global geometric structures of a shape in manifold space. By using histogram descriptors, the pairwise 3-D shape context has such characteristics as invariance to scale, invariance to orientation, and partial insensitivity to topological changes. Our results show that 3-D light poles and individual trees can be extracted from the RIEGL VMX-450 MLS point clouds and the performance achieved using our algorithm is much more accurate and effective than those of the other two existing algorithms. Yongtao Yu, Jonathan Li 0001, Jun Yu 0002, Haiyan Guan, Cheng Wang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2014 | Image clustering based on sparse patch alignment framework
Jun Yu 0002, Richang Hong, Meng Wang 0001, Jane You |
Pattern Recognit. | 1 |
| 2014 | High-Order Distance-Based Multiview Stochastic Learning in Image ClassificationabstractHow do we find all images in a larger set of images which have a specific content? Or estimate the position of a specific object relative to the camera? Image classification methods, like support vector machine (supervised) and transductive support vector machine (semi-supervised), are invaluable tools for the applications of content-based image retrieval, pose estimation, and optical character recognition. However, these methods only can handle the images represented by single feature. In many cases, different features (or multiview data) can be obtained, and how to efficiently utilize them is a challenge. It is inappropriate for the traditionally concatenating schema to link features of different views into a long vector. The reason is each view has its specific statistical property and physical interpretation. In this paper, we propose a high-order distance-based multiview stochastic learning (HD-MSL) method for image classification. HD-MSL effectively combines varied features into a unified representation and integrates the labeling information based on a probabilistic framework. In comparison with the existing strategies, our approach adopts the high-order distance obtained from the hypergraph to replace pairwise distance in estimating the probability matrix of data distribution. In addition, the proposed approach can automatically learn a combination coefficient for each view, which plays an important role in utilizing the complementary information of multiview data. An alternative optimization is designed to solve the objective functions of HD-MSL and obtain different views on coefficients and classification scores simultaneously. Experiments on two real world datasets demonstrate the effectiveness of HD-MSL in image classification. Jun Yu 0002, Yong Rui, Yuan Yan Tang, Dacheng Tao |
IEEE Trans. Cybern. | 1 |
| 2014 | Click Prediction for Web Image Reranking Using Multimodal Sparse CodingabstractImage reranking is effective for improving the performance of a text-based image search. However, existing reranking algorithms are limited for two main reasons: 1) the textual meta-data associated with images is often mismatched with their actual visual content and 2) the extracted visual features do not accurately describe the semantic similarities between images. Recently, user click information has been used in image reranking, because clicks have been shown to more accurately describe the relevance of retrieved images to search queries. However, a critical problem for click-based methods is the lack of click data, since only a small number of web images have actually been clicked on by users. Therefore, we aim to solve this problem by predicting image clicks. We propose a multimodal hypergraph learning-based sparse coding method for image click prediction, and apply the obtained click data to the reranking of images. We adopt a hypergraph to build a group of manifolds, which explore the complementarity of different features through a group of weights. Unlike a graph that has an edge between two vertices, a hyperedge in a hypergraph connects a set of vertices, and helps preserve the local smoothness of the constructed sparse codes. An alternating optimization procedure is then performed, and the weights of different modalities and the sparse codes are simultaneously obtained. Finally, a voting strategy is used to describe the predicted click as a binary event (click or no click), from the images' corresponding sparse codes. Thorough empirical studies on a large-scale database including nearly 330 K images demonstrate the effectiveness of our approach for click prediction when compared with several other methods. Additional image reranking experiments on real-world data show the use of click prediction is beneficial to improving the performance of prominent graph-based image reranking algorithms. Jun Yu 0002, Yong Rui, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2014 | Exploiting Click Constraints and Multi-view Features for Image Re-rankingabstractImage re-ranking is effective in improving performance of text-based image searches. However, improvements from existing re-ranking algorithms are limited by two factors: one is that the associated textual information of images often mismatches their actual visual contents; the other is that a visual's features cannot accurately describe the semantic similarities between images. In this paper, we adopt click data to bridge the semantic gap. We propose a novel multi-view hypergraph-based learning (MHL) method that adaptively integrates click data with varied visual features. In particular, MHL considers pairwise discriminative constraints from click data to maximally distinguish images with high click counts from images with no click counts, and a semantic manifold is constructed. It then adopts hypergraph learning to build multiple manifolds from varied visual features. Finally, MHL integrates the semantic manifold with visual manifolds through an iterative optimization procedure. The weights of different manifolds and the re-ranking score are simultaneously obtained after using this optimization strategy. We conduct experiments on real world datasets and the results demonstrate that MHL outperforms state-of-the-art image re-ranking methods. Jun Yu 0002, Yong Rui |
IEEE Trans. Multim. | 1 |
| 2013 | Image-Based 3D Human Pose Recovery with Locality Sensitive Sparse RetrievalabstractImage-based 3D human pose recovery is usually conducted by retrieving relevant poses with image features. However, it suffers from high dimensionality of image features and low efficiency of retrieving process. In this paper, we propose a novel approach to recover 3D human poses from silhouettes. This approach improves traditional methods by adopting locality sensitive sparse coding in the retrieving process. It incorporates a local similarity preserving term into the objective of sparse coding, which groups similar silhouettes to alleviate the instability of sparse codes. The experimental results demonstrate the effectiveness of the proposed method. Jun Yu 0002 |
SMC | 2 |
| 2013 | Multi-view hypergraph learning by patch alignment framework
Jun Yu 0002, Jonathan Li 0001 |
Neurocomputing | 2 |
| 2013 | Skeleton correspondence construction and its applications in animation style reusing
Zhijun Song, Jun Yu 0002, Changle Zhou, Dapeng Tao |
Neurocomputing | 2 |
| 2013 | Automatic cartoon matching in computer-assisted animation production
Zhijun Song, Jun Yu 0002, Changle Zhou, Meng Wang 0001 |
Neurocomputing | 2 |
| 2013 | High-level attributes modeling for indoor scenes classification
Chaojie Wang 0003, Jun Yu 0002, Dapeng Tao |
Neurocomputing | 2 |
| 2013 | Pairwise constraints based multiview features fusion for scene classification
Jun Yu 0002, Dacheng Tao, Yong Rui, Jun Cheng 0002 |
Pattern Recognit. | 1 |
| 2013 | Cartoon features selection using Diffusion Score
Jun Yu 0002 |
Signal Process. | 1 |
| 2012 | Transductive Cartoon Retrieval by Multiple Hypergraph Learning
Jun Yu 0002, Jun Cheng 0002, Jianmin Wang 0009, Dacheng Tao |
ICONIP (3) | 1 |
| 2012 | Graph based transductive learning for cartoon correspondence construction
Jun Yu 0002, Wei Bian 0003, Mingli Song, Jun Cheng 0002, Dacheng Tao |
Neurocomputing | 1 |
| 2012 | Semi-supervised distance metric learning based on local linear regression for data clustering
Jun Yu 0002, Meng Wang 0001, Yun Liu 0021 |
Neurocomputing | 2 |
| 2012 | Image classification by multimodal subspace learning
Jun Yu 0002, Feng Lin 0002, Seah Hock Soon, Cuihua Li, Ziyu Lin |
Pattern Recognit. Lett. | 1 |
| 2012 | Interactive cartoon reusing by transfer learning
Jun Yu 0002, Jun Cheng 0002, Dacheng Tao |
Signal Process. | 1 |
| 2012 | Adaptive Hypergraph Learning and its Application in Image ClassificationabstractRecent years have witnessed a surge of interest in graph-based transductive image classification. Existing simple graph-based transductive learning methods only model the pairwise relationship of images, however, and they are sensitive to the radius parameter used in similarity calculation. Hypergraph learning has been investigated to solve both difficulties. It models the high-order relationship of samples by using a hyperedge to link multiple samples. Nevertheless, the existing hypergraph learning methods face two problems, i.e., how to generate hyperedges and how to handle a large set of hyperedges. This paper proposes an adaptive hypergraph learning method for transductive image classification. In our method, we generate hyperedges by linking images and their nearest neighbors. By varying the size of the neighborhood, we are able to generate a set of hyperedges for each image and its visual neighbors. Our method simultaneously learns the labels of unlabeled images and the weights of hyperedges. In this way, we can automatically modulate the effects of different hyperedges. Thorough empirical studies show the effectiveness of our approach when compared with representative baselines. Jun Yu 0002, Dacheng Tao, Meng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | Semisupervised Multiview Distance Metric Learning for Cartoon SynthesisabstractIn image processing, cartoon character classification, retrieval, and synthesis are critical, so that cartoonists can effectively and efficiently make cartoons by reusing existing cartoon data. To successfully achieve these tasks, it is essential to extract visual features that comprehensively represent cartoon characters and to construct an accurate distance metric to precisely measure the dissimilarities between cartoon characters. In this paper, we introduce three visual features, color histogram, shape context, and skeleton, to characterize the color, shape, and action, respectively, of a cartoon character. These three features are complementary to each other, and each feature set is regarded as a single view. However, it is improper to concatenate these three features into a long vector, because they have different physical properties, and simply concatenating them into a high-dimensional feature vector will suffer from the so-called curse of dimensionality. Hence, we propose a semisupervised multiview distance metric learning (SSM-DML). SSM-DML learns the multiview distance metrics from multiple feature sets and from the labels of unlabeled cartoon characters simultaneously, under the umbrella of graph-based semisupervised learning. SSM-DML discovers complementary characteristics of different feature sets through an alternating optimization-based iterative algorithm. Therefore, SSM-DML can simultaneously accomplish cartoon character classification and dissimilarity measurement. On the basis of SSM-DML, we develop a novel system that composes the modules of multiview cartoon character classification, multiview graph-based cartoon synthesis, and multiview retrieval-based cartoon synthesis. Experimental evaluations based on the three modules suggest the effectiveness of SSM-DML in cartoon applications. Jun Yu 0002, Meng Wang 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2012 | On Combining Multiple Features for Cartoon Character Retrieval and Clip SynthesisabstractHow do we retrieve cartoon characters accurately? Or how to synthesize new cartoon clips smoothly and efficiently from the cartoon library? Both questions are important for animators and cartoon enthusiasts to design and create new cartoons by utilizing existing cartoon materials. The first key issue to answer those questions is to find a proper representation that describes the cartoon character effectively. In this paper, we consider multiple features from different views, i.e., color histogram, Hausdorff edge feature, and skeleton feature, to represent cartoon characters with different colors, shapes, and gestures. Each visual feature reflects a unique characteristic of a cartoon character, and they are complementary to each other for retrieval and synthesis. However, how to combine the three visual features is the second key issue of our application. By simply concatenating them into a long vector, it will end up with the so-called "curse of dimensionality," let alone their heterogeneity embedded in different visual feature spaces. Here, we introduce a semisupervised multiview subspace learning (semi-MSL) algorithm, to encode different features in a unified space. Specifically, under the patch alignment framework, semi-MSL uses the discriminative information from labeled cartoon characters in the construction of local patches where the manifold structure revealed by unlabeled cartoon characters is utilized to capture the geometric distribution. The experimental evaluations based on both cartoon character retrieval and clip synthesis demonstrate the effectiveness of the proposed method for cartoon application. Moreover, additional results of content-based image retrieval on benchmark data suggest the generality of semi-MSL for other applications. Jun Yu 0002, Dongquan Liu, Dacheng Tao, Seah Hock Soon |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2011 | Stroke Correspondence Construction Using Manifold LearningabstractAbstract Stroke correspondence construction is a precondition for generating inbetween frames from a set of key frames. In our case, each stroke in a key frame is a vector represented as a Disk B‐Spline Curve (DBSC) which is a flexible and compact vector format. However, it is not easy to construct correspondences between multiple DBSC strokes effectively because of the following points: (1) with the use of shape descriptors, the dimensionality of the feature space is high; (2) the number of strokes in different key frames is usually large and different from each other and (3) the length of corresponding strokes can be very different. The first point makes matching difficult. The other two points imply ‘many to many’ and ‘part to whole’ correspondences between strokes. To solve these problems, this paper presents a DBSC stroke correspondence construction approach, which introduces a manifold learning technique to the matching process. Moreover, in order to handle the mapping between unequal numbers of strokes with different lengths, a stroke reconstruction algorithm is developed to convert the ‘many to many’ and ‘part to whole’ stroke correspondences to ‘one to one’ compound stroke correspondence. Dongquan Liu, Jun Yu 0002, Huiqin Gu, Dacheng Tao, Seah Hock Soon |
Comput. Graph. Forum | 3 |
| 2011 | Fuzzy Diffusion Distance Learning for Cartoon Similarity Estimation
Jun Yu 0002, Seah Hock Soon |
J. Comput. Sci. Technol. | 1 |
| 2011 | Semi-automatic cartoon generation by motion planning
Jun Yu 0002, Dacheng Tao, Meng Wang 0001, Jun Cheng 0002 |
Multim. Syst. | 1 |
| 2011 | Cartoon synthesis using constrained spreading activation network
Jun Yu 0002, Seah Hock Soon, Yueting Zhuang |
Multim. Tools Appl. | 1 |
| 2011 | Complex Object Correspondence Construction in Two-Dimensional AnimationabstractCorrespondence construction of objects in key frames is the precondition for inbetweening and coloring in 2-D computer-assisted animation production. Since each frame of an animation consists of multiple layers, objects are complex in terms of shape and structure. Therefore, existing shape-matching algorithms specifically designed for simple structures such as a single closed contour cannot perform well on objects constructed by multiple contours with an open shape. This paper introduces a semisupervised patch alignment framework for complex object correspondence construction. In particular, the new framework constructs local patches for each point on an object and aligns these patches in a new feature space, in which correspondences between objects can be detected by the subsequent clustering. For local patch construction, pairwise constraints, which indicate the corresponding points (must link) or unfitting points (cannot link), are introduced by users to improve the performance of correspondence construction. This kind of input is convenient for animation software users via user-friendly interfaces. A dozen of experimental results on our cartoon data set that is built on industrial production suggest the effectiveness of the proposed framework for constructing correspondences of complex objects. As an extension of our framework, additional shape retrieval experiments on MPEG-7 data set show that its performance is comparable with that of a prominent algorithm published in T-PAMI 2009. Jun Yu 0002, Dongquan Liu, Dacheng Tao, Seah Hock Soon |
IEEE Trans. Image Process. | 1 |
| 2010 | Transductive graph based cartoon synthesisabstractAbstract To reduce tedious works in cartoon animation, some computer‐assisted systems including automatic inbetweening and cartoon reusing systems have been proposed. In existing automatic inbetweening systems, accurate correspondence construction, which is a prerequisite for inbetweening, cannot be achieved. For cartoon reusing systems, the lack of efficient similarity estimation method and reusing mechanism makes it impractical for the users. The Transductive Graph based Cartoon Synthesis (TGCS) approach proposed in this paper aims at synthesizing smooth cartoons from the existing data. In this approach, the similarity between cartoon frames can be accurately evaluated by calculating the distance based on local shape context, which is rotation, and scaling invariant. According to the similarity, the label propagation based graph transduction method is adopted to generate cartoon clips, which is smoother than the clips generated by the shortest path method used in previous cartoon reusing approaches. Besides, the synthesized cartoon clips can be applied in accurate correspondence building, based on which the inbetweening method can be used to refine the results. Experimental results on our cartoon dataset suggest the effectiveness of the proposed approach for cartoon synthesis. Additional experiments on correspondence show our approach's performance on accurate correspondence building. Copyright © 2010 John Wiley & Sons, Ltd. Jun Yu 0002, Dongquan Liu, Seah Hock Soon |
Comput. Animat. Virtual Worlds | 1 |
| 2010 | Recognizing Cartoon Image Gestures for Retrieval and Interactive Cartoon Clip SynthesisabstractIn this paper, we propose a new method to recognize gestures of cartoon images with two practical applications, i.e., content-based cartoon image retrieval and interactive cartoon clip synthesis. Upon analyzing the unique properties of four types of features including global color histogram, local color histogram (LCH), edge feature (EF), and motion direction feature (MDF), we propose to employ different features for different purposes and in various phases. We use EF to define a graph and then refine its local structure by LCH. Based on this graph, we adopt a transductive learning algorithm to construct local patches for each cartoon image. A spectral method is then proposed to optimize the local structure of each patch and then align these patches globally. MDF is fused with EF and LCH and a cartoon gesture space is constructed for cartoon image gesture recognition. We apply the proposed method to content-based cartoon image retrieval and interactive cartoon clip synthesis. The experiments demonstrate the effectiveness of our method. Yi Yang 0001, Yueting Zhuang, Dacheng Tao, Dong Xu 0001, Jun Yu 0002, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2008 | Perspective-aware cartoon clips synthesisabstractAbstract In this paper we propose an approach, which allows the users to synthesize cartoon clips according to the perspective of the background image. In order to construct the cartoons smoothly, the character's edge distance and motion direction distance are demonstrated to be the factors affecting the human perception in similarity evaluation, and utilized in cartoon clips synthesis. When applying the generated cartoons to the background image, in which the perspective exists, the size of the character is coordinated according to the scaling factor calculated from the vanishing line. The experiment results demonstrate that our approach can synthesize the cartoon clips more smoothly compared with other single frame reusing strategies. The generated cartoons, which are applied to the background image, can be accepted by the human perception well. Copyright © 2008 John Wiley & Sons, Ltd. Yueting Zhuang, Jun Yu 0002, Jun Xiao 0001, Cheng Chen 0023 |
Comput. Animat. Virtual Worlds | 2 |
| 2007 | Adaptive control in cartoon data reusingabstractAbstract In this paper, we propose a novel approach, which reuses Traditional Chinese Cartoon to create new animations. In order to extract the cartoon character precisely, a segmentation method based on edge detection is implemented. Before reusing the data, a lower‐ dimensional space of the cartoon data is constructed by ISOmap. The character's gesture difference calculated by optical flow is combined with character's edge difference through a novel distance function, which is controlled by a weight parameter. The animation is created by reordering the existing data into a sequence, which is the shortest path between two designated data in the space. Our approach utilizes image processing, computer vision, and machine learning in cartoon creation and the experiment results demonstrate that the animation's quality can be effectively improved by the fusion of these techniques. Copyright © 2007 John Wiley & Sons, Ltd. Jun Yu 0002, Yueting Zhuang, Jun Xiao 0001, Cheng Chen 0023 |
Comput. Animat. Virtual Worlds | 1 |
| 2001 | Spatiotemporal segmentation for compact video representation
Jianping Fan 0001, Jun Yu 0002, Gen Fujita, Takao Onoye, Lide Wu, Isao Shirakawa |
Signal Process. Image Commun. | 2 |