VLDB 2026 Research / reviewers in the wild / expert
Jun Yu 0001
dblp:50/5754-1
· DBLP profile ↗
155ranked-venue papers
46as first author
101since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 101 · 40 first-author · 56 since 2021Artificial intelligence and machine learning · 55 · 6 first-author · 45 since 2021Computer networks · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image ClassificationabstractThe booming remote sensing (RS) technology is giving rise to a novel multimodality generalization task, which requires the model to overcome data heterogeneity while possessing powerful cross-scene generalization ability. Moreover, most vision-language models usually describe surface materials using universal texts, lacking proprietary linguistic prior knowledge specific to different RS modalities. In this work, we formalize RS multimodality generalization (RSMG) as a learning paradigm, and propose a frequency-aware vision-language multimodality generalization network (FVMGN) for RS image classification. Specifically, a diffusion-based training-test-time augmentation (DTAug) strategy is designed to reconstruct multimodal land-cover distributions, enriching input information for FVMGN. Following that, to overcome multimodal heterogeneity, a multimodal wavelet disentanglement (MWDis) module is developed to learn cross-domain invariant features by resampling low and high frequency components in the frequency domain. Considering the characteristics of RS vision modalities, shared and proprietary class texts is designed as linguistic inputs for the transformer-based text encoder to extract diverse text features. For multimodal vision inputs, a spatial-frequency-aware image encoder (SFIE) is constructed to realize local-global feature reconstruction and representation. Finally, a multiscale spatial-frequency feature alignment (MSFFA) module is suggested to construct a unified semantic space, ensuring refined multiscale alignment of different text and vision features in spatial and frequency domains. Extensive experiments show that FVMGN has the excellent multimodality generalization ability compared with state-of-the-art methods. Junjie Zhang 0011, Feng Zhao 0005, Hanqiang Liu 0001, Jun Yu 0001 |
AAAI | 4 |
| 2026 | MetaDB: Metadata-Guided Diffusion Bridge Model for High-Fidelity Medical Image SynthesisabstractMedical image synthesis is pivotal in modern clinical workflows, addressing the issue of missing imaging modalities. While diffusion-based models have shown promise, existing approaches often neglect the rich clinical metadata, leading to synthesized images that lack semantic fidelity and fail to maintain strict consistency with the target modality. To address these challenges, we propose a metadata-guided diffusion bridge model, termed MetaDB, a novel framework that leverages textual clinical priors to steer the source-to-target translation process. Our method introduces two key innovations to ensure high-fidelity synthesis. First, we design a text-guided adaptive normalization layer, which dynamically modulates the feature statistics of the diffusion backbone using encoded clinical metadata. This mechanism explicitly aligns the synthesized features with the target modality’s attributes, ensuring semantic consistency throughout the generation process. Second, to prevent semantic degradation during the iterative denoising steps, we propose a semantics reconstruction network. This auxiliary module imposes a constraint that forces the network to preserve deep semantic representations, further reinforcing the semantic consistency between the generated output and the target description. Extensive experiments on multiple medical imaging datasets demonstrate that our approach achieves state-of-the-art performance in terms of quantitative metrics and visual quality, generating images that are both anatomically accurate and semantically faithful to clinical protocols. Yanjun Chi, Jiaen Liang, Jun Yu 0001 |
ICMR | 6 |
| 2026 | SAM-MPA: A SAM-Based Motion Perception and Aggregation Framework for Referring Video SegmentationabstractReferring video object segmentation relies on natural language descriptions to identify and segment target objects in videos, and has achieved substantial progress in recent years. However, most prior studies process videos in a frame-by-frame manner, failing to fully exploit temporal information. Recently, the large-scale segmentation model Segment Anything Model (SAM) has attracted considerable attention due to its strong segmentation capability and impressive zero-shot generalization. Nevertheless, SAM still exhibits limitations when handling complex action-oriented descriptions. Motivated by these observations, we propose a novel SAM-based Motion-Perception Aggregation framework for referring video object segmentation, termed SAM-MPA, which consists of four modules. DINO-SAM leverages the powerful segmentation ability of SAM to perform initial video segmentation guided by textual prompts, generating object-level masks. The Kalman Filtering Motion Modeling module injects explicit object motion modeling into DINO-SAM, improving segmentation robustness under occlusion and fast-motion scenarios. The motion-aware aggregation module effectively captures object action cues at multiple temporal scales, thereby enhancing global video understanding. The text-token matching module further enforces semantic consistency between the segmentation results and the referring expressions. Extensive experiments on challenging RVOS benchmarks demonstrate that SAM-MPA provides a competitive and efficient SAM-based solution for motion-centric referring video object segmentation, while offering a favorable trade-off between performance and computational cost compared with conventional non-MLLM baselines. The code is available at https://github.com/GXU-LIPE/SAM-MPA. Fang Gao 0001, Ao Lu, Qingbao Huang, Jun Yu 0001 |
IEEE Internet Things J. | 5 |
| 2026 | Strength-Adaptive Adversarial TrainingabstractAdversarial training (AT) has been shown to effectively enhance a network's resilience against adversarial attack. However, conventional AT, which relies on a fixed pre-specified perturbation budget, suffers from several limitations when training robust models. First, enforcing the same perturbation budget across networks with different capacities leads to varying levels of robustness disparity between natural and robust accuracies, which deviates from the desired outcome of a robust network. Second, because the perturbation budget is fixed throughout training, the attack strength fails to scale adaptively with the evolving robustness of the model. This mismatch often results in robust overfitting and further degradation of adversarial robustness. To address these limitations, we propose a novel technique called Strength-Adaptive Adversarial Training (SAAT). In SAAT, the adversary incorporates an adversarial-loss constraint to guide the generation of adversarial training data. This constraint allows the perturbation budget to adapt dynamically based on the current training state, which effectively mitigates robust overfitting. Moreover, by explicitly regulating the attack strength through the adversarial loss, SAAT enables precise control over the robustness disparity between natural accuracy and adversarial robustness. Extensive experiments demonstrate that SAAT substantially improves adversarial robustness over standard AT. Chaojian Yu, Dawei Zhou 0004, Li Shen 0008, Jun Yu 0001, Bo Han 0003, Mingming Gong, Nannan Wang 0001, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Bidirectional cross-modal image-guided point cloud completion with multi-scale progressive refinement
Fang Gao 0001, Yan Jin 0012, Shaodong Li, Hanbo Zheng, Jun Yu 0001 |
Pattern Recognit. | 6 |
| 2026 | Lafa: Unlocking Superior Memory Efficiency via Adaptive Metadata Strategy for Scalable Large-Scale Dataset LoadingabstractThe rapid growth of deep learning models and the increasing demand for large-scale datasets have posed unprece dented challenges for data loading and memory management. Existing frameworks (e.g., PyTorch, TensorFlow) often encounter performance bottlenecks when handling large datasets resulting in inefficiencies and excessive memory usage. To address these issues, we propose Lafa, a dynamic metadata loading mechanism optimized for efficient large-scale dataset processing. Lafa introduces the .Lafa format and an adaptive loading strategy with three modes to balance memory usage and loading performance, along with a local shuffle approach that reduces memory overhead and computational complexity while preserving data randomness. Experimental results on GPU (RTX 3090) and Ascend (910A) platforms demonstrate that Lafa significantly improves memory efficiency compared to existing frameworks. Specifically, for every 10 million samples loaded, Lafa reduces additional memory consumption by a factor of 1.33× to 31.34× across various dataset types, relative to the most memory-efficient baseline among PyTorch, TensorFlow, and MindSpore. Cong Wang 0039, Yang Luo 0006, Ke Wang 0065, Hui Zhang 0044, Naijie Gu, Wenzhuo Du, Fan Yu 0004, Jun Yu 0001 |
IEEE Trans. Big Data | 9 |
| 2026 | Language-Assisted Reconstruction for Self-Supervised Category-Level 6D Object Pose Estimation With Coarse-to-Fine Correspondence OptimizationabstractSelf-supervised category-level 6D pose estimation has emerged as a task of paramount significance within the field of computer vision. Despite recent advancements, current self-supervised methods grapple with two critical challenges. Primarily, the ability of existing networks to accurately reconstruct object models is constrained by pronounced part-level shape variations across specific categories. Additionally, the persistent many-to-one ambiguity within pixel-to-point cloud correspondences poses a significant barrier to achieving robust performance. To address these challenges, we propose a novel approach that includes a Language-Assisted Memory-Encoding Shape Reconstruction (LMR) module and a Coarse-to-Fine Correspondence Optimization (CFCO) module. In the LMR module, language descriptions are leveraged to bridge the gap between virtual and real images, thereby improving the alignment between learned representations and real-world object appearances. Additionally, a memory encoding mechanism is introduced to enhance reconstruction accuracy by capturing fine-grained shape variations. The CFCO module utilizes Hungarian matching to generate one-to-one pseudo labels at both region and pixel levels, providing explicit supervision for the corresponding similarity matrices. Furthermore, this process helps alleviate the many-to-one ambiguity to some extent, leading to more accurate correspondence learning. We evaluate our method on the REAL275 and WILD6D datasets. Extensive experiments demonstrate that our self-supervised approach outperforms existing methods and achieves new state-of-the-art results within the self-supervised framework. Jun Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Generative Information-Guided Heterogeneous Cross-Fusion Network With Contrastive Learning for Multimodal Remote Sensing Image ClassificationabstractMultimodal remote sensing (RS) images exhibit distinct structure and distribution characteristics, making it challenging to design an effective multimodal RS image classification algorithm. Moreover, although existing deep learning-based methods have become the darling in the multimodal RS image classification, they usually lack effective exploration and explicit integration for generative information from different modalities. Aiming at the above challenges, a generative information-guided heterogeneous cross-fusion network with contrastive learning (GIHCN) is proposed for multimodal RS image classification. Firstly, to simulate the land-cover distributions from different modal data, a multimodal generative information learning architecture (MGILA) is constructed to capture the unsupervised heterogeneous distribution features. Secondly, to achieve bidirectional modeling between heterogeneous data and the reconstructed land-cover distributions, a heterogeneous data & generative information cross-attention module (HGCM) is designed to explore the complementarity between multimodal data and the reconstructed land-cover distributions. HGCM can provide the heterogeneous generative information for current modal data or provide the heterogeneous data support for current modal generative information, thereby obtaining cross-fusion sources with different attributes. Furthermore, we achieve the effective feature extraction for different cross-fusion sources by a designed multimodal contrastive learning framework (MCLF). Notably, to capture local information and long-range dependencies, a hybrid classification network with convolutional neural network and Mamba (CMNet) is proposed as the feature extraction backbone of each cross-fusion source to further improve the classification performance. Finally, we construct a joint multimodality loss function for MCLF, which can reduce the distribution difference between modalities while focusing on the information flow within and across the modality. Experimental results on four multimodal RS datasets confirm the effectiveness of GIHCN compared with other state-of-the-art methods. The source code will be released at https://github.com/ZJier/GIHCN. Junjie Zhang 0011, Feng Zhao 0005, Hanqiang Liu 0001, Jun Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Guided Self-Attention: Find the Generalized Necessarily Distinct Vectors for Grain Size GradingabstractWith the development of steel materials, metallographic analysis has become increasingly important. Unfortunately, grain size analysis is a manual process that requires experts to evaluate metallographic photographs, which is unreliable and time-consuming. To resolve this problem, we propose a novel classification method based on deep learning, namely, GSNets, a family of hybrid models, which can effectively introduce guided self-attention for classifying grain size. Concretely, we build our models from three insights: 1) Introducing our novel guided self-attention module can assist the model in finding the generalized necessarily distinct vectors capable of retaining intricate relational connections and rich local feature information; 2) By improving the pixelwise linear independence of the feature map, the highly condensed semantic representation will be captured by the model; 3) Our novel triple-stream merging module can significantly improve the generalization capability and efficiency of the model. Experiments show that our GSNet yields a classification accuracy of 90.1%, surpassing the state-of-the-art Swin Transformer V2 by 1.9% on the steel grain size dataset, which comprises 3599 images with 14 grain size levels. Furthermore, we intuitively believe our approach is applicable to broader applications such as object detection and semantic segmentation. Fang Gao 0001, XueTao Li, Jiabao Wang 0004, Shengheng Ma, Jun Yu 0001 |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2026 | Image Diffusion Models With Multimodal Conditional Control in Zero-Shot Semantic Segmentation SynthesisabstractIn the era of booming large models, the significance of data scale in deep learning has been well-acknowledged. Nevertheless, obtaining large-scale datasets remains a challenging task. Large pre-trained diffusion models, boasting remarkable generative capabilities, offer a promising solution for dataset generation. This study zeroes in on the labor-intensive semantic segmentation task, where image annotation is a major bottleneck. We introduce a novel data generation pipeline leveraging multimodal conditional control. This pipeline not only enhances the universality and practicality of our method but also ensures high-precision alignment between generated images and corresponding masks. Specifically, we feed mask, Canny edge, and depth information into ControlNet and employ captions generated by the BLIP model as prompts to precisely control image generation. Notably, we pioneer the exploration of zero-shot generation in this context. This approach enables direct image generation without the need for fine-tuning or alignment with segmentation protocols. Experimental results demonstrate its effectiveness: in the CityScapes dataset, zero-shot generation leads to a 0.6% increase in the mean Intersection over Union (mIOU). Even without zero-shot generation, our method achieves significant mIOU improvements of up to 2.53% on the ADE20K dataset and 1.93% on the COCO-Stuff dataset. These findings highlight the potential of our proposed method in advancing semantic segmentation tasks. Leilei Wang, Renjie Lu 0001, Fengzhao Sun, Jun Yu 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Data-Free Knowledge Distillation with Diffusion ModelsabstractRecently Data-Free Knowledge Distillation (DFKD) has garnered attention and can transfer knowledge from a teacher neural network to a student neural network without requiring any access to training data. Although diffusion models are adept at synthesizing high-fidelity photorealistic images across various domains, existing methods cannot be easiliy implemented to DFKD. To bridge that gap, this paper proposes a novel approach based on diffusion models, DiffDFKD. Specifically, DiffDFKD involves targeted optimizations in two key areas. Firstly, DiffDFKD utilizes valuable information from teacher models to guide the pre-trained diffusion models’ data synthesis, generating datasets that mirror the training data distribution and effectively bridge domain gaps. Secondly, to reduce computational burdens, DiffDFKD introduces Latent CutMix Augmentation, an efficient technique, to enhance the diversity of diffusion model-generated images for DFKD while preserving key attributes for effective knowledge transfer. Extensive experiments validate the efficacy of DiffDFKD, yielding state-of-the-art results exceeding existing DFKD approaches.We release our code at https://github.com/xhqi0109/DiffDFKD. Xiaohua Qi, Renda Li, Qiang Ling 0001, Jun Yu 0001, Ziyi Chen 0005, Peng Chang 0002, Jing Xiao 0006 |
ICME | 5 |
| 2025 | Optimization of Multimodal Inputs Based on Diffusion Models: Zero-Shot Semantic Image GenerationabstractWith the continuous advancement of large models, the scale of data has become increasingly important in semantic segmentation tasks. However, the complexity and high cost of annotating semantic segmentation data pose significant challenges to the expansion of datasets. This study aims to leverage pre-trained diffusion generative models for conditional image generation, where labeled masks are used to generate corresponding synthetic images. This ensures a direct correspondence between input and output, effectively bypassing the annotation stage to reduce the cost of labor-intensive tasks. We employ multimodal conditions, to control the generation results. Additionally, we propose a multimodal alignment scheme to optimize the input control conditions, thereby improving the spatial structural accuracy of the generated results. Furthermore, we explore zero-shot generation tasks and successfully achieve zero-shot generation performance across multiple datasets, demonstrating the effectiveness of our approach. In the zero-shot generation experiments on the Cityscapes dataset, our method achieved a 0.6% improvement in the mIoU evaluation metric. On the ADE20K dataset, the performance improvement reached 2.52%, while on the COCO-Stuff dataset, the improvement was 2.43%. Leilei Wang, Renjie Lu 0001, Fengzhao Sun, Jun Yu 0001, Jianqing Sun, Jiaen Liang |
ICME | 5 |
| 2025 | Dualdiff: Dual-Branch Diffusion Model for Autonomous Driving with Semantic FusionabstractAccurate and high-fidelity driving scene reconstruction relies on fully leveraging scene information as conditioning. However, existing approaches, which primarily use 3D bounding boxes and binary maps for foreground and background control, fall short in capturing the complexity of the scene and integrating multi-modal information. In this paper, we propose DualDiff, a dual-branch conditional diffusion model designed to enhance multi-view driving scene generation. We introduce Occupancy Ray Sampling (ORS), a semantic-rich 3D representation, alongside numerical driving scene representation, for comprehensive foreground and background control. To improve cross-modal information integration, we propose a Semantic Fusion Attention (SFA) mechanism that aligns and fuses features across modalities. Furthermore, we design a foreground-aware masked (FGM) loss to enhance the generation of tiny objects. DualDiff achieves state-of-the-art performance in FID score, as well as consistently better results in downstream BEV segmentation and 3D object detection tasks. Haoteng Li, Zezhong Qian, Gongpeng Zhao, Jun Yu 0001, Huazheng Zhou, Longjun Liu |
ICRA | 6 |
| 2025 | Towards Robust Autonomous Driving: Conditional Multimodal Large Language Models for Fine-Grained PerceptionabstractMultimodal large language models (MLLMs) have shown remarkable performance across various visual understanding tasks. However, most existing MLLMs still lack image detail perception, limiting their effectiveness in tasks that require detailed visual information. In this paper, we introduce Percept-DriveLM, a novel MLLM designed to tackle the fine-grained perception challenges in autonomous driving tasks. At the core of our model is the Visual Fusion Module, which integrates several innovative components: a dynamic resolution mechanism that combines both high and low resolution features, and an RoI conditional mechanism to incorporate object/region-level features identified by offline detectors, further refining the model's fine-grained perception abilities. Trained in a two-stage process, our model demonstrates exceptional performance, outperforming existing MLLMs with comparable parameter sizes and excelling in both autonomous driving perception and general vision-language tasks. The effectiveness of our approach is validated through extensive empirical studies. Code will be available at https://github.com/DebuggerSunfz/PerceptDriveLM. Fengzhao Sun, Jun Yu 0001, Jiaming Hou, Xilong Lu, Heng Song, Fang Gao 0001 |
ICRA | 2 |
| 2025 | Logic Consistency Makes Large Language Models Personalized Reasoning TeachersabstractLarge Language Models (LLMs) have advanced natural language processing, particularly through Chain-of-Thought (CoT) reasoning, but their high computational costs limit deployment. We propose Personalized Chain-of-Thought Distillation (PeCoTD), a method that transfers CoT reasoning from LLMs to smaller models by addressing the distribution gap—the difference in how large and small models process information. To bridge this gap, PeCoTD introduces the Self Logic Consistency (SLC) metric, which helps small models evaluate and select LLM-generated rationales that align better with their reasoning abilities. PeCoTD iteratively refines these rationales, adjusting them to better fit the learning patterns of small models while preserving their original meaning. Experiments show PeCoTD significantly enhances the reasoning abilities of small models across datasets, making CoT distillation more practical and effective. Xulong Zhang 0001, Yong Zhang 0058, Jun Yu 0001, Jianzong Wang |
IJCNN | 4 |
| 2025 | LVLM-MIR: Large Vision-Language Model with Parameter-Efficient Fine-Tuning for Multimodal Interleaved ReasoningabstractMultimodal interleaved reasoning, which requires models to understand interleaved image-text sequences and multiple images, is a critical challenge in contemporary AI. This paper proposes a parameter-efficient fine-tuning framework based on Large Vision-Language Models, with Qwen2.5-VL as the backbone and Low-Rank Adaptation for task-specific adaptation. The framework integrates four stages: multimodal input preprocessing to align with pre-training distributions, visual feature extraction via a modified Vision Transformer, cross-modal fusion via attention mechanisms, and response generation via an autoregressive decoder. By freezing pre-trained weights and fine-tuning low-rank adapters in both visual and language modules, it balances preserving general multimodal knowledge with optimizing target tasks, achieving high performance with low computational overhead. On the MIRAGE Challenge Track A Dataset, it performs strongly across subtasks, achieving an aggregate score of 0.7857 and securing second place in the challenge. Ablation studies confirm that joint LoRA fine-tuning of visual and language modules yields optimal results; limitations in fine-grained visual difference tasks indicate future directions in enhancing subtle feature capture and adaptive cross-modal alignment. Jun Yu 0001, Xilong Lu, Cong Wang 0039, Qiang Ling 0001 |
ACM Multimedia | 1 |
| 2025 | CMA-VC: Large Vision-Language Model for Cross-Modal Alignment in Intention-Oriented Video CaptioningabstractTraditional video captioning methods often produce generic descriptions that fail to align with specific user intentions, limiting their applicability in scenarios requiring customized information extraction. This paper proposes a novel intention-oriented controllable video captioning approach, which leverages large-scale vision-language models (InternViT and InternLM) and achieves parameter-efficient fine-tuning through Low-Rank Adaptation (LoRA). The proposed framework processes both video content and user-specified intentions via a unified cross-modal pipeline, dynamically aligning visual features with intent semantics to generate focused and contextually accurate captions. Experiments on the IntentVC dataset validate the effectiveness of the proposed method in generating intention-aligned captions, with the following performance metrics: BLEU@4 scores 44.38 on the public test set and 40.21 on the private test set; METEOR scores 63.79 and 60.07 respectively; CIDEr scores 230.33 and 208.15 respectively; ROUGE-L scores 61.75 and 57.14 respectively. Ablation studies confirm the significant effectiveness of joint LoRA adaptation on vision and language modules, as well as the sensitivity of performance to LoRA parameters. This work advances the field by enabling precise control over caption generation, enhancing the practical utility of video understanding systems in applications such as accessibility services and targeted video retrieval. Jun Yu 0001, Xilong Lu, Qiang Ling 0001 |
ACM Multimedia | 1 |
| 2025 | LVLM-HBA: Large Vision-Language Model with Cross-Modal Alignment for Human Behavior AnalysisabstractBodily Behaviour Recognition (BBR) and Eye Contact Detection (ECD) in multi-person group conversations are critical for understanding social dynamics, but traditional methods often rely solely on visual cues, lacking integration with semantic context. To address this, we propose a novel framework based on Large Vision-Language Models (LVLMs), leveraging their cross-modal alignment capability to fuse visual features (e.g., body postures, gaze directions) and linguistic semantics (e.g., behavioral category descriptions). A parameter-efficient tuning strategy using Low-Rank Adaptation (LoRA) is adopted, adapting only a subset of parameters in both the Language Model (LM) and Vision Transformer (ViT) modules, thus retaining pre-trained knowledge while reducing computational costs. The framework incorporates multi-task output heads to simultaneously predict BBR and ECD results. Experiments on the MPIIGroupInteraction dataset demonstrate superior performance: our method achieves 0.65 accuracy on BBR and 0.82 accuracy on ECD, outperforming state-of-the-art approaches by 0.02-0.03 in absolute terms. Ablation studies validate that applying LoRA to both LM and ViT with optimal hyperparameters yields the best results, confirming the importance of cross-modal synergy. This work highlights the potential of LVLMs in social behavior analysis, providing a lightweight and effective solution for understanding complex group interactions. Jun Yu 0001, Xilong Lu, Lingsi Zhu, Qiang Ling 0001 |
ACM Multimedia | 1 |
| 2025 | Hierarchical Multi-Feature Extraction and Aggregation for Micro-Action RecognitionabstractMicro-action refers to subtle, low-intensity non-verbal behaviors that can provide insights into an individual's underlying emotions and intentions. Due to its brief duration and significant overlap, identifying these micro-actions poses a challenge for current models. In response to these challenges, this paper proposes a novel multi-feature fusion framework, which extracts coarse-grained body features and fine-grained action features separately. Specifically, we present Temporal Contextualization for fine-grained learning, a cross-frame injection mechanism designed to capture essential spatio-temporal information and introduce a 3D-ResNet Adapter for coarse-grained learning, which aggregates temporal data and facilitates parameter-efficient fine-tuning. In consideration of the task dataset distribution's long-tail nature, the implementation of Feature Decoupling is undertaken, adopting a two-stage training strategy. By conducting experiments, the aforementioned hierarchical multi-feature extraction and aggregation approach has been demonstrated to yield substantial enhancement in Micro-Action Recognition. Our method attains an F1-mean score of 77.75% on the MA-52 dataset, ranking 1st in the 2nd Micro-Action Analysis Grand Challenge in Conjunction with ACM MM'25. Zhichao Xia, Yanjun Chi, Lingsi Zhu, Mohan Jing, Jun Yu 0001 |
ACM Multimedia | 6 |
| 2025 | Unified Dual-Strategy Framework for Multi-Task Visual Question AnsweringabstractWe present a unified vision-language framework for the Responsible Multimodal AI Challenge 2025, tackling two related tasks: multimodal hallucination detection and factuality verification. Our approach builds on a Dynamic ViLBERT model enhanced with adaptive multi-branch attention fusion to effectively integrate visual and textual data. For Task A (Hallucination Detection), we employ four parallel answer-encoding branches with co-attentional transformers to compare AI-generated captions or answers against corresponding images, enabling accurate hallucination identification. For Task B (Factuality Verification), we fuse visual and textual features via element-wise addition, multiplication, and concatenation, feeding them into a classifier to assess claim validity. Inputs include visual features from Detectron2 and text embeddings from BERT tokenization. Experiments show notable accuracy gains over baselines, validated by official F1 scores. Extensive ablations, comparative studies, and architectural visualizations confirm the value of cross-modal attention and customized fusion strategies. Our final system ranks 3rd in hallucination detection (F1 = 0.80) and 2nd in factuality verification (F1 = 0.84), demonstrating the strength of our unified approach. Shuoping Yang, Jun Yu 0001 |
ACM Multimedia | 2 |
| 2025 | HierMEQA: A Relationship-Aware Hierarchical Framework for Consistent Micro-Expression Visual Question AnsweringabstractThe rise of Multimodal Large Language Models (MLLMs) offers new opportunities for Micro-Expression (ME) analysis. This paper introduces Micro-Expression Visual Question Answering (ME-VQA), a novel task reformulating ME annotations (e.g., emotion categories, action units) into QA pairs. To address key challenges-hardware limitations, context inconsistency, and compositional reasoning gaps-we propose a Relationship-Aware Hierarchical VQA Framework. Our approach leverages mined emotion correlations (e.g., coarse-to-fine label dependencies) and employs a two-stage process: 1) Coarse-grained anchoring for broad emotion categories, and 2) Fine-grained reasoning constrained by coarse outputs and statistical rules. We further optimize efficiency via a dual-phase video sampling strategy: during training, keyframes (onset/apex/offset) and random non-expression frames are used; uniform sampling is applied at inference. Experiments demonstrate significant improvements in answer consistency and accuracy. Lingsi Zhu, Yanjun Chi, Jun Yu 0001, Gongpeng Zhao, Yuefeng Zou, Fengzhao Sun, Xilong Lu |
ACM Multimedia | 3 |
| 2025 | Heterogeneous Encoder Fusion with KAN Decoder for Group Engagement Modeling via 8× Sliding PipelinesabstractEstimating engagement in group interactions is crucial for building socially intelligent systems, such as in human-agent and human-robot interaction. However, precisely modeling the continuous frame-level fluctuations of engagement remains challenging, particularly when considering the complex multi-party signal interactions within groups. Our method employs an encoder that integrates BiLSTM with Transformer to effectively capture both local and global temporal dependencies of multimodal features. Crucially, we explicitly fuse signals from both the target participant and their conversational partners in group to model the holistic group interaction dynamics. Furthermore, we introduce an 8x overlapped optimized sliding window strategy, constructing a ''sliding pipeline'', which significantly enhances the temporal smoothness, continuity, and stability of predictions. In the final regression stage, we replace the traditional multilayer perceptron(MLP) decoder with Kolmogorov-Arnold Network (KAN), leveraging their superior function approximation capability to achieve more accurate engagement predictions. Evaluated on the test sets of NoXi-base, NoXi-addition, MPIIGroupInteraction and NoXi-J datasets from the Multimediate'25 Engagement Challenge, our approach demonstrates significant performance improvements, achieving highly competitive Concordance Correlation Coefficients (CCC) of 0.678 for global, approximately 56.1% higher than the baseline, which shows a significant improvement. Yuefeng Zou, Hui Zhang 0044, Jun Yu 0001, Keda Lu, Lingsi Zhu, Fengzhao Sun, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 3 |
| 2025 | Breaking barriers in 3D point cloud data processing: A unified system for efficient storage and high-throughput loading
Cong Wang 0039, Yang Luo 0006, Ke Wang 0065, Yanfei Cao, Xiangzhi Tao, Dongjie Geng, Naijie Gu, Jun Yu 0001, Fan Yu 0004, Zhengdong Wang, Shouyang Dong |
Expert Syst. Appl. | 8 |
| 2025 | Winning Prize Comes from Losing Tickets: Improve Invariant Learning by Exploring Variant Parameters for Out-of-Distribution Generalization
Li Shen 0008, Jun Yu 0001, Chen Gong 0002, Bo Han 0003, Tongliang Liu |
Int. J. Comput. Vis. | 4 |
| 2025 | TVTracker: Target-Adaptive Text-Guided Visual Fusion for Multimodal RGB-T TrackingabstractCurrent multi-modal sensor trackers mainly rely on visual cues for target tracking. However, in challenging scenarios, the visual information acquired by multi-modal sensors has limited descriptive ability when the target state undergoes significant changes, which may lead to unsatisfactory tracker performance. In this work, we propose TVTracker, a two-stage tracking framework. It leverages semantic information of the target state to enhance visual cues, enabling effective RGB-T tracking. The first stage is the generation of target text descriptions. Utilizing the Bootstrapping Language-Image Pre-training (BLIP) model, we generate target textual descriptions that match the images in the dataset. In the second stage, text-guided visual fusion is performed for target tracking. Target textual and visual features are extracted separately using the text and visual branches. Then the target textual features are integrated with the visual features to localize the target position and predict the target bounding box. In the text branch, we design the Target Text Adaptation Enhancement (TTAE) module to mitigate the interference of low-quality target textual features on visual features. In the visual branch, we develop the multi-modal visual information prompters, which include the Multi-modal Visual Shared Information Prompter (MVSIP) and the Multi-modal Visual Shared and Complementary Information Prompter (MVSCIP), to facilitate learning of multi-modal shared or complementary visual prompts. Experiments on the LasHeR, RGBT210, RGBT234, and VTUAV datasets demonstrate the effectiveness of TVTracker. Fang Gao 0001, Yan Jin 0012, Jingfeng Tang, Hanbo Zheng, Shengheng Ma, Jun Yu 0001 |
IEEE Internet Things J. | 7 |
| 2025 | Faster and Stronger: Unleashing Data Processing Potential Through Hardware HeterogeneityabstractWith the rapid advancement of AI technology, there has been a substantial surge in the need for computational resources. Particularly in deep learning, machine learning, and large-scale data analysis, the processing of extensive datasets necessitates exceptionally high levels of computational efficacy and speed. Conventional homogeneous computing platforms, predominantly reliant on Central Processing Units (CPU), have encountered challenges in meeting the escalating demands for high-performance computing. Consequently, this study advocates for heterogeneous hardware acceleration technology, strategically migrating data operations from CPU to varied hardware components (e.g. GPU, NPU) to enhance processing efficiency and computational performance during the data preprocessing phase. We conducted experiments to evaluate the impact of utilizing hardware heterogeneous acceleration technologies on data processing speed under various workloads and system hardware configurations. By adjusting parameters like batch size and CPU utilization rates, we compared the performance of frameworks that support hardware heterogeneity with popular deep learning frameworks (e.g. PyTorch and TensorFlow) across various hardware configurations and neural network models. Empirical findings demonstrate that the system framework optimized through heterogeneous hardware acceleration technology (the preprocessing speed is improved in all the given experimental environment tests) exhibits commendable universality and superiority in performance. Codes are available at https://github.com/mindspore-ai/mindspore. Cong Wang 0039, Yang Luo 0006, Wenzhuo Du, Ke Wang 0065, Naijie Gu, Jun Yu 0001 |
IEEE Internet Things J. | 6 |
| 2025 | Joint Optic Disc and Cup Segmentation Via KNN-Based Transformer and Deformable Aggregation AttentionabstractDeep learning has advanced medical image segmentation, especially for optic disc and cup detection. While convolutional neural networks (CNNs) struggle with long-range dependencies, Transformer-based architectures have emerged to address this limitation.However, complete replacement of CNNs with Transformers may impair local feature extraction. Additionally, reliance on expert-annotated datasets makes supervised learning costly. To overcome these limitations, we present DAK-Former, a novel self-supervised contrastive learning approach for optic disc and cup segmentation. Our approach introduces a new attention mechanism that combines K-nearest neighbors (KNN) with Transformer components, along with a deformable aggregation attention module to improve global feature representation. Additionally, our method uses multiple types of medical imaging data, including MRI, CT scans, and X-rays. Using contrastive learning, we match encoded queries with a dictionary of encoded keys, allowing the network to learn meaningful unsupervised feature representations. We evaluated DAK-Former on two publicly available fundus image datasets and compared it with state-of-the-art methods. Experimental results show that DAK-Former is highly effective and consistently outperforms existing approaches. Jun Yu 0001, Shuoping Yang, Gongpeng Zhao, Lei Wang 0203 |
IEEE Internet Things J. | 1 |
| 2025 | Fine-grained hierarchical dynamics for image harmonization
Peng He 0004, Jun Yu 0001, Liuxue Ju, Fang Gao 0001 |
Neural Networks | 2 |
| 2025 | Improving the Instance-Dependent Transition Matrix Estimation by Exploiting Self-Supervised LearningabstractThe transition matrix reveals the transition relationship between clean labels and noisy labels. It plays an important role in building statistically consistent classifiers for learning with noisy labels. However, in real-world applications, the transition matrix is usually unknown and has to be estimated. It is a challenging task to accurately estimate the transition matrix which usually depends on the instance. With both instances and noisy labels at hand, the major difficulty of estimating the transition matrix comes from the absence of clean label information. Recent work suggests that self-supervised learning methods can effectively infer clean label information. These methods could even achieve comparable performance with supervised learning on many benchmark datasets but without requiring any labels. Motivated by this, our paper presents a practical approach that harnesses self-supervised learning to extract clean label information, which reduces the estimation error of the instance-dependent transition matrix. By exploiting the estimated transition matrix, the performance of classifiers is improved. Empirical results on different datasets illustrate that our proposed methodology outperforms existing state-of-the-art methods in terms of both classification accuracy and transition matrix estimation. Yexiong Lin, Yu Yao 0005, Zhaoqing Wang, Xu Shen 0001, Jun Yu 0001, Bo Han 0003, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Visual and Textual Commonsense-Enhanced Layout Learning for Vision-and-Language NavigationabstractIn the Vision-and-Language Navigation (VLN) task, an agent must comprehend natural language instructions and execute precise navigation in complex environments. While significant progress has been made in the VLN field, the limited availability of navigation data hinders existing methods from fully learning the commonsense relationships between rooms and landmarks, which are crucial for environmental understanding and successful navigation. To address this issue, this work proposes a Visual and Textual Commonsense-Enhanced Layout Learning Model (ViTeC). We leverage the open-world knowledge embedded in large models by utilizing ChatGPT and BLIP-2 to provide commonsense information about environments. Specifically, BLIP-2 analyzes the room type corresponding to each panoramic image, while ChatGPT infers and provides knowledge about the most common landmarks within each room type. Moreover, to compensate for the agent’s lack of commonsense at the visual level, we employ Stable Diffusion to generate commonsense-based visual images, enhancing the agent’s visual perception. To ensure the agent effectively learns commonsense about the environment, we designed a Text Commonsense Layout Learning Module and a Visual Commonsense Layout Learning Module. These modules help the agent acquire environmental commonsense from both linguistic and visual perspectives, enabling it to utilize commonsense information effectively during navigation, thereby improving its environmental understanding and reasoning capabilities. Experimental results demonstrate that ViTeC achieves strong performance on REVERIE, R2R, and SOON datasets, exhibiting good generalization ability in complex environments. This validates the effectiveness of ViTeC in enhancing the agent’s environmental understanding and navigation capabilities. Fang Gao 0001, Jingfeng Tang, Jiabao Wang 0004, Shaodong Li, Shengheng Ma, Jun Yu 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | Contrastive Learning With Multiple Prototypes for Unsupervised Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptive semantic segmentation aims to transfer knowledge from the annotated source domain to the unlabeled target domain. Recently, self-training methods have gained substantial attention, which leverage high-confidence predictions in the target domain as pseudo labels for supervision. However, limited exploration of intra-class variations across domains, including significant visual differences within each category, has led to misalignment between feature distribution across domains. In this article, we present a unified non-parametric distance-based online clustering method to efficiently maintain multiple centroid-based prototypes within each category subspace instead of one prototype for each category subspace, which enables prototypes to possess the capacity for richer feature representation. Then, considering the variance across different dimensions of a feature representation, we then extend the prototypes from centroid-based ones to distribution-based ones. Specifically, each subspace is modeled using a Gaussian mixture model which includes several anisotropic Gaussian distributions, aimed at prioritizing discriminative dimensions and obtaining a finer measurement of the pixel-to-prototype similarity. Meanwhile, a category-aware feature space is achieved through pixel-to-prototype contrastive learning to ensure the compactness of pixel features in the same subcategory and drive the separation between pixel features of different subcategories. What's more, multi-resolution features are utilized to promote diversity and robustness among intra-class prototypes. Experiments validate the competitiveness of our two prototype-based methods against existing state-of-the-art methods, with a mIoU of 76.8% on GTA$\rightarrow$Cityscapes, 68.4% on Synthia$\rightarrow$Cityscapes, 54.5% on Cityscapes$\rightarrow$DarkZurich and 56.4% on Cityscapes$\rightarrow$ACDC. Notably, our method is able to seamlessly integrate with existing UDA methods. Jun Yu 0001, Guochen Xie, Quansheng Liu, Zhen Kan, Lei Wang 0203, Qiang Ling 0001, Fang Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Domain-Separated Bottleneck Attention Fusion Framework for Multimodal Emotion RecognitionabstractAs a focal point of research in various fields, human body language understanding has long been a subject of intense interest. Within this realm, the exploration of emotion recognition through the analysis of facial expressions, voice patterns, and physiological signals holds significant practical value. Compared with unimodal approaches, multimodal emotion recognition models leverage complementary information from vision, acoustic, and language modalities to robust perceive the human sentiment attitudes. However, the heterogeneity among modality signals leads to significant domain shifts, posing challenges for achieving balanced fusion. In this article, we propose a Domain-Separated Bottleneck Attention (DBA) Fusion Framework for human multimodal emotion recognition with lower computational complexity. Specifically, we partition each modality into two distinct domains: the invariant/private domain. The invariant domain contains crucial shared information, while the private domain aims to capture modality-specific representations. For the decomposed features, we introduce two sets of bottleneck cross-attention modules to effectively utilize the complementarity between domains to reduce redundant information. In each module, we interweave two Fusion Adapter blocks into the Self-Attention Transformer backbone. Each Fusion Adapter block integrates a small group of latent tokens as bridges for inter-modal and inter-domain interactions, mitigating the adverse effects of modality distribution differences and lowering computational costs. Extensive experimental results demonstrate that our method outperforms State-of-the-Art (SOTA) approaches across three widely used benchmark datasets. Peng He 0004, Jun Yu 0001, Chengjie Ge, Lei Wang 0203, Zhen Kan |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | To-Former: semantic segmentation of transparent object with edge-enhanced transformer
Wen Su 0004, Mengjiao Ge, Jun Yu 0001 |
Vis. Comput. | 5 |
| 2024 | LMT-GP: Combined Latent Mean-Teacher and Gaussian Process for Semi-supervised Low-Light Image Enhancement
Fengxin Chen, Jun Yu 0001, Zhen Kan |
ECCV (83) | 3 |
| 2024 | Image Harmonization Based on Hierarchical DynamicsabstractImage harmonization is an essential technique in computer vision, aiming to generate visually consistent composite images by making the foreground compatible with the background. However, current methods primarily focus on applying a global transformation perspective, overlooking the fact that different regions in a real image can exhibit significant appearance variations. Yet, there is consistency within local regions. They also have limited representation ability by using fixed background statistics (e.g., mean, and standard deviation) for foreground normalization. Hence, we propose a hierarchical dynamics appearance translation strategy that adjusts the foreground appearance based on the corresponding background, adapting the model features and parameters from local to global view. To enhance the representation ability for targets, we employ a mixed attention mechanism for local dynamics, which adaptively modifies the features of different channels and positions. Additionally, we apply dynamic region-aware convolution guided by the foreground mask for global dynamics, which learns the adaptive representation of the foreground and background and correlations to global harmonization. To further improve the harmonization result, we integrate adversarial and perceptual loss into the model training. Experiments show our method significantly reduces parameters and achieves state-of-the-art performance compared with previous methods. Liuxue Ju, Chengdao Pu, Jun Yu 0001, Wen Su 0004 |
ICASSP | 3 |
| 2024 | EmoTalker: Emotionally Editable Talking Face Generation via Diffusion ModelabstractIn recent years, the field of talking faces generation has attracted considerable attention, with certain methods adept at generating virtual faces that convincingly imitate human expressions. However, existing methods face challenges related to limited generalization, particularly when dealing with challenging identities. Furthermore, methods for editing expressions are often confined to a singular emotion, failing to adapt to intricate emotions. To overcome these challenges, this paper proposes EmoTalker, an emotionally editable portraits animation approach based on the diffusion model. EmoTalker modifies the denoising process to ensure preservation of the original portrait’s identity during inference. To enhance emotion comprehension from text input, Emotion Intensity Block is introduced to analyze fine-grained emotions and strengths derived from prompts. Additionally, a crafted dataset is harnessed to enhance emotion comprehension within prompts. Experiments show the effectiveness of EmoTalker in generating high-quality, emotionally customizable facial expressions. Xulong Zhang 0001, Ning Cheng 0001, Jun Yu 0001, Jing Xiao 0006, Jianzong Wang |
ICASSP | 4 |
| 2024 | Rotated R-CNN: A Two-Stage Object Detection Method Adapted To Oriented Bounding BoxesabstractCurrently, oriented object detection, as an emerging subfield within object detection, has garnered significant attention. Besides encompassing directional information, datasets of oriented objects exhibit notable characteristics, including significant variations in object scales and a wide range of aspect ratios for ground-truth bounding boxes. Nevertheless, the current state-of-the-art two-stage rotating object detection models have not sufficiently addressed these characteristics, leading to inherent limitations in accuracy. In response to these challenges, we introduce the Rotated RCNN. Our model is the first to introduce trainable anchors in the field of oriented object detection to achieve anchor distributions similar to the ground truth boxes in the oriented object dataset. Furthermore, considering the distinctive traits of oriented ground truth boxes, we have devised a novel strategy for assigning labels to more effectively choose positive and negative samples specifically designed for oriented objects. In the regression phase of the RPN, we introduce shape constraints to alleviate accuracy losses stemming from mismatches between the encoding method and oriented objects. We comprehensively evaluate our model on the DOTAv1.0 and HRSC2016 datasets, demonstrating the effectiveness of our meticulously designed model. Chengdao Pu, Jun Yu 0001, Wen Su 0004 |
ICIP | 2 |
| 2024 | Neural Auto-designer for Enhanced Quantum KernelsabstractQuantum kernels hold great promise for offering computational advantages over classical learners, with the effectiveness of these kernels closely tied to the design of the feature map. However, the challenge of designing effective quantum feature maps for real-world datasets, particularly in the absence of sufficient prior information, remains a significant obstacle. In this study, we present a data-driven approach that automates the design of problem-specific quantum feature maps. Our approach leverages feature-selection techniques to handle high-dimensional data on near-term quantum machines with limited qubits, and incorporates a deep neural predictor to efficiently evaluate the performance of various candidate quantum kernels. Through extensive numerical simulations on different datasets, we demonstrate the superiority of our proposal over prior methods, especially for the capability of eliminating the kernel concentration issue and identifying the feature map with prediction advantages. Our work not only unlocks the potential of quantum kernels for enhancing real-world tasks, but also highlights the substantial role of deep learning in advancing quantum machine learning. Cong Lei, Peng Mi, Jun Yu 0001, Tongliang Liu |
ICLR | 4 |
| 2024 | Towards Realistic Model Selection for Semi-supervised LearningabstractSemi-supervised Learning (SSL) has shown remarkable success in applications with limited supervision. However, due to the scarcity of labels in the training process, SSL algorithms are known to be impaired by the lack of proper model selection, as splitting a validation set will further reduce the limited labeled data, and the size of the validation set could be too small to provide a reliable indication to the generalization error. Therefore, we seek alternatives that do not rely on validation data to probe the generalization performance of SSL models. Specifically, we find that the distinct margin distribution in SSL can be effectively utilized in conjunction with the model’s spectral complexity, to provide a non-vacuous indication of the generalization error. Built upon this, we propose a novel model selection method, specifically tailored for SSL, known as Spectral-normalized Labeled-margin Minimization (SLAM). We prove that the model selected by SLAM has upper-bounded differences w.r.t. the best model within the search space. In addition, comprehensive experiments showcase that SLAM can achieve significant improvements compared to its counterparts, verifying its efficacy from both theoretical and empirical standpoints. Xiaobo Xia, Runze Wu 0001, Fengming Huang, Jun Yu 0001, Bo Han 0003, Tongliang Liu |
ICML | 5 |
| 2024 | Mitigating Label Noise on Graphs via Topological Sample SelectionabstractDespite the success of the carefully-annotated benchmarks, the effectiveness of existing graph neural networks (GNNs) can be considerably impaired in practice when the real-world graph data is noisily labeled. Previous explorations in sample selection have been demonstrated as an effective way for robust learning with noisy labels, however, the conventional studies focus on i.i.d data, and when moving to non-iid graph data and GNNs, two notable challenges remain: (1) nodes located near topological class boundaries are very informative for classification but cannot be successfully distinguished by the heuristic sample selection. (2) there is no available measure that considers the graph topological information to promote sample selection in a graph. To address this dilemma, we propose a $\textit{Topological Sample Selection}$ (TSS) method that boosts the informative sample selection process in a graph by utilising topological information. We theoretically prove that our procedure minimizes an upper bound of the expected risk under target clean distribution, and experimentally show the superiority of our method compared with state-of-the-art baselines. Jiangchao Yao, Xiaobo Xia, Jun Yu 0001, Ruxin Wang 0002, Bo Han 0003, Tongliang Liu |
ICML | 4 |
| 2024 | Dialogue Cross-Enhanced Central Engagement Attention Model for Real-Time Engagement Estimation
Jun Yu 0001, Keda Lu, Ji Zhao 0020, Zhihong Wei, Iek-Heng Chu, Peng Chang 0002 |
IJCAI | 1 |
| 2024 | Boosting fairness for 3D face reconstructionabstractWith the increasing significance of 3D face reconstruction technology in various domains, including the metaverse, immersive communication, and medical cosmetology, the precise recovery of geometric shapes from 2D images, regardless of age, gender, or ethnicity, is essential. Recent attention to fairness concerns in 3D face reconstruction has primarily focused on skin color issues, such as albedo estimation, with little consideration for racial bias in facial geometric reconstruction. To address this gap, we first surveyed the most recent 3D face reconstruction methods and commonly used 3D face datasets, confirming the existence of racial bias in the accuracy of 3D face reconstruction. We then developed a fair multilevel 3D face reconstruction system by using data resampling and an asymmetric arc loss, which combines arc face loss and circle loss. Our experimental results show that the system achieves more accurate results on the REALY benchmark. Zeyu Cui, Jun Yu 0001 |
IJCNN | 2 |
| 2024 | Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
Haoxiang Shi, Ziqi Liang, Jun Yu 0001 |
INTERSPEECH | 3 |
| 2024 | Micro-Expression Spotting Based on Optical Flow Feature with Boundary Calibration
Jun Yu 0001, Gongpeng Zhao, Peng He 0004, Zhongpeng Cai, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 1 |
| 2024 | A Method for Visual Spatial Description Based on Large Language Model Fine-tuning
Jiabao Wang 0004, Fang Gao 0001, Jingfeng Tang, Shaodong Li, Hanbo Zheng, Shengheng Ma, Feng Shuang 0002, Jun Yu 0001 |
ACM Multimedia | 8 |
| 2024 | Building Robust Video-Level Deepfake Detection via Audio-Visual Local-Global InteractionsabstractThe continual advancements in Generative Artificial Intelligence have created substantial hurdles for accurate deepfake detection, leading to limitations of currently popular detection methods across content-driven video-level deepfake detection scenarios. In this paper, we present the solutions to the Video-Level Deepfake Detection task. Our empirical findings demonstrate that modeling correlations of audio-visual modalities is important for video-level deepfake detection. Therefore, we introduce the model denoted Audio-Visual Local-Global Neural Network (i.e., AV-LGNN) in which the core design is the proposed AV-LGI Module (Audio-Visual Local-Global Interaction Module). The AV-LGI Module is composed of three stages: Local Intra-Region Interaction, Global Inter-Region Interaction, and Local-Global Interaction, which can better capture detailed information at local-level and efficiently learn the fine-grained correlations of inter-modalities in video deepfake detection under lower computational overheads. We further propose an adaptive modality selection strategy to facilitate model learning. Besides, a variety of data augmentation techniques are incorporated for audio-visual branches to enhance the robustness of the AV-LGNN. The experimental results verify the effectiveness of our model. Jia Zhang 0016, Mohan Jing, Keda Lu, Jun Yu 0001, Wen Su 0004, Fang Gao 0001, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 6 |
| 2024 | End-to-end Spatio-Temporal Information Aggregation For Micro-Action DetectionabstractMicro-actions convey the emotions of characters in daily communication and offer richer semantic information compared to conventional actions. Accurate detection of these micro-actions is essential for video understanding. Due to their short duration, low intensity, and high overlap, micro-actions require more detailed video features, presenting a significant challenge for accurate detection. To address these challenges, we propose the 3D-SENet Adapter, which aggregates spatio-temporal information and enables end-to-end online video feature learning. We also find that incorporating background information significantly enhances the detection of small-scale micro-actions. Thus we develop the Cross-Attention Aggregation Detection Head, which integrates multi-scale features within the feature pyramid, thereby improving the detection accuracy of micro-actions occupying small regions in video frames. Our approach achieves first place in the Multi-label Micro-Action Detection (MMAD) and second place in the Micro-Action Recognition (MAR) of Micro-Action Analysis Grand Challenge. Jun Yu 0001, Mohan Jing, Guopeng Zhao, Keda Lu, Feng Zhao 0005, Jiaqing Sun, Jiaen Liang |
ACM Multimedia | 1 |
| 2024 | Temporal-Informative Adapters in VideoMAE V2 and Multi-Scale Feature Fusion for Micro-Expression Spotting-then-Recognize
Jun Yu 0001, Gongpeng Zhao, Peng He 0004, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 1 |
| 2024 | RAG-Guided Large Language Models for Visual Spatial Description with Adaptive Hallucination CorrectorabstractVisual Spatial Description (VSD) is an emerging image-to-text task which aims at generating descriptions of the spatial relationships between given objects in an image. In this paper, we apply Retrieval-Augmented Generation (RAG) technology in guiding Multimodal Large Language Models (MLLMs) for the task of VSD, complemented by an Adaptive Hallucination Corrector, and further fine-tuning them to bolster semantic understanding and overall model efficacy. We found that our approach demonstrated higher accuracy and fewer hallucination errors in both spatial relationship classification and visual language description tasks within the VSD task, achieving state-of-the-art results. Jun Yu 0001, Gongpeng Zhao, Fengzhao Sun, Fanrui Zhang, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 1 |
| 2024 | Part-level Reconstruction for Self-Supervised Category-level 6D Object Pose Estimation with Coarse-to-Fine Correspondence OptimizationabstractSelf-supervised category-level 6D pose estimation stands as a fundamental task in computer vision. However, current self-supervised methods face two major challenges. Firstly, existing networks struggle to reconstruct precise object models due to significant part-level shape variations among specific categories. Secondly, they are impacted by the many-to-one ambiguity in the correspondences between pixels and point clouds. To address these challenges, we propose a novel approach that includes a Part-level Shape Reconstruction (PSR) module and a Coarse-to-Fine Correspondence Optimization (CFCO) module. In the (PSR) module, we introduce a part-level discrete shape memory to capture more fine-grained shape variations of different objects and use it to perform precise reconstruction. In the (CFCO) module, we utilize Hungarian matching to generate one-to-one pseudo labels at both region and pixel levels, which provides explicit supervision for the corresponding similarity matrices. We evaluate our method on the REAL275 and WILD6D datasets. Our extensive experiments show that our self-supervised approach outperforms existing methods and achieves new state-of-the-art results within the self-supervised framework. Jun Yu 0001, Liangxian Cui, Qiang Ling 0001 |
ACM Multimedia | 2 |
| 2024 | Pin-CasNet: Detecting pin status in transmission lines based on cascade network
Fang Gao 0001, Rongwei Zhang, Jingfeng Tang, Shaomin Liu, Jun Yu 0001, Chang Wen Chen, Hanbo Zheng |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | Data and knowledge-driven deep multiview fusion network based on diffusion model for hyperspectral image classification
Junjie Zhang 0011, Feng Zhao 0005, Hanqiang Liu 0001, Jun Yu 0001 |
Expert Syst. Appl. | 4 |
| 2024 | ProtoSimi: label correction for fine-grained visual categorizationabstractAbstract Deep models trained by using clean data have achieved tremendous success in fine-grained image classification. Yet, they generally suffer from significant performance degradation when encountering noisy labels. Existing approaches to handle label noise, though proved to be effective for generic object recognition, usually fail on fine-grained data. The reason is that, on fine-grained data, the category difference is subtle and the training sample size is small. Then deep models could easily overfit the noisy labels. To improve the robustness of deep models on noisy data for fine-grained visual categorization, in this paper, we propose a novel learning framework named ProtoSimi. Our method employs an adaptive label correction strategy, ensuring effective learning on limited data. Specifically, our approach considers the criteria of exploring the effectiveness of both global class-prototype and part class-prototype similarities in identifying and correcting labels of samples. We evaluate our method on three standard benchmarks of fine-grained recognition. Experimental results show that our method outperforms the existing label noisy methods by a large margin. In ablation studies, we also verify that our method is non-sensitive to hyper-parameters selection and can be integrated with other FGVC methods to increase the generalization performance. Jialiang Shen, Yu Yao 0005, Shaoli Huang, Zhiyong Wang 0001, Jing Zhang 0037, Ruxing Wang, Jun Yu 0001, Tongliang Liu |
Mach. Learn. | 7 |
| 2024 | Tackling Noisy Labels With Network Parameter Additive DecompositionabstractGiven data with noisy labels, over-parameterized deep networks suffer overfitting mislabeled data, resulting in poor generalization. The memorization effect of deep networks shows that although the networks have the ability to memorize all noisy data, they would first memorize clean training data, and then gradually memorize mislabeled training data. A simple and effective method that exploits the memorization effect to combat noisy labels is early stopping. However, early stopping cannot distinguish the memorization of clean data and mislabeled data, resulting in the network still inevitably overfitting mislabeled data in the early training stage. In this paper, to decouple the memorization of clean data and mislabeled data, and further reduce the side effect of mislabeled data, we perform additive decomposition on network parameters. Namely, all parameters are additively decomposed into two groups, i.e., parameters w are decomposed as w=σ+γ. Afterward, the parameters σ are considered to memorize clean data, while the parameters γ are considered to memorize mislabeled data. Benefiting from the memorization effect, the updates of the parameters σ are encouraged to fully memorize clean data in early training, and then discouraged with the increase of training epochs to reduce interference of mislabeled data. The updates of the parameters γ are the opposite. In testing, only the parameters σ are employed to enhance generalization. Extensive experiments on both simulated and real-world benchmarks confirm the superior performance of our method. Xiaobo Xia, Long Lan, Xinghao Wu, Jun Yu 0001, Wenjing Yang 0002, Bo Han 0003, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | A Time-Consistency Curriculum for Learning From Instance-Dependent Noisy LabelsabstractMany machine learning algorithms are known to be fragile on simple instance-independent noisy labels. However, noisy labels in real-world data are more devastating since they are produced by more complicated mechanisms in an instance-dependent manner. In this paper, we target this practical challenge of Instance-Dependent Noisy Labels by jointly training (1) a model reversely engineering the noise generating mechanism, which produces an instance-dependent mapping between the clean label posterior and the observed noisy label and (2) a robust classifier that produces clean label posteriors. Compared to previous methods, the former model is novel and enables end-to-end learning of the latter directly from noisy labels. An extensive empirical study indicates that the time-consistency of data is critical to the success of training both models and motivates us to develop a curriculum selecting training data based on their dynamics on the two models' outputs over the course of training. We show that the curriculum-selected data provide both clean labels and high-quality input-output pairs for training the two models. Therefore, it leads to promising and robust classification performance even in notably challenging settings of instance-dependent noisy labels where many SoTA methods could easily fail. Extensive experimental comparisons and ablation studies further demonstrate the advantages and significance of the time-consistency curriculum in learning from instance-dependent noisy labels on multiple benchmark datasets. Songhua Wu, Tianyi Zhou 0001, Jun Yu 0001, Bo Han 0003, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Regularly Truncated M-Estimators for Learning With Noisy LabelsabstractThe sample selection approach is very popular in learning with noisy labels. As deep networks "learn pattern first", prior methods built on sample selection share a similar training procedure: the small-loss examples can be regarded as clean examples and used for helping generalization, while the large-loss examples are treated as mislabeled ones and excluded from network parameter updates. However, such a procedure is arguably debatable from two folds: (a) it does not consider the bad influence of noisy labels in selected small-loss examples; (b) it does not make good use of the discarded large-loss examples, which may be clean or have meaningful information for generalization. In this paper, we propose regularly truncated M-estimators (RTME) to address the above two issues simultaneously. Specifically, RTME can alternately switch modes between truncated M-estimators and original M-estimators. The former can adaptively select small-losses examples without knowing the noise rate and reduce the side-effects of noisy labels in them. The latter makes the possibly clean examples but with large losses involved to help generalization. Theoretically, we demonstrate that our strategies are label-noise-tolerant. Empirically, comprehensive experimental results show that our method can outperform multiple baselines and is robust to broad noise types and levels. Xiaobo Xia, Pengqian Lu, Chen Gong 0002, Bo Han 0003, Jun Yu 0001, Jun Yu 0002, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | A complex neural network model by Hilbert TransformabstractThe phase information of the optical wave plays a vital role in processing wave-related signals. In deep learning fields, complex-valued neural networks are put forward on the concept of complex amplitudes for full utilization of phase information. To build a complex-valued neural network, the common way is to exploit Fourier Transform of the observed signal to extract amplitude and phase information. However, this will lead to spectrum waste for a real-valued signal by introducing negative frequencies that have no physical meaning. To this end, we attempt to use Hilbert Transform as an alternative to yield a single sideband spectrum and avoid negative frequencies from interacting with positive ones. On the other hand, Fourier transform is a global analysis thus it tells nothing about the time domain. As our key insight, we further explore the usage of instantaneous frequency calculated by Hilbert Transform and propose a new method of constructing complex input from a time–frequency angle. Simple pixel-wise classification experiments are carried out on two hyperspectral datasets and MNIST dataset. Experimental results have demonstrated that Hilbert Transform with instantaneous frequency performs better by a large margin than Fourier Transform owing to the additional time information. Xinzhi Liu, Jun Yu 0001, Toru Kurihara, Congzhong Wu, Shu Zhan |
Pattern Recognit. Lett. | 2 |
| 2024 | Conditional Consistency Regularization for Semi-Supervised Multi-Label Image ClassificationabstractConsistency regularization has achieved great successes in Semi-Supervised Single-Label Image Classification (SS-SLC) with deep learning models, while few effort has been devoted to Semi-Supervised Multi-Label Image Classification (SS-MLC) with deep learning models. One intuitive solution for introducing consistency regularization to SS-MLC is to regularize model predictions to be invariant to different augmented data of the same input image. However, the solution lacks the consideration of label relations, which are key elements in multi-label image classification. In this article, we go beyond the consistency regularization for multi-view input images, and propose Conditional Consistency Regularization (CCR) that is tailored for SS-MLC. Specifically, for two augmented input images, we make the two model predictions conditioned on different label states (i.e., positive, negative, or unknown for each class). By encouraging the two predictions to be consistent, the model is able to build relations between the given two different label states, which helps to make use of label relations for boosting image classification. The experiments on large-scale real-world SS-MLC benchmarks demonstrate that the proposed method can surpass state-of-the-art methods by a large margin. Zhengning Wu, Tianyu He, Xiaobo Xia, Jun Yu 0001, Xu Shen 0001, Tongliang Liu |
IEEE Trans. Multim. | 4 |
| 2024 | Language-Guided Dual-Modal Local Correspondence for Single Object TrackingabstractThis paper focuses on the advancement of single-object tracking technologies in computer vision, which have broad applications including robotic vision, video surveillance, and sports video analysis. Current methods relying solely on the target's initial visual information encounter performance bottlenecks and limited applications, due to the scarcity of target semantics in appearance features and the continuous change in the target's appearance. To address these issues, we propose a novel approach, combining visual-language dual-modal single-object tracking, that leverages natural language descriptions to enrich the semantic information of the moving target. We introduce a dual-modal single-object tracking algorithm based on local correspondence modeling. The algorithm decomposes visual features into multiple local visual semantic features and pairs them with local language features extracted from natural language descriptions. In addition, we also propose a new global relocalization method that utilizes visual language bimodal information to perceive target disappearance and misalignment and adaptively reposition the target in the entire image. This improves the tracker's ability to adapt to changes in target appearance over long periods of time, enabling long-term single target tracking based on bimodal semantic and motion information. Experimental results show that our model outperforms state-of-the-art methods, which demonstrates the effectiveness and efficiency of our approach. Jun Yu 0001, Zhongpeng Cai, Lei Wang 0203, Fang Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Multimodal Visual-Semantic Representations Learning for Scene Text RecognitionabstractScene Text Recognition (STR), the critical step in OCR systems, has attracted much attention in computer vision. Recent research on modeling textual semantics with Language Model (LM) has witnessed remarkable progress. However, LM only optimizes the joint probability of the estimated characters generated from the Vision Model (VM) in a single language modality, ignoring the visual-semantic relations in different modalities. Thus, LM-based methods can hardly generalize well to some challenging conditions, in which the text has weak or multiple semantics, arbitrary shape, and so on. To migrate the above issue, in this paper, we propose Multimodal Visual-Semantic Representations Learning for Text Recognition Network (MVSTRN) to reason and combine the multimodal visual-semantic information for accurate Scene Text Recognition. Specifically, our MVSTRN builds a bridge between vision and language through its unified architecture and has the ability to reason visual semantics by guiding the network to reconstruct the original image from the latent text representation, breaking the structural gap between vision and language. Finally, the tailored multimodal Fusion (MMF) module is motivated to combine the multimodal visual and textual semantics from VM and LM to make the final predictions. Extensive experiments demonstrate our MVSTRN achieves state-of-the-art performance on several benchmarks. Xinjian Gao, Ye Pang, Yuyu Liu, Maokun Han, Jun Yu 0001, Wei Wang 0496, Yuanxu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Robust Generalization Against Photon-Limited Corruptions via Worst-Case Sharpness MinimizationabstractRobust generalization aims to tackle the most challenging data distributions which are rare in the training set and contain severe noises, i.e., photon-limited corruptions. Common solutions such as distributionally robust optimization (DRO) focus on the worst-case empirical risk to ensure low training error on the uncommon noisy distributions. However, due to the over-parameterized model being optimized on scarce worst-case data, DRO fails to produce a smooth loss landscape, thus struggling on generalizing well to the test set. Therefore, instead of focusing on the worst-case risk minimization, we propose SharpDRO by penalizing the sharpness of the worst-case distribution, which measures the loss changes around the neighbor of learning parameters. Through worst-case sharpness minimization, the proposed method successfully produces a flat loss curve on the corrupted distributions, thus achieving robust generalization. Moreover, by considering whether the distribution annotation is available, we apply SharpDRO to two problem settings and design a worst-case selection process for robust generalization. Theoretically, we show that SharpDRO has a great convergence guarantee. Experimentally, we simulate photon-limited corruptions using CIFAR10/100 and ImageNet30 datasets and show that SharpDRO exhibits a strong generalization ability against severe corruptions and exceeds well-known baseline methods with large performance gains. Miaoxi Zhu, Xiaobo Xia, Li Shen 0008, Jun Yu 0001, Chen Gong 0002, Bo Han 0003, Bo Du 0001, Tongliang Liu |
CVPR | 5 |
| 2023 | Cross-Domain Transformer with Adaptive Thresholding for Domain Adaptive Semantic Segmentation
Quansheng Liu, Lei Wang 0203, Jun Yu 0001, Fang Gao 0001 |
ICANN (8) | 3 |
| 2023 | Combating Noisy Labels with Sample Selection by Mining High-Discrepancy ExamplesabstractThe sample selection approach is popular in learning with noisy labels. The state-of-the-art methods train two deep networks simultaneously for sample selection, which aims to employ their different learning abilities. To prevent two networks from converging to a consensus, their divergence should be maintained. Prior work presents that the divergence can be kept by locating the disagreement data on which the prediction labels of the two networks are different. However, this procedure is sample-inefficient for generalization, which means that only a few clean examples can be utilized in training. In this paper, to address the issue, we propose a simple yet effective method called CoDis. In particular, we select possibly clean data that simultaneously have high-discrepancy prediction probabilities between two networks. As selected data have high discrepancies in probabilities, the divergence of two networks can be maintained by training on such data. In addition, the condition of high discrepancies is milder than disagreement, which allows more data to be considered for training, and makes our method more sample-efficient. Moreover, we show that the proposed method enables to mine hard clean examples to help generalization. Empirical results show that CoDis is superior to multiple baselines in the robustness of trained models. Xiaobo Xia, Bo Han 0003, Yibing Zhan, Jun Yu 0001, Mingming Gong, Chen Gong 0002, Tongliang Liu |
ICCV | 4 |
| 2023 | Adaptive Fine-Grained Region Matching for Image Harmonization
Liuxue Ju, Chengdao Pu, Fang Gao 0001, Jun Yu 0001 |
ICIG (3) | 4 |
| 2023 | RatiO R-CNN: An Efficient and Accurate Detection Method for Oriented Object Detection
Chengdao Pu, Liuxue Ju, Fang Gao 0001, Jun Yu 0001 |
ICIG (3) | 4 |
| 2023 | Mosaic Representation Learning for Self-supervised Visual Pre-training
Zhaoqing Wang, Yandong Guo, Jun Yu 0001, Mingming Gong, Tongliang Liu |
ICLR | 5 |
| 2023 | Moderate Coreset: A Universal Method of Data Selection for Real-world Data-efficient Deep Learning
Xiaobo Xia, Jun Yu 0001, Xu Shen 0001, Bo Han 0003, Tongliang Liu |
ICLR | 3 |
| 2023 | Which is Better for Learning with Noisy Labels: The Semi-supervised Method or Modeling Label Noise?abstractIn real life, accurately annotating large-scale datasets is sometimes difficult. Datasets used for training deep learning models are likely to contain label noise. To make use of the dataset containing label noise, two typical methods have been proposed. One is to employ the semi-supervised method by exploiting labeled confident examples and unlabeled unconfident examples. The other one is to model label noise and design statistically consistent classifiers. A natural question remains unsolved: which one should be used for a specific real-world application? In this paper, we answer the question from the perspective of causal data generative process. Specifically, the performance of the semi-supervised based method depends heavily on the data generative process while the method modeling label-noise is not influenced by the generation process. For example, for a given dataset, if it has a causal generative structure that the features cause the label, the semi-supervised based method would not be helpful. When the causal structure is unknown, we provide an intuitive method to discover the causal structure for a given dataset containing label noise. Yu Yao 0005, Mingming Gong, Jun Yu 0001, Bo Han 0003, Kun Zhang 0001, Tongliang Liu |
ICML | 4 |
| 2023 | Prototypical Contrastive Learning for Domain Adaptive Semantic SegmentationabstractThe goal of domain adaptive semantic segmentation is to train a model using labeled source domain data and produce accurate dense predictions on the unlabeled target domain. Previous methods adopt self-training, where reliable target domain predictions are used as pseudo labels for training. However, intra-class variations across domains, such as the varying visual appearance in each category, have not been fully explored, leading to misalignment in feature distribution between the source and target domains. In this paper, we propose to optimize the feature space with representative prototypes shared across domains. Specifically, we first adopt the non-parametric clustering to model multiple prototypes for each category feature space. Then, category-discriminative feature space is obtained via pixel-to-prototype contrastive learning. Through extensive experiments, our proposed method demonstrates competitive performance on GTA5→Cityscapes and Synthia→Cityscapes benchmark. It is noteworthy that our method is compatible with the existing UDA methods. Quansheng Liu, Chengdao Pu, Fang Gao 0001, Jun Yu 0001 |
IJCNN | 4 |
| 2023 | Lightweight Neural Path PlanningabstractLearning-based path planning is becoming a promising robot navigation methodology due to its adaptability to various environments. However, the expensive computing and storage associated with networks impose significant challenges for their deployment on low-cost robots. Motivated by this practical challenge, we develop a lightweight neural path planning architecture with a dual input network and a hybrid sampler for resource-constrained robotic systems. Our architecture is designed with efficient task feature extraction and fusion modules to translate the given planning instance into a guidance map. The hybrid sampler is then applied to restrict the planning within the prospective regions indicated by the guide map. To enable the network training, we further construct a publicly available dataset with various successful planning instances. Numerical simulations and physical experiments demonstrate that, compared with baseline approaches, our approach has nearly an order of magnitude fewer model size and five times lower computational while achieving promising performance. Besides, our approach can also accelerate the planning convergence process with fewer planning iterations compared to sample-based methods. Zhen Kan, Jun Yu 0001 |
IROS | 5 |
| 2023 | FSR-Net: Deep Fourier Network for Shadow RemovalabstractThe presence of shadows degrades the performance of various multimedia tasks. Image shadow removal aims at restoring the background of shadow regions, which is generally an open challenge. Unlike most existing deep learning-based methods that focus on restoring such degradations in the spatial domain, we introduce a novel shadow removal method that also exploits frequency domain information. Specifically, we firstly revisit the frequency characteristics of shadow images via Fourier transform, where amplitude components contain most lightness information and phase components are related to structure information. To this end, we propose a two-stage deep Fourier shadow removal network (FSR-Net) to enhance the brightness of shadow regions, and correspondingly improve the shadow removal performance of whole images. For each stage, it consists of an amplitude recovery network and a phase recovery network to progressively reconstruct the lightness and structure components. To facilitate the learning of these two representations, we introduce the frequency and spatial interaction blocks to process the local spatial features and the global frequency information separately. Extensive experiments demonstrate that FSR-Net achieves superior results than other approaches with fewer parameters. For example, our method obtains a 1.05dB improvement on ISTD[34] dataset over the previous state-of-the-art method [43] with 0.30M parameters. Jun Yu 0001, Peng He 0004, Ziqi Peng |
ACM Multimedia | 1 |
| 2023 | Efficient Micro-Expression Spotting Based on Main Directional Mean Optical Flow FeatureabstractHuman facial expressions can convey a great deal of information in daily life. Spotting macro-expression (MaE) and micro-expression (ME) intervals from long video sequences is a difficult challenge. In this paper, we propose an efficient framework for the expression spotting task. This framework consists of three main modules: Face Cropping and Alignment Module (FCAM), optical flow Feature Extraction Module (FEM), and expression Proposal Generation Module (PGM). The noise of optical flow features is reduced by face cropping and alignment, and the Main Directional Mean Optical Flow Feature of the regions of interest is extracted as the feature for expression spotting. Finally, the expression intervals are spotted by our designed expression proposal generation module. Our approach achieves very good results on the SAMM Long Videos and CAS(ME)^2. To demonstrate the transferability of our method, we tested it on the MEGC2023 unseen dataset and finally achieved the third place, proving the effectiveness of our method. Jun Yu 0001, Zhongpeng Cai, Shenshen Du, Xiaxin Shen, Lei Wang 0203, Fang Gao 0001 |
ACM Multimedia | 1 |
| 2023 | Answer-Based Entity Extraction and Alignment for Visual Text Question AnsweringabstractAs a variant of visual question answering (VQA), visual text question answering (VTQA) provides a text-image pair for each question. Text utilizes named entities to describe corresponding image. Consequently, the ability to perform multi-hop reasoning using named entities between text and image becomes critically important. However, existing models pay relatively less attention to this aspect. Therefore, we propose Answer-Based Entity Extraction and Alignment Model (AEEA) to enable a comprehensive understanding and support multi-hop reasoning. The core of AEEA lies in two main components: AKECMR and answer aware predictor. The former emphasizes the alignment of modalities and effectively distinguishes between intra-modal and inter-modal information, and the latter prioritizes the full utilization of intrinsic semantic information contained in answers during training. Our model outperforms the baseline by 2.24% on test-dev set and 1.06% on test set, securing the third place in VTQA2023(English). Jun Yu 0001, Mohan Jing, Weihao Liu 0004, Tongxu Luo, Keda Lu, Fangyu Lei, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 1 |
| 2023 | Sliding Window Seq2seq Modeling for Engagement EstimationabstractEngagement estimation in human conversations has been one of the most important research issues for natural human-robot interaction. However, previous datasets and studies mainly focus on the video-wise level of engagement estimation, therefore, can hardly reflect human's constantly changing engagement. Fortunately, the MultiMediate '23 challenge provides the frame-wise level of engagement estimation task. In this paper, we propose Sliding Window Seq2seq Modeling by BiLSTM and Transformer with powerful sequence modeling capabilities. Our method fully utilizes the global and local multi-modal feature information in the participants' videos and accurately expresses the engagement of the participants at each moment. Our method achieves the state-of-the-art CCC result of 0.71 for engagement estimation on the corresponding test sets. Jun Yu 0001, Keda Lu, Mohan Jing, Ziqi Liang, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 1 |
| 2023 | Leveraging the Latent Diffusion Models for Offline Facial Multiple Appropriate Reactions GenerationabstractOffline Multiple Appropriate Facial Reaction Generation (OMAFRG) aims to predict the reaction of different listeners given a speaker, which is useful in the senario of human-computer interaction and social media analysis. In recent years, the Offline Facial Reactions Generation (OFRG) task has been explored in different ways. However, most studies only focus on the deterministic reaction of the listeners. The research of the non-deterministic (i.e. OMAFRG) always lacks of sufficient attention and the results are far from satisfactory. Compared with the deterministic OFRG tasks, the OMAFRG task is closer to the true circumstance but corresponds to higher difficulty for its requirement of modeling stochasticity and context. In this paper, we propose a new model named FRDiff to tackle this issue. Our model is developed based on the diffusion model architecture with some modification to enhance its ability of aggregating the context features. And the inherent property of stochasticity in diffusion model enables our model to generate multiple reactions. We conduct experiments on the datasets provided by the ACM Multimedia REACT2023 and obtain the second place on the board, which demonstrates the effectiveness of our method. Jun Yu 0001, Ji Zhao 0020, Guochen Xie, Fengxin Chen, Minglei Li 0001, Zonghong Dai |
ACM Multimedia | 1 |
| 2023 | Subclass-Dominant Label Noise: A Counterexample for the Success of Early StoppingabstractIn this paper, we empirically investigate a previously overlooked and widespread type of label noise, subclass-dominant label noise (SDN). Our findings reveal that, during the early stages of training, deep neural networks can rapidly memorize mislabeled examples in SDN. This phenomenon poses challenges in effectively selecting confident examples using conventional early stopping techniques. To address this issue, we delve into the properties of SDN and observe that long-trained representations are superior at capturing the high-level semantics of mislabeled examples, leading to a clustering effect where similar examples are grouped together. Based on this observation, we propose a novel method called NoiseCluster that leverages the geometric structures of long-trained representations to identify and correct SDN. Our experiments demonstrate that NoiseCluster outperforms state-of-the-art baselines on both synthetic and real-world datasets, highlighting the importance of addressing SDN in learning with noisy labels. The code is available at https://github.com/tmllab/2023_NeurIPS_SDN. Yingbin Bai, Zhongyi Han, Erkun Yang, Jun Yu 0001, Bo Han 0003, Dadong Wang, Tongliang Liu |
NeurIPS | 4 |
| 2023 | FlatMatch: Bridging Labeled Data and Unlabeled Data with Cross-Sharpness for Semi-Supervised LearningabstractSemi-Supervised Learning (SSL) has been an effective way to leverage abundant unlabeled data with extremely scarce labeled data. However, most SSL methods are commonly based on instance-wise consistency between different data transformations. Therefore, the label guidance on labeled data is hard to be propagated to unlabeled data. Consequently, the learning process on labeled data is much faster than on unlabeled data which is likely to fall into a local minima that does not favor unlabeled data, leading to sub-optimal generalization performance. In this paper, we propose FlatMatch which minimizes a cross-sharpness measure to ensure consistent learning performance between the two datasets. Specifically, we increase the empirical risk on labeled data to obtain a worst-case model which is a failure case needing to be enhanced. Then, by leveraging the richness of unlabeled data, we penalize the prediction difference (i.e., cross-sharpness) between the worst-case model and the original model so that the learning direction is beneficial to generalization on unlabeled data. Therefore, we can calibrate the learning process without being limited to insufficient label information. As a result, the mismatched learning performance can be mitigated, further enabling the effective exploitation of unlabeled data and improving SSL performance. Through comprehensive validation, we show FlatMatch achieves state-of-the-art results in many SSL settings. Li Shen 0008, Jun Yu 0001, Bo Han 0003, Tongliang Liu |
NeurIPS | 3 |
| 2023 | InstanT: Semi-supervised Learning with Instance-dependent ThresholdsabstractSemi-supervised learning (SSL) has been a fundamental challenge in machine learning for decades. The primary family of SSL algorithms, known as pseudo-labeling, involves assigning pseudo-labels to confident unlabeled instances and incorporating them into the training set. Therefore, the selection criteria of confident instances are crucial to the success of SSL. Recently, there has been growing interest in the development of SSL methods that use dynamic or adaptive thresholds. Yet, these methods typically apply the same threshold to all samples, or use class-dependent thresholds for instances belonging to a certain class, while neglecting instance-level information. In this paper, we propose the study of instance-dependent thresholds, which has the highest degree of freedom compared with existing methods. Specifically, we devise a novel instance-dependent threshold function for all unlabeled instances by utilizing their instance-level ambiguity and the instance-dependent error rates of pseudo-labels, so instances that are more likely to have incorrect pseudo-labels will have higher thresholds. Furthermore, we demonstrate that our instance-dependent threshold function provides a bounded probabilistic guarantee for the correctness of the pseudo-labels it assigns. Runze Wu 0001, Haoyu Liu 0002, Jun Yu 0001, Xun Yang 0001, Bo Han 0003, Tongliang Liu |
NeurIPS | 4 |
| 2023 | Dual-scale point cloud completion network based on high-frequency feature fusionabstractFor many vision tasks and intelligent robotics applications, it is common that the scanned 3D point cloud is not complete, so inferring from the residual defect shape to the intact shape becomes an essential task. Previous 3D completion neural network models generally use voxel-based or point-based methods to learn and process 3D data. For the voxel-based models, the computational cost and memory increase exponentially with the improvement of input resolution, and fine-grained features cannot be guaranteed in the completed point cloud due to limited computational resources . Point-based models suffer from the lack of precision in feature acquisition and crude reconstruction of complicated structures, making it extremely hard to accomplish elaborated semantic shapes. Combining advantages of voxel-based and point-based feature extraction through the high-frequency feature fusion module, this paper proposes a dual-scale point cloud completion network called DSNet, which performs global feature analysis at the voxel scale, and local feature analysis at the point cloud scale. The fused features are then integrated into the decoding and generation process, so as to complete the point cloud completion task from coarse to fine. Experimental results, at both quantitative and qualitative perspectives, in several prevailing datasets demonstrate that our approach surpasses state-of-the-art point cloud completion networks and has a good generalization performance . Code is available at https://github.com/engqing/DSNet. Fang Gao 0001, Pengbo Shi, Yan Jin 0012, Jun Yu 0001, Shaodong Li |
Image Vis. Comput. | 5 |
| 2023 | A viable framework for semi-supervised learning on realistic dataset
Guochen Xie, Jun Yu 0001, Qiang Ling 0001, Fang Gao 0001 |
Mach. Learn. | 3 |
| 2023 | Multi-Object Tracking: Decoupling Features to Solve the Contradictory Dilemma of Feature RequirementsabstractMulti-object tracking achieves the acquisition of target location information and identity information through two subtasks, detection and re-identification (ReID). The existing commonly used one-shot framework has speed advantages, but the two subtasks have different feature requirements, which leads to competitive learning in the training and thus weakens the feature quality. We propose a feature decoupling based multi-object tracking framework FDTrack for contradictory feature requirements. Through the mutual inhibition of the two subtasks, the features of the backbone network are decoupled. Then the decoupled features are self-constrained to enhance effective features. Considering the instability of the target state and the different confidence of the detections, a more reasonable association strategy is employed to maximize the matchings between detections, thus recovering low-confidence targets. FDTrack is extensively tested on the MOT17 and MOT20 benchmarks. The experimental results show that FDTrack surpasses the previous state-of-the-art (SOTA) methods and has good anti-interference and real-time performance. Moreover, our proposed modules have good portability and can be applied in other one-shot trackers to achieve performance improvement. Yan Jin 0012, Fang Gao 0001, Jun Yu 0001, Jiabao Wang 0004, Feng Shuang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Sample Selection with Uncertainty of Losses for Learning with Noisy Labels
Xiaobo Xia, Tongliang Liu, Bo Han 0003, Mingming Gong, Jun Yu 0001, Gang Niu 0001, Masashi Sugiyama |
ICLR | 5 |
| 2022 | DBCAN: Dual-Branch Cross-Attention Network for Scene Text RecognitionabstractScene text recognition, especially irregular text recognition, is a challenging task due to the large variance in text appearance. Although some existing methods have achieved state-of-the-art performance with the attention-based encoder-decoder framework, they always perform poorly on some challenging text such as severely curved, blurred, and incomplete-semantic text. To address these issues, we propose a Dual-Branch Cross-Attention Network (DBCAN). Different from the previous methods heavily relying on semantic information, DBCAN can enhance the position clues and learn semantic relations with two separate branches and fuse them by a tailored Cross-Attention Module (CAM). Furthermore, a Convolution-Based 2D Positional Embedding (CBPE) is introduced to describe the 2D spatial dependencies of characters. Extensive experiments demonstrate our DBCAN is more accurate and robust than the previous methods and achieves state-of-the-art performance on several benchmarks, particularly CUTE (93.4%). Our code is made publicly available at https://github.com/GaoXinJian-USTC/DBCAN. Xinjian Gao, Ye Pang, Yuyu Liu, Jun Yu 0001, Maokun Han, Wei Wang 0496 |
ICME | 4 |
| 2022 | Understanding Robust Overfitting of Adversarial Training and BeyondabstractRobust overfitting widely exists in adversarial training of deep networks. The exact underlying reasons for this are still not completely understood. Here, we explore the causes of robust overfitting by comparing the data distribution of non-overfit (weak adversary) and overfitted (strong adversary) adversarial training, and observe that the distribution of the adversarial data generated by weak adversary mainly contain small-loss data. However, the adversarial data generated by strong adversary is more diversely distributed on the large-loss data and the small-loss data. Given these observations, we further designed data ablation adversarial training and identify that some small-loss data which are not worthy of the adversary strength cause robust overfitting in the strong adversary mode. To relieve this issue, we propose minimum loss constrained adversarial training (MLCAT): in a minibatch, we learn large-loss data as usual, and adopt additional measures to increase the loss of the small-loss data. Technically, MLCAT hinders data fitting when they become easy to learn to prevent robust overfitting; philosophically, MLCAT reflects the spirit of turning waste into treasure and making the best use of each adversarial data; algorithmically, we designed two realizations of MLCAT, and extensive experiments demonstrate that MLCAT can eliminate robust overfitting and further boost adversarial robustness. Chaojian Yu, Bo Han 0003, Li Shen 0008, Jun Yu 0001, Chen Gong 0002, Mingming Gong, Tongliang Liu |
ICML | 4 |
| 2022 | Facial Expression Spotting Based on Optical Flow FeaturesabstractThe purpose of micro expression (ME) and macro expression (MaE) spotting task is to locate the onset and offset frames of MaE and ME clips. Compared with MaEs, MEs are shorter in duration and lower in intensity, which makes MEs harder to be spotted. In this paper, we propose an efficient pipeline based on optical flow features to spot MEs and MaEs. We crop and align the faces and select the eyebrows area, nose area, and mouth area as our regions of interest to exclude the interference of extraneous factors on the face expression representation. Then, we extract the optical flow in these regions and enhance the features of the expressions in the optical flow with low-pass filter and EMD method. Finally, the sliding window method is used to locate the peaks of optical flow features and get the intervals containing MEs or MaEs. We evaluate the performance of our method on the MEGC2022-TestSet including 10 long videos from SAMM and CAS(ME)3 and achieve the first place in the MEGC2022 Challenge. The results prove the effectiveness of our method. Jun Yu 0001, Zhongpeng Cai, Guochen Xie, Peng He 0004 |
ACM Multimedia | 1 |
| 2022 | Micro Expression Generation with Thin-plate Spline Motion Model and Face ParsingabstractMicro-expression generation aims at transfering the expression from the driving videos to the source images, which can be viewed as a motion transfer task. Recently, several works have been proposed to tackle this problem and achieve great performance. However, due to the intrinsic complexity of the face motion and different attributes of face regions, the task still remains challenging. In this paper, we propose an end-to-end unsupervised motion transfer network to tackle this challenge. As the motion of the face is non-rigid, we adopt an effective and flexible thin-plate spline motion estimation method to estimate the optical flow of the face motion. What's more, we find that several faces with eyeglasses show weird deformation in motion transfering. Thus, we introduce face parsing method to pay specific attention to the eyeglasses regions to ensure the reasonability of the deformation. We conduct several experiments on the provided datasets of the ACM MM 2022 micro-expression grand challenge (MEGC2022) and compare our method with several other typical methods. In comparison, our method shows the best performance. We (Team: USTC-IAT-United) also compare our method with other competitors' in MEGC2022, and the expert evaluation results show that our method performs best, which verifies the effectiveness of our method. Our code is available at https://github.com/HowToNameMe/micro-expression Jun Yu 0001, Guochen Xie, Zhongpeng Cai, Peng He 0004, Fang Gao 0001, Qiang Ling 0001 |
ACM Multimedia | 1 |
| 2022 | Efficient 6D object pose estimation based on attentive multi-scale contextual informationabstractAbstract 6D pose estimation has been pervasively applied to various robotic applications, such as service robots, collaborative robots, and unmanned warehouses. However, accurate 6D pose estimation is still a challenge problem due to the complexity of application scenarios caused by illumination changes, occlusion and even truncation between objects, and additional refinement is required for accurate 6D object pose estimation in prior work. Aiming at the efficiency and accuracy of 6D object pose estimation in these complex scenes, this paper presents a novel end‐to‐end network, which effectively utilises the contextual information within a neighbourhood region of each pixel to estimate the 6D object pose from RGB‐D images. Specifically, our network first applies the attention mechanism to extract effective pixel‐wise dense multimodal features, which are then expanded to multi‐scale dense features by integrating pixel‐wise features at different scales for pose estimation. The proposed method is evaluated extensively on the LineMOD and YCB‐Video datasets, and the experimental results show that the proposed method is superior to several state‐of‐the‐art baselines in terms of average point distance and average closest point distance. Fang Gao 0001, Qingyi Sun, Shaodong Li, Yong Li 0028, Jun Yu 0001, Feng Shuang 0002 |
IET Comput. Vis. | 6 |
| 2022 | Dual feature fusion network: A dual feature fusion network for point cloud completionabstractAbstract Point cloud data in the real world is often affected by occlusion and light reflection, leading to incompleteness of the data. Large‐region missing point clouds will cause great deviations in downstream tasks. A dual feature fusion network (DFF‐Net) is proposed to improve the accuracy of the completion of a large missing region of the point cloud. First, a dual feature encoder is designed to extract and fuse the global and local features of the input point cloud. Subsequently, a decoder is used to directly generate a point cloud of missing region that retains local details. In order to make the generated point cloud more detailed, a loss function with multiple terms is employed to emphasise the distribution density and visual quality of the generated point cloud. A large number of experiments show that the authors’ DFF‐Net is better than the previous state‐of‐the‐art methods in the aspect of point cloud completion. Fang Gao 0001, Pengbo Shi, Jiabao Wang 0004, Yaoxiong Wang, Jun Yu 0001, Yong Li 0028, Feng Shuang 0002 |
IET Comput. Vis. | 6 |
| 2022 | Monocular depth estimation with spatially coherent sliced network
Wen Su 0004, Haifeng Zhang 0006, Yuan Su, Jun Yu 0001, Zengfu Wang |
Image Vis. Comput. | 4 |
| 2022 | Learning from Noisy Pairwise Similarity and Unlabeled DataabstractSU classification employs similar (S) data pairs (two examples belong to the same class) and unlabeled (U) data points to build a classifier, which can serve as an alternative to the standard supervised trained classifiers requiring data points with class labels. SU classification is advantageous because in the era of big data, more attention has been paid to data privacy. Datasets with specific class labels are often difficult to obtain in real-world classification applications regarding privacy-sensitive matters, such as politics and religion, which can be a bottleneck in supervised classification. Fortunately, similarity labels do not reveal the explicit information and inherently protect the privacy, e.g., collecting answers to “With whom do you share the same opinion on issue $\mathcal{I}$?" instead of “What is your opinion on issue $\mathcal{I}$?". Nevertheless, SU classification still has an obvious limitation: respondents might answer these questions in a manner that is viewed favorably by others instead of answering truthfully. Therefore, there exist some dissimilar data pairs labeled as similar, which significantly degenerates the performance of SU classification. In this paper, we study how to learn from noisy similar (nS) data pairs and unlabeled (U) data, which is called nSU classification. Specifically, we carefully model the similarity noise and estimate the noise rate by using the mixture proportion estimation technique. Then, a clean classifier can be learned by minimizing a denoised and unbiased classification risk estimator, which only involves the noisy data. Moreover, we further derive a theoretical generalization error bound for the proposed method. Experimental results demonstrate the effectiveness of the proposed algorithm on several benchmark datasets. Songhua Wu, Tongliang Liu, Bo Han 0003, Jun Yu 0001, Gang Niu 0001, Masashi Sugiyama |
J. Mach. Learn. Res. | 4 |
| 2022 | Embedding Pose Information for Multiview Vehicle Model RecognitionabstractVehicle model recognition is a typical fine-grained classification task that has a wide range of application prospects in safe cities and constitutes a research hotspot in the field of computer vision. Vehicles in images can appear at various angles, resulting in large differences in appearance. The existence of “multiviews” renders vehicle model recognition challenging. Recent research on vehicle model recognition has not fully explored the pose information of vehicles in different images, resulting in low model performance. In this study, we use vehicle pose information to solve the multiview vehicle model recognition (MV-VMR) problem and design a convolutional neural network (CNN) model with embedded vehicle pose information, known as the embedding pose CNN (EP-CNN). The proposed model includes two subnetworks: the pose estimation subnetwork (PE-SubNet) and vehicle model classification subnetwork (VMC-SubNet). PE-SubNet extracts the vehicle pose information, including the pose features and vehicle viewpoint. In VMC-SubNet, considering the scale variation of vehicles, an improved squeeze-and-excitation (SE) block, named the MultiSE block is implemented. We embed the vehicle viewpoint into the MultiSE block, which reweighs each channel such that the extracted features elicit different responses to different viewpoints. Subsequently, the pose features and classification features are integrated for classification. Experiments are conducted on the benchmark CompCars web-nature and Stanford Cars datasets. The results demonstrate that the proposed EP-CNN method can achieve higher recognition accuracy than most classic CNN models and several state-of-the-art fine-grained vehicle model classification algorithms. Code has been made available at:https://github.com/HFUT-CV/EP-CNN. Yuanzi Fu, Wei Jia 0001, Jun Yu 0001, Zhisheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | LR-SVM+: Learning Using Privileged Information with Noisy LabelsabstractThe paradigm of Learning Using Privileged Information (LUPI) always assumes that labels are annotated precisely. However, in practice, this assumption may be violated, as the labels may be heavily noisy, which inevitably degenerates the performance of learning algorithms in the LUPI paradigm. To handle the side effect of noisy labels, we propose a novel Label Noise Robust SVM+ (LR-SVM+) algorithm. Specifically, as the privileged information contains rich information of the latent labels, we first utilize it to infer underlying clean labels. Then we use the inference to modify the noisy labels. Comprehensive experiments demonstrate the necessity of studying label noise robust SVM+ and the effectiveness of the proposed method. Zhengning Wu, Xiaobo Xia, Ruxin Wang 0002, Jun Yu 0001, Yinian Mao, Tongliang Liu |
IEEE Trans. Multim. | 5 |
| 2022 | TWGAN: Twin Discriminator Generative Adversarial NetworksabstractGenerative Adversarial Networks (GAN) has become more and more popular these years. However, it is difficult to train and suffers from the training instability problem. To tackle this difficulty, this paper proposes a novel approach. Our idea is intuitive but proven to be very useful. In essence, it combines saturating loss and non-saturating loss into the loss function. Thus it will exploit the complementary statistical properties from two kinds of loss functions to effectively improve the training stability. We term our method twin discriminator Generative Adversarial Networks (TWGAN), which, unlike GAN, has a generator and a twin discriminator. The twin discriminator consists of two discriminators with identical architecture and both of them aim to distinguish whether the samples are from real data or fake data. We develop theoretical analysis to show that, given the optimal discriminators, optimizing the generator of TWGAN reduces to minimizing the Kullback-Leibler (KL) divergence between the distribution of generated data ($P_g$) and the distribution of real data ($P_data$), hence effectively addressing the training instability problem. Extensive experiments on MNIST, Fashion MNIST, CIFAR-10/100 and STL-10 datasets demonstrate that the competitive performance of our TWGAN in generating good quality and diverse samples over baselines. The obtained highest inception score (IS) and lowest Fr$\acute{e}$chet Inception Distance (FID), compared with other state-of-the-art GANs, show the superiority of our TWGAN. Zhaoyu Zhang 0001, Haonian Xie, Jun Yu 0001, Tongliang Liu, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2022 | Densely Enhanced Semantic Network for Conversation System in Social MediaabstractThe human–computer conversation system is a significant application in the field of multimedia. To select an appropriate response, retrieval-based systems model the matching between the dialogue history and response candidates. However, most of the existing methods cannot fully capture and utilize varied matching patterns, which may degrade the performance of the systems. To address the issue, a densely enhanced semantic network (DESN) is proposed in our work. Given a multi-turn dialogue history and a response candidate, DESN first constructs the semantic representations of sentences from the word perspective, the sentence perspective, and the dialogue perspective. In particular, the dialogue perspective is a novel one introduced in our work. The dependencies between a single sentence and the whole dialogue are modeled from the dialogue perspective. Then, the response candidate and each utterance in the dialogue history are made to interact with each other. The varied matching patterns are captured for each utterance–response pair by using a dense matching module. The matching patterns of all the utterance–response pairs are accumulated in chronological order to calculate the matching degree between the dialogue history and the response. The responses in the candidate pool are ranked with the matching degree, thereby returning the most appropriate candidate. Our model is evaluated on the benchmark datasets. The experimental results prove that our model achieves significant and consistent improvement when compared with other baselines. Zengfu Wang, Jun Yu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Deep Kinship Verification and Retrieval Based on Fusion Siamese Neural NetworkabstractAutomatic kinship analysis, which aims to judge the kinship of different individuals, has been widely used in many real world applications such as helping missing persons reunite with their families and social media analysis. In this work, we focus on three practical and challenging tasks related to kinship analysis, i.e., kinship verification, tri-subject kinship verification and kinship retrieval. A deep fusion Siamese neural network is proposed to address these tasks in a flexible and progressive manner. Firstly, we propose a basic deep Siamese neural network for kinship verification to judge the kinship between individuals based on face images. More specifically, the Siamese neural network takes two input face images and then outputs the similarity between them. To improve the performance, a jury system is also introduced for multi-model fusion. Secondly, we integrate two basic deep Siamese neural networks for tri-subject kinship verification(father, mother and child), which is intended to decide whether a child is related to a pair of parents or not. Specifically, the kinship similairty score of the triplet for verification is obtained by weighting the similarity scores of the father-child and mother-child ones. Thirdly, the proposed deep Siamese neural network can be used to quantify the similarity between any two persons. Thus, it is natural and easy to extend its application to the kinship retrieval task by sorting the similarities between the candidates and faces in the database. We conduct experiments on the RFIW2021 dataset, and final results validate the effectiveness of our solution. Jun Yu 0001, Guochen Xie, Xinlong Hao, Zeyu Cui, Zhongpeng Cai |
FG | 1 |
| 2021 | Removing Adversarial Noise in Class Activation Feature SpaceabstractDeep neural networks (DNNs) are vulnerable to adversarial noise. Pre-processing based defenses could largely remove adversarial noise by processing inputs. However, they are typically affected by the error amplification effect, especially in the front of continuously evolving attacks. To solve this problem, in this paper, we propose to remove adversarial noise by implementing a self-supervised adversarial training mechanism in a class activation feature space. To be specific, we first maximize the disruptions to class activation features of natural examples to craft adversarial examples. Then, we train a denoising model to minimize the distances between the adversarial examples and the natural examples in the class activation feature space. Empirical evaluations demonstrate that our method could significantly enhance adversarial robustness in comparison to previous state-of-the-art approaches, especially against unseen adversarial attacks and adaptive attacks. Dawei Zhou 0004, Nannan Wang 0001, Chunlei Peng, Xinbo Gao 0001, Xiaoyu Wang 0002, Jun Yu 0001, Tongliang Liu |
ICCV | 6 |
| 2021 | Radar Object Detection Using Data Merging, Enhancement and FusionabstractCompared to visible images, radar images are generally considered to be an active and robust solution, even in adverse driving situations, for object detection. However, the accuracy of radar object detection (ROD) is always poor. Owing to taking full advantage of data merging, enhancement and fusion, this paper proposes an effective ROD system with only radar images as the input. First, an aggregation module is designed to merge the data from all chirps in the same frame. Then, various gaussian noises with different parameters are employed to increase data diversity and reduce over-fitting based on the analysis of training data. Moreover, due to the process of inference with default parameters is not accurate enough, some hyperparameters are changed to increase the accuracy performance. Finally, a combination strategy is adopted to benefit from multi-model fusion. ROD2021 Challenge is supported by ACM ICMR 2021, and our team (ustc-nelslip) ranked 2nd in the test stage of this challenge. Diverse evaluations also verify the superiority of the proposed system. Jun Yu 0001, Xinlong Hao, Xinjian Gao, Yuyu Liu, Peng Chang 0002, Fang Gao 0001, Feng Shuang 0002 |
ICMR | 1 |
| 2021 | Fine-Grained Language Identification in Scene Text ImagesabstractIdentifying the language of the text in scene images is crucial for various applications. Studies that focus on identifying the script, which is a set of letters used for writing in a given language, in scene text images already exist. However, these works do not distinguish between different languages written in the same script and are thus unable to meet the needs of many applications. To address this challenge, we study a novel task: fine-grained language identification in scene text images, which aims to distinguish languages that share the same script. The datasets that include samples in seven languages, which are Dutch, English, French, Italian, German, Spanish, and Portuguese, are constructed. Furthermore, well-designed end-to-end trainable neural networks are proposed for fine-grained language identification, where semantic information concerning the text is mined and utilized to assist the language identification. We train the networks on the synthetic dataset and evaluate them with the collected real dataset. The experimental results demonstrate that the proposed frameworks are effective. Shilian Wu, Jun Yu 0001, Zengfu Wang |
ACM Multimedia | 3 |
| 2021 | Multimodal Inputs Driven Talking Face Generation With Spatial-Temporal DependencyabstractGiven an arbitrary speech clip or text information as input, the proposed work aims to generate a talking face video with accurate lip synchronization. Existing works mainly have three limitations. (1) A single-modal learning is adopted with either audio or text as input, hence it lacks the complementarity ofmultimodal inputs. (2) Each frame is generated independently, hence it ignores thetemporal dependencybetween consecutive frames. (3) Each face image is generated by the traditional convolution neural network (CNN) with a local receptive field, hence it cannot effectively capture thespatial dependencywithin internal representations of face images. To overcome these problems above, we decompose the talking face generation task into two steps: mouth landmarks prediction and video synthesis. First, a multimodal learning method is proposed to generate accurate mouth landmarks with multimedia inputs (both text and audio). Second, a network named Face2Vid is proposed to generate video frames conditioned on the predicted mouth landmarks. In Face2Vid, the optical flow is employed to model the temporal dependency between frames, meanwhile, a self-attention mechanism is introduced to model the spatial dependency across image regions. Extensive experiments demonstrate that our approach can generate photo-realistic video frames with the background, and exhibit the superiorities on accurate synchronization of lip movements and smooth transition of facial movements. Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | A Multilayer Pyramid Network Based on Learning for Vehicle Logo RecognitionabstractIn this paper, we present a novel learning-based scheme for vehicle logo recognition (VLR). This scheme is termed Multilayer Pyramid Network Based on Learning (MLPNL) and is based on the principle that considering multiple resolutions is helpful for extracting valuable features that benefit the final recognition performance. The innovations of this scheme include (1) a multilayer pyramid network, with pixel difference matrices (PDMs) as its input and output and feature parameters mapping one PDM to another; (2) an objective function and a corresponding optimization method designed to facilitate the learning of the feature parameters of the proposed multilayer pyramid network; and (3) a multi-codebook-based encoding method that makes best use of the features extracted from PDMs corresponding to different resolutions. Extensive experiments conducted with an open dataset, HFUT-VL, demonstrate that the proposed MLPNL scheme outperforms state-of-the-art handcrafted descriptors and non-deep-learning-based learning methods when fewer training samples exist. Experiments conducted with a benchmark dataset, XMU, demonstrate that MLPNL outperforms existing state-of-the-art VLR methods. Experiments conducted both on HFUT-VL and XMU demonstrate that MLPNL is faster than most deep-learning-based learning methods while maintaining nearly the same recognition rate. Code has been made available at:https://github.com/HFUT-CV/MLPNL. Jun Wang 0071, Hai Min, Wei Jia 0001, Jun Yu 0001, Chang Wen Chen |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Learning Face Image Super-Resolution Through Facial Semantic Attribute Transformation and Self-Attentive Structure EnhancementabstractFace super-resolution is a domain-specific super-resolution (SR) problem of generating high-resolution (HR) face images from low-resolution (LR) inputs. Even though existing face SR methods have achieved great performance on the global region evaluation, most of them cannot restore local attributes and structure reasonably, especially to ultra-resolve tiny LR face images (16 × 16 pixels) to its larger version (8 × upscaling factor). In this paper, we propose an open source face SR framework based on facial semantic attribute transformation and self-attentive structure enhancement. Specifically, the proposed framework introduces face semantic information (i.e., face attributes) and face structure information (i.e., face boundaries) in a successive two-stage fashion. In the first stage, an Attribute Transformation Network (AT-Net) is established. It upsamples LR face images to HR feature maps and then combines facial attributes with these features to generate the intermediate HR results with rational attributes. In the second stage, a Structure Enhancement Network (SE-Net) is built. It simultaneously extracts face features and estimates facial boundary heatmaps from the inputs, and then fuses them to output the final HR face images. Extensive experiments demonstrate that our method achieves superior super-resolved results and outperforms the state-of-the-art methods. Zhaoyu Zhang 0001, Jun Yu 0001, Chang Wen Chen |
IEEE Trans. Multim. | 3 |
| 2020 | Retrieval of Family Members Using Siamese Neural NetworkabstractRetrieval of family members in the wild aims at finding family members of the given subject in the dataset, which is useful in finding the lost children and analyzing the kinship. However, due to the diversity in age, gender, pose and illumination of the collected data, this task is always challenging. To solve this problem, we propose our solution with deep Siamese neural network. Our solution can be divided into two parts: similarity computation and ranking. In training procedure, the Siamese network firstly takes two candidate images as input and produces two feature vectors. And then, the similarity between the two vectors is computed with several fully connected layers. While in inference procedure, we try another similarity computing method by dropping the followed several fully connected layers and directly computing the cosine similarity of the two feature vectors. After similarity computation, we use the ranking algorithm to merge the similarity scores with the same identity and output the ordered list according to their similarities. To gain further improvement, we try different combinations of backbones, training methods and similarity computing methods. Finally, we submit the best combination as our solution and our team(ustc-nelslip) obtains favorable result in the track3 of the RFIW2020 challenge with the first runner-up, which verifies the effectiveness of our method. Our code is available at: https://github.com/gniknoil/FG2020-kinship. Jun Yu 0001, Guochen Xie, Xinlong Hao |
FG | 1 |
| 2020 | Deep Fusion Siamese Network for Automatic Kinship VerificationabstractAutomatic kinship verification aims to determine whether some individuals belong to the same family. It is of great research significance to help missing persons reunite with their families. In this Work, the challenging problem is progressively addressed in two respects. First, we propose a deep siamese network to quantify the relative similarity between two individuals. When given two input face images, the deep siamese network extracts the features from them and fuses these features by combining and concatenating. Then, the fused features are fed into a fully-connected network to obtain the similarity score between two faces, which is used to verify the kinship. To improve the performance, a jury system is also employed for multi-model fusion. Second, two deep siamese networks are integrated into a deep triplet network for tri-subject (i.e., father, mother and child) kinship verification, which is intended to decide whether a child is related to a pair of parents or not. Specifically, the obtained similarity scores of father-child and mother-child are weighted to generate the parent-child similarity score for kinship verification. Recognizing Families In the Wild (RFIW) is a challenging kinship recognition task with multiple tracks, which is based on Families in the Wild (FIW), a large-scale and comprehensive image database for automatic kinship recognition. The Kinship Verification (track I) and Tri-Subject Verification (track II) are supported during the ongoing RFIW2020 Challenge. Our team (ustc-nelslip) ranked 1st in track II, and 3rd in track L The code is available at https://github.com/gniknoil/FG2020-kinship. Jun Yu 0001, Xinlong Hao, Guochen Xie |
FG | 1 |
| 2020 | Weakly Supervised Local-Global Relation Network for Facial Expression RecognitionabstractTo extract crucial local features and enhance the complementary relation between local and global features, this paper proposes a Weakly Supervised Local-Global Relation Network (WS-LGRN), which uses the attention mechanism to deal with part location and feature fusion problems. Firstly, the Attention Map Generator quickly finds the local regions-of-interest under the supervision of image-level labels. Secondly, bilinear attention pooling is employed to generate and refine local features. Thirdly, Relational Reasoning Unit is designed to model the relation among all features before making classification. The weighted fusion mechanism in the Relational Reasoning Unit makes the model benefit from the complementary advantages between different features. In addition, contrastive losses are introduced for local and global features to increase the inter-class dispersion and intra-class compactness at different granularities. Experiments on lab-controlled and real-world facial expression dataset show that WS-LGRN achieves state-of-the-art performance, which demonstrates its superiority in FER. Haifeng Zhang 0006, Wen Su 0004, Jun Yu 0001, Zengfu Wang |
IJCAI | 3 |
| 2020 | Attention Based Beauty Product Retrieval Using Global and Local DescriptorsabstractBeauty product retrieval has drawn more and more attention for its wide application outlook and enormous economic benefits. However, this task is always challenging due to the variation of products, especially the disturbance of clustered background. In this paper, we first introduce attention mechanism into a global image descriptor, i.e., Maximum Activation of Convolutions (MAC), and propose Attention-based MAC (AMAC). With this enhancement, we can suppress the negative effect of background and highlight the foreground in an unsupervised manner. Then, AMAC and local descriptors are ensembled to complementarily increase the performance. Furthermore, we try to finetune multiple retrieval methods on the different datasets and adopt a query expansion strategy to obtain more improvements. Extensive experiments conducted on a dataset containing more the half million beauty products (Perfect-500K) demonstrate the effectiveness of the proposed method. Finally, our team (USTC-NELSLIP) wins the first place on the leaderboard of the 'AI Meets Beauty'Grand Challenge of ACM Multimedia 2020. The code is available at: https://github.com/gniknoil/Perfect500K-Beauty-Product-Retrieval-Challenge. Jun Yu 0001, Guochen Xie, Haonian Xie, Xinlong Hao, Fang Gao 0001, Feng Shuang 0002 |
ACM Multimedia | 1 |
| 2020 | A Deep Learning Approach for Face Hallucination Guided by Facial Boundary ResponsesabstractFace hallucination is a domain-specific super-resolution (SR) problem of learning a mapping between a low-resolution (LR) face image and its corresponding high-resolution (HR) image. Tremendous progress on deep learning has shown exciting potential for a variety of face hallucination tasks. However, most deep-learning–based methods are limited to handle facial appearance information without paying attention to facial structure priors. In this article, we propose an open source 1 Boundary-aware Dual-branch Network (BDN) for face hallucination, which simultaneously extracts face features and estimates facial boundary responses from LR inputs, ultimately fusing them to reconstruct HR results. Specifically, we first upsample LR face images to HR feature maps, and then feed the upsampled HR features into a memory unit and an attention unit synchronously to obtain the refined features and predict facial boundary responses. Next, they are fed into a feature map fusion unit to combine facial appearance and structure information by a spatial attention mechanism. Moreover, we employ a series of stacked units to boost performance before recovering HR face images. Finally, a discriminative network is developed to improve visual quality by introducing adversarial learning strategy. Extensive experiments show that the proposed approach achieves superior face hallucination results against the state-of-the-art ones. Zhaoyu Zhang 0001, Guochen Xie, Jun Yu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | Towards the Gradient Vanishing, Divergence Mismatching and Mode Collapse of Generative Adversarial NetsabstractGenerative adversarial network (GAN) is a powerful generative model. However, it suffers from gradient vanishing, divergence mismatching and mode collapse. To overcome these problems, we propose a novel GAN, which consists of one generator G and two discriminators (D1, D2). Focusing on the gradient vanishing, Spectral Normalization (SN) and ResBlock are first adopted in D1 and D2. Then, Scaled Exponential Linear Units (SELU) is adopted at last half layers of D2 to further address the problem. To divergence mismatching, relativistic discriminator is adopted in our GAN to make the loss function minimization in the training of generator equal to the theoretical divergence minimization. Concentrating on the mode collapse, D1 rewards high scores for the samples from the data distribution, while D2 favors the samples from the generator conversely. In addition, the minibatch discrimination is adopted in D1 to further address the problem. Extensive experiments on CIFAR-10/100 and ImageNet datasets demonstrate that our GAN can obtain the highest inception score (IS) and lowest Frechet Inception Distance (FID) compared with other state-of-the-art GANs. Zhaoyu Zhang 0001, Changwei Luo, Jun Yu 0001 |
CIKM | 3 |
| 2019 | D2PGGAN: Two Discriminators Used in Progressive Growing of GANSabstractGenerative adversarial network (GAN) is a powerful generative model. However, it suffers from two key problems which are convergence instability and mode collapse. Recently, progressive growing of GANs for improving quality, stability and variation (PGGAN) is proposed to better solve these two problems. Although the performance of PGGAN is good on these two problems, it is still not satisfied on mode collapse problem. In this paper, we propose a new architecture based on PGGAN called D2PGGAN to better solve the mode collapse problem. The key idea consists of one generator and two different discriminators in PGGAN. With the fact that GAN is the analogy of a minimax game, the proposed architecture is as follows. The generator (G) aims to produce realistic-looking samples to fool both of two discriminators. The first discriminator (D1) rewards high scores for samples from the data distribution, while the second one (D2) favors samples from the generator conversely. Specifically, a novel loss function is designed to optimize the proposed D2PGGAN. Extensive experiments on CIFAR-10 and CIFAR-100 datasets demonstrate that the proposed method is effective and obtains the highest inception scores compared with others state-of-the-art GANs. Zhaoyu Zhang 0001, Jun Yu 0001 |
ICASSP | 3 |
| 2019 | Dense Semantic Matching Network for Multi-turn ConversationabstractMining the semantic information in the text to model the multi-turn conversation has attracted great interests. Previous models ignore the capturing and utilizing of the matching patterns at different levels, which causes the loss of the valuable information for the calculation of the matching degree. To address the problem, we propose a dense semantic matching network (DMN). Given a context-response pair, DMN first constructs semantic representations for the response candidate and each utterance in the context. Then, DMN models the interaction between the response candidate and each utterance in the context to generate interactive matrices. By processing the interactive matrices with dense convolutional blocks, the hierarchical matching patterns are generated for each utterance-response pair. The matching patterns of all the utterance-response pairs are finally accumulated in chronological order with bidirectional long short-term memory network. The final matching score of the response candidate and the multi-turn context is finally calculated. We evaluate the performance of our network on the benchmark dataset. The results show that our network yields a significant performance gain compared with other methods. Jun Yu 0001, Zengfu Wang |
ICDM | 2 |
| 2019 | Mining Audio, Text and Visual Information for Talking Face GenerationabstractProviding methods to support audio-visual interaction with growing volumes of video data is an increasingly important challenge for data mining. To this end, there has been some success in speech-driven lip motion generation or talking face generation. Among them, talking face generation aims to generate realistic talking heads synchronized with the audio or text input. This task requires mining the relationship between audio signal/text and lip-sync video frames and ensures the temporal continuity between frames. Due to the issues such as polysemy, ambiguity, and fuzziness of sentences, creating visual images with lip synchronization is still challenging. To overcome the problems above, we present a data-mining framework to learn the synchronous pattern between different channels from large recorded audio/text dataset and visual dataset, and apply it to generate realistic talking face animations. Specifically, we decompose this task into two steps: mouth landmarks prediction and video synthesis. First, a multimodal learning method is proposed to generate accurate mouth landmarks with multimedia inputs (both text and audio). Second, a network named Face2Vid is proposed to generate video frames conditioned on the predicted mouth landmarks. In Face2Vid, optical flow is employed to model the temporal dependency between frames, meanwhile, a self-attention mechanism is introduced to model the spatial dependency across image regions. Extensive experiments demonstrate that our method can generate realistic videos with background, and exhibit the superiorities on accurate synchronization of lip movements and smooth transition of facial movements. Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001 |
ICDM | 2 |
| 2019 | Deep Learning Face Hallucination via Attributes Transfer and EnhancementabstractFace hallucination technique aims to generate high-resolution (HR) face images from low-resolution (LR) inputs. Even though existing face hallucination methods have achieved great performance on the global region evaluation, most of them cannot reasonably restore local attributes, especially when ultra-resolving tiny LR face image (16 × 16 pixels) to its larger version (8× upscaling factor). In this paper, we propose a novel attribute-guided face transfer and enhancement network for face hallucination. Specifically, we first construct a face transfer network, which upsamples LR face images to HR feature maps, and then fuses facial attributes and the upsampled features to generate HR face images with rational attributes. Finally, a face enhancement network is developed based on generative adversarial network (GAN) to improve visual quality by exploiting a composite loss that combines image color, texture and content. Extensive experiments demonstrate that our method achieves superior face hallucination results and outperforms the state-of-the-art. Yuechuan Sun, Zhaoyu Zhang 0001, Haonian Xie, Jun Yu 0001 |
ICME | 5 |
| 2019 | Beauty Product Retrieval Based on Regional Maximum Activation of Convolutions with Generalized AttentionabstractBeauty and Personal care product retrieval has attracted more and more research attention for its value in real life. However, suffering from data variants and complex background, this task has been very challenging. In this paper, we propose a novel Generalized-attention Regional Maximal Activation of Convolutions (GRMAC) descriptor which helps to generate image features for retrieval. This method introduces attention mechanism to reduce the influence of clustered background and highlight the target, and thus contributes to enhancing the effectiveness of features and boosting the retrieval performance. Different from other attention-based methods, our method supports adjusting mask with a hyperparameter p, which is more flexible and accurate in real application. To demonstrate its effectiveness, we conduct experiments on the dataset containing more than half million personal care products (Perfect-500K) and obtain remarkable results. Furthermore, we try to fuse multiple features from different models for more improvements. And finally, our team (USTC_NELSLIP) ranked 1st in the Grand Challenge of AI Meets Beauty in ACM Multimedia 2019 with a MAP score of 0.408614. Our code is available at: https://github.com/gniknoil/Perfect500K-Beauty-and-Personal-Care-Products-Retrieval-Challenge Jun Yu 0001, Guochen Xie, Haonian Xie, Lingyun Yu 0002 |
ACM Multimedia | 1 |
| 2019 | 3D Singing Head for Music VR: Learning External and Internal Articulatory Synchronicity from Lyric, Audio and NotesabstractWe propose a real-time 3D singing head system to enhance the talking head on model integrity, keyframe generation and song synchronicity. The individual head appearance meshes are first obtained by matching multi-view visible images with face prior for accuracy, and then used to reconstruct entire head model by integrating with generic internal articulatory meshes for efficiency. After embedding physiology, the keyframes of each phoneme-music note correspondence are substantially synthesized from real articulation data. The song synchronicity of articulators is learned using a deep neural network to train visual co-articulation model (VCM) on parallel audio-visual data. Finally, the keyframes of adjacent phoneme-music note correspondences are blended by VCM to produce song synchronized animation. Compared to state-of-the-art baselines, our system can not only clearly distinguish phonemes and notes, but also significantly reduce the dependence on training data. Jun Yu 0001, Chang Wen Chen, Zengfu Wang |
ACM Multimedia | 1 |
| 2019 | STDGAN: ResBlock Based Generative Adversarial Nets Using Spectral Normalization and Two Different DiscriminatorsabstractGenerative adversarial network (GAN) is a powerful generative model. However, it suffers from two key problems, which are convergence and mode collapse. To overcome these drawbacks, this paper presents a novel architecture of GAN, called STDGAN, which consists of one generator and two different discriminators. With the fact that GAN is the analogy of a minimax game, the proposed architecture is as follows. The generator G aims to produce realistic-looking samples to fool both of two discriminators. The first discriminator D1 rewards high scores for the samples from the data distribution, while the second one D2 favors the samples from the generator conversely. Specifically, the minibatch discrimination and Spectral Normalization (SN) are first adopted in D1. Then, based on the ResBlock architecture, Spectral Normalization (SN) and Scaled Exponential Linear Units (SELU) are adopted in the first and last half layers of D2 respectively. In particular, a novel loss function is designed to optimize the STDGAN by minimizing the KL divergence. Extensive experiments on CIFAR-10/100 and ImageNet datasets demonstrate that the proposed STDGAN can effectively solve the problems of convergence and mode collapse and obtain the higher inception score (IS) and lower Frechet Inception Distance (FID) compared with other state-of-the-art GANs. Zhaoyu Zhang 0001, Jun Yu 0001 |
ACM Multimedia | 2 |
| 2019 | Deep Neural Network Based 3D Articulatory Movement Prediction Using Both Text and Audio Inputs
Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001 |
MMM (1) | 2 |
| 2019 | Synthesizing 3D Trump: Predicting and Visualizing the Relationship Between Text, Speech, and Articulatory MovementsabstractThe movements of articulators, such as lips, tongue and teeth, play an important role in increasing the language expression capability by unmasking the information hid in text or speech. Hence, it is necessary to deeply mine and visualize the relationship between text, speech and articulatory movements for understanding language in multi-modality and multi-level. As a case study, given text and audio of President Donald John Trump, this paper synthesizes a high quality 3D animation of him speaking with accurate synchronicity between speech and articulators. First, visual co-articulation is modeled by predicting the mapping from text/speech to articulatory movements. Then, based on a reconstructed 3D head model, physiological characteristics and statistical learning are combined to visualize each phoneme. Finally, the visualization results of consecutive phonemes are fused by visual co-articulation model to generate synchronized articulatory animations. Experiments show that the system can not only produce photo-realistic results in front but also distinguish the visual differences among phonemes from unconstrained views. Jun Yu 0001, Qiang Ling 0001, Changwei Luo, Chang Wen Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Real-Time Head Pose Estimation and Face Modeling From a Depth ImageabstractWe address the issues of 3-D head pose estimation and face modeling from a depth image. Given a depth image, random forests are effective for estimating the location and orientation of a person's head. However, the accuracy of the estimation is not high enough. We propose using corrected regression votes. The corrected votes are obtained by considering the cooperation of all trees, leading to significant improvement of head pose estimation accuracy. Based on the head pose estimator, we present a face modeling system. In our system, the face model is generated by aligning a deformable face model to a depth image using an iterative closest point (ICP) algorithm. The novelty of our approach is that an optimal weight for each vertex is incorporated into the ICP algorithm with point to plane constraints. Experiments show that our system can automatically estimate the head pose and generate a realistic face model from a single depth image. We also provide a detailed evaluation that shows the benefits of our approach. Changwei Luo, Juyong Zhang, Jun Yu 0001, Chang Wen Chen, Shengjin Wang |
IEEE Trans. Multim. | 3 |
| 2019 | BLTRCNN-Based 3-D Articulatory Movement Prediction: Learning Articulatory Synchronicity From Both Text and Audio InputsabstractPredicting articulatory movements from audio or text has diverse applications, such as speech visualization. Various approaches have been proposed to solve the acoustic-articulatory mapping problem. However, their precision is not high enough with only acoustic features available. Recently, deep neural network (DNN) has brought tremendous success in various fields, like speech recognition and image processing. To increase the accuracy, we propose a new network architecture for articulatory movement prediction with both text and audio inputs, called a bottleneck long-term recurrent convolutional neural network (BLTRCNN). To the best of our knowledge, it is the first time to predict articulatory movements based on DNN by fusing text and audio inputs. Our BLTRCNN consists of two networks. The first is the bottleneck network, generating a compact bottleneck features of text information for each frame independently. The second, including convolutional neural network, long short-term memory and skip connection, is called the long-term recurrent convolutional neural network (LTRCNN). LTRCNN is used for articulatory movement prediction when bottleneck features, acoustic features, and text features are integrated as inputs together. Experiments show that the proposed BLTRCNN achieves the state-of-the-art root-mean-square error (RMSE) 0.528 mm and the correlation coefficient 0.961. Moreover, we also demonstrate how text information complements acoustic features in this prediction task. Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Deep Facial Attribute Detection in the Wild: From General to Specific
Yuechuan Sun, Jun Yu 0001 |
BMVC | 2 |
| 2018 | Synthesizing Photo-Realistic 3D Talking Head: Learning Lip Synchronicity and Emotion from Audio and VideoabstractWe propose a speech driven emotional 3D talking head. First, based on face prior knowledge, an individual head model is reconstructed by integrating the internal articulators from Magnetic Resonance Imaging slices with the appearance from visible images. Second, the blendshapes of each phoneme and the facial expression are synthesized by statistical learning. Third, visual coarticulation is modeled by training a deep neural network on parallel audio-visual data. Finally, the articulatory animations of continuous phonemes are fused by visual co articulation model and facial expressions to produce expressive speech synchronized animations. Experiments show that the system can not only distinguish the visual differences among phonemes in real-time, but also significantly increase the ability of human perception. Jun Yu 0001, Lingyun Yu 0002 |
ICIP | 1 |
| 2018 | A Coarse-to-Fine Face Hallucination Method by Exploiting Facial Prior KnowledgeabstractFace hallucination technique generates high-resolution (HR) face images from low-resolution (LR) ones. In this paper, we propose to use a coarse-to-fine method for face hallucination by constructing a two-branch network, which makes full use of the specific prior knowledge of face images and the advantages of generic image super-resolution (SR) methods. Specifically, we jointly build a deep neural network (DNN) with a face image SR branch and a semantic face parsing branch. The former branch implements the image upsampling and feature extraction using a cascade of convolutional layers. The latter branch extracts facial semantic parsing as prior knowledge. Then, we combine the image features and the prior know ledge to reconstruct HR face images. Finally, we optimize the DNN, by using adversarial training and a perceptual loss, in order to obtain high realism. Extensive experiments show that the proposed method outperforms the state-of-the-art alternatives in terms of accuracy and realism. Yuechuan Sun, Zhaoyu Zhang 0001, Jun Yu 0001 |
ICIP | 4 |
| 2018 | Simultaneous Facial Landmark and 3D Action Estimation Based on Probabilistic Random ForestabstractRandom forest is effective and efficient for detecting facial landmark from visual images. It has achieved the state-of-the-art performance, both in accuracy and speed, by regressing local binary features (LBF). This paper aims to increase the detection accuracy of random forest for facial landmarks and extends it to facial action estimation. First, probabilistic features with the improved regression measure are designed to overcome the weaknesses of LBF, namely feature sparseness and tracking jitter, for initial detection. Second, the initial detected facial landmarks and 3D facial actions are jointly refined and estimated. Specifically, a deformable facial model is registered to input images based on an optimized iterative closest point framework, in which an optimal weight is additionally assigned to each vertex with the point to plane constraints. Experiments show that the proposed methods significantly outperform the state-of-the-art ones in terms of accuracy, as well as achieve the excellent tracking stability and real-time ability at about 80 fps for estimating landmarks+actions on an ordinary PC. Jun Yu 0001, Yuechuan Sun |
ICIP | 1 |
| 2018 | A Cross-Layer Based Network for Faster Image GenerationabstractOwing to the great success of generative adversarial networks (GANs), unsupervised learning based image generation is popular currently. This paper presents a cross-layer architecture for the generators in GANs, which passes inputs to every subsequent layer and encourages the flow of information and gradients throughout the network. While traditional networks with L layers have L connections, the proposed network with cross-layer architecture has 2L-1 direct connections. For each layer, the network input and the output of its previous layer are used as inputs. Extensive experiments demonstrate that our network can generate images at a higher speed without introducing extra parameters by comparing with two state-of-the-art GANs, namely deep convolutional GAN and Wasserstein GAN-GP, on two datasets: Fashion-MNIST and CelebA. Zhaoyu Zhang 0001, Yuechuan Sun, Jun Yu 0001 |
ICIP | 3 |
| 2018 | Synthesizing 3D Acoustic-Articulatory Mapping Trajectories: Predicting Articulatory Movements by Long-Term Recurrent Convolutional Neural NetworkabstractRobust and accurate predicting of articulatory movements has various important applications, such as 3D articulatory animations and visual communication. Various approaches have been proposed to solve the acoustic-articulatory mapping problem. However, their precision is not high enough. Recently, deep neural network (DNN), especially convolutional neural network (CNN) and recurrent neural network (RNN), has brought tremendous success in speech recognition and synthesis. To increase the accuracy, we propose a new network architecture for acoustic-articulatory mapping, called long-term recurrent convolutional neural network (LTRCNN). The network consists of CNN, RNN and a skip connection. CNN can model the spectral correlation among acoustic features efficiently. RNN, like long short-term memory (LSTM), can learn the temporal context information from sequential data powerfully. Besides, skip connections can increase the input representation from different levels to preserve the feature information. Experiments show that LTRCNN achieves the state-of-the-art root-mean-squared error (RMSE) with 0.690 mm and the correlation coefficient with 0.949 in this prediction task. Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001 |
VCIP | 2 |
| 2018 | General-to-specific learning for facial attribute classification in the wild
Yuechuan Sun, Jun Yu 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Probability contour guided depth map inpainting and superresolution using non-local total generalized variation
Hai-Tao Zhang, Jun Yu 0001, Zengfu Wang |
Multim. Tools Appl. | 2 |
| 2018 | Real-Time 3-D Facial Animation: From Appearance to Internal ArticulatorsabstractA real-time 3-D facial animation system produces animation for appearance and internal articulators. For appearance, an anatomical model, including the skeleton, muscle, and skin, is built based on anatomical characteristics and a data-driven model is obtained by learning the mapping between texture and depth. Then, the two models are combined to produce animations with various strengths, since the anatomical model can control the animation strength directly and the data-driven model can capture the nuances of facial motion. For internal articulators, tongue tissue arrangements are obtained from medical data. Then, a nonlinear, quasi-incompressible, isotropic, hyperelastic biomechanical model is applied to describe tongue tissues and an anisotropic biomechanical model is applied to reflect the active and passive mechanical behavior of tongue muscle fibers. The tongue animation is simulated using the finite-element method for realism, while the collisions between the tongue and other articulators are simulated with a mass-spring model for efficiency. Experiments show that the system achieves high perceptual evaluation scores and quantitative improvements are demonstrated in the objective evaluation and user studies, compared with the outputs of other systems. Jun Yu 0001, Chen Jiang 0002, Rui Li 0021, Changwei Luo, Zengfu Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | A real-time 3D head mesh modeling and expressive articulatory animation systemabstractIn view of animated human computer interfaces, this paper proposes a 3D head mesh modeling and expressive articulatory animation system. The appearance mesh model is first reconstructed from multi-view visible images using inter-regional cooperative optimization and depth super-resolution, and the universal internal articulatory mesh is then integrated with the reconstructed appearance mesh by interpolation. After establishing the head mesh model, the anatomical and biomechanical characteristics of articulators are combined to synthesize articulatory animation. The evaluations demonstrate the system can build a realistic and vivid virtual head for animated interface in real-time. Jun Yu 0001, Zengfu Wang |
ICASSP | 1 |
| 2017 | HMM based speech-driven 3D tongue animationabstractWe propose a speech-driven 3D tongue animation system. Firstly, the input speech is analyzed to obtain the phoneme sequence. Next, articulatory movements are predicted from the phoneme sequence using a hidden Markov model (HMM) based framework. The HMMs are trained beforehand using a corpus of human articulatory movements, which are recorded by three electromagnetic articulograph (EMA) sensors glued on the tongue tip, tongue body, and tongue dorsum of a speaker respectively. Finally, the predicted articulatory movements are used to control the deformations of a 3D tongue model. The tongue model is a triangular mesh with three key vertices. The three vertices are chosen so that their positions are in correspondence with the three EMA sensors mentioned above. Our tongue model can achieve various tongue shapes with volume preservation. Experiments show that the generated tongue animations are realistic and synchronize well with the input speech. Changwei Luo, Jun Yu 0001 |
ICIP | 2 |
| 2017 | Depth map super-resolution using non-local higher-order regularization with classified weightsabstractHigh-order regularization in depth map super-resolution (SR) contributes to producing smoother depth map. However, assigning appropriate weights within regularization term is also important for preserving more detail information. In this paper, a novel and more adaptive depth SR model is proposed by using non-local total generalized variation (NLTGV) with classified weights. A random forest based classifier is trained to classify the pixels of depth map into four categories on the basis of several local structure features, such as gradient magnitude and texture energy extracted from color image, and then the weights within NLTGV are assigned with four groups of parameters corresponding to the four kinds of pixels. Evaluation results demonstrate that the local features make pixels have a good separability, and the classified weights can obviously improve the accuracy of depth map SR. Hai-Tao Zhang, Jun Yu 0001, Zengfu Wang |
ICIP | 2 |
| 2017 | From talking head to singing head: A significant enhancement for more natural human computer interactionabstractThis paper proposes a 3D virtual animating head system, which can not only talk but also sing. With a reconstructed head mesh model, including external/internal articulators, from multi-source images, biology information are first used to visualize each phoneme with a musical note. The synchronicity between songs and articulatory movements is then modeled by a deep neural network trained on an audio/articulatory corpus. Finally, the visualization results of phonemes are blended by the synchronicity model to produce the song synchronized articulatory animations. Quantitative and qualitative improvements of singing ability on human computer interaction are demonstrated by comparing with other state-of-the-art talking head systems. Jun Yu 0001, Chang Wen Chen |
ICME | 1 |
| 2017 | Adaptively Weighted Facial Expression Recognition by Feature Fusion Under Intense Illumination Condition
Yuechuan Sun, Jun Yu 0001 |
ICONIP (6) | 2 |
| 2017 | Facial Expression Recognition by Fusing Gabor and Local Binary Pattern Features
Yuechuan Sun, Jun Yu 0001 |
MMM (2) | 2 |
| 2017 | A Real-Time 3D Visual Singing Synthesis: From Appearance to Internal Articulators
Jun Yu 0001 |
MMM (1) | 1 |
| 2017 | Speech Synchronized Tongue Animation by Combining Physiology Modeling and X-ray Image Fitting
Jun Yu 0001 |
MMM (1) | 1 |
| 2017 | A Unified Framework for Monocular Video-Based Facial Motion Tracking and Expression Recognition
Jun Yu 0001 |
MMM (2) | 1 |
| 2017 | Multimodal 3D visible articulation system for syllable based Mandarin Chinese trainingabstractVisible of articulatory movements play a significant role in assistive language and articulation learning as well as training. In this paper, we develop a 3D visible articulation system based on speech signals and speech related vision information. The system is developed for the utmost important linguistic unit (syllable) of Mandarin Chinese. By 3D visualizing articulators in magnetic resonance images (MRI), with further parameterized modeling and electromagnetic articulography (EMA) data based animating, the developed 3D visible articulation system has the ability to generate speech synchronization and articulatory visible animation. Additionally, we propose a collision handling method to solve the penetration problem which occurs in the process of tongue motion. The accuracy of simulated articulatory motion was evaluated by root-mean-square error (RMSE). The performance of our system for articulation training was evaluated according to realism and expressiveness aspects. Rui Li 0021, Jun Yu 0001 |
VCIP | 2 |
| 2017 | Joint facial landmark detection and action estimation based on deep probabilistic random forestabstractRandom forest is effective and efficient for detecting facial landmark from visual images, and has achieved the state-of-the-art performance, both in accuracy and speed, by regressing local binary features (LBF). This paper aims to increase the detection accuracy of random forest for facial landmarks and extends it to facial action estimation. First, probabilistic features are designed to overcome the weaknesses of LBF, e.g., feature sparseness and tracking jitter. Second, a deep architecture is introduced to random forest for enhancing the capacity of representation learning. Third, the initial detected facial landmarks are refined and 3D facial actions are estimated jointly by registering a deformable facial model to images based on an optimized iterative closest point framework. Experiments show that the proposed methods significantly outperform the state-of-the-art ones in terms of accuracy, as well as achieve the excellent tracking stability and real-time ability at about 60 fps on an ordinary PC. Jun Yu 0001, Chang Wen Chen |
VCIP | 1 |
| 2017 | Image classification based on convolutional neural networks with cross-level strategy
Yu Liu 0023, Jun Yu 0001, Zengfu Wang |
Multim. Tools Appl. | 3 |
| 2017 | Creating and simulating a realistic physiological tongue model for speech production
Jun Yu 0001, Chen Jiang 0002, Zengfu Wang |
Multim. Tools Appl. | 1 |
| 2017 | Realistic emotion visualization by combining facial animation and hairstyle synthesis
Jun Yu 0001, Lingyan Li |
Multim. Tools Appl. | 1 |
| 2017 | A Video-Based Facial Motion Tracking and Expression Recognition System
Jun Yu 0001, Zengfu Wang |
Multim. Tools Appl. | 1 |
| 2017 | A realistic 3D articulatory animation system for emotional visual pronunciation
Lingyun Yu 0002, Jun Yu 0001, Zengfu Wang |
Multim. Tools Appl. | 2 |
| 2016 | A fast and precise speech-triggered tongue animation system by combining parameterized model and anatomical modelabstractA 3D realistic tongue system is proposed. Firstly, the muscle geometry and fiber arrangement are specified after a tongue mesh model is constructed from medical data. Secondly, with the target of the efficiency and realism of animation, the tongue tissues, including tongue muscles, are described by combining a fast parametric model and a precise anatomical model to simulate the active and passive mechanical characteristics. The experiments demonstrate the suitability of the system for speech visualization. Jun Yu 0001, Chen Jiang 0002, Zengfu Wang |
BIBM | 1 |
| 2016 | A realistic and reliable 3D pronunciation visualization instruction system for computer-assisted language learningabstractA text-driven 3D pronunciation visualization instruction system is proposed for computer-assisted language learning. Based on a 3D articulatory mesh model including appearance and internal articulators, both finite element method and anatomical model are used to synthesize the articulatory animation of phonemes by fitting the mesh model to the detected articulatory shapes in X-ray images. Visual co-articulation is modeled with a Hidden Markov Model trained on an articulatory speech corpus. Articulatory animations corresponding to all phonemes of a learned text are concatenated by visual co-articulation model to produce the speech synchronized articulatory animation. The experiments for Mandarin Chinese show the system can increase the pronunciation accuracy of learners. Jun Yu 0001, Zengfu Wang |
BIBM | 1 |
| 2016 | Facial video coding/decoding at ultra-low bit-rate: a 2D/3D model-based approach
Jun Yu 0001, Changwei Luo, Lingyun Yu 0002, Ling-yan Li, Zengfu Wang |
Multim. Tools Appl. | 1 |
| 2015 | Video Based Face Tracking and Animation
Changwei Luo, Jun Yu 0001, Zhigang Zheng, Lingyun Yu 0002, Zengfu Wang |
ICIG (3) | 2 |
| 2015 | Real-Time Robust Video Stabilization Based on Empirical Mode Decomposition and Multiple Evaluation Criteria
Jun Yu 0001, Changwei Luo, Chen Jiang 0002, Rui Li 0021, Ling-yan Li, Zengfu Wang |
ICIG (3) | 1 |
| 2015 | Locating Facial Landmarks Using Probabilistic Random ForestabstractRandom forest is a useful tool for face alignment/tracking. The method of regressing local binary features learned from random forest has achieved state-of-the-art performance both in fitting accuracy and speed. Despite the great success of this method, it has certain weaknesses: the number of available local binary features is rather limited and is not optimal for face alignment; the binary features inevitably lead to serious jitter when tracking a video sequence. To address these problems, we propose learning probability features from probabilistic random forest (PRF). The proposed PRF is the same as standard random forest except that it models the probability of a sample belonging to the nodes of a tree. By using the probability features, our method significantly outperforms the state-of-the-art in terms of accuracy. It also achieves about 60 fps for locating a few facial landmarks. In addition, our method shows excellent stability in face tracking. Changwei Luo, Zengfu Wang, Shaobiao Wang, Juyong Zhang, Jun Yu 0001 |
IEEE Signal Process. Lett. | 5 |
| 2015 | A Video, Text, and Speech-Driven Realistic 3-D Virtual Head for Human-Machine InterfaceabstractA multiple inputs-driven realistic facial animation system based on 3-D virtual head for human-machine interface is proposed. The system can be driven independently by video, text, and speech, thus can interact with humans through diverse interfaces. The combination of parameterized model and muscular model is used to obtain a tradeoff between computational efficiency and high realism of 3-D facial animation. The online appearance model is used to track 3-D facial motion from video in the framework of particle filtering, and multiple measurements, i.e., pixel color value of input image and Gabor wavelet coefficient of illumination ratio image, are infused to reduce the influence of lighting and person dependence for the construction of online appearance model. The tri-phone model is used to reduce the computational consumption of visual co-articulation in speech synchronized viseme synthesis without sacrificing any performance. The objective and subjective experiments show that the system is suitable for human-machine interaction. Jun Yu 0001, Zengfu Wang |
IEEE Trans. Cybern. | 1 |
| 2014 | Synthesizing real-time speech-driven facial animationabstractWe present a real-time speech-driven facial animation system. In this system, Gaussian Mixture Models (GMM) are employed to perform the audio-to-visual conversion. The conventional GMM-based method performs the conversion frame by frame using minimum mean square error (MMSE) estimation. The method is reasonably effective. However, discontinuities often appear in the sequences of estimated visual features. To solve this problem, we incorporate previous visual features into the conversion so that the conversion procedure is performed in the manner of a Markov chain. After audio-to-visual conversion, the estimated visual features are transformed to blendshape weights to synthesize facial animation. Experiments show that our system can accurately convert audio features into visual features. The conversion accuracy is comparable to a current state-of-the-art trajectory-based approach. Moreover, our system runs in real time and outputs high quality lip-sync animations. Changwei Luo, Jun Yu 0001, Zengfu Wang |
ICASSP | 2 |
| 2014 | Expressive facial animation from videosabstractWe address the issue of synthesizing real-time expressive facial animation from videos. Given video footage of a person's face, some existing methods track a few facial landmarks to drive the virtual character. Compared with these methods, ours has the following characteristics. 1) We incorporate global texture into a constrained local model to increase the accuracy of facial tracking. 2) To animate a blenshape face model, facial tracking results as well as facial expression recognition results are used to estimate blendshape weights. Experimental results demonstrate that our method is effective for producing realistic expressive facial animations. Moreover, the method does not require facial markers or complex offline pre-processing, these properties make it very easy to use for ordinary users. Changwei Luo, Chen Jiang 0002, Jun Yu 0001, Zengfu Wang |
ICIP | 3 |
| 2014 | Real-time control of 3D facial animationabstractFacial animation is useful in human-machine interaction, computer games and teleconferences. We propose a realtime performance-driven facial animation system for ordinary users. The system enables a user to animate an avatar by performing desired facial motions in front of a video camera. First, a constrained local model based approach is used to track facial features of a performer in the video. To increase the tracking accuracy, we propose an efficient method to build a user-specific local texture model. Next, a 3D blendshape face model is fitted to the tracked feature points. To improve the expressiveness of synthesized animations, facial expression recognition results and pre-recorded animation priors are incorporated into the fitting procedure. Finally, facial animations are created using blendshape interpolation. Experiments show that the synthetic facial motions are realistic and quite similar to the facial actions of the performer. By using an ordinary camera, our system provides the user complete control over the generated facial animations. Changwei Luo, Jun Yu 0001, Chen Jiang 0002, Rui Li 0021, Zengfu Wang |
ICME | 2 |
| 2014 | 3D facial motion tracking by combining online appearance model and cylinder head model in particle filtering
Jun Yu 0001, Zengfu Wang |
Sci. China Inf. Sci. | 1 |
| 2013 | 2D/3D Model-Based Facial Video Coding/Decoding at Ultra-Low Bit-Rate
Jun Yu 0001, Zengfu Wang, Yang Cao 0010 |
MMM (2) | 1 |