VLDB 2026 Research / reviewers in the wild / expert
Jia Li 0013
dblp:23/6950-13
· DBLP profile ↗
31ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0001-9446-249XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 19 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language NavigationabstractNavigating unseen environments based on natural language instructions remains difficult for egocentric agents in Vision-and-Language Navigation (VLN). Intuitively, humans inherently ground concrete semantic knowledge within spatial layouts during indoor navigation. Although previous studies have introduced diverse environmental representations to enhance reasoning, other co-occurrence modalities are often naively concatenated with RGB features, resulting in suboptimal utilization of each modality's distinct contribution. Inspired by this, we propose a hierarchical Semantic Understanding and Spatial Awareness (SUSA) architecture to enable agents to perceive and ground environments at diverse scales. Specifically, the Textual Semantic Understanding (TSU) module supports local action prediction by generating view-level descriptions, thereby capturing fine-grained environmental semantics and narrowing the modality gap between instructions and environments. Complementarily, the Depth-enhanced Spatial Perception (DSP) module incrementally constructs a trajectory-level depth exploration map, providing the agent with a coarse-grained comprehension of the global spatial layout. Extensive experiments demonstrate that SUSA's hierarchical representation enrichment not only boosts the navigation performance of the baseline on discrete VLN benchmarks (REVERIE, R2R, and SOON), but also exhibits superior generalization to the continuous R2R-CE. Yunbo Xu, Jia Li 0013, Zhenzhen Hu 0004 |
AAAI | 3 |
| 2026 | Fine-grained Text-Video Retrieval with Patch-level Temporal Difference and AggregationabstractExisting Text-Video Retrieval (TVR) methods predominantly rely on global frame representations, often disregarding the fine-grained temporal variations required for precise patch-level alignment. This is critical as video motion is inherently spatially localized; consequently, coarse frame-level modeling tends to be dominated by static backgrounds, overshadowing salient action cues. To address this limitation, we propose TRFG, a novel framework for text-video retrieval that addresses the challenges of modeling Temporal Reasoning and Fine-Grained cross-modal alignment. First, our Temporal Difference module captures frame-to-frame variations at the patch level, effectively suppressing static background noise to highlight "active" motion regions. Second, these differential signals are synthesized via a Temporal Aggregation module to form a coherent representation of the event’s trajectory. Finally, to ensure precise semantic matching, a fine-grained interaction module aligns these dynamic video tokens with textual details. Extensive experiments on MSRVTT, ActivityNet, and DiDeMo demonstrate that TRFG achieves state-of-the-art performance across multiple backbones and retrieval tasks. Ablation studies confirm the complementarity and generalizability of both modules, underscoring the importance of explicit temporal modeling and fine-grained interaction in bridging the modality gap. Jialong Hu, Zijie Song, Yang Wang 0023, Zhenzhen Hu 0004, Jia Li 0013, Richang Hong |
ICMR | 5 |
| 2026 | MetaPipe: Predicting Metaphoric Associations and Turning Your Metaphoric Imagination into RealityabstractVisual metaphors, in particular, are powerful tools to communicate complex ideas effectively. However, creating meaningful and visually compelling metaphors is inherently difficult, especially for amateur designers, who often lack the tools and expertise to form strong conceptual connections. To address these challenges, we introduce MetaPipe, a mobile-based design tool that assists users in crafting high-quality visual metaphors. By leveraging three interactive modules, MetaPipe simplifies the creation process, making it accessible to non-experts. We evaluated MetaPipe through two user studies: (1) a usability study with 16 non-design professionals rating their experience of the tool, and (2) a blinded experiment in which 74 participants assessing creativity and metaphorical strength of MetaPipe-generated images. The results demonstrate that MetaPipe significantly enhances the visual metaphor design process, enabling the creation of high-quality and creative images. The code is available: https://github.com/Qianvenh/MetaPipe. Wenhao Qian, Zhenzhen Hu 0004, Jialong Hu, Jia Li 0013, Richang Hong |
Int. J. Hum. Comput. Interact. | 4 |
| 2026 | CLAIP-Emo: Parameter-Efficient Adaptation of Language-Supervised Models for In-the-Wild Audiovisual Emotion RecognitionabstractAudiovisual emotion recognition (AVER) in the wild is still hindered by pose variation, occlusion, and background noise. Prevailing methods primarily rely on large-scale domain-specific pre-training, which is costly and often mismatched to real-world affective data. To address this, we present CLAIP-Emo, a modular framework that reframes in-the-wild AVER as a parameter-efficient adaptation of language-supervised foundation models (CLIP/CLAP). Specifically, it (i) preserves language-supervised priors by freezing CLIP/CLAP backbones and performing emotion-oriented adaptation via LoRA (updating \ensuremath{\le}4.0\% of the total parameters), (ii) allocates temporal modeling asymmetrically, employing a lightweight Transformer for visual dynamics while applying mean pooling for audio prosody, and (iii) applies a simple fusion head for prediction. On DFEW and MAFW, CLAIP-Emo (ViT-L/14) achieves 80.14\% and 61.18\% weighted average recall with only 8M training parameters, setting a new state of the art. Our findings suggest that parameter-efficient adaptation of language-supervised foundation models provides a scalable alternative to domain-specific pre-training for real-world AVER. The code and models will be available at \href{https://github.com/MSA-LMC/CLAIP-Emo}{https://github.com/MSA-LMC/CLAIP-Emo}. Jia Li 0013, Jinpeng Hu, Zhenzhen Hu 0004, Richang Hong |
IEEE Signal Process. Lett. | 2 |
| 2026 | Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression DataabstractDynamic facial expression recognition (DFER) infers emotions from the temporal evolution of expressions, unlike static facial expression recognition (SFER), which relies solely on a single snapshot. This temporal analysis provides richer information and promises greater recognition capability. However, current DFER methods often exhibit unsatisfied performance largely due to fewer training samples compared to SFER. Given the inherent correlation between static and dynamic expressions, we hypothesize that leveraging the abundant SFER data can enhance DFER. To this end, we propose Static-for-Dynamic (S4D), a unified dual-modal learning framework that integrates SFER data as a complementary resource for DFER. Specifically, S4D employs dual-modal self-supervised pre-training on facial images and videos using a shared Vision Transformer (ViT) encoder-decoder architecture, yielding improved spatiotemporal representations. The pre-trained encoder is then fine-tuned on static and dynamic expression datasets in a multi-task learning setup to facilitate emotional information interaction. Unfortunately, vanilla multi-task learning in our study results in negative transfer. To address this, we propose an innovative Mixture of Adapter Experts (MoAE) module that facilitates task-specific knowledge acquisition while effectively extracting shared knowledge from both static and dynamic expression data. Extensive experiments demonstrate that S4D achieves a deeper understanding of DFER, setting new state-of-the-art performance on FERV39K, MAFW, and DFEW benchmarks, with weighted average recall (WAR) of 53.65%, 58.44%, and 76.68%, respectively. Additionally, a systematic correlation analysis between SFER and DFER tasks is presented, which further elucidates the potential benefits of leveraging SFER. Jia Li 0013, Yu Zhang 0082, Zhenzhen Hu 0004, Shiguang Shan, Meng Wang 0001, Richang Hong |
IEEE Trans. Affect. Comput. | 2 |
| 2026 | PhysioSync: Temporal and Cross-Modal Contrastive Learning Inspired by Physiological Synchronization for EEG-Based Emotion RecognitionabstractElectroencephalography (EEG) signals provide a promising and involuntary reflection of brain activity related to emotional states, offering significant advantages over behavioral cues such as facial expressions. However, EEG signals are often noisy, affected by artifacts, and vary across individuals, complicating emotion recognition. While multimodal approaches have used peripheral physiological signals (PPS) such as galvanic skin response to complement EEG, they often overlook the dynamic synchronization and consistent semantics between the modalities. Additionally, the temporal dynamics of emotional fluctuations across different time resolutions in PPS remain underexplored. To address these challenges, we propose PhysioSync, a novel pretraining framework leveraging temporal and cross-modal contrastive learning (CM-CL), inspired by physiological synchronization phenomena. PhysioSync incorporates cross-modal consistency alignment (CM-CA) to model dynamic relationships between EEG and complementary PPS, enabling emotion-related synchronizations across modalities. Besides, it introduces long- and short-term temporal contrastive learning (LS-TCL) to capture emotional synchronization at different temporal resolutions within modalities. After pretraining, cross-resolution and cross-modal features are hierarchically fused and fine-tuned to enhance emotion recognition. Experiments on DEAP and DREAMER datasets demonstrate PhysioSync’s advanced performance under unimodal and cross-modal conditions, highlighting its effectiveness for EEG-centered emotion recognition. Jia Li 0013, Yu Liu 0023, Zhenzhen Hu 0004, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video RetrievalabstractText-video retrieval (TVR) has seen substantial advancements in recent years, fueled by the utilization of pre-trained models and large language models (LLMs). Despite these advancements, achieving accurate matching in TVR remains challenging due to inherent disparities between video and textual modalities and irregularities in data representation. In this paper, we propose Text-Video-ProxyNet (TV-ProxyNet), a novel framework designed to decompose the conventional 1-to-N relationship of TVR into N distinct 1-to-1 relationships. By replacing a single text query with a series of text proxies, TV-ProxyNet not only broadens the query scope but also achieves a more precise expansion. Each text proxy is crafted through a refined iterative process, controlled by mechanisms we term as the director and dash, which regulate the proxy's direction and distance relative to the original text query. This setup not only facilitates more precise semantic alignment but also effectively manages the disparities and noise inherent in multimodal data. Our experiments on three representative video-text retrieval benchmarks, MSRVTT, DiDeMo, and ActivityNet Captions, demonstrate the effectiveness of TV-ProxyNet. The results show an improvement of 2.0% to 3.3% in R@1 over the baseline. TV-ProxyNet achieved state-of-the-art performance on MSRVTT and ActivityNet Captions, and a 2.0% improvement on DiDeMo compared to existing methods, validating our approach's ability to enhance semantic mapping and reduce error propensity. Zhenzhen Hu 0004, Jia Li 0013, Richang Hong |
AAAI | 3 |
| 2025 | Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQAabstractVideo Question Answering (VideoQA) is a complex video-language task that demands a sophisticated understanding of both visual content and temporal dynamics. Traditional Transformer-style architectures, while effective in integrating multimodal data, often simplify temporal dynamics through positional encoding and fail to capture non-linear interactions within video sequences. In this paper, we introduce the Temporal Trio Transformer (T3T), a novel architecture that models time consistency and time variability. The T3T integrates three key components: Temporal Smoothing (TS), Temporal Difference (TD), and Temporal Fusion (TF). The TS module employs Brownian Bridge for capturing smooth, continuous temporal transitions, while the TD module identifies and encodes significant temporal variations and abrupt changes within the video content. Subsequently, the TF module synthesizes these temporal features with textual cues, facilitating a deeper contextual understanding and response accuracy. The efficacy of the T3T is demonstrated through extensive testing on multiple VideoQA benchmark datasets. Our results underscore the importance of a nuanced approach to temporal modeling in improving the accuracy and depth of video-based question answering. Zijie Song, Zhenzhen Hu 0004, Jia Li 0013, Richang Hong |
ICME | 4 |
| 2025 | Seeing is Believing? Enhancing Vision-Language Navigation using Visual PerturbationsabstractAutonomous navigation guided by natural language instructions in embodied environments remains a challenge for vision-language navigation (VLN) agents. Although recent advancements in learning diverse and fine-grained visual environmental representations have shown promise, the fragile performance improvements may not conclusively attribute to enhanced visual grounding—a limitation also observed in related vision-language tasks. In this work, we preliminarily investigate whether advanced VLN models genuinely comprehend the visual content of their environments by introducing varying levels of visual perturbations. These perturbations include ground-truth depth images, perturbed views and random noise. Surprisingly, we experimentally find that simple branch expansion, even with noisy visual inputs, paradoxically improves the navigational efficacy. Inspired by these insights, we further present a versatile Multi-Branch Architecture (MBA) designed to delve into the impact of both the branch quantity and visual quality. The proposed MBA extends a base agent into a multi-branch variant, where each branch processes a different visual input. This approach is embarrassingly simple yet agnostic to topology-based VLN agents. Extensive experiments on three VLN benchmarks (R2R, REVERIE, SOON) demonstrate that our method with optimal visual permutations matches or even surpasses state-of-the-art results. The source code is available at here. Jia Li 0013, Yunbo Xu, Zhenzhen Hu 0004, Richang Hong |
IJCNN | 2 |
| 2025 | Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor IdentificationabstractMetaphorical imagination, the ability to connect seemingly unrelated concepts, is fundamental to human cognition and communication. While understanding linguistic metaphors has advanced significantly, grasping multimodal metaphors, such as those found in internet memes, presents unique challenges due to their unconventional expressions and implied meanings. Existing methods for multimodal metaphor identification often struggle to bridge the gap between literal and figurative interpretations. Additionally, generative approaches that utilize large language models or text-to-image models, while promising, suffer from high computational costs. This paper introduces Concept Drift Guided LayerNorm Tuning (CDGLT), a novel and training-efficient framework for multimodal metaphor identification. CDGLT incorporates two key innovations: (1) Concept Drift, a mechanism that leverages Spherical Linear Interpolation (SLERP) of cross-modal embeddings from a CLIP encoder to generate a new, divergent concept embedding. This drifted concept helps to alleviate the gap between literal features and the figurative task. (2) A prompt construction strategy, that adapts the method of feature extraction and fusion using pre-trained language models for the multimodal metaphor identification task. CDGLT achieves state-of-the-art performance on the MET-Meme benchmark while significantly reducing training costs compared to existing generative methods. Ablation studies demonstrate the effectiveness of both Concept Drift and our adapted LN Tuning approach. Our method represents a significant step towards efficient and accurate multimodal metaphor understanding. The code is available: https://github.com/Qianvenh/CDGLT. Wenhao Qian, Zhenzhen Hu 0004, Zijie Song, Jia Li 0013 |
ICMR | 4 |
| 2025 | Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
Jia Li 0013, Yichao He, Jiacheng Xu 0008, Tianhao Luo, Zhenzhen Hu 0004, Richang Hong, Meng Wang 0001 |
ACM Multimedia | 1 |
| 2025 | Listening to the Unspoken: Exploring '365' Aspects of Multimodal Interview Performance AssessmentabstractInterview performance assessment is essential for determining candidates' suitability for professional positions. To ensure holistic and fair evaluations, we propose a novel and comprehensive framework that explores ''365'' aspects of interview performance by integrating three modalities (video, audio, and text), six responses per candidate, and five key evaluation dimensions. The framework employs modality-specific feature extractors to encode heterogeneous data streams and subsequently fused via a Shared Compression Multilayer Perceptron. This module compresses multimodal embeddings into a unified latent space, facilitating efficient feature interaction. To enhance prediction robustness, we incorporate a two-level ensemble learning strategy: (1) independent regression heads predict scores for each response, and (2) predictions are aggregated across responses using a mean-pooling mechanism to produce final scores for the five target dimensions. By listening to the unspoken, our approach captures both explicit and implicit cues from multimodal data, enabling comprehensive and unbiased assessments. Achieving a multi-dimensional average MSE of 0.1824, our framework secured first place in the AVI Challenge 2025, demonstrating its effectiveness and robustness in advancing automated and multimodal interview performance assessment. The full implementation is available at https://github.com/Qianvenh/AVI2025-Track2. Jia Li 0013, Yang Wang 0023, Wenhao Qian, Jialong Hu, Zhenzhen Hu 0004, Richang Hong, Meng Wang 0001 |
ACM Multimedia | 1 |
| 2025 | VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionabstractAudiovisual emotion recognition (AVER) aims to infer human emotions from nonverbal visual-audio (VA) cues, offering modality-complementary and language-agnostic advantages. However, AVER remains challenging due to the inherent ambiguity of emotional expressions, cross-modal expressive disparities, and the scarcity of reliably annotated data. Recent self-supervised AVER approaches have introduced strong multimodal representations, yet they predominantly rely on modality-specific encoders and coarse content-level alignment, limiting fine-grained emotional semantic modeling. To address these issues, we propose VAEmo, an efficient two-stage framework for emotion-centric joint VA representation learning with external knowledge injection. In Stage~1, a unified and lightweight representation network is pre-trained on large-scale speaker-centric VA corpora via masked reconstruction and contrastive objectives, mitigating the modality gap and learning expressive, complementary representations without emotion labels. In Stage~2, multimodal large language models automatically generate detailed affective descriptions according to our well-designed chain-of-thought prompting for only a small subset of VA samples; these rich textual semantics are then injected by aligning their corresponding embeddings with VA representations through dual-path contrastive learning, further bridging the emotion gap. Extensive experiments on multiple downstream AVER benchmarks show that VAEmo achieves state-of-the-art performance with a compact design, highlighting the benefit of unified cross-modal encoding and emotion-aware semantic guidance for efficient, generalizable VA emotion representations. Yichao He, Zhenzhen Hu 0004, Jia Li 0013, Meng Wang 0001, Richang Hong |
ACM Multimedia | 5 |
| 2025 | Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
Yangchen Yu, Jia Li 0013, Yu Zhang 0082, Zhenzhen Hu 0004, Meng Wang 0001, Richang Hong |
ACM Multimedia | 3 |
| 2025 | Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video RetrievalabstractRecent progress in text–video retrieval has been largely driven by contrastive learning.
However, existing methods often overlook the effect of the modality gap, which causes anchor representations to undergo in-place optimization (i.e., optimization tension) that limits their alignment capacity.
Moreover, noisy hard negatives further distort the semantics of anchors.
To address these issues, we propose GARE, a Gap-Aware Retrieval framework that introduces a learnable, pair-specific increment $\Delta_{ij}$ between text $t_i$ and video $v_j$, redistributing gradients to relieve optimization tension and absorb noise. We derive $\Delta_{ij}$ via a multivariate first-order Taylor expansion of the InfoNCE loss under a trust-region constraint, showing that it guides updates along locally consistent descent directions. A lightweight neural module conditioned on the semantic gap couples increments across batches for structure-aware correction. Furthermore, we regularize $\Delta$ through a variational information bottleneck with relaxed compression, enhancing stability and semantic consistency. Experiments on four benchmarks demonstrate that GARE consistently improves alignment accuracy and robustness, validating the effectiveness of gap-aware tension mitigation. Zijie Song, Jialong Hu, Zhenzhen Hu 0004, Jia Li 0013, Richang Hong |
NeurIPS | 6 |
| 2025 | From Static to Dynamic: Adapting Landmark-Aware Image Models for Facial Expression Recognition in VideosabstractDynamic facial expression recognition (DFER) in the wild is still hindered by data limitations, e.g., insufficient quantity and diversity of pose, occlusion and illumination, as well as the inherent ambiguity of facial expressions. In contrast, static facial expression recognition (SFER) currently shows much higher performance and can benefit from more abundant high-quality training data. Moreover, the appearance features and dynamic dependencies of DFER remain largely unexplored. Recognizing the potential in leveraging SFER knowledge for DFER, we introduce a novel Static-to-Dynamic model (S2D) that leverages existing SFER knowledge and dynamic information implicitly encoded in extracted facial landmark-aware features, thereby significantly improving DFER performance. First, we build and train an image model for SFER, which incorporates a standard Vision Transformer (ViT) and Multi-View Complementary Prompters (MCPs) only. Then, we obtain our video model (i.e., S2D), for DFER, by inserting Temporal-Modeling Adapters (TMAs) into the image model. MCPs enhance facial expression features with landmark-aware features inferred by an off-the-shelf facial landmark detector. And the TMAs capture and model the relationships of dynamic changes in facial expressions, effectively extending the pre-trained image model for videos. Notably, MCPs and TMAs only increase a fraction of trainable parameters (less than +10%) to the original image model. Moreover, we present a novel Emotion-Anchors (i.e., reference samples for each emotion category) based Self-Distillation Loss to reduce the detrimental influence of ambiguous emotion labels, further enhancing our S2D. Experiments conducted on popular SFER and DFER datasets show that we have achieved a new state of the art. Jia Li 0013, Shiguang Shan, Meng Wang 0001, Richang Hong |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Emotion Separation and Recognition From a Facial Expression by Generating the Poker Face With Vision TransformersabstractRepresentation learning and feature disentanglement have garnered significant research interest in the field of facial expression recognition (FER). The inherent ambiguity of emotion labels poses challenges for conventional supervised representation learning methods. Moreover, directly learning the mapping from a facial expression image to an emotion label lacks explicit supervision signals for capturing fine-grained facial features. In this article, we propose a novel FER model, named poker face vision transformer or PF-ViT, to address these challenges. PF-ViT aims to separate and recognize the disturbance-agnostic emotion from a static facial image by generating its corresponding poker face without the need for paired images. Inspired by the facial action coding system, we regard an expressive face as the combined result of a set of facial muscle movements on one's poker face (i.e., an emotionless face). PF-ViT utilizes vanilla vision transformers, and its components are first pretrained as masked autoencoders on a large facial expression dataset without emotion labels, yielding excellent representations. Subsequently, we train PF-ViT using a GAN framework. During training, the auxiliary task of poke face generation promotes the disentanglement between emotional and emotion-irrelevant components, guiding the FER model to holistically capture discriminative facial details. Quantitative and qualitative results demonstrate the effectiveness of our method, surpassing the state-of-the-art methods on four popular FER datasets. Jia Li 0013, Jiantao Nie, Dan Guo 0001, Richang Hong, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | Adaptive Dual Video Summarization: From Dynamic Keyframes to CaptionsabstractVideo summarization and captioning condense content by selecting keyframes and generating language descriptions, integrating both visual and textual perspectives. Existing video-and-language learning models typically select multiple frames as proxies rather than analyzing all frames, which improves computational efficiency but may not adequately represent the original content without redundancy. In this paper, we propose an adaptive dual video summarization framework and demonstrate its effectiveness within the context of video captioning. Given the video frames, we extract visual representations using a video-domain fine-tuned ViT model to narrow the domain shift. The keyframes are summarized based on the frame-level scores. To minimize the number of keyframes while ensuring captioning quality, we introduce a cross-modal video summarizer that selects the most semantically consistent frames according to pseudo score labels. Furthermore, we incorporate an adaptive keyframe selector that determines the optimal number of keyframes based on the video's complexity and content, enhancing the framework's adaptability and generalization. The proposed adaptive keyframe selector enables the framework to handle diverse video content, making it more generalizable and applicable to real-world scenarios.We designed a ranking scheme to assess the video's static appearance and temporal dynamics from score-based and time-based perspectives. To conclude, we use a lightweight LSTM decoder to generate descriptions. Experimental results on the MSR-VTT, MSVD and VATEX benchmarks demonstrate that our adaptive dual video summarization framework can effectively convey the same semantic information as the original video while using a significantly reduced number of keyframes, leading to improved video captioning performance. Zhenzhen Hu 0004, Zhenshan Wang, Jia Li 0013, Zijie Song, Richang Hong, Meng Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | DAT: Dialogue-Aware Transformer with Modality-Group Fusion for Human Engagement EstimationabstractEngagement estimation plays a crucial role in understanding human social behaviors, attracting increasing research interests in fields such as affective computing and human-computer interaction. In this paper, we propose a Dialogue-Aware Transformer framework (DAT) with Modality-Group Fusion (MGF), which relies solely on audio-visual input and is language-independent, for estimating human engagement in conversations. Specifically, our method employs a modality-group fusion strategy that independently fuses audio and visual features within each modality for each person before inferring the entire audio-visual content. This strategy significantly enhances the model's performance and robustness. Additionally, to better estimate the target participant's engagement levels, the introduced Dialogue-Aware Transformer considers both the participant's behavior and cues from their conversational partners. Our method was rigorously tested in the Multi-Domain Engagement Estimation Challenge held by MultiMediate'24, demonstrating notable improvements in engagement-level regression precision over the baseline model. Notably, our approach achieves a CCC score of 0.76 on the NoXi Base test set and an average CCC of 0.64 across the NoXi Base, NoXi-Add, and MPIIGI test sets. The source code will be available at https://github.com/MSA-LMC/DAT. Jia Li 0013, Yangchen Yu, Yu Zhang 0082, Yunbo Xu, Meng Wang 0001, Richang Hong |
ACM Multimedia | 1 |
| 2024 | Dual-Stream Keyframe Enhancement for Video Question AnsweringabstractThe redundancy in videos and the quadratic scaling with input length of Transformer models lead to the need for sampling and selection from input videos.During the selection process, differentiable Top-K algorithms are employed to ensure an end-to-end training process.However, these methods not only restrict the level at which temporal information is captured but also introduce sorting noise and inaccuracies.In this paper, we revisit the keyframe selection strategy for VideoQA and propose a novel framework named Dual-Stream Keyframe Enhancement (DSKE) incorporating the enhancement of temporal granularity.To balance end-to-end sorting and hard ranking, we employ a dual-stream keyframe selection strategy by fusing the differentiable and non-differentiable results together to achieve a unified approach.One stream is based on the approximate ranking obtained from the differentiable Top-K algorithm, while the other stream utilizes the results obtained from hard ranking.We separately train decoders on the outputs of each stream and then combine the decoder results to predict the final answer.By integrating both stream results, DSKE effectively balances the inclusion of relevant information while filtering out noise.Additionally, we capture temporal variation information by incorporating a series of overlapping sliding time windows to enrich the temporal granularity.To evaluate the effectiveness of DSKE, we conduct experiments on the NExT-QA and AGQA benchmarks.The results demonstrate that our framework significantly improves the performance of VideoQA by effectively incorporating temporal components and enhancing the keyframe ranking process. Zhenzhen Hu 0004, Jia Li 0013, Zijie Song, Richang Hong |
MMAsia | 3 |
| 2024 | Exploring and exploiting model uncertainty for robust visual question answering
Zhenzhen Hu 0004, Xun Yang 0001, Jia Li 0013, Richang Hong |
Multim. Syst. | 6 |
| 2024 | FTCM: Frequency-Temporal Collaborative Module for Efficient 3D Human Pose Estimation in VideoabstractCapturing cross-pose correlation from a sequence of frame-level 2D poses is essential for 3D human pose estimation (3D-HPE) in the video. Recent studies have shown the promising potential of modeling the pose relation with feature-mixing operations on the temporal domain. However, they seldom consider the interaction across poses in the frequency domain. This paper studies a Frequency-Temporal Collaborative Module (FTCM) to explore the feasibility of encoding the cross-pose correlations in both frequency and temporal domains. FTCM aims to jointly capture the global and local cross-pose correlations with a more lightweight network model. Specifically, FTCM splits the pose features into two groups along the channel dimension and separately models the frequency and temporal interactions across poses with different feature-mixing operations in parallel. To achieve this goal, we purposely design two pose-mixing units, i.e., the frequency pose-mixing (FPM) and the temporal pose-mixing (TPM). Particularly, FPM is designed to reap the global correlations among different pose frequencies with the representation obtained by converting the original pose signals with Fast Fourier transform (FFT). Unlike the pose-mixing used by previous methods like Transformers that influences an individual pose with all other poses, TPM locally calibrates the pose with dynamics aggregated within several adjacent poses in the temporal domain, explicitly weighting neighboring poses more with respect to the far-away ones so as to enforce a strict locality constraint. Besides, the group strategy significantly reduces the model complexity. To verify the effectiveness of FTCM, we conduct extensive experiments on two benchmarks (i.e., Human3.6M and MPI-INF-3DHP). Experimental results not only exhibit favorable accuracy/complexity trade-offs of our FTCM but also show superior or comparable performance to state-of-the-art methods on both datasets. The code and model are publicly available at:https://github.com/zhenhuat/FTCM. Zhenhua Tang 0001, Yanbin Hao, Jia Li 0013, Richang Hong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Robust facial expression recognition with global-local joint representation learning
Chunxiao Fan 0002, Jia Li 0013, Xiao Sun 0003 |
Multim. Syst. | 3 |
| 2023 | Exploring Sparse Spatial Relation in Graph Inference for Text-Based VQAabstractText-based visual question answering (TextVQA) faces the significant challenge of avoiding redundant relational inference. To be specific, a large number of detected objects and optical character recognition (OCR) tokens result in rich visual relationships. Existing works take all visual relationships into account for answer prediction. However, there are three observations: (1) a single subject in the images can be easily detected as multiple objects with distinct bounding boxes (considered repetitive objects). The associations between these repetitive objects are superfluous for answer reasoning; (2) two spatially distant OCR tokens detected in the image frequently have weak semantic dependencies for answer reasoning; and (3) the co-existence of nearby objects and tokens may be indicative of important visual cues for predicting answers. Rather than utilizing all of them for answer prediction, we make an effort to identify the most important connections or eliminate redundant ones. We propose a sparse spatial graph network (SSGN) that introduces a spatially aware relation pruning technique to this task. As spatial factors for relation measurement, we employ spatial distance, geometric dimension, overlap area, and DIoU for spatially aware pruning. We consider three visual relationships for graph learning: object-object, OCR-OCR tokens, and object-OCR token relationships. SSGN is a progressive graph learning architecture that verifies the pivotal relations in the correlated object-token sparse graph, and then in the respective object-based sparse graph and token-based sparse graph. Experiment results on TextVQA and ST-VQA datasets demonstrate that SSGN achieves promising performances. And some visualization results further demonstrate the interpretability of our method. Dan Guo 0001, Jia Li 0013, Xun Yang 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | MLP-JCG: Multi-Layer Perceptron With Joint-Coordinate Gating for Efficient 3D Human Pose EstimationabstractVarious structural relations/dependencies exist among human body joints, which makes it possible to estimate 3D poses from 2D sources. The current research on 3D human pose estimation (3D-HPE for short) mainly focuses on structural information from a specific perspective. However, this information cannot facilitate 2D-to-3D pose lifting. This paper presents a novel and efficient multi-layer perceptron with a joint-coordinate gating (MLP-JCG) model, exploring and utilizing both the local and global structural information to perform 3D pose estimations. Specifically, MLP-JCG contains two independent MLP blocks, i.e., joint-mixing MLP and coordinate-mixing MLP, which solely act on the joint and coordinate axes in modelling their local structural information. For the global structural information, we first explore two kinds of global statistics from the pose matrix embeddings, which are referred to as the dynamics aggregated along the joint/coordinate axis. Then, we propose two kinds of gating units to elementwisely contextualize the features learned from MLP blocks. All the model components are designed based on MLP, making the MLP-JCG easy to implement and train. We conduct experiments on three 3D-HPE benchmarks, and the results demonstrate the superior effectiveness and efficiency of the proposed approach. Zhenhua Tang 0001, Jia Li 0013, Yanbin Hao, Richang Hong |
IEEE Trans. Multim. | 2 |
| 2022 | Multi-stage and multi-branch network with similar expressions label distribution learning for facial expression recognition
Junjie Lang, Xiao Sun 0003, Jia Li 0013, Meng Wang 0001 |
Pattern Recognit. Lett. | 3 |
| 2022 | MAN: Mining Ambiguity and Noise for Facial Expression Recognition in the Wild
Xiao Sun 0003, Jia Li 0013, Meng Wang 0001 |
Pattern Recognit. Lett. | 3 |
| 2022 | Multi-Person Pose Estimation With Accurate Heatmap Regression and Greedy AssociationabstractMulti-person pose estimation aims at localizing the 2D keypoints (or body joints) for all the people in the image. There are mainly two paradigms to perform this task: top-down and bottom-up. In this paper, we present an advanced bottom-up approach based on accurate keypoint heatmap regression and greedy keypoint association. Firstly, we develop an encoding-decoding method with Gaussian heatmaps and guiding offset fields to represent multi-person pose information, encompassing keypoint positions and adjacent keypoint associations of all individuals in the scene. In particular, we analyze the deficiency of the Gaussian heatmap representation as regards keypoint localization precision if conventional element-wise$L_{2}$-type loss is employed merely for heatmap supervision. Therefore, we introduce a peak regularization loss to jointly supervise the heatmap regression. In addition, we present an improved Hourglass Network with multi-scale heatmap aggregation to simultaneously infer the said encoding. Finally, we propose a novel focal$L_{2}$loss to help the network cope with the imbalanced problem of keypoint detection in heatmaps. Our results show that the proposed approach surpasses other bottom-up approaches on COCO dataset, and even outperforms the top-down approaches on CrowdPose dataset containing more crowded scenes. Jia Li 0013, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Emotional Conversation Generation Based on a Bayesian Deep Neural NetworkabstractThe field of conversation generation using neural networks has attracted increasing attention from researchers for several years. However, traditional neural language models tend to generate a generic reply with poor semantic logic and no emotion. This article proposes an emotional conversation generation model based on a Bayesian deep neural network that can generate replies with rich emotions, clear themes, and diverse sentences. The topic and emotional keywords of the replies are pregenerated by introducing commonsense knowledge in the model. The reply is divided into multiple clauses, and then a multidimensional generator based on the transformer mechanism proposed in this article is used to iteratively generate clauses from two dimensions: sentence granularity and sentence structure. Subjective and objective experiments prove that compared with existing models, the proposed model effectively improves the semantic logic and emotional accuracy of replies. This model also significantly enhances the diversity of replies, largely overcoming the shortcomings of traditional models that generate safe replies. Xiao Sun 0003, Jia Li 0013, Xing Wei 0002, Changliang Li, Jianhua Tao 0001 |
ACM Trans. Inf. Syst. | 2 |
| 2019 | Monocular Depth Estimation as Regression of Classification using Piled Residual NetworksabstractPredicting depth from single monocular image is a challenging task in scene understanding. Most existing work predicts depth by regression or classification with features extracted from local neighborhood area. However, neither regression nor classification achieves the final satisfying solution and local context can be insufficient to predict the depth. This paper innovatively addresses this problem as regression of class related features on a piled residual convolutional neural network. Our framework works at two stages. First, a well-designed deep convolutional neural network model is employed to classify the depths in difference-scale invariance space. The model utilizes all scales of context though piled residual paths. The deeper layers that capture high-level semantic features with long-range context can be directly refined using fine-grained features with local context from earlier convolutions. We then apply centered information gain loss to the model to produce intra-class compact and inter-class discriminative features. Second, to obtain depths instead of class labels, we infer depth regression with convolutional layers which model the mapping from class discriminative features to continuous depth values. Experiments on the popular indoor and outdoor datasets show competitive results compared with the recent state of the art methods. Wen Su 0004, Haifeng Zhang 0006, Jia Li 0013, Wenzhen Yang, Zengfu Wang |
ACM Multimedia | 3 |
| 2019 | Real-Time Traffic Sign Recognition Based on Efficient CNNs in the WildabstractBoth unmanned vehicles and driver assistance systems require solving the problem of traffic sign recognition. A lot of work has been done in this area, but no approach has been presented to perform the task with high accuracy and high speed under various conditions until now. In this paper, we have designed and implemented a detector by adopting the framework of faster R-convolutional neural networks (CNN) and the structure of MobileNet. Here, color and shape information have been used to refine the localizations of small traffic signs, which are not easy to regress precisely. Finally, an efficient CNN with asymmetric kernels is used to be the classifier of traffic signs. Both the detector and the classifier have been trained on challenging public benchmarks. The results show that the proposed detector can detect all categories of traffic signs. The detector and the classifier proposed here are proved to be superior to the state-of-the-art method. Our code and results are available online. Jia Li 0013, Zengfu Wang |
IEEE Trans. Intell. Transp. Syst. | 1 |