VLDB 2026 Research / reviewers in the wild / expert
Zhenzhen Hu 0004
dblp:136/0899-4
· DBLP profile ↗
47ranked-venue papers
8as first author
34since 2021 · last 2026
0000-0003-1042-8361ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 first-author · 23 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 7 since 2021Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language NavigationabstractNavigating unseen environments based on natural language instructions remains difficult for egocentric agents in Vision-and-Language Navigation (VLN). Intuitively, humans inherently ground concrete semantic knowledge within spatial layouts during indoor navigation. Although previous studies have introduced diverse environmental representations to enhance reasoning, other co-occurrence modalities are often naively concatenated with RGB features, resulting in suboptimal utilization of each modality's distinct contribution. Inspired by this, we propose a hierarchical Semantic Understanding and Spatial Awareness (SUSA) architecture to enable agents to perceive and ground environments at diverse scales. Specifically, the Textual Semantic Understanding (TSU) module supports local action prediction by generating view-level descriptions, thereby capturing fine-grained environmental semantics and narrowing the modality gap between instructions and environments. Complementarily, the Depth-enhanced Spatial Perception (DSP) module incrementally constructs a trajectory-level depth exploration map, providing the agent with a coarse-grained comprehension of the global spatial layout. Extensive experiments demonstrate that SUSA's hierarchical representation enrichment not only boosts the navigation performance of the baseline on discrete VLN benchmarks (REVERIE, R2R, and SOON), but also exhibits superior generalization to the continuous R2R-CE. Yunbo Xu, Jia Li 0013, Zhenzhen Hu 0004 |
AAAI | 5 |
| 2026 | Fine-grained Text-Video Retrieval with Patch-level Temporal Difference and AggregationabstractExisting Text-Video Retrieval (TVR) methods predominantly rely on global frame representations, often disregarding the fine-grained temporal variations required for precise patch-level alignment. This is critical as video motion is inherently spatially localized; consequently, coarse frame-level modeling tends to be dominated by static backgrounds, overshadowing salient action cues. To address this limitation, we propose TRFG, a novel framework for text-video retrieval that addresses the challenges of modeling Temporal Reasoning and Fine-Grained cross-modal alignment. First, our Temporal Difference module captures frame-to-frame variations at the patch level, effectively suppressing static background noise to highlight "active" motion regions. Second, these differential signals are synthesized via a Temporal Aggregation module to form a coherent representation of the event’s trajectory. Finally, to ensure precise semantic matching, a fine-grained interaction module aligns these dynamic video tokens with textual details. Extensive experiments on MSRVTT, ActivityNet, and DiDeMo demonstrate that TRFG achieves state-of-the-art performance across multiple backbones and retrieval tasks. Ablation studies confirm the complementarity and generalizability of both modules, underscoring the importance of explicit temporal modeling and fine-grained interaction in bridging the modality gap. Jialong Hu, Zijie Song, Yang Wang 0023, Zhenzhen Hu 0004, Jia Li 0013, Richang Hong |
ICMR | 4 |
| 2026 | MetaPipe: Predicting Metaphoric Associations and Turning Your Metaphoric Imagination into RealityabstractVisual metaphors, in particular, are powerful tools to communicate complex ideas effectively. However, creating meaningful and visually compelling metaphors is inherently difficult, especially for amateur designers, who often lack the tools and expertise to form strong conceptual connections. To address these challenges, we introduce MetaPipe, a mobile-based design tool that assists users in crafting high-quality visual metaphors. By leveraging three interactive modules, MetaPipe simplifies the creation process, making it accessible to non-experts. We evaluated MetaPipe through two user studies: (1) a usability study with 16 non-design professionals rating their experience of the tool, and (2) a blinded experiment in which 74 participants assessing creativity and metaphorical strength of MetaPipe-generated images. The results demonstrate that MetaPipe significantly enhances the visual metaphor design process, enabling the creation of high-quality and creative images. The code is available: https://github.com/Qianvenh/MetaPipe. Wenhao Qian, Zhenzhen Hu 0004, Jialong Hu, Jia Li 0013, Richang Hong |
Int. J. Hum. Comput. Interact. | 2 |
| 2026 | CLAIP-Emo: Parameter-Efficient Adaptation of Language-Supervised Models for In-the-Wild Audiovisual Emotion RecognitionabstractAudiovisual emotion recognition (AVER) in the wild is still hindered by pose variation, occlusion, and background noise. Prevailing methods primarily rely on large-scale domain-specific pre-training, which is costly and often mismatched to real-world affective data. To address this, we present CLAIP-Emo, a modular framework that reframes in-the-wild AVER as a parameter-efficient adaptation of language-supervised foundation models (CLIP/CLAP). Specifically, it (i) preserves language-supervised priors by freezing CLIP/CLAP backbones and performing emotion-oriented adaptation via LoRA (updating \ensuremath{\le}4.0\% of the total parameters), (ii) allocates temporal modeling asymmetrically, employing a lightweight Transformer for visual dynamics while applying mean pooling for audio prosody, and (iii) applies a simple fusion head for prediction. On DFEW and MAFW, CLAIP-Emo (ViT-L/14) achieves 80.14\% and 61.18\% weighted average recall with only 8M training parameters, setting a new state of the art. Our findings suggest that parameter-efficient adaptation of language-supervised foundation models provides a scalable alternative to domain-specific pre-training for real-world AVER. The code and models will be available at \href{https://github.com/MSA-LMC/CLAIP-Emo}{https://github.com/MSA-LMC/CLAIP-Emo}. Jia Li 0013, Jinpeng Hu, Zhenzhen Hu 0004, Richang Hong |
IEEE Signal Process. Lett. | 4 |
| 2026 | Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression DataabstractDynamic facial expression recognition (DFER) infers emotions from the temporal evolution of expressions, unlike static facial expression recognition (SFER), which relies solely on a single snapshot. This temporal analysis provides richer information and promises greater recognition capability. However, current DFER methods often exhibit unsatisfied performance largely due to fewer training samples compared to SFER. Given the inherent correlation between static and dynamic expressions, we hypothesize that leveraging the abundant SFER data can enhance DFER. To this end, we propose Static-for-Dynamic (S4D), a unified dual-modal learning framework that integrates SFER data as a complementary resource for DFER. Specifically, S4D employs dual-modal self-supervised pre-training on facial images and videos using a shared Vision Transformer (ViT) encoder-decoder architecture, yielding improved spatiotemporal representations. The pre-trained encoder is then fine-tuned on static and dynamic expression datasets in a multi-task learning setup to facilitate emotional information interaction. Unfortunately, vanilla multi-task learning in our study results in negative transfer. To address this, we propose an innovative Mixture of Adapter Experts (MoAE) module that facilitates task-specific knowledge acquisition while effectively extracting shared knowledge from both static and dynamic expression data. Extensive experiments demonstrate that S4D achieves a deeper understanding of DFER, setting new state-of-the-art performance on FERV39K, MAFW, and DFEW benchmarks, with weighted average recall (WAR) of 53.65%, 58.44%, and 76.68%, respectively. Additionally, a systematic correlation analysis between SFER and DFER tasks is presented, which further elucidates the potential benefits of leveraging SFER. Jia Li 0013, Yu Zhang 0082, Zhenzhen Hu 0004, Shiguang Shan, Meng Wang 0001, Richang Hong |
IEEE Trans. Affect. Comput. | 4 |
| 2026 | PhysioSync: Temporal and Cross-Modal Contrastive Learning Inspired by Physiological Synchronization for EEG-Based Emotion RecognitionabstractElectroencephalography (EEG) signals provide a promising and involuntary reflection of brain activity related to emotional states, offering significant advantages over behavioral cues such as facial expressions. However, EEG signals are often noisy, affected by artifacts, and vary across individuals, complicating emotion recognition. While multimodal approaches have used peripheral physiological signals (PPS) such as galvanic skin response to complement EEG, they often overlook the dynamic synchronization and consistent semantics between the modalities. Additionally, the temporal dynamics of emotional fluctuations across different time resolutions in PPS remain underexplored. To address these challenges, we propose PhysioSync, a novel pretraining framework leveraging temporal and cross-modal contrastive learning (CM-CL), inspired by physiological synchronization phenomena. PhysioSync incorporates cross-modal consistency alignment (CM-CA) to model dynamic relationships between EEG and complementary PPS, enabling emotion-related synchronizations across modalities. Besides, it introduces long- and short-term temporal contrastive learning (LS-TCL) to capture emotional synchronization at different temporal resolutions within modalities. After pretraining, cross-resolution and cross-modal features are hierarchically fused and fine-tuned to enhance emotion recognition. Experiments on DEAP and DREAMER datasets demonstrate PhysioSync’s advanced performance under unimodal and cross-modal conditions, highlighting its effectiveness for EEG-centered emotion recognition. Jia Li 0013, Yu Liu 0023, Zhenzhen Hu 0004, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Image Super-Resolution Using Hierarchical Cross-Scale Self-SimilarityabstractPrevious studies have revealed that extending the spatial range of informative pixels offers positive performance gains for image Super-Resolution (SR). To activate more informative pixels, considerable efforts have been devoted to exploring various variants of non-local attention mechanisms for capturing image self-similarity. However, even the state-of-the-art non-local attention mechanisms ignore an inherent property of images, namely hierarchical cross-scale self-similarity. In this paper, we propose the first Hierarchical Cross-Scale Attention (HCSA). Specifically, we first extend the search space to multiple feature maps from a single feature map, and then model cross-scale feature correspondences among different layers. This allows HCSA to activate more informative pixels for image SR by adaptively rescaling and aggregating input pixels and large-scale patches within different feature maps. To ensure accurate cross scale feature matching, we propose to replace plain down sampling operations (e.g., interpolation, pooling) with Haar Wavelet Transform (HWT) encoding, which transfers spatial information of feature maps into the channel dimension, effectively avoiding important information loss. Considering that softmax normal ization in the standard non-local attention often leads to homogeneous feature aggregation due to the amplification of small similarity weights, we propose a simple yet effective Adaptive Selection (AS) operator. This operator generates a learnable sparse mask to remove redundant features, enabling HCSA to perform discriminative feature aggregation. As a generic building block, the proposed HCSA can be flexibly integrated into existing CNN- or Transformer-based SR models, significantly strengthening cross-layer information interaction and cross-scale feature representation. Quantitative and qualitative results demonstrate that our HCSA facilitates existing SR models to achieve superior accuracy and visual quality. Xiancheng Zhu, Detian Huang, Taiheng Zeng, Xiaoqian Huang, Zhenzhen Hu 0004, Huanqiang Zeng |
IEEE Trans. Multim. | 5 |
| 2025 | Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video RetrievalabstractText-video retrieval (TVR) has seen substantial advancements in recent years, fueled by the utilization of pre-trained models and large language models (LLMs). Despite these advancements, achieving accurate matching in TVR remains challenging due to inherent disparities between video and textual modalities and irregularities in data representation. In this paper, we propose Text-Video-ProxyNet (TV-ProxyNet), a novel framework designed to decompose the conventional 1-to-N relationship of TVR into N distinct 1-to-1 relationships. By replacing a single text query with a series of text proxies, TV-ProxyNet not only broadens the query scope but also achieves a more precise expansion. Each text proxy is crafted through a refined iterative process, controlled by mechanisms we term as the director and dash, which regulate the proxy's direction and distance relative to the original text query. This setup not only facilitates more precise semantic alignment but also effectively manages the disparities and noise inherent in multimodal data. Our experiments on three representative video-text retrieval benchmarks, MSRVTT, DiDeMo, and ActivityNet Captions, demonstrate the effectiveness of TV-ProxyNet. The results show an improvement of 2.0% to 3.3% in R@1 over the baseline. TV-ProxyNet achieved state-of-the-art performance on MSRVTT and ActivityNet Captions, and a 2.0% improvement on DiDeMo compared to existing methods, validating our approach's ability to enhance semantic mapping and reduce error propensity. Zhenzhen Hu 0004, Jia Li 0013, Richang Hong |
AAAI | 2 |
| 2025 | Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQAabstractVideo Question Answering (VideoQA) is a complex video-language task that demands a sophisticated understanding of both visual content and temporal dynamics. Traditional Transformer-style architectures, while effective in integrating multimodal data, often simplify temporal dynamics through positional encoding and fail to capture non-linear interactions within video sequences. In this paper, we introduce the Temporal Trio Transformer (T3T), a novel architecture that models time consistency and time variability. The T3T integrates three key components: Temporal Smoothing (TS), Temporal Difference (TD), and Temporal Fusion (TF). The TS module employs Brownian Bridge for capturing smooth, continuous temporal transitions, while the TD module identifies and encodes significant temporal variations and abrupt changes within the video content. Subsequently, the TF module synthesizes these temporal features with textual cues, facilitating a deeper contextual understanding and response accuracy. The efficacy of the T3T is demonstrated through extensive testing on multiple VideoQA benchmark datasets. Our results underscore the importance of a nuanced approach to temporal modeling in improving the accuracy and depth of video-based question answering. Zijie Song, Zhenzhen Hu 0004, Jia Li 0013, Richang Hong |
ICME | 2 |
| 2025 | Seeing is Believing? Enhancing Vision-Language Navigation using Visual PerturbationsabstractAutonomous navigation guided by natural language instructions in embodied environments remains a challenge for vision-language navigation (VLN) agents. Although recent advancements in learning diverse and fine-grained visual environmental representations have shown promise, the fragile performance improvements may not conclusively attribute to enhanced visual grounding—a limitation also observed in related vision-language tasks. In this work, we preliminarily investigate whether advanced VLN models genuinely comprehend the visual content of their environments by introducing varying levels of visual perturbations. These perturbations include ground-truth depth images, perturbed views and random noise. Surprisingly, we experimentally find that simple branch expansion, even with noisy visual inputs, paradoxically improves the navigational efficacy. Inspired by these insights, we further present a versatile Multi-Branch Architecture (MBA) designed to delve into the impact of both the branch quantity and visual quality. The proposed MBA extends a base agent into a multi-branch variant, where each branch processes a different visual input. This approach is embarrassingly simple yet agnostic to topology-based VLN agents. Extensive experiments on three VLN benchmarks (R2R, REVERIE, SOON) demonstrate that our method with optimal visual permutations matches or even surpasses state-of-the-art results. The source code is available at here. Jia Li 0013, Yunbo Xu, Zhenzhen Hu 0004, Richang Hong |
IJCNN | 4 |
| 2025 | Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor IdentificationabstractMetaphorical imagination, the ability to connect seemingly unrelated concepts, is fundamental to human cognition and communication. While understanding linguistic metaphors has advanced significantly, grasping multimodal metaphors, such as those found in internet memes, presents unique challenges due to their unconventional expressions and implied meanings. Existing methods for multimodal metaphor identification often struggle to bridge the gap between literal and figurative interpretations. Additionally, generative approaches that utilize large language models or text-to-image models, while promising, suffer from high computational costs. This paper introduces Concept Drift Guided LayerNorm Tuning (CDGLT), a novel and training-efficient framework for multimodal metaphor identification. CDGLT incorporates two key innovations: (1) Concept Drift, a mechanism that leverages Spherical Linear Interpolation (SLERP) of cross-modal embeddings from a CLIP encoder to generate a new, divergent concept embedding. This drifted concept helps to alleviate the gap between literal features and the figurative task. (2) A prompt construction strategy, that adapts the method of feature extraction and fusion using pre-trained language models for the multimodal metaphor identification task. CDGLT achieves state-of-the-art performance on the MET-Meme benchmark while significantly reducing training costs compared to existing generative methods. Ablation studies demonstrate the effectiveness of both Concept Drift and our adapted LN Tuning approach. Our method represents a significant step towards efficient and accurate multimodal metaphor understanding. The code is available: https://github.com/Qianvenh/CDGLT. Wenhao Qian, Zhenzhen Hu 0004, Zijie Song, Jia Li 0013 |
ICMR | 2 |
| 2025 | Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
Jia Li 0013, Yichao He, Jiacheng Xu 0008, Tianhao Luo, Zhenzhen Hu 0004, Richang Hong, Meng Wang 0001 |
ACM Multimedia | 5 |
| 2025 | Listening to the Unspoken: Exploring '365' Aspects of Multimodal Interview Performance AssessmentabstractInterview performance assessment is essential for determining candidates' suitability for professional positions. To ensure holistic and fair evaluations, we propose a novel and comprehensive framework that explores ''365'' aspects of interview performance by integrating three modalities (video, audio, and text), six responses per candidate, and five key evaluation dimensions. The framework employs modality-specific feature extractors to encode heterogeneous data streams and subsequently fused via a Shared Compression Multilayer Perceptron. This module compresses multimodal embeddings into a unified latent space, facilitating efficient feature interaction. To enhance prediction robustness, we incorporate a two-level ensemble learning strategy: (1) independent regression heads predict scores for each response, and (2) predictions are aggregated across responses using a mean-pooling mechanism to produce final scores for the five target dimensions. By listening to the unspoken, our approach captures both explicit and implicit cues from multimodal data, enabling comprehensive and unbiased assessments. Achieving a multi-dimensional average MSE of 0.1824, our framework secured first place in the AVI Challenge 2025, demonstrating its effectiveness and robustness in advancing automated and multimodal interview performance assessment. The full implementation is available at https://github.com/Qianvenh/AVI2025-Track2. Jia Li 0013, Yang Wang 0023, Wenhao Qian, Jialong Hu, Zhenzhen Hu 0004, Richang Hong, Meng Wang 0001 |
ACM Multimedia | 5 |
| 2025 | VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionabstractAudiovisual emotion recognition (AVER) aims to infer human emotions from nonverbal visual-audio (VA) cues, offering modality-complementary and language-agnostic advantages. However, AVER remains challenging due to the inherent ambiguity of emotional expressions, cross-modal expressive disparities, and the scarcity of reliably annotated data. Recent self-supervised AVER approaches have introduced strong multimodal representations, yet they predominantly rely on modality-specific encoders and coarse content-level alignment, limiting fine-grained emotional semantic modeling. To address these issues, we propose VAEmo, an efficient two-stage framework for emotion-centric joint VA representation learning with external knowledge injection. In Stage~1, a unified and lightweight representation network is pre-trained on large-scale speaker-centric VA corpora via masked reconstruction and contrastive objectives, mitigating the modality gap and learning expressive, complementary representations without emotion labels. In Stage~2, multimodal large language models automatically generate detailed affective descriptions according to our well-designed chain-of-thought prompting for only a small subset of VA samples; these rich textual semantics are then injected by aligning their corresponding embeddings with VA representations through dual-path contrastive learning, further bridging the emotion gap. Extensive experiments on multiple downstream AVER benchmarks show that VAEmo achieves state-of-the-art performance with a compact design, highlighting the benefit of unified cross-modal encoding and emotion-aware semantic guidance for efficient, generalizable VA emotion representations. Yichao He, Zhenzhen Hu 0004, Jia Li 0013, Meng Wang 0001, Richang Hong |
ACM Multimedia | 4 |
| 2025 | Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
Yangchen Yu, Jia Li 0013, Yu Zhang 0082, Zhenzhen Hu 0004, Meng Wang 0001, Richang Hong |
ACM Multimedia | 7 |
| 2025 | Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video RetrievalabstractRecent progress in text–video retrieval has been largely driven by contrastive learning.
However, existing methods often overlook the effect of the modality gap, which causes anchor representations to undergo in-place optimization (i.e., optimization tension) that limits their alignment capacity.
Moreover, noisy hard negatives further distort the semantics of anchors.
To address these issues, we propose GARE, a Gap-Aware Retrieval framework that introduces a learnable, pair-specific increment $\Delta_{ij}$ between text $t_i$ and video $v_j$, redistributing gradients to relieve optimization tension and absorb noise. We derive $\Delta_{ij}$ via a multivariate first-order Taylor expansion of the InfoNCE loss under a trust-region constraint, showing that it guides updates along locally consistent descent directions. A lightweight neural module conditioned on the semantic gap couples increments across batches for structure-aware correction. Furthermore, we regularize $\Delta$ through a variational information bottleneck with relaxed compression, enhancing stability and semantic consistency. Experiments on four benchmarks demonstrate that GARE consistently improves alignment accuracy and robustness, validating the effectiveness of gap-aware tension mitigation. Zijie Song, Jialong Hu, Zhenzhen Hu 0004, Jia Li 0013, Richang Hong |
NeurIPS | 5 |
| 2025 | EPDiff: Enhancing Prior-guided Diffusion model for Real-world Image Super-ResolutionabstractDiffusion Models (DMs) have achieved promising success in Real-world Image Super-Resolution (Real-ISR), where they reconstruct High-Resolution (HR) images from available Low-Resolution (LR) counterparts with unknown degradation by leveraging pre-trained Text-to-Image (T2I) diffusion models. However, due to the randomness nature of DMs and the severe degradation commonly presented in LR images, most DMs-based Real-ISR methods neglect the structure-level and semantic information, which results in reconstructed HR images suffering not only from important edge missing, but also from undesired regional information confusion. To tackle these challenges, we propose an Enhancing Prior-guided Diffusion model (EPDiff) for Real-ISR, which leverages high-frequency priors and semantic guidance to generate reconstructed images with realistic details. Firstly, we design a Guide Adapter (GA) module that extracts latent texture and edge features from LR images to provide high-frequency priors. Subsequently, we introduce a Semantic Prompt Extractor (SPE) that generates high-quality semantic prompts to enhance image understanding. Additionally, we build a Feature Rectify ControlNet (FRControlNet) to refine feature modulation, enabling realistic detail generation. Extensive experiments demonstrate that the proposed EPDiff outperforms state-of-the-art methods on both synthetic and real-world datasets. Detian Huang, Miaohua Ruan, Yaohui Guo, Zhenzhen Hu 0004, Huanqiang Zeng |
Comput. Vis. Image Underst. | 4 |
| 2025 | Grid Jigsaw Representation with CLIP: a new perspective on image clustering
Zijie Song, Zhenzhen Hu 0004, Richang Hong |
Multim. Syst. | 2 |
| 2025 | Multi-Modal Prior-Guided Diffusion Model for Blind Image Super-ResolutionabstractRecently, diffusion models have achieved remarkable success in blind image super-resolution. However, most existing methods rely solely on uni-modal degraded low-resolution images to guide diffusion models for restoring high-fidelity images, resulting in inferior realism. In this letter, we propose a Multi-modal Prior-Guided diffusion model for blind image Super-Resolution (MPGSR), which fine-tunes Stable Diffusion (SD) by utilizing the superior visual-and-textual guidance for restoring realistic high-resolution images. Specifically, our MPGSR involves two stages, i.e., multi-modal guidance extraction and adaptive guidance injection. For the former, we propose a composited transformer and further incorporate it with GPT-CLIP to extract the representative visual-and-textual guidance. For the latter, we design a feature calibration ControlNet to inject the visual guidance and employ the cross-attention layer provided by the frozen SD to inject the textual guidance, thus effectively activating the powerful text-to-image generation potential. Extensive experiments show that our MPGSR outperforms state-of-the-art methods in restoration quality and convergence time. Detian Huang, Jiaxun Song, Xiaoqian Huang, Zhenzhen Hu 0004, Huanqiang Zeng |
IEEE Signal Process. Lett. | 4 |
| 2025 | Adaptive Dual Video Summarization: From Dynamic Keyframes to CaptionsabstractVideo summarization and captioning condense content by selecting keyframes and generating language descriptions, integrating both visual and textual perspectives. Existing video-and-language learning models typically select multiple frames as proxies rather than analyzing all frames, which improves computational efficiency but may not adequately represent the original content without redundancy. In this paper, we propose an adaptive dual video summarization framework and demonstrate its effectiveness within the context of video captioning. Given the video frames, we extract visual representations using a video-domain fine-tuned ViT model to narrow the domain shift. The keyframes are summarized based on the frame-level scores. To minimize the number of keyframes while ensuring captioning quality, we introduce a cross-modal video summarizer that selects the most semantically consistent frames according to pseudo score labels. Furthermore, we incorporate an adaptive keyframe selector that determines the optimal number of keyframes based on the video's complexity and content, enhancing the framework's adaptability and generalization. The proposed adaptive keyframe selector enables the framework to handle diverse video content, making it more generalizable and applicable to real-world scenarios.We designed a ranking scheme to assess the video's static appearance and temporal dynamics from score-based and time-based perspectives. To conclude, we use a lightweight LSTM decoder to generate descriptions. Experimental results on the MSR-VTT, MSVD and VATEX benchmarks demonstrate that our adaptive dual video summarization framework can effectively convey the same semantic information as the original video while using a significantly reduced number of keyframes, leading to improved video captioning performance. Zhenzhen Hu 0004, Zhenshan Wang, Jia Li 0013, Zijie Song, Richang Hong, Meng Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Dual-Stream Keyframe Enhancement for Video Question AnsweringabstractThe redundancy in videos and the quadratic scaling with input length of Transformer models lead to the need for sampling and selection from input videos.During the selection process, differentiable Top-K algorithms are employed to ensure an end-to-end training process.However, these methods not only restrict the level at which temporal information is captured but also introduce sorting noise and inaccuracies.In this paper, we revisit the keyframe selection strategy for VideoQA and propose a novel framework named Dual-Stream Keyframe Enhancement (DSKE) incorporating the enhancement of temporal granularity.To balance end-to-end sorting and hard ranking, we employ a dual-stream keyframe selection strategy by fusing the differentiable and non-differentiable results together to achieve a unified approach.One stream is based on the approximate ranking obtained from the differentiable Top-K algorithm, while the other stream utilizes the results obtained from hard ranking.We separately train decoders on the outputs of each stream and then combine the decoder results to predict the final answer.By integrating both stream results, DSKE effectively balances the inclusion of relevant information while filtering out noise.Additionally, we capture temporal variation information by incorporating a series of overlapping sliding time windows to enrich the temporal granularity.To evaluate the effectiveness of DSKE, we conduct experiments on the NExT-QA and AGQA benchmarks.The results demonstrate that our framework significantly improves the performance of VideoQA by effectively incorporating temporal components and enhancing the keyframe ranking process. Zhenzhen Hu 0004, Jia Li 0013, Zijie Song, Richang Hong |
MMAsia | 1 |
| 2024 | Exploring and exploiting model uncertainty for robust visual question answering
Zhenzhen Hu 0004, Xun Yang 0001, Jia Li 0013, Richang Hong |
Multim. Syst. | 4 |
| 2024 | Efficiently Gluing Pre-Trained Language and Vision Models for Image CaptioningabstractVision-and-language pre-training models have achieved impressive performance for image captioning. But most of them are trained with millions of paired image-text data and require huge memory and computing overhead. To alleviate this, we try to stand on the shoulders of large-scale pre-trained language models (PLM) and pre-trained vision models (PVM) and efficiently connect them for image captioning. There are two major challenges: one is that language and vision modalities have different semantic granularity (e.g., a noun may cover many pixels), and the other is that the semantic gap still exists between the pre-trained language and vision models. To this end, we design a lightweight and efficient connector to glue PVM and PLM, which holds a criterion of selection-then-transformation . Specifically, in the selection phase, we treat each image as a set of patches instead of pixels. We select salient image patches and cluster them into visual regions to align with text. Then, to effectively reduce the semantic gap, we propose to map the selected image patches into text space through spatial and channel transformations. With training on image captioning datasets, the connector learns to bridge the semantic granularity and semantic gap via backpropagation, preparing for the PLM to generate descriptions. Experimental results on the MSCOCO and Flickr30k datasets demonstrate that our method yields comparable performance to existing works. By solely training the small connector, we achieve a CIDEr performance of 132.2% on the MSCOCO Karpathy test split. Moreover, our findings reveal that fine-tuning the PLM can further enhance performance potential, resulting in a CIDEr score of 140.6%. Code and models are available at https://github.com/YuanEZhou/PrefixCap . Peipei Song, Yuanen Zhou, Xun Yang 0001, Daqing Liu, Zhenzhen Hu 0004, Depeng Wang, Meng Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2024 | Math Word Problem Generation via Disentangled Memory RetrievalabstractThe task of math word problem (MWP) generation, which generates an MWP given an equation and relevant topic words, has increasingly attracted researchers’ attention. In this work, we introduce a simple memory retrieval module to search related training MWPs, which are used to augment the generation. To retrieve more relevant training data, we also propose a disentangled memory retrieval module based on the simple memory retrieval module. To this end, we first disentangle the training MWPs into logical description and scenario description and then record them in respective memory modules. Later, we use the given equation and topic words as queries to retrieve relevant logical descriptions and scenario descriptions from the corresponding memory modules, respectively. The retrieved results are then used to complement the process of the MWP generation. Extensive experiments and ablation studies verify the superior performance of our method and the effectiveness of each proposed module. The code is available at https://github.com/mwp-g/MWPG-DMR . Zhenzhen Hu 0004, Lei Wang 0185, Yunshi Lan, Richang Hong |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Embedded Heterogeneous Attention Transformer for Cross-Lingual Image CaptioningabstractCross-lingual image captioning is a challenging task that requires addressing both cross-lingual and cross-modal obstacles in multimedia analysis. The crucial issue in this task is to model the global and the local matching between the image and different languages. Existing cross-modal embedding methods based on the transformer architecture oversee the local matching between the image region and monolingual words, especially when dealing with diverse languages. To overcome these limitations, we propose an Embedded Heterogeneous Attention Transformer (EHAT) to establish cross-domain relationships and local correspondences between images and different languages by using a heterogeneous network. EHAT comprises Masked Heterogeneous Cross-attention (MHCA), Heterogeneous Attention Reasoning Network (HARN), and Heterogeneous Co-attention (HCA). The HARN serves as the core network and it captures cross-domain relationships by leveraging visual bounding box representation features to connect word features from two languages and to learn heterogeneous maps. MHCA and HCA facilitate cross-domain integration in the encoder through specialized heterogeneous attention mechanisms, enabling a single model to generate captions in two languages. We evaluate our approach on the MSCOCO dataset to generate captions in English and Chinese, two languages that exhibit significant differences in their language families. The experimental results demonstrate the superior performance of our method compared to existing advanced monolingual methods. Our proposed EHAT framework effectively addresses the challenges of cross-lingual image captioning, paving the way for improved multilingual image analysis and understanding. Zijie Song, Zhenzhen Hu 0004, Yuanen Zhou, Ye Zhao 0001, Richang Hong, Meng Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Dual Video Summarization: From Frames to CaptionsabstractVideo summarization and video captioning both condense the video content from the perspective of visual and text modes, i.e. the keyframe selection and language description generation. Existing video-and-language learning models commonly sample multiple frames for training instead of observing all. These sampled deputies greatly improve computational efficiency, but do they represent the original video content enough with no more redundancy? In this work, we propose a dual video summarization framework and verify it in the context of video captioning. Given the video frames, we firstly extract the visual representation based on the ViT model fine-tuned on the video-text domain. Then we summarize the keyframes according to the frame-lever score. To compress the number of keyframes as much as possible while ensuring the quality of captioning, we learn a cross-modal video summarizer to select the most semantically consistent frames according to the pseudo score label. Top K frames ( K is no more than 3% of the entire video.) are chosen to form the video representation. Moreover, to evaluate the static appearance and temporal information of video, we design the ranking scheme of video representation from two aspects: feature-oriented and sequence-oriented. Finally, we generate the descriptions with a lightweight LSTM decoder. The experiment results on the MSR-VTT and MSVD dataset reveal that, for the generative task as video captioning, a small number of keyframes can convey the same semantic information to perform well on captioning, or even better than the original sampling. Zhenzhen Hu 0004, Zhenshan Wang, Zijie Song, Richang Hong |
IJCAI | 1 |
| 2023 | Grid Feature Jigsaw for Self-supervised Image ClusteringabstractImage clustering is an essential unsupervised learning task in computer vision. The key issue for image clustering is to learn the representative visual features without annotations to extend class spacing. Jigsaw puzzles, as one pretext task of self-supervised visual representation learning to learn the relative spatial position of image tiles, has attracted the attention of many researchers. Most of existing jigsaw puzzle solving strategies are based on the raw image patches, i.e. original pixels, which makes them only concentrate on the low-level statistics. In this paper, we propose a novel learning strategy named Grid Feature Jigsaw (GFJ) for self-supervised image clustering to increase class spacing by deep mining of single sample feature. We train the model to learn the intra-grid representation via the self-supervised paradigm. By dividing the feature map into grids and arranging adjacent grids in each block, we implement a linear regression from the surrounding grids to represent the reference grid. The experiments of unsupervised computer vision benchmark show the effectiveness on the clustering task with respect to the ACC, NMI and ARI three metrics and we verify GFJ universal performance via various deep convolutional neural networks. Zijie Song, Zhenzhen Hu 0004, Richang Hong |
IJCNN | 2 |
| 2023 | Efficient and self-adaptive rationale knowledge base for visual commonsense reasoning
Zijie Song, Zhenzhen Hu 0004, Richang Hong |
Multim. Syst. | 2 |
| 2023 | A Text-Guided Generation and Refinement Model for Image CaptioningabstractA high-quality image description requires not only the logic and fluency of language but also the richness and accuracy ofcontent. However, due to the semantic gap between vision and language, most existing image captioning approaches thatdirectly learn the cross-modal mapping from vision to language are difficult to meet these two requirements simultaneously. Inspired by the progressive learning mechanism, we trace the “generating + refining” route and propose a novel Text-GuidedGeneration and Refinement (dubbed as TGGAR) model with assistance from the guide text to improve the quality of captions.The guide text is selected from the training set according to content similarity, then utilized to explore salient objects andextend candidate words. Specifically, we follow the encoderdecoder architecture, and design a Text-Guided Relation Encoder(TGRE) to learn the visual representation that is more consistent with human visual cognition. Besides, we divide the decoderpart into two sub-modules: a Generator for the primary sentence generation and a Refiner for the sentence refinement.Generator, consisting of a standard LSTM and a Gate on Attention (GOA) module, aims to generate the primary sentencelogically and fluently. Refiner contains a caption encoder module, an attentionbased LSTM and a GOA module, whichiteratively modifies the details in the primary caption to make captions rich and accurate. Extensive experiments on theMSCOCO captioning dataset demonstrate our framework with fewer parameters remains comparable to transformer-basedmethods, and achieves state-of-the-art performance compared with other relevant approaches. Depeng Wang, Zhenzhen Hu 0004, Yuanen Zhou, Richang Hong, Meng Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | OCR-oriented Master Object for Text Image CaptioningabstractText image captioning aims to understand the scene text in images for image caption generation. The key issue of this challenging task is to understand the relationship between the text OCR tokens and images. In this paper, we propose a novel text image captioning method by purifying the OCR-oriented scene graph with themaster object. The master object is the object to which the OCR is attached, which is the semantic relationship bridge between the OCR token and the image. We consider the master object as a proxy to connect OCR tokens and other regions in the image. By exploring the master object for each OCR token, we build the purified scene graph based on the master objects and then enrich the visual embedding by the Graph Convolution Network (GCN). Furthermore, we cluster the OCR tokens and feed the hierarchical information to provide a richer representation. Experiments on the TextCaps validation and test dataset demonstrate the effectiveness of the proposed method. Wenliang Tang, Zhenzhen Hu 0004, Zijie Song, Richang Hong |
ICMR | 2 |
| 2022 | Math Word Problem Generation with Memory Retrieval
Zhenzhen Hu 0004, Lei Wang 0185, Yunshi Lan, Richang Hong |
PRCV (3) | 3 |
| 2022 | Visual feature synthesis with semantic reconstructor for traditional and generalized zero-shot object classificationabstractZero-shot learning (ZSL) addresses the novel object recognition problem by leveraging semantic embedding to transfer knowledge from seen categories to unseen categories. Generative ZSL models synthesize the visual features of unseen classes and convert ZSL task into a classical supervised learning problem. These generative ZSL models are trained by using the seen classes. Although promising progress has been achieved in the ZSL and generalized zero-shot learning (GZSL) tasks. The existing approaches still suffer from a strong bias problem between unseen and seen classes, where unseen objects in the target domain tend to be recognized as seen classes in the source domain. To deal with the problem, we propose a novel named semantic consistent Wasserstein generative adversarial network (scWGAN), which uses a semantic reconstructor to reconstruct semantic embeddings from generated visual features by incorporating a novel Semantic Consistent Loss noted L rec . The Semantic Consistent Loss guides our proposed scWGAN to generate visual features that mirror the semantic relationships between seen and unseen classes. We also introduce a visual classifier to constrain visual feature generator. Extensive experiments show that the proposed approach is superior to previous state-of-the-art works under both traditional ZSL and challenging GZSL settings on six popular data sets AWA1, AWA2, CUB, APY, and SUN. Ye Zhao 0001, Xueliang Liu, Dan Guo 0001, Zhenzhen Hu 0004, Hengchang Liu, Yicong Li 0004 |
Int. J. Intell. Syst. | 5 |
| 2021 | Sequential image encoding for vision-to-language problems
Yuanen Zhou, Zhenzhen Hu 0004, Meng Wang 0001 |
Multim. Tools Appl. | 3 |
| 2021 | Adversarial co-distillation learning for image recognition
Zhenzhen Hu 0004, Mingliang Xu 0001, Meng Wang 0001 |
Pattern Recognit. | 2 |
| 2020 | More Grounded Image Captioning by Distilling Image-Text Matching ModelabstractVisual attention not only improves the performance of image captioners, but also serves as a visual interpretation to qualitatively measure the caption rationality and model transparency. Specifically, we expect that a captioner can fix its attentive gaze on the correct objects while generating the corresponding words. This ability is also known as grounded image captioning. However, the grounding accuracy of existing captioners is far from satisfactory.To improve the grounding accuracy while retaining the captioning quality, it is expensive to collect the word-region alignment as strong supervision.To this end, we propose a Part-of-Speech (POS) enhanced image-text matching model (SCAN[24]): POS-SCAN, as the effective knowledge distillation for more grounded image captioning. The benefits are two-fold: 1) given a sentence and an image, POS-SCAN can ground the objects more accurately than SCAN; 2) POS-SCAN serves as a word-region alignment regularization for the captioner's visual attention module. By showing benchmark experimental results, we demonstrate that conventional image captioners equipped with POS-SCAN can significantly improve the grounding accuracy without strong supervision. Last but not the least, we explore the indispensable Self-Critical Sequence Training (SCST[46]) in the context of grounded image captioning and show that the image-text matching score can serve as a reward for more grounded captioning1. Yuanen Zhou, Meng Wang 0001, Daqing Liu, Zhenzhen Hu 0004, Hanwang Zhang |
CVPR | 4 |
| 2020 | WFN-PSC: weighted-fusion network with poly-scale convolution for image dehazingabstractImage dehazing is a fundamental task for the computer vision and multimedia and usually in the face of the challenge from two aspects, i) the uneven distribution of arbitrary haze and ii) the distortion of image pixels caused by the hazed image. In this paper, we propose an end-to-end trainable framework, named Weighted-Fusion Network with Poly-Scale Convolution (WFN-PSC), to address these dehazing issues. The proposed method is designed based on the Poly-Scale Convolution (PSConv). It can extract the image feature from different scales without upsampling and downsampled, which avoids the image distortion. Beyond this, we design the spatial and channel weighted-fusion modules to make the WFN-PSC model focus on the hard dehazing parts of image from two dimensions. Specifically, we design three Part Architectures followed by the channel weighted-fusion module. Each Part Architecture consists of three PSConv residual blocks and a spatial weighted-fusion module. The experiments on the benchmark demonstrate the dehazing effectiveness of the proposed method. Furthermore, considering that image dehazing is a low-level task in the computer vision, we evaluate the dehazed image on the object detection task and the results show that the proposed method can be a good pre-processing to assist the high-level computer vision task. Lexuan Sun, Xueliang Liu, Zhenzhen Hu 0004, Richang Hong |
MMAsia | 3 |
| 2019 | Quality-Aware Unpaired Image-to-Image TranslationabstractGenerative adversarial networks (GANs) have been widely used for the image-to-image translation task. While these models rely heavily on the labeled image pairs, recently some GAN variants have been proposed to tackle the unpaired image translation task. These models exploited supervision at the domain level with a reconstruction process for unpaired image translation. On the other hand, parallel works have shown that leveraging perceptual loss functions based on high-level deep features could enhance the generated image quality. Nevertheless, as these GAN-based models either depended on the pretrained deep network structure or relied on the labeled image pairs, they could not be directly applied to the unpaired image translation task. Moreover, despite the improvement of the introduced perceptual losses from deep neural networks, few researchers have explored the possibility of improving the generated image quality from classical image quality measures. To tackle the above two challenges, in this paper, we propose a unified quality-aware GAN-based framework for unpaired image-to-image translation, where a quality-aware loss is explicitly incorporated by comparing each source image and the reconstructed image at the domain level. Specifically, we design two detailed implementations of the quality loss. The first method is based on a classical image quality assessment measure by defining a classical quality-aware loss to ensure similar quality score between an original image and the reconstructed image at the domain level. The second method proposes an adaptive deep network based loss that compares the high level content structure between each original image and its reconstructed image from the generator. Finally, extensive experimental results on many real-world datasets clearly show the quality improvement of our proposed framework, and the superiority of leveraging classical image quality measures for unpaired image translation compared to the deep network based model. Lei Chen 0051, Le Wu 0001, Zhenzhen Hu 0004, Meng Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Speeding-Up Age Estimation in Intelligent Demographics System via Network OptimizationabstractAge estimation is a difficult task which requires the automatic detection and interpretation of facial features. Recently, Convolutional Neural Networks (CNNs) have made remarkable improvement on learning age patterns from benchmark datasets. However, for a face ``in the wild'' (from a video frame or Internet), the existing algorithms are not as accurate as for a frontal and neutral face. In addition, with the increasing number of in-the- wild aging data, the computation speed of existing deep learning platforms becomes another crucial issue. In this paper, we propose a high-efficient age estimation system with joint optimization of age estimation algorithm and deep learning system. Cooperated with the city surveillance network, this system can provide age group analysis for intelligent demographics. First, we build a three- tier fog computing architecture including an edge, a fog and a cloud layer, which directly processes age estimation from raw videos. Second, we optimize the age estimation algorithm based on CNNs with label distribution and K-L divergence distance embedded in the fog layer and evaluate the model on the latest wild aging dataset. Experimental results demonstrate that: 1. our system collects the demographics data dynamically at far-distance without contact, and makes the city population analysis automatically; and 2. the age model training has been speed-up without losing training progress or model quality. To our best knowledge, this is the first intelligent demographics system which has potential applications in improving the efficiency of smart cities and urban living. Zhenzhen Hu 0004, Peng Sun 0006, Yonggang Wen 0001 |
ICC | 1 |
| 2018 | Semantic Image Inpainting with Progressive Generative NetworksabstractRecently, image inpainting task has revived with the help of deep learning techniques. Deep neural networks, especially the generative adversarial networks~(GANs) make it possible to recover the missing details in images. Due to the lack of sufficient context information, most existing methods fail to get satisfactory inpainting results. This work investigates a more challenging problem, e.g., the newly-emerging semantic image inpainting - a task to fill in large holes in natural images. In this paper, we propose an end-to-end framework named progressive generative networks~(PGN), which regards the semantic image inpainting task as a curriculum learning problem. Specifically, we divide the hole filling process into several different phases and each phase aims to finish a course of the entire curriculum. After that, an LSTM framework is used to string all the phases together. By introducing this learning strategy, our approach is able to progressively shrink the large corrupted regions in natural images and yields promising inpainting results. Moreover, the proposed approach is quite fast to evaluate as the entire hole filling is performed in a single forward pass. Extensive experiments on Paris Street View and ImageNet dataset clearly demonstrate the superiority of our approach. Code for our models is available at https://github.com/crashmoon/Progressive-Generative-Networks. Zhenzhen Hu 0004, Changzhi Luo, Wangmeng Zuo, Meng Wang 0001 |
ACM Multimedia | 2 |
| 2017 | Facial Age Estimation With Age DifferenceabstractAge estimation based on the human face remains a significant problem in computer vision and pattern recognition. In order to estimate an accurate age or age group of a facial image, most of the existing algorithms require a huge face data set attached with age labels. This imposes a constraint on the utilization of the immensely unlabeled or weakly labeled training data, e.g., the huge amount of human photos in the social networks. These images may provide no age label, but it is easy to derive the age difference for an image pair of the same person. To improve the age estimation accuracy, we propose a novel learning scheme to take advantage of these weakly labeled data through the deep convolutional neural networks. For each image pair, Kullback-Leibler divergence is employed to embed the age difference information. The entropy loss and the cross entropy loss are adaptively applied on each image to make the distribution exhibit a single peak value. The combination of these losses is designed to drive the neural network to understand the age gradually from only the age difference information. We also contribute a data set, including more than 100 000 face images attached with their taken dates. Each image is both labeled with the timestamp and people identity. Experimental results on two aging face databases show the advantages of the proposed age difference learning system, and the state-of-the-art performance is gained. Zhenzhen Hu 0004, Yonggang Wen 0001, Meng Wang 0001, Richang Hong, Shuicheng Yan |
IEEE Trans. Image Process. | 1 |
| 2017 | Visual Classification of Furniture StylesabstractFurniture style describes the discriminative appearance characteristics of furniture. It plays an important role in real-world indoor decoration. In this article, we explore the furniture style features and study the problem of furniture style classification. Differing from traditional object classification, furniture style classification aims at classifying different furniture in terms of the “style” that describes its appearance (e.g., American style, Gothic style, Rococo style, etc.) rather than the “kind” that is more related to its functional structure (e.g., bed, desk, etc.). To pursue efficient furniture style features, we construct a novel dataset of furniture styles that contains 16 common style categories and implement three strategies with respect to two categories of classification, that is, handcrafted classification and learning-based classification. First, we follow the typical image classification pipeline to extract the handcrafted features and train the classifier by support vector machine. Then we use the convolutional neural network to extract learning-based features from training images. To obtain comprehensive furniture style features, we finally combine the handcrafted image classification pipeline and the learning-based network. We experimentally evaluate the performances of handcrafted features and learning-based features of each strategy, and the results show the superiority of learning-based features and also the comprehensiveness of handcrafted features. Zhenzhen Hu 0004, Yonggang Wen 0001, Luoqi Liu, Richang Hong, Meng Wang 0001, Shuicheng Yan |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2016 | Multi-View Object Retrieval via Multi-Scale Topic ModelsabstractThe increasing number of 3D objects in various applications has increased the requirement for effective and efficient 3D object retrieval methods, which attracted extensive research efforts in recent years. Existing works mainly focus on how to extract features and conduct object matching. With the increasing applications, 3D objects come from different areas. In such circumstances, how to conduct object retrieval becomes more important. To address this issue, we propose a multi-view object retrieval method using multi-scale topic models in this paper. In our method, multiple views are first extracted from each object, and then the dense visual features are extracted to represent each view. To represent the 3D object, multi-scale topic models are employed to extract the hidden relationship among these features with respect to varied topic numbers in the topic model. In this way, each object can be represented by a set of bag of topics. To compare the objects, we first conduct topic clustering for the basic topics from two data sets, and then generate the common topic dictionary for new representation. Then, the two objects can be aligned to the same common feature space for comparison. To evaluate the performance of the proposed method, experiments are conducted on two data sets. The 3D object retrieval experimental results and comparison with existing methods demonstrate the effectiveness of the proposed method. Richang Hong, Zhenzhen Hu 0004, Ruxin Wang 0002, Meng Wang 0001, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2015 | Understanding Blooming Human Groups in Social NetworksabstractHuman group, which indicates the people who share similar characteristics, is used to categorize humans into distinct populations or groups. In recent years, with the explosive growth of image, new concepts of human group are blooming in social networks . People in the same human group can be categorized by their facial and clothes appearance characteristics. In this work, we propose an approach to understanding the new concepts of human group with few positive samples. To this end, we construct visual models crossing two modalities related to human images and surrounding texts. Two convolutional neural networks based on face and upper body are constructed separately. Two different convolutional neural networks (CNNs) architectures are explored for visual pre-traing. To assist the human group recognition, we also merge global convolutional feature of the image. The surrounding texts are represented by semantical vectors and utilized as image labels. We transform words in the text into fixed length vectors by the skip-gram model. Then the texts corresponding to each image are converted into one feature vector by sparse coding and max pooling. Given a few positive samples of new concepts of human group, the visual model can be improved to understand the semantical meaning of the new label. The experimental results demonstrate the effectiveness of the proposed visual model and show the excellent learning capacity with few samples. Richang Hong, Zhenzhen Hu 0004, Luoqi Liu, Meng Wang 0001, Shuicheng Yan, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2014 | PicWords: Render a Picture by Packing KeywordsabstractIn this paper, we propose a novel text-art system: input a source picture and some keywords introducing the information about the picture, and the output is the so-called PicWords in the form of the source picture composed of the introduction keywords. Different from traditional text-graphics which are created by highly skilled artists and involve a huge amount of tedious manual work, PicWords is an automatic non-photorealistic rendering (NPR) packing system. Given a source picture, we first generate its silhouette, which is a binary image containing a Yang part and a Yin part. Yang part is for keywords placing while the Yin part can be ignored. Next, the Yang part is further over-segmented into small patches, each of which serves as a container for one keyword. To make sure that more important keywords are put into more salient and larger image patches, we rank both the patches and keywords and construct a correspondence between the patch list and keyword list. Then, mean value coordinates method is used for the keyword-patch warping. Finally, certain post-processing techniques are adopted to improve the aesthetics of PicWords. Extensive experimental results well demonstrate the effectiveness of the proposed PicWords system. Zhenzhen Hu 0004, Si Liu 0001, Richang Hong, Meng Wang 0001, Shuicheng Yan |
IEEE Trans. Multim. | 1 |
| 2014 | Fashion Parsing With Weak Color-Category LabelsabstractIn this paper we address the problem of automatically parsing the fashion images with weak supervision from the user-generated color-category tags such as “red jeans” and “white T-shirt”. This problem is very challenging due to the large diversity of fashion items and the absence of pixel-level tags, which make the traditional fully supervised algorithms inapplicable. To solve the problem, we propose to combine the human pose estimation module, the MRF-based color and category inference module and the (super)pixel-level category classifier learning module to generate multiple well-performing category classifiers, which can be directly applied to parse the fashion items in the images. Besides, all the training images are parsed with color-category labels and the human poses of the images are estimated during the model learning phase in this work. We also construct a new fashion image dataset called Colorful-Fashion, in which all 2,682 images are labeled with pixel-level color-category labels. Extensive experiments on this dataset clearly show the effectiveness of the proposed method for the weakly supervised fashion parsing task. Si Liu 0001, Jiashi Feng, Csaba Domokos, Junshi Huang, Zhenzhen Hu 0004, Shuicheng Yan |
IEEE Trans. Multim. | 6 |
| 2013 | eHeritage of shadow puppetry: creation and manipulationabstractIn this demo, we propose the puppetry eHeritage, including a creator module and a manipulator module, to preserve the precious traditional heritage Chinese shadow puppetry. The creator module accepts a frontal face image and a profile face image of the user as input, and automatically generates the corresponding puppet, which looks like the original person and meanwhile preserves typical characteristics of traditional Chinese shadow puppetry. The manipulator module can accept the script provided by the user as input and automatically generate the motion sequences. For better visual effects, we propose the sparsity optimization over simplexes formulation. Zhenzhen Hu 0004, Si Liu 0001, Meng Wang 0001, Richang Hong, Shuicheng Yan |
ACM Multimedia | 1 |
| 2013 | eHeritage of shadow puppetry: creation and manipulationabstractTo preserve the precious traditional heritage Chinese shadow puppetry, we propose the puppetry eHeritage, including a creator module and a manipulator module. The creator module accepts a frontal view face image and a profile face image of the user as input, and automatically generates the corresponding puppet, which looks like the original person and meanwhile has some typical characteristics of traditional Chinese shadow puppetry. In order to create the puppet, we first extract the central profile curve and warp the reference puppet eye and eyebrow to the shape of the frontal view eye and eyebrow. Then we transfer the puppet texture to the real face area. The manipulator module can accept the script provided by the user as input and automatically generate the motion sequences. Technically, we first learn atomic motions from a set of shadow puppetry videos. A scripting system converts the user's input to atomic motions, and finally synthesizes the animation based on the atomic motion instances. For better visual effects, we propose the sparsity optimization over simplexes formulation to automatically assemble weighted instances of different atomic actions into a smooth shadow puppetry animation sequence. We evaluate the performance of the creator module and the manipulator module sequentially. Extensive experimental results on the creation of puppetry characters and puppetry plays well demonstrate the effectiveness of the proposed system. Zhenzhen Hu 0004, Si Liu 0001, Meng Wang 0001, Richang Hong, Shuicheng Yan |
ACM Multimedia | 2 |