VLDB 2026 Research / reviewers in the wild / expert
Tingting Han 0003
dblp:38/3003-3
· DBLP profile ↗
26ranked-venue papers
16as first author
15since 2021 · last 2026
0000-0002-2131-9200ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 8 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Emotional conflict adaptation for multimodal sentiment analysis
Tingting Han 0003, Lingyun Yu 0004, Min Tan 0005, Zhou Yu 0001, Hongxun Yao |
Pattern Recognit. | 1 |
| 2026 | HEART: Emotionally Grounded Video Captioning via Hierarchical Emotion-Aligned RepresentationabstractEmotional Video Captioning (EVC) seeks to generate video descriptions that are both factually accurate and emotionally expressive. However, existing approaches often lack structured semantic grounding and fine-grained temporal modeling, leading to incomplete or emotionally inconsistent captions. To address these issues, we proposeHEART(HierarchicalEmotion-AlignedRepresentation withTemporal structure), a unified framework that jointly models hierarchical visual semantics and multi-scale temporal context. Specifically, HEART introduces a Hierarchical Semantic Extraction Module that decomposes visual content into entity-, action-, and event-level representations, providing a rich foundation for multi-level emotional alignment. A Temporal Pyramid Module captures short- and long-range temporal dependencies through multi-scale convolution, enabling temporally coherent captioning. Together, these components enable HEART to generate captions that are both emotionally grounded and temporally complete. To support this framework, we construct EmoStruct, a new benchmark dataset with fine-grained emotional annotations at the subject and predicate levels. Experiments on EmoStruct and public datasets demonstrate that HEART significantly outperforms prior methods in both semantic and emotional dimensions. Tingting Han 0003, Yuxuan Gong, Sicheng Zhao, Min Tan 0005, Zhou Yu 0001, Hongxun Yao |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion AdaptationabstractVisual emotion recognition (VER), which aims at understanding humans' emotional reactions toward different visual stimuli, has attracted increasing attention. Given the subjective and ambiguous characteristics of emotion, annotating a reliable large-scale dataset is hard. For reducing reliance on data labeling, domain adaptation offers an alternative solution by adapting models trained on labeled source data to unlabeled target data. Conventional domain adaptation methods require access to source data. However, due to privacy concerns, source emotional data may be inaccessible. To address this issue, we propose an unexplored task: source-free domain adaptation (SFDA) for VER, which does not have access to source data during the adaptation process. To achieve this, we propose a novel framework termed Bridge then Begin Anew (BBA), which consists of two steps: domain-bridged model generation (DMG) and target-related model adaptation (TMA). First, the DMG bridges cross-domain gaps by generating an intermediate model, avoiding direct alignment between two VER datasets with significant differences. Then, the TMA begins training the target model anew to fit the target structure, avoiding the influence of source-specific knowledge. Extensive experiments are conducted on six SFDA settings for VER. The results demonstrate the effectiveness of BBA, which achieves remarkable performance gains compared with state-of-the-art SFDA methods and outperforms representative unsupervised domain adaptation approaches. Jiankun Zhu, Sicheng Zhao, Wenbo Tang, Zhaopan Xu, Tingting Han 0003, Pengfei Xu 0001, Hongxun Yao |
AAAI | 6 |
| 2025 | MMCNav: MLLM-empowered Multi-agent Collaboration for Outdoor Visual Language Navigation
Suguo Zhu, Tingting Han 0003, Zhou Yu 0001 |
ICMR | 4 |
| 2025 | FutureGS: Structured Gaussian Fields for Future-Aware Dynamic Scene Modeling
Mingyang Ding, Tingting Han 0003, Jiajun Ding, Min Tan 0005, Zhenzhong Kuang |
ACM Multimedia | 4 |
| 2025 | FedGA: Federated Learning via Gradient Adaptive AggregationabstractIn modern lives, the rapid proliferation of Internet of Things (IoT) devices has made them indispensable tools for data collection and analysis across various domains. However, growing concerns over data ownership and privacy have hindered effective data sharing among IoT devices, leading to the persistent challenge of data silos. Federated Learning (FL) has emerged as a promising solution to this problem by enabling collaborative model training without direct data exchange. Despite its potential, FL faces two critical limitations: severe catastrophic forgetting for historical knowledge and inefficient average aggregation. To address these challenges, this paper proposes FedGA, an innovative FL framework that leverages cosine similarity-based weighted aggregation to enhance model convergence speed. Furthermore, FedGA incorporates a mechanism to memorize historical models, thereby significantly alleviating catastrophic forgetting. Extensive experiments on three public datasets validate the effectiveness of FedGA, demonstrating its superior performance in both accuracy and training efficiency compared to state-of-the-art methods. The results highlight FedGA’s capability to overcome the key shortcomings of existing FL approaches, making it a robust solution for practical IoT applications. Changfeng Hu, Min Tan 0005, Tingting Han 0003, Zhenzhong Kuang |
SMC | 4 |
| 2025 | Modality-aware contrast and fusion for multi-modal summarization
Lixin Dai, Tingting Han 0003, Zhou Yu 0001, Jun Yu 0002, Min Tan 0005 |
Neurocomputing | 2 |
| 2025 | Adversarial temporal sentence grounding by learning from external data
Tingting Han 0003, Kai Wang 0036, Jun Yu 0002, Sicheng Zhao, Jianping Fan 0001 |
Pattern Recognit. | 1 |
| 2025 | Action-Driven Semantic Representation and Aggregation for Video CaptioningabstractVideo captioning, a challenging task that entails generating natural language descriptions of visual content, often fails to effectively grasp the essence of action semantics. To harness the power of action detection to facilitate a deeper understanding of the video content, we propose an action-driven method, named Hierarchical Semantic Representation and Aggregation (HSRA) network. This method explicitly exploits action clues with a hierarchical semantic representation module, which models visual semantics in a three-level structure: “object-action-event”. By employing learnable action queries, our approach injects extensive action semantics into the model, thereby enabling more accurate and context-rich captions. To further enhance semantic alignment and understanding, we introduce a semantic aggregation composed of a semantic interaction module and a semantic refinement module. This component facilitates the alignment of semantics across different levels and emphasizes key information, ultimately leading to significant improvements in semantic consistency between the video and generated captions. We performed extensive evaluations on two well-established public datasets, MSVD and MSR-VTT, and the findings consistently demonstrate that our proposed HSRA network outperforms contemporary state-of-the-art methods. Tingting Han 0003, Yaochen Xu, Jun Yu 0002, Zhou Yu 0001, Sicheng Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability
Buyu Liu, Kai Wang 0036, Jun Bao, Tingting Han 0003, Jun Yu 0002 |
ACM Multimedia | 5 |
| 2024 | Multimodal Federated Learning Via Local-Global FusionabstractThe number of Internet of Things (IoT) devices across diverse domains in modern life has made them significant sources for collecting and analyzing multi-modal data. However, concerns about ownership and data privacy associated with IoT devices make data sharing among multiple devices impractical. Recently, multimodal federated learning has emerged as an innovative solution where each device client can collectively train a satisfactory local model without exchanging local data. Nevertheless, most existing multimodal federated learning approaches prioritize training a powerful global server model while neglecting the performance of local client models. In this context, this paper introduces FedAF, a multimodal Federated learning approach via feature Fusion with adversarial representation learning, aimed at enhancing local representations and thereby improving local client models. Specifically, using the trained global model, FedAF integrates the global representation of each local client's data into its local feature obtained from the corresponding client model. Furthermore, domain adversarial learning is employed to align global and local representations by minimizing the discrepancy between local and global encoders, compelling the global encoder to adapt to local tasks. Comprehensive experiments on two unimodal classifications and one multimodal retrieval dataset demonstrate that FedAF achieves state-of-the-art performance compared to other federated learning methods and significantly improves local client models while maintaining the satisfactory performance of the global server model. Zilin Xia, Min Tan 0005, Lingqiang Chu, Tingting Han 0003 |
SMC | 5 |
| 2024 | Effective Video Summarization by Extracting Parameter-Free Motion AttentionabstractVideo summarization remains a challenging task despite increasing research efforts. Traditional methods focus solely on long-range temporal modeling of video frames, overlooking important local motion information that cannot be captured by frame-level video representations. In this article, we propose the Parameter-free Motion Attention Module (PMAM) to exploit the crucial motion clues potentially contained in adjacent video frames, using a multi-head attention architecture. The PMAM requires no additional training for model parameters, leading to an efficient and effective understanding of video dynamics. Moreover, we introduce the Multi-feature Motion Attention Network (MMAN), integrating the PMAM with local and global multi-head attention based on object-centric and scene-centric video representations. The synergistic combination of local motion information, extracted by the proposed PMAM, with long-range interactions modeled by the local and global multi-head attention mechanism, can significantly enhance the performance of video summarization. Extensive experimental results on the benchmark datasets, SumMe and TVSum, demonstrate that the proposed MMAN outperforms other state-of-the-art methods, resulting in remarkable performance gains. Tingting Han 0003, Jun Yu 0002, Zhou Yu 0001, Sicheng Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Contrastive Perturbation Network for Weakly Supervised Temporal Sentence Grounding
Tingting Han 0003, Yuanxin Lv, Zhou Yu 0001, Jun Yu 0002, Jianping Fan 0001 |
PRCV (1) | 1 |
| 2022 | Modeling long-term video semantic distribution for temporal action proposal generation
Tingting Han 0003, Sicheng Zhao, Xiaoshuai Sun, Jun Yu 0002 |
Neurocomputing | 1 |
| 2022 | Weakly supervised moment localization with natural language based on semantic reconstruction
Tingting Han 0003, Kai Wang 0036, Jun Yu 0002, Jianping Fan 0001 |
Image Vis. Comput. | 1 |
| 2020 | Actionness-pooled Deep-convolutional Descriptor for fine-grained action recognition
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Wenlong Xie, Sicheng Zhao, Wei Yu 0004 |
Neurocomputing | 1 |
| 2020 | TVENet: Temporal variance embedding network for fine-grained action representation
Tingting Han 0003, Hongxun Yao, Wenlong Xie, Xiaoshuai Sun, Sicheng Zhao, Jun Yu 0002 |
Pattern Recognit. | 1 |
| 2019 | Discovering Latent Discriminative Patterns for Multi-Mode Event RepresentationabstractRepresentation of videos is essential since it conveys an understanding of video content and enables many higher level tasks to be tackled efficiently. However, it is challenging to propose a rational representation for complex event videos, as most video information is either noisy or redundant. In this paper, we propose a compact event representation method that can concisely describe the inner modes of events. We deem that an optimal event representation scheme should reflect the long-term and high-level visual semantics (visual topics) of events, so different from previous frame-level video semantics representation methods and concept-based video representation methods, we investigate the problem from the perspective of segment-level video representations. We then present three appealing properties of segment-level visual semantics. Based on the observation, we propose different algorithms that rely on a novel deep-visual-word-based video encoding method to discover latent discriminative patterns of events. Finally, our multi-mode event representation is obtained by concatenating the discovered patterns as inner modes. We adopt our event representation for representative event parts mining, which can highlight the visual topics of events and remarkably prune the raw videos. We validate our event representation method based on complex event detection task. Experimental results on two standard benchmarking datasets, MED11 and CCV Dataset, show that the proposed method can significantly outperform the state-of-the-art approaches. Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Tingting Han 0003, Sicheng Zhao, Tat-Seng Chua |
IEEE Trans. Multim. | 4 |
| 2018 | Add: Actionness-Pooled Deep-Convolutional DescriptorabstractRecognition of general actions has achieved great breakthroughs in recent years. However, in real-world applications, finer-grained action classification is often needed. The major challenge is that fine-grained actions usually share high similarities in both appearance and motion pattern, making it difficult to distinguish them with existing general action representation. To solve this problem, we introduce visual attention mechanism into the proposed descriptor, termed as Actionness-pooled Deep-convolutional Descriptor (ADD). Instead of pooling features uniformly from the entire video, we aggregate features in sub-regions that are more likely to contain actions according to actionness maps, which endow ADD with the capability of capturing the subtle differences between fine-grained actions. We conduct experiments on HIT Dances dataset, one of the few existing datasets for fine-grained action analysis. Quantitative results have demonstrated that ADD remarkably outperforms traditional two-stream representation. Extensive experiments on two general action benchmarks, JHMDB and UCF101, have additionally proved that combining ADD with end-to-end ConvNet can further boost the recognition performance. Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Wenlong Xie, Yanhao Zhang 0001 |
ICME | 1 |
| 2018 | Event patches: Mining effective parts for event detection and understanding
Wenlong Xie, Hongxun Yao, Sicheng Zhao, Xiaoshuai Sun, Tingting Han 0003 |
Signal Process. | 5 |
| 2017 | Dancelets Mining for Video Recommendation Based on Dance StylesabstractDance is a unique and meaningful type of human expression, composed of abundant and various action elements. However, existing methods based on associated texts and spatial visual features have difficulty capturing the highly articulated motion patterns. To overcome this limitation, we propose to take advantage of the intrinsic motion information in dance videos to solve the video recommendation problem. We present a novel system that recommends dance videos based on a mid-level action representation, termed Dancelets. The Dancelets are used to bridge the semantic gap between video content and high-level concept, dance style, which plays a significant role in characterizing different types of dances. The proposed method executes automatic mining of dancelets with a concatenation of normalized cut clustering and linear discriminant analysis. This ensures that the discovered dancelets are both representative and discriminative. Additionally, to exploit the motion cues in videos, we employ motion boundaries as saliency priors to generate volumes of interest and extract C3D features to capture spatiotemporal information from the mid-level patches. Extensive experiments validated on our proposed large dance dataset, HIT Dances dataset, demonstrate the effectiveness of the proposed methods for dance style-based video recommendation. Tingting Han 0003, Hongxun Yao, Chenliang Xu, Xiaoshuai Sun, Yanhao Zhang 0001, Jason J. Corso |
IEEE Trans. Multim. | 1 |
| 2016 | Mining representative actions for actor identificationabstractPrevious works on actor identification mainly focused on static features based on face identification and costume detection, without considering the abundant dynamic information contained in videos. In this paper, we propose a novel method to mine representative actions of each actor, and show the remarkable power of such actions for actor identification task. Videos are firstly divided into shots and represented by BoW based on spatial-temporal features. Then we integrate the prototype theory with SVM to rank the shots and obtain the representative actions. Our method for actor identification combines representative actions with actors' appearance. We validate the method on episodes of the TV series "The Big Bang Theory". The experimental results show that the representative actions are consistent with human judgements and can greatly improve the matching performance as complementary to existing handcrafted static features for actor identification. Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Tingting Han 0003 |
ICASSP | 5 |
| 2016 | Unsupervised discovery of crowd activities by saliency-based clustering
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Yanhao Zhang 0001 |
Neurocomputing | 1 |
| 2015 | "Clustering of Dancelets": Towards Video Recommendation Based on Dance StylesabstractDance is a special and important type of action, composed of abundant and various action elements. However, the recommendation of dance videos on the web are still not well studied. It is hard to realize it in the way of traditional methods using associated texts or static features of video content. In this paper, we study the problem focusing on extraction and representation of action information in dances. We propose to recommend dance videos based on the automatically discovered ``Dance Styles'', which play a significant role in characterizing different types of dances. To bridge the semantic gap of video content and mid-level concept, style, we take advantage of a mid-level action representation method, and extract representative patches as ``Dancelets'', a sort of intermediation between videos and the concepts. Furthermore, we propose to employ Motion Boundaries as saliency priors and sparsely extract patches containing more representative information to generate a set of dancelet candidates. Dancelets are then discovered by Normalized-cut method, which is superior in grouping visually similar patterns into the same clusters. For the fast and effective recommendation, a random forest-based index is built, and the ranking results are derived according to the matching results in all the leaf notes. Extensive experiments validated on the web dance videos demonstrate the effectiveness of the proposed methods for dance style discovery and video recommendation based on styles. Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Yanhao Zhang 0001, Sicheng Zhao, Xiusheng Lu, Yinghao Huang, Wenlong Xie |
ACM Multimedia | 1 |
| 2014 | "Clustering by saliency" - Unsupervised discovery of crowd activitiesabstractIn this paper, we develop a novel unsupervised crowd activity discovery algorithm aiming to automatically explore latent action patterns among crowd activities and partition them into meaningful clusters. Inspired by computational model of human vision system, we present a spatiotemporal saliency-based representation to simulate visual attention mechanism and encode human-focused components in an activity stream. Combining with feature pooling, we could obtain a more compact and robust activity representation. Based on the affinity matrix of activities, N-cut is performed to generate clusters with meaningful activity patterns. We carry out experiments on our proposed HIT-BJUT dataset and another public UMN dataset. The experimental results demonstrate that the proposed unsupervised discovery method is capable of automatically mining meaningful activities from large-scale video data with mixed crowd activities. Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Yanhao Zhang 0001 |
ICIP | 1 |
| 2013 | A spatial-temporal constraint-based action recognition methodabstractIn this paper, we propose a spatial-temporal constraint-based action recognition method, in which two actions are compared by both the appearance features and the spatial-temporal structures. To represent the appearance information in videos, we utilize a random quantization method to obtain a more precise BoW-based representation. To calculate the similarity of two videos, we match the quantized interest point sets and map the matched pairs into the spatial-temporal offset space to compare the similarities of spatial-temporal structures. We leverage the KNN model to classify the actions. The experiment results on both KTH action dataset and YouTube action dataset demonstrate the effectiveness of our proposed method on action recognition. Tingting Han 0003, Hongxun Yao, Yanhao Zhang 0001, Pengfei Xu 0001 |
ICIP | 1 |