Tingting Gao

dblp:55/9453 · DBLP profile ↗
← Back
37ranked-venue papers
8as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 6 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Beyond Tokens: Dynamic Latent Reasoning via Semantic Residual Refinement
abstract
Chain-of-Thought prompting has remarkably advanced LLM reasoning by generating explicit step-by-step tokens, yet its discrete nature inherently limits expressiveness and efficiency, struggling with abstract, ambiguous, or semantically divergent cognition beyond linguistic tokens. Latent reasoning offers a promising alternative by operating in the model’s internal continuous space for richer cognitive representations. However, existing methods typically rely on finetuning or token interpolation to bridge latent and input spaces, introducing training difficulty or semantic degradation. To this end, we propose Dynamic Latent Reasoning (DyLaR), a training-free framework that preserves semantic fidelity to latent space. DyLaR introduces a Semantic Residual Refinement module that progressively refines latent inputs by integrating semantic residuals from prior hidden states, thus capturing expressive semantic hierarchies that closely approximate continuous latent representations. To enhance flexibility, DyLaR further incorporates a dynamic switching policy that allows LLMs to alternate between discrete and latent reasoning based on model uncertainty, favoring explicit reasoning when confident and latent exploration under ambiguity. Empirical experiments across knowledge- and reasoning-intensive tasks demonstrate that DyLaR consistently outperforms strong baselines in both effectiveness and token efficiency. Qualitative analyses further illustrate its interpretability and flexibility in navigating complex reasoning scenarios.
Fangrui Lv, Ruixin Hong, Tingting Gao, Guorui Zhou, Changshui Zhang
AAAI6
2026 TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs
abstract
Video large language models have achieved remarkable performance in tasks such as video question answering, however, their temporal understanding remains suboptimal. To address this limitation, we curate a dedicated instruction fine-tuning dataset that focuses on enhancing temporal comprehension across five key dimensions. In order to reduce reliance on costly temporal annotations, we introduce a multi-task prompt fine-tuning approach that seamlessly integrates temporal-sensitive tasks into existing instruction datasets without requiring additional annotations. Furthermore, we develop a novel benchmark for temporal-sensitive video understanding that not only fills the gaps in dimension coverage left by existing benchmarks but also rigorously filters out potential shortcuts, ensuring a more accurate evaluation. Extensive experimental results demonstrate that our approach significantly enhances the temporal understanding of video-LLMs while avoiding reliance on shortcuts.
Meng Liu 0006, Xuemeng Song, Fan Yang 0094, Tingting Gao, Di Zhang 0026, Guorui Zhou, Liqiang Nie
AAAI7
2026 Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
abstract
Da Li, Yuxiao Luo, Keping Bi, Jiafeng Guo, Wei Yuan, Biao Yang, Yan Wang, Fan Yang, Tingting Gao, Guorui Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Da Li 0003, Keping Bi, Jiafeng Guo, Fan Yang 0094, Tingting Gao, Guorui Zhou
ACL (1)9
2026 IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation
abstract
Yinghao Tang, Xueding Liu, Boyuan Zhang, Tingfeng Lan, Yupeng Xie, Jiale Lao, Yiyao Wang, Haoxuan Li, Tingting Gao, Bo Pan, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu, Yingchaojie Feng, Yuyu Luo, Wei Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yinghao Tang, Xueding Liu, Tingfeng Lan, Jiale Lao, Yiyao Wang, Tingting Gao, Bo Pan 0004, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu 0001, Yingchaojie Feng, Yuyu Luo, Wei Chen 0001
ACL (1)9
2026 The co-occurrence between depressive symptoms and smartphone addiction: a network analysis
abstract
The present study aimed to explore the co-occurrence between depression and smartphone addiction (SA) from the network perspective as well as the network invariance across different groups. A total of 8347 Chinese college students were included in the study. Network analysis was conducted to estimate the network structure of the co-occurrence between depression and SA symptoms. A network comparison test was utilized to explore sex, severity of depression and SA variations in the network structures and strengths. D18 ‘Sad’ was the central symptom for the estimated network. D5 ‘Mind’ may be the most important bridge symptom between depression and SA. Males and females differed in the distribution of edge weights (M = 0.103, p = 0.024). Students with or without depressive symptoms showed significant differences in the distribution of edge weights (M = 0.129, p = 0.001) and global strength (14.5 vs. 13.5, p = 0.006). In addition, there are variations in the distribution of edge weights for college students in different SA severity groups (M = 0.119, p < 0.001). Attention to these core and bridge symptoms may decrease the odds of co-occurrence of depressive symptoms and SA. Interventions to address the co-occurrence should also take into account sex and severity of depression and SA.
Tingting Gao, Yinna Xu, Chengchao Zhou, Xiangfei Meng
Behav. Inf. Technol.1
2026 A3Bench: an audience-aligned multilingual benchmark for video audience insights understanding
Yiming Lei 0001, Guozhen Peng, Zeming Liu, Hui Qiu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu 0001, Annan Li, Yunhong Wang 0001
Frontiers Comput. Sci.7
2025 GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
abstract
Video Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural and contextual subtleties. Although Multimodal Large Language Models (MLLMs) and Chain-of-Thought (CoT) have demonstrated strong reasoning abilities in STEM tasks (e.g. mathematics and coding), they still struggle to generate creative expressions such as resonant jokes and insightful satire. Moreover, existing benchmarks are constrained by their limited modalities and insufficient categories, hindering the exploration of comprehensive creativity in video-based Comment Art creation. To address these limitations, we introduce GODBench, a novel benchmark that integrates video and text modalities to systematically evaluate MLLMs’ abilities to compose Comment Art. Furthermore, inspired by the propagation patterns of waves in physics, we propose Ripple of Thought (RoT), a multi-step reasoning framework designed to enhance the creativity of MLLMs. Extensive experiments on GODBench reveal that existing MLLMs and CoT methods still face significant challenges in understanding and generating creative video comments. In contrast, RoT provides an effective approach to improving creative composing, highlighting its potential to drive meaningful advancements in MLLM-based creativity.
Yiming Lei 0001, Zeming Liu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu 0001, Yunhong Wang 0001
ACL (1)6
2025 CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
abstract
Interleaved image-text generation has emerged as a vital multimodal task aimed at creating sequences of interleaved visual and textual content given a query. Despite notable advancements in recent multimodal large language models (MLLMs), generating integrated image-text sequences that exhibit narrative coherence and entity and style consistency remains challenging due to poor training data quality. To this end, we introduce CoMM, a high-quality Coherent interleaved image-text MultiModal dataset designed to enhance the coherence, consistency, and alignment of generated multimodal content. Initially, CoMM harnesses raw data from diverse sources, focusing on instructional content and visual storytelling, establishing a foundation for coherent and consistent content. To further refine the data quality, we devise a multi-perspective filter strategy that leverages advanced pre-trained models to ensure the development of sentences, consistency of inserted images, and semantic alignment between them. Various quality evaluation metrics are designed to prove the high quality of the filtered dataset. Meanwhile, extensive few-shot experiments on various downstream tasks demonstrate CoMM’s effectiveness in significantly enhancing the in-context learning capabilities of MLLMs. Moreover, we propose four new tasks to evaluate MLLMs’ interleaved generation abilities, supported by a comprehensive evaluation framework. We believe CoMM opens a new avenue for advanced MLLMs with superior multimodal in-context learning and understanding ability.
Wei Chen 0070, Lin Li 0065, Yongqi Yang, Fan Yang 0094, Tingting Gao, Yu Wu 0011, Long Chen 0016
CVPR6
2025 Libra-Merging: Importance-redundancy and Pruning-merging Trade-off for Acceleration Plug-in in Large Vision-Language Model
abstract
Large Vision-Language Models (LVLMs) have achieved significant progress in recent years. However, the expensive inference cost limits the realistic deployment of LVLMs. Some works find that visual tokens are redundant and compress tokens to reduce the inference cost. These works identify important non-redundant tokens as target tokens, then prune the remaining tokens (non-target tokens) or merge them into target tokens. However, target token identification faces the token importance-redundancy dilemma. Besides, token merging and pruning face a dilemma between disrupting target token information and losing non-target token information. To solve these problems, we propose a novel visual token compression scheme, named Libra-Merging. In target token identification, Libra-Merging selects the most important tokens from spatially discrete intervals, achieving a more robust token importance-redundancy trade-off than relying on a hyper-parameter. In token compression, when non-target tokens are dissimilar to target tokens, Libra-Merging does not merge them into the target tokens, thus avoiding disrupting target token information. Meanwhile, Libra-Merging condenses these non-target tokens into an information compensation token to prevent losing important non-target token information. Our method can serve as a plug-in for diverse LVLMs, and extensive experimental results demonstrate its effectiveness. The code will be publicly available at https://github.com/longrongyang/Libra-Merging.
Longrong Yang, Dong Shen 0003, Chaoxiang Cai, Kaibing Chen, Fan Yang 0094, Tingting Gao, Di Zhang 0026
CVPR6
2025 SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
abstract
With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities of these models. However, these benchmarks focus on standalone videos and only assess "visual elements" like human actions and object states. In reality, contemporary videos often encompass complex and continuous narratives, typically presented as a series. To address this challenge, we propose SeriesBench, a benchmark consisting of 105 carefully curated narrative-driven series, covering 28 specialized tasks that require deep narrative understanding to solve. Specifically, we first select a diverse set of drama series spanning various genres. Then, we introduce a novel long-span narrative annotation method, combined with a full-information transformation approach to convert manual annotations into diverse task formats. To further enhance the model’s capacity for detailed analysis of plot structures and character relationships within series, we propose a novel narrative reasoning framework, PC-DCoT. Extensive results on SeriesBench indicate that existing MLLMs still face significant challenges in understanding narrative-driven series, while PC-DCoT enables these MLLMs to achieve performance improvements. Overall, our SeriesBench and PC-DCoT highlight the critical necessity of advancing model capabilities for understanding narrative-driven series, guiding future MLLMs development. SeriesBench is publicly available at https://github.com/zackhxn/SeriesBench-CVPR2025.
Yiming Lei 0001, Zeming Liu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu 0001, Yunhong Wang 0001
CVPR6
2025 MUSE: Multi-Subject Unified Synthesis Via Explicit Layout Semantic Expansion
abstract
Existing text-to-image diffusion models have demonstrated remarkable capabilities in generating high-quality images guided by textual prompts. However, achieving multi-subject compositional synthesis with precise spatial control remains a significant challenge. In this work, we address the task of layout-controllable multi-subject synthesis (LMS), which requires both faithful reconstruction of reference subjects and their accurate placement in specified regions within a unified image. While recent advancements have separately improved layout control and subject synthesis, existing approaches struggle to simultaneously satisfy the dual requirements of spatial precision and identity preservation in this composite task. To bridge this gap, we propose MUSE, a unified synthesis framework that employs concatenated cross-attention (CCA) to seamlessly integrate layout specifications with textual guidance through explicit semantic space expansion. The proposed CCA mechanism enables bidirectional modality alignment between spatial constraints and textual descriptions without interference. Furthermore, we design a progressive two-stage training strategy that decomposes the LMS task into learnable sub-objectives for effective optimization. Extensive experiments demonstrate that MUSE achieves zero-shot end-to-end generation with superior spatial accuracy and identity consistency compared to existing solutions, advancing the frontier of controllable image synthesis. Our code and model are available at https://github.com/pf0607/MUSE.
Fei Peng 0003, Junqiang Wu, Yan Li 0043, Tingting Gao, Di Zhang 0026, Huiyuan Fu
ICCV4
2025 TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
abstract
Multimodal visual language models are gaining prominence in open-world applications, driven by advancements in model architectures, training techniques, and high-quality data. However, their performance is often limited by insufficient task-specific data, leading to poor generalization and biased outputs. Existing efforts to increase task diversity in fine-tuning datasets are hindered by the labor-intensive process of manual task labeling, which typically produces only a few hundred task types. To address this, we propose TaskGalaxy, a large-scale multimodal instruction fine-tuning dataset comprising 19,227 hierarchical task types and 413,648 samples. TaskGalaxy utilizes GPT-4o to enrich task diversity by expanding from a small set of manually defined tasks, with CLIP and GPT-4o filtering those that best match open-source images, and generating relevant question-answer pairs. Multiple models are employed to ensure sample quality. This automated process enhances both task diversity and data quality, reducing manual intervention. Incorporating TaskGalaxy into LLaVA-v1.5 and InternVL-Chat-v1.0 models shows substantial performance improvements across 16 benchmarks, demonstrating the critical importance of task diversity. TaskGalaxy is publicly released at https://github.com/Kwai-YuanQi/TaskGalaxy.
Jiankang Chen, Tianke Zhang, Changyi Liu, Haojie Ding, Yaya Shi, Huihui Xiao, Fan Yang 0094, Tingting Gao, Di Zhang 0026
ICLR10
2025 Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model
abstract
The Mixture-of-Experts (MoE) has gained increasing attention in studying Large Vision-Language Models (LVLMs). It uses a sparse model to replace the dense model, achieving comparable performance while activating fewer parameters during inference, thus significantly reducing the inference cost. Existing MoE methods in LVLM encourage different experts to specialize in different tokens, and they usually employ a router to predict the routing of each token. However, the router is not optimized concerning distinct parameter optimization directions generated from tokens within an expert. This may lead to severe interference between tokens within an expert. To address this problem, we propose to use the token-level gradient analysis to Solving Token Gradient Conflict (STGC) in this paper. Specifically, we first use token-level gradients to identify conflicting tokens in experts. After that, we add a regularization loss tailored to encourage conflicting tokens routing from their current experts to other experts, for reducing interference between tokens within an expert. Our method can serve as a plug-in for diverse LVLM methods, and extensive experimental results demonstrate its effectiveness. demonstrate its effectiveness. The code will be publicly available at https://github.com/longrongyang/STGC.
Longrong Yang, Dong Shen 0003, Chaoxiang Cai, Fan Yang 0094, Tingting Gao, Di Zhang 0026
ICLR5
2025 MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
abstract
Existing efforts to align multimodal large language models (MLLMs) with human preferences have only achieved progress in narrow areas, such as hallucination reduction, but remain limited in practical applicability and generalizability. To this end, we introduce **MM-RLHF**, a dataset containing **120k** fine-grained, human-annotated preference comparison pairs. This dataset represents a substantial advancement over existing resources, offering superior size, diversity, annotation granularity, and quality. Leveraging this dataset, we propose several key innovations to improve both the quality of reward models and the efficiency of alignment algorithms. Notably, we introduce the **Critique-Based Reward Model**, which generates critiques of model outputs before assigning scores, offering enhanced interpretability and more informative feedback compared to traditional scalar reward mechanisms. Additionally, we propose **Dynamic Reward Scaling**, a method that adjusts the loss weight of each sample according to the reward signal, thereby optimizing the use of high-quality comparison pairs. Our approach is rigorously evaluated across **10** distinct dimensions, encompassing **27** benchmarks, with results demonstrating significant and consistent improvements in model performance (Figure.1).
Yifan Zhang 0004, Haochen Tian 0001, Chaoyou Fu, Peiyan Li 0001, Jianshu Zeng, Wulin Xie, Yang Shi 0009, Huanyu Zhang 0002, Junkang Wu, Xue Wang 0010, Yibo Hu 0001, Tingting Gao, Zhang Zhang 0001, Fan Yang 0094, Di Zhang 0026, Liang Wang 0001, Rong Jin 0001
ICML14
2025 VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform
Tianke Zhang, Chang Meng, Xiaobei Wang, Jinpeng Wang 0002, Yifan Zhang 0004, Shisong Tang, Changyi Liu, Haojie Ding, Kaiyu Jiang, Kaiyu Tang, Hai-Tao Zheng 0002, Fan Yang 0094, Tingting Gao, Di Zhang 0026, Kun Gai
KDD (2)15
2025 StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
abstract
The rapid growth of streaming video applications demands multimodal models with enhanced capabilities for temporal dynamics understanding and complex reasoning. However, current Video Question Answering (VideoQA) datasets suffer from two critical limitations: 1) Static annotation mechanisms fail to capture the evolving nature of answers in temporal video streams, and 2) The absence of explicit reasoning process annotations restricts model interpretability and logical deduction capabilities. To address these challenges, we introduce StreamingCoT, the first dataset explicitly designed for temporally evolving reasoning in streaming VideoQA and multimodal Chain-of-Thought (CoT) tasks. Our framework first establishes a dynamic hierarchical annotation architecture that generates per-second dense descriptions and constructs temporally-dependent semantic segments through similarity fusion, paired with question-answer sets constrained by temporal evolution patterns. We further propose an explicit reasoning chain generation paradigm that extracts spatiotemporal objects via keyframe semantic alignment, derives object state transition-based reasoning paths using large language models, and ensures logical coherence through human-verified validation. This dataset establishes a foundation for advancing research in streaming video understanding, complex temporal reasoning, and multimodal inference. Our StreamingCoT and its construction toolkit can be accessed at https://github.com/Fleeting-hyh/StreamingCoT.
Zhenyu Yang 0009, Shihan Wang 0006, Shengsheng Qian, Fan Yang 0094, Tingting Gao, Changsheng Xu
ACM Multimedia7
2025 Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models
abstract
Although multimodal large language models (MLLMs) exhibit remarkable reasoning capabilities on complex multimodal understanding tasks, they still suffer from the notorious 'hallucination' issue: generating outputs misaligned with obvious visual or factual evidence. Currently, training-based solutions, like direct preference optimization (DPO), leverage paired preference data to suppress hallucinations. However, they risk sacrificing general reasoning capabilities due to the likelihood displacement. Meanwhile, training-free solutions, like contrastive decoding, achieve this goal by subtracting the estimated hallucination pattern from a distorted input. Yet, these handcrafted perturbations (e.g., add noise to images) may poorly capture authentic hallucination patterns. To avoid these weaknesses of existing methods, and realize ``robust'' hallucination mitigation (\ie, maintaining general reasoning performance), we propose a novel framework: Decoupling Contrastive Decoding (DCD). Specifically, DCD decouples the learning of positive and negative samples in preference datasets, and trains separate positive and negative image projections within the MLLM. The negative projection implicitly models real hallucination patterns, which enables vision-aware negative images in the contrastive decoding inference stage. Our DCD alleviates likelihood displacement by avoiding pairwise optimization and generalizes robustly without handcrafted degradation. Extensive ablations across hallucination benchmarks and general reasoning tasks demonstrate the effectiveness of DCD, \ie, it matches DPO’s hallucination suppression while preserving general capabilities and outperforms the handcrafted contrastive decoding methods.
Wei Chen 0070, Fan Yang 0094, Tingting Gao, Di Zhang 0026, Long Chen 0016
NeurIPS5
2025 LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
abstract
Despite significant progress in Video Large Language Models (Video-LLMs) for offline video understanding, existing online Video-LLMs typically struggle to simultaneously process continuous frame-by-frame inputs and determine optimal response timing, often compromising real-time responsiveness and narrative coherence. To address these limitations, we introduce LiveStar, a pioneering live streaming assistant that achieves always-on proactive responses through adaptive streaming decoding. Specifically, LiveStar incorporates: (1) a training strategy enabling incremental video-language alignment for variable-length video streams, preserving temporal consistency across dynamically evolving frame sequences; (2) a response-silence decoding framework that determines optimal proactive response timing via a single forward pass verification; (3) memory-aware acceleration via peak-end memory compression for online inference on 10+ minute videos, combined with streaming key-value cache to achieve 1.53× faster inference. We also construct an OmniStar dataset, a comprehensive dataset for training and benchmarking that encompasses 15 diverse real-world scenarios and 5 evaluation tasks for online video understanding. Extensive experiments across three benchmarks demonstrate LiveStar's state-of-the-art performance, achieving an average 19.5\% improvement in semantic correctness with 18.1\% reduced timing difference compared to existing online Video-LLMs, while improving FPS by 12.0\% across all five OmniStar tasks. Our model and dataset can be accessed at https://github.com/yzy-bupt/LiveStar.
Zhenyu Yang 0009, Shengsheng Qian, Fan Yang 0094, Tingting Gao, Weiming Dong, Changsheng Xu
NeurIPS8
2025 Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
abstract
Preference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face challenges in handling noisy images of different timesteps and require complex transformations into pixel space. In this work, we show that pre-trained diffusion models are naturally suited for step-level reward modeling in the noisy latent space, as they are explicitly designed to process latent images at various noise levels. Accordingly, we propose the **Latent Reward Model (LRM)**, which repurposes components of the diffusion model to predict preferences of latent images at arbitrary timesteps. Building on LRM, we introduce **Latent Preference Optimization (LPO)**, a step-level preference optimization method conducted directly in the noisy latent space. Experimental results indicate that LPO significantly improves the model's alignment with general, aesthetic, and text-image alignment preferences, while achieving a 2.5-28x training speedup over existing preference optimization methods.
Cheng Da, Kun Ding 0001, Huan Yang 0005, Yan Li 0043, Tingting Gao, Di Zhang 0026, Shiming Xiang, Chunhong Pan
NeurIPS7
2025 Paragraph-to-Image Generation with Information-Enriched Diffusion Model
Weijia Wu 0001, Zhuang Li 0002, Yefei He, Zheng Shou 0001, Chunhua Shen, Lele Cheng, Tingting Gao
Int. J. Comput. Vis.8
2025 Event-Based Adaptive Consensus Control for Multiagent Systems With Asymmetric Multi-Information-Related Constraints on All States
abstract
In this article, the problem of adaptive event-triggered tracking control is investigated for a class of nonlinear multiagent systems (MASs) with asymmetric multi-information-related (MIR) constraints on all states. The fuzzy-logic systems (FLSs) are utilized to model the system unknown items by virtue of their universal approximation properties. The appropriate integral barrier Lyapunov functions (IBLFs) are selected to prevent the states from exceeding the asymmetric constraint boundaries associated with multi-information, which include historical states, time and neighbor outputs. The event-triggered mechanism (ETM) with varying threshold is employed to reduce the update frequency of the controller, thereby achieving the purpose of saving network resources, including communication bandwidth and computation abilities. Under the backstepping technique framework, the required control scheme is designed by integrating the adaptive controller with triggering mechanism. And it is proven that the controlled plant with state constraint conditions is stable, the consensus tracking errors can eventually remain near the origin, and the Zeno behavior does not exist. Finally, the simulation results corroborate the view that the designed control scheme is effective.
Tingting Gao, Tieshan Li 0001, Yan-Jun Liu 0003, Shaocheng Tong, Lei Liu 0006
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Decouple Content and Motion for Conditional Image-to-Video Generation
abstract
The goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text. The previous cI2V generation methods conventionally perform in RGB pixel space, with limitations in modeling motion consistency and visual continuity. Additionally, the efficiency of generating videos in pixel space is quite low. In this paper, we propose a novel approach to address these challenges by disentangling the target RGB pixels into two distinct components: spatial content and temporal motions. Specifically, we predict temporal motions which include motion vector and residual based on a 3D-UNet diffusion model. By explicitly modeling temporal motions and warping them to the starting image, we improve the temporal consistency of generated videos. This results in a reduction of spatial redundancy, emphasizing temporal details. Our proposed method achieves performance improvements by disentangling content and motion, all without introducing new structural complexities to the model. Extensive experiments on various datasets confirm our approach's superior performance over the majority of state-of-the-art methods in both effectiveness and efficiency.
Cuifeng Shen, Yulu Gan, Xiongwei Zhu, Lele Cheng, Tingting Gao, Jinzhi Wang
AAAI6
2024 Learning Multi-Dimensional Human Preference for Text-to-Image Generation
abstract
Current metrics for text-to-image models typically rely on statistical metrics which inadequately represent the real preference of humans. Although recent work attempts to learn these preferences via human annotated images, they reduce the rich tapestry of human preference to a single overall score. However, the preference results vary when humans evaluate images with different aspects. Therefore, to learn the multidimensional human preferences, we propose the Multi-dimensional Preference Score (MPS), the first multidimensional preference scoring model for the evaluation of text-to-image models. The MPS introduces the preference condition module upon CLIP model to learn these diverse preferences. It is trained based on our Multi-dimensional Human Preference (MHP) Dataset, which comprises 918,315 human preference choices across four dimensions (i.e., aesthetics, semantic alignment, detail quality and overall assessment) on 607,541 images. The images are generated by a wide range of latest text-to-image models. The MPS outperforms existing scoring methods across 3 datasets in 4 dimensions, enabling it a promising metric for evaluating and improving text-to-image generation. The model and dataset will be made publicly available to facilitate future research. Project page: htt ps: //wangbohan97.github.io/MPS/.
Sixian Zhang, Junqiang Wu, Yan Li 0043, Tingting Gao, Di Zhang 0026, Zhongyuan Wang 0006
CVPR5
2024 DragAnything: Motion Control for Anything Using Entity Representation
Weijia Wu 0001, Zhuang Li 0002, Yuchao Gu, Rui Zhao 0001, Yefei He, Junhao Zhang 0001, Zheng Shou 0001, Tingting Gao
ECCV (22)9
2023 A Unified Model for Video Understanding and Knowledge Embedding with Heterogeneous Knowledge Graph Dataset
abstract
Video understanding is an important task in short video business platforms and it has a wide application in video recommendation and classification. Most of the existing video understanding works only focus on the information that appeared within the video content, including the video frames, audio and text. However, introducing common sense knowledge from the external Knowledge Graph (KG) dataset is essential for video understanding when referring to the content which is less relevant to the video. Owing to the lack of video knowledge graph dataset, the work which integrates video understanding and KG is rare. In this paper, we propose a heterogeneous dataset that contains the multi-modal video entity and fruitful common sense relations. This dataset also provides multiple novel video inference tasks like the Video-Relation-Tag (VRT) and Video-Relation-Video (VRV) tasks. Furthermore, based on this dataset, we propose an end-to-end model that jointly optimizes the video understanding objective with knowledge graph embedding, which can not only better inject factual knowledge into video understanding but also generate effective multi-modal entity embedding for KG. Comprehensive experiments indicate that combining video understanding embedding with factual knowledge benefits the content-based video retrieval performance. Moreover, it also helps the model generate better knowledge graph embedding which outperforms traditional KGE-based methods on VRT and VRV tasks with at least 42.36% and 17.73% improvement in [email protected].
Jiaxin Deng, Dong Shen 0003, Haojie Pan, Ximan Liu, Gaofeng Meng, Fan Yang 0094, Tingting Gao, Ruiji Fu, Zhongyuan Wang 0006
ICMR8
2023 Observer-Based Adaptive Fuzzy Control of Nonstrict Feedback Nonlinear Systems With Function Constraints
abstract
In this article, an adaptive fuzzy tracking control scheme based on a fuzzy state observer is proposed for a class of uncertain, nonstrict feedback, nonlinear systems with function constraints. In the first place, based on the approximation characteristic of fuzzy logic systems (FLSs), a fuzzy state observer is designed to estimate the immeasurable state variables in the controlled system. Next, under the framework of adaptive backstepping control technology, FLSs are selected not only to approximate unknown nonlinear functions but also to avoid the algebraic loop problem caused by nonstrict feedback structure. At the same time, asymmetric Barrier Lyapunov functions are selected to solve the problem that system states are subject to function constraints, which are related to states and time. Then, Lyapunov stability theory is utilized to prove the stability of the controlled system, the realizability of function constraints, and the convergence of output tracking errors. Finally, a simulation is given to check the effectiveness of the proposed control scheme.
Tingting Gao, Tieshan Li 0001, Yan-Jun Liu 0003, Shaocheng Tong, Fuchun Sun 0001
IEEE Trans. Fuzzy Syst.1
2023 Adaptive Event-Triggered Fuzzy Control of State-Constrained Stochastic Nonlinear Systems Using IBLFs
abstract
In this article, an adaptive tracking control problem is addressed for nonstrict-feedback stochastic nonlinear systems subject to state constraints. Fuzzy logic systems (FLSs) are used to model unknown nonlinearities and avoid the algebraic loop arising from the system structure. Appropriate integral Barrier Lyapunov functions (IBLFs) are chosen so that time-varying full state constraints can be guaranteed directly rather than by transforming the constraint object. In the framework of backstepping technology, the relative threshold strategy is introduced to modify the adaptive control scheme so that the controller can be updated only after the trigger condition has been met, which reduces the update frequency of the controller and the loss of the actuator. Combined with Lyapunov stability theory, it is shown that all closed-loop signals are bounded in probability, in which the states remain within the specified constraints, and there is no Zeno behavior. A series of simulation results are given to reveal the effectiveness of the constructed control scheme.
Tingting Gao, Tieshan Li 0001, Yan-Jun Liu 0003, Shaocheng Tong, Lei Liu 0006
IEEE Trans. Fuzzy Syst.1
2022 Domain Generalization via Shuffled Style Assembly for Face Anti-Spoofing
abstract
With diverse presentation attacks emerging continually, generalizable face anti-spoofing (FAS) has drawn growing attention. Most existing methods implement domain generalization (DG) on the complete representations. However, different image statistics may have unique properties for the FAS tasks. In this work, we separate the complete representation into content and style ones. A novel Shuffled Style Assembly Network (SSAN) is proposed to extract and reassemble different content and style features for a stylized feature space. Then, to obtain a generalized representation, a contrastive learning strategy is developed to emphasize liveness-related style information while suppress the domain-specific one. Finally, the representations of the correct assemblies are used to distinguish between living and spoofing during the inferring. On the other hand, despite the decent performance, there still exists a gap between academia and industry, due to the difference in data quantity and distribution. Thus, a new large-scale benchmark for FAS is built up to further evaluate the performance of algorithms in reality. Both qualitative and quantitative results on existing and proposed benchmarks demonstrate the effectiveness of our methods. The codes will be available at https://github.com/wangzhuo2019/SSAN.
Zezheng Wang 0002, Zitong Yu, Weihong Deng, Tingting Gao, Zhongyuan Wang 0006
CVPR6
2022 IBLF-Based Adaptive Neural Control of State-Constrained Uncertain Stochastic Nonlinear Systems
abstract
In this article, the adaptive neural backstepping control approaches are designed for uncertain stochastic nonlinear systems with full-state constraints. According to the symmetry of constraint boundary, two cases of controlled systems subject to symmetric and asymmetric constraints are studied, respectively. Then, corresponding adaptive neural controllers are developed by virtue of backstepping design procedure and the learning ability of radial basis function neural network (RBFNN). It is worth mentioning that the integral Barrier Lyapunov function (IBLF), as an effective tool, is first applied to solve the above constraint problems. As a result, the state constraints are avoided from being transformed into error constraints via the proposed schemes. In addition, based on Lyapunov stability analysis, it is demonstrated that the errors can converge to a small neighborhood of zero, the full states do not exceed the given constraint bounds, and all signals in the closed-loop systems are semiglobally uniformly ultimately bounded (SGUUB) in probability. Finally, the numerical simulation results are provided to exhibit the effectiveness of the proposed control approaches.
Tingting Gao, Tieshan Li 0001, Yan-Jun Liu 0003, Shaocheng Tong
IEEE Trans. Neural Networks Learn. Syst.1
2021 Adaptive Neural Control Using Tangent Time-Varying BLFs for a Class of Uncertain Stochastic Nonlinear Systems With Full State Constraints
abstract
In this paper, an adaptive neural network (NN) control scheme is developed for a class of stochastic nonlinear systems with time-varying full state constraints. In the controller design, RBF NNs are employed to approximate the unknown terms, and the backtracking technique is introduced to overcome the restriction of matching conditions. At the same time, tangent type time-varying barrier Lyapunov functions (tan-TVBLFs) are constructed to ensure the full state constraints are never violated, where tan-TVBLFs are beneficial to integrate constraint analysis into a common method. Furthermore, the Lyapunov stability theory is used to prove that all closed-loop signals are semiglobal uniformly ultimately bounded in probability and error signals remain in the compact set do not violate the time-varying constraints. A simulation example will be used to exhibit the effectiveness of the proposed control scheme.
Tingting Gao, Yan-Jun Liu 0003, Dapeng Li 0004, Shaocheng Tong, Tieshan Li 0001
IEEE Trans. Cybern.1
2016 Development of two-dimensional materials for electronic applications
Tingting Gao, Yanqing Wu
Sci. China Inf. Sci.2
2014 Hysteresis modeling and compensation of PZT milliactuator in hard disk drives
abstract
Dual-stage actuation consisting of a PZT Microactuator (MA) and a Voice Coil Motor (VCM) has been used to improve the servo bandwidth and the disturbance rejections of Hard Disk Drivers (HDDs). However, the hysteresis in PZT MA limits the performance that can be achieved. In this paper, a Hammerstein model structure consisting of static hysteresis nonlinear block and dynamic linear block is used to model the PZT MA in HDDs and the identification method is also given. The nonlinear subsystem in the Hammerstein model is represented by Modified Prandtl-Ishlinskii (MPI) hysteresis model. It is proved that the proposed model is equivalent to the physical system. A hysteresis compensator is designed based on the proposed model. By the hysteresis compensation, the effects of hysteresis on the frequency responses of the PZT MA are measured. It is shown that apparent phase lead is achieved by hysteresis compensation, which is useful to improve the performance of control system.
Chunling Du, Tingting Gao, Lihua Xie 0001
ICARCV3
2014 Uncovering social network Sybils in the wild
abstract
Sybil accounts are fake identities created to unfairly increase the power or resources of a single malicious user. Researchers have long known about the existence of Sybil accounts in online communities such as file-sharing systems, but they have not been able to perform large-scale measurements to detect them or measure their activities. In this article, we describe our efforts to detect, characterize, and understand Sybil account activity in the Renren Online Social Network (OSN). We use ground truth provided by Renren Inc. to build measurement-based Sybil detectors and deploy them on Renren to detect more than 100,000 Sybil accounts. Using our full dataset of 650,000 Sybils, we examine several aspects of Sybil behavior. First, we study their link creation behavior and find that contrary to prior conjecture, Sybils in OSNs do not form tight-knit communities. Next, we examine the fine-grained behaviors of Sybils on Renren using clickstream data. Third, we investigate behind-the-scenes collusion between large groups of Sybils. Our results reveal that Sybils with no explicit social ties still act in concert to launch attacks. Finally, we investigate enhanced techniques to identify stealthy Sybils. In summary, our study advances the understanding of Sybil behavior on OSNs and shows that Sybils can effectively avoid existing community-based Sybil detectors. We hope that our results will foster new research on Sybil detection that is based on novel types of Sybil features.
Zhi Yang 0001, Christo Wilson, Xiao Wang 0018, Tingting Gao, Ben Y. Zhao, Yafei Dai
ACM Trans. Knowl. Discov. Data4
2012 Analysis of actuator in-phase property in terms of control performance and integrated plant/controller design using a novel model matching method
abstract
This paper is concerned with resonance in-phase property of a VCM (voice coil motor) plant system in the sense of control performance in HDDs (hard disk drives). Its relationships with the optimal performance level γopt, the stability margins and the disturbance rejection capability are revealed. It is found that the main resonance being in-phase is particularly beneficial to rejection of the narrow-band disturbances with frequencies near plant resonances. In order to meet the requirement on the inphase property, a partial model matching method is proposed. This model matching problem is solved by an H∞method using an linear matrix inequality approach. The partial model matching method is then applied to the VCM plant system. We especially take into account the in-phase case for the purpose to improve the system ability to attenuate high frequency disturbance. For the new system designed using the proposed model matching method, a feedback controller and a group peak filter are designed to attenuate the disturbance near the plant resonances. The advantages of the in-phase resonances are illustrated, when compared with the original plant.
Chunling Du, Tingting Gao, Lihua Xie 0001
ICARCV2
2012 Control performance comparison of PZT microactuator driven by voltage and current amplifiers in HDD dual-stage systems
abstract
In this paper, we investigate the effect of voltage and current amplifiers for PZT microactuators on the control performance of dual-stage servo systems in hard disk drives (HDDs), where the PZT microactuator is used as a secondary actuator and works together with the primary actuator of voice coil motor (VCM). First, the PZT microactuator's behavior in terms of motion linearization and frequency responses is experimentally studied and compared when it is driven by a conventional voltage amplifier and a charge or current amplifier. It is found that the PZT microactuator with current amplifier has less hysteresis than with voltage amplifier and its first resonance is relatively smaller. Inspired by this difference, the control performance of the dual-stage servo systems in track-seeking and track-following is then compared between the two driving methods for the PZT microactuator.
Tingting Gao, Chunling Du, Lihua Xie 0001
ICARCV1
2011 Uncovering social network sybils in the wild
abstract
Sybil accounts are fake identities created to unfairly increase the power or resources of a single user. Researchers have long known about the existence of Sybil accounts in online communities such as file-sharing systems, but have not been able to perform large scale measurements to detect them or measure their activities. In this paper, we describe our efforts to detect, characterize and understand Sybil account activity in the Renren online social network (OSN). We use ground truth provided by Renren Inc. to build measurement based Sybil account detectors, and deploy them on Renren to detect over 100,000 Sybil accounts. We study these Sybil accounts, as well as an additional 560,000 Sybil accounts caught by Renren, and analyze their link creation behavior. Most interestingly, we find that contrary to prior conjecture, Sybil accounts in OSNs do not form tight-knit communities. Instead, they integrate into the social graph just like normal users. Using link creation timestamps, we verify that the large majority of links between Sybil accounts are created accidentally, unbeknownst to the attacker. Overall, only a very small portion of Sybil accounts are connected to other Sybils with social links. Our study shows that existing Sybil defenses are unlikely to succeed in today's OSNs, and we must design new techniques to effectively detect and defend against Sybil attacks.
Zhi Yang 0001, Christo Wilson, Xiao Wang 0018, Tingting Gao, Ben Y. Zhao, Yafei Dai
Internet Measurement Conference4
2010 Impulsive disturbance rejection in hard disk drives
abstract
This paper proposes filtering methods to cancel the impulsive disturbance contained in position error signal (PES) in hard disk drives (HDDs). The impulsive disturbances may be observed as a few single sudden changes or some consecutive changes in PES. The filtering includes impulsive disturbance identification and estimation. A recursive method is used to determine the dynamic boundaries for identification. Two methods are proposed for estimation: a linear interpolation and an adaptive least mean square (LMS) algorithm. The former is used to estimate the normal PES directly and the later is used to estimate the impulsive disturbance for cancellation. It turns out that these methods are able to effectively cancel the impulsive disturbance and do not affect the servo performance.
Tingting Gao, Chunling Du, Lihua Xie 0001, Wen-Jian Cai
ICARCV1