EDBT 2026 Demo / reviewers in the wild / expert
Ting Liu 0018
dblp:52/5150-18
· DBLP profile ↗
21ranked-venue papers
3as first author
21since 2021 · last 2026
0009-0008-7988-5935ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression ComprehensionabstractReferring expression comprehension (REC) is a vision-language task to locate a target object in an image based on a language expression. Fully fine-tuning general-purpose pre-trained vision-language foundation models for REC yields impressive performance but becomes increasingly costly. Parameter-efficient transfer learning (PETL) methods have shown strong performance with fewer tunable parameters. However, directly applying PETL to REC faces two challenges: (1) insufficient multi-modal interaction between pre-trained vision-language foundation models, and (2) high GPU memory usage due to gradients passing through the heavy vision-language foundation models. To this end, we present M2IST: Multi-Modal Interactive Side-Tuning with M3ISAs: Mixture of Multi-Modal Interactive Side-Adapters. During fine-tuning, we fix the pre-trained uni-modal encoders and update M3ISAs to enable efficient vision-language alignment for REC. Empirical results reveal that M2IST achieves better performance-efficiency trade-off than full fine-tuning and other PETL methods, requiring only 2.11% tunable parameters, 39.61% GPU memory, and 63.46% training time while maintaining competitive performance. Our code is released at https://github.com/xuyang-liu16/M2IST. Xuyang Liu 0002, Ting Liu 0018, Siteng Huang, Yi Xin 0003, Yue Hu 0016, Long Qin 0004, Yuanyuan Wu 0001, Honggang Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Draw Like an Artist: Complex Scene Generation With Diffusion Model via Composition, Painting, and RetouchingabstractRecent advances in text-to-image diffusion models have demonstrated impressive capabilities in image quality. However, complex scene generation remains relatively unexplored, and even the definition of ‘complex scene’ itself remains unclear. In this paper, we address this gap by providing a precise definition of complex scenes and introducing a set of Complex Decomposition Criteria (CDC) based on this definition. Inspired by the artist’s painting process, we propose a training-free diffusion framework called Complex Diffusion (CxD), which divides the process into three stages: composition, painting and retouching. Our method leverages the powerful chain-of-thought capabilities of large language models (LLMs) to decompose complex prompts based on CDC and to manage composition and layout. We then develop an attention modulation method that guides simple prompts to specific regions to complete the complex scene painting. Finally, we inject the detailed output of the LLM into a retouching model to enhance the image details, thus implementing the retouching stage. Extensive experiments demonstrate that our method outperforms previous SOTA approaches, significantly improving the generation of high-quality, semantically consistent, and visually diverse images for complex scenes, even with intricate prompts. Minghao Liu 0022, Yingjie Tian 0001, Xiaochao Qu, Luoqi Liu, Ting Liu 0018 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Text-Guided Vision Token Reduction With Low-Rank Adaptation for Efficient Visual GroundingabstractTransformer-based pretrained models have advanced Visual Grounding (VG) significantly, but their scaling up has caused soaring training and inference costs. While efforts adapting Parameter-Efficient Fine-Tuning to VG have cut training expenses to some extent, inference costs are over-looked, stemming from Transformer’s computation costs growing quadratically with input token length. To address the above challenge, we propose a text-guided token prune method based on a one-stream architecture for VG, named OneSVG. Specifically, OneSVG transfers low-rank adaptation LoRA to VG, and calculates the relevance between each vision token and text semantics in multiple stages, and gradually prunes the vision tokens with low relevance. In addition, traditional VG methods use a single [REG] token to predict bounding boxes, and the [REG] token relies on complete vision tokens during training, pruning large number of vision tokens will disrupt the spatial orientation information of images. Therefore, the prediction head in OneSVG first restores the square structure of images by padding the missing area, and then accepts complete vision tokens. Experimental results on three widely-used benchmarks demonstrate that our OneSVG achieves state-of-the-art real-time speed while maintaining the best accuracy with only 34.3% tokens. Codes and models are available at https://github.com/ltShi/OneSVG. Liangtao Shi, Ting Liu 0018, Jinxia Xie, Ning Li 0044, Bineng Zhong 0001, Richang Hong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Memory Efficient Matting with Adaptive Token RoutingabstractTransformer-based models have recently achieved outstanding performance in image matting. However, their application to high-resolution images remains challenging due to the quadratic complexity of global self-attention. To address this issue, we propose MEMatte, a memory-efficient matting framework for processing high-resolution images. MEMatte incorporates a router before each global attention block, directing informative tokens to the global attention while routing other tokens to a Lightweight Token Refinement Module (LTRM). Specifically, the router employs a local-global strategy to predict the routing probability of each token, and the LTRM utilizes efficient modules to simulate global attention. Additionally, we introduce a Batch-constrained Adaptive Token Routing (BATR) mechanism, which allows each router to dynamically route tokens based on image content and the stages of attention block in the network. Furthermore, we construct an ultra high-resolution image matting dataset, UHR-395, comprising 35,500 training images and 1,000 test images, with an average resolution of 4872 × 6017. This dataset is created by compositing 395 different alpha mattes across 11 categories onto various backgrounds, all with high-quality manual annotation. Extensive experiments demonstrate that MEMatte outperforms existing methods on both high-resolution and real-world datasets, significantly reducing memory usage by approximately 88% and latency by 50% on the Composition-1K benchmark. Yiheng Lin 0002, Yihan Hu 0004, Chenyi Zhang 0004, Ting Liu 0018, Xiaochao Qu, Luoqi Liu, Yao Zhao 0001, Yunchao Wei |
AAAI | 4 |
| 2025 | NTClick: Achieving Precise Interactive Segmentation With Noise-tolerant ClicksabstractInteractive segmentation is a pivotal task in computer vision, focused on predicting precise masks with minimal user input. Although the click has recently become the most prevalent form of interaction due to its flexibility and efficiency, its advantages diminish as the complexity and details of target objects increase because it’s time-consuming and user-unfriendly to precisely locate and click on narrow, fine regions. To tackle this problem, we propose NTClick, a powerful click-based interactive segmentation method capable of predicting accurate masks even with imprecise user clicks when dealing with intricate targets. We first introduce a novel interaction form called noise-tolerant click, a type of click that does not require user’s precise localization when selecting fine regions. Then, we design a two-stage workflow, consisting of an Explicit Coarse Perception network for initial estimation and a High Resolution Refinement network for final classification. Quantitative results across extensive datasets demonstrate that NTClick not only maintains an efficient and user-friendly interaction mode but also significantly outperforms existing methods in segmentation accuracy. Chenyi Zhang 0004, Ting Liu 0018, Xiaochao Qu, Luoqi Liu, Yao Zhao 0001, Yunchao Wei |
CVPR | 2 |
| 2025 | EVPGS: Enhanced View Prior Guidance for Splatting-based Extrapolated View SynthesisabstractGaussian Splatting (GS)-based methods rely on sufficient training view coverage and perform synthesis on interpolated views. In this work, we tackle the more challenging and underexplored Extrapolated View Synthesis (EVS) task. Here we enable GS-based models trained with limited view coverage to generalize well to extrapolated views. To achieve our goal, we propose a view augmentation framework to guide training through a coarse-to-fine process. At the coarse stage, we reduce rendering artifacts due to insufficient view coverage by introducing a regularization strategy at both appearance and geometry levels. At the fine stage, we generate reliable view priors to provide further training guidance. To this end, we incorporate an occlusion awareness into the view prior generation process, and refine the view priors with the aid of coarse stage output. We call our framework Enhanced View Prior Guidance for Splatting (EVPGS). To comprehensively evaluate EVPGS on the EVS task, we collect a real-world dataset called Merchandise3D dedicated to the EVS scenario. Experiments on three datasets including both real and synthetic demonstrate EVPGS achieves state-of-the-art performance, while improving synthesis quality at extrapolated views for GS-based methods both qualitatively and quantitatively. Our code and dataset are available on the EVPGS Homepage. Jiahe Li 0006, Xiaochao Qu, Chengjing Wu, Luoqi Liu, Ting Liu 0018 |
CVPR | 6 |
| 2025 | MTADiffusion: Mask Text Alignment Diffusion Model for Object InpaintingabstractAdvancements in generative models have enabled image inpainting models to generate content within specific regions of an image based on provided prompts and masks. However, existing inpainting methods often suffer from problems such as semantic misalignment, structural distortion, and style inconsistency. In this work, we present MTADiffusion, a Mask-Text Alignment diffusion model designed for object inpainting. To enhance the semantic capabilities of the inpainting model, we introduce MTAPipeline, an automatic solution for annotating masks with detailed descriptions. Based on the MTAPipeline, we construct a new MTADataset comprising 5 million images and 25 million mask-text pairs. Furthermore, we propose a multi-task training strategy that integrates both inpainting and edge prediction tasks to improve structural stability. To promote style consistency, we present a novel inpainting style-consistency loss using a pre-trained VGG network and the Gram matrix. Comprehensive evaluations on BrushBench and EditBench demonstrate that MTADiffusion achieves state-of-the-art performance compared to other methods. Ting Liu 0018, Yihang Wu, Xiaochao Qu, Luoqi Liu |
CVPR | 2 |
| 2025 | GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text EditingabstractScene text editing, a subfield of image editing, requires modifying texts in images while preserving style consistency and visual coherence with the surrounding environment. While diffusion-based methods have shown promise in text generation, they still struggle to produce high-quality results. These methods often generate distorted or unrecognizable characters, particularly when dealing with complex characters like Chinese. In such systems, characters are composed of intricate stroke patterns and spatial relationships that must be precisely maintained. We present Glyph-Mastero, a specialized glyph encoder designed to guide the latent diffusion model for generating texts with stroke-level precision. Our key insight is that existing methods, despite using pretrained OCR models for feature extraction, fail to capture the hierarchical nature of text structures - from individual strokes to stroke-level interactions to overall character-level structure. To address this, our glyph encoder explicitly models and captures the cross-level interactions between local-level individual characters and global-level text lines through our novel glyph attention module. Meanwhile, our model implements a feature pyramid network to fuse the multi-scale OCR backbone features at the global-level. Through these cross-level and multi-scale fusions, we obtain more detailed glyph-aware guidance, enabling precise control over the scene text generation process. Our method achieves an 18.02% improvement in sentence accuracy over the state-of-the-art multi-lingual scene text editing baseline, while simultaneously reducing the text-region Fréchet inception distance by 53.28%. Ting Liu 0018, Xiaochao Qu, Chengjing Wu, Luoqi Liu |
CVPR | 2 |
| 2025 | SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything ModelabstractInteractive segmentation is to segment the mask of the target object according to the user’s interactive prompts. There are two mainstream strategies: early fusion and late fusion. Current specialist models utilize the early fusion strategy that encodes the combination of images and prompts to target the prompted objects, yet repetitive complex computations on the images result in high latency. Late fusion models extract image embeddings once and merge them with the prompts in later interactions. This strategy avoids redundant image feature extraction and improves efficiency significantly. A recent milestone is the Segment Anything Model (SAM). However, this strategy limits the models’ ability to extract detailed information from the prompted target zone. To address this issue, we propose SAM-REF, a two-stage refinement framework that fully integrates images and prompts by using a lightweight refiner into the interaction of late fusion, which combines the accuracy of early fusion and maintains the efficiency of late fusion. Through extensive experiments, we show that our SAM-REF model outperforms the current state-of-the-art method in most metrics on segmentation quality without compromising efficiency. Chongkai Yu, Ting Liu 0018, Xiaochao Qu, Chengjing Wu, Luoqi Liu |
CVPR | 2 |
| 2025 | Accelerating Diffusion Transformers with Token-wise Feature CachingabstractDiffusion transformers have shown significant effectiveness in both image and video synthesis at the expense of huge computation costs. To address this problem, feature caching methods have been introduced to accelerate diffusion transformers by caching the features in previous timesteps and reusing them in the following timesteps. However, previous caching methods ignore that different tokens exhibit different sensitivities to feature caching, and feature caching on some tokens may lead to 10$\times$ more destruction to the overall generation quality compared with other tokens. In this paper, we introduce token-wise feature caching, allowing us to adaptively select the most suitable tokens for caching, and further enable us to apply different caching ratios to neural layers in different types and depths. Extensive experiments on PixArt-alpha, OpenSora, and DiT demonstrate our effectiveness in both image and video generation with no requirements for training. For instance, 2.36$\times$ and 1.93$\times$ acceleration are achieved on OpenSora and PixArt-$\alpha$ with almost no drop in generation quality. Codes have been released in the supplementary material and Github. Chang Zou, Xuyang Liu 0002, Ting Liu 0018, Siteng Huang, Linfeng Zhang 0001 |
ICLR | 3 |
| 2025 | Consistent Image Layout Editing With Diffusion ModelsabstractDespite the great success of large-scale text-to-image diffusion models in image generation and image editing, existing methods still struggle with editing the layout of real-world images. Although a few works have been developed to address this issue, they either fail to adjust the image layout effectively or encounter challenges in preserving the visual appearance of objects after layout adjustment. To bridge this gap, this paper proposes a novel image layout editing method that not only re-arranges a real-world image to a specified layout, but also ensures that the visual appearance of the objects remains consistent with their original state prior to editing. Concretely, a Multi-Concept Learning scheme is developed to learn the concepts of different objects from a single image, which can be seen as a novel inversion scheme tailored for image layout editing. Then, we leverage the semantic consistency within intermediate features of diffusion models to project the appearance information of objects to the target regions to improve the fidelity of objects after editing. Additionally, a novel initialization noise design is adopted to facilitate the convergence and success rate of re-arranging the layout. The phenomenon of concept entanglement is also analyzed, and resolved by a novel asynchronous editing strategy. Extensive experimental results demonstrate that the proposed method outperforms existing methods in both layout alignment and visual consistency for the task of image layout editing. Ting Liu 0018, Lei Zhang 0021 |
IEEE Trans. Image Process. | 3 |
| 2025 | SwimVG: Step-Wise Multimodal Fusion and Adaption for Visual GroundingabstractVisual grounding aims to ground an image region through natural language, which heavily relies on cross-modal alignment. Most existing methods transfer visual/linguistic knowledge separately by fully fine-tuning uni-modal pre-trained models, followed by a simple stack of visual-language transformers for multimodal fusion. However, these approaches not only limit adequate interaction between visual and linguistic contexts, but also incur significant computational costs. Therefore, to address these issues, we explore a step-wise multimodal fusion and adaption framework, namely SwimVG. Specifically, SwimVG proposes step-wise multimodal prompts (Swip) and cross-modal interactive adapters (CIA) for visual grounding, replacing the cumbersome transformer stacks for multimodal fusion. Swip can improve the alignment between the vision and language representations step by step, in a token-level fusion manner. In addition, weight-level CIA further promotes multimodal fusion by cross-modal interaction. Swip and CIA are both parameter-efficient paradigms, and they fuse the cross-modal features from shallow to deep layers gradually. Experimental results on four widely-used benchmarks demonstrate that SwimVG achieves remarkable abilities and considerable benefits in terms of efficiency. Liangtao Shi, Ting Liu 0018, Xiantao Hu, Yue Hu 0016, Quanjun Yin, Richang Hong |
IEEE Trans. Multim. | 2 |
| 2024 | DAP: Domain-Aware Prompt Learning for Vision-and-Language NavigationabstractFollowing language instructions to navigate in unseen environments is a challenging task for autonomous embodied agents. With strong representation capabilities, pretrained vision-and-language models are widely used in VLN. However, most of them are trained on web-crawled generalpurpose datasets, which incurs a considerable domain gap when used for VLN tasks. To address the problem, we propose a novel and model-agnostic Domain-Aware Prompt learning (DAP) framework. For equipping the pretrained models with specific object-level and scene-level cross-modal alignment in VLN tasks, DAP applies a low-cost prompt tuning paradigm to learn soft visual prompts for extracting in-domain image semantics. Specifically, we first generate a set of in-domain image-text pairs with the help of the CLIP model. Then we introduce soft visual prompts in the input space of the visual encoder in a pretrained model. DAP injects in-domain visual knowledge into the visual encoder of the pretrained model in an efficient way. Experimental results on both R2R and REVERIE show the superiority of DAP compared to existing state-of-the-art methods. Ting Liu 0018, Yue Hu 0016, Wansen Wu, Youkai Wang, Kai Xu 0014, Quanjun Yin |
ICASSP | 1 |
| 2024 | DARA: Domain- and Relation-Aware Adapters Make Parameter-Efficient Tuning for Visual GroundingabstractVisual grounding (VG) is a challenging task to localize an object in an image based on a textual description. Recent surge in the scale of VG models has substantially improved performance, but also introduced a significant burden on computational costs during fine-tuning. In this paper, we explore applying parameter-efficient transfer learning (PETL) to efficiently transfer the pre-trained vision-language knowledge to VG. Specifically, we propose DARA, a novel PETL method comprising Domain-aware Adapters (DA Adapters) and Relation-aware Adapters (RA Adapters) for VG. DA Adapters first transfer intra-modality representations to be more fine-grained for the VG domain. Then RA Adapters share weights to bridge the relation between two modalities, improving spatial reasoning. Empirical results on widely-used benchmarks demonstrate that DARA achieves the best accuracy while saving numerous updated parameters compared to the full fine-tuning and other PETL methods. Notably, with only 2.13% tunable backbone parameters, DARA improves average accuracy by 0.81% across the three benchmarks compared to the baseline model. Our code is available at https://github.com/liuting20/DARA. Ting Liu 0018, Xuyang Liu 0002, Siteng Huang, Honggang Chen, Quanjun Yin, Long Qin 0004, Yue Hu 0016 |
ICME | 1 |
| 2024 | PANDA: Prompt-Based Context- and Indoor-Aware Pretraining for Vision and Language Navigation
Ting Liu 0018, Yue Hu 0016, Wansen Wu, Youkai Wang, Kai Xu 0014, Quanjun Yin |
MMM (1) | 1 |
| 2024 | ACT: Action-assoCiated and Target-Related Representations for Object Navigation
Youkai Wang, Yue Hu 0016, Wansen Wu, Ting Liu 0018, Yong Peng 0006 |
MMM (1) | 4 |
| 2024 | V-PETL Bench: A Unified Visual Parameter-Efficient Transfer Learning BenchmarkabstractParameter-efficient transfer learning (PETL) methods show promise in adapting a pre-trained model to various downstream tasks while training only a few parameters. In the computer vision (CV) domain, numerous PETL algorithms have been proposed, but their direct employment or comparison remains inconvenient. To address this challenge, we construct a Unified Visual PETL Benchmark (V-PETL Bench) for the CV domain by selecting 30 diverse, challenging, and comprehensive datasets from image recognition, video action recognition, and dense prediction tasks. On these datasets, we systematically evaluate 25 dominant PETL algorithms and open-source a modular and extensible codebase for fair evaluation of these algorithms. V-PETL Bench runs on NVIDIA A800 GPUs and requires approximately 310 GPU days. We release all the benchmark, making it more efficient and friendly to researchers. Additionally, V-PETL Bench will be continuously updated for new PETL algorithms and CV tasks. Yi Xin 0003, Xuyang Liu 0002, Yuntao Du 0001, Haodi Zhou, Christina E. Lee, Junlong Du, Haozhe Wang 0002, Mingcai Chen, Ting Liu 0018, Guimin Hu, Zhongwei Wan, Rongchao Zhang, Aoxue Li, Mingyang Yi, Xiaohong Liu 0001 |
NeurIPS | 11 |
| 2023 | Dynamic Multi-modal Prompting for Efficient Visual Grounding
Wansen Wu, Ting Liu 0018, Youkai Wang, Kai Xu 0014, Quanjun Yin, Yue Hu 0016 |
PRCV (7) | 2 |
| 2022 | You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object SegmentationabstractWe present YOFO (You Only inFer Once), a new paradigm for referring video object segmentation (RVOS) that operates in an one-stage manner. Our key insight is that the language descriptor should serve as target-specific guidance to identify the target object, while a direct feature fusion of image and language can increase feature complexity and thus may be sub-optimal for RVOS. To this end, we propose a meta-transfer module, which is trained in a learning-to-learn fashion and aims to transfer the target-specific information from the language domain to the image domain, while discarding the uncorrelated complex variations of language description. To bridge the gap between the image and language domains, we develop a multi-scale cross-modal feature mining block that aggregates all the essential features required by RVOS from both domains and generates regression labels for the meta-transfer module. The whole system can be trained in an end-to-end manner and shows competitive performance against state-of-the-art two-stage approaches. Dezhuang Li, Ruoqi Li, Lijun Wang 0001, Yifan Wang 0004, Jinqing Qi, Lu Zhang 0053, Ting Liu 0018, Qingquan Xu, Huchuan Lu |
AAAI | 7 |
| 2022 | Multi-Source Uncertainty Mining for Deep Unsupervised Saliency DetectionabstractDeep learning-based image salient object detection (SOD) heavily relies on large-scale training data with pixel-wise labeling. High-quality labels involve intensive labor and are expensive to acquire. In this paper, we propose a novel multi-source uncertainty mining method to facilitate unsupervised deep learning from multiple noisy labels generated by traditional handcrafted SOD methods. We design an Uncertainty Mining Network (UMNet) which consists of multiple Merge-and-Split (MS) modules to recursively analyze the commonality and difference among multiple noisy labels and infer pixel-wise uncertainty map for each label. Meanwhile, we model the noisy labels using Gibbs distribution and propose a weighted uncertainty loss to jointly train the UMNet with the SOD network. As a consequence, our UMNet can adaptively select reliable labels for SOD network learning. Extensive experiments on benchmark datasets demonstrate that our method not only outperforms existing unsupervised methods, but also is on par with fully-supervised state-of-the-art models. Yifan Wang 0004, Lijun Wang 0001, Ting Liu 0018, Huchuan Lu |
CVPR | 4 |
| 2022 | Adaptive Co-teaching for Unsupervised Monocular Depth Estimation
Weisong Ren, Lijun Wang 0001, Yongri Piao, Miao Zhang 0004, Huchuan Lu, Ting Liu 0018 |
ECCV (1) | 6 |