Ting Liu 0012

dblp:52/5150-12 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0003-3458-6567ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Efficient Decoupled Feature 3D Gaussian Splatting via Hierarchical Compression
abstract
Efficient 3D scene representation has become a key challenge with the rise of 3D Gaussian Splatting (3DGS), particularly when incorporating semantic information into the scene representation. Existing 3DGS-based methods embed both color and high-dimensional semantic features into a single field, leading to significant storage and computational overhead. To mitigate this, we propose Decoupled Feature 3D Gaussian Splatting (DF-3DGS), a novel method that decouples the color and semantic fields, thereby reducing the number of 3D Gaussians required for semantic representation. We then introduce a hierarchical compression strategy that first employs our novel quantization approach with dynamic codebook evolution to reduce data size, followed by a scene-specific autoencoder for further compression of the semantic feature dimensions. This multi-stage approach results in a compact representation that enhances both storage efficiency and reconstruction speed. Experimental results demonstrate that DF-3DGS outperforms previous 3DGS-based methods, achieving faster training and rendering times while requiring less storage, without sacrificing performance—in fact, it improves performance in the novel view semantic segmentation task. Specifically, DF-3DGS achieves remarkable improvements over Feature 3DGS, reducing training time by 10× and storage by 20×, while improving the mIoU of novel view semantic segmentation by 4%. Code is available at https://github.com/dai647/DF_3DGS.
Zhenqi Dai, Ting Liu 0012
CVPR2
2025 Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation
abstract
Recent advances in zero-shot referring image segmentation (RIS), driven by models such as the Segment Anything Model (SAM) and CLIP, have made substantial progress in aligning visual and textual information. Despite these successes, the extraction of precise and high-quality mask region representations remains a critical challenge, limiting the full potential of RIS tasks. In this paper, we introduce a training-free, hybrid global-local feature extraction approach that integrates detailed mask-specific features with contextual information from the surrounding area, enhancing mask region representation. To further strengthen alignment between mask regions and referring expressions, we propose a spatial guidance augmentation strategy that improves spatial coherence, which is essential for accurately localizing described areas. By incorporating multiple spatial cues, this approach facilitates more robust and precise referring segmentation. Extensive experiments on standard RIS benchmarks demonstrate that our method significantly outperforms existing zero-shot RIS models, achieving substantial performance gains. We believe our approach advances RIS tasks and establishes a versatile framework for region-text alignment, offering broader implications for cross-modal understanding and interaction. Code is available at https://github.com/fhgyuanshen/HybridGL.
Ting Liu 0012
CVPR1
2025 Effective Road Segmentation With Selective State-Space Model and Frequency Feature Compensation
abstract
Road segmentation from high-resolution remote sensing imagery is critical for tasks such as autonomous driving, urban planning, and geographic information systems. However, challenges such as intensity nonuniformity, pixel ambiguity, and the visual similarity between roads and natural features make accurate segmentation difficult. In this article, we propose a road segmentation framework built upon the Mamba architecture, integrating a novel frequency feature compensation (FFC) approach to improve segmentation performance. Specifically, we introduce a progressive FFC method, leveraging wavelet decomposition to capture fine-grained details by separating features into high- and low-frequency components. Multistage features extracted from the Mamba backbone are decomposed using this approach and progressively integrated to compensate for the essential details for accurate road segmentation. We also introduce a wavelet loss (WL) to improve the model’s ability to capture fine structural variations in the frequency domain. Furthermore, we develop a spatial perception Mamba block (SPMB) to enhance the capture of spatial relationships. By seamlessly integrating global context and local structures with the selective state-space model and FFC, our framework significantly boosts road segmentation accuracy. Extensive experiments on three publicly available road segmentation datasets demonstrate that our method achieves state-of-the-art performance, surpassing existing approaches in segmenting complex roads.
Ting Liu 0012, Lei Zhang 0054, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Hierarchical Grafting Network With Structural Alignment for Ultra-High Resolution Image Segmentation
Ting Liu 0012, Shikui Wei, Yanning Zhang 0001
IEEE Trans. Multim.1
2024 Lyapunov-Stable Deep Equilibrium Models
abstract
Deep equilibrium (DEQ) models have emerged as a promising class of implicit layer models, which abandon traditional depth by solving for the fixed points of a single nonlinear layer. Despite their success, the stability of the fixed points for these models remains poorly understood. By considering DEQ models as nonlinear dynamic systems, we propose a robust DEQ model named LyaDEQ with guaranteed provable stability via Lyapunov theory. The crux of our method is ensuring the Lyapunov stability of the DEQ model's fixed points, which enables the proposed model to resist minor initial perturbations. To avoid poor adversarial defense due to Lyapunov-stable fixed points being located near each other, we orthogonalize the layers after the Lyapunov stability module to separate different fixed points. We evaluate LyaDEQ models under well-known adversarial attacks, and experimental results demonstrate significant improvement in robustness. Furthermore, we show that the LyaDEQ model can be combined with other defense methods, such as adversarial training, to achieve even better adversarial robustness.
Haoyu Chu, Shikui Wei, Ting Liu 0012, Yao Zhao 0001, Yuto Miyatake
AAAI3
2024 One-shot In-context Part Segmentation
abstract
In this paper, we present the One-shot In-context Part Segmentation (OIParts) framework, designed to tackle the challenges of part segmentation by leveraging visual foundation models (VFMs). Existing training-based one-shot part segmentation methods that utilize VFMs encounter difficulties when faced with scenarios where the one-shot image and test image exhibit significant variance in appearance and perspective, or when the object in the test image is partially visible. We argue that training on the one-shot example often leads to overfitting, thereby compromising the model's generalization capability. Our framework offers a novel approach to part segmentation that is training-free, flexible, and data-efficient, requiring only a single in-context example for precise segmentation with superior generalization ability. By thoroughly exploring the complementary strengths of VFMs, specifically DINOv2 and Stable Diffusion, we introduce an adaptive channel selection approach by minimizing the intra-class distance for better exploiting these two features, thereby enhancing the discriminatory power of the extracted features for the fine-grained parts. We have achieved remarkable segmentation performance across diverse object categories. The OIParts framework not only eliminates the need for extensive labeled data but also demonstrates superior generalization ability. Through comprehensive experimentation on three benchmark datasets, we have demonstrated the superiority of our proposed method over existing part segmentation approaches in one-shot settings.
Zhenqi Dai, Ting Liu 0012, Xingxing Zhang 0001, Yunchao Wei, Yanning Zhang 0001
ACM Multimedia2
2024 Toward Accurate Human Parsing Through Edge Guided Diffusion
abstract
Existing human parsing frameworks commonly employ joint learning of semantic edge detection and human parsing to facilitate the localization around boundary regions. Nevertheless, the parsing prediction within the interior of the part contour may still exhibit inconsistencies due to the inherent ambiguity of fine-grained semantics. In contrast, binary edge detection does not suffer from such fine-grained semantic ambiguity, leading to a typical failure case where misclassification occurs inner the part contour while the semantic edge is accurately detected. To address these challenges, we develop a novel diffusion scheme that incorporates guidance from the detected semantic edge to mitigate this problem by propagating corrected classified semantics into the misclassified regions. Building upon this diffusion scheme, we present an Edge Guided Diffusion Network (EGDNet) for human parsing, which can progressively refine the parsing predictions to enhance the accuracy and coherence of human parsing results. Moreover, we design a horizontal-vertical aggregation to exploit inherent correlations among body parts along both the horizontal and vertical axes, which aims at enhancing the initial parsing results. Extensive experimental evaluations on various challenging datasets demonstrate the effectiveness of the proposed EGDNet. Remarkably, our EGDNet shows impressive performances on six benchmark datasets, including four human body parsing datasets (LIP, CIHP, ATR, and PASCAL-Person-Part), and two human face parsing datasets (CelebAMask-HQ and LaPa).
Ting Liu 0012, Hongkun Zhu, Yunchao Wei, Shikui Wei, Yao Zhao 0001, Yanning Zhang 0001
IEEE Trans. Image Process.1
2024 Exploring the Applicability of Spectral Recovery in Semantic Segmentation of RGB Images
abstract
Compared with RGB images, hyperspectral images (HSIs) offer a distinct advantage in that they can record continuous spectral bands of light reflectance in each pixel, reflecting the physical and chemical characteristics of materials. This capability enables differentiation between objects that may have similar textures but different spectral characteristics. It is desirable to recover spectral information from RGB images to improve semantic segmentation accuracy. Additionally, semantic information can serve as a guide for spectral information recovery, thereby ensuring the quality of the recovered spectral information. The two tasks are mutually beneficial in this regard. In light of these considerations, we propose a multi-task framework that exploits the complementary relationship between spectral recovery and semantic segmentation tasks, comprising a complementary spectral-semantic attentive fusion model (CSSF) that enables the two tasks to mutually facilitate each other by fusing information from both branches. Specifically, the proposed CSSF incorporates a window-based spectral-semantic attentive fusion (WSSAF) module to incorporate recovered spectral information into the segmentation process effectively, and a pixel-shuffle-based fusion (PSF) module to provide semantic guidance for spectral recovery. To evaluate the effectiveness of our approach, we built the first flower hyperspectral image dataset (FHRS) with corresponding segmentation annotations and RGB images. By doing so, we have made the first attempt to explore the complementary relationship between semantic segmentation and spectral recovery. Experimental results on both the FHRS dataset and the publicly available LIB-HSI dataset demonstrate that our proposed method has the ability to enhance both tasks by utilizing their complementary relationship, indicating the generalization ability of our method.
Zhuoran Du, Shikui Wei, Ting Liu 0012, Shunli Zhang 0005, Shiyin Zhang, Yao Zhao 0001
IEEE Trans. Multim.3
2023 Progressive Neighborhood Aggregation for Semantic Segmentation Refinement
abstract
Multi-scale features from backbone networks have been widely applied to recover object details in segmentation tasks. Generally, the multi-level features are fused in a certain manner for further pixel-level dense prediction. Whereas, the spatial structure information is not fully explored, that is similar nearby pixels can be used to complement each other. In this paper, we investigate a progressive neighborhood aggregation (PNA) framework to refine the semantic segmentation prediction, resulting in an end-to-end solution that can perform the coarse prediction and refinement in a unified network. Specifically, we first present a neighborhood aggregation module, the neighborhood similarity matrices for each pixel are estimated on multi-scale features, which are further used to progressively aggregate the high-level feature for recovering the spatial structure. In addition, to further integrate the high-resolution details into the aggregated feature, we apply a self-aggregation module on the low-level features to emphasize important semantic information for complementing losing spatial details. Extensive experiments on five segmentation datasets, including Pascal VOC 2012, CityScapes, COCO-Stuff 10k, DeepGlobe, and Trans10k, demonstrate that the proposed framework can be cascaded into existing segmentation models providing consistent improvements. In particular, our method achieves new state-of-the-art performances on two challenging datasets, DeepGlobe and Trans10k. The code is available at https://github.com/liutinglt/PNA.
Ting Liu 0012, Yunchao Wei, Yanning Zhang 0001
AAAI1
2022 Multi-Source Aggregation Transformer for Concealed Object Detection in Millimeter-Wave Images
abstract
The active millimeter wave scanner has been widely used for detecting objects concealed underneath a person’s clothing in the field of security inspection and anti-terrorism. However, the active millimeter wave (AMMW) images always suffer from low signal-noise ratio, motion blur, and small size objects, making it challenging to detect concealed objects efficiently and accurately. The scanner usually captures a sequence of images in different views around a human body at once, while the existing algorithms only utilize the single image without considering the relationships among images. In this paper, we design a multi-source aggregation transformer (MATR) with two different attention mechanisms to model spatial correlations within an image and contextual interactions across images. Specifically, a self-attention module is introduced to encode local relationships between the region proposals in each image, while a cross-attention mechanism is built to focus on modeling the cross-correlations between different images. Besides, to handle the problem of small objects in size and suppress the noise in AMMW images, we present a selective context module (SCM). It designs a dynamic selection mechanism to enhance the high-resolution feature with spatial details and make it more distinguishable from the noisy background. Experiments on two AMMW image datasets demonstrate that the proposed methods lead to a remarkable improvement compared to previous state-of-the-art and will benefit the concealed object detection in practice.
Ting Liu 0012, Shiyin Zhang, Yao Zhao 0001, Shikui Wei
IEEE Trans. Circuits Syst. Video Technol.2
2021 Spatial-Aware Texture Transformer for High-Fidelity Garment Transfer
abstract
Garment transfer aims to transfer the desired garment from a model image with the desired clothing to a target person, which has attracted a great deal of attention due to its wider potential applications. However, considering the model and target persons are often given at different views, body shapes and poses, realistic garment transfer is facing the following challenges that have not been well addressed: 1) deforming the garment; 2) inferring unobserved appearance; 3) preserving fine texture details. To tackle these challenges, we propose a novel SPatial-Aware Texture Transformer (SPATT) model. Different from existing models, SPATT establishes correspondence and infers unobserved clothing appearance by leveraging the spatial prior information of a UV-space. Specifically, the source image is transformed into a partial UV texture map guided by the extracted dense pose. To better infer the unseen appearance utilizing seen region, we first propose a novel coordinate-prior map that defines the spatial relationship between the coordinates in the UV texture map, and design an algorithm to compute it. Based on the proposed coordinate-prior map, we present a novel spatial-aware texture generation network to complete the partial UV texture. In the second stage, we first transform the completed UV texture to fit the target person. To polish the details and improve realism, we introduce a refinement generative network conditioned on the warped image and source input. Compared with existing frameworks as shown experimentally, the proposed framework can generate more realistic images with better-preserved texture details. Furthermore, difficult cases where two persons have large pose and view differences can also be well handled by SPATT.
Ting Liu 0012, Xuecheng Nie, Yunchao Wei, Shikui Wei, Yao Zhao 0001, Jiashi Feng
IEEE Trans. Image Process.1
2021 Blind Image Clustering for Camera Source Identification via Row-Sparsity Optimization
abstract
Given a set of images with the number of cameras providing those images unknown, how to blindly identify the sources of the images has been a critical problem in digital forensics. Although state-of-the-art methods have achieved impressive results, they have failed at suppressing outliers. When they deal with a noisy dataset, the performance is significantly degraded. To address this issue, we propose an optimization approach with sparsity constraints to simultaneously handle the how-many subproblem (i.e., the number of cameras) and the which-from-which subproblem (i.e., the image–camera relationship). In our approach, we first formulate the blind camera source clustering as a row-sparsity optimization problem, in which the representation errors are minimized and the outliers caused by noisy features are suppressed. Then, a new two-stage refinement method based on inter- and the intra-class differences is proposed to achieve a more accurate estimation of the number of cameras. Because strong sparsity constraints have been adopted and the interactive relationship among data points can be fully explored to distinguish the images originated from different cameras, the proposed method can effectively handle outliers. Extensive experiments on the popular Dresden dataset show that the proposed method outperforms existing methods in both identification accuracy and efficiency.
Xiang Jiang 0005, Shikui Wei, Ting Liu 0012, Ruizhen Zhao, Yao Zhao 0001, Heng Huang 0001
IEEE Trans. Multim.3
2019 Devil in the Details: Towards Accurate Single and Multiple Human Parsing
abstract
Human parsing has received considerable interest due to its wide application potentials. Nevertheless, it is still unclear how to develop an accurate human parsing system in an efficient and elegant way. In this paper, we identify several useful properties, including feature resolution, global context information and edge details, and perform rigorous analyses to reveal how to leverage them to benefit the human parsing task. The advantages of these useful properties finally result in a simple yet effective Context Embedding with Edge Perceiving (CE2P) framework for single human parsing. Our CE2P is end-to-end trainable and can be easily adopted for conducting multiple human parsing. Benefiting the superiority of CE2P, we won the 1st places on all three human parsing tracks in the 2nd Look into Person (LIP) Challenge. Without any bells and whistles, we achieved 56.50% (mIoU), 45.31% (mean APr) and 33.34% (APp0.5) in Track 1, Track 2 and Track 5, which outperform the state-of-the-arts more than 2.06%, 3.81% and 1.87%, respectively. We hope our CE2P will serve as a solid baseline and help ease future research in single/multiple human parsing. Code has been made available at https://github.com/liutinglt/CE2P.
Ting Liu 0012, Yunchao Wei, Shikui Wei, Yao Zhao 0001
AAAI2
2019 Adversarial task-specific learning
Xin Fu 0009, Yao Zhao 0001, Ting Liu 0012, Yunchao Wei, Jianan Li 0001, Shikui Wei
Neurocomputing3
2019 Magic-Wall: Visualizing Room Decoration by Enhanced Wall Segmentation
abstract
This paper presents an intelligent system named Magic-wall, which enables visualization of the effect of room decoration automatically. Concretely, given an image of the indoor scene and a preferred color, the Magic-wall can automatically locate the wall regions in the image and smoothly replace the existing wall with the required one. The key idea of the proposed Magic-wall is to leverage visual semantics to guide the entire process of color substitution, including wall segmentation and replacement. To strengthen the reality of visualization, we make the following contributions. First, we propose an edge-aware fully convolutional neural network (Edge-aware-FCN) for indoor semantic scene parsing, in which a novel edge-prior branch is introduced to identify the boundary of different semantic regions better. To further polish the details between the wall and other semantic regions, we leverage the output of Edge-aware-FCN as the prior knowledge, concatenating with the image to form a new input for the Enhanced-Net. In such a case, the Enhanced-Net is able to capture more semantic-aware information from the input and polish some ambiguous regions. Finally, to naturally replace the color of the original walls, a simple yet effective color space conversion method is proposed for replacement with brightness reserved. We build a new indoor scene dataset upon ADE20K for training and testing, which includes six semantic labels. Extensive experimental evaluations and visualizations well demonstrate that the proposed Magic-wall is effective and can automatically generate a set of visually pleasing results.
Ting Liu 0012, Yunchao Wei, Yao Zhao 0001, Si Liu 0001, Shikui Wei
IEEE Trans. Image Process.1
2017 Enhanced isomorphic semantic representation for cross-media retrieval
abstract
Nowadays cross-media retrieval is an useful technology that helps people find expected information from the huge amount of multimodal data more efficiently. A common cross-media retrieval framework is first to map features of different modalities into an isomorphic semantic space so that the similarity between heterogeneous data can be measured. For most of semantic space based methods, the mapping mechanism from original to semantic space of each modality is optimized independently, yet the more discriminative characteristic of a certain modality is not taken into account. In this paper, we propose a deep framework which introduces a latent embedding layer to learn joint parameters to obtain semantically meaningful representations of images and texts. Specifically, the discriminative characteristic embedded in the textual modality can be transferred to images through the latent embedding layer and joint parameters to enhance the consistency between semantic representations. Extensive experiments on the three popular publicly available datasets well demonstrate the superiority of the proposed method, which achieves the new state-of-the-arts.
Ting Liu 0012, Yao Zhao 0001, Shikui Wei, Yunchao Wei, Lixin Liao
ICME1
2017 Magic-wall: Visualizing Room Decoration
abstract
This work focuses on Magic-wall, an automatic system for visualizing the effect of room decoration. Given an image of the indoor scene and a preferred color, the Magic-wall can automatically locate the wall regions in the image and smoothly replace the existing color with the required one. The key idea of the proposed Magic-wall is to leverage visual semantics to guide the entire process of color substitution including wall segmentation and color replacement. We propose an edge-aware fully convolutional neural network (FCN) for indoor semantic scene parsing, in which a novel edge-prior branch is introduced to better identify the boundary of different semantic regions. To accurately localize the wall regions, we adapt a semantic-dependent optimized strategy, which pays more attention to those pixels belonging to the wall by adapting larger optimization weights compared with those from other semantic regions. Finally, to naturally replace the color of original walls, a simple yet effective color space conversion method is proposed for replacement with brightness reservation. We build a new indoor scene dataset upon ADE20K for training and testing, which includes 6 semantic labels. Extensive experimental evaluations and visualizations well demonstrate that the proposed Magic-wall is effective and can automatically generate a set of visually pleasing results.
Ting Liu 0012, Yunchao Wei, Yao Zhao 0001, Si Liu 0001, Shikui Wei
ACM Multimedia1
2017 Two-stream Attentive CNNs for Image Retrieval
abstract
In content-based image retrieval, the most challenging (and ambiguous) part is to define the similarity between images. For the human-being, such similarity can be defined with respect to where they pay attention to and what semantic attributes they understand. Inspired by this fact, this paper presents two-stream attentive CNNs for image retrieval. As the human-being does, the proposed network has two streams that simultaneously handle two tasks. The Main stream focuses on extracting discriminative visual features that are tightly correlated with semantic attributes. Meanwhile, the Auxiliary stream aims to facilitate the main stream by redirecting the feature extraction operation mainly to the image content that human may pay attention to. By fusing these two streams into the Main and Auxiliary CNNs (MAC), image similarity can be computed as the human-being does by reserving the conspicuous content and suppressing the irrelevant regions. Extensive experiments show that the proposed model achieves impressive performance in image retrieval on four public datasets.
Jia Li 0003, Shikui Wei, Qinjie Zheng, Ting Liu 0012, Yao Zhao 0001
ACM Multimedia5