Delong Liu

dblp:77/444 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval
abstract
As a challenging vision-language task, Zero-Shot Composed Image Retrieval (ZS-CIR) is designed to retrieve target images using bi-modal (image+text) queries. Typical ZS-CIR methods employ an inversion network to generate pseudo-word tokens that effectively represent the input semantics. However, the inversion-based methods suffer from two inherent issues: First, the task discrepancy exists because inversion training and CIR inference involve different objectives. Second, the modality discrepancy arises from the input feature distribution mismatch between training and inference. To this end, we propose a lightweight post-hoc framework, consisting of two components: (1) A new text-anchored triplet construction pipeline leverages a large language model (LLM) to transform a standard image-text dataset into a triplet dataset, where a textual description serves as the target of each triplet. (2) The MoTa-Adapter, a novel parameter-efficient fine-tuning method, adapts the dual encoder to the CIR task using our constructed triplet data. Specifically, on the text side, multiple sets of learnable task prompts are integrated via a Mixture-of-Experts (MoE) layer to capture task-specific priors and handle different types of modifications. On the image side, MoTa-Adapter modulates the inversion network's input to better match the downstream text encoder. In addition, an entropy-based optimization strategy is proposed to assign greater weight to challenging samples, thus improving adaptation efficiency. Experiments show that, with the incorporation of our proposed components, inversion-based methods achieve significant improvements, reaching state-of-the-art performance across four widely-used benchmarks.
Haiwen Li, Delong Liu, Zhaohui Hou, Zeliang Ma, Zhicheng Zhao 0001
AAAI2
2026 RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic Data
abstract
Precise and controllable image editing, especially object removal and insertion, represents one of the most common demands in image manipulation. However, existing methods suffer from severe limitations. Mask-based inpainting often introduces visual artifacts and semantic inconsistencies, while instruction-based approaches lack accurate spatial control and tend to unintentionally modify background regions. To address these issues, we propose two key contributions. First, we develop a fully automated and self-improving pipeline for synthetic data generation. This pipeline utilizes a Large Language Model (LLM) to generate diverse prompts, a Diffusion Transformer (DiT) fine-tuned evolutionarily to synthesize high-quality images, and a Multimodal LLM (MLLM) combined with open-set object detector for automated quality control and annotation. This process produces the Remove/Add Dataset (RAD), consisting of over 514,510 high-quality image pairs, each richly annotated with bounding boxes, segmentation masks, and a variety of editing instructions. Second, based on RAD, we introduce Remove/Add Anything (RAA), a novel editing framework with precise spatial control. Built upon a diffusion-based inpainting model, RAA achieves high editing accuracy by conditioning on both textual instructions and an explicitly defined region of interest (ROI), enabling efficient fine-tuning while maintaining global visual coherence. Extensive experiments demonstrate that RAA significantly outperforms existing open-source methods on both addition and removal tasks, and even slightly surpasses costly proprietary models.
Delong Liu, Haotian Hou, Zhaohui Hou, Shihao Han, Mingjie Zhan, Zhicheng Zhao 0001
AAAI1
2025 UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer
abstract
Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by aesthetic principles, we propose a non-invasive plug-in called Uniform Frame Organizer (UFO), which is compatible with any diffusion-based video generation model. The UFO comprises a series of adaptive adapters with adjustable intensities, which can significantly enhance the consistency between the foreground and background of videos and improve image quality without altering the original model parameters when integrated. The training for UFO is simple, efficient, requires minimal resources, and supports stylized training. Its modular design allows for the combination of multiple UFOs, enabling the customization of personalized video generation models. Furthermore, the UFO also supports direct transferability across different models of the same specification without the need for specific retraining. The experimental results indicate that UFO effectively enhances video generation quality and demonstrates its superiority in public video generation benchmarks.
Delong Liu, Zhaohui Hou, Mingjie Zhan, Shihao Han, Zhicheng Zhao 0001
AAAI1
2025 Object-Centric Discriminative Learning for Text-Based Person Retrieval
abstract
Text-based person retrieval (TBPR) is a vision-language task that aims to find specific pedestrians in a large image gallery using the textual description. However, due to the heterogeneity between modalities and the redundancy in visual representations, it remains a challenging task. Existing methods do not explicitly reduce the influence of the background regions in images, inevitably decreasing representation ability and reducing the image-text matching performance. In this paper, we propose a novel framework for text-based person retrieval, termed Object-Centric Discriminative Learning (OCDL), which incorporates person masks to indicate attentive regions, thereby enhancing the model’s focus on the pedestrians in images while suppressing the background noise. Additionally, a novel crossmodal matching loss, namely Soft Angular Distribution Matching (SADM), is introduced to learn discriminative visual and textual representations. Extensive experiments on three widely-used TBPR datasets demonstrate the effectiveness of our approach. The code is available at https://github.com/JThuge/OCDL.
Haiwen Li, Delong Liu, Zhicheng Zhao 0001
ICASSP2
2025 CE-LoRA: Consistent Person Synthesis by Exploring the Model's Spatial Consistency
abstract
In image generation, a large number of studies focus on advanced network architectures so as to adapt to various consistency requirements. However, the training of these models usually relies on labor-intensive real-world data collection. Additionally, they overlook the model’s inherent ability to generate consistent outputs, such as naturally producing identical objects within a single image. In contrast, we propose a scalable self-distillation-based consistency data generation pipeline, which progressively enhances spatial consistency and ultimately enables the automatic generation of high-quality consistency images. Taking person consistency as a case study, we first generate a high-quality open-source dataset named Per-400K. Secondly, based on this dataset, the Consistency-Enhanced Low-Rank Adaptation (CE-LoRA) module is presented to learn the spatial consistency by incorporating the guidance image into the same spatial dimension as the target person image, achieving high-fidelity generation of person images. Experimental results demonstrate that CE-LoRA achieves state-of-the-art performance across multiple consistency generation metrics.
Delong Liu, Zhicheng Zhao 0001
ICME1
2025 Automatic Synthetic Data and Fine-grained Adaptive Feature Alignment for Composed Person Retrieval
abstract
Person retrieval has attracted rising attention. Existing methods are mainly divided into two retrieval modes, namely image-only and text-only. However, they are unable to make full use of the available information and are difficult to meet diverse application requirements. To address the above limitations, we propose a new Composed Person Retrieval (CPR) task, which combines visual and textual queries to identify individuals of interest from large-scale person image databases. Nevertheless, the foremost difficulty of the CPR task is the lack of available annotated datasets. Therefore, we first introduce a scalable automatic data synthesis pipeline, which decomposes complex multimodal data generation into the creation of textual quadruples followed by identity-consistent image synthesis using fine-tuned generative models. Meanwhile, a multimodal filtering method is designed to ensure the resulting SynCPR dataset retains 1.15 million high-quality and fully synthetic triplets. Additionally, to improve the representation of composed person queries, we propose a novel Fine-grained Adaptive Feature Alignment (FAFA) framework through fine-grained dynamic alignment and masked feature reasoning. Moreover, for objective evaluation, we manually annotate the Image-Text Composed Person Retrieval (ITCPR) test set. The extensive experiments demonstrate the effectiveness of the SynCPR dataset and the superiority of the proposed FAFA framework when compared with the state-of-the-art methods. All code and data will be provided at https://github.com/Delong-liu-bupt/Composed_Person_Retrieval.
Delong Liu, Haiwen Li, Zhaohui Hou, Zhicheng Zhao 0001
NeurIPS1
2025 Automated text annotation: a new paradigm for generalizable text-to-image person retrieval
Delong Liu, Zhicheng Zhao 0001
Appl. Intell.1
2025 GloNeRF: Boosting NeRF capabilities and multi-view consistency in low-light environments
Zongqiang Liu, Qingyao Meng, Daying Lu, Gongzheng Li, Decai Guo, Delong Liu, Chuanxin Liu, Zongxu Yang
Comput. Graph.6
2025 OpenDriver: An open-road driver state detection benchmark
Delong Liu, Zhu Meng, Zhicheng Zhao 0001
J. Netw. Comput. Appl.1
2025 MindShot: A few-shot brain decoding framework via transferring cross-subject prior and distilling frequency domain knowledge
Zhu Meng, Haiwen Li, Delong Liu, Zhicheng Zhao 0001
Knowl. Based Syst.4
2025 Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval
Delong Liu, Haiwen Li, Zhicheng Zhao 0001
Neural Networks1
2024 Boundary-refined prototype generation: A general end-to-end paradigm for semi-supervised semantic segmentation
Junhao Dong 0002, Zhu Meng, Delong Liu, Zhicheng Zhao 0001
Eng. Appl. Artif. Intell.3
2023 An Adaptive Clustering Algorithm Based on Local-Density Peaks for Imbalanced Data Without Parameters
abstract
Imbalanced data clustering is a challenging problem in machine learning. The main difficulty is caused by the imbalance in both cluster size and data density distribution. To address this problem, we propose a novel clustering algorithm called LDPI based on local-density peaks in this study. First, an initial sub-cluster construction scheme is designed based on a 3-dimensional (3-D) decision graph that can easily detect the initial sub-cluster centers and identify the noise points. Second, a sub-cluster updating strategy is designed, which can automatically identify the false sub-cluster centers and update the initial sub-clusters. Third, a sub-cluster merging scheme is designed, which merges the updated initial sub-clusters into final clusters. Consequently, the proposed algorithm has three advantages: 1) It does not require any input parameters; 2) It can automatically determine the cluster centers and number of clusters; 3) It is suitable for imbalanced datasets and datasets with arbitrary shapes and distributions. The effectiveness of LDPI is demonstrated experimentally and the superiority of LDPI is identified by comparison with 5 state-of-the-art algorithms.
Wuning Tong, Yuping Wang 0003, Delong Liu
IEEE Trans. Knowl. Data Eng.3
2021 ASTS: attention based spatio-temporal sequential framework for movie trailer genre classification
Yitong Yu, Yang Li 0145, Delong Liu
Multim. Tools Appl.4
2018 Protein Complexes Detection Based on Global Network Representation Learning
Bo Xu 0009, Delong Liu, Yi-Jia Zhang 0001, Hongfei Lin, Jian Wang 0021, Feng Xia 0001
BIBM4
2016 Novel design and kinematics modeling for delta robot with improved end effector
abstract
A improved delta robot with a novel transmission mechanism to reduce abrasion and a new configuration of end effector is proposed in this paper. Specifically, the traditional telescopic chain which employs prismatic pair is replaced by parallel mechanism with four rods hinged. The frication in the connection can be significantly decreased. The motor that drives the telescopic chin is fixed on the base so that the moving platform can reach high speed and high acceleration. As for end effector, it is designed as mechanical claws based on the principle of plane thread. This structure makes it better adapted to multifarious shapes of load. Moreover, the rigidity and the stability of the end effector are distinctly improved. In addition, the analysis of kinematics including inverse kinematics and direct kinematics is conducted on the simplified system. The analytical results are validated by the co-simulation that combines ADAMS with MATLAB. The study of the kinematics is the solid foundation for parameters optimization and further control.
Liang Yan 0001, Delong Liu, Zongxia Jiao
IECON2
2006 Phase analysis of circadian-related genes in two tissues
abstract
BACKGROUND: Recent circadian clock studies using gene expression microarray in two different tissues of mouse have revealed not all circadian-related genes are synchronized in phase or peak expression times across tissues in vivo. Instead, some circadian-related genes may be delayed by 4-8 hrs in peak expression in one tissue relative to the other. These interesting biological observations prompt a statistical question regarding how to distinguish the synchronized genes from genes that are systematically lagged in phase/peak expression time across two tissues. RESULTS: We propose a set of techniques from circular statistics to analyze phase angles of circadian-related genes in two tissues. We first estimate the phases of a cycling gene separately in each tissue, which are then used to estimate the paired angular difference of the phase angles of the gene in the two tissues. These differences are modeled as a mixture of two von Mises distributions which enables us to cluster genes into two groups; one group having synchronized transcripts with the same phase in the two tissues, the other containing transcripts with a discrepancy in phase between the two tissues. For each cluster of genes we assess the association of phases across the tissue types using circular-circular regression. We also develop a bootstrap methodology based on a circular-circular regression model to evaluate the improvement in fit provided by allowing two components versus a one-component von-Mises model. CONCLUSION: We applied our proposed methodologies to the circadian-related genes common to heart and liver tissues in Storch et al. 2, and found that an estimated 80% of circadian-related transcripts common to heart and liver tissues were synchronized in phase, and the other 20% of transcripts were lagged about 8 hours in liver relative to heart. The bootstrap p-value for being one cluster is 0.063, which suggests the possibility of two clusters. Our methodologies can be extended to analyze peak expression times of circadian-related genes across more than two tissues, for example, kidney, heart, liver, and the suprachiasmatic nuclei (SCN) of the hypothalamus.
Delong Liu, Shyamal D. Peddada, Leping Li, Clarice R. Weinberg
BMC Bioinform.1
2004 A geometric approach to determine association and coherence of the activation times of cell-cycling genes under differing experimental conditions
abstract
Differing arresting agents and protocols can be used to synchronize cells in cultures to specific phases of the cell when studying cell-cycle gene expressions. Often, data derived from individual experiments are analyzed separately, since no appropriate statistical methodology is available at the moment to analyze the data from all such experiments simultaneously. The focus of this paper is to determine the association and coherence of the relative activation times of cell-cycling genes under different experimental conditions. Using a circular-circular regression model, we define two parameters, a rotation parameter for the angular difference between cells' arresting times (phases) in two cell-cycle experiments, and an association parameter to describe the correspondence between the cycle times of maximal expression (phase angles) for a set of genes studied in two experiments. Further, we propose a procedure to assess coherence across multiple experiments, i.e. to what extent the circular ordering of the phase angles of genes is maintained across multiple experiments. Coherence of genes across experiments suggests that functionally these genes tend to respond in a stereotypically sequenced way under different experimental conditions. Our proposed methodology is illustrated by applying it to a HeLa cell-cycle gene-expression data.
Delong Liu, Clarice R. Weinberg, Shyamal D. Peddada
Bioinform.1