Xinyan Liu 0008

dblp:93/8587-8 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-2638-4324ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection
abstract
As forgery types continue to emerge consistently, Incremental Face Forgery Detection (IFFD) has become a crucial paradigm. However, existing methods typically rely on data replay or coarse binary supervision, which fails to explicitly constrain the feature space, leading to severe feature drift and catastrophic forgetting. To address this, we propose AIFIND, Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection, which leverages semantic anchors to stabilize incremental learning. We design the Artifact-Driven Semantic Prior Generator to instantiate invariant semantic anchors, establishing a fixed coordinate system from low-level artifact cues. These anchors are injected into the image encoder via Artifact-Probe Attention, which explicitly constrains volatile visual features to align with stable semantic anchors. Adaptive Decision Harmonizer harmonizes the classifiers by preserving angular relationships of semantic anchors, maintaining geometric consistency across tasks. Extensive experiments on multiple incremental protocols validate the superiority of AIFIND.
Hao Wang 0035, Beichen Zhang 0006, Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu 0008, Weigang Zhang
ICMR7
2026 SeqCount: A sequence modeling framework for class-agnostic counting
abstract
Class-agnostic counting aims to count the number of objects in any category with only a few exemplars. It is crucial for solving the challenge of counting any visual class without re-finetuning, which in turn lowers deployment costs across diverse scenarios. Existing methods count the number of exemplar objects by integrating them over the density map smoothed by Gaussian kernels. However, designing generic kernels to generate density maps is challenging due to the different sizes and shapes of objects. To solve this problem, we propose SeqCount, which eliminates the need for density maps and treats object counting as a sequence generation problem. Specifically, we consider an input image as $N\times N$ patches and propose a serialization scheme. Then, we use an encoding-decoding structure to exploit the correlation among the patches. Experimental results on five challenging datasets demonstrate that our method performs favorably against the state-of-the-art models.
Guorong Li, Xinyan Liu 0008, Zhenjun Han, Yuankai Qi
J. Vis. Commun. Image Represent.3
2026 Distinguishing semantically similar queries in temporal video grounding via LLM-generated query
Yibo Dang, Zhaobo Qi, Xinyan Liu 0008, Xinzhe Han, Weigang Zhang
Multim. Syst.3
2026 Consistency-Aware Anchor Pyramid Network for Crowd Localization
abstract
Crowd localization aims to predict the positions of humans in images of crowded scenes. While existing methods have made significant progress, two primary challenges remain: (i) a fixed number of evenly distributed anchors can cause excessive or insufficient predictions across regions in an image with varying crowd densities, and (ii) ranking inconsistency of predictions between the testing and training phases leads to the model being sub-optimal in inference. To address these issues, we propose a Consistency-Aware Anchor Pyramid Network (CAAPN) comprising two key components: an Adaptive Anchor Generator (AAG) and a Localizer with Augmented Matching (LAM). The AAG module adaptively generates anchors based on estimated crowd density in local regions to alleviate the anchor deficiency or excess problem. It also considers the spatial distribution prior to heads for better performance. The LAM module is designed to augment the predictions which are used to optimize the neural network during training by introducing an extra set of target candidates and correctly matching them to the ground truth. The proposed method achieves favorable performance against state-of-the-art approaches on five challenging datasets: ShanghaiTech A and B, UCF-QNRF, JHU-CROWD++, and NWPU-Crowd.
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Zhenjun Han, Anton van den Hengel, Nicu Sebe, Ming-Hsuan Yang 0001, Qingming Huang
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Dynamic example network for class-agnostic object counting
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Ziheng Yan, Weigang Zhang, Laiyun Qing, Qingming Huang
Pattern Recognit.1
2026 Compactness driven Co-learning for crowd counting and localization
Ziheng Yan, Xinyan Liu 0008, Guorong Li, Weigang Zhang, Fang Wan 0001, Qingming Huang
Pattern Recognit.2
2025 Subpart Suppression Network for Few-Shot Object Counting
abstract
Few-shot object counting and detection aim to count objects along with their bounding boxes specified by exemplar bounding boxes. Current mainstream methods predict density maps by applying similarity between exemplar and image features to get counting results and detecting peak points from the density map as the positions of objects. However, the sub-parts of objects can also have high similarity to the exemplars, harming the counting and detection performance. To address these issues, we propose a two-stage Subpart Suppression Network (SSN), consisting of a Subpart Suppression Density map Predictor (SSDP), which enforces the model focus on whole objects rather than subparts, and a SAM-based Detector and Verifier, which refines the final prediction by a clustering method. Extensive experiments on two popular datasets show an advantage in performance over the state-of-the-art methods and prove the components’ effectiveness.
Lanxin Liu, Xinyan Liu 0008, Guorong Li
ICASSP2
2025 Combatting Data Imbalance and Noise in Micro-Action Recognition
Weidong Chen 0010, Zhaobo Qi, Pengqi Huang, Xinyan Liu 0008, Weigang Zhang
ACM Multimedia7
2025 Multi-round Mutual Emotion-Cause Pair Extraction for Emotion-Attributed Video Captioning
abstract
Emotional Video Captioning (EVC) is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. Existing EVC methods perceive global emotional cues through visual features at first, and then combine them with the video features to guide the emotional caption generation, which ignores the critical characteristic of the EVC task that emotional cues have intrinsic motivational causes reflected in the video content. Such video causes have a facilitative effect on both emotion perception and emotion-attributed caption generation. To this end, a multi-round mutual emotion-cause pair extraction network (MM-ECPE) is proposed in this paper for the joint extraction of emotional cues and visual causes through iterative mutual refinement. Specifically, in the 1st-round mutual learning, we propose a spatio-temporal disentangled visual adaptive refinement (ST-DVAR) and a multi-level video-guided emotion affine transformation (MV-EAT) to achieve preliminary refinement on video features and emotion lexicon to eliminate the noise caused by emotion-irrelevant visual information and video-irrelevant emotional information. Then, in the 2nd-round mutual learning, we exploit the cross-attention of the preliminary refined features and the original features to obtain the ultimate emotional cues and visual causes, and couple them in pair-wise extraction through contrastive loss. Overall, our approach optimizes complex semantic understanding and emotion perception of videos, leading to a promising performance in emotional captioning. Extensive experiments on three challenging datasets demonstrate the superiority of our approach and each proposed module, e.g., improving the latest records by +97.5% and +76.2% w.r.t. CIDEr and CFS, respectively, on the EVC-MSVD dataset.
Cheng Ye 0004, Weidong Chen 0013, Peipei Song, Xinyan Liu 0008, Lei Zhang 0119, Zhendong Mao 0001
ACM Multimedia4
2025 Boost Tracking by Natural Language With Prompt-Guided Grounding
abstract
TNL (Tracking by Natural Language) aims to locate the target described by a natural language sentence in a video. Most existing TNL methods are typically composed of three modules: object grounding, object tracking, and switching module, and their performance is limited by the poor performance of the grounding and switching modules due to the complex backgrounds and inaccurate information stored in the memory. This paper presents a global-local framework to address these issues, which includes a prompt-guided grounding module, a trained local tracking module, and a memory-based switcher module. The prompt-guided grounding module uses noun prompts to guide the CLIP model in focusing more on target regions and aligning visual features semantically with linguistic features, avoiding being misled by distractors and background. The memory-based switch module stores historical information with higher-quality memory, allowing the model to make more accurate decisions based on reliable data, thus improving the overall performance. Experiments on TNL2K, LaSOT, and OTB-Lang demonstrate the effectiveness and generalizability of the proposed framework.
Hengyou Li, Xinyan Liu 0008, Guorong Li, Shuhui Wang, Laiyun Qing, Qingming Huang
IEEE Trans. Intell. Transp. Syst.2
2024 Weakly Supervised Video Individual Counting
abstract
Video Individual Counting (VIC) aims to predict the number of unique individuals in a single video. Existing methods learn representations based on trajectory labels for individuals, which are annotation-expensive. To provide a more realistic reflection of the underlying practical challenge, we introduce a weakly supervised VIC task, wherein trajectory labels are not provided. Instead, two types of labels are provided to indicate traffic entering the field of view (inflow) and leaving the field view (outflow). We also propose the first solution as a baseline that formulates the task as a weakly supervised contrastive learning problem under group-level matching. In doing so, we devise an end-to-end trainable soft contrastive loss to drive the network to distin-guish inflow, outflow, and the remaining. To facilitate future study in this direction, we generate annotations from the existing VIC datasets Sense Crowd and CroHD and also build a new dataset, UAVVIC. Extensive results show that our baseline weakly supervised method outperforms supervised methods, and thus, little information is lost in the transition to the more practically relevant weakly supervised task. The code and trained model can be found at CGNet.
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Ziheng Yan, Zhenjun Han, Anton van den Hengel, Ming-Hsuan Yang 0001, Qingming Huang
CVPR1
2024 Directly Locating Actions in Video with Single Frame Annotation
abstract
We propose a novel method for point-supervised action localization.Differs from the common practice of locating actions by first categorizing each video frame, our method directly predicts actions' positions and length. Specifically, point-supervised action localization is achieved by a series of fully supervised action location iteratively. In each iteration, the input video are used as input tokens and fed into a transformer, where the encoder extracts global context of the clips, and the decoder generates queries containing information for action localization. Three MLP heads are built on each query to obtain the probability, the center, and the length of each action instance respectively. Experiments on three popular datasets prove the potential of our method.
Haoran Tong, Xinyan Liu 0008, Guorong Li, Laiyun Qing
ICMR2
2024 Progressive Multi-Resolution Loss for Crowd Counting
abstract
Crowd counting is usually handled in a density map regression fashion, which is supervised via an L2 loss between the predicted density map and ground truth. To effectively regulate models, various improved L2 loss functions have been developed to find a better correspondence between predicted density and annotation positions. In this paper, we propose to predict the density map at one resolution but measure its quality via a derived log-formed loss at multiple resolutions. Unlike existing methods that assume density maps at different resolutions are independent, our loss is obtained by modeling the likelihood function inspired by the relationship of density maps across multi-resolutions. We find that the traditional single-resolution L2 loss is a particular case of our derived log-likelihood. We mathematically prove it is superior to a single-resolution L2 loss. Without bells and whistles, the proposed loss substantially improves several baselines and performs favorably compared to state-of-the-art methods on five crowd counting datasets: NWPU-Crowd, ShanghaiTech A & B, UCF-QNRF, and JHU-Crowd++. The source code and trained models are released athttps://github.com/streamer-AP/PML_Loss.git.
Ziheng Yan, Yuankai Qi, Guorong Li, Xinyan Liu 0008, Weigang Zhang, Ming-Hsuan Yang 0001, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.4
2021 Exploiting sample correlation for crowd counting with multi-expert network
abstract
Crowd counting is a difficult task because of the diversity of scenes. Most of the existing crowd counting methods adopt complex structures with massive backbones to enhance the generalization ability. Unfortunately, the performance of existing methods on large-scale data sets is not satisfactory. In order to handle various scenarios with less complex network, we explored how to efficiently use the multi-expert model for crowd counting tasks. We mainly focus on how to train more efficient expert networks and how to choose the most suitable expert. Specifically, we propose a task-driven similarity metric based on sample’s mutual enhancement, referred as co-fine-tune similarity, which can find a more efficient subset of data for training the expert network. Similar samples are considered as a cluster which is used to obtain parameters of an expert. Besides, to make better use of the proposed method, we design a simple network called FPN with Deconvolution Counting Network, which is a more suitable base model for the multi-expert counting network. Experimental results show that multiple experts FDC (MFDC) achieves the best performance on four public data sets, including the large scale NWPU-Crowd data set. Furthermore, the MFDC trained on an extensive dense crowd data set can generalize well on the other data sets without extra training or fine-tuning.1
Xinyan Liu 0008, Guorong Li, Zhenjun Han, Weigang Zhang, Qingming Huang, Nicu Sebe
ICCV1