Guanglin Niu

dblp:240/6876 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0001-7260-7352ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object Tracking
abstract
Open-vocabulary multi-object tracking (OV-MOT) aims to track objects with unseen categories beyond the training set. While existing methods rely on pseudo video sequences synthesized from static images, they struggle to model realistic motion patterns, resulting in limited association performance in real-world scenarios. To alleviate these issues, we propose SAM2-OV, a novel association learning-free OV-MOT method that adopts a detection-only tuning paradigm, eliminating the need for synthetic sequences or spatiotemporal supervision and substantially reducing the overall learnable parameters. The core of our method is a Unified Detection Module (UDM), which effectively provides object-level prompts to enable SAM2 for OV-MOT. Enabled by UDM, SAM2-OV is the first to integrate SAM2 for OV-MOT, fully unleashing its zero-shot cross-frame association ability. To further enhance object association under occlusion and abrupt motion, we introduce a Motion Prior Assistance Module (MPAM) that incorporates motion cues into the mask selection process. In addition, a Semantic Enhancement Adapter (SEA) distilled from CLIP is used to improve classification generalization. A sparse prompting strategy is also adopted to reduce computational redundancy by triggering detection only on selected keyframes. As only the detection module is tuned on static images, the overall training process remains simple and efficient. Experiments on the TAO dataset demonstrate that SAM2-OV achieves state-of-the-art performance under the TETA metric, particularly on novel categories. Evaluations on the KITTI dataset show the strong zero-shot cross-domain transferability of our SAM2-OV.
Yangkai Chen, Qiangqiang Wu, Junlong Gao, Guanglin Niu, Hanzi Wang
AAAI5
2026 A Comprehensive Survey of Knowledge Graph Reasoning: Approaches and Applications
abstract
Knowledge graph reasoning (KGR) aims to infer novel knowledge based on existing facts in knowledge graphs (KGs), playing a crucial role in various cognition intelligence systems across diverse domains. The previous review works explore KGR models from specific perspectives such as KG types and embedding spaces. In contrast, this survey provides a more comprehensive perspective of KGR from foundational approaches and their applications. Notably, some seldom-attended approaches such as negative sampling strategies, popular open-source libraries, and rule-guided KGR paradigms are carefully reviewed. Besides, we explore advanced techniques, such as large language models (LLMs) and their impact on KGR. The comparison among foundational models are analyzed to declare their strengths and limitations. More interestingly, this is the first effort to provide a taxonomy of real-world KGR applications for both horizontal and vertical domains. Furthermore, we highlight the challenges and opportunities in the field of KGR, including trustworthiness, multimodal reasoning, continual learning, uncertainty and LLM-driven approaches. This work aims to bridge the gap between theoretical advancements and practical deployment of KGR models, and outline promising future directions.
Guanglin Niu, Bo Li 0006, Yangguang Lin
IEEE Trans. Big Data1
2025 TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
abstract
Recent advancements in Large Language Models (LLMs) have markedly enhanced the interpretation and processing of tabular data, introducing previously unimaginable capabilities. Despite these achievements, LLMs still encounter significant challenges when applied in industrial scenarios, particularly due to the increased complexity of reasoning required with real-world tabular data, underscoring a notable disparity between academic benchmarks and practical applications. To address this discrepancy, we conduct a detailed investigation into the application of tabular data in industrial scenarios and propose a comprehensive and complex benchmark TableBench, including 18 fields within four major categories of table question answering (TableQA) capabilities. Furthermore, we introduce TableLLM, trained on our meticulously constructed training set TableInstruct, achieving comparable performance with GPT-3.5. Massive experiments conducted on TableBench indicate that both open-source and proprietary LLMs still have significant room for improvement to meet real-world demands, where the most advanced model, GPT-4, achieves only a modest score compared to humans.
Xianjie Wu, Jian Yang 0030, Linzheng Chai, Ge Zhang 0009, Xeron Du, Di Liang, Daixin Shu, Xianfu Cheng, Tianzhen Sun, Tongliang Li, Zhoujun Li 0001, Guanglin Niu
AAAI13
2025 From Poses to Identity: Training-Free Person Re-Identification via Feature Centralization
abstract
Person re-identification (ReID) aims to extract accurate identity representation features. However, during feature extraction, individual samples are inevitably affected by noise (background, occlusions, and model limitations). Considering that features from the same identity follow a normal distribution around identity centers after training, we propose a Training-Free Feature Centralization ReID framework (Pose2ID) by aggregating the same identity features to reduce individual noise and enhance the stability of identity representation, which preserves the feature’s original distribution for following strategies such as re-ranking. Specifically, to obtain samples of the same identity, we introduce two components: ➀Identity-Guided Pedestrian Generation: by leveraging identity features to guide the generation process, we obtain high-quality images with diverse poses, ensuring identity consistency even in complex scenarios such as infrared, and occlusion. ➁Neighbor Feature Centralization: it explores each sample’s potential positive samples from its neighborhood. Experiments demonstrate that our generative model exhibits strong generalization capabilities and maintains high identity consistency. With the Feature Centralization framework, we achieve impressive performance even with an ImageNet pre-trained model without ReID training, reaching mAP/Rank-1 of 52.81/78.92 on Market1501. Moreover, our method sets new state-of-the-art results across standard, cross-modality, and occluded ReID tasks, showcasing strong adaptability.
Guiwei Zhang, Changxiao Ma, Guanglin Niu
CVPR5
2025 Diffusion-Based Hierarchical Negative Sampling for Multimodal Knowledge Graph Completion
Guanglin Niu
DASFAA (3)1
2025 Identity-aware Feature Decoupling Learning for Clothing-change Person Re-identification
abstract
Clothing-change person re-identification (CC Re-ID) has attracted increasing attention in recent years due to its application prospect. Most existing works struggle to adequately extract the ID-related information from the original RGB images. In this paper, we propose an Identity-aware Feature Decoupling (IFD) learning framework to mine identity-related features. Particularly, IFD exploits a dual stream architecture that consists of a main stream and an attention stream. The attention stream takes the clothing-masked images as inputs and derives the identity attention weights for effectively transferring the spatial knowledge to the main stream and highlighting the regions with abundant identity-related information. To eliminate the semantic gap between the inputs of two streams, we propose a clothing bias diminishing module specific to the main stream to regularize the features of clothing-relevant regions. Extensive experimental results demonstrate that our framework outperforms other baseline models on several widely-used CC Re-ID datasets.
Bo Li 0006, Guanglin Niu
ICASSP3
2025 Knowledge Distilled Group Prompts Learning for HOI Detection with Large Vision-Language Models
abstract
Large vision-language models (VLMs) have significantly advanced human-object interaction (HOI) detection. However, existing VLM-based HOI detectors primarily rely on simple text prompt paradigms, specifically in relation to knowledge hallucination, with limited exploration of the intrinsic attributes or extrinsic context. In this paper, we propose a knowledge distilled group prompts learning method for HOI detection, termed GPL-HOI, which transfer knowledge from vision-language models via group prompts and knowledge distillation. Specifically, we design visual-textual group prompts by combining scene-aware, region-aware, and pose-aware prompt to guide knowledge transfer from VLMs. Additionally, we introduce a cross-modal group distillation module,which aligns the semantic features of both the vision and text models via KL divergence, encouraging the visual encoder to generate similar probability distributions to the text encoder through the learnable prompts. Extensive experiments demonstrate that our method surpasses state-of-the-art approaches in both conventional and zero-shot settings, achieving improvements of +2.04 mAP and +1.84 mAP on HICO-DET, respectively. Code will be available at https://github.com/hxqstree/GPL-HOI.
Xiaoqian Han, Guanglin Niu, Mingliang Zhou 0001, Xiaowei Zhang 0003
ICME2
2025 CCUP: A Controllable Synthetic Data Generation Pipeline for Pretraining Cloth-Changing Person Re-Identification Models
abstract
Due to the high cost of constructing Cloth-changing person reidentification (CC-ReID) data, the existing data-driven models are hard to train efficiently on limited data, which causes the issue of overfitting. To address this challenge, we propose a low-cost and efficient pipeline specific to CC-ReID tasks for generating controllable and high-quality synthetic data simulating the surveillance scenarios. Particularly, we construct a new self-annotated CC-ReID dataset named Cloth-Changing Unreal Person (CCUP), containing 6,000 IDs, 1,179,976 images, 100 cameras, and 26.5 outfits per individual. Based on this large-scale dataset, we introduce an effective and scalable pretrain-finetune framework for enhancing the generalization of the traditional CC-ReID models. The extensive experimental results demonstrate that our framework could improve the original models such as two typical models TransReID and FIRe2after pretraining on CCUP and finetuning on a benchmark, and outperform other state-of-the-art models. The dataset is available at: https://github.com/yjzhao1019/CCUP.
Yujian Zhao, Chengru Wu, Yinong Xu, Xuanzheng Du, Ruiyu Li, Guanglin Niu
ICME6
2025 Try Harder: Hard Sample Generation and Learning for Cloth-Changing Person Re-ID
abstract
Hard samples pose a significant challenge in person re-identification (ReID) tasks, particularly in clothing-changing person Re-ID (CC-ReID). Their inherent ambiguity or similarity, coupled with the lack of explicit definitions, makes them a fundamental bottleneck. These issues not only limit the design of targeted learning strategies but also diminish the model's robustness under clothing or viewpoint changes. In this paper, we propose a novel multimodal-guided Hard Sample Generation and Learning (HSGL) framework, which is the first effort to unify textual and visual modalities to explicitly define, generate, and optimize hard samples within a unified paradigm. HSGL comprises two core components: (1) Dual-Granularity Hard Sample Generation (DGHSG), which leverages multimodal cues to synthesize semantically consistent samples, including both coarse- and fine-grained hard positives and negatives for effectively increasing the hardness and diversity of the training data. (2) Hard Sample Adaptive Learning (HSAL), which introduces a hardness-aware optimization strategy that adjusts feature distances based on textual semantic labels, encouraging the separation of hard positives and drawing hard negatives closer in the embedding space to enhance the model's discriminative capability and robustness to hard samples. Extensive experiments on multiple CC-ReID benchmarks demonstrate the effectiveness of our approach and highlight the potential of multimodal-guided hard sample generation and learning for robust CC-ReID. Notably, HSAL significantly accelerates the convergence of the targeted learning procedure and achieves state-of-the-art performance on both PRCC and LTCC datasets. The code is available at https://github.com/undooo/TryHarder-ACMMM25.
Hankun Liu, Yujian Zhao, Guanglin Niu
ACM Multimedia3
2025 Geometry-Guided Point Generation for 3D Object Detection
abstract
Point cloud completion 3D object detectors effectively tackle the challenge of incomplete shapes in sparse point clouds by generating pseudo points to improve detection performance. However, the absence of guidance provided by the heatmap information and the geometric shape information renders the precise recovery of object shapes an arduous task. To this end, we propose a Geometry-guided Point Generation for 3D Object Detection, named GgPG. Specifically, we first design a 3D heatmap auxiliary supervision subnetwork to enhance the quality of object proposals by capturing the actual size and position of the object within the 3D heatmap representation. Moreover, we introduce a density-aware point generation module that employs Kernel Density Estimation (KDE) to embed the point density into the grid point's feature representation, thereby enabling the completion of more precise object shapes. Our GgPG achieves progressive performance in both Waymo and KITTI benchmarks, notably GgPG outperforms PGRCNN by +1.02$\%$, +1.18$\%$, and +0.56$\%$on the vehicle, pedestrian, and cyclist under LEVEL$\_$2 mAPH classes on Waymo Open Dataset, respectively.
Mingliang Zhou 0001, Guanglin Niu, Xiaowei Zhang 0003
IEEE Signal Process. Lett.4
2025 A Pluggable Common Sense-Enhanced Framework for Knowledge Graph Completion
abstract
Knowledge graph completion (KGC) tasks aim to infer missing facts in a knowledge graph (KG) for many knowledgeintensive applications. However, existing embedding-based KGC approaches primarily rely on factual triples, potentially leading to outcomes inconsistent with common sense. Besides, generating explicit common sense is often impractical or costly for a KG. To address these challenges, we propose a pluggable common sense-enhanced KGC framework that incorporates both fact and common sense for KGC. This framework is adaptable to different KGs based on their entity concept richness and has the capability to automatically generate explicit or implicit common sense from factual triples. Furthermore, we introduce common senseguided negative sampling and a coarse-to-fine inference approach for KGs with rich entity concepts. For KGs without concepts, we propose a dual scoring scheme involving a relation-aware concept embedding mechanism. Importantly, our approach can be integrated as a pluggable module for many knowledge graph embedding (KGE) models, facilitating joint common sense and fact-driven training and inference. The experiments illustrate that our framework exhibits good scalability and outperforms existing models across various KGC tasks.
Guanglin Niu, Bo Li 0006, Siling Feng
IEEE Trans. Big Data1
2024 CAMEL: CAusal Motion Enhancement Tailored for Lifting Text-Driven Video Editing
abstract
Text-driven video editing poses significant challenges in exhibiting flicker-free visual continuity while preserving the inherent motion patterns of original videos. Existing methods operate under a paradigm where motion and appearance are intricately intertwined. This coupling leads to the network either over-fitting appearance content - failing to capture motion patterns - or focusing on motion patterns at the expense of content generalization to diverse textual scenarios. Inspired by the pivotal role of wavelet transform in dissecting video sequences, we propose CAusal Motion Enhancement tailored for Lifting text-driven video editing (CAMEL), a novel technique with two core designs. First, we introduce motion prompts, designed to summarize motion concepts from video templates through direct optimization. The optimized prompts are purposefully integrated into latent representations of diffusion models to enhance the motion fidelity of generated results. Second, to enhance motion coherence and extend the generalization of appearance content to creative textual prompts, we propose the causal motion-enhanced attention mechanism. This mechanism is implemented in tandem with a novel causal motion filter, synergistically enhancing the motion coherence of disentangled high-frequency components, and concurrently preserving the generalization of appearance content across various textual scenarios. Extensive experimental results show the superior performance of CAMEL.
Guiwei Zhang, Guanglin Niu, Zichang Tan, Yalong Bai, Qing Yang 0033
CVPR3
2024 Hierarchical bi-directional conceptual interaction for text-video retrieval
Wenpeng Han, Guanglin Niu, Mingliang Zhou 0001, Xiaowei Zhang 0003
Multim. Syst.2
2023 Logic and Commonsense-Guided Temporal Knowledge Graph Completion
abstract
A temporal knowledge graph (TKG) stores the events derived from the data involving time. Predicting events is extremely challenging due to the time-sensitive property of events. Besides, the previous TKG completion (TKGC) approaches cannot represent both the timeliness and the causality properties of events, simultaneously. To address these challenges, we propose a Logic and Commonsense-Guided Embedding model (LCGE) to jointly learn the time-sensitive representation involving timeliness and causality of events, together with the time-independent representation of events from the perspective of commonsense. Specifically, we design a temporal rule learning algorithm to construct a rule-guided predicate embedding regularization strategy for learning the causality among events. Furthermore, we could accurately evaluate the plausibility of events via auxiliary commonsense knowledge. The experimental results of TKGC task illustrate the significant performance improvements of our model compared with the existing approaches. More interestingly, our model is able to provide the explainability of the predicted results in the view of causal inference. The appendix, source code and datasets of this paper are available at https://github.com/ngl567/LCGE.
Guanglin Niu
AAAI1
2022 CAKE: A Scalable Commonsense-Aware Framework For Multi-View Knowledge Graph Completion
abstract
Knowledge graphs store a large number of factual triples while they are still incomplete, inevitably.The previous knowledge graph completion (KGC) models predict missing links between entities merely relying on fact-view data, ignoring the valuable commonsense knowledge.The previous knowledge graph embedding (KGE) techniques suffer from invalid negative sampling and the uncertainty of fact-view link prediction, limiting KGC's performance.To address the above challenges, we propose a novel and scalable Commonsense-Aware Knowledge Embedding (CAKE) framework to automatically extract commonsense from factual triples with entity concepts.The generated commonsense augments effective selfsupervision to facilitate both high-quality negative sampling (NS) and joint commonsense and fact-view link prediction.Experimental results 1 on the KGC task demonstrate that assembling our framework could enhance the performance of the original KGE models, and the proposed commonsense-aware NS module is superior to other NS techniques.Besides, our proposed framework could be easily adaptive to various KGE models and explain the predicted results.
Guanglin Niu, Bo Li 0006, Yongfei Zhang, Shiliang Pu
ACL (1)1
2022 Perform like an Engine: A Closed-Loop Neural-Symbolic Learning Framework for Knowledge Graph Inference
abstract
Knowledge graph (KG) inference aims to address the natural incompleteness of KGs, including rule learning-based and KG embedding (KGE) models. However, the rule learning-based models suffer from low efficiency and generalization while KGE models lack interpretability. To address these challenges, we propose a novel and effective closed-loop neural-symbolic learning framework EngineKG via incorporating our developed KGE and rule learning modules. KGE module exploits symbolic rules and paths to enhance the semantic association between entities and relations for improving KG embeddings and interpretability. A novel rule pruning mechanism is proposed in the rule learning module by leveraging paths as initial candidate rules and employing KG embeddings together with concepts for extracting more high-quality rules. Experimental results on four real-world datasets show that our model outperforms the relevant baselines on link prediction tasks, demonstrating the superiority of our KG inference model in a neural-symbolic learning fashion. The source code and datasets of this paper are available at https://github.com/ngl567/EngineKG.
Guanglin Niu, Bo Li 0006, Yongfei Zhang, Shiliang Pu
COLING1
2022 Joint semantics and data-driven path representation for knowledge graph reasoning
Guanglin Niu, Bo Li 0006, Yongfei Zhang, Yongpan Sheng, Chuan Shi 0001, Shiliang Pu
Neurocomputing1
2021 Relational Learning with Gated and Attentive Neighbor Aggregator for Few-Shot Knowledge Graph Completion
abstract
Aiming at expanding few-shot relations' coverage in knowledge graphs (KGs), few-shot knowledge graph completion (FKGC) has recently gained more research interests. Some existing models employ a few-shot relation's multi-hop neighbor information to enhance its semantic representation. However, noise neighbor information might be amplified when the neighborhood is excessively sparse and no neighbor is available to represent the few-shot relation. Moreover, modeling and inferring complex relations of one-to-many (1-N), many-to-one (N-1), and many-to-many (N-N) by previous knowledge graph completion approaches requires high model complexity and a large amount of training instances. Thus, inferring complex relations in the few-shot scenario is difficult for FKGC models due to limited training instances. In this paper, we propose a few-shot relational learning with global-local framework to address the above issues. At the global stage, a novel gated and attentive neighbor aggregator is built for accurately integrating the semantics of a few-shot relation's neighborhood, which helps filtering the noise neighbors even if a KG contains extremely sparse neighborhoods. For the local stage, a meta-learning based TransH (MTransH) method is designed to model complex relations and train our model in a few-shot learning fashion. Extensive experiments show that our model outperforms the state-of-the-art FKGC approaches on the frequently-used benchmark datasets NELL-One and Wiki-One. Compared with the strong baseline model MetaR, our model achieves 5-shot FKGC performance improvements of 8.0% on NELL-One and 2.8% on Wiki-One by the metric [email protected]
Guanglin Niu, Yang Li 0218, Chengguang Tang, Ruiying Geng, Hao Wang 0005, Jian Sun 0021, Fei Huang 0002, Luo Si
SIGIR1
2020 Rule-Guided Compositional Representation Learning on Knowledge Graphs
abstract
Representation learning on a knowledge graph (KG) is to embed entities and relations of a KG into low-dimensional continuous vector spaces. Early KG embedding methods only pay attention to structured information encoded in triples, which would cause limited performance due to the structure sparseness of KGs. Some recent attempts consider paths information to expand the structure of KGs but lack explainability in the process of obtaining the path representations. In this paper, we propose a novel Rule and Path-based Joint Embedding (RPJE) scheme, which takes full advantage of the explainability and accuracy of logic rules, the generalization of KG embedding as well as the supplementary semantic structure of paths. Specifically, logic rules of different lengths (the number of relations in rule body) in the form of Horn clauses are first mined from the KG and elaborately encoded for representation learning. Then, the rules of length 2 are applied to compose paths accurately while the rules of length 1 are explicitly employed to create semantic associations among relations and constrain relation embeddings. Moreover, the confidence level of each rule is also considered in optimization to guarantee the availability of applying the rule to representation learning. Extensive experimental results illustrate that RPJE outperforms other state-of-the-art baselines on KG completion task, which also demonstrate the superiority of utilizing logic rules as well as paths for improving the accuracy and explainability of representation learning.
Guanglin Niu, Yongfei Zhang, Bo Li 0006, Peng Cui 0001, Si Liu 0001, Xiaowei Zhang 0003
AAAI1
2019 Real-time object tracking via self-adaptive appearance modeling
Ming Xin 0003, Bo Li 0006, Guanglin Niu
Neurocomputing4