Ping Kuang

dblp:78/586 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0003-3088-7135ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs
abstract
Object hallucination remains a critical challenge in Large Vision-Language Models (LVLMs), where models generate content inconsistent with visual inputs. Existing language-decoder based mitigation approaches often regulate visual or textual attention independently, overlooking their interaction as two key causal factors. To address this, we propose Owl (Bi-mOdal attention reWeighting for Layer-wise hallucination mitigation), a causally-grounded framework that models hallucination process via a structural causal graph, treating decomposed visual and textual attentions as mediators. We introduce VTACR (Visual-to-Textual Attention Contribution Ratio), a novel metric that quantifies the modality contribution imbalance during decoding. Our analysis reveals that hallucinations frequently occur in low-VTACR scenarios, where textual priors dominate and visual grounding is weakened. To mitigate this, we design a fine-grained attention intervention mechanism that dynamically adjusts token- and layer-wise attention guided by VTACR signals. Finally, we propose a dual-path contrastive decoding strategy: one path emphasizes visually grounded predictions, while the other amplifies hallucinated ones -- letting visual truth shine and hallucination collapse. Experimental results on the POPE and CHAIR benchmarks show that Owl achieves significant hallucination reduction, setting a new SOTA in faithfulness while preserving vision-language understanding capability. Our code is available at https://github.com/CikZ2023/OWL
Liu Yu 0001, Zhonghao Chen, Ping Kuang, Zhikun Feng, Fan Zhou 0002, Gillian Dobbie
AAAI3
2026 Enhancing decision boundaries in continual learning through a decoupled Gaussian framework
abstract
The goal of continual learning (CL) is to acquire new knowledge while retaining previously learned information. CNN-based and prompt-based CL methods have achieved remarkable progress in recent years. However, most prior work has primarily focused on reducing forgetting from the perspective of the model itself. In this paper, we investigate CL from the perspective of decision boundaries, analyzing the impact of instance-level feature overlap. To address this issue, we propose a generic Decoupled Gaussian Softmax Classifier that enhances class discriminability during CL process. Specifically, we decouple the features extracted by the backbone into multiple Gaussian distributions, which are directly fused into the feature space through weighted integration. A regularization term is introduced to penalize the overlap of similar features, while an adaptive decision boundary is assigned to each class to encourage inter-class separation and intra-class compactness. Experiments on 4 widely used continual learning datasets and 12 CL scenarios show that our method has good plug-and-play capability. It improves the average accuracy by 1%–2.63% over the baseline models, while effectively reducing both the forgetting rate and the Expected Calibration Error. Our code is available at: https://anonymous.4open.science/r/DGSC-main-310D .
Zhikun Feng, Liu Yu 0001, Ping Kuang, Mian Zhou, Kang Dang, Yakun Ju
Inf. Process. Manag.5
2026 TIPS: Two-level prompt selection for more stability-plasticity balance in continual learning
Zhikun Feng, Kang Dang, Mian Zhou, Ping Kuang, Mingyu Wu 0011, Liu Yu 0001, Jionglong Su
Pattern Recognit.5
2025 Bridging the Fairness Gap: Enhancing Pre-trained Models with LLM-Generated Sentences
abstract
Pre-trained language models (PLMs) are trained on data that inherently contains gender biases, leading to undesirable impacts. Traditional debiasing methods often rely on external corpora, which may lack quality, diversity, or demographic balance, affecting the effectiveness of debiasing. With the rise of large language models and their extensive knowledge, we propose enhancing fairness (Fair-Gender) in PLMs by absorbing coherent, attribute-balanced, and semantically rich sentences. However, these sentences cannot be directly used for debiasing due to alignment issues and the risk of negative transfer. We address this by applying causal analysis to estimate causal effects, filtering out unaligned sentences, and identifying aligned ones for incorporation into PLMs, thereby ensuring positive transfer. Experiments show that our approach significantly reduces gender biases in PLMs while preserving their language expressiveness.
Liu Yu 0001, Ludie Guo, Ping Kuang, Fan Zhou 0002
ICASSP3
2025 Decoupling Overlapped Feature Spaces: When Continual Learning Meets Fine-Grain Classification
abstract
The goal of Class Incremental Learning (CIL) is to continuously learn new classes while preventing forgetting of old ones. Most previous works focused on reducing catastrophic forgetting from model’s perspective. However, the model is not the only factor contributing to forgetting. In this paper, we take the perspective of class instances and find that fine-grained class increments can lead to feature overlap between classes, further reducing instance margins. We call this interesting phenomenon as Fine-grained class confusion effect in CIL. Since preserving instance margins is crucial for resisting forgetting, it is beneficial to maintain the margin amount as much as possible. To achieve this, we propose a general Gaussian decoupling classifier to enhance the discriminability of similar classes during incremental learning. Specifically, we decouple the features of different classes extracted by the backbone network into multiple independent Gaussian distributions. By directly integrating them into the features with weighted fusion, we introduce a regularization penalty that encourages minimizing the overlap of similar features, thus increasing the feature distance between classes. Extensive experiments show that our method effectively improves class separation and better preserves instance margins, ultimately alleviating forgetting. The improved model achieves better performance on CUB-200 and CARS-196.
Zhikun Feng, Mingyu Wu 0011, Ping Kuang, Kang Dang, Mian Zhou, Liu Yu 0001
ICME3
2025 Delight-UPS: Uncalibrated Photometric Stereo via Diffusion Model-Based Relighting
abstract
Photometric stereo aims to recover detailed surface normal maps from images captured under varying illuminations. However, existing methods often rely on extensive real images under controlled lighting conditions to achieve accurate surface normal estimation and struggle to perform effectively when the number of input images is limited.Hence, we propose Diffusion Model based Relighting for Uncalibrated Photometric Stereo (Delight-UPS) to address the above limitations. Specifically, we first employ diffusion model to simplify complex illumination with wide-angle light. Then, narrow-angle light is applied to generate images with object illumination variations. Finally, we recover the normal map from the illumination-variant images.This enables photometric stereo using only a single image, even under complex lighting conditions.Experiments show that Delight-UPS outperforms the SOTA model SDM by 38.0% using limited input images.Moreover, Delight-UPS improves previous models’ performance by 18.4% on average and allows them to recover normal maps from synthetic images.* Our source code is anonymously hosted at ${\text{Delight - UPS}}$.
Zhenyu Qiao, Mingyun He, Liu Yu 0001, Rui Zhou 0012, Ping Kuang
ICME6
2025 Knowledge Graphs Acquisition via Forward-Reverse Relation Enhanced Contrastive Pretraining from Large-scale Models
abstract
Commonsense knowledge graph acquisition (CKGA) is essential for many knowledge-intensive applications, such as natural language understanding, question answering, and conversational system. Traditional methods that directly use word-level triplet for knowledge generation often lack sufficient context, leading to ambiguous relations and a limited ability to handle complex or abstract concepts. Moreover, they also rely on forward relations and struggle to fully capture the reverse connections between entities, which can lead to the "reversal curse". To address these, we firstly transform pre-defined relations into sentence templates, and introduce a new pipeline that enhances CKGA task via bi-directional relation-enhanced contrastive pretraining from recent large-scale foundation models. Our closed-loop Bi-REACT includes data preprocessing, contrastive pre-training, task-driven instruction tuning, filtering model, and evaluation system. Experiments show Bi-REACT can easily harvest extensive high-quality knowledge (390K), achieving up to 90.02% (ATOMIC) and 84.29% (ConceptNet) accuracy, approaching human-level performance for these resources, and effectively reducing the "reversal curse" issues. Our code is available at https://anonymous.4open.science/r/CKGA-345E.
Liu Yu 0001, Fenghui Tian, Ping Kuang, Zhikun Feng, Fan Zhou 0002
ICME3
2025 Bimodal Debiasing for Text-to-Image Diffusion: Adaptive Guidance in Textual and Visual Spaces
Liu Yu 0001, Ping Kuang, Rui Zhou 0012, Fan Zhou 0002, Zhikun Feng
ACM Multimedia3
2025 Amplifying commonsense knowledge via bi-directional relation integrated graph-based contrastive pre-training from large language models
Liu Yu 0001, Fenghui Tian, Ping Kuang, Fan Zhou 0002
Inf. Process. Manag.3
2024 Disentanglement-Guided Spatial-Temporal Graph Neural Network for Metro Flow Forecasting (Student Abstract)
abstract
In recent intelligent transportation applications, metro flow forecasting has received much attention from researchers. Most prior arts endeavor to explore spatial or temporal dependencies while ignoring the key characteristic patterns underlying historical flows, e.g., trend and periodicity. Although the multiple granularity distillations or spatial dependency correlation can promote the flow estimation. However, the potential noise and spatial dynamics are under-explored. To this end, we propose a novel Disentanglement-Guided Spatial-Temporal Graph Neural Network or DGST to address the above concerns. It contains a Disentanglement Pre-training procedure for characteristic pattern disentanglement learning, a Characteristic Pattern Prediction for different future characteristic explorations, and a Spatial-Temporal Correlation for spatial-temporal dynamic learning. Experiments on a real-world dataset demonstrate the superiority of our DGST.
Jinyu Hong, Ping Kuang, Qiang Gao 0003, Fan Zhou 0002
AAAI2
2024 Biases Mitigation and Expressiveness Preservation in Language Models: A Comprehensive Pipeline (Student Abstract)
abstract
Pre-trained language models (PLMs) have greatly transformed various downstream tasks, yet frequently display social biases from training data, raising fairness concerns. Recent efforts to debias PLMs come with limitations: they either fine-tune the entire parameters in PLMs, which is time-consuming and disregards the expressiveness of PLMs, or ignore the reintroducing biases from downstream tasks when applying debiased models to them. Hence, we propose a two-stage pipeline to mitigate biases from both internal and downstream contexts while preserving expressiveness in language models. Specifically, for the debiasing procedure, we resort to continuous prefix-tuning, not fully fine-tuning the PLM, in which we design a debiasing term for optimization and an alignment term to keep words’ relative distances and ensure the model's expressiveness. For downstream tasks, we perform causal intervention across different demographic groups for invariant predictions. Results on three GLUE tasks show our method alleviates biases from internal and downstream contexts, while keeping PLM expressiveness intact.
Liu Yu 0001, Ludie Guo, Ping Kuang, Fan Zhou 0002
AAAI3
2024 Amplifying Diversity and Quality in Commonsense Knowledge Graph Completion (Student Abstract)
abstract
Conventional commonsense knowledge graph completion (CKGC) methods provide inadequate sequence when fine-tuning or generating stages and incorporate full fine-tuning, which fail to align with the autoregressive model's pre-training patterns and have insufficient parameter efficiency. Moreover, decoding through beam or greedy search produces low diversity and high similarity in generated tail entities. Hence, we resort to prefix-tuning and propose a lightweight, effective pipeline to enhance the quality and diversity of extracted commonsense knowledge. Precisely, we measure head entity similarity to yield and then concatenate top-k tuples before each target tuple for prefix-tuning the source LM, thereby improving the efficiency and speed for pretrained models; then, we design a penalty-tailored diverse beam search (p-DBS) for decoding tail entities, producing a greater quantity and diversity of generated commonsense tuples; besides, a filter strategy is utilized to filter out invalid commonsense knowledge. Through extensive automatic evaluations, including ChatGPT scoring, our method can extract diverse, novel, and accurate commonsense knowledge (CK).
Liu Yu 0001, Fenghui Tian, Ping Kuang, Fan Zhou 0002
AAAI3
2024 Predicting Human Mobility via Self-Supervised Disentanglement Learning
abstract
Deep neural networks have recently achieved considerable improvements in learning human behavioral patterns and individual preferences from massive spatial-temporal trajectory data. However, most of the existing research concentrates on fusing different semantics underlying sequential trajectories for mobility pattern learning which, in turn, yields a narrow perspective on comprehending human intrinsic motions. In addition, the inherent sparsity and under-explored heterogeneous collaborative items pertaining to human check-ins hinder the potential exploitation of human diverse periodic regularities as well as common interests. Motivated by recent advances in disentanglement learning, we propose a novel disentangled solution called SSDL for tackling the next POI prediction problem. SSDL primarily seeks to disentangle the potential time-invariant and time-varying factors into different latent spaces from massive trajectories, providing an interpretable view to understand the intricate semantics underlying human diverse mobility representations. To address the data sparsity issue, we present two realistic trajectory augmentation approaches to enhance the understanding of both the human intrinsic periodicity/habits and constantly-changing intents. In addition, we devise a POI-centric graph structure to explore heterogeneous collaborative signals underlying historical check-ins. Extensive experiments conducted on four real-world datasets demonstrate that SSDL significantly outperforms the state-of-the-art approaches–for example, it yields up to 8.57% averaged improvement on ACC@1.
Qiang Gao 0003, Jinyu Hong, Xovee Xu, Ping Kuang, Fan Zhou 0002, Goce Trajcevski
IEEE Trans. Knowl. Data Eng.4
2023 Mobility Prediction via Sequential Trajectory Disentanglement (Student Abstract)
abstract
Accurately predicting human mobility is a critical task in location-based recommendation. Most prior approaches focus on fusing multiple semantics trajectories to forecast the future movement of people, and fail to consider the distinct relations in underlying context of human mobility, resulting in a narrow perspective to comprehend human motions. Inspired by recent advances in disentanglement learning, we propose a novel self-supervised method called SelfMove for next POI prediction. SelfMove seeks to disentangle the potential time-invariant and time-varying factors from massive trajectories, which provides an interpretable view to understand the complex semantics underlying human mobility representations. To address the data sparsity issue, we present two realistic trajectory augmentation approaches to help understand the intrinsic periodicity and constantly changing intents of humans. In addition, a POI-centric graph structure is proposed to explore both homogeneous and heterogeneous collaborative signals behind historical trajectories. Experiments on two real-world datasets demonstrate the superiority of SelfMove compared to the state-of-the-art baselines.
Jinyu Hong, Fan Zhou 0002, Qiang Gao 0003, Ping Kuang, Kunpeng Zhang 0001
AAAI4
2021 Robust and Dynamic Graph Convolutional Network For Multi-view Data Classification
abstract
Abstract Since graph learning could preserve the structure information of the samples to improve the learning ability, it has been widely applied in both shallow learning and deep learning. However, the current graph learning methods still suffer from the issues such as outlier influence and model robustness. In this paper, we propose a new dynamic graph neural network (DGCN) method to conduct semi-supervised classification on multi-view data by jointly conducting the graph learning and the classification task in a unified framework. Specifically, our method investigates three strategies to improve the quality of the graph before feeding it into the GCN model: (i) employing robust statistics to consider the sample importance for reducing the outlier influence, i.e. assigning every sample with soft weights so that the important samples are with large weights and outliers are with small or even zero weights; (ii) learning the common representation across all views to improve the quality of the graph for every view; and (iii) learning the complementary information from all initial graphs on multi-view data to further improve the learning of the graph for every view. As a result, each of the strategies could improve the robustness of the DGCN model. Moreover, they are complementary for reducing outlier influence from different aspects, i.e. the sample importance reduces the weights of the outliers, both the common representation and the complementary information improve the quality of the graph for every view. Experimental result on real data sets demonstrates the effectiveness of our method, compared to the comparison methods, in terms of multi-class classification performance.
Fei Kong, Chongzhi Liu, Ping Kuang
Comput. J.4
2021 An anchor-free object detector with novel corner matching method
Tingsong Ma, Wenhong Tian, Ping Kuang, Yuanlun Xie
Knowl. Based Syst.3
2020 An improved recurrent neural networks for 3d object reconstruction
Tingsong Ma, Ping Kuang, Wenhong Tian
Appl. Intell.2
2019 Image super-resolution with densely connected convolutional networks
Ping Kuang, Tingsong Ma
Appl. Intell.1
2019 SAVE: self-adaptive consolidation of virtual machines for energy efficiency of CPU-intensive applications in the cloud
Wenxia Guo, Ping Kuang, Yaqiu Jiang, Wenhong Tian
J. Supercomput.2
2018 Real-Time Pedestrian Detection Using Convolutional Neural Networks
abstract
Pedestrian detection provides manager of a smart city with a great opportunity to manage their city effectively and automatically. Specifically, pedestrian detection technology can improve our secure environment and make our traffic more efficient. In this paper, all of our work both modification and improvement are made based on YOLO, which is a real-time Convolutional Neural Network detector. In our work, we extend YOLO’s original network structure, and also give a new definition of loss function to boost the performance for pedestrian detection, especially when the targets are small, and that is exactly what YOLO is not good at. In our experiment, the proposed model is tested on INRIA, UCF YouTube Action Data Set and Caltech Pedestrian Detection Benchmark. Experimental results indicate that after our modification and improvement, the revised YOLO network outperforms the original version and also is better than other solutions.
Ping Kuang, Tingsong Ma
Int. J. Pattern Recognit. Artif. Intell.1
2004 Real-Time Strategy and Practice in Service Grid
abstract
The emerging service grids bring together various distributed application-level services to a 'market' for clients to request and enable the integration of services across distributed, heterogeneous, dynamic virtual organizations. However, there are a number of applications with the requirement of time constraints. We propose a real-time strategy in service grid architecture. We also extend the OGSI grid service semantics for fault-tolerance. The real-time and fault-tolerant strategies seem efficient through experiments.
Hai Jin 0001, Hanhua Chen, Jian Chen 0030, Ping Kuang, Deqing Zou
COMPSAC4