EDBT 2026 Demo / reviewers in the wild / expert
Qi Liu 0005
dblp:95/2446-5
· DBLP profile ↗
37ranked-venue papers
7as first author
32since 2021 · last 2026
0000-0001-5378-6404ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 20 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What-Meets-Where: Unified Learning of Action and Contact Localization in ImagesabstractPeople control their bodies to establish contact with the environment. To comprehensively understand actions across diverse visual contexts, it is essential to simultaneously consider what action is occurring and where it is happening. Current methodologies, however, often inadequately capture this duality, typically failing to jointly model both action semantics and their spatial contextualization within scenes. To bridge this gap, we introduce a novel vision task that simultaneously predicts high-level action semantics and fine-grained body-part contact regions. Our proposed framework, PaIR-Net, comprises three key components: the Contact Prior Aware Module (CPAM) for identifying contact-relevant body parts, the Prior-Guided Concat Segmenter (PGCS) for pixel-wise contact segmentation, and the Interaction Inference Module (IIM) responsible for integrating global interaction relationships. To facilitate this task, we present PaIR (Part-aware Interaction Representation), a comprehensive dataset containing 13,979 images that encompass 654 actions, 80 object categories, and 17 body parts. Experimental evaluation demonstrates that PaIR-Net significantly outperforms baseline approaches, while ablation studies confirm the efficacy of each architectural component. Yuxiao Wang 0003, Wolin Liang, Weiying Xue, Zhenao Wei, Nan Zhuang, Qi Liu 0005 |
AAAI | 7 |
| 2026 | QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions in images. Although DETR-based methods have recently emerged as the mainstream framework for HOI detection, they still suffer from a key limitation: Randomly initialized queries lack explicit semantics, leading to suboptimal detection performance. To address this challenge, we propose QueryCraft, a novel plug-and-play HOI detection framework that incorporates semantic priors and guided feature learning through transformer-based query initialization. Central to our approach is ACTOR (Action-aware Cross-modal TransfORmer), a cross-modal Transformer encoder that jointly attends to visual regions and textual prompts to extract action-relevant features. Rather than merely aligning modalities, ACTOR leverages language-guided attention to infer interaction semantics and produce semantically meaningful query representations. To further enhance object-level query quality, we introduce a Perceptual Distilled Query Decoder (PDQD), which distills object category awareness from a pre-trained detector to serve as object query initiation. This dual-branch query initialization enables the model to generate more interpretable and effective queries for HOI detection. Extensive experiments on HICO-Det and V-COCO benchmarks demonstrate that our method achieves state-of-the-art performance and strong generalization. Yuxiao Wang 0003, Wolin Liang, Weiying Xue, Nan Zhuang, Qi Liu 0005 |
AAAI | 6 |
| 2026 | Melodia: Training-Free Music Editing Guided by Attention Probing in Diffusion ModelsabstractText-to-music generation technology is progressing rapidly, creating new opportunities for musical composition and editing. However, existing music editing methods often fail to preserve the source music's temporal structure, including melody and rhythm, when altering particular attributes like instrument, genre, and mood. To address this challenge, this paper conducts an in-depth probing analysis on attention maps within AudioLDM 2, a diffusion-based model commonly used as the backbone for existing music editing methods. We reveal a key finding: cross-attention maps encompass details regarding distinct musical characteristics, and interventions on these maps frequently result in ineffective modifications. In contrast, self-attention maps are essential for preserving the temporal structure of the source music during its conversion into the target music. Building upon this understanding, we present Melodia, a training-free technique that selectively manipulates self-attention maps in particular layers during the denoising process and leverages an attention repository to store source music information, achieving accurate modification of musical characteristics while preserving the original structure without requiring textual descriptions of the source music. Additionally, we propose two novel metrics to better evaluate music editing methods. Both objective and subjective experiments demonstrate that our approach achieves superior results in terms of textual adherence and structural integrity across various datasets. This research enhances comprehension of internal mechanisms within music generation models and provides improved control for music creation. Haowen Li 0001, Boyu Cao, Qi Liu 0005 |
AAAI | 7 |
| 2026 | Depth-consistent 3D Gaussian Splatting via physical defocus modeling and multi-view geometric supervision
Baozhu Zhao, Junyan Su, Qi Liu 0005 |
Neural Networks | 5 |
| 2026 | Metamon-GS: Enhancing representability with variance-guided densification and light encoding
Junyan Su, Baozhu Zhao, Qi Liu 0005 |
Neural Networks | 4 |
| 2026 | PointCore: An efficient framework for unsupervised point cloud anomaly detection using joint local-global features
Baozhu Zhao, Qi Liu 0005 |
Neural Networks | 4 |
| 2026 | SAEN-BGS: Energy-efficient spiking autoencoder network for background subtraction
Xiaopeng Li 0005, Qi Liu 0005 |
Pattern Recognit. | 3 |
| 2026 | Beyond deceptive flatness: Dual-order solution for strengthening adversarial transferability
Pingyu Wang, Xingjian Zheng, Linbo Qing, Qi Liu 0005 |
Pattern Recognit. | 5 |
| 2026 | CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene GenerationabstractChallenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on scene interpretation. These factors hinder the creation of viewpoint-consistent and semantically aligned visual content, as well as the construction of accurate geometric structures. In this paper, we propose CGGS, a text-to-3D framework aiming to enhance 3D-content-awareness and address geometric distortions in ego-centric scene generation. Firstly, the Ego-centric Generator is proposed by fine-tuning a Multi-View Latent Diffusion Model with consistency-augmented loss to generate consistent, high-fidelity 2D content aligned with textual descriptions. Then, Layout Decorator leverages optical flow and point-track correspondence to estimate depth, therefore producing dense point clouds as coarse layouts from the ego-centric 2D priors. Building on this initialization, Geometric Refiner is proposed to enhance 3D Gaussian reconstruction via an entropy-based Mutual Information Depth Loss (MID) combined with a hierarchical optimization scheme for improving visual quality and geometric structure. Comprehensive experiments demonstrate that CGGS outperforms previous methods in generating coherent and accurate text-driven 3D scenes. Qi Liu 0005, Huan Wang 0014 |
IEEE Trans. Image Process. | 3 |
| 2025 | Semi-IIN: Semi-Supervised Intra-Inter Modal Interaction Learning Network for Multimodal Sentiment AnalysisabstractDespite multimodal sentiment analysis being a fertile research ground that merits further investigation, current approaches take up high annotation cost and suffer from label ambiguity, non-amicable to high-quality labeled data acquisition. Furthermore, choosing the right interactions is essential because the significance of intra- or inter-modal interactions can differ among various samples. To this end, we propose Semi-IIN, a Semi-supervised Intra-inter modal Interaction learning Network for multimodal sentiment analysis. Semi-IIN integrates masked attention and gating mechanisms, enabling effective dynamic selection after independently capturing intra- and inter-modal interactive information. Combined with the self-training approach, Semi-IIN fully utilizes the knowledge learned from unlabeled data. Experimental results on two public datasets, MOSI and MOSEI, demonstrate the effectiveness of Semi-IIN, establishing a new state-of-the-art on several metrics. Jinhao Lin, Yanwu Xu 0004, Qi Liu 0005 |
AAAI | 4 |
| 2025 | Precision-Enhanced Human-Object Contact Detection via Depth-Aware Perspective Interaction and Object Texture RestorationabstractHuman-object contact (HOT) is designed to accurately identify the areas where humans and objects come into contact. Current methods frequently fail to account for scenarios where objects are frequently blocking the view, resulting in inaccurate identification of contact areas. To tackle this problem, we suggest using a perspective interaction HOT detector called PIHOT, which utilizes a depth map generation model to offer depth information of humans and objects related to the camera, thereby preventing false interaction detection. Furthermore, we use mask dilatation and object restoration techniques to restore the texture details in covered areas, improve the boundaries between objects, and enhance the perception of humans interacting with objects. Moreover, a spatial awareness perception is intended to concentrate on the characteristic features close to the points of contact. The experimental results show that the PIHOT algorithm achieves state-of-the-art performance on three benchmark datasets for HOT detection tasks. Compared to the most recent DHOT, our method enjoys an average improvement of 13%, 27.5%, 16%, and 18.5% on SC-Acc., C-Acc., mIoU, and wIoU metrics, respectively. Yuxiao Wang 0003, Wenpeng Neng, Zhenao Wei, Weiying Xue, Nan Zhuang, Yanwu Xu 0004, Qi Liu 0005 |
AAAI | 9 |
| 2025 | Toy-GS: Assembling Local Gaussians for Precisely Rendering Large-Scale Free Camera TrajectoriesabstractCurrently, 3D rendering for large-scale free camera trajectories, namely, arbitrary input camera trajectories, poses significant challenges: 1) The distribution and observation angles of the cameras are irregular, and various types of scenes are included in the free trajectories; 2) Processing the entire point cloud and all images at once for large-scale scenes requires a substantial amount of GPU memory. This paper presents a Toy-GS method for accurately rendering large-scale free camera trajectories. Specifically, we propose an adaptive spatial division approach for free trajectories to divide cameras and the sparse point cloud of the entire scene into various regions according to camera poses. Training each local Gaussian in parallel for each area enables us to concentrate on texture details and minimize GPU memory usage. Next, we use the multi-view constraint and position-aware point adaptive control (PPAC) to improve the rendering quality of texture details. In addition, our regional fusion approach combines local and global Gaussians to enhance rendering quality with an increasing number of divided areas. Extensive experiments have been carried out to confirm the effectiveness and efficiency of Toy-GS, leading to state-of-the-art results on two public large-scale datasets as well as our SCUTic dataset. Our proposal demonstrates an enhancement of 1.19 dB in PSNR and conserves 7 G of GPU memory when compared to various benchmarks. Yukui Qiu, Junyan Su, Qi Liu 0005 |
AAAI | 5 |
| 2025 | LLM Agents Can Be Choice-Supportive Biased Evaluators: An Empirical StudyabstractWith Large Language Model (LLM) agents taking on more evaluation responsibilities in decision-making, it is essential to recognize their possible biases to guarantee fair and trustworthy AI-supported decisions. This study is the first to thoroughly examine the choice-supportive bias in LLM agents, a cognitive bias that is known to impact human decision-making and evaluation. We conduct experiments across 19 open/unopen-source LLM models in five scenarios at maximum, employing both memory-based and evaluation-based tasks adapted and redesigned from human cognitive studies. Our findings show that LLM agents may exhibit biased attribution or evaluation that supports their initial choices, and such bias may persist even if contextual hallucination is not observable. Key findings show that bias manifestation can differ greatly depending on prompt construction and context preservation, and the bias may be mitigated in larger models. Significantly, we observe that the bias increases when the agents perceive they are in control. Our extensive study involving 284 well-educated humans shows that, despite bias, certain LLM agents can still perform better than humans in similar evaluation tasks. This research contributes to the growing area of AI psychology, and the findings underscore the importance of addressing cognitive biases in LLM Agent systems, with wide-ranging implications spanning from improving AI-assisted decision-making to advancing AI safety and ethics. Nan Zhuang, Boyu Cao, Mingda Xu, Yuxiao Wang 0003, Qi Liu 0005 |
AAAI | 7 |
| 2025 | Prompt Guidance and Human Proximal Perception for HOT Prediction with Regional Joint LossabstractThe task of Human-Object conTact (HOT) detection involves identifying the specific areas of the human body that are touching objects. Nevertheless, current models are restricted to just one type of image, often leading to too much segmentation in areas with little interaction, and struggling to maintain category consistency within specific regions. To tackle this issue, a HOT framework, termed \textbf{P3HOT}, is proposed, which blends \textbf{P}rompt guidance and human \textbf{P}roximal \textbf{P}erception. To begin with, we utilize a semantic-driven prompt mechanism to direct the network's attention towards the relevant regions based on the correlation between image and text. Then a human proximal perception mechanism is employed to dynamically perceive key depth range around the human, using learnable parameters to effectively eliminate regions where interactions are not expected. Calculating depth resolves the uncertainty of the overlap between humans and objects in a 2D perspective, providing a quasi-3D viewpoint. Moreover, a Regional Joint Loss (RJLoss) has been created as a new loss to inhibit abnormal categories in the same area. A new evaluation metric called ``AD-Acc.'' is introduced to address the shortcomings of existing methods in addressing negative samples. Comprehensive experimental results demonstrate that our approach achieves state-of-the-art performance in four metrics across two benchmark datasets. Specifically, our model achieves an improvement of \textbf{0.7}$\uparrow$, \textbf{2.0}$\uparrow$, \textbf{1.6}$\uparrow$, and \textbf{11.0}$\uparrow$ in SC-Acc., mIoU, wIoU, and AD-Acc. metrics, respectively, on the HOT-Annotated dataset. The sources code are available at https://github.com/YuxiaoWang-AI/P3HOT. Yuxiao Wang 0003, Zhenao Wei, Weiying Xue, Nan Zhuang, Qi Liu 0005 |
ICCV | 7 |
| 2025 | Towards zero-shot human-object interaction detection via vision-language integration
Weiying Xue, Qi Liu 0005, Yuxiao Wang 0003, Zhenao Wei, Xiaofen Xing, Xiangmin Xu 0001 |
Neural Networks | 2 |
| 2025 | A Stable and Efficient Data-Free Model Attack With Label-Noise Data GenerationabstractThe objective of a data-free closed-box adversarial attack is to attack a victim model without using internal information, training datasets or semantically similar substitute datasets. Concerned about stricter attack scenarios, recent studies have tried employing generative networks to synthesize data for training substitute models. Nevertheless, these approaches concurrently encounter challenges associated with unstable training and diminished attack efficiency. In this paper, we propose a novel query-efficient data-free closed-box adversarial attack method. To mitigate unstable training, for the first time, we directly manipulate the intermediate-layer feature of a generator without relying on any substitute models. Specifically, a label noise-based generation module is created to enhance the intra-class patterns by incorporating partial historical information during the learning process. Additionally, we present a feature-disturbed diversity generation method to augment the inter-class distance. Meanwhile, we propose an adaptive intra-class attack strategy to heighten attack capability within a limited query budget. In this strategy, entropy-based distance is utilized to characterize the relative information from model outputs, while positive classes and negative samples are used to enhance low attack efficiency. The comprehensive experiments conducted on six datasets demonstrate the superior performance of our method compared to six state-of-the-art data-free closed-box competitors in both label-only and probability-only attack scenarios. Intriguingly, our method can realize the highest attack success rate on the online Microsoft Azure model under an extremely low query budget. Additionally, the proposed approach not only achieves more stable training but also significantly reduces the query count for a more balanced data generation. Furthermore, our method can maintain the best performance under the existing defense models and a limited query budget. Xingjian Zheng, Linbo Qing, Qi Liu 0005, Pingyu Wang, Yu Liu 0123, Jiyang Liao |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | R3CD: Scene Graph to Image Generation with Relation-Aware Compositional Contrastive Control DiffusionabstractImage generation tasks have achieved remarkable performance using large-scale diffusion models. However, these models are limited to capturing the abstract relations (viz., interactions excluding positional relations) among multiple entities of complex scene graphs. Two main problems exist: 1) fail to depict more concise and accurate interactions via abstract relations; 2) fail to generate complete entities. To address that, we propose a novel Relation-aware Compositional Contrastive Control Diffusion method, dubbed as R3CD, that leverages large-scale diffusion models to learn abstract interactions from scene graphs. Herein, a scene graph transformer based on node and edge encoding is first designed to perceive both local and global information from input scene graphs, whose embeddings are initialized by a T5 model. Then a joint contrastive loss based on attention maps and denoising steps is developed to control the diffusion model to understand and further generate images, whose spatial structures and interaction features are consistent with a priori relation. Extensive experiments are conducted on two datasets: Visual Genome and COCO-Stuff, and demonstrate that the proposal outperforms existing models both in quantitative and qualitative metrics to generate more realistic and diverse images according to different scene graph specifications. Jinxiu Liu, Qi Liu 0005 |
AAAI | 2 |
| 2024 | AG-NeRF: Attention-Guided Neural Radiance Fields for Multi-height Large-Scale Outdoor Scene Rendering
Baozhu Zhao, Qi Liu 0005 |
PRCV (6) | 4 |
| 2024 | TED-Net: Dispersal Attention for Perceiving Interaction Region in Indirectly-Contact HOI DetectionabstractHuman-Object Interaction (HOI) detection is a fertile research ground that merits further investigation in computer vision, and plays an important role in image high-level semantic information understanding. To achieve superior object detection performance, existing HOI models predominantly concentrate on the corresponding bounding box information of humans and objects, respectively, and ignore their surrounding information, thus it results in imprecise inference of instance interaction, which is severe for indirectly-contact interaction images (Intersection-over-Union = 0). To address that, a novel Triple stream Enhanced encoder-decoder Dispersal Network (TED-Net), equipped with human, object, and instance interaction decoding streams, is proposed to decouple instances’ relationships. Meanwhile, we design a dispersal attention mechanism to capture indirectly-contact interaction information and an auxiliary discrimination mechanism to improve the ability of instance interaction decoding stream for action category recognition. Experimental results show that the proposed TED-Net achieves the best performance among HOI models using the ResNet-50 backbone on the (big) HICO-Det dataset and comes third on the (small) V-COCO dataset in leaderboard1. Additionally, two indirectly-contact interaction datasets, namely, HICO-Det-IC and V-COCO-IC, are constructed to demonstrate the usefulness and effectiveness of our TED-Net in interacting between indirectly-contact instances, with an average of +3.80 mAP on HICO-Det-IC and +5.46 mAP on V-COCO-IC. Code is available at https://drliuqi.github.io/. Yuxiao Wang 0003, Qi Liu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Fast Robust Matrix Completion via Entry-Wise ℓ0-Norm MinimizationabstractMatrix completion (MC) aims at recovering missing entries, given an incomplete matrix. Existing algorithms for MC are mainly designed for noiseless or Gaussian noise scenarios and, thus, they are not robust to impulsive noise. For outlier resistance, entry-wise$\ell _{p}$-norm with$0 < p < 2$and M-estimation are two popular approaches. Yet the optimum selection of$p$for the entrywise$\ell _{p}$-norm-based methods is still an open problem. Besides, M-estimation is limited by a breakdown point, that is, the largest proportion of outliers. In this article, we adopt entrywise$\ell _{0}$-norm, namely, the number of nonzero entries in a matrix, to separate anomalies from the observed matrix. Prior to separation, the Laplacian kernel is exploited for outlier detection, which provides a strategy to automatically update the entrywise$\ell _{0}$-norm penalty parameter. The resultant multivariable optimization problem is addressed by block coordinate descent (BCD), yielding$\ell _{0}$-BCD and$\ell _{0}$-BCD-F. The former detects and separates outliers, as well as its convergence is guaranteed. In contrast, the latter attempts to treat outlier-contaminated elements as missing entries, which leads to higher computational efficiency. Making use of majorization–minimization (MM), we further propose$\ell _{0}$-BCD-MM and$\ell _{0}$-BCD-MM-F for robust non-negative MC where the nonnegativity constraint is handled by a closed-form update. Experimental results of image inpainting and hyperspectral image recovery demonstrate that the suggested algorithms outperform several state-of-the-art methods in terms of recovery accuracy and computational efficiency. Xiaopeng Li 0005, Zhanglei Shi, Qi Liu 0005, Hing-Cheung So |
IEEE Trans. Cybern. | 3 |
| 2022 | A Hybrid Learning Framework for Deep Spiking Neural Networks with One-Spike Temporal CodingabstractBio-inspired spiking neural networks (SNNs) are compelling candidates for spatio-temporal information processing on ultra-low power neuromorphic computing chips. However, the existing SNN training methods have not fully exploited the temporal information of spikes that plays a critical role in sparse information representation and communication. Hereby, we present a hybrid learning framework for deep SNNs with one-spike temporal coding to make full utilization of the spike timing. We first propose a novel ANN-to-SNN conversion method based on forward propagation mechanisms of ANNs and SNNs to offer a good initialization for SNNs. The performance of the converted SNN is further improved by training with a timing-based backpropagation (BP) method. Experimental results demonstrate that the proposed hybrid learning framework can achieve competitive accuracies on both visual and audio recognition tasks with significantly improved training efficiency over direct SNN BP methods. Jibin Wu, Malu Zhang, Qi Liu 0005, Haizhou Li 0001 |
ICASSP | 4 |
| 2022 | Knowledge distillation for In-memory keyword spotting model
Zeyang Song, Qi Liu 0005, Qu Yang, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2022 | Deep residual spiking neural network for keyword spotting in low-resource settings
Qu Yang, Qi Liu 0005, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2022 | Multistage Deep Transfer Learning for EmIoT-Enabled Human-Computer InteractionabstractEmotional Internet of Things (EmIoT), which provides Internet of Things (IoT) devices cognitive and socialization capabilities, has been regarded as a future direction to improve users’ experiences. With the development of intelligent techniques, the requirement of EmIoT is not only sensing the users’ emotional states but also providing emotional feedbacks. Human–computer interaction has been studied to achieve speech interaction with IoT devices. The recent advances in neural text-to-speech (TTS) have made “human parity” synthesized speech possible for IoT-enabled human–computer interaction. Furthermore, emotion control can be achieved by using the emotional codes in a unified model, referred to as emotional TTS (or ETTS for short). Such ETTS models have achieved promising emotional expressiveness using large-scale emotion-annotated English data set; however, they are not practical in IoT environments with other mainstream languages, especially for Chinese. In fact, the limited available large-scale emotion-annotated data set is challenging the development of Chinese ETTS. To address that we propose a multistage deep transfer learning scheme to design a high-quality Chinese ETTS system under a small-scale training corpus to achieve EmIoT in Mandarin environments. In this scheme, the pretrained knowledge from the former stages corresponding to a large-scale neutral English and a medium-scale emotional English corpora is transferred to a Mandarin ETTS model. Thereby, the trained model can achieve high-quality emotional speech with limited available emotional corpus, which is able to serve various EmIoT-oriented applications. The experiments have been conducted to demonstrate the effectiveness and superiority of the proposed model as compared to other counterparts in terms of naturalness and emotional expressiveness. We refer readers to visit our demo Webpage1enjoy the synthesized speech samples. Rui Liu 0008, Qi Liu 0005, Hongxu Zhu, Hui Cao 0004 |
IEEE Internet Things J. | 2 |
| 2022 | Ultralow Power Always-On Intelligent and Connected SNN-Based System for Multimedia IoT-Enabled ApplicationsabstractThe recent advances in artificial neural networks (ANNs) have created immense opportunities to achieve excellent results on the Internet of Things (IoT), which serve the uses of various real-time smart applications, such as in computer vision and speech recognition. However, the efficiency of ANNs comes at the expense of a huge number of computational resources. This tends to necessitate larger and wider ANNs unapplicable for embedded systems with limited hardware resources, e.g., mobile and wearable devices. To that end, biologically realistic spiking neural networks (SNNs) are first employed to build an always-on intelligent and connected integrated IoT-enabled system with ultralow power consumption, where we encode the temporal dynamic stimuli into effective, efficient, and reconstructable spike patterns to facilitate the subsequent processing. Herein, neural encoding plays a key role in faithfully describing the temporally rich patterns for downstream cognitive tasks. Therefore, a novel nonlinear piecewise latency coding approach for a fully event-driven SNN system is developed. Moreover, a surrogate postsynaptic potential kernel function is utilized to address the nondifferential nature of the spike generation scheme when using the error backpropagation learning method. The effectiveness of the proposal tandem with SNNs has been corroborated by indicative empirical results on different data sets serving cognitive tasks. Qi Liu 0005 |
IEEE Internet Things J. | 1 |
| 2022 | Efficient Low-Rank Matrix Factorization Based on ℓ1, ε-Norm for Online Background SubtractionabstractBackground subtraction refers to extracting the foreground from an observed video, and is the fundamental problem of various applications. There are two kinds of popular methods to deal with background separation, namely, robust principal component analysis (RPCA) and low-rank matrix factorization (LRMF). Nevertheless, the drawback of RPCA requires tuning penalty parameter to attain an ideal result. Compared with RPCA, the$\ell _{1}$-norm based LRMF does not involve extra parameters tuning, but it is challenging to optimize the$\ell _{1}$-norm based minimization because of the nonsmooth$\ell _{1}$-norm. In addition, it becomes time-consuming to find the optimal solution. In this work, we propose to employ smooth$\ell _{1,\epsilon }$-norm, an approximation of$\ell _{1}$-norm, to tackle background subtraction. Thus, the proposed model inherits the superiority of LRMF and even becomes tractable. Then the resultant optimization problem is solved by alternating minimization and gradient descent where the step-size of the gradient descent is adaptively updated via backtracking line searching approach. The proposed method is proved to be locally convergent. Experimental results on synthetic and real-world data demonstrate that our method outperforms the state-of-the-art algorithms in terms of reconstruction loss, computational speed and hardware performance. Qi Liu 0005, Xiaopeng Li 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | From Simulated to Visual Data: A Robust Low-Rank Tensor Completion Approach Using ℓp-Regression for Outlier ResistanceabstractLow-rank tensor completion (LRTC) that aims to restore the latent clean data from an incomplete and/or degraded observation, shows promising results in ubiquitous tensorial data completion applications. Most tensor completion approaches are vulnerable to outliers since their derivations are based on$\ell _{2}$-space to be robust against Gaussian noise. In this work, to tackle this issue,$\ell _{p}$-regression$(0 < p < 2)$is employed to achieve outlier resistance, where a factored form of tensor train (TT)-format representation is regularized by the low-TT-rank prior to exploit the inter-fibers correlation. On the basis of that, an effective iterative$\ell _{p}$-regression TT completion method (referred to$\ell _{p}$-TTC) is proposed, with the advantage of not requiring the hard-to-determine user-defined weights in TT rank model. Extensive experiment results are presented to demonstrate the outlier resistance of the proposed$\ell _{p}$-TTC, and showing the effective and superior performance in both bistatic MIMO radar localization and color image inpainting and denoising, compared with state-of-the-art tensor completion approaches. Qi Liu 0005, Xiaopeng Li 0005, Hui Cao 0004, Yuntao Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Optimum Codesign for Image Denoising Between Type-2 Fuzzy Identifier and Matrix Completion DenoiserabstractWith the wide deployment of digital image capturing equipment, the need of denoising to produce a crystal clear image from noisy capture environment has become indispensable. In this article, a novel type-2 fuzzy-based filter is proposed for denoising images corrupted by impulse noise, especially for the high density of salt-and-pepper noise. It operates two stages, namely, type-2 fuzzy identifier and matrix completion denoiser. In the proposed method, the type-2 fuzzy identifier is first employed to identify and trim the entries contaminated by impulse noise in the data matrix from fuzzy system. Then, the trimmed data matrix is utilized to retrieve the noiseless data matrix with the matrix completion technology. Herein, a novel matrix completion technique is developed without$a$$priori$rank information compared to its counterparts. Simulation results are presented, which vividly show the denoised images obtained by the proposed method can achieve crystal clear image with strong structural integrity, and are showing good performance in terms of peak signal-to-noise ratio. Qi Liu 0005, Xiaopeng Li 0005, Jicheng Yang |
IEEE Trans. Fuzzy Syst. | 1 |
| 2022 | A Neural-Inspired Architecture for EEG-Based Auditory Attention DetectionabstractHumans have the ability to focus on one of the sound sources in a noisy scene, which is critical for everyday communication. Auditory attention detection (AAD) seeks to detect selective attention from one’s brain signals. For AAD to be useful in brain–computer interface applications, new approaches with low computational cost, high classification performance, and low latency are required to be developed. In this study, we proposed a novel neural-inspired architecture to mimic the neural computation and coding strategy in the brain for electroencephalography-based AAD. We validated our model through data visualization, and conducted experiments on two publicly available databases. For both KUL and DTU databases, it outperforms both linear and convolutional neural network (CNN) models with consistent improvements from 1 s to 5 s decision windows in terms of detection accuracy. Although the accuracy of the proposed neural-inspired model is inferior to the state-of-the-art spatio-spectral feature (SSF)-CNN model, the computational cost of our model is less than 1% of SSF-CNN’s. Moreover, the neural-inspired decoder is more hardware friendly and energy-efficient due to its biological computing scheme. Overall, the proposed neural-inspired architecture realizes a fast, accurate, and low energy expenditure AAD, which is a big step forward towards practical neuro-steered hearing aids. Siqi Cai 0002, Peiwen Li, Enze Su, Qi Liu 0005, Longhan Xie |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2021 | HuRAI: A brain-inspired computational model for human-robot auditory interface
Jibin Wu, Qi Liu 0005, Malu Zhang, Zihan Pan, Haizhou Li 0001, Kay Chen Tan |
Neurocomputing | 2 |
| 2021 | Three-Dimensional Speaker Localization: Audio-Refined Visual Scaling Factor EstimationabstractNeither a monocular RGB camera nor a small-size microphone array is capable of accurate three-dimensional (3D) speaker localization. By taking advantage of accurate visual object detection, and audio-visual complementary sensor fusion, we formulate the three-dimensional (3D) speaker localization problem as a visual scaling factor estimation problem. As a result, we effectively reduce the traditional audio-only 3D speaker localization from an exhaustive grid search to a one-dimensional (1D) optimization problem. We propose a multi-modal perception system with two optimization approaches. We show that the proposed methods are effective, accurate, and robust against interference and, as corroborated by indicative empirical results on real dataset, competitive to the conventional uni-modal and the state-of-the-art audio-visual speaker localization approaches. Xinyuan Qian 0001, Qi Liu 0005, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Spike-Event-Driven Deep Spiking Neural Network With Temporal EncodingabstractFeature extractionplays an important role before pattern recognition takes place. The existing artificial neural networks (ANNs), however, ignoreto learn and represent temporal information, instead of only utilizing spatial information for recognition. Moreover, the substantial computational and energy costs resulted from the conventional ANN-based classifiers, limit their uses in mobile and embedded applications. In this work, we develop a sparse temporal encoding method which exploits both spatial and temporal information. On the basis of spike-timing-dependent plasticity and multi-scale structure, the resulting temporal feature representation integrates with a temporal spiking neural network (SNN) classifier to achieve high efficiency of parallel computing for feature extraction. Experimental evaluation on four benchmark datasets from image classification and speech recognition tasks show the proposed SNN model yielding state-of-the-art accuracy. Qi Liu 0005 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Rank-One Matrix Approximation With ℓp-Norm for Image InpaintingabstractIn the problem of image inpainting, one popular approach is based on low-rank matrix completion. Compared with other methods which need to convert the image into vectors or dividing the image into patches, matrix completion operates on the whole image directly. Therefore, it can preserve latent information of the two-dimensional image. An efficient method for low-rank matrix completion is to employ the matrix factorization technique. However, conventional low-rank matrix factorization-based methods often require a prespecified rank, which is challenging to determine in practice. The proposed method factorizes an image matrix as a sum of rank-one matrices so that it does not require rank information in advance as it can be automatically estimated by the algorithm itself when the algorithm has satisfactorily converged. In our study, matching pursuit is applied to search for the best rank-one matrix at each iteration. To be robust against impulsive noise, the residual error between the observed and estimated matrices is minimized by ℓp-norm with 0p-norm minimization is solved by the iteratively reweighted least squares method. The proposed model is beneficial for the robustness against outliers, and does not require rank information. Experimental results verify the effectiveness and higher accuracy of the proposed method with comparison to several state-of-the-art matrix completion-based image inpainting approaches. Xiaopeng Li 0005, Qi Liu 0005, Hing-Cheung So |
IEEE Signal Process. Lett. | 2 |
| 2019 | Smoothed sparse recovery via locally competitive algorithm and forward Euler discretization method
Qi Liu 0005, Yuantao Gu, Hing-Cheung So |
Signal Process. | 1 |
| 2018 | Robust sparse recovery via weakly convex optimization in impulsive noise
Qi Liu 0005, Chengzhu Yang, Yuantao Gu, Hing-Cheung So |
Signal Process. | 1 |
| 2017 | Off-grid DOA estimation with nonconvex regularization via joint sparse representation
Qi Liu 0005, Hing-Cheung So, Yuantao Gu |
Signal Process. | 1 |
| 2015 | Tensor-based real-valued subspace approach for angle estimation in bistatic MIMO radar with unknown mutual coupling
Xianpeng Wang 0001, Wei Wang 0076, Jing Liu 0041, Qi Liu 0005, Ben Wang 0002 |
Signal Process. | 4 |