Lingxiao Yang

dblp:56/7126 · DBLP profile ↗
← Back
74ranked-venue papers
23as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 13 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 5 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Communication-Aware Intelligent Cooperative Search for Multi-UAV Systems in Blocked Regions
Lingtao Xue, Xuewen Dong, Lingxiao Yang
ICIC (2)4
2026 Rethinking security of diffusion-based generative steganography
Jiahao Zhu 0005, Lingxiao Yang, Weiqi Luo 0001, Xiaohua Xie
Inf. Sci.4
2026 Neural Prediction Errors as a Unified Cue for Abstract Visual Reasoning
abstract
Humans exhibit remarkable abilities in recognizing relationships and performing complex reasoning. In contrast, deep neural networks have long been critiqued for their limitations in abstract visual reasoning (AVR), a key challenge in achieving artificial general intelligence. Drawing on the well-known concept of prediction errors from neuroscience, we propose that prediction errors can serve as a unified mechanism for both supervised and self-supervised learning in AVR. In our novel supervised learning model, AVR is framed as a prediction-and-matching process, where the central component is the discrepancy (i.e., prediction error) between a predicted feature based on abstract rules and candidate features within a reasoning context. In the self-supervised model, prediction errors as a key component unify the learning and inference processes. Both supervised and self-supervised prediction-based models achieve state-of-the-art performance on a broad range of AVR datasets and task conditions. Most notably, hierarchical prediction errors in the supervised model automatically decrease during training, an emergent phenomenon closely resembling the decrease of dopamine signals observed in biological learning. These findings underscore the critical role of prediction errors in AVR and highlight the potential of leveraging neuroscience theories to advance computational models for high-level cognition in artificial intelligence.
Lingxiao Yang, Xiaohua Xie, Wei-Shi Zheng 0001, Ru-Yuan Zhang
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 MMA++: Effective Multi-Modal Adaptation for Vision-Language Models
abstract
Large scale pre-trained Vision-Language Models (VLMs) have shown good generalization capabilities across diverse downstream tasks. However, adapting such large-scale models to few-shot generalization scenarios remains challenging due to the trade-off between preserving general knowledge and incorporating task-specific information. In this paper, we propose MMA++, an advanced and effective Multi-Modal Adapter framework for parameter-efficient VLM adaptation. Unlike prior works that independently inject adapters into each modality or uniformly across layers, MMA++ performs a dataset-level analysis to identify discriminative and generalizable features, and selectively applies adapters to the higher layers of both vision and text encoders. To bridge the modality gap, we further propose a shared feature projection space that enhances alignment between modalities. Beyond architecture design, we identify the fusion scale $\alpha$α-which controls the strength of adapter integration-as a key factor in few-shot generalization. We empirically and theoretically demonstrate that $\alpha$α should not be static, but adapted based on training data size. To reduce the effort of tuning this value across different datasets, we propose the $\alpha$α-consistency framework, consisting of: (1) a consistency training strategy under varying fusion scales; and (2) an $\alpha$α-decoupling strategy that uses a larger fusion scale during training and a smaller one at inference to account for sample size mismatch. We evaluate MMA++ on a wide range of few-shot generalization tasks, including base-to-novel generalization, cross-dataset transfer, and domain generalization. Our method consistently achieves leading performance.
Lingxiao Yang, Ru-Yuan Zhang, Yanchen Wang, Xiaohua Xie
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Tripartite Hybrid Game-Theoretic Optimization for Integrated Vehicle-Station-Grid System With Charging Station Heterogeneity
Ning Zhang 0037, Cungang Hu, Qiuye Sun, Lingxiao Yang, David Wenzhong Gao, Yushuai Li
IEEE Trans Autom. Sci. Eng.6
2026 Interpretable Hybrid Deep Reinforcement Learning-Based Energy Management in Low-Carbon Community Energy Systems With Temporal Attention Mechanism
abstract
This article proposes an interpretable deep reinforcement learning (DRL) method for energy management of low-carbon community energy systems (LCCES), which effectively addresses the transparency limitations caused by the black-box nature of traditional DRL neural network structures, thereby overcoming a key constraint in energy system applications. First, we develop a hybrid integer dynamic decision DRL algorithm to solve the low-carbon scheduling problem in community energy systems with continuous-discrete hybrid action spaces. Second, we construct an interpretable artificial intelligence framework, where the temporal attention mechanism is used to process and extract features to provide macro-level decision contribution analysis. These features are input into the decision tree for extracting device-level rules. Building upon this, we design an ensemble decision tree architecture with temporal attention mechanism to effectively identify critical time periods influenced by system inertia and energy fluctuations, thereby achieving interpretable optimization strategies while enhancing decision robustness under state fluctuations. Simulation results based on the independent test set demonstrate that, in comparison with alternative methods, the proposed approach yields a 22.1% cost reduction and a 32.4% carbon emission reduction rate relative to twin delayed deep deterministic policy gradient (TD3), and a 14.7% improvement in explanation accuracy compared with static decision trees.
Lingxiao Yang, Xiaoke Yuan, Ning Zhang 0037, Changyin Sun 0001
IEEE Trans. Comput. Soc. Syst.1
2026 Toward Multi-Source Sky-Ground Re-Identification: A New Benchmark and an Innovative Approach
abstract
Person re-identification (Re-ID) aims to accurately match pairs of person images across different cameras. Existing Re-ID methods primarily focus on associations within single-type camera networks (e.g., ground-ground or sky-sky matching), which are ineffective in addressing the significant viewpoint discrepancies in multi-type camera networks. One key reason for this is the absence of suitable large-scale datasets for algorithm evaluation, which limits the applicability of Re-ID across more diverse scenarios, despite its critical importance. To expand the scope of visual coverage and facilitate search operations in the special locations, we construct a novel benchmark: Multi-Source Sky-Land person Re-ID dataset (MSSL), including 66,928 images from 2,099 volunteers in nearly 20 unique scenes. Additionally, we observe that existing Re-ID systems struggle with drastic viewpoint variations in sky-ground Re-ID, especially on MSSL. To address these issues, we propose Multi-Source Prompts (MSP), separately learning finer cross-modal features of pedestrians from both sky and ground perspectives. These refined features better represent the true appearance of pedestrians from different viewpoints. Subsequently, we employ Multi-Source Alignment Loss to mitigate the impact of drastic viewpoint changes. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on our MSSL, as well as on other benchmarks such as AG-ReID dataset. Our MSSL dataset and the code will be available athttps://github.com/sysuchx/SkyGroundReID.
Zhizhi Lu, Nailong Zhao, Yuli Huang, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Multim.5
2026 MHCertChain: A Multi-CA Hierarchical Certificate Blockchain With Low Overhead
abstract
Blockchain-based certificate management schemes provide a distributed approach for Public Key Infrastructure through the integration of blockchain services, thereby enhancing transparency and security in identity authentication. However, existing schemes often suffer from limited scalability when accommodating new CAs, high overhead of blockchain systems, and a high false-positive rate in revocation status checks. To address above issues, we propose MHCertChain, a multi-CA hierarchical certificate blockchain with low overhead. Specifically, a two-layer blockchain structure including the main chain and sub-chains is designed to enhance scalability, while the main chain stores trust paths between CAs and each automatically deployed sub-chain records user certificates issued by an end-entity CA. Then, we propose a lightweight dual-signature certificate format that contains certificate location information without requiring any external certificate location method outside the blockchain. Considering time characteristics of certificates, a time-partitioned cuckoo filter is proposed with a low false positive rate and accelerates revocation status query. Moreover, we deduce an optimal global performance parameter of such filter through mathematical modeling. We present a thorough security analysis of our MHCertChain utilizing the universally composable framework, and extensive experiments demonstrate that the CPU overhead and false positive rate are reduced by 70% and 95%, respectively, compared to state-of-the-art schemes. Blockchain-based certificate management schemes provide a distributed approach for Public Key Infrastructure through the integration of blockchain services, thereby enhancing transparency and security in identity authentication. However, existing schemes seldom discuss hierarchical CA architecture, and often suffer from limited scalability and high overhead when considering new CAs, frequent certificate authentication and revocation status verification. To address above issues, we propose MHCertChain, a multi-CA hierarchical certificate blockchain with low overhead. Specifically, a two-layer blockchain structure including the main chain and sub-chains is designed, while the main chain stores trust paths between CAs and each automatically deployed sub-chain records user certificates issued by an end-entity CA. Then, we propose a lightweight dual-signature certificate format that contains certificate location information without requiring any external certificate location method outside the blockchain. Considering time characteristics of certificates, a time-partitioned cuckoo filter is proposed with a low false positive rate and accelerates revocation status query. Moreover, we deduce an optimal global performance parameter of such filter through mathematical modeling. We present a thorough security analysis of our MHCertChain utilizing the universally composable framework, and extensive experiments demonstrate that the CPU overhead and false positive rate are reduced by 70% and 95%, respectively, compared to state-of-the-art schemes. Blockchain-based certificate management schemes provide a distributed means to increase the transparency of PKI (Public Key Infrastructure) and prevent single point attacks in web communications. However, certificate management based on hierarchical multi-CA architecture often faces the problems of poor scalability and high CPU overhead in blockchain. In this article, we are the first to propose a multi-CA hierarchical certificate blockchain with low overhead. Specifically, a two-layer blockchain structure of the main chain and sub-chain is adopted, with the main chain storing the trust paths and the sub-chain storing the user certificates. By monitoring to the CA transactions of the main chain, the sub-chain is automatically deployed. Then, in the certificate operation, we consider the high CPU overhead of traditional certificate query in blockchain and propose a dual-signature certificate format. combined with the time characteristics of the certificate, a time-partitioned cuckoo filter is proposed with a low false positive rate for the revocation status query speeding, and we find a global performance optimal parameter through mathematical modeling. Finally, we use a general composable framework to prove the security of HiCertChain, and the experiments show that the CPU overhead and false positive rate are reduced by 90% and 95%, respectively, compared with those of state-of-the-arts.
Xuewen Dong, Qingsong Yao, Lingxiao Yang, Zhiwei Zhang 0004, Ning Xi 0002, Yulong Shen 0001
IEEE Trans. Serv. Comput.4
2026 AC-BaaS: An Asynchronous Cross-Blockchain as a Service for the Internet of Things
abstract
Cross-chain techniques improve blockchain scalability and interoperability, providing decentralized exchange and cross-chain collaboration services for Internet of Things (IoT) data across various domains. However, current state-of-the-art (SOTA) solutions for cross-chain data exchange across multiple domains are constrained by synchronous networks, hindering efficient data exchange in intermittent network environments. Furthermore, there is a lack of research on asynchronous cross-chain transaction pool mechanisms, which are crucial for optimizing system utility. In this paper, we propose AC-BaaS, anasynchronouscross-blockchainasaservice framework tailored for the multi-domain IoT. Built upon a specially designed asynchronous sidechain architecture, the system leverages a committee to provide AC-BaaS for data exchange across multiple IoT domains. To fulfill the need for asynchronous and efficient data exchange, we combine the ideas of aggregate signatures and verifiable delay functions to devise a novel cryptographic primitive called delayed aggregate signature (DAS), which constructs asynchronous cross-chain proofs (ACPs) that ensure the security of cross-chain interactions. To ensure the consistency of asynchronous transactions, we propose a multilevel buffered transaction pool that guarantees the transaction sequencing. We further propose a heuristic for optimizing the utility of the buffer pool mechanism to strike a balance between performance and resource consumption. We also examine DAS delay size settings to trade-off security and efficiency. We analyze and prove the security of AC-BaaS, simulate asynchronous communication environments under various security levels, and conduct a comprehensive evaluation. The results show that AC-BaaS outperforms SOTA schemes, improving throughput by an average of 1.71 to 5.09 times, reducing transaction latency by 64.36% to 85.49%, and maintaining comparable resource overhead.
Lingxiao Yang, Xuewen Dong, Zhiguo Wan, Sheng Gao 0002, Wei Tong 0003, Yong Yu 0002, Yulong Shen 0001
IEEE Trans. Serv. Comput.1
2025 Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
abstract
Fine-tuning pre-trained vision-language models has emerged as a powerful approach for enhancing open-vocabulary semantic segmentation (OVSS). However, the substantial computational and resource demands associated with training on large datasets have prompted interest in training-free methods for OVSS. Existing training-free approaches primarily focus on modifying model architectures and generating prototypes to improve segmentation performance. However, they often neglect the challenges posed by class redundancy, where multiple categories are not present in the current test image, and visual-language ambiguity, where semantic similarities among categories create confusion in class activation. These issues can lead to suboptimal class activation maps and affinity-refined activation maps. Motivated by these observations, we propose FreeCP, a novel training-free class purification framework designed to address these challenges. FreeCP focuses on purifying semantic categories and rectifying errors caused by redundancy and ambiguity. The purified class representations are then leveraged to produce final segmentation predictions. We conduct extensive experiments across eight benchmarks to validate FreeCP's effectiveness. Results demonstrate that FreeCP, as a plug-and-play module, significantly boosts segmentation performance when combined with other OVSS methods.
Qi Chen 0013, Lingxiao Yang, Nailong Zhao, Jian-Huang Lai, Xiaohua Xie
ICCV2
2025 AsyncSC: An Asynchronous Sidechain for Multi-Domain Data Exchange in Internet of Things
Lingxiao Yang, Xuewen Dong, Zhiguo Wan, Sheng Gao 0002, Wei Tong 0003, Di Lu 0001, Yulong Shen 0001, Xiaojiang Du
INFOCOM1
2025 OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
abstract
Understanding and synthesizing realistic 3D hand-object interactions (HOI) is critical for applications ranging from immersive AR/VR to dexterous robotics. Existing methods struggle with generalization, performing well on closed-set objects and predefined tasks but failing to handle unseen objects or open-vocabulary instructions. We introduce OpenHOI, the first framework for open-world HOI synthesis, capable of generating long-horizon manipulation sequences for novel objects guided by free-form language commands. Our approach integrates a 3D Multimodal Large Language Model (MLLM) fine-tuned for joint affordance grounding and semantic task decomposition, enabling precise localization of interaction regions (e.g., handles, buttons) and breakdown of complex instructions (e.g., “Find a water bottle and take a sip”) into executable sub-tasks. To synthesize physically plausible interactions, we propose an affordance-driven diffusion model paired with a training-free physics refinement stage that minimizes penetration and optimizes affordance alignment. Evaluations across diverse scenarios demonstrate OpenHOI’s superiority over state-of-the-art methods in generalizing to novel object categories, multi-stage tasks, and complex language instructions.
Ye Shi 0001, Lingxiao Yang, Suting Ni, Jingya Wang 0001
NeurIPS3
2025 Hard-Normal Example-Aware Template Mutual Matching for Industrial Anomaly Detection
Xiaohua Xie, Lingxiao Yang, Jian-Huang Lai
Int. J. Comput. Vis.3
2025 Learning with Enriched Inductive Biases for Vision-Language Models
Lingxiao Yang, Ru-Yuan Zhang, Xiaohua Xie
Int. J. Comput. Vis.1
2025 HiCoCS: High Concurrency Cross-Sharding on Permissioned Blockchains
abstract
As the foundation of the Web3 trust system, blockchain technology faces increasing demands for scalability. Sharding emerges as a promising solution, but it struggles to handle highly concurrent cross-shard transactions (CSTxs), primarily due to simultaneous ledger operations on the same account. Hyperledger Fabric, a permissioned blockchain, employs multi-version concurrency control for parallel processing. Existing solutions use channels and intermediaries to achieve cross-sharding in Hyperledger Fabric. However, the conflict problem caused by highly concurrent CSTxs has not been adequately resolved. To fill this gap, we propose HiCoCS, a high concurrency cross-shard scheme for permissioned blockchains. HiCoCS creates a unique virtual sub-broker for each CSTx by introducing a composite key structure, enabling conflict-free concurrent transaction processing while reducing resource overhead. The challenge lies in managing large numbers of composite keys and mitigating intermediary privacy risks. HiCoCS utilizes virtual sub-brokers to receive and process CSTxs concurrently while maintaining a transaction pool. Batch processing is employed to merge multiple CSTxs in the pool, improving efficiency. We explore composite key reuse to reduce the number of virtual sub-brokers and lower system overhead. Privacy preservation is enhanced using homomorphic encryption. Evaluations show that HiCoCS improves cross-shard transaction throughput by 3.5-20.2 times compared to the baselines.
Lingxiao Yang, Xuewen Dong, Zhiguo Wan, Di Lu 0001, Yushu Zhang 0001, Yulong Shen 0001
IEEE Trans. Computers1
2025 Angular Reconstructive Discrete Embedding With Fusion Similarity for Multi-View Clustering
abstract
Effectively and efficiently mining valuable clustering patterns is a challenging problem when handling large-scale data from diverse sources. Existing approaches adopt anchor graph learning or binary representation embedding to reduce computational complexity. Normally, anchor graph learning can not directly obtain the clustering assignment except adopt the post-processing stage, such as graph cut or k-means clustering. The binary representation embedding neglects the structure information in Hamming space. In order to overcome these limitations, this paper proposes a novel, effective, and efficient angular reconstructive discrete embedding method with fusion similarity for a multi-view clustering (AFMC) that can jointly learn the global and local structure preserving binary representation and clustering assignment. Specifically, we propose to use angular reconstructive error minimization to maintain the global similarity correlation of binary representations of heterogeneous features in a common Hamming space. Moreover, we design a multi-view discrete ridge regression with fusion similarity term to handle the out-of-sample problem and preserve the local manifold structure. In addition, we propose an efficient optimization algorithm with linear computational complexity to solve the non-convex and non-smooth objective function. The experimental results demonstrate that AFMC outperforms several state-of-the-art large-scale multi-view clustering methods.
Jintang Bian, Xiaohua Xie, Chang-Dong Wang 0001, Lingxiao Yang, Jian-Huang Lai, Feiping Nie 0001
IEEE Trans. Knowl. Data Eng.4
2025 Multilevel Contrastive Multiview Clustering With Dual Self-Supervised Learning
abstract
Multiview clustering (MVC) aims to integrate multiple related but different views of data to achieve more accurate clustering performance. Contrastive learning has found many applications in MVC due to its successful performance in unsupervised visual representation learning. However, existing MVC methods based on contrastive learning overlook the potential of high similarity nearest neighbors as positive pairs. In addition, these methods do not capture the multilevel (i.e., cluster, instance, and prototype levels) representational structure that naturally exists in multiview datasets. These limitations could further hinder the structural compactness of learned multiview representations. To address these issues, we propose a novel end-to-end deep MVC method called multilevel contrastive MVC (MCMC) with dual self-supervised learning (DSL). Specifically, we first treat the nearest neighbors of an object from the latent subspace as the positive pairs for multiview contrastive loss, which improves the compactness of the representation at the instance level. Second, we perform multilevel contrastive learning (MCL) on clusters, instances, and prototypes to capture the multilevel representational structure underlying the multiview data in the latent space. In addition, we learn consistent cluster assignments for MVC by adopting a DSL method to associate different level structural representations. The evaluation experiment showed that MCMC can achieve intracluster compactness, intercluster separability, and higher accuracy (ACC) in clustering performance. Our code is available at https://github.com/bianjt-morning/MCMC.
Jintang Bian, Yixiang Lin, Xiaohua Xie, Chang-Dong Wang 0001, Lingxiao Yang, Jian-Huang Lai, Feiping Nie 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 InfantCryNet: A Data-driven Framework for Intelligent Analysis of Infant Cries
Mengze Hong, Chen Zhang 0013, Lingxiao Yang, Yuanfeng Song, Di Jiang 0004
ACML3
2024 MMA: Multi-Modal Adapter for Vision-Language Models
abstract
Pretrained Vision-Language Models (VLMs) have served as excellent foundation models for transfer learning in diverse downstream tasks. However, tuning VLMs for few-shot generalization tasks faces a discrimination - generalization dilemma, i.e., general knowledge should be preserved and task-specific knowledge should be fine-tuned. How to precisely identify these two types of representations remains a challenge. In this paper, we propose a Multi-Modal Adapter (MMA) for VLMs to improve the alignment between representations from text and vision branches. MMA aggregates features from different branches into a shared feature space so that gradients can be communicated across branches. To determine how to incorporate MMA, we systematically analyze the discriminability and generalizability of features across diverse datasets in both the vision and language branches, and find that (1) higher lay-ers contain discriminable dataset-specific knowledge, while lower layers contain more generalizable knowledge, and (2) language features are more discriminable than visual features, and there are large semantic gaps between the features of the two modalities, especially in the lower layers. Therefore, we only incorporate MMA to a few higher lay-ers of transformers to achieve an optimal balance between discrimination and generalization. We evaluate the effectiveness of our approach on three tasks: generalization to novel classes, novel target datasets, and domain generalization. Compared to many state-of-the-art methods, our MMA achieves leading performance in all evaluations. Code is at https://github.com/ZjjConan/Multi-Modal-Adapter
Lingxiao Yang, Ru-Yuan Zhang, Yanchen Wang, Xiaohua Xie
CVPR1
2024 Spike-Temporal Latent Representation for Energy-Efficient Event-to-Video Reconstruction
Jianxiong Tang, Jian-Huang Lai, Lingxiao Yang, Xiaohua Xie
ECCV (42)3
2024 Guidance with Spherical Gaussian Constraint for Conditional Diffusion
abstract
Recent advances in diffusion models attempt to handle conditional generative tasks by utilizing a differentiable loss function for guidance without the need for additional training. While these methods achieved certain success, they often compromise on sample quality and require small guidance step sizes, leading to longer sampling processes. This paper reveals that the fundamental issue lies in the manifold deviation during the sampling process when loss guidance is employed. We theoretically show the existence of manifold deviation by establishing a certain lower bound for the estimation error of the loss guidance. To mitigate this problem, we propose Diffusion with Spherical Gaussian constraint (DSG), drawing inspiration from the concentration phenomenon in high-dimensional Gaussian distributions. DSG effectively constrains the guidance step within the intermediate data manifold through optimization and enables the use of larger guidance steps. Furthermore, we present a closed-form solution for DSG denoising with the Spherical Gaussian constraint. Notably, DSG can seamlessly integrate as a plugin module within existing training-free conditional diffusion methods. Implementing DSG merely involves a few lines of additional code with almost no extra computational overhead, yet it leads to significant performance improvements. Comprehensive experimental results in various conditional generation tasks validate the superiority and adaptability of DSG in terms of both sample quality and time efficiency.
Lingxiao Yang, Shutong Ding, Jingyi Yu 0001, Jingya Wang 0001, Ye Shi 0001
ICML1
2024 Usability of Cross-Device Interaction Interfaces for Augmented Reality in Physical Tasks
abstract
The shortcomings of established input methods for Augmented Reality (AR) head-mounted displays (HMDs) motivate us to investigate the use of AR HMDs with smartphones and smartwatches to improve AR interaction in physical tasks. However, it is unclear whether the cross-device interaction interfaces are efficient for AR systems on physical tasks because the physical tasks can break the interaction flow. In this work, we conducted a user study to explore this. The user study consists of three subtasks, respectively requiring one of the three representative user interface (UI) controls (buttons, sliders, or text input). We implemented them with mid-air gestures, a smartphone, and a smartwatch. We compared the three interfaces and found that the smartphone and the smartwatch interaction interfaces have greater usability, provide a better user experience, and have lower workloads than the mid-air gesture interface. However, there is no significant difference in terms of efficiency for simple interactions. We discuss the limitations of this research and directions for future work.
Weiping He, Mark Billinghurst, Daisong Liu, Lingxiao Yang, Yizhe Liu
Int. J. Hum. Comput. Interact.5
2024 Design and Evaluation of Bare-Hand Interaction for Precise Manipulation of Distant Objects in AR
abstract
Interaction with virtual objects is one of the essential features of Augmented Reality (AR) systems. One of its main issues is how to provide precise manipulation of distant virtual objects in AR. In this work, we explore bare-hand manipulation of distant objects in AR with DOF (degree-of-freedom) separation, motion scaling, and near-field metaphors. We developed two manipulation techniques: the distant widget-based metaphor (DWBM), and the near-field widget-based metaphor (NFWBM). We conducted a user study with 20 participants to compare the two techniques in terms of performance and user experience. We found that NFWBM has faster speed, lower mental effort, better ease-of-use, and a friendlier user experience. However, there is no significant difference in terms of precision; both techniques could manipulate distant objects precisely, with an average position error of less than 0.7 cm and an average orientational error of less than 1°. We also discussed the limitations of this research and directions for future work.
Weiping He, Mark Billinghurst, Lingxiao Yang, Daisong Liu
Int. J. Hum. Comput. Interact.4
2024 A Neuroinspired Contrast Mechanism enables Few-Shot Object Detection
Lingxiao Yang, Dapeng Chen, Yifei Chen 0010, Wei Peng 0011, Xiaohua Xie
Pattern Recognit.1
2024 Benchmarking deep models on salient object detection
Huajun Zhou, Lingxiao Yang, Jian-Huang Lai, Xiaohua Xie
Pattern Recognit.3
2024 Price-Matching-Based Regional Energy Market With Hierarchical Reinforcement Learning Algorithm
abstract
This article proposes a multienergy trading market model based on price matching, aiming to foster multienergy collaboration and enhance energy utilization through individual participation. With the ongoing advancements in energy distribution and marketization, the energy Internet necessitates improved applicability and efficiency for personalized energy responses. To address these requirements, a multienergy trading market model is proposed, which enables the avoidance of user information disclosure and guarantees user trading autonomy. In addition, a joint trading mechanism is designed that accounts for multiple time scales and energy types, consequently reducing trading failures caused by overlooking energy transmission processes. By performing the proposed trading mechanism, the market operator can match various energy types using conversion devices, thereby augmenting matching efficiency. An income mechanism is also established to deter the operator from purposefully evading potential trading opportunities for personal gain. To address the proposed model, an improved hierarchical reinforcement learning algorithm is employed, which effectively overcomes challenges associated with large state action spaces and sparse rewards. Numerical examples are provided to confirm the efficacy of the proposed approach.
Ning Zhang 0037, Cungang Hu, Qiuye Sun, Lingxiao Yang, David Wenzhong Gao, Josep M. Guerrero, Yushuai Li
IEEE Trans. Ind. Informatics5
2024 Pose Guided Person Image Generation Via Dual-Task Correlation and Affinity Learning
abstract
Pose Guided Person Image Generation (PGPIG) is the task of transforming a person's image from the source pose to a target pose. Existing PGPIG methods often tend to learn an end-to-end transformation between the source image and the target image, but do not seriously consider two issues: 1) the PGPIG is an ill-posed problem, and 2) the texture mapping requires effective supervision. In order to alleviate these two challenges, we propose a novel method by incorporating Dual-task Pose Transformer Network and Texture Affinity learning mechanism (DPTN-TA). To assist the ill-posed source-to-target task learning, DPTN-TA introduces an auxiliary task, i.e., source-to-source task, by a Siamese structure and further explores the dual-task correlation. Specifically, the correlation is built by the proposed Pose Transformer Module (PTM), which can adaptively capture the fine-grained mapping between sources and targets and can promote the source texture transmission to enhance the details of the generated images. Moreover, we propose a novel texture affinity loss to better supervise the learning of texture mapping. In this way, the network is able to learn complex spatial transformations effectively. Extensive experiments show that our DPTN-TA can produce perceptually realistic person images under significant pose changes. Furthermore, our DPTN-TA is not limited to processing human bodies but can be flexibly extended to view synthesis of other objects, i.e., faces and chairs, outperforming the state-of-the-arts in terms of both LPIPS and FID. Our code is available at: https://github.com/PangzeCheung/Dual-task-Pose-Transformer-Network.
Pengze Zhang, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Vis. Comput. Graph.2
2023 Texture-Guided Saliency Distilling for Unsupervised Salient Object Detection
abstract
Deep Learning-based Unsupervised Salient Object Detection (USOD) mainly relies on the noisy saliency pseudo labels that have been generated from traditional handcraft methods or pre-trained networks. To cope with the noisy labels problem, a class of methods focus on only easy samples with reliable labels but ignore valuable knowledge in hard samples. In this paper, we propose a novel USOD method to mine rich and accurate saliency knowledge from both easy and hard samples. First, we propose a Confidence-aware Saliency Distilling (CSD) strategy that scores samples conditioned on samples' confidences, which guides the model to distill saliency knowledge from easy samples to hard samples progressively. Second, we propose a Boundary-aware Texture Matching (BTM) strategy to refine the boundaries of noisy labels by matching the textures around the predicted boundaries. Extensive experiments on RGB, RGB-D, RGB-T, and video SOD benchmarks prove that our method achieves state-of-the-art USOD performance. Code is available at www.github.com/moothes/A2S-v2.
Huajun Zhou, Bo Qiao 0003, Lingxiao Yang, Jian-Huang Lai, Xiaohua Xie
CVPR3
2023 CuNeRF: Cube-Based Neural Radiance Field for Zero-Shot Medical Image Arbitrary-Scale Super Resolution
abstract
Medical image arbitrary-scale super-resolution (MIASSR) has recently gained widespread attention, aiming to supersample medical volumes at arbitrary scales via a single model. However, existing MIASSR methods face two major limitations: (i) reliance on high-resolution (HR) volumes and (ii) limited generalization ability, which restricts their applications in various scenarios. To overcome these limitations, we propose Cube-based Neural Radiance Field (CuNeRF), a zero-shot MIASSR framework that is able to yield medical images at arbitrary scales and free viewpoints in a continuous domain. Unlike existing MISR methods that only fit the mapping between low-resolution (LR) and HR volumes, CuNeRF focuses on building a continuous volumetric representation from each LR volume without the knowledge of the corresponding HR one. This is achieved by the proposed differentiable modules: cube-based sampling, isotropic volume rendering, and cube-based hierarchical rendering. Through extensive experiments on magnetic resource imaging (MRI) and computed tomography (CT) modalities, we demonstrate that CuNeRF can synthesize high-quality SR medical images, which outperforms state-of-the-art MISR methods, achieving better visual verisimilitude and fewer objectionable artifacts. Compared to existing MISR methods, our CuNeRF is more applicable in practice.
Lingxiao Yang, Jian-Huang Lai, Xiaohua Xie
ICCV2
2023 Optimal Hub Placement and Deadlock-Free Routing for Payment Channel Network Scalability
abstract
As a promising implementation model of payment channel network (PCN), payment channel hub (PCH) could achieve high throughput by providing stable off-chain transactions through powerful hubs. However, existing PCH schemes assume hubs are preplaced in advance, not considering payment requests' distribution and may affect network scalability, especially network load balancing. In addition, current source routing protocols with PCH allow each sender to make routing decision on his/her own request, which may have a bad effect on performance scalability (e.g., deadlock) for not considering other senders' requests. This paper proposes a novel multi-PCHs solution with high scalability. First, we are the first to study the PCH placement problem and propose optimal/approximation solutions with load balancing for small-scale and large-scale scenarios, by trading off communication costs among participants and turning the original NP-hard problem into a mixed-integer linear programming (MILP) problem solving by supermodular techniques. Then, on global network states and local directly connected clients' requests, a routing protocol is designed for each PCH with a dynamic adjustment strategy on request processing rates, enabling high-performance deadlock-free routing. Extensive experiments show that our work can effectively balance the network load, and improve the performance on throughput by 29.3% on average compared with state-of-the-arts.
Lingxiao Yang, Xuewen Dong, Sheng Gao 0002, Qiang Qu 0001, Xiaodong Zhang 0036, Wensheng Tian, Yulong Shen 0001
ICDCS1
2023 Self-supervised Cross-stage Regional Contrastive Learning for Object Detection
abstract
Cross-stage object similarity is a vital property of generic supervised object detectors, which maintains similar feature responses to the same object across feature maps of different intermediate stages of the backbone network. Since an object can be predicted by multiple stages, this similarity is beneficial for accurate object classification and localization. Inspired by this property, we introduce Cross-stage regional Contrastive Learning (CrossCL) to learn the cross-stage object similarity during the model pre-training. Since labels are unavailable in self-supervised learning, we treat the regions sharing the same position in different stages as the same object and constrain them to have similar feature responses across stages to achieve cross-stage object similarity. The learned feature representations of CrossCL share a similar property with supervised detectors, thus showing strong transfer capability to object detection tasks. Besides, we also provide in-depth discussions, ablation studies, and visualizations to understand better how CrossCL works. Code is available at https://github.com/yanjk3/CrossCL.
Junkai Yan, Lingxiao Yang, Yipeng Gao, Wei-Shi Zheng 0001
ICME2
2023 Neural Prediction Errors enable Analogical Visual Reasoning in Human Standard Intelligence Tests
abstract
Deep neural networks have long been criticized for lacking the ability to perform analogical visual reasoning. Here, we propose a neural network model to solve Raven's Progressive Matrices (RPM) - one of the standard intelligence tests in human psychology. Specifically, we design a reasoning block based on the well-known concept of prediction error (PE) in neuroscience. Our reasoning block uses convolution to extract abstract rules from high-level visual features of the 8 context images and generates the features of a predicted answer. PEs are then calculated between the predicted features and those of the 8 candidate answers, and are then passed to the next stage. We further integrate our novel reasoning blocks into a residual network and build a new Predictive Reasoning Network (PredRNet). Extensive experiments show that our proposed PredRNet achieves state-of-the-art average performance on several important RPM benchmarks. PredRNet also shows good generalization abilities in a variety of out-of-distribution scenarios and other visual reasoning tasks. Most importantly, our PredRNet forms low-dimensional representations of abstract rules and minimizes hierarchical prediction errors during model training, supporting the critical role of PE minimization in visual reasoning. Our work highlights the potential of using neuroscience theories to solve abstract visual reasoning problems in artificial intelligence. The code is available at https://github.com/ZjjConan/AVR-PredRNet.
Lingxiao Yang, Hongzhi You, Zonglei Zhen, Dahui Wang, Xiaohong Wan, Xiaohua Xie, Ru-Yuan Zhang
ICML1
2023 Spike Count Maximization for Neuromorphic Vision Recognition
abstract
Spiking Neural Networks (SNNs) are the promising models of neuromorphic vision recognition. The mean square error (MSE) and cross-entropy (CE) losses are widely applied to supervise the training of SNNs on neuromorphic datasets. However, the relevance between the output spike counts and predictions is not well modeled by the existing loss functions. This paper proposes a Spike Count Maximization (SCM) training approach for the SNN-based neuromorphic vision recognition model based on optimizing the output spike counts. The SCM is achieved by structural risk minimization (SRM) and a specially designed spike counting loss. The spike counting loss counts the output spikes of the SNN by using the L0-norm, and the SRM maximizes the distance between the margin boundaries of the classifier to ensure the generalization of the model. The SCM is non-smooth and non-differentiable, and we design a two-stage algorithm with fast convergence to solve the problem. Experiment results demonstrate that the SCM performs satisfactorily in most cases. Using the output spikes for prediction, the accuracies of SCM are 2.12%~16.50% higher than the popular training losses on the CIFAR10-DVS dataset. The code is available at https://github.com/TJXTT/SCM-SNN.
Jianxiong Tang, Jian-Huang Lai, Xiaohua Xie, Lingxiao Yang
IJCAI4
2023 RuleMatch: Matching Abstract Rules for Semi-supervised Learning of Human Standard Intelligence Tests
abstract
Raven's Progressive Matrices (RPM), one of the standard intelligence tests in human psychology, has recently emerged as a powerful tool for studying abstract visual reasoning (AVR) abilities in machines. Although existing computational models for RPM problems achieve good performance, they require a large number of labeled training examples for supervised learning. In contrast, humans can efficiently solve unlabeled RPM problems after learning from only a few example questions. Here, we develop a semi-supervised learning (SSL) method, called RuleMatch, to train deep models with a small number of labeled RPM questions along with other unlabeled questions. Moreover, instead of using pixel-level augmentation in object perception tasks, we exploit the nature of RPM problems and augment the data at the level of abstract rules. Specifically, we disrupt the possible rules contained among context images in an RPM question and force the two augmented variants of the same unlabeled sample to obey the same abstract rule and predict a common pseudo label for training. Extensive experiments show that the proposed RuleMatch achieves state-of-the-art performance on two popular RAVEN datasets. Our work makes an important stride in aligning abstract analogical visual reasoning abilities in machines and humans. Our Code is at https://github.com/ZjjConan/AVR-RuleMatch.
Yunlong Xu 0001, Lingxiao Yang, Hongzhi You, Zonglei Zhen, Da-Hui Wang, Xiaohong Wan, Xiaohua Xie, Ru-Yuan Zhang
IJCAI2
2023 Few Shot Object Detection with Incompletely Annotated Samples
abstract
Few shot object detection aims to generalize the model to previously unseen classes with only a few training samples, which has been attached great attention due to its practicability in real scenes. Many existing methods hold the missed detection issue, mainly due to the problem of incompletely annotated samples. Specifically, only partial objects in a sample are labeled, resulting in unlabeled objects being used as the background for training. This problem is especially serious for few-shot learning. To solve this noisy label problem, we first propose a label calibration method based on confidence to correct potential incorrect labels, and then introduce the class center library to eliminate the negative impact of unlabeled objects. We conduct extensive experiments on PASCAL VOC and MS-COCO benchmarks, which proves the effectiveness of our approach and achieves the state-of-the-art results.
Bo Qiao 0003, Huajun Zhou, Lingxiao Yang, Xiaohua Xie
IJCNN3
2023 Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
abstract
Emotional Voice Conversion aims to manipulate a speech according to a given emotion while preserving non-emotion components. Existing approaches cannot well express fine-grained emotional attributes. In this paper, we propose an Attention-based Interactive diseNtangling Network (AINN) that leverages instance-wise emotional knowledge for voice conversion. We introduce a two-stage pipeline to effectively train our network: Stage I utilizes inter-speech contrastive learning to model fine-grained emotion and intra-speech disentanglement learning to better separate emotion and content. In Stage II, we propose to regularize the conversion with a multi-view consistency mechanism. This technique helps us transfer fine-grained emotion and maintain speech content. Extensive experiments show that our AINN outperforms state-of-the-arts in both objective and subjective metrics.
Lingxiao Yang, Qi Chen 0013, Jian-Huang Lai, Xiaohua Xie
INTERSPEECH2
2023 ARCoA: Using the AR-Assisted Cooperative Assembly System to Visualize Key Information about the Occluded Partner
abstract
During component assembly, some operations must be completed by two or more workers due to the size or assembly mode. For instance, in manual riveting, two workers are positioned on either side of a steel plate, which blocks the view. Traditional collaborative approaches limit assembly efficiency and is difficult to ensure accurate and rapid interaction between workers. In this study, we developed an AR-Assisted Cooperative Assembly System (ARCoA) to address the issue. ARCoA allows users to view their partner’s key information that is occluded, including tools, gestures, orientations, and shared markers. Besides, we presented a user experiment that compared this method with the traditional approach. The results indicated that the new system could significantly improve assembly efficiency, system availability, and sense of social presence. Moreover, most users we surveyed preferred ARCoA. In the future, we will incorporate more functions and improve the accuracy of the system to solve complex multi-person collaborative problems.
Weiping He, Qianrui Zhang, Mark Billinghurst, Lingxiao Yang
Int. J. Hum. Comput. Interact.5
2023 TI-BIoV: Traffic Information Interaction for Blockchain-Based IoV With Trust and Incentive
abstract
Recent Blockchain-based Internet of Vehicles (BIoV) solutions are proposed to provide the capabilities of trust management and incentive distribution for traffic information interaction in decentralized trustless Internet of Vehicles (IoV). However, existing trust management methods in BIoV are designed based on subjective user feedback, which is vulnerable to bad-mouthing and collusion attacks. Besides, these incentive strategies achieve accurate information interaction based on the game theory, yet it is challenging for the practical IoV scenario without completely explicit parameters. To address these issues, we propose TI-BIoV, a traffic information interaction system based on three blockchains for IoV with the nonsubjective trust evaluation and optimal incentive with partial inexplicit parameters. Specifically, a nonsubjective trust mechanism is designed based on the traffic information offset calculated by other related traffic information, which ensures the change of vehicle trust value without any subjective factors. On this basis, a trust-based consensus protocol, which selects entities with high trust values as participants, is given to realize the reliable public audit of transactions. According to traffic information accuracy measurements, we develop a$Q$-learning-based algorithm to encourage vehicles continuously submit accurate traffic information and optimally schedule the incentive for both platform and vehicle via training with incompletely explicit parameters of TI-BIoV. Finally, we analyze the security properties and common attacks of TI-BIoV and implement a prototype. The experimental results show that TI-BIoV achieves reliable consensus with nonsubjective trust evaluation and runs stably for a long time with two-sided incentive strategies.
Wei Tong 0003, Xuewen Dong, Yushu Zhang 0001, Zongyang Zhang, Lingxiao Yang, Weidong Yang 0003, Yulong Shen 0001
IEEE Internet Things J.5
2023 AC2AS: Activation Consistency Coupled ANN-SNN framework for fast and memory-efficient SNN training
Jianxiong Tang, Jian-Huang Lai, Xiaohua Xie, Lingxiao Yang, Wei-Shi Zheng 0001
Pattern Recognit.4
2023 Activation to Saliency: Forming High-Quality Labels for Unsupervised Salient Object Detection
abstract
This paper focuses on the Unsupervised Salient Object Detection (USOD) issue. We come up with a two-stage Activation-to-Saliency (A2S) framework that effectively excavates saliency cues to train a robust saliency detector. It is worth noting that our method does not require any manual annotation in the whole process. In the first stage, we transform an unsupervisedly pre-trained network to aggregate multi-level features into a single activation map, where an Adaptive Decision Boundary (ADB) is proposed to assist the training of the transformed network. Moreover, a new loss function is proposed to facilitate the generation of high-quality pseudo labels. In the second stage, a self-rectification learning strategy is developed to train a saliency detector and refine the pseudo labels online. In addition, we construct a lightweight saliency detector using two Residual Attention Modules (RAMs) to learn robust saliency information. Extensive experiments on several SOD benchmarks prove that our framework reports significant performance compared with existing USOD methods. Moreover, training our framework on 3,000 images consumes about 1 hour, which is over 10 times faster than previous state-of-the-art methods. Code has been published athttps://github.com/moothes/A2S-USOD.
Huajun Zhou, Peijia Chen, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Circuits Syst. Video Technol.3
2023 Hybrid Policy-Based Reinforcement Learning of Adaptive Energy Management for the Energy Transmission-Constrained Island Group
abstract
This article proposes a hybrid policy-based reinforcement learning (HPRL) adaptive energy management to realize the optimal operation for the island group energy system with energy transmission-constrained environment. An island energy hub (IEH) model that can realize the energy cascade utilization is proposed. Compared with the traditional model, the IEH can satisfy the special energy demand of island, meanwhile, ensure the energy supply of island. Moreover, an energy management model of islands group (EMIG) based on the IEH is formulated which comprehensively considers the inverse distribution of energy demand and resources, as well as the limited energy transmission. Since the environment model of the island is difficult to construct due to the increase of proportion of renewable energy generation and civilian load, the EMIG is transformed into a reinforcement learning (RL) task which features model-free. Considering the limitations of traditional RL in discrete-continuous hybrid action space, HPRL is proposed to achieve optimal operation without simplifying the model. Numerical simulations demonstrate the effectiveness of the proposed adaptive energy management.
Lingxiao Yang, Xiaofeng Li 0014, Mengwei Sun, Changyin Sun 0001
IEEE Trans. Ind. Informatics1
2023 Toward Intrinsic Adversarial Robustness Through Probabilistic Training
abstract
Modern deep neural networks have made numerous breakthroughs in real-world applications, yet they remain vulnerable to some imperceptible adversarial perturbations. These tailored perturbations can severely disrupt the inference of current deep learning-based methods and may induce potential security hazards to artificial intelligence applications. So far, adversarial training methods have achieved excellent robustness against various adversarial attacks by involving adversarial examples during the training stage. However, existing methods primarily rely on optimizing injective adversarial examples correspondingly generated from natural examples, ignoring potential adversaries in the adversarial domain. This optimization bias can induce the overfitting of the suboptimal decision boundary, which heavily jeopardizes adversarial robustness. To address this issue, we propose Adversarial Probabilistic Training (APT) to bridge the distribution gap between the natural and adversarial examples via modeling the latent adversarial distribution. Instead of tedious and costly adversary sampling to form the probabilistic domain, we estimate the adversarial distribution parameters in the feature level for efficiency. Moreover, we decouple the distribution alignment based on the adversarial probability model and the original adversarial example. We then devise a novel reweighting mechanism for the distribution alignment by considering the adversarial strength and the domain uncertainty. Extensive experiments demonstrate the superiority of our adversarial probabilistic training method against various types of adversarial attacks in different datasets and scenarios.
Junhao Dong 0001, Lingxiao Yang, Yuan Wang 0030, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Image Process.2
2023 Learning Shadow Removal From Unpaired Samples via Reciprocal Learning
abstract
We focus on addressing the problem of shadow removal for an image, and attempt to make a weakly supervised learning model that does not depend on the pixelwise-paired training samples, but only uses the samples with image-level labels that indicate whether an image contains shadow or not. To this end, we propose a deep reciprocal learning model that interactively optimizes the shadow remover and the shadow detector to improve the overall capability of the model. On the one hand, shadow removal is modeled as an optimization problem with a latent variable of the detected shadow mask. On the other hand, a shadow detector can be trained using the prior from the shadow remover. A self-paced learning strategy is employed to avoid fitting to intermediate noisy annotation during the interactive optimization. Furthermore, a color-maintenance loss and a shadow-attention discriminator are both designed to facilitate model optimization. Extensive experiments on the pairwise ISTD dataset, SRD dataset, and unpaired USR dataset demonstrate the superiority of the proposed deep reciprocal model.
Xiaohua Xie, Kuoyu Deng, Lingxiao Yang, Jian-Huang Lai
IEEE Trans. Image Process.4
2023 Learning Weak Semantics by Feature Graph for Attribute-Based Person Search
abstract
Attribute-based person search aims to find the target person from the gallery images based on the given query text. It often plays an important role in surveillance systems when visual information is not reliable, such as identifying a criminal from a few witnesses. Although recent works have made great progress, most of them neglect the attribute labeling problems that exist in the current datasets. Moreover, these problems also increase the risk of non-alignment between attribute texts and visual images, leading to large semantic gaps. To address these issues, in this paper, we propose Weak Semantic Embeddings (WSEs), which can modify the data distribution of the original attribute texts and thus improve the representability of attribute features. We also introduce feature graphs to learn more collaborative and calibrated information. Furthermore, the relationship modeled by our feature graphs between all semantic embeddings can reduce the semantic gap in text-to-image retrieval. Extensive evaluations on three challenging benchmarks - PETA, Market-1501 Attribute, and PA100K, demonstrate the effectiveness of the proposed WSEs, and our method outperforms existing state-of-the-art methods.
Qiyang Peng, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Image Process.2
2022 Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation
abstract
Weakly Supervised Semantic Segmentation (WSSS) based on image-level labels has attracted much attention due to low annotation costs. Existing methods often rely on Class Activation Mapping (CAM) that measures the correlation between image pixels and classifier weight. However, the classifier focuses only on the discriminative regions while ignoring other useful information in each image, resulting in incomplete localization maps. To address this issue, we propose a Self-supervised Image-specific Prototype Exploration (SIPE) that consists of an Image-specific Prototype Exploration (IPE) and a General-Specific Consistency (GSC) loss. Specifically, IPE tailors prototypes for every image to capture complete regions, formed our Image-Specific CAM (IS-CAM), which is realized by two sequential steps. In addition, GSC is proposed to construct the consistency of general CAM and our specific IS-CAM, which further optimizes the feature representation and empowers a self-correction ability of prototype exploration. Extensive experiments are conducted on PASCAL VOC 2012 and MS COCO 2014 segmentation benchmark and results show our SIPE achieves new state-of-the-art performance using only image-level labels. The code is available at https://github.com/chenqi1126/SIPE.
Qi Chen 0013, Lingxiao Yang, Jian-Huang Lai, Xiaohua Xie
CVPR2
2022 Exploring Dual-task Correlation for Pose Guided Person Image Generation
abstract
Pose Guided Person Image Generation (PGPIG) is the task of transforming a person image from the source pose to a given target pose. Most of the existing methods only focus on the ill-posed source-to-target task and fail to capture reasonable texture mapping. To address this problem, we propose a novel Dual-task Pose Transformer Network (DPTN), which introduces an auxiliary task (i.e., source-to-source task) and exploits the dual-task correlation to promote the performance of PGPIG. The DPTN is of a Siamese structure, containing a source-to-source self-reconstruction branch, and a transformation branch for source-to-target generation. By sharing partial weights between them, the knowledge learned by the source-to-source task can effectively assist the source-to-target learning. Furthermore, we bridge the two branches with a proposed Pose Transformer Module (PTM) to adaptively explore the correlation between features from dual tasks. Such correlation can establish the fine-grained mapping of all the pixels between the sources and the targets, and promote the source texture transmission to enhance the details of the generated target images. Extensive experiments show that our DPTN outperforms state-of-the-arts in terms of both PSNR and LPIPS. In addition, our DPTN only contains 9.79 million parameters, which is significantly smaller than other approaches. Our code is available at: https://github.com/PangzeCheung/Dual-task-Pose-Transformer-Network.
Pengze Zhang, Lingxiao Yang, Jian-Huang Lai, Xiaohua Xie
CVPR2
2022 AcroFOD: An Adaptive Method for Cross-Domain Few-Shot Object Detection
Yipeng Gao, Lingxiao Yang, Yunmu Huang, Song Xie, Wei-Shi Zheng 0001
ECCV (33)2
2022 A Weighting-Based Tabu Search Algorithm for the p-Next Center Problem
abstract
The p-next center problem (pNCP) is an extension of the classical p-center problem. It consists of locating p centers from a set of candidate centers and allocating both a reference and a backup center to each client, to minimize the maximum cost, which is the length of the path from a client to its reference center and then to its backup center. Among them, the reference center is the closest center to a client and serves it under normal circumstances, while the backup center is the closest center to the reference center and serves the client when the reference center is out of service. In this paper, we propose a weighting-based tabu search algorithm called WTS for solving pNCP. WTS optimizes the pNCP by solving its decision subproblems with given assignment costs with an efficient swap-based neighborhood structure and a hierarchical penalty strategy for neighborhood evaluation. Extensive experimental studies on 413 benchmark instances demonstrate that WTS outperforms the state-of-the-art methods in the literature. Specifically, WTS improves 12 previous best known results and matches the optimal results for all remaining 401 ones in a much shorter time than other algorithms. More importantly, WTS reaches the lower bounds for 10 instances for the first time.
Zhouxing Su, Zhipeng Lü, Lingxiao Yang
IJCAI4
2022 Re4: Learning to Re-contrast, Re-attend, Re-construct for Multi-interest Recommendation
abstract
Effectively representing users lie at the core of modern recommender systems. Since users’ interests naturally exhibit multiple aspects, it is of increasing interest to develop multi-interest frameworks for recommendation, rather than represent each user with an overall embedding. Despite their effectiveness, existing methods solely exploit the encoder (the forward flow) to represent multiple aspects of interests. However, without explicit regularization, the interest embeddings may not be distinct from each other nor semantically reflect representative historical items. Towards this end, we propose the Re4 framework, which leverages the backward flow to reexamine each interest embedding. Specifically, Re4 encapsulates three backward flows, i.e., 1) Re-contrast, which drives each interest embedding to be distinct from other interests using contrastive learning; 2) Re-attend, which ensures the interest-item correlation estimation in the forward flow to be consistent with the criterion used in final recommendation; and 3) Re-construct, which ensures that each interest embedding can semantically reflect the information of representative items that relate to the corresponding interest. We demonstrate the novel forward-backward multi-interest paradigm on ComiRec, and perform extensive experiments on three real-world datasets. Empirical studies validate that Re4 helps to learn learning distinct and effective multi-interest representations.
Shengyu Zhang 0001, Lingxiao Yang, Dong Yao, Fuli Feng, Zhou Zhao 0001, Tat-Seng Chua, Fei Wu 0001
WWW2
2022 Relaxation LIF: A gradient-based spiking neuron for direct training deep spiking neural networks
Jianxiong Tang, Jian-Huang Lai, Wei-Shi Zheng 0001, Lingxiao Yang, Xiaohua Xie
Neurocomputing4
2022 Lightweight Texture Correlation Network for Pose Guided Person Image Generation
abstract
Pose Guided Person Image Generation (PGPIG) is a popular task in deepfake, which aims at generating a person image with the given pose based on the source image. However, existing methods cannot comprehensively model the correlation between the source and the target domain. Most of them only focus on the correlation of the keypoints but ignore detail textures. In this paper, we propose a novel Texture Correlation Network (TCN) to simultaneously build pose and texture correlations. Specifically, our TCN adopts a two-stage design, including two networks: Pose Guided Person Alignment Network (PGPAN) and Texture Correlation Attention Network (TCAN). The PGPAN generates a coarse person image aligned with the target pose, while the TCAN produces a target generated image with the guidance of multiple correlations. The key component of TCAN is our new module, Texture Correlation Attention Module (TCAM), which explicitly builds geometry and texture correlation between the source image and the coarse target image. Those kinds of correlations facilitate to transfer real textures from the source to the target. Extensive experiments on the DeepFashion and Market1501 benchmarks demonstrate the superior performance of the proposed method. In addition, our model only uses 8.5 million parameters, which is significantly smaller than other methods.
Pengze Zhang, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Circuits Syst. Video Technol.2
2022 Selective Intra-Image Similarity for Personalized Fixation-Based Object Segmentation
abstract
Personalized Fixation-based Object Segmentation (PFOS) aims at segmenting the gazed objects in images conditioned on personalized fixations. However, the performances of existing PFOS methods are degraded when facing anomalous fixation maps (some fixations fall in the background) or enormous objects because of their poor localization ability. In this paper, we propose a novel Selective Intra-image Similarity Network (SISNet) that achieves significant performance by precisely localizing the gazed objects. First, we propose a Response Purifying Module (RPM) to eliminate the false response regions caused by anomalous fixations in the background. By suppressing these false responses, we can significantly reduce the negative impacts caused by anomalous fixations. Second, we propose an intra-image similarity module (ISM) to better localize large objects by integrating more long-range information. In addition, we propose a new Discriminative Intersection-over-Union metric that evaluates whether PFOS methods can produce distinctive predictions for varying fixations. Experiments on the PFOS and our proposed OSIE-CFPS-UN datasets prove that our network achieves remarkable improvements and outperforms existing state-of-the-art methods. Code has been published athttps://www.github.com/moothes/SISNet.
Huajun Zhou, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Circuits Syst. Video Technol.2
2022 Event-Triggered Distributed Hybrid Control Scheme for the Integrated Energy System
abstract
For the integrated energy system (IES) formed by a cluster of energy hubs (EHs), the outputs of EHs and system parameters, which greatly influence the security performance, should be properly adjusted. In this article, an event-triggered distributed hybrid control scheme is proposed to achieve security and economic operation for the IES. First, the EH output control is designed based on the features of energy networks as well as the containment and consensus algorithms. According to the control of outputs, the electricity and heat load power can be accurately shared without knowing the network parameters. Second, the pressure is bounded within an acceptable range, and the frequency is reverted to the reference value by implementing the proposed control method. Third, an event-triggered communication strategy is employed to design the corresponding protocols, resulting in reduced communication cost. Finally, the control of devices based on the equal incremental principle is proposed to achieve the minimal economic cost for each EH with considering energy prices. The results of numerical case studies are presented to validate the performance of the proposed control method.
Ning Zhang 0037, Qiuye Sun, Lingxiao Yang, Yushuai Li
IEEE Trans. Ind. Informatics3
2022 Optimal Energy Operation Strategy for We-Energy of Energy Internet Based on Hybrid Reinforcement Learning With Human-in-the-Loop
abstract
This article investigates the energy operation problem based on We-Energy (WE), a novel full-duplex model in Energy Internet (EI). A dual-objective optimal energy operation model of WE is formulated with the consideration of economical benefit and security operation under different time scenarios. Due to the inaccurate model of distributed generation devices and loads, a multipolicy convex hull reinforcement learning (MCRL) algorithm is proposed. It can find the multiobjective strategy set with model-free feature. Moreover, considering the limitations of artificial intelligence technology and the human advantages in information processing for complex task, a two-channel Human-in-the-loop (HITL) method is designed to combine with MCRL to avoid decision-making risks. The one channel of HITL can evaluate the operation strategy by human under normal conditions so that the understanding of human for complex operating conditions can be incorporated into the machine learning algorithms to improve the confidence of intelligent systems. The other channel of HITL can allow human to participate in real-time adjustment under abnormal conditions to avoid system out of control. Simulation studies of modified EI are confirmed that the proposed algorithm can improve system performance effectively.
Lingxiao Yang, Qiuye Sun, Ning Zhang 0037, Zhenwei Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2021 SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks
abstract
In this paper, we propose a conceptually simple but very effective attention module for Convolutional Neural Networks (ConvNets). In contrast to existing channel-wise and spatial-wise attention modules, our module instead infers 3-D attention weights for the feature map in a layer without adding parameters to the original networks. Specifically, we base on some well-known neuroscience theories and propose to optimize an energy function to find the importance of each neuron. We further derive a fast closed-form solution for the energy function, and show that the solution can be implemented in less than ten lines of code. Another advantage of the module is that most of the operators are selected based on the solution to the defined energy function, avoiding too many efforts for structure tuning. Quantitative evaluations on various visual tasks demonstrate that the proposed module is flexible and effective to improve the representation ability of many ConvNets. Our code is available at Pytorch-SimAM.
Lingxiao Yang, Ru-Yuan Zhang, Lida Li, Xiaohua Xie
ICML1
2021 Distributed Adaptive Dual Control via Consensus Algorithm in the Energy Internet
abstract
This article investigates a distributed adaptive dual control that employs both consensus algorithm and improved equal incremental principle (IEIP), to guarantee the security operation while reducing the energy consumption of the energy Internet (EI). Since the security operation is a critical factor of the EI with the energy hub (EH), it is necessary for the EI to adjust the outputs of EHs and the system parameters, which have a great influence on the performance. The consensus-based control of the EI can resolve the energy-coupling issue and accurately share the electricity and heat loads power without requiring the information of the network parameters, which is hard to know. Meanwhile, the variations of system parameters caused by the droop action greatly impact the security operation of the EI. It can be adaptively recovered by executing the proposed control strategy. Furthermore, in order to reduce the energy consumption, the equipment in hubs should also be managed in view of the features of EH. The minimal loss of the energy can be achieved by utilizing the control of devices via the IEIP by considering the coupling characteristics of devices. The performance of the proposed method is demonstrated through numerical case studies.
Ning Zhang 0037, Qiuye Sun, Jiawei Wang 0015, Lingxiao Yang
IEEE Trans. Ind. Informatics4
2021 Contour-Aware Loss: Boundary-Aware Learning for Salient Object Segmentation
abstract
We present a learning model that makes full use of boundary information for salient object segmentation. Specifically, we come up with a novel loss function, i.e., Contour Loss, which leverages object contours to guide models to perceive salient object boundaries. Such a boundary-aware network can learn boundary-wise distinctions between salient objects and background, hence effectively facilitating the salient object segmentation. Yet the Contour Loss emphasizes the boundaries to capture the contextual details in the local range. We further propose the hierarchical global attention module (HGAM), which forces the model hierarchically to attend to global contexts, thus captures the global visual saliency. Comprehensive experiments on six benchmark datasets show that our method achieves superior performance over state-of-the-art ones. Moreover, our model has a real-time speed of 26 fps on a TITAN X GPU.
Huajun Zhou, Jian-Huang Lai, Lingxiao Yang, Xiaohua Xie
IEEE Trans. Image Process.4
2021 Layer-Output Guided Complementary Attention Learning for Image Defocus Blur Detection
abstract
Defocus blur detection (DBD), which has been widely applied to various fields, aims to detect the out-of-focus or in-focus pixels from a single image. Despite the fact that the deep learning based methods applied to DBD have outperformed the hand-crafted feature based methods, the performance cannot still meet our requirement. In this paper, a novel network is established for DBD. Unlike existing methods which only learn the projection from the in-focus part to the ground-truth, both in-focus and out-of-focus pixels, which are completely and symmetrically complementary, are taken into account. Specifically, two symmetric branches are designed to jointly estimate the probability of focus and defocus pixels, respectively. Due to their complementary constraint, each layer in a branch is affected by an attention obtained from another branch, effectively learning the detailed information which may be ignored in one branch. The feature maps from these two branches are then passed through a unique fusion block to simultaneously get the two-channel output measured by a complementary loss. Additionally, instead of estimating only one binary map from a specific layer, each layer is encouraged to estimate the ground truth to guide the binary map estimation in its linked shallower layer followed by a top-to-bottom combination strategy, gradually exploiting the global and local information. Experimental results on released datasets demonstrate that our proposed method remarkably outperforms state-of-the-art algorithms.
Jinxing Li 0003, Lingxiao Yang, Shuhang Gu, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001
IEEE Trans. Image Process.3
2020 Interactive Two-Stream Decoder for Accurate and Fast Saliency Detection
abstract
Recently, contour information largely improves the performance of saliency detection. However, the discussion on the correlation between saliency and contour remains scarce. In this paper, we first analyze such correlation and then propose an interactive two-stream decoder to explore multiple cues, including saliency, contour and their correlation. Specifically, our decoder consists of two branches, a saliency branch and a contour branch. Each branch is assigned to learn distinctive features for predicting the corresponding map. Meanwhile, the intermediate connections are forced to learn the correlation by interactively transmitting the features from each branch to the other one. In addition, we develop an adaptive contour loss to automatically discriminate hard examples during learning process. Extensive experiments on six benchmarks well demonstrate that our network achieves competitive performance with a fast speed around 50 FPS. Moreover, our VGG-based model only contains 17.08 million parameters, which is significantly smaller than other VGG-based approaches. Code has been made available at: https://github.com/moothes/ITSD-pytorch.
Huajun Zhou, Xiaohua Xie, Jian-Huang Lai, Lingxiao Yang
CVPR5
2020 Nash Q-learning based equilibrium transfer for integrated energy management game with We-Energy
Lingxiao Yang, Qiuye Sun, Dazhong Ma, Qinglai Wei
Neurocomputing1
2019 Learning a Visual Tracker from a Single Movie without Annotation
abstract
The recent success of deep network in visual trackers learning largely relies on human labeled data, which are however expensive to annotate. Recently, some unsupervised methods have been proposed to explore the learning of visual trackers without labeled data, while their performance lags far behind the supervised methods. We identify the main bottleneck of these methods as inconsistent objectives between off-line training and online tracking stages. To address this problem, we propose a novel unsupervised learning pipeline which is based on the discriminative correlation filter network. Our method iteratively updates the tracker by alternating between target localization and network optimization. In particular, we propose to learn the network from a single movie, which could be easily obtained other than collecting thousands of video clips or millions of images. Extensive experiments demonstrate that our approach is insensitive to the employed movies, and the trained visual tracker achieves leading performance among existing unsupervised learning approaches. Even compared with the same network trained with human labeled bounding boxes, our tracker achieves similar results on many tracking benchmarks. Code is available at: https://github.com/ZjjConan/UL-Tracker-AAAI2019.
Lingxiao Yang, David Zhang 0001, Lei Zhang 0006
AAAI1
2019 Dynamic Anchor Feature Selection for Single-Shot Object Detection
abstract
The design of anchors is critical to the performance of one-stage detectors. Recently, the anchor refinement module (ARM) has been proposed to adjust the initialization of default anchors, providing the detector a better anchor reference. However, this module brings another problem: all pixels at a feature map have the same receptive field while the anchors associated with each pixel have different positions and sizes. This discordance may lead to a less effective detector. In this paper, we present a dynamic feature selection operation to select new pixels in a feature map for each refined anchor received from the ARM. The pixels are selected based on the new anchor position and size so that the receptive filed of these pixels can fit the anchor areas well, which makes the detector, especially the regression part, much easier to optimize. Furthermore, to enhance the representation ability of selected feature pixels, we design a bidirectional feature fusion module by combining features from early and deep layers. Extensive experiments on both PASCAL VOC and COCO demonstrate the effectiveness of our dynamic anchor feature selection (DAFS) operation. For the case of high IoU threshold, our DAFS can improve the mAP by a large margin.
Shuai Li 0014, Lingxiao Yang, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006
ICCV2
2018 Learning discriminative visual elements using part-based convolutional neural network
Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
Neurocomputing1
2017 Part-based convolutional neural network for visual recognition
abstract
Mid-level element based representations have been proven to be very effective for visual recognition. We present a method to discover discriminative elements based on deep Convolutional Neural Networks (CNNs), namely Part-based CNN (P-CNN), which acts as the role of encoding module in part-based representation. The P-CNN can be attached at arbitrary layer of a pre-trained CNN and be trained using image-level labels. The training of P-CNN essentially corresponds to the optimization and selection of discriminative mid-level visual elements. For an input image, the output of P-CNN is naturally the part-based coding and can be directly used for image recognition. By applying P-CNN to multiple layers of a pretrained CNN, more diverse visual elements can be obtained for visual recognitions. Experiments are conducted on two recognition tasks and their results demonstrate the effectiveness of the proposed method.
Lingxiao Yang, Xiaohua Xie, Peihua Li, David Zhang 0001, Lei Zhang 0006
ICIP1
2017 Learning a real-time generic tracker using convolutional neural networks
abstract
This paper presents a novel frame-pair based method for visual object tracking. Instead of adopting two-stream Convolutional Neural Networks (CNNs) to represent each frame, we stack frame pairs as the input, resulting in a single-stream CNN tracker with much fewer parameters. The proposed tracker can learn generic motion patterns of objects with much less annotated videos than previous methods. Besides, it is found that trackers trained using two successive frames tend to predict the centers of searching windows as the locations of tracked targets. To alleviate this problem, we propose a novel sampling strategy for off-line training. Specifically, we construct a pair by sampling two frames with a random offset. The offset controls the moving smoothness of objects. Experiments on the challenging VOT14 and OTB datasets show that the proposed tracker performs on par with recently developed generic trackers, but with much less memory. In addition, our tracker can run in a speed of over 100 (30) fps with a GPU (CPU), much faster than most deep neural network based trackers.
Linnan Zhu, Lingxiao Yang, David Zhang 0001, Lei Zhang 0006
ICME2
2017 Multi-Agent Q( \lambda ) Learning for Optimal Operation Management of Energy Internet
Lingxiao Yang, Qiuye Sun
ICONIP (6)1
2017 Deep Location-Specific Tracking
abstract
Convolutional Neural Network (CNN) based methods have shown significant performance gains in the problem of visual tracking in recent years. Due to many uncertain changes of objects online, such as abrupt motion, background clutter and large deformation, the visual tracking is still a challenging task. We propose a novel algorithm, namely Deep Location-Specific Tracking, which decomposes the tracking problem into a localization task and a classification task, and trains an individual network for each task. The localization network exploits the information in the current frame and provides a specific location to improve the probability of successful tracking, while the classification network finds the target among many examples generated around the target location in the previous frame, as well as the one estimated from the localization network in the current frame. CNN based trackers often have massive number of trainable parameters, and are prone to over-fitting to some particular object states, leading to less precision or tracking drift. We address this problem by learning a classification network based on 1 × 1 convolution and global average pooling. Extensive experimental results on popular benchmark datasets show that the proposed tracker achieves competitive results without using additional tracking videos for fine-tuning. The code is available at https://github.com/ZjjConan/DLST
Lingxiao Yang, Risheng Liu, David Zhang 0001, Lei Zhang 0006
ACM Multimedia1
2016 Learning object-specific DAGs for multi-label material recognition
Xiaohua Xie, Lingxiao Yang, Wei-Shi Zheng 0001
Comput. Vis. Image Underst.2
2016 Exploiting object semantic cues for Multi-label Material Recognition
Lingxiao Yang, Xiaohua Xie
Neurocomputing1
2016 Query intent inference via search engine log
Lingxiao Yang
Knowl. Inf. Syst.2
2015 Max-margin analysis based patch sampling for discovery of mid-level parts
abstract
Discovering representative, discriminative mid-level parts is crucial for visual recognition models such as Bag-Of-Parts. We present a weakly-supervised approach to learn class-specific mid-level parts from a database. In our approach, only the image-level labels but no additional human annotations are used. As a start, we employ a SVM-like model to sample discriminant patches from each image. The employed SVM-like model corresponds to a max-margin analysis between a specific image patch and other patches from the whole training set, which can be easily solved in a closed form. For each class, the sampled patches are then clustered in an agglomerative manner to generate the final semantic parts, in the meantime the less-representative patches are discarded. The proposed approach is effective since it sequentially discards the non-discriminative and non-representative patches. The approach is also efficient since the clustering operation only needs to handle a small number of discriminant patches. The state-of-the-art results are observed in scene classification benchmarks when using the learned parts as a visual codebook.
Lingxiao Yang, Xiaohua Xie
ICIP1
2015 TEII: Topic enhanced inverted index for top-k document retrieval
Kenneth Wai-Ting Leung, Lingxiao Yang, Wilfred Ng
Knowl. Based Syst.3
2015 Query suggestion with diversification and personalization
Kenneth Wai-Ting Leung, Lingxiao Yang, Wilfred Ng
Knowl. Based Syst.3
2015 SG-WSTD: A framework for scalable geographic web search topic discovery
Jan Vosecky, Kenneth Wai-Ting Leung, Lingxiao Yang, Wilfred Ng
Knowl. Based Syst.4