Luxin Zhang

dblp:218/7106 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 5 since 2021Computer networks · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Class-aware contrastive learning for radio signal generalized category discovery
Jie Chen 0090, Shilian Zheng, Luxin Zhang, Keqiang Yue, Zhijin Zhao
Eng. Appl. Artif. Intell.3
2026 MSM-Pnet: Multiscale-Masked Transformer Pretraining for FM-Based Positioning
abstract
To overcome the limitations of traditional satellite navigation technologies in complex and signal-obstructed industrial environments, this paper presents MSM-Pnet, a novel semi-supervised FM-based positioning framework leveraging FM signals of opportunity. By integrating wavelet packet decomposition with a multi-scale Vision Transformer and a hybrid masking strategy that combines random and time–frequency-aware masking, MSM-Pnet introduces a masked autoencoder architecture capable of robust positioning with limited labeled data. Experimental results demonstrate that MSM-Pnet consistently outperforms conventional supervised learning methods in both indoor and outdoor environments, while also significantly reducing model complexity. These results highlight the method’s potential as a cost-effective and scalable solution for seamless indoor–outdoor positioning for Internet of Things systems.
Shilian Zheng, Quan Lin, Luxin Zhang, Xinjiang Qiu, Keqiang Yue, Zhijin Zhao, Xiaoniu Yang
IEEE Internet Things J.3
2026 Adversarially Robust Wideband Spectrum Sensing in the Frequency Domain
Shilian Zheng, Zhihao Ye, Luxin Zhang, Keqiang Yue, Weiguo Shen, Zhijin Zhao
IEEE Trans. Commun.3
2025 MoCha: Towards Movie-Grade Talking Character Generation
abstract
Recent advancements in video generation have achieved impressive motion realism, yet they often overlook character-driven storytelling, a crucial task for automated film, animation generation. We introduce Talking Characters, a more realistic task to generate talking character animations directly from speech and text. Unlike talking head tasks, Talking Characters aims at generating the full portrait of one or more characters beyond the facial region. In this paper, we propose MoCha, the first of its kind to generate talking characters. To ensure precise synchronization between video and speech, we propose a localized audio attention mechanism that effectively aligns speech and video tokens. To address the scarcity of large-scale speech-labelled video datasets, we introduce a joint training strategy that leverages both speech-labelled and text-labelled video data, significantly improving generalization across diverse character actions. We also design structured prompt templates with character tags, enabling, for the first time, multi-character conversation with turn-based dialogue—allowing AI-generated characters to engage in context-aware conversations with cinematic coherence. Extensive qualitative and quantitative evaluations, including human evaluation studies and benchmark comparisons, demonstrate that MoCha sets a new standard for AI-generated cinematic storytelling, achieving superior realism, controllability and generalization.
Cong Wei 0001, Ji Hou, Felix Juefei-Xu, Zecheng He, Xiaoliang Dai, Luxin Zhang, Tingbo Hou, Animesh Sinha, Peter Vajda, Wenhu Chen
NeurIPS8
2025 BallPri: test cases prioritization for deep neuron networks via tolerant ball in variable space
Chengyu Jia 0001, Jinyin Chen, Xiaohao Li, Haibin Zheng, Luxin Zhang
Autom. Softw. Eng.5
2025 WK-Pnet: FM-Based Positioning via Wavelet Packet Decomposition and Knowledge Distillation
abstract
Accurate and efficient positioning in complex environments remains a critical challenge where satellite-based systems (e.g., GNSS) suffer from signal attenuation and multipath interference. This paper proposes WK-Pnet, a lightweight positioning framework that utilizes frequency modulation (FM) signals and integrates Wavelet Packet Decomposition (WPD) with knowledge distillation. WK-Pnet first decomposes raw FM IQ signals using WPD to extract fine-grained multi-scale time-frequency features, preserving both spectral and phase information. These features are then fed into a deep neural network for location estimation. To reduce computational complexity, we employ a knowledge distillation strategy that transfers knowledge from a large-capacity ResNeXt-based teacher model—enhanced with a spatial attention mechanism—to a compact student network with significantly fewer parameters and FLOPs. The proposed method is validated on publicly available indoor and outdoor datasets, showing that WK-Pnet achieves comparable positioning accuracy to the teacher model while reducing FLOPs by 95.9%, model parameters by 99.3%, and inference latency by 90.5% on edge devices. Experimental comparisons also reveal that WPD outperforms STFT and EMD in positioning stability and accuracy, especially in outdoor scenarios. WK-Pnet demonstrates strong robustness, low-latency inference, and high accuracy, making it highly suitable for real-time, resource-constrained mobile and IoT applications.
Shilian Zheng, Quan Lin, Peihan Qi, Luxin Zhang, Xinjiang Qiu, Zhijin Zhao, Xiaoniu Yang
IEEE Internet Things J.4
2024 AVID: Any-Length Video Inpainting with Diffusion Model
abstract
Recent advances in diffusion models have successfully enabled text-guided image inpainting. While it seems straightforward to extend such editing capability into the video domain, there have been fewer works regarding textguided video inpainting. Given a video, a masked region at its initial frame, and an editing prompt, it requires a model to do infilling at each frame following the editing guidance while keeping the out-of-mask region intact. There are three main challenges in text-guided video inpainting: (i) temporal consistency of the edited video, (ii) supporting different inpainting types at different structural fidelity levels, and (iii) dealing with variable video length. To address these challenges, we introduce Any-Length Video Inpainting with Diffusion Model, dubbed as AVID. At its core, our model is equipped with effective motion modules and adjustable structure guidance, for fixed-length video inpainting. Building on top of that, we propose a novel Temporal MultiDiffusion sampling pipeline with a middle-frame attention guidance mechanism, facilitating the generation of videos with any desired duration. Our comprehensive experiments show our model can robustly deal with various inpainting types at different video duration ranges, with high quality11More visualization results are made publicly available here.
Bichen Wu, Yaqiao Luo, Luxin Zhang, Peter Vajda, Dimitris N. Metaxas, Licheng Yu
CVPR5
2024 MASSnet: Deep-Learning-Based Multiple-Antenna Spectrum Sensing for Cognitive-Radio-Enabled Internet of Things
abstract
Cognitive radio-based Internet of Things (CR-IoTs) provide an efficient spectrum management for IoT networks with massive wireless access and data transmission needs. As one of the key technologies of CR-IoT, spectrum sensing is of great research significance. Motivated by the recent boom on applications of deep learning in wireless communications networks and IoT, several spectrum sensing methods based on deep learning have emerged. However these algorithms train the sensing models with the extracted features of received signals and require a retraining of sensing models when the number of sensing antennas changes. Thus, we develop multiple-antenna spectrum sensing methods based on convolutional neural networks (MASSnet) using the in-phase (I) and quadrature (Q) components of the signals as the input. The three schemes of MASSnet also provide the flexibility to choose between retraining the sensing models or using the obtained models for different sensing antenna configurations. Experiment results demonstrate the superior performance of the proposed methods over existing deep learning-based spectrum sensing methods in terms of probability of detection especially in very low signal-to-noise ratio (SNR) condition. Furthermore, the proposed methods have good generalization ability to new noise distribution, new fading channel, different frequency offsets, and detecting signals with a new modulation even without retraining.
Luxin Zhang, Shilian Zheng, Kunfeng Qiu, Caiyi Lou, Xiaoniu Yang
IEEE Internet Things J.1
2024 FM-Based Positioning via Deep Learning
abstract
Frequency Modulation (FM) broadcast signals, regarded as opportunistic signals, hold significant potential for indoor and outdoor positioning applications. The existing FM-based positioning methods primarily rely on Received Signal Strength (RSS) for positioning, the accuracy of which needs improvement. In this paper, we introduce FM-Pnet, an end-to-end FM-based positioning method that leverages deep learning. This method utilizes the time-frequency representation of FM signals as network input, enabling automatically learning of deep features for positioning. We also propose two strategies, noise injection and enriching training samples, to enhance the model’s generalization performance over long time spans. We construct datasets for both indoor and outdoor scenarios and conduct extensive experiments to validate the performance of our proposed method. Experimental results demonstrate that FM-Pnet significantly outperforms traditional RSS-based positioning methods in terms of both positioning accuracy and stability.
Shilian Zheng, Jiacheng Hu, Luxin Zhang, Kunfeng Qiu, Jie Chen 0090, Peihan Qi, Zhijin Zhao, Xiaoniu Yang
IEEE J. Sel. Areas Commun.3
2024 DeepSIG: A Hybrid Heterogeneous Deep Learning Framework for Radio Signal Classification
abstract
Deep learning has been widely used in automatic modulation classification (AMC) recently. Most of deep learning-based AMC uses a single network model to deal with radio signals with a single input format. In this paper, we propose a hybrid heterogeneous modulation classification architecture named DeepSIG, which integrates Recurrent Neural Network (RNN), Convolutional Neural Network (CNN) and Graph Neural Network (GNN) models in a single framework to process radio signals with heterogeneous input formats, i.e., in-phase (I) and quadrature (Q) sequences, images mapped from IQ signals and graphs converted from IQ signals, to extract and integrate the features from different perspectives. A fusion training mechanism is presented to train DeepSIG. We use three different radio signal datasets for simulations. Results show that our proposed DeepSIG performs the best in terms of classification accuracy compared with the three methods with single input, i.e., sequence, image or graph. The performance gain is larger in few-shot scenarios.
Kunfeng Qiu, Shilian Zheng, Luxin Zhang, Caiyi Lou, Xiaoniu Yang
IEEE Trans. Wirel. Commun.3
2023 Adversarial Attacks on Deep Learning-Based DOA Estimation With Covariance Input
abstract
Although deep learning methods have made significant advancements across various domains, recent research has shown that carefully crafted adversarial samples can lead to a significant degradation in the performance of deep learning models. Such adversarial examples raise concerns about the reliability and safety of deep learning-based models. Currently, there is a lack of research on the robustness of deep learning based DOA methods against adversarial samples. This letter aims to fill this research gap by leveraging the differentiability of the transformation process from the original signal to the covariance matrix. By utilizing this differentiability, the robustness of the DOA estimation model, which takes the covariance matrix as input, is investigated. Four different white-box attack methods are considered to generate adversarial samples to evaluate the resilience of the model. The experimental results demonstrate that all four methods employed significantly increase the estimation error of the DOA estimation model, posing a serious threat to the model's security.
Shilian Zheng, Luxin Zhang, Zhijin Zhao, Xiaoniu Yang
IEEE Signal Process. Lett.3
2022 Interpretable Domain Adaptation for Hidden Subdomain Alignment in the Context of Pre-trained Source Models
abstract
Domain adaptation aims to leverage source domain knowledge to predict target domain labels. Most domain adaptation methods tackle a single-source, single-target scenario, whereas source and target domain data can often be subdivided into data from different distributions in real-life applications (e.g., when the distribution of the collected data changes with time). However, such subdomains are rarely given and should be discovered automatically. To this end, some recent domain adaptation works seek separations of hidden subdomains, w.r.t. a known or fixed number of subdomains. In contrast, this paper introduces a new subdomain combination method that leverages a variable number of subdomains. Precisely, we propose to use an inter-subdomain divergence maximization criterion to exploit hidden subdomains. Besides, our proposition stands in a target-to-source domain adaptation scenario, where one exploits a pre-trained source model as a black box; thus, the proposed method is model-agnostic. By providing interpretability at two complementary levels (transformation and subdomain levels), our method can also be easily interpreted by practitioners with or without machine learning backgrounds. Experimental results over two fraud detection datasets demonstrate the efficiency of our method.
Luxin Zhang, Pascal Germain, Yacine Kessaci, Christophe Biernacki
AAAI1
2022 Interpretable domain adaptation using unsupervised feature selection on pre-trained source models
Luxin Zhang, Pascal Germain, Yacine Kessaci, Christophe Biernacki
Neurocomputing1
2020 Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset
abstract
Large-scale public datasets have been shown to benefit research in multiple areas of modern artificial intelligence. For decision-making research that requires human data, high-quality datasets serve as important benchmarks to facilitate the development of new methods by providing a common reproducible standard. Many human decision-making tasks require visual attention to obtain high levels of performance. Therefore, measuring eye movements can provide a rich source of information about the strategies that humans use to solve decision-making tasks. Here, we provide a large-scale, high-quality dataset of human actions with simultaneously recorded eye movements while humans play Atari video games. The dataset consists of 117 hours of gameplay data from a diverse set of 20 games, with 8 million action demonstrations and 328 million gaze samples. We introduce a novel form of gameplay, in which the human plays in a semi-frame-by-frame manner. This leads to near-optimal game decisions and game scores that are comparable or better than known human records. We demonstrate the usefulness of the dataset through two simple applications: predicting human gaze and imitating human demonstrated actions. The quality of the data leads to promising results in both tasks. Moreover, using a learned human gaze model to inform imitation learning leads to an 115% increase in game performance. We interpret these results as highlighting the importance of incorporating human visual attention in models of decision making and demonstrating the value of the current dataset to the research community. We hope that the scale and quality of this dataset can provide more opportunities to researchers in the areas of visual attention, imitation learning, and reinforcement learning.
Calen Walshe, Zhuode Liu, Lin Guan 0003, Karl S. Muller, Jake Alden Whritner, Luxin Zhang, Mary M. Hayhoe, Dana H. Ballard
AAAI7
2020 Cloth Region Segmentation for Robust Grasp Selection
abstract
Cloth detection and manipulation is a common task in domestic and industrial settings, yet such tasks remain a challenge for robots due to cloth deformability. Furthermore, in many cloth-related tasks like laundry folding and bed making, it is crucial to manipulate specific regions like edges and corners, as opposed to folds. In this work, we focus on the problem of segmenting and grasping these key regions. Our approach trains a network to segment the edges and corners of a cloth from a depth image, distinguishing such regions from wrinkles or folds. We also provide a novel algorithm for estimating the grasp location, direction, and directional uncertainty from the segmentation. We demonstrate our method on a real robot system and show that it outperforms baseline methods on grasping success. Video and other supplementary materials are available at: https://sites.google.com/view/cloth-segmentation.
Jianing Qian, Thomas Weng, Luxin Zhang, Brian Okorn, David Held
IROS3
2020 Target to Source Coordinate-Wise Adaptation of Pre-trained Models
Luxin Zhang, Pascal Germain, Yacine Kessaci, Christophe Biernacki
ECML/PKDD (1)1
2019 Optimal Communication Scheduling in the Smart Grid
abstract
This paper focuses on obtaining the optimal communication topology in the smart grid architecture, i.e., what is the optimal communication setup of smart meters in a smart building. The fact that smart meters also consume energy, more often than not, gets ignored by researchers and engineers. In this paper, we will show that smart meter networks can consume significantly less energy with optimal scheduling. Numerical results show that the overall energy consumption can be reduced by implementing the optimal communication architecture and transmission rate setup, rather than implementing a straightforward communication architecture with uniform channel bandwidth.
Luxin Zhang, Eric C. Kerrigan, Bikash C. Pal
IEEE Trans. Ind. Informatics1
2018 Learning Attention Model From Human for Visuomotor Tasks
abstract
A wealth of information regarding intelligent decision making is conveyed by human gaze and visual attention, hence, modeling and exploiting such information might be a promising way to strengthen algorithms like deep reinforcement learning. We collect high-quality human action and gaze data while playing Atari games. Using these data, we train a deep neural network that can predict human gaze positions and visual attention with high accuracy.
Luxin Zhang, Zhuode Liu, Mary M. Hayhoe, Dana H. Ballard
AAAI1
2018 AGIL: Learning Attention from Human for Visuomotor Tasks
Zhuode Liu, Luxin Zhang, Jake Alden Whritner, Karl S. Muller, Mary M. Hayhoe, Dana H. Ballard
ECCV (11)3