Yuxin Cheng

dblp:48/7054 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture
abstract
Low-rank adaptation (LoRA) is a predominant parameter-efficient finetuning method for adapting large language models (LLMs) to downstream tasks. Meanwhile, Compute-in-Memory (CIM) architectures demonstrate superior energy efficiency due to their array-level parallel in-memory computing designs. In this article, we propose deploying the LoRA-finetuned LLMs on the hybrid CIM architecture (i.e., pretrained weights onto energy-efficient Resistive Random-Access Memory (RRAM) and LoRA branches onto noise-free Static Random-Access Memory (SRAM)), reducing the energy cost to about 3% compared with the Nvidia A100 GPU. However, the inherent noise of RRAM on the saved weights leads to performance degradation, simultaneously. To address this issue, we design a novel Hardware-aware Low-rank Adaptation (HaLoRA) method. The key insight is to train a LoRA branch that is robust toward such noise and then deploy it on noise-free SRAM, while the extra cost is negligible since the parameters of LoRAs are much fewer than pretrained weights (e.g., 0.15% for LLaMA-3.2 1B model). To improve the robustness towards the noise, we theoretically analyze the gap between the optimization trajectories of the LoRA branch under both ideal and noisy conditions and further design an extra loss to minimize the upper bound of this gap. Therefore, we can enjoy both energy efficiency and accuracy during inference. Experiments finetuning the Qwen and LLaMA series demonstrate the effectiveness of HaLoRA across multiple reasoning tasks, achieving up to 22.7 improvement in average score while maintaining robustness at various noise types and noise levels.
Taiqiang Wu, Chenchen Ding, Wenyong Zhou, Yuxin Cheng, Xincheng Feng, Wendong Xu, Chufan Shi, Zhengwu Liu, Ngai Wong 0001
ACM Trans. Design Autom. Electr. Syst.4
2025 Tightening Robustness Verification of MaxPool-based Neural Networks via Minimizing the Over-Approximation Zone
abstract
The robustness of neural network classifiers is important in the safety-critical domain and can be quantified by robustness verification. At present, efficient and scalable verification techniques are always sound but incomplete, and thus, the improvement of verified robustness results is the key criterion to evaluate the performance of incomplete verification approaches. The multi-variate function MaxPool is widely adopted yet challenging to verify. In this paper, we present Ti-Lin, a robustness verifier for MaxPool-based CNNs with Tight Linear Approximation. Following the sequel of minimizing the over-approximation zone of the nonlinear function of CNNs, we are the first to propose the provably neuron-wise tightest linear bounds for the MaxPool function. By our proposed linear bounds, we can certify larger robustness results for CNNs. We evaluate the effectiveness of Ti-Lin on different verification frameworks with open-sourced benchmarks, including LeNet, PointNet, and networks trained on the MNIST, CIFAR-10, Tiny ImageNet and ModelNet40 datasets. Experimental results show that Ti-Lin significantly outperforms the state-of-the-art methods across all networks with up to 78.6% improvement in terms of the certified accuracy with almost the same time consumption as the fastest tool. Our code is available at https://github.com/xiaoyuanpigo/Ti-Lin-Hybrid-Lin.
Yuan Xiao 0003, Shiqing Ma, Chunrong Fang, Tongtong Bai, Mingzheng Gu, Yuxin Cheng, Zhenyu Chen 0001
CVPR7
2025 Enhancing Robustness of Implicit Neural Representations Against Weight Perturbations
abstract
Implicit Neural Representations (INRs) encode discrete signals in a continuous manner using neural networks, demonstrating significant value across various multimedia applications. However, the vulnerability of INRs presents a critical challenge for their real-world deployments, as the network weights might be subjected to unavoidable perturbations. In this work, we investigate the robustness of INRs for the first time and find that even minor perturbations can lead to substantial performance degradation in the quality of signal reconstruction. To mitigate this issue, we formulate the robustness problem in INRs by minimizing the difference between loss with and without weight perturbations. Furthermore, we derive a novel robust loss function to regulate the gradient of the reconstruction loss with respect to weights, thereby enhancing the robustness. Extensive experiments on reconstruction tasks across multiple modalities demonstrate that our method achieves up to a 7.5 dB improvement in peak signal-to-noise ratio (PSNR) values compared to original INRs under noisy conditions.
Wenyong Zhou, Yuxin Cheng, Zhengwu Liu, Taiqiang Wu, Ngai Wong 0001
ICASSP2
2025 MINR: Efficient Implicit Neural Representations for Multi-Image Encoding
abstract
Implicit Neural Representations (INRs) aim to parameterize discrete signals through implicit continuous functions. However, formulating each image with a separate neural network (typically, a Multi-Layer Perceptron (MLP)) leads to computational and storage inefficiencies when encoding multi-images. To address this issue, we propose MINR, sharing specific layers to encode multi-image efficiently. We first compare the layer-wise weight distributions for several trained INRs and find that corresponding intermediate layers follow highly similar distribution patterns. Motivated by this, we share these intermediate layers across multiple images while preserving the input and output layers as input-specific. In addition, we design an extra novel projection layer for each image to capture its unique features. Experimental results on image reconstruction and super-resolution tasks demonstrate that MINR can save up to 60% parameters while maintaining comparable performance. Particularly, MINR scales effectively to handle 100 images, maintaining an average peak signal-to-noise ratio (PSNR) of 34 dB. Further analysis of various backbones proves the robustness of the proposed MINR.
Wenyong Zhou, Taiqiang Wu, Zhengwu Liu, Yuxin Cheng, Ngai Wong 0001
ICASSP4
2025 Perspective-Aware 3D Gaussian Inpainting with Multi-View Consistency
abstract
3D Gaussian inpainting, a critical technique for numerous applications in virtual reality and multimedia, has made significant progress with pretrained diffusion models. However, ensuring multi-view consistency, an essential requirement for high-quality inpainting, remains a key challenge. In this work, we present PAInpainter, a novel approach designed to advance 3D Gaussian inpainting by leveraging perspective-aware content propagation and consistency verification across multi-view inpainted images. Our method iteratively refines inpainting and optimizes the 3D Gaussian representation with multiple views adaptively sampled from a perspective graph. By propagating inpainted images as prior information and verifying consistency across neighboring views, PAInpainter substantially enhances global consistency and texture fidelity in restored 3D scenes. Extensive experiments demonstrate the superiority of PAInpainter over existing methods. Our approach achieves superior 3D inpainting quality, with PSNR scores of 26.03 dB and 29.51 dB on the SPIn-NeRF and NeRFiller datasets, respectively, highlighting its effectiveness and generalization capability.
Yuxin Cheng, Binxiao Huang, Taiqiang Wu, Wenyong Zhou, Chenchen Ding, Zhengwu Liu, Graziano Chesi, Ngai Wong 0001
ICCV1
2025 Distribution-Aware Hadamard Quantization for Hardware-Efficient Implicit Neural Representations
abstract
Implicit Neural Representations (INRs) encode discrete signals using Multi-Layer Perceptrons (MLPs) with complex activation functions. While INRs achieve superior performance, they depend on full-precision number representation for accurate computation, resulting in significant hardware overhead. Previous INR quantization approaches have primarily focused on weight quantization, offering only limited hardware savings due to the lack of activation quantization. To fully exploit the hardware benefits of quantization, we propose DHQ, a novel distribution-aware Hadamard quantization scheme that targets both weights and activations in INRs. Our analysis shows that the weights in the first and last layers have distributions distinct from those in the intermediate layers, while the activations in the last layer differ significantly from those in the preceding layers. Instead of customizing quantizers individually, we utilize the Hadamard transformation to standardize these diverse distributions into a unified bell-shaped form, supported by both empirical evidence and theoretical analysis, before applying a standard quantizer. To demonstrate the practical advantages of our approach, we present an FPGA implementation of DHQ that highlights its hardware efficiency. Experiments on diverse image reconstruction tasks show that DHQ outperforms previous quantization methods, reducing latency by 32.7%, energy consumption by 40.1%, and resource utilization by up to 98.3% compared to full-precision counterparts.
Wenyong Zhou, Jiachen Ren, Taiqiang Wu, Yuxin Cheng, Zhengwu Liu, Ngai Wong 0001
ICME4
2025 Embedding-Based Sparse Retrieval in E-Commerce Search
Jiyuan He, Yunchuan Lin, Yuxin Cheng
ICONIP (1)5
2025 Hybrid Mesh-Gaussian Representation for Efficient Indoor Scene Reconstruction
abstract
3D Gaussian splatting (3DGS) has demonstrated exceptional performance in image-based 3D reconstruction and real-time rendering. However, regions with complex textures require numerous Gaussians to capture significant color variations accurately, leading to inefficiencies in rendering speed. To address this challenge, we introduce a hybrid representation for indoor scenes that combines 3DGS with textured meshes. Our approach uses textured meshes to handle texture-rich flat areas, while retaining Gaussians to model intricate geometries. The proposed method begins by pruning and refining the extracted mesh to eliminate geometrically complex regions. We then employ a joint optimization for 3DGS and mesh, incorporating a warm-up strategy and transmittance-aware supervision to balance their contributions seamlessly.Extensive experiments demonstrate that the hybrid representation maintains comparable rendering quality and achieves superior frames per second FPS with fewer Gaussian primitives.
Binxiao Huang, Zhihao Li 0002, Shiyong Liu, Jiajun Tang 0001, Yuxin Cheng, Ngai Wong 0001
IJCAI7
2025 Re-Activating Frozen Primitives for 3D Gaussian Splatting
Yuxin Cheng, Binxiao Huang, Wenyong Zhou, Taiqiang Wu, Zhengwu Liu, Graziano Chesi, Ngai Wong 0001
ACM Multimedia1
2025 AW-GBGAE: An Adaptive Weighted Graph Autoencoder Based on Granular-Balls for General Data Clustering
abstract
In the current scenario, a vast amount of unlabeled high-dimensional data exhibits intrinsic relationships, making it suitable for information extraction through graph-based clustering methods. However, these datasets often lack edge structure information and contain numerous irrelevant features. To address these challenges, we propose a comprehensive solution that involves: (1) applying a feature weighting approach to manage features, (2) constructing edges based on weighted granular-balls, and (3) integrating graph convolutional networks (GCNs) with edge generation to develop an autoencoder network. Our method significantly enhances the extraction of relevant information from high-dimensional, unlabeled data, improving the overall performance and reliability of the clustering process. Extensive experimental results demonstrate that our model, AW-GBGAE, excels in clustering tasks and exhibits strong competitiveness compared to baseline models. The code is publicly available at https://github.com/xjnine/AWGBGAE.
Jiang Xie 0002, Yuxin Cheng, Shuyin Xia, Chunfeng Hua, Guoyin Wang 0001, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 W-GBC: An Adaptive Weighted Clustering Method Based on Granular-Ball Structure
abstract
Existing weighted clustering algorithms often heavily rely on specific parameters. Specifically, in addition to the number of clusters (k), several other parameters need to be manually tuned, which greatly limits their practical applicability. The fundamental issue lies in the fact that most weighted clustering methods derive feature weights through global iterations. To address this challenge, this paper introduces a novel weighted granular-ball structure, continually optimizing weights during the ball splitting process and restricting the calculation of local data point weights to the corresponding weighted granular-ball. We employ local iterations within this structure as an approximation to global weight calculations. This method eliminates the need for parameter tuning during the weight calculation process and incidentally addresses the “curse of dimensionality” in traditional granular-ball computing model. When applied to complex real-world datasets, this method accurately represents high-dimensional data, thereby improving clustering precision and extending the adaptability of the granular-ball computing model in high-dimensional spaces. Comprehensive experimental analysis demonstrates that our W-GBC algorithm performs well in terms of clustering results and competes strongly with baseline algorithms. The code has been released and is now available at https://github.com/xjnine/W-GBC.
Jiang Xie 0002, Chunfeng Hua, Shuyin Xia, Yuxin Cheng, Guoyin Wang 0001, Xinbo Gao 0001
ICDE4
2023 Using IRS to Improve the Secrecy Rate of Millimeter Wave Communication System
abstract
With the development of 6G, millimeter wave communication has received extensive attention. Due to the characteristics of wireless transmission, information secrecy transmission is facing significant challenges. This paper uses the physical layer security (PLS) to explore the information secrecy transmission. Specifically, we use Intelligent Reflecting Surface (IRS) to control the wireless propagation environment and improve the secrecy rate of millimeter wave communication. The active beamforming matrix of the base station and the passive beamforming matrix of the IRS are optimized to achieve the maximum secrecy rate. We deduce the closed form solution of active beamforming and the approximate optimal solution of passive beamforming. An alternating optimization (AO) algorithm is applied to solve the non-convex optimization problem. Simulation results verify the convergence and effectiveness of the algorithm, which can obtain nearly twice the secrecy rate gain of the benchmark algorithm.
Kunpeng Song, Fangshu Ma, Zexian Chen, Yong Shang, Yuxin Cheng
VTC2023-Spring6
2023 Algorithm for diabetic retinal image analysis based on deep learning
Liwei Deng 0002, Yuxin Cheng, Guofu Zhao, Jiazhong Xu
Multim. Tools Appl.3
2022 A novel decomposition-based ensemble model for short-term load forecasting using hybrid artificial neural networks
Zhiyuan Liao, Jiehui Huang, Yuxin Cheng, Chunquan Li 0001, Peter Xiaoping Liu
Appl. Intell.3
2021 Physical Layer Security of Untrusted UAV-enabled Relaying NOMA Network Using SWIPT and the Cooperative Jamming
abstract
Unmanned aerial vehicles (UAVs) have been widely used in wireless communication network for its high flexibility and broad coverage. However, due to the wireless broadcast and high power consumption features of UAV networks, physical layer secrecy (PLS), spectrum efficiency and energy supply are of major importance. In this paper, we propose a novel untrusted simultaneous wireless information and power transfer (SWIPT) UAV relaying network in which cooperative jamming is utilized to enhance secrecy performance. We take advantage of nonorthogonal multiple access (NOMA) technique to improve the spectrum efficiency. For energy supply, the UAV relay harvests energy from source and jammer node using power splitting (PS) scheme as well as amplifies and forwards (AF) the signal. The closed expressions of connection outage probability (COP) and secrecy outage probability (SOP) are derived over Nakagami-m fading channel. Simulation results demonstrate that the communication network security has been improved compared with orthogonal frequency division multiple access (OFDMA) network. Finally, minimum COP and SOP value can be achieved with appropriate parameters.
Fangshu Ma, Yong Shang, Yuxin Cheng
VTC Fall4
2021 Modeling and Analyzing LTE Licensed Assisted Access Network with Capture Effect
abstract
The coexistence performance of LTE-licensed assisted access (LAA) and WiFi networks has been extensively investigated. However, these works ignore capture effect, which is the phenomenon that the strongest signal may still be successfully received when more than two signals are transmitted simultaneously on the same channel, and which may occur more frequently in the coexistence scenario than in the pure WiFi network. This may lead to very large deviation in the coexistence performance evaluation. In the paper, we deeply investigate the coexistence performance of LAA and WiFi networks with the capture effect. More specifically, a capture model for more than two signals is first proposed in the coexistence scenario, and the capture probability is derived. Then the LAA access schemes are modeled as a new two-dimensional discrete Markov model integrating with the capture effect. A large number of simulation and numerical results verify the validity of the proposed Markov chain and capture model. The results also show that the capture effect can not only significantly decrease the collision probability but also increase LAA and WiFi throughput as well as total throughput. All these results prove the necessity of considering the capture effect in coexistence performance evaluation.
Errong Pei, Lineng Zhou, Bingguang Deng, Yuxin Cheng, Yun Li 0001
VTC Spring4
2016 Centralized Control Plane for Passive Optical Top-of-Rack Interconnects in Data Centers
abstract
To efficiently handle the fast growing traffic inside data centers, several optical interconnect architectures have been recently proposed. However, most of them are targeting the aggregation and core tiers of the data center network, while relying on conventional electronic top-of-rack (ToR) switches to connect the servers inside the rack. The electronic ToR switches pose serious limitations on the data center network in terms of high cost and power consumption. To address this problem, we recently proposed a passive optical top-of-rack interconnect architecture, where we focused on the data plane design utilizing simple passive optical components to interconnect the servers within the rack. However, an appropriate control plane tailored for this architecture is needed to be able to analyze the network performance, e.g., packet delay, drop rate, etc., and also obtain a holistic network design for our passive optical top-of-rack interconnect, which we refer to as POTORI. To fill in this gap, this paper proposes the POTORI control plane design which relies on a centralized rack controller to manage the communications inside the rack. To achieve high network performance in POTORI, we also propose a centralized medium access control (MAC) protocol and two dynamic bandwidth allocation (DBA) algorithms, namely Largest First (LF)and Largest First with Void Filling (LFVF). Simulation results show that POTORI achieves packet delays in the order of microseconds and negligible packet loss probability under realistic data center traffic scenarios.
Yuxin Cheng, Matteo Fiorani, Lena Wosinska, Jiajia Chen 0001
GLOBECOM1
2013 Partial Noise Value Aided Reduced K-Best Sphere Decoding
abstract
This article focuses on reducing the complexity of K-best sphere decoding (K-best SD) algorithm for the detection of multiple-input multiple-output (MIMO) systems. One common reduction method is that one or more selected thresholds are set to cut excess nodes with partial Euclidean Distance (PED) larger than them. For a long time, statistical characteristic of noise has been well explored to generate thresholds. But the known noise in a certain specific transmission process is always overlooked. In this article, not only the statistical characteristic of noise is calculated, but also the known value of noise is considered. By adding a parameter determined by both noise and quality of service (QoS) to the smallest PED in each searching layer, a tighter and more suitable threshold can be calculated for this layer. Simulation results show that the proposed algorithm makes an efficient complexity reduction while the performance drops little. Specially, the proposed algorithm reduces the computational complexity about 90\% while the bit error ratio (BER) performance drops around 10\% in 4-by-4 MIMO systems employing 16-QAM or 64-QAM modulation. A new parameter, half complexity point, is proposed to evaluate the reduction effect, and half complexity points of the proposed algorithm are better than one selected original algorithm.
Yuxin Cheng, Haige Xiang
VTC Fall2
2013 Two Block Partitioned Dijkstra Algorithms
abstract
The Dijkstra algorithm (DA) is a kind of tree search algorithm. The biggest advantage is that it has the smallest number of visited nodes among all optimal tree search algorithms. But stack sizes required by the DA are always too large to achieve. By partitioning the searching tree into blocks, two modified algorithms are proposed in this article to shrink stack sizes. One, serial block partitioned DA, searches blocks one by one. Another, parallel block partitioned DA, searches blocks at the same time. Radii, which are updated when one block search is finished, are set to cut nodes with metrics larger than them in both algorithms. Simulation results show that the visited nodes number of serial block partitioned DA increase is very limited while the stack size is reduced exponentially. It also shows that stack sizes of the parallel block partitioned DA are reduced exponentially and the processing time is reduced efficiently. The performance of proposed algorithms is kept optimal in both proposed algorithms.
Yuxin Cheng, Haige Xiang
VTC Fall2
2012 Step Reduced K-Best Sphere Decoding
abstract
We propose an algorithm that reduces the complexity of the K-best sphere decoding (K-best SD) algorithm, which is a powerful parallel detection algorithm for multiple-input multiple-output systems (MIMO). By analyzing the probability of different nodes to be the final solution, the algorithm prunes some nodes during the tree search to reduce the complexity. Simulation results prove that compared with the K-best SD algorithm the proposed algorithm performance drops very little. Compared with the famous fixed-complexity sphere decoding (FSD) with the same complexity, the proposed algorithm has better performance.
Yuxin Cheng, Haige Xiang
VTC Fall2
2009 Improved turbo equalization based on soft ISI cancellation
Yong Shang, Yuxin Cheng, Haige Xiang
Signal Process.3