Ziyang Gao

dblp:229/0743 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 From Personal to Clinical: Personalisation and Depersonalisation for Explainable Depression Detection
Ziyang Gao, Linhai Zhang, Yulan He 0001
DASFAA (3)1
2026 ARFNet: Attentional refocusing fusion network for real-time semantic segmentation of urban road scenes
Fenglei Ren, Ziyang Gao, Lu Yang 0007, Muyu Li
Signal Process. Image Commun.2
2025 Dual-branch cross-modal fusion with local-to-global learning for UAV object detection
abstract
Due to the significant differences between unmanned aerial vehicle (UAV) images and natural scene images in terms of lighting, scale, and viewing angle, existing multispectral detection techniques often fail to fully utilize the remote dependencies between global and local information, resulting in poor performance in complex UAV scenarios. In this paper, we propose a novel two-branch cross-modal fusion network that integrates a dual cross attention transformer fusion block (CTF) for global feature dependency and an adaptive mask convolution fusion block (MCF) for underlying feature extraction. This achieves a unified representation with both global and local receptive fields. Our local-global training strategy utilizes a shallow global fusion network and a deep local fusion network, which operate on the entire image while also focusing on detailed local features. Additionally, we integrate an asymptotic feature pyramid network that employs adaptive spatial fusion to refine features, enhancing the accuracy of small object detection in UAV scenes. Evaluating our work with the DroneVehicle dataset for vehicle detection using infrared and visible light, our network outperformed existing methods, improving [email protected] by 7.01% compared to CAL-Net. APs for YOLOv8m-based single-modal infrared detection and visible light detection increased by 3.1% and 11.6%, respectively.
Binyi Fang, Yixin Yang 0004, Jingjing Chang, Ziyang Gao, Haibao Chen
ASP-DAC4
2025 DuQTTA: Dual Quantized Tensor-Train Adaptation with Decoupling Magnitude-Direction for Efficient Fine-Tuning of LLMs
abstract
Recent parameter-efficient fine-tuning (PEFT) techniques have enabled large language models (LLMs) to be efficiently fine-tuned for specific tasks, while maintaining model performance with minimal additional trainable parameters. However, existing PEFT techniques continue to face challenges in balancing both accuracy and efficiency, especially when addressing scalability and the demands of lightweight deployment for LLMs. In this paper, we propose an efficient fine-tuning method of LLMs based on dual quantized Tensor-Train adaptation with decoupling magnitude-direction (DuQTTA). The proposed DuQTTA method employs Tensor-Train decomposition and dual-stage quantization to minimize model size and resource consumption. Additionally, it employs an adaptive optimization strategy and a decoupled update mechanism to improve model performance, thereby minimizing suboptimal outcomes and ensuring alignment with the full-parameter fine-tuning goals. Experimental results indicate that the proposed DuQTTA method outperforms existing PEFT methods, achieving up to a $65 \times$ compression rate compared to the LLaMA2-7B models, meanwhile delivering improvements of $4.44 \%, 3.14 \%$, and 0.97% over LoRA on LLaMA2-7B, LLaMA3-8B, and LLaMA2-13B, respectively. The proposed DuQTTA method is effective in compressing LLMs for deployment on resource-constrained edge devices.
Haoyan Dong, Haibao Chen, Jingjing Chang, Yixin Yang 0004, Ziyang Gao, Zhigang Ji, Runsheng Wang, Ru Huang 0001
DAC5
2025 MemATr: An Efficient and Lightweight Memory-augmented Transformer for Video Anomaly Detection
abstract
Anomaly detection in videos is a long-standing and challenging problem. Previous methods often adopt deep and large neural networks to achieve the best detection accuracy; however, the high computational costs prevent them from being used in real-world applications with constrained computational resources. In this paper, we develop a mem ory- a ugmented tr ansformer named MemATr, which is capable of detecting video anomalies effectively. The proposed network is lightweight and can be easily deployed on mobile devices. Furthermore, we propose a memory transformer module to make predictions that are closer to normal inputs, thereby leading to a higher error for abnormal input patterns. Memory-attention is the main component of the proposed memory transformer, which can retrieve the features from learnable values rather than from the backbone like previous methods. Extensive experiments on the UCSD Ped2, CUHK Avenue, and ShanghaiTech benchmarks can demonstrate that our model has a significantly smaller model size while still achieving competitive detection accuracy. Our model has only 1/12 the number of parameters of the baseline model. Besides, our model achieves a 4.6% increase in accuracy on the ShanghaiTech dataset and has roughly the same accuracy compared with the baseline on the other two datasets. We validate the performance of the proposed model on the mobile device and the result shows it only has 49.8ms latency. The effectiveness of the proposed method on mobile devices is further supported by experimental results. A new quantitative parameter AMD (Applicability for Mobile Devices) is proposed to offer a novel approach to assist in making trade-offs for mobile devices. The proposed model obtains state-of-the-art results in terms of AMD.
Jingjing Chang, Peining Zhen, Xiaotao Yan, Yixin Yang 0004, Ziyang Gao, Haibao Chen
ACM Trans. Embed. Comput. Syst.5
2025 CaS2M: A Calibrated Single-to-Multiple Framework for Real-World Partial Fingerprint Recognition
abstract
With the reducing size of fingerprint collection modules in mobile devices, partial fingerprints are increasingly characterized by smaller overlapping areas and higher self-similarity. Existing methods either aggregate similarity scores from individual Single-to-Single recognition or directly employ a Single-to-Multiple network to verify the match between the query and templates. However, these methods either lack sufficient interaction between templates, or fail to provide adequate supervision for the alignment process, a crucial step in fingerprint recognition, thereby limiting overall accuracy. In this paper, we propose a novel partial fingerprint recognition strategy termed Calibrated Single-to-Multiple (CaS2M), which first calibrates template fingerprints individually, then combines them with the query fingerprint in a matcher network for feature fusion. Building upon this strategy, we develop a dual-stage framework tailored to real-world applications. During enrollment, a lightweight patch-based feature indexing algorithm and a template selection strategy are employed accounting for limited hardware resources. For authentication, independent calibration is first applied, followed by an attention-based matcher network to verify identity consistency. Experimental results on multiple public datasets (NIST 302, NIST SD4, SpoofGAN, FVC2002 DB1A & DB3A) and a self-build dataset demonstrate that our framework achieves superior performance over state-of-the-art algorithms, providing new insights for multi-template partial fingerprint recognition.
Ziyang Gao, Tianfan Peng, Jingjing Chang, Yixin Yang 0004, Haibao Chen
IEEE Trans. Inf. Forensics Secur.1
2024 Automated void detection in high resolution x-ray printed circuit boards (PCBs) images with deep segmentation neural network
Ho Yeung Ma, Minglu Xia, Ziyang Gao, Wenjing Ye
Eng. Appl. Artif. Intell.3
2023 Driver-Integrated Silicon Carbide Based Power Module with Self-Optimized Current-Sensorless Temperature-Driven Deadtime Control
abstract
Silicon carbide (SiC) MOSFETs half-bridge power modules are widely applicable in diversified advanced power electronics converters and systems due to higher breakdown voltage, operating temperature, and reduced dynamic parameters for improving power efficiency and power density concurrently. However, customize-designed external gate driver circuitries are required for commercially available$S$iC MOSFETs power modules resulting in higher parasitic parameters, leading to excess switching loss, overshoot and ringing on gate-source and drain-source voltages. Moreover, conventional complementary pulse width modulation signals deadtime control of half-bridge circuit relies on applying hall sensor based current sensing methodology, incurring extra power loss and possibly inaccuracy measurements. In this work, a gate driver integrated SiC half-bridge power module (1000 V/23 A), namely “Easy-SiC” power module, is proposed. The proposed solution is not only possessing gate driver integration to minimize gate driver circuitry parasitic parameters for switching performance optimization, but also introducing a novel self-optimized current-sensorless temperature-driven adaptive deadtime control for advanced power modules applications. Electrical and thermal simulation results revealed that the proposed solution could reduce the power module stray inductance and junction-to-case thermal resistance by 25% and 20% respectively, compared with commercially available product having similar specifications and form factor. Additionally, the proposed self-optimized temperature-driven deadtime control without current sensors is successfully demonstrated via preliminary experimental verifications.
Chun-Kit Cheung, Ziyang Gao
IECON2
2022 Deeply Tensor Compressed Transformers for End-to-End Object Detection
abstract
DEtection TRansformer (DETR) is a recently proposed method that streamlines the detection pipeline and achieves competitive results against two-stage detectors such as Faster-RCNN. The DETR models get rid of complex anchor generation and post-processing procedures thereby making the detection pipeline more intuitive. However, the numerous redundant parameters in transformers make the DETR models computation and storage intensive, which seriously hinder them to be deployed on the resources-constrained devices. In this paper, to obtain a compact end-to-end detection framework, we propose to deeply compress the transformers with low-rank tensor decomposition. The basic idea of the tensor-based compression is to represent the large-scale weight matrix in one network layer with a chain of low-order matrices. Furthermore, we propose a gated multi-head attention (GMHA) module to mitigate the accuracy drop of the tensor-compressed DETR models. In GMHA, each attention head has an independent gate to determine the passed attention value. The redundant attention information can be suppressed by adopting the normalized gates. Lastly, to obtain fully compressed DETR models, a low-bitwidth quantization technique is introduced for further reducing the model storage size. Based on the proposed methods, we can achieve significant parameter and model size reduction while maintaining high detection performance. We conduct extensive experiments on the COCO dataset to validate the effectiveness of our tensor-compressed (tensorized) DETR models. The experimental results show that we can attain 3.7 times full model compression with 482 times feed forward network (FFN) parameter reduction and only 0.6 points accuracy drop.
Peining Zhen, Ziyang Gao, Tianshu Hou, Haibao Chen
AAAI2
2022 Learning from Noisy Labels via Meta Credible Label Elicitation
abstract
Deep neural networks (DNNS) can easily overfit to noisy data, which leads to a significant degradation of performance. Previous efforts are primarily made by label correction or sample selection to alleviate supervision problem. To distinguish between noisy labels and clean labels, we propose a meta-learning framework which could gradually elicit credible labels via the meta-gradient descent step under the guidance of potentially non-noisy samples. Specifically, by exploiting the topological information of feature space, we can automatically estimate label confidence with a meta-learner. An iterative procedure is designed to select the most trustworthy noisy-labeled instances to generate pseudo labels. Then we train DNNs with pseudo supervision and original noisy super vision, which learns sufficiency and robustness properties in a joint learning objective. Experimental results on benchmark classification datasets show the superiority of our approach against the state-of-the-art methods.
Ziyang Gao, Yaping Yan, Xin Geng 0001
ICIP1