Chenyang Gao

dblp:177/7959 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dual-perspective filter pruning via diversity and independence collaboration
Chenyang Gao, Qinglong Cao, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003
Pattern Recognit.1
2025 SemiVisBooster: Boosting Semi-Supervised Learning for Fine-Grained Classification through Pseudo-Label Semantic Guidance
Chenyang Gao, Ivan Marsic
ICCV3
2025 ASELMAR: Active and semi-supervised learning-based framework to reduce multi-labeling efforts for activity recognition
abstract
Manual annotation of unlabeled data for model training is expensive and time-consuming, especially for visual datasets requiring domain-specific experience for multi-labeling, such as video records generated in hospital settings. There is a need to build frameworks to reduce human labeling efforts while improving training performance. Semi-supervised learning is widely used to generate predictions for unlabeled samples in a partially labeled datasets. Active learning can be used with semi-supervised learning to annotate unlabeled samples to reduce the sampling bias due to the label predictions. We developed the aselmar framework based on active and semi-supervised learning techniques to reduce the time and effort associated with multi-labeling of unlabeled samples for activity recognition. aselmar (i) categorizes the predictions for unlabeled data based on the confidence level in predictions using fixed and adaptive threshold settings, (ii) applies a label verification procedure for the samples with the ambiguous prediction, and (iii) retrains the model iteratively using samples with their high-confidence predictions or manual annotations. We also designed a software tool to guide domain experts in verifying ambiguous predictions. We applied aselmar to recognize eight selected activities from our trauma resuscitation video dataset and evaluated their performance based on the label verification time and the mean ap score metric. The label verification required by aselmar was 12.1% of the manual annotation effort for the unlabeled video records. The improvement in the mean ap score was 5.7% for the first iteration and 8.3% for the second iteration with the fixed threshold-based method compared to the baseline model . The p-values were below 0.05 for the target activities. Using an adaptive-threshold method, aselmar achieved a decrease in ap score deviation, implying an improvement in model robustness. For a speech-based case study , the word error rate decreased by 6.2%, and the average transcription factor increased 2.6 times, supporting the broad applicability of ASELMAR in reducing labeling efforts from domain experts.
Aydin Saribudak, Sifan Yuan, Chenyang Gao, Waverly Gestrich-Thompson, Zachary P. Milestone, Randall S. Burd, Ivan Marsic
Comput. Vis. Image Underst.3
2025 Less leakage and more precise: Efficient wildcard keyword search over encrypted data
abstract
Wildcard searchable encryption allows the server to efficiently perform wildcard-based keyword searches over encrypted data while maintaining data privacy. A promising solution to achieve wildcard SSE is to extract the characteristics of the queried keyword and check the existence based on a membership test structure. However, existing schemes have false positives of character order, that is, the server cannot identify the order between the first and the last wildcard character. Besides, the schemes also suffer from characteristic matching pattern leakage due to the one-by-one membership testing. In this paper, we present the first efficient wildcard SSE scheme to eliminate the false positives of character order and characteristic matching pattern leakage. To this end, we design a novel characteristic extraction technique that enables the client to exact the characteristics of the queried keyword maintaining the order between the first and the last wildcard character. Then, we utilize the primitive of Symmetric Subset Predicate Encryption, which supports checking if one set is a subset of another in one shot to reduce the characteristic matching pattern leakage. Finally, by performing a formal security analysis and implementing the scheme on a real-world database, we demonstrate that the desired security properties are achieved with high performance.
Yunling Wang, Chenyang Gao, Yong Yu 0002
High Confid. Comput.2
2024 Post-Training Embedding Alignment for Decoupling Enrollment and Runtime Speaker Recognition Models
abstract
Automated speaker identification (SID) is a crucial step for the personalization of a wide range of speech-enabled services. Typical SID systems use a symmetric enrollment-verification framework with a single model to derive embeddings both offline for voice profiles extracted from enrollment utterances, and online from runtime utterances. Due to the distinct circumstances of enrollment and runtime, such as different computation and latency constraints, several applications would benefit from an asymmetric enrollment-verification framework that uses different models for enrollment and runtime embedding generation. To support this asymmetric SID where each of the two models can be updated independently, we propose using a lightweight neural network to map the embeddings from the two independent models to a shared speaker embedding space. Our results show that this approach significantly outperforms cosine scoring in a shared speaker logit space for models that were trained with a contrastive loss on large datasets with many speaker identities. This proposed Neural Embedding Speaker Space Alignment (NESSA) combined with an asymmetric update of only one of the models delivers at least 60% of the performance gain achieved by updating both models in the standard symmetric SID approach.
Chenyang Gao, Brecht Desplanques, Chelsea J.-T. Ju, Aman Chadha, Andreas Stolcke
ICASSP1
2024 Knowledge Mining of Scene Text for Referring Expression Comprehension
Chenyang Gao, Wenwen Yu, Xiang Bai
ICDAR (5)1
2024 Metal-Organic Framework (MOF) Based Film Bulk Acoustic Resonator (FBAR) Sensor for Volatile Organic Compounds (VOCs) Detection
abstract
Recently, the detection of volatile organic compounds (VOCs) has received extensive attention in the field of industrial pollutants detection. As a microgravity sensor based on microelectromechanical system (MEMS), film bulk acoustic resonator (FBAR) plays an important role in the detection of micro-mass and composition. Combined with metal-organic framework (MOF), a type of porous crystal material, FBAR has the ability to adsorb micro-VOC vapor with high sensitivity. Herein, we proposed a MIL-101(Cr) MOF-based FBAR sensors with high resonant frequency and Q factor for five VOCs detection (acetone, ethanol, isopropanol, acetonitrile, and methanol), realizing high sensitivity and stable sensing performance. The experimental data shows the highest sensitivity is 35.15 Hz/ppm for acetone with the limit of detection (LOD) of 322 ppm, which can provide support in ensuring industrial production safety and VOC real-time in-situ detection.
Chenyang Gao, Mengyao Fu, Shuyu Fan, Dibo Hou, Yunqi Cao
IECON1
2023 Self-Supervised Speech Representation Learning for Keyword-Spotting With Light-Weight Transformers
abstract
Self-supervised speech representation learning (S3RL) is revolutionizing the way we leverage the ever-growing availability of data. While S3RL related studies typically use large models, we employ light-weight networks to comply with tight memory of compute-constrained devices. We demonstrate the effectiveness of S3RL on a keyword-spotting (KS) problem by using transformers with 330k parameters and propose a mechanism to enhance utterance-wise distinction, which proves crucial for improving performance on classification tasks. On the Google speech commands v2 dataset, the proposed method applied to the Auto-Regressive Predictive Coding S3RL led to a 1.2% accuracy improvement compared to training from scratch. On an in-house KS dataset with four different keywords, it provided 6% to 23.7% relative false accept improvement at fixed false reject rate. We argue this demonstrates the applicability of S3RL approaches to light-weight models for KS and confirms S3RL is a powerful alternative to traditional supervised learning for resource-constrained applications.
Chenyang Gao, Francesco Calivá, Yuzong Liu
ICASSP1
2023 ICDAR 2023 Competition on Recognition of Multi-line Handwritten Mathematical Expressions
Chenyang Gao, Shiyu Yao, Jinfeng Bai, Xiang Bai, Cheng-Lin Liu 0001
ICDAR (2)1
2023 TextREC: A Dataset for Referring Expression Comprehension with Reading Comprehension
Chenyang Gao, Hao Wang 0207, Wenwen Yu, Xiang Bai
ICDAR (3)1
2023 A Brake Pair Misalignment Detection Scheme Based on A Battery-Free Electromagnetic-Based Gap Sensor
abstract
The well-functioning of disc-pad brake subsystems concerns the safe operation of both automobiles and trains. However, there is a lack of effective constant condition monitoring methods for the vital disc-pad brake pairs in such brake subsystems of driving vehicles. Therefore, we propose a brake pair misalignment detection scheme based on a battery-free electromagnetic-based gap sensor, to deal with several common brake pair misalignments including distance changes, translational deviations, and rotary deflections. The alternating-polarity magnet array in the proposed sensor on the brake disc generates a unique distribution of the effective magnetic flux density within a relatively moving microfabricated planar coil sheet along with the brake pad, which is sensitive to the varied gap if misalignments occur between the disc-pad brake pair. Thus, discriminative voltage performance is induced by the inductive planar coils in response to different misalignment types and degrees. Comprehensive finite element simulation analysis and experiments are conducted to elaborate the misalignment detection capability of the proposed sensor. The proposed scheme has the potential to enhance the self-perception of intelligent vehicle systems, without adding the power consumption burden due to the battery-free nature of the electromagnetic-based sensor.
Shuyu Fan, Haozhen Chi, Chenyang Gao, Wangdi Du, Dibo Hou, Yunqi Cao
IECON3
2023 Improving Label Assignments Learning by Dynamic Sample Dropout Combined with Layer-wise Optimization in Speech Separation
abstract
In supervised speech separation, permutation invariant training (PIT) is widely used to handle label ambiguity by selecting the best permutation to update the model. Despite its success, previous studies showed that PIT is plagued by excessive label assignment switching in adjacent epochs, impeding the model to learn better label assignments. To address this issue, we propose a novel training strategy, dynamic sample dropout (DSD), which considers previous best label assignments and evaluation metrics to exclude the samples that may negatively impact the learned label assignments during training. Additionally, we include layer-wise optimization (LO) to improve the performance by solving layer-decoupling. Our experiments showed that combining DSD and LO outperforms the baseline and solves excessive label assignment switching and layer-decoupling issues. The proposed DSD and LO approach is easy to implement, requires no extra training sets or steps, and shows generality to various speech separation tasks.
Chenyang Gao, Ivan Marsic
INTERSPEECH1
2023 Mining High-Quality Pseudoinstance Soft Labels for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection in remote sensing image (RSI) is still a challenge because of the lack of instance-level labels, and many existing methods have two problems. Firstly, most of the existing methods usually mine the pseudo ground truth (PGT) instances solely relying on proposal class scores (PCS). Actually, the reliability of PCS is not enough because of the bird’s eye view imaging and large-scale chaotic background of RSIs, and the instances with high PCS incline to cover the discriminative region rather than the whole object. Secondly, the existing methods assign a one-hot label to each instance, and the label of PGT instance is copied to its neighbor instances, which induces the misclassification problem to some extent. Actually, the probability that the neighbor instances contain the object with the same category is smaller than the PGT instance. For the first problem, the proposal quality score (PQS) is proposed for mining high-quality PGT instances, which contains PCS and dual-context projection score (DCPS). The DCPS is calculated through semantic segmentation, and is employed to measure the completeness that each proposal covers an object. For the second problem, a pseudo soft label assignment (PSLA) strategy is proposed to assign more precise soft label for each instance, where the soft label is determined by the spatial distance between each instance and its nearest PGT instance. The ablation study validates the effectiveness of the PQS and PSLA. The comprehensive comparisons with other WSOD methods on three popular benchmarks show the excellent performance of our method.
Xiaoliang Qian, Yu Huo 0001, Gong Cheng 0003, Chenyang Gao, Xiwen Yao, Wei Wang 0245
IEEE Trans. Geosci. Remote. Sens.4
2022 Cosmos Propagation Network: Deep learning model for point cloud completion
abstract
Point clouds measured by 3D scanning devices often have partially missing data due to the view positioning of the scanner. The missing data can reduce the performance of a point cloud in downstream tasks such as segmentation, location, and pose estimation. Consequently, 3D point cloud completion aims to predict the missing regions of incomplete objects for these fundamental 3D vision tasks. However, predicting the complete object can easily diminish the detail or structure of a measured region, which usually does not require repair. This study proposes a novel neural network architecture, Cosmos Propagation Network (CP-Net), for 3D point cloud completion. CP-Net extracts latent features in different scales from incomplete point clouds used as input. For point cloud generation, we propose a novel point expand method using a Mirror Expand module. Compared with existing methods, our Mirror Expand module introduces less information redundancy, which makes the distribution of points more reliable. CP-Net predicts the details of missing regions and maintains a clear general structure. The performance of CP-Net on several benchmarks was compared to that of current baseline methods. Compared to the existing methods, CP-Net showed the best performance for various metrics. Thus, CP-Net is expected to help address various problems related to 3D point cloud completion. Its source code is available at https://github.com/ark1234/CP-Net.
Fangzhou Lin, Yajun Xu, Chenyang Gao, Kazunori D. Yamada
Neurocomputing4
2022 An Effective Convolutional Neural Network for Visualized Understanding Transboundary Air Pollution Based on Himawari-8 Satellite Images
abstract
Air pollution is a societal and cross-boundary environmental problem that can be visualized using a satellite. Satellite imaging is not only useful to the home country but also to the neighboring countries. Moreover, monitoring the movement of air pollution can help susceptible people avoid acid rain and photochemical smog. Using advanced remote sensing (RS) images, substantial information can be obtained, which can produce numerous effective methods for visualizing air pollution. In this article, a novel method for extracting air pollution has been proposed; it applies various pipeline networks along with a focus area method to exploit the spectral aspect information. Afterward, three indices with numerous modified fully convolutional networks (FCNs) were extracted. Then, by employing a multivote module, visualized air pollution can be presented. In the conducted experiments, five-year Himawari-8 satellite images have been utilized in the North–East Asia area to validate the frameworks. Furthermore, the experimental result indicating that the given methods could effectively visualize air pollution. Source code and data sets are available athttps://github.com/ark1234/Himawari-8-based-visualized-understanding.
Fangzhou Lin, Chenyang Gao, Kazunori D. Yamada
IEEE Geosci. Remote. Sens. Lett.2
2022 Conditional Feature Learning Based Transformer for Text-Based Person Search
abstract
Text-based person search aims at retrieving the target person in an image gallery using a descriptive sentence of that person. The core of this task is to calculate a similarity score between the pedestrian image and description, which requires inferring the complex latent correspondence between image sub-regions and textual phrases at different scales. Transformer is an intuitive way to model the complex alignment by its self-attention mechanism. Most previous Transformer-based methods simply concatenate image region features and text features as input and learn a cross-modal representation in a brute force manner. Such weakly supervised learning approaches fail to explicitly build alignment between image region features and text features, causing an inferior feature distribution. In this paper, we present CFLT, Conditional Feature Learning based Transformer. It maps the sub-regions and phrases into a unified latent space and explicitly aligns them by constructing conditional embeddings where the feature of data from one modality is dynamically adjusted based on the data from the other modality. The output of our CFLT is a set of similarity scores for each sub-region or phrase rather than a cross-modal representation. Furthermore, we propose a simple and effective multi-modal re-ranking method named Re-ranking scheme by Visual Conditional Feature (RVCF). Benefit from the visual conditional feature and better feature distribution in our CFLT, the proposed RVCF achieves significant performance improvement. Experimental results show that our CFLT outperforms the state-of-the-art methods by 7.03% in terms of top-1 accuracy and 5.01% in terms of top-5 accuracy on the text-based person search dataset.
Chenyang Gao, Guanyu Cai, Xinyang Jiang, Feng Zheng 0001, Jun Zhang 0018, Yifei Gong, Fangzhou Lin, Xing Sun 0001, Xiang Bai
IEEE Trans. Image Process.1