Mingyi Yang

dblp:222/5391 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 6 since 2021Computer networks · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Processing Network for Transcoding: Bridging Initial and Subsequent Encoding
abstract
Video transcoding is essential for multimedia processing as it enhances transmission efficiency, supports a variety of devices, and improves the user's experience. However, the output from the initial encoder is often unfriendly to subsequent transcoding. Existing transcoding optimization methods focus either concentrate on the initial encoding or the subsequent transcoding, neglecting the interplay between the two, even though both encoders significantly impact the overall transcoding process. In this work, we propose a processing network that bridges the initial encoding and subsequent transcoding, enabling a joint optimization of the transcoding process. For the initial encoder, considering the areas with lower residuals typically have smaller quantization losses, whereas areas with higher residuals do not, we employ residuals to guide the network in restoring compression distortion. In parallel, for the joint optimization of the subsequent encoder and the processing network, considering areas with large quantization losses typically indicate that the original region's distribution is either unsuitable for encoding or has complex textures, we have developed a corresponding mask in the DCT domain, and employ the quantified loss distribution from the subsequent encoder to fine-tune the loss training of the processing network. See Figure 1 for more details. Experiments show substantial enhancements in transcoding performance when transitioning from H.264 to H.265.
Miaojun Ni, Mingyi Yang
DCC2
2025 A Fast Fusion Algorithm for Cervical Cell Microscopic Images Based on the Gray-Scale Characteristics of Papanicolaou Staining
abstract
The early screening of cervical cancer relies on the accurate identification of cell nuclei in microscopic images, but the cells are distributed in different focal planes under high magnification, resulting in local blur in a single frame image. Traditional fusion algorithms are difficult to meet the real-time requirements due to high computational complexity, and deep learning-based methods are limited by hardware cost and data dependence. In this paper, according to the gray characteristics of Pap staining cervical nucleus (the pixel value in clear areas is lower), a lightweight multi-focus image fast fusion algorithm is proposed. The fusion image is directly synthesized by extracting the minimum gray value of the RGB channel of the multi-focus plane image pixel by pixel. In addition, a coarse-fine double-step search strategy is used to quickly locate the optimal focal plane, and a dynamic threshold mechanism is introduced to reduce redundant calculations. The experimental results demonstrate that the proposed algorithm achieves state-of-the-art multi-focus image fusion performance in cervical cytology applications.
Yinlong Zhang, Mingyi Yang
INDIN3
2025 Evolutionary neural architecture search for automatically designing CNN-GRU-Attention neural networks for turntable servo systems
Cheng Xie 0003, Bing Xue 0001, Mingyi Yang, Mengjie Zhang 0001, Yang Liu 0377
Expert Syst. Appl.3
2025 Broad Learning System Combined With Feature Weighted Stacked Target-Related Laplacian Autoencoder for Industrial Quality Prediction
abstract
In industrial processes, accurate quality prediction is essential for optimizing operations and ensuring product consistency. However, the inherent complexity of industrial processes typically involves a large number of monitoring variables, accompanied by significant challenges such as high nonlinearity and data redundancy. To address this issue, this paper proposes the Broad Learning System combined with a Feature Weighted Stacked Target-Related Laplacian Autoencoder (BLS-FW-STLapAE). The proposed model innovatively incorporates hidden layer features extracted by the Stacked Target-Related Laplacian Autoencoder (STLapAE) into the feature nodes of the Broad Learning System (BLS). This approach addresses the issue of insufficient feature extraction caused by random mapping in BLS. This integration enables the model to capture both local geometric structures and key target-related features, significantly enhancing predictive performance in complex industrial environments. Additionally, mutual information (MI) is introduced in STLapAE to evaluate the correlation between features and the target variable, further improving the model’s ability to capture relevant information about the target variable and ensuring that all relevant information is effectively utilized. Finally, the feasibility and practical value of the proposed algorithm were validated on benchmark public datasets and two real-world industrial applications: Penicillin Fermentation and Solid Propellant Manufacturing. The superior predictive performance demonstrated across these tests enables operators to proactively adjust process parameters and prevent quality deviations, thereby significantly improving production safety and reducing operational costs in such high-risk chemical engineering domains.
Mingyi Yang, Xuhang Chen 0004
IEEE Internet Things J.2
2024 Learning-Based Video Compression with Continuously Variable Bitrate Coding
abstract
In this paper, we propose a learning-based video compression which can perform continuously variable bitrate coding. The proposed method generates feature transformation parameters through a conditional network according to the input spatial quality map. These parameters are then used to adaptively transform the intermediate features of the encoder, decoder, and spatiotemporal entropy model in the codec, thus enabling variable bitrate coding. Additionally, to improve the compression efficiency of the codec, we propose incorporating the quality map of the preceding frame into the hyperprior encoder and leveraging the temporal prior encoder. A multi-stage training strategy is employed to jointly train the codec with a multi-frame rate-distortion loss function. The experimental results demonstrate that the proposed method can achieve continuously variable bitrate adaptation while maintaining rate-distortion performance comparable to the fixed bitrate model. Furthermore, the proposed method also supports ROI-based compression.
Mingyi Yang, Xionghui Mao, Yujie Yin, Defa Wang, Shuai Wan, Fuzheng Yang 0001
ICIP1
2024 Task-Switchable Pre-Processor for Image Compression for Multiple Machine Vision Tasks
abstract
Visual content is increasingly being processed by machines for various automated content analysis tasks instead of being consumed by humans. Despite the existence of several compression methods tailored for machine tasks, few consider real-world scenarios with multiple tasks. In this paper, we aim to address this gap by proposing a task-switchable pre-processor that optimizes input images specifically for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. The proposed task-switchable pre-processor adeptly maintains relevant semantic information based on the specific characteristics of different downstream tasks, while effectively suppressing irrelevant information to reduce bitrate. To enhance the processing of semantic information for diverse tasks, we leverage pre-extracted semantic features to modulate the pixel-to-pixel mapping within the pre-processor. By switching between different modulations, multiple tasks can be seamlessly incorporated into the system. Extensive experiments demonstrate the practicality and simplicity of our approach. It significantly reduces the number of parameters required for handling multiple tasks while still delivering impressive performance. Our method showcases the potential to achieve efficient and effective compression for machine vision tasks, supporting the evolving demands of real-world applications.
Mingyi Yang, Fei Yang 0004, Luka Murn, Marc Gorriz, Juil Sock, Shuai Wan, Fuzheng Yang 0001, Luis Herranz
IEEE Trans. Circuits Syst. Video Technol.1
2024 Joint Rate-Distortion Optimization for Video Coding and Learning-Based In-Loop Filtering
abstract
Learning-based in-loop filters (ILFs) have recently been widely deployed in the video codec to remove compression artifacts and to obtain better-quality reconstructed videos. However, in the existing codec, the impact of the learning-based ILF is not considered in the Rate-Distortion optimization (RDO) process. With the learning-based ILF, the set of coding parameters selected by the conventional RDO process may no longer be the best one, and the best overall Rate-Distortion (R-D) performance can not be guaranteed. In this article, we propose a joint RDO (JRDO) for Video Coding and learning-based in-loop filtering, which incorporates the effect of the learning-based ILF on the reconstructed video into the RDO process, aiming to achieve the best overall R-D performance of the reconstructed video after in-loop filtering. Furthermore, to realize the proposed JRDO in a standardized video codec, we propose practical strategies to efficiently estimate the effect of learning-based ILF during the RDO process, i.e., efficiently estimate the distortion of the reconstructed block after in-loop filtering during the RDO process. Extensive experiments demonstrate that the proposed joint RDO is standard-compliant and can improve the R-D performance without increasing the decoding time. Besides, the superiority of joint RDO is achieved in various ILFs, indicating the generality of the proposed work.
Mingyi Yang, Junyan Huo, Xile Zhou, Wenhan Qiao, Shuai Wan, Hao Wang 0184, Fuzheng Yang 0001
IEEE Trans. Multim.1
2023 Semantic Preprocessor for Image Compression for Machines
abstract
Visual content is being increasingly transmitted and consumed by machines rather than humans to perform automated content analysis tasks. In this paper, we propose an image preprocessor that optimizes the input image for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. To achieve a better trade-off between the accuracy of the machine analysis task and bitrate, we propose leveraging pre-extracted semantic information to improve the preprocessor’s ability to accurately identify and filter out task-irrelevant information. Furthermore, we propose a two-part loss function to optimize the preprocessor, consisted of a rate-task performance loss and a semantic distillation loss, which helps the reconstructed image obtain more information that contributes to the accuracy of the task. Experiments show that the proposed preprocessor can save up to 48.83% bitrate compared with the method without the preprocessor, and save up to 36.24% bitrate compared to existing preprocessors for machine vision.
Mingyi Yang, Luis Herranz, Fei Yang 0004, Luka Murn, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001, Marta Mrak
ICASSP1
2023 PIMnet: A quality enhancement network for compressed videos with prior information modulation
Mingyi Yang, Xile Zhou, Fuzheng Yang 0001, Mingcai Zhou, Hao Wang 0184
Signal Process. Image Commun.1
2022 abc4pwm: affinity based clustering for position weight matrices in applications of DNA sequence analysis
abstract
BACKGROUND: Transcription factor (TF) binding motifs are identified by high throughput sequencing technologies as means to capture Protein-DNA interactions. These motifs are often represented by consensus sequences in form of position weight matrices (PWMs). With ever-increasing pool of TF binding motifs from multiple sources, redundancy issues are difficult to avoid, especially when every source maintains its own database for collection. One solution can be to cluster biologically relevant or similar PWMs, whether coming from experimental detection or in silico predictions. However, there is a lack of efficient tools to cluster PWMs. Assessing quality of PWM clusters is yet another challenge. Therefore, new methods and tools are required to efficiently cluster PWMs and assess quality of clusters. RESULTS: A new Python package Affinity Based Clustering for Position Weight Matrices (abc4pwm) was developed. It efficiently clustered PWMs from multiple sources with or without using DNA-Binding Domain (DBD) information, generated a representative motif for each cluster, evaluated the clustering quality automatically, and filtered out incorrectly clustered PWMs. Additionally, it was able to update human DBD family database automatically, classified known human TF PWMs to the respective DBD family, and performed TF motif searching and motif discovery by a new ensemble learning approach. CONCLUSION: This work demonstrates applications of abc4pwm in the DNA sequence analysis for various high throughput sequencing data using ~ 1770 human TF PWMs. It recovered known TF motifs at gene promoters based on gene expression profiles (RNA-seq) and identified true TF binding targets for motifs predicted from ChIP-seq experiments. Abc4pwm is a useful tool for TF motif searching, clustering, quality assessment and integration in multiple types of sequence data analysis including RNA-seq, ChIP-seq and ATAC-seq.
Omer Ali, Amna Farooq, Mingyi Yang, Victor X. Jin, Magnar Bjørås, Junbai Wang
BMC Bioinform.3
2020 Towards the Instant Tile-Switching for Dash-Based Omnidirectional Video Streaming: Random Access Reference Frame
abstract
In this work, we propose an instant tile-switching mechanism for the tile-based omnidirectional DASH video streaming, which solves the problem that the client cannot switch to the high quality tiles immediately as the viewport changes. More specifically, we present the random access reference frame (RARF) on the server for each P frame as the random access point. When users change the viewport, the corresponding P frame at the viewport change moment will be replaced by RARF to transferred to the client for an immediately quality switching. By using the proposed instant tile-switching mechanism, the client can acquire the high quality video in the new viewport at any time and provide users a much better quality of experience.
Mingyi Yang, Wenjie Zou, Jiarun Song, Fuzheng Yang 0001
ICME1
2018 Spoofing Attack Detection Using Physical Layer Information in Cross-Technology Communication
abstract
Recent advances in Cross-Technology Communication (CTC) enable the coexistence and collaboration among heterogeneous wireless devices operating in the same ISM band (e.g., Wi-Fi, ZigBee, and Bluetooth in 2.4 GHz). However, state-of-the-art CTC schemes are vulnerable to spoofing attacks since there is no practice authentication mechanism yet. This paper proposes a scheme to enable the spoofing attack detection for CTC in heterogeneous wireless networks by using physical layer information. First, we propose a model to detect ZigBee packets and measure the corresponding Received Signal Strength (RSS) on Wi-Fi devices. Then, we design a collaborative mechanism between Wi-Fi and ZigBee devices to detect the spoofing attack. Finally, we implement and evaluate our methods through experiments on commercial off-the- shelf (COTS) Wi-Fi and ZigBee devices. Our results show that it is possible to measure the RSS of ZigBee packets on Wi-Fi device and detect spoofing attack with both a high detection rate and a low false positive rate in heterogeneous wireless networks.
Bingxian Lu, Zhenquan Qin, Mingyi Yang, Lei Wang 0005
SECON3