Minxi Yang

dblp:284/3171 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-0951-9672ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Variable Rate Image Compression Guided by Cross-Image Semantics in Visual in-Context
abstract
Existing image compression models often lack personalization capabilities, treating all image regions equally and failing to meet the compression needs of different users for specific Regions of Interest (ROI). To address this challenge, we propose an innovative variable rate image compression framework that achieves user-centric dynamic compression by introducing visual in-context learning. Our method extracts cross-image semantics from user-provided visual examples to understand their intent. This semantic information is then converted into visual semantic query tokens and spatial masks to effectively guide the bit allocation of the compression model. Furthermore, we design a novel Semantic Spatial Control Block (SSCB) to fully leverage these semantic and spatial cues, thereby achieving a balance between preserving user-specified details and overall image quality. Experimental results demonstrate that our method significantly improves performance on the ROI, achieving a 31.54 % BD-Rate reduction and a 2.7479 dB BD-PSNR gain over the baseline model.
Dahua Gao, Minxi Yang
DCC3
2026 Poster: DiSC2: Bandwidth-Efficient Distributed Semantic Communication for Cross-Modal Redundancy Suppression in the Internet of Bodies
Dahua Gao, Minxi Yang, Guangming Shi
SECON3
2026 AdaDCL: Adaptive noise enhanced diffusion contrastive learning for recommendation
Renjie Tian, Minxi Yang, Feng Xie 0009, Dahua Gao, Zhenyuan Lin
Neurocomputing2
2026 Decoupled visual concept token learning via two-stage hierarchical optimization
Yibo Zhang 0013, Minxi Yang, Dahua Gao, Feng Xie 0009, Ruichao Liu
Neurocomputing2
2026 GeoSIC: Geometric-aware stereo image compression via implicit 3D feature fields
Minxi Yang, Dahua Gao
Pattern Recognit.3
2025 Stable Control Visual AutoRegressive Model: Precise and Efficient Image Generation via Scale Alignment
abstract
Although diffusion models advance condition-based visual generation, they suffer from speed and cost issues, unlike faster AutoRegressive methods that are limited in performance. To address these, we introduce the Stable Control Visual AutoRegressive Model (SCVAR). SCVAR ensures stable control by aligning visual conditions on multiple scales. Rather than unfolding the 2D image into a 1D raster, SCVAR decouples it into multiple scales. This shifts the sequential representation in SCVAR from tokens to scales, satisfying the unidirectional dependency of the AR model while preserving the 2D structure of the image. Compared to indiscriminate conditional guidance, cross-scale alignment provides more precise constraints, enabling SCVAR to achieve state-of-the-art performance in experiments against diffusion models, with 10x faster generation speed. The decoupled condition also reduces training costs. Compared to end-to-end conditional computation, experiments demonstrate that SCVAR matches performance with only 40% additional parameters.
Feng Xie 0009, Dahua Gao, Ruichao Liu, Minxi Yang, Yibo Zhang 0013
ICASSP4
2025 FreeAlign: Superior Text-Image Alignment by Modulating Prompt Attention
abstract
In recent years, Text-to-Image (T2I) models have made remarkable advancements, yet accurate accurate association of attributes remains a key challenge. This paper presents FreeAlign, a novel training-free framework designed to enhance attribute alignment in T2I generation. By modulating attention and adapting U-Net components, FreeAlign achieves precise alignment between image attributes and textual descriptions. It strengthens attribute-target associations through refining attention maps, adjusts U-Net’s backbone and skip connections based on energy ratios, and reorders prompts to balance attribute focus. Large Language Models (LLMs) enrich prompts with diverse, contextually relevant text, enhancing diffusion models’ generative power and quality. Extensive experiments show that FreeAlign delivers superior alignment for diverse prompts while preserving intricate details and ensuring structural integrity, establishing a new benchmark for attribute precision in T2I generation.
Yibo Zhang 0013, Dahua Gao, Feng Xie 0009, Minxi Yang, Ruichao Liu
ICASSP4
2025 A Transmitter-Model Unaware Generative Image Compression Framework for Semantic Communication
abstract
Unlike traditional bit-level data transmission methods, semantic communication focuses on conveying the meaning behind the data. Though promising results have been achieved, existing end-to-end learning-based semantic communication frameworks often require a synchronization of deep models between the transmitter and the receiver. Such design leads to tens of thousands models to be stored at receiver since different manufactures may optimize their own models. To address this problem, we propose a novel model-unaware generative image compression framework for semantic communication. It features at employing human-understandable multi-modality representations as an intermediate layer to enhance information transmission efficiency and semantic consistency. Our framework introduces a mask-based rate-distortion optimization module, which effectively removes low-relevance information for image generation and reduces the bit rate while maintaining semantic consistency. Experimental results demonstrate that the framework can still reconstruct high-quality images at very low bit rates, showcasing its potential for applications in modern communication systems.
Rongcan Zheng, Xiaodan Song, Xuguang Zuo, Minxi Yang, Dahua Gao, Xuemei Xie
ICASSP4
2025 Semantic Communication Using Intent-guided Coarse- and Fine-grained Codec with Pre-trained Diffusion Models
abstract
In image semantic communication, the granularity of semantic descriptions required for different objects within an image varies based on the communication intent and the importance of the objects. However, current semantic codecs optimized for global assessment metrics fail to adapt to user intent and cannot provide differentiated semantic granularity for objects of different importance. Generative semantic codecs using representations such as edge maps or semantic segmentation maps are insufficient for capturing fine-grained semantic information. This paper proposes dividing the transmitted image semantics into global coarse-grained and key object fine-grained semantics to better align with sender intent and optimize bandwidth usage. We introduce a novel semantic codec scheme based on a pre-trained text-to-image diffusion model. Global coarse-grained semantics are represented using short textual descriptions. Fine-grained semantic information of key objects is extracted using the Denoising Diffusion Implicit Model (DDIM) inversion and compressed in the frequency domain. Experimental results demonstrate that the proposed semantic codec enables high-quality recovery of coarse- and fine-grained semantics in image transmission while significantly reducing data transmission requirements.
Dahua Gao, Minxi Yang
ICME3
2024 Mosic: Multimodal Semantic Integrated Communication for Health Monitoring in Iot Scenarios
abstract
Monitoring multimodal signals provides a more comprehensive understanding of health conditions compared to singlemode monitoring. In the face of the significant volumes of multimodal signals, existing IoT health monitoring systems primarily focus on high-fidelity signal transmission by encoding multimodal signals separately. However, due to the lack of consideration for the downstream applications and correlation between multimodal signals, a portion of bandwidth resources is wasted on task-irrelevant information and intermodal redundancy. To address this issue, we propose the Multimodal Semantic Integration Communication (MoSIC) framework composed of three levels: At the sensor level, multiple wearable sensors collect and send different modal signals to a mobile terminal; at the mobile terminal level, the terminal employs deep source-channel joint encoding for the received multimodal signals, extracting single-modal embedded features using a backbone network, and obtaining cross-modal features through a feature fusion network with contrastive constrain; at the cloud level, a decoding network symmetric to the encoding network reconstructs the multimodal signals, which are then used for downstream applications such as human activity recognition. MoSIC focuses on semantically integrating multimodal signals for downstream applications, resulting in improved encoding and transmission efficiency. It also reduces the radio-frequency power consumption and bandwidth requirements.
Minxi Yang, Dahua Gao, Xiaodan Song, Guangming Shi
ICASSP1
2024 SG2SC: A Generative Semantic Communication Framework for Scene Understanding-Oriented Image Transmission
abstract
In recent years, semantic communication based on deep learning for source-channel joint encoding has garnered significant attention. It utilizes network models trained end-to-end to represent signals as embedding vectors and has demonstrated superior performance compared to traditional methods. However, due to the significant disparity between embedding vectors and human language, it can be challenging to succinctly capture abstract semantics such as scenes. In this paper, we introduce the Scene Graph-based Generative Semantic Communication (SG2SC) framework, built upon structured semantics and conditional generative models for image transmission. SG2SC aims to faithfully convey abstract semantics like scenes. It begins by detecting object categories, spatial attributes, and inter-category relationships in the image, representing scene semantics in a graph structure. Subsequently, it employs graph neural networks for scene graph encoding, decoding, and transmission, and finally utilizes a conditional diffusion model for semantic decoding. Benefiting from its concise graph structure semantics, SG2SC outperforms traditional method, semantic communication based on deep joint source-channel coding, and segmentation-based generative semantic communication in terms of noise resistance and encoding efficiency.
Minxi Yang, Dahua Gao, Feng Xie 0009, Xiaodan Song, Guangming Shi
ICASSP1
2024 Superimposed Semantic Communication for IoT-Based Real-Time ECG Monitoring
abstract
Real-time electrocardiogram (ECG) monitoring and diagnosis through Internet of Things (IoT) are crucial for addressing the severity and timely treatment of cardiovascular diseases, enabling timely intervention and preventing life-threatening complications. However, current ECG monitoring research predominantly focuses on individual aspects such as signal compression, diagnostic analysis, or secure transmission, lacking joint optimization of various modules in IoT scenarios. To address this gap, this work proposes a novel framework based on superimposed semantic communication for real-time ECG monitoring in IoT. The framework comprises three hierarchical levels: the edge level for data collection and processing, the relay level for signal compression and coding, and the cloud level for data analysis and reconstruction. The proposed framework offers several unique advantages. By employing semantic encoding guided by ECG classification tasks, it selectively extracts crucial features within and between signals, improving compression ratio and adaptability to channel noise. The superimposed semantic encoding achieves content encryption without requiring any additional operations. Moreover, the framework utilizes lightweight anomaly detection neural networks, reducing edge device power consumption and conserving communication resources. Simulation and real experimental results demonstrate that the proposed method achieves real-time encoding and transmission of ECG signals with a compression ratio of 0.019 on the MIT-BIH dataset. Furthermore, it attains a heartbeat classification accuracy of 0.988 and a reconstruction error of 0.061.
Minxi Yang, Dahua Gao, Guangming Shi
IEEE J. Biomed. Health Informatics1
2023 Residual Degradation Learning Unfolding Framework with Mixing Priors Across Spectral and Spatial for Compressive Spectral Imaging
abstract
To acquire a snapshot spectral image, coded aperture snapshot spectral imaging (CASSI) is proposed. A core problem of the CASSI system is to recover the reliable and fine underlying 3D spectral cube from the 2D measurement. By alternately solving a data subproblem and a prior subproblem, deep unfolding methods achieve good performance. However, in the data subproblem, the used sensing matrix is ill-suited for the real degradation process due to the device errors caused by phase aberration, distortion; in the prior subproblem, it is important to design a suitable model to jointly exploit both spatial and spectral priors. In this paper, we propose a Residual Degradation Learning Unfolding Framework (RDLUF), which bridges the gap between the sensing matrix and the degradation process. Moreover, a MixS2Transformer is designed via mixing priors across spectral and spatial to strengthen the spectral-spatial representation capability. Finally, plugging the MixS2Transformer into the RDLUF leads to an end-to-end trainable neural network RDLUF-MixS2. Experimental results establish the superior performance of the proposed method over existing ones. Code is available: https://github.com/ShawnDong98/RDLUF_MixS2
Yubo Dong, Dahua Gao, Minxi Yang, Guangming Shi
CVPR5
2022 A novel intrinsically explainable model with semantic manifolds established via transformed priors
Guangming Shi, Minxi Yang, Dahua Gao
Knowl. Based Syst.2