Xiaojun Mo

dblp:419/4069 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › speech coding
neural speech codec
1.012026
SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs · AAAI 2026
Audio and music processing
speech coding
1.012026
SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs · AAAI 2026
Audio and music processing › audio representation learning
speech representation
1.012026
SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs · AAAI 2026
Audio and music processing
speech synthesis
0.312026
SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs · AAAI 2026

Methods — techniques the papers use, named apart from their topics

semantic anchoring · 1.0residual vector quantization · 1.0mHuBERT · 1.0asymmetric dual quantization · 1.0
YearPublicationVenuePosition
2026 SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
abstract
Neural Speech Codecs face a fundamental trade-off at low bitrates: preserving acoustic fidelity often compromises semantic richness. To address this, we introduce SACodec, a novel codec built upon an asymmetric dual-quantizer that employs our proposed Semantic Anchoring mechanism. This design strategically decouples the quantization of Semantic and Acoustic details. The semantic anchoring is achieved via a lightweight projector that aligns acoustic features with a frozen, large-scale mHuBERT codebook, injecting linguistic priors while guaranteeing full codebook utilization. Sequentially, for acoustic details, a residual activation module with SimVQ enables a single-layer quantizer (acoustic path) to faithfully recover fine-grained information. At just 1.5 kbps, SACodec establishes a new state of the art by excelling in both fidelity and semantics: subjective listening tests confirm that its reconstruction quality is perceptually highly comparable to ground-truth audio, while its tokens demonstrate substantially improved semantic richness in downstream tasks. This work suggests that assigning specialized semantic quantizers to distinct information streams offers an effective path to reconcile the long-standing trade-off between fidelity, semantics, and modeling simplicity in low-bitrate speech tokenization.
Zhongren Dong, Jing Han 0010, Xiaojun Mo, Yimin Cao, Zixing Zhang 0001
AAAI5
2025 PSFD: Proactive Spatial-Frequency Defense against Malicious Exemplar-Guided Image Editing
abstract
Diffusion models has threatened image authenticity by enabling highly realistic fakes. Proactive defense offers protection by adding a "protective layer" that resists such manipulation. However, current proactive defenses mainly focus on text-guided editing but are less effective for the more challenging exemplar-guided tasks. To bridge this gap, we propose Proactive Spatial-Frequency Defense (PSFD), a novel proactive defense for exemplar-guided image editing. PSFD leverages adversarial attack to add subtle perturbations to make images immune to editing. We apply protections in both the frequency and spatial domains. Spatial perturbation disrupts feature extraction by forcing visual encoders to map the image to "bad" representations. Frequency perturbation tweaks the high-frequency components to distort the image’s texture information. We design two optimization strategies: PSFD-U that aims to generate maximal variation and PSFD-T that seeks to achieve specific editing styles. Extensive experiments on MS-COCO and ImageNet demonstrate PSFD’s strong defensive capabilities and transferability.
Xiaojun Mo, Meng Xie, Hangtao Zhang, Yixiang Liu, Yezhuo Peng, Yanchun Li
ICME2