VLDB 2026 Research / reviewers in the wild / expert
Xiaojun Mo
dblp:419/4069
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › speech coding
neural speech codec |
1.0 | 1 | 2026 | SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs · AAAI 2026 |
Audio and music processing
speech coding |
1.0 | 1 | 2026 | SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs · AAAI 2026 |
Audio and music processing › audio representation learning
speech representation |
1.0 | 1 | 2026 | SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs · AAAI 2026 |
Audio and music processing
speech synthesis |
0.3 | 1 | 2026 | SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
semantic anchoring · 1.0residual vector quantization · 1.0mHuBERT · 1.0asymmetric dual quantization · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech CodecsabstractNeural Speech Codecs face a fundamental trade-off at low bitrates: preserving acoustic fidelity often compromises semantic richness. To address this, we introduce SACodec, a novel codec built upon an asymmetric dual-quantizer that employs our proposed Semantic Anchoring mechanism. This design strategically decouples the quantization of Semantic and Acoustic details. The semantic anchoring is achieved via a lightweight projector that aligns acoustic features with a frozen, large-scale mHuBERT codebook, injecting linguistic priors while guaranteeing full codebook utilization. Sequentially, for acoustic details, a residual activation module with SimVQ enables a single-layer quantizer (acoustic path) to faithfully recover fine-grained information. At just 1.5 kbps, SACodec establishes a new state of the art by excelling in both fidelity and semantics: subjective listening tests confirm that its reconstruction quality is perceptually highly comparable to ground-truth audio, while its tokens demonstrate substantially improved semantic richness in downstream tasks. This work suggests that assigning specialized semantic quantizers to distinct information streams offers an effective path to reconcile the long-standing trade-off between fidelity, semantics, and modeling simplicity in low-bitrate speech tokenization. Zhongren Dong, Jing Han 0010, Xiaojun Mo, Yimin Cao, Zixing Zhang 0001 |
AAAI | 5 |
| 2025 | PSFD: Proactive Spatial-Frequency Defense against Malicious Exemplar-Guided Image EditingabstractDiffusion models has threatened image authenticity by enabling highly realistic fakes. Proactive defense offers protection by adding a "protective layer" that resists such manipulation. However, current proactive defenses mainly focus on text-guided editing but are less effective for the more challenging exemplar-guided tasks. To bridge this gap, we propose Proactive Spatial-Frequency Defense (PSFD), a novel proactive defense for exemplar-guided image editing. PSFD leverages adversarial attack to add subtle perturbations to make images immune to editing. We apply protections in both the frequency and spatial domains. Spatial perturbation disrupts feature extraction by forcing visual encoders to map the image to "bad" representations. Frequency perturbation tweaks the high-frequency components to distort the image’s texture information. We design two optimization strategies: PSFD-U that aims to generate maximal variation and PSFD-T that seeks to achieve specific editing styles. Extensive experiments on MS-COCO and ImageNet demonstrate PSFD’s strong defensive capabilities and transferability. Xiaojun Mo, Meng Xie, Hangtao Zhang, Yixiang Liu, Yezhuo Peng, Yanchun Li |
ICME | 2 |