Xinyuan Song 0002

dblp:134/9104-2 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2026
0009-0005-4209-5671ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 17% Deep learning architectures and training · 17% Vision and language · 17%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 60% Image and video coding · 40%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.912025
TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning · NeurIPS 2025
Machine learning › Graph learning › limited supervision › multi-view semi-supervised learning
co-training
0.912025
TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model
0.912025
The Curse of Depth in Large Language Models · NeurIPS 2025
Machine learning › Deep learning architectures and training › normalization
layer normalization
0.912025
The Curse of Depth in Large Language Models · NeurIPS 2025
Machine learning › Representation and self-supervised learning
pre-training
0.912025
The Curse of Depth in Large Language Models · NeurIPS 2025
Computer vision › Vision and language › vision-language model
prompt learning
0.912025
Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training · ACM Multimedia 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning · NeurIPS 2025
Machine learning › Learning paradigms
semi-supervised learning
0.912025
TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
The Curse of Depth in Large Language Models · NeurIPS 2025
Computer vision › Vision and language
vision-language model
0.912025
Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training · ACM Multimedia 2025
Visual content generation and editing
image editing
0.912025
Free-Mask: A Novel Paradigm of Integration Between the Segmentation Diffusion Model and Image Editing · ACM Multimedia 2025
Image and video coding
image quality assessment
0.912025
Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training · ACM Multimedia 2025
Image and video coding › image quality assessment
no-reference image quality assessment
0.912025
Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training · ACM Multimedia 2025
Visual content generation and editing › image generation
text-to-image generation
0.912025
Twin Co-Adaptive Dialogue for Progressive Image Generation · ACM Multimedia 2025
Machine learning › Generative modeling
diffusion model
0.312025
Free-Mask: A Novel Paradigm of Integration Between the Segmentation Diffusion Model and Image Editing · ACM Multimedia 2025
Knowledge, reasoning and agents › Multi-agent systems
game theory
0.312025
TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation › multi-source learning
multi-dataset training
0.312025
Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training · ACM Multimedia 2025
Knowledge, reasoning and agents › Multi-agent systems › game theory
stackelberg game
0.312025
TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.312025
The Curse of Depth in Large Language Models · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

segmentation diffusion model · 1.7learnable prompt vectors · 1.7iterative refinement · 1.7dual weight adjustment · 1.7dialogue agent · 1.7conditional fusion · 1.7pseudo-labeling · 0.9mutual information · 0.9meta-learning · 0.9layer normalization scaling · 0.9
YearPublicationVenuePosition
2026 Semi-supervised multi-label feature selection with consistent sparse graph learning
Yan Zhong 0001, Xinping Zhao, Li Zhang 0104, Xinyuan Song 0002, Lei Shi 0030, Bingbing Jiang 0001
Neural Networks5
2026 Bridging NIP and MLM: A Unified Meta-Learning Framework for Sequential Recommendation
abstract
Sequential Recommender Systems (SRSs) predict items of interest for users based on their historical interactions. Two key popular paradigms for SRs are unidirectional Next Item Prediction (NIP) and bidirectional Masked Language Modeling (MLM). NIP performs well in recommendation tasks but is constrained by its reliance on prior information and a rigid temporal order assumption, limiting its ability to capture dynamic personalized preferences. MLM, on the other hand, enhances sequence representation and personalization by leveraging bidirectional context, but its objective is misaligned with recommendation tasks. To this end, we develop MetaSR, a meta-learning-based SR approach that ingeniously combines both paradigms into a unified framework to enhance personalized sequential recommendation. Specifically, we treat the mask prediction task in MLM as the inner-loop task (i.e., the support set), adjusting mask items dynamically to generate sequences that match individual behavioral patterns to capture personalized user preferences. In the outer loop (i.e., the query set), we adopt the NIP paradigm to directly train MetaSR toward the SR target. The joint training bridges the gap between MLM and SRs tasks, also overcoming NIP’s limitations in capturing dynamic personalized preferences. Furthermore, to improve the overall efficiency, we develop a reinforcement learning-based Adaptive Masked Sequence Selection (AMSS) mechanism, automatically selecting the optimal masking prediction tasks within the meta-learning support set to accelerate the meta-learning and achieve better personalization. Experiments on multiple public benchmark datasets show that MetaSR significantly outperforms existing SRs models. The dataset and codes are available at MetaSR .
Youhua Li, Ersheng Ni, Sibo Xu, Mingxuan Wu, Junchen Fu, Yuanqi He, Xinyuan Song 0002, Yongxin Ni
ACM Trans. Knowl. Discov. Data10
2025 Multi-Scale Volumetric Transformers with Adaptive Uncertainty Modeling for Robust Bacterial Flagellar Motor Localization in Cryo-Electron Tomography
abstract
Automated identification of bacterial flagellar motors in cryo-electron tomography (cryo-ET) remains a significant challenge due to low signal-to-noise ratios, imaging artifacts, and the complexity of capturing both local and global structural contexts. In this work, we propose MSVT+AUM, a hybrid framework that combines a Multi-Scale Volumetric Transformer (MSVT) with Adaptive Uncertainty Modeling (AUM) to address these limitations. MSVT employs a 3D ResNet-50 backbone to extract hierarchical volumetric features, which are integrated through a volume-level self-attention mechanism and refined via 3D Vision Transformer blocks with positional encoding. To enhance discriminative capability in densely packed molecular environments, we introduce a Contextual Spatial Attention module that adaptively reweights spatial and channel-wise representations. To reduce false positives and provide calibrated confidence estimates, AUM jointly models aleatoric and epistemic uncertainties using heteroscedastic regression and Monte Carlo dropout, with detection thresholds dynamically adjusted based on uncertainty estimates. Evaluation on the BYU flagellar motor dataset and an extended subset of the CryoET Data Portal shows that our method achieves an F 2 score of 0.891, representing a 1.5 % improvement over existing approaches, while reducing in-ference time by 23%. MSVT+AUM exhibits robust generalization across diverse bacterial species and noise conditions, offering a scalable and reliable solution for high-throughput structural analysis in cryo-ET.
Junqiao Wang, Yangfan He, Xinyuan Song 0002, Peilai Yu, Xunfei Zhu
BIBM4
2025 AMBER: Adaptive Meta Balanced Paradigm for Heterogeneous Graph-Based Knowledge Tracing
abstract
Knowledge Tracing (KT) is a fundamental task in personalized education, aiming to predict student performance by modeling their evolving concept mastery. Recent state-of-the-art approaches adopt multi-graph architectures to capture diverse concept and behavior relations. However, such models often suffer from graph imbalance, where one graph branch dominates training, undermining the benefits of structural integration. To address this, we propose AMBER (Adaptive Meta-Balanced Ensemble Representation learning), a KT framework designed to promote balanced learning across heterogeneous graph structures. AMBER introduces an external dual-graph teacher to guide the learning of ensemble representations. As the teacher itself may encode graph imbalance bias, we further incorporate a meta-distillation strategy that adaptively adjusts the teacher using student feedback, amplifying signals beneficial to underperforming branches. In addition, an adaptive graph rebalancing strategy is introduced to balance the optimization of different graph branches in real time, preventing dominance by any single structure. Experiments on three real-world datasets show that AMBER consistently outperforms competitive baselines. By promoting balanced optimization across graphs, AMBER enables more effective integration of heterogeneous learning signals in KT, providing a robust and scalable solution for personalized education. Code is available at https://github.com/AMBER2025KT/AMBER2025CIKM.
Lifan Sun, Zichen Yuan, Ersheng Ni, Weihua Cheng, Xinyuan Song 0002, Linkun Dai, Sibo Xu, Yucen Zhuang, Yongxin Ni, Youhua Li
CIKM5
2025 ContrastiveGaussian: High-Fidelity 3D Generation with Contrastive Learning and Gaussian Splatting
abstract
Creating 3D content from single-view images is a challenging problem that has attracted considerable attention in recent years. Current approaches typically utilize score distillation sampling (SDS) from pre-trained 2D diffusion models to generate multi-view 3D representations. Although some methods have made notable progress by balancing generation speed and model quality, their performance is often limited by the visual inconsistencies of the diffusion model outputs. In this work, we propose ContrastiveGaussian, which integrates contrastive learning into the generative process. By using a perceptual loss, we effectively differentiate between positive and negative samples, leveraging the visual inconsistencies to improve 3D generation quality. To further enhance sample differentiation and improve contrastive learning, we incorporate a super-resolution model and introduce another Quantity-Aware Triplet Loss to address varying sample distributions during training. Our experiments demonstrate that our approach achieves superior texture fidelity and improved geometric consistency. Code will be available at https://github.com/YaNLlan-Ljb/ContrastiveGaussian.
Junbang Liu, Enpei Huang, Dongxing Mao, Hui Zhang 0062, Xinyuan Song 0002, Yongxin Ni
ICME5
2025 Free-Mask: A Novel Paradigm of Integration Between the Segmentation Diffusion Model and Image Editing
Bo Gao 0004, Jianhui Wang 0001, Xinyuan Song 0002, Yangfan He, Fangxu Xing, Tianyu Shi 0003
ACM Multimedia3
2025 Twin Co-Adaptive Dialogue for Progressive Image Generation
abstract
Modern text-to-image generation systems have enabled the creation of remarkably realistic and high-quality visuals, yet they often falter when handling the inherent ambiguities in user prompts. In this work, we present Twin-Co, a framework that leverages synchronized, co-adaptive dialogue to progressively refine image generation. Instead of a static generation process, Twin-Co employs a dynamic, iterative workflow where an intelligent dialogue agent continuously interacts with the user. Initially, a base image is generated from the user's prompt. Then, through a series of synchronized dialogue exchanges, the system adapts and optimizes the image according to evolving user feedback. The co-adaptive process allows the system to progressively narrow down ambiguities and better align with user intent. Experiments demonstrate that Twin-Co not only enhances user experience by reducing trial-and-error iterations but also improves the quality of the generated images, streamlining creative process across various applications.
Jianhui Wang 0001, Yangfan He, Yan Zhong 0001, Xinyuan Song 0002, Jiayi Su, Yuheng Feng, Hongyang He, Wenyu Zhu, Xinhang Yuan, Miao Zhang 0010, Tianyu Shi 0003, Xueqian Wang 0001
ACM Multimedia4
2025 Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training
abstract
Due to the high cost and small scale of Image Quality Assessment (IQA) datasets, achieving robust generalization remains challenging for prevalent Blind IQA (BIQA) methods. Traditional deep learning-based methods emphasize visual information to capture quality features, while recent developments in Vision-Language Models (VLMs) demonstrate strong potential in learning generalizable representations through textual information. However, applying VLMs to BIQA poses three major Challenges: (1) How to make full use of the multi-modal information. (2) The prompt engineering for appropriate quality description is extremely time-consuming. (3) How to use mixed data for joint training to enhance the generalization of VLM-based BIQA model. To this end, we propose a Multi-modal BIQA method with prompt learning, named MMP-IQA. For (1), we propose a conditional fusion module to better utilize the cross-modality information. By jointly adjusting visual and textual features, our model can capture quality information with a stronger representation ability. For (2), we model the quality prompt's context words with learnable vectors during the training process, which can be adaptively updated for superior performances. For (3), we jointly train a linearity-induced quality evaluator, a relative quality evaluator, and a dataset-specific absolute quality evaluator. In addition, we propose a dual automatic weight adjustment strategy to adaptively balance the loss weights between different datasets and among various losses within the same dataset. Extensive experiments illustrate the superior effectiveness of MMP-IQA.
Yan Zhong 0001, Xinping Zhao, Li Zhang 0104, Xinyuan Song 0002, Tingting Jiang 0001
ACM Multimedia4
2025 GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
abstract
Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers, causing the shortcut to dominate over sub-layer outputs in the residual connection and limiting the learning capacity of deeper layers. To mitigate this issue, we propose Gradient-Preserving Activation Scaling (GPAS), a simple technique that can be used in combination with existing approaches. GPAS works by scaling down the intermediate activations while keeping their gradients unchanged. This leaves information in the activations intact, and avoids the gradient vanishing problem associated with gradient downscaling. Extensive experiments across various model sizes from 71M to 1B show that GPAS achieves consistent performance gains. Beyond enhancing Pre-LN Transformers, GPAS also shows promise in improving alternative architectures such as Sandwich-LN and DeepNorm, demonstrating its versatility and potential for improving training dynamics in a wide range of settings. Our code is available at https://github.com/dandingsky/GPAS.
Tianhao Chen, Xin Xu 0001, Zijing Liu, Xinyuan Song 0002, Ajay Jaiswal, Jishan Hu, Yang Wang 0020, Hao Chen 0103, Shizhe Diao, Shiwei Liu 0003, Lu Yin 0006, Can Yang 0002
NeurIPS5
2025 TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning
abstract
We introduce TRiCo, a novel triadic game-theoretic co-training framework that rethinks the structure of semi-supervised learning by incorporating a teacher, two students, and an adversarial generator into a unified training paradigm. Unlike existing co-training or teacher-student approaches, TRiCo formulates SSL as a structured interaction among three roles: (i) two student classifiers trained on frozen, complementary representations, (ii) a meta-learned teacher that adaptively regulates pseudo-label selection and loss balancing via validation-based feedback, and (iii) a non-parametric generator that perturbs embeddings to uncover decision boundary weaknesses. Pseudo-labels are selected based on mutual information rather than confidence, providing a more robust measure of epistemic uncertainty. This triadic interaction is formalized as a Stackelberg game, where the teacher leads strategy optimization and students follow under adversarial perturbations. By addressing key limitations in existing SSL frameworks—such as static view interactions, unreliable pseudo-labels, and lack of hard sample modeling—TRiCo provides a principled and generalizable solution. Extensive experiments on CIFAR-10, SVHN, STL-10, and ImageNet demonstrate that TRiCo consistently achieves state-of-the-art performance in low-label regimes, while remaining architecture-agnostic and compatible with frozen vision backbones.
Hongyang He, Xinyuan Song 0002, Yangfan He, Yanshu Li, Haochen You, Lifan Sun, Wenqiao Zhang
NeurIPS2
2025 The Curse of Depth in Large Language Models
abstract
In this paper, we re-introduce the Curse of Depth, a concept that re-introduces, explains, and addresses the recent observation in modern Large Language Models (LLMs) where deeper layers are much less effective than expected. We first confirm the wide existence of this phenomenon across the most popular families of LLMs, such as Llama, Mistral, DeepSeek, and Qwen. Our analysis, theoretically and empirically, identifies that the underlying reason for the ineffectiveness of deep layers in LLMs is the widespread usage of Pre-Layer Normalization (Pre-LN). While Pre-LN stabilizes the training of Transformer LLMs, its output variance exponentially grows with the model depth, which undesirably causes the derivative of the deep Transformer blocks to be an identity matrix, and therefore barely contributes to the training. To resolve this training pitfall, we propose LayerNorm Scaling, which scales the variance of output of the layer normalization inversely by the square root of its depth. This simple modification mitigates the output variance explosion of deeper Transformer layers, improving their contribution. Our experimental results, spanning model sizes from 130M to 7B, demonstrate that \ours significantly enhances LLM pre-training performance compared to Pre-LN. Moreover, this improvement seamlessly carries over to supervised fine-tuning. All these gains can be attributed to the fact that LayerNorm Scaling enables deeper layers to contribute more effectively during training.
Wenfang Sun, Xinyuan Song 0002, Lu Yin 0006, Yefeng Zheng 0001, Shiwei Liu 0003
NeurIPS2