Ruixiao Shi

dblp:367/6999 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 25% Knowledge representation and reasoning · 25% Efficient and distributed learning · 25%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › image generation
controllable image generation
1.012026
DivControl: Knowledge Diversion for Controllable Image Generation · AAAI 2026
Visual content generation and editing
image generation
1.012026
DivControl: Knowledge Diversion for Controllable Image Generation · AAAI 2026
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › markov random field
decomposable models
0.912025
KIND: Knowledge Integration and Diversion for Training Decomposable Models · ICML 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge engineering
knowledge integration
0.912025
KIND: Knowledge Integration and Diversion for Training Decomposable Models · ICML 2025
Machine learning › Transfer learning and domain adaptation
knowledge transfer
0.912025
ECO: Evolving Core Knowledge for Efficient Transfer · NeurIPS 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
KIND: Knowledge Integration and Diversion for Training Decomposable Models · ICML 2025

Methods — techniques the papers use, named apart from their topics

representation alignment loss · 1.0dynamic gating · 1.0diffusion model · 1.0SVD factorization · 1.0singular value decomposition · 0.9genetic transfer learning · 0.9evolutionary algorithm · 0.9class gate mechanism · 0.9
YearPublicationVenuePosition
2026 DivControl: Knowledge Diversion for Controllable Image Generation
abstract
Diffusion models have advanced from text-to-image (T2I) to image-to-image (I2I) generation by incorporating structured inputs such as depth maps, enabling fine-grained spatial control. However, existing methods either train separate models for each condition or rely on unified architectures with entangled representations, resulting in poor generalization and high adaptation costs for novel conditions. To this end, we propose DivControl, a decomposable pretraining framework for unified controllable generation and efficient adaptation. DivControl factorizes ControlNet via SVD into basic components—pairs of singular vectors—which are disentangled into condition-agnostic learngenes and condition-specific tailors through knowledge diversion during multi-condition training. Knowledge diversion is implemented via a dynamic gate that performs soft routing over tailors based on the semantics of condition instructions, enabling zero-shot generalization and parameter-efficient adaptation to novel conditions. To further improve condition fidelity and training efficiency, we introduce a representation alignment loss that aligns condition embeddings with early diffusion features. Extensive experiments demonstrate that DivControl achieves state-of-the-art controllability with 36.4× less training cost, while simultaneously improving average performance on basic conditions. It also delivers strong zero-shot and few-shot performance on unseen conditions, demonstrating superior scalability, modularity, and transferability.
Yucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang 0113, Yong Rui, Xin Geng 0001
AAAI3
2025 KIND: Knowledge Integration and Diversion for Training Decomposable Models
abstract
Pre-trained models have become the preferred backbone due to the increasing complexity of model parameters. However, traditional pre-trained models often face deployment challenges due to their fixed sizes, and are prone to negative transfer when discrepancies arise between training tasks and target tasks. To address this, we propose **KIND**, a novel pre-training method designed to construct decomposable models. KIND integrates knowledge by incorporating Singular Value Decomposition (SVD) as a structural constraint, with each basic component represented as a combination of a column vector, singular value, and row vector from $U$, $\Sigma$, and $V^\top$ matrices. These components are categorized into **learngenes** for encapsulating class-agnostic knowledge and \textbf{tailors} for capturing class-specific knowledge, with knowledge diversion facilitated by a class gate mechanism during training. Extensive experiments demonstrate that models pre-trained with KIND can be decomposed into learngenes and tailors, which can be adaptively recombined for diverse resource-constrained deployments. Moreover, for tasks with large domain shifts, transferring only learngenes with task-agnostic knowledge, when combined with randomly initialized tailors, effectively mitigates domain shifts. Code will be made available at https://github.com/Te4P0t/KIND.
Yucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang 0113, Yong Rui, Xin Geng 0001
ICML3
2025 ECO: Evolving Core Knowledge for Efficient Transfer
abstract
Knowledge in modern neural networks is often entangled and structurally opaque, making current transfer methods—typically based on reusing entire parameter sets—inefficient and inflexible. Efforts to improve flexibility by reusing partial parameters frequently depend on handcrafted heuristics or rigid structural assumptions, which constrain generalization. In contrast, biological evolution enables efficient knowledge transfer by encoding only essential information into genes through iterative refinement under environmental pressure. Inspired by this principle, we propose **ECO**, a framework that **E**volves **CO**re knowledge into modular, reusable neural components—termed *learngenes*—through similar evolutionary dynamics. To this end, we redefine learngenes as neural circuits and introduce Genetic Transfer Learning (GTL), a biologically inspired paradigm that establishes a genetic mechanism within neural networks in the context of supervised learning. GTL simulates evolutionary processes by generating diverse network populations, selecting high-performing individuals, and transferring their learngenes to subsequent generations. Through iterative refinement, GTL enables learngenes to accumulate transferable common knowledge. Extensive experiments show that ECO achieves efficient initialization and strong generalization across diverse models and tasks, while significantly reducing computational and memory costs compared to conventional methods.
Fu Feng, Yucheng Xie, Ruixiao Shi, Jianlu Shen, Jing Wang 0113, Xin Geng 0001
NeurIPS3
2024 Multimodal-Enhanced Objectness Learner For Corner Case Detection In Autonomous Driving
abstract
Previous works on object detection have achieved high accuracy in closed-set scenarios, but their performance in open-world scenarios is not satisfactory. One of the challenging open-world problems is corner case detection in autonomous driving. Existing detectors struggle with these cases, relying heavily on visual appearance and exhibiting poor generalization ability. In this paper, we propose a solution by reducing the discrepancy between known and unknown classes and introduce a multimodal-enhanced objectness notion learner. Leveraging both vision-centric and vision-language multiple modalities, our semi-supervised learning framework imparts objectness knowledge to the student model, enabling class-aware detection. Our approach, Multimodal-Enhanced Objectness Learner (MENOL) for Corner Case Detection, significantly improves recall for novel classes with lower training costs. By achieving a $76.6 \% \mathrm{mAR}$-corner and $79.8 \%$ mAR-agnostic on the CODA-val dataset with just 5100 labeled training images, MENOL outperforms the baseline ORE by $71.3 \%$ and $60.6 \%$, respectively. The code will be available at https://github.com/tryhiseyyysum/ MENOL.
Lixing Xiao, Ruixiao Shi, Xiaoyang Tang
ICIP2