VLDB 2026 Research / reviewers in the wild / expert
Sicheng Mo
dblp:319/6786
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Generative modeling · 45% Vision and language · 32% Language models and text generation · 8% | |
| Computer graphics and multimedia
3 papers |
Visual content generation and editing · 61% Computational photography and imaging · 39% |
Topics — the 19 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.3 | 3 | 2024 | SimGen: Simulator-conditioned Driving Scene Generation · NeurIPS 2024 Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance · NeurIPS 2024 FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition · CVPR 2024 |
Visual content generation and editing › image generation
text-to-image generation |
1.5 | 2 | 2024 | Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance · NeurIPS 2024 FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition · CVPR 2024 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | X-Fusion: Introducing New Modality to Frozen Large Language Models · ICCV 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | X-Fusion: Introducing New Modality to Frozen Large Language Models · ICCV 2025 |
Computational photography and imaging
high-speed imaging |
0.9 | 1 | 2025 | Physics to the Rescue: Deep Non-Line-of-Sight Reconstruction for High-Speed Imaging · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computational photography and imaging
non-line-of-sight imaging |
0.9 | 1 | 2025 | Physics to the Rescue: Deep Non-Line-of-Sight Reconstruction for High-Speed Imaging · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
controllable diffusion generation |
0.8 | 1 | 2024 | FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition · CVPR 2024 |
Machine learning › Generative modeling › diffusion model
controllable generation |
0.8 | 1 | 2024 | Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model › controllable generation
controllable scene generation |
0.8 | 1 | 2024 | SimGen: Simulator-conditioned Driving Scene Generation · NeurIPS 2024 |
Robotics › Autonomous driving › scenario generation
driving scene generation |
0.8 | 1 | 2024 | SimGen: Simulator-conditioned Driving Scene Generation · NeurIPS 2024 |
Machine learning › Kernel, tree and ensemble methods › classifier combination
late fusion |
0.8 | 1 | 2024 | SnAG: Scalable and Accurate Video Grounding · CVPR 2024 |
Computer vision › Vision and language
multimodal fusion |
0.8 | 1 | 2024 | SnAG: Scalable and Accurate Video Grounding · CVPR 2024 |
Computer vision › Vision and language
temporal grounding |
0.8 | 1 | 2024 | SnAG: Scalable and Accurate Video Grounding · CVPR 2024 |
Computer vision › Vision and language
video grounding |
0.8 | 1 | 2024 | SnAG: Scalable and Accurate Video Grounding · CVPR 2024 |
Visual content generation and editing › image editing › controllable image editing
structure and appearance control |
0.8 | 1 | 2024 | Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance · NeurIPS 2024 |
Computer vision › Vision and language › vision-language generation
image-to-text generation |
0.3 | 1 | 2025 | X-Fusion: Introducing New Modality to Frozen Large Language Models · ICCV 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.3 | 1 | 2025 | X-Fusion: Introducing New Modality to Frozen Large Language Models · ICCV 2025 |
Visual content generation and editing
image editing |
0.2 | 1 | 2024 | Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance · NeurIPS 2024 |
Visual content generation and editing
image generation |
0.2 | 1 | 2024 | FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
training-free guidance · 1.5structure guidance · 1.5feed-forward structure control · 1.5diffusion model · 1.5appearance transfer · 1.5appearance guidance · 1.5wave propagation prior · 0.9volume rendering · 0.9neural network · 0.9frozen LLM · 0.9feature alignment · 0.9dual-tower architecture · 0.9video-centric sampling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | X-Fusion: Introducing New Modality to Frozen Large Language ModelsabstractWe propose X-Fusion, a framework that extends pretrained Large Language Models (LLMs) for multimodal tasks while preserving their language capabilities. X-Fusion employs a dual-tower design with modality-specific weights, keeping the LLM's parameters frozen while integrating vision-specific information for both understanding and generation. Our experiments demonstrate that X-Fusion consistently outperforms alternative architectures on both image-to-text and text-to-image tasks. We find that incorporating understanding-focused data improves generation quality, reducing image data noise enhances overall performance, and feature alignment accelerates convergence for smaller models but has minimal impact on larger ones. Our findings provide valuable insights into building efficient unified multimodal models. Sicheng Mo, Siddharth Srinivasan Iyer, Yijun Li 0001, Yuchen Liu 0002, Abhishek Tandon, Eli Shechtman, Krishna Kumar Singh, Yong Jae Lee, Bolei Zhou |
ICCV | 1 |
| 2025 | Physics to the Rescue: Deep Non-Line-of-Sight Reconstruction for High-Speed ImagingabstractComputational approach to imaging around the corner, or non-line-of-sight (NLOS) imaging, is becoming a reality thanks to major advances in imaging hardware and reconstruction algorithms. A recent development towards practical NLOS imaging, (Nam et al. 2021) demonstrated a high-speed non-confocal imaging system that operates at 5Hz, 100x faster than the prior art. This enormous gain in acquisition rate, however, necessitates numerous approximations in light transport, breaking many existing NLOS reconstruction methods that assume an idealized image formation model. To bridge the gap, we present a novel deep model that incorporates the complementary physics priors of wave propagation and volume rendering into a neural network for high-quality and robust NLOS reconstruction. This orchestrated design regularizes the solution space by relaxing the image formation model, resulting in a deep model that generalizes well on real captures despite being exclusively trained on synthetic data. Further, we devise a unified learning framework that enables our model to be flexibly trained using diverse supervision signals, including target intensity images or even raw NLOS transient measurements. Once trained, our model renders both intensity and depth images at inference time in a single forward pass, capable of processing more than 5 captures per second on a high-end GPU. Through extensive qualitative and quantitative experiments, we show that our method outperforms prior physics and learning based approaches on both synthetic and real measurements. We anticipate that our method along with the fast capturing system will accelerate future development of NLOS imaging for real world applications that require high-speed imaging. Fangzhou Mu, Sicheng Mo, Jiayong Peng, Xiaochun Liu, Ji Hyun Nam, Siddeshwar Raghavan, Andreas Velten, Yin Li 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any ConditionabstractRecent approaches such as ControlNet [59] offer users fine-grained spatial control over text-to-image (T2I) diffusion models. However, auxiliary modules have to be trained for each spatial condition type, model architecture, and checkpoint, putting them at odds with the diverse intents and preferences a human designer would like to convey to the AI models during the content creation process. In this work, we present FreeControl, a training-free approach for controllable T2I generation that supports multiple conditions, architectures, and checkpoints simultaneously. Free Control enforces structure guidance to facilitate the global alignment with a guidance image, and appearance guidance to collect visual details from images generated without control. Extensive qualitative and quantitative experiments demonstrate the superior performance of Free Control across a variety of pre-trained T2I models. In particular, FreeControl enables convenient training-free control over many different architectures and checkpoints, allows the challenging input conditions on which most of the existing training-free methods fail, and achieves competitive synthesis quality compared to training-based approaches. Project page: https://genforce.github.io/freecontrol/. Sicheng Mo, Fangzhou Mu, Kuan Heng Lin, Bochen Guan, Yin Li 0003, Bolei Zhou |
CVPR | 1 |
| 2024 | SnAG: Scalable and Accurate Video GroundingabstractTemporal grounding of text descriptions in videos is a central problem in vision-language learning and video understanding. Existing methods often prioritize accuracy over scalability - they have been optimized for grounding only a few text queries within short videos, and fail to scale up to long videos with hundreds of queries. In this paper, we study the effect of cross-modal fusion on the scalability of video grounding models. Our analysis establishes late fusion as a more cost-effective fusion scheme for long-form videos with many text queries. Moreover, it leads us to a novel, video-centric sampling scheme for efficient training. Based on these findings, we present SnAG, a simple baseline for scalable and accurate video grounding. Without bells and whistles, SnAG is 43% more accurate and$l.5\times$faster than CONE, a state of the art for long-form video grounding on the challenging MAD dataset, while achieving highly competitive results on short videos. Our code is available at https://github.com/fmu2/snag_release. Fangzhou Mu, Sicheng Mo, Yin Li 0003 |
CVPR | 2 |
| 2024 | Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without GuidanceabstractRecent controllable generation approaches such as FreeControl and Diffusion Self-Guidance bring fine-grained spatial and appearance control to text-to-image (T2I) diffusion models without training auxiliary modules. However, these methods optimize the latent embedding for each type of score function with longer diffusion steps, making the generation process time-consuming and limiting their flexibility and use. This work presents *Ctrl-X*, a simple framework for T2I diffusion controlling structure and appearance without additional training or guidance. Ctrl-X designs feed-forward structure control to enable the structure alignment with a structure image and semantic-aware appearance transfer to facilitate the appearance transfer from a user-input image. Extensive qualitative and quantitative experiments illustrate the superior performance of Ctrl-X on various condition inputs and model checkpoints. In particular, Ctrl-X supports novel structure and appearance control with arbitrary condition images of any modality, exhibits superior image quality and appearance transfer compared to existing works, and provides instant plug-and-play functionality to any T2I and text-to-video (T2V) diffusion model. See our project page for the code and an overview of the results: https://genforce.github.io/ctrl-x Kuan Heng Lin, Sicheng Mo, Ben Klingher, Fangzhou Mu, Bolei Zhou |
NeurIPS | 2 |
| 2024 | SimGen: Simulator-conditioned Driving Scene GenerationabstractControllable synthetic data generation can substantially lower the annotation cost of training data. Prior works use diffusion models to generate driving images conditioned on the 3D object layout. However, those models are trained on small-scale datasets like nuScenes, which lack appearance and layout diversity. Moreover, overfitting often happens, where the trained models can only generate images based on the layout data from the validation set of the same dataset. In this work, we introduce a simulator-conditioned scene generation framework called SimGen that can learn to generate diverse driving scenes by mixing data from the simulator and the real world. It uses a novel cascade diffusion pipeline to address challenging sim-to-real gaps and multi-condition conflicts. A driving video dataset DIVA is collected to enhance the generative diversity of SimGen, which contains over 147.5 hours of real-world driving videos from 73 locations worldwide and simulated driving data from the MetaDrive simulator. SimGen achieves superior generation quality and diversity while preserving controllability based on the text prompt and the layout pulled from a simulator. We further demonstrate the improvements brought by SimGen for synthetic data augmentation on the BEV detection and segmentation task and showcase its capability in safety-critical data generation. Yunsong Zhou, Michael Simon, Zhenghao Peng, Sicheng Mo, Hongzi Zhu, Minyi Guo, Bolei Zhou |
NeurIPS | 4 |