Mohan Zhang

dblp:218/0654 · DBLP profile ↗
← Back
28ranked-venue papers
13as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 CommSAR: Enabling Bidirectional Communication in SAR Imaging Satellites via Shared Waveform
abstract
Low Earth Orbit (LEO) Synthetic Aperture Radar (SAR) satellites conventionally rely on dedicated communication links, which impose prohibitive hardware, spectrum, and power overhead as satellite constellations scale. This paper presents CommSAR, a novel system that reuses existing SAR imaging waveforms to enable bidirectional communication without modifying satellite hardware or compromising imaging performance. For the downlink, data are embedded by modulating the starting frequency offset of the imaging waveform, preserving the waveform structure and imaging quality. For the uplink, we propose a compact, low-cost programmable metasurface to replace conventional large, expensive antennas, significantly lowering the barrier for dense ground station deployment. To handle extreme satellite dynamics, we employ an opposite-slope waveform as a pilot to compensate for mobility-induced effects. We implement a ground station prototype of CommSAR and validate its performance using an in-orbit commercial SAR satellite and a UAV SAR platform. Experimental results show that CommSAR preserves imaging performance without degradation while achieving downlink and uplink data rates of up to 105 kbps and 112 kbps, respectively, significantly outperforming the state of the art and demonstrating utility-grade performance.
Hao Pan 0003, Minhao Cui, Jie Xiong 0001, Yihai Wei, Yang Liu 0387, Mohan Zhang, Guihai Chen, Kaiyu Liu, Linghe Kong
SIGCOMM10
2026 Hello!AI: An Interactive Rhyme-Based Game for Children AI Literacy Education
abstract
AI literacy is critical for young children as AI is rapidly integrated into people’s daily lives. However, the complexity of AI knowledge presents significant learning challenges, and there is currently a lack of effective approaches for converting complex AI concepts into easy-comprehend content. Based on the formative analysis, we propose Hello!AI, an interactive rhyme-based AI literacy education game targeted at children in Grades 2–6 of primary schools. Hello!AI comprises 3 modules: (i) Algorithm Adventure, focusing on basic AI concept learning, (ii) Algorithm Handbook, promoting thinking and reflection on AI algorithms, and (iii) City Builder, emphasizing the application of AI algorithm to solve real-life problems. We developed the prototype system, iterated it through pilot study, and then conducted a user study. The results demonstrate that Hello!AI can effectively engage children and, to a certain extent, improve their ability to understand and apply AI knowledge, as well as their thinking and reflective capabilities regarding AI technologies.
Mohan Zhang, Changjuan Ran, Fang Liu 0002, Ming Yin 0001, Shenglan Cui, Chuhan Li, Biyao Li
Int. J. Hum. Comput. Interact.1
2026 OMPT: One-stage multiple prompts transfer learning
Yangyang Yu, Keru Wang, Mohan Zhang
Neurocomputing3
2026 Beyond static cues: Detecting fine-grained forgeries via temporal inconsistencies in facial dynamics
Peixu Zhang, Mohan Zhang
Inf. Sci.2
2026 X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering
abstract
We present X2Video, the first diffusion model for rendering photorealistic videos guided by a sequence of intrinsic channels including albedo, normal, roughness, metallicity, and irradiance, while supporting intuitive multi-modal controls with reference images and text prompts for both global and local regions. The intrinsic guidance allows accurate manipulation of color, material, geometry, and lighting, while reference images and text prompts provide intuitive adjustments in the absence of intrinsic information. To enable these functionalities, we extend the intrinsic-guided image generation model XRGB to video generation by employing a novel and efficient Hybrid Self-Attention, which ensures temporal consistency across video frames and also enhances fidelity to reference images. We further develop a Masked Cross-Attention to disentangle global and local text prompts, applying them effectively onto respective local and global regions. For generating long videos, our novel Recursive Sampling method incorporates progressive frame sampling, combining keyframe prediction and frame interpolation to maintain long-range temporal consistency while preventing error accumulation. To support the training of X2Video, we assembled a video dataset named InteriorVideo, featuring 1,154 rooms from 295 interior scenes, complete with reliable ground-truth intrinsic channel sequences and smooth camera trajectories. Both qualitative and quantitative evaluations demonstrate that X2Video can produce long, temporally consistent, and photorealistic videos guided by intrinsic conditions. Additionally, X2Video effectively accommodates multi-modal controls with reference images, global and local text prompts, and simultaneously supports editing on color, material, geometry, and lighting through parametric tuning. Upon acceptance, we will publicly release our model and dataset.
Zhitong Huang, Mohan Zhang, Renhan Wang, Rui Tang 0015, Hao Zhu 0004, Jing Liao 0001
IEEE Trans. Vis. Comput. Graph.2
2025 Modalities Contribute Unequally: Enhancing Medical Multi-modal Learning through Adaptive Modality Token Re-balancing
abstract
Medical multi-modal learning requires an effective fusion capability of various heterogeneous modalities. One vital challenge is how to effectively fuse modalities when their data quality varies across different modalities and patients. For example, in the TCGA benchmark, the performance of the same modality can differ between types of cancer. Moreover, data collected at different times, locations, and with varying reagents can introduce inter-modal data quality differences ($i.e.$, $\textbf{Modality Batch Effect}$). In response, we propose ${\textbf{A}}$daptive ${\textbf{M}}$odality Token Re-Balan${\textbf{C}}$ing ($\texttt{AMC}$), a novel top-down dynamic multi-modal fusion approach. The core of $\texttt{AMC}$ is to quantify the significance of each modality (Top) and then fuse them according to the modality importance (Down). Specifically, we access the quality of each input modality and then replace uninformative tokens with inter-modal tokens, accordingly. The more important a modality is, the more informative tokens are retained from that modality. The self-attention will further integrate these mixed tokens to fuse multi-modal knowledge. Comprehensive experiments on both medical and general multi-modal datasets demonstrate the effectiveness and generalizability of $\texttt{AMC}$.
Jie Peng 0002, Jenna L. Ballard, Mohan Zhang, Sukwon Yun, Jiayi Xin, Qi Long, Yanyong Zhang, Tianlong Chen 0001
ICML3
2025 Experimental and Analytical Analysis of Turn-on Parasitic Oscillation and Dynamic Current Sharing of the Si/SiC Hybrid Switch
abstract
The Si/SiC hybrid switch (HyS), consisting of a low current rated SiC MOSFET paralleled with a high current rated IGBT, can achieve high output current at low cost and reduce switching losses caused by the tail current during the IGBT turn-off transition. When the device turns on, high-frequency ringing at the gate can induce voltage overshoots and current spikes, thereby increasing the device’s current stress. The high-frequency oscillation occurs between the transistor’s capacitance and the board’s parasitic inductance. Such high-frequency oscillations also increase electromagnetic interference (EMI), which can disrupt the performance of the parallel hybrid switch. When the load current is concentrated in the MOSFET, the generated heat accelerates its degradation. Reducing the transition time needed to achieve current balancing can enhance system reliability. In this paper, a detailed turn-on analytical model has been proposed to analyze the effect of parasitic inductances and capacitances, and the current dynamic sharing transition speed to offer guidelines for Si/SiC hybrid power module design. The symmetry of the power loop inductance is also analyzed in the model, providing a more systematic and comprehensive analysis of the turn-on characteristics of the hybrid switch compared to existing models. The validity of the model is verified through simulation and experiments.
Mohan Zhang, Ian Laird, Saeed Jahdi, Wenzhi Zhou, Zhaobo Zhang
IECON1
2025 An Aesthetic Cultural Relic Poster Generation Framework Based on Multi-target Learning and Multimodal Large Language Model
abstract
This paper presents CrePoster, a data-driven framework to generate aesthetic posters for Chinese cultural relics, aiming to enhance the exhibition experience and promote cultural spread. CrePoster comprises three modules: (1) object segmentation module, (2) content generation module, and (3) poster generation module. Upon processing a cultural relic image, the object segmentation module first leverages a cascaded U2Net-SAM structure to obtain the visual target. Secondly, the content generation module utilizes a multi-target learning-enabled caption generator to produce professional captions. Thirdly, the Multimodal Large Language Model (MLLM) based poster generation module adaptively creates aesthetic parameters, including layout and color scheme, ultimately rendering them into refined posters.
Mohan Zhang, Qianqian Hu, Chuhan Li, Yanxiu Dan, Shenglan Cui, Fang Liu 0002
ACM Multimedia1
2025 Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert Parallelism Design
abstract
Mohan Zhang, Pingzhi Li, Jie Peng, Mufan Qiu, Tianlong Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Mohan Zhang, Pingzhi Li, Jie Peng 0002, Mufan Qiu, Tianlong Chen 0001
NAACL (Long Papers)1
2025 RF-Agent: Automated Reward Function Design via Language Agent Tree Search
abstract
Designing efficient reward functions for low-level control tasks is a challenging problem. Recent research aims to reduce reliance on expert experience by using Large Language Models (LLMs) with task information to generate dense reward functions. These methods typically rely on training results as feedback, iteratively generating new reward functions with greedy or evolutionary algorithms. However, they suffer from poor utilization of historical feedback and inefficient search, resulting in limited improvements in complex control tasks. To address this challenge, we propose RF-Agent, a framework that treats LLMs as language agents and frames reward function design as a sequential decision-making process, enhancing optimization through better contextual reasoning. RF-Agent integrates Monte Carlo Tree Search (MCTS) to manage the reward design and optimization process, leveraging the multi-stage contextual reasoning ability of LLM. This approach better utilizes historical information and improves search efficiency to identify promising reward functions. Outstanding experimental results in 17 diverse low-level control tasks demonstrate the effectiveness of our method.
Ning Gao 0004, Xiuhui Zhang, Xingyu Jiang 0003, Mukang You, Mohan Zhang, Yue Deng 0001
NeurIPS5
2025 Progress Reward Model for Reinforcement Learning via Large Language Models
abstract
Traditional reinforcement learning (RL) algorithms face significant limitations in handling long-term tasks with sparse rewards. Recent advancements have leveraged large language models (LLMs) to enhance RL by utilizing their world knowledge for task planning and reward generation. However, planning-based approaches often depend on pre-defined skill libraries and fail to optimize low-level control policies, while reward-based methods require extensive human feedback or exhaustive searching due to the complexity of tasks. In this paper, we propose the Progress Reward Model for RL (PRM4RL), a novel framework that integrates task planning and dense reward to enhance RL. For high-level planning, a complex task is decomposed into a series of simple manageable subtasks, with a subtask-oriented, fine-grained progress function designed to monitor task execution progress. For low-level reward generation, inspired by potential-based reward shaping, we use the progress function to construct a Progress Reward Model (PRM), providing theoretically grounded optimality and convergence guarantees, thereby enabling effective policy optimization. Experimental results on robotics control tasks demonstrate that our approach outperforms both LLM-based planning and reward methods, achieving state-of-the-art performance.
Xiuhui Zhang, Ning Gao 0004, Xingyu Jiang 0003, Yuheng Pan, Mohan Zhang, Yue Deng 0001
NeurIPS6
2025 Reasoning Is Not a Race: When Stopping Early Beats Going Deeper
abstract
We study the use of Process Reward Models (PRMs) for guiding Long Chain-of-Thought (CoT) reasoning in large language models. Although PRMs deliver fine-grained feedback in standard tasks, PRM-guided beam search does not consistently outperform PRM-free approaches in long CoT reasoning. We trace this shortfall to a "step quality degradation''—the expected step quality shows concave behavior, yielding unimodal or monotonically declining trends. To counteract this, we propose Z-Score Guided Early Stopping (ZGES), which halts search at the detected quality peak using local PRM-reward z-scores. Across multiple math benchmarks and model scales, ZGES outperforms both standard PRM-guided beam search and the PRM-free methods. Ablation studies further highlight the advantages and robustness of ZGES’s adaptive stopping mechanism.
Mohan Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu 0013
NeurIPS1
2025 One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
abstract
Modern large reasoning models (LRMs) exhibit impressive multi-step problem-solving via chain-of-thought (CoT) reasoning. However, this iterative thinking mechanism introduces a new vulnerability surface. We present the Deadlock Attack, a resource exhaustion method that hijacks an LRM's generative control flow by training a malicious adversarial embedding to induce perpetual reasoning loops. Specifically, the optimized embedding encourages transitional tokens (e.g., “Wait”, “But”) after reasoning steps, preventing the model from concluding its answer. A key challenge we identify is the continuous-to-discrete projection gap: naïve projections of adversarial embeddings to token sequences nullify the attack. To overcome this, we introduce a backdoor implantation strategy, enabling reliable activation through specific trigger tokens. Our method achieves a 100\% attack success rate across four advanced LRMs (Phi-RM, Nemotron-Nano, R1-Qwen, R1-Llama) and three math reasoning benchmarks, forcing models to generate up to their maximum token limits. The attack is also stealthy (in terms of causing negligible utility loss on benign user inputs) and remains robust against existing strategies trying to mitigate the overthinking issue. Our findings expose a critical and underexplored security vulnerability in LRMs from the perspective of reasoning (in)efficiency.
Mohan Zhang, Jinghan Jia, Zhangyang Wang, Sijia Liu 0001, Tianlong Chen 0001
NeurIPS1
2025 Integrating conceptual and visual representations with domain expertise for scalable visual plagiarism detection
Shenglan Cui, Fang Liu 0002, Yunfan Ye, Mohan Zhang
Expert Syst. Appl.5
2025 ACIH-VQT: aesthetic constraints incorporated hierarchical VQ-transformer for text logo synthesis
Fang Liu 0002, Mohan Zhang, Shenglan Cui
Multim. Syst.3
2024 Intelligent Graphic Layout Generation: Current Status and Future Perspectives
abstract
Graphic Layout Generation focuses on providing layout references for various visual design tasks, such as advertisement design and poster design. Researchers have conducted exploratory studies from different dimensions, such as human-computer interaction and layout representation. With the success of deep learning techniques, there has been a recent surge in research on graphic layout generation using deep generative models. Yet there are still no comprehensive studies on graphic layout generation. In this paper, we review methods for graphic layout generation from two perspectives: implementation and interactivity. We analyze the current status, the advantageous application scenarios, and the challenges of different methods. Based on the analysis results, we summarize the workflow of designers collaborating with graphic automatic layout systems and propose potential directions for future research.
Fang Liu 0002, Mohan Zhang
CSCWD3
2024 CrePoster: Leveraging multi-level features for cultural relic poster generation via attention-based framework
Mohan Zhang, Fang Liu 0002, Biyao Li, Wentao Ma 0003, Changjuan Ran
Expert Syst. Appl.1
2024 Intelligent-paint: a Chinese painting process generation method based on vision transformer
Zunfu Wang, Fang Liu 0002, Changjuan Ran, Mohan Zhang
Multim. Syst.5
2024 LVCD: Reference-based Lineart Video Colorization with Diffusion Models
abstract
We propose the first video diffusion framework for reference-based lineart video colorization. Unlike previous works that rely solely on image generative models to colorize lineart frame by frame, our approach leverages a large-scale pretrained video diffusion model to generate colorized animation videos. This approach leads to more temporally consistent results and is better equipped to handle large motions. Firstly, we introduce Sketch-guided ControlNet which provides additional control to finetune an image-to-video diffusion model for controllable video synthesis, enabling the generation of animation videos conditioned on lineart. We then propose Reference Attention to facilitate the transfer of colors from the reference frame to other frames containing fast and expansive motions. Finally, we present a novel scheme for sequential sampling, incorporating the Overlapped Blending Module and Prev-Reference Attention , to extend the video diffusion model beyond its original fixed-length limitation for long video colorization. Both qualitative and quantitative results demonstrate that our method significantly outperforms state-of-the-art techniques in terms of frame and video quality, as well as temporal consistency. Moreover, our method is capable of generating high-quality, long temporal-consistent animation videos with large motions, which is not achievable in previous works. Our code and model are available at https://luckyhzt.github.io/lvcd.
Zhitong Huang, Mohan Zhang, Jing Liao 0001
ACM Trans. Graph.2
2023 Image captioning for cultural artworks: a case study on ceramics
Baoying Zheng, Fang Liu 0002, Mohan Zhang, Tongqing Zhou, Shenglan Cui, Yunfan Ye, Yeting Guo
Multim. Syst.3
2022 Understanding and Identifying Artwork Plagiarism with the Wisdom of Designers: A Case Study on Poster Artworks
abstract
The wide sharing and rapid dissemination of digital artworks has aggravated the issues of plagiarism, raising significant concerns in cultural preservation and copyright protection. Yet, modes of plagiarism are formally uncharted, causing rough plagiarism detection practices with duplicate checking. This work is thus devoted to understanding artwork plagiarism, with poster design as the running case, for building more dedicated detection techniques. As the first study of such, we elaborate on 8 elements that form unique posters and 6 judgement criteria for plagiarism using an exploratory study with designers. Second, we build a novel poster dataset with plagiarism annotations according to the criteria. Third, we propose models, leveraging the combination of primary elements and criteria of plagiarism, to find suspect instances in a retrieval process. The models are trained under the context of modern artwork and evaluated on the poster plagiarism dataset. The proposal is shown to outperform the baseline with superior Top-K accuracy (~33%) and retrieval performance (~42%).
Shenglan Cui, Fang Liu 0002, Tongqing Zhou, Mohan Zhang
ACM Multimedia4
2022 SMPL: Simulated Industrial Manufacturing and Process Control Learning Environments
abstract
Traditional biological and pharmaceutical manufacturing plants are controlled by human workers or pre-defined thresholds. Modernized factories have advanced process control algorithms such as model predictive control (MPC). However, there is little exploration of applying deep reinforcement learning to control manufacturing plants. One of the reasons is the lack of high fidelity simulations and standard APIs for benchmarking. To bridge this gap, we develop an easy-to-use library that includes five high-fidelity simulation environments: BeerFMTEnv, ReactorEnv, AtropineEnv, PenSimEnv and mAbEnv, which cover a wide range of manufacturing processes. We build these environments on published dynamics models. Furthermore, we benchmark online and offline, model-based and model-free reinforcement learning algorithms for comparisons of follow-up research.
Mohan Zhang, Xiaozhou Wang, Benjamin Decardi-Nelson, Song Bo, An Zhang 0007, Jinfeng Liu 0001, Sile Tao, Jiayi Cheng, Xiaohong Liu 0001, Dengdeng Yu, Matthew Poon, Animesh Garg
NeurIPS1
2022 Deep Exemplar-Based Color Transfer for 3D Model
abstract
Recoloring 3D models is a challenging task that often requires professional knowledge and tedious manual efforts. In this article, we present the first deep-learning framework for exemplar-based 3D model recolor, which can automatically transfer the colors from a reference image to the 3D model texture. Our framework consists of two modules to solve two major challenges in the 3D color transfer. First, we propose a new feed-forward Color Transfer Network to achieve high-quality semantic-level color transfer by finding dense semantic correspondences between images. Second, considering 3D model constraints such as UV mapping, we design a novel 3D Texture Optimization Module which can generate a seamless and coherent texture by combining color transferred results rendered in multiple views. Experiments show that our method performs robustly and generalizes well to various kinds of models.
Mohan Zhang, Jing Liao 0001
IEEE Trans. Vis. Comput. Graph.1
2021 Don't Change Me! User-Controllable Selective Paraphrase Generation
abstract
Mohan Zhang, Luchen Tan, Zihang Fu, Kun Xiong, Jimmy Lin, Ming Li, Zhengkai Tu. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Mohan Zhang, Luchen Tan, Zihang Fu, Kun Xiong, Jimmy Lin, Ming Li 0001, Zhengkai Tu
EACL1
2020 RT-VENet: A Convolutional Network for Real-time Video Enhancement
abstract
Real-time video enhancement is in great demand due to the extensive usage of live video applications, but existing approaches are far from satisfying the strict requirements of speed and stability. We present a novel convolutional network that can perform high-quality enhancement on 1080p videos at 45 FPS with a single CPU, which has high potential for real-world deployment. The proposed network is designed based on a light-weight image network and further consolidated for temporal consistency with a temporal feature aggregation (TFA) module. Unlike most image translation networks that use decoders to generate target images, our network discards decoders and employs only an encoder and a small head. The network predicts color mapping functions instead of pixel values in a grid-like container which fits the CNN structure well and also advances the enhancement to be scalable to any video resolution. Furthermore, the temporal consistency of the output will be enforced by the TFA module which utilizes the learned temporal coherence of semantics across frames. We also demonstrate that the mapping representation is general to various enhancement tasks, such as relighting, retouching and dehazing, on benchmark datasets. Our approach achieves the state-of-the-art performance and performs about 10 times faster than the current real-time method on high-resolution videos.
Mohan Zhang, Jinglu Wang, Henrik Turbell, Yan Lu 0001
ACM Multimedia1
2020 SCODED: Statistical Constraint Oriented Data Error Detection
abstract
Statistical Constraints (SCs) play an important role in statistical modeling and analysis. This paper brings the concept to data cleaning and studies how to leverage SCs for error detection. SCs provide a novel approach that has various application scenarios and works harmoniously with downstream statistical modeling. Entailment relationships between SCs and integrity constraints provide analytical insight into SCs. We develop SCODED, an SC-Oriented Data Error Detection system, comprising two key components: (1) SC Violation Detection : checks whether an SC is violated on a given dataset, and (2) Error Drill Down : identifies the top-k records that contribute most to the violation of an SC. Experiments on synthetic and real-world data show that SCs are effective in detecting data errors that violate them, compared to state-of-the-art approaches.
Jing Nathan Yan, Oliver Schulte, Mohan Zhang, Jiannan Wang 0001, Reynold Cheng
SIGMOD Conference3
2019 Artistic Augmentation of Photographs with Droplets
Mohan Zhang, Kang Zhang 0001, Junsong Zhang
J. Comput. Sci. Technol.1
2018 Computer Simulation and Generation of Moving Sand Pictures
abstract
Moving sand pictures are interesting devices that can be used to generate an infinite number of unique scenes when repeatedly being flipped over. However, little work has been done on attempting to simulate the process of picture formulation. In this paper, we present an approach capable of generating images in the style of moving sand pictures. Our system defines moving sand pictures in a few steps, such as initialization, segmentation and physical simulation, so that a variety of moving sand pictures including mountain ridges, desert, clouds and even regular patterns can be generated by either automatic or semi-automatic via interaction during initialization and segmentation. Potential applications of our approach range from advertisements, posters, post cards, packaging, to digital arts.
Mohan Zhang, Kang Zhang 0001
IEEE Trans. Vis. Comput. Graph.1