Chenguang Ma

dblp:28/9718 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-3627-2740ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SPEED-Q: Staged Processing with Enhanced Distillation Towards Efficient Low-Bit On-Device VLM Quantization
abstract
Deploying Vision-Language Models (VLMs) on edge devices (e.g., smartphones and robots) is crucial for enabling low-latency and privacy-preserving intelligent applications. Given the resource constraints of these devices, quantization offers a promising solution by improving memory efficiency and reducing bandwidth requirements, thereby facilitating the deployment of VLMs. However, existing research has rarely explored aggressive quantization on VLMs, particularly for the models ranging from 1B to 2B parameters, which are more suitable for resource-constrained edge devices. In this paper, we propose SPEED-Q, a novel Staged Processing with EnhancEd Distillation framework for VLM low-bit weight-only quantization that systematically addresses the following two critical obstacles: (1) significant discrepancies in quantization sensitivity between vision (ViT) and language (LLM) components in VLMs; (2) training instability arising from the reduced numerical precision inherent in low-bit quantization. In SPEED-Q, a staged sensitivity adaptive mechanism is introduced to effectively harmonize performance across different modalities. We further propose a distillation-enhanced quantization strategy to stabilize the training process and reduce data dependence. Together, SPEED-Q enables accurate, stable, and data-efficient quantization of complex VLMs. SPEED-Q is the first framework tailored for quantizing entire small-scale billion-parameter VLMs to low bits. Extensive experiments across multiple benchmarks demonstrate that SPEED-Q achieves up to 6x higher accuracy than existing quantization methods under 2-bit settings and consistently outperforms prior on-device VLMs under both 2-bit and 4-bit settings.
Tianyu Guo 0013, Shanwei Zhao, Shiai Zhu, Chenguang Ma
AAAI4
2026 EchoMimicV3: 1.3B Parameters Are All You Need for Unified Multi-Modal and Multi-Task Human Animation
abstract
Recent work on human animation usually incorporates large-scale video models, thereby achieving more vivid performance. However, the practical use of such methods is hindered by the slow inference speed and high computational demands. Moreover, traditional work typically employs separate models for each animation task, increasing costs in multi-task scenarios and worsening the dilemma. To address these limitations, we introduce EchoMimicV3, an efficient framework that unifies multi-task and multi-modal human animation. At the core of EchoMimicV3 lies a threefold design: a Soup-of-Tasks paradigm, a Soup-of-Modals paradigm, and a novel training and inference strategy. The Soup-of-Tasks leverages multi-task mask inputs and a counter-intuitive task allocation strategy to achieve multi-task gains without multi-model pains. Meanwhile, the Soup-of-Modals introduces a Coupled-Decoupled Multi-Modal Cross Attention module to inject multi-modal conditions, complemented by a Timestep Phase-aware Multi-Modal Allocation mechanism to dynamically modulate multi-modal mixtures. Besides, we propose Negative Direct Preference Optimization and Phase-aware Negative Classifier-Free Guidance, which ensure stable training and inference. Extensive experiments and analyses demonstrate that EchoMimicV3, with a minimal model size of 1.3 billion parameters, achieves competitive performance in both quantitative and qualitative evaluations. We are committed to open-sourcing our code for community use.
Rang Meng, Weipeng Wu, Ruobing Zheng, Chenguang Ma
AAAI6
2026 GroupPortrait: Multi-ID Portrait Generation with High Identity Preservation and Fine-Grained Control
abstract
Identity-preserving portrait generation has achieved tremendous advancements with the development of diffusion models. However, multi-ID generation remains challenging due to degraded identity fidelity and insufficient control over layout, pose, and expression. To address these challenges, we propose GroupPortrait, a novel approach for multi-ID portrait generation with three key innovations:(1) LatentID for high-fidelity identity preservation, (2) Facial Controller enabling layout guidance and fine-grained facial control, and (3) Mask-Attention Controller allocating identity embeddings to specific facial regions. First, the LatentID module improves identity preservation by adding LatentID loss during training. It maps latent representations to identity features and uses ID consistency loss for feedback training to improve identity retention. Since LatentID loss is calculated in latent space, it is more efficient in terms of time and GPU usage compared to the method that calculates ID loss in pixel space. Second, to enhance layout and facial controllability, the Facial Controller utilizes 3D Morphable Models (3DMM) to acquire facial shapes, poses, and expressions for each individual, imposing strong spatial conditions during the diffusion process. Finally, we propose a novel Mask-Attention Controller for multi-ID generation, which distributes ID embeddings into target facial regions by aligning the cross-attention map of LatentID with the given facial region masks. Extensive experiments demonstrate that GroupPortrait can generate group portraits with high fidelity, local harmony, and controllability.
Meijia Huang, Ruida Li, Liangwei Jiang, Shuo Fang, Chenguang Ma
WACV6
2026 EmojiDiff: Advanced Facial Expression Control with High Identity Preservation in Portrait Generation
abstract
This paper aims to bring fine-grained expression control while maintaining high-fidelity identity in portrait generation. This is challenging due to the mutual interference between expression and identity. On one hand, fine expression control signals inevitably introduce appearance-related semantics (e.g., facial contours, and ratio), which impact the identity of the generated portrait. On the other hand, even coarse-grained expression control can cause facial changes that compromise identity, since they all act on the face. Here, we introduce EmojiDiff, the first end-to-end solution that enables simultaneous control of extremely detailed expression (RGB-level) and high-fidelity identity in portrait generation. To address the above challenges, EmojiDiff adopts a two-stage scheme involving decoupled training and fine-tuning. For decoupled training, we innovate ID-irrelevant Data Iteration (IDI) to synthesize high-quality cross-identity expression pairs by separating and optimizing the processes of maintaining expression and altering identity. Training the model with this data, we effectively disentangle fine expression features in the expression template from other extraneous information (e.g., identity, skin). Subsequently, we present ID-enhanced Contrast Alignment (ICA) for further fine-tuning. ICA achieves rapid reconstruction and joint supervision of identity and expression information, thus aligning identity representations of images with and without expression control. Experimental results demonstrate that our method significantly outperforms its counterparts, achieving precise expression control with highly maintained identity, and generalizing well to various diffusion models. Project page: https://emojidiff.github.io/.
Liangwei Jiang, Ruida Li, Shuo Fang, Chenguang Ma
WACV5
2026 Spatial-Temporal Multimodal Large Language Model for Generative Recommendation in Alipay
abstract
Despite the encouraging achievements, the practical application of recommendation systems still faces two key issues. The first is how to better understand the multimodal real-time requests that are the more mainstream request behavior in industrial scenarios; the other is how to effectively capture users' dynamic needs that change with temporal and spatial conditions. The breakthroughs in text understanding and generation capabilities of Large Language Models (LLMs) have demonstrated their tremendous potential in precise recommendation systems, particularly through the enhancement of the understanding of user intent. To address these issues, we propose a novel Spatial-Temporal Multimodal LLM for generative recommendation. Specifically, on the basis of the behavior data constructed from Alipay, spatial-temporal knowledge-guided fine-tuning module is proposed to capture specific needs in user real-time requests. Furthermore, a preference discovery module is developed to learn user preferences in visual queries from multimodal request perspective. Meanwhile, a personalized recommendation module is designed to aggregate spatial-temporal knowledge and user preferences for generative recommendation. Experimental results on a real-world deployed generative recommendation task from the ‘Explore' scenario in Alipay have demonstrated the effectiveness of the proposed framework.
Yunhui Xu, Youru Li, Zhenfeng Zhu, Zujian Weng, Jingjuan Zhao, Chenguang Ma, Jieping Ye, Yao Zhao 0001
IEEE Trans. Knowl. Data Eng.7
2025 EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions
abstract
The area of portrait image animation, propelled by audio input, has witnessed notable progress in the generation of lifelike and dynamic portraits. Conventional methods are limited to utilizing either audios or facial key points to drive images into videos, while they can yield satisfactory results, certain issues exist. For instance, methods driven solely by audios can be unstable at times due to the relatively weaker audio signal, while methods driven exclusively by facial key points, although more stable in driving, can result in unnatural outcomes due to the excessive control of key point information. In addressing the previously mentioned challenges, in this paper, we introduce a novel approach which we named EchoMimic. EchoMimic is concurrently trained using both audios and facial landmarks. Through the implementation of a novel training strategy, EchoMimic is capable of generating portrait videos not only by audios and facial landmarks individually, but also by a combination of both audios and selected facial landmarks. EchoMimic has been comprehensively compared with alternative algorithms across various public datasets and our collected dataset, showcasing superior performance in both quantitative and qualitative evaluations. The code and models are available on the project page.
Jiajiong Cao, Zhiquan Chen, Chenguang Ma
AAAI5
2025 EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
abstract
Recent work on human animation usually involves audio, pose, or movement maps conditions, thereby achieves vivid animation quality. However, these methods often face practical challenges due to extra control conditions, cumbersome condition injection modules, or limitation to head region driving. Hence, we ask if it is possible to achieve striking half-body human animation while simplifying unnecessary conditions. To this end, we propose a half-body human animation method, dubbed EchoMimicV2, that leverages a novel Audio-Pose Dynamic Harmonization strategy, including Pose Sampling and Audio Diffusion, to enhance half-body details, facial and gestural expressiveness, and meanwhile reduce conditions redundancy. To compensate for the scarcity of half-body data, we utilize Head Partial Attention to seamlessly accommodate headshot data into our training framework, which can be omitted during inference, providing a free lunch for animation. Furthermore, we design the Phase-specific Denoising Loss to guide motion, detail, and low-level quality for animation in specific phases, respectively. Besides, we also present a novel benchmark for evaluating the effectiveness of half-body human animation. Extensive experiments and analyses demonstrate that EchoMimicV2 surpasses existing methods in both quantitative and qualitative evaluations.
Rang Meng, Chenguang Ma
CVPR4
2025 Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency
abstract
As a very common type of video, face videos often appear in movies, talk shows, live broadcasts, and other scenes. Real-World online videos are often plagued by degradations such as blurring and quantization noise, due to the high compression ratio caused by high communication costs and limited transmission bandwidth. These degradations have a particularly serious impact on face videos because the human visual system is highly sensitive to facial details. Despite the significant advancement in video face enhancement, current methods still suffer from i) long processing time and ii) inconsistent spatial-temporal visual effects (e.g., flickering). This study proposes a novel and efficient blind video face enhancement method to overcome the above two challenges, restoring high-quality videos from their compressed low-quality versions with an effective de-flickering mechanism. In particular, the proposed method develops upon a 3D-VQGAN backbone associated with spatial-temporal codebooks recording high-quality portrait features and residual-based temporal information. We develop a two-stage learning framework for the model. In Stage I, we learn the model with a regularizer mitigating the codebook collapse problem. In Stage II, we learn two transformers to look up code from the codebooks and further update the encoder of low-quality videos. Experiments conducted on the VFHQ-Test dataset demonstrate that our method surpasses the current state-of-the-art blind face video restoration and de-flickering methods on both efficiency and effectiveness. Code is available at https://github.com/Dixin-Lab/BFVR-STC.
Jiajie Teng, Jiajiong Cao, Chenguang Ma, Hongteng Xu, Dixin Luo
CVPR5
2025 Noise-Optimized Distribution Distillation for Dataset Condensation
Tongfei Liu, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Chenguang Ma
ACM Multimedia6
2024 SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models
Weilong Chai, Jiajiong Cao, Zhiquan Chen, Changbao Wang, Chenguang Ma
ECCV (43)6
2023 Multistability Analysis and Digital Circuit Implementation of a New Conformable Fractional-Order Chaotic System
Chenguang Ma, Jun Mou, Peng Li 0037, Tianming Liu 0005
Mob. Networks Appl.1
2023 A New Meminductor Based Hyperchaotic Circuit and its Implementation
Xujiong Ma, Jun Mou, Chenguang Ma, Jieyang Wang, Tianming Liu 0005
Mob. Networks Appl.3
2022 FaceVerse: a Fine-grained and Detail-controllable 3D Face Morphable Model from a Hybrid Dataset
abstract
We present FaceVerse, a fine-grained 3D Neural Face Model, which is built from hybrid East Asian face datasets containing 60K fused RGB-D images and 2K high-fidelity 3D head scan models. A novel coarse-to-fine structure is proposed to take better advantage of our hybrid dataset. In the coarse module, we generate a base parametric model from large-scale RGB-D images, which is able to predict accurate rough 3D face models in different genders, ages, etc. Then in the fine module, a conditional StyleGAN architecture trained with high-fidelity scan models is introduced to enrich elaborate facial geometric and texture details. Note that different from previous methods, our base and detailed modules are both changeable, which enables an innovative application of adjusting both the basic attributes and the facial details of 3D face models. Furthermore, we propose a single-image fitting framework based on differentiable rendering. Rich experiments show that our method outperforms the state-of-the-art methods.
Lizhen Wang 0002, Tao Yu 0007, Chenguang Ma, Yebin Liu
CVPR4
2022 Adaptive Rectangle Loss for Speaker Verification
Ruida Li, Shuo Fang, Chenguang Ma
INTERSPEECH3
2022 m3Track: mmwave-based multi-user 3D posture tracking
abstract
Nowadays, the market of 3D human posture tracking has extended to a broad range of application scenarios. As current mainstream solutions, vision-based posture tracking systems suffer from privacy leakage concerns and depend on lighting conditions. Towards more privacy-preserving and robust tracking manner, recent works have exploited commodity radio frequency signals to realize 3D human posture tracking. However, these studies cannot handle the case where multiple users are in the same space. In this paper, we present a mmWave-based multi-user 3D posture tracking system, m3Track, which leverages a single commercial off-the-shelf (COTS) mmWave radar to track multiple users' postures simultaneously as they move, walk, or sit. Based on the sensing signals from a mmWave radar in multi-user scenarios, m3Track first separates all the users on mmWave signals. Then, m3Track extracts shape and motion features of each user, and reconstructs 3D human posture for each user through a designed deep learning model. Furthermore. m3Track maps the reconstructed 3D postures of all users into 3D space, and tracks users' positions through a coordinate-corrected tracking method, realizing practical multi-user 3D posture tracking with a COTS mmWave radar. Experiments conducted in real-world multi-user scenarios validate the accuracy and robustness of m3Track on multi-user 3D posture tracking.
Hao Kong 0004, Xiangyu Xu 0001, Jiadi Yu, Qilin Chen, Chenguang Ma, Yingying Chen 0001, Yi-Chao Chen 0001, Linghe Kong
MobiSys5
2021 Coexistence of infinite attractors in a fractional-order chaotic system with two nonlinear functions and its DSP implementation
Xintong Han, Jun Mou, Li Xiong 0016, Chenguang Ma, Tianming Liu 0005, Yinghong Cao
Integr.4
2020 Characteristic analysis of the fractional-order hyperchaotic complex system and its image encryption application
Jun Mou, Jian Liu 0023, Chenguang Ma, Huizhen Yan
Signal Process.4
2014 Transparent Object Reconstruction via Coded Transport of Intensity
abstract
Capturing and understanding visual signals is one of the core interests of computer vision. Much progress has been made w.r.t. many aspects of imaging, but the reconstruc-tion of refractive phenomena, such as turbulence, gas and heat flows, liquids, or transparent solids, has remained a challenging problem. In this paper, we derive an intuitive formulation of light transport in refractive media using light fields and the transport of intensity equation. We show how coded illumination in combination with pairs of recorded images allow for robust computational reconstruction of dy-namic two and three-dimensional refractive phenomena. 1.
Chenguang Ma, Jin-Li Suo, Qionghai Dai, Gordon Wetzstein
CVPR1
2014 Acquisition of High Spatial and Spectral Resolution Video with a Hybrid Camera System
Chenguang Ma, Xun Cao, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001
Int. J. Comput. Vis.1
2013 High-rank coded aperture projection for extended depth of field
abstract
Projectors require large apertures to maximize light throughput. Unfortunately, this leads to shallow depths of field (DOF), hence blurry images, when projecting on non-planar surfaces, such as cultural heritage sites, curved screens, or when sharing visual information in everyday environments. We introduce high-rank coded aperture projectors - a new computational display technology that combines optical designs with computational processing to overcome depth of field limitations of conventional devices. In particular, we employ high-speed spatial light modulators (SLMs) on the image plane and in the aperture of modified projectors. The patterns displayed on these SLMs are computed with a new mathematical framework that uses high-rank light field factorizations and directly exploits the limited temporal resolution and contrast sensitivity of the human visual system. With an experimental prototype projector, we demonstrate significantly increased DOF as compared to conventional technology.
Chenguang Ma, Jin-Li Suo, Qionghai Dai, Ramesh Raskar, Gordon Wetzstein
ICCP1
2010 Enhanced floor control protocol for PoC application in data packet voice communication
abstract
Push-to-Talk over Cellular (PoC) is a one-to-one or one-to-many Push-To-Talk (PTT) service designed to work over the cellular network. In the PTT, a participant may speak as soon as he/she has obtained the permission through pushing a “talk” button. The typical telephone setup delay is avoided. In the traditional floor control protocol defined in the Open Mobile Alliance standard (FCP-OMA) for PoC, at any time only one participant is permitted to speak. The permitted speaker is totally “deaf” and the listeners are fully “dumb” until the speaker releases or is revoked the floor. In this article, we propose an enhanced floor control protocol (EFCP) for PoC application in a packet voice scenario where full-duplex communications can be supported. The proposed protocol is divided into two categories: EFCP-M and EFCP-NM depending on whether media mixing is needed or supported. Through a queuing model, we compare the performance of EFCP against FCP-OMA. Experimental results show that our scheme can significantly reduce the probability that floor requests are denied and the expected waiting time in queue.
Hai Jiang 0004, Albert Kai-Sun Wong, Vincent Wing-Hei Luk, Xiyu Duan, Jun Li 0002, Chenguang Ma
ISCC6