EDBT 2026 Demo / reviewers in the wild / expert
Ruobing Zheng
dblp:229/7131
· DBLP profile ↗
14ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EchoMimicV3: 1.3B Parameters Are All You Need for Unified Multi-Modal and Multi-Task Human AnimationabstractRecent work on human animation usually incorporates large-scale video models, thereby achieving more vivid performance. However, the practical use of such methods is hindered by the slow inference speed and high computational demands. Moreover, traditional work typically employs separate models for each animation task, increasing costs in multi-task scenarios and worsening the dilemma. To address these limitations, we introduce EchoMimicV3, an efficient framework that unifies multi-task and multi-modal human animation. At the core of EchoMimicV3 lies a threefold design: a Soup-of-Tasks paradigm, a Soup-of-Modals paradigm, and a novel training and inference strategy. The Soup-of-Tasks leverages multi-task mask inputs and a counter-intuitive task allocation strategy to achieve multi-task gains without multi-model pains. Meanwhile, the Soup-of-Modals introduces a Coupled-Decoupled Multi-Modal Cross Attention module to inject multi-modal conditions, complemented by a Timestep Phase-aware Multi-Modal Allocation mechanism to dynamically modulate multi-modal mixtures. Besides, we propose Negative Direct Preference Optimization and Phase-aware Negative Classifier-Free Guidance, which ensure stable training and inference. Extensive experiments and analyses demonstrate that EchoMimicV3, with a minimal model size of 1.3 billion parameters, achieves competitive performance in both quantitative and qualitative evaluations. We are committed to open-sourcing our code for community use. Rang Meng, Weipeng Wu, Ruobing Zheng, Chenguang Ma |
AAAI | 4 |
| 2026 | HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMsabstractWhile Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of empathetic, context-aware responses. Here we introduce HumanSense, a comprehensive benchmark designed to evaluate the human-centered perception and interaction capabilities of MLLMs, with a particular focus on deep understanding of extended multimodal contexts and the formulation of rational feedback. Our evaluation reveals that leading MLLMs still have considerable room for improvement, particularly for advanced interaction-oriented tasks. Supplementing visual input with audio and text information yields substantial improvements, and Omni-modal models show advantages on these tasks.Furthermore, grounded in the observation that appropriate feedback stems from a contextual analysis of the interlocutor's needs and emotions, we posit that reasoning ability serves as the key to unlocking it. We devise a multi-stage, modality-progressive reinforcement learning approach, resulting in HumanSense-Omni-Reasoning, which substantially enhances performance on higher-level understanding and interactive tasks. Additionally, we observe that successful reasoning processes appear to exhibit consistent thought patterns. By designing corresponding prompts, we also enhance the performance of non-reasoning models in a training-free manner. Ruobing Zheng, Jingdong Chen, Le Wang 0003 |
AAAI | 2 |
| 2025 | Animate-X: Universal Character Image Animation with Enhanced Motion RepresentationabstractCharacter image animation, which generates high-quality videos from a reference image and target pose sequence, has seen significant progress in recent years. However, most existing methods only apply to human figures, which usually do not generalize well on anthropomorphic characters commonly used in industries like gaming and entertainment. Our in-depth analysis suggests to attribute this limitation to their insufficient modeling of motion, which is unable to comprehend the movement pattern of the driving video, thus imposing a pose sequence rigidly onto the target character. To this end, this paper proposes $\texttt{Animate-X}$, a universal animation framework based on LDM for various character types (collectively named $\texttt{X}$), including anthropomorphic characters. To enhance motion representation, we introduce the Pose Indicator, which captures comprehensive motion pattern from the driving video through both implicit and explicit manner. The former leverages CLIP visual features of a driving video to extract its gist of motion, like the overall movement pattern and temporal relations among motions, while the latter strengthens the generalization of LDM by simulating possible inputs in advance that may arise during inference. Moreover, we introduce a new Animated Anthropomorphic Benchmark ($\texttt{$A^2$Bench}$) to evaluate the performance of $\texttt{Animate-X}$ on universal and widely applicable animation images. Extensive experiments demonstrate the superiority and effectiveness of $\texttt{Animate-X}$ compared to state-of-the-art methods. Biao Gong, Xiang Wang 0012, Shiwei Zhang 0001, Ruobing Zheng, Kecheng Zheng, Jingdong Chen, Ming Yang 0007 |
ICLR | 6 |
| 2025 | Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
Ruobing Zheng, Jingdong Chen, Ming Yang 0007 |
ACM Multimedia | 2 |
| 2025 | Versatile Multimodal Controls for Expressive Talking Human Animation
Ruobing Zheng, Zixin Zhu, Sanping Zhou, Ming Yang 0007, Le Wang 0003 |
ACM Multimedia | 2 |
| 2024 | Learning Dynamic Tetrahedra for High-Quality Talking Head SynthesisabstractRecent works in implicit representations, such as Neural Radiance Fields (NeRF), have advanced the generation of realistic and animatable head avatars from video sequences. These implicit methods are still confronted by visual artifacts and jitters, since the lack of explicit geometric constraints poses a fundamental challenge in accurately modeling complex facial deformations. In this paper, we introduce Dynamic Tetrahedra (DynTet), a novel hybrid representation that encodes explicit dynamic meshes by neural networks to ensure geometric consistency across various motions and viewpoints. DynTet is parameterized by the coordinate-based networks which learn signed distance, deformation, and material texture, anchoring the training data into a predefined tetrahedra grid. Leveraging Marching Tetrahedra, DynTet efficiently decodes textured meshes with a consistent topology, enabling fast rendering through a differentiable rasterizer and supervision via a pixel loss. To enhance training efficiency, we incorporate classical 3D Morphable Models to facilitate geometry learning and define a canonical space for simplifying texture learning. These advantages are readily achievable owing to the effective geometric representation employed in DynTet. Compared with prior works, DynTet demonstrates significant improvements in fidelity, lip synchronization, and real-time performance according to various metrics. Beyond producing stable and visually appealing synthesis videos, our method also outputs the dynamic meshes which is promising to enable many emerging applications. Code is available at https://github.com/zhangzc21/DynTet. Ruobing Zheng, Bonan Li, Congying Han, Tiande Guo, Jingdong Chen, Ziwen Liu 0001, Ming Yang 0007 |
CVPR | 2 |
| 2024 | StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models
Wen Li 0024, Muyuan Fang, Biao Gong, Ruobing Zheng, Jingdong Chen, Ming Yang 0007 |
ECCV (28) | 5 |
| 2023 | DC-Former: Diverse and Compact Transformer for Person Re-identificationabstractIn person re-identification (ReID) task, it is still challenging to learn discriminative representation by deep learning, due to limited data. Generally speaking, the model will get better performance when increasing the amount of data. The addition of similar classes strengthens the ability of the classifier to identify similar identities, thereby improving the discrimination of representation. In this paper, we propose a Diverse and Compact Transformer (DC-Former) that can achieve a similar effect by splitting embedding space into multiple diverse and compact subspaces. Compact embedding subspace helps model learn more robust and discriminative embedding to identify similar classes. And the fusion of these diverse embeddings containing more fine-grained information can further improve the effect of ReID. Specifically, multiple class tokens are used in vision transformer to represent multiple embedding spaces. Then, a self-diverse constraint (SDC) is applied to these spaces to push them away from each other, which makes each embedding space diverse and compact. Further, a dynamic weight controller (DWC) is further designed for balancing the relative importance among them during training. The experimental results of our method are promising, which surpass previous state-of-the-art methods on several commonly used person ReID benchmarks. Our code is available at https://github.com/ant-research/Diverse-and-Compact-Transformer. Wen Li 0024, Furong Xu, Jianan Zhao 0010, Ruobing Zheng |
AAAI | 6 |
| 2021 | Learning Pose-Adaptive Lip Sync with Cascaded Temporal Convolutional NetworkabstractSpeech-driven lip sync has become a promising technique for generating and editing talking-head videos. These studies mainly use 3D morphable models or 2D facial landmarks as the intermediate face representations. However, 2D-based methods have been stagnant recently due to their inability to handle out-of-plane rotations, even though the 2D landmarks have the advantage of fast and accurate extraction. In this paper, we design a cascaded temporal convolutional network to successively generate mouth shapes and corresponding jawlines based on audio signals and template headposes. Instead of explicitly calibrating the rotation between the predicted mouth and the template face, we employ neural networks to learn the pose-adaptive mapping implicitly. We also propose an image-to-image translation-based neural rendering method for producing high-resolution and photo-realistic videos. Experiments show our solution improves both the mapping accuracy and visual performance than baselines. This work could benefit many real-world applications like virtual anchors, telepresence, and conversational agents. Ruobing Zheng, Changjiang Ji |
ICASSP | 1 |
| 2020 | A Neural Lip-Sync Framework for Synthesizing Photorealistic Virtual News AnchorsabstractLip sync has emerged as a promising technique for generating mouth movements from audio signals. However, synthesizing a high-resolution and photorealistic virtual news anchor is still challenging. Lack of natural appearance, visual consistency, and processing efficiency are the main problems with existing methods. In this paper, we present a novel lip-sync framework specially designed for producing high fidelity virtual news anchors. A pair of Temporal Convolutional Networks are used to learn the cross-modal sequential mapping from audio signals to mouth movements, followed by a neural rendering network that translates the synthetic facial map into high-resolution and photorealistic appearance. This fully-trainable framework provides an end-to-end processing that outperforms traditional graphics-based methods in many low-delay applications. Experiments also show the framework has advantages over modern neural-based methods in both visual appearance and efficiency. Ruobing Zheng, Zhou Zhu, Changjiang Ji |
ICPR | 1 |
| 2020 | Joint Ranking SVM and Binary Relevance with robust Low-rank learning for multi-label classification
Guoqiang Wu, Ruobing Zheng, Yingjie Tian 0001, Dalian Liu |
Neural Networks | 2 |
| 2019 | Exploiting Time-Series Image-to-Image Translation to Expand the Range of Wildlife Habitat AnalysisabstractCharacterizing wildlife habitat is one of the main topics in animal ecology. Locational data obtained from radio tracking and field observation are widely used in habitat analysis. However, such sampling methods are costly and laborious, and insufficient relocations often prevent scientists from conducting large-range and long-term research. In this paper, we innovatively exploit the image-to-image translation technology to expand the range of wildlife habitat analysis. We proposed a novel approach for implementing time-series imageto-image translation via metric embedding. A siamese neural network is used to learn the Euclidean temporal embedding from the image space. This embedding produces temporal vectors which bring time information into the adversarial network. The well-trained framework could effectively map the probabilistic habitat models from remote sensing imagery, helping scientists get rid of the persistent dependence on animal relocations. We illustrate our approach in a real-world application for mapping the habitats of Bar-headed Geese at Qinghai Lake breeding ground. We compare our model against several baselines and achieve promising results. Ruobing Zheng, Ze Luo, Baoping Yan |
AAAI | 1 |
| 2019 | Building and Dynamically Managing Workflows for Processing Remote Sensing Data in Distributed High-Throughput EnvironmentabstractThe growing number of remote sensing data create more opportunity for scientific studies in various disciplines, along with new challenges to data processing techniques. Most of existing solutions show limitations at either throughput or automation level. In this paper, we propose an efficient framework for building and dynamically managing remote sensing data processing workflows in a distributed high-throughput environment. We use HTCondor system to allocate computing resource and schedule tasks while developing Python modules for processing operations and workflow management. The proposed framework not only illustrate a general way to process mass remote sensing data in a high-throughput manner but also demonstrates a novel solution to dynamically manage scientific workflows in the HTCondor system. Ruobing Zheng, Yingchao Piao, Ze Luo, Baoping Yan, Miron Livny |
IGARSS | 1 |
| 2018 | Investigating Waterfowl Habitat-Use Patterns with Multi-Source Remote Sensing DataabstractWaterfowl habitat analysis is significant to understand species behavior and make conservation plans, especially for Bar-headed Geese, which was involved in the large-scale outbreak of highly pathogenic avian influenza H5Nl in the year 2005 in China. Many studies have demonstrated there is a significant correlation between wildlife habitat and remote sensing data. The various reflectance data contain substantial ecological information that is valuable to model the habitat selection of wildlife. In this paper, we investigate the habitat use patterns of Bar-headed Geese by combining multi-source satellite images with bird GPS records, using Log-likelihood chi-square test to explore the waterfowl habitat preferences. The results show the bird's favorites are significant in various habitatcat-egories, which confirm previous surveys. This work helps to manage species and make disease control strategies for this sensitive waterfowl. Ruobing Zheng, Ze Luo, Baoping Yan |
IGARSS | 1 |