EDBT 2026 Demo / reviewers in the wild / expert
Anni Tang
dblp:311/2494
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-9772-3293ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hybrid Scheme for Face Video CompressionabstractWith the rapid development of social media, the amount of face video data has grown rapidly, making face video compression a hot research topic. Traditional video coding techniques do not discriminate video content and compress all videos in the same way, while talking head video compression should have more potential. Existing generative compression methods mostly adopt static reference frames, resulting in a decrease in fidelity caused by dynamic background or large pose change. In this article, we propose a hybrid compression scheme for face videos which combines traditional coding with generative compression. On the one hand, we sample and encode key frames with traditional codecs to provide dynamic reference frames which contain real-time background and motion information. On the other hand, we devise a deep video generation model to synthesize smooth video frames according to the extracted sparse keypoints. Combining the pixel-level recovery capability of traditional coding with the detail generation capability of deep generative models, our proposed hybrid scheme is able to implement high-fidelity face video compression at low bitrate in real time. Additionally, we also devise a Portrait Recovery module to recover the low-quality key frames, improving the reconstruction quality in low-bitrate scenarios. Extensive experiments show that our method has advantages over traditional codecs and existing generative compression methods in terms of both rate-distortion performance and coding complexity. Anni Tang, Zhiyu Zhang 0010, Jun Ling, Rong Xie 0004, Li Song 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | SingAvatar: High-fidelity Audio-driven Singing Avatar SynthesisabstractGenerating photo-realistic avatars from audio plays an important role in extended reality (XR) and metaverse. In this paper, we lift the input audio from speech to singing, which has been rarely studied. The significant distinction between singing and talking poses great challenges for adapting talking face generation methods to the singing regime. To address this, we propose a high-fidelity singing avatar synthesis method called SingAvatar. Besides the audio, we incorporate vocal conditions involving phonemes and variance to alleviate the ambiguity of learning the singing-to-face mapping. Concretely, we tailor a two-stage pipeline: singing voice synthesis and portrait generation from the synthesized audio and auxiliary vocal conditions. Further, we curate a fine-grained singing head dataset containing singing videos with synchronized audio and accurate vocal conditions. In experiments, SingAvatar outperforms competing methods regarding audio-mouth synchronization, the naturalness of head movements, and controllability over the results. The code and dataset will be made publicly available. Anni Tang, Jun Ling, Huiheng Liao, Yunhui Zhu, Li Song 0001 |
ICME | 2 |
| 2024 | Efficient Dynamic-NeRF Based Volumetric Video Coding with Rate Distortion OptimizationabstractVolumetric videos, benefiting from immersive 3D realism and interactivity, hold vast potential for various applications, while the tremendous data volume poses significant challenges for compression. Recently, NeRF has demonstrated remarkable potential in volumetric video compression thanks to its simple representation and powerful 3D modeling capabilities, where a notable work is ReRF. However, ReRF separates the modeling from compression process, resulting in suboptimal compression efficiency. In contrast, in this paper, we propose a volumetric video compression method based on dynamic NeRF in a more compact manner. Specifically, we decompose the NeRF representation into the coefficient fields and the basis fields, incrementally updating the basis fields in the temporal domain to achieve dynamic modeling. Additionally, we perform end-to-end joint optimization on the modeling and compression process to further improve the compression efficiency. Extensive experiments demonstrate that our method achieves higher compression efficiency compared to ReRF on various datasets. Zhiyu Zhang 0010, Guo Lu, Huanxiong Liang, Anni Tang, Qiang Hu 0003, Li Song 0001 |
ICME | 4 |
| 2024 | Rate-aware Compression for NeRF-based Volumetric VideoabstractThe neural radiance fields (NeRF) have advanced the development of 3D volumetric video technology, but the large data volumes they involve pose significant challenges for storage and transmission. To address these problems, the existing solutions typically compress these NeRF representations after the training stage, leading to a separation between representation training and compression. In this paper, we try to directly learn a compact NeRF representation for volumetric video in the training stage based on the proposed rate-aware compression framework. Specifically, for volumetric video, we use a simple yet effective modeling strategy to reduce temporal redundancy for the NeRF representation. Then, during the training phase, an implicit entropy model is utilized to estimate the bitrate of the NeRF representation. This entropy model is then encoded into the bitstream to assist in the decoding of the NeRF representation. This approach enables precise bitrate estimation, thereby leading to a compact NeRF representation.Furthermore, we propose an adaptive quantization strategy and learn the optimal quantization step for the NeRF representations. Finally, the NeRF representation can be optimized by using the rate-distortion trade-off. Our proposed compression framework can be used for different representations and experimental results demonstrate that our approach significantly reduces the storage size with marginal distortion and achieves state-of-the-art rate-distortion performance for volumetric video on the HumanRF and ReRF datasets. Compared to the previous state-of-the-art method TeTriRF, we achieved an approximately -80% BD-rate on the HumanRF dataset and -60% BD-rate on the ReRF dataset. Zhiyu Zhang 0010, Guo Lu, Huanxiong Liang, Zhengxue Cheng, Anni Tang, Li Song 0001 |
ACM Multimedia | 5 |
| 2024 | Compositional 3D-aware Video Generation with LLM DirectorabstractSignificant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual elements within the generated video, such as the movement and appearance of specific characters and the manipulation of viewpoints. In this work, we propose a novel paradigm that generates each element in 3D representation separately and then composites them with priors from Large Language Models (LLMs) and 2D diffusion models. Specifically, given an input textual query, our scheme consists of four stages: 1) we leverage the LLMs as the director to first decompose the complex query into several sub-queries, where each sub-query describes each element of the generated video; 2) to generate each element, pre-trained models are invoked by the LLMs to obtain the corresponding 3D representation; 3) to composite the generated 3D representations, we prompt multi-modal LLMs to produce coarse guidance on the scale, location, and trajectory of different objects; 4) to make the results adhere to natural distribution, we further leverage 2D diffusion priors and use score distillation sampling to refine the composition. Extensive experiments demonstrate that our method can generate high-fidelity videos from text with flexible control over each element. Hanxin Zhu, Tianyu He, Anni Tang, Junliang Guo, Zhibo Chen 0001, Jiang Bian 0002 |
NeurIPS | 3 |
| 2024 | Memories are One-to-Many Mapping Alleviators in Talking Face GenerationabstractTalking face generation aims at generating photo-realistic video portraits of a target person driven by input audio. According to the nature of audio to lip motions mapping, the same speech content may have different appearances even for the same person at different occasions. Such one-to-many mapping problem brings ambiguity during training and thus causes inferior visual results. Although this one-to-many mapping could be alleviated in part by a two-stage framework (i.e., an audio-to-expression model followed by a neural-rendering model), it is still insufficient since the prediction is produced without enough information (e.g., emotions, wrinkles, etc.). In this paper, we propose MemFace to complement the missing information with an implicit memory and an explicit memory that follow the sense of the two stages respectively. More specifically, the implicit memory is employed in the audio-to-expression model to capture high-level semantics in the audio-expression shared space, while the explicit memory is employed in the neural-rendering model to help synthesize pixel-level details. Our experimental results show that our proposed MemFace surpasses all the state-of-the-art results across multiple scenarios consistently and significantly. Anni Tang, Tianyu He, Xu Tan 0003, Jun Ling, Runnan Li, Sheng Zhao 0002, Jiang Bian 0002, Li Song 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | ViCoFace: Learning Disentangled Latent Motion Representations for Visual-Consistent Face ReenactmentabstractUnsupervised face reenactment aims to animate a source image to imitate the motions of a target image while retaining the source portrait’s attributes like facial geometry, identity, hair texture, and background. While prior methods can extract the motion from the target image via compact representations (e.g., keypoints or latent motion bases [ 50 ]), they are not robust in predicting motions that are disentangled with portrait attributes, thus failing to preserve portrait attributes in the cross-subject reenactment. In this work, we propose an effective and cost-efficient face reenactment approach to address this issue. Our approach is highlighted by two major strengths. First, based on the theory of latent motion bases, we disentangle the full-head motion into two parts: the transferable motion and preservable motion and then compose the full motion representation using latent motions from the source image and the target image. Second, to optimize and learn disentangled motions, we introduce an efficient training framework, which features two training strategies: (1) a mixture training strategy that encompasses self-reenactment training and cross-subject training for better motion disentanglement and (2) a multi-path training strategy that improves the visual consistency of portrait attributes. Extensive experiments on widely used benchmarks demonstrate that our method exhibits a remarkable generalization ability compared to state-of-the-art baselines. Project and demos are available at https://junleen.github.io/projects/vicoface . Jun Ling, Anni Tang, Rong Xie 0004, Li Song 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | High-Fidelity Free-View Talking Head Synthesis for Low-Bandwidth Video ConferenceabstractAs video conferencing becomes an indispensable part of human’s daliy life, how to achieve a high-fidelity calling experience under low bandwidth has been a popular and challenging issue. Deep generative models have great potential in low-bandwidth facial video compression due to the excellent generation capability based on abridged information. Nevertheless, exsiting deep generation-based compression methods tend to handle motion information in pure 2D or pseudo 3D space, causing facial distortion when large head poses are encountered. In this paper, we propose a 3D-aware high-fidelity facial video conferencing system based on a parameterized NeRF-based face model. Through the compression of the parameterized face model and the transmisstion of extracted facial parameters, we implement high-fidelity talking head synthesis for video conferencing at an ultra-low bitrate. Additionally, the 3D perception capability of the system allows for viewpoint control over the head, achieving higher interactivity and practicability. Extensive experiments verify the effectiveness of the proposed 3D-aware high-fidelity free-view facial video conferencing system. Zhiyu Zhang 0010, Anni Tang, Guo Lu, Rong Xie 0004, Li Song 0001 |
VCIP | 2 |
| 2023 | High-Fidelity Face Reenactment Via Identity-Matched Correspondence LearningabstractFace reenactment aims to generate an animation of a source face using the poses and expressions from a target face. Although recent methods have made remarkable progress by exploiting generative adversarial networks, they are limited in generating high-fidelity and identity-preserving results due to the inappropriate driving information and insufficiently effective animating strategies. In this work, we propose a novel face reenactment framework that achieves both high-fidelity generation and identity preservation. Instead of sparse face representations (e.g., facial landmarks and keypoints), we utilize the Projected Normalized Coordinate Code (PNCC) to better preserve facial details. We propose to reconstruct the PNCC with the source identity parameters and the target pose and expression parameters estimated by 3D face reconstruction to factor out the target identity. By adopting the reconstructed representation as the driving information, we address the problem of identity mismatch. To effectively utilize the driving information, we establish the correspondence between the reconstructed representation and the source representation based on the features extracted by an encoder network. This identity-matched correspondence is then utilized to animate the source face using a novel feature transformation strategy. The generator network is further enhanced by the proposed geometry-aware skip connection. Once trained, our model can be applied to previously unseen faces without further training or fine-tuning. Through extensive experiments, we demonstrate the effectiveness of our method in face reenactment and show that our model outperforms state-of-the-art approaches both qualitatively and quantitatively. Additionally, the proposed PNCC reconstruction module can be easily inserted into other methods and improve their performance in cross-identity face reenactment. Jun Ling, Anni Tang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Generative Compression for Face Video: A Hybrid SchemeabstractAs the latest video coding standard, versatile video coding (VVC) has shown its ability in retaining pixel quality. To excavate more compression potential for video conference scenarios under ultra-low bitrate, this paper proposes a bitrate-adjustable hybrid compression scheme for face video. This hybrid scheme combines the pixel-level precise recovery capability of traditional coding with the generation capability of deep learning based on abridged information, where Pixel-wise Bi-Prediction, Low-Bitrate-FOM and Lossless Keypoint Encoder collaborate to achieve PSNR up to 36.23 dB at a low bitrate of 1.47 KB/s. Without introducing any additional bi-trate, our method has a clear advantage over VVC under a completely fair comparative experiment, which proves the effectiveness of our proposed scheme. Moreover, our scheme can adapt to any existing encoder/configuration to deal with different encoding requirements, and the bitrate can be dynamically adjusted according to the network condition. Anni Tang, Yan Huang 0033, Jun Ling, Zhiyu Zhang 0010, Rong Xie 0004, Li Song 0001 |
ICME | 1 |
| 2021 | Dense 3D Coordinate Code Prior Guidance for High-Fidelity Face Swapping and Face ReenactmentabstractIn face synthesis tasks, commonly used 2D face representations (e.g. 2D landmarks, segmentation maps, etc.) are usually sparse and discontinuous. To combat these shortcomings, we utilize a dense and continuous representation, named Projected Normalized Coordinate Code (PNCC), as the guidance and develop a PNCC-Spatio-Normalization (PSN) method to achieve face synthesis regarding arbitrary head poses and expressions. Based on PSN, we provide an effective framework for face reenactment and face swapping task. To ensure a harmonious and seamless face swapping, a simple yet effective Appearance-Blending Module (ABM) is proposed to fit the synthesized face to the target face. Our method is subject-agnostic and can be applied to any pair of faces without extra fine-tuning. Both qualitative and quantitative experiments are conducted to demonstrate the superiority of the proposed method in comparisons to existing state-of-the-art systems. Anni Tang, Jun Ling, Rong Xie 0004, Li Song 0001 |
FG | 1 |