EDBT 2026 Demo / reviewers in the wild / expert
Yuanqi Chen
dblp:172/2769
· DBLP profile ↗
17ranked-venue papers
6as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Semantic Facial Descriptors for Accurate Face AnimationabstractFace animation is a challenging task. Existing model-based methods (utilizing 3DMMs or landmarks) often result in a model-like reconstruction effect, which doesn't effectively preserve identity. Conversely, model-free approaches face challenges in attaining a decoupled and semantically rich feature space, thereby making accurate motion transfer difficult to achieve. We introduce the semantic facial descriptors in learnable disentangled vector space to address the dilemma. The approach involves decoupling the facial space into identity and motion subspaces while endowing each of them with semantics by learning complete orthogonal basis vectors. We obtain basis vector coefficients by employing an encoder on the source and driving faces, leading to effective facial descriptors in the identity and motion subspaces. Ultimately, these descriptors can be recombined as latent codes to animate faces. Our approach successfully addresses the issue of model-based methods' limitations in high-fidelity identity and the challenges faced by model-free methods in accurate motion transfer. Extensive experiments are conducted on three challenging benchmarks (i.e. VoxCeleb, HDTF, CelebV). Comprehensive quantitative and qualitative results demonstrate that our model outperforms SOTA methods with superior identity preservation and motion transfer. Yuanqi Chen, Thomas H. Li |
ICASSP | 2 |
| 2024 | Closing the Gap Between Theory and Practice During Alternating Optimization for GANsabstractSynthesizing high-quality and diverse samples is the main goal of generative models. Despite recent great progress in generative adversarial networks (GANs), mode collapse is still an open problem, and mitigating it will benefit the generator to better capture the target data distribution. This article rethinks alternating optimization in GANs, which is a classic approach to training GANs in practice. We find that the theory presented in the original GANs does not accommodate this practical solution. Under the alternating optimization manner, the vanilla loss function provides an inappropriate objective for the generator. This objective forces the generator to produce the output with the highest discriminative probability of the discriminator, which leads to mode collapse in GANs. To address this problem, we introduce a novel loss function for the generator to adapt to the alternating optimization nature. When updating the generator by the proposed loss function, the reverse Kullback-Leibler divergence between the model distribution and the target distribution is theoretically optimized, which encourages the model to learn the target distribution. The results of extensive experiments demonstrate that our approach can consistently boost model performance on various datasets and network structures. Yuanqi Chen, Shangkun Sun, Ge Li 0002, Wei Gao 0003, Thomas H. Li |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Deep Traffic Benchmark: Aerial Perception and Driven Behavior Dataset
Guoxing Zhang, Zhanpeng Wang, Yuanqi Chen, Weiye Zhang, Bingting Guo, Jiasong Zhu |
ACML | 5 |
| 2023 | Flow-Guided Attention Deformation for Person Image GenerationabstractPose-guided person image generation aims to transfer reference images to target poses while preserving the source appearance. Recent approaches achieve considerable improvement by using spatial transformation modules such as attention operation. However, the commonly used vanilla attention tends to generate a dense correlation matrix which means that the value of a target position is the weighted sum of many source positions, resulting in blurry appearance. In this paper, we propose a novel model named Flow-guided Attention Deformation (FAD) to perform the spatial transformation. Our model first establishes the correlation between sources and targets with a flow-guided attention operation. Then, with the obtained correlation matrix, we perform an accurate deformation for source features to generate the predicted image. Extensive results demonstrate the superiority of the proposed method, outperforming state-of-the-art methods quantitatively and qualitatively. Ablation studies clarify the efficiency of the proposed modules and verify our hypothesis. Yubo Wu, Yurui Ren, Yuanqi Chen |
ICME | 3 |
| 2023 | IPFR: Identity-Preserving Face Reenactment with Enhanced Domain Adversarial Training and Multi-level Identity Priors
Ge Li 0002, Yuanqi Chen, Thomas H. Li |
PRCV (10) | 3 |
| 2023 | Mitigating Label Noise in GANs via Enhanced Spectral NormalizationabstractLabel noise is a ubiquitous issue in GANs, which degrades the generalization ability of the discriminator and usually leads to instability when training GANs. This issue stems from both real data and generated data. Previous works either only consider one of these two sources, or are not robust enough to noisy labels. In this paper, we revisit spectral normalization in robust learning with noisy labels. Based on its pros and cons, we propose to combine spectral normalization and weight decay to regularize the discriminator, which enjoys a more robust training process. To extend to conditional GANs, we propose to balance the relative importance of marginal matching and conditional matching in the projection discriminator. The proposed Enhanced Spectral Normalization for Generative Adversarial Networks (ESNGAN) can be easily integrated into various existing GANs frameworks without excessive additional cost. The effectiveness of the proposed method is validated on the CIFAR10, LSUN Church, CelebA, and ImageNet datasets, including the unconditional image generation task and the class-conditional image generation task. We also show that the proposed method can further improve the performance of the high-resolution image generation task. Yuanqi Chen, Cece Jin, Ge Li 0002, Thomas H. Li, Wei Gao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | SKFlow: Learning Optical Flow with Super KernelsabstractOptical flow estimation is a classical yet challenging task in computer vision. One of the essential factors in accurately predicting optical flow is to alleviate occlusions between frames. However, it is still a thorny problem for current top-performing optical flow estimation methods due to insufficient local evidence to model occluded areas. In this paper, we propose the Super Kernel Flow Network (SKFlow), a CNN architecture to ameliorate the impacts of occlusions on optical flow estimation. SKFlow benefits from the super kernels which bring enlarged receptive fields to complement the absent matching information and recover the occluded motions. We present efficient super kernel designs by utilizing conical connections and hybrid depth-wise convolutions. Extensive experiments demonstrate the effectiveness of SKFlow on multiple benchmarks, especially in the occluded areas. Without pre-trained backbones on ImageNet and with a modest increase in computation, SKFlow achieves compelling performance and ranks $\textbf{1st}$ among currently published methods on the Sintel benchmark. On the challenging Sintel clean and final passes (test), SKFlow surpasses the best-published result in the unmatched areas ($7.96$ and $12.50$) by $9.09\%$ and $7.92\%$. The code is available at https://github.com/littlespray/SKFlow. Shangkun Sun, Yuanqi Chen, Yu Zhu 0006, Guodong Guo |
NeurIPS | 2 |
| 2022 | Zero-shot unsupervised image-to-image translation via exploiting semantic attributes
Yuanqi Chen, Xiaoming Yu, Shan Liu 0001, Wei Gao 0003, Ge Li 0002 |
Image Vis. Comput. | 1 |
| 2022 | Large-Scale Spatio-Temporal Person Re-Identification: Algorithms and BenchmarkabstractPerson re-identification (re-ID) in the scenario with large spatial and temporal spans has not been fully explored. This fact partially occurs because existing benchmark datasets were mainly collected with limited spatial and temporal ranges,e.g.,using videos recorded in a few days by cameras in a specific region of the campus. Such limited spatial and temporal ranges make it hard to simulate the difficulties of person re-ID in real scenarios. In this work, we contribute a novel Large-scale Spatio-Temporal (LaST) person re-ID dataset, including 10,862 identities with more than 228k images. Compared with existing datasets, LaST presents more challenging and high-diversity re-ID settings and significantly larger spatial and temporal ranges. For instance, each person can appear in different cities or countries, and in various time slots from day to evening, and in different seasons from spring to winter. To our best knowledge, LaST is a novel person re-ID dataset with the largest spatio-temporal ranges. Based on LaST, we verified its challenge by conducting a comprehensive performance evaluation of 14 re-ID algorithms. We further propose an easy-to-implement baseline that works well in such challenging re-ID settings. We also verified that models pre-trained on LaST can generalize well on existing datasets with short-term and cloth-changing scenarios. We expect LaST to inspire future works toward more realistic and challenging re-ID tasks. More information about the dataset is available athttps://github.com/shuxjweb/last.git. Xiujun Shu, Xiao Wang 0014, Xianghao Zang, Shiliang Zhang, Yuanqi Chen, Ge Li 0002, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | SSD-GAN: Measuring the Realness in the Spatial and Spectral DomainsabstractThis paper observes that there is an issue of high frequencies missing in the discriminator of standard GAN, and we reveal it stems from downsampling layers employed in the network architecture. This issue makes the generator lack the incentive from the discriminator to learn high-frequency content of data, resulting in a significant spectrum discrepancy between generated images and real images. Since the Fourier transform is a bijective mapping, we argue that reducing this spectrum discrepancy would boost the performance of GANs. To this end, we introduce SSD-GAN, an enhancement of GANs to alleviate the spectral information loss in the discriminator. Specifically, we propose to embed a frequency-aware classifier into the discriminator to measure the realness of the input in both the spatial and spectral domains. With the enhanced discriminator, the generator of SSD-GAN is encouraged to learn high-frequency content of real data and generate exact details. The proposed method is general and can be easily integrated into most existing GANs framework without excessive cost. The effectiveness of SSD-GAN is validated on various network architectures, objective functions, and datasets. Code is available at https://github.com/cyq373/SSD-GAN. Yuanqi Chen, Ge Li 0002, Cece Jin, Shan Liu 0001, Thomas H. Li |
AAAI | 1 |
| 2021 | PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingabstractGenerating portrait images by controlling the motions of existing faces is an important task of great consequence to social media industries. For easy use and intuitive control, semantically meaningful and fully disentangled parameters should be used as modifications. However, many existing techniques do not provide such fine-grained controls or use indirect editing methods i.e. mimic motions of other individuals. In this paper, a Portrait Image Neural Renderer (PIRenderer) is proposed to control the face motions with the parameters of three-dimensional morphable face models (3DMMs). The proposed model can generate photo-realistic portrait images with accurate movements according to intuitive modifications. Experiments on both direct and indirect editing tasks demonstrate the superiority of this model. Meanwhile, we further extend this model to tackle the audio-driven facial reenactment task by extracting sequential motions from audio inputs. We show that our model can generate coherent videos with convincing movements from only a single reference image and a driving audio stream. Our source code is available at https://github.com/RenYurui/PIRender. Yurui Ren, Ge Li 0002, Yuanqi Chen, Thomas H. Li, Shan Liu 0001 |
ICCV | 3 |
| 2021 | Structure-transformed Texture-enhanced Network for Person Image SynthesisabstractPose-guided virtual try-on task aims to modify the fashion item based on pose transfer task. These two tasks that belong to person image synthesis have strong correlations and similarities. However, existing methods treat them as two individual tasks and do not explore correlations between them. Moreover, these two tasks are challenging due to large misalignment and occlusions, thus most of these methods are prone to generate unclear human body structure and blurry fine-grained textures. In this paper, we devise a structure-transformed texture-enhanced network to generate high-quality person images and construct the relationships between two tasks. It consists of two modules: structure-transformed renderer and texture-enhanced stylizer. The structure-transformed renderer is introduced to transform the source person structure to the target one, while the texture-enhanced stylizer is served to enhance detailed textures and controllably inject the fashion style founded on the structural transformation. With the two modules, our model can generate photorealistic person images in diverse poses and even with various fashion styles. Extensive experiments demonstrate that our approach achieves state-of-the-art results on two tasks. Munan Xu, Yuanqi Chen, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
ICCV | 2 |
| 2020 | Distributed Video Analysis for Mobile Live Broadcasting ServicesabstractWhile webcast platforms on mobile devices are becoming more and more prevalent, inspection for irregularities is getting harder and harder. To solve this problem, the convolution neural network(CNN) has been applied to recognize or detect specified objections in pictures and videos. However, when supervising large platforms, it isn’t very easy to collect mountain piles of video data and send them to the computation center. Other problems like long time delay and the high computational burden will reduce system performance, especially when dealing with data from live streams. This paper presents a method to coordinate mobile devices with remote servers(computers or embedded systems) to achieve real-time monitoring of live streams. The system can make use of computational capacity on mobile devices and reduce the cost of sending data while guaranteeing accuracy for supervision. Yuanqi Chen, Yongjie Guan, Tao Han 0002 |
WCNC | 1 |
| 2020 | ThermoBench: A thermal efficiency benchmark for clusters in data centers
Yi Zhou 0009, Yuanqi Chen, Shubbhi Taneja, Ajit Chavan, Xiao Qin 0001, Jifu Zhang |
Parallel Comput. | 2 |
| 2019 | ResGAN: A Low-Level Image Processing Network to Restore Original Quality of JPEG Compressed ImagesabstractLow-level image processing is mainly concerned with extracting descriptions (that are usually represented as images themselves) from images. With the rapid development of neural networks, many deep learning-based low-level image processing tasks have shown outstanding performance. In this paper, we describe a unified deep learning based approach for low-level image processing, in particular, image denoising, image deblurring, and compressed image restoration. The proposed method is composed of deep convolutional neural and conditional generative adversarial networks. For the discriminator network, we present a new network architecture with bi-skip connections to address hard training and details losing issues. In the generative network, a multi-objective optimization is derived to solve the problem of common conditions being non-identical. Through extensive experiments on three low-level image processing tasks on both qualitative and quantitative criteria, we demonstrate that our proposed method performs favorably against all current state-of-the-art approaches. Chunbiao Zhu, Yuanqi Chen, Shan Liu 0001, Ge Li 0002 |
DCC | 2 |
| 2019 | Multi-mapping Image-to-Image Translation via Learning DisentanglementabstractRecent advances of image-to-image translation focus on learning the one-to-many mapping from two aspects: multi-modal translation and multi-domain translation. However, the existing methods only consider one of the two perspectives, which makes them unable to solve each other's problem. To address this issue, we propose a novel unified model, which bridges these two objectives. First, we disentangle the input images into the latent representations by an encoder-decoder architecture with a conditional adversarial training in the feature space. Then, we encourage the generator to learn multi-mappings by a random cross-domain translation. As a result, we can manipulate different parts of the latent representations to perform multi-modal and multi-domain translations simultaneously. Experiments demonstrate that our method outperforms state-of-the-art methods. Xiaoming Yu, Yuanqi Chen, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
NeurIPS | 2 |
| 2017 | aHDFS: An Erasure-Coded Data Archival System for Hadoop ClustersabstractIn this paper, we propose an erasure-coded data archival system called aHDFS for Hadoop clusters, where RS(k + r; k) codes are employed to archive data replicas in the Hadoop distributed file system or HDFS. We develop two archival strategies (i.e., aHDFS-Grouping and aHDFS-Pipeline) in aHDFSto speed up the data archival process. aHDFS-Groupinga MapReduce-based data archiving scheme - keeps each mapper's intermediate output Key-Value pairs in a local key-value store. With the local store in place, aHDFS-Grouping merges all the intermediate key-value pairs with the same key into one single key-value pair, followed by shuffling the single Key-Value pair to reducers to generate final parity blocks. aHDFS-Pipeline forms a data archival pipeline using multiple data node in a Hadoop cluster. aHDFS-Pipeline delivers the merged single key-value pair to a subsequent node's local key-value store. Last node in the pipeline is responsible for outputting parity blocks. We implement aHDFS in a real-world Hadoop cluster. The experimental results show that aHDFS-Grouping and aHDFS-Pipeline speed up Baseline's shuffle and reduce phases by a factor of 10 and 5, respectively. When block size is larger than 32 MB, aHDFS improves the performance of HDFS-RAID and HDFS-EC by approximately 31.8 and 15.7 percent, respectively. Yuanqi Chen, Yi Zhou 0009, Shubbhi Taneja, Xiao Qin 0001, Jianzhong Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |