VLDB 2026 Research / reviewers in the wild / expert
Xintian Wu
dblp:06/7160
· DBLP profile ↗
11ranked-venue papers
6as first author
5since 2021 · last 2023
0000-0002-2988-5215ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | D3T-GAN: Data-Dependent Domain Transfer GANs for Image Generation with Limited DataabstractAs an important and challenging problem, image generation with limited data aims at generating realistic images through training a GAN model given few samples. A typical solution is to transfer a well-trained GAN model from a data-rich source domain to the data-deficient target domain. In this paper, we propose a novel self-supervised transfer scheme termed D 3 T-GAN, addressing the cross-domain GANs transfer in limited image generation. Specifically, we design two individual strategies to transfer knowledge between generators and discriminators, respectively. To transfer knowledge between generators, we conduct a data-dependent transformation, which projects target samples into the latent space of source generator and reconstructs them back. Then, we perform knowledge transfer from transformed samples to generated samples. To transfer knowledge between discriminators, we design a multi-level discriminant knowledge distillation from the source discriminator to the target discriminator on both the real and fake samples. Extensive experiments show that our method improves the quality of generated images and achieves the state-of-the-art FID scores on commonly used datasets. Xintian Wu, Yiming Wu 0005, Xi Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Saliency Hierarchy Modeling via Generative Kernels for Salient Object Detection
Wenhu Zhang, Liangli Zheng, Xintian Wu, Xi Li 0001 |
ECCV (28) | 4 |
| 2022 | Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image GenerationabstractAs a challenging task, text-to-image generation aims to generate photo-realistic and semantically consistent images according to the given text descriptions. Existing methods mainly extract the text information from only one sentence to represent an image and the text representation effects the quality of the generated image well. However, directly utilizing the limited information in one sentence misses some key attribute descriptions, which are the crucial factors to describe an image accurately. To alleviate the above problem, we propose an effective text representation method with the complements of attribute information. Firstly, we construct an attribute memory to jointly control the text-to-image generation with sentence input. Secondly, we explore two update mechanisms, sample-aware and sample-joint mechanisms, to dynamically optimize a generalized attribute memory. Furthermore, we design an attribute-sentence-joint conditional generator learning scheme to align the feature embeddings among multiple representations, which promotes the cross-modal network training. Experimental results illustrate that the proposed method obtains substantial performance improvements on both the CUB (FID from 14.81 to 8.57) and COCO (FID from 21.42 to 12.39) datasets. Xintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding, Xi Li 0001 |
ACM Multimedia | 1 |
| 2021 | MGH: Metadata Guided Hypergraph Modeling for Unsupervised Person Re-identificationabstractAs a challenging task, unsupervised person ReID aims to match the same identity with query images which does not require any labeled information. In general, most existing approaches focus on the visual cues only, leaving potentially valuable auxiliary metadata information (e.g., spatio-temporal context) unexplored. In the real world, such metadata is normally available alongside captured images, and thus plays an important role in separating several hard ReID matches. With this motivation in mind, we propose MGH, a novel unsupervised person ReID approach that uses meta information to construct a hypergraph for feature learning and label refinement. In principle, the hypergraph is composed of camera-topology-aware hyperedges, which can model the heterogeneous data correlations across cameras. Taking advantage of label propagation on the hypergraph, the proposed approach is able to effectively refine the ReID results, such as correcting the wrong labels or smoothing the noisy labels. Given the refined results, we further present a memory-based listwise loss to directly optimize the average precision in an approximate manner. Extensive experiments on three benchmarks demonstrate the effectiveness of the proposed approach against the state-of-the-art. Yiming Wu 0005, Xintian Wu, Xi Li 0001 |
ACM Multimedia | 2 |
| 2021 | F³A-GAN: Facial Flow for Face Animation With Generative Adversarial NetworksabstractFormulated as a conditional generation problem, face animation aims at synthesizing continuous face images from a single source image driven by a set of conditional face motion. Previous works mainly model the face motion as conditions with 1D or 2D representation (e.g., action units, emotion codes, landmark), which often leads to low-quality results in some complicated scenarios such as continuous generation and large-pose transformation. To tackle this problem, the conditions are supposed to meet two requirements, i.e., motion information preserving and geometric continuity. To this end, we propose a novel representation based on a 3D geometric flow, termed facial flow, to represent the natural motion of the human face at any pose. Compared with other previous conditions, the proposed facial flow well controls the continuous changes to the face. After that, in order to utilize the facial flow for face editing, we build a synthesis framework generating continuous images with conditional facial flows. To fully take advantage of the motion information of facial flows, a hierarchical conditional framework is designed to combine the extracted multi-scale appearance features from images and motion features from flows in a hierarchical manner. The framework then decodes multiple fused features back to images progressively. Experimental results demonstrate the effectiveness of our method compared to other state-of-the-art methods. Xintian Wu, Qihang Zhang, Yiming Wu 0005, Lingyun Sun, Xi Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Semantic Neighborhood-Aware Deep Facial Expression RecognitionabstractDifferent from many other attributes, facial expression can change in a continuous way, and therefore, a slight semantic change of input should also lead to the output fluctuation limited in a small scale. This consistency is important. However, current Facial Expression Recognition (FER) datasets may have the extreme imbalance problem, as well as the lack of data and the excessive amounts of noise, hindering this consistency and leading to a performance decreasing when testing. In this paper, we not only consider the prediction accuracy on sample points, but also take the neighborhood smoothness of them into consideration, focusing on the stability of the output with respect to slight semantic perturbations of the input. A novel method is proposed to formulate semantic perturbation and select unreliable samples during training, reducing the bad effect of them. Experiments show the effectiveness of the proposed method and state-of-the-art results are reported, getting closer to an upper limit than the state-of-the-art methods by a factor of 30% in AffectNet, the largest in-the-wild FER database by now. Yongjian Fu 0002, Xintian Wu, Xi Li 0001, Daxin Luo |
IEEE Trans. Image Process. | 2 |
| 2004 | Speaker adaptation using constrained transformationabstractIn speech recognition research, transformation-based adaptation algorithms provide an effective way of adapting acoustic models to improve the recognition accuracy. However, when only limited amounts of adaptation data are available, the transformation is often poorly estimated, which may cause performance degradation. This paper presents the Markov Random Field Linear Regression (MRFLR) algorithm, which constrains the transformation-based adaptation by the correlations among acoustic parameters. The Markov Random Field theory is used to model the correlations. The correlations are estimated from the training corpus and hypothesized as prior knowledge of acoustic models. By explicitly incorporating them into adaptation, robust and fast adaptation can be achieved. The hypothesis is tested by comparing MRFLR with MLLR (Maximum Likelihood Linear Regression), a widely used transformation-based adaptation algorithm. Experimental results show that MRFLR outperforms MLLR when adaptation data are sparse, and converges to the MLLR performance when more adaptation data are available. Xintian Wu, Yonghong Yan 0002 |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Linear regression under maximum a posteriori criterion with Markov random field priorabstractSpeaker adaptation using linear transformations under the maximum a posteriori (MAP) criterion has been studied in this paper. The purpose is to improve the matrix estimation in the widely used maximum likelihood linear regression (MLLR) adaptation, which might generate poorly structured transform matrices when adaptation data are sparse. Unlike traditional MAP based adaptations, many known prior distributions of HMM parameters, such as normal-Washart priors, do not have a close form solution in the transform estimation. In Markov random field linear regression (MRFLR), the prior distribution of HMM parameters is modeled by Markov random field, which leads to a close form solution of estimating the linear transforms. Experimental results show that MRFLR outperforms MLLR when adaptation data are sparse, and converges to the MLLR performances when more adaptation data are available. Xintian Wu, Yonghong Yan 0002 |
ICASSP | 1 |
| 1999 | High accuracy acoustic modeling based on multi-stage decision tree
Chaojun Liu, Xintian Wu, Yonghong Yan 0002 |
EUROSPEECH | 3 |
| 1999 | High accuracy acoustic modeling using two-level decision-tree based state-tyingabstractThis paper addresses the problem of language modeling for the transcription of broadcast news data. Different approaches for language model training were explored and tested in the context of a complete transcription system. Language model efficiency was investigated for the following aspects: mixing of different training material (sources and epoch); approach for mixing (interpolation vs count merging); and using class-based language models. The experimental results indicate that judicious selection of the training source and epoch is important, and that given sufficient broadcast new transcriptions, newspaper and newswire texts are not necessary. Results are given in terms of perplexity and word error rates. The combined improvements in text selection, interpolation, 4-gram and class-based LMs led to a 20% reduction in the perplexity of the LM of the final pass (3-gram class interpolated with a word 4-gram) compared with the 3-gram LM used in the the LIMSI Nov’97 BN system. Chaojun Liu, Xintian Wu, Yonghong Yan 0002 |
EUROSPEECH | 2 |
| 1999 | Development of the 1998 OGI-FONIX broadcast news transcription system
Xintian Wu, Yonghong Yan 0002 |
EUROSPEECH | 1 |