EDBT 2026 Demo / reviewers in the wild / expert
Xing Di
dblp:191/9363
· DBLP profile ↗
12ranked-venue papers
4as first author
10since 2021 · last 2024
0000-0001-7232-2330ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Fine-Grained Rebalancing of Datasets for Correct Demographic ClassificationabstractThe use of face biometrics to automatically recognize people, though fascinating, often raises ethical concerns related to the composition of the datasets used for performance evaluation of the recognition approaches and the bias that stems from the possible demographic imbalance of age, gender, and ethnicity classes during the training phase. This study tackles such imbalance in face datasets and proposes an approach to fair age, gender, and ethnicity classification by training on finely rebalanced cohorts of face images. Special attention is devoted to the ethical aspects related to having face samples of real people vs. having synthetic face samples (generated with a Generative Adversarial Network - GAN - model). Therefore, the dataset rebalancing approach exploits synthetic images instead of real ones in order to decrease the possible privacy concerns raised by new image captures. The work further aims to demonstrate that gross re-balancing is insufficient to solve all the problems related to fair demographic classification, but a finer strategy is worth adopting. The experiments compare rebalancing the single demographic classes with a finer strategy considering classes characterized by combinations of features. This entails analyzing the imbalance of the different cohorts in the dataset and appropriately rebalancing them to evaluate the new performance. Andrea Bozzitelli, Pia Cavasinni di Benedetto, Maria De Marsico, Xing Di, Vishal M. Patel |
CBMI | 4 |
| 2024 | ProS: Facial Omni-Representation Learning via Prototype-based Self-DistillationabstractThis paper presents a novel approach, called Prototype-based Self-Distillation (ProS), for unsupervised face representation learning. The existing supervised methods heavily rely on a large amount of annotated training facial data, which poses challenges in terms of data collection and privacy concerns. To address these issues, we propose ProS, which leverages a vast collection of unlabeled face images to learn a comprehensive facial omni-representation. In particular, ProS consists of two vision-transformers (teacher and student models) that are trained with different augmented images (cropping, blurring, coloring, etc.). Besides, we build a face-aware retrieval system along with augmentations to obtain the curated images comprising predominantly facial areas. To enhance the discrimination of learned features, we introduce a prototype-based matching loss that aligns the similarity distributions between features (teacher or student) and a set of learnable prototypes. After pre-training, the teacher vision transformer serves as a backbone for downstream tasks, including attribute estimation, expression recognition, and landmark alignment, achieved through simple fine-tuning with additional layers. Extensive experiments demonstrate that our method achieves state-of-the-art performance on various tasks, both in full and few-shot settings. Further, we investigate pre-training with synthetic face images, and ProS exhibits promising performance in this scenario as well. Xing Di, Yiyu Zheng, Xiaoming Liu 0002, Yu Cheng 0001 |
WACV | 1 |
| 2024 | Transform-Equivariant Consistency Learning for Temporal Sentence GroundingabstractThis paper addresses the temporal sentence grounding (TSG). Although existing methods have made decent achievements in this task, they not only severely rely on abundant video-query paired data for training, but also easily fail into the dataset distribution bias. To alleviate these limitations, we introduce a novel Equivariant Consistency Regulation Learning (ECRL) framework to learn more discriminative query-related frame-wise representations for each video, in a self-supervised manner. Our motivation comes from that the temporal boundary of the query-guided activity should be consistently predicted under various video-level transformations. Concretely, we first design a series of spatio-temporal augmentations on both foreground and background video segments to generate a set of synthetic video samples. In particular, we devise a self-refine module to enhance the completeness and smoothness of the augmented video. Then, we present a novel self-supervised consistency loss (SSCL) applied on the original and augmented videos to capture their invariant query-related semantic by minimizing the KL-divergence between the sequence similarity of two videos and a prior Gaussian distribution of timestamp distance. At last, a shared grounding head is introduced to predict the transform-equivariant query-guided segment boundaries for both the original and augmented videos. Extensive experiments on three challenging datasets (ActivityNet, TACoS, and Charades-STA) demonstrate both effectiveness and efficiency of our proposed ECRL framework. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Zichuan Xu, Haozhao Wang, Xing Di, Weining Lu, Yu Cheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | Hypotheses Tree Building for One-Shot Temporal Sentence LocalizationabstractGiven an untrimmed video, temporal sentence localization (TSL) aims to localize a specific segment according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on dense video frame annotations, which require a tremendous amount of human effort to collect. In this paper, we target another more practical and challenging setting: one-shot temporal sentence localization (one-shot TSL), which learns to retrieve the query information among the entire video with only one annotated frame. Particularly, we propose an effective and novel tree-structure baseline for one-shot TSL, called Multiple Hypotheses Segment Tree (MHST), to capture the query-aware discriminative frame-wise information under the insufficient annotations. Each video frame is taken as the leaf-node, and the adjacent frames sharing the same visual-linguistic semantics will be merged into the upper non-leaf node for tree building. At last, each root node is an individual segment hypothesis containing the consecutive frames of its leaf-nodes. During the tree construction, we also introduce a pruning strategy to eliminate the interference of query-irrelevant nodes. With our designed self-supervised loss functions, our MHST is able to generate high-quality segment hypotheses for ranking and selection with the query. Experiments on two challenging datasets demonstrate that MHST achieves competitive performance compared to existing methods. Daizong Liu, Pan Zhou 0001, Xing Di, Weining Lu, Yu Cheng 0001 |
AAAI | 4 |
| 2022 | Memory-Guided Semantic Learning Network for Temporal Sentence GroundingabstractTemporal sentence grounding (TSG) is crucial and fundamental for video understanding. Although existing methods train well-designed deep networks with large amount of data, we find that they can easily forget the rarely appeared cases during training due to the off-balance data distribution, which influences the model generalization and leads to unsatisfactory performance. To tackle this issue, we propose a memory-augmented network, called Memory-Guided Semantic Learning Network (MGSL-Net), that learns and memorizes the rarely appeared content in TSG task. Specifically, our proposed model consists of three main parts: cross-modal interaction module, memory augmentation module, and heterogeneous attention module. We first align the given video-query pair by a cross-modal graph convolutional network, and then utilize memory module to record the cross-modal shared semantic features in the domain-specific persistent memory. During training, the memory slots are dynamically associated with both common and rare cases, alleviating the forgetting issue. In testing, the rare cases can thus be enhanced by retrieving the stored memories, leading to better generalization. At last, the heterogeneous attention module is utilized to integrate the enhanced multi-modal features in both video and query domains. Experimental results on three benchmarks show the superiority of our method on both effectiveness and efficiency, which substantially improves the accuracy not only on the entire dataset but also on the rare cases. Daizong Liu, Xiaoye Qu, Xing Di, Yu Cheng 0001, Zichuan Xu, Pan Zhou 0001 |
AAAI | 3 |
| 2022 | Unsupervised Temporal Video Grounding with Deep Semantic ClusteringabstractTemporal video grounding (TVG) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on abundant video-query paired data, which is expensive to collect in real-world scenarios. In this paper, we explore whether a video grounding model can be learned without any paired annotations. To the best of our knowledge, this paper is the first work trying to address TVG in an unsupervised setting. Considering there is no paired supervision, we propose a novel Deep Semantic Clustering Network (DSCNet) to leverage all semantic information from the whole query set to compose the possible activity in each video for grounding. Specifically, we first develop a language semantic mining module, which extracts implicit semantic features from the whole query set. Then, these language semantic features serve as the guidance to compose the activity in video via a video-based semantic aggregation module. Finally, we utilize a foreground attention branch to filter out the redundant background activities and refine the grounding results. To validate the effectiveness of our DSCNet, we conduct experiments on both ActivityNet Captions and Charades-STA datasets. The results demonstrate that our DSCNet achieves competitive performance, and even outperforms most weakly-supervised approaches. Daizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di, Yu Cheng 0001, Zichuan Xu, Pan Zhou 0001 |
AAAI | 4 |
| 2022 | Backdoor Attacks on Crowd CountingabstractCrowd counting is a regression task that estimates the number of people in a scene image, which plays a vital role in a range of safety-critical applications, such as video surveillance, traffic monitoring and flow control. In this paper, we investigate the vulnerability of deep learning based crowd counting models to backdoor attacks, a major security threat to deep learning. A backdoor attack implants a backdoor trigger into a target model via data poisoning so as to control the model's predictions at test time. Different from image classification models on which most of existing backdoor attacks have been developed and tested, crowd counting models are regression models that output multi-dimensional density maps, thus requiring different techniques to manipulate. In this paper, we propose two novel Density Manipulation Backdoor Attacks (DMBA- and DMBA+) to attack the model to produce arbitrarily large or small density estimations. Experimental results demonstrate the effectiveness of our DMBA attacks on five classic crowd counting models and four types of datasets. We also provide an in-depth analysis of the unique challenges of backdooring crowd counting models and reveal two key elements of effective attacks: 1) full and dense triggers and 2) manipulation of the ground truth counts or density maps. Our work could help evaluate the vulnerability of crowd counting models to potential backdoor attacks. Tailai Zhang, Xingjun Ma, Pan Zhou 0001, Jian Lou 0001, Zichuan Xu, Xing Di, Yu Cheng 0001, Lichao Sun 0001 |
ACM Multimedia | 7 |
| 2021 | Heterogeneous Face Frontalization via Domain Agnostic LearningabstractRecent advances in deep convolutional neural networks (DCNNs) have shown impressive performance improvements on thermal to visible face synthesis and matching problems. However, current DCNN-based synthesis models do not perform well on thermal faces with large pose variations. In order to deal with this problem, heterogeneous face frontal-ization methods are needed in which a model takes a thermal profile face image and generates a frontal visible face. This is an extremely difficult problem due to the large domain as well as large pose discrepancies between the two modalities. Despite its applications in biometrics and surveillance, this problem is relatively unexplored in the literature. We propose a domain agnostic learning-based generative adversarial network (DAL-GAN) which can synthesize frontal views in the visible domain from thermal faces with pose variations. DAL-GAN consists of a generator with an auxiliary classifier and two discriminators which capture both local and global texture discriminations for better synthesis. A contrastive constraint is enforced in the latent space of the generator with the help of a dual-path training strategy, which improves the feature vector's discrimination. Finally, a multi-purpose loss function is utilized to guide the network in synthesizing identity-preserving cross-domain frontalization. Extensive experimental results demonstrate that DAL-GAN can generate better quality frontal views compared to the other baseline methods. Xing Di, Shuowen Hu, Vishal M. Patel |
FG | 1 |
| 2021 | DeepStationing: Thoracic Lymph Node Station Parsing in CT Scans Using Anatomical Context Encoding and Key Organ Auto-Search
Dazhou Guo, Xianghua Ye, Jia Ge, Xing Di, Le Lu 0001, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Zhongjie Lu, Senxiang Yan, Dakai Jin |
MICCAI (5) | 4 |
| 2021 | A Large-Scale, Time-Synchronized Visible and Thermal Face DatasetabstractThermal face imagery, which captures the naturally emitted heat from the face, is limited in availability compared to face imagery in the visible spectrum. To help address this scarcity of thermal face imagery for research and algorithm development, we present the DEVCOM Army Research Laboratory Visible-Thermal Face Dataset (ARL-VTF). With over 500,000 images from 395 subjects, the ARL-VTF dataset represents, to the best of our knowledge, the largest collection of paired visible and thermal face images to date. The data was captured using a modern long wave infrared (LWIR) camera mounted alongside a stereo setup of three visible spectrum cameras. Variability in expressions, pose, and eyewear has been systematically recorded. The dataset has been curated with extensive annotations, metadata, and standardized protocols for evaluation. Furthermore, this paper presents extensive benchmark results and analysis on thermal face landmark detection and thermal-to-visible face verification by evaluating state-of-the-art models on the ARL-VTF dataset. Domenick Poster, Matthew Thielke, Robert Nguyen, Srinivasan Rajaraman, Xing Di, Cedric Nimpa Fondje, Vishal M. Patel, Nathan J. Short, Benjamin S. Riggan, Nasser M. Nasrabadi, Shuowen Hu |
WACV | 5 |
| 2018 | GP-GAN: Gender Preserving GAN for Synthesizing Faces from LandmarksabstractFacial landmarks constitute the most compressed representation of faces and are known to preserve information such as pose, gender and facial structure present in the faces. Several works exist that attempt to perform high-level face-related analysis tasks based on landmarks alone without the aid of face images. In contrast, in this work, an attempt is made to tackle the inverse problem of synthesizing faces from their respective landmarks. The primary aim of this work is to demonstrate that information preserved by landmarks (gender in particular) can be further accentuated by leveraging generative models to synthesize corresponding faces. Though the problem is particularly challenging due to its ill-posed nature, we believe that successful synthesis will enable several applications such as boosting performance of high-level face related tasks using landmark points and performing dataset augmentation. To this end, a novel face-synthesis method known as Gender Preserving Generative Adversarial Network (GP-GAN) that is guided by adversarial loss, perceptual loss and a gender preserving loss is presented. Further, we propose a novel generator sub-network UDeNet for GP-GAN that leverages advantages of U-Net and DenseNet architectures. Extensive experiments and comparison with recent methods are performed to verify the effectiveness of the proposed method. Our code is available at: https://github.com/DetionDXlGP-GAN-Gender-Preserving-GAN-for-Synthesizing-Faces-from-Landmarks Xing Di, Vishwanath A. Sindagi, Vishal M. Patel |
ICPR | 1 |
| 2017 | Large Margin Multi-Modal Triplet Metric LearningabstractDistance metric learning is a significant technique that can improve the similarity accuracy in verification systems. In this paper, we propose a multi-metric learning algorithm with the triplet distance constraints for multi-modal verification problems. The main feature of our algorithm is that when learning multi-metric, we not only enforce the distance between the anchor and the positive samples to be less than the distance between the anchor and the negative samples but we also make the distance between the anchor and the positive samples as small as possible. A simple iterative procedure is introduced to solve the proposed optimization problem. Extensive experiments on three publicly available multi-modal datasets show that our method can perform significantly better than many state-ofthe- art multi-modal metric learning methods. Xing Di, Vishal M. Patel |
FG | 1 |