Site Li

dblp:221/8793 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Ordinal Unsupervised Domain Adaptation With Recursively Conditional Gaussian Imposed Variational Disentanglement
abstract
There has been a growing interest in unsupervised domain adaptation (UDA) to alleviate the data scalability issue, while the existing works usually focus on classifying independently discrete labels. However, in many tasks (e.g., medical diagnosis), the labels are discrete and successively distributed. The UDA for ordinal classification requires inducing non-trivial ordinal distribution prior to the latent space. Target for this, the partially ordered set (poset) is defined for constraining the latent vector. Instead of the typically i.i.d. Gaussian latent prior, in this work, a recursively conditional Gaussian (RCG) set is proposed for ordered constraint modeling, which admits a tractable joint distribution prior. Furthermore, we are able to control the density of content vectors that violate the poset constraint by a simple "three-sigma rule." We explicitly disentangle the cross-domain images into a shared ordinal prior induced ordinal content space and two separate source/target ordinal-unrelated spaces, and the self-training is worked on the shared space exclusively for ordinal-aware domain alignment. Extensive experiments on UDA medical diagnoses and facial age estimation demonstrate its effectiveness.
Xiaofeng Liu 0001, Site Li, Yubin Ge, Pengyi Ye, Jane You, Jun Lu 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 5IDER: Unified Query Rewriting for Steering, Intent Carryover, Disfluencies, Entity Carryover and Repair
abstract
Providing voice assistants the ability to navigate multi-turn conversations is a challenging problem.Handling multi-turn interactions requires the system to understand various conversational use-cases, such as steering, intent carryover, disfluencies, entity carryover, and repair.The complexity of this problem is compounded by the fact that these use-cases mix with each other, often appearing simultaneously in natural language.This work proposes a non-autoregressive query rewriting architecture that can handle not only the five aforementioned tasks, but also complex compositions of these use-cases.We show that our proposed model has competitive single task performance compared to the baseline approach, and even outperforms a fine-tuned T5 model in use-case compositions, despite being 15 times smaller in parameters and 25 times faster in latency.
Jiarui Lu, Bo-Hsiang Tseng, Joel Ruben Antony Moniz, Site Li, Xueyun Zhu, Murat Akbacak
INTERSPEECH4
2023 Inducing semantic hierarchy structure in empirical risk minimization with optimal transport measures
Wanqing Xie, Yubin Ge, Site Li, Zhenhua Guo 0001, Xiaofeng Liu 0001
Neurocomputing3
2022 Towards Modern Card Games with Large-Scale Action Spaces Through Action Representation
abstract
Axie infinity is a complicated card game with a huge-scale action space. This makes it difficult to solve this challenge using generic Reinforcement Learning (RL) algorithms. We propose a hybrid RL framework to learn action representations and game strategies. To avoid evaluating every action in the large feasible action set, our method evaluates actions in a fixed-size set which is determined using action representations. We compare the performance of our method with two baseline methods in terms of their sample efficiency and the winning rates of the trained models. We empirically show that our method achieves an overall best winning rate and the best sample efficiency among the three methods.
Zhiyuan Yao 0001, Site Li, Yiting Xie, Yuanyuan Qin, Xiongjie Xie, Huan Lu
CoG3
2022 Wasserstein Loss With Alternative Reinforcement Learning for Severity-Aware Semantic Segmentation
abstract
Semantic segmentation is important for many real-world systems, e.g., autonomous vehicles, which predict the class of each pixel. Recently, deep networks achieved significant progress w.r.t. the mean Intersection-over Union (mIoU) with the cross-entropy loss. However, the cross entropy loss can essentially ignore the difference of severity for an autonomous car with different wrong prediction mistakes. For example, predicting the car to the road is much more servery than recognize it as the bus. Targeting for this difficulty, we develop a Wasserstein training framework to explore the inter-class correlation by defining its ground metric as misclassification severity. The ground metric of Wasserstein distance can be pre-defined following the experience on a specific task. From the optimization perspective, we further propose to set the ground metric as an increasing function of the pre-defined ground metric. Furthermore, an adaptively learning scheme of the ground matrix is proposed to utilize the high-fidelity CARLA simulator. Specifically, we follow a reinforcement alternative learning scheme. The experiments on both CamVid and Cityscapes datasets evidenced the effectiveness of our Wasserstein loss. The SegNet, ENet, FCN and Deeplab networks can be adapted following a plug in manner. We achieve significant improves on the predefined important classes, and much longer continuous play time in our simulator.
Xiaofeng Liu 0001, Yunhong Lu, Xiongchang Liu, Site Li, Jane You
IEEE Trans. Intell. Transp. Syst.5
2021 Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models
abstract
AI Safety is a major concern in many deep learning applications such as autonomous driving. Given a trained deep learning model, an important natural problem is how to reliably verify the model's prediction. In this paper, we propose a novel framework --- deep verifier networks (DVN) to detect unreliable inputs or predictions of deep discriminative models, using separately trained deep generative models. Our proposed model is based on conditional variational auto-encoders with disentanglement constraints to separate the label information from the latent representation. We give both intuitive and theoretical justifications for the model. Our verifier network is trained independently with the prediction model, which eliminates the need of retraining the verifier network for a new model. We test the verifier network on both out-of-distribution detection and adversarial example detection problems, as well as anomaly detection problems in structured prediction tasks such as image caption generation. We achieve state-of-the-art results in all of these problems.
Tong Che, Xiaofeng Liu 0001, Site Li, Yubin Ge, Ruixiang Zhang, Caiming Xiong, Yoshua Bengio
AAAI3
2021 Embedding Semantic Hierarchy in Discrete Optimal Transport for Risk Minimization
abstract
The widely-used cross-entropy (CE) loss-based deep networks achieved significant progress w.r.t. the classification accuracy. However, the CE loss can essentially ignore the risk of misclassification which is usually measured by the distance between the prediction and label in a semantic hierarchical tree. In this paper, we propose to incorporate the risk-aware inter-class correlation in a discrete optimal transport (DOT) training framework by configuring its ground distance matrix. The ground distance matrix can be pre-defined following a priori of hierarchical semantic risk. Specifically, we define the tree induced error (TIE) on a hierarchical semantic tree and extend it to its increasing function from the optimization perspective. The semantic similarity in each level of a tree is integrated with the information gain. We achieve promising results on several large scale image classification tasks with a semantic tree structure in a plug and play manner.
Yubin Ge, Site Li, Wanqing Xie, Jane You, Xiaofeng Liu 0001
ICASSP2
2021 Adversarial Unsupervised Domain Adaptation with Conditional and Label Shift: Infer, Align and Iterate
abstract
In this work, we propose an adversarial unsupervised domain adaptation (UDA) method under inherent conditional and label shifts, in which we aim to align the distributions w.r.t. both p(x|y) and p(y). Since labels are inaccessible in a target domain, conventional adversarial UDA methods assume that p(y) is invariant across domains and rely on aligning p(x) as an alternative to the p(x|y) alignment. To address this, we provide a thorough theoretical and empirical analysis of the conventional adversarial UDA methods under both conditional and label shifts, and propose a novel and practical alternative optimization scheme for adversarial UDA. Specifically, we infer the marginal p(y) and align p(x|y) iteratively at the training stage, and precisely align the posterior p(y|x) at the testing stage. Our experimental results demonstrate its effectiveness on both classification and segmentation UDA and partial UDA.
Xiaofeng Liu 0001, Zhenhua Guo 0001, Site Li, Fangxu Xing, Jane You, C.-C. Jay Kuo, Georges El Fakhri, Jonghye Woo
ICCV3
2021 Recursively Conditional Gaussian for Ordinal Unsupervised Domain Adaptation
abstract
The unsupervised domain adaptation (UDA) has been widely adopted to alleviate the data scalability issue, while the existing works usually focus on classifying independently discrete labels. However, in many tasks (e.g., medical diagnosis), the labels are discrete and successively distributed. The UDA for ordinal classification requires inducing non-trivial ordinal distribution prior to the latent space. Target for this, the partially ordered set (poset) is defined for constraining the latent vector Instead of the typically i.i.d. Gaussian latent prior, in this work, a recursively conditional Gaussian (RCG) set is adapted for ordered constraint modeling, which admits a tractable joint distribution prior Furthermore, we are able to control the density of content vector that violates the poset constraints by a simple "three-sigma rule". We explicitly disentangle the cross-domain images into a shared ordinal prior induced ordinal content space and two separate source/target ordinal-unrelated spaces, and the self-training is worked on the shared space exclusively for ordinal-aware domain alignment. Extensive experiments on UDA medical diagnoses and facial age estimation demonstrate its effectiveness.
Xiaofeng Liu 0001, Site Li, Yubin Ge, Pengyi Ye, Jane You, Jun Lu 0002
ICCV2
2021 Band Selection for Specific Target Detection of Hyperspectral Imagery
abstract
Band selection (BS) is considered as an effective method for dimensionality reduction of hyperspectral data. As an important application for hyperspectral remote sensing, target detection is widely concerned. Therefore, how to select more representational band subset for specific target to improve performance of detection is worth discussing. This letter proposed a BS method for specific target detection, called constrained target band selection under adaptive subspace partitioning (CTASPBS). Firstly, all bands are partitioned into multiple weakly correlated subsets via adaptive subspace partition strategy (ASPS). Then, according to a target-constrained band prioritization (BP) criterion, the band with the highest priority in each subset is selected to form the optimal band subset. Due to the application of different BP criterion, two BS methods ASPS_MinV and ASPS_MaxV are proposed. Finally, experimental results on real hyperspectral data show that CTASPBS is an effective BS method for specific target detection.
Xudong Sun 0009, Site Li, Fengqiang Xu, Xianping Fu
IGARSS2
2020 Importance-Aware Semantic Segmentation in Self-Driving with Discrete Wasserstein Training
abstract
Semantic segmentation (SS) is an important perception manner for self-driving cars and robotics, which classifies each pixel into a pre-determined class. The widely-used cross entropy (CE) loss-based deep networks has achieved significant progress w.r.t. the mean Intersection-over Union (mIoU). However, the cross entropy loss can not take the different importance of each class in an self-driving system into account. For example, pedestrians in the image should be much more important than the surrounding buildings when make a decisions in the driving, so their segmentation results are expected to be as accurate as possible. In this paper, we propose to incorporate the importance-aware inter-class correlation in a Wasserstein training framework by configuring its ground distance matrix. The ground distance matrix can be pre-defined following a priori in a specific task, and the previous importance-ignored methods can be the particular cases. From an optimization perspective, we also extend our ground metric to a linear, convex or concave increasing function w.r.t. pre-defined ground distance. We evaluate our method on CamVid and Cityscapes datasets with different backbones (SegNet, ENet, FCN and Deeplab) in a plug and play fashion. In our extenssive experiments, Wasserstein loss demonstrates superior segmentation performance on the predefined critical classes for safe-driving.
Xiaofeng Liu 0001, Yuzhuo Han, Yi Ge, Tianxing Wang 0003, Site Li, Jane You, Jun Lu 0002
AAAI7
2020 AUTO3D: Novel View Synthesis Through Unsupervisely Learned Variational Viewpoint and Global 3D Representation
Xiaofeng Liu 0001, Tong Che, Yiqun Lu, Chao Yang 0011, Site Li, Jane You
ECCV (9)5
2019 Feature-Level Frankenstein: Eliminating Variations for Discriminative Recognition
abstract
Recent successes of deep learning-based recognition rely on maintaining the content related to the main-task label. However, how to explicitly dispel the noisy signals for better generalization remains an open issue. We systematically summarize the detrimental factors as task-relevant/irrelevant semantic variations and unspecified latent variation. In this paper, we cast these problems as an adversarial minimax game in the latent space. Specifically, we propose equipping an end-to-end conditional adversarial network with the ability to decompose an input sample into three complementary parts. The discriminative representation inherits the desired invariance property guided by prior knowledge of the task, which is marginally independent to the task-relevant/irrelevant semantic and latent variations. Our proposed framework achieves top performance on a serial of tasks, including digits recognition, lighting, makeup, disguise-tolerant face recognition, and facial attributes recognition.
Xiaofeng Liu 0001, Site Li, Lingsheng Kong, Wanqing Xie, Ping Jia, Jane You, B. V. K. Vijaya Kumar
CVPR2
2019 Permutation-Invariant Feature Restructuring for Correlation-Aware Image Set-Based Recognition
abstract
We consider the problem of comparing the similarity of image sets with variable-quantity, quality and un-ordered heterogeneous images. We use feature restructuring to exploit the correlations of both inner&inter-set images. Specifically, the residual self-attention can effectively restructure the features using the other features within a set to emphasize the discriminative images and eliminate the redundancy. Then, a sparse/collaborative learning-based dependency-guided representation scheme reconstructs the probe features conditional to the gallery features in order to adaptively align the two sets. This enables our framework to be compatible with both verification and open-set identification. We show that the parametric self-attention network and non-parametric dictionary learning can be trained end-to-end by a unified alternative optimization scheme, and that the full framework is permutation-invariant. In the numerical experiments we conducted, our method achieves top performance on competitive image set/video-based face recognition and person re-identification benchmarks.
Xiaofeng Liu 0001, Zhenhua Guo 0001, Site Li, Ping Jia, Lingsheng Kong, Jane You, B. V. K. Vijaya Kumar
ICCV3
2018 Data Augmentation via Latent Space Interpolation for Image Classification
abstract
Effective training of the deep neural networks requires much data to avoid underdetermined and poor generalization. Data Augmentation alleviates this by using existing data more effectively. However standard data augmentation produces only limited plausible alternative data by for example, flipping, distorting, adding noise to, cropping a patch from the original samples. In this paper, we introduce the adversarial autoencoder (AAE) to impose the feature representations with uniform distribution and apply the linear interpolation on latent space, which is potential to generate a much broader set of augmentations for image classification. As a possible “recognition via generation” framework, it has potentials for several other classification tasks. Our experiments on the ILSVRC 2012, CIFAR-10 datasets show that the latent space interpolation (LSI) improves the generalization and performance of state-of-the-art deep neural networks.
Xiaofeng Liu 0001, Yang Zou 0003, Lingsheng Kong, Zhihui Diao, Junliang Yan, Site Li, Ping Jia, Jane You
ICPR7