VLDB 2026 Research / reviewers in the wild / expert
Yanbin Liu 0003
dblp:44/1448-3
· DBLP profile ↗
24ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0003-4724-8065ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ideological isolation in online social networks: A survey of computational definitions, metrics, and mitigationabstractIdeological isolation in online social networks, including selective exposure, echo chambers, filter bubbles, tunnel vision, and polarization, has become a central concern for computational and neural modeling of information ecosystems. With the rapid adoption of graph learning, representation learning, and feedback-driven recommender systems, a growing body of work has proposed diverse metrics and models to quantify and mitigate these phenomena. However, existing studies and surveys rely on heterogeneous definitions and incompatible measurements, making empirical findings difficult to compare and obscuring how different forms of ideological isolation arise in learning-based systems. This survey provides a computationally grounded and comprehensive review of existing approaches to defining, analyzing, measuring, and mitigating ideological isolation in online social networks. We examine the mechanisms underlying content personalization, user behavior, and network structure that drive exposure concentration and attention narrowing. We then systematically review methodological approaches for detecting and quantifying ideological isolation, covering network-, content-, and behavior-based metrics, and synthesize empirical findings across platforms to assess their applicability and limitations. We further organize computational mitigation strategies, including network-topological interventions and recommendation-level controls, compare mitigation families and their trade-offs, and examine the dual role of large language models. The key ethical considerations in the design and deployment of diversity-aware systems are also discussed. By resolving the definition-metric-intervention mismatch that characterizes existing work, this survey provides a principled foundation for the design, evaluation, and deployment of neural and learning-based systems aimed at diagnosing and mitigating ideological isolation in online social networks. Yanbin Liu 0003, Shiqing Wu 0001, Ziying Zhao, Yuxuan Hu 0002, Weihua Li 0007, Quan Bai 0001 |
Neurocomputing | 2 |
| 2026 | Compression-Oriented Video Super-ResolutionabstractCurrent compressed video super-resolution methods have achieved promising performance, but they often assume that an input video is compressed under low-delay configurations. However, under random access configurations, those methods might struggle to leverage the metadata effectively due to the large variations of metadata in different compression configurations. In this work, we propose a Compression-Oriented Video Super-Resolution (COVSR) method that can address video super-resolution for both low-delay and random-access configurations. Specifically, we first introduce an efficient compression-aware propagation (ECAP) module that dynamically adjusts propagation routes in accordance with the compression configurations. Since existing methods require reconstructing frames in a frame-by-frame manner, it is difficult to achieve efficient parallelization. However, we find that by slightly relaxing sequential dependencies, our ECAP can significantly improve inference speed. Furthermore, existing methods typically perform alignment between adjacent frames or adjacent features. However, since ECAP may propagate features along non-adjacent reference routes, it introduces new challenges for accurate cross-frame feature alignment. In response, we propose a metadata-driven alignment (MDA) module that refines cross-frame motion vectors into dense, feature-level flow offsets, enabling precise alignment across temporally distant features. Extensive experimental results demonstrate that our COVSR not only achieves efficient and superior super-resolution performance but also is generalizable to various compression configurations. Our code will be available at https://covsr.github.io. Yanbin Liu 0003, Ming Lu 0002, Zhuojie Wu, Senmao Tian, Yandong Guo, Xin Yu 0002 |
IEEE Trans. Image Process. | 2 |
| 2025 | Tunnel Vision in Online Discourse: Formalization and Entropy-Based Quantification with LLM-Simulated Agents
Yanbin Liu 0003, Weihua Li 0007, Quan Bai 0001 |
PRICAI | 2 |
| 2024 | Class-Aware Sample Weight Learning for Cross-Modal Unsupervised Domain Adaptation in Cross-User Wearable Human Activity RecognitionabstractExisting unsupervised domain adaption approaches for cross-user Wearable Human Activity Recognition (WHAR) typically assume that users utilize the uni-modal sensor deployment configuration and cannot transfer across different sensor modalities. In this paper, we consider the more realistic cross-modal wearable human activity recognition setting to investigate the unsupervised domain adaptation task. This new context presents two formidable challenges: (1) how to alleviate modality heterogeneity across users, and (2) how to explore cross-modal domain correlation for better unsupervised domain adaptation. We propose a cross-modal unsupervised domain Adaptation model with Class-Aware Sample Weight Learning (CASWL-Adapt) to address both challenges. First, a spherical modality discriminator is designed to capture modal-specific discriminative features of each user during domain adaptation, thus achieving a reduction of sample variance caused by modal heterogeneity. Given a user-specific modal, modality-independent domain-invariant features can be efficiently generated by the well-developed modality discrimination loss and adversarial training. Second, a class-aware weight network is devised to calculate sample weights through classification loss and activity class similarity for each sample. Furthermore, the network leverage end-to-end learning and meta-optimization update rules to explore inter-domain correlations. Cross-modal activity classes are expected to adaptively implement different weighting schemes based on their intrinsic bias characteristics to select the most appropriate samples for domain knowledge transfer. We demonstrate that CASWL-Adapt achieves state-of-the-art results on three challenging benchmarks: Epic-Kitchens, Multimodal-EA and RealWorld, especially effective for new users of unseen modality. Yanbin Liu 0003, Sira Yongchareon |
ECAI | 2 |
| 2024 | Unsupervised Dense Prediction Using Differentiable Normalized Cuts
Yanbin Liu 0003, Stephen Gould |
ECCV (40) | 1 |
| 2024 | CRA-Eformer: Cross-Scale Residual Attention-Based Edge-Guide Transformer for Low-Dose CT Denoising
Yanbin Liu 0003, Sira Yongchareon |
ICONIP (5) | 2 |
| 2024 | TPR: Topology-Preserving Reservoirs for Generalized Zero-Shot LearningabstractPre-trained vision-language models (VLMs) such as CLIP have shown excellent performance for zero-shot classification. Based on CLIP, recent methods design various learnable prompts to evaluate the zero-shot generalization capability on a base-to-novel setting. This setting assumes test samples are already divided into either base or novel classes, limiting its application to realistic scenarios. In this paper, we focus on a more challenging and practical setting: generalized zero-shot learning (GZSL), i.e., testing with no information about the base/novel division. To address this challenging zero-shot problem, we introduce two unique designs that enable us to classify an image without the need of knowing whether it comes from seen or unseen classes. Firstly, most existing methods only adopt a single latent space to align visual and linguistic features, which has a limited ability to represent complex visual-linguistic patterns, especially for fine-grained tasks. Instead, we propose a dual-space feature alignment module that effectively augments the latent space with a novel attribute space induced by a well-devised attribute reservoir. In particular, the attribute reservoir consists of a static vocabulary and learnable tokens complementing each other for flexible control over feature granularity. Secondly, finetuning CLIP models (e.g., prompt learning) on seen base classes usually sacrifices the model's original generalization capability on unseen novel classes. To mitigate this issue, we present a new topology-preserving objective that can enforce feature topology structures of the combined base and novel classes to resemble the topology of CLIP. In this manner, our model will inherit the generalization ability of CLIP through maintaining the pairwise class angles in the attribute space. Extensive experiments on twelve object recognition datasets demonstrate that our model, termed Topology-Preserving Reservoir (TPR), outperforms strong baselines including both prompt learning and conventional generative-based zero-shot methods. Hui Chen 0036, Yanbin Liu 0003, Nanning Zheng 0001, Xin Yu 0002 |
NeurIPS | 2 |
| 2024 | Multivariate Traffic Demand Prediction via 2D Spectral Learning and Global Spatial Optimization
Changlu Chen, Yanbin Liu 0003, Ling Chen 0006, Chengqi Zhang |
ECML/PKDD (2) | 2 |
| 2024 | Test-Time Training for Spatial-Temporal ForecastingabstractDespite the recent success of deep neural networks in spatial-temporal forecasting, existing methods suffer from distribution shifts between the training and test data, failing to address the non-stationary and abrupt changes at test time. To solve this problem, we propose a novel test-time training framework for spatial-temporal forecasting. Instead of employing a fixed trained model, we adapt the trained model with only one or a mini-batch of test examples to address the test data shifts. The unique spatial structure with hundreds of geographical locations offers an effective batch size to explore the test-time distribution and avoid overfitting. Changlu Chen, Yanbin Liu 0003, Ling Chen 0006, Chengqi Zhang |
SDM | 2 |
| 2024 | NeRFEditor: Differentiable Style Decomposition for 3D Scene EditingabstractWe present NeRFEditor, an efficient learning framework for 3D scene editing, which takes a video as input and outputs a high-quality, identity-preserving stylized 3D scene. Our goal is to bridge the gap between 2D and 3D editing, catering to a wide array of creative modifications such as reference-guided alterations, text-based prompts, and user interactions. We achieve this by encouraging a pre-trained StyleGAN model and a NeRF model to learn mutually consistent renderings. Specifically, we use NeRF to generate numerous (image, camera pose)-pairs to train an adjustor module, which adapts the StyleGAN latent code for generating high-fidelity stylized images from any given viewing angle. To extrapolate edits to novel views, i.e., those not seen by StyleGAN pre-training, while maintaining 360° consistency, we propose a second self-supervised module that maps these views into the hidden space of StyleGAN. Together these two modules produce sufficient guidance for NeRF to learn consistent stylization effects across the full range of views. Experiments show that NeRFEditor outperforms prior work on benchmark and real-world scenes with better editability, fidelity, and identity preservation. Chunyi Sun, Yanbin Liu 0003, Junlin Han, Stephen Gould |
WACV | 2 |
| 2024 | Bilaterally Normalized Scale-Consistent Sinkhorn Distance for Few-Shot Image ClassificationabstractFew-shot image classification aims at exploring transferable features from base classes to recognize images of the unseen novel classes with only a few labeled images. Existing methods usually compare the support features and query features, which are implemented by either matching the global feature vectors or matching the local feature maps at the same position. However, few labeled images fail to capture all the diverse context and intraclass variations, leading to mismatch issues for existing methods. On one hand, due to the misaligned position and cluttered background, existing methods suffer from the object mismatch issue. On the other hand, due to the scale inconsistency between images, existing methods suffer from the scale mismatch issue. In this article, we propose the bilaterally normalized scale-consistent Sinkhorn distance (BSSD) to solve these issues. First, instead of same-position matching, we use the Sinkhorn distance to find an optimal matching between images, mitigating the object mismatch caused by misaligned position. Meanwhile, we propose the intraimage and interimage attentions as the bilateral normalization on the Sinkhorn distance to suppress the object mismatch caused by background clutter. Second, local feature maps are enhanced with the multiscale pooling strategy, making the Sinkhorn distance possible to find a consistent matching scale between images. Experimental results show the effectiveness of the proposed approach, and we achieve the state-of-the-art on three few-shot benchmarks. Yanbin Liu 0003, Linchao Zhu, Makoto Yamada, Yi Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | RiskContra: A Contrastive Approach to Forecast Traffic Risks with Multi-Kernel Networks
Changlu Chen, Yanbin Liu 0003, Ling Chen 0006, Chengqi Zhang |
PAKDD (4) | 2 |
| 2023 | Bidirectional Spatial-Temporal Adaptive Transformer for Urban Traffic Flow ForecastingabstractUrban traffic forecasting is the cornerstone of the intelligent transportation system (ITS). Existing methods focus on spatial-temporal dependency modeling, while two intrinsic properties of the traffic forecasting problem are overlooked. First, the complexity of diverse forecasting tasks is nonuniformly distributed across various spaces (e.g., suburb versus downtown) and times (e.g., rush hour versus off-peak). Second, the recollection of past traffic conditions is beneficial to the prediction of future traffic conditions. Based on these properties, we propose a bidirectional spatial-temporal adaptive transformer (Bi-STAT) for accurate traffic forecasting. Bi-STAT adopts an encoder-decoder architecture, where both the encoder and the decoder maintain a spatial-adaptive transformer and a temporal-adaptive transformer structure. Inspired by the first property, each transformer is designed to dynamically process the traffic streams according to their task complexities. Specifically, we realize this by the recurrent mechanism with a novel dynamic halting module (DHM). Each transformer performs iterative computation with shared parameters until DHM emits a stopping signal. Motivated by the second property, Bi-STAT utilizes one decoder to perform the present → past recollection task and the other decoder to perform the present → future prediction task. The recollection task supplies complementary information to assist and regularize the prediction task for a better generalization. Through extensive experiments, we show the effectiveness of each module in Bi-STAT and demonstrate the superiority of Bi-STAT over the state-of-the-art baselines on four benchmark datasets. The code is available at https://github.com/chenchl19941118/Bi-STAT.git. Changlu Chen, Yanbin Liu 0003, Ling Chen 0006, Chengqi Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Diminishing Empirical Risk Minimization for Unsupervised Anomaly DetectionabstractUnsupervised anomaly detection (AD) is a challenging task in realistic applications. Recently, there is an increasing trend to detect anomalies with deep neural networks (DNN). However, most popular deep AD detectors cannot protect the network from learning contaminated information brought by anomalous data, resulting in unsatisfactory detection performance and overfitting issues. In this work, we identify one reason that hinders most existing DNN-based anomaly detection methods from performing is the wide adoption of the Empirical Risk Minimization (ERM). ERM assumes that the performance of an algorithm on an unknown distribution can be approximated by averaging losses on the known training set. This averaging scheme thus ignores the distinctions between normal and anomalous instances. To break through the limitations of ERM, we propose a novel Diminishing Empirical Risk Minimization (DERM) framework. Specifically, DERM adaptively adjusts the impact of individual losses through a well-devised aggregation strategy. Theoretically, our proposed DERM can directly modify the gradient contribution of each individual loss in the optimization process to suppress the influence of outliers, leading to a robust anomaly detector. Empirically, DERM outperformed the state-of-the-art on the unsupervised AD benchmark consisting of 18 datasets. Shaoshen Wang, Yanbin Liu 0003, Ling Chen 0006, Chengqi Zhang |
IJCNN | 2 |
| 2022 | Feature-Robust Optimal Transport for High-Dimensional Data
Mathis Petrovich, Chao Liang 0002, Ryoma Sato, Yanbin Liu 0003, Yao-Hung Tsai, Linchao Zhu, Yi Yang 0001, Ruslan Salakhutdinov, Makoto Yamada |
ECML/PKDD (5) | 4 |
| 2021 | A Multi-Mode Modulator for Multi-Domain Few-Shot ClassificationabstractMost existing few-shot classification methods only consider generalization on one dataset (i.e., single-domain), failing to transfer across various seen and unseen domains. In this paper, we consider the more realistic multi-domain few-shot classification problem to investigate the cross-domain generalization. Two challenges exist in this new setting: (1) how to efficiently generate multi-domain feature representation, and (2) how to explore domain correlations for better cross-domain generalization. We propose a parameter-efficient multi-mode modulator to address both challenges. First, the modulator is designed to maintain multiple modulation parameters (one for each domain) in a single network, thus achieving single-network multi-domain representation. Given a particular domain, domain-aware features can be efficiently generated with the well-devised separative selection module and cooperative query module. Second, we further divide the modulation parameters into the domain-specific set and the domain-cooperative set to explore the intra-domain information and inter-domain correlations, respectively. The intra-domain information describes each domain independently to prevent negative interference. The inter-domain correlations guide information sharing among relevant domains to enrich their own representation. Moreover, unseen domains can utilize the correlations to obtain an adaptive combination of seen domains for extrapolation. We demonstrate that the proposed multi-mode modulator achieves state-of-the-art results on the challenging META-DATASET benchmark, especially for unseen test domains. Yanbin Liu 0003, Juho Lee 0001, Linchao Zhu, Ling Chen 0006, Humphrey Shi, Yi Yang 0001 |
ICCV | 1 |
| 2021 | Cross-aligned and Gumbel-refactored Autoencoders for Multi-view Anomaly DetectionabstractMulti-view anomaly detection (AD) is a challenging task due to the complicated data distributions across different views. Specifically, there exist two types of anomalies in multi-view distributions: attribute anomaly that exhibits consistent anomalous pattern in each view and class anomaly that exhibits inconsistent traits (e.g., semantic label) across multiple views. Existing methods detect these anomalies in an unsupervised manner with the clustering assumption: normal data share consistent clustering structure across views while anomalous data exhibits inconsistent clusters across views. However, these methods would fail for complex multi-view data distributions where there is no obvious clusters. Moreover, existing models suffer from robustness since they are undermined by anomalies during training time. To get rid of the clustering assumption, we propose a Cross-aligned and Gumbel-refactored AutoEncoders (CGAEs) model to effectively detect two types of multi-view anomalies. In CGAEs, we devise a cross-reconstruction module to detect class anomaly by recovering one view from another view. Class anomalies would lead to high cross-reconstruction loss since they do not have the correct information in one view to generate another. We further design a view-alignment module to detect attribute anomaly by the alignment distance among multiple views in the latent space. Attribute anomalies possess large distances since they are less aligned due to fewer anomalous training instances. To handle the robustness issue, we propose a Gumbel-refactored reconstruction loss to replace the mean square error (MSE) in original autoencoders. The cross entropy loss is calculated between the discreterized input and Gumbel-sampled output, thus disregarding the irrelevant details to achieve model robustness. Experimental results validate the superiority of the proposed CGAEs model on both the benchmark datasets and real world datasets. Shaoshen Wang, Yanbin Liu 0003, Ling Chen 0006, Chengqi Zhang |
ICTAI | 2 |
| 2021 | LSMI-Sinkhorn: Semi-supervised Mutual Information Estimation with Optimal Transport
Yanbin Liu 0003, Makoto Yamada, Yao-Hung Tsai, Tam Le, Ruslan Salakhutdinov, Yi Yang 0001 |
ECML/PKDD (1) | 1 |
| 2020 | Semantic Correspondence as an Optimal Transport ProblemabstractEstablishing dense correspondences across semantically similar images is a challenging task. Due to the large intra-class variation and background clutter, two common issues occur in current approaches. First, many pixels in a source image are assigned to one target pixel, i.e., many to one matching. Second, some object pixels are assigned to the background pixels, i.e., background matching. We solve the first issue by global feature matching, which maximizes the total matching correlations between images to obtain a global optimal matching matrix. The row sum and column sum constraints are enforced on the matching matrix to induce a balanced solution, thus suppressing the many to one matching. We solve the second issue by applying a staircase function on the class activation maps to re-weight the importance of pixels into four levels from foreground to background. The whole procedure is combined into a unified optimal transport algorithm by converting the maximization problem to the optimal transport formulation and incorporating the staircase weights into optimal transport algorithm to act as empirical distributions. The proposed algorithm achieves state-of-the-art performance on four benchmark datasets. Notably, a 26\% relative improvement is achieved on the large-scale SPair-71k dataset. Yanbin Liu 0003, Linchao Zhu, Makoto Yamada, Yi Yang 0001 |
CVPR | 1 |
| 2019 | Adaptive Sparse Confidence-Weighted Learning for Online Feature SelectionabstractIn this paper, we propose a new online feature selection algorithm for streaming data. We aim to focus on the following two problems which remain unaddressed in literature. First, most existing online feature selection algorithms merely utilize the first-order information of the data streams, regardless of the fact that second-order information explores the correlations between features and significantly improves the performance. Second, most online feature selection algorithms are based on the balanced data presumption, which is not true in many real-world applications. For example, in fraud detection, the number of positive examples are much less than negative examples because most cases are not fraud. The balanced assumption will make the selected features biased towards the majority class and fail to detect the fraud cases. We propose an Adaptive Sparse Confidence-Weighted (ASCW) algorithm to solve the aforementioned two problems. We first introduce an `0-norm constraint into the second-order confidence-weighted (CW) learning for feature selection. Then the original loss is substituted with a cost-sensitive loss function to address the imbalanced data issue. Furthermore, our algorithm maintains multiple sparse CW learner with the corresponding cost vector to dynamically select an optimal cost. We theoretically enhance the theory of sparse CW learning and analyze the performance behavior in F-measure. Empirical studies show the superior performance over the stateof-the-art online learning methods in the online-batch setting. Yanbin Liu 0003, Yan Yan 0006, Ling Chen 0006, Yahong Han, Yi Yang 0001 |
AAAI | 1 |
| 2019 | Learning to Propagate Labels: Transductive Propagation Network for Few-Shot Learning
Yanbin Liu 0003, Juho Lee 0001, Minseop Park, Saehoon Kim, Eunho Yang, Sung Ju Hwang, Yi Yang 0001 |
ICLR (Poster) | 1 |
| 2018 | Tensor learning and automated rank selection for regression-based video classification
Jianguang Zhang, Yanbin Liu 0003, Jianmin Jiang |
Multim. Tools Appl. | 2 |
| 2018 | Pooling the Convolutional Layers in Deep ConvNets for Video Action RecognitionabstractDeep ConvNets have shown their good performance in image classification tasks. However, there still remains problems in deep video representations for action recognition. On one hand, current video ConvNets are relatively shallow compared with image ConvNets, which limits their capability of capturing the complex video action information; on the other hand, temporal information of videos is not properly utilized to pool and encode the video sequences. Toward these issues, in this paper we utilize two state-of-the-art ConvNets, i.e., the very deep spatial net (VGGNet [1]) and the temporal net from Two-Stream ConvNets [2], for action representation. The convolutional layers and the proposed new layer, called frame-diff layer, are extracted and pooled with two temporal pooling strategies: Trajectory pooling and Line pooling. The pooled local descriptors are then encoded with vector of locally aggregated descriptors (VLAD) [3] to form the video representations. In order to verify the effectiveness of the proposed framework, we conduct experiments on UCF101 and HMDB51 data sets. It achieves accuracy of 92.08% on UCF101, which is the state-of-the-art, and the accuracy of 65.62% on HMDB51, which is comparable to the state-of-the-art. In addition, we propose the new Line pooling strategy, which can speed up the extraction of feature and achieve the comparable performance of the Trajectory pooling. Shichao Zhao, Yanbin Liu 0003, Yahong Han, Richang Hong, Qinghua Hu, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Discriminative multi-view feature selection and fusionabstractIn computer vision tasks such as action recognition and image classification, combining multiple visual feature sets is proven to be an effective strategy. However, simply combing these features may cause high dimensionality and lead to noises. Feature selection and fusion are common choices for multiple feature representation. In this paper, we propose a multi-view feature selection and fusion method which chooses and fuses discriminative features from multiple feature sets. For discriminative feature selection, we learn the selection matrix W by the minimization of the trace ratio objective function with ℓ2,1norm regularization. For multiple feature fusion, we incorporate local structures of each view in the Laplacian matrix. Since the Laplacian matrix is constructed in unsupervised manner and scaled category indicator matrix is solved iteratively, our work is fully unsupervised. Experimental results on four action recognition datasets and two large-scale image classification datasets demonstrate the effectiveness of multi-view feature selection and fusion. Yanbin Liu 0003, Binbing Liao, Yahong Han |
ICME | 1 |