EDBT 2026 Demo / reviewers in the wild / expert
Yanyu Xu 0001
dblp:188/7560-1
· DBLP profile ↗
39ranked-venue papers
11as first author
29since 2021 · last 2026
0000-0001-8926-7833ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RTS-LLM: Restoring time structure for time series forecasting with LLMs
Taihua Chen, Xiang Ma 0006, Yanyu Xu 0001, Shuyuan Qian, Li-Zhen Cui 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Test-time thinking: Reinforcement adaptation for multimodal LLM reasoning
Haotian Chen 0002, Yanyu Xu 0001, Fang Wang 0017, Li-Zhen Cui 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Annotation-efficient medical image segmentation via cross-latent graphs and vector-quantized memory
Yanyu Xu 0001, Menghan Zhou, Xinxing Xu, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026, Li-Zhen Cui 0001 |
Medical Image Anal. | 1 |
| 2026 | Toward Effective Model Merging in Semantic SegmentationabstractModel merging has become a popular approach for combining individual models into a single model that inherits their capabilities and achieves improved performance. However, its success has not yet been transferred to semantic segmentation tasks due to two major challenges: 1) current model merging methods predominantly employ static merging strategies with fixed coefficients, limiting their ability to incorporate task-specific prior knowledge and 2) semantic segmentation faces large distribution shifts across multiple domains, causing negative transfer in the merged model. In this article, we propose an effective model merging approach for semantic segmentation, named M2Seg. To dramatically integrate relevant priors based on the input data, we propose a novel SVD-structured MoE module for adaptive merging. To address the severe distribution shifts, we further introduce a test-time dynamic calibration function designed to minimize discrepancies between training and test statistics. Additionally, historical information is leveraged to refine activation statistics during inference. Recognizing that unreliable data can negatively impact update directions, we develop a pixel-efficient entropy minimization mechanism to filter unstable pixels, thus stabilizing the merging process and enhancing segmentation performance. Extensive experiments on both seen and unseen semantic segmentation tasks demonstrate the superior effectiveness and generalization capability of our proposed method. The source code and pretrained checkpoints are available athttps://github.com/cht619/MMSeg Haotian Chen 0002, Yanyu Xu 0001, Fang Wang 0017, Li-Zhen Cui 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2026 | Learning Prediction-aware Prior in Transformer Network for Accurate Spatio-Temporal Video GroundingabstractSpatio-temporal video grounding (STVG) aims to precisely locate a spatio-temporal tube in an untrimmed video corresponding to a given language description. Many existing methods decouple spatial and temporal grounding as separate tasks, missing the strong interdependencies between the two, which are crucial for accurately aligning spatial regions (such as objects) with their motion over time. Thus, to enhance spatio-temporal associations, we introduce a new Prior-Driven Transformer Network (PDTNet) with predicted temporal boundaries as priors to guide object bounding boxes for improved spatial grounding over time. Firstly, PDTNet employs a temporal prior, termed reference query, to enhance discriminability between language-related and language-irrelevant visual content, improving temporal boundary localization. Further, the context within predicted temporal boundaries serves as another prior knowledge to modulate spatial features. We also introduce a prediction-aware Gaussian prior to precise object localization. This ensures consistent tube construction and accurate object localization. Extensive experiments on STVG benchmarks validate the effectiveness of PDTNet. Code is available at https://github.com/tongzhang111/PDTNet . Yongshun Gong, Jialin Gao, Yanyu Xu 0001, Xiushan Nie, Li-Zhen Cui 0001, Chengqi Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | AdaHet-MKD: An Adaptive Heterogeneous Multi-teacher Knowledge Distillation for Medical Image AnalysisabstractContrastive Language-Image Pre-training (CLIP) has emerged as an effective framework for multi-modal representation learning, achieving notable success in diverse tasks such as medical image analysis. CLIP's growing prominence in medical image applications is restricted by its significant computational demands, creating implementation challenges in resource-constrained clinical environments. While knowledge distillation offers an effective approach for model compression with preserved accuracy, existing methods suffer from two fundamental limitations. Firstly, existing methods focus on learning better information from single models while ignoring the fact that student models can generalize well under the guidance of multiple teachers. Secondly, they overlook the complementary information in the CLIP model where the text encoder and image encoder can be leveraged as heterogeneous information to teach one single modality. To tackle these challenges, we propose an Adaptive Heterogeneous Multi-teacher Knowledge Distillation (AdaHet-MKD) framework for effective knowledge transfer across heterogeneous text-image models and among multiple teacher models. The key innovations include: (i) adaptively determining the contribution of each teacher model to specific instances, thereby generating integrated soft logits, and (ii) enabling the student model to operate independently of the teacher model's architecture, which enhances flexibility in teacher-student pairings. Experimental evaluations on publicly available medical datasets demonstrate that our approach has achieved the state-of-the-art performance compared to baselines. Helin Wang, Wei Du 0010, Ning Liu 0014, Qian Li 0043, Yanyu Xu 0001, Li-Zhen Cui 0001 |
CIKM | 5 |
| 2025 | SAC-B: Soft Actor-Critic with Bias for Suppressing Q-valueOverestimation in Off-Policy Reinforcement LearningabstractSoft Actor-Critic (SAC) and other off-policy reinforcement learning methods have emerged as powerful tools for complex dynamic decision-making. However, they are susceptible to detrimental Q-value overestimation bias in stochastic environments. This bias critically impairs policy optimization, propagates estimation errors, and can lead to suboptimal policies. To address this challenge, we introduce Soft Actor-Critic with Bias (SAC-B), a novel framework designed to mitigate Q-value overestimation while preserving the strengths of maximum entropy RL. SAC-B incorporates a dynamic bias compensator, implemented as a learnable network, to heuristically counter Q-value overestimation. Crucially, we employ an error-decoupled stabilized architecture that isolates this bias prediction module from the critic networks, thereby eliminating error propagation and enhancing stability. Furthermore, SAC-B utilizes a hybrid objective function that strategically balances temporal difference learning with entropy regularization, enabling rapid adaptation. These synergistic components collectively enable SAC-B to effectively reduce temporal bootstrapping errors while maintaining the desirable exploration properties of the original SAC algorithm. Experiments across six continuous control benchmarks, including BipedalWalkerHardcore-v3 and Quadruped, validate the superiority and enhanced stability of our approach compared to existing methods. Wei Du 0010, Yanyu Xu 0001, Li-Zhen Cui 0002 |
DAI | 3 |
| 2025 | Parameterized Diffusion Optimization Enabled Autoregressive Ordinal Regression for Diabetic Retinopathy Grading
Qinkai Yu, Wei Zhou 0021, Hantao Liu, Yanyu Xu 0001, Meng Wang 0038, Yitian Zhao, Huazhu Fu, Xujiong Ye, Yalin Zheng, Yanda Meng |
MICCAI (15) | 4 |
| 2025 | Learning Invariant Discriminative Patterns for Unified Anomaly DetectionabstractUnified Anomaly Detection (UAD) aims to identify anomalies across diverse domains without access to target domain data during training. Unlike traditional anomaly detection methods that rely on training separate models for each domain, UAD employs a single model to generalize across multiple categories. A key challenge lies in the domain shift between seen and unseen data, which requires capturing invariant discriminative patterns between reference and query images across different domains during in-context learning for unified anomaly detection. To tackle this, we propose a novel UAD framework to learn the invariant discriminative patterns through pre-, in- and post-processing modules. First, a pre-processing VLM-guided data augmentation module generates diverse and semantically consist images, followed by a latent-space filtering mechanism. Second, an in-processing Adaptive VQ memory module stores representative discriminative patterns to enable robust residual comparison. Third, a post-processing GUR (Geometric distributions Upgrade Representation) feature augmentation module models geometric feature distributions to synthesize informative prompts, improving the quality of feature delta estimation for anomaly scoring. Extensive experiments on benchmark datasets demonstrate that our method achieves superior generalization in detecting anomalies across unseen domains, outperforming existing state-of-the-art approaches. Chengcheng Xing, Yanyu Xu 0001, Li-Zhen Cui 0001 |
ACM Multimedia | 2 |
| 2025 | Diffusion-based Data Augmentation via Noise Concatenation for Chest X-Ray ClassificationabstractThe automated analysis of chest X-rays using deep learning technologies has emerged as a critical trend in healthcare, offering the potential to enhance diagnostic efficiency and reduce clinician workload. However, the prohibitive cost of annotation and privacy concerns restrict the expansion of chest X-ray datasets, hindering the development of chest X-ray analysis models. One solution is to introduce generative models that synthesize new samples to expand datasets. Recent advances in diffusion models have demonstrated impressive capabilities, yet prevalent diffusion-based data augmentation techniques typically depend on alterations in color and shape, which are not always appropriate for medical imaging tasks. Consequently, we propose a diffusion-based data augmentation method that generates noise incorporating only the essential information, derived from the region of interest and guided by textual prompts. This noise is then combined with noise that encapsulates global information through concatenation, resulting in a synthesized version of chest X-rays that satisfy the fine-grained characteristics of medical images. We fine-tune the diffusion model using a publicly available dataset and evaluate our method on typical chest X-ray classification tasks. The results show that our method outperforms common data augmentation methods on multiple tasks. Yujing Xin, Ning Liu 0014, Yanyu Xu 0001, Li-Zhen Cui 0001 |
SMC | 4 |
| 2025 | Ad2Mix: Adversarial and Adaptive Mixup for Unsupervised Domain AdaptationabstractTransformer has recently gained tremendous popularity in unsupervised domain adaptation tasks due to its superior generalization ability. State-of-the-art methods leverage mixup to build an intermediate domain to reduce domain gap. However, such strategy becomes less effective when the domain gap becomes large, as the domain gap between intermediate domain and source domain is not minimized and the constructed intermediate domain is non informative. How to address the adaptation problem when domain gap becomes large is an important research problem in domain adaptation. In this paper, we propose an adversarial and adaptive mixup (Ad2mix) framework which gradually aligns the intermediate domain towards source domain to fully unleash the potential of both the transformer architecture and mixup to address the large domain gap problem. Specifically, we formulate a general framework for intermediate domain learning with mixup. We propose adversarial mixup with a specially designed mixup alike adversarial adaptation operation to reduce the domain gap between the intermediate domain and source domain. To construct an informative intermediate domain, unlike existing methods which utilize a Beta distribution to generate mixup coefficients to interpolate source and target data, we adaptively assign mixup coefficient for each target data instance based on their transferability and discriminativity information. Our framework creates a natural curriculum of intermediate domains from near source domain to near target domain for gradual adaptation. Extensive experimental studies and evaluations on three public domain adaptation benchmark datasets and one medical domain adaptation task demonstrate the superiority of our framework. Lei Zhu 0003, Yanyu Xu 0001, Yong Liu 0026, Rick Siow Mong Goh, Xinxing Xu |
WACV | 2 |
| 2025 | Effective test-time personalization for federated semantic segmentation
Haotian Chen 0002, Yanyu Xu 0001, Yibowen Zhao, Li-Zhen Cui 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Test-Time Medical Image Segmentation Using CLIP-Guided SAM AdaptationabstractTest-time medical image segmentation is a critical component in clinical practice, enabling pre-trained medical segmentation models to effectively adapt unseen medical samples with potential distribution shifts. However, existing methods are typically task-specific and restricted to certain diseases, with limited research focusing on test-time adaptation for universal segmentation, i.e. Segment Anything Model (SAM). Moreover, it is difficult to generate suitable prompts that can be effectively utilized by SAM for unseen test data without any label information. To address these challenges, we propose TTCS, a novel Test-Time medical image segmentation method using CLIP-guided SAM adaption, which achieves effective universal segmentation across diverse medical segmentation tasks. Specifically, we introduce a test-time prompt tuning strategy that leverages the semantic information from CLIP to generate precise prompts for each test data, effectively addressing the issue of poor prompt quality for SAM due to label scarcity. After generating the prompt for SAM, we implement an adaptive self-training strategy to further increase SAM’s robustness under distribution shifts. Our proposed method is inherently task-agnostic, and extensive experiments demonstrate the superior performance of our TTCS approach. Haotian Chen 0002, Yanyu Xu 0001, Li-Zhen Cui 0001 |
BIBM | 3 |
| 2024 | MeshSegmenter: Zero-Shot Mesh Semantic Segmentation via Texture Synthesis
Ziming Zhong, Yanyu Xu 0001, Jing Li 0117, Chaohui Yu, Shenghua Gao |
ECCV (77) | 2 |
| 2024 | Parameter-Efficient Fine-Tuning with ControlsabstractIn contrast to the prevailing interpretation of Low-Rank Adaptation (LoRA) as a means of simulating weight changes in model adaptation, this paper introduces an alternative perspective by framing it as a control process. Specifically, we conceptualize lightweight matrices in LoRA as control modules tasked with perturbing the original, complex, yet frozen blocks on downstream tasks. Building upon this new understanding, we conduct a thorough analysis on the controllability of these modules, where we identify and establish sufficient conditions that facilitate their effective integration into downstream controls. Moreover, the control modules are redesigned by incorporating nonlinearities through a parameter-free attention mechanism. This modification allows for the intermingling of tokens within the controllers, enhancing the adaptability and performance of the system. Empirical findings substantiate that, without introducing any additional parameters, this approach surpasses the existing LoRA algorithms across all assessed datasets and rank configurations. Chi Zhang 0123, Jingpu Cheng, Yanyu Xu 0001, Qianxiao Li |
ICML | 3 |
| 2024 | Multi-Scale Region-Aware Implicit Neural Network for Medical Images Matting
Yanyu Xu 0001, Yingzhi Xia, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026, Xinxing Xu |
MICCAI (9) | 1 |
| 2024 | MedMLP: An Efficient MLP-Like Network for Zero-Shot Retinal Image Classification
Menghan Zhou, Yanyu Xu 0001, Zhi Da Soh, Huazhu Fu, Rick Siow Mong Goh, Ching Yu Cheng, Yong Liu 0026, Liangli Zhen |
MICCAI (3) | 2 |
| 2024 | BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-RaysabstractMedical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and facilitate adapting task-specific models to new setups using fewer examples. However, existing MedVLP methods often differ in terms of datasets, preprocessing, and finetuning implementations. This pose great challenges in evaluating how well a MedVLP method generalizes to various clinically-relevant tasks due to the lack of unified, standardized, and comprehensive benchmark. To fill this gap, we propose BenchX, a unified benchmark framework that enables head-to-head comparison and systematical analysis between MedVLP methods using public chest X-ray datasets. Specifically, BenchX is composed of three components: 1) Comprehensive datasets covering nine datasets and four medical tasks; 2) Benchmark suites to standardize data preprocessing, train-test splits, and parameter selection; 3) Unified finetuning protocols that accommodate heterogeneous MedVLP methods for consistent task adaptation in classification, segmentation, and report generation, respectively. Utilizing BenchX, we establish baselines for nine state-of-the-art MedVLP methods and found that the performance of some early MedVLP methods can be enhanced to surpass more recent ones, prompting a revisiting of the developments and conclusions from prior works in MedVLP. Our code are available at https://github.com/yangzhou12/BenchX. Yang Zhou 0017, Tan Li Hui Faith, Yanyu Xu 0001, Sicong Leng, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh |
NeurIPS | 3 |
| 2023 | Generative Gradient Inversion via Over-Parameterized Networks in Federated LearningabstractFederated learning has gained recognitions as a secure approach for safeguarding local private data in collaborative learning. But the advent of gradient inversion research has posed significant challenges to this premise by enabling a third-party to recover groundtruth images via gradients. While prior research has predominantly focused on low-resolution images and small batch sizes, this study highlights the feasibility of reconstructing complex images with high resolutions and large batch sizes. The success of the proposed method is contingent on constructing an over-parameterized convolutional network, so that images are generated before fitting to the gradient matching requirement. Practical experiments demonstrate that the proposed algorithm achieves high-fidelity image recovery, surpassing state-of-the-art competitors that commonly fail in more intricate scenarios. Consequently, our study shows that local participants in a federated learning system are vulnerable to potential data leakage issues. Source code is available at https://github.com/czhang024/CI-Net. Chi Zhang 0123, Xiaoman Zhang, Ekanut Sotthiwat, Yanyu Xu 0001, Ping Liu 0004, Liangli Zhen, Yong Liu 0026 |
ICCV | 4 |
| 2023 | Minimal-Supervised Medical Image Segmentation via Vector Quantization Memory
Yanyu Xu 0001, Menghan Zhou, Yangqin Feng, Xinxing Xu, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026 |
MICCAI (3) | 1 |
| 2023 | GAMMA challenge: Glaucoma grAding from Multi-Modality imAges
Huihui Fang, Fei Li 0021, Huazhu Fu, Fengbin Lin, Jiongcheng Li, Yue Huang 0001, Qinji Yu, Sifan Song, Xinxing Xu, Yanyu Xu 0001, Wensai Wang, Shuai Lu 0003, Huiqi Li, Shihua Huang, Zhichao Lu, Chubin Ou, Xifei Wei, Bingyuan Liu, Riadh Kobbi, Xiaoying Tang 0001, Li Lin 0006, Hrvoje Bogunovic, José Ignacio Orlando, Xiulan Zhang, Yanwu Xu 0001 |
Medical Image Anal. | 11 |
| 2022 | Spherical DNNs and Their Applications in 360$^\circ$∘ Images and VideosabstractSpherical images or videos, as typical non-euclidean data, are usually stored in the form of 2D panoramas obtained through an equirectangular projection, which is neither equal area nor conformal. The distortion caused by the projection limits the performance of vanilla Deep Neural Networks (DNNs) designed for traditional euclidean data. In this paper, we design a novel Spherical Deep Neural Network (DNN) to deal with the distortion caused by the equirectangular projection. Specifically, we customize a set of components, including a spherical convolution, a spherical pooling, a spherical ConvLSTM cell and a spherical MSE loss, as the replacements of their counterparts in vanilla DNNs for spherical data. The core idea is to change the identical behavior of the conventional operations in vanilla DNNs across different feature patches so that they will be adjusted to the distortion caused by the variance of sampling rate among different feature patches. We demonstrate the effectiveness of our Spherical DNNs for saliency detection and gaze estimation in$360^\circ$videos. For saliency detection, we take the temporal coherence of an observer’s viewing process into consideration and propose to use a Spherical U-Net and a Spherical ConvLSTM to predict the saliency maps for each frame sequentially. As for gaze prediction, we propose to leverage a Spherical Encoder Module to extract spatial panoramic features, then we combine them with the gaze trajectory feature extracted by an LSTM for future gaze prediction. To facilitate the study of the$360^\circ$videos saliency detection, we further construct a large-scale$360^\circ$video saliency detection dataset that consists of 104$360^\circ$videos viewed by 20+ human subjects. Comprehensive experiments validate the effectiveness of our proposed Spherical DNNs for 360$^\circ$handwritten digit classification and sport classification, saliency detection and gaze tracking in$360^\circ$videos. We also visualize the regions contributing to the classification decisions in our proposed Spherical DNNs via the Grad-CAM technique in the classification task, and the results show that our Spherical DNNs constantly leverage reasonable and important regions for decision making, regardless the large distortions. All codes and dataset are available onhttps://github.com/svip-lab/SphericalDNNs. Yanyu Xu 0001, Shenghua Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Amodal Segmentation Based on Visible Region Segmentation and Shape PriorabstractAlmost all existing amodal segmentation methods make the inferences of occluded regions by using features corresponding to the whole image. This is against the human's amodal perception, where human uses the visible part and the shape prior knowledge of the target to infer the occluded region. To mimic the behavior of human and solve the ambiguity in the learning, we propose a framework, it firstly estimates a coarse visible mask and a coarse amodal mask. Then based on the coarse prediction, our model infers the amodal mask by concentrating on the visible region and utilizing the shape prior in the memory. In this way, features corresponding to background and occlusion can be suppressed for amodal mask estimation. Consequently, the amodal mask would not be affected by what the occlusion is given the same visible regions. The leverage of shape prior makes the amodal mask estimation more robust and reasonable. Our proposed model is evaluated on three datasets. Experiments show that our proposed model outperforms existing state-of-the-art methods. The visualization of shape prior indicates that the category-specific feature in the codebook has certain interpretability. The code is available at https://github.com/YutingXiao/Amodal-Segmentation-Based-on-Visible-Region-Segmentation-and-Shape-Prior. Yanyu Xu 0001, Ziming Zhong, Weixin Luo, Shenghua Gao |
AAAI | 2 |
| 2021 | Layout-Guided Novel View Synthesis From a Single Indoor PanoramaabstractExisting view synthesis methods mainly focus on the perspective images and have shown promising results. How-ever, due to the limited field-of-view of the pinhole cam-era, the performance quickly degrades when large cam-era movements are adopted. In this paper, we make the first attempt to generate novel views from a single indoor panorama and take the large camera translations into consideration. To tackle this challenging problem, we first use Convolutional Neural Networks (CNNs) to extract the deep features and estimate the depth map from the source-view image. Then, we leverage the room layout prior, a strong structural constraint of the indoor scene, to guide the generation of target views. More concretely, we estimate the room layout in the source view and transform it into the target viewpoint as guidance. Meanwhile, we also con-strain the room layout of the generated target-view images to enforce geometric consistency. To validate the effectiveness of our method, we further build a large-scale photo-realistic dataset containing both small and large camera translations. The experimental results on our challenging dataset demonstrate that our method achieves state-of-the-art performance. The project page is at https://github.com/bluestyle97/PNVS. Jia Zheng 0002, Yanyu Xu 0001, Rui Tang 0015, Shenghua Gao |
CVPR | 3 |
| 2021 | Prior Based Human CompletionabstractWe study a very challenging task, human image completion, which tries to recover the human body part with a reasonable human shape from the corrupted region. Since each human body part is unique, it is infeasible to restore the missing part by borrowing textures from other visible regions. Thus, we propose two types of learned priors to compensate for the damaged region. One is a structure prior, it uses a human parsing map to represent the human body structure. The other is a structure-texture correlation prior. It learns a structure and a texture memory bank, which encodes the common body structures and texture patterns, respectively. With the aid of these memory banks, the model could utilize the visible pattern to query and fetch a similar structure and texture pattern to introduce additional reasonable structures and textures for the corrupted region. Besides, since multiple potential human shapes are underlying the corrupted region, we propose multi-scale structure discriminators to further restore a plausible topological structure. Experiments on various large-scale benchmarks demonstrate the effectiveness of our proposed method. Zibo Zhao 0001, Wen Liu 0003, Yanyu Xu 0001, Xianing Chen, Weixin Luo, Bohui Zhu, Tong Liu 0037, Binqiang Zhao, Shenghua Gao |
CVPR | 3 |
| 2021 | Crowd Counting With Partial Annotations in an ImageabstractTo fully leverage the data captured from different scenes with different view angles while reducing the annotation cost, this paper studies a novel crowd counting setting, i.e. only using partial annotations in each image as training data. Inspired by the repetitive patterns in the annotated and unannotated regions as well as the ones between them, we design a network with three components to tackle those unannotated regions: i) in an Unannotated Regions Characterization (URC) module, we employ a memory bank to only store the annotated features, which could help the visual features extracted from these annotated regions flow to these unannotated regions; ii) For each image, Feature Distribution Consistency (FDC) regularizes the feature distributions of annotated head and unannotated head regions to be consistent; iii) a Cross-regressor Consistency Regularization (CCR) module is designed to learn the visual features of unannotated regions in a self-supervised style. The experimental results validate the effectiveness of our proposed model under the partial annotation setting for several datasets, such as ShanghaiTech, UCF-CC-50, UCF-QNRF, NWPU-Crowd and JHU-CROWD++. With only 10% annotated regions in each image, our proposed model achieves better performance than the recent methods and baselines under semi-supervised or active learning settings on all datasets. The code is https://github.com/svip-lab/CrwodCountingPAL. Yanyu Xu 0001, Ziming Zhong, Dongze Lian, Jing Li 0117, Xinxing Xu, Shenghua Gao |
ICCV | 1 |
| 2021 | Accurate depth estimation from a hybrid event-RGB stereo setupabstractEvent-based visual perception is becoming increasingly popular owing to interesting sensor characteristics enabling the handling of difficult conditions such as highly dynamic motion or challenging illumination. The mostly complementary nature of event cameras however still means that best results are achieved if the sensor is paired with a regular frame-based sensor. The present work aims at answering a simple question: Assuming that both cameras do not share a common optical center, is it possible to exploit the hybrid stereo setup's baseline to perform accurate stereo depth estimation? We present a learning based solution to this problem leveraging modern spatio-temporal input representations as well as a novel hybrid pyramid attention module. Results on real data demonstrate competitive performance against pure frame-based stereo alternatives as well as the ability to maintain the advantageous properties of event-based sensors. Xin Peng 0005, Yanyu Xu 0001, Shenghua Gao, Xia Wang 0002, Laurent Kneip |
IROS | 4 |
| 2021 | Partially-Supervised Learning for Vessel Segmentation in Ocular Images
Yanyu Xu 0001, Xinxing Xu, Shenghua Gao, Rick Siow Mong Goh, Daniel S. W. Ting, Yong Liu 0026 |
MICCAI (1) | 1 |
| 2021 | SUNNet: A novel framework for simultaneous human parsing and pose estimation
Yanyu Xu 0001, Zhixin Piao, Wen Liu 0003, Shenghua Gao |
Neurocomputing | 1 |
| 2020 | Geometric Structure Based and Regularized Depth Estimation From 360 Indoor ImageryabstractMotivated by the correlation between the depth and the geometric structure of a 360 indoor image, we propose a novel learning-based depth estimation framework that leverages the geometric structure of a scene to conduct depth estimation. Specifically, we represent the geometric structure of an indoor scene as a collection of corners, boundaries and planes. On the one hand, once a depth map is estimated, this geometric structure can be inferred from the estimated depth map; thus, the geometric structure functions as a regularizer for depth estimation. On the other hand, this estimation also benefits from the geometric structure of a scene estimated from an image where the structure functions as a prior. However, furniture in indoor scenes makes it challenging to infer geometric structure from depth or image data. An attention map is inferred to facilitate both depth estimation from features of the geometric structure and also geometric inferences from the estimated depth map. To validate the effectiveness of each component in our framework under controlled conditions, we render a synthetic dataset, Shanghaitech-Kujiale Indoor 360 dataset with 3550 360 indoor images. Extensive experiments on popular datasets validate the effectiveness of our solution. We also demonstrate that our method can also be applied to counterfactual depth. Yanyu Xu 0001, Jia Zheng 0002, Rui Tang 0015, Shugong Xu, Jingyi Yu 0001, Shenghua Gao |
CVPR | 2 |
| 2020 | SIRI: Spatial Relation Induced Network For Spatial Description ResolutionabstractSpatial Description Resolution, as a language-guided localization task, is proposed for target location in a panoramic street view, given corresponding language descriptions. Explicitly characterizing an object-level relationship while distilling spatial relationships are currently absent but crucial to this task. Mimicking humans, who sequentially traverse spatial relationship words and objects with a first-person view to locate their target, we propose a novel spatial relationship induced (SIRI) network. Specifically, visual features are firstly correlated at an implicit object-level in a projected latent space; then they are distilled by each spatial relationship word, resulting in each differently activated feature representing each spatial relationship. Further, we introduce global position priors to fix the absence of positional information, which may result in global positional reasoning ambiguities. Both the linguistic and visual features are concatenated to finalize the target localization. Experimental results on the Touchdown show that our method is around 24\% better than the state-of-the-art method in terms of accuracy, measured by an 80-pixel radius. Our method also generalizes well on our proposed extended dataset collected using the same settings as Touchdown. The code for this project is publicly available at https://github.com/wong-puiyiu/siri-sdr. Weixin Luo, Yanyu Xu 0001, Shugong Xu, Shenghua Gao |
NeurIPS | 3 |
| 2019 | PPGNet: Learning Point-Pair Graph for Line Segment DetectionabstractIn this paper, we present a novel framework to detect line segments in man-made environments. Specifically, we propose to describe junctions, line segments and relationships between them with a simple graph, which is more structured and informative than end-point representation used in existing line segment detection methods. In order to extract a line segment graph from an image, we further introduce the PPGNet, a convolutional neural network that directly infers a graph from an image. We evaluate our method on published benchmarks including York Urban and Wireframe datasets. The results demonstrate that our method achieves satisfactory performance and generalizes well on all the benchmarks. The source code of our work is available at https://github.com/svip-lab/PPGNet. Ning Bi, Jia Zheng 0002, Kun Huang 0001, Weixin Luo, Yanyu Xu 0001, Shenghua Gao |
CVPR | 8 |
| 2019 | Personalized Saliency and Its PredictionabstractNearly all existing visual saliency models by far have focused on predicting a universal saliency map across all observers. Yet psychology studies suggest that visual attention of different observers can vary significantly under specific circumstances, especially a scene is composed of multiple salient objects. To study such heterogenous visual attention pattern across observers, we first construct a personalized saliency dataset and explore correlations between visual attention, personal preferences, and image contents. Specifically, we propose to decompose a personalized saliency map (referred to as PSM) into a universal saliency map (referred to as USM) predictable by existing saliency detection models and a new discrepancy map across users that characterizes personalized saliency. We then present two solutions towards predicting such discrepancy maps, i.e., a multi-task convolutional neural network (CNN) framework and an extended CNN with Person-specific Information Encoded Filters (CNN-PIEF). Extensive experimental results demonstrate the effectiveness of our models for PSM prediction as well their generalization capability for unseen observers. Yanyu Xu 0001, Shenghua Gao, Nianyi Li, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Multiview Multitask Gaze Estimation With Deep Convolutional Neural NetworksabstractGaze estimation, which aims to predict gaze points with given eye images, is an important task in computer vision because of its applications in human visual attention understanding. Many existing methods are based on a single camera, and most of them only focus on either the gaze point estimation or gaze direction estimation. In this paper, we propose a novel multitask method for the gaze point estimation using multiview cameras. Specifically, we analyze the close relationship between the gaze point estimation and gaze direction estimation, and we use a partially shared convolutional neural networks architecture to simultaneously estimate the gaze direction and gaze point. Furthermore, we also introduce a new multiview gaze tracking data set that consists of multiview eye images of different subjects. As far as we know, it is the largest multiview gaze tracking data set. Comprehensive experiments on our multiview gaze tracking data set and existing data sets demonstrate that our multiview multitask gaze point estimation solution consistently outperforms existing methods. Dongze Lian, Lina Hu, Weixin Luo, Yanyu Xu 0001, Lixin Duan, Jingyi Yu 0001, Shenghua Gao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Gaze Prediction in Dynamic 360° Immersive VideosabstractThis paper explores gaze prediction in dynamic 360° immersive videos, i.e., based on the history scan path and VR contents, we predict where a viewer will look at an upcoming time. To tackle this problem, we first present the large-scale eye-tracking in dynamic VR scene dataset. Our dataset contains 208 360° videos captured in dynamic scenes, and each video is viewed by at least 31 subjects. Our analysis shows that gaze prediction depends on its history scan path and image contents. In terms of the image contents, those salient objects easily attract viewers' attention. On the one hand, the saliency is related to both appearance and motion of the objects. Considering that the saliency measured at different scales is different, we propose to compute saliency maps at different spatial scales: the sub-image patch centered at current gaze point, the sub-image corresponding to the Field of View (FoV), and the panorama image. Then we feed both the saliency maps and the corresponding images into a Convolutional Neural Network (CNN) for feature extraction. Meanwhile, we also use a Long-Short-Term-Memory (LSTM) to encode the history scan path. Then we combine the CNN features and LSTM features for gaze displacement prediction between gaze point at a current time and gaze point at an upcoming time. Extensive experiments validate the effectiveness of our method for gaze prediction in dynamic VR scenes. Yanyu Xu 0001, Yanbing Dong, Zhengzhong Sun, Zhiru Shi, Jingyi Yu 0001, Shenghua Gao |
CVPR | 1 |
| 2018 | Encoding Crowd Interaction With Deep Neural Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a challenging task because of the complex nature of humans. In this paper, we tackle the problem within a deep learning framework by considering motion information of each pedestrian and its interaction with the crowd. Specifically, motivated by the residual learning in deep learning, we propose to predict displacement between neighboring frames for each pedestrian sequentially. To predict such displacement, we design a crowd interaction deep neural network (CIDNN) which considers the different importance of different pedestrians for the displacement prediction of a target pedestrian. Specifically, we use an LSTM to model motion information for all pedestrians and use a multi-layer perceptron to map the location of each pedestrian to a high dimensional feature space where the inner product between features is used as a measurement for the spatial affinity between two pedestrians. Then we weight the motion features of all pedestrians based on their spatial affinity to the target pedestrian for location displacement prediction. Extensive experiments on publicly available datasets validate the effectiveness of our method for trajectory prediction. Yanyu Xu 0001, Zhixin Piao, Shenghua Gao |
CVPR | 1 |
| 2018 | Saliency Detection in 360 ^\circ ∘ Videos
Yanyu Xu 0001, Jingyi Yu 0001, Shenghua Gao |
ECCV (7) | 2 |
| 2018 | Semantic Human MattingabstractHuman matting, high quality extraction of humans from natural images, is crucial for a wide variety of applications. Since the matting problem is severely under-constrained, most previous methods require user interactions to take user designated trimaps or scribbles as constraints. This user-in-the-loop nature makes them difficult to be applied to large scale data or time-sensitive scenarios. In this paper, instead of using explicit user input constraints, we employ implicit semantic constraints learned from data and propose an automatic human matting algorithm Semantic Human Matting(SHM). SHM is the first algorithm that learns to jointly fit both semantic information and high quality details with deep networks. In practice, simultaneously learning both coarse semantics and fine details is challenging. We propose a novel fusion strategy which naturally gives a probabilistic estimation of the alpha matte. We also construct a very large dataset with high quality annotations consisting of 35,513 unique foregrounds to facilitate the learning and evaluation of human matting. Extensive experiments on this dataset and plenty of real images show that SHM achieves comparable results with state-of-the-art interactive matting methods. Quan Chen 0006, Tiezheng Ge, Yanyu Xu 0001, Zhiqiang Zhang 0011, Kun Gai |
ACM Multimedia | 3 |
| 2017 | Beyond Universal Saliency: Personalized Saliency Prediction with Multi-task CNNabstractSaliency detection is a long standing problem in computer vision. Tremendous efforts have been focused on exploring a universal saliency model across users despite their differences in gender, race, age, etc. Yet recent psychology studies suggest that saliency is highly specific than universal: individuals exhibit heterogeneous gaze patterns when viewing an identical scene containing multiple salient objects. In this paper, we first show that such heterogeneity is common and critical for reliable saliency prediction. Our study also produces the first database of personalized saliency maps (PSMs). We model PSM based on universal saliency map (USM) shared by different participants and adopt a multi-task CNN framework to estimate the discrepancy between PSM and USM. Comprehensive experiments demonstrate that our new PSM model and prediction scheme are effective and reliable. Yanyu Xu 0001, Nianyi Li, Jingyi Yu 0001, Shenghua Gao |
IJCAI | 1 |