VLDB 2026 Research / reviewers in the wild / expert
Wei Ji 0011
dblp:52/3220-11
· DBLP profile ↗
47ranked-venue papers
8as first author
44since 2021 · last 2026
0000-0003-4059-5902ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 6 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 6 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAM3-I: Segment Anything with InstructionsabstractJingjing Li, Yue Feng, Yuchen Guo, Jincai Huang, Wei Ji, Qi Bi, Yongri Piao, Miao Zhang, Xiaoqi Zhao, Qiang Chen, Shihao Zou, Huchuan Lu, Li Cheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jincai Huang 0003, Wei Ji 0011, Qi Bi, Yongri Piao, Miao Zhang 0004, Xiaoqi Zhao 0003, Qiang Chen 0007, Shihao Zou, Huchuan Lu, Li Cheng 0001 |
ACL (1) | 5 |
| 2026 | Anatomically and metabolically informed diffusion for unified denoising and segmentation in low-count PET imaging
Menghua Xia, Kuan-Yin Ko, Der-Shiun Wang, Mingkai Chen 0003, Huidong Xie, Wei Ji 0011, Jinsong Ouyang, Reimund Bayerlein, Benjamin A. Spencer, Quanzheng Li, Ramsey Derek Badawi, Georges El Fakhri, Chi Liu 0001 |
Medical Image Anal. | 8 |
| 2026 | Dose-aware diffusion model for 3D PET image denoising: Multi-institutional validation with reader study and real low-dose data
Huidong Xie, Weijie Gan, Reimund Bayerlein, Bo Zhou 0009, Mingkai Chen 0003, Michal Kulon, Annemarie Boustani, Kuan-Yin Ko, Der-Shiun Wang, Benjamin A. Spencer, Wei Ji 0011, Xiongchao Chen, Xueqi Guo, Menghua Xia, Yinchi Zhou, Hongyu An, Ulugbek Kamilov, Hanzhong Wang, Axel Rominger, Kuangyu Shi, Ge Wang 0001, Ramsey Derek Badawi, Chi Liu 0001 |
Medical Image Anal. | 11 |
| 2026 | Harmonized medical federated learning via redundancy-aware client consistency
Jingjun Yi, Yuexiang Li, Qi Bi, Wei Ji 0011, Huimin Huang 0002, Yawen Huang, Yefeng Zheng 0001, Feiyue Huang |
Pattern Recognit. | 5 |
| 2026 | Revisiting Fine-Grained Image Analysis by Semantic-Part AlignmentabstractFine-grained image analysis is widely recognized as highly challenging, since distinguishing individual differences within a certain category, species, or type often depends on tiny, subtle patterns. However, learning fine-grained semantic categories from these subtle part patterns is inherently fragile, as they can easily be overwhelmed by the dominant patterns resting in the coarse-category information. Therefore, how to enhance the relation between the fine-grained semantics and these subtle patterns is the key. To push this frontier, a novel semantic-part alignment (SPA) learning scheme is proposed in this paper. Its general idea is to firstly measure the relevance of each part to the fine-grained semantics, and then regularize the fine-grained visual representation learning. Specifically, it consists of three key components, namely, joint semantic-part modeling, semantic-part set modeling, and optimal semantic-part transport. The joint semantic-part modeling associates each part in an image with the fine-grained semantics in a latent space. Then, the optimal semantic-part transport component is devised to enhance the relation between fine-grained semantic embeddings and the discriminative part embeddings. Notably, the proposed SPA is plug-in-and-play, easy-to-implement, and insensitive to the latent embedding dimension and loss weight. Experiments show the proposed method can substantially boost performance on multiple fine-grained image analysis tasks across various baselines. Qi Bi, Jingjun Yi, Haolan Zhan, Wei Ji 0011, Gui-Song Xia |
IEEE Trans. Image Process. | 4 |
| 2025 | DGFamba: Learning Flow Factorized State Space for Visual Domain GeneralizationabstractDomain generalization aims to learn a representation from the source domain, which can be generalized to arbitrary unseen target domains. A fundamental challenge for visual domain generalization is the domain gap caused by the dramatic style variation whereas the image content is stable. The realm of selective state space, exemplified by VMamba, demonstrates its global receptive field in representing the content. However, the way exploiting the domain-invariant property for selective state space is rarely explored. In this paper, we propose a novel Flow Factorized State Space model, dubbed as DGFamba, for visual domain generalization. To maintain domain consistency, we innovatively map the style-augmented and the original state embeddings by flow factorization. In this latent flow space, each state embedding from a certain style is specified by a latent probability path. By aligning these probability paths in the latent space, the state embeddings are able to represent the same content distribution regardless of the style differences. Extensive experiments conducted on various visual domain generalization settings show its state-of-the-art performance. Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li |
AAAI | 5 |
| 2025 | Learning Fine-grained Domain Generalization via Hyperbolic State Space HallucinationabstractFine-grained domain generalization (FGDG) aims to learn a fine-grained representation that can be well generalized to unseen target domains when only trained on the source domain data. Compared with generic domain generalization, FGDG is particularly challenging in that the fine-grained category can be only discerned by some subtle and tiny patterns. Such patterns are particularly fragile under the cross-domain style shifts caused by illumination, color and etc. To push this frontier, this paper presents a novel Hyperbolic State Space Hallucination (HSSH) method. It consists of two key components, namely, state space hallucination (SSH) and hyperbolic manifold consistency (HMC). SSH enriches the style diversity for the state embeddings by firstly extrapolating and then hallucinating the source images. Then, the pre- and post- style hallucinate state embeddings are projected into the hyperbolic manifold. The hyperbolic state space models the high-order statistics, and allows a better discernment of the fine-grained patterns. Finally, the hyperbolic distance is minimized, so that the impact of style variation on fine-grained patterns can be eliminated. Experiments on three FGDG benchmarks demonstrate its state-of-the-art performance. Qi Bi, Jingjun Yi, Haolan Zhan, Wei Ji 0011, Gui-Song Xia |
AAAI | 4 |
| 2025 | DefMamba: Deformable Visual State Space ModelabstractRecently, state space models (SSM), particularly Mamba, have attracted significant attention from scholars due to their ability to effectively balance computational efficiency and performance. However, most existing visual Mamba methods flatten images into 1D sequences using predefined scan orders, which results the model being less capable of utilizing the spatial structural information of the image during the feature extraction process. To address this issue, we proposed a novel visual foundation model called Def-Mamba. This model includes a multi-scale backbone structure and deformable mamba (DM) blocks, which dynamically adjust the scanning path to prioritize important information, thus enhancing the capture and processing of relevant input features. By combining a deformable scanning (DS) strategy, this model significantly improves its ability to learn image structures and detects changes in object details. Numerous experiments have shown that Def-Mamba achieves state-of-the-art performance in various visual tasks, including image classification, object detection, instance segmentation, and semantic segmentation. The code is open source on DefMamba . Leiye Liu, Miao Zhang 0004, Jihao Yin, Tingwei Liu, Wei Ji 0011, Yongri Piao, Huchuan Lu |
CVPR | 5 |
| 2025 | A Simple Yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision Tasks
Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
ICCV | 6 |
| 2025 | D-CAM: Learning Generalizable Weakly-Supervised Medical Image Segmentation from Domain-Invariant CAM
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Huimin Huang 0002, Yuexiang Li, Shaoxin Li 0001, Xian Wu 0001, Yefeng Zheng 0001, Feiyue Huang |
MICCAI (5) | 5 |
| 2025 | AtlantisGS: Underwater Sparse-View Scene Reconstruction via Gaussian Splatting
Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Haolan Zhan, Yixian Shen, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
ACM Multimedia | 7 |
| 2025 | MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMsabstractThe advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark for evaluating Multi-Video Understanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1,824 meticulously curated question-answer pairs spanning 4,959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos.The benchmark will be made publicly available to foster future research. Yuanxing Zhang, Noah Wang, Ge Zhang 0009, Jian Yang 0037, Yanghai Wang, Xintao Wang 0002, Houyi Li, Wei Ji 0011, Pengfei Wan 0001, Wenhao Huang 0001, Zhaoxiang Zhang 0001 |
NeurIPS | 12 |
| 2025 | Degradation-Aware Dynamic Schrödinger Bridge for Unpaired Image RestorationabstractImage restoration is a fundamental task in computer vision and machine learning, which learns a mapping between the clear images and the degraded images under various conditions (e.g., blur, low-light, haze).
Yet, most existing image restoration methods are highly restricted by the requirement of degraded and clear image pairs, which limits the generalization and feasibility to enormous real-world scenarios without paired images.
To address this bottleneck, we propose a Degradation-aware Dynamic Schr\"{o}dinger Bridge (DDSB) for unpaired image restoration.
Its general idea is to learn a Schr\"{o}dinger Bridge between clear and degraded image distribution,
while at the same time emphasizing the physical degradation priors to reduce the accumulation of errors during the restoration process.
A Degradation-aware Optimal Transport (DOT) learning scheme is accordingly devised.
Training a degradation model to learn the inverse restoration process is particularly challenging, as it must be applicable across different stages of the iterative restoration process.
A Dynamic Transport with Consistency (DTC) learning objective is further proposed to reduce the loss of image details in the early iterations and therefore refine the degradation model.
Extensive experiments on multiple image degradation tasks show its state-of-the-art performance over the prior arts. Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Yixian Shen, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
NeurIPS | 7 |
| 2025 | A generalizable diffusion framework for 3D low-dose and few-view cardiac SPECT imaging
Huidong Xie, Weijie Gan, Wei Ji 0011, Xiongchao Chen, Alaa Alashi, Stephanie Thorn, Bo Zhou 0009, Menghua Xia, Xueqi Guo, Yi-Hwa Liu, Hongyu An, Ulugbek Kamilov, Ge Wang 0001, Albert J. Sinusas, Chi Liu 0001 |
Medical Image Anal. | 3 |
| 2025 | Learning Generalized Medical Image Representation by Decoupled Feature QueriesabstractMedical images are usually collected from multiple clinical centers with various types of scanners. When confronted with such significant cross-domain distribution discrepancy, a deep network tends to capture similar patterns by multiple channels, while different cross-domain patterns are also allowed to rest in the same channel. Such channel redundancy limits the expressive capability of a representation, resulting in less preferable generalization ability. To address this fundamental yet challenging issue, we propose a novel decoupled feature as query (DFQ) framework for domain generalized medical image representation learning. Its general idea is to leverage the channel-wise decoupled deep features as queries. Particularly, a deep instance whitening transform with restricted isometry is proposed, which enforces each channel orthogonal to the rest channels after decoupling. Besides, the long-range dependency between decoupled deep and shallow features is implicitly constrained to minimize channel redundancy throughout training. Extensive experiments show its state-of-the-art performance on three medical domain generalization tasks with four modalities. Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | GAD: Domain generalized diabetic retinopathy grading by grade-aware de-stylization
Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Wei Ji 0011, Yuexiang Li, Yefeng Zheng 0001 |
Pattern Recognit. | 6 |
| 2025 | Highly Efficient 3D Human Pose Tracking From Events With Spiking Spatiotemporal TransformerabstractEvent camera, as an asynchronous vision sensor capturing scene dynamics, presents new opportunities for highly efficient 3D human pose tracking. Existing approaches typically adopt modern-day Artificial Neural Networks (ANNs), such as CNNs or Transformer, where sparse events are converted into dense images or paired with additional gray-scale images as input. Such practices, however, ignore the inherent sparsity of events, resulting in redundant computations, increased energy consumption, and potentially degraded performance. Motivated by these observations, we introduce the first sparse Spiking Neural Networks (SNNs) framework for 3D human pose tracking based solely on events. Our approach eliminates the need to convert sparse data to dense formats or incorporate additional images, thereby fully exploiting the innate sparsity of input events. Central to our framework is a novel Spiking Spatio-temporal Transformer, which enables bi-directional spatio-temporal fusion of spike pose features and provides a guaranteed similarity measurement between binary spike features in spiking attention. Moreover, we have constructed a largescale synthetic dataset, SynEventHPD, that features a broad and diverse set of 3D human motions, as well as much longer hours of event streams. Empirical experiments demonstrate the superiority of our approach over existing state-of-the-art (SOTA) ANN-based methods, requiring only 19.1% FLOPs and 3.6% energy cost. Furthermore, our approach outperforms existing SNN-based benchmarks in this task, highlighting the effectiveness of our proposed SNN framework. The dataset will be released upon acceptance, and code can be found at https://github.com/JimmyZou/HumanPoseTracking_SNN. Shihao Zou, Yuxuan Mu, Wei Ji 0011, Zi-An Wang, Xinxin Zuo, Sen Wang 0003, Weixin Si, Li Cheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Cross-Level Multi-Instance Distillation for Self-Supervised Fine-Grained Visual CategorizationabstractHigh-quality annotation of fine-grained visual categories demands great expert knowledge, which is taxing and time consuming. Alternatively, learning fine-grained visual representation from enormous unlabeled images (e.g., species, brands) by self-supervised learning becomes a feasible solution. However, recent investigations find that existing self-supervised learning methods are less qualified to represent fine-grained categories. The bottleneck lies in that the pre-trained class-agnostic representation is built from every patch-wise embedding, while fine-grained categories are only determined by several key patches of an image. In this paper, we propose a Cross-level Multi-instance Distillation (CMD) framework to tackle this challenge. Our key idea is to consider the importance of each image patch in determining the fine-grained representation by multiple instance learning. To comprehensively learn the relation between informative patches and fine-grained semantics, the multi-instance knowledge distillation is implemented on both the region/image crop pairs from the teacher and student net, and the region-image crops inside the teacher / student net, which we term as intra-level multi-instance distillation and inter-level multi-instance distillation. Extensive experiments on several commonly used datasets, including CUB-200-2011, Stanford Cars and FGVC Aircraft, demonstrate that the proposed method outperforms the contemporary methods by up to 10.14% and existing state-of-the-art self-supervised learning approaches by up to 19.78% on both top-1 accuracy and Rank-1 retrieval metric. Source code is available at https://github.com/BiQiWHU/CMD. Qi Bi, Wei Ji 0011, Jingjun Yi, Haolan Zhan, Gui-Song Xia |
IEEE Trans. Image Process. | 2 |
| 2025 | Universal Fine-Grained Visual Categorization by Concept Guided LearningabstractExisting fine-grained visual categorization (FGVC) methods assume that the fine-grained semantics rest in the informative parts of an image. This assumption works well on favorable front-view object-centric images, but can face great challenges in many real-world scenarios, such as scene-centric images (e.g., street view) and adverse viewpoint (e.g., object reidentification, remote sensing). In such scenarios, the mis-/over-feature activation is likely to confuse the part selection and degrade the fine-grained representation. In this paper, we are motivated to design a universal FGVC framework for real-world scenarios. More precisely, we propose a concept guided learning (CGL), which models concepts of a certain fine-grained category as a combination of inherited concepts from its subordinate coarse-grained category and discriminative concepts from its own. The discriminative concepts is utilized to guide the fine-grained representation learning. Specifically, three key steps are designed, namely, concept mining, concept fusion, and concept constraint. On the other hand, to bridge the FGVC dataset gap under scene-centric and adverse viewpoint scenarios, a Fine-grained Land-cover Categorization Dataset (FGLCD) with 59,994 fine-grained samples is proposed. Extensive experiments show the proposed CGL: 1) has a competitive performance on conventional FGVC; 2) achieves state-of-the-art performance on fine-grained aerial scenes & scene-centric street scenes; 3) good generalization on object re-identification and fine-grained aerial object detection. The dataset and source code will be available at https://github.com/BiQiWHU/CGL. Qi Bi, Beichen Zhou, Wei Ji 0011, Gui-Song Xia |
IEEE Trans. Image Process. | 3 |
| 2024 | Learning Generalized Medical Image Segmentation from Decoupled Feature QueriesabstractDomain generalized medical image segmentation requires models to learn from multiple source domains and generalize well to arbitrary unseen target domain. Such a task is both technically challenging and clinically practical, due to the domain shift problem (i.e., images are collected from different hospitals and scanners). Existing methods focused on either learning shape-invariant representation or reaching consensus among the source domains. An ideal generalized representation is supposed to show similar pattern responses within the same channel for cross-domain images. However, to deal with the significant distribution discrepancy, the network tends to capture similar patterns by multiple channels, while different cross-domain patterns are also allowed to rest in the same channel. To address this issue, we propose to leverage channel-wise decoupled deep features as queries. With the aid of cross-attention mechanism, the long-range dependency between deep and shallow features can be fully mined via self-attention and then guides the learning of generalized representation. Besides, a relaxed deep whitening transformation is proposed to learn channel-wise decoupled features in a feasible way. The proposed decoupled fea- ture query (DFQ) scheme can be seamlessly integrate into the Transformer segmentation model in an end-to-end manner. Extensive experiments show its state-of-the-art performance, notably outperforming the runner-up by 1.31% and 1.98% with DSC metric on generalized fundus and prostate benchmarks, respectively. Source code is available at https://github.com/BiQiWHU/DFQ. Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
AAAI | 4 |
| 2024 | MedSegDiff-V2: Diffusion-Based Medical Image Segmentation with TransformerabstractThe Diffusion Probabilistic Model (DPM) has recently gained popularity in the field of computer vision, thanks to its image generation applications, such as Imagen, Latent Diffusion Models, and Stable Diffusion, which have demonstrated impressive capabilities and sparked much discussion within the community. Recent investigations have further unveiled the utility of DPM in the domain of medical image analysis, as underscored by the commendable performance exhibited by the medical image segmentation model across various tasks. Although these models were originally underpinned by a UNet architecture, there exists a potential avenue for enhancing their performance through the integration of vision transformer mechanisms. However, we discovered that simply combining these two models resulted in subpar performance. To effectively integrate these two cutting-edge techniques for the Medical image segmentation, we propose a novel Transformer-based Diffusion framework, called MedSegDiff-V2. We verify its effectiveness on 20 medical image segmentation tasks with different image modalities. Through comprehensive evaluation, our approach demonstrates superiority over prior state-of-the-art (SOTA) methodologies. Code is released at https://github.com/KidsWithTokens/MedSegDiff. Wei Ji 0011, Huazhu Fu, Min Xu 0009, Yueming Jin, Yanwu Xu 0001 |
AAAI | 2 |
| 2024 | Spider: A Unified Framework for Context-dependent Concept SegmentationabstractDifferent from the context-independent (CI) concepts such as human, car, and airplane, context-dependent (CD) concepts require higher visual understanding ability, such as camouflaged object and medical lesion. Despite the rapid advance of many CD understanding tasks in respective branches, the isolated evolution leads to their limited cross-domain generalisation and repetitive technique innovation. Since there is a strong coupling relationship between foreground and background context in CD tasks, existing methods require to train separate models in their focused domains. This restricts their real-world CD concept understanding towards artificial general intelligence (AGI). We propose a unified model with a single set of parameters, Spider, which only needs to be trained once. With the help of the proposed concept filter driven by the image-mask group prompt, Spider is able to understand and distinguish diverse strong context-dependent concepts to accurately capture the Prompter's intention. Without bells and whistles, Spider significantly outperforms the state-of-the-art specialized models in 8 different context-dependent segmentation tasks, including 4 natural scenes (salient, camouflaged, and transparent objects and shadow) and 4 medical lesions (COVID-19, polyp, breast, and skin lesion with color colonoscopy, CT, ultrasound, and dermoscopy modalities). Besides, Spider shows obvious advantages in continuous learning. It can easily complete the training of new tasks by fine-tuning parameters less than 1% and bring a tolerable performance degradation of less than 5% for all old tasks. The source code will be publicly available at https://github.com/Xiaoqi-Zhao-DLUT/Spider-UniCDSeg. Xiaoqi Zhao 0003, Youwei Pang, Wei Ji 0011, Baicheng Sheng, Jiaming Zuo, Lihe Zhang, Huchuan Lu |
ICML | 3 |
| 2024 | Hallucinated Style Distillation for Single Domain Generalization in Medical Image Segmentation
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Shaoxin Li 0001, Yuexiang Li, Yefeng Zheng 0001, Feiyue Huang |
MICCAI (10) | 5 |
| 2024 | Learning Spectral-Decomposited Tokens for Domain Generalized Semantic SegmentationabstractThe rapid development of Vision Foundation Model (VFM) brings inherent out-domain generalization for a variety of down-stream tasks. Among them, domain generalized semantic segmentation (DGSS) holds unique challenges as the cross-domain images share common pixel-wise content information but vary greatly in terms of the style. In this paper, we present a novel Spectral-dEcomposed Token (SET) learning framework to advance the frontier. Delving into further than existing fine-tuning token & frozen backbone paradigm, the proposed SET especially focuses on the way learning style-invariant features from these learnable tokens. Particularly, the frozen VFM features are first decomposed into the phase and amplitude components in the frequency space, which mainly contain the information of content and style, respectively, and then separately processed by learnable tokens for task-specific information extraction. Particularly, the frozen VFM features are first decomposed into the phase and amplitude components in the frequency space, which mainly contain the information of content and style, respectively, and then separately processed by learnable tokens for task-specific information extraction.After the decomposition, style variation primarily impacts the token-based feature enhancement within the amplitude branch. To address this issue, we further develop an attention optimization method to bridge the gap between style-affected representation and static tokens during inference. Extensive cross-domain experiments show its state-of-the-art performance. Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
ACM Multimedia | 5 |
| 2024 | Samba: Severity-aware Recurrent Modeling for Cross-domain Medical Image GradingabstractDisease grading is a crucial task in medical image analysis. Due to the continuous progression of diseases, i.e., the variability within the same level and the similarity between adjacent stages, accurate grading is highly challenging.
Furthermore, in real-world scenarios, models trained on limited source domain datasets should also be capable of handling data from unseen target domains.
Due to the cross-domain variants, the feature distribution between source and unseen target domains can be dramatically different, leading to a substantial decrease in model performance.
To address these challenges in cross-domain disease grading, we propose a Severity-aware Recurrent Modeling (Samba) method in this paper.
As the core objective of most staging tasks is to identify the most severe lesions, which may only occupy a small portion of the image, we propose to encode image patches in a sequential and recurrent manner.
Specifically, a state space model is tailored to store and transport the severity information by hidden states.
Moreover, to mitigate the impact of cross-domain variants, an Expectation-Maximization (EM) based state recalibration mechanism is designed to map the patch embeddings into a more compact space.
We model the feature distributions of different lesions through the Gaussian Mixture Model (GMM) and reconstruct the intermediate features based on learnable severity bases.
Extensive experiments show the proposed Samba outperforms the VMamba baseline by an average accuracy of 23.5\%, 5.6\% and 4.1\% on the cross-domain grading of fatigue fracture, breast cancer and diabetic retinopathy, respectively.
Source code is available at \url{https://github.com/BiQiWHU/Samba}. Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Haolan Zhan, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
NeurIPS | 4 |
| 2024 | Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic SegmentationabstractThe emerging vision foundation model (VFM) has inherited the ability to generalize to unseen images.
Nevertheless, the key challenge of domain-generalized semantic segmentation (DGSS) lies in the domain gap attributed to the cross-domain styles, i.e., the variance of urban landscape and environment dependencies.
Hence, maintaining the style-invariant property with varying domain styles becomes the key bottleneck in harnessing VFM for DGSS.
The frequency space after Haar wavelet transformation provides a feasible way to decouple the style information from the domain-invariant content, since the content and style information are retained in the low- and high- frequency components of the space, respectively.
To this end, we propose a novel Frequency-Adapted (FADA) learning scheme to advance the frontier.
Its overall idea is to separately tackle the content and style information by frequency tokens throughout the learning process.
Particularly, the proposed FADA consists of two branches, i.e., low- and high- frequency branches. The former one is able to stabilize the scene content, while the latter one learns the scene styles and eliminates its impact to DGSS.
Experiments conducted on various DGSS settings show the state-of-the-art performance of our FADA and its versatility to a variety of VFMs.
Source code is available at \url{https://github.com/BiQiWHU/FADA}. Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Wei Ji 0011, Yuexiang Li, Yefeng Zheng 0001 |
NeurIPS | 6 |
| 2024 | Unleashing Multispectral Video's Potential in Semantic Segmentation: A Semi-supervised Viewpoint and New UAV-View BenchmarkabstractThanks to the rapid progress in RGB & thermal imaging, also known as multispectral imaging, the task of multispectral video semantic segmentation, or MVSS in short, has recently drawn significant attentions. Noticeably, it offers new opportunities in improving segmentation performance under unfavorable visual conditions such as poor light or overexposure. Unfortunately, there are currently very few datasets available, including for example MVSeg dataset that focuses purely toward eye-level view; and it features the sparse annotation nature due to the intensive demands of labeling process. To address these key challenges of the MVSS task, this paper presents two major contributions: the introduction of MVUAV, a new MVSS benchmark dataset, and the development of a dedicated semi-supervised MVSS baseline - SemiMV. Our MVUAV dataset is captured via Unmanned Aerial Vehicles (UAV), which offers a unique oblique bird’s-eye view complementary to the existing MVSS datasets; it also encompasses a broad range of day/night lighting conditions and over 30 semantic categories. In the meantime, to better leverage the sparse annotations and extra unlabeled RGB-Thermal videos, a semi-supervised learning baseline, SemiMV, is proposed to enforce consistency regularization through a dedicated Cross-collaborative Consistency Learning (C3L) module and a denoised temporal aggregation strategy. Comprehensive empirical evaluations on both MVSeg and MVUAV benchmark datasets have showcased the efficacy of our SemiMV baseline. Wei Ji 0011, Wenbo Li 0001, Yilin Shen, Li Cheng 0001, Hongxia Jin |
NeurIPS | 1 |
| 2023 | Multispectral Video Semantic Segmentation: A Benchmark Dataset and BaselineabstractRobust and reliable semantic segmentation in complex scenes is crucial for many real-life applications such as autonomous safe driving and nighttime rescue. In most approaches, it is typical to make use of RGB images as input. They however work well only in preferred weather conditions; when facing adverse conditions such as rainy, overexposure, or low-light, they often fail to deliver satisfactory results. This has led to the recent investigation into multispectral semantic segmentation, where RGB and thermal infrared (RGBT) images are both utilized as input. This gives rise to significantly more robust segmentation of image objects in complex scenes and under adverse conditions. Nevertheless, the present focus in single RGBT image input restricts existing methods from well addressing dynamic real-world scenes. Motivated by the above observations, in this paper, we set out to address a relatively new task of semantic segmentation of multispectral video input, which we refer to as Multispectral Video Semantic Segmentation, or MVSS in short. An in-house MVSeg dataset is thus curated, consisting of 738 calibrated RGB and thermal videos, accompanied by 3,545 fine-grained pixel-level semantic annotations of 26 categories. Our dataset contains a wide range of challenging urban scenes in both daytime and nighttime. Moreover, we propose an effective MVSS baseline, dubbed MVNet, which is to our knowledge the first model to jointly learn semantic representations from multispectral and temporal contexts. Comprehensive experiments are conducted using various semantic segmentation models on the MVSeg dataset. Empirically, the engagement of multispectral video input is shown to lead to significant improvement in semantic segmentation; the effectiveness of our MVNet baseline has also been verified. Wei Ji 0011, Cheng Bian, Zongwei Zhou, Jiaying Zhao, Alan L. Yuille, Li Cheng 0001 |
CVPR | 1 |
| 2023 | Animal3D: A Comprehensive Dataset of 3D Animal Pose and ShapeabstractAccurately estimating the 3D pose and shape is an essential step towards understanding animal behavior, and can potentially benefit many downstream applications, such as wildlife conservation. However, research in this area is held back by the lack of a comprehensive and diverse dataset with high-quality 3D pose and shape annotations. In this paper, we propose Animal3D, the first comprehensive dataset for mammal animal 3D pose and shape estimation. Animal3D consists of 3379 images collected from 40 mammal species, high-quality annotations of 26 key-points, and importantly the pose and shape parameters of the SMAL [50] model. All annotations were labeled and checked manually in a multi-stage process to ensure highest quality results. Based on the Animal3D dataset, we benchmark representative shape and pose estimation models at: (1) supervised learning from only the Animal3D data, (2) synthetic to real transfer from synthetically generated images, and (3) fine-tuning human pose and shape estimation models. Our experimental results demonstrate that predicting the 3D shape and pose of animals across species remains a very challenging task, despite significant advances in human pose estimation. Our results further demonstrate that synthetic pre-training is a viable strategy to boost the model performance. Overall, Animal3D opens new directions for facilitating future research in animal 3D pose and shape estimation, and is publicly available. Jiacong Xu, Yi Zhang 0099, Wufei Ma, Artur Jesslen, Pengliang Ji, Qixin Hu, Qihao Liu, Jiahao Wang 0001, Wei Ji 0011, Chen Wang 0049, Xiaoding Yuan, Prakhar Kaushik, Guofeng Zhang 0020, Jie Liu 0044, Yushan Xie, Yawen Cui, Alan L. Yuille, Adam Kortylewski |
ICCV | 11 |
| 2023 | SemanticRT: A Large-Scale Dataset and Method for Robust Semantic Segmentation in Multispectral ImagesabstractGrowing interests in multispectral semantic segmentation (MSS) have been witnessed in recent years, thanks to the unique advantages of combining RGB and thermal infrared images to tackle challenging scenarios with adverse conditions. However, unlike traditional RGB-only semantic segmentation, the lack of a large-scale MSS dataset has become a hindrance to the progress of this field. To address this issue, we introduce a SemanticRT dataset - the largest MSS dataset to date, comprising 11,371 high-quality, pixel-level annotated RGB-thermal image pairs. It is 7 times larger than the existing MFNet dataset, and covers a wide variety of challenging scenarios in adverse lighting conditions such as low-light and pitch black. Further, a novel Explicit Complement Modeling (ECM) framework is developed to extract modality-specific information, which is propagated through a robust cross-modal feature encoding and fusion process. Extensive experiments demonstrate the advantages of our approach and dataset over the existing counterparts. Our new dataset may also facilitate further development and evaluation of existing and new MSS algorithms. Wei Ji 0011, Cheng Bian, Zhicheng Zhang 0005, Li Cheng 0001 |
ACM Multimedia | 1 |
| 2023 | DVSOD: RGB-D Video Salient Object DetectionabstractSalient object detection (SOD) aims to identify standout elements in a scene, with recent advancements primarily focused on integrating depth data (RGB-D) or temporal data from videos to enhance SOD in complex scenes. However, the unison of two types of crucial information remains largely underexplored due to data constraints. To bridge this gap, we in this work introduce the DViSal dataset, fueling further research in the emerging field of RGB-D video salient object detection (DVSOD). Our dataset features 237 diverse RGB-D videos alongside comprehensive annotations, including object and instance-level markings, as well as bounding boxes and scribbles. These resources enable a broad scope for potential research directions. We also conduct benchmarking experiments using various SOD models, affirming the efficacy of multimodal video input for salient object detection. Lastly, we highlight some intriguing findings and promising future research avenues. To foster growth in this field, our dataset and benchmark results are publicly accessible at: https://dvsod.github.io/. Wei Ji 0011, Size Wang, Wenbo Li 0001, Li Cheng 0001 |
NeurIPS | 2 |
| 2023 | Learning rotation equivalent scene representation from instance-level semantics: A novel top-down perspectiveabstractThis paper focuses on rotation variant scene recognition. Different from existing rotation invariant recognition approaches which learn from either rotated images or rotated convolutional filters in a bottom-up manner, a new top-down perspective by learning is explored from instance-level semantic representation. The goal is to eliminate the convolutional feature differences in bottom-up feature propagation caused by the rotation sensitive nature of convolution operation. Our rotation equivalent convolutional neural network (RE-CNN) scheme consists of three components. Firstly, our key instance selection module highlights the instances strongly related to the scene scheme regardless of their orientation. Secondly, our key instance aggregation module builds a scene representation invariant to the position change of each instance caused by rotation. Finally, our semantic fusion module allows the framework to be organized as a whole and implements rotation regularization. Notably, our RE-CNN scheme can be adapted to existing CNNs in a plug-in-and-play manner. Extensive experiments on rotation variant scene recognition benchmarks from four domains demonstrate the state-of-the-art performance and generalization capability of the proposed RE-CNN. Qi Bi, Shaodi You, Wei Ji 0011, Theo Gevers |
Comput. Vis. Image Underst. | 3 |
| 2023 | Delving into Calibrated Depth for Accurate RGB-D Salient Object Detection
Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Huchuan Lu, Li Cheng 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | C$^{2}$DFNet: Criss-Cross Dynamic Filter Network for RGB-D Salient Object DetectionabstractThe ability to deal with intra and inter-modality features has been critical to the development of RGB-D salient object detection. While many works have advanced in leaps and bounds in this field, most existing methods have not taken their way down into the inherent differences between the RGB and depth data due to widely adopted conventional convolution in which fixed parameter kernels are applied during inference. To promote intra and inter-modality interaction conditioned on various scenarios, as RGB and depth data are processed independently and later fused interactively, we develop a new insight and a better model. In this paper, we introduce a criss-cross dynamic filter network by decoupling dynamic convolution. First, we propose a Model-specific Dynamic Enhanced Module (MDEM) that dynamically enhances the intra-modality features with global context guidance. Second, we propose a Scene-aware Dynamic Fusion Module (SDFM) to realize dynamic feature selection between two modalities. As a result, our model achieves accurate predictions of salient objects. Extensive experiments demonstrate that our method achieves competitive performance over 28 state-of-the-art RGB-D methods on 7 public datasets. Miao Zhang 0004, Shunyu Yao 0004, Beiqi Hu, Yongri Piao, Wei Ji 0011 |
IEEE Trans. Multim. | 5 |
| 2022 | Generating Diverse and Natural 3D Human Motions from TextabstractAutomated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem with a two-stage approach: text2length sampling and text2motion generation. Text2length involves sampling from the learned distribution function of motion lengths conditioned on the input text. This is followed by our text2motion module using temporal variational autoen-coder to synthesize a diverse set of human motions of the sampled lengths. Instead of directly engaging with pose sequences, we propose motion snippet code as our internal motion representation, which captures local semantic motion contexts and is empirically shown to facilitate the generation of plausible motions faithful to the input text. Moreover, a large-scale dataset of scripted 3D Human motions, HumanML3D, is constructed, consisting of 14,616 motion clips and 44,970 text descriptions. Chuan Guo 0002, Shihao Zou, Xinxin Zuo, Sen Wang 0003, Wei Ji 0011, Li Cheng 0001 |
CVPR | 5 |
| 2022 | Exploring Denoised Cross-video Contrast for Weakly-supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization aims to localize actions in untrimmed videos with only video-level labels. Most existing methods address this problem with a “localization-by-classification” pipeline that localizes action regions based on snippet-wise classification sequences. Snippet-wise classifications are unfortunately error prone due to the sparsity of video-level labels. Inspired by recent success in unsupervised contrastive representation learning, we propose a novel denoised cross-video contrastive algorithm, aiming to enhance the feature discrimination ability of video snippets for accurate temporal action localization in the weakly-supervised setting. This is enabled by three key designs: 1) an effective pseudo-label denoising module to alleviate the side effects caused by noisy contrastive features, 2) an efficient region-level feature contrast strategy with a region-level memory bank to capture “global” contrast across the entire dataset, and 3) a diverse contrastive learning strategy to enable action-background separation as well as intra-class compactness & inter-class separability. Extensive experiments on THUMOS14 and ActivityNet v1.3 demonstrate the superior performance of our approach. Tianyu Yang 0003, Wei Ji 0011, Jue Wang 0001, Li Cheng 0001 |
CVPR | 3 |
| 2022 | Promoting Saliency From Depth: Deep Unsupervised RGB-D Saliency Detection
Wei Ji 0011, Qi Bi, Chuan Guo 0002, Jie Liu 0044, Li Cheng 0001 |
ICLR | 1 |
| 2022 | TW-GAN: Topology and width aware GAN for retinal artery/vein classification
Wenting Chen, Kai Ma 0002, Wei Ji 0011, Cheng Bian, Chunyan Chu, LinLin Shen, Yefeng Zheng 0001 |
Medical Image Anal. | 4 |
| 2022 | DMRA: Depth-Induced Multi-Scale Recurrent Attention Network for RGB-D Saliency DetectionabstractIn this work, we propose a novel depth-induced multi-scale recurrent attention network for RGB-D saliency detection, named as DMRA. It achieves dramatic performance especially in complex scenarios. There are four main contributions of our network that are experimentally demonstrated to have significant practical merits. First, we design an effective depth refinement block using residual connections to fully extract and fuse cross-modal complementary cues from RGB and depth streams. Second, depth cues with abundant spatial information are innovatively combined with multi-scale contextual features for accurately locating salient objects. Third, a novel recurrent attention module inspired by Internal Generative Mechanism of human brain is designed to generate more accurate saliency results via comprehensively learning the internal semantic relation of the fused feature and progressively optimizing local details with memory-oriented scene understanding. Finally, a cascaded hierarchical feature fusion strategy is designed to promote efficient information interaction of multi-level contextual features and further improve the contextual representability of model. In addition, we introduce a new real-life RGB-D saliency dataset containing a variety of complex scenarios that has been widely used as a benchmark dataset in recent RGB-D saliency detection research. Extensive empirical experiments demonstrate that our method can accurately identify salient objects and achieve appealing performance against 18 state-of-the-art RGB-D saliency models on nine benchmark datasets. Wei Ji 0011, Ge Yan 0006, Yongri Piao, Shunyu Yao 0004, Miao Zhang 0004, Li Cheng 0001, Huchuan Lu |
IEEE Trans. Image Process. | 1 |
| 2021 | Calibrated RGB-D Salient Object DetectionabstractComplex backgrounds and similar appearances between objects and their surroundings are generally recognized as challenging scenarios in Salient Object Detection (SOD). This naturally leads to the incorporation of depth information in addition to the conventional RGB image as input, known as RGB-D SOD or depth-aware SOD. Meanwhile, this emerging line of research has been considerably hindered by the noise and ambiguity that prevail in raw depth images. To address the aforementioned issues, we propose a Depth Calibration and Fusion (DCF) framework that contains two novel components: 1) a learning strategy to calibrate the latent bias in the original depth maps towards boosting the SOD performance; 2) a simple yet effective cross reference module to fuse features from both RGB and depth modalities. Extensive empirical experiments demonstrate that the proposed approach achieves superior performance against 27 state-of-the-art methods. Moreover, our depth calibration strategy alone can work as a preprocessing step; empirically it results in noticeable improvements when being applied to existing cutting-edge RGB-D SOD models. Source code is available at https://github.com/jiwei0921/DCF. Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Shunyu Yao 0004, Qi Bi, Kai Ma 0002, Yefeng Zheng 0001, Huchuan Lu, Li Cheng 0001 |
CVPR | 1 |
| 2021 | Learning Calibrated Medical Image Segmentation via Multi-Rater Agreement ModelingabstractIn medical image analysis, it is typical to collect multiple annotations, each from a different clinical expert or rater, in the expectation that possible diagnostic errors could be mitigated. Meanwhile, from the computer vision practitioner viewpoint, it has been a common practice to adopt the ground-truth labels obtained via either the majority-vote or simply one annotation from a preferred rater. This process, however, tends to overlook the rich information of agreement or disagreement ingrained in the raw multi-rater annotations. To address this issue, we propose to explicitly model the multi-rater (dis-)agreement, dubbed MRNet, which has two main contributions. First, an expertise-aware inferring module or EIM is devised to embed the expertise level of individual raters as prior knowledge, to form high-level semantic features. Second, our approach is capable of reconstructing multi-rater gradings from coarse predictions, with the multi-rater (dis-)agreement cues being further exploited to improve the segmentation performance. To our knowledge, our work is the first in producing calibrated predictions under different expertise levels for medical image segmentation. Extensive empirical experiments are conducted across five medical segmentation tasks of diverse imaging modalities. In these experiments, superior performance of our MRNet is observed comparing to the state-of-the-arts, indicating the effectiveness and applicability of our MRNet toward a wide range of medical segmentation tasks. Source code is publicly available. Wei Ji 0011, Kai Ma 0002, Cheng Bian, Qi Bi, Hanruo Liu, Li Cheng 0001, Yefeng Zheng 0001 |
CVPR | 1 |
| 2021 | Dynamic Context-Sensitive Filtering Network for Video Salient Object DetectionabstractThe ability to capture inter-frame dynamics has been critical to the development of video salient object detection (VSOD). While many works have achieved great success in this field, a deeper insight into its dynamic nature should be developed. In this work, we aim to answer the following questions: How can a model adjust itself to dynamic variations as well as perceive fine differences in the real-world environment; How are the temporal dynamics well introduced into spatial information over time? To this end, we propose a dynamic context-sensitive filtering network (DCFNet) equipped with a dynamic context-sensitive filtering module (DCFM) and an effective bidirectional dynamic fusion strategy. The proposed DCFM sheds new light on dynamic filter generation by extracting location-related affinities between consecutive frames. Our bidirectional dynamic fusion strategy encourages the interaction of spatial and temporal information in a dynamic manner. Experimental results demonstrate that our proposed method can achieve state-of-the-art performance on most VSOD datasets while ensuring a real-time speed of 28 fps. The source code is publicly available at https://github.com/OIPLab-DUT/DCFNet. Miao Zhang 0004, Jie Liu 0044, Yongri Piao, Shunyu Yao 0004, Wei Ji 0011, Huchuan Lu, Zhongxuan Luo |
ICCV | 6 |
| 2021 | Local-Global Dual Perception Based Deep Multiple Instance Learning for Retinal Disease Classification
Qi Bi, Wei Ji 0011, Cheng Bian, Lijun Gong, Hanruo Liu, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (8) | 3 |
| 2021 | Joint Semantic Mining for Weakly Supervised RGB-D Salient Object DetectionabstractTraining saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when only weak supervision signals are available. This paper is set to tackle the problem of weakly-supervised RGB-D salient object detection. The key insight in this effort is the idea of maintaining per-pixel pseudo-labels with iterative refinements by reconciling the multimodal input signals in our joint semantic mining (JSM). Considering the large variations in the raw depth map and the lack of explicit pixel-level supervisions, we propose spatial semantic modeling (SSM) to capture saliency-specific depth cues from the raw depth and produce depth-refined pseudo-labels. Moreover, tags and captions are incorporated via a fill-in-the-blank training in our textual semantic modeling (TSM) to estimate the confidences of competing pseudo-labels. At test time, our model involves only a light-weight sub-network of the training pipeline, i.e., it requires only an RGB image as input, thus allowing efficient inference. Extensive evaluations demonstrate the effectiveness of our approach under the weakly-supervised setting. Importantly, our method could also be adapted to work in both fully-supervised and unsupervised paradigms. In each of these scenarios, superior performance has been attained by our approach with comparing to the state-of-the-art dedicated methods. As a by-product, a CapS dataset is constructed by augmenting existing benchmark training set with additional image tags and captions. Wei Ji 0011, Qi Bi, Miao Zhang 0004, Yongri Piao, Huchuan Lu, Li Cheng 0001 |
NeurIPS | 2 |
| 2020 | Accurate RGB-D Salient Object Detection via Collaborative Learning
Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Huchuan Lu |
ECCV (18) | 1 |
| 2020 | LFNet: Light Field Fusion Network for Salient Object DetectionabstractIn this work, we propose a novel light field fusion network-LFNet, a CNNs-based light field saliency model using 4D light field data containing abundant spatial and contextual information. The proposed method can reliably locate and identify salient objects even in a complex scene. Our LFNet contains a light field refinement module (LFRM) and a light field integration module (LFIM) which can fully refine and integrate focusness, depths and objectness cues from light field image. The LFRM learns the light field residual between light field and RGB images for refining features with useful light field cues, and then the LFIM weights each refined light field feature and learns spatial correlation between them to predict saliency maps. Our method can take full advantage of light field information and achieve excellent performance especially in complex scenes, e.g., similar foreground and background, multiple or transparent objects and low-contrast environment. Experiments show our method outperforms the state-of-the-art 2D, 3D and 4D methods across three light field datasets. Miao Zhang 0004, Wei Ji 0011, Yongri Piao, Yu Zhang 0165, Huchuan Lu |
IEEE Trans. Image Process. | 2 |
| 2019 | Depth-Induced Multi-Scale Recurrent Attention Network for Saliency DetectionabstractIn this work, we propose a novel depth-induced multi-scale recurrent attention network for saliency detection. It achieves dramatic performance especially in complex scenarios. There are three main contributions of our network that are experimentally demonstrated to have significant practical merits. First, we design an effective depth refinement block using residual connections to fully extract and fuse multi-level paired complementary cues from RGB and depth streams. Second, depth cues with abundant spatial information are innovatively combined with multi-scale context features for accurately locating salient objects. Third, we boost our model's performance by a novel recurrent attention module inspired by Internal Generative Mechanism of human brain. This module can generate more accurate saliency results via comprehensively learning the internal semantic relation of the fused feature and progressively optimizing local details with memory-oriented scene understanding. In addition, we create a large scale RGB-D dataset containing more complex scenarios, which can contribute to comprehensively evaluating saliency models. Extensive experiments on six public datasets and ours demonstrate that our method can accurately identify salient objects and achieve consistently superior performance over 16 state-of-the-art RGB and RGB-D approaches. Yongri Piao, Wei Ji 0011, Miao Zhang 0004, Huchuan Lu |
ICCV | 2 |