VLDB 2026 Research / reviewers in the wild / expert
Huijie Fan
dblp:36/8729
· DBLP profile ↗
47ranked-venue papers
6as first author
35since 2021 · last 2026
0000-0002-8548-861XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 21 since 2021Artificial intelligence and machine learning · 23 · 3 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unleashing the Potential of Large Language Models for Text-to-Image Generation Through Autoregressive Representation AlignmentabstractWe present Autoregressive Representation Alignment (ARRA), a new training framework that unlocks global-coherent text-to-image generation in autoregressive LLMs without architectural modifications. Different from prior works that require complex architectural redesigns, ARRA aligns LLM's hidden states with visual representations from external visual foundational models via a global visual alignment loss and a hybrid token, . This token enforces dual constraints: local next-token prediction and global semantic distillation, enabling LLMs to implicitly learn spatial and contextual coherence while retaining their original autoregressive paradigm. Extensive experiments validate ARRA's plug-and-play versatility. When training T2I LLMs from scratch, ARRA reduces FID by 16.6% (ImageNet), 12.0% (LAION-COCO) for autoregressive LLMs like LlamaGen, without modifying original architecture and inference mechanism. For training from text-generation-only LLMs, ARRA reduces FID by 25.5% (MIMIC-CXR), 8.8% (DeepEyeNet) for advanced LLMs like Chameleon. For domain adaptation, ARRA aligns general-purpose LLMs with specialized models (e.g., BioMedCLIP), achieving an 18.6% FID reduction over direct fine-tuning on medical imaging (MIMIC-CXR). These results demonstrate that training objective redesign, rather than architectural modifications, can resolve cross-modal global coherence challenges. ARRA offers a complementary paradigm for advancing autoregressive models. Jiawei Liu 0003, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu |
AAAI | 4 |
| 2026 | MGAF: LiDAR-Camera 3D Object Detection With Multiple Guidance and Adaptive FusionabstractRecent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However, most of them overlook the complementary interaction and guidance between LiDAR and camera. In this work, we propose a novel multi-modality 3D objection detection method, with multi-guided global interaction and LiDAR-guided adaptive fusion, named MGAF. Specifically, we introduce sparse depth guidance (SDG) and LiDAR occupancy guidance (LOG) to generate 3D features with sufficient depth and spatial information. The designed semantic segmentation network captures category and orientation prior information for raw point clouds. In the following, an Adaptive Fusion Dual Transformer (AFDT) is developed to adaptively enhance the interaction of different modal BEV features from both global and bidirectional perspectives. Meanwhile, additional downsampling with sparse height compression and multi-scale dual-path transformer (MSDPT) are designed in order to enlarge the receptive fields of different modal features. Finally, a temporal fusion module is introduced to aggregate features from previous frames. Notably, the proposed AFDT is general, which also shows superior performance on other models. Our framework has undergone extensive experimentation on the large-scale nuScenes dataset, Waymo Open Dataset, and long-range Argoverse2 dataset, consistently demonstrating state-of-the-art performance. Baojie Fan, Caixia Xia, Huijie Fan, Fengyu Xu 0001, Jiandong Tian |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Domain Consistency Representation Learning for Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) exhibits a contradictory relationship between intra-domain discrimination and inter-domain gaps when learning from continuous data. Intra-domain discrimination focuses on individual nuances (i.e., clothing type, accessories,etc.), while inter-domain gaps emphasize domain consistency. Achieving a trade-off between maximizing intra-domain discrimination and minimizing inter-domain gaps is a crucial challenge for improving LReID performance. Most existing methods strive to reduce inter-domain gaps through knowledge distillation to maintain domain consistency. However, they often ignore intra-domain discrimination. To address this challenge, we propose a novel domain consistency representation learning (DCR) model that explores global and attribute-wise representations as a bridge to balance intra-domain discrimination and inter-domain gaps. At the intra-domain level, we explore the complementary relationship between global and attribute-wise representations to improve discrimination among similar identities. Excessive learning intra-domain discrimination can lead to catastrophic forgetting. We further develop an attribute-oriented anti-forgetting (AF) strategy that explores attribute-wise representations to enhance inter-domain consistency, and propose a knowledge consolidation (KC) strategy to facilitate knowledge transfer. Extensive experiments show that our DCR achieves superior performance compared to state-of-the-art LReID methods. Our code is available at https://github.com/LiuShiBen/DCR. Shiben Liu, Huijie Fan, Qiang Wang 0015, Weihong Ren, Yandong Tang, Yang Cong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | DVG-Diffusion: Dual-View-Guided Diffusion Model for CT Reconstruction From X-RaysabstractDirectly reconstructing 3D CT volume from few-view 2D X-rays using an end-to-end deep learning network is a challenging task, as X-ray images are merely projection views of the 3D CT volume. In this work, we facilitate complex 2D X-ray image to 3D CT mapping by incorporating new view synthesis, and reduce the learning difficulty through view-guided feature alignment. Specifically, we propose a dual-view guided diffusion model (DVG-Diffusion), which couples a real input X-ray view and a synthesized new X-ray view to jointly guide CT reconstruction. First, a novel view parameter-guided encoder captures features from X-rays that are spatially aligned with CT. Next, we concatenate the extracted dual-view features as conditions for the latent diffusion model to learn and refine the CT latent representation. Finally, the CT latent representation is decoded into a CT volume in pixel space. By incorporating view parameter guided encoding and dual-view guided CT reconstruction, our DVG-Diffusion can achieve an effective balance between high fidelity and perceptual quality for CT reconstruction. Experimental results demonstrate our method outperforms state-of-the-art methods. Based on experiments, the comprehensive analysis and discussions for views and reconstruction are also presented. The model and code are available at https://github.com/xiexing0916/DVG-Diffusion. Jiawei Liu 0003, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu |
IEEE Trans. Image Process. | 3 |
| 2025 | PGFormer: Prompt guide network for underwater image enhancementabstractUnderwater images are often influenced by light scattering and refraction, which leads to color deviation and poor quality. The enhancement of underwater images is significant for high-level semantic learning but also challenging. In this paper, we introduce PGFormer, a novel underwater image enhancement network that leverages prompt priors by integrating global and local prior information to improve underwater image quality. PGFormer comprises a global local enhancement module (GLEM) and a prompt-guided forward feedback network (PGFN). The GLEM extracts robust feature information through global and local feature modulation, whereas PGFN introduces prompt information into the local optimization process to further enhance local expression and refinement. Extensive experiments on various underwater datasets show that our method outperforms existing state-of-the-art techniques in terms of both visual quality and quantitative performance. Xin Luan, Huijie Fan, Qiang Wang 0015, Yandong Tang |
CEC | 2 |
| 2025 | Joint Attention Mechanism and Multi-task Learning for Weakly Supervised Skin Image SegmentationabstractIn recent years, the application of deep learning technology in the field of medical image segmentation has become increasingly mature and has made certain progress. Applying this technique requires the use of a large number of medical image datasets with pixel-level annotations, however, the cost of pixel-level annotation of medical images is high. For this situation, we propose a weakly supervised medical image semantic segmentation model based on attention mechanism and multi-task learning. The attention mechanism can be used to focus on the information that is more critical to the current task in a large number of input information, and multi-task learning can share information among related tasks to promote network learning. As a result, the model can achieve better segmentation performance by using only image-level groundtruth labels. Since semantic segmentation task and saliency detection task are dense pixel prediction tasks,and class activation map (CAM) extracted from classification task is also used to represent pixel-level semantic information, it shows that these two tasks are highly related to the semantic segmentation task. Based on this premise,We propose the module of finding similarity from attention(SFA), which is used to calculate the similarity between multi-task feature maps to learn the correlation between tasks, and use similarity to optimize the prediction results of specific tasks. In addition, in order to extensively explore context relationships in dense pixel prediction tasks to achieve more accurate prediction results, we also propose a global context module (GCM) that pays attention to the global information of dense prediction tasks, captures long-distance dependencies in images, and improves the segmentation performance of the network. We did a lot of experiments, our method achieves 68.56% and 62.32% Miou on ISBI2016 and ISIC2017 datasets, respectively, and F1 -score is 78.25% on PH2 data set, significantly outperforming several recent state-of-the-art weakly supervised semantic segmentation methods. Yujianing Wang, Qiang Wang 0015, Huijie Fan |
CEC | 3 |
| 2025 | All-Day Multi-Camera Multi-Target TrackingabstractThe capability of tracking objects in low-light environments like nighttime is crucial for numerous real-world applications. However, previous Multi-Camera Multi-Target(MCMT) tracking methods are primarily focused on tracking during daytime with favorable lighting, overlooking the challenge posed by low-light conditions. The main difficulty of tracking under low-light condition is the lack of detailed visible appearance features. To address this issue, we incorporate the infrared modality into MCMT tracking framework to provide more useful information. We constructed the first Multi-modality (RGBT) Multi-camera Multi-target tracking dataset named M3Track, which contains sequences captured in low-light environments, laying a solid foundation for all-day multi-camera tracking. Based on the proposed dataset, we propose All-Day Multi-Camera Multi-Target tracking network, termed as ADM-CMT. Specifically, we propose an All-Day Mamba Fusion(ADMF) module to fuse information from different modalities adaptively. Within ADMF, the Lighting Guidance Model(LGM) extracts lighting relevant information to guide the fusion process. Furthermore, the Nearby Target Collection(NTC) strategy is designed to enhance tracking accuracy by leveraging information derived from surrounding objects of targets. Experiments conducted on M3Track demonstrate that ADMCMT exhibits strong generalization across different lighting conditions. The code will be released at https://github.com/QTRACKY/ADMCMT. Huijie Fan, Yihao Zhen, Tinghui Zhao, Baojie Fan, Qiang Wang 0015 |
CVPR | 1 |
| 2025 | RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy Prediction
Baojie Fan, Yuyu Jiang, Jiandong Tian, Huijie Fan |
ICCV | 6 |
| 2025 | Selective Aggregation for Low-Rank Adaptation in Federated LearningabstractWe investigate LoRA in federated learning through the lens of the asymmetry analysis of the learned $A$ and $B$ matrices. In doing so, we uncover that $A$ matrices are responsible for learning general knowledge, while $B$ matrices focus on capturing client-specific knowledge. Based on this finding, we introduce Federated Share-A Low-Rank Adaptation (FedSA-LoRA), which employs two low-rank trainable matrices $A$ and $B$ to model the weight update, but only $A$ matrices are shared with the server for aggregation. Moreover, we delve into the relationship between the learned $A$ and $B$ matrices in other LoRA variants, such as rsLoRA and VeRA, revealing a consistent pattern. Consequently, we extend our FedSA-LoRA method to these LoRA variants, resulting in FedSA-rsLoRA and FedSA-VeRA. In this way, we establish a general paradigm for integrating LoRA with FL, offering guidance for future work on subsequent LoRA variants combined with FL. Extensive experimental results on natural language understanding and generation tasks demonstrate the effectiveness of the proposed method. Our code is available at https://github.com/Pengxin-Guo/FedSA-LoRA. Pengxin Guo 0001, Shuang Zeng, Huijie Fan, Liangqiong Qu |
ICLR | 4 |
| 2025 | Learning Generalizable 3D Manipulation With 10 DemonstrationsabstractLearning robust and generalizable manipulation skills from few demonstrations remains a key challenge in robotics, with broad applications in industrial automation and service robotics. Although recent imitation learning methods have achieved impressive results, they often require a large amount of demonstration data and struggle to generalize across different spatial variants. In this work, we propose a framework that learns 3D manipulation policies from only 10 demonstrations while achieving robust generalization to unseen spatial configurations through semantic-guided perception and spatial-equivariant policy learning. Our framework consists of two key modules: a Semantic Guided Perception module that extracts task-aware 3D representations from RGB-D inputs using semantic priors and a Spatial Generalized Decision module implementing a diffusion-based policy that preserves spatial equivariance through denoising. Central to our framework is a spatially equivariant training strategy, which adapts 2D data augmentation principles to 3D manipulation by maintaining gripper-object spatial relationships during trajectory augmentation. We validate our framework through extensive experiments on both simulation benchmarks and real-world robotic systems. Our method demonstrates a significant improvement in success rates over state-of-the-art approaches on a series of challenging tasks, particularly under significant object pose variations. This work shows significant potential to advance efficient and generalizable manipulation skill learning in real-world applications. Yang Cong, Bohao Huang, Jiahao Long, Ronghan Chen, Huijie Fan |
IROS | 7 |
| 2025 | TJCMNet: An Efficient Vision-Text Joint Identity Clues Mining Network for Visible-Infrared Person Re-IdentificationabstractRetrieving images for Visible-Infrared Person Re-identification task is challenging, because of the huge modality discrepancy caused by the different imaging principle of RGB and infrared cameras. Existing approaches rely on seeking distinctive information within unified visual feature space, ignoring the stable identity information brought by textual description. To overcome these problems, this letter propose a novel Text-vision Joint Clue Mining (TJCM) network to aggregate vision and text features, then distill the joint knowledge for enhancing the modality-shared branch. Specifically, we first extract modality-shared and textual features using a parameter-shared vision encoder and a text encoder. Then, a text-vision co-refinement module is proposed to refine the implicit information within vision feature and text feature, then aggregate them into joint feature. Finally, introduce the heterogeneous distillation alignment loss provides enhancement for modality-shared feature through joint knowledge distillation at feature-level and logit-level. Our TJCMNet achieves significant improvements over the state-of-the-art methods on three mainstream datasets. Zhuxuan Cheng, Zhijia Zhang, Huijie Fan, XingQi Na |
IEEE Signal Process. Lett. | 3 |
| 2025 | FMambaIR: A Hybrid State-Space Model and Frequency Domain for Image RestorationabstractWith the development of deep learning, impressive progress has been made in the field of image restoration. The existing methods mainly rely on CNN and Transformer to obtain multi-scale feature information. However, these methods rarely integrate frequency domain information effectively during feature extraction, limiting their performance in image restoration. Additionally, few have combined Mamba with the Fourier domain for image restoration, which limits Mamba’s ability to perceive global degradation in the frequency domain. Therefore, we propose a new image restoration model called FMambaIR, which utilizes the complementarity between frequency and Mamba for image restoration. The core of FMambaIR is the F-Mamba block, which combines Fourier transform and Mamba for global degradation perception modeling. Specifically, F-Mamba adopts a dual branch complementary structure, including spatial Mamba branches and Fourier frequency domain global modeling. Mamba models the long-range dependencies of the entire image features, and the frequency branch utilizes Fourier to extract global degraded features from the image. Finally, we use a forward feedback network to integrate local information, which is beneficial for improving the recovery details. We comprehensively evaluate FMambaIR on several image restoration tasks, including underwater image enhancement, remote sensing image dehazing, and low-light image enhancement. The experimental results demonstrate that FMambaIR not only achieves superior performance compared to state-of-the-art methods but also significantly reduces computational complexity. Our code is available at https://github.com/mickoluan/FMambaIR. Xin Luan, Huijie Fan, Qiang Wang 0015, Shiben Liu, Xiaofeng Li 0001, Yandong Tang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Synergistic Prompting Learning for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection, as a foundational task in human-centric understanding, aims to detect interactive triplets in real-world scenarios. To better distinguish diverse HOIs within an open-world context, current HOI detectors utilize pre-trained Visual-Language Models (VLMs) to extract prior knowledge through textual prompts (i.e., descriptive texts for each HOI instance). However, relying on predetermined descriptive texts, such approaches only acquire a fixed set of textual knowledge for HOI prediction, consequently resulting in inferior performance and limited generalization. To remedy this, we propose a novel VLM-based method, which jointly performs prompting learning from both visual and textual perspectives and synergizes visual-textual prompting for HOI detection. Initially, we design a hierarchical adaptation architecture to perform progressive prompting: visual prompting is facilitated through gradual token migration from VLM's image encoder, while textual prompting is initialized with progressively leveled interaction descriptions. In addition, to synergize the visual-textual prompting learning, a text-supervising and image-tuning loop is introduced, in which the text-supervising stage guides visual prompting learning through contrastive learning and the image-tuning stage refines textual prompting by modal matching. Finally, we employ an interaction-aware knowledge merging mechanism to effectively transfer visual-textual knowledge encapsulated within synergistic prompting for HOI detection. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones, under both supervised and zero-shot settings. Jinguo Luo, Weihong Ren, Zhiyong Wang 0009, Xi'ai Chen, Huijie Fan, Zhi Han, Honghai Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Diverse Representations Embedding for Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) aims to continuously learn from sequential data streams, enabling cross-camera matching of individuals over time. A critical challenge in LReID lies in balancing the preservation of previously acquired knowledge with the incremental acquisition of new information, due to task-level gaps and limited representation capacity. Conventional methods relying on CNN backbones struggle to fully capture the diverse perspectives of each instance, leading to suboptimal model performance. To tackle these limitations, we propose a diverse representation embedding (DRE) framework that balances preserving old knowledge with adapting to new information. Specifically, our DRE incorporates a robust Transformer-based backbone that utilizes maximum embedding (ME) and multiple class tokens to generate overlapping representations for each instance. To further enhance the model's representation capacity, we design an adaptive constraint module (ACM), which performs integration and discrimination operations on overlapping representations to yield diverse yet diverse representations. Furthermore, we propose two strategies: knowledge update (KU) and knowledge preservation (KP), implemented within the adjustment and learner models, respectively. The KU strategy enhances the learner model's ability to adapt to new information by leveraging prior knowledge from the adjustment model. The KP strategy ensures the retention of historical knowledge while maintaining the model's adaptability. Extensive experiments validate that our DRE surpasses state-of-the-art approaches across large-scale, occluded, and holistic datasets, demonstrating significant performance gains. Our code is available at https://github.com/LiuShiBen/DRE. Shiben Liu, Huijie Fan, Qiang Wang 0015, Xi'ai Chen, Zhi Han, Yandong Tang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object DetectionabstractRecent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However, most of them overlook the complementary interaction and guidance be-tween LiDAR and camera. In this work, we propose a novel multi-modality 3D objection detection method, named GA-Fusion, with LiDAR-guided global interaction and adaptive fusion. Specifically, we introduce sparse depth guidance (SDG) and LiDAR occupancy guidance (LOG) to generate 3D features with sufficient depth information. In the following, LiDAR-guided adaptive fusion transformer (LGAFT) is developed to adaptively enhance the interaction of different modal BEV features from a global perspective. Meanwhile, additional downsampling with sparse height compression and multi-scale dual-path transformer (MSDPT) are de-signed to enlarge the receptive fields of different modal features. Finally, a temporal fusion module is introduced to ag-gregate features from previous frames. GAFusion achieves state-of-the-art 3D object detection results with 73.6% mAP and 74.9% NDS on the nuScenes test set. Baojie Fan, Jiandong Tian, Huijie Fan |
CVPR | 4 |
| 2024 | Residual Denoising Diffusion ModelsabstractWe propose residual denoising diffusion models (RDDM), a novel dual diffusion process that decouples the traditional single denoising diffusion process into residual diffusion and noise diffusion. This dual diffusion framework expands the denoising-based diffusion models, initially uninterpretable for image restoration, into a unified and interpretable model for both image generation and restoration by introducing residuals. Specifically, our residual diffusion represents directional diffusion from the target image to the degraded input image and explicitly guides the reverse generation process for image restoration, while noise diffusion represents random perturbations in the diffusion process. The residual prioritizes certainty, while the noise emphasizes diversity, enabling RDDM to effectively unify tasks with varying certainty or diversity requirements, such as image generation and restoration. We demonstrate that our sampling process is consistent with that of DDPM and DDIM through coefficient transformation, and propose a partially path-independent generation process to better understand the reverse process. Notably, our RDDM enables a generic UNet, trained with only an L1 loss and a batch size of 1, to compete with state-of-the-art image restoration methods. We provide code and pre-trained models to encourage further exploration, application, and development of our innovative framework (https://github.com/nachifurlRDDM). Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Yandong Tang, Liangqiong Qu |
CVPR | 3 |
| 2024 | Uni-YOLO: Vision-Language Model-Guided YOLO for Robust and Fast Universal Detection in the Open WorldabstractUniversal object detectors aim to detect any object in any scene without human annotation, exhibiting superior generalization. However, the current universal object detectors show degraded performance in harsh weather, and their insufficient real-time capabilities limit their application. In this paper, we present Uni-YOLO, a universal detector designed for complex scenes with real-time performance. Uni-YOLO is a one-stage object detector that uses general object confidence to distinguish between objects and backgrounds, and employs a grid cell regression method for real-time detection. To improve its robustness in harsh weather conditions, the input of Uni-YOLO is adaptively enhanced with a physical model-based enhancement module. During training and inference, Uni-YOLO is guided by the extensive knowledge of the vision-language model CLIP. An object augmentation method is proposed to improve generalization in training by utilizing multiple source datasets with heterogeneous annotations. Furthermore, an online self-enhancement method is proposed to allow Uni-YOLO to further focus on specific objects through self-supervised fine-tuning in a given scene. Extensive experiments on public benchmarks and a UAV deployment are conducted to validate its superiority and practical value. Weihong Ren, Xi'ai Chen, Huijie Fan, Yandong Tang, Zhi Han |
ACM Multimedia | 4 |
| 2024 | Feature distillation and guide network for unsupervised underwater image enhancement
Xin Luan, Qiang Wang 0015, Huijie Fan, Xiai Chen, Zhi Han, Yandong Tang |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Unsupervised person re-identification based on adaptive information supplementation and foreground enhancementabstractAbstract Unsupervised person re‐identification has attracted vital interest because of its ability to protect privacy, significantly lower the expense of manual annotation, and eliminate the need for data labels. General unsupervised methods train the network only through global features, which causes the fine‐grained information contained in local features to be ignored in the recognition process, resulting in large amounts of label noise and affecting the recognition accuracy. Moreover, more robust pedestrian features can also improve the accuracy of clustering and enable unsupervised person re‐identification to obtain better results. To address these issues, first, a dual‐branch structure was proposed, which separately obtains the global features of the pedestrian and the local features by dividing the global features into a few equal sections. Then, an adaptive information supplementation (AIS) method based on the k‐nearest neighbor algorithm is designed to ascertain each local feature's relevance to the global features, calculating adaptive weight scores for information supplementation. Finally, these weight scores are used to reallocate the weights of the global features in each part, acquiring features that contain more pedestrian information during the representation learning process. These better features are used to reduce label noise to obtain more accurate pseudo‐labels. Second, an adaptive foreground enhancement module (AFEM) was proposed and inserted before clustering to increase the robustness of pedestrian features, which increases the precision of the pseudo‐labels that are produced after clustering. Experiments on Market‐1501, DukeMTMC‐reID, and MSMT17 demonstrate that the proposed method achieves better results than state‐of‐the‐art methods in fully unsupervised person re‐identification tasks. Qiang Wang 0015, Huijie Fan, Shengpeng Fu, Yandong Tang |
IET Image Process. | 3 |
| 2024 | Wavelet-pixel domain progressive fusion network for underwater image enhancement
Shiben Liu, Huijie Fan, Qiang Wang 0015, Zhi Han, Yandong Tang |
Knowl. Based Syst. | 2 |
| 2024 | Skip Connection Aggregation Transformer for Occluded Person ReidentificationabstractThe occlusion problem is a significant challenge for person reidentification. Recently, transformer-based methods have been introduced to solve the occlusion problem and achieve performance improvements. However, the existing methods only apply the features of the last transformer layer and fail to consider the alignment of visible body parts. They also ignore fine-grained local features. Thus, they usually suffer from misalignment in occluded image matching. We observe that features from the high layers of the transformer focus on classification information and global features, while those from the middle layers pay more attention to pedestrians. We think that making full use of the features of different layers will facilitate alignment and then will promote reidentification accuracy. Therefore, we propose a novel skip connection aggregation transformer (SCAT) network by utilizing features from different transformer layers to increase the diversity of features and align visible body parts in occluded images. The diverse features include the following: first, features of the middle layer, which focus on the pedestrian in nonoccluded regions and favor alignment, second, features of high layers, which focus on global information, third fine-grained local features, which are obtained by the part pooling encoder and the fusion reconstruction module. The part pooling encoder and the fusion reconstruction module are proposed to obtain part-based local features and fused local features, respectively. The experimental results on the occluded, partial, and holistic benchmarks demonstrate that our method can significantly promote the accuracy of occluded person reidentification. Huijie Fan, Qiang Wang 0015, Sheng-Peng Fu, Yandong Tang |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | QueryTrack: Joint-Modality Query Fusion Network for RGBT TrackingabstractExisting RGB-Thermal trackers usually treat intra-modal feature extraction and inter-modal feature fusion as two separate processes, therefore the mutual promotion of extraction and fusion is neglected. Then, the complementary advantages of RGB-T fusion are not fully exploited, and the independent feature extraction is not adaptive to modal quality fluctuation during tracking. To address the limitations, we design a joint-modality query fusion network, in which the intra-modal feature extraction and the inter-modal fusion are coupled together and promote each other via joint-modality queries. The queries are initialized based on the multimodal features of the current frame, making the subsequent fusion adaptive to modal quality fluctuation during tracking. Then the joint-modality query fusion (JQF) utilizes the queries to interact with RGB-T features, allowing the intra-modal enhancement and the inter-modal interactions to be unified for mutual promotion. In this way, JQF can distinguish and enhance the complementary modality features, while filtering out redundant information. For real-time tracking, we propose regional cross-attention for cross-modal interactions to reduce computational cost. Our end-to-end tracker sets a new state-of-the-art performance on multiple RGBT tracking benchmarks including LasHeR, VTUAV, RGBT234 and GTOT, while running at a real-time speed. Huijie Fan, Zhencheng Yu, Qiang Wang 0015, Baojie Fan, Yandong Tang |
IEEE Trans. Image Process. | 1 |
| 2024 | Online Video Sparse Noise Removing via Nonlocal Robust PCAabstractOnline schemes and nonlocal similarity are two effective approaches for strengthening robust principal component analysis (RPCA) techniques in video denoising. However, their limitations are also evident. The online scheme is usually highly efficient but lacks consideration of regional appearance information, thus it cannot effectively handle videos with complex dynamics such as object movements. On the other hand, nonlocal similarity is used to better utilize regional information but incurs a heavy computational cost. Moreover, these two techniques are incompatible and challenging to work together. To overcome this barrier and harness the advantages of both approaches, this paper proposes a novel online nonlocal RPCA method. 1) A clustering based nonlocal strategy (ClusNonlocal) is adopted, which not only greatly reduces the computation cost, but also forms low-dimensional subspaces for online processing; 2) a new weighted RPCA model is proposed, which regards samples with different importances and improves the performance of subspace pursuit and video recovery; 3) a multi-level subspace updating scheme and weighted projection method is proposed, which keeps the performance of online video data processing at a high level at all time. A series of video denoising experiments are carried out to demonstrate the overall advantages of our procedure over several other ones, in terms of both visual quality and running speed. Zhi Han, Huijie Fan, Yandong Tang, Yao Wang 0003 |
IEEE Trans. Multim. | 4 |
| 2024 | A Shadow Imaging Bilinear Model and Three-Branch Residual Network for Shadow RemovalabstractThe current shadow removal pipeline relies on the detected shadow masks, which have limitations for penumbras and tiny shadows, and results in an excessively long pipeline. To address these issues, we propose a shadow imaging bilinear model and design a novel three-branch residual (TBR) network for shadow removal. Our bilinear model reveals the single-image shadow removal process and can explain why simply increasing the brightness of shadow areas cannot remove shadows without artifacts. We considerably shorten the shadow removal pipeline by modeling illumination compensation and developing a single-stage shadow removal network without additional detection and refinement networks. Specifically, our network consists of three task branches, i.e., shadow image reconstruction, shadow matte estimation, and shadow removal. To merge these three branches and enhance the shadow removal branch, we design a model-based TBR module. Multiple TBR modules are cascaded to generate an intensive information flow and facilitate feature integration among the three branches. Thus, our network ensures the fidelity of nonshadow areas and restores the light intensity of shadow areas through three-branch collaboration. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. The model and code are available at https://github.com/nachifur/TBRNet. Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Jiandong Tian, Yandong Tang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Underwater image enhancement via a channel-wise transmission estimation networkabstractAbstract Underwater image enhancement for image processing and underwater robotic vision have recently attracted much academic attention. However, in most existing methods, underwater image enhancement is completed with a simple assumption: the attenuation coefficients are unified across the color channels. This assumption leads to unstable and visually unpleasing enhancement results. Moreover, these methods cannot be successfully applied to explore relatively independent transmissions from multiple color channels with complimentary feature information. To address these challenges, a novel channel‐wise transmission estimation network (CTEN) is proposed, which aims to pioneer the exploration of the transmission difference across the color channels in an underwater scene. Specifically, a color‐specific correction module is proposed to automatically quantify the transmission ability of multiple color channels in the underwater environment. Furthermore, a channel‐wise transmission estimation module is designed to simultaneously explore the relative independence of multi‐color channels and estimate the medium transmissions for each color channel, which represents the attenuation degree of different color radiances after reflecting in the water. Then, a novel residual strategy is introduced to integrate these two modules to complete the underwater enhancement. Using the model, the authors are able to provide an answer as to why channel‐wise transmission estimation are better than single transmission estimation and establish a generalization theory to show the effect of the independent transmission estimation model for each color channel. Experiments on several underwater image datasets verify the superiority of the proposed CTEN model. Qiang Wang 0015, Huijie Fan |
IET Image Process. | 3 |
| 2023 | Region Selective Fusion Network for Robust RGB-T TrackingabstractRGB-T tracking utilizes thermal infrared images as a complement to visible light images in order to perform more robust visual tracking in various scenarios. However, the highly aligned RGB-T image pairs introduces redundant information, the modal quality fluctuation during tracking also brings unreliable information. Existing RGB-T trackers usually use channelwise multi-modal feature fusion in which the low-quality features degrades the fused features and causes trackers to drift. In this work, we propose a region selective fusion network that first evaluates each image region by cross-modal and cross-region modeling, then removes low-quality redundant region features to alleviate the negative effects caused by unreliable information in multi-modal fusion. Besides, the region removal scheme brings a efficiency boost as redundant features are removed progressively, this enables the tracker to run at a high tracking speed.Extensive experiments show that the proposed tracker achieves competitive performance with a real-time tracking speed on multiple RGB-T tracking benchmarks including LasHeR, RGBT234 and GTOT. Zhencheng Yu, Huijie Fan, Qiang Wang 0015, Ziwan Li, Yandong Tang |
IEEE Signal Process. Lett. | 2 |
| 2023 | A Decoupled Multi-Task Network for Shadow RemovalabstractShadow removal, which aims to restore the illumination in shadow regions, is challenging due to the diversity of shadows in terms of location, intensity, shape, and size. Different from most multi-task methods, which design elaborate multi-branch or multi-stage structures for better shadow removal, we introduce feature decomposition to learn better feature representations. Specifically, we propose a single-stage and decoupled multi-task network (DMTN) to explicitly learn the decomposed features for shadow removal, shadow matte estimation, and shadow image reconstruction. First, we propose several coarse-to-fine semi-convolution (SMC) modules to capture features sufficient for joint learning of these three tasks. Second, we design a theoretically supported feature decoupling layer to explicitly decouple the learned features into shadow image features and shadow matte features via weight reassignment. Last, these features are converted to a target shadow-free image, affiliated shadow matte, and shadow image, supervised by multi-task joint loss functions. With multi-task collaboration, DMTN effectively recovers the illumination in shadow areas while ensuring the fidelity of non-shadow areas. Experimental results show that DMTN competes favorably with state-of-the-art multi-branch/multi-stage shadow removal methods, while maintaining the simplicity of single-stage methods. We have released our code to encourage future exploration in powerful feature representation for shadow removalhttps://github.com/nachifur/DMTN Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Liangqiong Qu, Yandong Tang |
IEEE Trans. Multim. | 3 |
| 2022 | SemanticGAN: Facial Image Editing with Semantic to Realize Consistency
Xin Luan, Huijie Fan, Yandong Tang |
PRCV (3) | 3 |
| 2022 | Towards collaborative appearance and semantic adaptation for medical image segmentation
Qiang Wang 0015, Yingkui Du, Huijie Fan |
Neurocomputing | 3 |
| 2022 | APAN: Across-Scale Progressive Attention Network for Single Image DerainingabstractRecent single image deraining works have achieved significant improvement using convolutional neural networks. However, the rain streaks in the rain image share similar patterns with its multi-scale versions, which are not fully exploited in recent works. In this paper, we propose anAcross-scaleProgressiveAttentionNetwork (i.e.,APAN) to explore the multi-scale collaborative representation for single image deraining. Specifically, we represent each rainy image via a multi-scale module. An across-scale attention module is then used to capture long-range feature correspondences from multi-scale features, which can model the rain streaks at an enlarging feature dimension. Afterwards, we construct a pyramid structure and further predict the rain streak progressively, which also guides the across-scale attention module to refine the feature representation from coarse to fine. The proposed model exploits self-similarity of features via an across-scale attention between different scales, which can well model the rain streak with long-range information. Experiments on several datasets show that our model achieves significant improvement compared with most state-of-the-art deraining models. Qiang Wang 0015, Gan Sun, Huijie Fan, Yandong Tang |
IEEE Signal Process. Lett. | 3 |
| 2022 | Dual Aligned Siamese Dense Regression TrackerabstractAnchor or anchor-free based Siamese trackers have achieved the astonishing advancement. However, their parallel regression and classification branches lack the tracked target information link and interaction, and the corresponding independent optimization maybe lead to task-misalignment, such as the reliable classification prediction with imprecisely localization and vice versa. To address this problem, we develop a general Siamese dense regression tracker (SDRT) with both task and feature alignments. It consists of two cooperative and mutual-guidance core branches: dense local regression with RepPoint representation, the global and local multi-classifier fusion with aligned features. They complement and boost each other to constrain the results with well-localized followed to also be well-classified. Specifically, a dense local regression with RepPoint representation, directly estimates and averages multiple dense local bounding box offsets for accurate localization. And then, the refined bounding boxes can be used to learn the global and local affine alignment features for reliable multi-classifier fusion. The classified scores in turn guide the assigned positive bounding boxes for the regression task. The mutual guidance operations can bridge the connection between classification and regression substantially, since the assigned labels of one task depend on the prediction quality of the other task. The proposed tracking module is general, and it can boost both the anchor or anchor-free based Siamese trackers to some extent. The extensive tracking comparisons on six tracking benchmarks verify its favorable and competitive performance over states-of-the-arts tracking modules. Baojie Fan, Hui Zhang 0023, Yang Cong, Yandong Tang, Huijie Fan, Jiandong Tian |
IEEE Trans. Image Process. | 5 |
| 2021 | CAB-Net: Channel Attention Block Network for Pathological Image Cell Nucleus Segmentation
Huijie Fan |
ICIG (2) | 2 |
| 2021 | Temporal pyramid attention-based spatiotemporal fusion model for Parkinson's disease diagnosis from gait dataabstractAbstract Parkinson's disease (PD) is currently an ongoing challenge in daily clinical medicine. To reduce diagnosis time and arduousness and even assess PD levels, a temporal pyramid attention‐based spatiotemporal (PAST) fusion model for diagnosis of PD is produced by using gait data from ground reaction forces. This model is innovative in two aspects. First, by using the temporal pyramid attention module, multiscale temporal attention is obtained from raw sequences. Second, 1D convolutional neural network and bidirectional long short‐term memory layers are used together to learn spatial fusion features from multiple channels in the spatial domain to obtain multichannel, multiscale fusion features. Experiments are performed on the PhysioBank data set, and the results show that the proposed PAST model outperforms other state‐of‐the‐art methods on classification results. This model can assist in the diagnosis and treatment of PD by using gait data. Xiaomin Pei, Huijie Fan, Yandong Tang |
IET Signal Process. | 2 |
| 2021 | Multi-Scale Context-Guided Deep Network for Automated Lesion Segmentation With Endoscopy Images of Gastrointestinal TractabstractAccurate lesion segmentation based on endoscopy images is a fundamental task for the automated diagnosis of gastrointestinal tract (GI Tract) diseases. Previous studies usually use hand-crafted features for representing endoscopy images, while feature definition and lesion segmentation are treated as two standalone tasks. Due to the possible heterogeneity between features and segmentation models, these methods often result in sub-optimal performance. Several fully convolutional networks have been recently developed to jointly perform feature learning and model training for GI Tract disease diagnosis. However, they generally ignore local spatial details of endoscopy images, as down-sampling operations (e.g., pooling and convolutional striding) may result in irreversible loss of image spatial information. To this end, we propose a multi-scale context-guided deep network (MCNet) for end-to-end lesion segmentation of endoscopy images in GI Tract, where both global and local contexts are captured as guidance for model training. Specifically, one global subnetwork is designed to extract the global structure and high-level semantic context of each input image. Then we further design two cascaded local subnetworks based on output feature maps of the global subnetwork, aiming to capture both local appearance information and relatively high-level semantic information in a multi-scale manner. Those feature maps learned by three subnetworks are further fused for the subsequent task of lesion segmentation. We have evaluated the proposed MCNet on 1,310 endoscopy images from the public EndoVis-Ab and CVC-ClinicDB datasets for abnormal segmentation and polyp segmentation, respectively. Experimental results demonstrate that MCNet achieves [Formula: see text] and [Formula: see text] mean intersection over union (mIoU) on two datasets, respectively, outperforming several state-of-the-art approaches in automated lesion segmentation with endoscopy images of GI Tract. Shuai Wang 0003, Yang Cong, Hancan Zhu, Xianyi Chen, Liangqiong Qu, Huijie Fan, Qiang Zhang 0008, Mingxia Liu 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Recurrent Generative Adversarial Network for Face CompletionabstractMost recently-proposed face completion algorithms use high-level features extracted from convolutional neural networks (CNNs) to recover semantic texture content. Although the completed face is natural-looking, the synthesized content still lacks lots of high-frequency details, since the high-level features cannot supply sufficient spatial information for details recovery. To tackle this limitation, in this paper, we propose aRecurrentGenerativeAdversarialNetwork (RGAN) for face completion. Unlike previous algorithms, RGAN can take full advantage of multi-level features, and further provide advanced representations from multiple perspectives, which can well restore spatial information and details in face completion. Specifically, our RGAN model is composed of a CompletionNet and a DisctiminationNet, where the CompletionNet consists of two deep CNNs and a recurrent neural network (RNN). The first deep CNN is presented to learn the internal regulations of a masked image and represent it with multi-level features. The RNN model then exploits the relationships among the multi-level features and transfers these features in another domain, which can be used to complete the face image. Benefiting from bidirectional short links, another CNN is used to fuse multi-level features transferred from RNN and reconstruct the face image in different scales. Meanwhile, two context discrimination networks in the DisctiminationNet are adopted to ensure the completed image consistency globally and locally. Experimental results on benchmark datasets demonstrate qualitatively and quantitatively that our model performs better than the state-of-the-art face completion models, and simultaneously generates realistic image content and high-frequency details. The code will be released available soon. Qiang Wang 0015, Huijie Fan, Gan Sun, Weihong Ren, Yandong Tang |
IEEE Trans. Multim. | 2 |
| 2020 | Heart sound classification based on improved MFCC features and convolutional recurrent neural networks
Muqing Deng, Tingting Meng, Jiuwen Cao, Shimin Wang, Jing Zhang 0037, Huijie Fan |
Neural Networks | 6 |
| 2019 | Novel event analysis for human-machine collaborative underwater exploration
Yang Cong, Baojie Fan, Dongdong Hou, Huijie Fan, Kaizhou Liu, Jiebo Luo 0001 |
Pattern Recognit. | 4 |
| 2019 | Laplacian pyramid adversarial network for face completion
Qiang Wang 0015, Huijie Fan, Gan Sun, Yang Cong, Yandong Tang |
Pattern Recognit. | 2 |
| 2019 | Deeply Supervised Face Completion With Multi-Context Generative Adversarial NetworkabstractRecent face completion works have achieved significant improvement using generative adversarial networks (GANs). There are still two important issues in this challenging task: first, semantic understanding; and second, high-frequency details prediction. In this letter, we propose a unified model by introducing multi-context structures within GANs. Our model, named multi-context generative adversarial networks (MCGAN), automatically learns the hierarchical appearances of a corrupted image and predicted the missing regions from different perspectives. In this model, semantic understanding and high-frequency details are both taken into account and modeled with two parallel networks, respectively. While one learns the semantic understanding of the input face image at a high level, the other extracts low-level features for high-frequency details prediction. Our MCGAN takes full advantage of multi-scale features learned from two complementary networks and generates semantically new pixels for the missing region with fine details. Extensive quantitative and qualitative experiments on benchmark datasets show that the proposed model outperforms several state-of-the-art models. Qiang Wang 0015, Huijie Fan, Yandong Tang |
IEEE Signal Process. Lett. | 2 |
| 2018 | Evaluation of shadow featuresabstractShadow features such as colour ratio, texture, and chromaticity have proved to be quite effective in shadow detection. Many shadow detection methods have been proposed on the basis of different features. However, previous works for shadow detection mainly focus on designing an effective classifier for existing shadow features, but pay less attention on the analysis of shadow features themselves. The majority of studies simply report the final shadow detection results rather than make an evaluation on each feature. Readers often do not know which features are more effective or whether these shadow features are complementary. The following problems are still unsolved: the robustness of each feature, which feature plays the most important role in a detection method, and what is the best performance that current features can reach. The purpose of this study is to answer these questions, and the authors hope that this study can offer guidance for future shadow detection algorithms via the evaluation of frequently used shadow features. Several useful and interesting conclusions are obtained after conducting extensive comparison experiments on a large dataset. Liangqiong Qu, Jiandong Tian, Huijie Fan, Yandong Tang |
IET Comput. Vis. | 3 |
| 2017 | Large receptive field convolutional neural network for image super-resolutionabstractThis paper presents a new approach to Single Image Super Resolution (SISR), based upon Convolutional Neural Network (CNN). Although the SISR is ill-posed which can be seen as finding a non-linear mapping from a low to high-dimensional space. Deep learning techniques have been successfully applied in many areas of computer vision, including low-level image restoration and non-linear mapping problems. We consider the single image Super-Resolution (SR) problem as convolution operators and develop a CNN to capture the characteristics of Low-Resolution (LR) input image. We find that increasing the receptive field shows the improvement in accuracy. Our solution is to establish the connection between traditional optimization-based schemes and neural network architectures. In the paper a novel, separable structure is introduced as a reliable support for robust convolution against artifacts. Our proposed method performs better than existing methods in terms of accuracy and visual improvements in our results are easily noticeable. Qiang Wang 0015, Huijie Fan, Yang Cong, Yandong Tang |
ICIP | 2 |
| 2017 | Deep learning of directional truncated signed distance function for robust 3D object recognitionabstractIn this paper, we develop a novel 3D object recognition algorithm to perform detection and pose estimation jointly. We focus on analyzing the advantages of the 3D point cloud relative to the RGB-D image and try to eliminate the unpredictability of output values that inevitably occurs in regression tasks. To achieve this, we first adopt the Truncated Signed Distance Function (TSDF) to encode the point cloud and extract low compact discriminative feature via unsupervised deep learning network. This approach can not only eliminate the dense scale sampling for offline model training but also reduce the distortion by mapping the 3D shape to the 2D plane and overcome the dependence on color cues. Then, we train a Hough forests to achieve multi-object detection and 6-DoF pose estimation simultaneously. In addition, we propose a robust multilevel verification strategy that effectively reduces the unpredictability of output values which occurs in the hough regression module. Experiments on public datasets demonstrate that our approach provides effective results comparable to the state-of-the-arts. Hongsen Liu, Yang Cong, Shuai Wang 0003, Huijie Fan, Dongying Tian, Yandong Tang |
IROS | 4 |
| 2017 | Multi-Class Latent Concept Pooling for Computer-Aided Endoscopy DiagnosisabstractSuccessful computer-aided diagnosis systems typically rely on training datasets containing sufficient and richly annotated images. However, detailed image annotation is often time consuming and subjective, especially for medical images, which becomes the bottleneck for the collection of large datasets and then building computer-aided diagnosis systems. In this article, we design a novel computer-aided endoscopy diagnosis system to deal with the multi-classification problem of electronic endoscopy medical records (EEMRs) containing sets of frames, while labels of EEMRs can be mined from the corresponding text records using an automatic text-matching strategy without human special labeling. With unambiguous EEMR labels and ambiguous frame labels, we propose a simple but effective pooling scheme called Multi-class Latent Concept Pooling, which learns a codebook from EEMRs with different classes step by step and encodes EEMRs based on a soft weighting strategy. In our method, a computer-aided diagnosis system can be extended to new unseen classes with ease and applied to the standard single-instance classification problem even though detailed annotated images are unavailable. In order to validate our system, we collect 1,889 EEMRs with more than 59K frames and successfully mine labels for 348 of them. The experimental results show that our proposed system significantly outperforms the state-of-the-art methods. Moreover, we apply the learned latent concept codebook to detect the abnormalities in endoscopy images and compare it with a supervised learning classifier, and the evaluation shows that our codebook learning method can effectively extract the true prototypes related to different classes from the ambiguous data. Shuai Wang 0003, Yang Cong, Huijie Fan, Baojie Fan, Lianqing Liu, Yunsheng Yang, Yandong Tang, Huaici Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2015 | Computer aided endoscope diagnosis via weakly labeled data miningabstractIn comparison to most computer aided endoscope diagnosis methods using pixel-wise groundtruth by physicians manually, it is easy to get lots of endoscope images with corresponding diagnostic reports. In this paper, we intend to mine pixel-wise label information from these reports with weak frame-level labels automatically. To achieve this, we formulate our computer aided diagnosis problem as a Multiple Instance Learning (MIL) issue, where we represent each image as superpixels. Each image and each superpixel is cast as bag and instance, respectively. We then evaluate and select the most positive instances from positive bags automatically which helps us transform the frame-level classification problem into a standard supervised learning problem. In the experiment, we build a new gastroscopic image dataset with more than 3000 weakly labeled images, and ours outperforms the state-of-the-art methods, which verifies the effectiveness of our model. Shuai Wang 0003, Yang Cong, Huijie Fan, Yunsheng Yang, Yandong Tang, Huaici Zhao |
ICIP | 3 |
| 2015 | Object detection based on scale-invariant partial shape matching
Huijie Fan, Yang Cong, Yandong Tang |
Mach. Vis. Appl. | 1 |
| 2012 | Self-closed partial shape descriptor for shape retrievalabstractWe propose a discriminative partial-based algorithm for shape recognition and retrieval. A key distinction of our approach is that we use pairwise geometric relations between contour fragments containing important and salient shape information to establish self-closed partial descriptor (SCPD), it can capture similar local parts in matching shape contours and meanwhile overcome part occlusion and distortion. We establish local coordinate system for each fragment to make sure SCPD is invariant to RST (rotation, scaling, and translation) transformation. In the matching stage, a scale approximation scheme is used to get rid of invalid matches. We experiment on MPEG7 shape database, and experimental results illustrate that our algorithm performs well on shape retrieval. Huijie Fan, Yang Cong, Yandong Tang |
ICIP | 1 |
| 2010 | Skew detection in document images based on rectangular active contour
Huijie Fan, Yandong Tang |
Int. J. Document Anal. Recognit. | 1 |