VLDB 2026 Research / reviewers in the wild / expert
Jie Chen 0001
dblp:92/6289-1
· DBLP profile ↗
134ranked-venue papers
14as first author
93since 2021 · last 2026
0000-0002-9765-4523ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 94 · 9 first-author · 67 since 2021Artificial intelligence and machine learning · 82 · 11 first-author · 57 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 11 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProAR: Probabilistic Autoregressive Modeling for Molecular DynamicsabstractUnderstanding the structural dynamics of biomolecules is crucial for uncovering biological functions. As molecular dynamics (MD) simulation data becomes more available, deep generative models have been developed to synthesize realistic MD trajectories. However, existing methods produce fixed-length trajectories by jointly denoising high-dimensional spatiotemporal representations, which conflicts with MD’s frame-by-frame integration process and fails to capture time-dependent conformational diversity. Inspired by MD's sequential nature, we introduce a new probabilistic autoregressive (ProAR) framework for trajectory generation. ProAR uses a dual-network system that models each frame as a multivariate Gaussian distribution and employs an anti-drifting sampling strategy to reduce cumulative errors. This approach captures conformational uncertainty and time-coupled structural changes while allowing flexible generation of trajectories of arbitrary length. Experiments on ATLAS, a large-scale protein MD dataset, demonstrate that for long trajectory generation, our model achieves a 7.5% reduction in reconstruction RMSE and an average 25.8% improvement in conformation change accuracy compared to previous state-of-the-art methods. For conformation sampling task, it performs comparably to specialized time-independent models, providing a flexible and dependable alternative to standard MD simulations. Kaiwen Cheng, Yutian Liu 0004, Zhiwei Nie, Mujie Lin, Yanzhen Hou, Yiheng Tao, Jie Chen 0001, Youdong Mao, Yonghong Tian 0001 |
AAAI | 8 |
| 2026 | WaveFormer: Frequency-Time Decoupled Vision Modeling with Wave EquationabstractVision modeling has advanced rapidly with Transformers, whose attention mechanisms capture visual dependencies but lack a principled account of how semantic information propagates spatially. We revisit this problem from a wave-based perspective: feature maps are treated as spatial signals whose evolution over an internal propagation time (aligned with network depth) is governed by an underdamped wave equation. In this formulation, spatial frequency—from low-frequency global layout to high-frequency edges and textures—is modeled explicitly, and its interaction with propagation time is controlled rather than implicitly fixed. We derive a closed-form, frequency–time decoupled solution and implement it as the Wave Propagation Operator (WPO), a lightweight module that models global interactions in O(NlogN) time—far lower than attention. Building on WPO, we propose a family of WaveFormer models as drop-in replacements for standard ViTs and CNNs, achieving competitive accuracy across image classification, object detection, and semantic segmentation, while delivering up to 1.6× higher throughput and 30% fewer FLOPs than attention-based alternatives. Furthermore, our results demonstrate that wave propagation introduces a complementary modeling bias to heat-based methods, effectively capturing both global coherence and high-frequency details essential for rich visual semantics. Zishan Shu, Juntong Wu, Xudong Liu 0001, Hongyu Zhang 0002, Chang Liu 0030, Youdong Mao, Jie Chen 0001 |
AAAI | 8 |
| 2026 | BiHiTo: Biomolecular Hierarchy-inspired TokenizationabstractThree-dimensional atomic arrangements of biomolecules are key to demystifying biological functions. The rapid expansion of accessible structural data, driven by advances in AI for science, highlights the critical challenge of efficiently modeling large-scale biomolecular structures, which are high-dimensional systems shaped by biological assembly principles. To address this, we introduce BiHiTo, a multi-level Biomolecular Hierarchy-inspired Tokenizer that intrinsically mimics natural biological assembly hierarchies. Specifically, we design a multi-codebook quantizer that mirrors the natural hierarchy of biomolecular structure, enabling simultaneous capture of representations spanning atomic motifs to global conformational variations. This hierarchical alignment markedly improves the biological interpretability and reconstruction fidelity of biomolecular structure.Extensive experiments demonstrate that BiHiTo delivers state-of-the-art performance and robust generalization across molecular dynamics trajectories and macromolecular complexes, facilitating advances in structure generation and dynamic conformation exploration. In the reconstruction of the CASP14 and OOD test set FastFolding protein multi-conformation data, our method achieves a 17% and 51% reduction in RMSD compared to Bio2Token, respectively. Ruochong Zheng, Yutian Liu 0004, Yian Zhao, Zhiwei Nie, Xuehan Hou, Youdong Mao, Jie Chen 0001 |
AAAI | 9 |
| 2026 | Knowing Where to Focus: Attention-Guided Alignment for Text-based Person Search
Pingyang Dai, Jie Chen 0001, Liujuan Cao, Rongrong Ji |
Int. J. Comput. Vis. | 4 |
| 2026 | Switch-UMamba: Dynamic scanning vision Mamba UNet for medical image segmentation
Ziyao Zhang 0003, Qiankun Ma, Tong Zhang 0017, Jie Chen 0001, Hairong Zheng, Wen Gao 0001 |
Medical Image Anal. | 4 |
| 2026 | Unified Granularity Controller for Interactive SegmentationabstractInteractive Segmentation (IS) segments specific objects or parts by deducing human intent from sparse input prompts. However, the sparse-to-dense mapping is ambiguous, making it challenging for users to obtain segmentations at the desired granularity and causing them to engage in trial-and-error cycles. Although existing multi-granularity IS models (e.g., SAM) alleviate the ambiguity of single-granularity methods by predicting multiple masks simultaneously, this approach has limited scalability and produces redundant results. To address this issue, we introduce a creative granularity-controllable IS paradigm that resolves ambiguity by enabling users to precisely control the segmentation granularity. Specifically, we propose a Unified Granularity Controller (UniGraCo) that supports multi-type optional granularity control signals to pursue unified control over diverse segmentation requirements, effectively overcoming the limitation of single-type control in adapting to different needs, thus boosting the system efficiency and practicality. To mitigate the excessive cost of annotating the multi-granularity masks and the corresponding granularity control signals for training UniGraCo, we construct an automated data engine capable of generating high-quality and granularity-abundant mask-granularity data pairs at low cost. To enable UniGraCo to learn unified granularity controllability in an efficient and stable manner, we further design a granularity-controllable learning strategy. This strategy leverages the generated data pairs to incrementally equip the pre-trained IS model with granularity controllability while preserving its segmentation capability. Extensive experiments on intricate scenarios at both instance and part level demonstrate that our UniGraCo has significant advantages over previous methods, highlighting its potential as a practical interactive tool. Yian Zhao, Kehan Li 0002, Pengchong Qiao, Chang Liu 0047, Rongrong Ji, Jie Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | InterTeach: A Novel Approach for Semi-Supervised Medical Image Segmentation Using Cooperative Teacher-Student NetworksabstractIn medical image segmentation, the reliance on extensive, high-quality labeled datasets poses a significant challenge, especially considering the associated costs and the requirement for specialized expertise. In response, the field has progressively embraced semi-supervised learning (SSL) methods that leverage both labeled and unlabeled data. Nonetheless, these methods frequently encounter issues related to inconsistent label quality and constrained generalizability of models. To surmount these obstacles, we present InterTeach, an innovative SSL framework that seamlessly integrates cross-supervision with the mean teacher model. This framework facilitates effective knowledge transfer and boosts model performance through the implementation of two unique teacher-student training configurations. Herein, knowledge is exchanged between models via their respective teacher counterparts, facilitating mutual learning and enhancement. This strategy diverges from traditional SSL approaches, which mainly depend on mutual learning between two models updated through gradient descent. Furthermore, the incorporation of Feature Divergence Loss (FDL) in InterTeach encourages the transfer of diverse and complementary knowledge between models, thereby enriching the overall learning dynamics. The evaluation results revealed that our method could approach or even match the performance of fully supervised learning methods on certain evaluation metrics. This finding further confirms the effectiveness and wide applicability of the IntraTeach method in handling multi-modal and multi-dimensional medical image segmentation tasks. Ziyao Zhang 0003, Qiankun Ma, Jie Chen 0001, Hairong Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Practical Lossless Volumetric Medical Image Compression via Tri-Plane Context Tree LearningabstractLossless compression of volumetric medical images is of paramount importance for clinical and research applications where data fidelity is essential. Traditional compression methods are often limited in efficiency due to rigid, handcrafted models. Conversely, deep neural network (DNN)-based compression methods, while effective, demand substantial computational resources, hindering deployment in resource-constrained settings. To address these challenges, we propose a novel tri-plane context tree (TCT)-based method for lossless volumetric medical image compression that delivers high performance without relying on DNNs or external training data. To exploit intra-slice and inter-slice redundancies, we introduce a compact tri-plane context representation that decomposes complex 3D context modeling into efficient 2D modeling on three orthogonal planes. By integrating this representation with a context tree framework, we develop an input-specific TCT model employing an adaptive binary tree structure. At each tree node, the model dynamically selects from a suite of tri-plane based predictors and contextual feature extractors, enabling data-adaptive context modeling tailored to local structural characteristics. Instead of offline training, we sample a subset of the input volume to learn the TCT model by optimizing the minimum description length (MDL) through iterative construction and pruning. With the learned TCT model, each pixel retrieves its corresponding context, computes the prediction residual using the predictor dictated by the context, and performs entropy encoding based on the associated histograms. Experimental results demonstrate that the proposed method achieves compression performance on par with recent DNN-based methods on multiple datasets, while maintaining low computational cost and fast coding speeds, making it highly applicable in practice. Yuanchao Bai, Kai Wang 0070, Yuanbo Du, Jie Chen 0001, Teng Fang, Xianming Liu 0005, Wen Gao 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Disentangled Concept Matching for Text-video Retrieval through Perception ImitationabstractText-video retrieval plays a pivotal role in cross-modal tasks, aiming to match textual descriptions with corresponding video content accurately. Existing methods often employ fine-grained feature matching to improve retrieval accuracy, but such approaches consume extensive computational resources. Conversely, coarse-grained feature matching between entire sentences and videos offers computational efficiency but may overlook the heterogeneous semantic concepts embedded within the data. To overcome these challenges, we develop the Disentangled Concept Matching (DCM) framework, designed as an imitation of human semantic perception processes. The framework utilizes Disentangled Representation Learning (DRL) to divide coarse-grained features into distinct semantic concepts represented as latent factors, effectively generating finer-grained features while reducing computational demands. To improve the accuracy of retrieval, we first propose the Composed Spatial-temporal Module (CSTM) to optimize the quality of multimodal feature extraction. Utilizing a branch-structured temporal modeling approach, CSTM effectively enhances the DCM model’s comprehension of video content and temporal information, leading to the extraction of refined video features. Second, building on the optimized features, we propose the Adaptive Pooling Module (APM) to measure the confidence level of each latent factor matching during the process of decoupling concepts. APM enhances the fidelity of text and video concepts, thereby further ensuring the accuracy of matching after decoupling. With CSTM and APM, DCM accurately matches latent factors in lower dimensions, achieving significant improvements in computing efficiency and retrieval performance. Our experimental evaluations across standard datasets, namely MSR-VTT, LSMDC, MSVD, ActivityNet, and DiDeMo, demonstrate that the DCM framework achieves state-of-the-art performance, with Recall@1 scores of 48.7%, 25.6%, 48.4%, 45.0%, and 48.6%, respectively. Compared to our previous model, the DCM framework shows improvements of 2.54%, 0.08%, 2.11%, 6.89%, and 6.35%, respectively. Peng Jin 0001, Chunyu Zou, Ziyao Zhang 0003, Jie Chen 0001, Wen Gao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Aligning Instance Brownian Bridge with Texts for Open-Vocabulary Video Instance SegmentationabstractTemporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video data, previous methods leverage the image-text pretraining model for recognizing object instances by separately aligning each frame with class texts. As a result, the separation breaks the instance movement context of videos and requires a lot of inference overhead. To tackle these issues, we propose BridgeText Alignment (BTA) to link frame-level instance representations as a Brownian Bridge. On one hand, we can calculate the global descriptor of a Brownian bridge for capturing instance dynamics, which enables extra considering temporal information rather than only static information of each frame for aligning with texts. On the other hand, according to the goal-conditioned property of the Brownian bridge, we can estimate the middle frame features via the start and the end frame features so the global feature calculation of a Brownian bridge only needs to infer a few frames, which largely reduces inference overhead. We term our overall pipeline as BriVIS. Following the training settings of previous works, BriVIS surpasses the SOTA (OV2Seg) by a clear margin. For example, on the challenging large-vocabulary datasets (BURST, LVVIS), BriVIS achieves 5.7 and 20.9 mAP, which exhibits +2.2∼+6.7 mAP improvement compared to OV2Seg. Furthermore, after training via BTA, using only the head and the tail frames for alignment improves the speed by 32% (2.77 → 1.88 s/iter) while just decreasing the performance by 0.2 mAP (21.1 → 20.9 mAP). Zesen Cheng, Kehan Li 0002, Hao Li 0073, Peng Jin 0001, Xiawu Zheng, Jie Chen 0001 |
AAAI | 7 |
| 2025 | DigitalLLaVA: Incorporating Digital Cognition Capability for Physical World Comprehension in Multimodal LLMsabstractMultimodal Large Language Models (MLLMs) have shown remarkable cognitive capabilities in various cross-modal tasks.However, existing MLLMs struggle with tasks that require physical digital cognition, such as accurately reading an electric meter or pressure gauge. This limitation significantly reduces their effectiveness in practical applications like industrial monitoring and home energy management, where digital sensors are not feasible. For humans, physical digits are artificially defined quantities presented on specific carriers, which require training to recognize. As existing MLLMs are only pre-trained in the manner of object recognition, they fail to comprehend the relationship between digital carriers and their reading. To this end, referring to human behavior, we propose a novel DigitalLLaVA method to explicitly inject digital cognitive abilities into MLLMs in a two-step manner. In the first step, to improve the MLLM's understanding of physical digit carriers, we propose a digit carrier mapping method. This step utilizes object-level text-image pairs to enhance the model's comprehension of objects containing physical digits. For the second step, unlike previous methods that rely on sequential digital prediction or digit regression, we propose a 32 bit floating point simulation approach that treats digit prediction as a whole. Using digit-level text-image pairs, we train three float heads to predict 32-bit floating-point numbers using 0/1 binary classification. This step significantly reduces the search space, making the prediction process more robust and straightforward. Being simple but effective, our method can identify very precise metrics (i.e., accurate to ±0.001) and provide floating-point results, showing its applicability in digital carrier domains. Pengxu Wei, Pengchong Qiao, Chang Liu 0030, Jie Chen 0001 |
AAAI | 5 |
| 2025 | Adversarial Diffusion Compression for Real-World Image Super-ResolutionabstractReal-world image super-resolution (Real-ISR) aims to reconstruct high-resolution images from low-resolution inputs degraded by complex, unknown processes. While many Stable Diffusion (SD)-based Real-ISR methods have achieved remarkable success, their slow, multi-step inference hinders practical deployment. Recent SD-based one-step networks like OSEDiff and S3Diff alleviate this issue but still incur high computational costs due to their reliance on large pretrained SD models. This paper proposes a novel Real-ISR method, AdcSR, by distilling the one-step diffusion network OSEDiff into a streamlined diffusion-GAN model under our Adversarial Diffusion Compression (ADC) framework. We meticulously examine the modules of OSEDiff, categorizing them into two types: (1) Removable (VAE encoder, prompt extractor, text encoder, etc.) and (2) Prunable (denoising UNet and VAE decoder). Since direct removal and pruning can degrade the model’s generation capability, we pretrain our pruned VAE decoder to restore its ability to decode images and employ adversarial distillation to compensate for performance loss. This ADC-based diffusion-GAN hybrid design effectively reduces complexity by 73% in inference time, 78% in computation, and 74% in parameters, while preserving the model’s generation capability. Experiments manifest that our proposed AdcSR achieves competitive recovery quality on both synthetic and real-world datasets, offering up to 9.3× speedup over previous one-step diffusion-based methods. Code and models are available at https://github.com/Guaishou74851/AdcSR. Bin Chen 0006, Gehui Li, Rongyuan Wu, Jie Chen 0001, Jian Zhang 0018, Lei Zhang 0001 |
CVPR | 5 |
| 2025 | DASH: 4D Hash Encoding with Self-Supervised Decomposition for Real-Time Dynamic Scene Rendering
Jie Chen 0001, Zhangchi Hu, Peixi Wu, Huyue Zhu, Hebei Li, Xiaoyan Sun 0001 |
ICCV | 1 |
| 2025 | Temporal-Aware Query Routing for Real-Time Video Instance Segmentation
Zesen Cheng, Kehan Li 0002, Yian Zhao, Jie Chen 0001 |
ICCV | 6 |
| 2025 | Efficient Spiking Point Mamba for Point Cloud Analysis
Peixi Wu, Bosong Chai, Menghua Zheng, Zhangchi Hu, Jie Chen 0001, Zheyu Zhang 0002, Hebei Li, Xiaoyan Sun 0001 |
ICCV | 6 |
| 2025 | Tune-Your-Style: Intensity-Tunable 3D Style Transfer with Gaussian Splatting
Yian Zhao, Rushi Ye, Ruochong Zheng, Zesen Cheng, Jiashu Yang, Pengchong Qiao, Jie Chen 0001 |
ICCV | 9 |
| 2025 | MTPNet: Multi-Grained Target Perception for Unified Activity Cliff PredictionabstractActivity cliff prediction is a critical task in drug discovery and material design. Existing computational methods are limited to handling single binding targets, which restricts the applicability of these prediction models. In this paper, we present the Multi-Grained Target Perception network (MTPNet) to incorporate the prior knowledge of interactions between the molecules and their target proteins. Specifically, MTPNet is a unified framework for activity cliff prediction, which consists of two components: Macro-level Target Semantic (MTS) guidance and Micro-level Pocket Semantic (MPS) guidance. By this way, MTPNet dynamically optimizes molecular representations through multi-grained protein semantic conditions. To our knowledge, it is the first time to employ the receptor proteins as guiding information to effectively capture critical interaction details. Extensive experiments on 30 representative activity cliff datasets demonstrate that MTPNet significantly outperforms previous approaches, achieving an average RMSE improvement of 18.95% on top of several mainstream GNN architectures. Overall, MTPNet internalizes interaction patterns through conditional deep learning to achieve unified predictions of activity cliffs, helping to accelerate compound optimization and design. Codes are available at: https://github.com/ZishanShu/MTPNet. Zishan Shu, Yufan Deng, Hongyu Zhang 0002, Zhiwei Nie, Jie Chen 0001 |
IJCAI | 5 |
| 2025 | Dome-DETR: DETR with Density-Oriented Feature-Query Manipulation for Efficient Tiny Object DetectionabstractTiny object detection plays a vital role in drone surveillance, remote sensing, and autonomous systems, enabling the identification of small targets across vast landscapes. However, existing methods suffer from inefficient feature leverage and high computational costs due to redundant feature processing and rigid query allocation. To address these challenges, we propose Dome-DETR, a novel framework with Density-Oriented Feature-Query Manipulation for Efficient Tiny Object Detection. To reduce feature redundancies, we introduce a lightweight Density-Focal Extractor (DeFE) to produce clustered compact foreground masks. Leveraging these masks, we incorporate Masked Window Attention Sparsification (MWAS) to focus computational resources on the most informative regions via sparse attention. Besides, we propose Progressive Adaptive Query Initialization (PAQI), which adaptively modulates query density across spatial areas for better query allocation. Extensive experiments demonstrate that Dome-DETR achieves state-of-the-art performance (+3.3 AP on AI-TOD-V2 and +2.5 AP on VisDrone) while maintaining low computational complexity and a compact model size. Code is available at https://github.com/RicePasteM/Dome-DETR. Zhangchi Hu, Peixi Wu, Jie Chen 0001, Huyue Zhu, Yansong Peng, Hebei Li, Xiaoyan Sun 0001 |
ACM Multimedia | 3 |
| 2025 | Generative prediction of real-world prevalent SARS-CoV-2 mutation with in silico virus evolutionabstractPredicting the mutation prevalence trends of emerging viruses in the real world is an efficient means to update vaccines or drugs in advance. It is crucial to develop a computational method for the prediction of real-world prevalent SARS-CoV-2 mutations considering the impact of multiple selective pressures within and between hosts. Here, a deep-learning generative framework for real-world prevalent SARS-CoV-2 mutation prediction, named ViralForesight, is developed on top of protein language models and in silico virus evolution. Through the paradigm of host-to-herd in silico virus evolution, ViralForesight reproduced previous real-world prevalent SARS-CoV-2 mutations for multiple lineages with superior performance. More importantly, ViralForesight correctly predicted the future prevalent mutations that dominated the COVID-19 pandemic in the real world more than half a year in advance with in vitro experimental validation. Overall, ViralForesight demonstrates a proactive approach to the prevention of emerging viral infections, accelerating the process of discovering future prevalent mutations with the power of generative deep learning. Xudong Liu 0001, Zhiwei Nie, Haorui Si, Xurui Shen, Yutian Liu 0004, Xiansong Huang, Tianyi Dong, Zhixiang Ren, Jie Chen 0001 |
Briefings Bioinform. | 11 |
| 2025 | Predicting protein stability changes upon mutations with dual-view ensemble learning from single sequenceabstractPredicting the protein stability changes upon mutations is one of the effective ways to improve the efficiency of protein engineering. Here, we propose a dual-view ensemble learning-based framework, DVE-stability, for mutation-induced protein stability change prediction from single sequence. DVE-stability integrates the global and local dependencies of mutations to capture the intramolecular interactions from two views through ensemble learning, in which a structural microenvironment simulation module is designed to indirectly introduce the information of structural microenvironment at the sequence level. DVE-stability achieved state-of-the-art prediction performance on seven single-point mutation benchmark datasets, and comprehensively surpassed other methods on five of them. Furthermore, DVE-stability outperformed other methods comprehensively through zero-shot inference on multiple-point mutation prediction task, demonstrating superior model generalizability to capture the epistasis of multiple-point mutations. More importantly, DVE-stability exhibited superior generalization performance in predicting rare beneficial mutations that are crucial for practical protein directed evolution scenarios. In addition, DVE-stability identified important intramolecular interactions via attention scores, demonstrating interpretable. Overall, DVE-stability provides a flexible and efficient tool for mutation-induced protein stability change prediction in an interpretable ensemble learning manner. Zhiwei Nie, Yutian Liu 0004, Xiansong Huang, Peng Yang 0001, Zigang Li, Jie Fu 0001, Zhixiang Ren, Jie Chen 0001 |
Briefings Bioinform. | 13 |
| 2025 | Adaptive Fuzzy Positive Learning for Annotation-Scarce Semantic Segmentation
Pengchong Qiao, Yu Wang 0027, Chang Liu 0030, Baigui Sun, Zhennan Wang 0001, Xiawu Zheng, Rongrong Ji, Jie Chen 0001 |
Int. J. Comput. Vis. | 9 |
| 2025 | An Information Theory-Inspired Strategy for Automated Network Pruning
Xiawu Zheng, Yuexiao Ma, Teng Xi, Errui Ding, Jie Chen 0001, Yonghong Tian 0001, Rongrong Ji |
Int. J. Comput. Vis. | 7 |
| 2025 | SAFA: Lifelong Person Re-Identification learning by statistics-aware feature alignment
Qiankun Gao, Mengxi Jia, Jie Chen 0001, Jian Zhang 0018 |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | Hardware-friendly rate estimation algorithm and architecture design for AVS3
Yunyao Yan, Guoqing Xiang, Jie Chen 0001, Xiaofeng Huang, Peng Zhang 0007, Huizhu Jia |
Multim. Tools Appl. | 3 |
| 2025 | Invertible Diffusion Models for Compressed SensingabstractWhile deep neural networks (NNs) significantly advance image compressed sensing (CS) by improving reconstruction quality, the necessity of training current CS NNs from scratch constrains their effectiveness and hampers rapid deployment. Although recent methods utilize pre-trained diffusion models for image reconstruction, they struggle with slow inference and restricted adaptability to CS. To tackle these challenges, this paper proposes Invertible Diffusion Models (IDM), a novel efficient, end-to-end diffusion-based CS method. IDM repurposes a large-scale diffusion sampling process as a reconstruction model, and fine-tunes it end-to-end to recover original images directly from CS measurements, moving beyond the traditional paradigm of one-step noise estimation learning. To enable such memory-intensive end-to-end fine-tuning, we propose a novel two-level invertible design to transform both 1) multi-step sampling process and 2) noise estimation U-Net in each step into invertible networks. As a result, most intermediate features are cleared during training to reduce up to 93.8% GPU memory. In addition, we develop a set of lightweight modules to inject measurements into noise estimator to further facilitate reconstruction. Experiments demonstrate that IDM outperforms existing state-of-the-art CS networks by up to 2.64 dB in PSNR. Compared to the recent diffusion-based approach DDNM, our IDM achieves up to 10.09 dB PSNR gain and 14.54 times faster inference. Bin Chen 0006, Zhenyu Zhang 0005, Chen Zhao 0002, Jiwen Yu, Shijie Zhao 0001, Jie Chen 0001, Jian Zhang 0018 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Hierarchical Banzhaf Interaction for General Video-Language Representation LearningabstractMultimodal representation learning, with contrastive learning, plays an important role in the artificial intelligence domain. As an important subfield, video-language representation learning focuses on learning representations using global semantic interactions between pre-defined video-text pairs. However, to enhance and refine such coarse-grained global interactions, more detailed interactions are necessary for fine-grained multimodal learning. In this study, we introduce a creative approach that models video-text as game players using multivariate cooperative game theory to handle uncertainty during fine-grained semantic interactions with diverse granularity, flexible combination, and vague intensity. Specifically, we design the Hierarchical Banzhaf Interaction to simulate the finegrained correspondence between video clips and textual words from hierarchical perspectives. Furthermore, to mitigate the bias in calculations within Banzhaf Interaction, we propose reconstructing the representation through a fusion of single-modal and crossmodal components. This reconstructed representation ensures fine granularity comparable to that of the single-modal representation, while also preserving the adaptive encoding characteristics of cross-modal representation. Additionally, we extend our original structure into a flexible encoder-decoder framework, enabling the model to adapt to various downstream tasks. Extensive experiments on commonly used text-video retrieval, video-question answering, and video captioning benchmarks, with superior performance, validate the effectiveness and generalization of our method. The code is available at https://github.com/jpthu17/HBI. Peng Jin 0001, Hao Li 0073, Li Yuan 0007, Shuicheng Yan, Jie Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Oriented-Derivative Representation for Boundary-Aware Polyp SegmentationabstractThe diagnosis of colon polyps is important for the prevention of colorectal cancer. Polyp segmentation, however, is still a challenging problem given that recent medical computer-aided equipment suffers from situations of polyp variations in terms of size, color, texture, and poor illuminations in endoscopy videos. These obstacles hinder the prediction of polyp boundaries. Inspired by the observation that the values of pixels on the border region change more sharply than others, we propose the oriented-derivative (OD) representation to capture the relationship between pixels and the boundary region given distance and orientation. To adaptively use the proposed representation in arbitrary frameworks, we design plug-in modules to learn the representation and aggregate features to improve the accuracy of boundary predictions in the polyp segmentation task, which can be implemented in frameworks including the encoder-decoder and top-down architectures. Extensive experimental results show the improvement from the proposed oriented-derivative representation for the polyp segmentation task and the extendibility of our proposed modules in different architectures. Our methods achieved an improvement ranging from 0.3% to 2.5% (mDice) compared with the baseline on five publicly available datasets, includingKvasir, CVC-ClinicDB, EndoScene, CVC-ColonDB, andETIS. Mengjun Cheng, Xiawu Zheng, Rongrong Ji, Jie Chen 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Dual-Level Masked Semantic Inference for Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation pursues a holistic pixel-wise understanding of unseen images with limited annotation. To this end, existing methods focus on regularizing per-pixel prediction consistency within unlabeled data, while rarely modeling contextual relationships. But in fact, contextual semantics can provide valuable clues for scene understanding like inner-object continuity and spatial relationships' causality. Thus, in this paper, we propose a Dual-level Masked Semantics Inference (DMSI) that takes the initiative to explicitly learn contextual relationships via enforcing our model to infer the semantics of a pixel according to its surrounding contexts. This allows our model to exhaust accurate semantics by incorporating inter-pixel context clues, further leading to comprehensive segmentation. Specifically, DMSI comprises two main components. 1) Dual-level mask consistency regularization (DMCR) that learns the ability of semantics inference by aligning the predictions of masked views with the prediction of the complete view. The masked views here come from both the image level and feature level, where our model captures low-level attributes and high-level representations respectively. 2) AdaMask that provides a proper mask position and ratio for each image, guiding our model to focus on semantic-rich regions while providing balanced training between hard and easy samples. Through learning the ability of semantic inferring, DMSI remarkably enhances the interaction between pixels, further progressively intensifying the understanding of semantics. Extensive experiments under various settings on Cityscapes and Pascal VOC 2012 show that DMSI achieves new state-of-the-art performances. Furthermore, analysis indicates that our method has superiority in mining inter-pixel semantic relationships and improving robustness facing noise corruption. Qiankun Ma, Ziyao Zhang 0003, Pengchong Qiao, Yu Wang 0027, Rongrong Ji, Chang Liu 0047, Jie Chen 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Parallel Vertex Diffusion for Unified Visual GroundingabstractUnified visual grounding (UVG) capitalizes on a wealth of task-related knowledge across various grounding tasks via one-shot training, which curtails retraining costs and task-specific architecture design efforts. Vertex generation-based UVG methods achieve this versatility by unified modeling object box and contour prediction and provide a text-powered interface to vast related multi-modal tasks, e.g., visual question answering and captioning. However, these methods typically generate vertexes sequentially through autoregression, which is prone to be trapped in error accumulation and heavy computation, especially for high-dimension sequence generation in complex scenarios. In this paper, we develop Parallel Vertex Diffusion (PVD) based on the parallelizability of diffusion models to accurately and efficiently generate vertexes in a parallel and scalable manner. Since the coordinates fluctuate greatly, it typically encounters slow convergence when training diffusion models without geometry constraints. Therefore, we consummate our PVD by two critical components, i.e., center anchor mechanism and angle summation loss, which serve to normalize coordinates and adopt a differentiable geometry descriptor from the point-in-polygon problem of computational geometry to constrain the overall difference of prediction and label vertexes. These innovative designs empower our PVD to demonstrate its superiority with state-of-the-art performance across various grounding tasks. Zesen Cheng, Kehan Li 0002, Peng Jin 0001, Siheng Li, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
AAAI | 8 |
| 2024 | FaceChain-SuDe: Building Derived Class to Inherit Category Attributes for One-Shot Subject-Driven GenerationabstractRecently, subject-driven generation has garnered significant interest due to its ability to personalize text-to-image generation. Typical works focus on learning the new subject's private attributes. However, an important fact has not been taken seriously that a subject is not an isolated new concept but should be a specialization of a certain category in the pre-trained model. This results in the subject failing to comprehensively inherit the attributes in its category, causing poor attribute-related generations. In this paper, motivated by object-oriented programming, we model the subject as a derived class whose base class is its semantic category. This modeling enables the subject to inherit public attributes from its category while learning its private attributes from the user-provided example. Specifically, we propose a plug-and-play method, Subject-Derived regularization (SuDe). It constructs the base-derived class modeling by constraining the subject-driven generated images to semantically belong to the subject's category. Extensive experiments under three baselines and two backbones on various subjects show that our SuDe enables imaginative attribute-related generations while maintaining subject fidelity. For the codes, please refer to FaceChain. Pengchong Qiao, Chang Liu 0030, Baigui Sun, Xiangyang Ji, Jie Chen 0001 |
CVPR | 6 |
| 2024 | GraCo: Granularity-Controllable Interactive SegmentationabstractInteractive Segmentation (IS) segments specific objects or parts in the image according to user input. Current IS pipelines fall into two categories: single-granularity out-put and multi-granularity output. The latter aims to allevi-ate the spatial ambiguity present in the former. However, the multi-granularity output pipeline suffers from limited interaction flexibility and produces redundant results. In this work, we introduce Granularity-Controllable Interactive Segmentation (GraCo), a novel approach that allows precise control of prediction granularity by introducing ad-ditional parameters to input. This enhances the customization of the interactive system and eliminates redundancy while resolving ambiguity. Nevertheless, the exorbitant cost of annotating multi-granularity masks and the lack of avail-able datasets with granularity annotations make it difficult for models to acquire the necessary guidance to control out-put granularity. To address this problem, we design an any-granularity mask generator that exploits the semantic property of the pre-trained IS model to automatically gen-erate abundant mask-granularity pairs without requiring additional manual annotation. Based on these pairs, we propose a granularity-controllable learning strategy that efficiently imparts the granularity controllability to the IS model. Extensive experiments on intricate scenarios at ob-ject and part levels demonstrate that our GraCo has signifi-cant advantages over previous methods. This highlights the potential of GraCo to be a flexible annotation tool, capable of adapting to diverse segmentation scenarios. The project page: https://zhao-yian.github.io/GraCo. Yian Zhao, Kehan Li 0002, Zesen Cheng, Pengchong Qiao, Xiawu Zheng, Rongrong Ji, Chang Liu 0030, Li Yuan 0007, Jie Chen 0001 |
CVPR | 9 |
| 2024 | Textual Grounding for Open-Vocabulary Visual Information Extraction in Layout-Diversified Documents
Mengjun Cheng, Chengquan Zhang, Chang Liu 0047, Xiawu Zheng, Rongrong Ji, Jie Chen 0001 |
ECCV (45) | 9 |
| 2024 | Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
Peng Jin 0001, Hao Li 0073, Zesen Cheng, Kehan Li 0002, Runyi Yu 0002, Chang Liu 0047, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
ECCV (25) | 9 |
| 2024 | Learning Pseudo 3D Guidance for View-Consistent Texturing with 2D Diffusion
Kehan Li 0002, Yanbo Fan, Yang Wu 0001, Zhongqian Sun, Wei Yang 0019, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
ECCV (86) | 8 |
| 2024 | ParCo: Part-Coordinating Text-to-Motion Synthesis
Qiran Zou, Shangyuan Yuan, Shian Du, Yu Wang 0027, Chang Liu 0030, Yi Xu 0008, Jie Chen 0001, Xiangyang Ji |
ECCV (56) | 7 |
| 2024 | Protein-Ligand Interaction Prior for Binding-aware 3D Molecule Diffusion ModelsabstractGenerating 3D ligand molecules that bind to specific protein targets via diffusion models has shown great promise for structure-based drug design. The key idea is to disrupt molecules into noise through a fixed forward process and learn its reverse process to generate molecules from noise in a denoising way. However, existing diffusion models primarily focus on incorporating protein-ligand interaction information solely in the reverse process, and neglect the interactions in the forward process. The inconsistency between forward and reverse processes may impair the binding affinity of generated molecules towards target protein. In this paper, we propose a novel Interaction Prior-guided Diffusion model (IPDiff) for the protein-specific 3D molecular generation by introducing geometric protein-ligand interactions into both diffusion and sampling process. Specifically, we begin by pretraining a protein-ligand interaction prior network (IPNet) by utilizing the binding affinity signals as supervision. Subsequently, we leverage the pretrained prior network to (1) integrate interactions between the target protein and the molecular ligand into the forward process for adapting the molecule diffusion trajectories (prior-shifting), and (2) enhance the binding-aware molecule sampling process (prior-conditioning). Empirical studies on CrossDocked2020 dataset show IPDiff can generate molecules with more realistic 3D structures and state-of-the-art binding affinities towards the protein targets, with up to -6.42 Avg. Vina Score, while maintaining proper molecular properties. https://github.com/YangLing0818/IPDiff Zhilin Huang, Ling Yang 0006, Xiangxin Zhou, Wentao Zhang 0001, Xiawu Zheng, Jie Chen 0001, Yu Wang 0008, Bin Cui 0001, Wenming Yang |
ICLR | 7 |
| 2024 | Training-Free Transformer Architecture Search With Zero-Cost Proxy Guided EvolutionabstractTransformers have shown remarkable performance, however, their architecture design is a time-consuming process that demands expertise and trial-and-error. Thus, it is worthwhile to investigate efficient methods for automatically searching high-performance Transformers via Transformer Architecture Search (TAS). In order to improve the search efficiency, training-free proxy based methods have been widely adopted in Neural Architecture Search (NAS). Whereas, these proxies have been found to be inadequate in generalizing well to Transformer search spaces, as confirmed by several studies and our own experiments. This paper presents an effective scheme for TAS called TRansformer Architecture search with ZerO-cost pRoxy guided evolution (T-Razor) that achieves exceptional efficiency. First, through theoretical analysis, we discover that the synaptic diversity of multi-head self-attention (MSA) and the saliency of multi-layer perceptron (MLP) are correlated with the performance of corresponding Transformers. The properties of synaptic diversity and synaptic saliency motivate us to introduce the ranks of synaptic diversity and saliency that denoted as DSS++ for evaluating and ranking Transformers. DSS++ incorporates correlation information among sampled Transformers to provide unified scores for both synaptic diversity and synaptic saliency. We then propose a block-wise evolution search guided by DSS++ to find optimal Transformers. DSS++ determines the positions for mutation and crossover, enhancing the exploration ability. Experimental results demonstrate that our T-Razor performs competitively against the state-of-the-art manually or automatically designed Transformer architectures across four popular Transformer search spaces. Significantly, T-Razor improves the searching efficiency across different Transformer search spaces, e.g., reducing required GPU days from more than 24 to less than 0.4 and outperforming existing zero-cost approaches. We also apply T-Razor to the BERT search space and find that the searched Transformers achieve competitive GLUE results on several Neural Language Processing (NLP) datasets. This work provides insights into training-free TAS, revealing the usefulness of evaluating Transformers based on the properties of their different blocks. Qinqin Zhou 0001, Kekai Sheng, Xiawu Zheng, Ke Li 0015, Yonghong Tian 0001, Jie Chen 0001, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | An Organ-Aware Diagnosis Framework for Radiology Report GenerationabstractRadiology report generation (RRG) is crucial to save the valuable time of radiologists in drafting the report, therefore increasing their work efficiency. Compared to typical methods that directly transfer image captioning technologies to RRG, our approach incorporates organ-wise priors into the report generation. Specifically, in this paper, we propose Organ-aware Diagnosis (OaD) to generate diagnostic reports containing descriptions of each physiological organ. During training, we first develop a task distillation (TD) module to extract organ-level descriptions from reports. We then introduce an organ-aware report generation module that, for one thing, provides a specific description for each organ, and for another, simulates clinical situations to provide short descriptions for normal cases. Furthermore, we design an auto-balance mask loss to ensure balanced training for normal/abnormal descriptions and various organs simultaneously. Being intuitively reasonable and practically simple, our OaD outperforms SOTA alternatives by large margins on commonly used IU-Xray and MIMIC-CXR datasets, as evidenced by a 3.4% BLEU-1 improvement on MIMIC-CXR and 2.0% BLEU-2 improvement on IU-Xray. Pengchong Qiao, Lin Wang 0026, Munan Ning, Li Yuan 0007, Yefeng Zheng 0001, Jie Chen 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Flexible Alignment Super-Resolution Network for Multi-Contrast Magnetic Resonance ImagingabstractSuper-resolution is essential in improving the image quality of Magnetic Resonance Imaging (MRI). Existing MRI Super-Resolution methods leverage multi-contrast MRI and achieve satisfied effects. However, these methods perform alignment by calculating the similarity of single-scale semantic features between reference images and low-resolution images, which causes misalignment and limits the performance of MRI Super-Resolution. To tackle this problem, we propose the Flexible Alignment Super-resolution Network (FASR-Net) for multi-contrast MRI Super-resolution, which explores the interaction of multi-scale features. To this end, we first use the feature extractor to generate multi-scale features, including hierarchical features and semantic pyramid features. Subsequently, we introduce the Hierarchical-Feature Alignment (HF) module and the Semantic-Pyramid-Feature Alignment (SF) module to align hierarchical features and semantic pyramid features, respectively. Finally, the Cross-Hierarchical Progressive Fusion (CHPF) module fuses these aligned features at different scales, which further improves the model's performance. Extensive experiments on FastMRI and IXI datasets show that FASR-net achieves the most competitive results over state-of-the-art approaches. Our code will be available atFASR-Net. Bo Hou 0001, Jie Chen 0001, Heqing Lian |
IEEE Trans. Multim. | 6 |
| 2024 | Two-Stage Perceptual Quality Oriented Rate Control Algorithm for HEVCabstractAs a practical technique in mainstream video coding applications, rate control dominates important to ensure compression quality with limited bitrates constraints. However, most rate control methods mainly focus on objective quality while ignoring the perceptual quality improvement for human eyes. In this paper, we propose a two-stage rate control algorithm to optimize the perceptual quality at the frame encoding stage and the coding tree unit (CTU) encoding stage for high efficiency video coding (HEVC), respectively. Firstly, for the frame encoding stage, with inter-frame distortion dependency consideration, a frame-level rate control method is presented by adjusting the frame-level Lagrange multiplier adaptively with a preprocessing method. Secondly, for the CTU encoding stage, we propose a saliency-based CTU-level perceptual quality rate control algorithm, which employs CTU-level saliency weight to adjust the perceptual rate-distortion (R-D) model. We conduct the CTU-level rate control by an optimized Lagrange multiplier and quantization parameter (QP) to achieve perceptual quality optimization. Extensive experimental results reveal that, compared with state-of-the-art rate control methods on HEVC, our algorithm achieves significant perceptual coding performance with improved subjective visual quality. Yunyao Yan, Guoqing Xiang, Huizhu Jia, Jie Chen 0001, Xiaofeng Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | ACSeg: Adaptive Conceptualization for Unsupervised Semantic SegmentationabstractRecently, self-supervised large-scale visual pre-training models have shown great promise in representing pixel-level semantic relationships, significantly promoting the development of unsupervised dense prediction tasks, e.g., unsupervised semantic segmentation (USS). The extracted relationship among pixel-level representations typically contains rich class-aware information that semantically identical pixel embeddings in the representation space gather together to form sophisticated concepts. However, leveraging the learned models to ascertain semantically consistent pixel groups or regions in the image is non-trivial since over/ under-clustering overwhelms the conceptualization procedure under various semantic distributions of different images. In this work, we investigate the pixel-level semantic aggregation in self-supervised ViT pre-trained models as image Segmentation and propose the Adaptive Conceptualization approach for USS, termed ACSeg. Concretely, we explicitly encode concepts into learnable prototypes and design the Adaptive Concept Generator (ACG), which adaptively maps these prototypes to informative concepts for each image. Meanwhile, considering the scene complexity of different images, we propose the modularity loss to optimize ACG independent of the concept number based on estimating the intensity of pixel pairs belonging to the same concept. Finally, we turn the USS task into classifying the discovered concepts in an unsupervised manner. Extensive experiments with state-of-the-art results demonstrate the effectiveness of the proposed ACSeg. Kehan Li 0002, Zhennan Wang 0001, Zesen Cheng, Runyi Yu 0002, Yian Zhao, Guoli Song, Chang Liu 0030, Li Yuan 0007, Jie Chen 0001 |
CVPR | 9 |
| 2023 | From Node Interaction to Hop Interaction: New Effective and Scalable Graph Learning ParadigmabstractExisting Graph Neural Networks (GNNs) follow the message-passing mechanism that conducts information interaction among nodes iteratively. While considerable progress has been made, such node interaction paradigms still have the following limitation. First, the scalability limitation precludes the broad application of GNNs in large-scale industrial settings since the node interaction among rapidly expanding neighbors incurs high computation and memory costs. Second, the over-smoothing problem restricts the discrimination ability of nodes, i.e., node representations of different classes will converge to indistinguishable after repeated node interactions. In this work, we propose a novel hop interaction paradigm to address these limitations simultaneously. The core idea is to convert the interaction target among nodes to pre-processed multi-hop features inside each node. We design a simple yet effective HopGNN framework that can easily utilize existing GNNs to achieve hop interaction. Furthermore, we propose a multi-task learning strategy with a self-supervised learning objective to enhance HopGNN. We conduct extensive experiments on 12 benchmark datasets in a wide range of domains, scales, and smoothness of graphs. Experimental results show that our methods achieve superior performance while maintaining high scalability and efficiency. The code is at https://github.com/JC-202/HopGNN. Jie Chen 0001, Zilong Li 0001, Junping Zhang, Jian Pu |
CVPR | 1 |
| 2023 | Out-of-Candidate Rectification for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation is typically inspired by class activation maps, which serve as pseudo masks with class-discriminative regions highlighted. Although tremendous efforts have been made to recall precise and complete locations for each class, existing methods still commonly suffer from the unsolicited Out-of-Candidate (OC) error predictions that do not belong to the label candidates, which could be avoidable since the contradiction with image-level class tags is easy to be detected. In this paper, we develop a group ranking-based Out-of-f;Candidate Rectification (OCR) mechanism in a plug-and-play fashion. Firstly, we adaptively split the semantic categories into In-Candidate (IC) and OC groups for each OC pixel according to their prior annotation correlation and posterior prediction correlation. Then, we derive a differentiable rectification loss to force OC pixels to shift to the IC group. Incorporating OCR with seminal baselines (e.g., AffinityNet, SEAM, MCTformer), we can achieve remarkable performance gains on both Pascal VOC (+3.2%, +3.3%, +0.8% mIoU) and MS COCO (+1.0%, +1.3%, +0.5% mIoU) datasets with negligible extra training overhead, which jus-tifies the effectiveness and generality of OCR.††Ŋ github.com/sennnnn/Out-of-Candidate-Rectification Zesen Cheng, Pengchong Qiao, Kehan Li 0002, Siheng Li, Pengxu Wei, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
CVPR | 9 |
| 2023 | Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation LearningabstractContrastive learning-based video-language representation learning approaches, e.g., CLIP, have achieved outstanding performance, which pursue semantic interaction upon pre-defined video-text pairs. To clarify this coarse-grained global interaction and move a step further, we have to encounter challenging shell-breaking interactions for fine-grained cross-modal learning. In this paper, we creatively model video-text as game players with multivariate cooperative game theory to wisely handle the uncertainty during fine-grained semantic interaction with diverse granularity, flexible combination, and vague intensity. Concretely, we propose Hierarchical Banzhaf Interaction (HBI) to value possible correspondence between video frames and text words for sensitive and explainable cross-modal contrast. To efficiently realize the cooperative game of multiple video frames and multiple text words, the proposed method clusters the original video frames (text words) and computes the Banzhaf Interaction between the merged tokens. By stacking token merge modules, we achieve cooperative games at different semantic levels. Extensive experiments on commonly used text-video retrieval and video-question answering bench-marks with superior performances justify the efficacy of our HBI. More encouragingly, it can also serve as a visualization tool to promote the understanding of cross-modal interaction, which have a far-reaching impact on the community. Project page is available at https://jpthu17.github.io/HBI/. Peng Jin 0001, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu 0030, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
CVPR | 8 |
| 2023 | Fuzzy Positive Learning for Semi-Supervised Semantic SegmentationabstractSemi-supervised learning (SSL) essentially pursues class boundary exploration with less dependence on human annotations. Although typical attempts focus on ameliorating the inevitable error-prone pseudo-labeling, we think differently and resort to exhausting informative semantics from multiple probably correct candidate labels. In this paper, we introduce Fuzzy Positive Learning (FPL) for accurate SSL semantic segmentation in a plug-and-play fashion, targeting adaptively encouraging fuzzy positive predictions and suppressing highly-probable negatives. Being conceptually simple yet practically effective, FPL can remarkably alleviate interference from wrong pseudo labels and progressively achieve clear pixel-level semantic discrimination. Concretely, our FPL approach consists of two main components, including fuzzy positive assignment (FPA) to provide an adaptive number of labels for each pixel and fuzzy positive regularization (FPR) to restrict the predictions of fuzzy positive categories to be larger than the rest under different perturbations. Theoretical analysis and extensive experiments on Cityscapes and VOC 2012 with consistent performance gain justify the superiority of our approach. Codes are provided in https://github.com/qpc1611094/FPL. Pengchong Qiao, Zhidan Wei, Yu Wang 0027, Zhennan Wang 0001, Guoli Song, Xiangyang Ji, Chang Liu 0030, Jie Chen 0001 |
CVPR | 9 |
| 2023 | Out-of-Distributed Semantic Pruning for Robust Semi-Supervised LearningabstractRecent advances in robust semi-supervised learning (SSL) typically filter out-of-distribution (OOD) information at the sample level. We argue that an overlooked problem of robust SSL is its corrupted information on semantic level, practically limiting the development of the field. In this paper, we take an initial step to explore and propose a unified framework termed OOD Semantic Pruning (OSP), which aims at pruning OOD semantics out from in-distribution (ID) features. Specifically, (i) we propose an aliasing OOD matching module to pair each ID sample with an OOD sample with semantic overlap. (ii) We design a soft orthogonality regularization, which first transforms each ID feature by suppressing its semantic component that is collinear with paired OOD sample. It then forces the predictions before and after soft orthogonality decomposition to be consistent. Being practically simple, our method shows a strong performance in OOD detection and ID classification on challenging benchmarks. In particular, OSP surpasses the previous state-of-the-art by 13.7% on accuracy for ID classification and 5.9% on AUROC for OOD detection on TinyImageNet dataset. The source codes are publicly available at https://github.com/rain305f/OSP. Yu Wang 0027, Pengchong Qiao, Chang Liu 0030, Guoli Song, Xiawu Zheng, Jie Chen 0001 |
CVPR | 6 |
| 2023 | Learning Task-Aligned Mask Query for Instance SegmentationabstractRecently, query-based instance segmentation methods have achieved comparable performance to previous state-of-the-art methods. However, the query lacks the learning of the consistency between classification and segmentation tasks, which may lead to misalignment between classification score and mask quality (i.e., mask IoU) and can not result in a reliable ranking for predictions. In this work, we propose a novel instance segmentation method, termed AlignMask, which effectively learns task-aligned mask queries for instance end-toend. Specifically, we propose Aligned Query Learning (AQL) to learn task-aligned features for pixel embedding and transformer decoder, which helps segmentation quality estimation of the mask query. We also use Aligned Label Assignment to explicitly align the optimization goals for classification score and mask quality of the query. Extensive experiments on MSCOCO show that our proposed AlignMask achieves competitive performance with state-of-the-art models. Pengxu Wei, Jie Chen 0001 |
ICASSP | 4 |
| 2023 | Recurrent Fine-Grained Self-Attention Network for Video Crowd CountingabstractStriking a balance between exploring the spatio-temporal correlation and controlling model complexity is vital for video-based crowd counting methods. In this paper, we propose a Recurrent Fine-Grained Self-Attention Network (RFSNet) to achieve efficient and accurate counting in video scenes via the self-attention mechanism and a recurrent fine-tuning strategy. Specifically, we design a decoder which consists of patch-wise spatial self-attention and temporal self-attention. Compared with vanilla self-attention, it effectively leverages the dependencies in spatial and temporal domain respectively, while significantly reducing computational complexity. Moreover, the RFSNet recurrently feeds the features into the decoder to enhance the spatio-temporal representations. This strategy not only simplifies the model structure and reduces the number of parameters, but also improves the quality of estimated density maps. Our RFSNet achieves state-of-the-art performance on three video crowd counting benchmarks, and outperforms other methods by more than 20% on the challenging FDST dataset. Jifan Zhang, Zhe Wu 0006, Xinfeng Zhang 0001, Guoli Song, Yaowei Wang 0001, Jie Chen 0001 |
ICASSP | 6 |
| 2023 | DiffusionRet: Generative Text-Video Retrieval with Diffusion ModelabstractExisting text-video retrieval solutions are, in essence, discriminant models focused on maximizing the conditional likelihood, i.e., p(candidates|query). While straightforward, this de facto paradigm overlooks the underlying data distribution p(query), which makes it challenging to identify out-of-distribution data. To address this limitation, we creatively tackle this task from a generative viewpoint and model the correlation between the text and the video as their joint probability p(candidates,query). This is accomplished through a diffusion-based text-video retrieval framework (Diffusion-Ret), which models the retrieval task as a process of gradually generating joint distribution from noise. During training, DiffusionRet is optimized from both the generation and discrimination perspectives, with the generator being optimized by generation loss and the feature extractor trained with contrastive loss. In this way, DiffusionRet cleverly leverages the strengths of both generative and discriminative methods. Extensive experiments on five commonly used text-video retrieval benchmarks, including MSRVTT, LSMDC, MSVD, ActivityNet Captions, and DiDeMo, with superior performances, justify the efficacy of our method. More encouragingly, without any modification, DiffusionRet even performs well in out-domain retrieval settings. We believe this work brings fundamental insights into the related fields. Code is available at https://github.com/jpthu17/DiffusionRet. Peng Jin 0001, Hao Li 0073, Zesen Cheng, Kehan Li 0002, Xiangyang Ji, Chang Liu 0030, Li Yuan 0007, Jie Chen 0001 |
ICCV | 8 |
| 2023 | Multi-granularity Interaction Simulation for Unsupervised Interactive SegmentationabstractInteractive segmentation enables users to segment as needed by providing cues of objects, which introduces human-computer interaction for many fields, such as image editing and medical image analysis. Typically, massive and expansive pixel-level annotations are spent to train deep models by object-oriented interactions with manually labeled object masks. In this work, we reveal that informative interactions can be made by simulation with semantic-consistent yet diverse region exploration in an unsupervised paradigm. Concretely, we introduce a Multi-granularity Interaction Simulation (MIS) approach to open up a promising direction for unsupervised interactive segmentation. Drawing on the high-quality dense features produced by recent self-supervised models, we propose to gradually merge patches or regions with similar features to form more extensive regions and thus, every merged region serves as a semantic-meaningful multi-granularity proposal. By randomly sampling these proposals and simulating possible interactions based on them, we provide meaningful interaction at multiple granularities to teach the model to understand interactions. Our MIS significantly outperforms non-deep learning unsupervised methods and is even comparable with some previous deep-supervised methods without any annotation. Kehan Li 0002, Yian Zhao, Zhennan Wang 0001, Zesen Cheng, Peng Jin 0001, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
ICCV | 9 |
| 2023 | TopoSeg: Topology-Aware Nuclear Instance SegmentationabstractNuclear instance segmentation has been critical for pathology image analysis in medical science, e.g., cancer diagnosis. Current methods typically adopt pixel-wise optimization for nuclei boundary exploration, where rich structural information could be lost for subsequent quantitative morphology assessment. To address this issue, we develop a topology-aware segmentation approach, termed TopoSeg, which exploits topological structure information to keep the predictions rational, especially in common situations with densely touching and overlapping nucleus instances. Concretely, TopoSeg builds on a topology-aware module (TAM), which encodes dynamic changes of different topology structures within the three-class probability maps (inside, boundary, and background) of the nuclei to persistence barcodes and makes the topology-aware loss function. To efficiently focus on regions with high topological errors, we propose an adaptive topology-aware selection (ATS) strategy to enhance the topology-aware optimization procedure further. Experiments on three nuclear instance segmentation datasets justify the superiority of TopoSeg, which achieves state-of-the-art performance. The code is available at https://github.com/hhlisme/toposeg. Pengxu Wei, Xiangyang Ji, Chang Liu 0030, Jie Chen 0001 |
ICCV | 7 |
| 2023 | Learning to Distill Global Representation for Sparse-View CTabstractSparse-view computed tomography (CT)—using a small number of projections for tomographic reconstruction—enables much lower radiation dose to patients and accelerated data acquisition. The reconstructed images, however, suffer from strong artifacts, greatly limiting their diagnostic value. Current trends for sparse-view CT turn to the raw data for better information recovery. The resultant dual-domain methods, nonetheless, suffer from secondary artifacts, especially in ultra-sparse view scenarios, and their generalization to other scanners/protocols is greatly limited. A crucial question arises: have the image post-processing methods reached the limit? Our answer is not yet. In this paper, we stick to image post-processing methods due to great flexibility and propose global representation(GloRe) distillation framework for sparse-view CT, termed GloReDi. First, we propose to learn GloRe with Fourier convolution, so each element in GloRe has an image-wide receptive field. Second, unlike methods that only use the full-view images for supervision, we propose to distill GloRe from intermediate-view reconstructed images that are readily available but not explored in previous literature. The success of GloRe distillation is attributed to two key components: representation directional distillation to align the GloRe directions, and band-pass-specific contrastive distillation to gain clinically important details. Extensive experiments demonstrate the superiority of the proposed GloReDi over the state-of-the-art methods, including dual-domain ones. The source code is available at https://github.com/longzilicart/GloReDi. Zilong Li 0001, Chenglong Ma 0002, Jie Chen 0001, Junping Zhang, Hongming Shan |
ICCV | 3 |
| 2023 | Towards Real-World Burst Image Super-Resolution: Benchmark and MethodabstractDespite substantial advances, single-image super-resolution (SISR) is always in a dilemma to reconstruct high-quality images with limited information from one input image, especially in realistic scenarios. In this paper, we establish a large-scale real-world burst super-resolution dataset, i.e., RealBSR, to explore the faithful reconstruction of image details from multiple frames. Furthermore, we introduce a Federated Burst Affinity network (FBAnet) to investigate non-trivial pixel-wise displacements among images under real-world image degradation. Specifically, rather than using pixel-wise alignment, our FBAnet employs a simple homography alignment from a structural geometry aspect and a Federated Affinity Fusion (FAF) strategy to aggregate the complementary information among frames. Those fused informative representations are fed to a Transformer-based module of burst representation decoding. Besides, we have conducted extensive experiments on two versions of our datasets, i.e., RealBSR-RAW and RealBSR-RGB. Experimental results demonstrate that our FBAnet outperforms existing state-of-the-art burst SR methods and also achieves visually-pleasant SR image predictions with model details. Our dataset, codes, and models are publicly available at https://github.com/yjsunnn/FBANet. Pengxu Wei, Yujing Sun 0004, Xingbei Guo, Chang Liu 0030, Guanbin Li, Jie Chen 0001, Xiangyang Ji, Liang Lin 0004 |
ICCV | 6 |
| 2023 | LaPE: Layer-adaptive Position Embedding for Vision Transformers with Independent Layer NormalizationabstractPosition information is critical for Vision Transformers (VTs) due to the permutation-invariance of self-attention operations. A typical way to introduce position information is adding the absolute Position Embedding (PE) to patch embedding before entering VTs. However, this approach operates the same Layer Normalization (LN) to token embedding and PE, and delivers the same PE to each layer. This results in restricted and monotonic PE across layers, as the shared LN affine parameters are not dedicated to PE, and the PE cannot be adjusted on a per-layer basis. To overcome these limitations, we propose using two independent LNs for token embeddings and PE in each layer, and progressively delivering PE across layers. By implementing this approach, VTs will receive layer-adaptive and hierarchical PE. We name our method as Layer-adaptive Position Embedding, abbreviated as LaPE, which is simple, effective, and robust. Extensive experiments on image classification, object detection, and semantic segmentation demonstrate that LaPE significantly outperforms the default PE method. For example, LaPE improves +1.06% for CCT on CIFAR100, +1.57% for DeiT-Ti on ImageNet-1K, +0.7 box AP and +0.5 mask AP for ViT-Adapter-Ti on COCO, and +1.37 mIoU for tiny Segmenter on ADE20K. This is remarkable considering LaPE only increases negligible parameters, memory, and computational cost. Runyi Yu 0002, Zhennan Wang 0001, Yinhuai Wang, Kehan Li 0002, Chang Liu 0030, Haoyi Duan, Xiangyang Ji, Jie Chen 0001 |
ICCV | 8 |
| 2023 | WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image SegmentationabstractThe top-down and bottom-up methods are two mainstreams of referring segmentation, while both methods have their own intrinsic weaknesses. Top-down methods are chiefly disturbed by Polar Negative (PN) errors owing to the lack of fine-grained cross-modal alignment. Bottom-up methods are mainly perturbed by Inferior Positive (IP) errors due to the lack of prior object information. Nevertheless, we discover that two types of methods are highly complementary for restraining respective weaknesses but the direct average combination leads to harmful interference. In this context, we build Win-win Cooperation (WiCo) to exploit complementary nature of two types of methods on both interaction and integration aspects for achieving a win-win improvement. For the interaction aspect, Complementary Feature Interaction (CFI) introduces prior object information to bottom-up branch and provides fine-grained information to top-down branch for complementary feature enhancement. For the integration aspect, Gaussian Scoring Integration (GSI) models the gaussian performance distributions of two branches and weighted integrates results by sampling confident scores from the distributions. With our WiCo, several prominent bottom-up and top-down combinations achieve remarkable improvements on three common datasets with reasonable extra costs, which justifies effectiveness and generality of our method. Zesen Cheng, Peng Jin 0001, Hao Li 0073, Kehan Li 0002, Siheng Li, Xiangyang Ji, Chang Liu 0030, Jie Chen 0001 |
IJCAI | 8 |
| 2023 | Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set AlignmentabstractText-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local details or are computationally expensive. What's worse, they fail to leverage the heterogeneous concepts in data. In this paper, we propose the Disentangled Conceptualization and Set-to-set Alignment (DiCoSA) to simulate the conceptualizing and reasoning process of human beings. For disentangled conceptualization, we divide the coarse feature into multiple latent factors related to semantic concepts. For set-to-set alignment, where a set of visual concepts correspond to a set of textual concepts, we propose an adaptive pooling method to aggregate semantic concepts to address the partial matching. In particular, since we encode concepts independently in only a few dimensions, DiCoSA is superior at efficiency and granularity, ensuring fine-grained interactions using a similar computational complexity as coarse-grained alignment. Extensive experiments on five datasets, including MSR-VTT, LSMDC, MSVD, ActivityNet, and DiDeMo, demonstrate that our method outperforms the existing state-of-the-art methods. Peng Jin 0001, Hao Li 0073, Zesen Cheng, Jinfa Huang, Zhennan Wang 0001, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
IJCAI | 8 |
| 2023 | TG-VQA: Ternary Game of Video Question AnsweringabstractVideo question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions, i.e., annotations or priors, current contrastive learning-based VideoQA methods remains challenging to perform fine-grained visual-linguistic alignments. In this work, we innovatively resort to game theory, which can simulate complicated relationships among multiple players with specific interaction strategies, e.g., video, question, and answer as ternary players, to achieve fine-grained alignment for VideoQA task. Specifically, we carefully design a VideoQA-specific interaction strategy to tailor the characteristics of VideoQA, which can mathematically generate the fine-grained visual-linguistic alignment label without label-intensive efforts. Our TG-VQA outperforms existing state-of-the-art by a large margin (more than 5%) on long-term and short-term VideoQA datasets, verifying its effectiveness and generalization ability. Thanks to the guidance of game-theoretic interaction, our model impressively convergences well on limited data (10^4 videos), surpassing most of those pre-trained on large-scale data (10^7 videos). Hao Li 0073, Peng Jin 0001, Zesen Cheng, Songyang Zhang 0001, Kai Chen 0026, Zhennan Wang 0001, Chang Liu 0030, Jie Chen 0001 |
IJCAI | 8 |
| 2023 | LocLoc: Low-level Cues and Local-area Guides for Weakly Supervised Object LocalizationabstractWeakly Supervised Object Localization (WSOL) aims to localize objects using only image-level labels while ensuring competitive classification performance. However, previous efforts have prioritized localization over classification accuracy in discriminative features, in which low-level information is neglected. We argue that low-level image representations, such as edges, color, texture, and motions are crucial for accurate detection. That is, using such information further achieves more refined localization, which can be used to promote classification accuracy. In this paper, we propose a unified framework that simultaneously improves localization and classification accuracy, termed as LocLoc (Low-level Cues and Local-area Guides). It leverages low-level image cues to explore global and local representations for accurate localization and classification. Specifically, we introduce a GrabCut-Enhanced Generator (GEG) to learn global semantic representations for localization based on graph cuts to enhance low-level information based on long-range dependencies captured by the transformer. We further design a Local Feature Digging Module (LFDM) that utilizes low-level cues to guide the learning route of local feature representations for accurate classification. Extensive experiments demonstrate the effectiveness of LocLoc with 84.4%(↑5.2%) Top-1 Loc., 85.8% Top-1 Cls. on CUB-200-2011 and 57.6% (↑1.5%) Top-1 Loc., 78.6% Top-1Cls. on ILSVRC 2012, indicating that our method achieves competitive performance with a large margin compared to previous approaches. Code and models are available at https://github.com/Cliffia123/LocLoc. Xinzi Cao, Xiawu Zheng, Yunhang Shen, Ke Li 0015, Jie Chen 0001, Yutong Lu, Yonghong Tian 0001 |
ACM Multimedia | 5 |
| 2023 | Event-Diffusion: Event-Based Image Reconstruction and Restoration with Diffusion ModelsabstractEvent cameras offer the advantages of low latency, high temporal resolution and HDR compared to conventional cameras. Due to the asynchronous and sparse nature of events, many existing algorithms cannot be directly applied, necessitating the reconstruction of intensity frames. However, existing reconstruction methods often result in artifacts and edge blurring due to noise and event accumulation. In this paper, we argue that the key to event-based image reconstruction is to enhance the edge information of objects and restore the artifacts in the reconstructed images. To explain, edge information is one of the most important features in the event stream, providing information on the shape and contour of objects. Considering the extraordinary capabilities of Denoising Diffusion Probabilistic Models (DDPMs) in image generation, reconstruction, and restoration, we propose a new framework which incorporate it into the reconstruction pipeline to obtain high-quality results which effectively remove artifacts and blur in reconstructed images. Specifically, we first extract edge information from the event stream using the proposed event-based denoising method. It employs the contrast maximization framework to remove noise from the event stream and extract clear object edge information. And then, the edge information is further adopted to our diffusion model, which is used to enhance the edges of objects in the reconstructed images, thus improving the restoration effect. Experimental results show that our method achieves significant improvements in the mean squared error (MSE), the structural similarity (SSIM), and the perceptual similarity (LPIPS) metrics, with average improvements of 40%, 15%, and 25%, respectively, compared to previous state-of-the-art models, and has good generalization performance. Quanmin Liang, Xiawu Zheng, Kai Huang 0001, Yan Zhang 0109, Jie Chen 0001, Yonghong Tian 0001 |
ACM Multimedia | 5 |
| 2023 | Discover and Align Taxonomic Context Priors for Open-world Semi-Supervised LearningabstractOpen-world Semi-Supervised Learning (OSSL) is a realistic and challenging task, aiming to classify unlabeled samples from both seen and novel classes using partially labeled samples from the seen classes.
Previous works typically explore the relationship of samples as priors on the pre-defined single-granularity labels to help novel class recognition. In fact, classes follow a taxonomy and samples can be classified at multiple levels of granularity, which contains more underlying relationships for supervision. We thus argue that learning with single-granularity labels results in sub-optimal representation learning and inaccurate pseudo labels, especially with unknown classes. In this paper, we take the initiative to explore and propose a uniformed framework, called Taxonomic context prIors Discovering and Aligning (TIDA), which exploits the relationship of samples under various granularity. It allows us to discover multi-granularity semantic concepts as taxonomic context priors (i.e., sub-class, target-class, and super-class), and then collaboratively leverage them to enhance representation learning and improve the quality of pseudo labels.
Specifically, TIDA comprises two components: i) A taxonomic context discovery module that constructs a set of hierarchical prototypes in the latent space to discover the underlying taxonomic context priors; ii) A taxonomic context-based prediction alignment module that enforces consistency across hierarchical predictions to build the reliable relationship between classes among various granularity and provide additions supervision. We demonstrate that these two components are mutually beneficial for an effective OSSL framework, which is theoretically explained from the perspective of the EM algorithm. Extensive experiments on seven commonly used datasets show that TIDA can significantly improve the performance and achieve a new state of the art. The source codes are publicly available at https://github.com/rain305f/TIDA. Yu Wang 0027, Zhun Zhong, Pengchong Qiao, Xuxin Cheng, Xiawu Zheng, Chang Liu 0030, Nicu Sebe, Rongrong Ji, Jie Chen 0001 |
NeurIPS | 9 |
| 2023 | Object-Aware Transfer-Based Black-Box Adversarial Attack on Object Detector
Zhuo Leng, Zesen Cheng, Pengxu Wei, Jie Chen 0001 |
PRCV (12) | 4 |
| 2023 | Two-Stage Deep Learning Segmentation for Tiny Brain Regions
Xiawu Zheng, Rongrong Ji, Jie Chen 0001 |
PRCV (13) | 4 |
| 2023 | AFP-MFL: accurate identification of antifungal peptides using multi-view feature learningabstractRecently, peptide-based drugs have gained unprecedented interest in discovering and developing antifungal drugs due to their high efficacy, broad-spectrum activity, low toxicity and few side effects. However, it is time-consuming and expensive to identify antifungal peptides (AFPs) experimentally. Therefore, computational methods for accurately predicting AFPs are highly required. In this work, we develop AFP-MFL, a novel deep learning model that predicts AFPs only relying on peptide sequences without using any structural information. AFP-MFL first constructs comprehensive feature profiles of AFPs, including contextual semantic information derived from a pre-trained protein language model, evolutionary information, and physicochemical properties. Subsequently, the co-attention mechanism is utilized to integrate contextual semantic information with evolutionary information and physicochemical properties separately. Extensive experiments show that AFP-MFL outperforms state-of-the-art models on four independent test datasets. Furthermore, the SHAP method is employed to explore each feature contribution to the AFPs prediction. Finally, a user-friendly web server of the proposed AFP-MFL is developed and freely accessible at http://inner.wei-group.net/AFPMFL/, which can be considered as a powerful tool for the rapid screening and identification of novel AFPs. Yitian Fang, Lesong Wei, Jie Chen 0001, Leyi Wei |
Briefings Bioinform. | 5 |
| 2023 | Learning Representation for Clustering Via Prototype Scattering and Positive SamplingabstractExisting deep clustering methods rely on either contrastive or non-contrastive representation learning for downstream clustering task. Contrastive-based methods thanks to negative pairs learn uniform representations for clustering, in which negative pairs, however, may inevitably lead to the class collision issue and consequently compromise the clustering performance. Non-contrastive-based methods, on the other hand, avoid class collision issue, but the resulting non-uniform representations may cause the collapse of clustering. To enjoy the strengths of both worlds, this paper presents a novel end-to-end deep clustering method with prototype scattering and positive sampling, termed ProPos. Specifically, we first maximize the distance between prototypical representations, named prototype scattering loss, which improves the uniformity of representations. Second, we align one augmented view of instance with the sampled neighbors of another view-assumed to be truly positive pair in the embedding space-to improve the within-cluster compactness, termed positive sampling alignment. The strengths of ProPos are avoidable class collision issue, uniform representations, well-separated clusters, and within-cluster compactness. By optimizing ProPos in an end-to-end expectation-maximization framework, extensive experimental results demonstrate that ProPos achieves competing performance on moderate-scale clustering benchmark datasets and establishes new state-of-the-art performance on large-scale datasets. Source code is available at https://github.com/Hzzone/ProPos. Zhizhong Huang, Jie Chen 0001, Junping Zhang, Hongming Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Weakly-Supervised 3D Spatial Reasoning for Text-Based Visual Question AnsweringabstractText-based Visual Question Answering (TextVQA) aims to produce correct answers for given questions about the images with multiple scene texts. In most cases, the texts naturally attach to the surface of the objects. Therefore, spatial reasoning between texts and objects is crucial in TextVQA. However, existing approaches are constrained within 2D spatial information learned from the input images and rely on transformer-based architectures to reason implicitly during the fusion process. Under this setting, these 2D spatial reasoning approaches cannot distinguish the fine-grained spatial relations between visual objects and scene texts on the same image plane, thereby impairing the interpretability and performance of TextVQA models. In this paper, we introduce 3D geometric information into the spatial reasoning process to capture the contextual knowledge of key objects step-by-step. Specifically, (i) we propose a relation prediction module for accurately locating the region of interest of critical objects; (ii) we design a depth-aware attention calibration module for calibrating the OCR tokens' attention according to critical objects. Extensive experiments show that our method achieves state-of-the-art performance on TextVQA and ST-VQA datasets. More encouragingly, our model surpasses others by clear margins of 5.7% and 12.1% on questions that involve spatial reasoning in TextVQA and ST-VQA valid split. Besides, we also verify the generalizability of our model on the text-based image captioning task. Hao Li 0073, Jinfa Huang, Peng Jin 0001, Guoli Song, Qi Wu 0001, Jie Chen 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | Robust and Hierarchical Spatial Relation Analysis for Traffic ForecastingabstractHow to model the complex spatial-temporal relation in traffic data is an important problem for precisely predicting the future status of a city traffic system. Existing traffic forecasting methods rarely consider the traffic state trend, and the robust spatial relation has not been well explored. To tackle these issues, we design a novel Robust And Hierarchical spatial Relation Analysis (RAHRA) method to calculate the local-period spatial relation, which applies temporal context information in both traffic state and trend similarities. This could capture abundant traffic patterns and learn stable and comprehensive spatial relations for accurate traffic forecasting. Furthermore, we introduce a Temporal Attention Module (TAM) to capture the temporal features and propose a Future Feature Inference Module (FFIM) to infer the future traffic information. Experiments on four real-world traffic datasets demonstrate that the proposed method outperforms the other state-of-the-art methods. Zhe Wu 0006, Xinfeng Zhang 0001, Guoli Song, Yaowei Wang 0001, Jie Chen 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Semi-Supervised CT Lesion Segmentation Using Uncertainty-Based Data Pairing and SwapMixabstractSemi-supervised learning (SSL) methods show their powerful performance to deal with the issue of data shortage in the field of medical image segmentation. However, existing SSL methods still suffer from the problem of unreliable predictions on unannotated data due to the lack of manual annotations for them. In this paper, we propose an unreliability-diluted consistency training (UDiCT) mechanism to dilute the unreliability in SSL by assembling reliable annotated data into unreliable unannotated data. Specifically, we first propose an uncertainty-based data pairing module to pair annotated data with unannotated data based on a complementary uncertainty pairing rule, which avoids two hard samples being paired off. Secondly, we develop SwapMix, a mixed sample data augmentation method, to integrate annotated data into unannotated data for training our model in a low-unreliability manner. Finally, UDiCT is trained by minimizing a supervised loss and an unreliability-diluted consistency loss, which makes our model robust to diverse backgrounds. Extensive experiments on three chest CT datasets show the effectiveness of our method for semi-supervised CT lesion segmentation. Pengchong Qiao, Guoli Song, Hu Han 0001, Yonghong Tian 0001, Yongsheng Liang 0001, Xi Li 0011, Shaohua Kevin Zhou, Jie Chen 0001 |
IEEE Trans. Medical Imaging | 10 |
| 2023 | Carrying Out CNN Channel Pruning in a White BoxabstractChannel pruning has been long studied to compress convolutional neural networks (CNNs), which significantly reduces the overall computation. Prior works implement channel pruning in an unexplainable manner, which tends to reduce the final classification errors while failing to consider the internal influence of each channel. In this article, we conduct channel pruning in a white box. Through deep visualization of feature maps activated by different channels, we observe that different channels have a varying contribution to different categories in image classification. Inspired by this, we choose to preserve channels contributing to most categories. Specifically, to model the contribution of each channel to differentiating categories, we develop a class-wise mask for each channel, implemented in a dynamic training manner with respect to the input image's category. On the basis of the learned class-wise mask, we perform a global voting mechanism to remove channels with less category discrimination. Lastly, a fine-tuning process is conducted to recover the performance of the pruned model. To our best knowledge, it is the first time that CNN interpretability theory is considered to guide channel pruning. Extensive experiments on representative image classification tasks demonstrate the superiority of our White-Box over many state-of-the-arts (SOTAs). For instance, on CIFAR-10, it reduces 65.23% floating point operations per seconds (FLOPs) with even 0.62% accuracy improvement for ResNet-110. On ILSVRC-2012, White-Box achieves a 45.6% FLOP reduction with only a small loss of 0.83% in the top-1 accuracy for ResNet-50. Code is available at https://github.com/zyxxmu/White-Box. Yuxin Zhang 0002, Mingbao Lin, Chia-Wen Lin, Jie Chen 0001, Yongjian Wu 0001, Yonghong Tian 0001, Rongrong Ji |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | ConformerDTI: Local Features Coupling Global Representations for Drug-Target Interaction PredictionabstractDrug-target interaction(DTI) prediction is one of the most important topics in drug design and drug development, and deep learning approaches have achieved state-of-the-art performance in this field. However, the current methods are difficult to successfully combine the local and global features of drug molecules and protein sequences, while ignoring the modeling of complicated interaction mechanisms, which leads to a certain limitation of prediction performance. To overcome this barrier, we propose an end-to-end method based on Convolutional Neural Network (CNN) and Transformer to predict DTI problems, named ConformerDTI. The CNN and Transformer branches extract features from the simplified molecular input line entry system (SMILES) string of drugs and the amino acid sequence of proteins, respectively. The local and global features are coupled by the mutual transfer of the two branches through cross attention. Decoupling of local and global features in parallel leverages CNN’s power in extracting local features as well as the efficiency of Transformer at global processing. I n addition, ConformerDTI exploits the convolutional interaction network to model the interaction mechanism, both drugs and targets are convoluted by dynamic filters generated based on each other. Experimental results demonstrate that our model has better prediction performance than the most advanced deep learning methods on three different datasets. Furthermore, this performance improvement was validated by ablation experiments. Wenming Yang, Jie Chen 0001, Yonghong Tian 0001 |
BIBM | 3 |
| 2022 | ViSTA: Vision and Scene Text Aggregation for Cross-Modal RetrievalabstractVisual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of scene text information and directly adding this information may lead to performance degradation in scene text free scenarios. To address this issue, we propose a full transformer architecture to unify these cross-modal retrieval scenarios in a single Vision and Scene Text Aggregation framework (ViSTA). Specifically, ViSTA utilizes transformer blocks to directly encode image patches and fuse scene text embedding to learn an aggregated visual representation for cross-modal retrieval. To tackle the modality missing problem of scene text, we propose a novel fusion token based transformer aggregation approach to exchange the necessary scene text information only through the fusion token and concentrate on the most important features in each modality. To further strengthen the visual modality, we develop dual contrastive learning losses to embed both image-text pairs and fusion-text pairs into a common cross-modal space. Compared to existing methods, ViSTA enables to aggregate relevant scene text semantics with visual appearance, and hence improve results under both scene text free and scene text aware scenarios. Experimental results show that ViSTA outperforms other methods by at least 8.4% at Recall@ 1 for scene text aware retrieval task. Compared with state-of-the-art scene text free retrieval methods, ViSTA can achieve better accuracy on Flicker30K and MSCOCO while running at least three times faster during the inference stage, which validates the effectiveness of the proposed framework. Mengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu, Jie Chen 0001, Guoli Song, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
CVPR | 6 |
| 2022 | Training-free Transformer Architecture SearchabstractRecently, Vision Transformer (ViT) has achieved remarkable success in several computer vision tasks. The progresses are highly relevant to the architecture design, then it is worthwhile to propose Transformer Architecture Search (TAS) to search for better ViTs automatically. However, current TAS methods are time-consuming and existing zero-cost proxies in CNN do not generalize well to the ViT search space according to our experimental observations. In this paper, for the first time, we investigate how to conduct TAS in a training-free manner and devise an effective training-free TAS (TF-TAS) scheme. Firstly, we observe that the properties of multi-head self-attention (MSA) and multi-layer perceptron (MLP) in ViTs are quite different and that the synaptic diversity of MSA affects the performance notably. Secondly, based on the observation, we devise a modular strategy in TF-TAS that evaluates and ranks ViT architectures from two theoretical perspectives: synaptic diversity and synaptic saliency, termed as DSS-indicator. With DSS-indicator, evaluation results are strongly corre-lated with the test accuracies of ViT models. Experimental results demonstrate that our TF- TAS achieves a competitive performance against the state-of-the-art manually or automatically design ViT architectures, and it promotes the searching efficiency in ViT search space greatly: from about 24 GPU days to less than 0.5 GPU days. Moreover, the proposed DSS-indicator outperforms the existing cutting-edge zero-cost approaches (e.g., TE-score and NASWOT). Qinqin Zhou 0001, Kekai Sheng, Xiawu Zheng, Ke Li 0015, Xing Sun 0001, Yonghong Tian 0001, Jie Chen 0001, Rongrong Ji |
CVPR | 7 |
| 2022 | Locality Guidance for Improving Vision Transformers on Tiny Datasets
Kehan Li 0002, Runyi Yu 0002, Zhennan Wang 0001, Li Yuan 0007, Guoli Song, Jie Chen 0001 |
ECCV (24) | 6 |
| 2022 | Joint Learning of Object Graph and Relation Graph for Visual Question AnsweringabstractModeling visual question answering (VQA) through scene graphs can significantly improve the reasoning accuracy and interpretability. However, existing models answer poorly for complex reasoning questions with attributes or relations, which causes false attribute selection or missing relation in Figure 1(a). It is because these models cannot balance all kinds of information in scene graphs, neglecting relation and attribute information. In this paper, we introduce a novel Dual Message-passing enhanced Graph Neural Net-work (DM-GNN), which can obtain a balanced represen-tation by properly encoding multi-scale scene graph infor-mation. Specifically, we (i) transform the scene graph into two graphs with diversified focuses on objects and relations; Then we design a dual structure to encode them, which in-creases the weights from relations (ii) fuse the encoder out-put with attribute features, which increases the weights from attributes; (iii) propose a message-passing mechanism to en-hance the information transfer between objects, relations and attributes. We conduct extensive experiments on datasets in-cluding GQA, VG, motif-VG and achieve new state of the art. Hao Li 0073, Xu Li 0001, Belhal Karimi, Jie Chen 0001, Mingming Sun 0001 |
ICME | 4 |
| 2022 | Efficient Algorithm and Hardware Architecture for Rate Estimation in Mode Decision of AVS3abstractTowards enabling advanced video coding for emerging ap-plications, the AVS3 standard has been developed recently, achieving twice the coding efficiency of the AVS2 stan-dard through complex coding tools including advanced rate-distortion optimization (RDO) to select the best mode. The bit-rates are produced with the Advanced-Entropy-Coding (AEC) in the RDO process of AVS3. However, AEC dom-inates the time complexity of RDO and among the steps, con-text updating and interval subdivision are performed recur-sively, which is not conducive to real-time application, espe-cially for the hardware implementation. Thus this paper pro-poses an adaptive rate estimation algorithm with a piece-wise linear function that is very friendly to hardware implemen-tation to accelerate the rate estimation process in the RDO for AVS3 practical applications. The proposed architecture can meet the requirement of 4K@120fps ultra-high-definition videos at 200 MHz, whereas the BD-Rate increases only by 0.67% under the All-Intra (AI) configuration. Yunyao Yan, Guoqing Xiang, Huizhu Jia, Yuan Li 0014, Peng Zhang 0007, Jie Chen 0001 |
ICME | 7 |
| 2022 | An Inclusive Task-Aware Framework for Radiology Report Generation
Lin Wang 0026, Munan Ning, Donghuan Lu, Dong Wei 0004, Yefeng Zheng 0001, Jie Chen 0001 |
MICCAI (8) | 6 |
| 2022 | Expectation-Maximization Contrastive Learning for Compact Video-and-Language RepresentationsabstractMost video-and-language representation learning approaches employ contrastive learning, e.g., CLIP, to project the video and text features into a common latent space according to the semantic similarities of text-video pairs. However, such learned shared latent spaces are not often optimal, and the modality gap between visual and textual representation can not be fully eliminated. In this paper, we propose Expectation-Maximization Contrastive Learning (EMCL) to learn compact video-and-language representations. Specifically, we use the Expectation-Maximization algorithm to find a compact set of bases for the latent space, where the features could be concisely represented as the linear combinations of these bases. Such feature decomposition of video-and-language representations reduces the rank of the latent space, resulting in increased representing power for the semantics. Extensive experiments on three benchmark text-video retrieval datasets prove that our EMCL can learn more discriminative video-and-language representations than previous methods, and significantly outperform previous state-of-the-art methods across all metrics. More encouragingly, the proposed method can be applied to boost the performance of existing approaches either as a jointly training layer or an out-of-the-box inference module with no extra training, making it easy to be incorporated into any existing methods. Peng Jin 0001, Jinfa Huang, Xian Wu 0001, Shen Ge, Guoli Song, David A. Clifton, Jie Chen 0001 |
NeurIPS | 8 |
| 2022 | OpenMedIA: Open-Source Medical Image Analysis Toolbox and Benchmark Under Heterogeneous AI Computing Platforms
Jiaxin Zhuang, Xiansong Huang, Yang Yang 0002, Jiancong Chen, Yue Yu 0001, Wei Gao 0003, Ge Li 0002, Jie Chen 0001, Tong Zhang 0017 |
PRCV (1) | 8 |
| 2022 | Dynamic Perception Framework for Fine-Grained RecognitionabstractFine-grained recognition poses the challenge of discriminating categories with only small subtle visual differences, which can be easily overwhelmed by diverse appearance within categories. Conventional approaches generally locate discriminative parts and then recognize the part-based features. However, we find that tuning the effective receptive field (ERF) of the network to the task plays the key role, which enables significant regions to contribute more to the output. Inspired by the receptive field stimulation mechanism of the visual cortex, we propose a Dynamic Perception framework as a solution. Our framework adapts the ERF by considering the image space and the kernel space simultaneously. In the image space, the Spatial Selective Sampling module is adopted to enlarge informative regions locally. In the kernel space, Spatial Selective Kernel convolution is introduced to adapt different kernel sizes for regions of interest and backgrounds by embedding spatial attention in the multi-path convolution. Extensive experiments on challenging benchmarks, including CUB-200-2011, FGVC-Aircraft, and Stanford Cars, demonstrate that our method yields a performance boost over the state-of-the-art methods. Yao Ding 0006, Zhenjun Han, Yanzhao Zhou, Yi Zhu 0004, Jie Chen 0001, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | RR-Net: Relation Reasoning for End-to-End Human-Object Interaction DetectionabstractThe task of Human-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects via inferring fine-grained triplets of$\langle $human, verb, object$\rangle $. Most HOI feature learning techniques are dependent on pre-detected instance regions or human body-part regions, which are computationally expensive and hardly applicable to end-to-end detectors in real applications. In this paper, based on an end-to-end HOI detector, we make a first try to explore region-independent relation reasoning for HOI detection. We first present a Relation-aware Frame, which brings a progressive structure for interaction inference. Upon the Relation-aware Frame, an Interaction Intensifier Module and a Correlation Parsing Module are carefully designed, where: a) interactive semantics from humans can be exploited and passed to objects to intensify interactions, b) interactive correlations among humans, objects and interactions are integrated to promote predictions. Based on modules above, we construct a fully differentiable and end-to-end trainable network named Relation Reasoning Network (abbr. RR-Net). Extensive experiments show that our proposed RR-Net leads to competitive results compared with the state-of-the-art methods on both V-COCO and HICO-DET benchmarks and improves the baseline about 7.6% and 11.1% relatively, validating that this first effort in exploring region-independent relation reasoning has brought obvious improvement for end-to-end HOI detection. Dongming Yang, Yuexian Zou, Can Zhang 0001, Meng Cao 0002, Jie Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | MDAN: Mirror Difference Aware Network for Brain Stroke Lesion SegmentationabstractBrain stroke lesion segmentation is of great importance for stroke rehabilitation neuroimaging analysis. Due to the large variance of stroke lesion shapes and similarities of tissue intensity distribution, it remains a challenging task. To help detect abnormalities, the anatomical symmetries of brain magnetic resonance (MR) images have been widely used as visual cues for clinical practices. However, most methods for brain images segmentation do not fully utilize structural symmetry information. This paper presents a novel mirror difference aware network (MDAN) for stroke lesion segmentation. The network uses an encoder-decoder architecture, aiming at holistically exploiting the symmetries of image features. Specifically, a differential feature augmentation (DFA) module is developed in the encoding path to highlight the semantically pathological asymmetries of features in abnormalities. In the DFA module, a Siamese contrastive supervised loss is designed to enhance discriminative features, and a mirror position-based difference augmentation (MDA) module is used to further magnify the discrepancy. Moreover, mirror feature fusion (MFF) modules are applied to efficiently fuse and transfer the information both of the original input and the horizontally flipped features to the decoding path. Extensive experiments on the Anatomical Tracings of Lesions After Stroke (ATLAS) dataset show the proposed MDAN outperforms the state-of-the-art methods. Qiqi Bao 0001, Shiyu Mi, Bowen Gang, Wenming Yang, Jie Chen 0001, Qingmin Liao |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Memory Attention Networks for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition has been extensively studied, but it remains an unsolved problem because of the complex variations of skeleton joints in 3-D spatiotemporal space. To handle this issue, we propose a newly temporal-then-spatial recalibration method named memory attention networks (MANs) and deploy MANs using the temporal attention recalibration module (TARM) and spatiotemporal convolution module (STCM). In the TARM, a novel temporal attention mechanism is built based on residual learning to recalibrate frames of skeleton data temporally. In the STCM, the recalibrated sequence is transformed or encoded as the input of CNNs to further model the spatiotemporal information of skeleton sequence. Based on MANs, a new collaborative memory fusion module (CMFM) is proposed to further improve the efficiency, leading to the collaborative MANs (C-MANs), trained with two streams of base MANs. TARM, STCM, and CMFM form a single network seamlessly and enable the whole network to be trained in an end-to-end fashion. Comparing with the state-of-the-art methods, MANs and C-MANs improve the performance significantly and achieve the best results on six data sets for action recognition. The source code has been made publicly available at https://github.com/memory-attention-networks. Ce Li 0002, Chunyu Xie, Baochang Zhang 0001, Jungong Han, Xiantong Zhen, Jie Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | Discover Cross-Modality Nuances for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (Re-ID) aims to match the pedestrian images of the same identity from different modalities. Existing works mainly focus on alleviating the modality discrepancy by aligning the distributions of features from different modalities. However, nuanced but discriminative information, such as glasses, shoes, and the length of clothes, has not been fully explored, especially in the infrared modality. Without discovering nuances, it is challenging to match pedestrians across modalities using modality alignment solely, which inevitably reduces feature distinctiveness. In this paper, we propose a joint Modality and Pattern Alignment Network (MPANet) to discover cross-modality nuances in different patterns for visible-infrared person Re-ID, which introduces a modality alleviation module and a pattern alignment module to jointly extract discriminative features. Specifically, we first propose a modality alleviation module to dislodge the modality information from the extracted feature maps. Then, We devise a pattern alignment module, which generates multiple pattern maps for the diverse patterns of a person, to discover nuances. Finally, we introduce a mutual mean learning fashion to alleviate the modality discrepancy and propose a center cluster loss to guide both identity learning and nuances discovering. Extensive experiments on the public SYSU-MM01 and RegDB datasets demonstrate the superiority of MPANet over state-of-the-arts. Qiong Wu 0012, Pingyang Dai, Jie Chen 0001, Chia-Wen Lin, Yongjian Wu 0001, Feiyue Huang, Bineng Zhong 0001, Rongrong Ji |
CVPR | 3 |
| 2021 | CoLA: Weakly-Supervised Temporal Action Localization With Snippet Contrastive LearningabstractWeakly-supervised temporal action localization (WS-TAL) aims to localize actions in untrimmed videos with only video-level labels. Most existing models follow the "localization by classification" procedure: locate temporal regions contributing most to the video-level classification. Generally, they process each snippet (or frame) individually and thus overlook the fruitful temporal context relation. Here arises the single snippet cheating issue: "hard" snippets are too vague to be classified. In this paper, we argue that learning by comparing helps identify these hard snip-pets and we propose to utilize snippet Contrastive learning to Localize Actions, CoLA for short. Specifically, we propose a Snippet Contrast (SniCo) Loss to refine the hard snippet representation in feature space, which guides the network to perceive precise temporal boundaries and avoid the temporal interval interruption. Besides, since it is in-feasible to access frame-level annotations, we introduce a Hard Snippet Mining algorithm to locate the potential hard snippets. Substantial analyses verify that this mining strategy efficaciously captures the hard snippets and SniCo Loss leads to more informative feature representation. Extensive experiments show that CoLA achieves state-of-the-art results on THUMOS’14 and ActivityNet v1.2 datasets. Can Zhang 0001, Meng Cao 0002, Dongming Yang, Jie Chen 0001, Yuexian Zou |
CVPR | 4 |
| 2021 | CDNet: Centripetal Direction Network for Nuclear Instance SegmentationabstractNuclear instance segmentation is a challenging task due to a large number of touching and overlapping nuclei in pathological images. Existing methods cannot effectively recognize the accurate boundary owing to neglecting the relationship between pixels (e.g., direction information). In this paper, we propose a novel Centripetal Direction Net-work (CDNet) for nuclear instance segmentation. Specifically, we define centripetal direction feature as a class of adjacent directions pointing to the nuclear center to rep-resent the spatial relationship between pixels within the nucleus. These direction features are then used to construct a direction difference map to represent the similarity within instances and the differences between instances. Finally, we propose a direction-guided refinement module, which acts as a plug-and-play module to effectively integrate auxiliary tasks and aggregate the features of different branches. Experiments on MoNuSeg and CPM17 datasets show that CDNet is significantly better than the other methods and achieves the state-of-the-art performance. The code is available at https://github.com/honglianghe/CDNet. Yao Ding 0006, Guoli Song, Lin Wang 0026, Qian Ren, Pengxu Wei, Jie Chen 0001 |
ICCV | 9 |
| 2021 | ReCU: Reviving the Dead Weights in Binary Neural Networks
Mingbao Lin, Jianzhuang Liu, Jie Chen 0001, Ling Shao 0001, Yue Gao 0002, Yonghong Tian 0001, Rongrong Ji |
ICCV | 4 |
| 2021 | RR-Net: Injecting Interactive Semantics in Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects. Latest end-to-end HOI detectors are short of relation reasoning, which leads to inability to learn HOI-specific interactive semantics for predictions. In this paper, we therefore propose novel relation reasoning for HOI detection. We first present a progressive Relation-aware Frame, which brings a new structure and parameter sharing pattern for interaction inference. Upon the frame, an Interaction Intensifier Module and a Correlation Parsing Module are carefully designed, where: a) interactive semantics from humans can be exploited and passed to objects to intensify interactions, b) interactive correlations among humans, objects and interactions are integrated to promote predictions. Based on modules above, we construct an end-to-end trainable framework named Relation Reasoning Network (abbr. RR-Net). Extensive experiments show that our proposed RR-Net sets a new state-of-the-art on both V-COCO and HICO-DET benchmarks and improves the baseline about 5.5% and 9.8% relatively, validating that this first effort in exploring relation reasoning and integrating interactive semantics has brought obvious improvement for end-to-end HOI detection. Dongming Yang, Yuexian Zou, Can Zhang 0001, Meng Cao 0002, Jie Chen 0001 |
IJCAI | 5 |
| 2021 | Learnable Oriented-Derivative Network for Polyp Segmentation
Mengjun Cheng, Zishang Kong, Guoli Song, Yonghong Tian 0001, Yongsheng Liang 0001, Jie Chen 0001 |
MICCAI (1) | 6 |
| 2021 | Few-shot Learning for Multi-Modality TasksabstractRecent deep learning methods rely on a large amount of labeled data to achieve high performance. These methods may be impractical in some scenarios, where manual data annotation is costly or the samples of certain categories are scarce (e.g., tumor lesions, endangered animals and rare individual activities). When only limited annotated samples are available, these methods usually suffer from the overfitting problem severely, which degrades the performance significantly. In contrast, humans can recognize the objects in the images rapidly and correctly with their prior knowledge after exposed to only a few annotated samples. To simulate the learning schema of humans and relieve the reliance on the large-scale annotation benchmarks, researchers start shifting towards the few-shot learning problem: they try to learn a model to correctly recognize novel categories with only a few annotated samples. Jie Chen 0001, Qixiang Ye, Xiaoshan Yang, Shaohua Kevin Zhou, Xiaopeng Hong, Li Zhang 0040 |
ACM Multimedia | 1 |
| 2021 | Hard-Boundary Attention Network for Nuclei Instance SegmentationabstractImage segmentation plays an important role in medical image analysis, and accurate segmentation of nuclei is especially crucial to clinical diagnosis. However, existing methods fail to segment dense nuclei due to the hard-boundary which has similar texture to nuclear inside. To this end, we propose a Hard-Boundary Attention Network (HBANet) for nuclei instance segmentation. Specifically, we propose a Background Weaken Module (BWM) to weaken the attention of our model to the nucleus background by integrating low-level features into high-level features. To improve the robustness of the model to the hard-boundary of nuclei, we further design a Gradient-based boundary adaptive Strategy (GS) which generates boundary-weakened data for model training in an adversarial manner. We conduct extensive experiments on MoNuSeg and CPM-17 datasets, and experimental results show that our HBANet outperforms the state-of-the-art methods. Yalu Cheng, Pengchong Qiao, Guoli Song, Jie Chen 0001 |
MMAsia | 5 |
| 2021 | MIGO-NAS: Towards Fast and Generalizable Neural Architecture SearchabstractNeural architecture search (NAS) has achieved unprecedented performance in various computer vision tasks. However, most existing NAS methods are defected in search efficiency and model generalizability. In this paper, we propose a novel NAS framework, termed MIGO-NAS, with the aim to guarantee the efficiency and generalizability in arbitrary search spaces. On the one hand, we formulate the search space as a multivariate probabilistic distribution, which is then optimized by a novel multivariate information-geometric optimization (MIGO). By approximating the distribution with a sampling, training, and testing pipeline, MIGO guarantees the memory efficiency, training efficiency, and search flexibility. Besides, MIGO is the first time to decrease the estimation error of natural gradient in multivariate distribution. On the other hand, for a set of specific constraints, the neural architectures are generated by a novel dynamic programming network generation (DPNG), which significantly reduces the training cost under various hardware environments. Experiments validate the advantages of our approach over existing methods by establishing a superior accuracy and efficiency i.e., 2.39 test error on CIFAR-10 benchmark and 21.7 on ImageNet benchmark, with only 1.5 GPU hours and 96 GPU hours for searching, respectively. Besides, the searched architectures can be well generalize to computer vision tasks including object detection and semantic segmentation, i.e., 25× FLOPs compression, with 6.4 mAP gain over Pascal VOC dataset, and 29.9× FLOPs compression, with only 1.41 percent performance drop over Cityscapes dataset. The code is publicly available. Xiawu Zheng, Rongrong Ji, Qiang Wang 0060, Baochang Zhang 0001, Jie Chen 0001, Qixiang Ye, Feiyue Huang, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Joint segmentation and detection of COVID-19 via a sequential region generation network
Jipeng Wu, Shengchuan Zhang, Xi Li 0011, Jie Chen 0001, Jiawen Zheng, Yue Gao 0002, Yonghong Tian 0001, Yongsheng Liang 0001, Rongrong Ji |
Pattern Recognit. | 5 |
| 2021 | Semi-Supervised Natural Face De-OcclusionabstractOcclusions are often present in face images in the wild, e.g., under video surveillance and forensic scenarios. Existing face de-occlusion methods are limited as they require the knowledge of an occlusion mask. To overcome this limitation, we propose in this paper a new generative adversarial network (named OA-GAN) for natural face de-occlusion without an occlusion mask, enabled by learning in a semi-supervised fashion using (i) paired images with known masks of artificial occlusions and (ii) natural images without occlusion masks. The generator of our approach first predicts an occlusion mask, which is used for filtering the feature maps of the input image as a semantic cue for de-occlusion. The filtered feature maps are then used for face completion to recover a non-occluded face image. The initial occlusion mask prediction might not be accurate enough, but it gradually converges to the accurate one because of the adversarial loss we use to perceive which regions in a face image need to be recovered. The discriminator of our approach consists of an adversarial loss, distinguishing the recovered face images from natural face images, and an attribute preserving loss, ensuring that the face image after de-occlusion can retain the attributes of the input face image. Experimental evaluations on the widely used CelebA dataset and a dataset with natural occlusions we collected show that the proposed approach can outperform the state of the art methods in natural face de-occlusion. Jiancheng Cai, Hu Han 0001, Jiyun Cui, Jie Chen 0001, Li Liu 0002, Shaohua Kevin Zhou |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | SRN: Side-Output Residual Network for Object Reflection Symmetry Detection and BeyondabstractThis article establishes a baseline for object reflection symmetry detection in natural images by releasing a new benchmark named Sym-PASCAL and proposing an end-to-end deep learning approach for reflection symmetry. Sym-PASCAL spans challenges of multiobjects, object diversity, part invisibility, and clustered backgrounds, which is far beyond those in existing data sets. The end-to-end deep learning approach, referred to as a side-output residual network (SRN), leverages the output residual units (RUs) to fit the errors between the symmetry ground truth and the side outputs of multiple stages of a trunk network. By cascading RUs from deep to shallow, SRN exploits the "flow" of errors along multiple stages to effectively matching object symmetry at different scales and suppress the clustered backgrounds. SRN is interpreted as a boosting-like algorithm, which assembles features using RUs during network forward and backward propagations. SRN is further upgraded to a multitask SRN (MT-SRN) for joint symmetry and edge detection, demonstrating its generality to image-to-mask learning tasks. Experimental results verify that the Sym-PASCAL benchmark is challenging related to real-world images, SRN achieves state-of-the-art performance, and MT-SRN has the capability to simultaneously predict edge and symmetry mask without loss of performance. Wei Ke 0003, Jie Chen 0001, Jianbin Jiao, Guoying Zhao 0001, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | AD-Cluster: Augmented Discriminative Clustering for Domain Adaptive Person Re-IdentificationabstractDomain adaptive person re-identification (re-ID) is a challenging task, especially when person identities in target domains are unknown. Existing methods attempt to address this challenge by transferring image styles or aligning feature distributions across domains, whereas the rich unlabeled samples in target domains are not sufficiently exploited. This paper presents a novel augmented discriminative clustering (AD-Cluster) technique that estimates and augments person clusters in target domains and enforces the discrimination ability of re-ID models with the augmented clusters. AD-Cluster is trained by iterative density-based clustering, adaptive sample augmentation, and discriminative feature learning. It learns an image generator and a feature encoder which aim to maximize the intra-cluster diversity in the sample space and minimize the intra-cluster distance in the feature space in an adversarial min-max manner. Finally, AD-Cluster increases the diversity of sample clusters and improves the discrimination capability of re-ID models greatly. Extensive experiments over Market-1501 and DukeMTMC-reID show that AD-Cluster outperforms the state-of-the-art with large margins. Yunpeng Zhai, Shijian Lu, Qixiang Ye, Xuebo Shan, Jie Chen 0001, Rongrong Ji, Yonghong Tian 0001 |
CVPR | 5 |
| 2020 | Exploring Entity-Level Spatial Relationships for Image-Text MatchingabstractExploring the entity-level (i.e., objects in an image, words in a text) spatial relationship contributes to understanding multimedia content precisely. The ignorance of spatial information in previous works probably leads to misunderstandings of image contents. For instance, sentences `Boats are on the water' and `Boats are under the water' describe the same objects, but correspond to different sceneries. To this end, we utilize the relative position of objects to capture entity-level spatial relationships for image-text matching. Specifically, we fuse semantic and spatial relationships of image objects in a visual intra-modal relation module. The module performs promisingly to understand image contents and improve object representation learning. It contributes to capturing entity-level latent correspondence of image-text pairs. Then the query (text) plays a role of textual context to refine the interpretable alignments of image-text pairs in the inter-modal relation module. Our proposed method achieves state-of-the-art results on MSCOCO and Flickr30K datasets. Yaxian Xia, Lun Huang, Wenmin Wang 0001, Xiaoyong Wei, Jie Chen 0001 |
ICASSP | 5 |
| 2020 | BCData: A Large-Scale Dataset and Benchmark for Cell Detection and Counting
Yao Ding 0006, Guoli Song, Lin Wang 0026, Ruizhe Geng, Yonghong Tian 0001, Yongsheng Liang 0001, Shaohua Kevin Zhou, Jie Chen 0001 |
MICCAI (5) | 12 |
| 2020 | An automated method with anchor-free detection and U-shaped segmentation for nuclei instance segmentationabstractNuclei segmentation plays an important role in cancer diagnosis. Automated methods for digital pathology become popular due to the developments of deep learning and neural networks. However, this task still faces challenges. Most of current techniques cannot be applied directly because of the clustered state and the large number of nuclei in images. Moreover, anchor-based methods for object detection lead a huge amount of calculation, which is even worse on pathological images with a large target density. To address these issues, we propose a novel network with an anchor-free detection and a U-shaped segmentation. An altered feature enhancement module is attached to improve the performance in dense target detection. Meanwhile, the U-Shaped structure in segmentation block ensures the aggregation of features in different dimensions generated from the backbone network. We evaluate our work on a Multi-Organ Nuclei Segmentation dataset from MICCAI 2018 challenge. In comparisons with others, our proposed method achieves state-of-the-art performance. Lijuan Duan, Jie Chen 0001 |
MMAsia | 3 |
| 2020 | Adaptive feature aggregation network for nuclei segmentationabstractNuclei instance segmentation is essential for cell morphometrics and analysis, playing a crucial role in digital pathology. The problem of variability in nuclei characteristics among diverse cell types makes this task more challenging. Recently, proposal-based segmentation methods with feature pyramid network (FPN) has shown good performance because FPN integrates multi-scale features with strong semantics. However, FPN has information loss of the highest-level feature map and sub-optimal feature fusion strategies. This paper proposes a proposal-based adaptive feature aggregation methods (AANet) to make full use of multi-scale features. Specifically, AANet consists of two components: Context Augmentation Module (CAM) and Feature Adaptive Selection Module (ASM). In feature fusion, CAM focus on exploring extensive contextual information and capturing discriminative semantics to reduce the information loss of feature map at the highest pyramid level. The enhanced features are then sent to ASM to get a combined feature representation adaptively over all feature levels for each RoI. The experiments show our model's effectiveness on two publicly available datasets: the Kaggle 2018 Data Science Bowl dataset and the Multi-Organ nuclei segmentation dataset. Ruizhe Geng, Jie Chen 0001 |
MMAsia | 3 |
| 2020 | An Automated Method with Feature Pyramid Encoder and Dual-Path Decoder for Nuclei Segmentation
Lijuan Duan, Jie Chen 0001 |
PRCV (1) | 3 |
| 2020 | Deep Learning for Generic Object Detection: A SurveyabstractAbstract Object detection, one of the most fundamental and challenging problems in computer vision, seeks to locate object instances from a large number of predefined categories in natural images. Deep learning techniques have emerged as a powerful strategy for learning feature representations directly from data and have led to remarkable breakthroughs in the field of generic object detection. Given this period of rapid evolution, the goal of this paper is to provide a comprehensive survey of the recent achievements in this field brought about by deep learning techniques. More than 300 research contributions are included in this survey, covering many aspects of generic object detection: detection frameworks, object feature representation, object proposal generation, context modeling, training strategies, and evaluation metrics. We finish the survey by identifying promising directions for future research. Li Liu 0002, Wanli Ouyang, Xiaogang Wang 0001, Paul W. Fieguth, Jie Chen 0001, Xinwang Liu 0002, Matti Pietikäinen |
Int. J. Comput. Vis. | 5 |
| 2019 | Attention on Attention for Image CaptioningabstractAttention mechanisms are widely used in current encoder/decoder frameworks of image captioning, where a weighted average on encoded vectors is generated at each time step to guide the caption decoding process. However, the decoder has little idea of whether or how well the attended vector and the given attention query are related, which could make the decoder give misled results. In this paper, we propose an Attention on Attention (AoA) module, which extends the conventional attention mechanisms to determine the relevance between attention results and queries. AoA first generates an information vector and an attention gate using the attention result and the current context, then adds another attention by applying element-wise multiplication to them and finally obtains the attended information, the expected useful knowledge. We apply AoA to both the encoder and the decoder of our image captioning model, which we name as AoA Network (AoANet). Experiments show that AoANet outperforms all previously published methods and achieves a new state-of-the-art performance of 129.8 CIDEr-D score on MS COCO Karpathy offline test split and 129.6 CIDEr-D (C40) score on the official online testing server. Code is available at https://github.com/husthuaan/AoANet. Lun Huang, Wenmin Wang 0001, Jie Chen 0001, Xiaoyong Wei |
ICCV | 3 |
| 2019 | Tumor Tissue Segmentation for Histopathological ImagesabstractHistopathological image analysis is considered as a gold standard for cancer identification and diagnosis. Tumor segmentation for histopathological images is one of the most important research topics and its performance directly affects the diagnosis judgment of doctors for cancer categories and their periods. With the remarkable development of deep learning methods, extensive methods have been proposed for tumor segmentation. However, there are few researches on analysis of specific pipeline of tumor segmentation. Moreover, few studies have done detailed research on the hard example mining of tumor segmentation. In order to bridge this gap, this study firstly summarize a specific pipeline of tumor segmentation. Then, hard example mining in tumor segmentation is also explored. Finally, experiments are conducted for evaluating segmentation performance of our method, demonstrating the effects of our method and hard example mining. Xiansong Huang, Pengxu Wei, Juncen Zhang, Jie Chen 0001 |
MMAsia | 6 |
| 2019 | Adaptively Aligned Image Captioning via Adaptive Attention TimeabstractRecent neural models for image captioning usually employ an encoder-decoder framework with an attention mechanism. However, the attention mechanism in such a framework aligns one single (attended) image feature vector to one caption word, assuming one-to-one mapping from source image regions and target caption words, which is never possible. In this paper, we propose a novel attention model, namely Adaptive Attention Time (AAT), to align the source and the target adaptively for image captioning. AAT allows the framework to learn how many attention steps to take to output a caption word at each decoding step. With AAT, an image region can be mapped to an arbitrary number of caption words while a caption word can also attend to an arbitrary number of image regions. AAT is deterministic and differentiable, and doesn't introduce any noise to the parameter gradients. In this paper, we empirically show that AAT improves over state-of-the-art methods on the task of image captioning. Code is available at https://github.com/husthuaan/AAT. Lun Huang, Wenmin Wang 0001, Yaxian Xia, Jie Chen 0001 |
NeurIPS | 4 |
| 2019 | From BoW to CNN: Two Decades of Texture Representation for Texture ClassificationabstractTexture is a fundamental characteristic of many types of images, and texture representation is one of the essential and challenging problems in computer vision and pattern recognition which has attracted extensive research attention over several decades. Since 2000, texture representations based on Bag of Words and on Convolutional Neural Networks have been extensively studied with impressive performance. Given this period of remarkable evolution, this paper aims to present a comprehensive survey of advances in texture representation over the last two decades. More than 250 major publications are cited in this survey covering different aspects of the research, including benchmark datasets and state of the art results. In retrospect of what has been achieved so far, the survey discusses open challenges and directions for future research. Li Liu 0002, Jie Chen 0001, Paul W. Fieguth, Guoying Zhao 0001, Rama Chellappa, Matti Pietikäinen |
Int. J. Comput. Vis. | 2 |
| 2019 | Guest Editors' Introduction to the Special Section on Compact and Efficient Feature Representation and Learning in Computer VisionabstractThe papers in this special section examine compact and efficient feature representation and learning in computer vision. Li Liu 0002, Matti Pietikäinen, Jie Chen 0001, Guoying Zhao 0001, Xiaogang Wang 0001, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Deep contour and symmetry scored object proposal
Wei Ke 0003, Jie Chen 0001, Qixiang Ye |
Pattern Recognit. Lett. | 2 |
| 2019 | Texture Classification in Extreme Scale Variations Using GANetabstractResearch in texture recognition often concentrates on recognizing textures with intraclass variations, such as illumination, rotation, viewpoint, and small-scale changes. In contrast, in real-world applications, a change in scale can have a dramatic impact on texture appearance to the point of changing completely from one texture category to another. As a result, texture variations due to changes in scale are among the hardest to handle. In this paper, we conduct the first study of classifying textures with extreme variations in scale. To address this issue, we first propose and then reduce scale proposals on the basis of dominant texture patterns. Motivated by the challenges posed by this problem, we propose a new GANet network where we use a genetic algorithm to change the filters in the hidden layers during network training in order to promote the learning of more informative semantic texture patterns. Finally, we adopt a Fisher vector pooling of a convolutional neural network filter bank feature encoder for global texture representation. Because extreme scale variations are not necessarily present in most standard texture databases, to support the proposed extreme-scale aspects of texture understanding, we are developing a new dataset, the extreme scale variation textures (ESVaT), to test the performance of our framework. It is demonstrated that the proposed framework significantly outperforms the gold-standard texture features by more than 10% on ESVaT. We also test the performance of our proposed approach on the KTHTIPS2b and OS datasets and a further dataset synthetically derived from Forrest, showing the superior performance compared with the state-of-the-art. Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Paul W. Fieguth, Xilin Chen 0001, Matti Pietikäinen |
IEEE Trans. Image Process. | 2 |
| 2019 | Saliency Integration: An Arbitrator ModelabstractSaliency integration has attracted much attention on unifying saliency maps from multiple saliency models. Previous offline integration methods usually face two challenges: 1) if most of the candidate saliency models misjudge the saliency on an image, the integration result will lean heavily on those inferior candidate models; and 2) an unawareness of the ground truth saliency labels brings difficulty in estimating the expertise of each candidate model. To address these problems, in this paper, we propose an arbitrator model (AM) for saliency integration. First, we incorporate the consensus of multiple saliency models and the external knowledge into a reference map to effectively rectify the misleading by candidate models. Second, our quest for ways of estimating the expertise of the saliency models without ground truth labels gives rise to two distinct online model-expertise estimation methods. Finally, we derive a Bayesian integration framework to reconcile the saliency models of varying expertise and the reference map. To extensively evaluate the proposed AM model, we test 27 state-of-the-art saliency models, covering both traditional and deep learning ones, on various combinations over four datasets. The evaluation results show that the AM model improves the performance substantially compared to the existing state-of-the-art integration methods, regardless of the chosen candidate saliency models. Yingyue Xu, Xiaopeng Hong, Fatih Porikli, Xin Liu 0012, Jie Chen 0001, Guoying Zhao 0001 |
IEEE Trans. Multim. | 5 |
| 2018 | Super Wide Regression Network for Unsupervised Cross-Database Facial Expression RecognitionabstractUnsupervised cross-database facial expression recognition (FER) is a challenging problem, in which the training and testing samples belong to different facial expression databases. For this reason, the training (source) and testing (target) facial expression samples would have different feature distributions and hence the performance of lots of existing FER methods may decrease. To solve this problem, in this paper we propose a novel super wide regression network (SWiRN) model, which serves as the regression parameter to bridge the original feature space and the label space and herein in each layer the maximum mean discrepancy (MMD) criterion is used to enforce the source and target facial expression samples to share the same or similar feature distributions. Consequently, the learned SWiRN is able to predict the expression categories of the target samples although we have no access to any label information of target samples. We conduct extensive cross-database FER experiments on CK+, eNTERFACE, and Oulu-CASIA VIS facial expression databases to evaluate the proposed SWiRN. Experimental results show that our SWiRN model achieves more promising performance than recent proposed cross-database emotion recognition methods. Baofeng Zhang, Yuan Zong, Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Junchao Zhu |
ICASSP | 5 |
| 2018 | Unsupervised Cross-Corpus Speech Emotion Recognition Using Domain-Adaptive Subspace LearningabstractIn this paper, we investigate an interesting problem, i.e., unsupervised cross-corpus speech emotion recognition (SER), in which the training and testing speech signals come from two different speech emotion corpora. Meanwhile, the training speech signals are labeled, while the label information of the testing speech signals is entirely unknown. Due to this setting, the training (source) and testing (target) speech signals may have different feature distributions and therefore lots of existing SER methods would not work. To deal with this problem, we propose a domain-adaptive subspace learning (DoSL) method for learning a projection matrix with which we can transform the source and target speech signals from the original feature space to the label space. The transformed source and target speech signals in the label space would have similar feature distributions. Consequently, the classifier learned on the labeled source speech signals can effectively predict the emotional states of the unlabeled target speech signals. To evaluate the performance of the proposed DoSL method, we carry out extensive cross-corpus SER experiments on three speech emotion corpora including EmoDB, eNTERFACE, and AFEW 4.0. Compared with recent state-of-the-art cross-corpus SER methods, the proposed DoSL can achieve more satisfactory overall results. Yuan Zong, Baofeng Zhang, Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Junchao Zhu |
ICASSP | 5 |
| 2017 | SRN: Side-Output Residual Network for Object Symmetry Detection in the WildabstractIn this paper, we establish a baseline for object symmetry detection in complex backgrounds by presenting a new benchmark and an end-to-end deep learning approach, opening up a promising direction for symmetry detection in the wild. The new benchmark, named Sym-PASCAL, spans challenges including object diversity, multi-objects, part-invisibility, and various complex backgrounds that are far beyond those in existing datasets. The proposed symmetry detection approach, named Side-output Residual Network (SRN), leverages output Residual Units (RUs) to fit the errors between the object symmetry ground-truth and the outputs of RUs. By stacking RUs in a deep-to-shallow manner, SRN exploits the flow of errors among multiple scales to ease the problems of fitting complex outputs with limited layers, suppressing the complex backgrounds, and effectively matching object symmetry of different scales. Experimental results validate both the benchmark and its challenging aspects related to real-world images, and the state-of-the-art performance of our symmetry detection approach. The benchmark and the code for SRN are publicly available at https://github.com/KevinKecc/SRN. Wei Ke 0003, Jie Chen 0001, Jianbin Jiao, Guoying Zhao 0001, Qixiang Ye |
CVPR | 2 |
| 2017 | Self-Learning Scene-Specific Pedestrian Detectors Using a Progressive Latent ModelabstractIn this paper, a self-learning approach is proposed towards solving scene-specific pedestrian detection problem without any human annotation involved. The self-learning approach is deployed as progressive steps of object discovery, object enforcement, and label propagation. In the learning procedure, object locations in each frame are treated as latent variables that are solved with a progressive latent model (PLM). Compared with conventional latent models, the proposed PLM incorporates a spatial regularization term to reduce ambiguities in object proposals and to enforce object localization, and also a graph-based label propagation to discover harder instances in adjacent frames. With the difference of convex (DC) objective functions, PLM can be efficiently optimized with a concave-convex programming and thus guaranteeing the stability of self-learning. Extensive experiments demonstrate that even without annotation the proposed self-learning approach outperforms weakly supervised learning approaches, while achieving comparable performance with transfer learning and fully supervised approaches. Qixiang Ye, Tianliang Zhang 0003, Wei Ke 0003, Qiang Qiu 0001, Jie Chen 0001, Guillermo Sapiro, Baochang Zhang 0001 |
CVPR | 5 |
| 2017 | Robust local features for remote face recognition
Jie Chen 0001, Vishal M. Patel, Li Liu 0002, Vili Kellokumpu, Guoying Zhao 0001, Matti Pietikäinen, Rama Chellappa |
Image Vis. Comput. | 1 |
| 2017 | Hierarchical Contour Closure-Based Holistic Salient Object DetectionabstractMost existing salient object detection methods compute the saliency for pixels, patches, or superpixels by contrast. Such fine-grained contrast-based salient object detection methods are stuck with saliency attenuation of the salient object and saliency overestimation of the background when the image is complicated. To better compute the saliency for complicated images, we propose a hierarchical contour closure-based holistic salient object detection method, in which two saliency cues, i.e., closure completeness and closure reliability, are thoroughly exploited. The former pops out the holistic homogeneous regions bounded by completely closed outer contours, and the latter highlights the holistic homogeneous regions bounded by averagely highly reliable outer contours. Accordingly, we propose two computational schemes to compute the corresponding saliency maps in a hierarchical segmentation space. Finally, we propose a framework to combine the two saliency maps, obtaining the final saliency map. Experimental results on three publicly available datasets show that even each single saliency map is able to reach the state-of-the-art performance. Furthermore, our framework, which combines two saliency maps, outperforms the state of the arts. Additionally, we show that the proposed framework can be easily used to extend existing methods and further improve their performances substantially. Qing Liu 0003, Xiaopeng Hong, Beiji Zou 0001, Jie Chen 0001, Zailiang Chen 0001, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | RoLoD: Robust local descriptors for computer vision
Jie Chen 0001, Zhen Lei 0001, Li Liu 0002, Guoying Zhao 0001, Matti Pietikäinen |
Neurocomputing | 1 |
| 2016 | Exploring illumination robust descriptors for human epithelial type 2 cell classification
Xianbiao Qi, Guoying Zhao 0001, Jie Chen 0001, Matti Pietikäinen |
Pattern Recognit. | 3 |
| 2016 | HEp-2 cell classification: The role of Gaussian Scale Space Theory as a pre-processing approach
Xianbiao Qi, Guoying Zhao 0001, Jie Chen 0001, Matti Pietikäinen |
Pattern Recognit. Lett. | 3 |
| 2014 | Remote Heart Rate Measurement from Face Videos under Realistic SituationsabstractHeart rate is an important indicator of people's physiological state. Recently, several papers reported methods to measure heart rate remotely from face videos. Those methods work well on stationary subjects under well controlled conditions, but their performance significantly degrades if the videos are recorded under more challenging conditions, specifically when subjects' motions and illumination variations are involved. We propose a framework which utilizes face tracking and Normalized Least Mean Square adaptive filtering methods to counter their influences. We test our framework on a large difficult and public database MAHNOB-HCI and demonstrate that our method substantially outperforms all previous methods. We also use our method for long term heart rate monitoring in a game evaluation scenario and achieve promising results. Jie Chen 0001, Guoying Zhao 0001, Matti Pietikäinen |
CVPR | 2 |
| 2013 | RLBP: Robust Local Binary PatternabstractIn this paper, we propose a simple and robust local descriptor, called the robust local binary pattern (RLBP). The local binary pattern (LBP) works very successfully in many domains, such as texture classification, human detection and face recognition. However, an issue of LBP is that it is not so robust to the noise present in the image. We improve the robustness of LBP by changing the coding bit of LBP. Experimental results on the Brodatz and UIUC texture databases show that RLBP impressively outperforms the other widely used descriptors (e.g., SIFT, Gabor, MR8 and LBP) and other variants of LBP (e.g., completed LBP), especially when we add noise in the images. In addition, experimental results on human face recognition also show a promising performance comparable to the best known results on the Face Recognition Grand Challenge (FRGC) face dataset. Jie Chen 0001, Vili Kellokumpu, Guoying Zhao 0001, Matti Pietikäinen |
BMVC | 1 |
| 2013 | Large Margin Subspace Learning for feature selection
Bin Fang 0001, Xinwang Liu 0002, Jie Chen 0001, Zhenghong Huang, Xiping He |
Pattern Recognit. | 4 |
| 2013 | Automatic Dynamic Texture Segmentation Using Local Descriptors and Optical FlowabstractA dynamic texture (DT) is an extension of the texture to the temporal domain. How to segment a DT is a challenging problem. In this paper, we address the problem of segmenting a DT into disjoint regions. A DT might be different from its spatial mode (i.e., appearance) and/or temporal mode (i.e., motion field). To this end, we develop a framework based on the appearance and motion modes. For the appearance mode, we use a new local spatial texture descriptor to describe the spatial mode of the DT; for the motion mode, we use the optical flow and the local temporal texture descriptor to represent the temporal variations of the DT. In addition, for the optical flow, we use the histogram of oriented optical flow (HOOF) to organize them. To compute the distance between two HOOFs, we develop a simple effective and efficient distance measure based on Weber's law. Furthermore, we also address the problem of threshold selection by proposing a method for determining thresholds for the segmentation method by an offline supervised statistical learning. The experimental results show that our method provides very good segmentation results compared to the state-of-the-art methods in segmenting regions that differ in their dynamics. Jie Chen 0001, Guoying Zhao 0001, Mikko Salo, Esa Rahtu, Matti Pietikäinen |
IEEE Trans. Image Process. | 1 |
| 2012 | Unsupervised dynamic texture segmentation using local descriptors in volumes
Jie Chen 0001, Guoying Zhao 0001, Matti Pietikäinen |
ICPR | 1 |
| 2012 | Evaluation of local feature descriptors and their combination for pedestrian representation
Jixiang Liang, Qixiang Ye, Jie Chen 0001, Jianbin Jiao |
ICPR | 3 |
| 2011 | Maximal Linear Embedding for Dimensionality ReductionabstractOver the past few decades, dimensionality reduction has been widely exploited in computer vision and pattern analysis. This paper proposes a simple but effective nonlinear dimensionality reduction algorithm, named Maximal Linear Embedding (MLE). MLE learns a parametric mapping to recover a single global low-dimensional coordinate space and yields an isometric embedding for the manifold. Inspired by geometric intuition, we introduce a reasonable definition of locally linear patch, Maximal Linear Patch (MLP), which seeks to maximize the local neighborhood in which linearity holds. The input data are first decomposed into a collection of local linear models, each depicting an MLP. These local models are then aligned into a global coordinate space, which is achieved by applying MDS to some randomly selected landmarks. The proposed alignment method, called Landmarks-based Global Alignment (LGA), can efficiently produce a closed-form solution with no risk of local optima. It just involves some small-scale eigenvalue problems, while most previous aligning techniques employ time-consuming iterative optimization. Compared with traditional methods such as ISOMAP and LLE, our MLE yields an explicit modeling of the intrinsic variation modes of the observation data. Extensive experiments on both synthetic and real data indicate the effectivity and efficiency of the proposed algorithm. Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001, Jie Chen 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | WLD: A Robust Local Image DescriptorabstractInspired by Weber's Law, this paper proposes a simple, yet very powerful and robust local descriptor, called the Weber Local Descriptor (WLD). It is based on the fact that human perception of a pattern depends not only on the change of a stimulus (such as sound, lighting) but also on the original intensity of the stimulus. Specifically, WLD consists of two components: differential excitation and orientation. The differential excitation component is a function of the ratio between two terms: One is the relative intensity differences of a current pixel against its neighbors, the other is the intensity of the current pixel. The orientation component is the gradient orientation of the current pixel. For a given image, we use the two components to construct a concatenated WLD histogram. Experimental results on the Brodatz and KTH-TIPS2-a texture databases show that WLD impressively outperforms the other widely used descriptors (e.g., Gabor and SIFT). In addition, experimental results on human face detection also show a promising performance comparable to the best known results on the MIT+CMU frontal face test set, the AR face data set, and the CMU profile test set. Jie Chen 0001, Shiguang Shan, Chu He, Guoying Zhao 0001, Matti Pietikäinen, Xilin Chen 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Fusing Local Patterns of Gabor Magnitude and Phase for Face RecognitionabstractGabor features have been known to be effective for face recognition. However, only a few approaches utilize phase feature and they usually perform worse than those using magnitude feature. To investigate the potential of Gabor phase and its fusion with magnitude for face recognition, in this paper, we first propose local Gabor XOR patterns (LGXP), which encodes the Gabor phase by using the local XOR pattern (LXP) operator. Then, we introduce block-based Fisher's linear discriminant (BFLD) to reduce the dimensionality of the proposed descriptor and at the same time enhance its discriminative power. Finally, by using BFLD, we fuse local patterns of Gabor magnitude and phase for face recognition. We evaluate our approach on FERET and FRGC 2.0 databases. In particular, we perform comparative experimental studies of different local Gabor patterns. We also make a detailed comparison of their combinations with BFLD, as well as the fusion of different descriptors by using BFLD. Extensive experimental results verify the effectiveness of our LGXP descriptor and also show that our fusion approach outperforms most of the state-of-the-art approaches. Shufu Xie, Shiguang Shan, Xilin Chen 0001, Jie Chen 0001 |
IEEE Trans. Image Process. | 4 |
| 2009 | A New Gabor Phase Difference Pattern for Face and Ear Recognition
Yimo Guo, Guoying Zhao 0001, Jie Chen 0001, Matti Pietikäinen, Zhengguang Xu |
CAIP | 3 |
| 2009 | Learning mappings for face synthesis from near infrared to visual light imagesabstractThis paper deals with a new problem in face recognition research, in which the enrollment and query face samples are captured under different lighting conditions. In our case, the enrollment samples are visual light (VIS) images, whereas the query samples are taken under near infrared (NIR) condition. It is very difficult to directly match the face samples captured under these two lighting conditions due to their different visual appearances. In this paper, we propose a novel method for synthesizing VIS images from NIR images based on learning the mappings between images of different spectra (i.e., NIR and VIS). In our approach, we reduce the inter-spectral differences significantly, thus allowing effective matching between faces taken under different imaging conditions. Face recognition experiments clearly show the efficacy of the proposed approach. Jie Chen 0001, Dong Yi, Jimei Yang, Guoying Zhao 0001, Stan Z. Li, Matti Pietikäinen |
CVPR | 1 |
| 2009 | Dynamic texture synthesis using a spatial temporal descriptorabstractDynamic textures are image sequences with visual pattern repetition in time and space, such as smoke, flames, moving objects and so on. Dynamic texture synthesis is to provide a continuous and infinitely varying stream of images by doing operations on dynamic textures. Considering that the previous video texture method provides high-quality visual results, but its representation does not well explore the temporal correlation among frames, we develop a novel spatial temporal descriptor for frame description accompanied with a similarity measure on the basis of the video texture method. Compared with the previous one, our method considers both the spatial and temporal domains of video sequences in representation; moreover, combines the local and global description on each spatial-temporal plane. From experimental results, the proposed method achieves better performance in both the syntheses of natural scene and human motion. Especially, it has the characteristic to be robust to noise in remodeling videos into infinite time domain. Yimo Guo, Guoying Zhao 0001, Jie Chen 0001, Matti Pietikäinen, Zhengguang Xu |
ICIP | 3 |
| 2009 | Optimization of a training set for more robust face detection
Jie Chen 0001, Xilin Chen 0001, Jie Yang 0001, Shiguang Shan, Ruiping Wang 0001, Wen Gao 0001 |
Pattern Recognit. | 1 |
| 2008 | A robust descriptor based on Weber's LawabstractInspired by Weber’s Law, this paper proposes a simple, yet very powerful and robust local descriptor, Weber Local Descriptor (WLD). It is based on the fact that human perception of a pattern depends on not only the change of a stimulus (such as sound, lighting, et al.) but also the original intensity of the stimulus. Specifically, WLD consists of two components: its differential excitation and orientation. A differential excitation is a function of the ratio between two terms: One is the relative intensity differences of its neighbors against a current pixel; the other is the intensity of the current pixel. An orientation is the gradient orientation of the current pixel. For a given image, we use the differential excitation and the orientation components to construct a concatenated WLD histogram feature. Experimental results on Brodatz textures show that WLD impressively outperforms the other classical descriptors (e.g., Gabor). Especially, experimental results on face detection show a promising performance. Although we train only one classifier based on WLD features, the classifier obtains a comparable performance to state-of-the-art methods on MIT+CMU frontal face test set, AR face dataset and CMU profile test set. Jie Chen 0001, Shiguang Shan, Guoying Zhao 0001, Xilin Chen 0001, Wen Gao 0001, Matti Pietikäinen |
CVPR | 1 |
| 2008 | Unsupervised dynamic texture segmentation using local spatiotemporal descriptorsabstractDynamic texture (DT) is an extension of texture to the temporal domain. In this paper, we address the problem of segmenting DT into disjoint regions in an unsupervised way. Each region is characterized by histograms of local binary patterns and contrast in a spatiotemporal mode. It combines the motion and appearance of DT together. Experimental results show that our method is effective in segmenting regions that differ in their dynamics. Jie Chen 0001, Guoying Zhao 0001, Matti Pietikäinen |
ICPR | 1 |
| 2007 | Matrix-Structural Learning (MSL) of Cascaded Classifier from Enormous Training SetabstractAiming at the problem when both positive and negative training set are enormous, this paper proposes a novel matrix-structural learning (MSL) method, as an extension to Viola and Jones' cascade learning method for object detection. Briefly speaking, unlike Viola and Jones' method that learn linearly by bootstrapping only negative samples, the proposed MSL method bootstraps both positive and negative samples in a matrix-like structure. Moreover, an accumulative way is further presented to improve the training efficiency of MSL by inheriting features learned previously during training procedure. The proposed method is evaluated on face detection problem. On a positive set containing 230000 face samples, only 12 hours are needed on a common PC with a 3.20 GHz Pentium IV processor to learn a classifier with false alarm rate less than 1/1000000. What's more, the accuracy of the learned detector exceeds the state-of-the-art results on the CMU+MIT frontal face test set. Shengye Yan, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001, Jie Chen 0001 |
CVPR | 5 |
| 2007 | Enhancing Human Face Detection by Resampling Examples Through ManifoldsabstractAs a large-scale database of hundreds of thousands of face images collected from the Internet and digital cameras becomes available, how to utilize it to train a well-performed face detector is a quite challenging problem. In this paper, we propose a method to resample a representative training set from a collected large-scale database to train a robust human face detector. First, in a high-dimensional space, we estimate geodesic distances between pairs of face samples/examples inside the collected face set by isometric feature mapping (Isomap) and then subsample the face set. After that, we embed the face set to a low-dimensional manifold space and obtain the low-dimensional embedding. Subsequently, in the embedding, we interweave the face set based on the weights computed by locally linear embedding (LLE). Furthermore, we resample nonfaces by Isomap and LLE likewise. Using the resulting face and nonface samples, we train an AdaBoost-based face detector and run it on a large database to collect false alarms. We then use the false detections to train a one-class support vector machine (SVM). Combining the AdaBoost and one-class SVM-based face detector, we obtain a stronger detector. The experimental results on the MIT + CMU frontal face test set demonstrated that the proposed method significantly outperforms the other state-of-the-art methods. Jie Chen 0001, Ruiping Wang 0001, Shengye Yan, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 1 |