VLDB 2026 Research / reviewers in the wild / expert
Hao Chen 0041
dblp:175/3324-41
· DBLP profile ↗
79ranked-venue papers
3as first author
69since 2021 · last 2026
0000-0003-4417-614XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 71 · 2 first-author · 62 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 1 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon TasksabstractLanguage-guided long-horizon mobile manipulation has long been a grand challenge in embodied semantic reasoning, generalizable manipulation, and adaptive locomotion. Three fundamental limitations hinder progress: First, although large language models have shown promise in enhancing spatial reasoning and task planning through learned semantic priors, existing implementations remain confined to tabletop scenarios, failing to address the constrained perception and limited actuation ranges characteristic of mobile platforms. Second, current manipulation strategies exhibit insufficient generalization when confronted with the diverse object configurations encountered in open-world environments. Third, while crucial for practical deployment, the dual requirement of maintaining high platform maneuverability alongside precise end-effector control in unstructured settings remains understudied in the literature. In this work, we present ODYSSEY, a unified mobile manipulation framework for agile quadruped robots equipped with manipulators, which seamlessly integrates high-level task planning with low-level whole-body control. To address the challenge of egocentric perception in language-conditioned tasks, we introduce a hierarchical planner powered by a vision-language model, enabling long-horizon instruction decomposition and precise action execution. At the control level, our novel whole-body policy achieves robust coordination of locomotion and manipulation across challenging terrains. We further present the first comprehensive benchmark for long-horizon mobile manipulation, evaluating diverse indoor and outdoor scenarios. Through successful sim-to-real transfer, we demonstrate the system’s generalization and robustness in real-world deployments, underscoring the practicality of legged manipulators in unstructured environments. Our work advances the feasibility of generalized robotic assistants capable of complex, dynamic tasks. Kaijun Wang, Liqin Lu, Jianuo Jiang, Zeju Li, Wancai Zheng, Hao Chen 0041, Chunhua Shen |
AAAI | 9 |
| 2026 | Beyond Hard Masks: Progressive Token Evolution for Diffusion Language ModelsabstractLinhao Zhong, Linyu Wu, Bozhen Fang, Tianjian Feng, Chenchen Jing, Wen Wang, Jiaheng Zhang, Hao Chen, Chunhua Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Linhao Zhong 0001, Linyu Wu, Bozhen Fang, Tianjian Feng, Chenchen Jing, Wen Wang 0015, Jiaheng Zhang, Hao Chen 0041, Chunhua Shen |
ACL (1) | 8 |
| 2026 | Efficient Self-Evaluation for Diffusion Language Models via Sequence RegenerationabstractLinhao Zhong, Linyu Wu, Wen Wang, Yuling Xi, Chenchen Jing, Jiaheng Zhang, Hao Chen, Chunhua Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Linhao Zhong 0001, Linyu Wu, Wen Wang 0015, Yuling Xi, Chenchen Jing, Jiaheng Zhang, Hao Chen 0041, Chunhua Shen |
ACL (1) | 7 |
| 2026 | Embracing the overlooked: harnessing feature disentanglement for cross-domain learning
Hao Chen 0041 |
Frontiers Comput. Sci. | 1 |
| 2026 | Multi-Modal Primitive Retrieval for Compositional Zero-Shot Learning
Chenchen Jing, Haozhe Zhang 0002, Junbo Lu, Yang Liu 0357, Hao Chen 0041, Xiaoqin Zhang 0002, Chunhua Shen |
Int. J. Comput. Vis. | 5 |
| 2026 | FreerCustom: Training-Free Multi-Concept Customization for Image and Video Generation
Canyu Zhao, Ganggui Ding, Wen Wang 0015, Zhen Yang 0009, Zide Liu, Hao Chen 0041, Chunhua Shen |
Int. J. Comput. Vis. | 6 |
| 2026 | HumanRecon: Neural reconstruction of dynamic human using geometric cues and physical priors
Junhui Yin, Wei Yin 0006, Hao Chen 0041, Xuqian Ren, Zhanyu Ma, Jun Guo 0002, Yifan Liu 0001 |
Pattern Recognit. | 3 |
| 2026 | Towards efficient pixel labeling for industrial anomaly detection and localization
Jingqi Wu, Lin Wu 0001, Hao Chen 0041, Deyin Liu, Haiqiang Jin |
Pattern Recognit. Lett. | 4 |
| 2026 | Retrieval-Enhanced Visual Prompt Learning for Few-Shot ClassificationabstractThe Contrastive Language-Image Pretraining (CLIP) model has been widely used in various downstream vision tasks. The few-shot learning paradigm has been widely adopted to augment its capacity for these tasks. However, current paradigms may struggle with fine-grained classification, such as satellite image recognition, due to widening domain gaps. To address this limitation, we propose retrieval-enhanced visual prompt learning (RePrompt), which introduces retrieval mechanisms to cache and reuse the knowledge of downstream tasks. RePrompt constructs a retrieval database from either training examples or external data if available, and uses a retrieval mechanism to enhance multiple stages of a simple prompt learning baseline, thus narrowing the domain gap. During inference, our enhanced model can reference similar samples brought by retrieval to make more accurate predictions. A detailed analysis reveals that retrieval helps to improve the distribution of late features, thus, improving generalization for downstream tasks. RePrompt attains state-of-the-art performance on a wide range of vision datasets, including 11 image datasets, 3 video datasets, 1 multi-view dataset, and 4 domain generalization benchmarks. Jintao Rong 0001, Hao Chen 0041, Linlin Ou, Tianxiao Chen, Yifan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Zippo: RGB-Alpha Joint Modeling With a Unified Diffusion ModelabstractRecent advances in generative models have sparked growing interest in moving beyond pure image generation toward transparent image generation, i.e., joint generation of image and its alpha mask. However, most existing approaches adopt a two-stage pipeline, where a diffusion-based model first generates an RGB image and a subsequent matting head predicts the alpha mask. This separation not only leads to error accumulation and inaccurate predictions but also overlooks the intrinsic correlation between the cross-modal data. In this work, we introduce Zippo, a unified diffusion framework, zipping color and transparency distributions into a single diffusion model, by learning joint distribution of RGB image and alpha mask. Zippo not only generates high-fidelity images but also produces plausible and sharp alpha masks. In practice, Zippo inflates the latent space into a unified representation that encodes cross-modal data, and builds upon it with a modality-aware diffusion process that flexibly switches between RGB and alpha domains. In this process, conditioning on one modality while denoising the other allows the model to generate RGB images from alpha masks and predict transparency from input images. In addition to single-modality prediction, we further design a modality-aware noise reassignment strategy to empower Zippo with the joint generation capability of RGB images and their corresponding alpha masks under text guidance. With these techniques, Zippo supports a wide range of transparent image generation tasks, including image-alpha joint generation, image matting, and alpha mask conditioned image generation. Extensive experiments demonstrate that Zippo not only delivers superior visual fidelity but also achieves competitive performance in visual downstream prediction, highlighting joint image-alpha modeling as a powerful alternative to traditional paradigms. Kangyang Xie, Chenchen Jing, Cheng Peng 0011, Ming Yang 0007, Heqian Qiu, Hongliang Li 0001, Hao Chen 0041 |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2026 | Accurate Industrial Anomaly Detection and Localization Using Weakly-Supervised Residual TransformersabstractRecent advancements in industrial anomaly detection (AD) have demonstrated that incorporating a small number of anomalous samples during training can significantly enhance accuracy. However, this improvement often comes at the cost of extensive annotation efforts, which are impractical for many real-world applications. In this paper, we introduce a novel framework, "Weakly-supervised RESidual $T$ ransformer" (WeakREST), designed to achieve high anomaly detection accuracy while minimizing the reliance on manual annotations. First, we reformulate the pixel-wise anomaly localization task into a block-wise classification problem. Second, we introduce a residual-based feature representation called "Positional $F$ ast $A$ nomaly $R$ esiduals" (PosFAR) which captures anomalous patterns more effectively. To leverage this feature, we adapt the Swin Transformer for enhanced anomaly detection and localization. Additionally, we propose a weak annotation approach utilizing bounding boxes and image tags to define anomalous regions. This approach establishes a semi-supervised learning context that reduces the dependency on precise pixel-level labels. To further improve the learning process, we develop a novel ResMixMatch algorithm, capable of handling the interplay between weak labels and residual-based representations. On the benchmark dataset MVTec-AD, our method achieves an Average Precision (AP) of 83.0%, surpassing the previous best result of 82.7% in the unsupervised setting. In the supervised AD setting, WeakREST attains an AP of 87.6%, outperforming the previous best of 86.0%. Notably, even when using weaker annotations such as bounding boxes, WeakREST exceeds the performance of leading methods relying on pixel-wise supervision, achieving an AP of 87.1% compared to the prior best of 86.0% on MVTec-AD. This superior performance is consistently replicated across other well-established AD datasets, including MVTec 3D, KSDD2 and Real-IAD. Code is available at: https://github.com/BeJane/Semi_REST. Jingqi Wu, Deyin Liu, Lin Wu 0001, Hao Chen 0041, Chunhua Shen |
IEEE Trans. Image Process. | 5 |
| 2025 | A Denoising Pre-training Framework for Accelerating Novel Material DiscoveryabstractCrystal materials play an important role in the development of society. The discovery of new materials is critical to achieving sustainable development goals (SDGs), such as climate change mitigation, affordable and clean energy, and fostering innovation in industry and infrastructure. Recent advances in deep learning for crystal property prediction have accelerated material discovery, but these methods typically rely on labeled data, which is often limited and varies across different properties. This limitation hinders the full utilization of the vast amount of unlabeled data in materials science. To overcome this challenge, we introduce an unsupervised Denoising Pre-training Framework (DPF) tailored for crystal structures. DPF trains a model to reconstruct the original crystal structure by recovering the masked atom types, perturbed atom positions, and perturbed crystal lattices. Through pre-training, models learn the intrinsic features of crystal structures and capture the key features influencing crystal properties. We pre-train models on a dataset of 380,743 unlabeled crystal structures and fine-tune them on downstream property prediction tasks. Extensive experiments demonstrate the effectiveness of our framework, showing its potential to significantly advance material science and contribute to the development of society by accelerating the discovery of materials crucial for sustainable technologies. Shuaike Shen, Ke Liu 0012, Muzhi Zhu, Hao Chen 0041 |
AAAI | 4 |
| 2025 | TG-LLaVA: Text Guided LLaVA via Learnable Latent EmbeddingsabstractCurrently, inspired by the success of vision-language models (VLMs), an increasing number of researchers are focusing on improving VLMs and have achieved promising results. However, most existing methods concentrate on optimizing the connector and enhancing the language model component, while neglecting improvements to the vision encoder itself. In contrast, we propose Text Guided LLaVA (TG-LLaVA) in this paper, which optimizes VLMs by guiding the vision encoder with text, offering a new and orthogonal optimization direction. Specifically, inspired by the purpose-driven logic inherent in human behavior, we use learnable latent embeddings as a bridge to analyze textual instruction and add the analysis results to the vision encoder as guidance, refining it. Subsequently, another set of latent embeddings extracts additional detailed text-guided information from high-resolution local patches as auxiliary information. Finally, with the guidance of text, the vision encoder can extract text-related features, similar to how humans focus on the most relevant parts of an image when considering a question. This results in generating better answers. Experiments on various datasets validate the effectiveness of the proposed method. Remarkably, without the need for additional training data, our proposed method can bring more benefits to the baseline (LLaVA-1.5) compared with other concurrent methods. Furthermore, the proposed method consistently brings improvement in different settings. Dawei Yan 0001, Hao Chen 0041, Weihua Luo, Wei Dong 0010, Qingsen Yan, Haokui Zhang, Chunhua Shen |
AAAI | 4 |
| 2025 | SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator TrajectoriesabstractWhile MLLMs have demonstrated impressive image understanding capabilities, they still struggle with pixel-level comprehension, limiting their practical applications. Current evaluation tasks such as VQA and visual grounding remain too coarse to assess fine-grained pixel comprehension accurately. Although segmentation is foundational for pixel-level understanding, existing methods often require MLLMs to generate implicit tokens, decoded through external pixel decoders. This approach disrupts the MLLM’s text output space, potentially compromising language capabilities and reducing flexibility and extensibility while failing to reflect the model’s intrinsic pixel-level understanding. Thus, we introduce the Human-Like Mask Annotation Task (HLMAT), a new paradigm where MLLMs mimic human annotators using interactive segmentation tools. Modelling segmentation as a multi-step Markov Decision Process, HLMAT enables MLLMs to iteratively generate text-based click points, achieving high-quality masks without architectural changes or implicit tokens. Through this setup, we develop SegAgent, a model fine-tuned on human-like annotation trajectories, which achieves performance comparable to SoTA methods and supports additional tasks like mask refinement and annotation filtering. HLMAT provides a protocol for assessing fine-grained pixel understanding in MLLMs and introduces a vision-centric, multi-step decision-making task that facilitates the exploration of MLLMs’ visual reasoning abilities. Our adaptations of policy improvement method StaR and PRM guided tree search further enhance model robustness in complex segmentation tasks, laying a foundation for future advancements in fine-grained visual perception and multi-step decision-making for MLLMs. Code can be found at https://github.com/aim-uofa/SegAgent. Muzhi Zhu, Yuzhuo Tian, Hao Chen 0041, Chunluan Zhou, Qingpei Guo, Yang Liu 0357, Ming Yang 0007, Chunhua Shen |
CVPR | 3 |
| 2025 | SurfaceSplat: Connecting Surface Reconstruction and Gaussian SplattingabstractSurface reconstruction and novel view rendering from sparse-view images are challenging. Signed Distance Function (SDF)-based methods struggle with fine details, while 3D Gaussian Splatting (3DGS)-based approaches lack global geometry coherence. We propose a novel hybrid method that combines the strengths of both approaches: SDF captures coarse geometry to enhance 3DGS-based rendering, while newly rendered images from 3DGS refine the details of SDF for accurate surface reconstruction. As a result, our method surpasses state-of-the-art approaches in surface reconstruction and novel view synthesis on the DTU and MobileBrick datasets. Code will be released at https://github.com/aim-uofa/SurfaceSplat. Zihui Gao, Jiawang Bian, Guosheng Lin, Hao Chen 0041, Chunhua Shen |
ICCV | 4 |
| 2025 | Unified Open-World Segmentation with Multi-Modal Prompts
Yang Liu 0357, Yufei Yin, Chenchen Jing, Muzhi Zhu, Hao Chen 0041, Yuling Xi, Hao Wang 0052, Chunhua Shen |
ICCV | 5 |
| 2025 | POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction
Songyan Zhang, Yongtao Ge, Jinyuan Tian, Guangkai Xu, Hao Chen 0041, Chunhua Shen |
ICCV | 5 |
| 2025 | Framer: Interactive Frame InterpolationabstractWe propose Framer for interactive frame interpolation, which targets producing smoothly transitioning frames between two images as per user creativity. Concretely, besides taking the start and end frames as inputs, our approach supports customizing the transition process by tailoring the trajectory of some selected keypoints. Such a design enjoys two clear benefits. First, incorporating human interaction mitigates the issue arising from numerous possibilities of transforming one image to another, and in turn enables finer control of local motions. Second, as the most basic form of interaction, keypoints help establish the correspondence across frames, enhancing the model to handle challenging cases (e.g., objects on the start and end frames are of different shapes and styles). It is noteworthy that our system also offers an "autopilot" mode, where we introduce a module to estimate the keypoints and refine the trajectory automatically, to simplify the usage in practice. Extensive experimental results demonstrate the appealing performance of Framer on various applications, such as image morphing, time-lapse video generation, cartoon interpolation, etc. The code, model, and interface are publicly accessible at https://github.com/aim-uofa/Framer. Wen Wang 0015, Qiuyu Wang, Kecheng Zheng, Hao Ouyang, Zhekai Chen, Biao Gong, Hao Chen 0041, Yujun Shen, Chunhua Shen |
ICLR | 7 |
| 2025 | Revisiting Convolution Architecture in the Realm of DNA Foundation ModelsabstractIn recent years, A variety of methods based on Transformer and state space model (SSM) architectures have been proposed, advancing foundational DNA language models.
However, there is a lack of comparison between these recent approaches and the classical architecture—convolutional networks (CNNs)—on foundation model benchmarks.
This raises the question: are CNNs truly being surpassed by these recent approaches based on transformer and SSM architectures? In this paper, we develop a simple but well-designed CNN-based method, termed ConvNova. ConvNova identifies and proposes three effective designs: 1) dilated convolutions, 2) gated convolutions, and 3) a dual-branch framework for gating mechanisms.
Through extensive empirical experiments, we demonstrate that ConvNova significantly outperforms recent methods on more than half of the tasks across several foundation model benchmarks. For example, in histone-related tasks, ConvNova exceeds the second-best method by an average of 5.8\%, while generally utilizing fewer parameters and enabling faster computation. In addition, the experiments observed findings that may be related to biological characteristics. This indicates that CNNs are still a strong competitor compared to Transformers and SSMs. We anticipate that this work will spark renewed interest in CNN-based methods for DNA foundation models. Yu Bo, Weian Mao, Yanjun Shao, Weiqiang Bai, Peng Ye 0006, Xinzhu Ma, Hao Chen 0041, Chunhua Shen |
ICLR | 8 |
| 2025 | PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual TrainingabstractThis paper aims to address the challenge of hallucinations in Multimodal Large Language Models (MLLMs) particularly for dense image captioning tasks. To tackle the challenge, we identify the current lack of a metric that finely measures the caption quality in concept level. We hereby introduce HalFscore, a novel metric built upon the language graph and is designed to evaluate both the accuracy and completeness of dense captions at a
granular level. Additionally, we identify the root cause of hallucination as the model's over-reliance on its language prior. To address this, we propose PerturboLLaVA, which reduces the model's reliance on the language prior by incorporating adversarially perturbed text during training. This method enhances the model's focus on visual inputs, effectively reducing hallucinations and producing accurate, image-grounded descriptions without incurring additional computational overhead. PerturboLLaVA significantly improves the fidelity of generated captions, outperforming existing approaches in handling multimodal hallucinations and achieving improved performance across general multimodal benchmarks. Chenchen Jing, Yizhou Zhou, Fengyun Rao, Hao Chen 0041, Bo Zhang 0046, Chunhua Shen |
ICLR | 6 |
| 2025 | Boltzmann-Aligned Inverse Folding Model as a Predictor of Mutational Effects on Protein-Protein InteractionsabstractPredicting the change in binding free energy ($\Delta \Delta G$) is crucial for understanding and modulating protein-protein interactions, which are critical in drug design.
Due to the scarcity of experimental $\Delta\Delta G$ data,
existing methods focus on pre-training,
while neglecting the importance of alignment.
In this work, we propose Boltzmann Alignment technique to transfer knowledge from pre-trained inverse folding models to prediction of $\Delta\Delta G$.
We begin by analyzing the thermodynamic definition of $\Delta\Delta G$ and introducing the Boltzmann distribution to connect energy to the protein conformational distribution.
However, the protein conformational distribution is intractable. Therefore, we employ Bayes’ theorem to circumvent direct estimation and instead utilize the log-likelihood provided by protein inverse folding models for the estimation of $\Delta\Delta G$.
Compared to previous methods based on inverse folding, our method explicitly accounts for the unbound state of the protein complex in the $\Delta \Delta G$ thermodynamic cycle, introducing a physical inductive bias and achieving supervised and unsupervised state-of-the-art (SoTA) performance.
Experimental results on SKEMPI v2 indicate that our method achieves Spearman coefficients of 0.3201 (unsupervised) and 0.5134 (supervised) on SKEMPI v2, significantly surpassing the previously reported
%SoTA values
SoTA results
of 0.2632 and 0.4324, respectively.
Furthermore, we demonstrate the capability of our method in binding
energy prediction, protein-protein docking, and antibody optimization tasks.
Code is available at [https://github.com/aim-uofa/BA-DDG](https://github.com/aim-uofa/BA-DDG) Xiaoran Jiao, Weian Mao, Wengong Jin, Peiyuan Yang, Hao Chen 0041, Chunhua Shen |
ICLR | 5 |
| 2025 | What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?abstractExtensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I diffusion models for a few dense perception tasks. However, several crucial design decisions in this process still lack comprehensive justification, encompassing the necessity of the multi-step diffusion mechanism, training strategy, inference ensemble strategy, and fine-tuning data quality. In this work, we conduct a thorough investigation into critical factors that affect transfer efficiency and performance when using diffusion priors. Our key findings are: 1) High-quality fine-tuning data is paramount for both semantic and geometry perception tasks. 2) As a special case of the diffusion scheduler by setting its hyper-parameters, the multi-step generation can be simplified to a one-step fine-tuning paradigm without any loss of performance, while significantly speeding up inference. 3) Apart from fine-tuning the diffusion model with only latent space supervision, task-specific supervision can be beneficial to enhance fine-grained details. These observations culminate in the development of GenPercept, an effective deterministic one-step fine-tuning paradigm tailored for dense visual perception tasks exploiting diffusion priors. Different from the previous multi-step methods, our paradigm offers a much faster inference speed, and can be seamlessly integrated with customized perception decoders and loss functions for task-specific supervision, which can be critical for improving the fine-grained details of predictions. Comprehensive experiments on a diverse set of dense visual perceptual tasks, including monocular depth estimation, surface normal estimation, image segmentation, and matting, are performed to demonstrate the remarkable adaptability and effectiveness of our proposed method. Code: https://github.com/aim-uofa/GenPercept Guangkai Xu, Yongtao Ge, Chengxiang Fan, Kangyang Xie, Zhiyue Zhao, Hao Chen 0041, Chunhua Shen |
ICLR | 7 |
| 2025 | MovieDreamer: Hierarchical Generation for Coherent Long Visual SequencesabstractRecent advancements in video generation have primarily leveraged diffusion models for short-duration content. However, these approaches often fall short in modeling complex narratives and maintaining character consistency over extended periods, which is essential for long-form video production like movies. We propose MovieDreamer, a novel hierarchical framework that integrates the strengths of autoregressive models with diffusion-based rendering to pioneer long-duration video generation with intricate plot progressions and high visual fidelity. Our approach utilizes autoregressive models for global narrative coherence, predicting sequences of visual tokens that are subsequently transformed into high-quality video frames through diffusion rendering. This method is akin to traditional movie production processes, where complex stories are factorized down into manageable scene capturing. Further, we employ a multimodal script that enriches scene descriptions with detailed character information and visual style, enhancing continuity and character identity across scenes. We present extensive experiments across various movie genres, demonstrating that our approach not only achieves superior visual and narrative quality but also effectively extends the duration of generated content significantly beyond current capabilities. Canyu Zhao, Wen Wang 0015, Fan Wang 0019, Hao Chen 0041, Bo Zhang 0025, Chunhua Shen |
ICLR | 6 |
| 2025 | Seeing the Unseen: Composing Outliers for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) is to recognize unseen attribute-object compositions by learning from seen compositions. The distribution shift between unseen compositions and seen compositions poses challenges to CZSL models, especially when test images are mixed with both seen and unseen compositions. The challenge will be addressed more easily if a model can distinguish unseen/seen compositions and treat them with specific recognition strategies. However, identifying images with unseen compositions is non-trivial, considering that unseen compositions are absent in training and usually contain only subtle differences from seen compositions. In this paper, we propose a novel compositional zero-shot learning method called COMO, which composes outliers in training for distinguishing seen and unseen compositions and further applying specific strategies for them. Specifically, we compose attribute-object representations for unseen compositions based on primitive representations of training images as outliers to enable the model to identify unseen compositions in inference. At test time, the method distinguishes images containing seen/unseen compositions and uses different weights for composition classification and primitive classification to recognize seen/unseen compositions. Experimental results on three datasets show the effectiveness of our method in both the closed-world setting and the open-world setting. Chenchen Jing, Hao Chen 0041, Yuling Xi, Xingyuan Bu, Dong Gong, Chunhua Shen |
IJCAI | 3 |
| 2025 | DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksabstractThis paper's primary objective is to develop a robust generalist perception model capable of addressing multiple tasks under constraints of computational resources and limited training data. We leverage text-to-image diffusion models pre-trained on billions of images and successfully introduce our DICEPTION, a visual generalist model. Exhaustive evaluations demonstrate that DICEPTION effectively tackles diverse perception tasks, even achieving performance comparable to SOTA single-task specialist models. Specifically, we achieve results on par with SAM-vit-h using only 0.06% of their data (e.g., 600K vs.\ 1B pixel-level annotated images). We designed comprehensive experiments on architectures and input paradigms, demonstrating that the key to successfully re-purposing a single diffusion model for multiple perception tasks lies in maximizing the preservation of the pre-trained model's prior knowledge. Consequently, DICEPTION can be trained with substantially lower computational costs than conventional models requiring training from scratch. Furthermore, adapting DICEPTION to novel tasks is highly efficient, necessitating fine-tuning on as few as 50 images and approximately 1% of its parameters. Finally, we demonstrate that a subtle application of classifier-free guidance can improve the model's performance on depth and normal estimation. We also show that pixel-aligned training, as is characteristic of perception tasks, significantly enhances the model's ability to preserve fine details. DICEPTION offers valuable insights and presents a promising direction for the development of advanced diffusion-based visual generalist models. Canyu Zhao, Yanlong Sun, Huanyi Zheng, Muzhi Zhu, Zhiyue Zhao, Hao Chen 0041, Tong He 0001, Chunhua Shen |
NeurIPS | 7 |
| 2025 | Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System CollaborationabstractLong-horizon video-audio reasoning and fine-grained pixel understanding impose conflicting requirements on omnimodal models: dense temporal coverage demands many low-resolution frames, whereas precise grounding calls for high-resolution inputs. We tackle this trade-off with a two-system architecture: a Global Reasoning System selects informative keyframes and rewrites the task at low spatial cost, while a Detail Understanding System performs pixel-level grounding on the selected high-resolution snippets.
Because "optimal" keyframe selection and reformulation are ambiguous and hard to supervise, we formulate them as a reinforcement-learning (RL) problem and present Omni-R1, an end-to-end RL framework built on Group Relative Policy Optimization.
Omni-R1 trains the Global Reasoning System through hierarchical rewards obtained via online collaboration with the Detail Understanding System, requiring only one epoch of RL on small task splits.
Experiments on two challenging benchmarks, Referring Audio-Visual Segmentation (RefAVS) and Reasoning Video Object Segmentation (REVOS), show that Omni-R1 not only surpasses strong supervised baselines but also outperforms specialized state-of-the-art models, while substantially improving out-of-domain generalization and mitigating multimodal hallucination.
Our results demonstrate the first successful application of RL to large-scale omnimodal reasoning and highlight a scalable path toward universally foundation models. Muzhi Zhu, Zongze Du, Canyu Zhao, Wen Wang 0015, Hao Chen 0041, Chunhua Shen |
NeurIPS | 8 |
| 2025 | Segment Anything in Context with Vision Foundation Models
Yang Liu 0357, Muzhi Zhu, Hao Chen 0041, Hao Wang 0052, Raviteja Vemulapalli, Chunhua Shen |
Int. J. Comput. Vis. | 3 |
| 2025 | AutoStory: Generating Diverse Storytelling Images with Minimal Human Efforts
Wen Wang 0015, Canyu Zhao, Hao Chen 0041, Zhekai Chen, Kecheng Zheng, Chunhua Shen |
Int. J. Comput. Vis. | 3 |
| 2025 | CoLeCLIP: Open-Domain Continual Learning via Joint Task Prompt and Vocabulary LearningabstractThis article investigates the problem of continual learning (CL) of vision-language models (VLMs) in open domains, where models are required to perform continual updating and inference on a stream of datasets from diverse seen and unseen domains with novel classes. Such a capability is crucial for various applications in open environments, e.g., AI assistants, autonomous driving systems, and robotics. Current CL studies mostly focus on closed-set scenarios in a single domain with known classes. Large pretrained VLMs such as CLIP have showcased exceptional zero-shot recognition capabilities, and several recent studies have leveraged the unique characteristics of VLMs to mitigate catastrophic forgetting in CL. However, they primarily focus on closed-set CL in a single-domain dataset. Open-domain CL of large VLMs is significantly more challenging due to 1) large class correlations and domain gaps across the datasets and 2) the forgetting of zero-shot knowledge in the pretrained VLMs and the knowledge learned from the newly adapted datasets. In this work, we introduce a novel approach, termed CoLeCLIP, which learns an open-domain CL model based on CLIP. It addresses these challenges through joint learning of a set of task prompts and a cross-domain class vocabulary. Extensive experiments on 11 domain datasets show that CoLeCLIP achieves new state-of-the-art performance for open-domain CL under both task- and class-incremental learning (CIL) settings. Guansong Pang, Wei Suo, Chenchen Jing, Yuling Xi, Lingqiao Liu, Hao Chen 0041, Guoqiang Liang 0001, Peng Wang 0015 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Retrieval-Augmented Primitive Representations for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) aims to recognize unseen attribute-object compositions by learning from seen compositions. Composing the learned knowledge of seen primitives, i.e., attributes or objects, into novel compositions is critical for CZSL. In this work, we propose to explicitly retrieve knowledge of seen primitives for compositional zero-shot learning. We present a retrieval-augmented method, which augments standard multi-path classification methods with two retrieval modules. Specifically, we construct two databases storing the attribute and object representations of training images, respectively. For an input training/testing image, we use two retrieval modules to retrieve representations of training images with the same attribute and object, respectively. The primitive representations of the input image are augmented by using the retrieved representations, for composition recognition. By referencing semantically similar images, the proposed method is capable of recalling knowledge of seen primitives for compositional generalization. Experiments on three widely-used datasets show the effectiveness of the proposed method. Chenchen Jing, Hao Chen 0041, Chunhua Shen |
AAAI | 3 |
| 2024 | Revisiting Open-Set Panoptic SegmentationabstractIn this paper, we focus on the open-set panoptic segmentation (OPS) task to circumvent the data explosion problem. Different from the close-set setting, OPS targets to detect both known and unknown categories, where the latter is not annotated during training. Different from existing work that only selects a few common categories as unknown ones, we move forward to the real-world scenario by considering the various tail categories (~1k). To this end, we first build a new dataset with long-tail distribution for the OPS task. Based on this dataset, we additionally add a new class type for unknown classes and re-define the training annotations to make the OPS definition more complete and reasonable. Moreover, we analyze the influence of several significant factors in the OPS task and explore the upper bound of performance on unknown classes with different settings. Furthermore, based on the analyses, we design an effective two-phase framework for the OPS task, including thing-agnostic map generation and unknown segment mining. We further adopt semi-supervised learning to improve the OPS performance. Experimental results on different datasets validate the effectiveness of our method. Yufei Yin, Hao Chen 0041, Wengang Zhou 0001, Jiajun Deng, Houqiang Li |
AAAI | 2 |
| 2024 | FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept CompositionabstractBenefiting from large-scale pre-trained text-to-image (T2I) generative models, impressive progress has been achieved in customized image generation, which aims to generate user-specified concepts. Existing approaches have extensively focused on single-concept customization and still encounter challenges when it comes to complex scenarios that involve combining multiple concepts. These approaches often require retraining/fine-tuning using a few images, leading to time-consuming training processes and impeding their swift implementation. Furthermore, the reliance on multiple images to represent a singular concept increases the difficulty of customization. To this end, we propose FreeCustom, a novel tuning-free method to generate customized images of multi-concept composition based on reference concepts, using only one image per concept as input. Specifically, we introduce a new multi-reference self-attention (MRSA) mechanism and a weighted mask strategy that enables the generated image to access and focus more on the reference concepts. In addition, MRSA leverages our key finding that input concepts are better preserved when providing images with context interactions. Experiments show that our method's produced images are consistent with the given concepts and better aligned with the input text. Our method outperforms or performs on par with other training-based methods in terms of multi-concept composition and single-concept customization, but is simpler. Codes can be found here. Ganggui Ding, Canyu Zhao, Wen Wang 0015, Zhen Yang 0009, Zide Liu, Hao Chen 0041, Chunhua Shen |
CVPR | 6 |
| 2024 | DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative DataabstractInstance segmentation is data-hungry, and as model capacity increases, data scale becomes crucial for improving the accuracy. Most instance segmentation datasets today require costly manual annotation, limiting their data scale. Models trained on such data are prone to overfitting on the training set, especially for those rare categories. While recent works have delved into exploiting generative models to create synthetic datasets for data augmentation, these approaches do not efficiently harness the full potential of generative models. To address these issues, we introduce a more efficient strategy to construct generative datasets for data augmentation, termed DiverGen. Firstly, we provide an explanation of the role of generative data from the perspective of distribution discrepancy. We investigate the impact of different data on the distribution learned by the model. We argue that generative data can expand the data distribution that the model can learn, thus mitigating overfitting. Additionally, we find that the diversity of generative data is crucial for improving model performance and enhance it through various strategies, including category diversity, prompt diversity, and generative model diversity. With these strategies, we can scale the data to millions while maintaining the trend of model performance improvement. On the LVIS dataset, DiverGen significantly outperforms the strong model X-Paste, achieving +1.1 box AP and +1.1 mask AP across all categories, and +1.9 box AP and +2.5 mask AP for rare categories. Our codes are available at https://github.com/aim-uofa/DiverGen. Chengxiang Fan, Muzhi Zhu, Hao Chen 0041, Yang Liu 0357, Weijia Wu 0001, Huaqi Zhang, Chunhua Shen |
CVPR | 3 |
| 2024 | FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior
Zhekai Chen, Wen Wang 0015, Zhen Yang 0009, Zeqing Yuan, Hao Chen 0041, Chunhua Shen |
ECCV (17) | 5 |
| 2024 | Object-Aware Inversion and Reassembly for Image EditingabstractDiffusion-based image editing methods have achieved remarkable advances in text-driven image editing. The editing task aims to convert an input image with the original text prompt into the desired image that is well-aligned with the target text prompt. By comparing the original and target prompts, we can obtain numerous editing pairs, each comprising an object and its corresponding editing target. To allow editability while maintaining fidelity to the input image, existing editing methods typically involve a fixed number of inversion steps that project the whole input image to its noisier latent representation, followed by a denoising process guided by the target prompt. However, we find that the optimal number of inversion steps for achieving ideal editing results varies significantly among different editing pairs, owing to varying editing difficulties. Therefore, the current literature, which relies on a fixed number of inversion steps, produces sub-optimal generation quality, especially when handling multiple editing pairs in a natural image.
To this end, we propose a new image editing paradigm, dubbed Object-aware Inversion and Reassembly (OIR), to enable object-level fine-grained editing. Specifically, we design a new search metric, which determines the optimal inversion steps for each editing pair, by jointly considering the editability of the target and the fidelity of the non-editing region. We use our search metric to find the optimal inversion step for each editing pair when editing an image. We then edit these editing pairs separately to avoid \concept. Subsequently, we propose an additional reassembly step to seamlessly integrate the respective editing results and the non-editing region to obtain the final edited image. To systematically evaluate the effectiveness of our method, we collect two datasets called OIRBench for benchmarking single- and multi-object editing, respectively. Experiments demonstrate that our method achieves superior performance in editing object shapes, colors, materials, categories, \textit{etc.}, especially in multi-object editing scenarios.
The project page can be found in https://aim-uofa.github.io/OIR-Diffusion/. Zhen Yang 0009, Ganggui Ding, Wen Wang 0015, Hao Chen 0041, Bohan Zhuang, Chunhua Shen |
ICLR | 4 |
| 2024 | Matcher: Segment Anything with One Shot Using All-Purpose Feature MatchingabstractPowered by large-scale pre-training, vision foundation models exhibit significant potential in open-world image understanding. However, unlike large language models that excel at directly tackling various language tasks, vision foundation models require a task-specific model structure followed by fine-tuning on specific tasks. In this work, we present $\textbf{Matcher}$, a novel perception paradigm that utilizes off-the-shelf vision foundation models to address various perception tasks. Matcher can segment anything by using an in-context example without training. Additionally, we design three effective components within the Matcher framework to collaborate with these foundation models and unleash their full potential in diverse perception tasks. Matcher demonstrates impressive generalization performance across various segmentation tasks, all without training. For example, it achieves 52.7% mIoU on COCO-20$^i$ with one example, surpassing the state-of-the-art specialist model by 1.6%. In addition, Matcher achieves 33.0% mIoU on the proposed LVIS-92$^i$ for one-shot semantic segmentation, outperforming the state-of-the-art generalist model by 14.4%. Our visualization results further showcase the open-world generality and flexibility of Matcher when applied to images in the wild. Yang Liu 0357, Muzhi Zhu, Hengtao Li, Hao Chen 0041, Chunhua Shen |
ICLR | 4 |
| 2024 | De novo Protein Design Using Geometric Vector Field NetworksabstractAdvances like protein diffusion have marked revolutionary progress in $\textit{de novo}$ protein design, a central topic in life science. These methods typically depend on protein structure encoders to model residue backbone frames, where atoms do not exist. Most prior encoders rely on atom-wise features, such as angles and distances between atoms, which are not available in this context. Only a few basic encoders, like IPA, have been proposed for this scenario, exposing the frame modeling as a bottleneck. In this work, we introduce the Vector Field Network (VFN), that enables network layers to perform learnable vector computations between coordinates of frame-anchored virtual atoms, thus achieving a higher capability for modeling frames. The vector computation operates in a manner similar to a linear layer, with each input channel receiving 3D virtual atom coordinates instead of scalar values. The multiple feature vectors output by the vector computation are then used to update the residue representations and virtual atom coordinates via attention aggregation. Remarkably, VFN also excels in modeling both frames and atoms, as the real atoms can be treated as the virtual atoms for modeling, positioning VFN as a potential $\textit{universal encoder}$. In protein diffusion (frame modeling), VFN exhibits a impressive performance advantage over IPA, excelling in terms of both designability ($\textbf{67.04}$\% vs. 53.58\%) and diversity ($\textbf{66.54}$\% vs. 51.98\%). In inverse folding(frame and atom modeling), VFN outperforms the previous SoTA model, PiFold ($\textbf{54.7}$\% vs. 51.66\%), on sequence recovery rate; we also propose a method of equipping VFN with the ESM model, which significantly surpasses the previous ESM-based SoTA ($\textbf{62.67}$\% vs. 55.65\%), LM-Design, by a substantial margin. Code is available at https://github.com/aim-uofa/VFN Weian Mao, Muzhi Zhu, Shuaike Shen, Lin Wu 0001, Hao Chen 0041, Chunhua Shen |
ICLR | 6 |
| 2024 | Generative Active Learning for Long-tailed Instance SegmentationabstractRecently, large-scale language-image generative models have gained widespread attention and many works have utilized generated data from these models to further enhance the performance of perception tasks. However, not all generated data can positively impact downstream models, and these methods do not thoroughly explore how to better select and utilize generated data. On the other hand, there is still a lack of research oriented towards active learning on generated data. In this paper, we explore how to perform active learning specifically for generated data in the long-tailed instance segmentation task. Subsequently, we propose BSGAL, a new algorithm that estimates the contribution of the current batch-generated data based on gradient cache. BSGAL is meticulously designed to cater for unlimited generated data and complex downstream segmentation tasks. BSGAL outperforms the baseline approach and effectually improves the performance of long-tailed segmentation. Muzhi Zhu, Chengxiang Fan, Hao Chen 0041, Yang Liu 0357, Weian Mao, Xiaogang Xu 0002, Chunhua Shen |
ICML | 3 |
| 2024 | Unveiling Universal Forensics of Diffusion Models with Adversarial PerturbationsabstractWith state-of-the-art performance in image synthesis especially in text-to-image generation, diffusion models (DMs) have received unprecedented attention. Despite promising application prospects, fake images generated from diffusion models are causing potential security concerns. To this end, in this work, we aim to investigate whether the generated images of DMs are different from other generated models e.g. generative adversarial networks (GANs) and whether a universal forensic classifier exists. To perform this work, we first collected a dataset consisting of 409k fake images generated from different types of DMs. Through a comprehensive analysis on this benchmark, we showcased that common forensic artifacts are shared among DMs and a forensic classifier trained for one model can generalize well for other agnostic generative models. Specifically, we first demonstrated that despite photorealism of the images generated by DMs, they still contain artifacts namely non-robust visual features which are hard for human but easy for machine to recognize. Then we studied the characteristics of the artifacts from the view of adversarial attack and unexpectedly found there exists a universal adversarial perturbation to fool the classifier. Furthermore, we devised visualization and analysis tools focusing on the spectral properties of the generated samples and adversarial features which demonstrates augmentations in the frequency domain greatly affect the performance of the detectors. Kangyang Xie, Jiaan Liu, Muzhi Zhu, Ganggui Ding, Zide Liu, Hao Chen 0041, Hangyue Chen |
IJCNN | 6 |
| 2024 | Transformer Doctor: Diagnosing and Treating Vision TransformersabstractDue to its powerful representational capabilities, Transformers have gradually become the mainstream model in the field of machine vision. However, the vast and complex parameters of Transformers impede researchers from gaining a deep understanding of their internal mechanisms, especially error mechanisms. Existing methods for interpreting Transformers mainly focus on understanding them from the perspectives of the importance of input tokens or internal modules, as well as the formation and meaning of features. In contrast, inspired by research on information integration mechanisms and conjunctive errors in the biological visual system, this paper conducts an in-depth exploration of the internal error mechanisms of Transformers. We first propose an information integration hypothesis for Transformers in the machine vision domain and provide substantial experimental evidence to support this hypothesis. This includes the dynamic integration of information among tokens and the static integration of information within tokens in Transformers, as well as the presence of conjunctive errors therein. Addressing these errors, we further propose heuristic dynamic integration constraint methods and rule-based static integration constraint methods to rectify errors and ultimately improve model performance. The entire methodology framework is termed as Transformer Doctor, designed for diagnosing and treating internal errors within transformers. Through a plethora of quantitative and qualitative experiments, it has been demonstrated that Transformer Doctor can effectively address internal errors in transformers, thereby enhancing model performance. Jiacong Hu, Hao Chen 0041, Kejia Chen 0007, Yang Gao 0001, Jingwen Ye, Xingen Wang, Mingli Song, Zunlei Feng |
NeurIPS | 2 |
| 2024 | A Simple Image Segmentation Framework via In-Context ExamplesabstractRecently, there have been explorations of generalist segmentation models that can effectively tackle a variety of image segmentation tasks within a unified in-context learning framework. However, these methods still struggle with task ambiguity in in-context segmentation, as not all in-context examples can accurately convey the task information. In order to address this issue, we present SINE, a simple image $\textbf{S}$egmentation framework utilizing $\textbf{in}$-context $\textbf{e}$xamples. Our approach leverages a Transformer encoder-decoder structure, where the encoder provides high-quality image representations, and the decoder is designed to yield multiple task-specific output masks to eliminate task ambiguity effectively. Specifically, we introduce an In-context Interaction module to complement in-context information and produce correlations between the target image and the in-context example and a Matching Transformer that uses fixed matching and a Hungarian algorithm to eliminate differences between different tasks. In addition, we have further perfected the current evaluation system for in-context image segmentation, aiming to facilitate a holistic appraisal of these models. Experiments on various segmentation tasks show the effectiveness of the proposed method. Yang Liu 0357, Chenchen Jing, Hengtao Li, Muzhi Zhu, Hao Chen 0041, Chunhua Shen |
NeurIPS | 5 |
| 2024 | Unleashing the Potential of the Diffusion Model in Few-shot Semantic SegmentationabstractThe Diffusion Model has not only garnered noteworthy achievements in the realm of image generation
but has also demonstrated its potential as an effective pretraining method utilizing unlabeled data.
Drawing from the extensive potential unveiled by the Diffusion Model in both semantic correspondence and open vocabulary segmentation, our work initiates an investigation into employing the Latent Diffusion Model for Few-shot Semantic Segmentation.
Recently, inspired by the in-context learning ability of large language models, Few-shot Semantic Segmentation has evolved into In-context Segmentation tasks, morphing into a crucial element in assessing generalist segmentation models.
In this context, we concentrate
on Few-shot Semantic Segmentation,
establishing a solid foundation for the future development of a Diffusion-based generalist model for segmentation. Our initial focus lies in understanding how to facilitate interaction between the query image and the support image, resulting in the proposal of a KV fusion method within the self-attention framework.
Subsequently, we delve deeper into optimizing the infusion of information from the support mask and simultaneously re-evaluating how to provide reasonable supervision from the query mask.
Based on our analysis, we establish a simple and effective framework named DiffewS, maximally retaining the original Latent Diffusion Model's generative framework and effectively utilizing the pre-training prior. Experimental results demonstrate that our method significantly outperforms the previous SOTA models in multiple settings. Muzhi Zhu, Yang Liu 0357, Zekai Luo, Chenchen Jing, Hao Chen 0041, Guangkai Xu, Chunhua Shen |
NeurIPS | 5 |
| 2024 | Metric3D v2: A Versatile Monocular Geometric Foundation Model for Zero-Shot Metric Depth and Surface Normal EstimationabstractWe introduce Metric3D v2, a geometric foundation model designed for zero-shot metric depth and surface normal estimation from single images, critical for accurate 3D recovery. Depth and normal estimation, though complementary, present distinct challenges. State-of-the-art monocular depth methods achieve zero-shot generalization through affine-invariant depths, but fail to recover real-world metric scale. Conversely, current normal estimation techniques struggle with zero-shot performance due to insufficient labeled data. We propose targeted solutions for both metric depth and normal estimation. For metric depth, we present a canonical camera space transformation module that resolves metric ambiguity across various camera models and large-scale datasets, which can be easily integrated into existing monocular models. For surface normal estimation, we introduce a joint depth-normal optimization module that leverages diverse data from metric depth, allowing normal estimators to improve beyond traditional labels. Our model, trained on over 16 million images from thousands of camera models with varied annotations, excels in zero-shot generalization to new camera settings. As shown in Fig. 1, It ranks the 1st in multiple zero-shot and standard benchmarks for metric depth and surface normal prediction. Our method enables the accurate recovery of metric 3D structures on randomly collected internet images, paving the way for plausible single-image metrology. Our model also relieves the scale drift issues of monocular-SLAM (Fig. 3), leading to high-quality metric scale dense mapping. Such applications highlight the versatility of Metric3D v2 models as geometric foundation models. Mu Hu, Wei Yin 0006, Chi Zhang 0007, Zhipeng Cai 0003, Xiaoxiao Long, Hao Chen 0041, Gang Yu 0002, Chunhua Shen, Shaojie Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | AODet: Aerial Object Detection Using Transformers for Foreground RegionsabstractAerial object detection is an important task and has received significant attention in recent years. Aerial images typically depict small and sparse instances against a simple background. Nevertheless, the simple background can only provide limited information. Based on the observation, we present a new transformer-based framework for aerial object detection. In contrast to previous methods that address sparsity through multi-stage pipelines involving Region-of-Interest (RoI) techniques or Sparse Convolutions, our method, referred as AODet, enjoy two significant advantages: 1) AODet is a simple yet accurate object detector which is specialized for aerial object detection. AODet identifies the background regions earlier and then only operates on the regions which most likely include the foreground objects, thereby significantly reducing the redundant computations. The utilization of transformer exploits more context information between foreground regions, helping to retain high-quality detection results. 2) Instead of involving the sparse operations like Sparse Convolutions or Clustering algorithms/ROI operations, AODet employs transformer to detect objects from foreground proposals. Our approach is simpler and can be easily implemented with simple tensor manipulations. Extensive experiments have conducted on VisDrone and DOTA. AODet achieves 40.9 AP on Visdrone and 79.6 mAP DOTA, demonstrating the effectiveness of AODet. Xiaoming Wang 0010, Hao Chen 0041, Xiangxiang Chu, Peng Wang 0015 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | CI3D: Context Interaction for Dynamic Objects and Static Map Elements in 3D Driving ScenesabstractMulti-view 3D visual perception including 3D object detection and Birds'-eye-view (BEV) map segmentation is essential for autonomous driving. However, there has been little discussion about 3D context attention between dynamic objects and static elements with multi-view camera inputs, due to the challenging nature of recovering the 3D spatial information from images and performing effective 3D context interaction. 3D context information is expected to provide more cues to enhance 3D visual perception for autonomous driving. We thus propose a new transformer-based framework named CI3D in an attempt to implicitly model 3D context interaction between dynamic objects and static map elements. To achieve this, we use dynamic object queries and static map queries to gather information from multi-view image features, which are represented sparsely in 3D space. Moreover, a dynamic 3D position encoder is utilized to precisely generate queries' positional embeddings. With accurate positional embeddings, the queries effectively aggregate 3D context information via a multi-head attention mechanism to model 3D context interaction. We further reveal that sparse supervision signals from the limited number of queries result in the issue of rough and vague image features. To overcome this challenge, we introduce a panoptic segmentation head as an auxiliary task and a 3D-to-2D deformable cross-attention module, greatly enhancing the robustness of spatial feature learning and sampling. Our approach has been extensively evaluated on two large-scale datasets, nuScenes and Waymo, and significantly outperforms the baseline method on both benchmarks. Feipeng Cai, Hao Chen 0041, Liuyuan Deng |
IEEE Trans. Image Process. | 2 |
| 2024 | Target Before Shooting: Accurate Anomaly Detection and Localization Under One Millisecond via Cascade Patch RetrievalabstractIn this work, by re-examining the "matching" nature of Anomaly Detection (AD), we propose a novel AD framework that simultaneously enjoys new records of AD accuracy and dramatically high running speed. In this framework, the anomaly detection problem is solved via a cascade patch retrieval procedure that retrieves the nearest neighbors for each test image patch in a coarse-to-fine fashion. Given a test sample, the top-K most similar training images are first selected based on a robust histogram matching process. Secondly, the nearest neighbor of each test patch is retrieved over the similar geometrical locations on those "most similar images", by using a carefully trained local metric. Finally, the anomaly score of each test image patch is calculated based on the distance to its "nearest neighbor" and the "non-background" probability. The proposed method is termed "Cascade Patch Retrieval" (CPR) in this work. Different from the previous patch-matching-based AD algorithms, CPR selects proper "targets" (reference images and patches) before "shooting" (patch-matching). On the well-acknowledged MVTec AD, BTAD and MVTec-3D AD datasets, the proposed algorithm consistently outperforms all the comparing SOTA methods by remarkable margins, measured by various AD metrics. Furthermore, CPR is extremely efficient. It runs at the speed of 113 FPS with the standard setting while its simplified version only requires less than 1 ms to process an image at the cost of a trivial accuracy drop. The code of CPR is available at https://github.com/flyinghu123/CPR. Jianfei Hu, Bo Li 0090, Hao Chen 0041, Yongbin Zheng, Chunhua Shen |
IEEE Trans. Image Process. | 4 |
| 2024 | DualGCN: Exploring Syntactic and Semantic Information for Aspect-Based Sentiment AnalysisabstractThe task of aspect-based sentiment analysis aims to identify sentiment polarities of given aspects in a sentence. Recent advances have demonstrated the advantage of incorporating the syntactic dependency structure with graph convolutional networks (GCNs). However, their performance of these GCN-based methods largely depends on the dependency parsers, which would produce diverse parsing results for a sentence. In this article, we propose a dual GCN (DualGCN) that jointly considers the syntax structures and semantic correlations. Our DualGCN model mainly comprises four modules: 1) SynGCN: instead of explicitly encoding syntactic structure, the SynGCN module uses the dependency probability matrix as a graph structure to implicitly integrate the syntactic information; 2) SemGCN: we design the SemGCN module with multihead attention to enhance the performance of the syntactic structure with the semantic information; 3) Regularizers: we propose orthogonal and differential regularizers to precisely capture semantic correlations between words by constraining attention scores in the SemGCN module; and 4) Mutual BiAffine: we use the BiAffine module to bridge relevant information between the SynGCN and SemGCN modules. Extensive experiments are conducted compared with up-to-date pretrained language encoders on two groups of datasets, one including Restaurant14, Laptop14, and Twitter and the other including Restaurant15 and Restaurant16. The experimental results demonstrate that the parsing results of various dependency parsers affect their performance of the GCN-based models. Our DualGCN model achieves superior performance compared with the state-of-the-art approaches. The source code and preprocessed datasets are provided and publicly available on GitHub (see https://github.com/CCChenhao997/DualGCN-ABSA). Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | USSA: A Unified Table Filling Scheme for Structured Sentiment AnalysisabstractMost previous studies on Structured Sentiment Analysis (SSA) have cast it as a problem of bi-lexical dependency parsing, which cannot address issues of overlap and discontinuity simultaneously.In this paper, we propose a nichetargeting and effective solution.Our approach involves creating a novel bi-lexical dependency parsing graph, which is then converted to a unified 2D table-filling scheme, namely USSA.The proposed scheme resolves the kernel bottleneck of previous SSA methods by utilizing 13 different types of relations.In addition, to closely collaborate with the USSA scheme, we have developed a model that includes a proposed bi-axial attention module to effectively capture the correlations among relations in the rows and columns of the table.Extensive experimental results on benchmark datasets demonstrate the effectiveness and robustness of our proposed framework, outperforming state-ofthe-art methods consistently 1 . Zepeng Zhai, Hao Chen 0041, Ruifan Li, Xiaojie Wang 0006 |
ACL (1) | 2 |
| 2023 | Learning to Fuse Monocular and Multi-view Cues for Multi-frame Depth Estimation in Dynamic ScenesabstractMulti-frame depth estimation generally achieves high accuracy relying on the multi-view geometric consistency. When applied in dynamic scenes, e.g., autonomous driving, this consistency is usually violated in the dynamic areas, leading to corrupted estimations. Many multi-frame methods handle dynamic areas by identifying them with explicit masks and compensating the multi-view cues with monocular cues represented as local monocular depth or features. The improvements are limited due to the uncontrolled quality of the masks and the underutilized benefits of the fusion of the two types of cues. In this paper, we propose a novel method to learn to fuse the multi-view and monocular cues encoded as volumes without needing the heuristically crafted masks. As unveiled in our analyses, the multiview cues capture more accurate geometric information in static areas, and the monocular cues capture more useful contexts in dynamic areas. To let the geometric perception learned from multi-view cues in static areas propagate to the monocular representation in dynamic areas and let monocular cues enhance the representation of multi-view cost volume, we propose a cross-cue fusion (CCF) module, which includes the cross-cue attention (CCA) to encode the spatially non-local relative intra-relations from each source to enhance the representation of the other. Experiments on real-world datasets prove the significant effectiveness and generalization ability of the proposed method. Rui Li 0013, Dong Gong, Wei Yin 0006, Hao Chen 0041, Yu Zhu 0004, Xiaozhi Chen, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 4 |
| 2023 | Learning Conditional Attributes for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to train models to recognize novel compositional concepts based on learned concepts such as attribute-object combinations. One of the challenges is to model attributes interacted with different objects, e.g., the attribute “wet” in “wet apple” and “wet cat” is different. As a solution, we provide analysis and argue that attributes are conditioned on the recognized object and input image and explore learning conditional attribute embeddings by a proposed attribute learning framework containing an attribute hyper learner and an attribute base learner. By encoding conditional attributes, our model enables to generate flexible attribute embeddings for generalization from seen to unseen compositions. Experiments on CZSL benchmarks, including the more challenging C-GQA dataset, demonstrate better performances compared with other state-of-the-art approaches and validate the importance of learning conditional attributes. Code‡1Gllee:https://gitee.com/wqshmzh/canet-czsl is available at https://github.com/wqshmzh/CANet-CZSL. Qingsheng Wang, Lingqiao Liu, Chenchen Jing, Hao Chen 0041, Guoqiang Liang 0001, Peng Wang 0015, Chunhua Shen |
CVPR | 4 |
| 2023 | Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageabstractReconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are built upon multi-view geometry. State-of-the-art (SOTA) monocular metric depth estimation methods can only handle a single camera model and are unable to perform mixed-data training due to metric ambiguity. Meanwhile, SOTA monocular methods trained on large mixed datasets achieve zero-shot generalization by learning affine-invariant depths, which cannot recover real-world metrics. In this work, we show that the key to a zero-shot single-view metric depth model lies in the combination of large-scale data training and resolving the metric ambiguity from various camera models. We propose a canonical camera space transformation module, which explicitly addresses the ambiguity problems and can be effortlessly plugged into existing monocular models. Equipped with our module, monocular models can be stably trained over 8 millions of images with thousands of camera models, resulting in zero-shot generalization to in-the-wild images with unseen camera settings. Experiments demonstrate SOTA performance of our method on 7 zero-shot benchmarks. Notably, our method won the championship in the 2nd Monocular Depth Estimation Challenge. Our method enables the accurate recovery of metric 3D structures on randomly collected internet images, paving the way for plausible single-image metrology. The potential benefits extend to downstream tasks, which can be significantly improved by simply plugging in our model. For example, our model relieves the scale drift issues of monocular-SLAM (Fig. 1), leading to high-quality metric scale dense mapping. The code is available at https://github.com/YvanYin/Metric3D. Wei Yin 0006, Chi Zhang 0007, Hao Chen 0041, Zhipeng Cai 0003, Gang Yu 0002, Xiaozhi Chen, Chunhua Shen |
ICCV | 3 |
| 2023 | FrozenRecon: Pose-free 3D Scene Reconstruction with Frozen Depth Modelsabstract3D scene reconstruction is a long-standing vision task. Existing approaches can be categorized into geometry-based and learning-based methods. The former leverages multi-view geometry but may face catastrophic failures due to the reliance on accurate pixel correspondence across views, while the latter mitigates these issues by learning 2D or 3D representation directly. However, without a largescale video or 3D training data, it can hardly be generalized to diverse real-world scenarios due to the presence of tens of millions or even billions of optimization parameters in the deep network.Recently, robust monocular depth estimation models trained with large-scale datasets have been proven to possess weak 3D geometry prior, but they are insufficient for reconstruction due to the unknown camera parameters, the affine-invariant property, and inter-frame inconsistency. To address these issues, we propose a novel test-time optimization approach that can transfer the robustness of affine- invariant depth models such as LeReS to challenging diverse scenes while ensuring inter-frame consistency, with only dozens of parameters to optimize per video frame. Specifically, our approach involves freezing the pre-trained affine-invariant depth model’s depth predictions, rectifying them by optimizing the unknown scale-shift values with a geometric consistency alignment module, and employing the resulting scale-consistent depth maps to robustly obtain camera poses and achieve dense scene reconstruction, even in low-texture regions. Experiments show that our method achieves state-of-the-art cross-dataset reconstruction on five zero-shot testing datasets. Code is available at: https://aim-uofa.github.io/FrozenRecon/ Guangkai Xu, Wei Yin 0006, Hao Chen 0041, Chunhua Shen, Feng Zhao 0004 |
ICCV | 3 |
| 2023 | CTVIS: Consistent Training for Online Video Instance SegmentationabstractThe discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/negative embeddings. Recent online VIS methods leverage CIs sourced from one reference frame only, which we argue is insufficient for learning highly discriminative embeddings. Intuitively, a possible strategy to enhance CIs is replicating the inference phase during training. To this end, we propose a simple yet effective training strategy, called Consistent Training for Online VIS (CTVIS), which devotes to aligning the training and inference pipelines in terms of building CIs. Specifically, CTVIS constructs CIs by referring inference the momentum-averaged embedding and the memory bank storage mechanisms, and adding noise to the relevant embeddings. Such an extension allows a reliable comparison between embeddings of current instances and the stable representations of historical instances, thereby conferring an advantage in modeling VIS challenges such as occlusion, re-identification, and deformation. Empirically, CTVIS outstrips the SOTA VIS models by up to +5.0 points on three VIS benchmarks, including YTVIS19 (55.1% AP), YTVIS21 (50.1% AP) and OVIS (35.5% AP). Furthermore, we find that pseudo-videos transformed from images can train robust models surpassing fully-supervised ones. Kaining Ying, Weian Mao, Zhenhua Wang 0003, Hao Chen 0041, Lin Wu 0001, Yifan Liu 0001, Chengxiang Fan, Yunzhi Zhuge, Chunhua Shen |
ICCV | 5 |
| 2023 | SegPrompt: Boosting Open-world Segmentation via Category-level Prompt LearningabstractCurrent closed-set instance segmentation models rely on pre-defined class labels for each mask during training and evaluation, largely limiting their ability to detect novel objects. Open-world instance segmentation (OWIS) models address this challenge by detecting unknown objects in a class-agnostic manner. However, previous OWIS approaches completely erase category information during training to keep the model’s ability to generalize to unknown objects. In this work, we propose a novel training mechanism termed SegPrompt that uses category information to improve the model’s class-agnostic segmentation ability for both known and unknown categories. In addition, the previous OWIS training setting exposes the unknown classes to the training set and brings information leakage, which is unreasonable in the real world. Therefore, we provide a new open-world benchmark closer to a real-world scenario by dividing the dataset classes into known-seen-unseen parts. For the first time, we focus on the model’s ability to discover objects that never appear in the training set images.Experiments show that SegPrompt can improve the overall and unseen detection performance by 5.6% and 6.1% in AR on our new benchmark without affecting the inference efficiency. We further demonstrate the effectiveness of our method on existing cross-dataset transfer and strongly supervised settings, leading to 5.5% and 12.3% relative improvement. Code and data are released at: https://github.com/aim-uofa/SegPrompt Muzhi Zhu, Hengtao Li, Hao Chen 0041, Chengxiang Fan, Weian Mao, Chenchen Jing, Yifan Liu 0001, Chunhua Shen |
ICCV | 3 |
| 2023 | DatasetDM: Synthesizing Data with Perception Annotations Using Diffusion ModelsabstractCurrent deep networks are very data-hungry and benefit from training on large-scale datasets, which are often time-consuming to collect and annotate. By contrast, synthetic data can be generated infinitely using generative models such as DALL-E and diffusion models, with minimal effort and cost. In this paper, we present DatasetDM, a generic dataset generation model that can produce diverse synthetic
images and the corresponding high-quality perception annotations (e.g., segmentation masks, and depth). Our method builds upon the pre-trained diffusion model and extends text-guided image synthesis to perception data generation. We show that the rich latent code of the diffusion model can be effectively decoded as accurate perception annotations using a decoder module. Training the decoder only needs less than 1% (around 100 images) of manually labeled images, enabling the generation of an infinitely large annotated dataset. Then these synthetic data can be used for training various perception models on downstream tasks. To showcase the power of the proposed approach, we generate datasets with rich dense pixel-wise labels for a wide range of downstream tasks, including semantic15
segmentation, instance segmentation, and depth estimation. Notably, it achieves 1) state-of-the-art results on semantic segmentation and instance segmentation; 2) significantly more efficient and robust in domain generalization than the real data; 3) state-of-the-art results in zero-shot segmentation setting; and 4) flexibility for efficient application and novel task composition (e.g., image editing) Weijia Wu 0001, Yuzhong Zhao, Hao Chen 0041, Yuchao Gu, Rui Zhao 0001, Yefei He, Zheng Shou 0001, Chunhua Shen |
NeurIPS | 3 |
| 2023 | A Dynamic Feature Interaction Framework for Multi-task Visual Perception
Yuling Xi, Hao Chen 0041, Ning Wang 0020, Peng Wang 0015, Yanning Zhang 0001, Chunhua Shen, Yifan Liu 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | Instance and Panoptic Segmentation Using Conditional ConvolutionsabstractWe propose a simple yet effective framework for instance and panoptic segmentation, termed CondInst (conditional convolutions for instance and panoptic segmentation). In the literature, top-performing instance segmentation methods typically follow the paradigm of Mask R-CNN and rely on ROI operations (typically ROIAlign) to attend to each instance. In contrast, we propose to attend to the instances with dynamic conditional convolutions. Instead of using instance-wise ROIs as inputs to the instance mask head of fixed weights, we design dynamic instance-aware mask heads, conditioned on the instances to be predicted. CondInst enjoys three advantages: 1) Instance and panoptic segmentation are unified into a fully convolutional network, eliminating the need for ROI cropping and feature alignment. 2) The elimination of the ROI cropping also significantly improves the output instance mask resolution. 3) Due to the much improved capacity of dynamically-generated conditional convolutions, the mask head can be very compact (e.g., 3 conv. layers, each having only 8 channels), leading to significantly faster inference time per instance and making the overall inference time less relevant to the number of instances. We demonstrate a simpler method that can achieve improved accuracy and inference speed on both instance and panoptic segmentation tasks. On the COCO dataset, we outperform a few state-of-the-art methods. We hope that CondInst can be a strong baseline for instance and panoptic segmentation. Code is available at: https://git.io/AdelaiDet. Zhi Tian, Bowen Zhang 0009, Hao Chen 0041, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Enhanced Multi-Channel Graph Convolutional Network for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE) is an emerging sentiment analysis task.Most of the existing studies focus on devising a new tagging scheme that enables the model to extract the sentiment triplets in an end-to-end fashion.However, these methods ignore the relations between words for ASTE task.In this paper, we propose an Enhanced Multi-Channel Graph Convolutional Network model (EMC-GCN) to fully utilize the relations between words.Specifically, we first define ten types of relations for ASTE task, and then adopt a biaffine attention module to embed these relations as an adjacent tensor between words in a sentence.After that, our EMC-GCN transforms the sentence into a multi-channel graph by treating words and the relation adjacent tensor as nodes and edges, respectively.Thus, relationaware node representations can be learnt.Furthermore, we consider diverse linguistic features to enhance our EMC-GCN model.Finally, we design an effective refining strategy on EMC-GCN for word-pair representation refinement, which considers the implicit results of aspect and opinion extraction when determining whether word pairs match or not.Extensive experimental results on the benchmark datasets demonstrate that the effectiveness and robustness of our proposed model, which outperforms state-of-the-art methods significantly. Hao Chen 0041, Zepeng Zhai, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006 |
ACL (1) | 1 |
| 2022 | Dual Decision Improves Open-Set Panoptic Segmentation
Hao Chen 0041, Lingqiao Liu, Yufei Yin |
BMVC | 2 |
| 2022 | COM-MRC: A COntext-Masked Machine Reading Comprehension Framework for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE) aims to extract sentiment triplets from sentences, which was recently formalized as an effective machine reading comprehension (MRC) based framework.However, when facing multiple aspect terms, the MRC-based methods could fail due to the interference from other aspect terms.In this paper, we propose a novel COntext-Masked MRC (COM-MRC) framework for ASTE.Our COM-MRC framework comprises three closely-related components: a context augmentation strategy, a discriminative model, and an inference method.Specifically, a context augmentation strategy is designed by enumerating all masked contexts for each aspect term.The discriminative model comprises four modules, i.e., aspect and opinion extraction modules, sentiment classification and aspect detection modules.In addition, a two-stage inference method first extracts all aspects and then identifies their opinions and sentiment through iteratively masking the aspects.Extensive experimental results on benchmark datasets show the effectiveness of our proposed COM-MRC framework, which outperforms state-of-the-art methods consistently 1 . Zepeng Zhai, Hao Chen 0041, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006 |
EMNLP | 2 |
| 2022 | Memory-Efficient Hierarchical Neural Architecture Search for Image Restoration
Haokui Zhang, Ying Li 0017, Hao Chen 0041, Chengrong Gong, Zongwen Bai, Chunhua Shen |
Int. J. Comput. Vis. | 3 |
| 2022 | ABCNet v2: Adaptive Bezier-Curve Network for Real-Time End-to-End Text SpottingabstractEnd-to-end text-spotting, which aims to integrate detection and recognition in a unified framework, has attracted increasing attention due to its simplicity of the two complimentary tasks. It remains an open problem especially when processing arbitrarily-shaped text instances. Previous methods can be roughly categorized into two groups: character-based and segmentation-based, which often require character-level annotations and/or complex post-processing due to the unstructured output. Here, we tackle end-to-end text spotting by presenting Adaptive Bezier Curve Network v2 (ABCNet v2). Our main contributions are four-fold: 1) For the first time, we adaptively fit arbitrarily-shaped text by a parameterized Bezier curve, which, compared with segmentation-based methods, can not only provide structured output but also controllable representation. 2) We design a novel BezierAlign layer for extracting accurate convolution features of a text instance of arbitrary shapes, significantly improving the precision of recognition over previous methods. 3) Different from previous methods, which often suffer from complex post-processing and sensitive hyper-parameters, our ABCNet v2 maintains a simple pipeline with the only post-processing non-maximum suppression (NMS). 4) As the performance of text recognition closely depends on feature alignment, ABCNet v2 further adopts a simple yet effective coordinate convolution to encode the position of the convolutional filters, which leads to a considerable improvement with negligible computation overhead. Comprehensive experiments conducted on various bilingual (English and Chinese) benchmark datasets demonstrate that ABCNet v2 can achieve state-of-the-art performance while maintaining very high efficiency. More importantly, as there is little work on quantization of text spotting models, we quantize our models to improve the inference time of the proposed ABCNet v2. This can be valuable for real-time applications. Code and model are available at: https://git.io/AdelaiDet. Chunhua Shen, Tong He 0001, Peng Chen 0037, Chongyu Liu, Hao Chen 0041 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | FCOS: A Simple and Strong Anchor-Free Object DetectorabstractIn computer vision, object detection is one of most important tasks, which underpins a few instance-level recognition tasks and many downstream applications. Recently one-stage methods have gained much attention over two-stage approaches due to their simpler design and competitive performance. Here we propose a fully convolutional one-stage object detector (FCOS) to solve object detection in a per-pixel prediction fashion, analogue to other dense prediction problems such as semantic segmentation. Almost all state-of-the-art object detectors such as RetinaNet, SSD, YOLOv3, and Faster R-CNN rely on pre-defined anchor boxes. In contrast, our proposed detector FCOS is anchor box free, as well as proposal free. By eliminating the pre-defined set of anchor boxes, FCOS completely avoids the complicated computation related to anchor boxes such as calculating the intersection over union (IoU) scores during training. More importantly, we also avoid all hyper-parameters related to anchor boxes, which are often sensitive to the final detection performance. With the only post-processing non-maximum suppression (NMS), we demonstrate a much simpler and flexible detection framework achieving improved detection accuracy. We hope that the proposed FCOS framework can serve as a simple and strong alternative for many other instance-level tasks. Code is available at: git.io/AdelaiDet. Zhi Tian, Chunhua Shen, Hao Chen 0041, Tong He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Dual Graph Convolutional Networks for Aspect-based Sentiment AnalysisabstractRuifan Li, Hao Chen, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang, Eduard Hovy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy |
ACL/IJCNLP (1) | 2 |
| 2021 | Generic Perceptual Loss for Modeling Structured Output DependenciesabstractThe perceptual loss has been widely used as an effective loss term in image synthesis tasks including image super-resolution [16], and style transfer [14]. It was believed that the success lies in the high-level perceptual feature representations extracted from CNNs pretrained with a large set of images. Here we reveal that, what matters is the network structure instead of the trained weights. Without any learning, the structure of a deep network is sufficient to capture the dependencies between multiple levels of variable statistics using multiple layers of CNNs. This insight removes the requirements of pre-training and a particular network structure (commonly, VGG) that are previously assumed for the perceptual loss, thus enabling a significantly wider range of applications. To this end, we demonstrate that a randomly-weighted deep CNN can be used to model the structured dependencies of outputs. On a few dense per-pixel prediction tasks such as semantic segmentation, depth estimation and instance segmentation, we show improved results of using the extended randomized perceptual loss, compared to the baselines using pixel-wise loss alone. We hope that this simple, extended perceptual loss may serve as a generic structured-output loss that is applicable to most structured output learning tasks. Yifan Liu 0001, Hao Chen 0041, Yu Chen 0037, Wei Yin 0006, Chunhua Shen |
CVPR | 2 |
| 2021 | BoxInst: High-Performance Instance Segmentation With Box AnnotationsabstractWe present a high-performance method that can achieve mask-level instance segmentation with only bounding-box annotations for training. While this setting has been studied in the literature, here we show significantly stronger performance with a simple design (e.g., dramatically improving previous best reported mask AP of 21.1% [13] to 31.6% on the COCO dataset). Our core idea is to redesign the loss of learning masks in instance segmentation, with no modification to the segmentation network itself. The new loss functions can supervise the mask training without relying on mask annotations. This is made possible with two loss terms, namely, 1) a surrogate term that minimizes the discrepancy between the projections of the ground-truth box and the predicted mask; 2) a pairwise loss that can exploit the prior that proximal pixels with similar colors are very likely to have the same category label.Experiments demonstrate that the redesigned mask loss can yield surprisingly high-quality instance masks with only box annotations. For example, without using any mask annotations, with a ResNet-101 backbone and 3× training schedule, we achieve 33.2% mask AP on COCO test-dev split (vs. 39.1% of the fully supervised counterpart). Our excellent experiment results on COCO and Pascal VOC indicate that our method dramatically narrows the performance gap between weakly and fully supervised instance segmentation.Code is available at: https://git.io/AdelaiDet Zhi Tian, Chunhua Shen, Hao Chen 0041 |
CVPR | 4 |
| 2021 | Exploring the Capacity of an Orderless Box Discretization Network for Multi-orientation Scene Text Detection
Tong He 0001, Hao Chen 0041, Xinyu Wang 0010, Canjie Luo, Shuaitao Zhang, Chunhua Shen |
Int. J. Comput. Vis. | 3 |
| 2021 | NAS-FCOS: Efficient Search for Object Detection Architectures
Ning Wang 0020, Yang Gao 0001, Hao Chen 0041, Peng Wang 0015, Zhi Tian, Chunhua Shen, Yanning Zhang 0001 |
Int. J. Comput. Vis. | 3 |
| 2021 | Learning deep part-aware embedding for person retrieval
Yang Zhao 0019, Chunhua Shen, Xiaohan Yu 0001, Hao Chen 0041, Yongsheng Gao 0001, Shengwu Xiong 0001 |
Pattern Recognit. | 4 |
| 2020 | BlendMask: Top-Down Meets Bottom-Up for Instance SegmentationabstractInstance segmentation is one of the fundamental vision tasks. Recently, fully convolutional instance segmentation methods have drawn much attention as they are often simpler and more efficient than two-stage approaches like Mask R-CNN. To date, almost all such approaches fall behind the two-stage Mask R-CNN method in mask precision when models have similar computation complexity, leaving great room for improvement. In this work, we achieve improved mask prediction by effectively combining instance-level information with semantic information with lower-level fine-granularity. Our main contribution is a blender module which draws inspiration from both top-down and bottom-up instance segmentation approaches. The proposed BlendMask can effectively predict dense per-pixel position-sensitive instance features with very few channels, and learn attention maps for each instance with merely one convolution layer, thus being fast in inference. BlendMask can be easily incorporate with the state-of-the-art one-stage detection frameworks and outperforms Mask R-CNN under the same training schedule while being faster. A light-weight version of BlendMask achieves 36.0 mAP at 27 FPS evaluated on a single 1080Ti. Because of its simplicity and efficacy, we hope that our BlendMask could serve as a simple yet strong baseline for a wide range of instance-wise prediction tasks. Hao Chen 0041, Kunyang Sun, Zhi Tian, Chunhua Shen, Yongming Huang 0001, Youliang Yan |
CVPR | 1 |
| 2020 | ABCNet: Real-Time Scene Text Spotting With Adaptive Bezier-Curve NetworkabstractScene text detection and recognition has received increasing research attention. Existing methods can be roughly categorized into two groups: character-based and segmentation-based. These methods either are costly for character annotation or need to maintain a complex pipeline, which is often not suitable for real-time applications. Here we address the problem by proposing the Adaptive Bezier-Curve Network (\BeCan). Our contributions are three-fold: 1) For the first time, we adaptively fit oriented or curved text by a parameterized Bezier curve. 2) We design a novel BezierAlign layer for extracting accurate convolution features of a text instance with arbitrary shapes, significantly improving the precision compared with previous methods. 3) Compared with standard bounding box detection, our Bezier curve detection introduces negligible computation overhead, resulting in superiority of our method in both efficiency and accuracy. Experiments on oriented or curved benchmark datasets, namely Total-Text and CTW1500, demonstrate that \BeCan achieves state-of-the-art accuracy, meanwhile significantly improving the speed. In particular, on Total-Text, our real-time version is over 10 times faster than recent state-of-the-art methods with a competitive recognition accuracy. Code is available at \url{https://git.io/AdelaiDet}. Hao Chen 0041, Chunhua Shen, Tong He 0001 |
CVPR | 2 |
| 2020 | NAS-FCOS: Fast Neural Architecture Search for Object DetectionabstractThe success of deep neural networks relies on significant architecture engineering. Recently neural architecture search (NAS) has emerged as a promise to greatly reduce manual effort in network design by automatically searching for optimal architectures, although typically such algorithms need an excessive amount of computational resources, e.g., a few thousand GPU-days. To date, on challenging vision tasks such as object detection, NAS, especially fast versions of NAS, is less studied. Here we propose to search for the decoder structure of object detectors with search efficiency being taken into consideration. To be more specific, we aim to efficiently search for the feature pyramid network (FPN) as well as the prediction head of a simple anchor-free object detector, namely FCOS, using a tailored reinforcement learning paradigm. With carefully designed search space, search algorithms and strategies for evaluating network quality, we are able to efficiently search a top-performing detection architecture within 4 days using 8 V100 GPUs. The discovered architecture surpasses state-of-the-art object detection models (such as Faster R-CNN, RetinaNet and FCOS) by 1.5 to 3.5 points in AP on the COCO dataset, with comparable computation complexity and memory footprint, demonstrating the efficacy of the proposed NAS for object detection. Ning Wang 0020, Yang Gao 0001, Hao Chen 0041, Peng Wang 0015, Zhi Tian, Chunhua Shen, Yanning Zhang 0001 |
CVPR | 3 |
| 2020 | Memory-Efficient Hierarchical Neural Architecture Search for Image DenoisingabstractRecently, neural architecture search (NAS) methods have attracted much attention and outperformed manually designed architectures on a few high-level vision tasks. In this paper, we propose HiNAS (Hierarchical NAS), an effort towards employing NAS to automatically design effective neural network architectures for image denoising. HiNAS adopts gradient based search strategies and employs operations with adaptive receptive field to build an flexible hierarchical search space. During the search stage, HiNAS shares cells across different feature levels to save memory and employ an early stopping strategy to avoid the collapse issue in NAS, and considerably accelerate the search speed. The proposed HiNAS is both memory and computation efficient, which takes only about 4.5 hours for searching using a single GPU. We evaluate the effectiveness of our proposed HiNAS on two different datasets, namely an additive white Gaussian noise dataset BSD500, and a realistic noise dataset SIM1800. Experimental results show that the architecture found by HiNAS has fewer parameters and enjoys a faster inference speed, while achieving highly competitive performance compared with state-of-the-art methods. We also present analysis on the architectures found by NAS. HiNAS also shows good performance on experiments for image de-raining. Haokui Zhang, Ying Li 0017, Hao Chen 0041, Chunhua Shen |
CVPR | 3 |
| 2020 | Conditional Convolutions for Instance Segmentation
Zhi Tian, Chunhua Shen, Hao Chen 0041 |
ECCV (1) | 3 |
| 2020 | Architecture Search of Dynamic Cells for Semantic Video SegmentationabstractIn semantic video segmentation the goal is to acquire consistent dense semantic labelling across image frames. To this end, recent approaches have been reliant on manually arranged operations applied on top of static semantic segmentation networks - with the most prominent building block being the optical flow able to provide information about scene dynamics. Related to that is the line of research concerned with speeding up static networks by approximating expensive parts of them with cheaper alternatives, while propagating information from previous frames. In this work we attempt to come up with generalisation of those methods, and instead of manually designing contextual blocks that connect per-frame outputs, we propose a neural architecture search solution, where the choice of operations together with their sequential arrangement are being predicted by a separate neural network. We showcase that such generalisation leads to stable and accurate results across common benchmarks, such as CityScapes and CamVid datasets. Importantly, the proposed methodology takes only 2 GPU-days, finds high-performing cells and does not rely on the expensive optical flow computation. Vladimir Nekrasov, Hao Chen 0041, Chunhua Shen, Ian D. Reid 0001 |
WACV | 2 |
| 2020 | Adversarial Learning of Structure-Aware Fully Convolutional Networks for Landmark LocalizationabstractLandmark/pose estimation in single monocular images has received much effort in computer vision due to its important applications. It remains a challenging task when input images come with severe occlusions caused by, e.g., adverse camera views. Under such circumstances, biologically implausible pose predictions may be produced. In contrast, human vision is able to predict poses by exploiting geometric constraints of landmark point inter-connectivity. To address the problem, by incorporating priors about the structure of pose components, we propose a novel structure-aware fully convolutional network to implicitly take such priors into account during training of the deep network. Explicit learning of such constraints is typically challenging. Instead, inspired by how human identifies implausible poses, we design discriminators to distinguish the real poses from the fake ones (such as biologically implausible ones). If the pose generator G generates results that the discriminator fails to distinguish from real ones, the network successfully learns the priors. Training of the network follows the strategy of conditional Generative Adversarial Networks (GANs). The effectiveness of the proposed network is evaluated on three pose-related tasks: 2D human pose estimation, 2D facial landmark estimation and 3D human pose estimation. The proposed approach significantly outperforms several state-of-the-art methods and almost always generates plausible pose predictions, demonstrating the usefulness of implicit learning of structures using GANs. Yu Chen 0037, Chunhua Shen, Hao Chen 0041, Xiu-Shen Wei, Lingqiao Liu, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Fast Neural Architecture Search of Compact Semantic Segmentation Models via Auxiliary CellsabstractAutomated design of neural network architectures tailored for a specific task is an extremely promising, albeit inherently difficult, avenue to explore. While most results in this domain have been achieved on image classification and language modelling problems, here we concentrate on dense per-pixel tasks, in particular, semantic image segmentation using fully convolutional networks. In contrast to the aforementioned areas, the design choices of a fully convolutional network require several changes, ranging from the sort of operations that need to be used - e.g., dilated convolutions - to a solving of a more difficult optimisation problem. In this work, we are particularly interested in searching for high-performance compact segmentation architectures, able to run in real-time using limited resources. To achieve that, we intentionally over-parameterise the architecture during the training time via a set of auxiliary cells that provide an intermediate supervisory signal and can be omitted during the evaluation phase. The design of the auxiliary cell is emitted by a controller, a neural network with the fixed structure trained using reinforcement learning. More crucially, we demonstrate how to efficiently search for these architectures within limited time and computational budgets. In particular, we rely on a progressive strategy that terminates non-promising architectures from being further trained, and on Polyak averaging coupled with knowledge distillation to speed-up the convergence. Quantitatively, in 8 GPU-days our approach discovers a set of architectures performing on-par with state-of-the-art among compact models on the semantic segmentation, pose estimation and depth prediction tasks. Code will be made available here: https://github.com/drsleep/nas-segm-pytorch. Vladimir Nekrasov, Hao Chen 0041, Chunhua Shen, Ian D. Reid 0001 |
CVPR | 2 |
| 2019 | FCOS: Fully Convolutional One-Stage Object DetectionabstractWe propose a fully convolutional one-stage object detector (FCOS) to solve object detection in a per-pixel prediction fashion, analogue to semantic segmentation. Almost all state-of-the-art object detectors such as RetinaNet, SSD, YOLOv3, and Faster R-CNN rely on pre-defined anchor boxes. In contrast, our proposed detector FCOS is anchor box free, as well as proposal free. By eliminating the pre-defined set of anchor boxes, FCOS completely avoids the complicated computation related to anchor boxes such as calculating overlapping during training. More importantly, we also avoid all hyper-parameters related to anchor boxes, which are often very sensitive to the final detection performance. With the only post-processing non-maximum suppression (NMS), FCOS with ResNeXt-64x4d-101 achieves 44.7% in AP with single-model and single-scale testing, surpassing previous one-stage detectors with the advantage of being much simpler. For the first time, we demonstrate a much simpler and flexible detection framework achieving improved detection accuracy. We hope that the proposed FCOS framework can serve as a simple and strong alternative for many other instance-level tasks. Code is available at: https://tinyurl.com/FCOSv1. Zhi Tian, Chunhua Shen, Hao Chen 0041, Tong He 0001 |
ICCV | 3 |
| 2019 | Light-Weight Hybrid Convolutional Network for Liver Tumor SegmentationabstractAutomated segmentation of liver tumors in contrast-enhanced abdominal computed tomography (CT) scans is essential in assisting medical professionals to evaluate tumor development and make fast therapeutic schedule. Although deep convolutional neural networks (DCNNs) have contributed many breakthroughs in image segmentation, this task remains challenging, since 2D DCNNs are incapable of exploring the inter-slice information and 3D DCNNs are too complex to be trained with the available small dataset. In this paper, we propose the light-weight hybrid convolutional network (LW-HCN) to segment the liver and its tumors in CT volumes. Instead of combining a 2D and a 3D networks for coarse-to-fine segmentation, LW-HCN has a encoder-decoder structure, in which 2D convolutions used at the bottom of the encoder decreases the complexity and 3D convolutions used in other layers explore both spatial and temporal information. To further reduce the complexity, we design the depthwise and spatiotemporal separate (DSTS) factorization for 3D convolutions, which not only reduces parameters dramatically but also improves the performance. We evaluated the proposed LW-HCN model against several recent methods on the LiTS and 3D-IRCADb datasets and achieved, respectively, the Dice per case of 73.0% and 94.1% for tumor segmentation, setting a new state of the art. Yutong Xie 0001, Hao Chen 0041, Yong Xia 0001, Chunhua Shen |
IJCAI | 4 |