VLDB 2026 Research / reviewers in the wild / expert
Dong Gong
dblp:125/5032
· DBLP profile ↗
63ranked-venue papers
10as first author
40since 2021 · last 2026
0000-0002-2668-9630ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 8 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 8 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-Based Test OraclesabstractLarge Language Models (LLMs) have achieved significant progress in language understanding and reasoning. Evaluating and analyzing their logical reasoning abilities has therefore become essential. However, existing datasets and benchmarks are often limited to overly simplistic, unnatural, or contextually constrained examples. In response to the growing demand, we introduce SMARTYPAT-BENCH, a challenging, naturally expressed, and systematically labeled benchmark derived from real-world high-quality Reddit posts containing subtle logical fallacies. Unlike existing datasets and benchmarks, it provides more detailed annotations of logical fallacies and features more diverse data. To further scale up the study and address the limitations of manual data collection and labeling, such as fallacy-type imbalance and labor-intensive annotation, we introduce SMARTYPAT, an automated framework powered by logic programming-based oracles. SMARTYPAT utilizes Prolog rules to systematically generate logically fallacious statements, which are then refined into fluent natural language sentences by LLMs, ensuring precise fallacy rep- resentation. Extensive evaluation demonstrates that SMARTYPAT produces fallacies comparable in subtlety and quality to human-generated content and significantly outperforms baseline methods. Finally, experiments reveal insights into LLM capabilities, highlighting that while excessive reasoning steps hinder fallacy detection accuracy, structured reasoning enhances fallacy categorization performance. Junchen Ding, Yiling Lou, Dong Gong, Yuekang Li |
AAAI | 5 |
| 2026 | CSTutorBench: Benchmarking Large Language Models for Realistic Computer Science TutoringabstractLarge Language Models (LLMs) show promise for CS educational assistance; however, the absence of comprehensive benchmarks limits our ability to assess their effectiveness in real-world teaching scenarios accurately. To fill this gap, we present CSTutorBench, a dataset from authentic course discussion forums with 2,970 multimodal question–answer pairs. Additionally, we propose an evaluation framework across five dimensions—accuracy, clarity, conciseness, personalization, and engagement—to gauge the performance of various models in tutoring settings. We benchmark leading LLMs—including GPT-4o, Claude, Llama 4, and others—using both automated metrics and expert human assessments. Across these real CS tutoring exchanges, we found leading LLMs approach human performance in terms of accuracy and clarity, but fall notably short on personalization and interactive scaffolding, often producing fluent yet less learner-adaptive guidance. Zekai Cheng, Yunfeng Wan, Daijiao Liu, Yuekang Li, Dong Gong |
SIGCSE (2) | 6 |
| 2026 | Identifying Weight-Variant Latent Causal ModelsabstractThe task of causal representation learning aims to uncover latent higher-level causal variables that affect lower-level observations. Identifying the true latent causal variables from observed data, while allowing instantaneous causal relations among latent variables, remains a challenge, however. To this end, we start with the analysis of three intrinsic indeterminacies in identifying latent variables from observations: transitivity, permutation indeterminacy, and scaling indeterminacy. We find that transitivity acts as a key role in impeding the identifiability of latent causal variables. To address the unidentifiable issue due to transitivity, we introduce a novel identifiability condition where the underlying latent causal model satisfies a linear-Gaussian model, in which the causal coefficients and the distribution of Gaussian noise are modulated by an additional observed variable. Under certain assumptions, including the existence of a reference condition under which latent causal influences vanish, we can show that the latent causal variables can be identified up to trivial permutation and scaling, and that partial identifiability results can still be obtained when this reference condition is violated for a subset of latent variables. Furthermore, based on these theoretical results, we propose a novel method, termed Structural caUsAl Variational autoEncoder (SuaVE), which directly learns causal representations and causal relationships among them, together with the mapping from the latent causal variables to the observed ones. Experimental results on synthetic and real data demonstrate the identifiability and consistency results and the efficacy of SuaVE in learning causal representations. Yuhang Liu 0002, Zhen Zhang 0008, Dong Gong, Mingming Gong, Biwei Huang, Anton van den Hengel, Kun Zhang 0001, Qinfeng Shi |
J. Mach. Learn. Res. | 3 |
| 2026 | Link prediction on multi-relational graphs from an influence propagation perspective
Zidu Yin, Yuankai Qi, Dong Gong, Ehsan Abbasnejad, Kun Yue, Qinfeng Shi |
Pattern Recognit. | 3 |
| 2025 | Self-Expansion of Pre-trained Models with Mixture of Adapters for Continual LearningabstractContinual learning (CL) aims to continually accumulate knowledge from a non-stationary data stream without catastrophic forgetting of learned knowledge, requiring a balance between stability and adaptability. Relying on the generalizable representation in pre-trained models (PTMs), PTM-based CL methods perform effective continual adaptation on downstream tasks by adding learnable adapters or prompts upon the frozen PTMs. However, many existing PTM-based CL methods use restricted adaptation on a fixed set of these modules to avoid forgetting, suffering from limited CL ability. Periodically adding task-specific modules results in linear model growth rate and impaired knowledge reuse. We propose Self-Expansion of pre-trained models with Modularized Adaptation (SEMA), a novel approach to enhance the control of stability-plasticity balance in PTM-based CL. SEMA automatically decides to reuse or add adapter modules on demand in CL, depending on whether significant distribution shift that cannot be handled is detected at different representation levels. We design modular adapter consisting of a functional adapter and a representation descriptor. The representation descriptors are trained as a distribution shift indicator and used to trigger self-expansion signals. For better composing the adapters, an expandable weighting router is learned jointly for mixture of adapter outputs. SEMA enables better knowledge reuse and sub-linear expansion rate. Extensive experiments demonstrate the effectiveness of the proposed self-expansion method, achieving state-of-the-art performance compared to PTM-based CL methods without memory rehearsal. Code is available at https://github.com/huiyiwang01/SEMA-CL. Huiyi Wang 0001, Haodong Lu 0002, Lina Yao 0001, Dong Gong |
CVPR | 4 |
| 2025 | DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective SchedulingabstractText-to-image diffusion model alignment is critical for improving the alignment between the generated images and human preferences. While training-based methods are constrained by high computational costs and dataset requirements, training-free alignment methods remain underexplored and are often limited by inaccurate guidance. We propose a plug-and-play training-free alignment method, DyMO, for aligning the generated images and human preferences during inference. Apart from text-aware human preference scores, we introduce a semantic alignment objective for enhancing the semantic alignment in the early stages of diffusion, relying on the fact that the attention maps are effective reflections of the semantics in noisy images. We propose dynamic scheduling of multiple objectives and intermediate recurrent steps to reflect the requirements at different steps. Experiments with diverse pre-trained diffusion models and metrics demonstrate the effectiveness and robustness of the proposed method. The project page: dymo.github.io. Dong Gong |
CVPR | 2 |
| 2025 | D^3CTTA: Domain-Dependent Decorrelation for Continual Test-Time Adaption of 3D LiDAR SegmentationabstractAdapting pre-trained LiDAR segmentation models to dynamic domain shifts during testing is of paramount importance for the safety of autonomous driving. Most existing methods neglect the influence of domain changes and point density in continual test-time adaption (CTTA), relying on backpropagation and large batch sizes for stability. We approach this problem with three insights: 1) Point clouds at different distances usually have different densities resulting in distribution disparities; 2) The feature distribution of different domains varies, and domain-aware parameters can alleviate domain gaps; 3) Features are highly correlated and make segmentation of different labels confusing. To this end, this work presents D3CTTA, an online backpropagation-free framework for 3D continual test-time adaption for LiDAR segmentation. D3CTTA consists of a distance-aware prototype learning module to integrate LiDAR-based geometry prior and a domain-dependent decorrelation module to reduce feature correlations among different domains and different categories. Extensive experiments on three benchmarks showcase that our method achieves a state-of-the-art performance compared to both backpropagation-based methods and backpropagation-free methods. Code is available at https://github.com/ZhaoJichun1/D3CTTA. Jichun Zhao, Haiyong Jiang, Haoxuan Song, Jun Xiao 0005, Dong Gong |
CVPR | 5 |
| 2025 | Is Less More? Exploring Token Condensation as Training-Free Test-Time Adaptation
Dong Gong, Sen Wang 0001, Zi Huang, Yadan Luo |
ICCV | 2 |
| 2025 | Analytic DAG Constraints for Differentiable DAG LearningabstractRecovering the underlying Directed Acyclic Graph (DAG)
structures from observational data presents a formidable challenge, partly due
to the combinatorial nature of the DAG-constrained optimization
problem. Recently, researchers have identified gradient vanishing as
one of the primary obstacles in differentiable DAG learning and have
proposed several DAG constraints to mitigate this issue. By developing
the necessary theory to establish a connection between analytic
functions and DAG constraints, we demonstrate that analytic functions
from the set $\\{f(x) = c_0 + \\sum_{i=1}^{\infty}c_ix^i | \\forall i > 0, c_i > 0; r = \\lim_{i\\rightarrow \\infty}c_{i}/c_{i+1} > 0\\}$ can be employed to
formulate effective DAG constraints. Furthermore, we establish that
this set of functions is closed under several functional operators,
including differentiation, summation, and
multiplication. Consequently, these operators can be leveraged to
create novel DAG constraints based on existing ones. Using these
properties, we design a series of DAG constraints and develop an
efficient algorithm to evaluate them. Experiments
in various settings demonstrate that our DAG constraints
outperform previous state-of-the-art comparators. Our implementation is available at https://github.com/zzhang1987/AnalyticDAGLearning. Zhen Zhang 0008, Ignavier Ng, Dong Gong, Yuhang Liu 0002, Mingming Gong, Biwei Huang, Kun Zhang 0001, Anton van den Hengel, Qinfeng Shi |
ICLR | 3 |
| 2025 | Mining your own secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion ModelsabstractPersonalized text-to-image diffusion models have grown popular for their ability to efficiently acquire a new concept from user-defined text descriptions and a few images. However, in the real world, a user may wish to personalize a model on multiple concepts but one at a time, with no access to the data from previous concepts due to storage/privacy concerns. When faced with this continual learning (CL) setup, most personalization methods fail to find a balance between acquiring new concepts and retaining previous ones -- a challenge that *continual personalization* (CP) aims to solve.
Inspired by the successful CL methods that rely on class-specific information for regularization, we resort to the inherent class-conditioned density estimates, also known as *diffusion classifier* (DC) scores, for CP of text-to-image diffusion models.
Namely, we propose using DC scores for regularizing the parameter-space and function-space of text-to-image diffusion models.
Using several diverse evaluation setups, datasets, and metrics, we show that our proposed regularization-based CP methods outperform the state-of-the-art C-LoRA, and other baselines. Finally, by operating in the replay-free CL setup and on low-rank adapters, our method incurs zero storage and parameter overhead, respectively, over the state-of-the-art. Saurav Jha, Masato Ishii, Christian Simon, Muhammad Jehanzeb Mirza, Dong Gong, Lina Yao 0001, Shusuke Takahashi, Yuki Mitsufuji |
ICLR | 7 |
| 2025 | Coreset Selection via Reducible Loss in Continual LearningabstractRehearsal-based continual learning (CL) aims to mitigate catastrophic forgetting by maintaining a subset of samples from previous tasks and replaying them. The rehearsal memory can be naturally constructed as a coreset, designed to form a compact subset that enables training with performance comparable to using the full dataset. The coreset selection task can be formulated as bilevel optimization that solves for the subset to minimize the outer objective of the learning task. Existing methods primarily rely on inefficient probabilistic sampling or local gradient-based scoring to approximate sample importance through an iterative process that can be susceptible to ambiguity or noise. Specifically, non-representative samples like ambiguous or noisy samples are difficult to learn and incur high loss values even when training on the full dataset. However, existing methods relying on local gradient tend to highlight these samples in an attempt to minimize the outer loss, leading to a suboptimal coreset. To enhance coreset selection, especially in CL where high-quality samples are essential, we propose a coreset selection method that measures sample importance using reducible loss (ReL) that quantifies the impact of adding a sample to model performance. By leveraging ReL and a process derived from bilevel optimization, we identify and retain samples that yield the highest performance gain. They are shown to be informative and representative. Furthermore, ReL requires only forward computation, making it significantly more efficient than previous methods. To better apply coreset selection in CL, we extend our method to address key challenges such as task interference, streaming data, and knowledge distillation. Experiments on data summarization and continual learning demonstrate the effectiveness and efficiency of our approach. Ruilin Tong, Yuhang Liu 0002, Qinfeng Shi, Dong Gong |
ICLR | 4 |
| 2025 | Seeing the Unseen: Composing Outliers for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) is to recognize unseen attribute-object compositions by learning from seen compositions. The distribution shift between unseen compositions and seen compositions poses challenges to CZSL models, especially when test images are mixed with both seen and unseen compositions. The challenge will be addressed more easily if a model can distinguish unseen/seen compositions and treat them with specific recognition strategies. However, identifying images with unseen compositions is non-trivial, considering that unseen compositions are absent in training and usually contain only subtle differences from seen compositions. In this paper, we propose a novel compositional zero-shot learning method called COMO, which composes outliers in training for distinguishing seen and unseen compositions and further applying specific strategies for them. Specifically, we compose attribute-object representations for unseen compositions based on primitive representations of training images as outliers to enable the model to identify unseen compositions in inference. At test time, the method distinguishes images containing seen/unseen compositions and uses different weights for composition classification and primitive classification to recognize seen/unseen compositions. Experimental results on three datasets show the effectiveness of our method in both the closed-world setting and the open-world setting. Chenchen Jing, Hao Chen 0041, Yuling Xi, Xingyuan Bu, Dong Gong, Chunhua Shen |
IJCAI | 6 |
| 2025 | Let Your Video Listen to Your Music! - Beat-Aligned, Content-Preserving Video Editing with Arbitrary MusicabstractAligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing. Effective alignment between motion and musical beats enhances viewer engagement and visual appeal, particularly in music videos, promotional content, and cinematic editing. Existing methods typically depend on labor-intensive manual cutting, speed adjustments, or heuristic-based editing techniques to achieve synchronization. While some generative models handle joint video and music generation, they often entangle the two modalities, limiting flexibility in aligning video to music beats while preserving the full visual content. In this paper, we propose a novel and efficient framework-termed MVAA (Music-Video Auto-Alignment)-that automatically edits video to align with the rhythm of a given music track while preserving the original visual content. To enhance flexibility, we modularize the task into a two-step process in our MVAA: aligning motion keyframes with audio beats, followed by rhythm-aware video inpainting. Specifically, we first insert keyframes at timestamps aligned with musical beats, then use a frame-conditioned diffusion model to generate coherent intermediate frames, preserving the original video's semantic content. Since comprehensive test-time training can be time-consuming, we adopt a two-stage strategy: pretraining the inpainting module on a small video set to learn general motion priors, followed by rapid inference-time fine-tuning for video-specific adaptation. This hybrid approach enables adaptation within ~10 minutes with one epoch on a single NVIDIA 4090 GPU using CogVideoX-5b-I2V [77] as the backbone. Extensive experiments show that our approach can achieve high-quality beat alignment and visual smoothness. User studies further validate the natural rhythmic quality of the results, confirming their effectiveness for practical music-video editing. The code is available at: zhangxinyu-xyz.github.io/MVAA Xinyu Zhang 0017, Dong Gong, Zicheng Duan, Anton van den Hengel, Lingqiao Liu |
ACM Multimedia | 2 |
| 2025 | Model Inversion with Layer-Specific Modeling and Alignment for Data-Free Continual LearningabstractContinual learning (CL) aims to incrementally train a model to a sequence of tasks while maintaining performance on previously seen ones. Despite effectiveness in mitigating forgetting, data storage and replay may be infeasible due to privacy or security constraints, and are impractical or unavailable for arbitrary pre-trained models. Data-free or examplar-free CL aims to continually update models with new
tasks without storing previous data. In addition to regularizing updates, we employ model inversion to synthesize data from the trained model, anchoring learned knowledge through replay without retaining old data. However, model inversion in predictive models faces two key challenges. First, generating inputs (e.g., images) solely from highly compressed output labels (e.g., classes) often causes drift between synthetic and real data. Replaying on such synthetic data can contaminate and erode knowledge learned from real data, further degrading inversion quality over time. Second, performing inversion is usually computationally expensive, as each iteration requires backpropagation through the entire model and many steps are needed for convergence. These problems are more severe with large pre-trained models such as Contrastive Language-Image Pre-training (CLIP) models. To improve model inversion efficiency, we propose Per-layer Model Inversion (PMI) approach inspired by the faster convergence of single-layer optimization. The inputs optimized from PMI provide strong initialization for full-model inversion, significantly reducing the number of iterations required for convergence. To address feature distribution shift, we model class-wise feature distribution using a Gaussian distribution and preserve distributional information with a contrastive model. Sampling features for inversion ensures alignment between synthetic and real feature distributions. Combining PMI and feature modeling, we demonstrate the feasibility of incrementally training models on new classes by generating data from pseudo image features mapped through semantic-aware feature projection. Our method shows strong effectiveness and compatibility across multiple CL settings. Ruilin Tong, Yuhang Liu 0002, Dong Gong |
NeurIPS | 4 |
| 2025 | FlashMo: Geometric Interpolants and Frequency-Aware Sparsity for Scalable Efficient Motion GenerationabstractDiffusion models have recently advanced 3D human motion generation by producing smoother and more realistic sequences from natural language. However, existing approaches face two major challenges: high computational cost during training and inference, and limited scalability due to reliance on U-Net inductive bias. To address these challenges, we propose **FlashMo**, a frequency-aware sparse motion diffusion model that prunes low-frequency tokens to enhance efficiency without custom kernel design. We further introduce *MotionSiT*, a scalable diffusion transformer based on a joint-temporal factorized interpolant with Lie group geodesics over $\mathrm{SO}(3)$ manifolds, enabling principled generation of joint rotations. Extensive experiments on the large-scale MotionHub V2 dataset and standard benchmarks including HumanML3D and KIT-ML demonstrate that our method significantly outperforms previous approaches in motion quality, efficiency, and scalability. Compared to the state-of-the-art 1-step distillation baseline, FlashMo reduces **12.9%** inference time and FID by **34.1%**. Project website: https://steve-zeyu-zhang.github.io/FlashMo. Zeyu Zhang 0006, Danning Li, Dong Gong, Ian D. Reid 0001, Richard I. Hartley |
NeurIPS | 4 |
| 2024 | Learning with Mixture of Prototypes for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection aims to detect testing samples far away from the in-distribution (ID) training data, which is crucial for the safe deployment of machine learning models in the real world. Distance-based OOD detection methods have emerged with enhanced deep representation learning. They identify unseen OOD samples by measuring their distances from ID class centroids or prototypes. However, existing approaches learn the representation relying on oversimplified data assumptions, e.g. modeling ID data of each class with one centroid class prototype or using loss functions not designed for OOD detection, which overlook the natural diversities within the data. Naively enforcing data samples of each class to be compact around only one prototype leads to inadequate modeling of realistic data and limited performance. To tackle these issues, we propose PrototypicAl Learning with a Mixture of prototypes (PALM) that models each class with multiple prototypes to capture the sample diversities, which learns more faithful and compact samples embeddings for enhanching OOD detection. Our method automatically identifies and dynamically updates prototypes, assigning each sample to a subset of prototypes via reciprocal neighbor soft assignment weights. To learn embeddings with multiple prototypes, PALM optimizes a maximum likelihood estimation (MLE) loss to encourage the sample embeddings to compact around the associated prototypes, as well as a contrastive loss on all prototypes to enhance intra-class compactness and inter-class discrimination at the prototype level. Compared to previous methods with prototypes, the proposed mixture prototype modeling of PALM promotes the representations of each ID class to be more compact and separable from others and the unseen OOD samples, resulting in more reliable OOD detection. Moreover, the automatic estimation of prototypes enables our approach to be extended to the challenging OOD detection task with unlabelled ID data. Extensive experiments demonstrate the superiority of PALM over previous methods, achieving state-of-the-art average AUROC performance of 93.82 on the challenging CIFAR-100 benchmark. Haodong Lu 0002, Dong Gong, Shuo Wang 0012, Minhui Xue 0001, Lina Yao 0001, Kristen Moore |
ICLR | 2 |
| 2024 | Identifiable Latent Polynomial Causal Models through the Lens of ChangeabstractCausal representation learning aims to unveil latent high-level causal representations from observed low-level data. One of its primary tasks is to provide reliable assurance of identifying these latent causal models, known as \textit{identifiability}. A recent breakthrough explores identifiability by leveraging the change of causal influences among latent causal variables across multiple environments \citep{liu2022identifying}. However, this progress rests on the assumption that the causal relationships among latent causal variables adhere strictly to linear Gaussian models. In this paper, we extend the scope of latent causal models to involve nonlinear causal relationships, represented by polynomial models, and general noise distributions conforming to the exponential family. Additionally, we investigate the necessity of imposing changes on all causal parameters and present partial identifiability results when part of them remains unchanged. Further, we propose a novel empirical estimation method, grounded in our theoretical finding, that enables learning consistent latent causal representations. Our experimental results, obtained from both synthetic and real-world data, validate our theoretical contributions concerning identifiability and consistency. Yuhang Liu 0002, Zhen Zhang 0008, Dong Gong, Mingming Gong, Biwei Huang, Anton van den Hengel, Kun Zhang 0001, Qinfeng Shi |
ICLR | 3 |
| 2024 | SDGE: Stereo Guided Depth Estimation for 360°Camera SetsabstractDepth estimation is a critical technology in autonomous driving, and multi-camera systems are often used to achieve a 360° perception. These 360° camera sets often have limited or low-quality overlap regions, making multi-view stereo methods infeasible for the entire image. Alternatively, monocular methods may not produce consistent cross-view predictions. To address these issues, we propose the Stereo Guided Depth Estimation (SGDE) method, which enhances depth estimation of the full image by explicitly utilizing multi-view stereo results on the overlap. We suggest building virtual pinhole cameras to resolve the distortion problem of fisheye cameras and unify the processing for the two types of 360° cameras. For handling the varying noise on camera poses caused by unstable movement, the approach employs a self-calibration method to obtain highly accurate relative poses of the adjacent cameras with minor overlap. These enable the use of robust stereo methods to obtain a high-quality depth prior in the overlap region. This prior serves not only as an additional input but also as pseudo-labels that enhance the accuracy of depth estimation methods and improve cross-view prediction consistency. The effectiveness of SGDE is evaluated on one fisheye camera dataset, Synthetic Urban, and two pinhole camera datasets, DDAD and nuScenes. Our experiments demonstrate that SGDE is effective for both supervised and self-supervised depth estimation, and highlight the potential of our method for advancing autonomous driving technology. Our project page is at https://github.com/JialeiXu/SGDE. Jialei Xu, Dong Gong, Junjun Jiang, Xianming Liu 0005 |
IROS | 3 |
| 2024 | EGGen: Image Generation with Multi-entity Prior Learning through Entity GuidanceabstractDiffusion models have shown remarkable prowess in text-to-image synthesis and editing, yet they often stumble when tasked with interpreting complex prompts that describe multiple entities with specific attributes and interrelations. The generated images often contain inconsistent multi-entity representation (IMR), reflected as inaccurate presentations of the multiple entities and their attributes. Although providing spatial layout guidance improves the multi-entity generation quality in existing works, it is still challenging to handle the leakage attributes and avoid unnatural characteristics. To address the IMR challenge, we first conduct in-depth analyses of the diffusion process and attention operation, revealing that the IMR challenges largely stem from the process of cross-attention mechanisms. According to the analyses, we introduce the entity guidance generation mechanism, which maintains the integrity of the original diffusion model parameters by integrating plug-in networks. Our work advances the stable diffusion model by segmenting comprehensive prompts into distinct entity-specific prompts with bounding boxes, enabling a transition from multi-entity to single-entity generation in cross-attention layers. More importantly, we introduce entity-centric cross-attention layers that focus on individual entities to preserve their uniqueness and accuracy, alongside global entity alignment layers that refine cross-attention maps using multi-entity priors for precise positioning and attribute accuracy. Additionally, a linear attenuation module is integrated to progressively reduce the influence of these layers during inference, preventing oversaturation and preserving generation fidelity. Our comprehensive experiments demonstrate that this entity guidance generation enhances existing text-to-image models in generating detailed, multi-entity images. Zhenhong Sun, Junyan Wang 0001, Zhiyu Tan, Daoyi Dong, Hailan Ma, Hao Li 0030, Dong Gong |
ACM Multimedia | 7 |
| 2024 | CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language ModelsabstractContinual learning (CL) aims to help deep neural networks to learn new knowledge while retaining what has been learned. Owing to their powerful generalizability, pre-trained vision-language models such as Contrastive Language-Image Pre-training (CLIP) have lately gained traction as practical CL candidates. However, the domain mismatch between the pre-training and the downstream CL tasks calls for finetuning of the CLIP on the latter. The deterministic nature of the existing finetuning methods makes them overlook the many possible interactions across the modalities and deems them unsafe for high-risk tasks requiring reliable uncertainty estimation. To address these, our work proposes **C**ontinual **L**e**A**rning with **P**robabilistic finetuning (CLAP) - a probabilistic modeling framework over visual-guided text features per task, thus providing more calibrated CL finetuning. Unlike recent data-hungry anti-forgetting CL techniques, CLAP alleviates forgetting by exploiting the rich pre-trained knowledge of CLIP for weight initialization and distribution regularization of task-specific parameters. Cooperating with the diverse range of existing prompting methods, CLAP can surpass the predominant deterministic finetuning approaches for CL with CLIP. We conclude with out-of-the-box applications of superior uncertainty estimation abilities of CLAP including novel data detection and exemplar selection within the existing CL setups. Our code is available at https://github.com/srvCodes/clap4clip. Saurav Jha, Dong Gong, Lina Yao 0001 |
NeurIPS | 2 |
| 2024 | Continual All-in-One Adverse Weather Removal With Knowledge Replay on a Unified Network StructureabstractIn real-world applications, image degeneration caused by adverse weather is always complex and changes with different weather conditions from days and seasons. Systems in real-world environments constantly encounter adverse weather conditions that are not previously observed. Therefore, it practically requires adverse weather removal models to continually learn from incrementally collected data reflecting various degeneration types. Existing adverse weather removal approaches, for either single or multiple adverse weathers, are mainly designed for a static learning paradigm, which assumes that the data of all types of degenerations to handle can be finely collected at one time before a single-phase learning process. They thus cannot directly handle the incremental learning requirements. To address this issue, we made the earliest effort to investigate the continual all-in-one adverse weather removal task, in a setting closer to real-world applications. Specifically, we develop a novel continual learning framework with effective knowledge replay (KR) on a unified network structure. Equipped with a principal component projection and an effective knowledge distillation mechanism, the proposed KR techniques are tailored for the all-in-one weather removal task. It considers the characteristics of the image restoration task with multiple degenerations in continual learning, and the knowledge for different degenerations can be shared and accumulated in the unified network structure. Extensive experimental results demonstrate the effectiveness of the proposed method to deal with this challenging task, which performs competitively to existing dedicated or joint training image restoration methods. Our code is available athttps://github.com/xiaojihh/CL_all-in-one. De Cheng, Yanling Ji, Dong Gong, Yan Li 0125, Nannan Wang 0001, Junwei Han 0001, Dingwen Zhang |
IEEE Trans. Multim. | 3 |
| 2023 | Learning to Fuse Monocular and Multi-view Cues for Multi-frame Depth Estimation in Dynamic ScenesabstractMulti-frame depth estimation generally achieves high accuracy relying on the multi-view geometric consistency. When applied in dynamic scenes, e.g., autonomous driving, this consistency is usually violated in the dynamic areas, leading to corrupted estimations. Many multi-frame methods handle dynamic areas by identifying them with explicit masks and compensating the multi-view cues with monocular cues represented as local monocular depth or features. The improvements are limited due to the uncontrolled quality of the masks and the underutilized benefits of the fusion of the two types of cues. In this paper, we propose a novel method to learn to fuse the multi-view and monocular cues encoded as volumes without needing the heuristically crafted masks. As unveiled in our analyses, the multiview cues capture more accurate geometric information in static areas, and the monocular cues capture more useful contexts in dynamic areas. To let the geometric perception learned from multi-view cues in static areas propagate to the monocular representation in dynamic areas and let monocular cues enhance the representation of multi-view cost volume, we propose a cross-cue fusion (CCF) module, which includes the cross-cue attention (CCA) to encode the spatially non-local relative intra-relations from each source to enhance the representation of the other. Experiments on real-world datasets prove the significant effectiveness and generalization ability of the proposed method. Rui Li 0013, Dong Gong, Wei Yin 0006, Hao Chen 0041, Yu Zhu 0004, Xiaozhi Chen, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 2 |
| 2023 | Maximizing Spatio-Temporal Entropy of Deep 3D CNNs for Efficient Video Recognition
Junyan Wang 0001, Zhenhong Sun, Yichen Qian, Dong Gong, Xiuyu Sun, Ming Lin 0002, Maurice Pagnucco, Yang Song 0001 |
ICLR | 4 |
| 2023 | NPCL: Neural Processes for Uncertainty-Aware Continual LearningabstractContinual learning (CL) aims to train deep neural networks efficiently on streaming data while limiting the forgetting caused by new tasks. However, learning transferable knowledge with less interference between tasks is difficult, and real-world deployment of CL models is limited by their inability to measure predictive uncertainties. To address these issues, we propose handling CL tasks with neural processes (NPs), a class of meta-learners that encode different tasks into probabilistic distributions over functions all while providing reliable uncertainty estimates. Specifically, we propose an NP-based CL approach (NPCL) with task-specific modules arranged in a hierarchical latent variable model. We tailor regularizers on the learned latent distributions to alleviate forgetting. The uncertainty estimation capabilities of the NPCL can also be used to handle the task head/module inference challenge in CL. Our experiments show that the NPCL outperforms previous CL approaches. We validate the effectiveness of uncertainty estimation in the NPCL for identifying novel data and evaluating instance-level model confidence. Code is available at https://github.com/srvCodes/NPCL. Saurav Jha, Dong Gong, He Zhao 0001, Lina Yao 0001 |
NeurIPS | 2 |
| 2023 | RanPAC: Random Projections and Pre-trained Models for Continual LearningabstractContinual learning (CL) aims to incrementally learn different tasks (such as classification) in a non-stationary data stream without forgetting old ones. Most CL works focus on tackling catastrophic forgetting under a learning-from-scratch paradigm. However, with the increasing prominence of foundation models, pre-trained models equipped with informative representations have become available for various downstream requirements. Several CL methods based on pre-trained models have been explored, either utilizing pre-extracted features directly (which makes bridging distribution gaps challenging) or incorporating adaptors (which may be subject to forgetting). In this paper, we propose a concise and effective approach for CL with pre-trained models. Given that forgetting occurs during parameter updating, we contemplate an alternative approach that exploits training-free random projectors and class-prototype accumulation, which thus bypasses the issue. Specifically, we inject a frozen Random Projection layer with nonlinear activation between the pre-trained model's feature representations and output head, which captures interactions between features with expanded dimensionality, providing enhanced linear separability for class-prototype-based CL. We also demonstrate the importance of decorrelating the class-prototypes to reduce the distribution disparity when using pre-trained representations. These techniques prove to be effective and circumvent the problem of forgetting for both class- and domain-incremental continual learning. Compared to previous methods applied to pre-trained ViT-B/16 models, we reduce final error rates by between 20% and 62% on seven class-incremental benchmark datasets, despite not using any rehearsal memory. We conclude that the full potential of pre-trained models for simple, effective, and fast continual learning has not hitherto been fully tapped. Code is available at https://github.com/RanPAC/RanPAC. Mark D. McDonnell, Dong Gong, Amin Parvaneh, Ehsan Abbasnejad, Anton van den Hengel |
NeurIPS | 2 |
| 2023 | SharpFormer: Learning Local Feature Preserving Global Representations for Image DeblurringabstractThe goal of dynamic scene deblurring is to remove the motion blur presented in a given image. To recover the details from the severe blurs, conventional convolutional neural networks (CNNs) based methods typically increase the number of convolution layers, kernel-size, or different scale images to enlarge the receptive field. However, these methods neglect the non-uniform nature of blurs, and cannot extract varied local and global information. Unlike the CNNs-based methods, we propose a Transformer-based model for image deblurring, named SharpFormer, that directly learns long-range dependencies via a novel Transformer module to overcome large blur variations. Transformer is good at learning global information but is poor at capturing local information. To overcome this issue, we design a novel Locality preserving Transformer (LTransformer) block to integrate sufficient local information into global features. In addition, to effectively apply LTransformer to the medium-resolution features, a hybrid block is introduced to capture intermediate mixed features. Furthermore, we use a dynamic convolution (DyConv) block, which aggregates multiple parallel convolution kernels to handle the non-uniform blur of inputs. We leverage a powerful two-stage attentive framework composed of the above blocks to learn the global, hybrid, and local features effectively. Extensive experiments on the GoPro and REDS datasets show that the proposed SharpFormer performs favourably against the state-of-the-art methods in blurred image restoration. Qingsen Yan, Dong Gong, Zhen Zhang 0008, Yanning Zhang 0001, Qinfeng Shi |
IEEE Trans. Image Process. | 2 |
| 2022 | Learning Bayesian Sparse Networks with Full Experience Replay for Continual LearningabstractContinual Learning (CL) methods aim to enable machine learning models to learn new tasks without catastrophic forgetting of those that have been previously mastered. Existing CL approaches often keep a buffer of previously-seen samples, perform knowledge distillation, or use regularization techniques towards this goal. Despite their performance, they still suffer from interference across tasks which leads to catastrophic forgetting. To ameliorate this problem, we propose to only activate and select sparse neurons for learning current and past tasks at any stage. More parameters space and model capacity can thus be reserved for the future tasks. This minimizes the interference between parameters for different tasks. To do so, we propose a Sparse neural Network for Continual Learning (SNCL), which employs variational Bayesian sparsity priors on the activations of the neurons in all layers. Full Experience Replay (FER) provides effective supervision in learning the sparse activations of the neurons in different layers. A loss-aware reservoir-sampling strategy is developed to maintain the memory buffer. The proposed method is agnostic as to the network structures and the task boundaries. Experiments on different datasets show that SNCL achieves state-of-the-art result for mitigating forgetting. Qingsen Yan, Dong Gong, Yuhang Liu 0002, Anton van den Hengel, Qinfeng Shi |
CVPR | 2 |
| 2022 | Truncated Matrix Power Iteration for Differentiable DAG LearningabstractRecovering underlying Directed Acyclic Graph (DAG) structures from observational data is highly challenging due to the combinatorial nature of the DAG-constrained optimization problem. Recently, DAG learning has been cast as a continuous optimization problem by characterizing the DAG constraint as a smooth equality one, generally based on polynomials over adjacency matrices. Existing methods place very small coefficients on high-order polynomial terms for stabilization, since they argue that large coefficients on the higher-order terms are harmful due to numeric exploding. On the contrary, we discover that large coefficients on higher-order terms are beneficial for DAG learning, when the spectral radiuses of the adjacency matrices are small, and that larger coefficients for higher-order terms can approximate the DAG constraints much better than the small counterparts. Based on this, we propose a novel DAG learning method with efficient truncated matrix power iteration to approximate geometric series based DAG constraints. Empirically, our DAG learning method outperforms the previous state-of-the-arts in various settings, often by a factor of $3$ or more in terms of structural Hamming distance. Zhen Zhang 0008, Ignavier Ng, Dong Gong, Yuhang Liu 0002, Ehsan Abbasnejad, Mingming Gong, Kun Zhang 0001, Qinfeng Shi |
NeurIPS | 3 |
| 2022 | Dual-Attention-Guided Network for Ghost-Free High Dynamic Range Imaging
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Ian D. Reid 0001, Yanning Zhang 0001 |
Int. J. Comput. Vis. | 2 |
| 2022 | Video super-resolution via mixed spatial-temporal convolution and selective fusion
Wei Sun 0036, Dong Gong, Qinfeng Shi, Anton van den Hengel, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2022 | High dynamic range imaging via gradient-aware context aggregation network
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Jinqiu Sun, Yu Zhu 0004, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2022 | Part-Guided Attention Learning for Vehicle Instance RetrievalabstractVehicle instance retrieval (IR) often requires one to recognize the fine-grained visual differences between vehicles. Besides the holistic appearance of vehicles which is easily affected by the viewpoint variation and distortion, vehicle parts also provide crucial cues to differentiate near-identical vehicles. Motivated by these observations, we introduce aPart-Guided Attention Network(PGAN) to pinpoint the prominent part regions and effectively combine the global and local information for discriminative feature learning. PGAN first detects the locations of different part components and salient regions regardless of the vehicle identity, which serves as thebottom-up attentionto narrow down the possible searching regions. To estimate the importance of detected parts, we propose aPart Attention Module(PAM) to adaptively locate the most discriminative regions with high-attention weights and suppress the distraction of irrelevant parts with relatively low weights. The PAM is guided by the identification loss and therefore providestop-down attentionthat enables attention to be calculated at the level of car parts and other salient regions. Finally, we aggregate the global appearance and local features together to improve the feature performance further. The PGAN combines part-guided bottom-up and top-down attention, global and local visual features in an end-to-end framework. Extensive experiments demonstrate that the proposed method achieves new state-of-the-art vehicle IR performance on four large-scale benchmark datasets.1 Xinyu Zhang 0015, Rufeng Zhang, Jiewei Cao, Dong Gong, Mingyu You, Chunhua Shen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | A Comprehensive CT Dataset for Liver Computer Assisted Diagnosis
Qingsen Yan, Bo Wang 0011, Dong Gong, Dingwen Zhang, Yang Yang 0009, Zheng You, Yanning Zhang 0001, Qinfeng Shi |
BMVC | 3 |
| 2021 | Memory-augmented Dynamic Neural Relational InferenceabstractDynamic interacting systems are prevalent in vision tasks. These interactions are usually difficult to observe and measure directly, and yet understanding latent interactions is essential for performing inference tasks on dynamic systems like forecasting. Neural relational inference (NRI) techniques are thus introduced to explicitly estimate interpretable relations between the entities in the system for trajectory prediction. However, NRI assumes static relations; thus, dynamic neural relational inference (DNRI) was proposed to handle dynamic relations using LSTM. Unfortunately, the older information will be washed away when the LSTM updates the latent variable as a whole, which is why DNRI struggles with modeling long-term dependences and forecasting long sequences. This motivates us to propose a memory-augmented dynamic neural relational inference method, which maintains two associative memory pools: one for the interactive relations and the other for the individual entities. The two memory pools help retain useful relation features and node features for the estimation in the future steps. Our model dynamically estimates the relations by learning better embeddings and utilizing the long-range information stored in the memory. With the novel memory modules and customized structures, our memory-augmented DNRI can update and access the memory adaptively as required. The memory pools also serve as global latent variables across time to maintain detailed long-term temporal relations readily available for other components to use. Experiments on synthetic and real-world datasets show the effectiveness of the proposed method on modeling dynamic relations and forecasting complex trajectories. Dong Gong, Zhen Zhang 0008, Qinfeng Shi, Anton van den Hengel |
ICCV | 1 |
| 2021 | Image editing with varying intensities of processing
Yasi Wang, Yuankai Qi, Hongxun Yao, Dong Gong, Qi Wu 0001 |
Comput. Vis. Image Underst. | 4 |
| 2021 | Boundarymix: Generating pseudo-training images for improving segmentation with scribble annotations
Wanxuan Lu, Dong Gong, Kun Fu 0001, Xian Sun 0001, Wenhui Diao, Lingqiao Liu |
Pattern Recognit. | 2 |
| 2021 | COVID-19 Chest CT Image Segmentation Network by Multi-Scale Fusion and Enhancement OperationsabstractA novel coronavirus disease 2019 (COVID-19) was detected and has spread rapidly across various countries around the world since the end of the year 2019. Computed Tomography (CT) images have been used as a crucial alternative to the time-consuming RT-PCR test. However, pure manual segmentation of CT images faces a serious challenge with the increase of suspected cases, resulting in urgent requirements for accurate and automatic segmentation of COVID-19 infections. Unfortunately, since the imaging characteristics of the COVID-19 infection are diverse and similar to the backgrounds, existing medical image segmentation methods cannot achieve satisfactory performance. In this article, we try to establish a new deep convolutional neural network tailored for segmenting the chest CT images with COVID-19 infections. We first maintain a large and new chest CT image dataset consisting of 165,667 annotated chest CT images from 861 patients with confirmed COVID-19. Inspired by the observation that the boundary of the infected lung can be enhanced by adjusting the global intensity, in the proposed deep CNN, we introduce a feature variation block which adaptively adjusts the global properties of the features for segmenting COVID-19 infection. The proposed FV block can enhance the capability of feature representation effectively and adaptively for diverse cases. We fuse features at different scales by proposing Progressive Atrous Spatial Pyramid Pooling to handle the sophisticated infection areas with diverse appearance and shapes. The proposed method achieves state-of-the-art performance. Dice similarity coefficients are 0.987 and 0.726 for lung and COVID-19 segmentation, respectively. We conducted experiments on the data collected in China and Germany and show that the proposed deep CNN can produce impressive performance effectively. The proposed network enhances the segmentation ability of the COVID-19 infection, makes the connection with other techniques and contributes to the development of remedying COVID-19 infection. Qingsen Yan, Bo Wang 0011, Dong Gong, Chuan Luo 0003, Jianhu Shen, Jingyang Ai, Qinfeng Shi, Yanning Zhang 0001, Liang Zhang 0010, Zheng You |
IEEE Trans. Big Data | 3 |
| 2021 | Learning to Zoom-In via Learning to Zoom-Out: Real-World Super-Resolution by Generating and Adapting DegradationabstractMost learning-based super-resolution (SR) methods aim to recover high-resolution (HR) image from a given low-resolution (LR) image via learning on LR-HR image pairs. The SR methods learned on synthetic data do not perform well in real-world, due to the domain gap between the artificially synthesized and real LR images. Some efforts are thus taken to capture real-world image pairs. However, the captured LR-HR image pairs usually suffer from unavoidable misalignment, which hampers the performance of end- to-end learning. Here, focusing on the real-world SR, we ask a different question: since misalignment is unavoidable, can we propose a method that does not need LR-HR image pairing and alignment at all and utilizes real images as they are? Hence we propose a framework to learn SR from an arbitrary set of unpaired LR and HR images and see how far a step can go in such a realistic and "unsupervised" setting. To do so, we firstly train a degradation generation network to generate realistic LR images and, more importantly, to capture their distribution (i.e., learning to zoom out). Instead of assuming the domain gap has been eliminated, we minimize the discrepancy between the generated data and real data while learning a degradation adaptive SR network (i.e., learning to zoom in). The proposed unpaired method achieves state-of- the-art SR results on real-world images, even in the datasets that favour the paired-learning methods more. Wei Sun 0036, Dong Gong, Qinfeng Shi, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Learning Sparse PCA with Stabilized ADMM Method on Stiefel ManifoldabstractSparse principal component analysis (SPCA) produces principal components with sparse loadings, which is very important for handling data with many irrelevant features and also critical to interpret the results. To deal with orthogonal constraints, most previous approaches address SPCA with several components using techniques such as deflation technique and convex relaxations. However, the deflation technique usually suffers from suboptimal solutions due to poor approximations. On the other hand, the convex relaxations are often computationally expensive. To address the above issues, in this paper, we propose to address SPCA over the Stiefel manifold directly, and develop a stabilized Alternating Direction Method of Multipliers (SADMM) to handle the nonconvex orthogonal constraints. Compared to traditional ADMM, the proposed SADMM method converges well with a wide range of parameters and obtains a better solution. We also theoretically study the convergence property of the proposed SADMM method. Furthermore, most existing methods ignore an inherent drawback of SPCA - the importance of different components is not considered when doing feature selection, which often makes the selected features nonoptimal. To address this, we further propose a two-stage method which considers the importance of different components to select the most important features. Empirical studies on both synthetic and real-world datasets show that the proposed algorithms achieve better performance compared to existing state-of-the-art methods. Mingkui Tan, Zhibin Hu, Yuguang Yan, Jiezhang Cao, Dong Gong, Qingyao Wu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2021 | Deep Single Image Deraining via Modeling Haze-Like EffectabstractRemoving rain from images is of a great importance to various applications such as autonomous driving, drone piloting, and photo editing. Conventional methods rely on some heuristics to handcraft various priors to remove or separate rain from images. Recently, deep learning models are proposed to learn various end-to-end methods to complete this task. However, these methods might fail in obtaining satisfactory results in some real-world scenarios, especially when the captured images suffer from heavy rain that brings not only rain streaks but also a haze-like effect (caused by the accumulation of tiny raindrops). Different from most of the existing deep learning deraining methods that focus on handling rain streaks, we add a new variable to model the haze-like effect in a general model for rain, based on which a deep neural network is designed accordingly. Specifically, in our method, two branches are designed to handle rain streaks and the haze-like effect, respectively. The output of such branch structure is fed to an additional module to further enhance the performance. Three modules are trained jointly, leading to an end-to-end network, which supports a adjustment to the strength of removing the haze-like effect. Extensive experiments on several datasets show that our method outperforms several state-of-the-art methods in both objective assessment and visual quality. Yinglong Wang 0002, Dong Gong, Jie Yang 0002, Qinfeng Shi, Anton van den Hengel, Dehua Xie, Bing Zeng 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Learning and Memorizing Representative Prototypes for 3D Point Cloud Semantic and Instance Segmentation
Tong He 0001, Dong Gong, Zhi Tian, Chunhua Shen |
ECCV (18) | 2 |
| 2020 | Learning Distilled Graph for Large-Scale Social Network Data ClusteringabstractSpectral analysis is critical in social network analysis. As a vital step of the spectral analysis, the graph construction in many existing works utilizes content data only. Unfortunately, the content data often consists of noisy, sparse, and redundant features, which makes the resulting graph unstable and unreliable. In practice, besides the content data, social network data also contain link information, which provides additional information for graph construction. Some of previous works utilize the link data. However, the link data is often incomplete, which makes the resulting graph incomplete. To address these issues, we propose a novel Distilled Graph Clustering (DGC) method. It pursuits adistilled graphbased on both the content data and the link data. The proposed algorithm alternates between two steps: in the feature selection step, it finds the most representative feature subset w.r.t. an intermediate graph initialized with link data; in graph distillation step, the proposed method updates and refines the graph based on only the selected features. The final resulting graph, which is referred to as the distilled graph, is then utilized for spectral clustering on the large-scale social network data. Extensive experiments demonstrate the superiority of the proposed method. Wenhe Liu, Dong Gong, Mingkui Tan, Qinfeng Shi, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Learning Deep Gradient Descent Optimization for Image DeconvolutionabstractAs an integral component of blind image deblurring, non-blind deconvolution removes image blur with a given blur kernel, which is essential but difficult due to the ill-posed nature of the inverse problem. The predominant approach is based on optimization subject to regularization functions that are either manually designed or learned from examples. Existing learning-based methods have shown superior restoration quality but are not practical enough due to their restricted and static model design. They solely focus on learning a prior and require to know the noise level for deconvolution. We address the gap between the optimization- and learning-based approaches by learning a universal gradient descent optimizer. We propose a recurrent gradient descent network (RGDN) by systematically incorporating deep neural networks into a fully parameterized gradient descent scheme. A hyperparameter-free update unit shared across steps is used to generate the updates from the current estimates based on a convolutional neural network. By training on diverse examples, the RGDN learns an implicit image prior and a universal update rule through recursive supervision. The learned optimizer can be repeatedly used to improve the quality of diverse degenerated observations. The proposed method possesses strong interpretability and high generalization. Extensive experiments on synthetic benchmarks and challenging real-world images demonstrate that the proposed deep optimization method is effective and robust to produce favorable results as well as practical for real-world image deblurring applications. Dong Gong, Zhen Zhang 0008, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Knowledge Adaptation for Efficient Semantic SegmentationabstractBoth accuracy and efficiency are of significant importance to the task of semantic segmentation. Existing deep FCNs suffer from heavy computations due to a series of high-resolution feature maps for preserving the detailed knowledge in dense estimation. Although reducing the feature map resolution (i.e., applying a large overall stride) via subsampling operations (e.g., polling and convolution striding) can instantly increase the efficiency, it dramatically decreases the estimation accuracy. To tackle this dilemma, we propose a knowledge distillation method tailored for semantic segmentation to improve the performance of the compact FCNs with large overall stride. To handle the inconsistency between the features of the student and teacher network, we optimize the feature similarity in a transferred latent domain formulated by utilizing a pre-trained autoencoder. Moreover, an affinity distillation module is proposed to capture the long-range dependency by calculating the non local interactions across the whole image. To validate the effectiveness of our proposed method, extensive experiments have been conducted on three popular benchmarks: Pascal VOC, Cityscapes and Pascal Context. Built upon a highly competitive baseline, our proposed method can improve the performance of a student network by 2.5% (mIOU boosts from 70.2 to 72.7 on the cityscapes test set) and can train a better compact model with only 8% float operations (FLOPS) of a model that achieves comparable performances. Tong He 0001, Chunhua Shen, Zhi Tian, Dong Gong, Changming Sun, Youliang Yan |
CVPR | 4 |
| 2019 | RGBD Based Dimensional Decomposition Residual Network for 3D Semantic Scene CompletionabstractRGB images differentiate from depth as they carry more details about the color and texture information, which can be utilized as a vital complement to depth for boosting the performance of 3D semantic scene completion (SSC). SSC is composed of 3D shape completion (SC) and semantic scene labeling while most of the existing approaches use depth as the sole input which causes the performance bottleneck. Moreover, the state-of-the-art methods employ 3D CNNs which have cumbersome networks and tremendous parameters. We introduce a light-weight Dimensional Decomposition Residual network (DDR) for 3D dense prediction tasks. The novel factorized convolution layer is effective for reducing the network parameters, and the proposed multi-scale fusion mechanism for depth and color image can improve the completion and segmentation accuracy simultaneously. Our method demonstrates excellent performance on two public datasets. Compared with the latest method SSCNet, we achieve 5.9% gains in SC-IoU and 5.7% gains in SSC-IOU, albeit with only 21% network parameters and 16.6% FLOPs employed compared with that of SSCNet. Jie Li 0040, Yu Liu 0029, Dong Gong, Qinfeng Shi, Xia Yuan, Chunxia Zhao, Ian D. Reid 0001 |
CVPR | 3 |
| 2019 | Variational Bayesian Dropout With a Hierarchical PriorabstractVariational dropout (VD) is a generalization of Gaussian dropout, which aims at inferring the posterior of network weights based on a log-uniform prior on them to learn these weights as well as dropout rate simultaneously. The log-uniform prior not only interprets the regularization capacity of Gaussian dropout in network training, but also underpins the inference of such posterior. However, the log-uniform prior is an improper prior (i.e., its integral is infinite), which causes the inference of posterior to be ill-posed, thus restricting the regularization performance of VD. To address this problem, we present a new generalization of Gaussian dropout, termed variational Bayesian dropout (VBD), which turns to exploit a hierarchical prior on the network weights and infer a new joint posterior. Specifically, we implement the hierarchical prior as a zero-mean Gaussian distribution with variance sampled from a uniform hyper-prior. Then, we incorporate such a prior into inferring the joint posterior over network weights and the variance in the hierarchical prior, with which both the network training and dropout rate estimation can be cast into a joint optimization problem. More importantly, the hierarchical prior is a proper prior which enables the inference of posterior to be well-posed. In addition, we further show that the proposed VBD can be seamlessly applied to network compression. Experiments on classification and network compression demonstrate the superior performance of the proposed VBD in regularizing network training. Yuhang Liu 0002, Wenyong Dong, Lei Zhang 0054, Dong Gong, Qinfeng Shi |
CVPR | 4 |
| 2019 | Attention-Guided Network for Ghost-Free High Dynamic Range ImagingabstractGhosting artifacts caused by moving objects or misalignments is a key challenge in high dynamic range (HDR) imaging for dynamic scenes. Previous methods first register the input low dynamic range (LDR) images using optical flow before merging them, which are error-prone and cause ghosts in results. A very recent work tries to bypass optical flows via a deep network with skip-connections, however, which still suffers from ghosting artifacts for severe movement. To avoid the ghosting from the source, we propose a novel attention-guided end-to-end deep neural network (AHDRNet) to produce high-quality ghost-free HDR images. Unlike previous methods directly stacking the LDR images or features for merging, we use attention modules to guide the merging according to the reference image. The attention modules automatically suppress undesired components caused by misalignments and saturation and enhance desirable fine details in the non-reference images. In addition to the attention model, we use dilated residual dense block (DRDB) to make full use of the hierarchical features and increase the receptive field for hallucinating the missing details. The proposed AHDRNet is a non-flow-based method, which can also avoid the artifacts generated by optical-flow estimation error. Experiments on different datasets show that the proposed AHDRNet can achieve state-of-the-art quantitative and qualitative results. Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Ian D. Reid 0001, Yanning Zhang 0001 |
CVPR | 2 |
| 2019 | Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly DetectionabstractDeep autoencoder has been extensively used for anomaly detection. Training on the normal data, the autoencoder is expected to produce higher reconstruction error for the abnormal inputs than the normal ones, which is adopted as a criterion for identifying anomalies. However, this assumption does not always hold in practice. It has been observed that sometimes the autoencoder "generalizes" so well that it can also reconstruct anomalies well, leading to the miss detection of anomalies. To mitigate this drawback for autoencoder based anomaly detector, we propose to augment the autoencoder with a memory module and develop an improved autoencoder called memory-augmented autoencoder, i.e. MemAE. Given an input, MemAE firstly obtains the encoding from the encoder and then uses it as a query to retrieve the most relevant memory items for reconstruction. At the training stage, the memory contents are updated and are encouraged to represent the prototypical elements of the normal data. At the test stage, the learned memory will be fixed, and the reconstruction is obtained from a few selected memory records of the normal data. The reconstruction will thus tend to be close to a normal sample. Thus the reconstructed errors on anomalies will be strengthened for anomaly detection. MemAE is free of assumptions on the data type and thus general to be applied to different tasks. Experiments on various datasets prove the excellent generalization and high effectiveness of the proposed MemAE. Dong Gong, Lingqiao Liu, Vuong Le, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh, Anton van den Hengel |
ICCV | 1 |
| 2019 | Robust and Accurate Hybrid Structure-From-MotiabstractIn this paper, we propose a hybrid Structure-from-Motion scheme which combines the strength of both global and local incremental SfM methods to get a drift-free and accurate estimation with lower time consumption. More specifically, we propose to construct a robust maximum leaf spanning tree (RMLST) from the initial scene graph and further expand it to a robust graph (RG) to grasp the global picture of camera distribution and scene structure. Then the views in the robust graph are solved in global manner as an initial estimation. After that, the remaining views are estimated with the proposed community-based local incremental approach to guarantee local accuracy and scalability. Bundle adjustment is conducted to optimize the estimation. Experiments show that our method is robust and free from the scene drift as global SfM, and shows much better efficiency than incremental approaches. Besides, our algorithm achieves higher accuracy compared with the state-of-the-art methods. Rui Li 0013, Dong Gong, Jinqiu Sun, Yu Zhu 0004, Ziwei Wei, Yanning Zhang 0001 |
ICIP | 2 |
| 2019 | Multi-Scale Dense Networks for Deep High Dynamic Range ImagingabstractGenerating a high dynamic range (HDR) image from a set of sequential exposures is a challenging task for dynamic scenes. The most common approaches are aligning the input images to a reference image before merging them into an HDR image, but artifacts often appear in cases of large scene motion. The state-of-the-art method using deep learning can solve this problem effectively. In this paper, we propose a novel deep convolutional neural network to generate HDR, which attempts to produce more vivid images. The key idea of our method is using the coarse-to-fine scheme to gradually reconstruct the HDR image with the multi-scale architecture and residual network. By learning the relative changes of inputs and ground truth, our method can produce not only artificial free image but also restore missing information. Furthermore, we compare to existing methods for HDR reconstruction, and show high-quality results from a set of low dynamic range (LDR) images. We evaluate the results in qualitative and quantitative experiments, our method consistently produces excellent results than existing state-of-the-art approaches in challenging scenes. Qingsen Yan, Dong Gong, Qinfeng Shi, Jinqiu Sun, Ian D. Reid 0001, Yanning Zhang 0001 |
WACV | 2 |
| 2019 | ARSAC: Efficient model estimation via adaptively ranked sample consensus
Rui Li 0013, Jinqiu Sun, Dong Gong, Yu Zhu 0004, Haisen Li, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2019 | MPTV: Matching Pursuit-Based Total Variation Minimization for Image DeconvolutionabstractTotal variation (TV) regularization has proven effective for a range of computer vision tasks through its preferential weighting of sharp image edges. Existing TV-based methods, however, often suffer from the over-smoothing issue and solution bias caused by the homogeneous penalization. In this paper, we consider addressing these issues by applying inhomogeneous regularization on different image components. We formulate the inhomogeneous TV minimization problem as a convex quadratic constrained linear programming problem. Relying on this new model, we propose a matching pursuit-based total variation minimization method (MPTV), specifically for image deconvolution. The proposed MPTV method is essentially a cutting-plane method that iteratively activates a subset of nonzero image gradients and then solves a subproblem focusing on those activated gradients only. Compared with existing methods, the MPTV is less sensitive to the choice of the trade-off parameter between data fitting and regularization. Moreover, the inhomogeneity of MPTV alleviates the over-smoothing and ringing artifacts and improves the robustness to errors in blur kernel. Extensive experiments on different tasks demonstrate the superiority of the proposed method over the current state of the art. Dong Gong, Mingkui Tan, Qinfeng Shi, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | An Adaptive Markov Random Field for Structured Compressive SensingabstractExploiting intrinsic structures in sparse signals underpins the recent progress in compressive sensing (CS). The key for exploiting such structures is to achieve two desirable properties: generality (i.e., the ability to fit a wide range of signals with diverse structures) and adaptability (i.e., being adaptive to a specific signal). Most existing approaches, however, often only achieve one of these two properties. In this study, we propose a novel adaptive Markov random field sparsity prior for CS, which not only is able to capture a broad range of sparsity structures, but also can adapt to each sparse signal through refining the parameters of the sparsity prior with respect to the compressed measurements. To maximize the adaptability, we also propose a new sparse signal estimation where the sparse signals, support, noise and signal parameter estimation are unified into a variational optimization problem, which can be effectively solved with an alternative minimization scheme. Extensive experiments on three real-world datasets demonstrate the effectiveness of the proposed method in recovery accuracy, noise tolerance, and runtime. Suwichaya Suwanwimolkul, Lei Zhang 0054, Dong Gong, Zhen Zhang 0008, Chao Chen 0012, Damith Chinthana Ranasinghe, Qinfeng Shi |
IEEE Trans. Image Process. | 3 |
| 2019 | Two-Stream Convolutional Networks for Blind Image Quality AssessmentabstractTraditional image quality assessment (IQA) methods do not perform robustly due to the shallow hand-designed features. It has been demonstrated that deep neural network can learn more effective features than ever. In this paper, we describe a new deep neural network to predict the image quality accurately without relying on the reference image. To learn more effective feature representations for non-reference IQA, we propose a two-stream convolution network that includes two subcomponents for image and gradient image. The motivation for this design is using a two-stream scheme to capture different-level information of inputs and easing the difficulty of extracting features from one steam. The gradient stream focuses on extracting structure features in details, and the image stream pays more attention to the information in intensity. In addition, to consider the locally non-uniform distribution of distortion in images, we add a region-based fully convolutional layer for using the information around the center of the input image patch. The final score of the overall image is calculated by averaging of the patch scores. The proposed network performs in an end-to-end manner in both the training and testing phases. The experimental results on a series of benchmark datasets, e.g., LIVE, CISQ, IVC, TID2013, and Waterloo Exploration Database, show that the proposed algorithm outperforms the state-of-the-art methods, which verifies the effectiveness of our network architecture. Qingsen Yan, Dong Gong, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Deblurring Natural Image Using Super-Gaussian Fields
Yuhang Liu 0002, Wenyong Dong, Dong Gong, Lei Zhang 0054, Qinfeng Shi |
ECCV (1) | 3 |
| 2018 | Seeing Deeply and Bidirectionally: A Deep Learning Approach for Single Image Reflection Removal
Jie Yang 0002, Dong Gong, Lingqiao Liu, Qinfeng Shi |
ECCV (3) | 2 |
| 2018 | Blind image deblurring by promoting group sparsity
Dong Gong, Rui Li 0013, Yu Zhu 0004, Haisen Li, Jinqiu Sun, Yanning Zhang 0001 |
Neurocomputing | 1 |
| 2017 | MPGL: An Efficient Matching Pursuit Method for Generalized LASSOabstractUnlike traditional LASSO enforcing sparsity on the variables, Generalized LASSO (GL) enforces sparsity on a linear transformation of the variables, gaining flexibility and success in many applications. However, many existing GL algorithms do not scale up to high-dimensional problems, and/or only work well for a specific choice of the transformation. We propose an efficient Matching Pursuit Generalized LASSO (MPGL) method, which overcomes these issues, and is guaranteed to converge to a global optimum. We formulate the GL problem as a convex quadratic constrained linear programming (QCLP) problem and tailor-make a cutting plane method. More specifically, our MPGL iteratively activates a subset of nonzero elements of the transformed variables, and solves a subproblem involving only the activated elements thus gaining significant speed-up. Moreover, MPGL is less sensitive to the choice of the trade-off hyper-parameter between data fitting and regularization, and mitigates the long-standing hyper-parameter tuning issue in many existing methods. Experiments demonstrate the superior efficiency and accuracy of the proposed method over the state-of-the-arts in both classification and image processing tasks. Dong Gong, Mingkui Tan, Yanning Zhang 0001, Anton van den Hengel, Qinfeng Shi |
AAAI | 1 |
| 2017 | From Motion Blur to Motion Flow: A Deep Learning Solution for Removing Heterogeneous Motion BlurabstractRemoving pixel-wise heterogeneous motion blur is challenging due to the ill-posed nature of the problem. The predominant solution is to estimate the blur kernel by adding a prior, but extensive literature on the subject indicates the difficulty in identifying a prior which is suitably informative, and general. Rather than imposing a prior based on theory, we propose instead to learn one from the data. Learning a prior over the latent image would require modeling all possible image content. The critical observation underpinning our approach, however, is that learning the motion flow instead allows the model to focus on the cause of the blur, irrespective of the image content. This is a much easier learning task, but it also avoids the iterative process through which latent image priors are typically applied. Our approach directly estimates the motion flow from the blurred image through a fully-convolutional deep neural network (FCN) and recovers the unblurred image from the estimated motion flow. Our FCN is the first universal end-to-end mapping from the blurred image to the dense motion flow. To train the FCN, we simulate motion flows to generate synthetic blurred-image-motion-flow pairs thus avoiding the need for human labeling. Extensive experiments on challenging realistic blurred images demonstrate that the proposed method outperforms the state-of-the-art. Dong Gong, Jie Yang 0002, Lingqiao Liu, Yanning Zhang 0001, Ian D. Reid 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi |
CVPR | 1 |
| 2017 | Self-Paced Kernel Estimation for Robust Blind Image DeblurringabstractThe challenge in blind image deblurring is to remove the effects of blur with limited prior information about the nature of the blur process. Existing methods often assume that the blur image is produced by linear convolution with additive Gaussian noise. However, including even a small number of outliers to this model in the kernel estimation process can significantly reduce the resulting image quality. Previous methods mainly rely on some simple but unreliable heuristics to identify outliers for kernel estimation. Rather than attempt to identify outliers to the model a priori, we instead propose to sequentially identify inliers, and gradually incorporate them into the estimation process. The self-paced kernel estimation scheme we propose represents a generalization of existing self-paced learning approaches, in which we gradually detect and include reliable inlier pixel sets in a blurred image for kernel estimation. Moreover, we automatically activate a subset of significant gradients w.r.t. the reliable inlier pixels, and then update the intermediate sharp image and the kernel accordingly. Experiments on both synthetic data and real-world images with various kinds of outliers demonstrate the effectiveness and robustness of the proposed method compared to the stateof- the-art methods. Dong Gong, Mingkui Tan, Yanning Zhang 0001, Anton van den Hengel, Qinfeng Shi |
ICCV | 1 |
| 2016 | Blind Image Deconvolution by Automatic Gradient ActivationabstractBlind image deconvolution is an ill-posed inverse problem which is often addressed through the application of appropriate prior. Although some priors are informative in general, many images do not strictly conform to this, leading to degraded performance in the kernel estimation. More critically, real images may be contaminated by nonuniform noise such as saturation and outliers. Methods for removing specific image areas based on some priors have been proposed, but they operate either manually or by defining fixed criteria. We show here that a subset of the image gradients are adequate to estimate the blur kernel robustly, no matter the gradient image is sparse or not. We thus introduce a gradient activation method to automatically select a subset of gradients of the latent image in a cutting-plane-based optimization scheme for kernel estimation. No extra assumption is used in our model, which greatly improves the accuracy and flexibility. More importantly, the proposed method affords great convenience for handling noise and outliers. Experiments on both synthetic data and real-world images demonstrate the effectiveness and robustness of the proposed method in comparison with the state-of-the-art methods. Dong Gong, Mingkui Tan, Yanning Zhang 0001, Anton van den Hengel, Qinfeng Shi |
CVPR | 1 |
| 2014 | Joint Motion Deblurring with Blurred/Noisy Image PairabstractMotion blurred images are widely existing when using a hand-held camera especially under the dim lighting conditions. Since edge information contained in the noisy image may be blurred by the motion blur, a blurred/noisy image pair captured under different exposure time can help to restore a sharp image. In the traditional deblurring methods based on blurred/noisy image pair, the deblurring process is in series with the denoising process, so that restoration result is sensitive to the denoised result. In this paper, we propose a robust algorithm to obtain the sharp image by fusing the blurred image and noisy image. By joint modeling the deblurring model and denoising model, the restoration result can be optimized via estimating the sharp image and blur kernel alternately in the proposed methods, and it is not sensitive to the denoised result benefited by the joint model. Experimental results demonstrated that the proposed method can achieve better performance compared with the state-of-the-art single image denoising methods, single image deblurring methods and blurred/noisy pair deblurring methods. Haisen Li, Yanning Zhang 0001, Jinqiu Sun, Dong Gong |
ICPR | 4 |
| 2013 | Neighbor combination for atmospheric turbulence image reconstructionabstractIn this paper, we propose a novel neighbor combination framework for the reconstruction of the atmospheric turbulence degenerated image sequence. To utilize the spatial and temporal redundancy, a neighbor vector sampling strategy in spatial and temporal domain is conducted relying on the modeling of the registered sequence. Then, a combinator of neighbor vectors is developed based on a resampling maximum likelihood model and a relative approximation. Relying on the neighbor combination and spatial-invariant deconvolution, a clear image is reconstructed. Experiments on real data sets demonstrate the effectiveness of this framework. Dong Gong, Yanning Zhang 0001, Shaobo Dang, Jinqiu Sun |
ICIP | 1 |