EDBT 2026 Demo / reviewers in the wild / expert
Zhen Zhang 0008
dblp:19/5112-8
· DBLP profile ↗
39ranked-venue papers
11as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 11 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Truth as a Trajectory: What Internal Representations Reveal About Large Language Model ReasoningabstractHamed Damirchi, Imezadelajara, Ehsan Abbasnejad, Afshar Shamsi, Zhen Zhang, Javen Qinfeng Shi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hamed Damirchi, Imezadelajara, Ehsan Abbasnejad, Afshar Shamsi, Zhen Zhang 0008, Qinfeng Shi |
ACL (1) | 5 |
| 2026 | Identifying Weight-Variant Latent Causal ModelsabstractThe task of causal representation learning aims to uncover latent higher-level causal variables that affect lower-level observations. Identifying the true latent causal variables from observed data, while allowing instantaneous causal relations among latent variables, remains a challenge, however. To this end, we start with the analysis of three intrinsic indeterminacies in identifying latent variables from observations: transitivity, permutation indeterminacy, and scaling indeterminacy. We find that transitivity acts as a key role in impeding the identifiability of latent causal variables. To address the unidentifiable issue due to transitivity, we introduce a novel identifiability condition where the underlying latent causal model satisfies a linear-Gaussian model, in which the causal coefficients and the distribution of Gaussian noise are modulated by an additional observed variable. Under certain assumptions, including the existence of a reference condition under which latent causal influences vanish, we can show that the latent causal variables can be identified up to trivial permutation and scaling, and that partial identifiability results can still be obtained when this reference condition is violated for a subset of latent variables. Furthermore, based on these theoretical results, we propose a novel method, termed Structural caUsAl Variational autoEncoder (SuaVE), which directly learns causal representations and causal relationships among them, together with the mapping from the latent causal variables to the observed ones. Experimental results on synthetic and real data demonstrate the identifiability and consistency results and the efficacy of SuaVE in learning causal representations. Yuhang Liu 0002, Zhen Zhang 0008, Dong Gong, Mingming Gong, Biwei Huang, Anton van den Hengel, Kun Zhang 0001, Qinfeng Shi |
J. Mach. Learn. Res. | 2 |
| 2026 | Do it yourself dynamic single image super resolution network via ODE
Xiao Zhang 0058, Zhen Zhang 0008, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2025 | Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based InferenceabstractDespite the advancements made in Vision Large Language Models (VLLMs), like text Large Language Models (LLMs), they have limitations in addressing questions that require real-time information or are knowledgeintensive.Indiscriminately adopting Retrieval Augmented Generation (RAG) techniques is an effective yet expensive way to enable models to answer queries beyond their knowledge scopes.To mitigate the dependence on retrieval and simultaneously maintain, or even improve, the performance benefits provided by retrieval, we propose a method to detect the knowledge boundary of VLLMs, allowing for more efficient use of techniques like RAG.Specifically, we propose a method with two variants that finetune a VLLM on an automatically constructed dataset for boundary identification.Experimental results on various types of Visual Question Answering datasets show that our method successfully depicts a VLLM's knowledge boundary, based on which we are able to reduce indiscriminate retrieval while maintaining or improving the performance.In addition, we show that the knowledge boundary identified by our method for one VLLM can be used as a surrogate boundary for other VLLMs.Code will be released at https://github.com/Chord-Che n-30/VLLM-KnowledgeBoundary Xinyu Wang 0013, Yong Jiang 0005, Zhen Zhang 0008, Xinyu Geng, Pengjun Xie, Fei Huang 0002, Kewei Tu |
EMNLP | 4 |
| 2025 | Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
Tianjiao Jiang, Zhen Zhang 0008, Yuhang Liu 0002, Qinfeng Shi |
ICCV | 2 |
| 2025 | Analytic DAG Constraints for Differentiable DAG LearningabstractRecovering the underlying Directed Acyclic Graph (DAG)
structures from observational data presents a formidable challenge, partly due
to the combinatorial nature of the DAG-constrained optimization
problem. Recently, researchers have identified gradient vanishing as
one of the primary obstacles in differentiable DAG learning and have
proposed several DAG constraints to mitigate this issue. By developing
the necessary theory to establish a connection between analytic
functions and DAG constraints, we demonstrate that analytic functions
from the set $\\{f(x) = c_0 + \\sum_{i=1}^{\infty}c_ix^i | \\forall i > 0, c_i > 0; r = \\lim_{i\\rightarrow \\infty}c_{i}/c_{i+1} > 0\\}$ can be employed to
formulate effective DAG constraints. Furthermore, we establish that
this set of functions is closed under several functional operators,
including differentiation, summation, and
multiplication. Consequently, these operators can be leveraged to
create novel DAG constraints based on existing ones. Using these
properties, we design a series of DAG constraints and develop an
efficient algorithm to evaluate them. Experiments
in various settings demonstrate that our DAG constraints
outperform previous state-of-the-art comparators. Our implementation is available at https://github.com/zzhang1987/AnalyticDAGLearning. Zhen Zhang 0008, Ignavier Ng, Dong Gong, Yuhang Liu 0002, Mingming Gong, Biwei Huang, Kun Zhang 0001, Anton van den Hengel, Qinfeng Shi |
ICLR | 1 |
| 2025 | Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning AgentabstractMultimodal Retrieval Augmented Generation (mRAG) plays an important role in mitigating the “hallucination” issue inherent in multimodal large language models (MLLMs). Although promising, existing heuristic mRAGs typically predefined fixed retrieval processes, which causes two issues: (1) Non-adaptive Retrieval Queries. (2) Overloaded Retrieval Queries. However, these flaws cannot be adequately reflected by current knowledge-seeking visual question answering (VQA) datasets, since the most required knowledge can be readily obtained with a standard two-step retrieval. To bridge the dataset gap, we first construct Dyn-VQA dataset, consisting of three types of ``dynamic'' questions, which require complex knowledge retrieval strategies variable in query, tool, and time: (1) Questions with rapidly changing answers. (2) Questions requiring multi-modal knowledge. (3) Multi-hop questions. Experiments on Dyn-VQA reveal that existing heuristic mRAGs struggle to provide sufficient and precisely relevant knowledge for dynamic questions due to their rigid retrieval processes. Hence, we further propose the first self-adaptive planning agent for multimodal retrieval, **OmniSearch**. The underlying idea is to emulate the human behavior in question solution which dynamically decomposes complex multimodal questions into sub-question chains with retrieval action. Extensive experiments prove the effectiveness of our OmniSearch, also provide direction for advancing mRAG. Code and dataset will be open-sourced. Yangning Li, Xinyu Wang 0013, Yong Jiang 0005, Zhen Zhang 0008, Xinran Zheng, Hui Wang 0030, Hai-Tao Zheng 0002, Fei Huang 0002, Jingren Zhou 0001, Philip S. Yu |
ICLR | 5 |
| 2025 | On the Value of Cross-Modal Misalignment in Multimodal Representation LearningabstractMultimodal representation learning, exemplified by multimodal contrastive learning (MMCL) using image-text pairs, aims to learn powerful representations by aligning cues across modalities. This approach relies on the core assumption that the exemplar image-text pairs constitute two representations of an identical concept. However, recent research has revealed that real-world datasets often exhibit cross-modal misalignment. There are two distinct viewpoints on how to address this issue: one suggests mitigating the misalignment, and the other leveraging it. We seek here to reconcile these seemingly opposing perspectives, and to provide a practical guide for practitioners. Using latent variable models we thus formalize cross-modal misalignment by introducing two specific mechanisms: Selection bias, where some semantic variables are absent in the text, and perturbation bias, where semantic variables are altered—both leading to misalignment in data pairs. Our theoretical analysis demonstrates that, under mild assumptions, the representations learned by MMCL capture exactly the information related to the subset of the semantic variables invariant to selection and perturbation biases. This provides a unified perspective for understanding misalignment. Based on this, we further offer actionable insights into how misalignment should inform the design of real-world ML systems. We validate our theoretical findings via extensive empirical studies on both synthetic data and real image-text datasets, shedding light on the nuanced impact of cross-modal misalignment on multimodal representation learning. Yichao Cai 0001, Yuhang Liu 0002, Erdun Gao, Tianjiao Jiang, Zhen Zhang 0008, Anton van den Hengel, Qinfeng Shi |
NeurIPS | 5 |
| 2025 | Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept SpaceabstractHuman cognition typically involves thinking through abstract, fluid concepts rather than strictly using discrete linguistic tokens. Current Large Language Models (LLMs), however, are constrained to reasoning within the boundaries of human language, processing discrete token embeddings that represent fixed points in semantic space. This discrete constraint restricts the expressive power and upper potential of such reasoning models, often causing incomplete exploration of reasoning paths, as standard Chain-of-Thought (CoT) methods rely on sampling one token per step. In this work, we introduce Soft Thinking, a training-free method that emulates human-like ``soft'' reasoning by generating abstract concept tokens in a continuous concept space. These concept tokens are created by the probability-weighted mixture of token embeddings, which span the continuous concept space, enabling smooth transitions and richer representations that transcend traditional discrete boundaries. In essence, each generated concept token encapsulates multiple meanings from related discrete tokens, implicitly exploring various reasoning paths to converge effectively toward the correct answer. Empirical evaluations on diverse mathematical and coding benchmarks consistently demonstrate the effectiveness and efficiency of Soft Thinking, improving pass@1 accuracy by up to 2.48 points while simultaneously reducing token usage by up to 22.4\% compared to standard CoT. Qualitative analysis further reveals that Soft Thinking outputs remain highly interpretable and readable, highlighting the potential of Soft Thinking to break the inherent limits of discrete language-based reasoning. Zhen Zhang 0008, Xuehai He, Weixiang Yan, Xin Wang 0061 |
NeurIPS | 1 |
| 2025 | Solving the Asymmetric Traveling Salesman Problem via Trace-Guided Cost AugmentationabstractThe Asymmetric Traveling Salesman Problem (ATSP) ranks among the most fundamental and notoriously difficult problems in combinatorial optimization. We propose a novel continuous relaxation framework for the Asymmetric Traveling Salesman Problem (ATSP) by leveraging differentiable constraints that encourage acyclic structures and valid permutations. Our approach integrates a differentiable trace-based Directed Acyclic Graph (DAG) constraint with a doubly stochastic matrix relaxation of the assignment problem, enabling gradient-based optimization over soft permutations. We develop a projected exponentiated gradient method with adaptive step size to minimize tour cost while satisfying the relaxed constraints. To recover high-quality discrete tours, we introduce a greedy post-processing procedure that iteratively corrects subtours using cost-aware cycle merging. Our method achieves state-of-the-art performance on standard asymmetric TSP benchmarks and demonstrates competitive scalability and accuracy, particularly on large or asymmetric instances where heuristic solvers such as LKH-3 struggle. Zhen Zhang 0008, Qinfeng Shi, Wee Sun Lee |
NeurIPS | 1 |
| 2025 | Exploring Training and Inference Scaling Laws in Generative RetrievalabstractGenerative retrieval reformulates retrieval as an autoregressive generation task, where large language models (LLMs) generate target documents directly from a query. As a novel paradigm, the mechanisms that underpin its performance and scalability remain largely unexplored. We systematically investigate training and inference scaling laws in generative retrieval, exploring how model size, training data scale, and inference-time compute jointly influence performance. We propose a novel evaluation metric inspired by contrastive entropy and generation loss, providing a continuous performance signal that enables robust comparisons across diverse generative retrieval methods. Our experiments show that n-gram-based methods align strongly with training and inference scaling laws. We find that increasing model size, training data scale, and inference-time compute all contribute to improved performance, highlighting the complementary roles of these factors in enhancing generative retrieval. Across these settings, LLaMA models consistently outperform T5 models, suggesting a particular advantage for larger decoder-only models in generative retrieval. Our findings underscore that model sizes, data availability, and inference computation interact to unlock the full potential of generative retrieval, offering new insights for designing and optimizing future systems. We release code at SLGR GitHub repository. Hongru Cai, Yongqi Li 0001, Ruifeng Yuan, Wenjie Wang 0007, Zhen Zhang 0008, Wenjie Li 0002, Tat-Seng Chua |
SIGIR | 5 |
| 2025 | Leapfrog Polymorphic Neural Ordinary Differential EquationabstractNeural Ordinary Differential Equations (NODEs) revolutionize the way we view residual networks as solvers for initial value problems (IVPs), with layer depth serving as the time step. In this study, we propose a more efficient extension of NODEs called Leap-Frog Polymorphic Neural ODEs (LF-NODEs). LF-NODEs introduce the leap-frog updating scheme, breaking free from the specific structure of time-evolving mixtures of multiple dynamical systems. In each time step (corresponding to each residual network layer), LF-NODEs integrate the features of multiple dynamical systems (residual network layers) by employing time-variant weights from different dynamical systems. Our LF-NODEs not only overcome the limitations in the representation power of individual residual networks but also provide a more flexible and effective structure for residual networks based on multiple dynamical systems. The efficiency of the leap-frog updating scheme is theoretically derived and demonstrated. Furthermore, we expand the model by incorporating a multi-scale approach (LF-MSNODEs) to achieve even more accurate model representations. Empirical results showcase the improvements provided by our proposed methods in 2D concentric spheres task, irregularly sampled time series prediction, and classification tasks. Additionally, the results provide empirical evidence that the learned feature space enhances system efficiency with fewer parameters and function evaluations compared to the baseline methods. Xiao Zhang 0058, Wei Wei 0008, Zhen Zhang 0008, Lei Zhang 0054 |
IEEE Signal Process. Lett. | 3 |
| 2024 | CLAP: Isolating Content from Style Through Contrastive Learning with Augmented Prompts
Yichao Cai 0001, Yuhang Liu 0002, Zhen Zhang 0008, Qinfeng Shi |
ECCV (21) | 3 |
| 2024 | Identifiable Latent Polynomial Causal Models through the Lens of ChangeabstractCausal representation learning aims to unveil latent high-level causal representations from observed low-level data. One of its primary tasks is to provide reliable assurance of identifying these latent causal models, known as \textit{identifiability}. A recent breakthrough explores identifiability by leveraging the change of causal influences among latent causal variables across multiple environments \citep{liu2022identifying}. However, this progress rests on the assumption that the causal relationships among latent causal variables adhere strictly to linear Gaussian models. In this paper, we extend the scope of latent causal models to involve nonlinear causal relationships, represented by polynomial models, and general noise distributions conforming to the exponential family. Additionally, we investigate the necessity of imposing changes on all causal parameters and present partial identifiability results when part of them remains unchanged. Further, we propose a novel empirical estimation method, grounded in our theoretical finding, that enables learning consistent latent causal representations. Our experimental results, obtained from both synthetic and real-world data, validate our theoretical contributions concerning identifiability and consistency. Yuhang Liu 0002, Zhen Zhang 0008, Dong Gong, Mingming Gong, Biwei Huang, Anton van den Hengel, Kun Zhang 0001, Qinfeng Shi |
ICLR | 2 |
| 2024 | A Causal Inspired Early-Branching Structure for Domain Generalization
Liang Chen 0030, Yong Zhang 0034, Yibing Song, Zhen Zhang 0008, Lingqiao Liu |
Int. J. Comput. Vis. | 4 |
| 2023 | Factor Graph Neural NetworksabstractIn recent years, we have witnessed a surge of Graph Neural Networks (GNNs), most of which can learn powerful representations in an end-to-end fashion with great success in many real-world applications. They have resemblance to Probabilistic Graphical Models (PGMs), but break free from some limitations of PGMs. By aiming to provide expressive methods for representation learning instead of computing marginals or most likely configurations, GNNs provide flexibility in the choice of information flowing rules while maintaining good performance. Despite their success and inspirations, they lack efficient ways to represent and learn higher-order relations among variables/nodes. More expressive higher-order GNNs which operate on k-tuples of nodes need increased computational resources in order to process higher-order tensors. We propose Factor Graph Neural Networks (FGNNs) to effectively capture higher-order relations for inference and learning. To do so, we first derive an efficient approximate Sum-Product loopy belief propagation inference algorithm for discrete higher-order PGMs. We then neuralize the novel message passing scheme into a Factor Graph Neural Network (FGNN) module by allowing richer representations of the message update rules; this facilitates both efficient inference and powerful end-to-end learning. We further show that with a suitable choice of message aggregation operators, our FGNN is also able to represent Max-Product belief propagation, providing a single family of architecture that can represent both Max and Sum-Product loopy belief propagation. Our extensive experimental evaluation on synthetic as well as real datasets demonstrates the potential of the proposed model. Zhen Zhang 0008, Mohammed Haroon Dupty, Fan Wu 0011, Qinfeng Shi, Wee Sun Lee |
J. Mach. Learn. Res. | 1 |
| 2023 | Human Interaction Understanding With Consistency-Aware LearningabstractCompared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main causation is that recent approaches learn human interactive relations via shallow graphical representations, which are inadequate to model complicated human interactive-relations. This paper proposes a deep consistency-aware framework aiming at tackling the grouping and labelling inconsistencies in HIU. This framework consists of three components, including a backbone CNN to extract image features, a factor graph network to implicitly learn higher-order consistencies among labelling and grouping variables, and a consistency-aware reasoning module to explicitly enforcing consistencies. The last module is inspired by our key observation that the consistency-aware reasoning bias can be embedded into an energy function or a particular loss function, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained in an end-to-end fashion. Experimental results demonstrate that the two proposed consistency-learning modules complement each other, and both make considerable contributions in achieving leading performance on three benchmarks of HIU. The effectiveness of the proposed approach is further validated by experiments on detecting human-object interactions. Jiajun Meng, Zhenhua Wang 0003, Kaining Ying, Jianhua Zhang 0002, Dongyan Guo, Zhen Zhang 0008, Qinfeng Shi, Shengyong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Milstein-driven neural stochastic differential equation model with uncertainty estimates
Xiao Zhang 0058, Wei Wei 0008, Zhen Zhang 0008, Lei Zhang 0054, Wei Li 0219 |
Pattern Recognit. Lett. | 3 |
| 2023 | SharpFormer: Learning Local Feature Preserving Global Representations for Image DeblurringabstractThe goal of dynamic scene deblurring is to remove the motion blur presented in a given image. To recover the details from the severe blurs, conventional convolutional neural networks (CNNs) based methods typically increase the number of convolution layers, kernel-size, or different scale images to enlarge the receptive field. However, these methods neglect the non-uniform nature of blurs, and cannot extract varied local and global information. Unlike the CNNs-based methods, we propose a Transformer-based model for image deblurring, named SharpFormer, that directly learns long-range dependencies via a novel Transformer module to overcome large blur variations. Transformer is good at learning global information but is poor at capturing local information. To overcome this issue, we design a novel Locality preserving Transformer (LTransformer) block to integrate sufficient local information into global features. In addition, to effectively apply LTransformer to the medium-resolution features, a hybrid block is introduced to capture intermediate mixed features. Furthermore, we use a dynamic convolution (DyConv) block, which aggregates multiple parallel convolution kernels to handle the non-uniform blur of inputs. We leverage a powerful two-stage attentive framework composed of the above blocks to learn the global, hybrid, and local features effectively. Extensive experiments on the GoPro and REDS datasets show that the proposed SharpFormer performs favourably against the state-of-the-art methods in blurred image restoration. Qingsen Yan, Dong Gong, Zhen Zhang 0008, Yanning Zhang 0001, Qinfeng Shi |
IEEE Trans. Image Process. | 4 |
| 2022 | Truncated Matrix Power Iteration for Differentiable DAG LearningabstractRecovering underlying Directed Acyclic Graph (DAG) structures from observational data is highly challenging due to the combinatorial nature of the DAG-constrained optimization problem. Recently, DAG learning has been cast as a continuous optimization problem by characterizing the DAG constraint as a smooth equality one, generally based on polynomials over adjacency matrices. Existing methods place very small coefficients on high-order polynomial terms for stabilization, since they argue that large coefficients on the higher-order terms are harmful due to numeric exploding. On the contrary, we discover that large coefficients on higher-order terms are beneficial for DAG learning, when the spectral radiuses of the adjacency matrices are small, and that larger coefficients for higher-order terms can approximate the DAG constraints much better than the small counterparts. Based on this, we propose a novel DAG learning method with efficient truncated matrix power iteration to approximate geometric series based DAG constraints. Empirically, our DAG learning method outperforms the previous state-of-the-arts in various settings, often by a factor of $3$ or more in terms of structural Hamming distance. Zhen Zhang 0008, Ignavier Ng, Dong Gong, Yuhang Liu 0002, Ehsan Abbasnejad, Mingming Gong, Kun Zhang 0001, Qinfeng Shi |
NeurIPS | 1 |
| 2022 | Sparse Structure Search for Delta TuningabstractAdapting large pre-trained models (PTMs) through fine-tuning imposes prohibitive computational and storage burdens. Recent studies of delta tuning (DT), i.e., parameter-efficient tuning, find that only optimizing a small portion of parameters conditioned on PTMs could yield on-par performance compared to conventional fine-tuning. Generally, DT methods exquisitely design delta modules (DT modules) which could be applied to arbitrary fine-grained positions inside PTMs. However, the effectiveness of these fine-grained positions largely relies on sophisticated manual designation, thereby usually producing sub-optimal results. In contrast to the manual designation, we explore constructing DT modules in an automatic manner. We automatically \textbf{S}earch for the \textbf{S}parse \textbf{S}tructure of \textbf{Delta} Tuning (S$^3$Delta). Based on a unified framework of various DT methods, S$^3$Delta conducts the differentiable DT structure search through bi-level optimization and proposes shifted global sigmoid method to explicitly control the number of trainable parameters. Extensive experiments show that S$^3$Delta surpasses manual and random structures with less trainable parameters. The searched structures preserve more than 99\% fine-tuning performance with 0.01\% trainable parameters. Moreover, the advantage of S$^3$Delta is amplified with extremely low trainable parameters budgets (0.0009\%$\sim$0.01\%). The searched structures are transferable and explainable, providing suggestions and guidance for the future design of DT methods. Our codes are publicly available at \url{https://github.com/thunlp/S3Delta}. Shengding Hu, Zhen Zhang 0008, Ning Ding 0002, Yadao Wang, Yasheng Wang, Zhiyuan Liu 0001, Maosong Sun 0001 |
NeurIPS | 2 |
| 2021 | Memory-augmented Dynamic Neural Relational InferenceabstractDynamic interacting systems are prevalent in vision tasks. These interactions are usually difficult to observe and measure directly, and yet understanding latent interactions is essential for performing inference tasks on dynamic systems like forecasting. Neural relational inference (NRI) techniques are thus introduced to explicitly estimate interpretable relations between the entities in the system for trajectory prediction. However, NRI assumes static relations; thus, dynamic neural relational inference (DNRI) was proposed to handle dynamic relations using LSTM. Unfortunately, the older information will be washed away when the LSTM updates the latent variable as a whole, which is why DNRI struggles with modeling long-term dependences and forecasting long sequences. This motivates us to propose a memory-augmented dynamic neural relational inference method, which maintains two associative memory pools: one for the interactive relations and the other for the individual entities. The two memory pools help retain useful relation features and node features for the estimation in the future steps. Our model dynamically estimates the relations by learning better embeddings and utilizing the long-range information stored in the memory. With the novel memory modules and customized structures, our memory-augmented DNRI can update and access the memory adaptively as required. The memory pools also serve as global latent variables across time to maintain detailed long-term temporal relations readily available for other components to use. Experiments on synthetic and real-world datasets show the effectiveness of the proposed method on modeling dynamic relations and forecasting complex trajectories. Dong Gong, Zhen Zhang 0008, Qinfeng Shi, Anton van den Hengel |
ICCV | 2 |
| 2021 | A bidirectional graph neural network for traveling salesman problems on arbitrary symmetric graphs
Yujiao Hu, Zhen Zhang 0008, Yuan Yao 0004, Xingpeng Huyan, Xingshe Zhou 0001, Wee Sun Lee |
Eng. Appl. Artif. Intell. | 2 |
| 2020 | Visual Relationship Detection with Low Rank Non-Negative Tensor DecompositionabstractWe address the problem of Visual Relationship Detection (VRD) which aims to describe the relationships between pairs of objects in the form of triplets of (subject, predicate, object). We observe that given a pair of bounding box proposals, objects often participate in multiple relations implying the distribution of triplets is multimodal. We leverage the strong correlations within triplets to learn the joint distribution of triplet variables conditioned on the image and the bounding box proposals, doing away with the hitherto used independent distribution of triplets. To make learning the triplet joint distribution feasible, we introduce a novel technique of learning conditional triplet distributions in the form of their normalized low rank non-negative tensor decompositions. Normalized tensor decompositions take form of mixture distributions of discrete variables and thus are able to capture multimodality. This allows us to efficiently learn higher order discrete multimodal distributions and at the same time keep the parameter size manageable. We further model the probability of selecting an object proposal pair and include a relation triplet prior in our model. We show that each part of the model improves performance and the combination outperforms state-of-the-art score on the Visual Genome (VG) and Visual Relationship Detection (VRD) datasets. Mohammed Haroon Dupty, Zhen Zhang 0008, Wee Sun Lee |
AAAI | 2 |
| 2020 | Multiplicative Gaussian Particle FilterabstractWe propose a new sampling-based approach for approximate inference in filtering problems. Instead of approximating conditional distributions with a finite set of states, as done in particle filters, our approach approximates the distribution with a weighted sum of functions from a set of continuous functions. Central to the approach is the use of sampling to approximate multiplications in the Bayes filter. We provide theoretical analysis, giving conditions for sampling to give good approximation. We next specialize to the case of weighted sums of Gaussians, and show how properties of Gaussians enable closed-form transition and efficient multiplication. Lastly, we conduct preliminary experiments on a robot localization problem and compare performance with the particle filter, to demonstrate the potential of the proposed method. Xuan Su, Wee Sun Lee, Zhen Zhang 0008 |
AISTATS | 3 |
| 2020 | Factor Graph Neural NetworksabstractMost of the successful deep neural network architectures are structured, often consisting of elements like convolutional neural networks and gated recurrent neural networks. Recently, graph neural networks (GNNs) have been successfully applied to graph-structured data such as point cloud and molecular data. These networks often only consider pairwise dependencies, as they operate on a graph structure. We generalize the GNN into a factor graph neural network (FGNN) providing a simple way to incorporate dependencies among multiple variables. We show that FGNN is able to represent Max-Product belief propagation, an approximate inference method on probabilistic graphical models, providing a theoretical understanding on the capabilities of FGNN and related GNNs. Experiments on synthetic and real datasets demonstrate the potential of the proposed architecture. Zhen Zhang 0008, Fan Wu 0011, Wee Sun Lee |
NeurIPS | 1 |
| 2020 | Learning Deep Gradient Descent Optimization for Image DeconvolutionabstractAs an integral component of blind image deblurring, non-blind deconvolution removes image blur with a given blur kernel, which is essential but difficult due to the ill-posed nature of the inverse problem. The predominant approach is based on optimization subject to regularization functions that are either manually designed or learned from examples. Existing learning-based methods have shown superior restoration quality but are not practical enough due to their restricted and static model design. They solely focus on learning a prior and require to know the noise level for deconvolution. We address the gap between the optimization- and learning-based approaches by learning a universal gradient descent optimizer. We propose a recurrent gradient descent network (RGDN) by systematically incorporating deep neural networks into a fully parameterized gradient descent scheme. A hyperparameter-free update unit shared across steps is used to generate the updates from the current estimates based on a convolutional neural network. By training on diverse examples, the RGDN learns an implicit image prior and a universal update rule through recursive supervision. The learned optimizer can be repeatedly used to improve the quality of diverse degenerated observations. The proposed method possesses strong interpretability and high generalization. Extensive experiments on synthetic benchmarks and challenging real-world images demonstrate that the proposed deep optimization method is effective and robust to produce favorable results as well as practical for real-world image deblurring applications. Dong Gong, Zhen Zhang 0008, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Deep Graphical Feature Learning for the Feature Matching ProblemabstractThe feature matching problem is a fundamental problem in various areas of computer vision including image registration, tracking and motion analysis. Rich local representation is a key part of efficient feature matching methods. However, when the local features are limited to the coordinate of key points, it becomes challenging to extract rich local representations. Traditional approaches use pairwise or higher order handcrafted geometric features to get robust matching; this requires solving NP-hard assignment problems. In this paper, we address this problem by proposing a graph neural network model to transform coordinates of feature points into local features. With our local features, the traditional NP-hard assignment problems are replaced with a simple assignment problem which can be solved efficiently. Promising results on both synthetic and real datasets demonstrate the effectiveness of the proposed method. Zhen Zhang 0008, Wee Sun Lee |
ICCV | 1 |
| 2019 | An Adaptive Markov Random Field for Structured Compressive SensingabstractExploiting intrinsic structures in sparse signals underpins the recent progress in compressive sensing (CS). The key for exploiting such structures is to achieve two desirable properties: generality (i.e., the ability to fit a wide range of signals with diverse structures) and adaptability (i.e., being adaptive to a specific signal). Most existing approaches, however, often only achieve one of these two properties. In this study, we propose a novel adaptive Markov random field sparsity prior for CS, which not only is able to capture a broad range of sparsity structures, but also can adapt to each sparse signal through refining the parameters of the sparsity prior with respect to the compressed measurements. To maximize the adaptability, we also propose a new sparse signal estimation where the sparse signals, support, noise and signal parameter estimation are unified into a variational optimization problem, which can be effectively solved with an alternative minimization scheme. Extensive experiments on three real-world datasets demonstrate the effectiveness of the proposed method in recovery accuracy, noise tolerance, and runtime. Suwichaya Suwanwimolkul, Lei Zhang 0054, Dong Gong, Zhen Zhang 0008, Chao Chen 0012, Damith Chinthana Ranasinghe, Qinfeng Shi |
IEEE Trans. Image Process. | 4 |
| 2018 | Convolutional Sequence to Sequence Model for Human DynamicsabstractHuman motion modeling is a classic problem in computer vision and graphics. Challenges in modeling human motion include high dimensional prediction as well as extremely complicated dynamics.We present a novel approach to human motion modeling based on convolutional neural networks (CNN). The hierarchical structure of CNN makes it capable of capturing both spatial and temporal correlations effectively. In our proposed approach, a convolutional long-term encoder is used to encode the whole given motion sequence into a long-term hidden variable, which is used with a decoder to predict the remainder of the sequence. The decoder itself also has an encoder-decoder structure, in which the short-term encoder encodes a shorter sequence to a short-term hidden variable, and the spatial decoder maps the long and short-term hidden variable to motion predictions. By using such a model, we are able to capture both invariant and dynamic information of human motion, which results in more accurate predictions. Experiments show that our algorithm outperforms the state-of-the-art methods on the Human3.6M and CMU Motion Capture datasets. Our code is available at the project website. Chen Li 0038, Zhen Zhang 0008, Wee Sun Lee, Gim Hee Lee |
CVPR | 2 |
| 2018 | Understanding human activities in videos: A joint action and interaction learning approach
Zhenhua Wang 0003, Jiali Jin, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Zhen Zhang 0008, Dongyan Guo, Zhanpeng Shao |
Neurocomputing | 7 |
| 2017 | Solving Constrained Combinatorial Optimisation Problems via MAP Inference without High-Order PenaltiesabstractSolving constrained combinatorial optimisation problems via MAP inference is often achieved by introducing extra potential functions for each constraint. This can result in very high order potentials, e.g. a 2nd-order objective with pairwise potentials and a quadratic constraint over all N variables would correspond to an unconstrained objective with an order-N potential. This limits the practicality of such an approach, since inference with high order potentials is tractable only for a few special classes of functions. We propose an approach which is able to solve constrained combinatorial problems using belief propagation without increasing the order. For example, in our scheme the 2nd-order problem above remains order 2 instead of order N. Experiments on applications ranging from foreground detection, image reconstruction, quadratic knapsack, and the M-best solutions problem demonstrate the effectiveness and efficiency of our method. Moreover, we show several situations in which our approach outperforms commercial solvers like CPLEX and others designed for specific constrained MAP inference problems. Zhen Zhang 0008, Qinfeng Shi, Julian J. McAuley, Wei Wei 0008, Yanning Zhang 0001, Rui Yao 0006, Anton van den Hengel |
AAAI | 1 |
| 2017 | Dynamic Programming Bipartite Belief Propagation For Hyper Graph MatchingabstractHyper graph matching problems have drawn attention recently due to their ability to embed higher order relations between nodes. In this paper, we formulate hyper graph matching problems as constrained MAP inference problems in graphical models. Whereas previous discrete approaches introduce several global correspondence vectors, we introduce only one global correspondence vector, but several local correspondence vectors. This allows us to decompose the problem into a (linear) bipartite matching problem and several belief propagation sub-problems. Bipartite matching can be solved by traditional approaches, while the belief propagation sub-problem is further decomposed as two sub-problems with optimal substructure. Then a newly proposed dynamic programming procedure is used to solve the belief propagation sub-problem. Experiments show that the proposed methods outperform state-of-the-art techniques for hyper graph matching. Zhen Zhang 0008, Julian J. McAuley, Wei Wei 0008, Yanning Zhang 0001, Qinfeng Shi |
IJCAI | 1 |
| 2017 | Real-Time Correlation Filter Tracking by Efficient Dense Belief Propagation With Structure PreservingabstractPatch-based models that combine local image features or regions into loose geometric assemblies are a powerful paradigm for visual object tracking, and they present favorable properties such as robustness to partial occlusion, deformation, and the ability to address viewpoint changes. However, effectively exploiting the spatial-temporal confidence scores of each patch to construct a robust tracker while ensuring a low computational cost with dense discrete search remains a challenging problem. In this paper, we propose a unified Markov random field (MRF) model that can effectively capture spatio-temporal intrapatch relations and occlusion priors to enhance the tracking performance, and we derive a highly efficient dense belief propagation for inference of the proposed MRF model. We propose a tracker that models the tracking object with a constellation topology (i.e., a global object and several local patches), where the graph structure model describes the pairwise spatial structure and the image observation model corresponding to the correlation filter with occlusion handling measuring the appearance similarity. Furthermore, these two models are updated online by exploiting the flexibility of local patches. Extensive experimental results on single- and multiple-object tracking show that the proposed algorithm performs favorably against state-of-the-art methods and runs in real time. Rui Yao 0006, Shixiong Xia, Zhen Zhang 0008, Yanning Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | Joint Probabilistic Matching Using m-Best SolutionsabstractMatching between two sets of objects is typically approached by finding the object pairs that collectively maximize the joint matching score. In this paper, we argue that this single solution does not necessarily lead to the optimal matching accuracy and that general one-to-one assignment problems can be improved by considering multiple hypotheses before computing the final similarity measure. To that end, we propose to utilize the marginal distributionsfor each entity. Previously, this idea has been neglected mainly because exact marginalization is intractable due to a combinatorial number of all possible matching permutations. Here, we propose a generic approach to efficiently approximate the marginal distributions by exploiting the m-best solutions of the original problem. This approach not only improves the matching solution, but also provides more accurate ranking of the results, because of the extra information included in the marginal distribution. We validate our claim on two distinct objectives: (i) person re-identification and temporal matching modeled as an integer linear program, and (ii) feature point matching using a quadratic cost function. Our experiments confirm that marginalization indeed leads to superior performance compared to the single (nearly) optimal solution, yielding state-of-the-art results in both applications on standard benchmarks. Seyed Hamid Rezatofighi, Anton Milan, Zhen Zhang 0008, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001 |
CVPR | 3 |
| 2016 | Pairwise Matching through Max-Weight Bipartite Belief PropagationabstractFeature matching is a key problem in computer vision and pattern recognition. One way to encode the essential interdependence between potential feature matches is to cast the problem as inference in a graphical model, though recently alternatives such as spectral methods, or approaches based on the convex-concave procedure have achieved the state-of-the-art. Here we revisit the use of graphical models for feature matching, and propose a belief propagation scheme which exhibits the following advantages: (1) we explicitly enforce one-to-one matching constraints, (2) we offer a tighter relaxation of the original cost function than previous graphical-model-based approaches, and (3) our sub-problems decompose into max-weight bipartite matching, which can be solved efficiently, leading to orders-of-magnitude reductions in execution time. Experimental results show that the proposed algorithm produces results superior to those of the current state-of-the-art. Zhen Zhang 0008, Qinfeng Shi, Julian J. McAuley, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel |
CVPR | 1 |
| 2015 | Learning graph structure for multi-label image classification via clique generationabstractExploiting label dependency for multi-label image classification can significantly improve classification performance. Probabilistic Graphical Models are one of the primary methods for representing such dependencies. The structure of graphical models, however, is either determined heuristically or learned from very limited information. Moreover, neither of these approaches scales well to large or complex graphs. We propose a principled way to learn the structure of a graphical model by considering input features and labels, together with loss functions. We formulate this problem into a max-margin framework initially, and then transform it into a convex programming problem. Finally, we propose a highly scalable procedure that activates a set of cliques iteratively. Our approach exhibits both strong theoretical properties and a significant performance improvement over state-of-the-art methods on both synthetic and real-world data sets. Mingkui Tan, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Junbin Gao, Fuyuan Hu, Zhen Zhang 0008 |
CVPR | 7 |
| 2015 | Joint Probabilistic Data Association RevisitedabstractIn this paper, we revisit the joint probabilistic data association (JPDA) technique and propose a novel solution based on recent developments in finding the m-best solutions to an integer linear program. The key advantage of this approach is that it makes JPDA computationally tractable in applications with high target and/or clutter density, such as spot tracking in fluorescence microscopy sequences and pedestrian tracking in surveillance footage. We also show that our JPDA algorithm embedded in a simple tracking framework is surprisingly competitive with state-of-the-art global tracking methods in these two applications, while needing considerably less processing time. Seyed Hamid Rezatofighi, Anton Milan, Zhen Zhang 0008, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001 |
ICCV | 3 |
| 2014 | Variable kernel density estimation based robust regression and its applications
Zhen Zhang 0008, Yanning Zhang 0001 |
Neurocomputing | 1 |