VLDB 2026 Research / reviewers in the wild / expert
Tianshu Yu 0001
dblp:152/6675
· DBLP profile ↗
32ranked-venue papers
11as first author
20since 2021 · last 2026
0000-0002-6537-1924ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 10 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physically-Informed Flow Matching with Graph Neural Networks for Complex Fluid DynamicsabstractComputational fluid dynamics (CFD) simulations traditionally require extensive computational resources, limiting their utility in many scientific and engineering applications at scale. We introduce Physically-Informed Flow Matching Graph Networks (PIFM-GN), a novel generative framework that directly samples fluid states under specified physical conditions without requiring expensive time-stepping simulations. The key innovation of our approach is the incorporation of incompressibility constraints directly into the flow matching transport process by parameterizing velocity fields through vector potentials, with graph-based curl operators ensuring divergence-free predictions without requiring global pressure-Poisson solves. Experiments on diverse fluid dynamics problems -- ranging from two-dimensional surface pressure distributions and complete flow fields, to complex three-dimensional airflow fields -- demonstrate that PIFM-GN generates high-fidelity samples with significantly fewer sampling steps than diffusion-based alternatives. Most notably, our model maintains competitive performance even with a single sampling step, a regime where diffusion models completely fail. Our generated samples accurately reproduce the statistical characteristics of target flows, successfully capturing multi-modal pressure distributions across various flow conditions, while achieving significant computational speedups compared to diffusion-based methods. PIFM-GN thus enables efficient generation of fluid states for downstream analysis and design tasks in scientific and engineering applications. Xiaozhuang Song, Tianshu Yu 0001 |
AAAI | 2 |
| 2025 | Enhancing Generalizability in Molecular Conformation Generation with METRIZATION-Informed Geometric Diffusion PretrainingabstractDiffusion-based generative models have recently excelled in generating molecular conformations but struggled with the generalization issue -- models trained on one dataset may produce meaningless conformations on out-of-distribution molecules. On the other hand, distance geometry serves as a generalizable tool for the traditional computational chemistry methods of molecular conformation, which is predicated on the assumption that it is possible to adequately define the set of all potential conformations of any non-rigid molecular system using purely geometric constraints. In this work, we for the first time explicitly incorporate distance geometry constraints into pretraining phase of diffusion-based molecular generation models to improve the generalizability. Inspired by the classical distance geometry solution designed for solving the molecular distance geometry problem, we propose MiGDiff, a Metrization-Informed Geometric Diffusion framework. MiGDiff injects distance geometry constraints by pretraining the deep geometric diffusion backbone within the Metrization sampling approach, yielding a "Metrization-driven pretraining + Data-driven finetuning" paradigm. Experimental results demonstrate that MiGDiff outperforms state-of-the-art methods and possesses strong generalization capabilities, particularly on generating previously unseen molecules, revealing the vast untapped potential of combining traditional computational methods with deep generative models for 3D molecular generation. Xiaozhuang Song, Yuzhao Tu, Hangting Ye, Wei Fan 0010, Tianshu Yu 0001 |
AAAI | 7 |
| 2025 | Generating Physical Dynamics under PriorsabstractGenerating physically feasible dynamics in a data-driven context is challenging, especially when adhering to physical priors expressed in specific equations or formulas. Existing methodologies often overlook the integration of ''physical priors'', resulting in violation of basic physical laws and suboptimal performance. In this paper, we introduce a novel framework that seamlessly incorporates physical priors into diffusion-based generative models to address this limitation. Our approach leverages two categories of priors: 1) distributional priors, such as roto-translational invariance, and 2) physical feasibility priors, including energy and momentum conservation laws and PDE constraints. By embedding these priors into the generative process, our method can efficiently generate physically realistic dynamics, encompassing trajectories and flows. Empirical evaluations demonstrate that our method produces high-quality dynamics across a diverse array of physical phenomena with remarkable robustness, underscoring its potential to advance data-driven studies in AI4Physics. Our contributions signify a substantial advancement in the field of generative modeling, offering a robust solution to generate accurate and physically consistent dynamics. Zihan Zhou 0024, Tianshu Yu 0001 |
ICLR | 3 |
| 2025 | Sampling from Binary Quadratic Distributions via Stochastic LocalizationabstractSampling from binary quadratic distributions (BQDs) is a fundamental but challenging problem in discrete optimization and probabilistic inference. Previous work established theoretical guarantees for stochastic localization (SL) in continuous domains, where MCMC methods efficiently estimate the required posterior expectations during SL iterations. However, achieving similar convergence guarantees for discrete MCMC samplers in posterior estimation presents unique theoretical challenges.
In this work, we present the first application of SL to general BQDs, proving that after a certain number of iterations, the external field of posterior distributions constructed by SL tends to infinity almost everywhere, hence satisfy Poincaré inequalities with probability near to 1, leading to polynomial-time mixing. This theoretical breakthrough enables efficient sampling from general BQDs, even those that may not originally possess fast mixing properties. Furthermore, our analysis, covering enormous discrete MCMC samplers based on Glauber dynamics and Metropolis-Hastings algorithms, demonstrates the broad applicability of our theoretical framework.
Experiments on instances with quadratic unconstrained binary objectives, including maximum independent set, maximum cut, and maximum clique problems, demonstrate consistent improvements in sampling efficiency across different discrete MCMC samplers. Kaiyuan Cui, Tianshu Yu 0001 |
ICML | 4 |
| 2025 | Understanding Oversmoothing in Diffusion-Based GNNs From the Perspective of Operator Semigroup TheoryabstractThis paper presents an analytical study of the oversmoothing issue in diffusion-based Graph Neural Networks (GNNs). Generalizing beyond extant approaches grounded in random walk analysis or particle systems, we approach this problem through operator semigroup theory. This theoretical framework allows us to rigorously prove that oversmoothing is intrinsically linked to the ergodicity of the diffusion operator. Relying on semigroup method, we can quantitatively analyze the dynamic of graph diffusion and give a specific mathematical form of the smoothing feature by ergodicity and invariant measure of operator, which improves previous works only show existence of oversmoothing. This finding further poses a general and mild ergodicity-breaking condition, encompassing the various specific solutions previously offered, thereby presenting a more universal and theoretically grounded approach to relieve oversmoothing in diffusion-based GNNs. Additionally, we offer a probabilistic interpretation of our theory, forging a link with prior works and broadening the theoretical horizon. Our experimental results reveal that this ergodicity-breaking term effectively mitigates oversmoothing measured by Dirichlet energy, and simultaneously enhances performance in node classification tasks. Chenguang Wang 0001, Xinyan Wang 0004, Congying Han, Tiande Guo, Tianshu Yu 0001 |
KDD (1) | 6 |
| 2025 | Graph Learning with Distributional Edge LayoutsabstractGraph Neural Networks (GNNs) learn from graph-structured data by passing messages between neighboring nodes along edges on certain topological layouts. While layouts can be essential to GNNs' performance, extant methods generally consider obtaining layouts from limited perspectives. In this paper, we introduce Distributional Edge Layouts (DELs), a first-of-its-kind method to sample a collection of topological layouts from a Boltzmann distribution under physical energies. By integrating DELs into GNNs, a wide landscape of feasible graph layouts can be captured from a holistic perspective, overcoming the intrinsic drawbacks in existing GNN designs.In practice, DELs can complement various GNN architectures with high versatility. Our theoretical analysis proves that GNNs equipped with DELs maintain at least the same expressive as their original counterparts, with empirical potential offering extra expressivity. Extensive experiments demonstrate that DELs consistently and substantially improve the performance of a wide range of GNN baselines across multiple datasets, achieving state-of-the-art results. This improvement suggests that DELs capture important distributional information previously overlooked by traditional GNN approaches. DEL is open-sourced at https://github.com/LOGO-CUHKSZ/DEL. Xinjian Zhao, Chaolong Ying, Yaoyao Xu, Tianshu Yu 0001 |
KDD (1) | 4 |
| 2025 | TEMPO: Temporal Multi-scale Autoregressive Generation of Protein Conformational EnsemblesabstractUnderstanding the dynamic behavior of proteins is critical to elucidating their functional mechanisms, yet generating realistic, temporally coherent trajectories of protein ensembles remains a significant challenge. In this work, we introduce a novel hierarchical autoregressive framework for modeling protein dynamics that leverages the intrinsic multi-scale organization of molecular motions. Unlike existing methods that focus on generating static conformational ensembles or treat dynamic sampling as an independent process, our approach characterizes protein dynamics as a Markovian process. The framework employs a two-scale architecture: a low-resolution model captures slow, collective motions driving major conformational transitions, while a high-resolution model generates detailed local fluctuations conditioned on these large-scale movements. This hierarchical design ensures that the causal dependencies inherent in protein dynamics are preserved, enabling the generation of temporally coherent and physically realistic trajectories. By bridging high-level biophysical principles with state-of-the-art generative modeling, our approach provides an efficient framework for simulating protein dynamics that balances computational efficiency with physical accuracy. Yaoyao Xu, Zihan Zhou 0024, Tianshu Yu 0001, Mingchen Chen |
NeurIPS | 4 |
| 2025 | Improving Task-Specific Multimodal Sentiment Analysis with General MLLMs via PromptingabstractMultimodal Sentiment Analysis (MSA) aims to predict sentiment from diverse data types, such as video, audio, and language. Recent progress in Multimodal Large Language Models (MLLMs) have demonstrated impressive performance across various tasks. However, in MSA, the increase in computational costs does not always correspond to a significant improvement in performance, raising concerns about the cost-effectiveness of applying MLLMs to MSA. This paper introduces the MLLM-Guided Multimodal Sentiment Learning Framework (MMSLF). It improves the performance of task-specific MSA models by leveraging the generalized knowledge of MLLMs through a teacher-student framework, rather than directly using MLLMs for sentiment prediction. First, the proposed teacher built upon a powerful MLLM (e.g., GPT-4o-mini), guides the student model to align multimodal representations through MLLM-generated context-aware prompts. Then, knowledge distillation enables the student to mimic the teacher’s predictions, thus allowing it to predict sentiment independently without relying on the context-aware prompts. Extensive experiments on the SIMS, MOSI, and MOSEI datasets demonstrate that our framework enables task-specific models to achieve state-of-the-art performance across most metrics. This also provides new insights into the application of general MLLMs for improving MSA. Haoyu Zhang 0001, Chaolong Ying, Tianshu Yu 0001 |
NeurIPS | 5 |
| 2025 | The Underappreciated Power of Vision Models for Graph Structural UnderstandingabstractGraph Neural Networks operate through bottom-up message-passing, fundamentally differing from human visual perception, which intuitively captures global structures first. We investigate the underappreciated potential of vision models for graph understanding, finding they achieve performance comparable to GNNs on established benchmarks while exhibiting distinctly different learning patterns.
These divergent behaviors, combined with limitations of existing benchmarks that conflate domain features with topological understanding, motivate our introduction of GraphAbstract. This benchmark evaluates models' ability to perceive global graph properties as humans do: recognizing organizational archetypes, detecting symmetry, sensing connectivity strength, and identifying critical elements. Our results reveal that vision models significantly outperform GNNs on tasks requiring holistic structural understanding and
maintain generalizability across varying graph scales, while GNNs struggle with global pattern abstraction and degrade with increasing graph size. This work demonstrates that vision models possess remarkable yet underutilized capabilities for graph structural understanding, particularly for problems requiring global topological awareness and scale-invariant reasoning. These findings open new avenues to leverage this underappreciated potential for developing more effective graph foundation models for tasks dominated by holistic pattern recognition. Xinjian Zhao, Zhongkai Xue, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Tianshu Yu 0001 |
NeurIPS | 9 |
| 2024 | Single Cell Gene Expression Prediction via Prototype-based Proximal Neural FactorizationabstractIn the realm of single-cell analysis, accurately predicting gene expressions is crucial for understanding cellular functions and interactions. Traditional approaches often face significant challenges due to intrinsic noise, high dimensionality, and limited data availability in single-cell datasets. On the other hand, deep learning methods are prone to overfitting and perform poorly with limited data. This paper introduces a novel framework, Prototype-based Proximal Neural Factorization (PPNF), which harnesses the power of prototype learning and neural factorization to address these issues. Our method leverages a robust learning paradigm that identifies representative prototypes from single-cell data, facilitating a more resilient and interpretable data representations. We validate our approach using a diverse set of single-cell datasets, demonstrating that our method significantly outperforms existing techniques in terms of both robustness and accuracy. PPNF shows its effectiveness even with limited data, thereby reducing the financial and computational burden associated with high-throughput technologies. By enhancing the robustness and generalizability of single-cell gene expression predictions, our framework provides significant benefits for advancing the analysis and interpretation of single-cell gene expression data, particularly in data-limited scenarios, demonstrating its potential for more cost-effective applications. Xiaozhuang Song, Hangting Ye, Yaoyao Xu, Wei Fan 0010, Tianshu Yu 0001 |
BIBM | 7 |
| 2024 | Neural Enhanced Variational Bayesian Inference on Graphs for Localized Statistical Channel ModelingabstractThis paper proposes an innovative graph neural network (GNN)-based approach to address the challenge of recovering ill-conditioned sparse signals within the task of multi-grid localized statistical channel modeling (LSCM). Our proposed GNN architecture captures the structural sparsity inherent in the channel angular power spectrum (APS) by leveraging reference signal receiving power (RSRP) measured from multiple grids. It can effectively mitigate the severe coherence in the measurement matrix. Furthermore, we present a novel online unsupervised training scheme that enables real-time adaptability for multi-grid LSCM applications. Through extensive simulations, we demonstrate the superior performance of our GNN-based method in the context of multi-grid LSCM, showcasing its advantages over existing sparse recovery techniques. Ye Xue, Tianshu Yu 0001, Qingjiang Shi, Tsung-Hui Chang |
ICC | 4 |
| 2024 | Boosting Graph Pooling with Persistent HomologyabstractRecently, there has been an emerging trend to integrate persistent homology (PH) into graph neural networks (GNNs) to enrich expressive power. However, naively plugging PH features into GNN layers always results in marginal improvement with low interpretability. In this paper, we investigate a novel mechanism for injecting global topological invariance into pooling layers using PH, motivated by the observation that filtration operation in PH naturally aligns graph pooling in a cut-off manner. In this fashion, message passing in the coarsened graph acts along persistent pooled topology, leading to improved performance. Experimentally, we apply our mechanism to a collection of graph pooling methods and observe consistent and substantial performance gain over several popular datasets, demonstrating its wide applicability and flexibility. Chaolong Ying, Xinjian Zhao, Tianshu Yu 0001 |
NeurIPS | 3 |
| 2024 | Towards Robust Multimodal Sentiment Analysis with Incomplete DataabstractThe field of Multimodal Sentiment Analysis (MSA) has recently witnessed an emerging direction seeking to tackle the issue of data incompleteness. Recognizing that the language modality typically contains dense sentiment information, we consider it as the dominant modality and present an innovative Language-dominated Noise-resistant Learning Network (LNLN) to achieve robust MSA. The proposed LNLN features a dominant modality correction (DMC) module and dominant modality based multimodal learning (DMML) module, which enhances the model's robustness across various noise scenarios by ensuring the quality of dominant modality representations. Aside from the methodical design, we perform comprehensive experiments under random data missing scenarios, utilizing diverse and meaningful settings on several popular datasets (e.g., MOSI, MOSEI, and SIMS), providing additional uniformity, transparency, and fairness compared to existing evaluations in the literature. Empirically, LNLN consistently outperforms existing baselines, demonstrating superior performance across these challenging and extensive evaluation metrics. Haoyu Zhang 0001, Wenbin Wang 0001, Tianshu Yu 0001 |
NeurIPS | 3 |
| 2024 | Boosting Protein Language Models with Negative Sample Mining
Yaoyao Xu, Xinjian Zhao, Xiaozhuang Song, Benyou Wang, Tianshu Yu 0001 |
ECML/PKDD (10) | 5 |
| 2024 | Efficient Neural Collaborative Search for Pickup and Delivery ProblemsabstractIn this paper, we introduce Neural Collaborative Search (NCS), a novel learning-based framework for efficiently solving pickup and delivery problems (PDPs). NCS pioneers the collaboration between the latest prevalent neural construction and neural improvement models, establishing a collaborative framework where an improvement model iteratively refines solutions initiated by a construction model. Our NCS collaboratively trains the two models via reinforcement learning with an effective shared-critic mechanism. In addition, the construction model enhances the improvement model with high-quality initial solutions via curriculum learning, while the improvement model accelerates the convergence of the construction model through imitation learning. Besides the new framework design, we also propose the efficient Neural Neighborhood Search (N2S), an efficient improvement model employed within the NCS framework. N2S exploits a tailored Markov decision process formulation and two customized decoders for removing and then reinserting a pair of pickup-delivery nodes, thereby learning a ruin-repair search process for addressing the precedence constraints in PDPs efficiently. To balance the computation cost between encoders and decoders, N2S streamlines the existing encoder design through a light Synthesis Attention mechanism that allows the vanilla self-attention to synthesize various features regarding a route solution. Moreover, a diversity enhancement scheme is further leveraged to ameliorate the performance during the inference of N2S. Our NCS and N2S are both generic, and extensive experiments on two canonical PDP variants show that they can produce state-of-the-art results among existing neural methods. Remarkably, our NCS and N2S could surpass the well-known LKH3 solver especially on the more constrained PDP variant. Detian Kong, Yining Ma 0001, Zhiguang Cao, Tianshu Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | ASP: Learn a Universal Neural Solver!abstractApplying machine learning to combinatorial optimization problems has the potential to improve both efficiency and accuracy. However, existing learning-based solvers often struggle with generalization when faced with changes in problem distributions and scales. In this paper, we propose a new approach called ASP: Adaptive Staircase Policy Space Response Oracle to address these generalization issues and learn a universal neural solver. ASP consists of two components: Distributional Exploration, which enhances the solver's ability to handle unknown distributions using Policy Space Response Oracles, and Persistent Scale Adaption, which improves scalability through curriculum learning. We have tested ASP on several challenging COPs, including the traveling salesman problem, the vehicle routing problem, and the prize collecting TSP, as well as the real-world instances from TSPLib and CVRPLib. Our results show that even with the same model size and weak training signal, ASP can help neural solvers explore and adapt to unseen distributions and varying scales, achieving superior performance. In particular, compared with the same neural solvers under a standard training pipeline, ASP produces a remarkable decrease in terms of the optimality gap with 90.9% and 47.43% on generated instances and real-world instances for TSP, and a decrease of 19% and 45.57% for CVRP. Chenguang Wang 0011, Zhouliang Yu, Stephen McAleer, Tianshu Yu 0001, Yaodong Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment AnalysisabstractThough Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved.To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales.With the obtained hypermodality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA.In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism. Haoyu Zhang 0001, Yu Wang 0246, Guanghao Yin, Kejun Liu, Yuanyuan Liu 0004, Tianshu Yu 0001 |
EMNLP | 6 |
| 2023 | Learning to Decouple Complex SystemsabstractA complex system with cluttered observations may be a coupled mixture of multiple simple sub-systems corresponding to latent entities. Such sub-systems may hold distinct dynamics in the continuous-time domain; therein, complicated interactions between sub-systems also evolve over time. This setting is fairly common in the real world but has been less considered. In this paper, we propose a sequential learning approach under this setting by decoupling a complex system for handling irregularly sampled and cluttered sequential observations. Such decoupling brings about not only subsystems describing the dynamics of each latent entity but also a meta-system capturing the interaction between entities over time. Specifically, we argue that the meta-system evolving within a simplex is governed by projected differential equations (ProjDEs). We further analyze and provide neural-friendly projection operators in the context of Bregman divergence. Experimental results on synthetic and real-world datasets show the advantages of our approach when facing complex and cluttered sequential data compared to the state-of-the-art. Zihan Zhou 0024, Tianshu Yu 0001 |
ICML | 2 |
| 2021 | Combinatorial Learning of Graph Edit Distance via Dynamic EmbeddingabstractGraph Edit Distance (GED) is a popular similarity measurement for pairwise graphs and it also refers to the recovery of the edit path from the source graph to the target graph. Traditional A* algorithm suffers scalability issues due to its exhaustive nature, whose search heuristics heavily rely on human prior knowledge. This paper presents a hybrid approach by combing the interpretability of traditional search-based techniques for producing the edit path, as well as the efficiency and adaptivity of deep embedding models to achieve a cost-effective GED solver. Inspired by dynamic programming, node-level embedding is designated in a dynamic reuse fashion and suboptimal branches are encouraged to be pruned. To this end, our method can be readily integrated into A* procedure in a dynamic fashion, as well as significantly reduce the computational burden with a learned heuristic. Experimental results on different graph datasets show that our approach can remarkably ease the search process of A* without sacrificing much accuracy. To our best knowledge, this work is also the first deep learning-based GED method for recovering the edit path. Runzhong Wang, Tianshu Yu 0001, Junchi Yan, Xiaokang Yang 0001 |
CVPR | 3 |
| 2021 | Deep Latent Graph MatchingabstractDeep learning for graph matching (GM) has emerged as an important research topic due to its superior performance over traditional methods and insights it provides for solving other combinatorial problems on graph. While recent deep methods for GM extensively investigated effective node/edge feature learning or downstream GM solvers given such learned features, there is little existing work questioning if the fixed connectivity/topology typically constructed using heuristics (e.g., Delaunay or k-nearest) is indeed suitable for GM. From a learning perspective, we argue that the fixed topology may restrict the model capacity and thus potentially hinder the performance. To address this, we propose to learn the (distribution of) latent topology, which can better support the downstream GM task. We devise two latent graph generation procedures, one deterministic and one generative. Particularly, the generative procedure emphasizes the across-graph consistency and thus can be viewed as a matching-guided co-generative model. Our methods deliver superior performance over previous state-of-the-arts on public benchmarks, hence supporting our hypothesis. Tianshu Yu 0001, Runzhong Wang, Junchi Yan, Baoxin Li |
ICML | 1 |
| 2020 | Determinant Regularization for Gradient-Efficient Graph MatchingabstractGraph matching refers to finding vertex correspondence for a pair of graphs, which plays a fundamental role in many vision and learning related tasks. Directly applying gradient-based continuous optimization on graph matching can be attractive for its simplicity but calls for effective ways of converting the continuous solution to the discrete one under the matching constraint. In this paper, we show a novel regularization technique with the tool of determinant analysis on the matching matrix which is relaxed into continuous domain with gradient based optimization. Meanwhile we present a theoretical study on the property of our relaxation technique. Our paper strikes an attempt to understand the geometric properties of different regularization techniques and the gradient behavior during the optimization. We show that the proposed regularization is more gradient-efficient than traditional ones during early update stages. The analysis will also bring about insights for other problems under bijection constraints. The algorithm procedure is simple and empirical results on public benchmark show its effectiveness on both synthetic and real-world data. Tianshu Yu 0001, Junchi Yan, Baoxin Li |
CVPR | 1 |
| 2020 | RhyRNN: Rhythmic RNN for Recognizing Events in Long and Complex Videos
Tianshu Yu 0001, Yikang Li 0001, Baoxin Li |
ECCV (10) | 1 |
| 2020 | Deep Learning of Determinantal Point Processes via Proper Spectral Sub-gradient
Tianshu Yu 0001, Yikang Li 0001, Baoxin Li |
ICLR | 1 |
| 2020 | Learning deep graph matching with channel-independent embedding and Hungarian attention
Tianshu Yu 0001, Runzhong Wang, Junchi Yan, Baoxin Li |
ICLR | 1 |
| 2018 | Simultaneous Event Localization and Recognition in Surveillance VideoabstractThe ubiquity of video-based surveillance demands automated approaches to analysis of ever-increasing video footages. Action/Event localization and recognition are two critical capabilities in surveillance video analysis, which have been largely addressed separately in the literature. In this paper, we propose an approach to simultaneously localize and recognize visual events from raw surveillance videos, employing an end-to-end learning strategy. Our approach formulates the task as weakly-supervised sequential semantic segmentation, in which we utilize a specific convolutional RNN to capture not only the appearance and the motion information but also their temporal evolution patterns. We tested our approach on the VIRAT 2.0 dataset. The experimental results, in comparison with relevant existing state-of-the-art, suggest that the proposed approach is promising in delivering a practical solution. Yikang Li 0001, Tianshu Yu 0001, Baoxin Li |
AVSS | 2 |
| 2018 | Joint Cuts and Matching of Partitions in One GraphabstractAs two fundamental problems, graph cuts and graph matching have been intensively investigated over the decades, resulting in vast literature in these two topics respectively. However the way of jointly applying and solving graph cuts and matching receives few attention. In this paper, we first formalize the problem of simultaneously cutting a graph into two partitions i.e. graph cuts and establishing their correspondence i.e. graph matching. Then we develop an optimization algorithm by updating matching and cutting alternatively, provided with theoretical analysis. The efficacy of our algorithm is verified on both synthetic dataset and real-world images containing similar regions or structures. Tianshu Yu 0001, Junchi Yan, Jieyi Zhao, Baoxin Li |
CVPR | 1 |
| 2018 | Incremental Multi-graph Matching via Diversity and Randomness Based Graph Clustering
Tianshu Yu 0001, Junchi Yan, Wei Liu 0005, Baoxin Li |
ECCV (13) | 1 |
| 2018 | Generalizing Graph Matching beyond Quadratic Assignment ModelabstractGraph matching has received persistent attention over decades, which can be formulated as a quadratic assignment problem (QAP). We show that a large family of functions, which we define as Separable Functions, can approximate discrete graph matching in the continuous domain asymptotically by varying the approximation controlling parameters. We also study the properties of global optimality and devise convex/concave-preserving extensions to the widely used Lawler's QAP form. Our theoretical findings show the potential for deriving new algorithms and techniques for graph matching. We deliver solvers based on two specific instances of Separable Functions, and the state-of-the-art performance of our method is verified on popular benchmarks. Tianshu Yu 0001, Junchi Yan, Yilin Wang 0002, Wei Liu 0005, Baoxin Li |
NeurIPS | 1 |
| 2016 | Enhancing scene parsing by transferring structures via efficient low-rank graph matchingabstractScene parsing has attracted significant attention for its practical and theoretical value in computer vision. A typical scene parsing algorithm seeks to densely label pixels or 3-dimensional points from a scene. Traditionally, this procedure relies on a pre-trained classifier to identify the label information, and a smoothing step via Markov Random Field to enhance the consistency. LabelTranfer is a category of scene parsing algorithms to enhance traditional scene parsing framework, by finding dense correspondence and transferring labels across scenes. In this paper, we present a novel scene parsing algorithm which matches maximal similar structures between scenes via efficient low-rank graph matching. The inputs of the algorithm are images, and well- aligned point clouds if available. The images and the point clouds are processed in separate pipelines. The pipeline of images is to learn a reliable classifier and to match local structures via graph matching. The pipeline of point clouds is to conduct preliminary segmentation and to generate feasible label sets. The two pipelines are merged at inference step, in which we elaborate effective and efficient potential functions. We propose a new graph matching model incorporating low-rank and Frobenius regularization, which not only guarantees an accurate solution, but also provides high optimization efficiency via an eigen-decomposition strategy. Several challenging experiments are conducted, showing competitive performance of the proposed method compared to state-of-the-art LabelTransfer algorithm. Further, with point clouds, the performance can be significantly enhanced. Tianshu Yu 0001, Ruisheng Wang 0001 |
SIGSPATIAL/GIS | 1 |
| 2016 | Graph matching with low-rank regularizationabstractGraph matching is a widely researched topic which has been utilized in various applications of computer vision. Due to the combinatorial nature of graph matching, it is NP-hard to find an exact solution. So exact graph matching is always relaxed to inexact graph matching which seeks to find an approximate solution for the original problem. For a matching problem in quadratic form, semidefinite programming (SDP) relaxation is proven to be effective. However, previous SDP relaxation methods discard the constraint that the solution matrix is rank one, because the rank of a matrix is non-convex. In this paper, we explore some good properties of the solution matrix. By relaxing the rank into convex form using the properties, we propose to reformulate the graph matching with low rank constraint into a standard SDP, which can be easily solved. We test our method on both synthetic and real world data. The experimental results demonstrate that our method effectively handles low rank constraint and achieves competitive performance on robustness test against state-of-the-art counterparts. Tianshu Yu 0001, Ruisheng Wang 0001 |
WACV | 1 |
| 2016 | Scene parsing using graph matching on street-view data
Tianshu Yu 0001, Ruisheng Wang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2015 | Face recognition based on subset selection via metric learning on manifoldabstractWith the development of face recognition using sparse representation based classification (SRC), many relevant methods have been proposed and investigated. However, when the dictionary is large and the representation is sparse, only a small proportion of the elements contributes to the l 1-minimization. Under this observation, several approaches have been developed to carry out an efficient element selection procedure before SRC. In this paper, we employ a metric learning approach which helps find the active elements correctly by taking into account the interclass/intraclass relationship and manifold structure of face images. After the metric has been learned, a neighborhood graph is constructed in the projected space. A fast marching algorithm is used to rapidly select the subset from the graph, and SRC is implemented for classification. Experimental results show that our method achieves promising performance and significant efficiency enhancement. Hong Shao, Shuang Chen 0006, Jieyi Zhao, Wencheng Cui, Tianshu Yu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 5 |