VLDB 2026 Research / reviewers in the wild / expert
Tianfu Wu 0001
dblp:08/4148-1 · also Matt Tianfu Wu
· DBLP profile ↗
67ranked-venue papers
7as first author
30since 2021 · last 2025
0000-0001-8911-5506ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 7 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 48 · 3 first-author · 18 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploratory Analysis of the Regulation of Long Non-Coding RNA Transcription with Nucleotide Large Language ModelsabstractLarge language models (LLMs) have emerged as powerful tools for biological sequence analysis. However, their applicability to the transcriptional regulation of long non-coding RNAs (IncRNAs) remains underexplored due to the complexity and diversity of IncRNA sequences, combined with limited knowledge of their regulatory mechanisms and functional characteristics. In this study, we systematically evaluate both singletask and multi-task fine-tuning strategies of genome foundation models across four tasks designed to capture increasing biological complexity. By fine-tuning genome foundation models on a series of progressively complex tasks, each designed to closely mimic the complexities of IncRNA classification, we explore how task complexity impacts model performance and biological interpretability. Our findings reveal that while foundation models capture promoter-specific signals, task complexity, input length, and data quality significantly influence performance. Notably, the most challenging task, promoter sequence classification of IncRNA from coding genes, remains limited by primary sequence features alone. Through performance evaluation and attentionbased interpretability analysis, this work provides key insights into the potential and limitations of current LLM-based models and offers a foundation for future development of integrative models in IncRNA regulation analysis. The dataset is available at https://github.com/wangwe90/IncRNA-promoter-sequence-detection-dataset Wei Wang 0010, Zhichao Hou, Tianfu Wu 0001, Xinxia Peng |
BIBM | 3 |
| 2025 | ScaleLSD: Scalable Deep Line Segment Detection StreamlinedabstractThis paper studies the problem of Line Segment Detection (LSD) for the characterization of line geometry in images, with the aim of learning a domain-agnostic robust LSD model that works well for any natural images. With the focus of scalable self-supervised learning of LSD, we revisit and streamline the fundamental designs of (deep and non-deep) LSD approaches to have a high-performing and efficient LSD learner, dubbed as ScaleLSD, for the curation of line geometry at scale from over 10M unlabeled real-world images. Our ScaleLSD works very well to detect much more number of line segments from any natural images even than the pioneered non-deep LSD approach, having a more complete and accurate geometric characterization of images using line segments. Experimentally, our proposed ScaleLSD is comprehensively testified under zero-shot protocols in detection performance, single-view 3D geometry estimation, two-view line segment matching, and multiview 3D line mapping, all with excellent performance obtained. Based on the thorough evaluation, our ScaleLSD is observed to be the first deep approach that outperforms the pioneered non-deep LSD in all aspects we have tested, significantly expanding and reinforcing the versatility of the line geometry of images. Zeran Ke, Bin Tan 0002, Xianwei Zheng, Yujun Shen, Tianfu Wu 0001, Nan Xue 0001 |
CVPR | 5 |
| 2025 | GSBAK: top-K Geometric Score-based Black-box AttackabstractExisting score-based adversarial attacks mainly focus on crafting $top$-1 adversarial examples against classifiers with single-label classification. Their attack success rate and query efficiency are often less than satisfactory, particularly under small perturbation requirements; moreover, the vulnerability of classifiers with multi-label learning is yet to be studied. In this paper, we propose a comprehensive surrogate free score-based attack, named \b geometric \b score-based \b black-box \b attack (GSBA$^K$), to craft adversarial examples in an aggressive $top$-$K$ setting for both untargeted and targeted attacks, where the goal is to change the $top$-$K$ predictions of the target classifier. We introduce novel gradient-based methods to find a good initial boundary point to attack. Our iterative method employs novel gradient estimation techniques, particularly effective in $top$-$K$ setting, on the decision boundary to effectively exploit the geometry of the decision boundary. Additionally, GSBA$^K$ can be used to attack against classifiers with $top$-$K$ multi-label learning. Extensive experiential results on ImageNet and PASCAL VOC datasets validate the effectiveness of GSBA$^K$ in crafting $top$-$K$ adversarial examples. Md Farhamdur Reza, Richeng Jin, Tianfu Wu 0001, Huaiyu Dai |
ICLR | 3 |
| 2025 | Adversarial Perturbations Are Formed by Iteratively Learning Linear Combinations of the Right Singular Vectors of the Adversarial JacobianabstractWhite-box targeted adversarial attacks reveal core vulnerabilities in Deep Neural Networks (DNNs), yet two key challenges persist: (i) How many target classes can be attacked simultaneously in a specified order, known as the ordered top-$K$ attack problem ($K \geq 1$)? (ii) How to compute the corresponding adversarial perturbations for a given benign image directly in the image space? We address both by showing that ordered top-$K$ perturbations can be learned via iteratively optimizing linear combinations of the $\underline{ri}ght\text{ } \underline{sing}ular$ vectors of the adversarial Jacobian (i.e., the logit-to-image Jacobian constrained by target ranking). These vectors span an orthogonal, informative subspace in the image domain. We introduce RisingAttacK, a novel Sequential Quadratic Programming (SQP)-based method that exploits this structure. We propose a holistic figure-of-merits (FoM) metric combining attack success rates (ASRs) and $\ell_p$-norms ($p=1,2,\infty$). Extensive experiments on ImageNet-1k across six ordered top-$K$ levels ($K=1, 5, 10, 15, 20, 25, 30$) and four models (ResNet-50, DenseNet-121, ViT-B, DEiT-B) show RisingAttacK consistently surpasses the state-of-the-art QuadAttacK. Thomas Paniagua, Chinmay Savadikar, Tianfu Wu 0001 |
ICML | 3 |
| 2025 | WeGeFT: Weight‑Generative Fine-Tuning for Multi-Faceted Efficient Adaptation of Large ModelsabstractFine-tuning large pretrained Transformer models can focus on either introducing a small number of new learnable parameters (parameter efficiency) or editing representations of a small number of tokens using lightweight modules (representation efficiency). While the pioneering method LoRA (Low-Rank Adaptation) inherently balances parameter, compute, and memory efficiency, many subsequent variants trade off compute and memory efficiency and/or performance to further reduce fine-tuning parameters. To address this limitation and unify parameter-efficient and representation-efficient fine-tuning, we propose Weight-Generative Fine-Tuning (WeGeFT, pronounced wee-gift), a novel approach that learns to generate fine-tuning weights directly from the pretrained weights. WeGeFT employs a simple low-rank formulation consisting of two linear layers, either shared across multiple layers of the pretrained model or individually learned for different layers. This design achieves multi-faceted efficiency in parameters, representations, compute, and memory, while maintaining or exceeding the performance of LoRA and its variants. Extensive experiments on commonsense reasoning, arithmetic reasoning, instruction following, code generation, and visual recognition verify the effectiveness of our proposed WeGeFT. Chinmay Savadikar, Tianfu Wu 0001 |
ICML | 3 |
| 2025 | Hierarchical Ensemble Based Clustering For Networked MicrogridsabstractTo accommodate the exponential integration of distributed energy resources (DERs) into the power grid, there is a pressing need for a scalable and computationally efficient networked microgrid energy management framework. In this context, a Hierarchical Distributed Consensus (HDC) based approach has emerged as a promising solution. A key prerequisite to ensure an effective implementation of the HDC framework is to obtain an optimal hierarchical partition from the network. This paper presents a framework for determining the optimal number of clusters using the ensemble method to form a hierarchical structured networked microgrid. The proposed algorithm incorporates the consensus matrix and the elbow method to analyze and pinpoint the best possible network configuration. The simulation results highlight the importance of obtaining the optimal number of clusters and interdependence of cluster configuration to speed of consensus convergence in the network. Aditya Joshi 0003, Tianfu Wu 0001, Mo-Yuen Chow |
IECON | 2 |
| 2025 | DiffMesh: A Motion-Aware Diffusion Framework for Human Mesh Recovery from VideosabstractHuman mesh recovery (HMR) provides rich human body information for various real-world applications such as gaming, human-computer interaction, and virtual reality. While image-based HMR methods have achieved impressive results, they often struggle to recover humans in dynamic scenarios, leading to temporal inconsistencies and non-smooth 3D motion predictions due to the absence of human motion. In contrast, video-based approaches leverage temporal information to mitigate this issue. In this paper, we present DiffMesh, an innovative motion-aware diffusion framework for video-based HMR. DiffMesh establishes a bridge between diffusion models and human motion, efficiently generating accurate and smooth output mesh sequences by incorporating human motion within the forward process and reverse process in the diffusion model. Extensive experiments are conducted on the widely used datasets (Human3.6M [15] and 3DPW [48]), which demonstrate the effectiveness and efficiency of our DiffMesh. Visual comparisons in real-world scenarios further highlight DiffMesh's suitability for practical applications. The project webpage is: https://zczcwh.github.io/diffmesh_page/ Xianpeng Liu, Qucheng Peng, Tianfu Wu 0001, Pu Wang 0001, Chen Chen 0001 |
WACV | 4 |
| 2025 | Sign-Based Gradient Descent With Heterogeneous Data: Convergence and Byzantine ResilienceabstractCommunication overhead has become one of the major bottlenecks in the distributed training of modern deep neural networks. With such consideration, various quantization-based stochastic gradient descent (SGD) solvers have been proposed and widely adopted, among which SignSGD with majority vote shows a promising direction because of its communication efficiency and robustness against Byzantine attackers. However, SignSGD fails to converge in the presence of data heterogeneity, which is commonly observed in the emerging federated learning (FL) paradigm. In this article, a sufficient condition for the convergence of the sign-based gradient descent method is derived, based on which a novel magnitude-driven stochastic-sign-based gradient compressor is proposed to address the non-convergence issue of SignSGD. The convergence of the proposed method is established in the presence of arbitrary data heterogeneity. The Byzantine resilience of sign-based gradient descent methods is quantified, and the error-feedback mechanism is further incorporated to boost the learning performance. Experimental results on the MNIST dataset, the CIFAR-10 dataset, and the Tiny-ImageNet dataset corroborate the effectiveness of the proposed methods. Richeng Jin, Yuding Liu, Yufan Huang, Xiaofan He, Tianfu Wu 0001, Huaiyu Dai |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | NEAT: Distilling 3D Wireframes from Neural Attraction FieldsabstractThis paper studies the problem of structured 3D reconstruction using wireframes that consist of line segments and junctions, focusing on the computation of structured boundary geometries of scenes. Instead of leveraging matching-based solutions from 2D wireframes (or line segments) for 3D wireframe reconstruction as done in prior arts, we present NEAT, a rendering-distilling formulation using neural fields to represent 3D line segments with 2D observations, and bipartite matching for perceiving and distilling of a sparse set of 3D global junctions. The proposed NEAT enjoys the joint optimization of the neural fields and the global junctions from scratch, using view-dependent 2D observations without precomputed cross-view feature matching. Comprehensive experiments on the DTU and BlendedMVS datasets demonstrate our NEAT's superiority over state-of-the-art alternatives for 3D wireframe reconstruction. Moreover, the distilled 3D global junctions by NEAT, are a better initialization than SfM points, for the recently-emerged 3D Gaussian Splatting for high-fidelity novel view synthesis using about 20 times fewer initial 3D points. Project page: https://xuenan.net/neat. Nan Xue 0001, Bin Tan 0002, Yuxi Xiao, Gui-Song Xia, Tianfu Wu 0001, Yujun Shen |
CVPR | 6 |
| 2024 | Multi-View Attentive Contextualization for Multi-View 3D Object DetectionabstractWe present Multi-View Attentive Contextualization (MvACon), a simple yet effective method for improving 2D-to-3D feature lifting in query-based multi-view 3D (MV3D) object detection. Despite remarkable progress witnessed in the field of query-based MV3D object detection, prior art often suffers from either the lack of exploiting high-resolution 2D features in dense attention-based lifting, due to high computational costs, or from insufficiently dense grounding of 3D queries to multi-scale 2D features in sparse attention-based lifting. Our proposed MvACon hits the two birds with one stone using a representationally dense yet computationally sparse attentive feature contextualization scheme that is agnostic to specific 2D-to-3D feature lifting approaches. In experiments, the proposed MvA-Con is thoroughly tested on the nuScenes benchmark, using both the BEVFormer and its recent 3D deformable attention (DFA3D) variant, as well as the PETR, showing consistent detection performance improvement, especially in enhancing performance in location, orientation, and velocity prediction. It is also tested on the Waymo-mini benchmark using BEVFormer with similar improvement. We qualitatively and quantitatively show that global cluster-based contexts effectively encode dense scene-level contexts for MV3D object detection. The promising results of our proposed MvA-Con reinforces the adage in computer vision - “(contextualized) feature matters”. Xianpeng Liu, Ming Qian, Nan Xue 0001, Chen Chen 0001, Zhebin Zhang, Tianfu Wu 0001 |
CVPR | 8 |
| 2023 | PaCa-ViT: Learning Patch-to-Cluster Attention in Vision TransformersabstractVision Transformers (ViTs) are built on the assumption of treating image patches as “visual tokens” and learn patch-to-patch attention. The patch embedding based tokenizer has a semantic gap with respect to its counterpart, the textual tokenizer. The patch-to-patch attention suffers from the quadratic complexity issue, and also makes it non-trivial to explain learned ViTs. To address these issues in ViT, this paper proposes to learn Patch-to-Cluster attention (PaCa) in ViT. Queries in our PaCa-ViT starts with patches, while keys and values are directly based on clustering (with a predefined small number of clusters). The clusters are learned end-to-end, leading to better tokenizers and inducing joint clustering-for-attention and attention-for-clustering for better and interpretable models. The quadratic complexity is relaxed to linear complexity. The proposed PaCa module is used in designing efficient and interpretable ViT backbones and semantic segmentation head networks. In experiments, the proposed methods are tested on ImageNet-1k image classification, MS-COCO object detection and instance segmentation and MIT-ADE20k semantic segmentation. Compared with the prior art, it obtains better performance in all the three benchmarks than the SWin [32] and the PVTs [47], [48] by significant margins in ImageNet-1k and MIT-ADE20k. It is also significantly more efficient than PVT models in MS-COCO and MIT-ADE20k due to the linear complexity. The learned clusters are semantically meaningful. Code and model checkpoints are available at https:/github.com/iVMCL/PaCaViT. Ryan Grainger, Thomas Paniagua, Naresh P. Cuntoor, Mun Wai Lee, Tianfu Wu 0001 |
CVPR | 6 |
| 2023 | Level-S2fM: Structure from Motion on Neural Level Set of Implicit SurfacesabstractThis paper presents a neural incremental Structure-from-Motion (SfM) approach, Level-S2fM, which estimates the camera poses and scene geometry from a set of uncalibrated images by learning coordinate MLPs for the implicit surfaces and the radiance fields from the established key-point correspondences. Our novel formulation poses some new challenges due to inevitable two-view and few-view configurations in the incremental SfM pipeline, which complicates the optimization of coordinate MLPs for volumetric neural rendering with unknown camera poses. Nevertheless, we demonstrate that the strong inductive basis conveying in the 2D correspondences is promising to tackle those challenges by exploiting the relationship between the ray sampling schemes. Based on this, we revisit the pipeline of incremental SfM and renew the key components, including two-view geometry initialization, the camera poses registration, the 3D points triangulation, and Bundle Adjustment, with a fresh perspective based on neural implicit surfaces. By unifying the scene geometry in small MLP networks through coordinate MLPs, our Level-S2fM treats the zero-level set of the implicit surface as an informative top-down regularization to manage the reconstructed 3D points, reject the outliers in correspondences via querying SDF, and refine the estimated geometries by NBA (Neural BA). Not only does our Level-S2fM lead to promising results on camera pose estimation and scene geometry reconstruction, but it also shows a promising way for neural implicit rendering without knowing camera extrinsic beforehand. Yuxi Xiao, Nan Xue 0001, Tianfu Wu 0001, Gui-Song Xia |
CVPR | 3 |
| 2023 | Implicit Bayes Adaptation: A Collaborative Transport ApproachabstractThe power and flexibility of Optimal Transport (OT) have pervaded a wide spectrum of problems, including recent Machine Learning challenges such as unsupervised domain adaptation. Its essence of quantitatively relating two probability distributions by some optimal metric, has been creatively exploited and shown to hold promise for many real-world data challenges. In a related theme in the present work, we posit that domain adaptation robustness is rooted in the intrinsic (latent) representations of the respective data, which are inherently lying in a non-linear submanifold embedded in a higher dimensional Euclidean space. We account for the geometric properties by refining the l2Euclidean metric to better reflect the geodesic distance between two distinct representations. We integrate a metric correction term as well as a prior cluster structure in the source data of the OT-driven adaptation. We show that this is tantamount to an implicit Bayesian framework, which we demonstrate to be viable for a more robust and better-performing approach to domain adaptation. Substantiating experiments are also included for validation purposes. Hamid Krim, Tianfu Wu 0001, Derya Cansever |
ICASSP | 3 |
| 2023 | Monocular 3D Object Detection with Bounding Box Denoising in 3D by PerceiverabstractThe main challenge of monocular 3D object detection is the accurate localization of 3D center. Motivated by a new and strong observation that this challenge can be remedied by a 3D-space local-grid search scheme in an ideal case, we propose a stage-wise approach, which combines the information flow from 2D-to-3D (3D bounding box proposal generation with a single 2D image) and 3D-to-2D (proposal verification by denoising with 3D-to-2D contexts) in a top-down manner. Specifically, we first obtain initial proposals from off-the-shelf backbone monocular 3D detectors. Then, we generate a 3D anchor space by local-grid sampling from the initial proposals. Finally, we perform 3D bounding box denoising at the 3D-to-2D proposal verification stage. To effectively learn discriminative features for denoising highly overlapped proposals, this paper presents a method of using the Perceiver I/O model [20] to fuse the 3D-to-2D geometric information and the 2D appearance information. With the encoded latent representation of a proposal, the verification head is implemented with a self-attention module. Our method, named as MonoXiver, is generic and can be easily adapted to any backbone monocular 3D detectors. Experimental results on the well-established KITTI dataset and the challenging large-scale Waymo dataset show that MonoXiver consistently achieves improvement with limited computation overhead. Xianpeng Liu, Kelvin Cheng 0003, Nan Xue 0001, Guo-Jun Qi, Tianfu Wu 0001 |
ICCV | 6 |
| 2023 | CGBA: Curvature-aware Geometric Black-box AttackabstractDecision-based black-box attacks often necessitate a large number of queries to craft an adversarial example. Moreover, decision-based attacks based on querying boundary points in the estimated normal vector direction often suffer from inefficiency and convergence issues. In this paper, we propose a novel query-efficient curvature-aware geometric decision-based black-box attack (CGBA) that conducts boundary search along a semicircular path on a restricted 2D plane to ensure finding a boundary point successfully irrespective of the boundary curvature. While the proposed CGBA attack can work effectively for an arbitrary decision boundary, it is particularly efficient in exploiting the low curvature to craft high-quality adversarial examples, which is widely seen and experimentally verified in commonly used classifiers under non-targeted attacks. In contrast, the decision boundaries often exhibit higher curvature under targeted attacks. Thus, we develop a new query-efficient variant, CGBA-H, that is adapted for the targeted attack. In addition, we further design an algorithm to obtain a better initial boundary point at the expense of some extra queries, which considerably enhances the performance of the targeted attack. Extensive experiments are conducted to evaluate the performance of our proposed methods against some well-known classifiers on the ImageNet and CIFAR10 datasets, demonstrating the superiority of CGBA and CGBA-H over state-of-the-art non-targeted and targeted attacks, respectively. The source code is available at https://github.com/Farhamdur/CGBA. Md Farhamdur Reza, Ali Rahmati, Tianfu Wu 0001, Huaiyu Dai |
ICCV | 3 |
| 2023 | Learning Spatially-Adaptive Style-Modulation Networks for Single Image SynthesisabstractRecently there has been a growing interest in learning generative models from a single image. This task is important as in many real world applications, collecting large dataset is not feasible. Existing work like SinGAN is able to synthesize novel images that resemble the patch distribution of the training image. However, SinGAN cannot learn high level semantics of the image, and thus their synthesized samples tend to have unrealistic spatial layouts. To address this issue, this paper proposes a spatially adaptive style-modulation (SASM) module that learns to preserve realistic spatial configuration of images. Specifically, it extracts style vector (in the form of channel-wise attention) and latent spatial mask (in the form of spatial attention) from a coarse level feature separately. The style vector and spatial mask are then aggregated to modulate features of deeper layers. The disentangled modulation of spatial and style attributes enables the model to preserve the spatial structure of the image without overfitting. Experimental results show that the proposed module learns to generate samples with better fidelity than prior works. Jianghao Shen, Tianfu Wu 0001 |
ICIP | 2 |
| 2023 | Learning Spatially-Adaptive Squeeze-Excitation Networks for Few Shot Image SynthesisabstractLearning light-weight yet expressive deep networks for image synthesis is a challenging problem. Inspired by a recent observation that it is the data-specificity that makes the multi-head self-attention (MHSA) in the Transformer model so powerful, this paper proposes to extend the widely adopted light-weight Squeeze-Excitation (SE) module to be spatially-adaptive to reinforce its data specificity, as a convolutional alternative of the MHSA, while retaining the efficiency of SE and the inductive bias of convolution. It proposes a spatially-adaptive squeeze-excitation (SASE) module for image synthesis task.SASE is tested in low-shot image generative learning task, and shows better performance than prior arts. Jianghao Shen, Tianfu Wu 0001 |
ICIP | 2 |
| 2023 | QuadAttacK: A Quadratic Programming Approach to Learning Ordered Top-K Adversarial AttacksabstractThe adversarial vulnerability of Deep Neural Networks (DNNs) has been well-known and widely concerned, often under the context of learning top-$1$ attacks (e.g., fooling a DNN to classify a cat image as dog). This paper shows that the concern is much more serious by learning significantly more aggressive ordered top-$K$ clear-box targeted attacks proposed in~\citep{zhang2020learning}. We propose a novel and rigorous quadratic programming (QP) method of learning ordered top-$K$ attacks with low computing cost, dubbed as \textbf{QuadAttac$K$}. Our QuadAttac$K$ directly solves the QP to satisfy the attack constraint in the feature embedding space (i.e., the input space to the final linear classifier), which thus exploits the semantics of the feature embedding space (i.e., the principle of class coherence). With the optimized feature embedding vector perturbation, it then computes the adversarial perturbation in the data space via the vanilla one-step back-propagation. In experiments, the proposed QuadAttac$K$ is tested in the ImageNet-1k classification using ResNet-50, DenseNet-121, and Vision Transformers (ViT-B and DEiT-S). It successfully pushes the boundary of successful ordered top-$K$ attacks from $K=10$ up to $K=20$ at a cheap budget ($1\times 60$) and further improves attack success rates for $K=5$ for all tested models, while retaining the performance for $K=1$. Thomas Paniagua, Ryan Grainger, Tianfu Wu 0001 |
NeurIPS | 3 |
| 2023 | NOPE-SAC: Neural One-Plane RANSAC for Sparse-View Planar 3D ReconstructionabstractThis article studies the challenging two-view 3D reconstruction problem in a rigorous sparse-view configuration, which is suffering from insufficient correspondences in the input image pairs for camera pose estimation. We present a novel Neural One-PlanE RANSAC framework (termed NOPE-SAC in short) that exerts excellent capability of neural networks to learn one-plane pose hypotheses from 3D plane correspondences. Building on the top of a Siamese network for plane detection, our NOPE-SAC first generates putative plane correspondences with a coarse initial pose. It then feeds the learned 3D plane correspondences into shared MLPs to estimate the one-plane camera pose hypotheses, which are subsequently reweighed in a RANSAC manner to obtain the final camera pose. Because the neural one-plane pose minimizes the number of plane correspondences for adaptive pose hypotheses generation, it enables stable pose voting and reliable pose refinement with a few of plane correspondences for the sparse-view inputs. In the experiments, we demonstrate that our NOPE-SAC significantly improves the camera pose estimation for the two-view inputs with severe viewpoint changes, setting several new state-of-the-art performances on two challenging benchmarks, i.e., MatterPort3D and ScanNet, for sparse-view 3D reconstruction. Bin Tan 0002, Nan Xue 0001, Tianfu Wu 0001, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Holistically-Attracted Wireframe Parsing: From Supervised to Self-Supervised LearningabstractThis article presents Holistically-Attracted Wireframe Parsing (HAWP), a method for geometric analysis of 2D images containing wireframes formed by line segments and junctions. HAWP utilizes a parsimonious Holistic Attraction (HAT) field representation that encodes line segments using a closed-form 4D geometric vector field. The proposed HAWP consists of three sequential components empowered by end-to-end and HAT-driven designs: 1) generating a dense set of line segments from HAT fields and endpoint proposals from heatmaps, 2) binding the dense line segments to sparse endpoint proposals to produce initial wireframes, and 3) filtering false positive proposals through a novel endpoint-decoupled line-of-interest aligning (EPD LOIAlign) module that captures the co-occurrence between endpoint proposals and HAT fields for better verification. Thanks to our novel designs, HAWPv2 shows strong performance in fully supervised learning, while HAWPv3 excels in self-supervised learning, achieving superior repeatability scores and efficient training (24 GPU hours on a single GPU). Furthermore, HAWPv3 exhibits a promising potential for wireframe parsing in out-of-distribution images without providing ground truth labels of wireframes. Nan Xue 0001, Tianfu Wu 0001, Song Bai 0001, Fudong Wang 0001, Gui-Song Xia, Liangpei Zhang 0001, Philip Torr 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | HoW-3D: Holistic 3D Wireframe Perception from a Single ImageabstractThis paper studies the problem of holistic 3D wireframe perception (HoW-3D), a new task of perceiving both the visible 3D wireframes and the invisible ones from single-view 2D images. As the non-front surfaces of an object cannot be directly observed in a single view, estimating the nonline-of-sight (NLOS) geometries in HoW-3D is a fundamentally challenging problem and remains open in computer vision. We study the problem of HoW-3D by proposing an ABC-HoW benchmark, which is created on top of CAD models sourced from the ABC-dataset with 12k single-view images and the corresponding holistic 3D wireframe models. With our large-scale ABC-HoW benchmark available, we present a novel Deep Spatial Gestalt (DSG) model to learn the visible junctions and line segments as the basis and then infer the NLOS 3D structures from the visible cues by following the Gestalt principles of human vision systems. In our experiments, we demonstrate that our DSG model performs very well in inferring the holistic 3D wireframes from single-view images. Compared with the strong baseline methods, our DSG model outperforms the previous wire-frame detectors in detecting the invisible line geometry in single-view images and is even very competitive with prior arts that take high-fidelity PointCloud as inputs on reconstructing 3D wireframes. Bin Tan 0002, Nan Xue 0001, Tianfu Wu 0001, Xianwei Zheng, Gui-Song Xia |
3DV | 4 |
| 2022 | Learning Auxiliary Monocular Contexts Helps Monocular 3D Object DetectionabstractMonocular 3D object detection aims to localize 3D bounding boxes in an input single 2D image. It is a highly challenging problem and remains open, especially when no extra information (e.g., depth, lidar and/or multi-frames) can be leveraged in training and/or inference. This paper proposes a simple yet effective formulation for monocular 3D object detection without exploiting any extra information. It presents the MonoCon method which learns Monocular Contexts, as auxiliary tasks in training, to help monocular 3D object detection. The key idea is that with the annotated 3D bounding boxes of objects in an image, there is a rich set of well-posed projected 2D supervision signals available in training, such as the projected corner keypoints and their associated offset vectors with respect to the center of 2D bounding box, which should be exploited as auxiliary tasks in training. The proposed MonoCon is motivated by the Cramer–Wold theorem in measure theory at a high level. In implementation, it utilizes a very simple end-to-end design to justify the effectiveness of learning auxiliary monocular contexts, which consists of three components: a Deep Neural Network (DNN) based feature backbone, a number of regression head branches for learning the essential parameters used in the 3D bounding box prediction, and a number of regression head branches for learning auxiliary contexts. After training, the auxiliary context regression branches are discarded for better inference efficiency. In experiments, the proposed MonoCon is tested in the KITTI benchmark (car, pedestrian and cyclist). It outperforms all prior arts in the leaderboard on the car category and obtains comparable performance on pedestrian and cyclist in terms of accuracy. Thanks to the simple design, the proposed MonoCon method obtains the fastest inference speed with 38.7 fps in comparisons. Our code is released at https://git.io/MonoCon. Xianpeng Liu, Nan Xue 0001, Tianfu Wu 0001 |
AAAI | 3 |
| 2022 | Learning Local-Global Contextual Adaptation for Multi-Person Pose EstimationabstractThis paper studies the problem of multi-person pose estimation in a bottom-up fashion. With a new and strong observation that the localization issue of the center-offset formulation can be remedied in a local-window search scheme in an ideal situation, we propose a multi-person pose estimation approach, dubbed as LOGO-CAP, by learning the LOcal-GlObal Contextual Adaptation for human Pose. Specifically, our approach learns the keypoint attraction maps (KAMs) from the local keypoints expansion maps (KEMs) in small local windows in the first step, which are subsequently treated as dynamic convolutional kernels on the keypoints-focused global heatmaps for contextual adaptation, achieving accurate multi-person pose estimation. Our method is end-to-end trainable with near real-time inference speed in a single forward pass, obtaining state-of-the-art performance on the COCO keypoint benchmark for bottom-up human pose estimation. With the COCO trained model, our method also outperforms prior arts by a large margin on the challenging OCHuman dataset. Nan Xue 0001, Tianfu Wu 0001, Gui-Song Xia, Liangpei Zhang 0001 |
CVPR | 2 |
| 2022 | Refining Self-Supervised Learning in Imaging: Beyond Linear MetricabstractWe introduce in this paper a new statistical perspective, exploiting the Jaccard similarity metric, as a measure-based metric to effectively invoke non-linear features in the loss of self-supervised contrastive learning. Specifically, our proposed metric may be interpreted as a dependence measure between two adapted projections learned from the so-called latent representations. This is in contrast to the cosine similarity measure in the conventional contrastive learning model, which accounts for correlation information. To the best of our knowledge, this effectively non-linearly fused information embedded in the Jaccard similarity, is novel to self-supervision learning with promising results. The proposed approach is compared to two state-of-the-art self-supervised contrastive learning methods on three image datasets. We not only demonstrate its amenable applicability in current ML problems, but also its improved performance and training efficiency. Hamid Krim, Tianfu Wu 0001, Derya Cansever |
ICIP | 3 |
| 2022 | Revisiting Non-Parametric Matching Cost Volumes for Robust and Generalizable Stereo MatchingabstractStereo matching is a classic challenging problem in computer vision, which has recently witnessed remarkable progress by Deep Neural Networks (DNNs). This paradigm shift leads to two interesting and entangled questions that have not been addressed well. First, it is unclear whether stereo matching DNNs that are trained from scratch really learn to perform matching well. This paper studies this problem from the lens of white-box adversarial attacks. It presents a method of learning stereo-constrained photometrically-consistent attacks, which by design are weaker adversarial attacks, and yet can cause catastrophic performance drop for those DNNs. This observation suggests that they may not actually learn to perform matching well in the sense that they should otherwise achieve potentially even better after stereo-constrained perturbations are introduced. Second, stereo matching DNNs are typically trained under the simulation-to-real (Sim2Real) pipeline due to the data hungriness of DNNs. Thus, alleviating the impacts of the Sim2Real photometric gap in stereo matching DNNs becomes a pressing need. Towards joint adversarially robust and domain generalizable stereo matching, this paper proposes to learn DNN-contextualized binary-pattern-driven non-parametric cost-volumes. It leverages the perspective of learning the cost aggregation via DNNs, and presents a simple yet expressive design that is fully end-to-end trainable, without resorting to specific aggregation inductive biases. In experiments, the proposed method is tested in the SceneFlow dataset, the KITTI2015 dataset, and the Middlebury dataset. It significantly improves the adversarial robustness, while retaining accuracy performance comparable to state-of-the-art methods. It also shows a better Sim2Real generalizability. Our code and pretrained models are released at \href{https://github.com/kelkelcheng/AdversariallyRobustStereo}{this Github Repo}. Kelvin Cheng 0003, Tianfu Wu 0001, Christopher G. Healey |
NeurIPS | 2 |
| 2022 | Learning Layout and Style Reconfigurable GANs for Controllable Image SynthesisabstractWith the remarkable recent progress on learning deep generative models, it becomes increasingly interesting to develop models for controllable image synthesis from reconfigurable structured inputs. This paper focuses on a recently emerged task, layout-to-image, whose goal is to learn generative models for synthesizing photo-realistic images from a spatial layout (i.e., object bounding boxes configured in an image lattice) and its style codes (i.e., structural and appearance variations encoded by latent vectors). This paper first proposes an intuitive paradigm for the task, layout-to-mask-to-image, which learns to unfold object masks in a weakly-supervised way based on an input layout and object style codes. The layout-to-mask component deeply interacts with layers in the generator network to bridge the gap between an input layout and synthesized images. Then, this paper presents a method built on Generative Adversarial Networks (GANs) for the proposed layout-to-mask-to-image synthesis with layout and style control at both image and object levels. The controllability is realized by a proposed novel Instance-Sensitive and Layout-Aware Normalization (ISLA-Norm) scheme. A layout semi-supervised version of the proposed method is further developed without sacrificing performance. In experiments, the proposed method is tested in the COCO-Stuff dataset and the Visual Genome dataset with state-of-the-art performance obtained. Wei Sun 0033, Tianfu Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | PlaneTR: Structure-Guided Transformers for 3D Plane RecoveryabstractThis paper presents a neural network built upon Transformers, namely PlaneTR, to simultaneously detect and reconstruct planes from a single image. Different from previous methods, PlaneTR jointly leverages the context information and the geometric structures in a sequence-to-sequence way to holistically detect plane instances in one forward pass. Specifically, we represent the geometric structures as line segments and conduct the network with three main components: (i) context and line segments encoders, (ii) a structure-guided plane decoder, (iii) a pixel-wise plane embedding decoder. Given an image and its detected line segments, PlaneTR generates the context and line segment sequences via two specially designed encoders and then feeds them into a Transformers-based decoder to directly predict a sequence of plane instances by simultaneously considering the context and global structure cues. Finally, the pixel-wise embeddings are computed to assign each pixel to one predicted plane instance which is nearest to it in embedding space. Comprehensive experiments demonstrate that PlaneTR achieves state-of-the-art performance on the ScanNet and NYUv2 datasets. Bin Tan 0002, Nan Xue 0001, Song Bai 0001, Tianfu Wu 0001, Gui-Song Xia |
ICCV | 4 |
| 2021 | Learning Regional Attraction for Line Segment DetectionabstractThis paper presents regional attraction of line segment maps, and hereby poses the problem of line segment detection (LSD) as a problem of region coloring. Given a line segment map, the proposed regional attraction first establishes the relationship between line segments and regions in the image lattice. Based on this, the line segment map is equivalently transformed to an attraction field map (AFM), which can be remapped to a set of line segments without loss of information. Accordingly, we develop an end-to-end framework to learn attraction field maps for raw input images, followed by a squeeze module to detect line segments. Apart from existing works, the proposed detector properly handles the local ambiguity and does not rely on the accurate identification of edge pixels. Comprehensive experiments on the Wireframe dataset and the YorkUrban dataset demonstrate the superiority of our method. In particular, we achieve an F-measure of 0.831 on the Wireframe dataset, advancing the state-of-the-art performance by 10.3 percent. Nan Xue 0001, Song Bai 0001, Fudong Wang 0001, Gui-Song Xia, Tianfu Wu 0001, Liangpei Zhang 0001, Philip Torr 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Event driven sensor fusion
Siddharth Roheda, Hamid Krim, Zhi-Quan Luo, Tianfu Wu 0001 |
Signal Process. | 4 |
| 2021 | A Bottom-Up and Top-Down Integration Framework for Online Object TrackingabstractRobust online object tracking entails integrating short-term memory based trackers and long-term memory based trackers in an elegant framework to handle structural and appearance variations of unknown objects in an online manner. The integration and synergy between short-term and long-term memory based trackers have yet studied well in the literature, especially in pre-training free settings. To address this issue, this paper presents a bottom-up and top-down integration framework. The bottom-up component realizes a data-driven approach for particle generation. It exploits a short-term memory based tracker to generate bounding box proposals in a new frame. In the top-down component, this paper presents a graph regularized sparse coding scheme as the long-term memory based tracker. The over-complete bases for sparse coding are composed of part-based representations learned from earlier tracking results and new observations to form a space with rich temporal context information. A particle graph is computed whose nodes are the bottom-up discriminative particles and edges are formed on-the-fly in terms of appearance and spatial-temporal similarities between particles. The particle graph induces a regularization term in optimizing the sparse coding coefficients for bottom-up particles. In experiments, the proposed method is tested on the widely used OTB-100 benchmark and the VOT2016 benchmark with better performance obtained than baselines including deep learning based trackers. In addition, the outputs from the top-down sparse coding are potentially useful for downstream tasks such as action recognition, multiple-object tracking, and object re-identification. Meihui Li, Lingbing Peng, Tianfu Wu 0001, Zhenming Peng |
IEEE Trans. Multim. | 3 |
| 2020 | Inducing Hierarchical Compositional Model by Sparsifying Generator NetworkabstractThis paper proposes to learn hierarchical compositional AND-OR model for interpretable image synthesis by sparsifying the generator network. The proposed method adopts the scene-objects-parts-subparts-primitives hierarchy in image representation. A scene has different types (i.e., OR) each of which consists of a number of objects (i.e., AND). This can be recursively formulated across the scene-objects-parts-subparts hierarchy and is terminated at the primitive level (e.g., wavelets-like basis). To realize this AND-OR hierarchy in image synthesis, we learn a generator network that consists of the following two components: (i) Each layer of the hierarchy is represented by an over-complete set of convolutional basis functions. Off-the-shelf convolutional neural architectures are exploited to implement the hierarchy. (ii) Sparsity-inducing constraints are introduced in end-to-end training, which induces a sparsely activated and sparsely connected AND-OR model from the initially densely connected generator network. A straightforward sparsity-inducing constraint is utilized, that is to only allow the top-k basis functions to be activated at each layer (where k is a hyper-parameter). The learned basis functions are also capable of image reconstruction to explain the input images. In experiments, the proposed method is tested on four benchmark datasets. The results show that meaningful and interpretable hierarchical representations are learned with better qualities of image synthesis and reconstruction obtained than baselines. Xianglei Xing, Tianfu Wu 0001, Song-Chun Zhu, Ying Nian Wu |
CVPR | 2 |
| 2020 | Holistically-Attracted Wireframe ParsingabstractThis paper presents a fast and parsimonious parsing method to accurately and robustly detect a vectorized wireframe in an input image with a single forward pass. The proposed method is end-to-end trainable, consisting of three components: (i) line segment and junction proposal generation, (ii) line segment and junction matching, and (iii) line segment and junction verification. For computing line segment proposals, a novel exact dual representation is proposed which exploits a parsimonious geometric reparameterization for line segments and forms a holistic 4-dimensional attraction field map for an input image. Junctions can be treated as the “basins” in the attraction field. The proposed method is thus called Holistically-Attracted Wireframe Parser (HAWP). In experiments, the proposed method is tested on two benchmarks, the Wireframe dataset [14] and the YorkUrban dataset [8]. On both benchmarks, it obtains state-of-the-art performance in terms of accuracy and efficiency. For example, on the Wireframe dataset, compared to the previous state-of-the-art method L-CNN [36], it improves the challenging mean structural average precision (msAP) by a large margin (2.8% absolute improvements), and achieves 29.5 FPS on a single GPU (89% relative improvement). A systematic ablation study is performed to further justify the proposed method. Nan Xue 0001, Tianfu Wu 0001, Song Bai 0001, Fudong Wang 0001, Gui-Song Xia, Liangpei Zhang 0001, Philip Torr 0001 |
CVPR | 2 |
| 2020 | Attentive Normalization
Xilai Li, Wei Sun 0033, Tianfu Wu 0001 |
ECCV (17) | 3 |
| 2020 | Local Clustering with Mean Teacher for Semi-supervised learningabstractThe Mean Teacher (MT) model of Tarvainen and Valpola has shown good performance on several semi-supervised benchmark datasets. MT maintains a teacher model's weights as the exponential moving average of a student model's weights and minimizes the divergence between their probability predictions under diverse perturbations of the inputs. However, MT is known to suffer from confirmation bias, that is, reinforcing incorrect teacher model predictions. In this work, we propose a simple yet effective method called Local Clustering (LC) to mitigate the effect of confirmation bias. In MT, each data point is considered independent of other points during training; however, data points are likely to be close to each other in feature space if they share similar features. Motivated by this, we cluster data points locally by minimizing the pairwise distance between neighboring data points in feature space. Combined with a standard classification cross-entropy objective on labeled data points, the misclassified unlabeled data points are pulled towards high-density regions of their correct class with the help of their neighbors, thus improving model performance. We demonstrate on semi-supervised benchmark datasets SVHN and CIFAR-10 that adding our LC loss to MT yields significant improvements compared to MT and performance comparable to the state of the art in semi-supervised learning11The code is available at: https://github.com/jay1204/local_clustering_with_mt_for_ssl. Zexi Chen, Benjamin Dutton, Bharathkumar Ramachandra, Tianfu Wu 0001, Ranga Raju Vatsavai |
ICPR | 4 |
| 2019 | AOGNets: Compositional Grammatical Architectures for Deep LearningabstractNeural architectures are the foundation for improving performance of deep neural networks (DNNs). This paper presents deep compositional grammatical architectures which harness the best of two worlds: grammar models and DNNs. The proposed architectures integrate compositionality and reconfigurability of the former and the capability of learning rich features of the latter in a principled way. We utilize AND-OR Grammar (AOG) as network generator in this paper and call the resulting networks AOGNets. An AOGNet consists of a number of stages each of which is composed of a number of AOG building blocks. An AOG building block splits its input feature map into N groups along feature channels and then treat it as a sentence of N words. It then jointly realizes a phrase structure grammar and a dependency grammar in bottom-up parsing the “sentence” for better feature exploration and reuse. It provides a unified framework for the best practices developed in state-of-the-art DNNs. In experiments, AOGNet is tested in the ImageNet-1K classification benchmark and the MS-COCO object detection and segmentation benchmark. In ImageNet-1K, AOGNet obtains better performance than ResNet and most of its variants, ResNeXt and its attention based variants such as SENet, DenseNet and DualPathNet. AOGNet also obtains the best model interpretability score using network dissection. AOGNet further shows better potential in adversarial defense. In MS-COCO, AOGNet obtains better performance than the ResNet and ResNeXt backbones in Mask R-CNN. Xilai Li, Tianfu Wu 0001 |
CVPR | 3 |
| 2019 | Learning Attraction Field Representation for Robust Line Segment DetectionabstractThis paper presents a region-partition based attraction field dual representation for line segment maps, and thus poses the problem of line segment detection (LSD) as the region coloring problem. The latter is then addressed by learning deep convolutional neural networks (ConvNets) for accuracy, robustness and efficiency. For a 2D line segment map, our dual representation consists of three components: (i) A region-partition map in which every pixel is assigned to one and only one line segment; (ii) An attraction field map in which every pixel in a partition region is encoded by its 2D projection vector w.r.t. the associated line segment; and (iii) A squeeze module which squashes the attraction field to a line segment map that almost perfectly recovers the input one. By leveraging the duality, we learn ConvNets to compute the attraction field maps for raw in-put images, followed by the squeeze module for LSD, in an end-to-end manner. Our method rigorously addresses several challenges in LSD such as local ambiguity and class imbalance. Our method also harnesses the best practices developed in ConvNets based semantic segmentation methods such as the encoder-decoder architecture and the a-trous convolution. In experiments, our method is tested on the WireFrame dataset and the YorkUrban dataset with state-of-the-art performance obtained. Especially, we advance the performance by 4.5 percents on the WireFramedataset. Our method is also fast with 6.6∼10.4 FPS, outperforming most of existing line segment detectors. Nan Xue 0001, Song Bai 0001, Fudong Wang 0001, Gui-Song Xia, Tianfu Wu 0001, Liangpei Zhang 0001 |
CVPR | 5 |
| 2019 | Image Synthesis From Reconfigurable Layout and StyleabstractDespite remarkable recent progress on both unconditional and conditional image synthesis, it remains a long- standing problem to learn generative models that are capable of synthesizing realistic and sharp images from re- configurable spatial layout (i.e., bounding boxes + class labels in an image lattice) and style (i.e., structural and appearance variations encoded by latent vectors), especially at high resolution. By reconfigurable, it means that a model can preserve the intrinsic one-to-many mapping from a given layout to multiple plausible images with different styles, and is adaptive with respect to perturbations of a layout and style latent code. In this paper, we present a layout- and style-based architecture for generative adversarial networks (termed LostGANs) that can be trained end-to-end to generate images from reconfigurable layout and style. Inspired by the vanilla StyleGAN, the proposed LostGAN consists of two new components: (i) learning fine-grained mask maps in a weakly-supervised manner to bridge the gap between layouts and images, and (ii) learning object instance-specific layout-aware feature normalization (ISLA-Norm) in the generator to realize multi-object style generation. In experiments, the proposed method is tested on the COCO-Stuff dataset and the Visual Genome dataset with state-of-the-art performance obtained. The code and pretrained models are available at https://github.com/iVMCL/LostGANs. Wei Sun 0033, Tianfu Wu 0001 |
ICCV | 2 |
| 2019 | Towards Interpretable Object Detection by Unfolding Latent StructuresabstractThis paper first proposes a method of formulating model interpretability in visual understanding tasks based on the idea of unfolding latent structures. It then presents a case study in object detection using popular two-stage region-based convolutional network (i.e., R-CNN) detection systems. The proposed method focuses on weakly-supervised extractive rationale generation, that is learning to unfold latent discriminative part configurations of object instances automatically and simultaneously in detection without using any supervision for part configurations. It utilizes a top-down hierarchical and compositional grammar model embedded in a directed acyclic AND-OR Graph (AOG) to explore and unfold the space of latent part configurations of regions of interest (RoIs). It presents an AOGParsing operator that seamlessly integrates with the RoIPooling/RoIAlign operator widely used in R-CNN and is trained end-to-end. In object detection, a bounding box is interpreted by the best parse tree derived from the AOG on-the-fly, which is treated as the qualitatively extractive rationale generated for interpreting detection. In experiments, Faster R-CNN is used to test the proposed method on the PASCAL VOC 2007 and the COCO 2017 object detection datasets. The experimental results show that the proposed method can compute promising latent structures without hurting the performance. The code and pretrained models are available at https://github.com/iVMCL/iRCNN. Tianfu Wu 0001 |
ICCV | 1 |
| 2019 | Learn to Grow: A Continual Structure Learning Framework for Overcoming Catastrophic ForgettingabstractAddressing catastrophic forgetting is one of the key challenges in continual learning where machine learning systems are trained with sequential or streaming tasks. Despite recent remarkable progress in state-of-the-art deep learning, deep neural networks (DNNs) are still plagued with the catastrophic forgetting problem. This paper presents a conceptually simple yet general and effective framework for handling catastrophic forgetting in continual learning with DNNs. The proposed method consists of two components: a neural structure optimization component and a parameter learning and/or fine-tuning component. By separating the explicit neural structure learning and the parameter estimation, not only is the proposed method capable of evolving neural structures in an intuitively meaningful way, but also shows strong capabilities of alleviating catastrophic forgetting in experiments. Furthermore, the proposed method outperforms all other baselines on the permuted MNIST dataset, the split CIFAR100 dataset and the Visual Domain Decathlon dataset in continual learning setting. Xilai Li, Yingbo Zhou 0002, Tianfu Wu 0001, Richard Socher, Caiming Xiong |
ICML | 3 |
| 2019 | Jointly social grouping and identification in visual dynamics with causality-induced hierarchical Bayesian model
Zhao Xie, Tianfu Wu 0001, Xingming Yang, Kewei Wu |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Scene-Centric Joint Parsing of Cross-View VideosabstractCross-view video understanding is an important yet under-explored area in computer vision. In this paper, we introduce a joint parsing framework that integrates view-centric proposals into scene-centric parse graphs that represent a coherent scene-centric understanding of cross-view scenes. Our key observations are that overlapping fields of views embed rich appearance and geometry correlations and that knowledge fragments corresponding to individual vision tasks are governed by consistency constraints available in commonsense knowledge. The proposed joint parsing framework represents such correlations and constraints explicitly and generates semantic scene-centric parse graphs. Quantitative experiments show that scene-centric predictions in the parse graph outperform view-centric predictions. Hang Qi 0001, Yuanlu Xu, Tianfu Wu 0001, Song-Chun Zhu |
AAAI | 4 |
| 2018 | Neural Abstract Style Transfer for Chinese Traditional Painting
Bo Li 0031, Caiming Xiong, Tianfu Wu 0001, Yu Zhou 0016, Rufeng Chu |
ACCV (2) | 3 |
| 2017 | Online Object Tracking, Learning and Parsing with And-Or GraphsabstractThis paper presents a method, called AOGTracker, for simultaneously tracking, learning and parsing (TLP) of unknown objects in video sequences with a hierarchical and compositional And-Or graph (AOG) representation. The TLP method is formulated in the Bayesian framework with a spatial and a temporal dynamic programming (DP) algorithms inferring object bounding boxes on-the-fly. During online learning, the AOG is discriminatively learned using latent SVM [1] to account for appearance (e.g., lighting and partial occlusion) and structural (e.g., different poses and viewpoints) variations of a tracked object, as well as distractors (e.g., similar objects) in background. Three key issues in online inference and learning are addressed: (i) maintaining purity of positive and negative examples collected online, (ii) controling model complexity in latent structure learning, and (iii) identifying critical moments to re-learn the structure of AOG based on its intrackability. The intrackability measures uncertainty of an AOG based on its score maps in a frame. In experiments, our AOGTracker is tested on two popular tracking benchmarks with the same parameter setting: the TB-100/50/CVPR2013 benchmarks , [3] , and the VOT benchmarks [4] -VOT 2013, 2014, 2015 and TIR2015 (thermal imagery tracking). In the former, our AOGTracker outperforms state-of-the-art tracking algorithms including two trackers based on deep convolutional network [5] , [6] . In the latter, our AOGTracker outperforms all other trackers in VOT2013 and is comparable to the state-of-the-art methods in VOT2014, 2015 and TIR2015. Tianfu Wu 0001, Yang Lu 0006, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Face Video Retrieval via Deep Learning of Binary Hash RepresentationsabstractRetrieving faces from large mess of videos is an attractive research topic with wide range of applications. Its challenging problems are large intra-class variations, and tremendous time and space complexity. In this paper, we develop a new deep convolutional neural network (deep CNN) to learn discriminative and compact binary representations of faces for face video retrieval. The network integrates feature extraction and hash learning into a unified optimization framework for the optimal compatibility of feature extractor and hash functions. In order to better initialize the network, the low-rank discriminative binary hashing is proposed to pre-learn hash functions during the training procedure. Our method achieves excellent performances on two challenging TV-Series datasets. Zhen Dong 0002, Su Jia, Tianfu Wu 0001, Mingtao Pei |
AAAI | 3 |
| 2016 | Recognizing Car Fluents from VideoabstractPhysical fluents, a term originally used by Newton, refers to time-varying object states in dynamic scenes. In this paper, we are interested in inferring the fluents of vehicles from video. For example, a door (hood, trunk) is open or closed through various actions, light is blinking to turn. Recognizing these fluents has broad applications, yet have received scant attention in the computer vision literature. Car fluent recognition entails a unified framework for car detection, car part localization and part status recognition, which is made difficult by large structural and appearance variations, low resolutions and occlusions. This paper learns a spatial-temporal And-Or hierarchical model to represent car fluents. The learning of this model is formulated under the latent structural SVM framework. Since there are no publicly related dataset, we collect and annotate a car fluent dataset consisting of car videos with diverse fluents. In experiments, the proposed method outperforms several highly related baseline methods in terms of car fluent recognition and car part localization. Bo Li 0031, Tianfu Wu 0001, Caiming Xiong, Song-Chun Zhu |
CVPR | 2 |
| 2016 | Face Detection with End-to-End Integration of a ConvNet and a 3D Model
Yunzhu Li, Benyuan Sun, Tianfu Wu 0001, Yizhou Wang 0001 |
ECCV (3) | 3 |
| 2016 | Learning And-Or Model to Represent Context and Occlusion for Car Detection and Viewpoint EstimationabstractThis paper presents a method for learning an And-Or model to represent context and occlusion for car detection and viewpoint estimation. The learned And-Or model represents car-to-car context and occlusion configurations at three levels: (i) spatially-aligned cars, (ii) single car under different occlusion configurations, and (iii) a small number of parts. The And-Or model embeds a grammar for representing large structural and appearance variations in a reconfigurable hierarchy. The learning process consists of two stages in a weakly supervised way (i.e., only bounding boxes of single cars are annotated). Firstly, the structure of the And-Or model is learned with three components: (a) mining multi-car contextual patterns based on layouts of annotated single car bounding boxes, (b) mining occlusion configurations between single cars, and (c) learning different combinations of part visibility based on CAD simulations. The And-Or model is organized in a directed and acyclic graph which can be inferred by Dynamic Programming. Secondly, the model parameters (for appearance, deformation and bias) are jointly trained using Weak-Label Structural SVM. In experiments, we test our model on four car detection datasets - the KITTI dataset [1], the PASCAL VOC2007 car dataset [2], and two self-collected car datasets, namely the Street-Parking car dataset and the Parking-Lot car dataset, and three datasets for car viewpoint estimation - the PASCAL VOC2006 car dataset [2], the 3D car dataset [3], and the PASCAL3D+ car dataset [4]. Compared with state-of-the-art variants of deformable part-based models and other methods, our model achieves significant improvement consistently on the four detection datasets, and comparable performance on car viewpoint estimation. Tianfu Wu 0001, Bo Li 0031, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | A Reconfigurable Tangram Model for Scene Representation and CategorizationabstractThis paper presents a hierarchical and compositional scene layout (i.e., spatial configuration) representation and a method of learning reconfigurable model for scene categorization. Three types of shape primitives (i.e., triangle, parallelogram, and trapezoid), called tans, are used to tile scene image lattice in a hierarchical and compositional way, and a directed acyclic AND-OR graph (AOG) is proposed to organize the overcomplete dictionary of tan instances placed in image lattice, exploring a very large number of scene layouts. With certain off-the-shelf appearance features used for grounding terminal-nodes (i.e., tan instances) in the AOG, a scene layout is represented by the globally optimal parse tree learned via a dynamic programming algorithm from the AOG, which we call tangram model. Then, a scene category is represented by a mixture of tangram models discovered with an exemplar-based clustering method. On basis of the tangram model, we address scene categorization in two aspects: 1) building a tangram bank representation for linear classifiers, which utilizes a collection of tangram models learned from all categories and 2) building a tangram matching kernel for kernel-based classification, which accounts for all hidden spatial configurations in the AOG. In experiments, our methods are evaluated on three scene data sets for both the configuration-level and semantic-level scene categorization, and outperform the spatial pyramid model consistently. Tianfu Wu 0001, Song-Chun Zhu, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Learning Near-Optimal Cost-Sensitive Decision Policy for Object DetectionabstractMany popular object detectors, such as AdaBoost, SVM and deformable part-based models (DPM), compute additive scoring functions at a large number of windows in an image pyramid, thus computational efficiency is an important consideration in real time applications besides accuracy. In this paper, a decision policy refers to a sequence of two-sided thresholds to execute early reject and early accept based on the cumulative scores at each step. We formulate an empirical risk function as the weighted sum of the cost of computation and the loss of false alarm and missing detection. Then a policy is said to be cost-sensitive and optimal if it minimizes the risk function. While the risk function is complex due to high-order correlations among the two-sided thresholds, we find that its upper bound can be optimized by dynamic programming efficiently. We show that the upper bound is very tight empirically and thus the resulting policy is said to be near-optimal. In experiments, we show that the decision policy outperforms state-of-the-art cascade methods significantly, in several popular detection tasks and benchmarks, in terms of computational efficiency with similar accuracy of detection. Tianfu Wu 0001, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Online Object Tracking, Learning, and Parsing with And-Or GraphsabstractThis paper presents a framework for simultaneously tracking, learning and parsing objects with a hierarchical and compositional and-or graph (AOG) representation. The AOG is discriminatively learned online to account for the appearance (e.g., lighting and partial occlusion) and structural (e.g., different poses and viewpoints) variations of the object itself, as well as the distractors (e.g., similar objects) in the scene background. In tracking, the state of the object (i.e., bounding box) is inferred by parsing with the current AOG using a spatial-temporal dynamic programming (DP) algorithm. When the AOG grows big for handling objects with large variations in long-term tracking, we propose a bottom-up/top-down scheduling scheme for efficient inference, which performs focused inference with the most stable and discriminative small sub-AOG. During online learning, the AOG is re-learned iteratively with two steps: (i) Identifying the false positives and false negatives of the current AOG in a new frame by exploiting the spatial and temporal constraints observed in the trajectory, (ii) Updating the structure of the AOG, and re-estimating the parameters based on the augmented training dataset. In experiments, the proposed method outperforms state-of-the-art tracking algorithms on a recent public tracking benchmark with 50 testing videos and 30 publicly available trackers evaluated [34]. Yang Lu 0006, Tianfu Wu 0001, Song-Chun Zhu |
CVPR | 2 |
| 2014 | Integrating Context and Occlusion for Car Detection by Hierarchical And-Or Model
Bo Li 0031, Tianfu Wu 0001, Song-Chun Zhu |
ECCV (6) | 2 |
| 2014 | A global energy optimization framework for 2.1D sketch extraction from monocular images
Cheng-Chi Yu, Yong-Jin Liu 0001, Tianfu Wu 0001, Kai-Yun Li, Xiaolan Fu |
Graph. Model. | 3 |
| 2014 | Coupling-and-decoupling: A hierarchical model for occlusion-free object detection
Bo Li 0031, Tianfu Wu 0001, Wenze Hu, Mingtao Pei |
Pattern Recognit. | 3 |
| 2013 | Discriminatively Trained And-Or Tree Models for Object DetectionabstractThis paper presents a method of learning reconfigurable And-Or Tree (AOT) models discriminatively from weakly annotated data for object detection. To explore the appearance and geometry space of latent structures effectively, we first quantize the image lattice using an over complete set of shape primitives, and then organize them into a directed a cyclic And-Or Graph (AOG) by exploiting their compositional relations. We allow overlaps between child nodes when combining them into a parent node, which is equivalent to introducing an appearance Or-node implicitly for the overlapped portion. The learning of an AOT model consists of three components: (i) Unsupervised sub-category learning (i.e., branches of an object Or-node) with the latent structures in AOG being integrated out. (ii) Weakly supervised part configuration learning (i.e., seeking the globally optimal parse trees in AOG for each sub-category). To search the globally optimal parse tree in AOG efficiently, we propose a dynamic programming (DP) algorithm. (iii) Joint appearance and structural parameters training under latent structural SVM framework. In experiments, our method is tested on PASCAL VOC 2007 and 2010 detection benchmarks of 20 object classes and outperforms comparable state-of-the-art methods. Tianfu Wu 0001, Yunde Jia, Song-Chun Zhu |
CVPR | 2 |
| 2013 | Modeling Occlusion by Discriminative AND-OR StructuresabstractOcclusion presents a challenge for detecting objects in real world applications. To address this issue, this paper models object occlusion with an AND-OR structure which (i) represents occlusion at semantic part level, and (ii) captures the regularities of different occlusion configurations (i.e., the different combinations of object part visibilities). This paper focuses on car detection on street. Since annotating part occlusion on real images is time-consuming and error-prone, we propose to learn the the AND-OR structure automatically using synthetic images of CAD models placed at different relative positions. The model parameters are learned from real images under the latent structural SVM (LSSVM) framework. In inference, an efficient dynamic programming (DP) algorithm is utilized. In experiments, we test our method on both car detection and car view estimation. Experimental results show that (i) Our CAD simulation strategy is capable of generating occlusion patterns for real scenarios, (ii) The proposed AND-OR structure model is effective for modeling occlusions, which outperforms the deformable part-based model (DPM) DPM, voc5 in car detection on both our self-collected street parking dataset and the Pascal VOC 2007 car dataset pascal-voc-2007}, (iii) The learned model is on-par with the state-of-the-art methods on car view estimation tested on two public datasets. Bo Li 0031, Wenze Hu, Tianfu Wu 0001, Song-Chun Zhu |
ICCV | 3 |
| 2013 | Learning Near-Optimal Cost-Sensitive Decision Policy for Object DetectionabstractMany object detectors, such as AdaBoost, SVM and deformable part-based models (DPM), compute additive scoring functions at a large number of windows scanned over image pyramid, thus computational efficiency is an important consideration beside accuracy performance. In this paper, we present a framework of learning cost-sensitive decision policy which is a sequence of two-sided thresholds to execute early rejection or early acceptance based on the accumulative scores at each step. A decision policy is said to be optimal if it minimizes an empirical global risk function that sums over the loss of false negatives (FN) and false positives (FP), and the cost of computation. While the risk function is very complex due to high-order connections among the two-sided thresholds, we find its upper bound can be optimized by dynamic programming (DP) efficiently and thus say the learned policy is near-optimal. Given the loss of FN and FP and the cost in three numbers, our method can produce a policy on-the-fly for Adaboost, SVM and DPM. In experiments, we show that our decision policy outperforms state-of-the-art cascade methods significantly in terms of speed with similar accuracy performance. Tianfu Wu 0001, Song-Chun Zhu |
ICCV | 1 |
| 2012 | Coupling-and-Decoupling: A Hierarchical Model for Occlusion-Free Car Detection
Bo Li 0031, Tianfu Wu 0001, Wenze Hu, Mingtao Pei |
ACCV (1) | 2 |
| 2012 | Tracking Pedestrian with Multi-component Online Deformable Part-Based Model
Yi Xie 0006, Mingtao Pei, Tianfu Wu 0001 |
ACCV (3) | 4 |
| 2012 | Learning Global and Reconfigurable Part-Based Models for Object DetectionabstractThis paper presents a method of learning global and reconfigurable part-based models (RPM) for object detection. Recently, deformable part-based model (DPM) is widely used. A DPM consists of a root node and a collection of part nodes, which is learned under the latent SVM formulation by treating part nodes as hidden variables. Although the configuration of parts (i.e., the shapes, sizes and locations of parts) plays a major role in improving performance of object detection, it has not been addressed well in the literature. In this paper, we propose RPM to tackle it. A dictionary of part types is defined by enumerating rectangular shapes of different aspect ratios and sizes given the whole lattice (often at twice resolution of the root node), and each part type has a set of part instances when placed in the lattice. So, the configuration space of parts is quantized by the part types and part instances, and then organized into a hierarchical And-Or directed a cyclic graph (AOG). The AOG consists of three types of nodes: terminal nodes (i.e., part instances), And-nodes (representing decompositions of a part instance into two smaller ones) and Or-nodes (representing alternative ways of decompositions). The globally optimal configuration in the AOG is solved using dynamic programming (DP) where the classification error rates of terminal nodes and And-nodes are used as their figures of merit. In experiments, we test our method on the 20 object categories in the PASCAL VOC2007 dataset and obtain comparable performance with state-of-the-art methods. Tianfu Wu 0001, Yi Xie 0006, Yunde Jia |
ICME | 2 |
| 2012 | Learning reconfigurable scene representation by tangram modelabstractThis paper proposes a method to learn reconfigurable and sparse scene representation in the joint space of spatial configuration and appearance in a principled way. We call it the tangram model, which has three properties: (1) Unlike fixed structure of the spatial pyramid widely used in the literature, we propose a compositional shape dictionary organized in an And-Or directed acyclic graph (AOG) to quantize the space of spatial configurations. (2) The shape primitives (called tans) in the dictionary can be described by using any “off-the-shelf” appearance features according to different tasks. (3) A dynamic programming (DP) algorithm is utilized to learn the globally optimal parse tree in the joint space of spatial configuration and appearance. We demonstrate the tangram model in both a generative learning formulation and a discriminative matching kernel. In experiments, we show that the tangram model is capable of capturing meaningful spatial configurations as well as appearance for various scene categories, and achieves state-of-the-art classification performance on the LSP 15-class scene dataset and the MIT 67-class indoor scene dataset. Tianfu Wu 0001, Song-Chun Zhu, Xiaokang Yang 0001, Wenjun Zhang 0001 |
WACV | 2 |
| 2011 | A Numerical Study of the Bottom-Up and Top-Down Inference Processes in And-Or Graphs
Tianfu Wu 0001, Song-Chun Zhu |
Int. J. Comput. Vis. | 1 |
| 2010 | Discovering scene categories by information projection and cluster samplingabstractThis paper presents a method for unsupervised scene categorization. Our method aims at two objectives: (1) automatic feature selection for different scene categories. We represent images in a heterogeneous feature space to account for the large variabilities of different scene categories. Then, we use the information projection strategy to pursue features which are both informative and discriminative, and simultaneously learn a generative model for each category. (2) automatic cluster number selection for the whole image set to be categorized. By treating each image as a vertex in a graph, we formulate unsupervised scene categorization as a graph partition problem under the Bayesian framework. Then, we use a cluster sampling strategy to do the partition (i.e. categorization) in which the cluster number is selected automatically for the globally optimal clustering in terms of maximizing a Bayesian posterior probability. In experiments, we test two datasets, LHI 8 scene categories and MIT 8 scene categories, and obtain state-of-the-art results. Dengxin Dai, Tianfu Wu 0001, Song-Chun Zhu |
CVPR | 2 |
| 2010 | Three-layer Spatial Sparse Coding for Image ClassificationabstractIn this paper, we propose a three-layer spatial sparse coding (TSSC) for image classification, aiming at three objectives: naturally recognizing image categories without learning phase, naturally involving spatial configurations of images, and naturally counteracting the intra-class variances. The method begins by representing the test images in a spatial pyramid as the to-be-recovered signals, and taking all sampled image patches at multiple scales from the labeled images as the bases. Then, three sets of coefficients are involved into the cardinal sparse coding to get the TSSC, one to penalize spatial inconsistencies of the pyramid cells and the corresponding selected bases, one to guarantee the sparsity of selected images, and the other to guarantee the sparsity of selected categories. Finally, the test images are classified according to a simple image-to-category similarity defined on the coding coefficients. In experiments, we test our method on two publicly available datasets and achieve significantly more accurate results than the conventional sparse coding with only a modest increase in computational complexity. Dengxin Dai, Wen Yang 0001, Tianfu Wu 0001 |
ICPR | 3 |
| 2009 | Evaluating information contributions of bottom-up and top-down processesabstractThis paper presents a method to quantitatively evaluate information contributions of individual bottom-up and top-down computing processes in object recognition. Our objective is to start a discovery on how to schedule bottom-up and top-down processes. (1) We identify two bottom-up processes and one top-down process in hierarchical models, termed α, β and γ channels respectively ; (2) We formulate the three channels under an unified Bayesian framework; (3) We use a blocking control strategy to isolate the three channels to separately train them and individually measure their information contributions in typical recognition tasks; (4) Based on the evaluated results, we integrate the three channels to detect objects with performance improvements obtained. Our experiments are performed in both low-middle level tasks, such as detecting edges/bars and junctions, and high level tasks, such as detecting human faces and cars, together with a group of human study designed to compare computer and human perception. Tianfu Wu 0001, Song-Chun Zhu |
ICCV | 2 |
| 2009 | A stochastic graph grammar for compositional object representation and recognition
Tianfu Wu 0001, Jake Porway, Zijian Xu 0001 |
Pattern Recognit. | 2 |
| 2008 | Design sparse features for age estimation using hierarchical face modelabstractA key point in automatic age estimation is to design feature set essential to age perception. To achieve this goal, this paper builds up a hierarchical graphical face model for faces appearing at low, middle and high resolution respectively. Along the hierarchy, a face image is decomposed into detailed parts from coarse to fine. Then four types of features are extracted from this graph representation guided by the priors of aging process embedded in the graphical model: topology, geometry, photometry and configuration. On age estimation, this paper follows the popular regression formulation for mapping feature vectors to its age label. The effectiveness of the presented feature set is justified by testing results on two datasets using different kinds of regression methods. The experimental results in this paper show that designing feature set for age estimation under the guidance of hierarchical face model is a promising method and a flexible framework as well. Jin-Li Suo, Tianfu Wu 0001, Song-Chun Zhu, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
FG | 2 |
| 2007 | Compositional Boosting for Computing Hierarchical Image StructuresabstractIn this paper, we present a compositional boosting algorithm for detecting and recognizing 17 common image structures in low-middle level vision tasks. These structures, called "graphlets", are the most frequently occurring primitives, junctions and composite junctions in natural images, and are arranged in a 3-layer And-Or graph representation. In this hierarchic model, larger graphlets are decomposed (in And-nodes) into smaller graphlets in multiple alternative ways (at Or-nodes), and parts are shared and re-used between graphlets. Then we present a compositional boosting algorithm for computing the 17 graphlets categories collectively in the Bayesian framework. The algorithm runs recursively for each node A in the And-Or graph and iterates between two steps -bottom-up proposal and top-down validation. The bottom-up step includes two types of boosting methods, (i) Detecting instances of A (often in low resolutions) using Adaboosting method through a sequence of tests (weak classifiers) image feature, (ii) Proposing instances of A (often in high resolution) by binding existing children nodes of A through a sequence of compatibility tests on their attributes (e.g angles, relative size etc). The Adaboosting and binding methods generate a number of candidates for node A which are verified by a top-down process in a way similar to Data-Driven Markov Chain Monte Carlo [18]. Both the Adaboosting and binding methods are trained off-line for each graphlet category, and the compositional nature of the model means the algorithm is recursive and can be learned from a small training set. We apply this algorithm to a wide range of indoor and outdoor images with satisfactory results. Tianfu Wu 0001, Gui-Song Xia, Song-Chun Zhu |
CVPR | 1 |