EDBT 2026 Demo / reviewers in the wild / expert
Yuqiang Fang
dblp:124/5793
· DBLP profile ↗
24ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | COXNet: Cross-Layer Fusion With Adaptive Alignment and Scale Integration for RGBT Tiny Object DetectionabstractDetecting tiny objects in multimodal Red-Green-Blue-Thermal (RGBT) imagery is a critical challenge in computer vision, particularly in surveillance, search and rescue, and autonomous navigation. Drone-based scenarios exacerbate these challenges due to spatial misalignment, low-light conditions, occlusion, and cluttered backgrounds. Current methods struggle to leverage the complementary information between visible and thermal modalities effectively. We propose COXNet, a novel framework for RGBT tiny object detection, addressing these issues through three core innovations: i) the Cross-Layer Fusion Module, fusing high-level visible and low-level thermal features for enhanced semantic and spatial accuracy; ii) the Dynamic Alignment and Scale Refinement module, correcting cross-modal spatial misalignments and preserving multi-scale features; and iii) an optimized label assignment strategy using the GeoShape Similarity Measure for better localization. COXNet achieves a 3.32% mAP50improvement on the RGBTDronePerson dataset over state-of-the-art methods, demonstrating its effectiveness for robust detection in complex environments. Peiran Peng, Tingfa Xu, Liqiang Song, Mengqi Zhu, Yuqiang Fang, Jianan Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | VETD220: A Visible-Event Benchmark and Baseline for Tracking DronesabstractAnti-drone tracking faces challenges from targets’ fast movement, complex backgrounds, and extreme lighting, which often cause overexposure and motion blur in the visible modality, leading to insufficient tracking robustness. Owing to the High Dynamic Range (HDR) and high temporal resolution of event cameras, they can provide sufficient auxiliary information to the visible modality, and their complementary features offer a new solution for anti-drone tracking. In this paper, we first introduce event modality to support visible data. Specifically, we propose the multi-modal anti-drone tracking dataset, providing both visible sequences and event streams, namely VETD220. It contains 220 videos and 68,379 visible-event sequence pairs, covering multiple scenarios and target sizes with fine-grained annotations.The proposed VETD220 provides sufficient data for the anti-drone tracking community, which facilitates model training and fair comparison. Meanwhile, we propose EMFETrack, an event-guided multi-modal feature enhancement algorithm for drone tracking. It extracts spatiotemporal features of events via sparse convolution and designs two key modules: Event-Guided Dynamic Feature Enhancement (EG-DFE) and Motion-aware Position Encoding (MAPE). These modules achieve effective modality interactions and improve feature discriminability in complex scenes like background clustering and target blur. We conduct comparative experiments with 17 mainstream trackers on VETD220, where our method achieves a Success Rate (SR) of 72.7% and a Precision Rate (PR) of 96.8%. It also performs competitively or superiorly on FE240hz and COESOT datasets, verifying its effectiveness. VETD220 is available at https://github.com/QiuuuJY/VETD220. Jiayu Qiu, Jianan Li 0001, Gege Sun, Yuqiang Fang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | A Review of Neural Radiation Field Based 3D Reconstruction Methods for Spatial Targets
Wanyun Li, Yuqiang Fang, Gege Sun |
ICIG (3) | 2 |
| 2024 | MappingFormer: Learning Cross-modal Feature Mapping for Visible-to-infrared Image TranslationabstractDue to the limitations of infrared image acquisition conditions, many essential tasks currently rely on visible images as the main source of training data. However, single-modal data makes it difficult for downstream networks to show optimal performance. Therefore, converting the more easily obtainable visible images into infrared images emerges as an effective remedy to alleviate the critical shortage of infrared data. Yet current methods typically focus solely on transferring visible images to infrared style, while overlooking the crucial infrared thermal feature during cross-modal translation. To elevate the authenticity of cross-model translation at the feature level, this paper introduces a translation network based on frequency feature mapping and dual patches contrast, MappingFormer, which can achieve cross-modal image generation from visible to infrared. Specifically, the generator incorporates two branches: low-frequency feature mapping (LFM) and high-frequency feature refinement (HFR), both have embedded the Swin Transformer blocks. The LFM branch captures the fuzzy structural from visible images, while the HFR focuses on mapping edge and texture features. The extracted dual-branch frequency features undergo refinement and fusion through cross-attention mechanisms. Additionally, a dual contrast learning mechanism based on feature patch (DFPC) is designed to infer effective mappings between unaligned cross-modal data. Numerous experimental results prove the effectiveness of this method in cross-modal feature mapping and image generation from visible to infrared. This method holds significant potential for downstream tasks where infrared data is limited. Haining Wang 0007, Na Li 0014, Huijie Zhao, Yuqiang Fang |
ACM Multimedia | 6 |
| 2024 | Modality adaptation via feature difference learning for depth human parsing
Shaofei Huang 0001, Tianrui Hui, Fengguang Peng, Yuqiang Fang, Bin Ma 0028, Xiaoming Wei, Jizhong Han |
Comput. Vis. Image Underst. | 5 |
| 2024 | Multi-Step Temporal Modeling for UAV TrackingabstractIn the realm of unmanned aerial vehicle (UAV) tracking, Siamese-based approaches have gained traction due to their optimal balance between efficiency and precision. However, UAV scenarios often present challenges such as insufficient sampling resolution, fast motion and small objects with limited feature information. As a result, temporal context in UAV tracking tasks plays a pivotal role in target location, overshadowing the target’s precise features. In this paper, we introduce MT-Track, a streamlined and efficient multi-step temporal modeling framework designed to harness the temporal context from historical frames for enhanced UAV tracking. This temporal integration occurs in two steps: correlation map generation and correlation map refinement. Specifically, we unveil a unique temporal correlation module that dynamically assesses the interplay between the template and search region features. This module leverages temporal information to refresh the template feature, yielding a more precise correlation map. Subsequently, we propose a mutual transformer module to refine the correlation maps of historical and current frames by modeling the temporal knowledge in the tracking sequence. This method significantly trims computational demands compared to the raw transformer. The compact yet potent nature of our tracking framework ensures commendable tracking outcomes, particularly in extended tracking scenarios. Comprehensive tests across four renowned UAV benchmarks substantiate the superior efficacy of our approach, delivering real-time performance at 84.7 FPS on a single GPU. Real-world test on the NVIDIA AGX hardware platform achieves a speed exceeding 30 FPS, validating the practicality of our method. Xiaoying Yuan, Tingfa Xu, Xincong Liu, Ying Wang 0064, Haolin Qin, Yuqiang Fang, Jianan Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | A Lightweight Fusion Strategy With Enhanced Interlayer Feature Correlation for Small Object DetectionabstractDetecting small objects in drone imagery is challenging due to low resolution and background blending, leading to limited feature information. Multiscale feature fusion can enhance detection by capturing information at different scales, but traditional strategies fall short. Simple concatenation or addition operations do not fully utilize multiscale fusion advantages, resulting in insufficient correlation between features. This inadequacy hinders the detection of small objects, especially in complex backgrounds and densely populated areas. To address this issue and efficiently utilize the limited computational resources, we propose a lightweight fusion strategy based on enhanced interlayer feature correlation (EFC) to replace the traditional feature fusion strategy in feature pyramid network (FPN). The semantic expressions of different layers in the feature pyramid are inconsistent. In EFC, the grouped feature focus unit (GFF) enhances the feature correlation of each layer by focusing on the contextual information of different features. The multilevel feature reconstruction module (MFR) effectively reconstructs and transforms the strength and weakness information of each layer in the pyramid to reduce redundant feature fusion and retain more information about small targets in deep networks. It is noteworthy that the proposed method is plug-and-play and can be widely applied to various base networks. Extensive experiments and comprehensive evaluations on VisDrone, unmanned aerial vehicle benchmark object detection and tracking (UAVDT), and microsoft common objects in context (COCO) demonstrate the effectiveness. Using generalized focal loss (GFL) as the baseline on the VisDrone dataset with a large number of small targets, the proposed method improves the detection mean average precision (mAP) by 1.7%, surpassing many lightweight state-of-the-art methods and significantly reducing the Params and GFLOPs at the neck end. The code will be available athttps://github.com/nuliweixiao/EFC.git. Tingfa Xu, Yuqiang Fang, Jianan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReIDabstractNeural Architecture Search (NAS) has been increasingly appealing to the society of object Re-Identification (ReID), for that task-specific architectures significantly improve the retrieval performance. Previous works explore new optimizing targets and search spaces for NAS ReID, yet they neglect the difference of training schemes between image classification and ReID. In this work, we propose a novel Twins Contrastive Mechanism (TCM) to provide more appropriate supervision for ReID architecture search. TCM reduces the category overlaps between the training and validation data, and assists NAS in simulating real-world ReID training schemes. We then design a Multi-Scale Interaction (MSI) search space to search for rational interaction operations between multi-scale features. In addition, we introduce a Spatial Alignment Module (SAM) to further enhance the attention consistency confronted with images from different sources. Under the proposed NAS scheme, a specific architecture is automatically searched, named as MSINet. Extensive experiments demonstrate that our method surpasses state-of-the-art ReID methods on both indomain and cross-domain scenarios. Source code available in https://github.com/vimar-gu/MSINet. Jianyang Gu, Kai Wang 0036, Hao Luo 0004, Chen Chen 0114, Wei Jiang 0009, Yuqiang Fang, Shanghang Zhang, Yang You 0001, Jian Zhao 0006 |
CVPR | 6 |
| 2023 | Single-Stage Multi-human Parsing via Point Sets and Center-based OffsetsabstractThis work studies the multi-human parsing problem. Existing methods, either following top-down or bottom-up two-stage paradigms, usually involve expensive computational costs. We instead present a high-performance Single-stage Multi-human Parsing (SMP) deep architecture that decouples the multi-human parsing problem into two fine-grained sub-problems,i.e., locating the human body and parts. SMP leverages the point features in the barycenter positions to obtain their segmentation and then generates a series of offsets from the barycenter of the human body to the barycenters of parts, thus performing human body and parts matching without the grouping process. Within the SMP architecture, we propose a Refined Feature Retain module to extract the global feature of instances through generated mask attention and a Mask of Interest Reclassify module as a trainable plug-in module to refine the classification results with the predicted segmentation. Extensive experiments on the MHPv2.0 dataset demonstrate the best effectiveness and efficiency of the proposed method, surpassing the state-of-the-art method by 2.1% in AP50p, 1.0% in APvolpsup>, and 1.2% in PCP50. Moreover, SMP also achieves superior performance in DensePose-COCO, verifying generalization of the model. In particular, the proposed method requires fewer training epochs and a less complex model architecture. Our codes are released in https://github.com/cjm-sfw/SMP. Jiaming Chu, Lei Jin 0003, Xiaojin Fan, Yinglei Teng, Yunchao Wei, Yuqiang Fang, Junliang Xing, Jian Zhao 0006 |
ACM Multimedia | 6 |
| 2023 | Alternating Direction Method of Multipliers for Convolutive Non-Negative Matrix FactorizationabstractNon-negative matrix factorization (NMF) has become a popular method for learning interpretable patterns from data. As one of the variants of standard NMF, convolutive NMF (CNMF) incorporates an extra time dimension to each basis, known as convolutive bases, which is well suited for representing sequential patterns. Previously proposed algorithms for solving CNMF use multiplicative updates which can be derived by either heuristic or majorization-minimization (MM) methods. However, these algorithms suffer from problems, such as low convergence rates, difficulty to reach exact zeroes during iterations and prone to poor local optima. Inspired by the success of alternating direction method of multipliers (ADMMs) on solving NMF, we explore variable splitting (i.e., the core idea of ADMM) for CNMF in this article. New closed-form algorithms of CNMF are derived with the commonly used β -divergences as optimization objectives. Experimental results have demonstrated the efficacy of the proposed algorithms on their faster convergence, better optima, and sparser results than state-of-the-art baselines. Yinan Li 0006, Ruili Wang 0001, Yuqiang Fang, Meng Sun 0001, Zhangkai Luo |
IEEE Trans. Cybern. | 3 |
| 2020 | Trainable TV-L1 model as recurrent nets for low-level vision
Yuqiang Fang, Wanting Ji |
Neural Comput. Appl. | 1 |
| 2019 | From TV-L1 to Gated Recurrent NetsabstractTV-L1is a classical diffusion-reaction model for low-level vision tasks, which can be solved by a duality based iterative algorithm. Considering the recent success of end-to-end learned representations, we propose a TV-LSTM network to unfold the duality based iterations into long short-term memory (LSTM) cells. To provide a trainable network, we relax the difference operators in the gate and cell update of TV-LSTM to trainable parameters. Then, the proposed end-to-end trainable TV-LSTMs can be naturally connected with various task-specific networks, e.g., optical flow estimation and image decomposition. Extensive experiments on optical flow estimation and structure + texture decomposition have demonstrated the effectiveness and efficiency of the proposed method. Yuqiang Fang, Haiyan Fan, Lin Sun 0004, Yulan Guo |
ICASSP | 1 |
| 2018 | Hybrid conditional random field based camera-LIDAR fusion for road detection
Liang Xiao 0007, Ruili Wang 0001, Bin Dai 0001, Yuqiang Fang, Daxue Liu, Tao Wu 0001 |
Inf. Sci. | 4 |
| 2017 | Design and Analysis of Electrical Resistance Feedback for Automated Patch Clamp on Adherent CellsabstractConventional patch clamp techniques require complex manual work. Automated patch clamp systems have been recently developed to eliminate the need of manual handling, but these systems can only perform recordings on suspended cells and require cell dissociation for recordings on adherent cells. Here, we develop a systematic approach to apply electrical resistance feedback for automated patch clamp recording on adherent cells. By developing resistance thresholds in the electrical resistance feedback algorithm, the engaging condition of the adherent cell was automatically determined. We analyze the critical parameters that affect the performance of the automated patch clamp recording. By using the new algorithm for automated engaging and formation, the cell-attached and whole-cell recordings on adherent cells were performed in less than 5 min with the yield of over 70%. The system is designed to directly measure the adherent cells and so, cell dissociation (e.g., trypsin treatment and cell scrapers) can be avoided and the cell properties can be maintained. By using the system, the efficiency of performing the automated patch clamp on adherent cells can be improved, which is of great importance in electrophysiological studies of single cells. Runhuai Yang, Yuqiang Fang, Jie Yang 0004, King Wai Chiu Lai |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2017 | Cross-Scale Cost Aggregation for Stereo MatchingabstractThis paper proposes a generic framework that enables a multiscale interaction in the cost aggregation step of stereo matching algorithms. Inspired by the formulation of image filters, we first reformulate cost aggregation from a weighted least-squares (WLS) optimization perspective and show that different cost aggregation methods essentially differ in the choices of similarity kernels. Our key motivation is that while the human stereo vision system processes information at both coarse and fine scales interactively for the correspondence search, state-of-the-art approaches aggregate costs at the finest scale of the input stereo images only, ignoring inter-consistency across multiple scales. This motivation leads us to introduce an inter-scale regularizer into the WLS optimization objective to enforce the consistency of the cost volume among the neighboring scales. The new optimization objective with the inter-scale regularization is convex, and thus, it is easily and analytically solved. Minimizing this new objective leads to the proposed framework. Since the regularization term is independent of the similarity kernel, various cost aggregation approaches, including discrete and continuous parameterization methods, can be easily integrated into the proposed framework. We show that the cross-scale framework is important as it effectively and efficiently expands state-of-the-art cost aggregation methods and leads to significant improvements, when evaluated on Middlebury, Middlebury Third, KITTI, and New Tsukuba data sets. Kang Zhang 0004, Yuqiang Fang, Dongbo Min, Lifeng Sun, Shiqiang Yang, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Traffic Sign Recognition Using Kernel Extreme Learning Machines With Deep Perceptual FeaturesabstractTraffic sign recognition plays an important role in autonomous vehicles as well as advanced driver assistance systems. Although various methods have been developed, it is still difficult for the state-of-the-art algorithms to obtain high recognition precision with low computational costs. In this paper, based on the investigation on the influence that color spaces have on the representation learning of convolutional neural network, a novel traffic sign recognition approach called DP-KELM is proposed by using a kernel-based extreme learning machine (KELM) classifier with deep perceptual features. Unlike the previous approaches, the representation learning process in DP-KELM is implemented in the perceptual Lab color space. Based on the learned deep perceptual feature, a kernel-based ELM classifier is trained with high computational efficiency and generalization performance. Through the experiments on the German traffic sign recognition benchmark, the proposed method is demonstrated to have higher precision than most of the state-of-the-art approaches. In particular, when compared with the hinge loss stochastic gradient descent method which has the highest precision, the proposed method can achieve a comparable recognition rate with significantly fewer computational costs. Yujun Zeng, Xin Xu 0001, Dayong Shen, Yuqiang Fang, Zhipeng Xiao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2016 | Learning deep compact channel features for object detection in traffic scenesabstractIn this work, we present a new multiple channel feature called Deep Compact Channel Feature (DCCF), which generates a compact, discriminative feature representation by a pre-trained deep encoder-decoder. With the combination of DCCF and boosted decision trees, a new object detector is proposed which achieved outstanding performance on standard pedestrian dataset INRIA and Caltech. Furthermore, a large scale and challenging Chinese Traffic Sign Detection benchmark is constructed. DCCF and other related methods are evaluated on this dataset. The dataset and baselines are available online. Yuqiang Fang, Lin Sun 0004, Hao Fu 0001, Tao Wu 0001, Ruili Wang 0001, Bin Dai 0001 |
ICIP | 1 |
| 2015 | Graph-Based Learning via Auto-Grouped Sparse Regularization and Kernelized ExtensionabstractThe key task in developing graph-based learning algorithms is constructing an informative graph to express the contextual information of a data manifold. Since traditional graph construction methods are sensitive to noise and less datum-adaptive to changes in density, a new method called$\ell^1$-graph was proposed recently. A graph construction needs to have two important properties: sparsity and locality. The$\ell^1$-graph has a strong sparsity property, but a weak locality property. Thus, we propose a new method of constructing an informative graph using auto-grouped sparse regularization based on the$\ell^1$-graph, which is called as Group Sparse graph (GS-graph). We also show how to efficiently construct a GS-graph in reproducing kernel Hilbert space with the kernel trick. The new methods, the GS-graph and its kernelized version (KGS-graph), have the same noise-insensitive property as that of$\ell^1$-graph and also can successively preserve the properties of sparsity and locality simultaneously. Furthermore, we integrate the proposed graph with several graph-based learning algorithms to demonstrate the effectiveness of our method. The empirical studies on benchmarks show that the proposed methods outperform the$\ell^1$-graph and other traditional graph construction methods in various learning tasks. Yuqiang Fang, Ruili Wang 0001, Bin Dai 0001, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | DL-SFA: Deeply-Learned Slow Feature Analysis for Action RecognitionabstractMost of the previous work on video action recognition use complex hand-designed local features, such as SIFT, HOG and SURF, but these approaches are implemented sophisticatedly and difficult to be extended to other sensor modalities. Recent studies discover that there are no universally best hand-engineered features for all datasets, and learning features directly from the data may be more advantageous. One such endeavor is Slow Feature Analysis (SFA) proposed by Wiskott and Sejnowski [33]. SFA can learn the invariant and slowly varying features from input signals and has been proved to be valuable in human action recognition [34]. It is also observed that the multi-layer feature representation has succeeded remarkably in widespread machine learning applications. In this paper, we propose to combine SFA with deep learning techniques to learn hierarchical representations from the video data itself. Specifically, we use a two-layered SFA learning structure with 3D convolution and max pooling operations to scale up the method to large inputs and capture abstract and structural features from the video. Thus, the proposed method is suitable for action recognition. At the same time, sharing the same merits of deep learning, the proposed method is generic and fully automated. Our classification results on Hollywood2, KTH and UCF Sports are competitive with previously published results. To highlight some, on the KTH dataset, our recognition rate shows approximately 1% improvement in comparison to state-of-the-art methods even without supervision or dense sampling. Lin Sun 0004, Kui Jia, Tsung-Han Chan, Yuqiang Fang, Gang Wang 0012, Shuicheng Yan |
CVPR | 4 |
| 2014 | Cross-Scale Cost Aggregation for Stereo MatchingabstractHuman beings process stereoscopic correspondence across multiple scales. However, this bio-inspiration is ignored by state-of-the-art cost aggregation methods for dense stereo correspondence. In this paper, a generic cross-scale cost aggregation framework is proposed to allow multi-scale interaction in cost aggregation. We firstly reformulate cost aggregation from a unified optimization perspective and show that different cost aggregation methods essentially differ in the choices of similarity kernels. Then, an inter-scale regularizer is introduced into optimization and solving this new optimization problem leads to the proposed framework. Since the regularization term is independent of the similarity kernel, various cost aggregation methods can be integrated into the proposed general framework. We show that the cross-scale framework is important as it effectively and efficiently expands state-of-the-art cost aggregation methods and leads to significant improvements, when evaluated on Middlebury, KITTI and New Tsukuba datasets. Kang Zhang 0004, Yuqiang Fang, Dongbo Min, Lifeng Sun, Shiqiang Yang, Shuicheng Yan, Qi Tian 0001 |
CVPR | 2 |
| 2014 | Efficient Vehicle Localization Based on Road-Boundary Maps
Dawei Zhao 0003, Tao Wu 0001, Yuqiang Fang, Ruili Wang 0001, Bin Dai 0001 |
PRICAI | 3 |
| 2014 | Decomposition and Extraction: A New Framework for Visual ClassificationabstractIn this paper, we present a novel framework for visual classification based on hierarchical image decomposition and hybrid midlevel feature extraction. Unlike most midlevel feature learning methods, which focus on the process of coding or pooling, we emphasize that the mechanism of image composition also strongly influences the feature extraction. To effectively explore the image content for the feature extraction, we model a multiplicity feature representation mechanism through meaningful hierarchical image decomposition followed by a fusion step. In particularly, we first propose a new hierarchical image decomposition approach in which each image is decomposed into a series of hierarchical semantical components, i.e, the structure and texture images. Then, different feature extraction schemes can be adopted to match the decomposed structure and texture processes in a dissociative manner. Here, two schemes are explored to produce property related feature representations. One is based on a single-stage network over hand-crafted features and the other is based on a multistage network, which can learn features from raw pixels automatically. Finally, those multiple midlevel features are incorporated by solving a multiple kernel learning task. Extensive experiments are conducted on several challenging data sets for visual classification, and experimental results demonstrate the effectiveness of the proposed method. Yuqiang Fang, Qiang Chen 0007, Lin Sun 0004, Bin Dai 0001, Shuicheng Yan |
IEEE Trans. Image Process. | 1 |
| 2014 | Unified Structured Learning for Simultaneous Human Pose Estimation and Garment Attribute ClassificationabstractIn this paper, we utilize structured learning to simultaneously address two intertwined problems: 1) human pose estimation (HPE) and 2) garment attribute classification (GAC), which are valuable for a variety of computer vision and multimedia applications. Unlike previous works that usually handle the two problems separately, our approach aims to produce an optimal joint estimation for both HPE and GAC via a unified inference procedure. To this end, we adopt a preprocessing step to detect potential human parts from each image (i.e., a set of candidates) that allows us to have a manageable input space. In this way, the simultaneous inference of HPE and GAC is converted to a structured learning problem, where the inputs are the collections of candidate ensembles, outputs are the joint labels of human parts and garment attributes, and joint feature representation involves various cues such as pose-specific features, garment-specific features, and cross-task features that encode correlations between human parts and garment attributes. Furthermore, we explore the strong edge evidence around the potential human parts so as to derive more powerful representations for oriented human parts. Such evidences can be seamlessly integrated into our structured learning model as a kind of energy function, and the learning process could be performed by standard structured support vector machines algorithm. However, the joint structure of the two problems is a cyclic graph, which hinders efficient inference. To resolve this issue, we compute instead approximate optima using an iterative procedure, where in each iteration, the variables of one problem are fixed. In this way, satisfactory solutions can be efficiently computed by dynamic programming. Experimental results on two benchmark data sets show the state-of-the-art performance of our approach. Jie Shen 0005, Guangcan Liu, Jia Chen 0001, Yuqiang Fang, Jianbin Xie, Yong Yu 0001, Shuicheng Yan |
IEEE Trans. Image Process. | 4 |
| 2012 | Graph-Oriented Learning via Automatic Group Sparsity for Data AnalysisabstractThe key task in graph-oriented learning is constructing an informative graph to model the geometrical and discriminant structure of a data manifold. Since traditional graph construction methods are sensitive to noise and less datum-adaptive to changes in density, a new graph construction method so-called ℓ1-Graph has been proposed [1] recently. A graph construction method needs to have two important properties: sparsity and locality. However, the ℓ1-Graph is strong in sparsity property, but weak in locality. In order to overcome such limitation, we propose a new method of constructing an informative graph using automatic group sparse regularization based on the work of ℓ1-Graph, which is called as group sparse graph (GroupSp-Graph). The newly developed GroupSp-Graph has the same noise-insensitive property as ℓ1-Graph, and also can successively preserve the group and local information in the graph. In other words, the proposed group sparse graph has both properties of sparsity and locality simultaneously. Furthermore, we integrate the proposed graph with several graph-oriented learning algorithms: spectral embedding, spectral clustering, subspace learning and manifold regularized non-negative matrix factorization. The empirical studies on benchmark data sets show that the proposed algorithms achieve considerable improvement over classic graph constructing methods and the ℓ1-Graph method in various learning task. Yuqiang Fang, Ruili Wang 0001, Bin Dai 0001 |
ICDM | 1 |