EDBT 2026 Demo / reviewers in the wild / expert
Hongxing Wang 0001
dblp:74/7298-1
· DBLP profile ↗
40ranked-venue papers
6as first author
19since 2021 · last 2025
0000-0001-7799-1023ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 21 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Modality-Aware Shot Relating and Comparing for Video Scene DetectionabstractVideo scene detection involves assessing whether each shot and its surroundings belong to the same scene. Achieving this requires meticulously correlating multi-modal cues, e.g., visual entity and place modalities, among shots and comparing semantic changes around each shot. However, most methods treat multi-modal semantics equally and do not examine contextual differences between the two sides of a shot, leading to sub-optimal detection performance. In this paper, we propose the Modality-Aware Shot Relating and Comparing approach (MASRC), which enables relating shots per their own characteristics of visual entity and place modalities, as well as comparing multi-shots similarities to have scene changes explicitly encoded. Specifically, to fully harness the potential of visual entity and place modalities in modeling shot relations, we mine long-term shot correlations from entity semantics while simultaneously revealing short-term shot correlations from place semantics. In this way, we can learn distinctive shot features that consolidate coherence within scenes and amplify distinguishability across scenes. Once equipped with distinctive shot features, we further encode the relations between preceding and succeeding shots of each target shot by similarity convolution, aiding in the identification of scene ending shots. We validate the broad applicability of the proposed components in MASRC. Extensive experimental results on public benchmark datasets demonstrate that the proposed MASRC significantly advances video scene detection. Jiawei Tan, Hongxing Wang 0001, Kang Dang, Zhilong Ou |
AAAI | 2 |
| 2025 | Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2DabstractVideo moment retrieval aims to locate specific moments from a video according to the query text. This task presents two main challenges: i) aligning the query and video frames at the feature level, and ii) projecting the query-aligned frame features to the start and end boundaries of the matching interval. Previous work commonly involves all frames in feature alignment, easy to cause aligning irrelevant frames with the query. Furthermore, they forcibly map visual features to interval boundaries but ignoring the information gap between them, yielding suboptimal performance. In this study, to reduce distraction from irrelevant frames, we designate an anchor frame as that with the maximum query-frame relevance measured by the established Vision-Language Model. Via similarity comparison between the anchor frame and the others, we produce a semantically compact segment around the anchor frame, which serves as a guide to align features of query and related frames. We observe that such a feature alignment will make similarity cohesive between target frames, which enables us to predict the interval boundaries by a single point detection in the 2D semantic similarity space of frames, thus well bridging the information gap between frame semantics and temporal boundaries. Experimental results across various datasets demonstrate that our approach significantly improves the alignment between queries and video frames while effectively predicting temporal moment boundaries. Especially, on QVHighlights Test and ActivityNet Captions datasets, our proposed approach achieves 3.8% and 7.4% respectively higher than current state-of-the-art [email protected] performance. The code is available at https://github.com/ExMorgan-Alter/AFAFSGD. Jiawei Tan, Hongxing Wang 0001, Junwu Weng, Zhilong Ou, Kang Dang |
CVPR | 2 |
| 2025 | NCD: Normal-Guided Chamfer Distance Loss for Watertight Mesh Reconstruction from Unoriented Point CloudsabstractAbstract As a widely used loss function in learnable watertight mesh reconstruction from unoriented point clouds, Chamfer Distance (CD) efficiently quantifies the alignment between the sampled point cloud from the reconstructed mesh and its corresponding input point cloud. Occasionally, to enhance reconstruction fidelity, CD incorporates a normal consistency term, albeit at the cost of efficiency. In this context, normal estimation for unoriented point clouds requires computationally intensive matrix decomposition or specialized pre‐trained models, whereas deriving normals for mesh‐sampled points can be readily achieved using the cross product of mesh vertices. However, the reconstruction models employing CD and its variants typically rely solely on the spatial coordinates of the points, which omits normal information in favor of efficiency and deployability. To tackle this challenge, we propose a novel loss function for watertight mesh reconstruction from unoriented point clouds, termed Normal‐guided Chamfer Distance (NCD). Building upon CD, NCD introduces a normal‐steered weighting mechanism based on the angle between the normal at each mesh‐sampled point and the vector to its corresponding input point, offering several advantages: (i) it leverages readily available mesh‐sampled point normals to weight coordinate‐based Euclidean distances, thus extending the capability of CD; (ii) it eliminates the need for normal estimation from input unoriented point clouds; (iii) it incurs a negligible increase in computational complexity compared to CD. We employ NCD as the training loss for point‐to‐mesh reconstruction with multiple models and initial watertight meshes on benchmark datasets, demonstrating its superiority over state‐of‐the‐art CD variants. Jiawei Tan, Zhilong Ou, Hongxing Wang 0001 |
Comput. Graph. Forum | 4 |
| 2025 | Label refinement for change detection in remote sensing
Zhilong Ou, Hongxing Wang 0001, Jiawei Tan, Zhangbin Qian |
Image Vis. Comput. | 2 |
| 2025 | D$^{2}$2AE: Data-Decoupled Active Experts Promote One-Class Anomaly DiscoveryabstractAnomaly detection (AD) suffers from severe performance decrease when dealing with corrupted datasets. By querying limited annotations from an oracle, active learning is prevalent in mitigating this problem. However, previous work ignores the particularity of the AD task on the one-class setting, where wrong pseudo-annotations of anomaly noise will mislead the active inference results. To address this challenge, we propose D 2AE, a novel active AD framework through Decoupling Data pools between training and inference process for Active Experts. Specifically, we design a data-splitting module named as DSS to obtain diverse subsets and weaken the mutual interference of similar anomalies. To decouple the data, we propose an Independent Active Experts (IAE) module formed by multiple expert replications, on which each data subset is trained by one separate expert (squad) and inferred by the other non-training ones. To further improve the efficiency of data utilization, we propose Active Expert Squad (AES) beyond IAE by introducing Mixtureof-Experts. The commonality and specificity between expert squads promote model training and active query, respectively. We conduct extensive experiments on various image, tabular, and NLP datasets. Experimental results show the superiority of our solution compared with existing methods. Hongxing Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Aligning Instance-Semantic Sparse Representation Towards Unsupervised Object Segmentation and Shape Abstraction With Repeatable PrimitivesabstractUnderstanding 3D object shapes necessitates shape representation by object parts abstracted from results of instance and semantic segmentation. Promising shape representations enable computers to interpret a shape with meaningful parts and identify their repeatability. However, supervised shape representations depend on costly annotation efforts, while current unsupervised methods work under strong semantic priors and involve multi-stage training, thereby limiting their generalization and deployment in shape reasoning and understanding. Driven by the tendency of high-dimensional semantically similar features to lie in or near low-dimensional subspaces, we introduce a one-stage, fully unsupervised framework towards semantic-aware shape representation. This framework produces joint instance segmentation, semantic segmentation, and shape abstraction through sparse representation and feature alignment of object parts in a high-dimensional space. For sparse representation, we devise a sparse latent membership pursuit method that models each object part feature as a sparse convex combination of point features at either the semantic or instance level, promoting part features in the same subspace to exhibit similar semantics. For feature alignment, we customize an attention-based strategy in the feature space to align instance- and semantic-level object part features and reconstruct the input shape using both of them, ensuring geometric reusability and semantic consistency of object parts. To firm up semantic disambiguation, we construct cascade unfrozen learning on geometric parameters of object parts. Experiments conducted on benchmark datasets confirm that our approach results in instance- and semantic-level joint segmentation and shape abstraction with repeatable primitives, providing coherent semantic interpretations of 3D object shapes across categories in a one-stage, fully unsupervised manner, without relying on annotations or heuristic semantic priors. Hongxing Wang 0001, Jiawei Tan, Zhilong Ou, Junsong Yuan 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Neighbor Relations Matter in Video Scene DetectionabstractVideo scene detection aims to temporally link shots for obtaining semantically compact scenes. It is essential for this task to capture scene-distinguishable affinity among shots by similarity assessment. However, most methods relies on ordinary shot-to-shot similarities, which may inveigle similar shots into being linked even though they are from different scenes, and meanwhile hinder dissimilar shots from being blended into a complete scene. In this paper, we propose NeighborNet to inject shot contexts into shot-to-shot similarities through carefully exploring the relations between semantic/temporal neighbors of shots over a local time period. In this way, shot-to-shot similarities are remeasured as semantic/temporal neighbor-aware similarities so that NeighborNet can learn context embedding into shot features using graph convolutional network. As a result, not only do the learned shot features suppress the affinity among similar shots from different scenes, but they also promote the affinity among dissimilar shots in the same scene. Experimental results on public benchmark datasets show that our proposed NeighborNet yields substantial improvements in video scene detection, especially outperforms released state-of-the-arts by at least 6% in Average Precision (AP). The code is available at https://github.com/ExMorgan-Alter/NeighborNet. Jiawei Tan, Hongxing Wang 0001, Zhilong Ou, Zhangbin Qian |
CVPR | 2 |
| 2024 | CLIP-Driven Multi-Scale Instance Learning for Weakly Supervised Video Anomaly DetectionabstractExisting weakly supervised video anomaly detection methods mainly employ Multiple Instance Learning (MIL) to identify abnormal snippets in untrimmed videos. However, the semantics and presentations of anomalies frequently exhibit ambiguity that MIL is difficult to tackle. Moreover, MIL suffers from false alarms due to its independent optimization of each instance, neglecting temporal correlation between adjacent snippets. Consequently, we badly need to better connect abnormal presentations and their semantics, as well as to enable multi-temporal-scale anomaly discovery. This paper proposes a CLIP-Driven Multi-Scale Instance Learning (CMSIL) framework with two branches including Vision-Language (VL) and Multi-Scale Instance Learning (MSIL). The VL branch leverages the powerful visual concept priors from Contrastive Language-Image Pre-training (CLIP) to generate pseudo anomalies, thereby providing suspected anomaly cues for model training guidance. The MSIL branch utilizes a feature pyramid to fully mine fine-grained temporal dependencies by employing MIL within each pyramid level to learn anomalous patterns across different temporal scales. By collaborating with the two branches, the proposed CMSIL shows better proficiency in handling anomalies with varying durations. Extensive experiments on the XD-Violence and UCF-Crime datasets demonstrate the superior performance of our method. The code is available at https://github.com/casperZB/CMSIL. Zhangbin Qian, Jiawei Tan, Zhilong Ou, Hongxing Wang 0001 |
ICME | 4 |
| 2024 | Semantic Transition Detection for Self-supervised Video Scene Segmentation
Jiawei Tan, Pingan Yang, Hongxing Wang 0001 |
MMM (3) | 4 |
| 2024 | Object Recognition Consistency in Regression for Active Detection
Ming Jing, Zhilong Ou, Hongxing Wang 0001 |
Mach. Vis. Appl. | 3 |
| 2024 | Lifelong learning with selective attention over seen classes and memorized instances
Hongxing Wang 0001 |
Neural Comput. Appl. | 2 |
| 2024 | Shared Latent Membership Enables Joint Shape Abstraction and Segmentation With Deformable SuperquadricsabstractPart-level 3D shape representations are crucial to shape reasoning and understanding. Two key sub-tasks are: 1) shape abstraction, creating primitive-based object parts; and 2) shape segmentation, finding partition-based object parts. However, for 3D object point clouds, most advanced methods produce parts relying on task-specific priors, such as similarity metrics and primitive geometries, resulting in misleading parts that deviate from semantics. To address prior limitations, we establish a foundation for joint shape abstraction and shape segmentation as formal linear transformations within a shared latent space, encapsulating essential dual-purpose membership information linking points and object parts for mutual reinforcement. We demonstrate that the transformations are underpinned by a derivation based on k-means, Non-negative Matrix Factorization (NMF), and the attention mechanism. As a result, we introduce Latent Membership Pursuit (LMP) for joint optimization of shape abstraction and segmentation. LMP utilizes a shared latent representation of object part membership to autonomously identify common object parts in both tasks without any supervision and priors. Furthermore, we adapt deformable superquadrics (DSQs) for primitives to capture variable part-level geometric and semantic information. Experiments on benchmark datasets validate that our approach enables mutual learning of shape abstraction and segmentation, and promotes consistent interpretations of 3D object shapes across instances and even categories in a fully unsupervised manner. Hongxing Wang 0001, Jiawei Tan, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | DualStreamFoveaNet: A Dual Stream Fusion Architecture With Anatomical Awareness for Robust Fovea LocalizationabstractAccurate fovea localization is essential for analyzing retinal diseases to prevent irreversible vision loss. While current deep learning-based methods outperform traditional ones, they still face challenges such as the lack of local anatomical landmarks around the fovea, the inability to robustly handle diseased retinal images, and the variations in image conditions. In this paper, we propose a novel transformer-based architecture called DualStreamFoveaNet (DSFN) for multi-cue fusion. This architecture explicitly incorporates long-range connections and global features using retina and vessel distributions for robust fovea localization. We introduce a spatial attention mechanism in the dual-stream encoder to extract and fuse self-learned anatomical information, focusing more on features distributed along blood vessels and significantly reducing computational costs by decreasing token numbers. Our extensive experiments show that the proposed architecture achieves state-of-the-art performance on two public datasets and one large-scale private dataset. Furthermore, we demonstrate that the DSFN is more robust on both normal and diseased retina images and has better generalization capacity in cross-dataset experiments. Sifan Song, Jinfeng Wang 0008, Zilong Wang 0006, Hongxing Wang 0001, Jionglong Su, Kang Dang |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Characters Link Shots: Character Attention Network for Movie Scene SegmentationabstractMovie scene segmentation aims to automatically segment a movie into multiple story units, i.e., scenes, each of which is a series of semantically coherent and time-continual shots. Previous methods have continued efforts on shot semantic association, but few take into account the impact of different semantics on foreground characters and background scenes in movie shots. In particular, the background scene in the shot can adversely affect scene boundary classification. Motivated by the fact that it is the characters who drive the plot development of a movie scene, we build a Character Attention Network (CANet) to detect movie scene boundaries in a character-centric fashion. To eliminate the background clutter, we extract multi-view character semantics for each shot in terms of human bodies and faces. Furthermore, we equip our CANet with two stages of character attention. The first is Masked Shot Attention (MSA) through selective self-attention over similar temporal contexts from multi-view character semantics to yield an enhanced omni-view shot representation, by which the CANet can better handle the variations of characters in pose and appearance. The second is Key Character Attention (KCA) through temporal-aware attention on character reappearances for Bidirectional Long Short-Term Memory (Bi-LSTM) feature association so that linking shots can be focused on those with recurring key characters. We encourage the proposed CANet in learning boundary-discriminative shot features. Specifically, we formulate a Boundary-Aware circle Loss (BAL) to push far apart CANet-features between adjacent scenes, which is also coupled with the cross-entropy loss to drive CANet-features sensitive to scene boundaries. Experimental results on the MovieNet-SSeg and OVSD datasets show that our method achieves superior performance in temporal scene segmentation compared with state-of-the-art methods. Jiawei Tan, Hongxing Wang 0001, Junsong Yuan 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Temporal Scene Montage for Self-Supervised Video Scene Boundary DetectionabstractOnce a video sequence is organized as basic shot units, it is of great interest to temporally link shots into semantic-compact scene segments to facilitate long video understanding. However, it still challenges existing video scene boundary detection methods to handle various visual semantics and complex shot relations in video scenes. We proposed a novel self-supervised learning method, Video Scene Montage for Boundary Detection (VSMBD), to extract rich shot semantics and learn shot relations using unlabeled videos. More specifically, we present Video Scene Montage (VSM) to synthesize reliable pseudo scene boundaries, which learns task-related semantic relations between shots in a self-supervised manner. To lay a solid foundation for modeling semantic relations between shots, we decouple visual semantics of shots into foreground and background. Instead of costly learning from scratch as in most previous self-supervised learning methods, we build our model upon large-scale pre-trained visual encoders to extract the foreground and background features. Experimental results demonstrate VSMBD trains a model with strong capability in capturing shot relations, surpassing previous methods by significant margins. The code is available at https://github.com/mini-mind/VSMBD. Jiawei Tan, Pingan Yang, Hongxing Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Cascade Sampling via Dual Uncertainty for Active Entity Alignment
Jiye Xie, Jiawei Tan, Hongxing Wang 0001 |
KSEM (2) | 4 |
| 2022 | Multiple cross-attention for video-subtitle moment retrieval
Hongxing Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Low-resolution assisted three-stream network for person re-identification
Jiahong Xie, Yongxin Ge, Junyin Zhang, Sheng Huang 0001, Feiyu Chen 0002, Hongxing Wang 0001 |
Vis. Comput. | 6 |
| 2021 | Self-attention binary neural tree for video summarization
Hongxing Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2020 | Deep Adversarial Active Learning With Model Uncertainty For Image ClassificationabstractActive learning aims at selecting and labeling as few samples as possible to train a good task model. Most existing methods rely on various heuristics to iteratively select a single sample in each active learning loop, thus cannot tackle large datasets efficiently. In this paper, we propose a new batch-mode active learning method, which can plug model prediction uncertainty into adversarial batch selection to ensure the selected samples are representative in unlabeled data, complementary to labeled data, and beneficial for model training. Experiments on four benchmark image datasets validate the effectiveness and efficiency of the proposed method for active image classification in comparison with the state-of-the-art methods. Hongxing Wang 0001 |
ICIP | 2 |
| 2020 | Video Summarization with a Dual Attention Capsule NetworkabstractIn this paper, we address the problem of video summarization, which aims at selecting a subset of video frames as a summary to represent the original video contents compactly and completely. We propose a simple but effective supervised approach with a dual attention capsule network towards this end. Unlike existing LSTM based methods, it pays attention to short- and long-term dependencies among video frames through an elaborate dual self-attention architecture, which can handle longer-term dependencies and admit parallel computing. To reconcile the outputs of dual self-attention, we rely on a two-stream capsule network to learn the underlying frame selection criteria. Experiments on real-world datasets show the advantages of the proposed approach compared with state-of-the-art methods. Hongxing Wang 0001, Jianyu Yang 0002 |
ICPR | 2 |
| 2020 | Multi-Grained Selection and Fusion for Fine-Grained Image RepresentationabstractHow to learn a good fine-grained image representation is a key problem for fine-grained tasks. Most previous supervised methods suffer from insufficient training data, which require laborious annotations of fine-grained objects. In this paper, we propose an annotation-free method for fine-grained image representation, dubbed Multi-Grained Selection and Fusion (MGSF). The proposed MGSF extracts two types of visual features, i.e., fine-grained discriminative features that highlight informative convolutional parts by spatial selection and channel selection, and coarse-grained scene-level features that provide context information for fine-grained objects. Extensive experiments in fine-grained image retrieval demonstrate the superiority of our proposed representation compared with the state-of-the-art approaches on several fine-grained datasets. Jianrong Jiang, Hongxing Wang 0001 |
IJCNN | 2 |
| 2018 | Deep Multi-Metric Learning for Person Re-IdentificationabstractIn this paper, to exploit more discriminative information of the global-body and body-parts features, we present a novel deep multi-metric learning (DMML) network for person re-identification under the triplet framework. The main novelty of our learning framework lies in two aspects: 1) Unlike most existing metric learning-based approaches, which learn only one distance metric for comparison, our DMM-L method aims to learn different metrics for the global-body and body-parts features respectively by using convolutional neural network (CNN); 2) A new multi-metric loss function is proposed to train the DMML network, under which the distance of each negative pair is greater than that of each positive pair by a predefined margin, and the correlations of different metrics are maximized. Compared with the previous person re-identification methods that have shown state-of-the-art performances, our DMML approach can achieve competitive results on the challenging CUHK03, CUHKOl, VIPeR and iLIDS datasets. Yongxin Ge, Xinqian Gu, Min Chen 0016, Hongxing Wang 0001, Dan Yang 0001 |
ICME | 4 |
| 2018 | Joint Deep Learning for RGB-D Action RecognitionabstractRecent approaches in RGB-based and depth-based human action recognition achieved outstanding performance respectively, which demonstrate the effectiveness of RGB and depth modalities for action classification, however it is infrequent to consider them both. Currently, available multimodal-based methods of action recognition suffer from some limitations, including non-end-to-end training, violent fusion and inefficiency. In this paper, we propose a novel joint deep learning (JDL) model which is capable of: 1) jointly optimizing the object of classification and feature extraction through a novel end-to-end two-stream deep learning model, 2) refining common-specific features via introducing the constraint of similarity loss in high-level, and 3) using 2D convolution kernel instead of 3D convolution kernel during feature extraction for gaining the efficiency. The experiments on two challenging datasets show the promising performance of our architecture. Xiaolei Qin, Yongxin Ge, Liuwei Zhan, Guangrui Li 0003, Sheng Huang 0001, Hongxing Wang 0001, Feiyu Chen 0002 |
VCIP | 6 |
| 2018 | Improved hypergraph regularized Nonnegative Matrix Factorization with sparse representationabstractAs a commonly used data representation technique, Nonnegative Matrix Factorization (NMF) has received extensive attentions in the pattern recognition and machine learning communities over decades, since its working mechanism is in accordance with the way how the human brain recognizes objects. Inspired by the remarkable successes of manifold learning, more and more researchers attempt to incorporate the manifold learning into NMF for finding a compact representation ,which uncovers the hidden semantics and respects the intrinsic geometric structure simultaneously. Graph regularized Nonnegative Matrix Factorization (GNMF) is one of the representative approaches in this category. The core of such approach is the graph, since a good graph can accurately reveal the relations of samples which benefits the data geometric structure depiction. In this paper, we leverage the sparse representation to construct a sparse hypergraph for better capturing the manifold structure of data, and then impose the sparse hypergraph as a regularization to the NMF framework to present a novel GNMF algorithm called Sparse Hypergraph regularized Nonnegative Matrix Factorization (SHNMF). Since the sparse hypergraph inherits the merits of both the sparse representation and the hypergraph model, SHNMF enjoys more robustness and can better exploit the high-order discriminant manifold information for data representation . We apply our work to address the image clustering issue for evaluation. The experimental results on five popular image databases show the promising performances of the proposed approach in comparison with the state-of-the-art NMF algorithms. Sheng Huang 0001, Hongxing Wang 0001, Yongxin Ge, Luwen Huangfu, Xiaohong Zhang 0002, Dan Yang 0001 |
Pattern Recognit. Lett. | 2 |
| 2018 | Representative Selection on a HypersphereabstractFinding representative examples is important for pattern discovery and data analytics. In this letter, we propose a novel formulation for representative selection via center reconstruction on a hypersphere, which makes the selection not affect the center information of given data, thus, the overall data distribution can also be easily maintained by those selected representatives. We adopt the proximal gradient strategy and the fast iterative shrinkage-thresholding algorithm to solve the problem. Compared with most existing methods with cubic time complexity in the number of samples, our method is considerably more efficient, with time complexity reduced to being quadratic. Our formulation has only one parameter. We analyze the behavior of this parameter and analyze its bound theoretically. Experiments on synthesis and real-world datasets validate the effectiveness and efficiency of our method and demonstrate its robustness to noise compared with the state-of-the-art methods. Hongxing Wang 0001, Junsong Yuan 0001 |
IEEE Signal Process. Lett. | 1 |
| 2018 | Video Summarization Via Multiview Representative SelectionabstractVideo contents are inherently heterogeneous. To exploit different feature modalities in a diverse video collection for video summarization, we propose to formulate the task as a multiview representative selection problem. The goal is to select visual elements that are representative of a video consistently across different views (i.e., feature modalities). We present in this paper the multiview sparse dictionary selection with centroid co-regularization method, which optimizes the representative selection in each view, and enforces that the view-specific selections to be similar by regularizing them towards a consensus selection. We also introduce a diversity regularizer to favor a selection of diverse representatives. The problem can be efficiently solved by an alternating minimizing optimization with the fast iterative shrinkage thresholding algorithm. Experiments on synthetic data and benchmark video datasets validate the effectiveness of the proposed approach for video summarization, in comparison with other video summarization methods and representative selection methods such as K-medoids, sparse dictionary selection, and multiview clustering. Jingjing Meng, Suchen Wang, Hongxing Wang 0001, Junsong Yuan 0001, Yap-Peng Tan |
IEEE Trans. Image Process. | 3 |
| 2017 | Representative Selection with Structured SparsityabstractWe propose a novel formulation to find representatives in data samples via learning with structured sparsity . To find representatives with both diversity and representativeness , we formulate the problem as a structurally-regularized learning where the objective function consists of a reconstruction error and three structured regularizers: (1) group sparsity regularizer, (2) diversity regularizer, and (3) locality-sensitivity regularizer. For the optimization of the objective, we propose an accelerated proximal gradient algorithm, combined with the proximal-Dykstra method and the calculation of parametric maximum flows. Experiments on image and video data validate the effectiveness of our method in finding exemplars with diversity and representativeness and demonstrate its robustness to outliers. Hongxing Wang 0001, Yoshinobu Kawahara, Chaoqun Weng, Junsong Yuan 0001 |
Pattern Recognit. | 1 |
| 2017 | Discovering Class-Specific Spatial Layouts for Scene RecognitionabstractScene image is a spatial composition of objects and background contexts and finding discriminative spatial layouts is critical for scene recognition. In this letter, we propose an ℓ1-regularized max-margin formulation to discover class-specific spatial layouts by jointly learning the image classifier and the class-specific spatial layouts for scene recognition. Unlike previous methods that classify images into different categories either without considering the spatial layouts explicitly or only using class generic spatial layout, our proposed method can discover a sparse combination of class-specific spatial layouts for different scenes and boost the recognition performance. Experiments on scene-15, landuse-21, and MIT indoor-67 datasets validate the advantages of our proposed algorithm. Chaoqun Weng, Hongxing Wang 0001, Junsong Yuan 0001, Xudong Jiang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2016 | From Keyframes to Key Objects: Video Summarization by Representative Object Proposal SelectionabstractWe propose to summarize a video into a few key objects by selecting representative object proposals generated from video frames. This representative selection problem is formulated as a sparse dictionary selection problem, i.e., choosing a few representatives object proposals to reconstruct the whole proposal pool. Compared with existing sparse dictionary selection based representative selection methods, our new formulation can incorporate object proposal priors and locality prior in the feature space when selecting representatives. Consequently it can better locate key objects and suppress outlier proposals. We convert the optimization problem into a proximal gradient problem and solve it by the fast iterative shrinkage thresholding algorithm (FISTA). Experiments on synthetic data and real benchmark datasets show promising results of our key object summarization approach in video content mining and search. Comparisons with existing representative selection approaches such as K-mediod, sparse dictionary selection and density based selection validate that our formulation can better capture the key video objects despite appearance variations, cluttered backgrounds and camera motions. Jingjing Meng, Hongxing Wang 0001, Junsong Yuan 0001, Yap-Peng Tan |
CVPR | 2 |
| 2016 | Invariant multi-scale descriptor for shape representation, matching and retrieval
Jianyu Yang 0002, Hongxing Wang 0001, Junsong Yuan 0001, Youfu Li 0001, Jianyang Liu |
Comput. Vis. Image Underst. | 2 |
| 2015 | Laplacian Scale-Space Behavior of Planar Curve CornersabstractScale-space behavior of corners is important for developing an efficient corner detection algorithm. In this paper, we analyze the scale-space behavior with the Laplacian of Gaussian (LoG) operator on a planar curve which constructs Laplacian Scale Space (LSS). The analytical expression of a Laplacian Scale-Space map (LSS map) is obtained, demonstrating the Laplacian Scale-Space behavior of the planar curve corners, based on a newly defined unified corner model. With this formula, some Laplacian Scale-Space behavior is summarized. Although LSS demonstrates some similarities to Curvature Scale Space (CSS), there are still some differences. First, no new extreme points are generated in the LSS. Second, the behavior of different cases of a corner model is consistent and simple. This makes it easy to trace the corner in a scale space. At last, the behavior of LSS is verified in an experiment on a digital curve. Xiaohong Zhang 0002, Ying Qu 0007, Dan Yang 0001, Hongxing Wang 0001, Jeffrey D. Kymer |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Collaborative Multifeature Fusion for Transductive Spectral LearningabstractMuch existing work of multifeature learning relies on the agreement among different feature types to improve the clustering or classification performance. However, as different feature types could have different data characteristics, such a forced agreement among different feature types may not bring a satisfactory result. We propose a novel transductive learning approach that considers multiple feature types simultaneously to improve the classification performance. Instead of forcing different feature types to agree with each other, we perform spectral clustering in different feature types separately. Each data sample is then described by a co-occurrence of feature patterns among different feature types, and we apply these feature co-occurrence representations to perform transductive learning, such that data samples of similar feature co-occurrence pattern will share the same label. As the spectral clustering results in different feature types and the formed co-occurrence patterns influence each other under the transductive learning formulation, an iterative optimization approach is proposed to decouple these factors. Different from co-training that need to iteratively update individual feature type, our method allows all feature types to collaborate simultaneously. It can naturally handle multiple feature types together and is less sensitive to noisy feature types. The experimental results on synthetic, object, and action recognition datasets all validate the advantages of our method compared to state-of-the-art methods. Hongxing Wang 0001, Junsong Yuan 0001 |
IEEE Trans. Cybern. | 1 |
| 2014 | Multi-feature Spectral Clustering with Minimax OptimizationabstractIn this paper, we propose a novel formulation for multi-feature clustering using minimax optimization. To find a consensus clustering result that is agreeable to all feature modalities, our objective is to find a universal feature embedding, which not only fits each individual feature modality well, but also unifies different feature modalities by minimizing their pairwise disagreements. The loss function consists of both (1) unary embedding cost for each modality, and (2) pairwise disagreement cost for each pair of modalities, with weighting parameters automatically selected to maximize the loss. By performing minimax optimization, we can minimize the loss for the worst case with maximum disagreements, thus can better reconcile different feature modalities. To solve the minimax optimization, an iterative solution is proposed to update the universal embedding, individual embedding, and fusion weights, separately. Our minimax optimization has only one global parameter. The superior results on various multi-feature clustering tasks validate the effectiveness of our approach when compared with the state-of-the-art methods. Hongxing Wang 0001, Chaoqun Weng, Junsong Yuan 0001 |
CVPR | 1 |
| 2014 | Context-Aware Discovery of Visual Co-Occurrence PatternsabstractOnce an image is decomposed into a number of visual primitives, e.g., local interest points or regions, it is of great interests to discover meaningful visual patterns from them. Conventional clustering of visual primitives, however, usually ignores the spatial and feature structure among them, thus cannot discover high-level visual patterns of complex structure. To overcome this problem, we propose to consider spatial and feature contexts among visual primitives for pattern discovery. By discovering spatial co-occurrence patterns among visual primitives and feature co-occurrence patterns among different types of features, our method can better address the ambiguities of clustering visual primitives. We formulate the pattern discovery problem as a regularized k-means clustering where spatial and feature contexts are served as constraints to improve the pattern discovery results. A novel self-learning procedure is proposed to utilize the discovered spatial or feature patterns to gradually refine the clustering result. Our self-learning procedure is guaranteed to converge and experiments on real images validate the effectiveness of our method. Hongxing Wang 0001, Junsong Yuan 0001, Ying Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Learning weighted geometric pooling for image classificationabstractLocal feature extraction, coding, spatial pooling, and image classification are the four typical steps for state-of-the-art visual recognition systems. Unlike previous work that treats spatial pooling and image classification as separated steps, we propose to jointly learn the geometric pooling and image classifier such that class-specific geometric information of local descriptors can be incorporated to improve classification performance. Inspired by previous work of spatial pyramid matching and receptive field learning, we also propose spatial pyramid geometric pooling, receptive field geometric pooling and random partition geometric pooling approaches to further exploit the spatial structural information to boost classification performance. Experiments on 15-scene dataset validate the advantages of our proposed algorithms. Chaoqun Weng, Hongxing Wang 0001, Junsong Yuan 0001 |
ICIP | 2 |
| 2013 | Hierarchical sparse coding based on spatial pooling and multi-feature fusionabstractWe propose a novel hierarchical sparse coding algorithm with spatial pooling and multi-feature fusion, to construct the low-level visual primitives, e.g., local image patches or regions, into high-level visual phrases, e.g., image patterns. In the first layer we learn the sparse codes for the visual primitives and then pass them into the second layer by spatial pooling and multi-feature fusion. In the second layer we further learn the sparse codes for the visual phrases. In order to obtain the high-quality representations for visual phrases, our proposed algorithm iteratively optimizes over the two-layer sparse codes, as well as the two-layer codebooks. Since we have explored both the spatial and multi-feature contextual information, more representative sparse codes of the visual phrases can be obtained. The experiments on image pattern discovery, image scene clustering and image classification justify the advantages of the proposed algorithm. Chaoqun Weng, Hongxing Wang 0001, Junsong Yuan 0001 |
ICME | 2 |
| 2011 | Combining Feature Context and Spatial Context for Image Pattern DiscoveryabstractOnce an image is decomposed into a number of visual primitives, e.g., local interest points or salient image regions, it is of great interests to discover meaningful visual patterns from them. Conventional clustering (e.g., k-means) of visual primitives, however, usually ignores the spatial dependency among them, thus cannot discover the high-level visual patterns of complex spatial structure. To overcome this problem, we propose to consider both spatial and feature contexts among visual primitives for pattern discovery. By discovering both spatial co-occurrence patterns among visual primitives and feature co-occurrence patterns among different types of features, our method can better handle the ambiguities of visual primitives, by leveraging these co-occurrences. We formulate the problem as a regularized k-means clustering, and propose an iterative bottom-up/top-down self-learning procedure to gradually refine the result until it converges. The experiments of image text on discovery and image region clustering convince that combining spatial and feature contexts can significantly improve the pattern discovery results. Hongxing Wang 0001, Junsong Yuan 0001, Yap-Peng Tan |
ICDM | 1 |
| 2010 | Corner detection based on gradient correlation matrices of planar curves
Xiaohong Zhang 0002, Hongxing Wang 0001, Andrew W. B. Smith, Brian C. Lovell, Dan Yang 0001 |
Pattern Recognit. | 2 |
| 2009 | Robust image corner detection based on scale evolution difference of planar curves
Xiaohong Zhang 0002, Hongxing Wang 0001, Mingjian Hong, Dan Yang 0001, Brian C. Lovell |
Pattern Recognit. Lett. | 2 |