EDBT 2026 Demo / reviewers in the wild / expert
Jingyong Su
dblp:82/8615
· DBLP profile ↗
51ranked-venue papers
4as first author
37since 2021 · last 2026
0000-0003-3216-7027ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KAN or MLP? Point Cloud Shows the Way Forward
Qingdong He, Yijun Liu 0012, Jingyong Su |
ICMR | 4 |
| 2026 | Failure Detection in Image Segmentation Under Conditions of Semantic and Covariate ShiftsabstractWhether deep neural networks can provide reliable confidence is of great significance, especially in risk-sensitive scenarios. This work explores the impact of covariate and semantic shifts on segmentation tasks, an area which has not been extensively studied. Covariate shift refers to changes in the data distribution without alterations in the label space, while semantic shift involves changes in both data distribution and label space. We find that model-unknown distributional shifts in test data can transform an overconfidence problem into a situation of making random predictions with arbitrary confidence. The paper proposes a novel approach for effective failure detection that combines holistic image-level analysis and detailed pixel-level information. This approach involves the use of a Gray Level Co-occurrence Matrix (GLCM) to analyze the prediction randomness between adjacent pixels and a Magnitude-Direction Confidence Score Function (MD-CSF) for determining pixel acceptance or rejection. Furthermore, we introduce a new benchmark dataset, the Robot Inspection dataset for Semantic and Covariate shift in Segmentation (RISKS, the homophone of RISCS), to fill the need for datasets capable of evaluating the simultaneous impact of semantic and covariate shifts. Experimental results demonstrate that our method successfully detects image-level failures in segmentation, with MD-CSF outperforming other pluggable CSFs. The code and RISKS dataset will be available at https://github.com/liuyijungoon/MD-CSF. Yijun Liu 0012, Zhuotao Tian, Hang Zhao 0019, Zipeng Zhu, Jingyong Su |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Gradient of White Matter Functional Variability via fALFF Differential IdentifiabilityabstractFunctional variability in both gray matter (GM) and white matter (WM) is closely associated with human brain cognitive and developmental processes, and is commonly assessed using functional connectivity (FC). However, as a correlationbased approach, FC captures the co-fluctuation between brain regions rather than the intensity of neural activity in each region. Consequently, FC provides only a partial view of functional variability, and this limitation is particularly pronounced in WM, where functional signals are weaker and more susceptible to noise. To tackle this limitation, we introduce fractional amplitude of low-frequency fluctuation (fALFF) to measure the intensity of spontaneous neural activity and analyze functional variability in WM. Specifically, we propose a novel method to quantify WM functional variability by estimating the differential identifiability of fALFF. Higher differential identifiability is observed in WM fALFF compared to FC, which indicates that fALFF is more sensitive to WM functional variability. Through fALFF differential identifiability, we evaluate the functional variabilities of both WM and GM, and find the overall functional variability pattern is similar although WM shows slightly lower variability than GM. The regional functional variabilities of WM are associated with structural connectivity, where commissural fiber regions generally exhibit higher variability than projection fiber regions. Furthermore, we discover that WM functional variability demonstrates a spatial gradient ascending from the brainstem to the cortex by hypothesis testing, which aligns well with the evolutionary expansion. The gradient of functional variability in WM provides novel insights for understanding WM function. To the best of our knowledge, this is the first attempt to investigate WM functional variability via fALFF. Our code is available at https://github.com/Xinle-Chang/WM-fALFF-Idiff-Gradient. Xinle Chang, Yang Yang 0002, Yueran Li, Zhengcen Li, Haijin Zeng, Jingyong Su |
BIBM | 6 |
| 2025 | BrainCognizer: Brain Decoding with Human Visual Cognition Simulation for fMRI-to-Image ReconstructionabstractBrain decoding is a key neuroscience field that reconstructs the visual stimuli from brain activity with fMRI, which helps illuminate how the brain represents the world. fMRI-to-image reconstruction has achieved impressive progress by leveraging diffusion models. However, brain signals infused with prior knowledge and associations exhibit a significant information asymmetry when compared to raw visual features, still posing challenges for decoding fMRI representations under the supervision of images. Consequently, the reconstructed images often lack fine-grained visual fidelity, such as missing attributes and distorted spatial relationships. To tackle this challenge, we propose BrainCognizer, a novel brain decoding model inspired by human visual cognition, which explores multilevel semantics and correlations without fine-tuning of generative models. Specifically, BrainCognizer introduces two modules: the Cognitive Integration Module which incorporates prior human knowledge to extract hierarchical region semantics; and the Cognitive Correlation Module which captures contextual semantic relationships across regions. Incorporating these two modules enhances intra-region semantic consistency and maintains interregion contextual associations, thereby facilitating fine-grained brain decoding. Moreover, we quantitatively interpret our components from a neuroscience perspective and analyze the associations between different visual patterns and brain functions. Extensive quantitative and qualitative experiments demonstrate that BrainCognizer outperforms state-of-the-art approaches on multiple evaluation metrics. Our code is released publicly at https://github.com/Grace160/BrainCognizer. Guoying Sun, Weiyu Guo, Tong Shao, Yang Yang 0002, Haijin Zeng, Jingyong Su |
BIBM | 7 |
| 2025 | Vision-Language Gradient Descent-driven All-in-One Deep Unfolding NetworksabstractDynamic image degradations, including noise, blur and lighting inconsistencies, pose significant challenges in image restoration, often due to sensor limitations or adverse environmental conditions. Existing Deep Unfolding Networks (DUNs) offer stable restoration performance but require manual selection of degradation matrices for each degradation type, limiting their adaptability across diverse scenarios. To address this issue, we propose the Vision-Language-guided Unfolding Network (VLU-Net), a unified DUN framework for handling multiple degradation types simultaneously. VLU-Net leverages a VisionLanguage Model (VLM) refined on degraded image-text pairs to align image features with degradation descriptions, selecting the appropriate transform for target degradation. By integrating an automatic VLM-based gradient estimation strategy into the Proximal Gradient Descent (PGD) algorithm, VLU-Net effectively tackles complex multi-degradation restoration tasks while maintaining interpretability. Furthermore, we design a hierarchical feature unfolding structure to enhance VLU-Net framework, efficiently synthesizing degradation patterns across various levels. VLU-Net is the first all-in-one DUN framework and outperforms current leading one-by-one and all-in-one end- to-end methods by 3.74 dB on the SOTS dehazing dataset and 1.70 dB on the Rain100L deraining dataset. Haijin Zeng, Xiangming Wang, Yongyong Chen, Jingyong Su |
CVPR | 4 |
| 2025 | Binarized Mamba-Transformer for Lightweight Quad Bayer HybridEVS DemosaicingabstractQuad Bayer demosaicing is the central challenge for enabling the widespread application of Hybrid Event-based Vision Sensors (HybridEVS). Although existing learning-based methods that leverage long-range dependency modeling have achieved promising results, their complexity severely limits deployment on mobile devices for real-world applications. To address these limitations, we propose a lightweight Mamba-based binary neural network designed for efficient and high-performing demosaicing of HybridEVS RAW images. First, to effectively capture both global and local dependencies, we introduce a hybrid Binarized Mamba-Transformer architecture that combines the strengths of the Mamba and Swin Transformer architectures. Next, to significantly reduce computational complexity, we propose a binarized Mamba (Bi-Mamba), which binarizes all projections while retaining the core Selective Scan in full precision. Bi-Mamba also incorporates additional global visual information to enhance global context and mitigate precision loss. We conduct quantitative and qualitative experiments to demonstrate the effectiveness of BMTNet in both performance and computational efficiency, providing a lightweight demosaicing solution suited for real-world edge devices. Our codes and models are available at https://github.com/Clausy9/BMTNet. Haijin Zeng, Yunfan Lu, Tong Shao, Yongyong Chen, Jingyong Su |
CVPR | 8 |
| 2025 | C2AD: Dual Consistency Learning for Zero-Shot Anomaly DetectionabstractZero-shot anomaly detection (ZSAD) is dedicated to detecting anomalies without having any seen normal or abnormal samples for the target set. Existing approaches utilize the pre-trained CLIP to assess normality/abnormality by exploiting the similarity between images and text with the frozen visual encoder. However, the frozen CLIP visual encoder impedes performance improvements. Additionally, their representations of anomalies are sensitive to contextual variations, leading to poor localization of unseen abnormalities. Therefore, this paper introduces the Dual Consistency Learning for Zero-Shot Anomaly Detection (C2AD), comprising two components: semantic and contextual consistency. Semantic consistency enhances generalization by maintaining correlational semantic consistency, while contextual consistency encourages representations to be robust to contextual changes. C2AD improves the model training without adding extra computational overhead during inference. Comprehensive experiments demonstrate that C2AD can boost the performance of ZSAD in anomaly detection and localization, achieving state-of-the-art results. Ruilong Xing, Zhuotao Tian, Yijun Liu 0012, Senqiao Yang, Jingyong Su |
ICASSP | 7 |
| 2025 | Spectral Compressive Imaging via Unmixing-driven Subspace Diffusion RefinementabstractSpectral Compressive Imaging (SCI) reconstruction is inherently ill-posed because a single observation admits multiple plausible reconstructions. Traditional deterministic methods struggle to effectively recover high-frequency details. Although diffusion models offer promising solutions to this challenge, their application is constrained by the limited training data and high computational demands associated with multispectral images (MSIs), making direct diffusion training impractical. To address these issues, we propose a novel Predict-and-unmixing-driven-Subspace-Refine framework (PSR-SCI). This framework begins with a light-weight predictor that produces an initial, rough estimate of the MSI. Subsequently, we introduce a unmixing-driven reversible spectral embedding module that decomposes the MSI into subspace images and spectral coefficients. This compact representation facilitates the adaptation of pre-trained RGB diffusion models and focuses refinement processes on high-frequency details, thereby enabling efficient diffusion generation with minimal MSI data. Additionally, we design a high-dimensional guidance mechanism enforcing SCI consistency during sampling. The refined subspace image is then reconstructed back into an MSI using the reversible embedding, yielding the final MSI with full spectral resolution. Experimental results on the standard KAIST and zero-shot datasets NTIRE, ICVL, and Harvard show that PSR-SCI enhances overall visual quality and delivers PSNR and SSIM results competitive with state-of-the-art diffusion, transformer, and deep-unfolding baselines. This framework provides a robust alternative to traditional deterministic SCI reconstruction methods. Code and models are available at [https://github.com/SMARK2022/PSR-SCI](https://github.com/SMARK2022/PSR-SCI). Haijin Zeng, Benteng Sun, Yongyong Chen, Jingyong Su, Yong Xu 0001 |
ICLR | 4 |
| 2025 | Brain Activation Mapping Based on Regional Synchronization of fMRI Signals Embedded in Graph Eigenmodes
Yang Yang 0002, Jingyong Su |
MICCAI (12) | 5 |
| 2025 | Rectifying Soft-Label Entangled Bias in Long-Tailed Dataset DistillationabstractDataset distillation compresses large-scale datasets into compact, highly informative synthetic data, significantly reducing storage and training costs. However, existing research primarily focuses on balanced datasets and struggles to perform under real-world long-tailed distributions. In this work, we emphasize the critical role of soft labels in long-tailed dataset distillation and uncover the underlying mechanisms contributing to performance degradation.
Specifically, we derive an imbalance-aware generalization bound for model trained on distilled dataset. We then identify two primary sources of soft-label bias, which originate from the distillation model and the distilled images, through systematic perturbation of the data imbalance levels.
To address this, we propose ADSA, an Adaptive Soft-label Alignment module that calibrates the entangled biases. This lightweight module integrates seamlessly into existing distillation pipelines and consistently improves performance. On ImageNet-1k-LT with EDC and IPC=50, ADSA improves tail-class accuracy by up to 11.8\% and raises overall accuracy to 41.4\%.
Extensive experiments demonstrate that ADSA provides a robust and generalizable solution under limited label budgets and across a range of distillation techniques. Chenyang Jiang 0001, Hang Zhao 0021, Zhengcen Li, Qiben Shan, Shaocong Wu, Jingyong Su |
NeurIPS | 7 |
| 2025 | Semi-supervised medical image segmentation via weak-to-strong perturbation consistency and edge-aware contrastive representation
Yang Yang 0002, Guoying Sun, Tong Zhang 0017, Jingyong Su |
Medical Image Anal. | 5 |
| 2025 | Transformer-based material recognition via short-time contact sensing
Zhenyang Liu, Yitian Shao, Qiliang Li, Jingyong Su |
Pattern Recognit. | 4 |
| 2025 | BFRA: A Bi-Level Feature Relation Alignment Method for Cross-Domain Few-Shot LearningabstractWhile existing Few-Shot Learning (FSL) techniques demonstrate strong performance on uniform datasets, they encounter domain shift challenges when presented with domainagnostic queries in real-world scenarios. So we investigate it in Cross-Domain Few-Shot Learning (CD-FSL) and propose to learn more universal feature representations to enhance generalization on unseen domains. Toward this issue, we pinpoint two issues in current multi-model fusion approaches: 1) the entanglement of domain and class information, and 2) feature overlap across distinct domains. To address these challenges, we introduce a Bi-level Feature Relation Alignment method, BFRA, which facilitates the acquisition of more versatile features by decoupling domain-class relationships and aligning feature relations. Through the segregation of domain and class feature learning, we devise a smoothing layer prior for domain feature alignment to mitigate inter-domain discrepancies. This approach enables our model to acquire domain-consistent features, diminishing interference in subsequent class feature alignment procedures. During the class feature alignment, we notice that class feature representations from various in-domain models may intersect, leading to a diminished distinction between classes. To address this, we adopt a topological perspective to train our target model, by aligning feature relations instead of features between our target model and multiple in-domain models. The integration of these components results in the establishment of a bi-level feature relation alignment framework aimed at acquiring more universal features. Furthermore, we partially fine-tune the plug-in layer-wise affine adapter on domain-agnostic queries to expedite adaptation without impacting the known domains. Experiments of 21 datasets on meta-dataset and BSCD-FSL benchmark demonstrate the effectiveness of our method. The code are made publicly available at https://github.com/leaves162/BFRA. Tong Shao, Zhuotao Tian, Jingyong Su |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Smooth Tensor Qatar Riyal Decomposition for Dynamic MRI ReconstructionabstractDynamic magnetic resonance imaging (dMRI) speed and imaging quality have always been a crucial issue in medical imaging research. Most existing methods characterize the tensor rank-based minimization to reconstruct dMRI from sampling $\bf k$-$t$ space data. However, (1) these approaches that unfold the tensor along each dimension destroy the inherent structure of dMR images. (2) they focus on preserving global information only, while ignoring the local details reconstruction such as the spatial piece-wise smoothness and sharp boundaries. To overcome these obstacles, we suggest a novel low-rank tensor decomposition approach by integrating tensor Qatar Riyal (QR) decomposition, low-rank tensor nuclear norm, and asymmetric total variation to reconstruct dMRI, named TQRTV. Specifically, while preserving the tensor inherent structure by utilizing tensor nuclear norm minimization to approximate tensor rank, QR decomposition reduces the dimensions in the low-rank constraint term, thereby improving the reconstruction performance. TQRTV further exploits the asymmetric total variation regularizer to capture local details. Numerical experiments demonstrate that the proposed reconstruction approach is superior to the existing ones. Yongyong Chen, Haijin Zeng, Jingyong Su |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Boundary-Guided Contrastive Learning for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning methods, compared to fully supervised learning, offer significant potential to alleviate the burden of manual annotations on clinicians. By leveraging unlabeled data, these methods can aid in the development of medical image segmentation systems for improving efficiency. Boundary segmentation is crucial in medical image analysis. However, accurate segmentation of boundary regions is under-explored in existing methods since boundary pixels constitute only a small fraction of the overall image, resulting in suboptimal segmentation performance for boundary regions. In this paper, we introduce boundary-guided contrastive learning for semi-supervised medical image segmentation (BoCLIS). Specifically, we first propose conservative-to-radical teacher networks with an uncertainty-weighted aggregation strategy to generate higher quality pseudo-labels, enabling more efficient utilization of unlabeled data. To further improve the performance of segmentation in boundary regions, we propose a boundary-guided patch sampling strategy to guide the framework in learning discriminative representations for these regions. Lastly, the patch-based contrastive learning is proposed to simultaneously compute the (dis)similarities of the discriminative representations across intra- and inter-images. Extensive experiments on three public datasets show that our method consistently outperforms existing methods, especially in the boundary region, with DSC improvements of 20.47%, 16.75%, and 17.18%, respectively. A comprehensive analysis is further performed to demonstrate the effectiveness of our approach. Our code is released publicly at https://github.com/youngyzzZ/BoCLIS. Yang Yang 0002, Jiaxin Zhuang, Guoying Sun, Jingyong Su |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Feature Distribution Matching by Optimal Transport for Effective and Robust Coreset SelectionabstractTraining neural networks with good generalization requires large computational costs in many deep learning methods due to large-scale datasets and over-parameterized models. Despite the emergence of a number of coreset selection methods to reduce the computational costs, the problem of coreset distribution bias, i.e., the skewed distribution between the coreset and the entire dataset, has not been well studied. In this paper, we find that the closer the feature distribution of the coreset is to that of the entire dataset, the better the generalization performance of the coreset, particularly under extreme pruning. This motivates us to propose a simple yet effective method for coreset selection to alleviate the distribution bias between the coreset and the entire dataset, called feature distribution matching (FDMat). Unlike gradient-based methods, which selects samples with larger gradient values or approximates gradient values of the entire dataset, FDMat aims to select coreset that is closest to feature distribution of the entire dataset. Specifically, FDMat transfers coreset selection as an optimal transport problem from the coreset to the entire dataset in feature embedding spaces. Moreover, our method shows strong robustness due to the removal of samples far from the distribution, especially for the entire dataset containing noisy and class-imbalanced samples. Extensive experiments on multiple benchmarks show that FDMat can improve the performance of coreset selection than existing coreset methods. The code is available at https://github.com/successhaha/FDMat. Weiwei Xiao, Yongyong Chen, Qiben Shan, Jingyong Su |
AAAI | 5 |
| 2024 | Skeleton-Based Group Activity Recognition via Spatial-Temporal Panoramic Graph
Zhengcen Li, Xinle Chang, Yueran Li, Jingyong Su |
ECCV (59) | 4 |
| 2024 | Explore the Potential of CLIP for Training-Free Open Vocabulary Semantic Segmentation
Tong Shao, Zhuotao Tian, Hang Zhao 0019, Jingyong Su |
ECCV (86) | 4 |
| 2024 | SAH-SCI: Self-supervised Adapter for Efficient Hyperspectral Snapshot Compressive Imaging
Haijin Zeng, Yongyong Chen, Youfa Liu, Chong Peng 0001, Jingyong Su |
ECCV (64) | 6 |
| 2024 | Typicalness-Aware Learning for Failure DetectionabstractDeep neural networks (DNNs) often suffer from the overconfidence issue, where incorrect predictions are made with high confidence scores, hindering the applications in critical systems. In this paper, we propose a novel approach called Typicalness-Aware Learning (TAL) to address this issue and improve failure detection performance.
We observe that, with the cross-entropy loss, model predictions are optimized to align with the corresponding labels via increasing logit magnitude or refining logit direction. However, regarding atypical samples, the image content and their labels may exhibit disparities. This discrepancy can lead to overfitting on atypical samples, ultimately resulting in the overconfidence issue that we aim to address.
To address this issue, we have devised a metric that quantifies the typicalness of each sample, enabling the dynamic adjustment of the logit magnitude during the training process. By allowing relatively atypical samples to be adequately fitted while preserving reliable logit direction, the problem of overconfidence can be mitigated. TAL has been extensively evaluated on benchmark datasets, and the results demonstrate its superiority over existing failure detection methods. Specifically, TAL achieves a more than 5\% improvement on CIFAR100 in terms of the Area Under the Risk-Coverage Curve (AURC) compared to the state-of-the-art. Code is available at https://github.com/liuyijungoon/TAL. Yijun Liu 0012, Jiequan Cui, Zhuotao Tian, Senqiao Yang, Qingdong He, Jingyong Su |
NeurIPS | 7 |
| 2024 | Class-Agnostic Detection of Unknown Objects from Foreground Improves Robust Open World Object Detection
Yongyong Chen, Zimu Zheng, Jingyong Su |
PRCV (12) | 6 |
| 2024 | Tensor Learning Meets Dynamic Anchor Learning: From Complete to Incomplete Multiview ClusteringabstractMultiview clustering (MVC), which can dexterously uncover the underlying intrinsic clustering structures of the data, has been particularly attractive in recent years. However, previous methods are designed for either complete or incomplete multiview only, without a unified framework that handles both tasks simultaneously. To address this issue, we propose a unified framework to efficiently tackle both tasks in approximately linear complexity, which integrates tensor learning to explore the inter-view low-rankness and dynamic anchor learning to explore the intra-view low-rankness for scalable clustering (TDASC). Specifically, TDASC efficiently learns smaller view-specific graphs by anchor learning, which not only explores the diversity embedded in multiview data, but also yields approximately linear complexity. Meanwhile, unlike most current approaches that only focus on pair-wise relationships, the proposed TDASC incorporates multiple graphs into an inter-view low-rank tensor, which elegantly models the high-order correlations across views and further guides the anchor learning. Extensive experiments on both complete and incomplete multiview datasets clearly demonstrate the effectiveness and efficiency of TDASC compared with several state-of-the-art techniques. Yongyong Chen, Xiaojia Zhao, Zheng Zhang 0006, Youfa Liu, Jingyong Su, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Double High-Order Correlation Preserved Robust Multi-View Ensemble ClusteringabstractEnsemble clustering (EC), utilizing multiple basic partitions (BPs) to yield a robust consensus clustering, has shown promising clustering performance. Nevertheless, most current algorithms suffer from two challenging hurdles: (1) a surge of EC-based methods only focus on pair-wise sample correlation while fully ignoring the high-order correlations of diverse views. (2) they deal directly with the co-association (CA) matrices generated from BPs, which are inevitably corrupted by noise and thus degrade the clustering performance. To address these issues, we propose a novel Double High-Order Correlation Preserved Robust Multi-View Ensemble Clustering (DC-RMEC) method, which preserves the high-order inter-view correlation and the high-order correlation of original data simultaneously. Specifically, DC-RMEC constructs a hypergraph from BPs to fuse high-level complementary information from different algorithms and incorporates multiple CA-based representations into a low-rank tensor to discover the high-order relevance underlying CA matrices, such that double high-order correlation of multi-view features could be dexterously uncovered. Moreover, a marginalized denoiser is invoked to gain robust view-specific CA matrices. Furthermore, we develop a unified framework to jointly optimize the representation tensor and the result matrix. An effective iterative optimization algorithm is designed to optimize our DC-RMEC model by resorting to the alternating direction method of multipliers. Extensive experiments on seven real-world multi-view datasets have demonstrated the superiority of DC-RMEC compared with several state-of-the-art multi-view ensemble clustering methods. Xiaojia Zhao, Qiangqiang Shen, Youfa Liu, Yongyong Chen, Jingyong Su |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Semi-supervised Medical Image Segmentation via Feature-perturbed ConsistencyabstractAlthough deep convolutional neural networks have achieved satisfactory performance in many medical image segmentation tasks, a considerable annotation challenge still needs to be solved, which is expensive and time-consuming for radiologists. Most existing popular semi-supervised methods mainly impose data-level perturbations (e.g., rotation, noising) or feature-level perturbations (e.g., MC dropout) on unlabeled data. In this paper, we propose a novel semi-supervised segmentation strategy with meaningful perturbations at the feature level to leverage abundant useful information naturally embedded in the unlabeled data. Specifically, we develop a dual-task network where the segmentation head produces multiple predictions with a perturbation module, and the reconstruction head further utilizes the semantic information to enhance segmentation performance. The proposed framework subtly perturbs the network at the feature-level to generate predictions which should be similar and consistent. However, enforcing them roughly to be consistent at all pixels harms stable training and neglects much delicate information. To better utilize those predictions and estimate the uncertainty, we further propose feature-perturbed consistency to exploit reliable regions for our framework to learn from. Extensive experiments on the public BraTS2020 dataset and the 2017 ACDC dataset confirm the efficiency and effectiveness of our method. In particular, the proposed method demonstrates remarkable superiority in the segmentation of boundary regions. The project is available at https://github.com/youngyzzZ/SFPC. Yang Yang 0002, Tong Zhang 0017, Jingyong Su |
BIBM | 4 |
| 2023 | Hierarchical Dense Correlation Distillation for Few-Shot SegmentationabstractFew-shot semantic segmentation (FSS) aims to form class-agnostic models segmenting unseen classes with only a handful of annotations. Previous methods limited to the semantic feature and prototype representation suffer from coarse segmentation granularity and train-set overfitting. In this work, we design Hierarchically Decoupled Matching Network (HDMNet) mining pixel-level support correlation based on the transformer architecture. The self-attention modules are used to assist in establishing hierarchical dense features, as a means to accomplish the cascade matching between query and support features. Moreover, we propose a matching module to reduce train-set overfitting and introduce correlation distillation leveraging semantic correspondence from coarse resolution to boost fine-grained segmentation. Our method performs decently in experiments. We achieve 50.0% mIoU on COCO-20idataset one-shot setting and 56.0% on five-shot segmentation, respectively. The code is available on the project website11https://github.com/Pbihao/HDMNet. Bohao Peng, Zhuotao Tian, Xiaoyang Wu 0002, Chengyao Wang, Shu Liu 0005, Jingyong Su, Jiaya Jia |
CVPR | 6 |
| 2023 | Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object PredictionabstractIn this paper, we propose a new paradigm, named Historical Object Prediction (HoP) for multi-view 3D detection to leverage temporal information more effectively. The HoP approach is straightforward: given the current times-tamp t, we generate a pseudo Bird’s-Eye View (BEV) feature of timestamp t-k from its adjacent frames and utilize this feature to predict the object set at timestamp t-k. Our approach is motivated by the observation that enforcing the detector to capture both the spatial location and temporal motion of objects occurring at historical timestamps can lead to more accurate BEV feature learning. First, we elaborately design short-term and long-term temporal decoders, which can generate the pseudo BEV feature for timestamp t-k without the involvement of its corresponding camera images. Second, an additional object decoder is flexibly attached to predict the object targets using the generated pseudo BEV feature. Note that we only perform HoP during training, thus the proposed method does not introduce extra overheads during inference. As a plug-and-play approach, HoP can be easily incorporated into state-of-the-art BEV detection frameworks, including BEVFormer and BEVDet series. Furthermore, the auxiliary HoP approach is complementary to prevalent temporal modeling methods, leading to significant performance gains. Extensive experiments are conducted to evaluate the effectiveness of the proposed HoP on the nuScenes dataset. We choose the representative methods, including BEVFormer and BEVDet4D-Depth to evaluate our method. Surprisingly, HoP achieves 68.5% NDS and 62.4% mAP with ViT-L on nuScenes test, outperforming all the 3D object detectors on the leaderboard. Codes are available at https://github.com/Sense-X/HoP. Zhuofan Zong, Dongzhi Jiang, Guanglu Song, Zeyue Xue, Jingyong Su, Hongsheng Li 0001, Yu Liu 0015 |
ICCV | 5 |
| 2023 | Two-Person Graph Convolutional Network for Skeleton-Based Human Interaction RecognitionabstractGraph convolutional networks (GCNs) have been the predominant methods in skeleton-based human action recognition, including human-human interaction recognition. However, when dealing with interaction sequences, current GCN-based methods simply split the two-person skeleton into two discrete graphs and perform graph convolution separately as done for single-person action classification. Such operations ignore rich interactive information and hinder effective spatial inter-body relationship modeling. To overcome the above shortcoming, we introduce a novel unified two-person graph to represent inter-body and intra-body correlations between joints. Experimental results show accuracy improvements in recognizing both interactions and individual actions when utilizing the proposed two-person graph topology. In addition, several graph labeling strategies are designed to supervise the model to learn discriminant spatial-temporal interactive features. Finally, we propose a two-person graph convolutional network (2P-GCN). Our model outperforms state-of-the-art methods on four benchmarks of three interaction datasets: SBU, interaction subsets of NTU-RGB+D and NTU-RGB+D 120. Zhengcen Li, Yueran Li, Linlin Tang, Tong Zhang 0017, Jingyong Su |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Cross-Scale-Guided Fusion Transformer for Disaster Assessment Using Satellite ImageryabstractWhen a disaster strikes, accurate disaster information and effective response are critical for saving lives and properties. High-resolution satellite (HRS) imagery provides valuable geographical information that can assist experts in analyzing damage levels in different areas and enacting appropriate relief plans. However, analyzing large HRS images is both time-consuming and inefficient, requiring efficient automated methods to replace expert analysis. Fortunately, deep learning methods have achieved impressive performance on HRS image processing tasks, considerably increasing automation levels. Despite this progress, most HRS-based damage assessment methods only consider a single time series of post-disaster images or simply integrate pre- and post-disaster images, lacking the integration of effective information between pre- and post-disaster images. To alleviate this problem, we propose a two-stage multi-scale fusion network that fully exploits the information contained in pre- and post-disaster images. Specifically, we employ a hierarchical Transformer to accurately locate buildings by pre-disaster images in the first stage, and then propose the guided fusion and cross-scale guided fusion modules in the second stage to efficiently utilize both pre- and post-disaster images. Our method outperforms state-of-the-art methods in building segmentation and building damage assessment on the xBD dataset, and exhibits improved generalization across diverse geographic regions and disaster types. Weiwei Xiao, Jingyong Su, Yongyong Chen, Guofeng Cao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Multi-Attention Feature Fusion Network for Accurate Estimation of Finger Kinematics From Surface Electromyographic SignalsabstractSimultaneous and proportional control (SPC) based on surface electromyographic (sEMG) signals has led to a broad range of applications. However, due to the limitation in the generalization and stability of current machine learning algorithms, these methods can only estimate less than 15 simultaneuous and proportional (SP) categories of finger movement. In this article, a novel deep learning algorithm, named multiattention feature fusion network (MAFN), is proposed to estimate comprehensive finger movement (up to 28 categories SP movements) from sEMG signals. MAFN is based on the multihead attention mechanism, which adaptively extracts essential features for analyzing the joint angles from the extracted sEMG features. Furthermore, a real-time exponential smoothing algorithm is designed for further improvement of the prediction stability. MAFN was evaluated on 28 finger movements of 38 subjects in the Ninapro_db2 dataset, and benchmarked with the state-of-the-art methods, such as temporal convolutional network (TCN) and long short term memory network (LSTM). The results demonstrated that the average Pearson correlation coefficient, root mean squared error of MAFN (0.84 ± 0.03,0.09 ± 0.01) were significantly higher than those of TCN (0.52 ± 0.06,pppp< 0.001). These improvements led to more stable and accurate movement predictions. Additionally, the time delay and power consumption of MAFN when applied to sEMG signals on a portable device are only 83.4 ms and 3 W, which implies prospective commercial applications. Weiyu Guo, Ning Jiang 0001, Dario Farina, Jingyong Su, Zheng Wang 0027, Chuang Lin 0001, Hui Xiong 0001 |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2023 | Multi-view Ensemble Clustering via Low-rank and Sparse Decomposition: From Matrix to TensorabstractAs a significant extension of classical clustering methods, ensemble clustering first generates multiple basic clusterings and then fuses them into one consensus partition by solving a problem concerning graph partition with respect to the co-association matrix. Although the collaborative cluster structure among basic clusterings can be well discovered by ensemble clustering, most advanced ensemble clustering utilizes the self-representation strategy with the constraint of low-rank to explore a shared consensus representation matrix in multiple views. However, they still encounter two challenges: (1) high computational cost caused by both the matrix inversion operation and singular value decomposition of large-scale square matrices; (2) less considerable attention on high-order correlation attributed to the pursue of the two-dimensional pair-wise relationship matrix. In this article, based on low-rank and sparse decomposition from both matrix and tensor perspectives, we propose two novel multi-view ensemble clustering methods, which tangibly decrease computational complexity. Specifically, our first method utilizes low-rank and sparse matrix decomposition to learn one common co-association matrix, while our last method constructs all co-association matrices into one third-order tensor to investigate the high-order correlation among multiple views by low-rank and sparse tensor decomposition. We adopt the alternating direction method of multipliers to solve two convex models by dividing them into several subproblems with closed-form solution. Experimental results on ten real-world datasets prove the effectiveness and efficiency of the proposed two multi-view ensemble clustering methods by comparing them with other advanced ensemble clustering methods. Xuanqi Zhang, Qiangqiang Shen, Yongyong Chen, Zhongyun Hua, Jingyong Su |
ACM Trans. Knowl. Discov. Data | 6 |
| 2023 | Optimizing Spaced Repetition Schedule by Capturing the Dynamics of MemoryabstractSpaced repetition, namely, learners review items in a given schedule, has been proven powerful for memorization and practice of skills. Most current spaced repetition methods focus on either predicting student recall or designing an optimal review schedule, thus omitting the integrity of the spaced repetition system. In this work, we propose a novel spaced repetition schedule framework by capturing the dynamics of memory, which alternates memory prediction and schedule optimization to improve the efficiency of learners’ reviews. First, the framework collects logs from students’ reviews and builds memory models with Markov property to capture the dynamics of memory. Then, the spaced repetition optimization is transformed a stochastic shortest path problem and solved via the value iteration method. We also construct a new benchmark dataset for spaced repetition, which is the first to contain time-series information during learners’ memorization. Experimental results on the collected data from the real world and the simulated environment demonstrate that the proposed approach reduces 64% error and 17% cost in predicting recall rates and optimizing schedules compared to several baselines. We have publicly released the dataset containing 220 million rows and codes used in this paper at:https://github.com/maimemo/SSP-MMC-Plus. Jingyong Su, Junyao Ye, Liqiang Nie, Yilong Cao, Yongyong Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | A Riemannian Framework for Structurally Curated Functional Clustering of Brain White Matter FibersabstractWhite matter (WM) consists of fibers that transmit information from one brain region to another, and functional fiber clustering that combines diffusion and functional MRI provides a novel perspective for exploring the functional architecture of axonal fibers. However, existing methods only concern functional signals in gray matter (GM), whereas the connecting fibers may not transmit relevant functional signals. There has been growing evidence that neural activity is encoded in WM BOLD signals as well, which provides rich multimodal information for fiber clustering. In this paper, we develop a comprehensive Riemannian framework for functional fiber clustering using WM BOLD signals along fibers. Specifically, we derive a novel metric that is highly discriminative of different functional classes while reducing the variability within classes and, in the meantime, enables low-dimensional coding of high-dimensional data. Our in vivo experiments show that the proposed framework is able to achieve clustering results with inter-subject consistency and functional homogeneity. In addition, we develop an atlas of WM functional architecture for standardizable yet flexible use and exemplify a machine-learning-based application for the classification of autism spectrum disorders, which further demonstrates the great potential of our approach in practical applications. Yi Zhao 0018, Zhaohua Ding, Jingyong Su |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Correntropy-Induced Tensor Learning for Multi-view Subspace ClusteringabstractUsing some specific optimization problems with specific regularizers, multi-view subspace clustering has achieved better performance over single-view subspace clustering. However, they simply assume the noise obeys the Gaussian distribution only, and thus the dataset with non-Gaussian noise or outliers may not be accurately clustered. To address this issue, this paper proposes a novel correntropy-induced tensor learning method for multi-view subspace clustering (CTMSC). Specifically, CTMSC adopts the correntropy-induced metric to substitute the traditional mean square error (MSE) to handle non-Gaussian noise or outliers. Furthermore, the proposed objective function is optimized using an alternating direction method of multipliers with the aid of half-quadratic technology in the form of multiplication. Extensive experimental results on various real-world datasets demonstrate the effectiveness of the proposed method by comparing several state-of-the-art multi-view subspace clustering methods. Yongyong Chen, Shuqin Wang 0001, Jingyong Su, Junxin Chen 0001 |
ICDM | 3 |
| 2022 | A Stochastic Shortest Path Algorithm for Optimizing Spaced Repetition SchedulingabstractSpaced repetition is a mnemonic technique where long-term memory can be efficiently formed by following review schedules. For greater memorization efficiency, spaced repetition schedulers need to model students' long-term memory and optimize the review cost. We have collected 220 million students' memory behavior logs with time-series features and built a memory model with Markov property. Based on the model, we design a spaced repetition scheduler guaranteed to minimize the review cost by a stochastic shortest path algorithm. Experimental results have shown a 12.6% performance improvement over the state-of-the-art methods. The scheduler has been successfully deployed in the online language-learning app MaiMemo to help millions of students. Junyao Ye, Jingyong Su, Yilong Cao |
KDD | 2 |
| 2022 | Statistical analysis of the community lockdown for COVID-19 pandemicabstractAs the global pandemic of the COVID-19 continues, the statistical modeling and analysis of the spreading process of COVID-19 have attracted widespread attention. Various propagation simulation models have been proposed to predict the spread of the epidemic and the effectiveness of related control measures. These models play an indispensable role in understanding the complex dynamic situation of the epidemic. Most existing work studies the spread of epidemic at two levels including population and agent. However, there is no comprehensive statistical analysis of community lockdown measures and corresponding control effects. This paper performs a statistical analysis of the effectiveness of community lockdown based on the Agent-Level Pandemic Simulation (ALPS) model. We propose a statistical model to analyze multiple variables affecting the COVID-19 pandemic, which include the timings of implementing and lifting lockdown, the crowd mobility, and other factors. Specifically, a motion model followed by ALPS and related basic assumptions is discussed first. Then the model has been evaluated using the real data of COVID-19. The simulation study and comparison with real data have validated the effectiveness of our model. Shaocong Wu, Xiaolong Wang 0001, Jingyong Su |
Appl. Intell. | 3 |
| 2022 | Transformer-based Cross Reference Network for video salient object detection
Kan Huang, Chunwei Tian, Jingyong Su, Jerry Chun-Wei Lin |
Pattern Recognit. Lett. | 3 |
| 2021 | MVDRNet: Multi-view diabetic retinopathy detection by combining DCNNs and attention mechanisms
Xiaoling Luo 0001, Zuhui Pu, Yong Xu 0001, Wai Keung Wong, Jingyong Su, Xiaoyan Dou, Baikang Ye, Jiying Hu, Lisha Mou |
Pattern Recognit. | 5 |
| 2020 | A Riemannian Framework for Detecting Stimulus-Relevant Fiber PathwaysabstractFunctional MRI based on blood oxygenation level-dependent (BOLD) contrast is well established as a neuroimaging technique for detecting neural activity in the cortex of the human brain. Recent studies have shown that variations of BOLD signals in white matter are also related to neural activities both in resting state and under functional loading. We develop a comprehensive framework of detecting task-specific fiber pathways. We not only study fiber tracts as open curves with different physical features (shape, scale, orientation and position), but also incorporate the BOLD signals associated with them to find stimulus-relevant pathways. Specifically, we propose a novel Riemannian metric, which is a weighted sum of distances in product space of shapes and functions. This metric provides both a cost function for registration and a proper distance for comparison. Experimental results on real data have shown that we can cluster fiber pathways correctly by evaluating correlations between BOLD signals and stimuli, temporal variations and power spectra of them. Mengmeng Guo, Jingyong Su, Linlin Tang, Zhaohua Ding |
ICPR | 2 |
| 2019 | A distance weighted linear regression classifier based on optimized distance calculating approach for face recognition
Linlin Tang, Huifen Lu, Zhen Pang, Zhangyan Li, Jingyong Su |
Multim. Tools Appl. | 5 |
| 2019 | Kernel nearest-farthest subspace classifier for face recognition
Linlin Tang, Zuohua Li, Jingyong Su, Huifen Lu, Zhangyan Li, Zhen Pang, Yong Zhang 0023 |
Multim. Tools Appl. | 3 |
| 2017 | Spatially Coherent Interpretations of Videos Using Pattern Theory
Fillipe D. M. de Souza, Sudeep Sarkar, Anuj Srivastava, Jingyong Su |
Int. J. Comput. Vis. | 4 |
| 2017 | Elastic Functional Coding of Riemannian TrajectoriesabstractVisual observations of dynamic phenomena, such as human actions, are often represented as sequences of smoothly-varying features. In cases where the feature spaces can be structured as Riemannian manifolds, the corresponding representations become trajectories on manifolds. Analysis of these trajectories is challenging due to non-linearity of underlying spaces and high-dimensionality of trajectories. In vision problems, given the nature of physical systems involved, these phenomena are better characterized on a low-dimensional manifold compared to the space of Riemannian trajectories. For instance, if one does not impose physical constraints of the human body, in data involving human action analysis, the resulting representation space will have highly redundant features. Learning an effective, low-dimensional embedding for action representations will have a huge impact in the areas of search and retrieval, visualization, learning, and recognition. Traditional manifold learning addresses this problem for static points in the euclidean space, but its extension to Riemannian trajectories is non-trivial and remains unexplored. The difficulty lies in inherent non-linearity of the domain and temporal variability of actions that can distort any traditional metric between trajectories. To overcome these issues, we use the framework based on transported square-root velocity fields (TSRVF); this framework has several desirable properties, including a rate-invariant metric and vector space representations. We propose to learn an embedding such that each action trajectory is mapped to a single point in a low-dimensional euclidean space, and the trajectories that differ only in temporal rates map to the same point. We utilize the TSRVF representation, and accompanying statistical summaries of Riemannian trajectories, to extend existing coding methods such as PCA, KSVD and Label Consistent KSVD to Riemannian trajectories or more generally to Riemannian functions. We show that such coding efficiently captures trajectories in applications such as action recognition, stroke rehabilitation, visual speech recognition, clustering and diverse sequence sampling. Using this framework, we obtain state-of-the-art recognition results, while reducing the dimensionality/ complexity by a factor of 100-250x. Since these mappings and codes are invertible, they can also be used to interactively-visualize Riemannian trajectories and synthesize actions. Rushil Anirudh, Pavan Turaga, Jingyong Su, Anuj Srivastava |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Action Recognition Using Rate-Invariant Analysis of Skeletal Shape TrajectoriesabstractWe study the problem of classifying actions of human subjects using depth movies generated by Kinect or other depth sensors. Representing human body as dynamical skeletons, we study the evolution of their (skeletons’) shapes as trajectories on Kendall’s shape manifold. The action data is typically corrupted by large variability in execution rates within and across subjects and, thus, causing major problems in statistical analyses. To address that issue, we adopt a recently-developed framework of Su et al. [1], [2] to this problem domain. Here, the variable execution rates correspond to re-parameterizations of trajectories, and one uses a parameterization-invariant metric for aligning, comparing, averaging, and modeling trajectories. This is based on a combination of transported square-root vector fields (TSRVFs) of trajectories and the standard Euclidean norm, that allows computational efficiency. We develop a comprehensive suite of computational tools for this application domain: smoothing and denoising skeleton trajectories using median filtering, up- and down-sampling actions in time domain, simultaneous temporal-registration of multiple actions, and extracting invertible Euclidean representations of actions. Due to invertibility these Euclidean representations allow both discriminative and generative models for statistical analysis. For instance, they can be used in a SVM-based classification of original actions, as demonstrated here using MSR Action-3D, MSR Daily Activity and 3D Action Pairs datasets. Using only the skeletal information, we achieve state-of-the-art classification results on these datasets. Boulbaba Ben Amor, Jingyong Su, Anuj Srivastava |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Pattern theory for representation and inference of semantic structures in videos
Fillipe D. M. de Souza, Sudeep Sarkar, Anuj Srivastava, Jingyong Su |
Pattern Recognit. Lett. | 4 |
| 2015 | Elastic functional coding of human actions: From vector-fields to latent variablesabstractHuman activities observed from visual sensors often give rise to a sequence of smoothly varying features. In many cases, the space of features can be formally defined as a manifold, where the action becomes a trajectory on the manifold. Such trajectories are high dimensional in addition to being non-linear, which can severely limit computations on them. We also argue that by their nature, human actions themselves lie on a much lower dimensional manifold compared to the high dimensional feature space. Learning an accurate low dimensional embedding for actions could have a huge impact in the areas of efficient search and retrieval, visualization, learning, and recognition. Traditional manifold learning addresses this problem for static points in ℝn, but its extension to trajectories on Riemannian manifolds is non-trivial and has remained unexplored. The challenge arises due to the inherent non-linearity, and temporal variability that can significantly distort the distance metric between trajectories. To address these issues we use the transport square-root velocity function (TSRVF) space, a recently proposed representation that provides a metric which has favorable theoretical properties such as invariance to group action. We propose to learn the low dimensional embedding with a manifold functional variant of principal component analysis (mfPCA). We show that mf-PCA effectively models the manifold trajectories in several applications such as action recognition, clustering and diverse sequence sampling while reducing the dimensionality by a factor of ~ 250×. The mfPCA features can also be reconstructed back to the original manifold to allow for easy visualization of the latent variable space. Rushil Anirudh, Pavan Turaga, Jingyong Su, Anuj Srivastava |
CVPR | 3 |
| 2015 | Temporally coherent interpretations for long videos using pattern theoryabstractGraph-theoretical methods have successfully provided semantic and structural interpretations of images and videos. A recent paper introduced a pattern-theoretic approach that allows construction of flexible graphs for representing interactions of actors with objects and inference is accomplished by an efficient annealing algorithm. Actions and objects are termed generators and their interactions are termed bonds; together they form high-probability configurations, or interpretations, of observed scenes. This work and other structural methods have generally been limited to analyzing short videos involving isolated actions. Here we provide an extension that uses additional temporal bonds across individual actions to enable semantic interpretations of longer videos. Longer temporal connections improve scene interpretations as they help discard (temporally) local solutions in favor of globally superior ones. Using this extension, we demonstrate improvements in understanding longer videos, compared to individual interpretations of non-overlapping time segments. We verified the success of our approach by generating interpretations for more than 700 video segments from the YouCook data set, with intricate videos that exhibit cluttered background, scenarios of occlusion, viewpoint variations and changing conditions of illumination. Interpretations for long video segments were able to yield performance increases of about 70% and, in addition, proved to be more robust to different severe scenarios of classification errors. Fillipe D. M. de Souza, Sudeep Sarkar, Anuj Srivastava, Jingyong Su |
CVPR | 4 |
| 2014 | Rate-Invariant Analysis of Trajectories on Riemannian Manifolds with Application in Visual Speech RecognitionabstractIn statistical analysis of video sequences for speech recognition, and more generally activity recognition, it is natural to treat temporal evolutions of features as trajectories on Riemannian manifolds. However, different evolution patterns result in arbitrary parameterizations of these trajectories. We investigate a recent framework from statistics literature that handles this nuisance variability using a cost function/distance for temporal registration and statistical summarization & modeling of trajectories. It is based on a mathematical representation of trajectories, termed transported square-root vector field (TSRVF), and the L2 norm on the space of TSRVFs. We apply this framework to the problem of speech recognition using both audio and visual components. In each case, we extract features, form trajectories on corresponding manifolds, and compute parametrization-invariant distances using TSRVFs for speech classification. On the OuluVS database the classification performance under metric increases significantly, by nearly 100% under both modalities and for all choices of features. We obtained speaker-dependent classification rate of 70% and 96% for visual and audio components, respectively. Jingyong Su, Anuj Srivastava, Fillipe D. M. de Souza, Sudeep Sarkar |
CVPR | 1 |
| 2014 | Pattern Theory-Based Interpretation of ActivitiesabstractWe present a novel framework, based on Germander's pattern theoretic concepts, for high-level interpretation of video activities. This framework allows us to elegantly integrate ontological constraints and machine learning classifiers in one formalism to construct high-level semantic interpretations that describe video activity. The unit of analysis is a generator that could represent either an ontological label as well as a group of features from a video. These generators are linked using bonds with different constraints. An interpretation of a video is a configuration of these connected generators, which results in a graph structure that is richer than conventional graphs used in computer vision. The quality of the interpretation is quantified by an energy function that is optimized using Markov Chain Monte Carlo based simulated annealing. We demonstrate the superiority of our approach over a purely machine learning based approach (SVM) using more than 650 video shots from the You Cook dataset. This dataset is very challenging in terms of complexity of background, presence of camera motion, object occlusion, clutter, and actor variability. We find significantly improved performance in nearly all cases. Our results show that the pattern theory inference process is able to construct the correct interpretation by leveraging the ontological constraints even when the machine learning classifier is poor and the most confident labels are wrong. Fillipe D. M. de Souza, Sudeep Sarkar, Anuj Srivastava, Jingyong Su |
ICPR | 4 |
| 2013 | Statistical analysis of manual segmentations of structures in medical images
Sebastian Kurtek, Jingyong Su, Cindy Grimm, Michelle Vaughan, Ross T. Sowell, Anuj Srivastava |
Comput. Vis. Image Underst. | 2 |
| 2012 | Fitting smoothing splines to time-indexed, noisy points on nonlinear manifolds
Jingyong Su, Ian L. Dryden, Eric Klassen, Huiling Le, Anuj Srivastava |
Image Vis. Comput. | 1 |
| 2010 | Detection of Shapes in 2D Point Clouds Generated from ImagesabstractWe present a novel statistical framework for detecting pre-determined shape classes in 2D cluttered point clouds, which are in turn extracted from images. In this model based approach, we use a 1D Poisson process for sampling points on shapes, a 2D Poisson process for points from background clutter, and an additive Gaussian model for noise. Combining these with a past stochastic model on shapes of continuous 2D contours, and optimization over unknown pose and scale, we develop a generalized likelihood ratio test for shape detection. We demonstrate the efficiency of this method and its robustness to clutter using both simulated and real data. Jingyong Su, Zhiqiang Zhu, Anuj Srivastava, Fred W. Huffer |
ICPR | 1 |