Jiawei Tan

dblp:234/3931 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Modality-Aware Shot Relating and Comparing for Video Scene Detection
abstract
Video scene detection involves assessing whether each shot and its surroundings belong to the same scene. Achieving this requires meticulously correlating multi-modal cues, e.g., visual entity and place modalities, among shots and comparing semantic changes around each shot. However, most methods treat multi-modal semantics equally and do not examine contextual differences between the two sides of a shot, leading to sub-optimal detection performance. In this paper, we propose the Modality-Aware Shot Relating and Comparing approach (MASRC), which enables relating shots per their own characteristics of visual entity and place modalities, as well as comparing multi-shots similarities to have scene changes explicitly encoded. Specifically, to fully harness the potential of visual entity and place modalities in modeling shot relations, we mine long-term shot correlations from entity semantics while simultaneously revealing short-term shot correlations from place semantics. In this way, we can learn distinctive shot features that consolidate coherence within scenes and amplify distinguishability across scenes. Once equipped with distinctive shot features, we further encode the relations between preceding and succeeding shots of each target shot by similarity convolution, aiding in the identification of scene ending shots. We validate the broad applicability of the proposed components in MASRC. Extensive experimental results on public benchmark datasets demonstrate that the proposed MASRC significantly advances video scene detection.
Jiawei Tan, Hongxing Wang 0001, Kang Dang, Zhilong Ou
AAAI1
2025 Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2D
abstract
Video moment retrieval aims to locate specific moments from a video according to the query text. This task presents two main challenges: i) aligning the query and video frames at the feature level, and ii) projecting the query-aligned frame features to the start and end boundaries of the matching interval. Previous work commonly involves all frames in feature alignment, easy to cause aligning irrelevant frames with the query. Furthermore, they forcibly map visual features to interval boundaries but ignoring the information gap between them, yielding suboptimal performance. In this study, to reduce distraction from irrelevant frames, we designate an anchor frame as that with the maximum query-frame relevance measured by the established Vision-Language Model. Via similarity comparison between the anchor frame and the others, we produce a semantically compact segment around the anchor frame, which serves as a guide to align features of query and related frames. We observe that such a feature alignment will make similarity cohesive between target frames, which enables us to predict the interval boundaries by a single point detection in the 2D semantic similarity space of frames, thus well bridging the information gap between frame semantics and temporal boundaries. Experimental results across various datasets demonstrate that our approach significantly improves the alignment between queries and video frames while effectively predicting temporal moment boundaries. Especially, on QVHighlights Test and ActivityNet Captions datasets, our proposed approach achieves 3.8% and 7.4% respectively higher than current state-of-the-art [email protected] performance. The code is available at https://github.com/ExMorgan-Alter/AFAFSGD.
Jiawei Tan, Hongxing Wang 0001, Junwu Weng, Zhilong Ou, Kang Dang
CVPR1
2025 NCD: Normal-Guided Chamfer Distance Loss for Watertight Mesh Reconstruction from Unoriented Point Clouds
abstract
Abstract As a widely used loss function in learnable watertight mesh reconstruction from unoriented point clouds, Chamfer Distance (CD) efficiently quantifies the alignment between the sampled point cloud from the reconstructed mesh and its corresponding input point cloud. Occasionally, to enhance reconstruction fidelity, CD incorporates a normal consistency term, albeit at the cost of efficiency. In this context, normal estimation for unoriented point clouds requires computationally intensive matrix decomposition or specialized pre‐trained models, whereas deriving normals for mesh‐sampled points can be readily achieved using the cross product of mesh vertices. However, the reconstruction models employing CD and its variants typically rely solely on the spatial coordinates of the points, which omits normal information in favor of efficiency and deployability. To tackle this challenge, we propose a novel loss function for watertight mesh reconstruction from unoriented point clouds, termed Normal‐guided Chamfer Distance (NCD). Building upon CD, NCD introduces a normal‐steered weighting mechanism based on the angle between the normal at each mesh‐sampled point and the vector to its corresponding input point, offering several advantages: (i) it leverages readily available mesh‐sampled point normals to weight coordinate‐based Euclidean distances, thus extending the capability of CD; (ii) it eliminates the need for normal estimation from input unoriented point clouds; (iii) it incurs a negligible increase in computational complexity compared to CD. We employ NCD as the training loss for point‐to‐mesh reconstruction with multiple models and initial watertight meshes on benchmark datasets, demonstrating its superiority over state‐of‐the‐art CD variants.
Jiawei Tan, Zhilong Ou, Hongxing Wang 0001
Comput. Graph. Forum2
2025 Label refinement for change detection in remote sensing
Zhilong Ou, Hongxing Wang 0001, Jiawei Tan, Zhangbin Qian
Image Vis. Comput.3
2025 A Cost-Aware Operator Migration Approach for Distributed Stream Processing System
abstract
Stream processing is integral to edge computing due to its low-latency attributes. Nevertheless, variability in user group sizes and disparate computing capabilities of edge devices necessitate frequent operator migrations within the stream. Moreover, intricate dependencies among stream operators often obscure the detection of potential bottleneck operators until an identified bottleneck is migrated in the stream. To address this, we propose a Cost-Aware Operator Migration (CAOM) scheme. The CAOM scheme incorporates a bottleneck operator detection mechanism that directly identifies all bottleneck operators based on task running metrics. This approach avoids multiple consecutive operator migrations in complex tasks, reducing the number of task interruptions caused by operator migration. Moreover, CAOM takes into account the temporal variance in operator migration costs. By factoring in the fluctuating data generation rate from data sources at different time intervals, CAOM selects the optimal start time for operator migration to minimize the amount of accumulated data during task interruptions. Finally, we implemented CAOM on Apache Flink and evaluated its performance using the WordCount and Nexmark applications. Our experiments show that CAOM effectively reduces the number of necessary operator migrations in tasks with complex topologies and decreases the latency overhead associated with operator migration compared to state-of-the-art schemes.
Jiawei Tan, Zhuo Tang, Wentong Cai 0001, Wen Jun Tan, Jiapeng Zhang 0001, Kenli Li 0001
IEEE Trans. Cloud Comput.1
2025 A distributed skewed stream processing system based on scoring high-frequency key perception
Jiawei Tan, Yaolian Guo, Zhiwei Zuo
J. Supercomput.1
2025 Aligning Instance-Semantic Sparse Representation Towards Unsupervised Object Segmentation and Shape Abstraction With Repeatable Primitives
abstract
Understanding 3D object shapes necessitates shape representation by object parts abstracted from results of instance and semantic segmentation. Promising shape representations enable computers to interpret a shape with meaningful parts and identify their repeatability. However, supervised shape representations depend on costly annotation efforts, while current unsupervised methods work under strong semantic priors and involve multi-stage training, thereby limiting their generalization and deployment in shape reasoning and understanding. Driven by the tendency of high-dimensional semantically similar features to lie in or near low-dimensional subspaces, we introduce a one-stage, fully unsupervised framework towards semantic-aware shape representation. This framework produces joint instance segmentation, semantic segmentation, and shape abstraction through sparse representation and feature alignment of object parts in a high-dimensional space. For sparse representation, we devise a sparse latent membership pursuit method that models each object part feature as a sparse convex combination of point features at either the semantic or instance level, promoting part features in the same subspace to exhibit similar semantics. For feature alignment, we customize an attention-based strategy in the feature space to align instance- and semantic-level object part features and reconstruct the input shape using both of them, ensuring geometric reusability and semantic consistency of object parts. To firm up semantic disambiguation, we construct cascade unfrozen learning on geometric parameters of object parts. Experiments conducted on benchmark datasets confirm that our approach results in instance- and semantic-level joint segmentation and shape abstraction with repeatable primitives, providing coherent semantic interpretations of 3D object shapes across categories in a one-stage, fully unsupervised manner, without relying on annotations or heuristic semantic priors.
Hongxing Wang 0001, Jiawei Tan, Zhilong Ou, Junsong Yuan 0001
IEEE Trans. Vis. Comput. Graph.3
2024 Neighbor Relations Matter in Video Scene Detection
abstract
Video scene detection aims to temporally link shots for obtaining semantically compact scenes. It is essential for this task to capture scene-distinguishable affinity among shots by similarity assessment. However, most methods relies on ordinary shot-to-shot similarities, which may inveigle similar shots into being linked even though they are from different scenes, and meanwhile hinder dissimilar shots from being blended into a complete scene. In this paper, we propose NeighborNet to inject shot contexts into shot-to-shot similarities through carefully exploring the relations between semantic/temporal neighbors of shots over a local time period. In this way, shot-to-shot similarities are remeasured as semantic/temporal neighbor-aware similarities so that NeighborNet can learn context embedding into shot features using graph convolutional network. As a result, not only do the learned shot features suppress the affinity among similar shots from different scenes, but they also promote the affinity among dissimilar shots in the same scene. Experimental results on public benchmark datasets show that our proposed NeighborNet yields substantial improvements in video scene detection, especially outperforms released state-of-the-arts by at least 6% in Average Precision (AP). The code is available at https://github.com/ExMorgan-Alter/NeighborNet.
Jiawei Tan, Hongxing Wang 0001, Zhilong Ou, Zhangbin Qian
CVPR1
2024 CLIP-Driven Multi-Scale Instance Learning for Weakly Supervised Video Anomaly Detection
abstract
Existing weakly supervised video anomaly detection methods mainly employ Multiple Instance Learning (MIL) to identify abnormal snippets in untrimmed videos. However, the semantics and presentations of anomalies frequently exhibit ambiguity that MIL is difficult to tackle. Moreover, MIL suffers from false alarms due to its independent optimization of each instance, neglecting temporal correlation between adjacent snippets. Consequently, we badly need to better connect abnormal presentations and their semantics, as well as to enable multi-temporal-scale anomaly discovery. This paper proposes a CLIP-Driven Multi-Scale Instance Learning (CMSIL) framework with two branches including Vision-Language (VL) and Multi-Scale Instance Learning (MSIL). The VL branch leverages the powerful visual concept priors from Contrastive Language-Image Pre-training (CLIP) to generate pseudo anomalies, thereby providing suspected anomaly cues for model training guidance. The MSIL branch utilizes a feature pyramid to fully mine fine-grained temporal dependencies by employing MIL within each pyramid level to learn anomalous patterns across different temporal scales. By collaborating with the two branches, the proposed CMSIL shows better proficiency in handling anomalies with varying durations. Extensive experiments on the XD-Violence and UCF-Crime datasets demonstrate the superior performance of our method. The code is available at https://github.com/casperZB/CMSIL.
Zhangbin Qian, Jiawei Tan, Zhilong Ou, Hongxing Wang 0001
ICME2
2024 Semantic Transition Detection for Self-supervised Video Scene Segmentation
Jiawei Tan, Pingan Yang, Hongxing Wang 0001
MMM (3)2
2024 Shared Latent Membership Enables Joint Shape Abstraction and Segmentation With Deformable Superquadrics
abstract
Part-level 3D shape representations are crucial to shape reasoning and understanding. Two key sub-tasks are: 1) shape abstraction, creating primitive-based object parts; and 2) shape segmentation, finding partition-based object parts. However, for 3D object point clouds, most advanced methods produce parts relying on task-specific priors, such as similarity metrics and primitive geometries, resulting in misleading parts that deviate from semantics. To address prior limitations, we establish a foundation for joint shape abstraction and shape segmentation as formal linear transformations within a shared latent space, encapsulating essential dual-purpose membership information linking points and object parts for mutual reinforcement. We demonstrate that the transformations are underpinned by a derivation based on k-means, Non-negative Matrix Factorization (NMF), and the attention mechanism. As a result, we introduce Latent Membership Pursuit (LMP) for joint optimization of shape abstraction and segmentation. LMP utilizes a shared latent representation of object part membership to autonomously identify common object parts in both tasks without any supervision and priors. Furthermore, we adapt deformable superquadrics (DSQs) for primitives to capture variable part-level geometric and semantic information. Experiments on benchmark datasets validate that our approach enables mutual learning of shape abstraction and segmentation, and promotes consistent interpretations of 3D object shapes across instances and even categories in a fully unsupervised manner.
Hongxing Wang 0001, Jiawei Tan, Junsong Yuan 0001
IEEE Trans. Image Process.3
2024 Characters Link Shots: Character Attention Network for Movie Scene Segmentation
abstract
Movie scene segmentation aims to automatically segment a movie into multiple story units, i.e., scenes, each of which is a series of semantically coherent and time-continual shots. Previous methods have continued efforts on shot semantic association, but few take into account the impact of different semantics on foreground characters and background scenes in movie shots. In particular, the background scene in the shot can adversely affect scene boundary classification. Motivated by the fact that it is the characters who drive the plot development of a movie scene, we build a Character Attention Network (CANet) to detect movie scene boundaries in a character-centric fashion. To eliminate the background clutter, we extract multi-view character semantics for each shot in terms of human bodies and faces. Furthermore, we equip our CANet with two stages of character attention. The first is Masked Shot Attention (MSA) through selective self-attention over similar temporal contexts from multi-view character semantics to yield an enhanced omni-view shot representation, by which the CANet can better handle the variations of characters in pose and appearance. The second is Key Character Attention (KCA) through temporal-aware attention on character reappearances for Bidirectional Long Short-Term Memory (Bi-LSTM) feature association so that linking shots can be focused on those with recurring key characters. We encourage the proposed CANet in learning boundary-discriminative shot features. Specifically, we formulate a Boundary-Aware circle Loss (BAL) to push far apart CANet-features between adjacent scenes, which is also coupled with the cross-entropy loss to drive CANet-features sensitive to scene boundaries. Experimental results on the MovieNet-SSeg and OVSD datasets show that our method achieves superior performance in temporal scene segmentation compared with state-of-the-art methods.
Jiawei Tan, Hongxing Wang 0001, Junsong Yuan 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Temporal Scene Montage for Self-Supervised Video Scene Boundary Detection
abstract
Once a video sequence is organized as basic shot units, it is of great interest to temporally link shots into semantic-compact scene segments to facilitate long video understanding. However, it still challenges existing video scene boundary detection methods to handle various visual semantics and complex shot relations in video scenes. We proposed a novel self-supervised learning method, Video Scene Montage for Boundary Detection (VSMBD), to extract rich shot semantics and learn shot relations using unlabeled videos. More specifically, we present Video Scene Montage (VSM) to synthesize reliable pseudo scene boundaries, which learns task-related semantic relations between shots in a self-supervised manner. To lay a solid foundation for modeling semantic relations between shots, we decouple visual semantics of shots into foreground and background. Instead of costly learning from scratch as in most previous self-supervised learning methods, we build our model upon large-scale pre-trained visual encoders to extract the foreground and background features. Experimental results demonstrate VSMBD trains a model with strong capability in capturing shot relations, surpassing previous methods by significant margins. The code is available at https://github.com/mini-mind/VSMBD.
Jiawei Tan, Pingan Yang, Hongxing Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Cascade Sampling via Dual Uncertainty for Active Entity Alignment
Jiye Xie, Jiawei Tan, Hongxing Wang 0001
KSEM (2)3
2022 Quantum and classical query complexities for generalized Simon's problem
Zhenggang Wu, Daowen Qiu, Jiawei Tan, Guangya Cai
Theor. Comput. Sci.3
2021 Machine learning-based prediction of survival prognosis in cervical cancer
abstract
BACKGROUND: Accurately forecasting the prognosis could improve cervical cancer management, however, the currently used clinical features are difficult to provide enough information. The aim of this study is to improve forecasting capability by developing a miRNAs-based machine learning survival prediction model. RESULTS: The expression characteristics of miRNAs were chosen as features for model development. The cervical cancer miRNA expression data was obtained from The Cancer Genome Atlas database. Preprocessing, including unquantified data removal, missing value imputation, samples normalization, log transformation, and feature scaling, was performed. In total, 42 survival-related miRNAs were identified by Cox Proportional-Hazards analysis. The patients were optimally clustered into four groups with three different 5-years survival outcome (≥ 90%, ≈ 65%, ≤ 40%) by K-means clustering algorithm base on top 10 survival-related miRNAs. According to the K-means clustering result, a prediction model with high performance was established. The pathways analysis indicated that the miRNAs used play roles involved in the regulation of cancer stem cells. CONCLUSION: A miRNAs-based machine learning cervical cancer survival prediction model was developed that robustly stratifies cervical cancer patients into high survival rate (5-years survival rate ≥ 90%), moderate survival rate (5-years survival rate ≈ 65%), and low survival rate (5-years survival rate ≤ 40%).
Dongyan Ding, Tingyuan Lang, Dongling Zou, Jiawei Tan, Dong Wang 0024, Yunzhe Li 0001, Jingshu Liu, Cui Ma
BMC Bioinform.4
2021 A Novel UAV-Enabled Data Collection Scheme for Intelligent Transportation System Through UAV Speed Control
abstract
The rapid and convenient travel of people and the timely transportation of goods depend on the correct decision of the Intelligent Transportation Systems (ITS). Due to the decision-making of ITS requires a large amount of data to support, UAV-enabled periodic data collection is an effective method. However, due to the limited resources of UAV, UAV cannot directly collect data from all storage devices, resulting in unfair data collection. Therefore, we propose a UAV Speed Control based Fairness Data Collection (USCFDC) scheme. First, since the fairness of data collection will affect the decision-making of ITS, a framework for controlling the flight speed of the UAV is proposed to improve the fairness of data collection. The flight speed of UAV will slow down in areas with a large number of nodes, thereby improving the fairness of data collection. Second, a novel method is proposed to maximize the amount of data collected by UAV from each node. With this method, the value of the amount of data will be used as the dichotomous value in the dichotomy algorithm, and the UAV must collect a certain amount of data from each node. The upper and lower limits of the dichotomy algorithm are adjusted according to the time duration for UAV to collect data. Compared with previous schemes, the fairness of data collection can be improved by a maximum of 15.89% under the same flight time of UAV. Besides, the energy consumption is reduced by 49.31%-52.55% and the flight time of the UAV is reduced by 48%-62.38% when the amount of collected data is the same.
Xiong Li 0002, Jiawei Tan, Anfeng Liu, Pandi Vijayakumar, Neeraj Kumar 0001, Mamoun Alazab
IEEE Trans. Intell. Transp. Syst.2