Siyuan Yao

dblp:137/5732 · DBLP profile ↗
← Back
29ranked-venue papers
11as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 9 first-author · 21 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploring Inter-Domain Wasserstein Metric for Adaptive Object Detection
abstract
Cross-domain adaptation has achieved significant development in recent years. Nevertheless, the models' performance varies dramatically across different scenarios. How to quantitatively measure inter-domain discrepancies to guide model training remains a challenging problem. Existing methods mainly focus on the trivial pixel-wise differences between the cross domain images, while they ignore the holistic discrepancies of the domain-specific distributions in various scenarios. Thus their effectivenesses are greatly limited in realistic applications. In this paper, we propose a universal method for measuring inter domain discrepancies based on Wasserstein distance. It alleviates the impact of intra-domain discrepancies on measurements and enables precise and quantitative representation of inter-domain discrepancies. We further integrate this metric into the generative models, and propose an assistant domain to conduct domain knowledge transfer for cross-domain object detection task. Experiments on three benchmarks validate the effectiveness of the proposed inter-domain measurement metric. Specifically, we achieve 51.8% mAP on CityScapes, 45.3% mAP on Clipart and 58.9% mAP on Watercolor, which are 1.5%, 0.5% and 0.8% higher than the state-of-the-art schemes, respectively.
Yanlong Lin, Ziqian Zhu, Yitian Guo, Zhonghong Ou, Siyuan Yao, Meina Song
IEEE Trans. Multim.5
2026 VolSegGS: Segmentation and Tracking in Dynamic Volumetric Scenes via Deformable 3D Gaussians
abstract
Visualization of large-scale time-dependent simulation data is crucial for domain scientists to analyze complex phenomena, but it demands significant I/O bandwidth, storage, and computational resources. To enable effective visualization on local, low-end machines, recent advances in view synthesis techniques, such as neural radiance fields, utilize neural networks to generate novel visualizations for volumetric scenes. However, these methods focus on reconstruction quality rather than facilitating interactive visualization exploration, such as feature extraction and tracking. We introduce VolSegGS, a novel Gaussian splatting framework that supports interactive segmentation and tracking in dynamic volumetric scenes for exploratory visualization and analysis. Our approach utilizes deformable 3D Gaussians to represent a dynamic volumetric scene, allowing for real-time novel view synthesis. For accurate segmentation, we leverage the view-independent colors of Gaussians for coarse-level segmentation and refine the results with an affinity field network for fine-level segmentation. Additionally, by embedding segmentation results within the Gaussians, we ensure that their deformation enables continuous tracking of segmented regions over time. We demonstrate the effectiveness of VolSegGS with several time-varying datasets and compare our solutions against state-of-the-art methods. With the ability to interact with a dynamic scene in real time and provide flexible segmentation and tracking capabilities, VolSegGS offers a powerful solution under low computational demands. This framework unlocks exciting new possibilities for time-varying volumetric data analysis and visualization.
Siyuan Yao, Chaoli Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2025 IWRN: A Robust Blind Watermarking Method for Artwork Image Copyright Protection Against Noise Attack
abstract
Adding imperceptible watermarks to artwork images, such as paintings and photographs, can effectively safeguard the copyright of these images without compromising their usability. However, existing blind watermarking techniques encounter two major challenges in addressing this task: imperceptibility and robustness, particularly when subjected to various noise attacks. In this paper, we propose a blind watermarking method for artwork image copyright protection, IWRN, which can ensure both the Imperceptibility of the Watermark and Robustness against Noise attacks. For imperceptibility, we design a Learnable Wavelet Network (LWN) to adaptively embed the watermark into the high-frequency region where the watermark has better invisibility. For robustness, we establish a Deform-Attention based Invertible Neural Network (DA-INN) with a decoding optimization, which offers the advantage of computational reversion, and combines the deform-attention mechanism and decoding optimization to enhance the model's resistance against noises. Additionally, we design a Joint Contrast Learning (JCL) mechanism to improve imperceptibility and robustness simultaneously. Experiments show that our IWRN outperforms other state-of-the-art blind watermarking methods, achieves an average performance of 41.55 PSNR and 99.57% accuracy on the Coco2017, Wikiart, and Div2k datasets when facing 12 kinds of noise attacks.
Feifei Kou, Yuhan Yao 0001, Siyuan Yao, Lei Shi 0030, Yawen Li 0001, Xuejing Kang
AAAI3
2025 LVPTrack: High Performance Domain Adaptive UAV Tracking with Label Aligned Visual Prompt Tuning
abstract
Visual object tracking is essentially crucial for unmanned aerial vehicles (UAVs). Despite the substantial progress, most of the existing UAV trackers are designed for well-conditioned daytime data, while for the scenarios in challenging weather condition, e.g. foggy or nighttime environment, the tremendous domain gap leads to significant performance degradation. To address this issue, in this paper, we propose a novel robust UAV tracker termed LVPTrack, which conducts high quality label-aligned visual prompt tuning to adapt to various challenging weather conditions. Specifically, we first synthesize the sequential foggy and nighttime video frames to assist the model training. A domain adaptive teacher-student network is utilized to distill the hierarchical visual semantic of the target objects in cross-domain scenarios. Then we propose a target-aware pseudo-label voting (PLV) strategy to alleviate the target-level misalignment in the dual domains. Furthermore, we propose a dynamic aggregated prompt (DAP) module to facilitate the appearance variation adaptation of the target object in challenging scenarios. Extensive experiments demonstrate that our tracker achieves superior performance over existing state-of-the-art UAV trackers.
Hongjing Wu, Siyuan Yao, Feng Huang 0007, Linchao Zhang, Zhuoran Zheng, Wenqi Ren
AAAI2
2025 UMDATrack: Unified Multi-Domain Adaptive Tracking under Adverse Weather Conditions
abstract
Visual object tracking has gained promising progress in past decades. Most of the existing approaches focus on learning target representation in well-conditioned daytime data, while for the unconstrained real-world scenarios with adverse weather conditions, e.g. nighttime or foggy environment, the tremendous domain shift leads to significant performance degradation. In this paper, we propose UMDATrack, which is capable of maintaining high-quality target state prediction under various adverse weather conditions within a unified domain adaptation framework. Specifically, we first use a controllable scenario generator to synthesize a small amount of unlabeled videos (less than 2% frames in source daytime datasets) in multiple weather conditions under the guidance of different text prompts. Afterwards, we design a simple yet effective domain-customized adapter (DCA), allowing the target objects' representation to rapidly adapt to various weather conditions without redundant model updating. Furthermore, to enhance the localization consistency between source and target domains, we propose a target-aware confidence alignment module (TCA) following optimal transport theorem. Extensive experiments demonstrate that UMDATrack can surpass existing advanced visual trackers and lead new state-of-the-art performance by a significant margin. Our code is available at https://github.com/Z-Z188/UMDATrack.
Siyuan Yao, Wenqi Ren, Yanyang Yan, Xiaochun Cao
ICCV1
2025 DGFSD: Bridging the Gap between Dense and Sparse for Fully Sparse 3D Object Detection
abstract
Recently, LiDAR-based fully sparse 3D object detection has gained great attention, which utilizes point clouds to boost efficiency. Nevertheless, the relationship between well-studied dense representation and fully sparse representation is under-explored in existing studies, which focuses solely on building sparse representation by feature diffusion to solve the notorious center point missing problem. To this end, we propose a dense-guided fully sparse detection scheme, named DGFSD, to bridge the gap between dense and sparse features by dense-guided diffusion. Different from prior studies, we propose DgD (Dense-guided Diffusion) to overcome the center feature missing problem by dense knowledge transferring. Specifically, DgD transfers high-quality central point features from dense representations to endow sparse representations with dense knowledge. Moreover, we customize DFW (Dense Feature Weighting) to express uninformative representation and lift foreground representation. It makes high-quality dense feature contribute more to arcuate regression. To the best of our knowledge, we are the first to explore dense knowledge's impact on fully sparse framework. Extensive experiments conducted on nuScenes and Argoverse2 benchmark demonstrate the effectiveness of the proposed method. Specifically, DGFSD achieves 71.6% NDS and 67.3% mAP on the nuScenes test benchmark. On Argoverse2, DGFSD achieves 40.6% mAP, outperforming previous best hybrid and fully sparse methods. The code is available at https://github.com/Raiden-cn/DGFSD.
Zhonghong Ou, Kaiwen Xue 0001, Jiangfeng Sun 0003, Yifan Zhu 0001, Siyuan Yao, Yiran Shen 0007, Meina Song
ACM Multimedia6
2025 ViSNeRF: Efficient Multidimensional Neural Radiance Field Representation for Visualization Synthesis of Dynamic Volumetric Scenes
abstract
Domain scientists often face I/O and storage challenges when keeping raw data from large-scale simulations. Saving visualization images, albeit practical, is limited to preselected viewpoints, transfer functions, and simulation parameters. Recent advances in scientific visualization leverage deep learning techniques for visualization synthesis by offering effective ways to infer unseen visualizations when only image samples are given during training. However, due to the lack of 3D geometry awareness, existing methods typically require many training images and significant learning time to generate novel visualizations faithfully. To address these limitations, we propose ViSNeRF, a novel 3D-aware approach for visualization synthesis using neural radiance fields. Leveraging a multidimensional radiance field representation, ViSNeRF efficiently reconstructs visualizations of dynamic volumetric scenes from a sparse set of labeled image samples with flexible parameter exploration over transfer functions, isovalues, timesteps, or simulation parameters. Through qualitative and quantitative comparative evaluation, we demonstrate ViSNeRF’s superior performance over several representative baseline methods, positioning it as the state-of-the-art solution. The code is available at https://github.com/JCBreath/ViSNeRF.
Siyuan Yao, Yunfei Lu, Chaoli Wang 0001
PacificVis1
2025 ReVolVE: Neural reconstruction of volumes for visualization enhancement of direct volume rendering
Siyuan Yao, Chaoli Wang 0001
Comput. Graph.1
2025 Community Detection in Heterogeneous Information Networks Without Materialization
abstract
Community detection in heterogeneous information networks (HINs) poses significant challenges due to the diversity of entity types and the complexity of their interrelations. While traditional algorithms may perform adequately in some scenarios, many struggle with the high memory usage and computational demands of large-scale HINs. To address these challenges, we introduce a novel framework, SCAR, which efficiently uncovers community structures in HINs without requiring network materialization. SCAR leverages insights from meta-paths to interpret multi-relational data through compact vertex-based sketches, significantly reducing computational overhead and materialization overhead. We propose a sketch-based technique for estimating changes in modularity, improving both the precision and speed in community detection. Our extensive evaluations on diverse real-world datasets provide detailed comparative metrics, demonstrating that SCAR outperforms several state-of-the-art methods, including Gdy, Louvain, Leiden, Infomap, Walktrap, and Networkit, in execution time and memory consumption while maintaining competitive accuracy. Overall, SCAR offers a robust and scalable solution for revealing community structures in large HINs, with applications across various domains, including social networks, academic collaboration networks, and e-commerce platforms.
Siyuan Yao, Bingsheng He, Yudong Niu, Yuchen Li 0001, Shixuan Sun, Yongchao Liu 0004
Proc. ACM Manag. Data2
2025 Dupin: A Parallel Framework for Densest Subgraph Discovery in Fraud Detection on Massive Graphs
abstract
Detecting fraudulent activities in financial and e-commerce transaction networks is crucial. One effective method for this is Densest Subgraph Discovery (DSD). However, deploying DSD methods in production systems faces substantial scalability challenges due to the predominantly sequential nature of existing methods, which impedes their ability to handle large-scale transaction networks and results in significant detection delays. To address these challenges, we introduce Dupin, a novel parallel processing framework designed for efficient DSD processing in billion-scale graphs. Dupin is powered by a processing engine that exploits the unique properties of the peeling process, with theoretical guarantees on detection quality and efficiency. Dupin provides user-friendly APIs for flexible customization of DSD objectives and ensures robust adaptability to diverse fraud detection scenarios. Empirical evaluations indicate that Dupin consistently outperforms several existing DSD methods, achieving performance improvements of up to two orders of magnitude compared to traditional approaches. On billion-scale graphs, Dupin demonstrates the potential to enhance the prevention of fraudulent transactions by approximately 49.5 basis points and reduces density error from 30.3% to below 5.0%, as supported by our experimental results.
Siyuan Yao, Yuchen Li 0001, Qiange Wang, Bingsheng He, Min Chen 0018
Proc. ACM Manag. Data2
2025 CAN: Cascade Augmentations Against Noise for Image Restoration
abstract
Image restoration aims to recover the latent clean image from a degraded counterpart. In general, the prevailing state-of-the-art image restoration methods concentrate on solving only a specific degradation type according to the task, e.g., deblurring or deraining. However, if the corresponding well-trained frameworks confront other real-world image corruptions, i.e., the corruptions are not covered in the training phase, and state-of-the-art restoration models will suffer from a lack of generalization ability. We have observed that an image restoration model can be easily confused by noise corruption. Towards improving the robustness of image restoration networks, in this paper, we focus on alleviating the corruption of noise in various image restoration tasks, which is almost inevitable in real-world scenes. To this end, we devise a novel Cascade Augmentation strategy against Noise (CAN) to enhance the robustness of specific image restoration. Specifically, the given degraded images are sequentially augmented from different perspectives, i.e., noise-aware augmentation and model-aware augmentation. The noise-aware augmentation is proposed to enrich the samples by introducing various noise operations. Moreover, to adapt to more unknown corruptions, we propose a novel model-aware augmentation mechanism, which enhances the scalability by exploring useful both spatial and frequency clues with the help of model randomness. It is worth noting that the proposed augmentation scheme is model-agnostic, and it can plug and play into arbitrary state-of-the-art image restoration architectures. In addition, we construct noise corruption benchmark datasets, derived from the validation set of standard image restoration datasets, to assist us in evaluating the robustness of restoration networks. Extensive quantitative and qualitative evaluations demonstrate that the proposed method has strong generalization capability, which can enhance the robustness of various image restoration frameworks when facing diverse noises.
Yanyang Yan, Siyuan Yao, Wenqi Ren, Rui Zhang 0040, Qi Guo 0001, Xiaochun Cao
IEEE Trans. Image Process.2
2025 UncTrack: Reliable Visual Object Tracking With Uncertainty-Aware Prototype Memory Network
abstract
Transformer-based trackers have achieved promising success and become the dominant tracking paradigm because of their accuracy and efficiency. Despite the substantial progress, most of the existing approaches handle object tracking as a deterministic coordinate regression problem, while the target localization uncertainty has been largely overlooked, which hampers trackers' ability to maintain reliable target state prediction in challenging scenarios. To address this issue, we propose UncTrack, a novel uncertainty-aware transformer-based tracker that predicts the target localization uncertainty and incorporates this uncertainty information for accurate target state inference. Specifically, UncTrack uses a transformer encoder to perform feature interactions between the template and search images. The output features are passed into an uncertainty-aware localization decoder (ULD) to coarsely predict the corner-based localization and the corresponding localization uncertainty. Then, the localization uncertainty is sent into a prototype memory network (PMN) to excavate valuable historical information to identify whether the target state prediction is reliable. To enhance the template representation, the samples with high confidence are fed back into the prototype memory bank for memory updating, which makes the tracker more robust to challenging appearance variations. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. Our code is available at https://github.com/ManOfStory/UncTrack.
Siyuan Yao, Yanyang Yan, Wenqi Ren, Xiaochun Cao
IEEE Trans. Image Process.1
2025 iVR-GS: Inverse Volume Rendering for Explorable Visualization via Editable 3D Gaussian Splatting
abstract
In volume visualization, users can interactively explore the three-dimensional data by specifying color and opacity mappings in the transfer function (TF) or adjusting lighting parameters, facilitating meaningful interpretation of the underlying structure. However, rendering large-scale volumes demands powerful GPUs and high-speed memory access for real-time performance. While existing novel view synthesis (NVS) methods offer faster rendering speeds with lower hardware requirements, the visible parts of a reconstructed scene are fixed and constrained by preset TF settings, significantly limiting user exploration. This article introduces inverse volume rendering via Gaussian splatting (iVR-GS), an innovative NVS method that reduces the rendering cost while enabling scene editing for interactive volume exploration. Specifically, we compose multiple iVR-GS models associated with basic TFs covering disjoint visible parts to make the entire volumetric scene visible. Each basic model contains a collection of 3D editable Gaussians, where each Gaussian is a 3D spatial point that supports real-time scene rendering and editing. We demonstrate the superior reconstruction quality and composability of iVR-GS against other NVS solutions (Plenoxels, CCNeRF, and base 3DGS) on various volume datasets. The code is available at https://github.com/TouKaienn/iVR-GS.
Kaiyuan Tang 0001, Siyuan Yao, Chaoli Wang 0001
IEEE Trans. Vis. Comput. Graph.2
2024 USSC: Universal and Storage-Efficient Sidechains
abstract
Blockchain interoperability has become an essential functionality, which enables asset/data transfers across different blockchains. Sidechains have been deemed as a key technique to provide interoperability. However, sidechains are rarely used in practice, this is because sidechain technologies are impractical and non cost-efficient. To make sidechains practical, in this paper, we design a universal sidechain construction named USSC, which applies to a variety of blockchains without forking them. USSC also enables interoperability across heterogeneous blockchains regardless of underlying consensus. This is facilitated by three components: i) a committee selection method, ii) a cross-chain certificate, and iii) a cross-chain bridge based on smart contracts. The proposed committee-selection method guarantees an honest majority within a committee. Through a concrete implementation of USSC, we outline how the proof-of-stake (PoS) and the proof-of-work (PoW) blockchains enable asset transfers. Furthermore, our USSC is more storage-efficient because it produces a smaller size of certificate and only needs partial nodes instead of all sidechain nodes following a blockchain. Thus, USSC can reduce the overhead of storage and communication of nodes. In addition, we prove that USSC achieves a secure sidechain construction with desirable security properties. Finally, we develop a proof-of-concept implementation of USSC using Cardano and Ethereum. Experimental results demonstrate that USSC outperforms PoW and PoS sidechains, in terms of the certificate size.
Taotao Li, Huawei Huang, Lingyuan Yin, Siyuan Yao, Zibin Zheng
ICDCS4
2024 GridMask: An Efficient Scheme for Real Time Curved Scene Text Detection
Zhonghong Ou, Siyuan Yao, Meina Song
PRCV (7)3
2024 A Comparative Study of Neural Surface Reconstruction for Scientific Visualization
abstract
This comparative study evaluates various neural surface reconstruction methods, particularly focusing on their implications for scientific visualization through reconstructing 3D surfaces via multi-view rendering images. We categorize ten methods into neural radiance fields and neural implicit surfaces, uncovering the benefits of leveraging distance functions (i.e., SDFs and UDFs) to enhance the accuracy and smoothness of the reconstructed surfaces. Our findings highlight the efficiency and quality of NeuS2 for reconstructing closed surfaces and identify NeUDF as a promising candidate for reconstructing open surfaces despite some limitations. By sharing our benchmark dataset, we invite researchers to test the performance of their methods, contributing to the advancement of surface reconstruction solutions for scientific visualization.
Siyuan Yao, Weixi Song, Chaoli Wang 0001
IEEE VIS1
2024 uBlade: Efficient Batch Processing for Uncertainty Graph Queries
abstract
The study of uncertain graphs is crucial in diverse fields, including but not limited to protein interaction analysis, viral marketing, and network reliability. Processing queries on uncertain graphs presents formidable challenges due to the vast probabilistic space they encapsulate. While existing systems employ batch processing to address these challenges, their performance is often compromised by the suboptimal selection of parallel graph traversal methods, the excessive costs in random number generation, and additional workloads intrinsic to batch processing. In this paper, we introduce uBlade, an efficient batch-processing framework for uncertain graph queries on multi-core CPUs. uBlade utilizes the work-efficient graph traversal, achieving superior parallelism in the batch processing model. Additionally, our Quasi-Sampling technique reduces the random number generation cost by a factor of B, with O(B) denoting the batch size. We further examine the extra workload resulting from batch processing and introduce an efficient strategy to reorder possible worlds, minimizing this associated overhead. Through comprehensive evaluations, we showcase that uBlade achieves up to two orders of magnitude speedups against the state-of-the-art CPU and GPU-based solutions.
Siyuan Yao, Yuchen Li 0001, Shixuan Sun, Bingsheng He
Proc. ACM Manag. Data1
2024 Shadow-aware decomposed transformer network for shadow detection and removal
Xiao Wang 0025, Siyuan Yao, Sili Yang, Zhenbao Liu
Pattern Recognit.2
2024 Hierarchical Graph Interaction Transformer With Dynamic Token Clustering for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) aims to identify the objects that seamlessly blend into the surrounding backgrounds. Due to the intrinsic similarity between the camouflaged objects and the background region, it is extremely challenging to precisely distinguish the camouflaged objects by existing approaches. In this paper, we propose a hierarchical graph interaction network termed HGINet for camouflaged object detection, which is capable of discovering imperceptible objects via effective graph interaction among the hierarchical tokenized features. Specifically, we first design a region-aware token focusing attention (RTFA) with dynamic token clustering to excavate the potentially distinguishable tokens in the local region. Afterwards, a hierarchical graph interaction transformer (HGIT) is proposed to construct bi-directional aligned communication between hierarchical features in the latent interaction space for visual semantics enhancement. Furthermore, we propose a decoder network with confidence aggregated feature fusion (CAFF) modules, which progressively fuses the hierarchical interacted features to refine the local detail in ambiguous regions. Extensive experiments conducted on the prevalent datasets, i.e. COD10K, CAMO, NC4K and CHAMELEON demonstrate the superior performance of HGINet compared to existing state-of-the-art methods. Our code is available at https://github.com/Garyson1204/HGINet.
Siyuan Yao, Hao Sun 0019, Tian-Zhu Xiang, Xiao Wang 0017, Xiaochun Cao
IEEE Trans. Image Process.1
2023 From Artifacts to Outcomes: Comparison of HMD VR, Desktop, and Slides Lectures for Food Microbiology Laboratory Instruction
abstract
Despite the value of VR (Virtual Reality) for educational purposes, the instructional power of VR in Biology Laboratory education remains under-explored. Laboratory lectures can be challenging due to students’ low motivation to learn abstract scientific concepts and low retention rate. Therefore, we designed a VR-based lecture on fermentation and compared its effectiveness with lectures using PowerPoint slides and a desktop application. Grounded in the theory of distributed cognition and motivational theories, our study examined how learning happens in each condition from students’ learning outcomes, behaviors, and perceptions. Our result indicates that VR facilitates students’ long-term retention to learn by cultivating their longer visual attention and fostering a higher sense of immersion, though students’ short-term retention remains the same across all conditions. This study extends current research on VR studies by identifying the characteristics of each teaching artifact and providing design implications for integrating VR technology into higher education.
Rongchen Guo, Siyuan Yao, Luxin Wang, Kwan-Liu Ma
CHI3
2023 GMT: A deep learning approach to generalized multivariate translation for scientific data analysis and visualization
Siyuan Yao, Jun Han 0010, Chaoli Wang 0001
Comput. Graph.1
2023 Free$\rm ^{3}$Net: Gliding Free, Orientation Free, and Anchor Free Network for Oriented Object Detection
abstract
Object detection for aerial images has achieved remarkable progress in recent years. Nevertheless, most exiting studies do not differentiate oriented object detection from horizontal detection. Certain schemes ignore the ambiguity of oriented object representation and leverage label assignment designed for horizontal object detection directly. Consequently, it leads to unstable training and causes performance degradation, because high-quality samples surrounding the oriented bounding boxes can not be leveraged effectively. To address this problem, we propose a gliding Free, orientation Free, and anchor Free Network (Free$\rm ^{3}$Net) with high-efficiency for oriented object detection. Specifically, we propose an unambiguous oriented object representation scheme, named FreeGliding, by gliding the projection points of samples on each edge of horizontal bounding boxes. It makes the detection largely free from representation ambiguity and multi-task dependency. To overcome the restrictions of label assignment, we put forward a novel Loss-aware Outer Sample Selection (LOSS) scheme, which takes into consideration spatial information and localization capability to retain high-quality samples surrounding the objects. Moreover, we introduce an Oriented Feature Fusion (OFF) scheme to tackle feature alignment by adjusting the receptive field and fusing oriented features dynamically. Experimental results on two large-scale remote sensing datasets HRSC2016 and DOTA demonstrate that Free$\rm ^{3}$Net outperforms the state-of-the-art schemes with a large margin. We hope our work can inspire rethinking the design of anchor-free detectors, and serve as a strong baseline for oriented object detection.
Zhonghong Ou, Zhongjie Chen, Shengyi Shen, Lina Fan, Siyuan Yao, Meina Song, Pan Hui 0001
IEEE Trans. Multim.5
2022 Weakly-supervised Disentanglement Network for Video Fingerspelling Detection
abstract
Fingerspelling detection, which aims to localize and recognize fingerspelling gestures in raw, untrimmed videos, is a nascent but important research area that could help bridge the communication gap between deaf people and others. Many existing works tend to exploit additional knowledge, such as pose annotations, and newly datasets for performance improvement. However, in real-world applications, additional data collection and annotation require tremendous human efforts that are not always affordable. In this paper, we propose the Weakly-supervised Disentanglement Network, namely WED, that requires no additional knowledge, and better exploits the video-sentence weak supervisions. Specifically, WED incorporates two critical components: 1) Masked Disentanglement Module, which employs a Variational Autoencoder for signed letters disentanglement. Each latent factor in the VAE corresponds to a particular signed letter, and we mask latent factors corresponding to letters that do not appear in the video during decoding. Compared to the vanilla VAE, the masked reconstruction leverages the video-sentence weak supervision, leading to a better sign language oriented disentanglement; and 2) the Dynamic Memory Network module, which leverages the disentangled sign knowledge as prior knowledge and reference for sign-related frame identification and gesture recognition through a carefully designed memory reading component. We conduct extensive experiments on the benchmark ChicagoFSWild and ChicagoFSWild+ datasets. Empirical studies validate that the WED network achieves effective sign gesture disentanglement, contributing to the state-of-the-art performance for fingerspelling detection and recognition.
Ziqi Jiang, Shengyu Zhang 0001, Siyuan Yao, Wenqiao Zhang, Sihan Zhang, Juncheng Li 0006, Zhou Zhao 0001, Fei Wu 0001
ACM Multimedia3
2022 ACE: Anchor-Free Corner Evolution for Real-Time Arbitrarily-Oriented Object Detection
abstract
Objects with different orientations are ubiquitous in the real world (e.g., texts/hands in the scene image, objects in the aerial image, etc.), and the widely-used axis-aligned bounding box does not compactly enclose the oriented objects. Thus arbitrarily-oriented object detection has attracted rising attention in recent years. In this paper, we propose a novel and effective model to detect arbitrarily-oriented objects. Instead of directly predicting the angles of oriented bounding boxes like most existing methods, we evolve the axis-aligned bounding box to the oriented quadrilateral box with the assistance of dynamically gathering contour information. More specifically, we first obtain the axis-aligned bounding box in an anchor-free manner. After that, we set the key points based on the sampled contour points of the axis-aligned bounding box. To improve the localization performance, we enrich the feature representations of these key points by exploiting a dynamic information gathering mechanism. This technique propagates the geometrical and semantic information along the sampled contour points, and fuses the information from the semantic neighbors of each sampled point, which varies for different locations. Finally, we estimate the offsets between the axis-aligned bounding box key points and the oriented quadrilateral box corner points. Extensive experiments on two frequently-used aerial image benchmarks HRSC2016 and DOTA, as well as scene text/hand datasets ICDAR2015, TD500, and Oxford-Hand, demonstrate the effectiveness and advantage of our proposed model.
Pengwen Dai, Siyuan Yao, Zekun Li 0007, Sanyi Zhang, Xiaochun Cao
IEEE Trans. Image Process.2
2021 Updated Paired Regions for Shadow Detection from Single Image
Xiao Wang 0017, Siyuan Yao, Pengwen Dai, Rui Wang 0032, Xiaochun Cao
BMVC2
2021 A Comparison of the Fatigue Progression of Eye-Tracked and Motion-Controlled Interaction in Immersive Space
abstract
Eye-tracking enabled virtual reality (VR) headsets have recently become more widely available. This opens up opportunities to incorporate eye gaze interaction methods in VR applications. However, studies on the fatigue-induced performance fluctuations of these new input modalities are scarce and rarely provide a direct comparison with established interaction methods. We conduct a study to compare the selection-interaction performance between commonly used handheld motion control devices and emerging eye interaction technology in VR. We investigate each interaction’s unique fatigue progression pattern in study sessions with ten minutes of continuous engagement. The results support and extend previous findings regarding the progression of fatigue in eye-tracked interaction over prolonged periods. By directly comparing gaze-with motion-controlled interaction, we put the emerging eye-trackers into perspective with the state-of-the-art interaction method for immersive space. We then discuss potential implications for future extended reality (XR) interaction design based on our findings.
Lukas Maximilian Masopust, David Bauer, Siyuan Yao, Kwan-Liu Ma
ISMAR3
2021 Learning Deep Lucas-Kanade Siamese Network for Visual Tracking
abstract
In most recent years, Siamese trackers have drawn great attention because of their well-balanced accuracy and efficiency. Although these approaches have achieved great success, the discriminative power of the conventional Siamese trackers is still limited by the insufficient template-candidate representation. Most of the existing approaches take non-aligned features to learn a similarity function for template-candidate matching, while the target object's geometrical transformation is seldom explored. To address this problem, we propose a novel Siamese tracking framework, which enables to dynamically transform the template-candidate features to a more discriminative viewpoint for similarity matching. Specifically, we reformulate the template-candidate matching problem of the conventional Siamese tracker from the perspective of Lucas-Kanade (LK) image alignment approach. A Lucas-Kanade network (LKNet) is proposed and incorporated to the Siamese architecture to learn aligned feature representations in data-driven trainable manner, which is able to enhance the model adaptability in challenging scenarios. Within this framework, we propose two Siamese trackers named LK-Siam and LK-SiamRPN to validate the effectiveness. Extensive experiments conducted on the prevalent datasets show that the proposed method is more competitive over a number of state-of-the-art methods.
Siyuan Yao, Xiaoguang Han 0001, Hua Zhang 0008, Xiao Wang 0017, Xiaochun Cao
IEEE Trans. Image Process.1
2021 Robust Online Tracking via Contrastive Spatio-Temporal Aware Network
abstract
Existing tracking-by-detection approaches using deep features have achieved promising results in recent years. However, these methods mainly exploit feature representations learned from individual static frames, thus paying little attention to the temporal smoothness between frames. This easily leads trackers to drift in the presence of large appearance variations and occlusions. To address this issue, we propose a two-stream network to learn discriminative spatio-temporal feature representations to represent the target objects. The proposed network consists of a Spatial ConvNet module and a Temporal ConvNet module. Specifically, the Spatial ConvNet adopts 2D convolutions to encode the target-specific appearance in static frames, while the Temporal ConvNet models the temporal appearance variations using 3D convolutions and learns consistent temporal patterns in a short video clip. Then we propose a proposal refinement module to adjust the predicted bounding box, which can make the target localizing outputs to be more consistent in video sequences. In addition, to improve the model adaptation during online update, we propose a contrastive online hard example mining (OHEM) strategy, which selects hard negative samples and enforces them to be embedded in a more discriminative feature space. Extensive experiments conducted on the OTB, Temple Color and VOT benchmarks demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods.
Siyuan Yao, Hua Zhang 0008, Wenqi Ren, Chao Ma 0004, Xiaoguang Han 0001, Xiaochun Cao
IEEE Trans. Image Process.1
2020 Efficient Adversarial Attacks for Visual Object Tracking
Siyuan Liang 0004, Xingxing Wei 0001, Siyuan Yao, Xiaochun Cao
ECCV (26)3