EDBT 2026 Demo / reviewers in the wild / expert
Si-Yu Xia
dblp:65/7819 · also Siyu Xia
· DBLP profile ↗
67ranked-venue papers
6as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 4 first-author · 23 since 2021Artificial intelligence and machine learning · 37 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorComputer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A survey of recent advances in adversarial attack and defense on vision-language models
Neeresh Kumar Perla, Afia Sajeeda, Si-Yu Xia, Ming Shao |
Neural Networks | 4 |
| 2025 | Supportive Negatives Spectral Augmentation for Source-Free Cross-Domain SegmentationabstractSource-free domain adaptation (SFDA) aims to transfer knowledge from the well-trained source model and optimize it to adapt target data distribution. SFDA methods are suitable for medical image segmentation task due to its data-privacy protection and achieve promising performances. However, cross-domain distribution shift makes it difficult for the adapted model to provide accurate decisions on several hard instances and negatively affects model generalization. To overcome this limitation, a novel method `supportive negatives spectral augmentation' (SNSA) is presented in this work. Concretely, SNSA includes the instance selection mechanism to automatically discover a few hard samples for which source model produces incorrect predictions. And, active learning strategy is adopted to re-calibrate their predictive masks. Moreover, SNSA deploys the spectral augmentation between hard instances and others to encourage source model to gradually capture and adapt the attributions of target distribution. Considerable experimental studies demonstrate that annotating merely 4%~5% of negative instances from the target domain significantly improves segmentation performance over previous methods. Kexin Zheng, Haifeng Xia, Si-Yu Xia, Ming Shao, Zhengming Ding |
AAAI | 3 |
| 2025 | PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit RemeshingabstractPhotorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, existing methods for monocular full-body reconstruction, typically relying on front and/or predicted back view, still struggle with satisfactory performance due to the ill-posed nature of the problem and sophisticated self-occlusions. In this paper, we propose PSHuman, a novel framework that explicitly reconstructs human meshes utilizing priors from the multiview diffusion model. It is found that directly applying multiview diffusion on single-view human images leads to severe geometric distortions, especially on generated faces. To address it, we propose a cross-scale diffusion that models the joint probability distribution of global full-body shape and local facial characteristics, enabling identity-preserved novel-view generation without geometric distortion. Moreover, to enhance cross-view body shape consistency of varied human poses, we condition the generative model on parametric models (SMPL-X), which provide body priors and prevent unnatural views inconsistent with human anatomy. Leveraging the generated multiview normal and color images, we present SMPLX-initialized explicit human carving to recover realistic textured human meshes efficiently. Extensive experiments on CAPE and THuman2.1 demonstrate PSHuman’s superiority in geometry details, texture fidelity, and generalization capability. Wangguandong Zheng, Yuan Liu 0025, Tao Yu 0007, Yangguang Li 0001, Xingqun Qi, Xiaowei Chi, Si-Yu Xia, Yan-Pei Cao 0001, Wei Xue 0002, Wenhan Luo, Yike Guo |
CVPR | 8 |
| 2025 | Learning Macroeconomic Policies Through Dynamic Stackelberg Mean-Field GamesabstractMacroeconomic outcomes emerge from individuals’ decisions, making it essential to model how agents interact with macro policy via consumption, investment, and labor choices. We formulate this as a dynamic Stackelberg game: the government (leader) sets policies, and agents (followers) respond by optimizing their behavior over time. Unlike static models, this dynamic formulation captures temporal dependencies and strategic feedback critical to policy design. However, as the number of agents increases, explicitly simulating all agent–agent and agent–government interactions becomes computationally infeasible. To address this, we propose the Dynamic Stackelberg Mean Field Game (DSMFG) framework, which approximates these complex interactions via agent–population and government–population couplings. This approximation preserves individual-level feedback while ensuring scalability, enabling DSMFG to jointly model three core features of real-world policy-making: dynamic feedback, asymmetry, and large-scale. We further introduce Stackelberg Mean Field Reinforcement Learning (SMFRL), a data-driven algorithm that learns the leader’s optimal policies while maintaining personalized responses for individual agents. Empirically, we validate our approach in a large-scale simulated economy, where it scales to 1,000 agents (vs. 100 in prior work) and achieves a 4× GDP gain over classical economic methods and a 19× improvement over the static 2022 U.S. federal income tax policy. Qirui Mi, Chengdong Ma, Si-Yu Xia, Yan Song 0003, Mengyue Yang, Jun Wang 0012, Haifeng Zhang 0002 |
ECAI | 4 |
| 2025 | Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMsabstractWe show that large language models (LLMs) exhibit an internal chain-of-thought: they sequentially decompose and execute composite tasks layer-by-layer.Two claims ground our study: (i) distinct subtasks are learned at different network depths, and (ii) these subtasks are executed sequentially across layers.On a benchmark of 15 two-step composite tasks, we employ layer-from context-masking and propose a novel cross-task patching method, confirming (i).To examine claim (ii), we apply LogitLens to decode hidden states, revealing a consistent layerwise execution pattern.We further replicate our analysis on the real-world TRACE benchmark, observing the same stepwise dynamics.Together, our results enhance LLMs transparency by showing their capacity to internally plan and execute subtasks (or instructions), opening avenues for fine-grained, instruction-level activation steering. Junzhuo Li, Si-Yu Xia, Xuming Hu |
EMNLP | 3 |
| 2025 | IPNet: Interpretable Prototype Network for Multi-Source Domain AdaptationabstractMulti-source domain adaptation (MSDA) borrows intrinsic knowledge from well-annotated source domains to identify target visual signals. The main challenges are effectively mitigating cross-domain shift and extracting discriminative target features via the suitable source semantics. To overcome them, this paper proposes a novel Interpretable Prototype Network (IPNet) with channel-wise augmentation and multi-domain prototype mechanism. Specifically, IPNet explores the parameterized channel fusion paradigm across multiple source domains and target one to generate intermediate instances and achieve beneficial alignment. Moreover, IPNet analyzes contributions of source domains with interpretable learning approach and adjusts their effects on representations of target signals. Extensive experiments on three MSDA benchmark datasets suggest the advantages of our IPNet over others and exhibit the path of knowledge transfer. Haifeng Xia, Si-Yu Xia, Ming Shao, Zhengming Ding |
ICASSP | 3 |
| 2025 | UniMLVG: Unified Framework for Multi-View Long Video Generation with Comprehensive Control Capabilities for Autonomous DrivingabstractThe creation of diverse and realistic driving scenarios has become essential to enhance perception and planning capabilities of the autonomous driving system. However, generating long-duration, surround-view consistent driving videos remains a significant challenge. To address this, we present UniMLVG, a unified framework designed to generate extended street multi-perspective videos under precise control. By integrating single- and multi-view driving videos into the training data, our approach updates a DiT-based diffusion model equipped with cross-frame and cross-view modules across three stages with multi training objectives, substantially boosting the diversity and quality of generated visual content. Importantly, we propose an innovative explicit viewpoint modeling approach for multi-view video generation to effectively improve motion transition consistency. Capable of handling various input reference formats (e.g., text, images, or video), our UniMLVG generates high-quality multi-view videos according to the corresponding condition constraints such as 3D bounding boxes or frame-level text descriptions. Compared to the best models with similar capabilities, our framework achieves improvements of 48.2% in FID and 35.2% in FVD. Zehuan Wu, Jingcheng Ni, Haifeng Xia, Si-Yu Xia |
ICCV | 7 |
| 2025 | RoBiFusion: A Robust and Bidirectional Interaction Camera-LiDAR 3D Object Detection FrameworkabstractCamera-LiDAR 3D object detection is currently becoming a crucial component in the field of autonomous driving perception. However, previous models only performed feature fusion in the deep-level BEV hierarchy when dealing with camera-LiDAR feature fusion. This approach lacks interaction with the shallow-level sensor features, which is beneficial in constructing the corresponding BEV features. However, a simple shallow-level feature interaction can introduce sensor noise caused by intrinsic and extrinsic camera calibration errors. To address this, we propose RoBiFusion, a novel camera-LiDAR 3D object detection framework designed for effective sensor feature interaction and mitigating sensor noise interference. This framework consists of three submodules: the Camera-LiDAR Feature Matching module, the LiDAR-to-Camera module, and the Camera-to-LiDAR module. Firstly, in the Camera-LiDAR Feature Matching module, we use the cross-attention module to dynamically match the camera features and the LiDAR features, which solves the problem of feature inconsistency caused by noise in the camera's intrinsic and extrinsic parameters. Secondly, in the LiDAR-to-Camera module, we propose a novel depth representation that can effectively mitigate LiDAR noise interference. Thirdly, in the Camera-to-LiDAR module, we introduce deformable attention to help LiDAR feature capture instance-level semantic features. Additionally, we design a novel differentiable and efficient grid sample module to accelerate the process since the bilinear grid sample module in deformable attention is time-consuming and not deployment-friendly. We compared RoBiFusion to the state-of-the-art BEVFusion on the nuScenes dataset and found that RoBiFusion surpasses BEVFusion by 1.5% mAP and 2.4% NDS. Furthermore, we designed a series of ablation experiments to verify the effectiveness of the aforementioned modules. Xubin Wen, Haifeng Xia, Zhengming Ding, Si-Yu Xia |
ICRA | 4 |
| 2025 | QuantBEVFusion: A Fully Quantized Framework for LiDAR-camera 3D Object Detectionabstract3D object detection is essential for robust environmental perception in autonomous driving and robotics. While LiDAR-camera fusion methods offer high accuracy, their computational complexity hinders deployment on resource-constrained edge devices. To address this, we introduce QuantBEVFusion, a fully quantized 3D object detection framework that prioritizes both quantization and operator optimization. Our approach tackles the inherent asymmetry between LiDAR and camera data in pillar-based models by incorporating a novel pillar bird’s-eye view (BEV) encoder, significantly boosting performance. Furthermore, we introduce 1) an optimized LiDAR input processing method that filters noise and enables per-tensor quantization; 2) an improved sparse feature quantization process with log-histogram balancing, adaptive bin widths, and distillation loss for enhanced accuracy; and 3) a deployment-friendly 3D-to-2D transformation operator facilitating fixed-point implementation. Extensive experiments demonstrate that QuantBEVFusion achieves state-of-the-art quantization performance while maintaining accuracy suitable for real-time applications on edge devices. Xubin Wen, Ming Shao, Libo Sun 0001, Wenhu Qin, Si-Yu Xia |
IJCNN | 5 |
| 2025 | EconGym: A Scalable AI Testbed with Diverse Economic TasksabstractArtificial intelligence (AI) has become a powerful tool for economic research, enabling large-scale simulation and policy optimization. However, applying AI effectively requires simulation platforms for scalable training and evaluation—yet existing environments remain limited to simplified, narrowly scoped tasks, falling short of capturing complex economic challenges such as demographic shifts, multi-government coordination, and large-scale agent interactions.To address this gap, we introduce EconGym, a scalable and modular testbed that connects diverse economic tasks with AI algorithms. Grounded in rigorous economic modeling, EconGym implements 11 heterogeneous role types (e.g., households, firms, banks, governments), their interaction mechanisms, and agent models with well-defined observations, actions, and rewards. Users can flexibly compose economic roles with diverse agent algorithms to simulate rich multi-agent trajectories across 25+ economic tasks for AI-driven policy learning and analysis.Experiments show that EconGym supports diverse and cross-domain tasks—such as coordinating fiscal, pension, and monetary policies—and enables benchmarking across AI, economic methods, and hybrids. Results indicate that richer task composition and algorithm diversity expand the policy space, while AI agents guided by classical economic methods perform best in complex settings. EconGym also scales to 100k agents with high realism and efficiency. Qirui Mi, Qipeng Yang, Zijun Fan, Wentian Fan, Heyang Ma, Chengdong Ma, Si-Yu Xia, Bo An 0001, Jun Wang 0012, Haifeng Zhang 0002 |
NeurIPS | 7 |
| 2025 | ACL-SAR: model agnostic adversarial contrastive learning for robust skeleton-based action recognition
Jiaxuan Zhu, Ming Shao, Libo Sun 0001, Si-Yu Xia |
Vis. Comput. | 4 |
| 2024 | Discriminative Pattern Calibration Mechanism for Source-Free Domain AdaptationabstractSource-free domain adaptation (SFDA) assumes that model adaptation only accesses the well-learned source model and unlabeled target instances for knowledge trans-fer. However, cross-domain distribution shift easily triggers invalid discriminative semantics from source model on rec-ognizing the target samples. Hence, understanding the specific content of discriminative pattern and adjusting their representation in target domain become the important key to overcome SFDA. To achieve such a vision, this paper proposes a novel explanation paradigm “Discriminative Pattern Calibration (DPC)” mechanism on solving SFDA issue. Concretely, DPC first utilizes learning network to infer the discriminative regions on the target images and specifically emphasizes them in feature space to enhance their representation. Moreover, DPC relies on the attention-reversed mixup mechanism to augment more samples and improve the robustness of the classifier. Considerable experimental results and studies suggest that the effectiveness of our DPC in enhancing the performance of existing SFDA baselines. Haifeng Xia, Si-Yu Xia, Zhengming Ding |
CVPR | 2 |
| 2024 | CSTalk: Correlation Supervised Speech-driven 3D Emotional Facial Animation GenerationabstractSpeech-driven 3D facial animation technology has been developed for years, but its practical application still lacks expectations. The main challenges lie in data limitations, lip alignment, and the naturalness of facial expressions. Although lip alignment has seen many related studies, existing methods struggle to synthesize natural and realistic expressions, resulting in a mechanical and stiff appearance of facial animations. Even with some research extracting emotional features from speech, the randomness of facial movements limits the effective expression of emotions. To address this issue, this paper proposes a method called CSTalk (Correlation Supervised) that models the correlations among different regions of facial movements and supervises the training of the generative model to generate realistic expressions that conform to human facial motion patterns. To generate more intricate animations, we employ a rich set of control parameters based on the metahuman character model and capture a dataset for five different emotions. We train a generative network using an autoencoder structure and input an emotion embedding vector to achieve the generation of user-control expressions. Experimental results demonstrate that our method outperforms existing state-of-the-art methods. Xiangyu Liang, Wenlin Zhuang, Tianyong Wang, Guangxing Geng, Guangyue Geng, Haifeng Xia, Si-Yu Xia |
FG | 7 |
| 2024 | Cross-Block Fine-Grained Semantic Cascade for Skeleton-Based Sports Action RecognitionabstractHuman action video recognition has recently attracted more attention in applications such as video security and sports posture correction. Popular solutions, including graph convolutional networks (GCNs) that model the human skeleton as a spatiotemporal graph, have proven very effective. GCNs-based methods with stacked blocks usually utilize top-layer semantics for classification/annotation purposes. Although the global features learned through the procedure are suitable for the general classification, they have difficulty capturing fine-grained action change across adjacent frames - decisive factors in sports actions. In this paper, we propose a novel “Cross-block Fine-grained Semantic Cascade (CFSC)” module to overcome this challenge. In summary, the proposed CFSC progressively integrates shallow visual knowledge into high-level blocks to allow networks to focus on action details. In particular, the CFSC module utilizes the GCN feature maps produced at different levels, as well as aggregated features from proceeding levels to consolidate fine-grained features. In addition, a dedicated temporal convolution is applied at each level to learn short-term temporal features, which will be carried over from shallow to deep layers to maximize the leverage of low-level details. This cross-block feature aggregation methodology, capable of mitigating the loss of fine-grained information, has resulted in improved performance. Last, FD-7, a new action recognition dataset for fencing sports, was collected and will be made publicly available. Experimental results and empirical analysis on public benchmarks (FSD-10) and self-collected (FD-7) demonstrate the advantage of our CFSC module on learning discriminative patterns for action classification over others. Haifeng Xia, Libo Sun 0001, Ming Shao, Si-Yu Xia |
FG | 6 |
| 2024 | Embedded Representation Learning Network for Animating Styled Video PortraitabstractThe talking head generation recently attracted considerable attention due to its widespread application prospects, especially for digital avatars and 3D animation design. Inspired by this practical demand, several works explored Neural Radiance Fields (NeRF) to synthesize the talking heads. However, these methods based on NeRF face two challenges: (1) Difficulty in generating style-controllable talking heads. (2) Displacement artifacts around the neck in rendered images. To overcome these two challenges, we propose a novel generative paradigm Embedded Representation Learning Network (ERLNet) with two learning stages. First, the audio-driven FLAME (ADF) module is constructed to produce facial expression and head pose sequences synchronized with content audio and style video. Second, given the sequence deduced by the ADF, one novel dual-branch fusion NeRF (DBF-NeRF) explores these contents to render the final images. Extensive empirical studies demonstrate that the collaboration of these two stages effectively facilitates our method to render a more realistic talking head than the existing algorithms. Tianyong Wang, Xiangyu Liang, Wangguandong Zheng, Dan Niu, Haifeng Xia, Si-Yu Xia |
FG | 6 |
| 2024 | Autonomous Generative Feature Replay for Non-Exemplar Class-Incremental LearningabstractDeep neural networks have been successfully applied in many computer vision tasks. However, these models suffer catastrophic forgetting when learning new knowledge incrementally. To overcome the stability-plasticity dilemma, class incremental learning (CIL) has been widely discussed recently. The state-of-the-art CIL methods mainly leverage additional exemplar sets, thus memory costly and may raise privacy issues. To that end, we propose an autonomous generative feature replay (AGFR) framework without using exemplar sets. It consists of three modules: the feature extractor module, the feature generator module, and the unified classification module. First, to stabilize features over tasks, robust feature extractors are learned in a self-supervised manner and thus generalize well to unseen data. Second, instead of using exemplar sets or producing raw images, we propose an autonomous generative feature replay scheme to constantly update unified classifier in CIL without saving any image data. This strategy avoids overwhelming memory usage or poor quality of the generated raw images. Experiments demonstrate that our method achieves state-of-the-art performance in terms of average classification accuracy.⋆ Yinjie Zhang, Ming Shao, Wenlong Shi, Haifeng Xia, Si-Yu Xia |
ICASSP | 5 |
| 2024 | DEITalk: Speech-Driven 3D Facial Animation with Dynamic Emotional Intensity Modeling
Haifeng Xia, Guangxing Geng, Guangyue Geng, Si-Yu Xia, Zhengming Ding |
ACM Multimedia | 5 |
| 2024 | Sketch3D: Style-Consistent Guidance for Sketch-to-3D Generation
Wangguandong Zheng, Haifeng Xia, Libo Sun 0001, Ming Shao, Si-Yu Xia, Zhengming Ding |
ACM Multimedia | 6 |
| 2024 | Few-shot Shape Recognition by Learning Deep Shape-aware FeaturesabstractTraditional shape descriptors have been gradually replaced by convolutional neural networks due to their superior performance in feature extraction and classification. The state-of-the-art methods recognize object shapes via image reconstruction or pixel classification. However, these methods are biased toward texture information and overlook the essential shape descriptions, thus, they fail to generalize to unseen shapes. We are the first to propose a few-shot shape descriptor (FSSD) to recognize object shapes given only one or a few samples. We employ an embedding module for FSSD to extract transformation-invariant shape features. Secondly, we develop a dual attention mechanism to decompose and reconstruct the shape features via learnable shape primitives. In this way, any shape can be formed through a finite set basis, and the learned representation model is highly interpretable and extendable to unseen shapes. Thirdly, we propose a decoding module to include the supervision of shape masks and edges and align the original and reconstructed shape features, enforcing the learned features to be more shape-aware. Lastly, all the proposed modules are assembled into a few-shot shape recognition scheme. Experiments on five datasets show that our FSSD significantly improves the shape classification compared to the state-of-the-art under the few-shot setting. Wenlong Shi, Changsheng Lu, Ming Shao, Yinjie Zhang, Si-Yu Xia, Piotr Koniusz |
WACV | 5 |
| 2024 | Energy Minimization of RIS-Assisted Cooperative UAV-USV MEC NetworkabstractUnmanned surface vehicles (USVs) are becoming increasingly significant in fulfilling integrated sensing, computing, and communication with the emergence of bidirectional computation tasks. However, Quality-of-Service provisioning is still challenging since USVs are restricted with limited onboard resources and direct links between them and shore-based terrestrial base stations (TBSs) are frequently blocked. This article proposes a novel reconfigurable intelligent surface (RIS)-assisted cooperative unmanned aerial vehicle (UAV)–USV mobile-edge computing (MEC) network architecture, where RIS-mounted tethered UAV (TUAV) and rotary-wing UAVs (RUAVs) are collaboratively utilized to serve USVs. RUAVs energy minimization is formulated by jointly considering TUAV hovering altitude, RIS phase-shift vector, RUAV service selection indicator, and RUAVs turning points. A heuristic solution is proposed to tackle the formulated problem, where the original problem is first decoupled into three subproblems, e.g., the joint optimization of RIS phase-shift vector and TUAV hovering altitude subproblem, RUAVs service selection indicator subproblem, and RUAVs turning points subproblem, each of which is solved by the proposed modified alternative direction method of multiplier (ADMM) algorithm, the proposed enhanced simulated annealing (ESA) algorithm and the proposed successive convex approximation (SCA)-based algorithm. In this way, the challenging problem can be efficiently solved iteratively. The results show that the proposed solution can decrease RUAVs energy consumption by nearly 29% compared to numerous selected advanced algorithms. Moreover, the performance of the proposed solution regarding typical penalty coefficients and number of RIS reflecting elements is investigated. Yangzhe Liao, Yuanyan Song, Si-Yu Xia, Yi Han 0007, Ning Xu 0006, Xiaojun Zhai |
IEEE Internet Things J. | 3 |
| 2023 | A Graph Convolutional Siamese Network for the Assessment and Recognition of Physical Rehabilitation Exercises
Chengxian Li, Xichong Ling, Si-Yu Xia |
ICANN (4) | 3 |
| 2023 | Graphrpe: Relative Position Encoding Graph Transformer for 3d Human Pose EstimationabstractThe graph neural network has been playing increasingly important roles in 2D-3D lifting based single-frame human pose estimation. However, it still suffers from inferior modeling of local and global associations between 2D nodes. To this end, we introduce two relative position encoding approaches and propose a novel graph-transformer structure "GraphRPE." The model consists of two components: graph relative position encoding (GRPE) and universal relative position encoding (URPE). GRPE embeds graph structure prior information into the attention map to correct the self-attention weights and prompt appropriate interactions between 2D nodes, both locally and globally. On the other hand, URPE introduces the Toeplitz matrix to address the limited representation capability of the transformer. In addition, we investigate several graph edge types and their impacts on the results. Extensive experimental results demonstrate that our proposed method achieves SOTA performance on the Human3.6M dataset. Junjie Zou, Ming Shao, Si-Yu Xia |
ICIP | 3 |
| 2023 | Image Inpainting Network Based on Deep Fusion of Texture and Structure
Huilai Liang, Xichong Ling, Si-Yu Xia |
ICPRAM | 3 |
| 2023 | Performance Evaluation of Flight Energy Consumption of UAVs in IRS-assisted UAV SystemsabstractThe research regarding the cooperation between unmanned aerial vehicles (UAVs) and intelligent reflecting surface (IRS) has attracted great attenuation in recent years. However, one of the key technical weaknesses is the limited service duration of onboard battery-empowered UAVs. Aiming to prolong UAVs flight time, in this paper, UAVs flight energy consumption minimization problem is formulated along with a list of constraints. To solve the challenging problem, the original problem is decoupled into two subproblems, where the first one is solved by using the proposed modified differential evolution algorithm and the second subproblem is resolved by means of the proposed modified grey wolf algorithm. The results verify the effectiveness of the proposed solution. The outcome obtained in this paper can be further extended to reliable indoor communication services, such as library indoor positioning and navigation and library emergency communications. Xiuyi Luo, Chongrui Lu, Siyi Ouyang, Si-Yu Xia |
TrustCom | 4 |
| 2023 | High-precision skeleton-based human repetitive action countingabstractAbstract A novel counting model is presented by the authors to estimate the number of repetitive actions in temporal 3D skeleton data. As per the authors’ knowledge, this is the first work of this kind using skeleton data for high‐precision repetitive action counting. Different from existing works on RGB video data, the authors’ model follows a bottom‐up pipeline to clip the sub‐action first followed by robust aggregation in inference. First, novel counting loss functions and robust inference with backtracking is proposed to pursue precise per‐frame count as well as overall count with boundary frames. Second, an efficient synthetic approach is proposed to augment skeleton data in training and thus avoid time‐consuming repetitive action data collection work. Finally, a challenging human repetitive action counting dataset named VSRep is collected with various types of action to evaluate the proposed model. Experiments demonstrate that the proposed counting model outperforms existing video‐based methods by a large margin in terms of accuracy in real‐time inference. Chengxian Li, Ming Shao, Si-Yu Xia |
IET Comput. Vis. | 4 |
| 2022 | ElDet: An Anchor-Free General Ellipse Object Detector
Tian Wang 0001, Changsheng Lu, Ming Shao, Si-Yu Xia |
ACCV (3) | 5 |
| 2022 | TMCR: A Twin Matching Networks for Chinese Scene Text Retrieval
Zhiheng Peng, Ming Shao, Si-Yu Xia |
PRCV (3) | 3 |
| 2022 | Mask removal : Face inpainting via attributes
Yefan Jiang, Fan Yang 0080, Zhangxing Bian, Changsheng Lu, Si-Yu Xia |
Multim. Tools Appl. | 5 |
| 2022 | Evolutionary Multitask Optimization With Adaptive Knowledge TransferabstractEvolutionary multitask optimization (EMTO) studies how to simultaneously solve multiple optimization tasks via evolutionary algorithms (EAs) while making the useful knowledge acquired from solving one task to assist solving other tasks, aiming to improve the overall performance of solving each individual task. Recent years have seen a large body of EMTO works based on different kinds of EAs and studying one or more aspects in how to represent, extract, transfer, and reuse knowledge. A key challenge to EMTO is the occurrence of negative knowledge transfer between tasks, which becomes severer when the total number of tasks increases. To address this issue, we propose an adaptive EMTO (AEMTO) framework. This framework can adapt knowledge transfer frequency, knowledge source selection, and knowledge transfer intensity in a synergistic way to make the best use of knowledge transfer, especially when facing many tasks. We implement the proposed AEMTO framework and evaluate our implementation on three suites of MTO problems with 2, 10, and 50 tasks and one real-world MTO problem with 2000 tasks in comparison to several state-of-the-art EMTO methods with certain adaptation strategies regarding knowledge transfer and the single-task optimization counterpart of the proposed method. Experimental results have demonstrated the effectiveness of the adaptive knowledge transfer strategies used in AEMTO and the overall performance superiority of AEMTO. A. K. Qin 0001, Si-Yu Xia |
IEEE Trans. Evol. Comput. | 3 |
| 2022 | Music2Dance: DanceNet for Music-Driven Dance GenerationabstractSynthesize human motions from music (i.e., music to dance) is appealing and has attracted lots of research interests in recent years. It is challenging because of the requirement for realistic and complex human motions for dance, but more importantly, the synthesized motions should be consistent with the style, rhythm, and melody of the music. In this article, we propose a novel autoregressive generative model, DanceNet, to take the style, rhythm, and melody of music as the control signals to generate 3D dance motions with high realism and diversity. Due to the high long-term spatio-temporal complexity of dance, we propose the dilated convolution to improve the receptive field, and adopt the gated activation unit as well as separable convolution to enhance the fusion of motion features and control signals. To boost the performance of our proposed model, we capture several synchronized music-dance pairs by professional dancers and build a high-quality music-dance pair dataset. Experiments have demonstrated that the proposed method can achieve state-of-the-art results. Wenlin Zhuang, Congyi Wang, Jinxiang Chai, Yangang Wang 0001, Ming Shao, Si-Yu Xia |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | Adversarial Attacks on Kinship Verification using TransformerabstractVisual kinship verification is one of the key research problems in computer vision with significant progress made in the past decade. Meanwhile, the harnessing of visual kinship models may lead to personal privacy leaking and raise people's concerns, especially the heavy social media users. One promising countermeasure is to overlay additional noise on images through the adversarial attack on kinship verification models to protect personal privacy. Motivated by the recent success of Transformer models in visual tasks, we propose a novel Transformer-based adversarial attack method named “Kinship-advTransGAN” towards the attack on kinship verification model. Essentially, Kinship-advTransGAN replaces the well-established CNN structure in conventional advGAN by TransGAN to generate adversarial samples with more sparse noise but comparable successful attacking rate. We verify our proposed method on a few open benchmarks, including FIW datasets and Kaggle Kinship Verification Challenges. Among these challenging tasks, we achieved surprisingly good performance: over 90% successful attacking rate on FIW datasets and 76.86% successful rate of attacks on Kaggle Kinship Verification Challenge, but with less visually perceivable noise on face images. Jiaxuan Zhu, Ming Shao, Hong Pan 0001, Si-Yu Xia |
FG | 5 |
| 2021 | DNA-Net: Age and Gender Aware Kin Face SynthesizerabstractVisual kinship verification aims to detect blood relatives in facial images. Its practical application have motivated many researchers to focus on the topic as of recent. In this paper, we focus on a new view of visual kinship technology: kin-based face generation. Specifically, we propose a two-stage kin-face generation model to predict the appearance of a child given a pair of parents. The first stage includes a deep generative adversarial auto-encoder conditioned on ages and genders to map between facial appearance and high-level features. The second stage is our proposed DNA-Net, which serves as a transformation between the deep and genetic features based on a random selection process to fuse genes of a parent pair to form the genes of a child. We demonstrate the effectiveness of the proposed method quantitatively and qualitatively. Experiments validate that the proposed model synthesizes convincing kin-faces using both subjective and objective standards. Pengyu Gao, Joseph P. Robinson, Jiaxuan Zhu, Ming Shao, Si-Yu Xia |
ICME | 6 |
| 2021 | Multi-level Feature Selection for Oriented Object Detection
Yefan Jiang, Zhangxing Bian, Fan Yang 0080, Si-Yu Xia |
ICPRAM | 5 |
| 2021 | Joint exploring of risky labeled and unlabeled samples for safe semi-supervised clustering
Haitao Gan, Si-Yu Xia, Xiaobin Xu 0002 |
Expert Syst. Appl. | 3 |
| 2020 | Localin Reshuffle Net: Toward Naturally and Efficiently Facial Image Blending
Chengyao Zheng, Si-Yu Xia, Joseph P. Robinson, Changsheng Lu, Wayne Wu, Chen Qian 0006, Ming Shao |
ACCV (5) | 2 |
| 2020 | Recognizing Families In the Wild (RFIW): The 4th EditionabstractRecognizing Families In the Wild (RFIW)- an annual large-scale, multi-track automatic kinship recognition evaluation- supports various visual kin-based problems on scales much higher than ever before. Organized in conjunction with the as a Challenge, RFIW provides a platform for publishing original work and the gathering of experts for a discussion of the next steps. This paper summarizes the supported tasks (i.e., kinship verification, tri-subject verification, and search & retrieval of missing children) in the evaluation protocols, which include the practical motivation, technical background, data splits, metrics, and benchmark results. Furthermore, top submissions (i.e., leader-board stats) are listed and reviewed as a high-level analysis on the state of the problem. In the end, the purpose of this paper is to describe the 2020 RFIW challenge, end-to-end, along with forecasts in promising future directions. Joseph P. Robinson, Yu Yin 0001, Zaid Khan 0001, Ming Shao, Si-Yu Xia, Michael Stopa, Samson Timoner, Matthew Turk 0001, Rama Chellappa, Yun Fu 0001 |
FG | 5 |
| 2020 | Hand-3d-Studio: A New Multi-View System for 3d Hand ReconstructionabstractThis paper proposes a new system named as Hand-3D-Studio to capture the 3D hand pose and shape information. Our system includes 15 synchronized DSLR cameras, which can acquire high quality multi-view 4K resolution color images in a circular manner. We then introduce a 2D hand keypoints guided iterative pixel growth matching strategy for 3D reconstruction, where the 2D keypoints are obtained via convolution neural network. We find that the pre-detected 2D hand keypoints can greatly remove the matching noise, and thus improve the performance of reconstruction. After that, a non-rigid iterative closest points algorithm is performed to drive a template hand to fit the point clouds and register all the hand meshes. As a consequence, we captured more than 20K high quality hand color images, annotated 2D hand key-points, 3D point cloud as well as the registered hand meshes (>200). All the data are public on the website http://www.yangangwang.com for future research. Tianyao Wang, Si-Yu Xia, Yangang Wang 0001 |
ICASSP | 3 |
| 2020 | Finding Achilles' Heel: Adversarial Attack on Multi-modal Action RecognitionabstractNeural network-based models are notoriously known for their adversarial vulnerability. Recent adversarial machine learning mainly focused on images, where a small perturbation can be simply added to fool the learning model. Very recently, this practice has been explored in human action video attacks by adding perturbation to key frames. Unfortunately, frame selection is usually computationally expensive in run-time, and adding noises to all frames is unrealistic, either. In this paper, we present a novel yet efficient approach to address this issue. Multi-modal video data such as RGB, depth and skeleton data have been widely used for human action modeling, and they have been demonstrated with superior performance than a single modality. Interestingly, we observed that the skeleton data is more "vulnerable" under adversarial attack, and we propose to leverage this "Achilles' Heel" to attack multi-modal video data. In particular, first, an adversarial learning paradigm is designed to perturb skeleton data for a specific action under a black box setting, which highlights how body joints and key segments in videos are subject to attack. Second, we propose a graph attention model to explore the semantics between segments from different modalities and within a modality. Third, the attack will be launched in run-time on all modalities through the learned semantics. The proposed method has been extensively evaluated on multi-modal visual action datasets, including PKU-MMD and NTU-RGB+D to validate its effectiveness. Deepak Kumar 0008, Chetan Kumar, Chun-Wei Seah, Si-Yu Xia, Ming Shao |
ACM Multimedia | 4 |
| 2020 | Automatic Classification of Sleep Stages Based on Raw Single-Channel EEG
Kailin Xu, Si-Yu Xia, Guang Li 0001 |
PRCV (2) | 2 |
| 2020 | Highlight Removal in Facial Images
Si-Yu Xia, Zhangxing Bian, Changsheng Lu |
PRCV (1) | 2 |
| 2020 | A hybrid safe semi-supervised learning method
Haitao Gan, Si-Yu Xia |
Expert Syst. Appl. | 3 |
| 2020 | Deep transfer neural network using hybrid representations of domain discrepancy
Changsheng Lu, Chaochen Gu, Kaijie Wu 0002, Si-Yu Xia, Xin-Ping Guan |
Neurocomputing | 4 |
| 2020 | Arc-Support Line Segments Revisited: An Efficient High-Quality Ellipse DetectionabstractOver the years many ellipse detection algorithms spring up and are studied broadly, while the critical issue of detecting ellipses accurately and efficiently in real-world images remains a challenge. In this paper, we propose a valuable industry-oriented ellipse detector by arc-support line segments, which simultaneously reaches high detection accuracy and efficiency. To simplify the complicated curves in an image while retaining the general properties including convexity and polarity, the arc-support line segments are extracted, which grounds the successful detection of ellipses. The arc-support groups are formed by iteratively and robustly linking the arc-support line segments that latently belong to a common ellipse. Afterward, two complementary approaches, namely, locally selecting the arc-support group with higher saliency and globally searching all the valid paired groups, are adopted to fit the initial ellipses in a fast way. Then, the ellipse candidate set can be formulated by hierarchical clustering of 5D parameter space of initial ellipses. Finally, the salient ellipse candidates are selected and refined as detections subject to the stringent and effective verification. Extensive experiments on three public datasets are implemented and our method achieves the best F-measure scores compared to the state-of-the-art methods. The source code is available at https://github.com/AlanLuSun/High-quality-ellipse-detection. Changsheng Lu, Si-Yu Xia, Ming Shao, Yun Fu 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Weakly Supervised Vitiligo Segmentation in Skin Image through Saliency PropagationabstractVitiligo is a skin disorder where pale or white patches develop due to the lack or absence of melanocytes. Vitiligo affects around 0.5% to 1% of the world's population, and it may have a profound psychological impact on patients' quality of life. In this paper, we present a novel weakly supervised framework to segment vitiligo regions with high quality, which is a fundamental task for the assessment of vitiligo. The proposed framework starts with pre-training a classification network using only image-level labels. Then we observed that the activation map obtained from the image classification network could be further exploited and introduced into the saliency propagation process as useful information. Finally, the saliency propagation process is performed on the graph built on superpixels to obtain a meaningful saliency map. These three steps lead to a compelling yet elegant method. Moreover, we propose a new large vitiligo image dataset named Vit2019. To the best of our knowledge, this is currently the first dataset for image segmentation of vitiligo diseases. Experimental results demonstrate the superiority of the proposed model over state-of-the-arts. Zhangxing Bian, Si-Yu Xia, Ming Shao |
BIBM | 2 |
| 2019 | Fast Facial Image Analogy with Spatial GuidanceabstractThis paper proposes a novel method for fast facial image analogy with spatial guidance. Given a facial image A and another one B in a different style (color, tone, or texture), we are allowed to render the face in A with style B to output a stylized facial image A' in a timely manner. For traditional image analogy, such a process could take an unbearable time and considerable computing resources. In this paper, for the first time, we make it possible to do image analogy at a speed much faster than the state-of-the-art method. Specifically, we first extract deep image features using a VGG-19 encoder, and implement patch-match with the guidance of facial landmarks. Then Procrustes analysis is applied to accelerate the program as a coarse-to-fine strategy. We repeat the above process at each layer of VGG-19 and decode image A' from bottom to top. Experimental results show that our method not only spends much less time but also provides high-quality image analogy results. Moreover, our method can be naturally extended to many face-related applications including but not limited to face swapping, makeup and style to photo. Chengyao Zheng, Si-Yu Xia, Ming Shao, Yun Fu 0001 |
FG | 2 |
| 2019 | CircView: a visualization and exploration tool for circular RNAsabstractCircular RNAs (circRNAs) are novel rising stars of noncoding RNAs, which are highly abundant and evolutionarily conserved across species. Number of publications related to circRNAs increased sharply in recent years, representing emerging focuses in the field. Therefore, tools, pipelines and databases have been developed to identify and store circRNAs. However, there is no existing tool to visualize and explore circRNAs. Therefore, we introduce CircView, a user-friendly visualization tool for circRNAs detected from existing tools. CircView enables users to visualize circRNAs and to quantify number of samples with detected circRNAs. CircView allows users to explore circRNAs detected by unique or multiple tools. Furthermore, CircView allows users to view the regulatory elements, such as microRNA response elements and RNA-binding protein binding sites. CircView is a unique tool to visualize and explore circRNAs, which helps users to better understand potential functions of circRNAs and design the functional experiments. Jing Feng 0005, Si-Yu Xia, Jun Wang 0154, Fatma Muge Ozguc, Lijun Lei, Ruoshan Kong, Lixia Diao, Chunjiang He, Leng Han |
Briefings Bioinform. | 3 |
| 2018 | Album to Family Tree: A Graph Based Method for Family Relationship Recognition
Si-Yu Xia, Ming Shao, Yun Fu 0001 |
ACCV (2) | 2 |
| 2018 | Multi-scale Adaptive Structure Network for Human Pose Estimation from Color Images
Wenlin Zhuang, Cong Peng 0001, Si-Yu Xia, Yangang Wang 0001 |
ACCV (1) | 3 |
| 2018 | Deep Evolutionary 3D Diffusion Heat Maps for Large-pose Face Alignment
Bin Sun 0002, Ming Shao, Si-Yu Xia, Yun Fu 0001 |
BMVC | 3 |
| 2018 | Graph Based Family Relationship Recognition from a Single Image
Si-Yu Xia, Ming Shao |
PRICAI (1) | 2 |
| 2018 | Genome-wide characterization of lncRNAs in acute myeloid leukemiaabstractLong noncoding RNAs (lncRNAs) are a large family of noncoding RNAs that play a critical role in various normal bioprocesses as well as tumorigenesis. However, the expression patterns and biological functions of lncRNAs in acute leukemia have not been well studied. Here, we performed transcriptome-wide lncRNA expression profiling of acute myeloid leukemia (AML) patient samples, along with non-leukemia control hematopoietic samples. We found that lncRNAs were differentially expressed in AML samples relative to control samples. Notably, we identified that lncRNAs upregulated in AML (relative to the control samples) are associated with a lower degree of DNA methylation and a higher ratio of being bound by transcription factors such as SP1, STAT4, ATF-2 and ELK-1 compared with those downregulated in AML. Moreover, an enrichment of H3K4me3 and a depletion of H3K27me3 were observed in upregulated lncRNAs in AML. Expression patterns of three types of lncRNAs (antisense, enhancer and intergenic lncRNAs) have previously been characterized. Of the identified lncRNAs, we found that high expression level lncRNA LOC285758 is associated with the poor prognosis in AML patients. Furthermore, we found that LOC285758 regulates proliferation of AML cell lines by enhancing the expression of HDAC2, a key factor in carcinogenesis. Collectively, our study depicts a landscape of important lncRNAs in AML and provides novel potential therapeutic targets and prognostic markers for AML treatment. Lijun Lei, Si-Yu Xia, Jing Feng 0005, Yaqi Zhu, Linjian Xia, Lieping Guo, Ke Chen 0013, Hanyang Hu, Nupur Mittal, Guohua Yang, Zhijian Qian, Leng Han, Chunjiang He |
Briefings Bioinform. | 2 |
| 2017 | Circle detection by arc-support line segmentsabstractCircle detection is fundamental in both object detection and high accuracy localization in visual control systems. We propose a novel method for circle detection by analysing and refining arc-support line segments. The key idea is to use line segment detector to extract the arc-support line segments which are likely to make up the circle, instead of all line segments. Each couple of line segments is analyzed to form a valid pair and followed by generating initial circle set. Through the mean shift clustering, the circle candidates are generated and verified based on the geometric attributes of circle edge. Finally, twice circle fitting is applied to increase the accuracy for circle locating and radius measuring. The experimental results demonstrate that the proposed method performs better than other well known approaches on circles that are incomplete, occluded, blurry and over-illumination. Moreover, our method shows significant improvement in accuracy, robustness and efficiency on the industrial Printed Circuit Board (PCB) images as well as the synthesized, natural and complicated images. Changsheng Lu, Si-Yu Xia, Wanming Huang, Ming Shao, Yun Fu 0001 |
ICIP | 2 |
| 2017 | Family Photo Recognition via Multiple Instance LearningabstractFamily photo recognition is an important task in social media analytics. Previous methods use singleton global features and conventional binary classifiers to distinguish family group photos from non-family ones. Different from them, we propose a novel family recognition approach with three dedicated local representations under Multiple Instance Learning framework, where geometry, kinship and semantic features are integrated to overcome issues in the previous work. Experimental results show that our method achieves the state-of-the-art result among global-feature models. Junkang Zhang, Si-Yu Xia, Ming Shao, Yun Fu 0001 |
ICMR | 2 |
| 2017 | Comprehensive characterization of tissue-specific circular RNAs in the human and mouse genomesabstractCircular RNA (circRNA) is a group of RNA family generated by RNA circularization, which was discovered ubiquitously across different species and tissues. However, there is no global view of tissue specificity for circRNAs to date. Here we performed the comprehensive analysis to characterize the features of human and mouse tissue-specific (TS) circRNAs. We identified in total 302 853 TS circRNAs in the human and mouse genome, and showed that the brain has the highest abundance of TS circRNAs. We further confirmed the existence of circRNAs by reverse transcription polymerase chain reaction (RT-PCR). We also characterized the genomic location and conservation of these TS circRNAs and showed that the majority of TS circRNAs are generated from exonic regions. To further understand the potential functions of TS circRNAs, we identified microRNAs and RNA binding protein, which might bind to TS circRNAs. This process suggested their involvement in development and organ differentiation. Finally, we constructed an integrated database TSCD (Tissue-Specific CircRNA Database: http://gb.whu.edu.cn/TSCD) to deposit the features of TS circRNAs. This study is the first comprehensive view of TS circRNAs in human and mouse, which shed light on circRNA functions in organ development and disorders. Si-Yu Xia, Jing Feng 0005, Lijun Lei, Linjian Xia, Jun Wang 0154, Lingjun Liu, Leng Han, Chunjiang He |
Briefings Bioinform. | 1 |
| 2016 | A genetics-motivated unsupervised model for tri-subject kinship verificationabstractGiven a child's and a couple's facial photos, tri-subject kinship verification aims to determine the existence of blood relation between the child and the couple. Different from existing methods which model the kinship inheritance process among three persons in separate stages and only use simple features, this work establishes a simple model inspired by genetics to measure tri-subject kinship similarity in one step. Meanwhile, high-dimensional features are incorporated into this simple model to seek for better performance. Experiment results demonstrate the effectiveness of our approach. Junkang Zhang, Si-Yu Xia, Hong Pan 0001, A. K. Qin 0001 |
ICIP | 2 |
| 2016 | Robust road detection from a single imageabstractRoad detection from images is a challenging task in computer vision. Previous methods are not robust, because their features and classifiers cannot adapt to different circumstances. To overcome this problem, we propose to apply unsupervised feature learning for road detection. Specifically, we develop an improved encoding function and add a feature selection process to obtain robust and discriminative road features. Besides, a road segmentation algorithm is proposed to extract road regions from the learned feature maps, in which a tree structure is established to represent the hierarchical relations of various regions segmented by multiple thresholds, and a two-loop optimization is then employed to select the most stable regions as road areas. Experimental results on several challenging datasets justify the effectiveness of our method. Junkang Zhang, Si-Yu Xia, Kaiyue Lu, Hong Pan 0001, A. K. Qin 0001 |
ICPR | 2 |
| 2014 | Self-adaptive differential evolution with local search chains for real-parameter single-objective optimizationabstractDifferential evolution (DE), as a very powerful population-based stochastic optimizer, is one of the most active research topics in the field of evolutionary computation. Self-adaptive differential evolution (SaDE) is a well-known DE variant, which aims to relieve the practical difficulty faced by DE in selecting among many candidates the most effective search strategy and its associated parameters. SaDE operates with multiple candidate strategies and gradually adapts the employed strategy and its accompanying parameter setting via learning the preceding behavior of already applied strategies and their associated parameter settings. Although highly effective, SaDE concentrates more on exploration than exploitation. To enhance SaDE's exploitation capability while maintaining its exploration power, we incorporate local search chains into SaDE following two different paradigms (Lamarckian and Baldwinian) that differ in the ways of utilizing local search results in SaDE. Our experiments are conducted on the CEC-2014 real-parameter single-objective optimization testbed. The statistical comparison results demonstrate that SaDE with Baldwinian local search chains, armed with suitable parameter settings, can significantly outperform original SaDE as well as classic DE at any tested problem dimensionality. A. K. Qin 0001, Ke Tang 0001, Hong Pan 0001, Si-Yu Xia |
IEEE Congress on Evolutionary Computation | 4 |
| 2014 | Face Clustering in Photo AlbumabstractDigital photo management is becoming indispensable for the explosively growing family photo albums due to the rapid popularization of digital cameras and mobile phone cameras. An effective photo management system could accurately and efficiently group all faces of the same person into a small number of clusters. In this paper, we present a novel photo grouping method based on spectral theory. The key idea is to utilize prior information of family photo albums to improve the performance. First, an individual can only appear once in one photo, which works as the similarity constraint in our graph construction. Second, an individual cannot show more times than the number of photos in each album. That is, the size of a cluster for an individual is at most the number of photos in an album. We consider this constraint as a Minimum Cost Flow (MCF) linear network optimization problem and therefore propose a constrained K-Means for data clustering after graph embedding. Two metrics, i.e., accuracy (AC) and normalized mutual information metric (NMI), are used to evaluate the clustering performance. Extensive experimental results demonstrate the effectiveness of the proposed method. Si-Yu Xia, Hong Pan 0001, A. K. Qin 0001 |
ICPR | 1 |
| 2013 | Investigation of self-adaptive differential evolution on the CEC-2013 real-parameter single-objective optimization testbedabstractSelf-adaptive differential evolution (SaDE) is a wellknown DE variant, which has received considerable attention since it was developed. SaDE gradually adapts its trial vector generation strategy and the accompanying parameter setting via learning the preceding performance of multiple candidate strategies and their associated parameter settings. This work systematically investigates SaDE on the CEC-2013 real-parameter single-objective optimization testbed. Parameter sensitivity analysis is carried out by using advanced statistical hypothesis testing methods, aiming to detect statistically significantly superior parameter settings. This analysis reveals that SaDE is actually less sensitive to the parameter choice since quite a number of parameter settings can lead to the statistically significantly better performance than the other settings. Based on this finding, we report SaDE's performance using one of the parameter settings advocated by sensitivity analysis and statistically compare this performance with that of a widely used classic DE (DE/rand/1/bin). The comparison results significantly favor SaDE. A. K. Qin 0001, Xiaodong Li 0001, Hong Pan 0001, Si-Yu Xia |
IEEE Congress on Evolutionary Computation | 4 |
| 2012 | Improved generic categorical object detection fusing depth cue with 2D appearance and shape features
Hong Pan 0001, Si-Yu Xia, A. K. Qin 0001 |
ICPR | 3 |
| 2012 | Toward kinship verification using visual attributes
Si-Yu Xia, Ming Shao, Yun Fu 0001 |
ICPR | 1 |
| 2012 | Understanding Kin Relationships in a PhotoabstractThere is an urgent need to organize and manage images of people automatically due to the recent explosion of such data on the Web in general and in social media in particular. Beyond face detection and face recognition, which have been extensively studied over the past decade, perhaps the most interesting aspect related to human-centered images is the relationship of people in the image. In this work, we focus on a novel solution to the latter problem, in particular the kin relationships. To this end, we constructed two databases: the first one named UB KinFace Ver2.0, which consists of images of children, their young parents and old parents, and the second one named FamilyFace. Next, we develop a transfer subspace learning based algorithm in order to reduce the significant differences in the appearance distributions between children and old parents facial images. Moreover, by exploring the semantic relevance of the associated metadata, we propose an algorithm to predict the most likely kin relationships embedded in an image. In addition, human subjects are used in a baseline study on both databases. Experimental results have shown that the proposed algorithms can effectively annotate the kin relationships among people in an image and semantic context can further improve the accuracy. Si-Yu Xia, Ming Shao, Jiebo Luo 0001, Yun Fu 0001 |
IEEE Trans. Multim. | 1 |
| 2011 | Kinship Verification through Transfer Learning
Si-Yu Xia, Ming Shao, Yun Fu 0001 |
IJCAI | 1 |
| 2010 | Context-based embedded image compression using binary wavelet transform
Hong Pan 0001, Lizuo Jin, Xiao-Hui Yuan, Si-Yu Xia, Liang-Zheng Xia |
Image Vis. Comput. | 4 |
| 2009 | A binary wavelet-based scheme for grayscale image compressionabstractThis paper proposes a novel grayscale image compression approach using the binary wavelet transform (BWT) and context-based arithmetic coding, namely the context-based binary wavelet transform coding algorithm (CBWTC). In our CBWTC, in order to alleviate the degradation of predictability caused by the BWT and eliminate the correlation within the same level subbands, three highpass wavelet coefficients at the same location are combined to form an octave symbol and then encoded with a ternary arithmetic coder. The conditional context of the CBWTC is properly modeled by exploiting the properties of the BWT as well as taking the advantages of non-causal adaptive context modeling. Experimental results show that the coding performance of the CBWTC is better than that of the state-of-the-art grayscale image coders except for images containing rich texture, and always outperforms the JBIG2 algorithm and other BWT-based binary coding technique. Hong Pan 0001, Lizuo Jin, Xiao-Hui Yuan, Si-Yu Xia, Jiuxian Li, Liang-Zheng Xia |
ICASSP | 4 |
| 2008 | Wavelet neural network for 2D object classificationabstractIn this paper, a wavelet neural network (WNN)-based approach for invariant 2D object classification is proposed. The method employs the WNN characterizing the singularities of the object curvature representation and performing the classification at the same time and in an automatic way. The discriminative time-frequency attributes of the singularities on the object boundary are firstly captured by the continuous wavelet transform (CWT) and then stored by the WNN as its initial scale-translation parameters. These parameters are trained to the optimum status during the learning stage. Thus, only a few convolutions at the optimum scale-translation grids are involved during the classification, which makes our method suitable for real-time recognition tasks. Compared with the artificial neural network (ANN)-based approach preceded by a wavelet filter bank with fixed scale-translation parameters as well as the traditional methods like Fourier descriptors and moment invariants, our scheme demonstrates the best discrimination performance under various noisy and affine conditions. Hong Pan 0001, Lizuo Jin, Xiao-Hui Yuan, Si-Yu Xia, Jiuxian Li, Liang-Zheng Xia |
ICASSP | 4 |
| 2007 | Tree-Structured Support Vector Machines for Multi-class Classification
Si-Yu Xia, Jiuxian Li, Liang-Zheng Xia, Chunhua Ju |
ISNN (3) | 1 |