EDBT 2026 Demo / reviewers in the wild / expert
Shiyuan He
dblp:146/2829
· DBLP profile ↗
20ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image GenerationabstractMultimodal Large Language Models (MLLMs) with unified architectures excel across a wide range of vision-language tasks, yet aligning them with personalized image generation remains a significant challenge. Existing methods for MLLMs are frequently subject-specific, demanding a data-intensive fine-tuning process for every new subject, which limits their scalability. In this paper, we introduce MM-R1, a framework that integrates a cross-modal Chain-of-Thought (X-CoT) reasoning strategy to unlock the inherent potential of unified MLLMs for personalized image generation. Specifically, we structure personalization as an integrated visual reasoning and generation process: (1) grounding subject concepts by interpreting and understanding user-provided images and contextual cues, and (2) generating personalized images conditioned on both the extracted subject representations and user prompts. To further enhance the reasoning capability, we adopt Grouped Reward Proximal Policy Optimization(GRPO) to explicitly align the generation. Experiments demonstrate that MM-R1 unleashes the personalization capability of unified MLLMs to generate images with high subject fidelity and strong text alignment in a zero-shot manner. Yujia Wu, Kuncheng Li, Jiwei Wei, Shiyuan He, Jinyu Guo, Ning Xie 0003 |
AAAI | 5 |
| 2026 | DWTSG: Parameter-Efficient Fine-Tuning of Large Pre-trained Models via Discrete Wavelet Transform and Subband GuidanceabstractFully fine-tuning large pre-trained models for each downstream task is impractical due to prohibitive memory, computation, and storage costs. Although parameter-efficient fine-tuning (PEFT) methods address this issue, leading methods like LoRA still exhibit linear scaling of trainable parameters with hidden size. Recent studies have explored PEFT in the frequency domain to reduce computational costs by employing fast Fourier transform and discrete cosine transform with sparse frequency selection. These methods rely on global frequency representations that lack spatial locality and disperse energy across the domain. As a result, sparse coefficient selection struggles to preserve fine-grained structural information and often introduces artifacts such as ringing near boundaries. To address these limitations, we propose DWTSG, a novel PEFT framework based on discrete wavelet transform (DWT) and subband guidance. DWTSG decomposes pre-trained weights into four wavelet subbands that jointly encode global context and local details. It fine-tunes only the most informative coefficients in each subband through an energy-based selection strategy that prioritizes coefficients based on their individual importance and interactions. Finally, inverse DWT reconstructs the updated weights, enabling efficient and precise adaptation. Extensive experiments on natural language understanding, commonsense reasoning, and image classification demonstrate that DWTSG outperforms existing PEFT methods, achieving superior performance and higher parameter efficiency. Chengwei Sun, Jiwei Wei, Shiyuan He, Zeyu Ma 0002, Ran Ran 0001, Jie Zou 0001, Yang Yang 0002 |
AAAI | 3 |
| 2026 | GASE: Generalized adaptive static enhancement for temporal sentence grounding
Ran Ran 0001, Kaiwen Shen, Jiwei Wei, Ruikun Chai, Shiyuan He, Zeyu Ma 0002, Malu Zhang, Yang Yang 0002 |
Knowl. Based Syst. | 5 |
| 2026 | Mismatched Pairs Dynamic Correction for Cross-Modal Alignment in Video Moment Retrieval
Hongxiao Hu, Ran Ran 0001, Jiwei Wei, Cheng-Wei Sun, Shiyuan He, Kuien Liu, Yang Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | LoLDU: Low-Rank Adaptation via Lower-Diag-Upper Decomposition for Parameter-Efficient Fine-TuningabstractThe rapid growth of model scale has necessitated substantial computational resources for fine-tuning. Existing approach such as low-rank adaptation (LoRA) has sought to address the problem of handling the large updated parameters in full fine-tuning (FT). However, LoRA utilize random initialization and optimization of low-rank matrices to approximate updated weights, which can result in suboptimal convergence and an accuracy gap compared to full fine-tuning (FT). To address these issues, we propose low-rank LDU (LoLDU), a parameter-efficient fine-tuning (PEFT) approach that significantly reduces trainable parameters by 2600 times compared to regular PEFT methods while maintaining comparable performance. LoLDU leverages lower-diag-upper (LDU) decomposition to initialize low-rank matrices for faster convergence and nonsingularity. We focus on optimizing the diagonal matrix for scaling transformations. To the best of our knowledge, LoLDU has the fewest parameters among all PEFT approaches. We conducted extensive experiments across 4 instruction-following datasets, six natural language understanding (NLU) datasets, eight image classification datasets, and image generation datasets with multiple model types [LLaMA2, RoBERTa, ViT, and stable diffusion (SD)], providing a comprehensive and detailed analysis. Our open-source code can be accessed at https://anonymous.4open.science/r/LoLDU-B5A6. Yiming Shi, Yujia Wu, Jiwei Wei, Ran Ran 0001, Cheng-Wei Sun, Shiyuan He, Yang Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | KDA: Knowledge Diffusion Alignment with Enhanced Context for Video Temporal Grounding
Ran Ran 0001, Jiwei Wei, Shiyuan He, Zeyu Ma 0002, Chaoning Zhang, Ning Xie 0003, Yang Yang 0002 |
ICCV | 3 |
| 2025 | An Improved Clique-Picking Algorithm for Counting Markov Equivalent DAGs via Super Cliques TransferabstractEfficiently counting Markov equivalent directed acyclic graphs (DAGs) is crucial in graphical causal analysis. Wienöbst et al. (2023) introduced a polynomial-time algorithm, known as the Clique-Picking algorithm, to count the number of Markov equivalent DAGs for a given completed partially directed acyclic graph (CPDAG). This algorithm iteratively selects a root clique, determines fixed orientations with outgoing edges from the clique, and generates the unresolved undirected connected components (UCCGs). In this work, we propose a more efficient approach to UCCG generation by utilizing previously computed results for different root cliques. Our method introduces the concept of super cliques within rooted clique trees, enabling their efficient transfer between trees with different root cliques. The proposed algorithm effectively reduces the computational complexity of the Clique-Picking method, particularly when the number of cliques is substantially smaller than the number of vertices and edges. Lifu Liu, Shiyuan He |
ICML | 2 |
| 2025 | SyncGaussian: Stable 3D Gaussian-Based Talking Head Generation with Enhanced Lip Sync via Discriminative Speech FeaturesabstractGenerating high-fidelity talking heads that maintain stable head poses and achieve robust lip sync remains a significant challenge. Although methods based on 3D Gaussian Splatting (3DGS) offer a promising solution via point-based deformation, they suffer from inconsistent head dynamics and mismatched mouth movements due to unstable Gaussian initialization and incomplete speech features. To overcome these limitations, we introduce SyncGaussian, a 3DGS-based framework that ensures stable head poses, enhanced lip sync, and realistic appearances with real-time rendering. SyncGaussian employs a stable head Gaussian initialization strategy to mitigate head jitter by optimizing commonly used rough head pose parameters. To enhance lip sync, we propose a sync-enhanced encoder that leverages audio-to-text and audio-to-visual speech features. Guided by a tailored cosine similarity loss function, the encoder integrates discriminative speech features through a multi-level sync adaptation mechanism, enabling the learning of an adaptive speech feature space. Extensive experiments demonstrate that SyncGaussian outperforms state-of-the-art methods in image quality, dynamic motion, and lip sync, with the potential for real-time applications. Jiwei Wei, Shiyuan He, Zeyu Ma 0002, Chaoning Zhang, Ning Xie 0003, Yang Yang 0002 |
IJCAI | 3 |
| 2025 | DSCS: Fast CPDAG-Based Verification of Collapsible Submodels in High-Dimensional Bayesian NetworksabstractBayesian networks (BNs), represented by directed acyclic graphs (DAGs), provide a principled framework for modeling complex dependencies among random variables. As data dimensionality increases into the tens of thousands, fitting and marginalizing a full BN becomes computationally prohibitive—particularly when inference is only needed for a small subset of variables. Estimation-collapsibility addresses this challenge by ensuring that directly fitting a submodel, obtained by ignoring non-essential variables, still yields exact inference on target variables. However, current DAG-based criterion for checking estimation-collapsibility is computationally intensive, involving exhaustive vertex searches and iterative removals. Additionally, practical applications typically identify the underlying DAG only up to its Markov equivalence class, represented by a completed partially directed acyclic graph (CPDAG). To bridge this gap, we introduce sequential $c$-simplicial sets—a novel graphical characterization of estimation-collapsibility applicable directly to CPDAGs. We further propose DSCS, a computationally efficient algorithm for verifying estimation-collapsibility within CPDAG framework that scales effectively to high-dimensional BNs. Extensive numerical experiments demonstrate the practicality, scalability, and efficiency of our proposed approach. Shiyuan He |
NeurIPS | 2 |
| 2025 | Text-guided dynamic mouth motion capturing for person-generic talking face generation
Jiwei Wei, Ruiqi Yuan, Ruikun Chai, Shiyuan He, Zeyu Ma 0002, Yang Yang 0002 |
Knowl. Based Syst. | 5 |
| 2025 | Fine-Grained Alignment and Interaction for Video Grounding With Cross-Modal Semantic Hierarchical GraphabstractVideo grounding tasks have recently gained significant attention. However, existing methods failed to fully comprehend the semantics within queries and videos, often overlooking key content. Moreover, the lack of fine-grained cross-modal alignment and interaction to guide the semantic matching of complex texts and videos lead to inconsistent representational modeling. To address this issue, we propose a Semantic Hierarchical Grounding model, referred to as SHG, and design a cross-modal semantic hierarchical graph to achieve fine-grained semantic understanding. SHG decomposes both the query and each video moment into three levels: global, action, and element. This topology, ranging from global to local, establishes multigranularity intrinsic connections between the two modalities, fostering a comprehensive understanding of dynamic semantics and fine-grained cross-modal matching. Accordingly, to fully leverage the rich information within the cross-modal semantic hierarchical graph, we employ contrastive learning by seeking samples with the same action and element semantics, then achieve node-moment cross-modal hierarchical matching for global alignment. This approach can unearth fine-grained clues and align semantics across multiple granularities. Moreover, we combine the designed hierarchical graph interaction for coarse-to-fine fusion of text and video, thereby enabling highly accurate video grounding. Extensive experiments conducted on three challenging public datasets (ActivityNet-Captions, TACoS, and Charades-STA) demonstrate that the proposed approach outperforms state-of-the-art techniques, validating its effectiveness. Ran Ran 0001, Jiwei Wei, Shiyuan He, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Hyperspectral and Multispectral Image Fusion With Functional Data Analysis TechniquesabstractThe fusion of hyperspectral images (HSI) and multispectral images (MSI) aims to synthesize high-resolution hyperspectral images (HR-HSI) from observable low-resolution hyperspectral (LR-HSI) and high-resolution multispectral (HR-MSI) data. Existing methodologies often overlook the importance of the degradation model within imaging systems and the inherent spectral and spatial smoothness properties of hyperspectral data, resulting in suboptimal outcomes in real-world applications. We introduce a novel approach that utilizes functional data analysis techniques to precisely estimate degradation parameters and effectively fuse HSI-MSI pairs. Based on the physical generation mechanisms of both HSI and MSI, alongside the intrinsic smoothness of degradation parameters, we propose an iterative function estimation technique to recover the spectral response and point spread functions. Additionally, we introduce a spatial content-aware spectral subspace regularizer (SSR) that promotes low-dimensional functional subspaces for spectral data within similar spatial contexts. Implemented within an ADMM optimization framework, our fusion model integrates the precisely estimated degradation parameters and SSR to accurately construct the target HR-HSI data. Extensive experimental evaluations on multiple public datasets demonstrate that our approach achieves superior performance in both quantitative metrics and qualitative assessments. The code for this work is available https://github.com/lingxingongye/FDAssr.git. Lingxin GongYe, Shiyuan He |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Density-Aware Cloud Removal of Remote Sensing Imagery Using a Global-Local Fusion TransformerabstractCloud cover poses a significant challenge in remote sensing image processing, affecting the extraction and analysis of terrestrial features. Despite advancements in multitemporal cloud removal methods, single-image declouding remains crucial for emergency response and disaster management, where rapid acquisition of cloud-free imagery is essential. Traditional approaches often rely on synthetic aperture radar (SAR) or cloud masks as guidance for cloud removal, introducing additional complexities and dependencies on extensive data. To address these limitations, we propose a density-aware cloud removal using a global-local fusion Transformer (DCR-GLFT), which leverages density information as guidance and does not rely on extensive data. Specifically, our method employs density labels to guide the cloud removal process through two primary stages: cloud density estimation and density-guided cloud removal. A cloud density classifier is proposed in the first stage, trained with roughly estimated ground truth, to generate density labels for guiding subsequent removal processes. The second stage integrates cloud density information with cloud-ground image features using a Transformer-based network, enabling precise and nuanced cloud removal while preserving underlying surface details through the integration of both global and local features. The proposed method achieved the state-of-the-art results (peak signal-to-noise ratio (PSNR) of 28.93 and structural similarity index measure (SSIM) of 0.84) on the renowned cloud-removal dataset SEN12MS-CR, even without utilizing SAR data for guidance. This accomplishment highlights its significant advancement in the single-image cloud removal task. Our code will be made available athttps://github.com/ruiquan1214/DCR-GLFT.git. Quan Rui, Shiyuan He, Tianyu Li 0003, Guoqing Wang 0001, Ningjuan Ruan, Lin Mei 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Boosting Adversarial Training with Hardness-Guided Attack StrategyabstractThe susceptibility of deep neural networks (DNNs) to adversarial examples has raised significant concerns regarding the security and reliability of artificial intelligence systems. These examples contain maliciously crafted perturbations not perceptible to the human eye but can cause the model to make wrong predictions. Adversarial training (AT) is the de facto standard method for enhancing adversarial robustness. However, the improved robustness is often at the cost of a significant drop in standard accuracy for clean samples. Numerous works have attempted to alleviate this trade-off by identifying its causes. A key factor lies in the variability of clean samples, which leads to different adversarial examples being generated using the same attack strategy. The other factor is the disruption of the underlying data structure caused by adversarial perturbations. To overcome these challenges, we propose a novel adversarial training framework named Hardness-Guided Sample-Dependent Adversarial Training (HGSD-AT), which dynamically adjusts the attack strategy based on the hardness of the current adversarial sample to further improve the robustness of the model. By utilizing the two types of constraints which construct from a temporal perspective and spatial distribution perspective, our method directly learns the impact of attack methods on the model, rather than the indirect effects associated with sample distribution. This approach aims to improve the generation of adversarial examples while simultaneously enhancing the robustness and accuracy of DNNs. Our approach exhibits superior performance in terms of both robustness and natural accuracy compared to state-of-the-art defense methods, as validated through comprehensive experiments conducted on three benchmark datasets. Shiyuan He, Jiwei Wei, Chaoning Zhang, Xing Xu 0001, Jingkuan Song, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 1 |
| 2024 | Towards Robust Person Re-Identification by Adversarial Training With Dynamic Attack StrategyabstractRecently, person re-identification has gained significant attention from both academic and industry fields due to its potential applications in surveillance and security. However, the security of re-identification systems has not been widely investigated, and they are vulnerable to adversarial attacks, which can significantly degrade their performance. Although numerous sophisticated adversarial training methods have been proposed for image classification, metric analysis systems such as person re-identification have not been fully explored. In this paper, we develop a novel adversarial training framework with a dynamic attack strategy for person re-identification, to further enhance the robustness of the model. Specifically, we gradually increase the perturbation budget during the generation until the generated adversarial examples reach a certain level of attack strength. As the iterations progress, the model becomes more robust, and our framework can generate stronger adversarial examples to continuously explore the robustness bounds of the model. Moreover, to alleviate the conflict between the adversarial robustness and natural generalization of the model, we design a novel performance alignment loss to further constrain the adversarial example generation process, which can make the generated adversarial examples as close as possible to the clean samples in terms of performance. Experiments on two widely used person re-ID benchmark datasets demonstrate the effectiveness and superiority of our proposed method. Jiwei Wei, Shiyuan He, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 3 |
| 2023 | A Unified Analysis of Multi-task Functional Linear Regression Models with Manifold Constraint and Composite Quadratic PenaltyabstractThis work studies the multi-task functional linear regression models where both the covariates and the unknown regression coefficients (called slope functions) are curves. For slope function estimation, we employ penalized splines to balance bias, variance, and computational complexity. The power of multi-task learning is brought in by imposing additional structures over the slope functions. We propose a general model with double regularization over the spline coefficient matrix: i) a matrix manifold constraint, and ii) a composite penalty as a summation of quadratic terms. Many multi-task learning approaches can be treated as special cases of this proposed model, such as a reduced-rank model and a graph Laplacian regularized model. We show the composite penalty induces a specific norm, which helps quantify the manifold curvature and determine the corresponding proper subset in the manifold tangent space. The complexity of tangent space subset is then bridged to the complexity of geodesic neighbor via generic chaining. A unified upper bound of the convergence rate is obtained and specifically applied to the reduced-rank model and the graph Laplacian regularized model. The phase transition behaviors for the estimators are examined as we vary the configurations of model parameters. Shiyuan He, Hanxuan Ye, Kejun He |
J. Mach. Learn. Res. | 1 |
| 2023 | Category Alignment Adversarial Learning for Cross-Modal RetrievalabstractCross-modal retrieval aims to retrieve one semantically similar media from multiple media types based on queries entered by another type of media. An intuitive idea is to map different media data into a common space and then directly measure content similarity between different types of data. In this paper, we present a novel method, called Category Alignment Adversarial Learning (CAAL) for cross-modal retrieval. It aims to find a common representation space supervised by category information, in which the samples from different modalities can be compared directly. Specifically, CAAL firstly employs two parallel encoders to generate common representations for image and text features respectively. Furthermore, we employ two parallel GANs with category information to generate fake image and text features which next will be utilized with already generated embedding to reconstruct the common representation. At last, two joint discriminators are utilized to reduce the gap between the mapping of the first stage and the embedding of the second stage. Comprehensive experimental results on four widely-used benchmark datasets demonstrate the superior performance of our proposed method compared with the state-of-the-art approaches. Shiyuan He, Weiyang Wang, Zheng Wang 0044, Xing Xu 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Self-supervised adversarial learning for cross-modal retrievalabstractCross-modal retrieval aims at enabling flexible retrieval across different modalities. The core of cross-modal retrieval is to learn projections for different modalities and make instances in the learned common subspace comparable to each other. Self-supervised learning automatically creates a supervision signal by transformation of input data and learns semantic features by training to predict the artificial labels. In this paper, we proposed a novel method named Self-Supervised Adversarial Learning (SSAL) for Cross-Modal Retrieval, which deploys self-supervised learning and adversarial learning to seek an effective common subspace. A feature projector tries to generate modality-invariant representations in the common subspace that can confuse an adversarial discriminator consists of two classifiers. One of the classifiers aims to predict rotation angle from image representations, while the other classifier tries to discriminate between different modalities from the learned embeddings. By confusing the self-supervised adversarial model, feature projector filters out the abundant high-level visual semantics and learns image embeddings that are better aligned with text modality in the common subspace. Through the joint exploitation of the above, an effective common subspace is learned, in which representations of different modlities are aligned better and common information of different modalities is well preserved. Comprehensive experimental results on three widely-used benchmark datasets show that the proposed method is superior in cross-modal retrieval and significantly outperforms the existing cross-modal retrieval methods. Yangchao Wang, Shiyuan He, Xing Xu 0001, Yang Yang 0002, Jingjing Li 0001, Heng Tao Shen |
MMAsia | 2 |
| 2020 | Bidirectional Discrete Matrix Factorization Hashing for Image SearchabstractUnsupervised image hashing has recently gained significant momentum due to the scarcity of reliable supervision knowledge, such as class labels and pairwise relationship. Previous unsupervised methods heavily rely on constructing sufficiently large affinity matrix for exploring the geometric structure of data. Nevertheless, due to lack of adequately preserving the intrinsic information of original visual data, satisfactory performance can hardly be achieved. In this article, we propose a novel approach, called bidirectional discrete matrix factorization hashing (BDMFH), which alternates two mutually promoted processes of 1) learning binary codes from data and 2) recovering data from the binary codes. In particular, we design the inverse factorization model, which enforces the learned binary codes inheriting intrinsic structure from the original visual data. Moreover, we develop an efficient discrete optimization algorithm for the proposed BDMFH. Comprehensive experimental results on three large-scale benchmark datasets show that the proposed BDMFH not only significantly outperforms the state-of-the-arts but also provides the satisfactory computational efficiency. Shiyuan He, Bokun Wang, Zheng Wang 0044, Yang Yang 0002, Fumin Shen, Zi Huang, Heng Tao Shen |
IEEE Trans. Cybern. | 1 |
| 2018 | Learning binary codes with local and inner data structure
Shiyuan He, Guo Ye, Mengqiu Hu, Yang Yang 0002, Fumin Shen, Heng Tao Shen, Xuelong Li 0001 |
Neurocomputing | 1 |