Sida Huang

dblp:271/8070 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Rectified Noise: A Generative Model Using Positive-incentive Noise
abstract
Rectified Flow (RF) has been widely used as an effective generative model. Although RF is primarily based on probability flow Ordinary Differential Equations (ODE), recent studies have shown that injecting noise through reverse-time Stochastic Differential Equations (SDE) for sampling can achieve superior generative performance. Inspired by Positive-incentive Noise (Pi-noise), we propose an innovative generative algorithm to train Pi-noise generators, namely Rectified Noise (RN), which improves the generative performance by injecting Pi-noise into the velocity field of pre-trained RF models. After introducing the Rectified Noise pipeline, pre-trained RF models can be efficiently transformed into Pi-noise generators. We validate Rectified Noise by conducting extensive experiments across various model architectures on different datasets. Notably, we find that: (1) RF models using Rectified Noise reduce FID from10.16 to 9.05 on ImageNet-1k. (2) The models of Pi-noise generators achieve improved performance with only 0.39% additional training parameters.
Yanchen Xu, Sida Huang, Yubin Guo, Hongyuan Zhang 0001
AAAI3
2026 Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
abstract
With the development of diffusion models, enhancing spatial controllability in text-to-image generation has become a vital challenge. As a representative task for addressing this challenge, layout-to-image generation aims to generate images that are spatially consistent with the given layout condition. Existing layout-to-image methods typically introduce the layout condition by integrating adapter modules into the base generative model. However, the generated images often exhibit low visual quality and stylistic inconsistency with the base model, indicating a loss of pretrained knowledge. To alleviate this issue, we construct the Layout Synthesis (LaySyn) dataset, which leverages images synthesized by the base model itself to mitigate the distribution shift from the pretraining data. Moreover, we propose the Layout Control (Laytrol) Network, in which parameters are inherited from MM-DiT to preserve the pretrained knowledge of the base model. To effectively activate the copied parameters and avoid disturbance from unstable control conditions, we adopt a dedicated initialization scheme for Laytrol. In this scheme, the layout encoder is initialized as a pure text encoder to ensure that its output tokens remain within the data domain of MM-DiT. Meanwhile, the outputs of the layout control network are initialized to zero. In addition, we apply Object-level Rotary Position Embedding to the layout tokens to provide coarse positional information. Qualitative and quantitative experiments demonstrate the effectiveness of our method.
Sida Huang, Ping Luo 0002, Hongyuan Zhang 0001
AAAI1
2026 CoLM: Collaborative Large Models via a Client-Server Paradigm
abstract
Large models have achieved remarkable performance across a range of reasoning and understanding tasks. Prior work often utilizes model ensembles or multi-agent systems to collaboratively generate responses, effectively operating in a server-to-server paradigm. However, such approaches do not align well with practical deployment settings, where a limited number of server-side models are shared by many clients under modern internet architectures. In this paper, we introduce CoLM (Collaboration in Large-Models), a novel framework for collaborative reasoning that redefines cooperation among large models from a client-server perspective. Unlike traditional ensemble methods that rely on simultaneous inference from multiple models to produce a single output, CoLM allows the outputs of multiple models to be aggregated or shared, enabling each client model to independently refine and update its own generation based on these high-quality outputs. This design enables collaborative benefits by fully leveraging both client-side and shared server-side models. We further extend CoLM to vision-language models (VLMs), demonstrating its applicability beyond language tasks. Experimental results across multiple benchmarks show that CoLM consistently improves model performance on previously failed queries, highlighting the effectiveness of collaborative guidance in enhancing single-model capabilities.
Sida Huang, Hongyuan Zhang 0001
AAAI2
2026 Explore How to Inject Beneficial Noise in MLLMs
abstract
Multimodal Large Language Models (MLLMs) have played an increasingly important role in multimodal intelligence. However, the existing fine-tuning methods often ignore cross-modal heterogeneity, limiting their full potential. In this work, we propose a novel fine-tuning strategy by injecting beneficial random noise, which outperforms previous methods and even surpasses full fine-tuning, with minimal additional parameters. The proposed Multimodal Noise Generator (MuNG) enables efficient modality fine-tuning by injecting customized noise into the frozen MLLMs. Specifically, we reformulate the reasoning process of MLLMs from a variational inference perspective, upon which we design a multimodal noise generator that dynamically analyzes cross-modal relationships in image-text pairs to generate task adaptive beneficial noise. Injecting this type of noise into the MLLMs effectively suppresses irrelevant semantic components, leading to significantly improved cross-modal representation alignment and enhanced performance on downstream tasks. Experiments on two mainstream MLLMs, QwenVL and LLaVA, demonstrate that our method surpasses full parameter fine-tuning and other existing fine-tuning approaches, while requiring adjustments to only about 1~2% additional parameters.
Ruishu Zhu, Sida Huang, Ziheng Jiao, Hongyuan Zhang 0001
AAAI2
2026 Trustworthy Data Equity: A Retrospective Risk-Hedging Protocol for High-Entropy IoT Data Assets
abstract
The mobilization of Internet of Things (IoT) data is currently stifled by a fundamental market friction: the inherent value uncertainty of data assets. Economically, data functions as an experience good; its true utility is highly context-dependent and verifiable only after consumption. In dynamic environments, this opacity creates High-Entropy characteristics, where the variance of realized utility, stemming from factors ranging from sensor noise to model incompatibility, renders traditional exante pricing inefficient. Risk-averse buyers, unable to assess quality prior to purchase, rationally exit the market, leading to a “Market for Lemons.” To resolve this deadlock, we propose Trustworthy Data Equity, a decentralized protocol that shifts the transaction paradigm from static pricing to retrospective risk-hedging. We model the trading process as a Stackelberg game in a bilateral setting (single buyer–single seller) and derive a closed-form equilibrium for an optimal equity share (α∗), which achieves the Pareto-efficient risk allocation between participants based on their relative risk aversion. While α∗is invariant to the estimated variance, as in classical insurance theory, the base fee (p∗) adapts directly to market risk conditions, governing both value capture and market feasibility. To implement this trustlessly, we introduce a hybrid architecture leveraging Blockchain for immutable settlement and Trusted Execution Environments for privacy-preserving, ex-post utility verification. Our empirical evaluation highlights three core contributions. First, the protocol exhibits strong Economic Robustness by significantly extending market viability under extreme uncertainty, effectively preventing market collapse where traditional mechanisms fail. Second, we identify and mitigate the “Scarcity Multiplier,” a structural inequity that penalizes data-scarce small and medium-sized enterprises, thereby promoting Market Fairness and Inclusivity. Finally, our prototype confirms System Feasibility, achieving constant-time (O(1)) on-chain settlement complexity and sub-second latency, proving the architecture’s scalability for high-frequency IoT data markets.
Yuji Dong, Yueqi Su, Sida Huang, Jie Zhang 0030
IEEE Internet Things J.3
2025 Enhance Vision-Language Alignment with Noise
abstract
With the advancement of pre-trained vision-language (VL) models, enhancing the alignment between visual and linguistic modalities in downstream tasks has emerged as a critical challenge. Different from existing fine-tuning methods that add extra modules to these two modalities, we investigate whether the frozen model can be fine-tuned by customized noise. Our approach is motivated by the scientific study of beneficial noise, namely Positive-incentive Noise (Pi-noise) , which quantitatively analyzes the impact of noise. It therefore implies a new scheme to learn beneficial noise distribution that can be employed to fine-tune VL models. Focusing on few-shot classification tasks based on CLIP, we reformulate the inference process of CLIP and apply variational inference, demonstrating how to generate Pi-noise towards visual and linguistic modalities. Then, we propose Positive-incentive Noise Injector (PiNI), which can fine-tune CLIP via injecting noise into both visual and text encoders. Since the proposed method can learn the distribution of beneficial noise, we can obtain more diverse embeddings of vision and language to better align these two modalities for specific downstream tasks within limited computational resources. We evaluate different noise incorporation approaches and network architectures of PiNI. The evaluation across 11 datasets demonstrates its effectiveness.
Sida Huang, Hongyuan Zhang 0001, Xuelong Li 0001
AAAI1
2025 Blockchain-Enabled Personalized Travel Recommendations with Semantic Search and Transparent Data
abstract
The ever-increasing demand for different Internet of Things (IoT) platforms while satisfying the personalised quality of services has attracted both industry and academic interests. Especially for the personalised travel system of the tourism area, the existing tourism recommendation systems often struggle to provide highly relevant suggestions due to limitations in understanding complex and varied user preferences. Therefore, developing a personalized tourism recommendation platform that satisfies user privacy-protecting requirements presents a necessity. In this paper, we propose a blockchain-supported personalized tourism recommendation platform that integrates semantic search models with blockchain technology. Our approach combines semantic similarity and contextual similarity using advanced natural language processing (NLP) techniques, such as word embeddings, RoBERTa models, and attention mechanisms, to align entities effectively across multiple datasets. This integration ensures a deeper understanding of user inputs, overcoming the limitations of traditional keyword-based matching. Preliminary experiments suggest that our system significantly improves recommendation performance while maintaining transparency and accountability, offering a novel solution to the challenges facing current tourism recommendation systems.
Yihan Huang, Sida Huang, Jialuoyi Tan, Shuangyao Huang, Yuji Dong
ICCCN3
2025 A Blockchain-Enhanced Deep Learning Platform for Secure Semantic Alignment and Sharing of Chemical-Biological Data
abstract
The integration and sharing of chemical-biological data have become increasingly crucial in advancing research in drug discovery, materials science, and catalysis. However, several key challenges persist, including ensuring accurate semantic alignment, protecting data privacy, and providing reliable decision support. Addressing these challenges is essential to improving the efficiency and security of data-sharing and analysis processes. This paper proposes a comprehensive solution that leverages deep learning for semantic alignment and blockchain technology for secure data sharing. Our platform utilizes natural language processing (NLP) and graph neural networks (GNN) to align heterogeneous chemical-biological datasets, ensuring consistency and completeness across different sources. Additionally, blockchain technology is employed to establish a decentralized and tamper-resistant data-sharing framework, enhancing security and trust among stakeholders. Through extensive experimentation using the Open Catalyst dataset, our results demonstrate the effectiveness of the proposed approach in achieving high-precision data alignment, secure data transactions, and reliable decision support. This work presents an innovative and integrated platform that addresses long-standing challenges in chemical-biological data integration and sharing, paving the way for more efficient and secure collaborative research.
Sida Huang, Jialuoyi Tan, Zhiran Wang, Wenzhang Zhang, Yuji Dong
ICCCN1
2025 Variational Positive-Incentive Noise: How Noise Benefits Models
abstract
A large number of works aim to alleviate the impact of noise due to an underlying conventional assumption of the negative role of noise. However, some existing works show that the assumption does not always hold. In this paper, we investigate how to benefit the classical models by random noise under the framework of Positive-incentive Noise (Pi-Noise) (Li, 2024). Since the ideal objective of Pi-Noise is intractable, we propose to optimize its variational bound instead, namely variational Pi-Noise (VPN). With the variational inference, a VPN generator implemented by neural networks is designed for enhancing base models and simplifying the inference of base models, without changing the architecture of base models. Benefiting from the independent design of base models and VPN generators, the VPN generator can work with most existing models. From the extensive experiments on different base models (including linear models, ResNet, ViT, etc.) it is shown that the proposed VPN generator can improve the base models. It is appealing that the trained VPN generator prefers to blur the irrelevant ingredients in complicated images, which meets our expectations.
Hongyuan Zhang 0001, Sida Huang, Yubin Guo, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Web3.0 Literary Landscape: Deep Learning and Blockchain for Nobel Prize Predictions
abstract
This research introduces a cutting-edge Web3 literary analysis platform, harnessing the power of blockchain and deep learning technologies. By employing the immutable and transparent nature of blockchain, the platform ensures robust copyright protection while offering readers enhanced interactive features. It applies deep learning techniques for comprehensive analyses of sentiment, topic, and stylistic elements, which are instrumental in predicting potential Nobel Prize laureates. This methodology not only enhances the accuracy of predictions but also sheds light on the evaluation criteria and historical trends associated with the Nobel Prize. Moreover, the platform adopts a directed graph model alongside the struc2vec algorithm to create text vectors for comparative studies, uncovering similarities between works that have won awards and those that have been nominated. Utilizing the LESS model for detailed content examination, the platform delves into sequence relationships within semantic networks, thus improving interpretability and visualization. The integration of blockchain technology guarantees access to unbiased datasets, enabling more precise literary analyses and predictions. This innovative approach has been validated using works that have either won or been nominated for the Nobel Prize, proving its efficacy in identifying the textual characteristics favored by the Nobel Prize committee.
Sida Huang, Jialuoyi Tan, Yuji Dong, Jie Zhang 0030
ICCCN1
2021 Distribution Adaptive INT8 Quantization for Training CNNs
abstract
Researches have demonstrated that low bit-width (e.g., INT8) quantization can be employed to accelerate the inference process. It makes the gradient quantization very promising since the backward propagation requires approximately twice more computation than forward one. Due to the variability and uncertainty of gradient distribution, a lot of methods have been proposed to attain training stability. However, most of them ignore the channel-wise gradient distributions and the impact of gradients with different magnitudes, resulting in the degradation of final accuracy. In this paper, we propose a novel INT8 quantization training framework for convolutional neural network to address the above issues. Specifically, we adopt Gradient Vectorized Quantization to quantize the gradient, based on the observation that layer-wise gradients contain multiple distributions along the channel dimension. Then, Magnitude-aware Clipping Strategy is introduced by taking the magnitudes of gradients into consideration when minimizing the quantization error, and we present a theoretical derivation to solve the quantization parameters of different distributions. Experimental results on broad range of computer vision tasks, such as image classification, object detection and video classification, demonstrate that the proposed Distribution Adaptive INT8 Quantization training method has achieved almost lossless training accuracy for different backbones, including ResNet, MobileNetV2, InceptionV3, VGG and AlexNet, which is superior to the state-of-the-art techniques. Moreover, we further implement the INT8 kernel that can accelerate the training iteration more than 200% under the latest Turing architecture, i.e., our method excels on both training accuracy and speed.
Sida Huang, Yingya Zhang
AAAI2
2020 Rear-End Collision Avoidance-Based on Multi-Channel Detection
abstract
With multi-sensor-based collision avoidance systems (CASs) being adopted in today’s automobiles, a new method that enables collaborative decision-making with preceding vehicle detection under various external environments is needed. In this paper, spatial–temporal correlations of multi-channel signals that are collected by multiple sensors on the host vehicle are considered, and a multi-channel detection technique with a stochastic model is introduced for automobile collision avoidance. We propose an accurate and robust multi-channel, generalized likelihood ratio test (GLRT)-based detection and collaborative decision-making scheme, with a vehicle kinematic analysis for avoiding rear-end collisions. The results of simulations and physical experiments demonstrated that our detector expands the detection range with a high detection rate and that our proposed scheme obtains good performance under varying operating and environmental conditions.
Yu Xiang 0002, Sida Huang, Jin Li 0026
IEEE Trans. Intell. Transp. Syst.2