Siyao Li

dblp:163/8053 · DBLP profile ↗
← Back
40ranked-venue papers
14as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Computer networks · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Theory of computation · 3 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Semantic Communication Performance Optimization with Channel and Content Preference Feedbacks
Defeng Zhou, Dongyu Wei, Siyao Li, Mingzhe Chen
ICC4
2026 Perceptive scale and selective attention few-shot learning network for hyperspectral and light detection and ranging fusion classification
Xiang-Hai Wang 0001, Tingting Geng, Xiaohan Xie, Xiao-Yang Zhao 0003, Siyao Li
Eng. Appl. Artif. Intell.6
2026 HiF2-FSLF: Hierarchical frequency fusion few-Shot learning framework for hyperspectral and lidar classification
Xiang-Hai Wang 0001, Xiaohan Xie, Xiao-Yang Zhao 0003, Siyao Li
Expert Syst. Appl.5
2026 Cross-Modal Visual Perception Consistency: A Language-Enhanced Approach for Heterogeneous Change Detection
abstract
Heterogeneous remote sensing image change detection (HRSICD) seeks to identify surface changes by comparing images captured at different times. However, CD faces significant challenges due to heterogeneity arising from varying sensor types and imaging conditions. Recently, powerful vision-language models like CLIP have emerged, with strong semantic decoding abilities. Opening new possibilities for using linguistic information as an auxiliary in visual tasks, potentially driving breakthroughs in HCD. Capitalizing on this prospect, we investigate graph learning with vision-language features and introduce LEVPC, the first language-enhanced visual perception consistency framework for HCD. First, we create a mutual information-guided graph aggregation module. Specifically, it builds modality-invariant structured relationships among visual nodes by using language features as connecting bridges, providing a consistent foundation for comparing changes. To reduce modeling bias from heterogeneity, language is used as an anchor to aggregate features, ensuring a unified expression of visual representations. In summary, language guides the generation and aggregation of multiple subgraphs from visual inputs, ultimately building robust representations of structural relationships within a shared semantic space. Moreover, a change semantic compensation module is introduced, which analyses the change intensity between bi-temporal data from a vision-language perspective. And then adds change-related semantic descriptions for salient change regions, enhancing the expressiveness of visual change features. Experiments on multiple datasets validate the superior performance of LEVPC in HCD, achieving an average increase of 2.6% in Kappa. The code will be publicly available at https://github.com/sylXIDIAN/LEVPC.
Siyao Li, Weiying Xie, Jitao Ma, Leyuan Fang, Yunsong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Toward Fine-Grained Load Balancing With Congested-Flow Isolation in Lossless Datacenters
abstract
Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) cooperating with Priority Flow Control (PFC) has been widely deployed in production datacenters to enable low latency, lossless transmission. At the same time, modern datacenters typically offer parallel transmission paths between any pair of end-hosts, underscoring the importance of load balancing. However, the well-studied load balancing mechanisms designed for lossy datacenter networks (DCNs) are ill-suited for such lossless environments. Through extensive experiments, we are among the first to comprehensively inspect the interactions between PFC and load balancing, and uncover that existing fine-grained rerouting schemes can be counterproductive to spread the congested flows among more paths, further aggravating PFC’s head-of-line (HoL) blocking. Motivated by this, we present FLB, a Fine-grained Load Balancing scheme for lossless DCNs. At its core, FLB employs threshold-free rerouting to effectively balance traffic load and improve link utilization during normal conditions and leverages timely congested flow isolation to eliminate HoL blocking on non-congested flows when congestion occurs. To handle complex multi-bottleneck scenarios, we further introduce FLB*, which incorporates an enhanced congestion-point-aware isolation mechanism using Congestion Point Identifiers (CPI) to eliminate HoL blocking among different congested flows.We have fully implemented a FLB prototype, and our evaluation results show that FLB reduces PFC PAUSE rate by up to 96% and avoids HoL blocking, translating to up to 45% improvement in goodput over CONGA+DCQCN and 40%, 36%, 29% and 18% reduction in average flow completion time (FCT) over LetFlow+Swift, MP-RDMA, Proteus+DCQCN and LetFlow+PCN, respectively.
Jinbin Hu 0001, Siyao Li, Wenxue Li 0004, Xiangzhou Liu, Bowen Liu 0002, Ping Yin, Mengyu Ma, Jin Wang 0001, Jianxin Wang 0001, Jiawei Huang 0001, Kai Chen 0005
IEEE Trans. Netw.2
2025 On the Sensing Capacity of Gaussian "Beam-Pointing" Channels with Block Memory and Feedback
Siyao Li, Shuangyang Li, Giuseppe Caire
ISIT1
2025 CGM: Intrusion Detection Based on a Multi-head Attention Optimization Model
Siyao Li, Yong Wang 0055, Zhen Wang 0042
KSEM (3)1
2025 VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
abstract
With the rapid development of MLLMs, evaluating their visual capabilities has become increasingly crucial. Current benchmarks primarily fall into two main types: basic perception benchmarks,which focus on local details but lack deep reasoning (e.g., ''what is in the image?''), and mainstream reasoning benchmarks, which concentrate on prominent image elements but may fail to assess subtle clues requiring intricate analysis. However, profound visual understanding and complex reasoning depend more on interpreting subtle, inconspicuous local details than on perceiving salient, macro-level objects. These details, though occupying minimal image area, often contain richer, more critical information for robust analysis. To bridge this gap, we introduce the VER-Bench, a novel framework to evaluate MLLMs' ability to: 1) identify fine-grained visual clues, often occupying, on average, just 0.25% of the image area; 2) integrate these clues with world knowledge for complex reasoning. Comprising 374 carefully designed questions across Geospatial, Temporal, Situational, Intent, System State, and Symbolic reasoning, each question in VER-Bench is accompanied by structured evidence: visual clues and question-related reasoning derived from them. VER-Bench reveals current models' limitations in extracting subtle visual evidence and constructing evidence-based reasoning chains, highlighting the need to enhance models' capabilities in fine-grained visual evidence extraction, integration, and reasoning for genuine visual understanding and human-like analysis. The dataset is available at https://github.com/verbta/ACMMM-25-Materials.
Chenhui Qiang, Zhaoyang Wei, Xumeng Han, Siyao Li, Xiangyuan Lan, Jianbin Jiao, Zhenjun Han
ACM Multimedia5
2025 FactCG: Enhancing Fact Checkers with Graph-Based Multi-Hop Data
abstract
Deren Lei, Yaxi Li, Siyao Li, Mengya Hu, Rui Xu, Ken Archer, Mingyu Wang, Emily Ching, Alex Deng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Deren Lei, Yaxi Li, Siyao Li, Mengya Hu, Ken Archer, Emily Ching, Alex Deng
NAACL (Long Papers)3
2025 EXP-PnP: General Extended Model for Detail-Injected Pan-Sharpening With Plug-and-Play Residual Optimization
abstract
Detail injection model-based methods are the mainstream pan-sharpening techniques for multispectral (MS) images. In recent years, the research on this type of method mainly focuses on optimizing the extraction and injection of panchromatic (PAN) image details, while paying less attention to the adaptive enhancement of MS image details. Due to the differences in spectral responses of different sources, it is difficult to effectively enhance or recover the multispectral details in the fused image. In this article, we analyze the limitations of the existing interpolation enhancement work from the perspective of model derivation, and propose an extended model of pan-sharpening detail injection based on “plug-and-play” (PnP) residual optimization. The model not only focuses on the flexibility in the choice of optimization routes and interpolation enhancement methods, but also emphasizes the universality of interpolation enhancement schemes across detail injection models. Our main contributions include the proposed residual interpolation optimization-based pan-sharpening extension model for detail injection oriented to additive and multiplicative rules, which successfully solves the problem of ineffective optimization in earlier related studies, especially for important multiresolution analysis (MRA) methods such as generalized Laplacian pyramid (GLP), and achieves universally effective optimization. In addition, through large-scale adaptive experiments, we selected 17 PnP methods including optimization model (OM) and deep learning (DL) methods for optimization tests, and comprehensively evaluated the applicability and effectiveness of the model. Comprehensive tests on 12 sets of images from five types of sensors on two public datasets show that our methods can achieve significant improvements in the main evaluation metrics, and the average ERGAS, Q2n, and HQNR metrics can reach 8.9%, 2.3%, and 6.3%, respectively. The source code of the proposed method can be downloaded fromhttps://github.com/JZ-Tao/EXP-PnP/.
Jingzhe Tao, Tingting Geng, Chunmei Han, Siyao Li, Chuanming Song 0001, Xiang-Hai Wang 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Patch- and Class-Wise Hyperspectral Knowledge Learning: A Composite Consistency-Constrained Self-Ensemble Framework for Change Detection
abstract
Obtaining fine land surface change information from multitemporal hyperspectral images (HSIs) is a key goal pursued in remote sensing image processing. Recently, HSI change detection (HSI-CD) methods based on convolutional neural networks (CNNs) have achieved surprising detection results. One of the reasons is the support of large-scale labeled samples for network learning. However, the existence of mixed pixels greatly increases the difficulty of HSI interpretation, resulting in accurate pixel-level labeling work with a heavy burden and unable to meet the needs of time-sensitive applications. For this reason, achieving stable and high-precision CD with fewer samples is a difficult issue in this field. To address the above problems, a composite consistency-constrained self-ensemble framework (C3SelF) for HSI-CD is proposed, to alleviate the problems of low detection accuracy and instability caused by small samples. The framework mainly comprises two lightweight networks with the same structure aiming at accelerating the model inference process and thus improving the processing timeliness. The composite learning mode implements patch-wise classification loss, class-wise consistency loss on labeled samples, and patch-wise consistency loss on unlabeled samples under a multilevel noise perturbation strategy, which improves the classification results and reduces the labeling cost. Moreover, to exploit the multidimensional features contained in HSIs, a lightweight selective spatial-spectral feature joint network (S3Net) is designed to overcome over-fitting, and to deeply mine the discriminative information in unlabeled samples, a new sample screening strategy is designed to ensure the stability of the network during training unlabeled samples. Extensive experiments prove that the proposed C3SelF outperforms the state-of-the-art (SOTA) methods at a sampling rate of 0.1%, reaching 93.44% Kappa and 97.26% overall accuracy (OA) on the Farmland dataset. The source code of the proposed framework will be released athttps://github.com/zxylnnu/C3SelF.
Xiao-Yang Zhao 0003, Siyao Li, Chuanming Song 0001, Xiang-Hai Wang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Dynamic Fronthaul Load Optimization for Uplink Scalable Cell-Free User-Centric Massive MIMO
abstract
This study investigates scalable uplink cell-free massive multiple-input multiple-output networks, comprising user equipments (UEs), radio units (RUs), data routers, and decentralized processing units (DUs). In our model, UEs are served by dynamically allocated user-centric clusters of RUs. The corresponding cluster processors, implementing the physical layer for each user, are hosted as software-defined virtual network functions by the DUs. In our paradigm, RUs, data routers, and DUs are not fully connected, necessitating a holistic approach to address the joint challenges of cluster processor placement (at one of the DUs) and the allocation of fronthaul data links among RUs, routers, and DUs. We simultaneously consider the fronthaul topology, the limited fronthaul communication capacity, and computation constraints at the DUs. Specifically, we formulate the joint optimization of fronthaul load balancing and cluster processor placement as a mixed-integer linear problem. Furthermore, we present numerical results that shed light on the interplay between these elements under finite resolution of the A/D quantization at the RUs.
Zhiyang Li 0001, Fabian Goettsch, Siyao Li, Ming Chen 0001, Giuseppe Caire
ICC3
2024 Compressed Sensing Inspired User Acquisition for Downlink Integrated Sensing and Communication Transmissions
abstract
This paper investigates radar-assisted user acquisition for downlink multi-user multiple-input multiple-output (MIMO) transmission using Orthogonal Frequency Division Multiplexing (OFDM) signals. Specifically, we formulate a concise mathematical model for the user acquisition problem, where each user is characterized by its delay and beamspace response. Therefore, we propose a two-stage method for user acquisition, where the Multiple Signal Classification (MUSIC) algorithm is adopted for delay estimation, and then a least absolute shrinkage and selection operator (LASSO) is applied for estimating the user response in the beamspace. Furthermore, we also provide a comprehensive performance analysis of the considered problem based on the pair-wise error probability (PEP). Particularly, we show that the rank and the geometric mean of non-zero eigenvalues of the squared beamspace difference matrix determines the user acquisition performance. More importantly, we reveal that simultaneously probing multiple beams outperforms concentrating power on a specific beam direction in each time slot under the power constraint, when only limited OFDM symbols are transmitted. Our numerical results confirm our conclusions and also demonstrate a promising acquisition performance of the proposed two-stage method.
Yi Song 0011, Fernando Pedraza, Shuangyang Li, Siyao Li, Han Yu 0010, Giuseppe Caire
ICC4
2024 On the Capacity of Gaussian "Beam-Pointing" Channels with Block Memory and Feedback
abstract
Motivated by wireless communications at high carrier frequencies in 5G and 6G systems (mmWaves, sub-THz), we consider a state-dependent channel model with in-block memory referred to as the Gaussian beam-pointing (GBP) channel. A transmitter equipped with a large antenna array wishes to communicate with a receiver located at an unknown angle of departure (AoD). The AoD defines discrete channel states, taking values in a discrete set of$M$possible values (quantized beam “directions”), constant within a coherence block and changing independently across blocks. Each block spans$Q$time slots of length$q$channel uses (also referred to as signal dimension). At the end of each slot, the transmitter receives a (strictly causal) feedback signal which may represent either the detection result of some radar sensor, or an explicit feedback signal from the receiver. The GBP model, a realistic extension of a binary beam-pointing channel studied in the authors' previous paper, offers a sufficiently simple yet insightful model for understanding channel capacity in beamforming-based communication systems. We establish both an upper bound and an approximate inner bound on capacity that can be calculated by solving carefully designed optimization problems. Numerical examples demonstrate that our proposed transmission strategy achieves a near-optimal achievable rate when the signal dimension$q$is large enough.
Siyao Li, Fernando Pedraza, Giuseppe Caire
ISIT1
2024 Reinforcement Learning Based Interference Coordination for Port Communications
abstract
Reliable port communications support maritime applications such as vessel navigation and cargo tracking for a large number of mobile users on ships, but the quality of services (QoS) such as data rate and energy consumption is severely degraded by inter-cell interference. In this paper, we propose a deep reinforcement learning (RL)-based interference coordination scheme for port communications to reduce the transmission latency and energy consumption, and improve the data rate. Based on the signal-to-interference plus noise ratio, the channel gains, the estimated interference levels and the transmission latency, the base station chooses the transmit power and downlink bandwidth constraint to avoid choosing risk policies that cause the communication performance degradation. In addition, a two-level hierarchical structure with two convolution networks and four fully connected layers is designed to reduce the algorithm complexity and enhance the convergence speed. Simulation results verify the performance gain of the proposed scheme in terms of the data rate, the transmission latency, and the energy consumption compared with the benchmark.
Siyao Li, Chuhuan Liu, Liang Xiao 0003, Helin Yang
VTC Spring1
2024 Interdomain Collaboration Between Hyperspectral and VHR Remote Sensing Images: A Cross-Scene Few-Shot Learning Framework for Change Detection
abstract
Hyperspectral image change detection (HSI-CD) based on deep learning (DL) has made significant progress. However, these methods rely significantly on the number of labeled data. Annotating HSI is a highly complex task that requires professional knowledge for guidance, resulting in a scarcity of high-quality labeled samples. The emergence of few-shot learning (FSL), which supports model learning from limited labeled samples, can address this issue. However, FSL-based methods still face some challenges: 1) existing methods mainly rely on single-source or homogenous cross-domain HSI data, which is difficult to adequately cope with the problem of scarcity of HSI labeled data; 2) most existing methods usually only focus on local features within patches and neglect interrelationships between patches, which is also important for model learning; and 3) transformers modeling long-range relationships rely on extensive labeled data, making it difficult to perform well in few-shot scenarios. Therefore we propose a cross-scene FSL framework based on interdomain collaboration (CSIDC-FSL) for HSI-CD. Specifically, the following is proposed: 1) FSL is performed on very high-resolution image (VHRI) and HSI, aiming to use the learnable information in VHRI with low annotation cost to help HSI-CD, reducing the dependence of the model on HSI annotation data while enabling multilevel feature hybrid perceptual CD; 2) a dual-information integrated mapping module (DI2M) is proposed, which designs a CNN and transformer integrated structure that can simultaneously focus on local features and class-wise long-range relationships to break the constraints of local perception of CNN while optimizing the performance of transformer under few-shot situations; and 3) the interdomain joint information allocation module (IDM) is designed to capture cross-scene domain-wise distribution features, and mitigate the impact of distribution differences in cross-scene data (VHRI and HSI) on knowledge learning and migration through the collaboratively consistent interdomain features. Under the condition of five samples per class, the CD results of CSIDC-FSL are better than those of recently advanced algorithms, with average improvements of 1.46%–1.5% for overall accuracy (OA) and average accuracy (AA), respectively. The code will be made available athttps://github.com/lsylnnu/CSIDC-FSL.
Xiang-Hai Wang 0001, Siyao Li, Xiao-Yang Zhao 0003, Yuetong Zhao
IEEE Trans. Geosci. Remote. Sens.2
2024 GTransCD: Graph Transformer-Guided Multitemporal Information United Framework for Hyperspectral Image Change Detection
abstract
Using multitemporal hyperspectral images (HSIs) to obtain fine-grained land cover change information is an essential task in remote sensing (RS) image processing. Convolutional neural networks (CNNs), which have strong feature extraction and nonlinear regression capabilities, have recently aided in the advancement of this subject. However, the performance of these supervised methods is usually limited by small receptive field and less labeled samples. To this end, how to break through the aforementioned bottleneck and build a more suitable change detection (CD) framework for HSI is a crucial and challenging issue. To this end, a graph transformer-guided multitemporal information united framework for HSI-CD (GTransCD) is proposed, which mainly consists of the following three components: 1) applying transformer to the graph structure, a salient relationship strengthening graph transformer (GTrans) module is created, making it possible for the network to capture distant change information, and on this basis, local- and global-range information are aggregated simultaneously; 2) a gated change information fusion (GCF) unit is designed to inject the GTrans-guided change features into the original bitemporal concatenated features to further enhance the representation of change information in the network; and 3) a general HSI-CD framework that can organically blend change features guided by GTrans module with original features is proposed, with the intention of reducing the reliance on training samples by utilizing the semi-supervised learning mode of graph neural networks. numerous experiments demonstrate the proposed GTransCD surpasses the state-of-the-art methods and has a high level even at low sampling rates. The source code of the proposed framework will be released athttps://github.com/zxylnnu/GTransCD.
Xiao-Yang Zhao 0003, Siyao Li, Tingting Geng, Xiang-Hai Wang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 GaMPF: A Full-Scale Gated Message Passing Framework Based on Collaborative Estimation for VHR Remote Sensing Image Change Detection
abstract
With the maturity and popularization of high-performance sensor technology, it is now possible to acquire huge amounts of very high-resolution (VHR) remote sensing images. The change detection (CD) for VHR images is currently receiving special attention for remote sensing earth observation applications, however, as a hot research field, it needs to be studied in depth to improve the detection accuracy of fine changes. To this end, a full-scale gated message passing framework (GaMPF) based on collaborative estimation for VHR remote sensing image change detection is proposed in this paper. On one hand, the key embedding representation is generated for each feature map by means of the collaborative estimation (CE) strategy; On the other hand, grounded in timing analysis, bitemporal features are sent selectively on dual paths according to the full-scale gated (FsG) mechanism. Specifically, this framework consists of the following four components: 1) Taking shared-weights Siamese network as an encoder to extract multi-scale features; 2) Generate a set of shared compact bases under the CE strategy and infer the key embedding representations on the basis of the shared bases for feature maps at the same level, considering the representations as the gated switches; 3) FsG mechanism is used as the mode of message passing between bitemporal images, which guides the information can be transmitted simultaneously on both within-and cross-temporal paths. 4) Creating a stepwise dense fusion module (DFM) as a decoder for predicting the change map. Experimental results show that the GaMPF proposed in this paper outperforms existing SOTA methods, and is particularly good at detecting edges and small objects. The source code will be released at https://github.com/zxylnnu/GaMPF.
Xiao-Yang Zhao 0003, Keyun Zhao, Siyao Li, Chuanming Song 0001, Xiang-Hai Wang 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Joint Fronthaul Load Balancing and Computation Resource Allocation in Cell-Free User-Centric Massive MIMO Networks
abstract
We consider scalable cell-free massive multiple-input multiple-output networks under an open radio access network paradigm comprising user equipments (UEs), radio units (RUs), and decentralized processing units (DUs). UEs are served by dynamically allocated user-centric clusters of RUs. The corresponding cluster processors (implementing the physical layer for each user) are hosted by the DUs as software-defined virtual network functions. Unlike the current literature, mainly focused on the characterization of the user rates under unrestricted fronthaul communication and computation, in this work we explicitly take into account the fronthaul topology, the limited fronthaul communication capacity, and computation constraints at the DUs. In particular, we systematically address the new problem of joint fronthaul load balancing and allocation of the computation resource. As a consequence of our new optimization framework, we present representative numerical results highlighting the existence of an optimal number of quantization bits in the analog-to-digital conversion at the RUs.
Zhiyang Li 0002, Fabian Goettsch, Siyao Li, Ming Chen 0001, Giuseppe Caire
IEEE Trans. Wirel. Commun.3
2023 On the Capacity and State Estimation Error of Binary "Beam-Pointing" Channels with Block Memory and Feedback
abstract
To counter the large isotropic pathloss in Millimeter-wave (mmWave) communications, high beamforming gain by means of large antenna arrays is required. Joint Communication and Sensing (JCAS) is predicted to be a major feature of future communication systems, but suitable channel models and performance metrics are still in development. This paper investigates the information-theoretic limits of JCAS using a channel model proposed in [1], which consists of a binary state-dependent channel with unit-delayed feedback and an in-block memory (iBM) [2] of the fixed block length. This model is sufficiently simple to treat, yet captures some fundamental aspects of beam acquisition (BA) where the feedback models the backscatter signal as in a radar system, thus fitting the paradigm of JCAS. When the transmission cost is bounded on average, we simplify the capacity computation as an optimization problem by exploring a small set of optimal average input costs at each channel use based on feedback. The converse proof inspires a capacity-achieving strategy that uniformly explores a given number of beam directions under the cost constraint. The sensing performance is characterized in terms of the state estimation error. We present the minimum distortion without communication constraint as another optimization problem, which can be solved by the successive convex approximation (SCA) method similar to the capacity computation problem. Numerical examples are provided to illustrate the performance of the proposed strategies and show a very small gap between the estimation error achieved by the capacity-achieving strategy and the minimum possible error. Hence, for the model at hand sensing comes essentially "for free", i.e., there exists a JCAS strategy that pays very little in terms of sensing performance while being optimal with respect to the communication rate.
Siyao Li, Giuseppe Caire
ISIT1
2023 On the State Estimation Error of "Beam-Pointing" Channels: The Binary Case
abstract
Sensing capabilities as an integral part of the network have been identified as a novel feature of sixth-generation (6G) wireless networks. As a key driver, millimeter-wave (mmWave) communication largely boosts speed, capacities, and connectivity. In order to maximize the potential of mmWave communication, precise and fast beam acquisition (BA) is crucial, since it compensates for a high pathloss and provides a large beamforming gain. Practically, the angle-of-departure (AoD) remains almost constant over numerous consecutive time slots, the backscatter signal experiences some delay, and the hardware is restricted under the peak power constraint. This work captures these main features by a simple binary beam-pointing (BBP) channel model with in-block memory (iBM) [1], peak cost constraint, and one unit-delayed feedback. In particular, we focus on the sensing capabilities of such a model and characterize the performance of the BA process in terms of the Hamming distortion of the estimated channel state. We encode the position of the AoD and derive the minimum distortion of the BBP channel under the peak cost constraint with no communication constraint. Our previous work [2] proposed a joint communication and sensing (JCAS) algorithm, which achieves the capacity of the same channel model. Herein, we show that by employing this JCAS transmission strategy, optimal data communication and channel estimation can be accomplished simultaneously. This yields the complete characterization of the capacity-distortion tradeoff for this model.
Siyao Li, Giuseppe Caire
ITW1
2023 GTMSiam: Gated Transmitting-Based Multiscale Siamese Network for Hyperspectral Image Change Detection
abstract
Hyperspectral image change detection (HSI-CD) is a technique that detects changes in land cover occurring in a specific area within a closed time. At present, most existing methods for HSI-CD employ exceedingly intricate network architectures, leading to a high model complexity that hampers the achievement of a favorable trade-off between change detection accuracy and timeliness. Furthermore, existing methods often confine the feature extraction process to a single scale rather than multiple diverse scales. However, employing a multiscale approach for feature extraction allows for capturing finer-grained features encompassing more intricate details, as well as coarser-grained features that aggregate local information over a larger range. On the other hand, most existing methods overemphasize the complexity of the feature extraction process and underestimate the importance of the conversion process from bi-temporal features to valuable change features. To this end, a gated transmitting based multiscale siamese network (GTMSiam) is proposed, which mainly contains the following two portions: 1) dual branches with the siamese structure, which capture spatial features of the HSIs at multiple scales while preserving rich spectral information. Moreover, the siamese design effectively reduces the network parameters, thereby alleviating the computational complexity of the model. 2) gated change information transmitting module (GTM), which utilizes gated neural units to transform bi-temporal image features into land cover change information, while progressively transmitting change information at different scales. This enables the network to leverage diverse scale change information for comprehensive discrimination of land object changes. Experimental results on three publicly available datasets demonstrate the superior performance of the proposed GTMSiam. Simultaneously, the complexity analysis experiment proves that the GTMSiam can give consideration to both detection performance and timeliness. The source code of this letter will be released at https://github.com/zkylnnu/GTMSiam.
Xiang-Hai Wang 0001, Keyun Zhao, Xiao-Yang Zhao 0003, Siyao Li
IEEE Geosci. Remote. Sens. Lett.4
2023 BiG-FSLF: A Cross Heterogeneous Domain Few-Shot Learning Framework Based on Bidirectional Generation for Hyperspectral Image Change Detection
abstract
In recent years, hyperspectral image change detection (HSI-CD) based on deep learning has achieved high detection accuracy, but these methods obtain excellent detection results usually rely on having sufficient labeled samples to train the network. However, the production of HSI label is difficult, costly and inefficient. In practical tasks, often only a limited number of labeled samples can be obtained due to the limitation of timeliness. To address this problem, a cross heterogeneous domain few-shot learning framework based on bidirectional generation (BiG-FSLF) is proposed for HSI-CD, which aims to solve the few-shot problem of HSI-CD by few-shot learning (FSL), and to assist HSI-FSL perform better by obtaining learnable changed information (i.e., empirical knowledge) from another remote sensing data. Specifically, a multitask generation encoder (MLGenE) is designed to take on both the tasks of FSL and domain adaptation to achieve HSI-CD under the condition of cross heterogeneous domain few-shot. First, we take any pair of image data in a very high resolution image (VHRI) CD dataset as the source domain and HSI is used as the target domain, using sufficient labeled samples in source domain and a small number of labeled samples in target domain for FSL. Meanwhile, a bidirectional generation domain adaptation (BiGDA) method based on generative adversarial strategy is proposed to achieve adaptive alignment of the two heterogeneous domains (source and target domains) feature distributions, to mitigate the impact of the domain shift problem inherent to cross domain data on FSL. Abundant experiments with only five training samples on the publicly available popular HSI-CD datasets confirm that the proposed method can show great detection performance. The source code of the proposed framework will be released at https://github.com/lsylnnu/BiG-FSLF.
Xiang-Hai Wang 0001, Siyao Li, Xiao-Yang Zhao 0003, Keyun Zhao
IEEE Trans. Geosci. Remote. Sens.2
2023 TriTF: A Triplet Transformer Framework Based on Parents and Brother Attention for Hyperspectral Image Change Detection
abstract
Hyperspectral image (HSI) change detection (CD) is a technique to accurately detect land cover changes by using HSIs with rich spatial-spectral information. In recent years, the HSI-CD methods based on convolutional neural networks (CNNs) have achieved great success because of their flexible and effective feature extraction ability. However, these methods often take the HSI patches as the input of the networks, which undoubtedly hinders the overall perception of the HSIs. Meanwhile, the valuable temporal information in HSIs is often underutilized. For this end, a triplet transformer framework (TriTF) based on parents-temporal attention and brother-spatial attention is proposed for HSI-CD. The proposed framework mainly contains the following three parts: 1) Transformer-based network backbone, which uses the self-attention to capture the correlation between arbitrarily two pixels in the same patch and extracts the global spatial correlation in the unit of encoded input patches; 2) parents-temporal attention (PTA) branch. Unlike the previous cross-temporal attention mechanisms of the “T1↔T2” mode which only consider the interaction between bi-temporal HSIs, this paper constructs a novel PTA of the “T1→T3←T2” mode which takes the difference-temporal image T3 as the core. The impact of bi-temporal HSIs on the land cover changes is more concerned in the PTA; 3) brother-spatial attention (BSA) branch. The most similar patch in the current training batch of each patch is defined as its brother patch. Furthermore, cross-spatial attention is applied to propagate the features of the brother patch to the current patch. Thus, the middle- and long-range dependencies can be utilized and the scope of feature propagation can be extended. In this paper, the experiments under low and high sampling rates are conducted and proved the outstanding change detection performance of the proposed TriTF when compared with abundant state-of-the-art (SOTA) CD algorithms. The source code of this paper will be released at https://github.com/zkylnnu/TriTF.
Xiang-Hai Wang 0001, Keyun Zhao, Xiao-Yang Zhao 0003, Siyao Li
IEEE Trans. Geosci. Remote. Sens.4
2023 GeSANet: Geospatial-Awareness Network for VHR Remote Sensing Image Change Detection
abstract
The characteristics of very high resolution (VHR) remote sensing images (RSIs) have higher spatial resolution inherently, and are easier to obtain globally compared with hyperspectral images (HSIs), making it possible to detect small-scale land cover changes in multiple applications. RSI change detection (RSI-CD) based on deep learning has been paid attention to and become a frontier research field in recent years, and is currently facing two challenging problems: The first is high dependence on registration between bi-temporal images caused by high spatial resolution; The other is high pseudo-change information response caused by low spectral resolution. In order to address the above-mentioned two problems, a novel RSI-CD framework called Geospatial-Awareness Network (GeSANet) based on the geospatial Position Matching Mechanism (PMM) with multi-level adjustment and the geo-spatial Content Reasoning Mechanism (CRM) with diverse pseudo-change information filtering is proposed. First of all, the PMM assigns independent two-dimensional offset coordinates to each position in the previous temporal image, afterwards, bilinear interpolation is employed to obtain the subpixel feature value after the offset, and the sparse results based on the difference are transmitted to the next level prediction to realize multi-level geospatial correction. The CRM extracts global features from the corrected sparse feature map in terms of dimensions, implementing effective discriminant feature extraction on basis of the original feature map in a stepwise refinement manner through the cross-dimension exchange mechanism, to filter out various pseudo-change information as well as maintain real change information. Comparison experiments with five recent SOTA methods are carried out on two popular datasets with diverse changes, the results show that the proposed method has good robustness and validity for multi-temporal RSI-CD. In particular, it has a strong comparative advantage in detecting small entity changes and edge details. The source code of the proposed framework can be downloaded from https://github.com/zxylnnu/GeSANet.
Xiao-Yang Zhao 0003, Keyun Zhao, Siyao Li, Xiang-Hai Wang 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 On the Capacity and State Estimation Error of "Beam-Pointing" Channels: The Binary Case
abstract
Motivated by beamformed millimeter-wave (mmWave) communication, we consider the optimal tradeoff between reliable communication rate and state estimation error for a new state-dependent channel model with in-block memory referred to as the binary beam-pointing (BBP) channel. The multiantenna base station (BS) uses a finite beamforming codebook (i.e., discrete beam directions) and associates the target user to the most favorable beam, i.e., the beam directed along the strongest propagation path from BS to the user in 5G and IEEE 802.11ad mmWave communication systems. We model the target user’s Angle-of-Departure (AoD) as the state of a state-dependent channel. Since the AoD remains almost constant over several consecutive time slots, the channel has memory. In addition, we assume the BS receives implicit causal feedback (e.g., modeling a backscatter signal as in radar) and consider a joint communication and sensing (JCAS) problem where the BS is also interested in explicitly estimate the channel state (quantized AoD). We derive a closed-form solution to the capacity of the BBP channel model under the peak input cost constraint. Under the average constraint, we present an implicit capacity result, where the capacity is given as the solution of an optimization problem, and we provide an algorithm for its numerical computation. Finally, for the JCAS problem at hand, we provide the minimum distortion under the peak input constraint and show that this coincides with that obtained by the capacity-achieving strategy. This completely characterizes the capacity-distortion tradeoff for the BBP channel under peak input constraint. For the average constraint case, numerical results show that our capacity-achieving strategy yields a state estimation error very close to the theoretical minimum, showing the near-optimality of the proposed strategy for the JCAS problem under the average cost constraint.
Siyao Li, Giuseppe Caire
IEEE Trans. Inf. Theory1
2022 Recommendations for Visualization Recommendations: Exploring Preferences and Priorities in Public Health
abstract
The promise of visualization recommendation systems is that analysts will be automatically provided with relevant and high-quality visualizations that will reduce the work of manual exploration or chart creation. However, little research to date has focused on what analysts value in the design of visualization recommendations. We interviewed 18 analysts in the public health sector and explored how they made sense of a popular in-domain dataset1 in service of generating visualizations to recommend to others. We also explored how they interacted with a corpus of both automatically- and manually-generated visualization recommendations, with the goal of uncovering how the design values of these analysts are reflected in current visualization recommendation systems. We find that analysts champion simple charts with clear takeaways that are nonetheless connected with existing semantic information or domain hypotheses. We conclude by recommending that visualization recommendation designers explore ways of integrating context and expectation into their systems.
Calvin Bao, Siyao Li, Sarah G. Flores, Michael Correll, Leilani Battle
CHI2
2022 Deep Learning-Aided Coding for the Fading Broadcast Channel with Feedback
abstract
We consider the design of practical codes for a symmetric two-user fading Gaussian Broadcast Channel (BC) with feedback. We construct a two-phase coding scheme with the help of deep Neural Networks (NNs) that seeks to optimize the encoder and decoders jointly. Interpreting a communication system as an autoencoder (denoted by AE), we train the AE under various scenarios of noiseless feedback signals. Performance evaluation is presented for Rayleigh distributed channel state, which reveals the existence of a trained NN-based two-phase model that outperforms state-of-the-art codes in the low SNR regime. Considering the availability of feedback signals, we train the AE with different inputs, and observe that feedback consisting of received signals appears to be more beneficial than channel states to boost reliability under the proposed scheme. We provide initial interpretations of the encoding scheme which uses channel state feedback.
Siyao Li, Daniela Tuninetti, Natasha Devroye
ICC1
2022 A reverse personnel assignment method with duration re-inferring for smart "IOT+ blockchain" project
abstract
As a decentralized and distrusted distributed ledger technology, blockchain is gradually applied in the IOT. Cost overrun are inherent part of most smart “IOT+ blockchain” projects. In order to guarantee a successful delivery of a smart “IOT+ blockchain” project with the ideal budget, with respect to the minimum cost of the forward problem is still higher than the approved budget, this research proposes a re-verse optimization method of 0-1 mixed-integer, bi-level programming model for reverse-inferring duration and personnel re-assignment. Based on a numerical experiment to a “IOT+ blockchain” construction project, the comparative results show that the reverse optimization method is superior to the forward method in terms of total cost reduction and can further shorten the duration. The result indicates that the reverse optimization methodology can be applied in scenarios which need to guarantee the objective value achieved through the proposed reverse modelling methodology by optimizing parameters and decision variables.
Di Su, Siyao Li
KES5
2022 GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention Refinement
abstract
Grounded Situation Recognition (GSR) aims to generate structured semantic summaries of images for "human-like'' event understanding. Specifically, GSR task not only detects the salient activity verb (e.g. buying), but also predicts all corresponding semantic roles (e.g. agent and goods). Inspired by object detection and image captioning tasks, existing methods typically employ a two-stage framework: 1) detect the activity verb, and then 2) predict semantic roles based on the detected verb. Obviously, this illogical framework constitutes a huge obstacle to semantic understanding. First, pre-detecting verbs solely without semantic roles inevitably fails to distinguish many similar daily activities (e.g., offering and giving, buying and selling). Second, predicting semantic roles in a closed auto-regressive manner can hardly exploit the semantic relations among the verb and roles. To this end, in this paper we propose a novel two-stage framework that focuses on utilizing such bidirectional relations within verbs and roles. In the first stage, instead of pre-detecting the verb, we postpone the detection step and assume a pseudo label, where an intermediate representation for each corresponding semantic role is learned from images. In the second stage, we exploit transformer layers to unearth the potential semantic relations within both verbs and semantic roles. With the help of a set of support images, an alternate learning scheme is designed to simultaneously optimize the results: update the verb using nouns corresponding to the image, and update nouns using verbs from support images. Extensive experimental results on challenging SWiG benchmarks show that our renovated framework outperforms other state-of-the-art methods under various metrics.
Zhi-Qi Cheng, Qi Dai 0001, Siyao Li, Teruko Mitamura, Alex Hauptmann 0001
ACM Multimedia3
2022 Demonstration of VegaPlus: Optimizing Declarative Visualization Languages
abstract
While many visualization specification languages are user-friendly, they tend to have one critical drawback: they are designed for small data on the client-side and, as a result, perform poorly at scale. We propose a system that takes declarative visualization specifications as input and automatically optimizes the resulting visualization execution plans by offloading computational-intensive operations to a separate database management system (DBMS). Our demo emphasizes live programming of visualizations over big data, enabling users to write or import Vega specifications, view the optimized plans from our system, and even modify these plans and compare their performance via a dedicated performance dashboard.
Junran Yang, Hyekang Joo, Sai S. Yerramreddy, Siyao Li, Dominik Moritz, Leilani Battle
SIGMOD Conference4
2022 CSDBF: Dual-Branch Framework Based on Temporal-Spatial Joint Graph Attention With Complement Strategy for Hyperspectral Image Change Detection
abstract
Hyperspectral image (HSI) change detection (CD) aims at obtaining internal components’ change information of land cover and land use. In recent years, the development of convolutional neural networks (CNNs) has greatly promoted the research progress in this field. However, the fixed small-size convolution kernels used by CNNs have severely limited the receptive field of information. Another defect of most CNN-based models is their strong dependence on samples, and they are not competent for tasks with a small number of samples. Besides, the traditional CNN-based models can only perform convolution to learn the spatial–spectral features in the Euclidean space, which is not conducive to capturing the geometric changes in land covers in the HSIs. Differently, the graph attention network (GAT) has come into prominence due to its ability to capture the holistic topology structure of images flexibly, and the attention coefficients can be used to effectively model the long-range correlations between land covers. The semi-supervised nature of GAT is also well-suited to handle HSI-CD tasks with limited samples. Nevertheless, the pixel-level topology structure often generates expensive computational costs. To this end, a dual-branch framework based on temporal–spatial joint graph attention (TSJGAT) with complement strategy (CSDBF) is proposed for HSI-CD, which extracts superpixel- and pixel-level features from bitemporal HSIs in parallel and enables them to complement each other. The proposed CSDBF mainly consists of two branches: superpixel-level feature extraction branch (S-branch) and pixel-level feature extraction branch (P-branch). In the S-branch, we introduce the idea of GAT into HSI-CD for the first time and propose a novel TSJGAT module. Thus, the temporal–spatial features of HSIs are propagated and aggregated on the nonlinear graph structure, which makes the changed regions more discriminable. In the P-branch, pixel-level features are obtained by CNNs to correct uncertain factors caused by superpixel segmentation in the S-branch, which is complementary to the S-branch and lays a foundation for more accurate CD. Abundant experiments show that compared with other pioneer methods, the proposed CSDBF can improve the Kappa coefficient by more than 1.9% and 2.5% on average in general sampling rate situations and a low sampling rate situation, respectively, which shows better robustness and better detection accuracy than most existing state-of-the-art methods. The source code of this article can be downloaded fromhttps://github.com/zkylnnu/CSDBF.
Xiang-Hai Wang 0001, Keyun Zhao, Xiao-Yang Zhao 0003, Siyao Li
IEEE Trans. Geosci. Remote. Sens.4
2021 Guiding the Growth: Difficulty-Controllable Question Generation through Step-by-Step Rewriting
abstract
Yi Cheng, Siyao Li, Bang Liu, Ruihui Zhao, Sujian Li, Chenghua Lin, Yefeng Zheng. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Siyao Li, Bang Liu 0003, Ruihui Zhao, Sujian Li, Chenghua Lin 0002, Yefeng Zheng 0001
ACL/IJCNLP (1)2
2021 A Control-Theoretic Linear Coding Scheme for the Fading Gaussian Broadcast Channel with Feedback
abstract
This paper proposes a linear coding scheme for the two-user fading additive white Gaussian noise broadcast channel, under the assumptions that: (i) perfect Channel State Information (CSI) is available at the receivers; and (ii) unit delayed CSI along with channel output feedback (COF) is available at the transmitter. The proposed scheme is derived from a control-theoretic perspective that generalizes the communication scheme for the point-to-point (P2P) fading Gaussian channel under the same assumptions by Liu et al. [1]. The proposed scheme asymptotically achieves the rates of a posterior matching scheme, from the same authors, for a certain choice of parameters.
Siyao Li, Daniela Tuninetti, Natasha Devroye
ISIT1
2020 The Fading Gaussian Broadcast Channel with Channel State Information and Output Feedback
abstract
The fading broadcast channel (BC) with additive white Gaussian noise (AWGN) channel, channel output feedback (COF) and channel state information (CSI) is considered. Perfect CSI is available at the receivers, and unit delayed CSI along with COF at the transmitter. Under the assumption of memoryless fading, a posterior matching scheme that incorporates the additional CSI feedback into the coding scheme is presented. With COF, the achievable rates depend on the joint distribution of the fading process. Numerical examples show that the capacity region of two-user fading AWGN-BC is enlarged by COF. The coding scheme is however suboptimal since some parts of the achievable rate region are outperformed by superposition coding without COF.
Siyao Li, Daniela Tuninetti, Natasha Devroye
ISIT1
2019 Deep Reinforcement Learning with Distributional Semantic Rewards for Abstractive Summarization
abstract
Siyao Li, Deren Lei, Pengda Qin, William Yang Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Siyao Li, Deren Lei, Pengda Qin, William Yang Wang
EMNLP/IJCNLP (1)1
2019 On the Capacity Region of the Layered Packet Erasure Broadcast Channel with Feedback
abstract
In this paper the capacity region of the Layered Packet Erasure Broadcast Channel (LPE-BC) with Channel Output Feedback (COF) available at the transmitter is investigated. The LPE-BC is a high-SNR approximation of the fading Gaussian BC recently proposed by Tse and Yates, who characterized the capacity region for any number of users and any number of layers when there is no COF. This paper derives capacity inner and outer bounds for the LPE-BC with COF for the case of two users and any number of layers. The inner bounds generalize past results for the two-user erasure BC, which is a special case of the LPE-BC with COF with only one layer. The novelty lies in the use of inter-user & inter-layer network coding retransmissions (for those packets that have only been received by the unintended user), where each random linear combination may involve packets intended for any user originally sent on any of the layers. Analytical and numerical examples show that the proposed outer bound is optimal for some LPE-BCs.
Siyao Li, Daniela Tuninetti, Natasha Devroye
ICC1
2019 On The Stability Region of the Layered Packet Erasure Broadcast Channel with Output Feedback
abstract
This paper studies the Layered Packet Erasure Broadcast Channel (LPE-BC) with Channel Output Feedback (COF), which is a high-SNR approximation of the fading Gaussian BC, proposed by Tse and Yates in 2012 for the case without COF. This model is also a multi-layer generalization of the Binary Erasure Channel (BEC). In a past work, the Authors derived inner and outer bounds to the rate region (set of achievable rates with backlogged arrivals) of the LPE-BC with COF; here, the arrival region (set of exogenous arrival rates for which packet arrival queues are stable) for the same model is analyzed. For the case of K=2 users and Q ≥ 1 layers, the known achievable rate region and the derived arrival region coincide; both strategically employ a. For the case of Q = 2 layers, sufficient conditions are given for the achievable arrival region to coincide with the known converse rate region, thus showing that in those cases the optimal rate and arrival regions coincide.
Siyao Li, Hulya Seferoglu, Daniela Tuninetti, Natasha Devroye
ITW1
2016 Multi-Agent System Development MADE Easy
abstract
Agent-Oriented Software Engineering (AOSE) is an emerging software engineering paradigm that advocates the application of best practices in the development of Multi-Agent Systems (MAS) through the use of agents and organizations of agents. This paper outlines the MADE system, which provides an interactive platform for people who are not well-versed in AOSE to contribute to the rapid prototyping of MASs with ease.
Zhiqi Shen 0001, Han Yu 0001, Chunyan Miao, Siyao Li, Yiqiang Chen 0001
AAAI4
2015 Teachable Agents with Intrinsic Motivation
Ailiya, Chunyan Miao, Su Fang Lim, Siyao Li, Zhiqi Shen 0001
AIED4