VLDB 2026 Research / reviewers in the wild / expert
Danlan Huang
dblp:198/9583
· DBLP profile ↗
17ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaJSCCV: Adaptive Prompt-Tuning for Semantic Video Transmission
Danlan Huang, Zhixin Qi, Xiyang Wang 0009 |
WCNC | 1 |
| 2026 | Integrated sensing, communication, and control for multi-agent networked formation control
Zhiyong Feng 0001, Zhiqing Wei, Dingyou Ma, Danlan Huang, Zeyang Meng, Yinglong Fan, Jie Xu 0002, Ping Zhang 0003 |
Sci. China Inf. Sci. | 5 |
| 2025 | Efficient Wireless Video Transmission via Adaptive Spatio-Temporal Token MergingabstractThe rapid proliferation of emerging video services has substantially increased the demand for real-time transmission of high-definition video content. However, optimizing the trade-off between video transmission quality and communication bandwidth remains a significant challenge. To address this issue, we propose RAJSCC, a rate-adaptive video deep joint source-channel coding (DeepJSCC) framework. RAJSCC incorporates a spatio-temporal token merging mechanism to aggregate semantically similar tokens, effectively reducing redundancy both within and across frames. Additionally, an adaptive token merging predictor, designed based on simple statistical features of the input videos, enables dynamic rate control at the group of pictures (GoP) level, ensuring smooth and continuous variation in the overall video coding rate. Extensive experiments demonstrate that RAJSCC significantly outperforms traditional video transmission schemes, such as H.264 and H.265 with low-density parity-check (LDPC), as well as existing DeepJSCC methods, in terms of reconstruction quality. More importantly, the proposed adaptive spatio-temporal token merging mechanism reduces bandwidth consumption by 63.5% and computational cost by 13.0%, while incurring only a marginal 1–2 dB degradation in reconstruction quality. These findings highlight the effectiveness of RAJSCC in achieving a superior balance between transmission efficiency and video quality, making it a promising solution for real-time high-definition video communication in bandwidth-constrained environments. Xinyi Zhou 0015, Danlan Huang, Zhixin Qi, Ting Jiang 0008 |
GLOBECOM | 2 |
| 2025 | Static-dynamic class-level perception consistency in video semantic segmentation
Zhigang Cen, Ningyan Guo, Zhiyong Feng 0001, Danlan Huang |
Neural Networks | 5 |
| 2024 | Wi-Mapping: A WiFi-based Respiration Detection System Using Complex Plane MappingabstractRespiration serves as a critical indicator of human health, reflecting the well-being of various bodily organs. The potential of device-free WiFi signals for respiration detection has been shown in recent investigations. Nevertheless, there are several disadvantages to the current WiFi signal-based respiration detection methods. These include inadequate utilization of the respiratory component within the signal, and the high-quality subcarriers cannot be screened out effectively. In response to these issues, we propose a novel respiration detection system named Wi-Mapping. Firstly, to enhance the utilization of the respiratory component within the signal, we introduce a novel complex plane mapping approach. This method reconstructs a signal with a more noticeable respiratory component by integrating the original amplitude and phase of Channel State Information (CSI). Wi-Mapping then concentrates on utilizing features in the frequency domain and subcarrier dimension. Additionally, a subcarrier screening method is proposed, which combines Respiration Energy Ratio (RER) and correlation between subcarriers. This method can effectively screen out subcarriers with higher quality. Finally, a dataset containing various scenes was established to confirm Wi-Mapping’s respiration detecting capabilities. The detection error achieved by Wi-Mapping is 0.382, with a detection rate of 86.91%, surpassing that of state-of-the-art methods. Ting Jiang 0008, Xinyi Zhou 0015, Danlan Huang |
GLOBECOM | 4 |
| 2024 | Wi-locind: Location-Independent Respiration Sensing Based on WIFI CSIabstractCurrent WiFi-based respiration detection algorithms may experience performance degradation due to variations in user location, as the relationship between user location and patterns of respiration has not been adequately considered. To overcome this limitation, this paper proposes a spatially directional respiration detection approach, named Wi-locind. Wi-locind employs antenna arrays on commercial WiFi receivers to achieve directional enhancement of respiration signals. Combined with post-filtering techniques, Wi-locind is capable of extracting respiration patterns that are independent of changes in the user's location. Specifically, the Minimum Variance Distortionless Response algorithm is used to identify the arrival angle of the target user and directionally enhance the received signal in the corresponding direction. The Empirical Mode Decomposition algorithm is subsequently utilized to suppress the environmental noise and time domain artifacts caused by the enhancement method, enabling the extraction of the target's respiration pat-tern. Our results show that the proposed approach consistently achieves an average absolute error of less than 0.3 breaths per minute across all positions, significantly outperforming the baseline approaches. Ting Jiang 0008, Xinyi Zhou 0015, Danlan Huang |
WCNC | 4 |
| 2024 | Boosting Scene Graph Generation with Contextual InformationabstractScene graph generation (SGG) has been developed to detect objects and their relationships from the visual data and has attracted increasing attention in recent years. Existing works have focused on extracting object context for SGG. However, very few works have attempted to exploit implicit contextual correlations among relationships of the objects. Furthermore, most existing SGG schemes rely on high-level features to predict the predicates while overlooking the potential inherent association of low-level features with the object relationships. We present in this article a novel scheme to capture enhanced contextual information for both objects and relationships. We design a Dual-branch Context Analysis Transformer (DCAT) architecture to extract both object context and relationship context from the visual data with dual transformer branches and then effectively fuse both high-level and low-level features by an adaptive approach to facilitate relationship prediction. Specifically, we first conduct feature representation learning to enrich relation representations by the visual, spatial, and linguistic feature extractors. Next, two transformer branches are designed to leverage the modeling of global associative interaction and mine the hidden association among objects and relationships. Then, we devise a novel feature disentangling method to decouple contextualized high-level features with guidance from the visual semantics. Finally, we develop a refined attention module to perform low-level feature recalibration for the refinement of the final predicate prediction. Experiments on Visual Genome and Action Genome datasets demonstrate the effectiveness of DCAT for both image and video SGG settings. Moreover, we also test the quality of the generated image scene graphs to verify the generalizability on downstream tasks like sentence-to-graph retrieval and image retrieval. Shiqi Sun 0002, Danlan Huang, Xiaoming Tao 0001, Chengkang Pan, Guangyi Liu 0001, Chang Wen Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | USGG: Union Message Based Scene Graph GenerationabstractScene graph generation (SGG) is designed to represent images by objects and their relationships. Existing works mainly attempt to strengthen object pair representations for SGG. However, most methods ignore the significant semantic information implied in union regions, which refers to the surrounding area of object pairs. In this paper, we propose a new union message based architecture, named as USGG, to profoundly exploit the relational semantics of unions to facilitate SGG. Concretely, we employ sufficient feature extraction to enhance the features of objects and unions. Next, we devise the Union Embedding Network to model the relational representations through two symmetric encoder-decoder branches. Moreover, the Union Fusion Network is designed to integrate the refined semantics by two-stage feature fusion. Extensive experiments are conducted on Visual Genome dataset, which demonstrates that the proposed approach achieves competitive performance against state-of-the-art methods on Recall, mean Recall and Zero Shot Recall metrics. Shiqi Sun 0002, Danlan Huang, Zhijin Qin, Xiaoming Tao 0001, Chengkang Pan, Guangyi Liu 0001 |
ICIP | 2 |
| 2023 | DeformSg2im: Scene graph based multi-instance image generation with a deformable geometric layout
Yuxiao Li 0001, Danlan Huang, Juan Wang 0012, Ning Ge 0001, Jianhua Lu |
Neurocomputing | 3 |
| 2023 | Toward Semantic Communications: Deep Learning-Based Image Semantic CodingabstractSemantic communications has received growing interest since it can remarkably reduce the amount of data to be transmitted without missing critical information. Most existing works explore the semantic encoding and transmission for text and apply techniques in Natural Language Processing (NLP) to interpret the meaning of the text. In this paper, we conceive the semantic communications for image data that is much more richer in semantics and bandwidth sensitive. We propose an reinforcement learning based adaptive semantic coding (RL-ASC) approach that encodes images beyond pixel level. Firstly, we define the semantic concept of image data that includes the category, spatial arrangement, and visual feature as the representation unit, and propose a convolutional semantic encoder to extract semantic concepts. Secondly, we propose the image reconstruction criterion that evolves from the traditional pixel similarity to semantic similarity and perceptual performance. Thirdly, we design a novel RL-based semantic bit allocation model, whose reward is the increase in rate-semantic-perceptual performance after encoding a certain semantic concept with adaptive quantization level. Thus, the task-related information is preserved and reconstructed properly while less important data is discarded. Finally, we propose the Generative Adversarial Nets (GANs) based semantic decoder that fuses both locally and globally features via an attention module. Experimental results demonstrate that the proposed RL-ASC is noise robust and could reconstruct visually pleasant and semantic consistent image in low bit rate condition. Danlan Huang, Feifei Gao 0001, Xiaoming Tao 0001, Qiyuan Du, Jianhua Lu |
IEEE J. Sel. Areas Commun. | 1 |
| 2022 | A Robust Deep Learning Enabled Semantic Communication System for TextabstractWith the advent of the 6G era, the concept of semantic communication has attracted increasing attention. Compared with conventional communication systems, semantic communication systems are not only affected by physical noise existing in the wireless communication environment, e.g., additional white Gaussian noise, but also by semantic noise due to the source and the nature of deep learning-based systems. In this paper, we elaborate on the mechanism of semantic noise. In particular, we categorize semantic noise into two categories: literal semantic noise and adversarial semantic noise. The former is caused by written errors or expression ambiguity, while the latter is caused by perturbations or attacks added to the embedding layer via the semantic channel. To prevent semantic noise from influencing semantic communication systems, we present a robust deep learning enabled semantic communication system (R-DeepSC) that leverages a calibrated self-attention mechanism and adversarial training to tackle semantic noise. Compared with baseline models that only consider physical noise for text transmission, the proposed R-DeepSC achieves remarkable performance in dealing with semantic noise under different signal-to-noise ratios. Zhijin Qin, Danlan Huang, Xiaoming Tao 0001, Jianhua Lu, Guangyi Liu 0001, Chengkang Pan |
GLOBECOM | 3 |
| 2022 | Category-Adaptive Domain Adaptation for Semantic SegmentationabstractUnsupervised domain adaptation (UDA) becomes more and more popular in tackling real-world problems without ground truths of the target domain. Though tedious annotation work is not required, UDA unavoidably faces two problems: 1) how to narrow the domain discrepancy to boost the transferring performance; 2) how to improve the pseudo annotation producing mechanism for self-supervised learning (SSL). In this paper, we focus on UDA for semantic segmentation tasks. Firstly, we introduce adversarial learning into style gap bridging mechanism to keep the style information from two domains in a similar space. Secondly, to keep the balance of pseudo labels on each category, we propose a category-adaptive threshold mechanism to choose category-wise pseudo labels for SSL. The experiments are conducted using GTA5 as the source domain, Cityscapes as the target domain. The results show that our model outperforms the state-of-the-arts with a noticeable gain on cross-domain adaptation tasks. Yantian Luo, Danlan Huang, Ning Ge 0001, Jianhua Lu |
ICASSP | 3 |
| 2021 | Deep Learning-Based Image Semantic Coding for Semantic CommunicationsabstractThis paper presents the Generative Adversarial Networks (GANs)-based image semantic coding, the goal of which is semantic exchange rather than symbol transmission. State-of-the-art visually pleasing reconstruction and semantic preserving performance are obtained in extreme low bitrate via a rate-perception-distortion optimization framework. In particular, we investigate convolutional encoder, quantizer, conditional SPADE generator, residual coding as well as perceptual losses. In contrast to previous work, we designed a coarse-to-fine image semantic coding model for multimedia semantic communication system. The base layer of the image is fully generated and preserves semantic information while the enhancement layer restores the fine details. We explore the perception and distortion performance trade-off by tuning the rate of base layer and enhancement layer. Different from the existing methods that adopt pixel accuracy as distortion metric, we train and evaluate the proposed image semantic coding model with multiple perception metrics, in line with the purpose of semantic communications. Experimental results demonstrate that our model could achieve visually pleasant and semantic consistent reconstruction, as well as saving times of bitrate, compared to BPG, WebP, JPEG2000, JPEG, and other deep learning-based image codecs. Danlan Huang, Xiaoming Tao 0001, Feifei Gao 0001, Jianhua Lu |
GLOBECOM | 1 |
| 2021 | Deformable Geometry based Semantic Reconstruction from Scene GraphsabstractStructural scene graph based image generation provides a new paradigm for image-oriented semantic communications, whose goal is the semantic level rather than pixel-level reconstruction. The challenges include capturing relationships between objects and producing a reasonable geometric layout for each object accordingly. However, category information alone is not instructive enough for the generation process at the receiver side. Moreover, it is worth effort to extract the spatial dependencies among different objects in an image, therefore determine the object layouts on the whole instead of in an independent manner. In this paper, a deformable geometry framework for scene graph based image generation is proposed, in order to reconstruct images with higher semantic fidelity and visual pleasure. In particular, we introduce shape and appearance information to guide the generation process, from the scope of statistic modeling. Furthermore, we apply a spatial warping network to conduct geometric deformations on the layouts of different objects. Qualitative and quantitative experiments illustrate the superiority of our model compared to the state-of-the-art Sg2im method. Yuxiao Li 0001, Danlan Huang, Yantian Luo, Ning Ge 0001, Jianhua Lu |
GLOBECOM | 3 |
| 2020 | Trace-Driven QoE-Aware Proactive Caching for Mobile Video Streaming in MetropolisabstractTo meet the ever-increasing demands for mobile video streaming, proactive caching over the network edge has been proposed as a promising solution for next generation wireless networks. In this paper, we consider the trace-driven cache-enabled video streaming design in the scenario of a metropolis to boost the spectral efficiency on the system side and the quality of experience (QoE) on the user side. A novel scheme to jointly provide proactive caching, power allocation, user association and adaptive video streaming is designed via the formation of a QoE-aware throughput maximization problem. Specifically, the caches are refreshed in the content placement phase according to the resource status and expected traffic, which is obtained by exploring the traces collected over a big city. In addition, users need to be associated with a proper small base station (SBS) in the content delivering phase to provide the highest attainable rate. We demonstrate the effectiveness of the proposed scheme via experiments conducted over real user trace datasets. Danlan Huang, Xiaoming Tao 0001, Chunxiao Jiang, Shuguang Cui, Jianhua Lu |
IEEE Trans. Wirel. Commun. | 1 |
| 2019 | Geometry-Aware GAN for Face Attribute TransferabstractIn this paper, the geometry-aware GAN is proposed to address the issue of facial attribute transfer with unpaired data. To tackle the unpaired training sample problem, the CycleGAN architecture is applied, where the bilateral mappings between the source and target domains are learned. The deformation flow is learned to capture the geometric variation between two domains. We first warp the source face into desired pose and shape according to the flow. Then, the transfer sub-network is designed to refine the results by hallucinating new components on the warped image. The attribute is removed by the reconstruction sub-network, coupled with the warping process. Experiments on benchmark demonstrate the advantages of our method compared to baselines. Danlan Huang, Xiaoming Tao 0001, Jianhua Lu, Minh N. Do |
ICIP | 1 |
| 2017 | Latency-Efficient Video Streaming in Metropolis: A Caching FrameworkabstractThis paper presents a latency-efficient mobile video streaming design in the context of metropolis by incorporating caching. It shows that video traffic can be substantially offloaded from backhaul by caching predictable demands in the network edge. Notably, exploiting the spatial and temporal characteristics of video popularity, we focus on two sub problems: how to cache the content and how to associate users. Firstly, we investigate cache deployment strategy based on clients' viewing behavior in both downtown and suburb. The proposed hybrid collaborative filtering (CF)-based scheme guarantees high hit rate utilizing the available storage capacity in small base stations (SBS). Further, we formulate the dynamic user equipment and SBS (UE- SBS) optimal association problem into a convex optimization problem, so as to maximize the sum transmission rate of SBSs under the resource and quality-of-service constraint. Performance evaluation of real trace data demonstrates the significant advantage of our proposed framework. Danlan Huang, Xiaoming Tao 0001, Chunxiao Jiang, Yong Li 0008, Jianhua Lu |
GLOBECOM | 1 |