VLDB 2026 Research / reviewers in the wild / expert
Jianyang Shi
dblp:157/9282
· DBLP profile ↗
15ranked-venue papers
2as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InGaN-based high-speed mini laser diode surpasses PAM-4 visible light links by over 30 Gbps
Junhui Hu, Zhenqian Gu, Zengxin Li, Jiabin Wu, Leihao Sun, Aolong Sun, Ouhan Huang, Changsheng Xia, Boon S. Ooi, Jianyang Shi, Junwen Zhang 0001, Shaohua Yu, Nan Chi, Chao Shen 0009 |
Sci. China Inf. Sci. | 12 |
| 2026 | Multi-modal integration with adversarial mutual distribution matching
Ouhan Huang, Jianyang Shi, Siyuan Ye, Chao Shen 0009, Junwen Zhang 0001, Haiwen Cai, Nan Chi, Feng Bao 0002 |
Pattern Recognit. | 2 |
| 2026 | Blind Nonlinear MIMO Vector-Quantized Variational Autoencoder Equalizer for Free-Space Coherent Optical Transmission
Guojin Qin, Ziqi Tang, Junwen Zhang 0001, Guowei Jiang, Jianyang Shi, Nan Chi |
IEEE Trans. Commun. | 7 |
| 2026 | A Semantic-Empowered Free-Space Optical Communication System With Turbulence-Resilient Vector BeamsabstractHigh-altitude and low Earth orbit (LEO) platforms assisted free-space optical (FSO) communication systems have emerged as a promising solution to facilitate the development of future integrated space-air-ground networks, owing to their expansive coverage, high bandwidth, unlicensed spectrum, and large-capacity. However, these systems are highly susceptible to the adverse effects of atmospheric turbulence, degrading overall system performance and posing a risk of signal outages. In this paper, we propose a novel semantic communication paradigm, termed prior probability-based vector quantized autoencoder (PVQAE), to enable uninterrupted and efficient image transmission in free-space optical communication systems under turbulent conditions. Specifically, it utilizes a neural network-based encoder-decoder architecture for semantic information extraction and image reconstruction, and achieves uninterrupted transmission of semantic information in turbulent channels via vector beams shift keying without the need of any adaptive optics for beam compensation. A shared discrete codebook, designed for both the transmitter and receiver, facilitates latent feature representation and enables the quantization of semantic information for light-field shift keying transmission. Moreover, a power differential detection (PDD) algorithm is derived to reduce the complexity associated with receiver-side detection. Simulation results demonstrate the superior performance of PVQAE over benchmark schemes, such as conventional methods using JPEG for image source coding and Reed-Solomon (RS) for channel coding, in terms of structural similarity index (SSIM) and compression ratio (CR) product, and classification accuracy. These results highlight the effectiveness of the proposed framework in achieving robust and efficient transmission without interruption in turbulent channels. Chaoxu Chen, Fujie Li, Guowei Jiang, Feng Bao 0002, Yingjun Zhou, Chao Shen 0009, Junwen Zhang 0001, Nan Chi, Jianyang Shi |
IEEE Trans. Wirel. Commun. | 15 |
| 2025 | GiVE: Guiding Visual Encoder to Perceive Overlooked InformationabstractMultimodal Large Language Models have advanced AI in applications like text-to-video generation and visual question answering. These models rely on visual encoders to convert non-text data into vectors, but current encoders either lack semantic alignment or overlook non-salient objects. We propose the Guiding Visual Encoder to Perceive Overlooked Information (GiVE) approach. GiVE enhances visual representation with an Attention-Guided Adapter (AG-Adapter) module and an Object-focused Visual Semantic Learning module. These incorporate three novel loss terms: Object-focused Image-Text Contrast (OITC) loss, Object-focused Image-Image Contrast (OIIC) loss, and Object-focused Image Discrimination (OID) loss, improving object consideration, retrieval accuracy, and comprehensiveness. Our contributions include dynamic visual focus adjustment, novel loss functions to enhance object retrieval, and the Multi-Object Instruction (MOInst) dataset. Experiments show our approach achieves state-of-the-art performance. Jianghong Ma, Xiaofeng Zhang 0002, Jianyang Shi |
ICME | 5 |
| 2025 | SeqPose: An End-to-End Framework to Unify Single-frame and Video-based RGB Category-Level Pose EstimationabstractCategory-level object pose estimation is a longstanding and fundamental task crucial for augmented reality and robotic manipulation applications. Existing RGB-based approaches struggle with multi-stage settings and heavily rely on off-the-shelf techniques, such as object detectors, depth estimators, non-differentiable NOCS shape alignment, etc. Extra dependencies lead to the accumulation of errors and complicate the whole pipeline, limiting the deployment of these approaches in practical applications. This paper streamlined an end-to-end framework unifying the single-frame and video-based category-level pose estimation. Specifically, instead of explicitly introducing extra dependencies, the DINOv2 encoder and depth decoder, as robust semantic and geometric prior extractors, are leveraged to produce intra-frame hierarchical semantic and geometric features. A spatial-temporal sparse query network is developed to model the implicit correspondence and inter-frame correlations between a set of implicit 3D query anchors and intra-frame features. Finally, a pose prediction head is employed using the bipartite matching algorithm. Experimental results demonstrate that our model achieves state-of-the-art performance compared with RGB-based categorical pose estimation methods on the REAL275 and CAMERA25 datasets. Our code is available at https://andrewchiyz.github.io/vision.3dv.seqpose/. Yuzhu Ji, Mingshan Sun, Jianyang Shi, Xiaoke Jiang, Yiqun Zhang 0006, Haijun Zhang 0002 |
IJCAI | 3 |
| 2025 | A flexible-rate mid-long-distance visible light communication system utilizing time-domain sub-constellation probabilistic shaping modulation
Guojin Qin, Chaoxu Chen, Junwen Zhang 0001, Chao Shen 0009, Jianyang Shi, Nan Chi |
Sci. China Inf. Sci. | 6 |
| 2025 | Next-Generation Access Network Based on Coherent Optics With Hybrid Transceivers, Multi Formats, and Flexible RatesabstractCoherent detection has emerged as a key technology for advancing passive optical networks (PON) beyond 100 Gbps per wavelength, due to its advantages over the intensity and direct-detection (IM/DD) approach, which was previously dominant in PON systems. In coherent PON architectures, the cost of the optical network unit (ONU) is substantial due to the complexity of its transceivers. Consequently, numerous studies have focused on simplifying the transmitters and receivers on the ONU side. Yet, simplifying the ONU transceiver introduces issues like reduced dynamic range and lower data rates. How to make a balance between the performance and the cost, particularly in terms of component complexity, represents a critical challenge for the advancement of coherent PON. Therefore, we propose and demonstrate a hybrid, multi-format, and flexible-rate bidirectional coherent PON system, supporting a compatible OLT and ONUs with different types of transceivers. ONUs are categorized into different tiers based on their performance and cost requirements, with transceivers of varying complexity allocated accordingly. As a demonstration of concept, we have successfully conducted experimental transmissions of 25-GBaud 4/16/64- quadrature-amplitude-modulation signals across 20-km fiber in a bidirectional setup, achieving data rates ranging from 50-Gbps to 300-Gbps. The low-end, middle-end, and high-end ONUs attain power budgets of 39/31/21 dB, 43/36/26 dB, and 40/33/23 dB, respectively. This architecture serves as an effective bridge from the current 50G IM/DD PON to the anticipated 200G coherent PON, meeting the varied requirements of users at different service levels. Aolong Sun, Sizhe Xing, Guoqiang Li 0010, Wangwei Shen, Yongzhu Hu, Junhao Zhao, Ouhan Huang, Jifan Cai, Jianyang Shi, Nan Chi, Junwen Zhang 0001 |
IEEE J. Sel. Areas Commun. | 10 |
| 2025 | Research on the impact of pointing gestures based on computer vision technology on classroom concentration
Jianyang Shi, Zhangze Chen, Jia Zhu 0003 |
Neural Comput. Appl. | 1 |
| 2024 | BC-GAN: A Generative Adversarial Network for Synthesizing a Batch of Collocated ClothingabstractCollocated clothing synthesis using generative networks has become an emerging topic in the field of fashion intelligence, as it has significant potential economic value to increase revenue in the fashion industry. In previous studies, several works have attempted to synthesize visually-collocated clothing based on a given clothing item using generative adversarial networks (GANs) with promising results. These works, however, can only accomplish the synthesis of one collocated clothing item each time. Nevertheless, users may require different clothing items to meet their multiple choices due to their personal tastes and different dressing scenarios. To address this limitation, we introduce a novel batch clothing generation framework, named BC-GAN, which is able to synthesize multiple visually-collocated clothing images simultaneously. In particular, to further improve the fashion compatibility of synthetic results, BC-GAN proposes a new fashion compatibility discriminator in a contrastive learning perspective by fully exploiting the collocation relationship among all clothing items. Our model was examined in a large-scale dataset with compatible outfits constructed by ourselves. Extensive experiment results confirmed the effectiveness of our proposed BC-GAN in comparison to state-of-the-art methods in terms of diversity, visual authenticity, and fashion compatibility. Dongliang Zhou, Haijun Zhang 0002, Jianghong Ma, Jianyang Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Toward Intelligent Interactive Design: A Generation Framework Based on Cross-domain Fashion ElementsabstractTraditional fashion design typically requires the expertise of designers, which limits the involvement of ordinary users during the design process. While it would be desirable for users to participate in the preliminary design phase, their lack of basic design knowledge may render them too inexperienced to produce satisfactory designs. To improve the design efficiency for common users, we present a novel interactive fashion design framework based on generative adversarial network (GAN). This framework can assist users in designing fashion items by drawing only rough scribbles and providing simple fashion styles. Specifically, we propose a new cross-domain feature fusion encoder network that maps design image features from different domains into a series of style vectors which are then fed into a generator. We demonstrate that the learned style vectors can decouple the representations of cross-domain design elements and control the design results through scribbles and style images. Furthermore, we propose a method for rewriting our model with scribbles and style images, to allow designers to train our model more easily. To examine the effectiveness of our proposed model, we constructed a large-scale dataset containing 90,000 pairs of fashion item images. Experimental results show that our proposed method outperforms state-of-the-art methods and can effectively control cross-domain image features, suggesting the potential of our model for providing users with an intelligence-driven interactive design tool. Jianyang Shi, Haijun Zhang 0002, Dongliang Zhou, Zhao Zhang 0001 |
ACM Multimedia | 1 |
| 2023 | Texture Brush for Fashion Inspiration Transfer: A Generative Adversarial Network With Heatmap-Guided Semantic DisentanglementabstractAutomatically accomplishing intelligent fashion design with certain ‘inspiration’ images can greatly facilitate a designer’s design process, as well as allow users to interactively participate in the process. In this research, we propose a generative adversarial network with heatmap-guided semantic disentanglement (HSD-GAN) to perform an ‘intelligent’ design with ‘inspiration’ transfer. Our model aims to learn how to integrate the feature representations, from the styles of both source fashion items and target fashion items, in an unsupervised manner. Specifically, a semantic disentanglement attention-based encoder is proposed to capture the most discriminative regions of different input fashion items and disentangle the features into two key factors: attribute and texture. A generator is then developed to synthesize mixed-style fashion items by utilizing the two factors. In addition, a heatmap-based patch loss is introduced to evaluate the visual-semantic matching degree between the texture of the generated fashion items and the input texture information. Extensive experimental results show that our proposed HSD-GAN consistently achieves superior performance, compared to other state-of-the-art methods. Han Yan 0003, Haijun Zhang 0002, Jianyang Shi, Jianghong Ma |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | DETA: A Point-Based Tracker With Deformable Transformer and Task-Aligned LearningabstractCurrent point-based trackers are usually implemented by the following two branches: a classification branch for predicting the target candidate locations and a regression branch for regressing the tracking box, which may lead to a spatial misalignment between the two tasks. Meanwhile, they ignore a meaningful exploration on how to define positive and negative samples during training and explicit border information for accurate box prediction. In this research, we investigate the key issues of point-based trackers and unlock their key limitations. First, we design a novel task-aligned component and a new loss function, named task-aligned loss, to learn the alignment of the classification and regression tasks. Second, we introduce a border alignment (BorderAlign) component in both the classification and regression branches to effectively exploit the border features of a tracking target. Third, we develop an adaptive training sample assignment (ATSA) to adaptively divide the positive and negative samples based on the statistical characteristics of the tracking object. Finally, a deformable transformer is developed to enhance the representations of search features and explore rich temporal contexts among video frames. Extensive experimental results demonstrate that the proposed tracker achieves state-of-the-art performance on six tracking benchmark datasets. Kai Yang 0018, Haijun Zhang 0002, Feng Gao 0015, Jianyang Shi, Q. M. Jonathan Wu |
IEEE Trans. Multim. | 4 |
| 2023 | Toward Intelligent Fashion Design: A Texture and Shape Disentangled Generative Adversarial NetworkabstractTexture and shape in fashion, constituting essential elements of garments, characterize the body and surface of the fabric and outline the silhouette of clothing, respectively. The selection of texture and shape plays a critical role in the design process, as they largely determine the success of a new design for fashion items. In this research, we propose a texture and shape disentangled generative adversarial network (TSD-GAN) to perform “intelligent” design with the transformation of texture and shape in fashion items. Our TSD-GAN aims to learn how to disentangle the features of texture and shape of different fashion items in an unsupervised manner. Specifically, a fashion attribute encoder is developed to decompose the input fashion items into independent representations of texture and shape. Then, to learn the coarse or fine styles hidden in the features of texture and shape, a texture mapping network and a shape mapping network are proposed to disentangle the features into different hierarchical representations. The different hierarchical representations of texture and shape are then fed into a multi-factor-based generator to generate mixed-style fashion items. In addition, a multi-discriminator framework is developed to distinguish the authenticity and texture similarity between the generated images and the real images. Experimental results on different fashion categories demonstrate that our proposed TSD-GAN may be useful for assisting designers to accomplish the design process by transforming the texture and shape of fashion items. Han Yan 0003, Haijun Zhang 0002, Jianyang Shi, Jianghong Ma, Xiaofei Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Logic tensor network with massive learned knowledge for aspect-based sentiment analysis
Hu Huang 0009, Bowen Zhang 0005, Liwen Jing 0001, Xianghua Fu, Xiaojun Chen 0006, Jianyang Shi |
Knowl. Based Syst. | 6 |