Kaihui Wang

dblp:63/8494 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0001-8898-8989ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
abstract
Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on Vision-Language Action (VLA) models achieve strong results with monocular input, yet they still lag behind methods using panoramic RGB-D information. We present MonoDream, a lightweight VLA framework that enables monocular agents to learn a Unified Navigation Representation (UNR). This shared feature representation jointly aligns navigation-relevant visual semantics (e.g., global layout, depth, and future cues) and language-grounded action intent, enabling more reliable action prediction. MonoDream further introduces Latent Panoramic Dreaming (LPD) tasks to supervise the UNR, which train the model to predict latent features of panoramic RGB and depth observations at both current and future steps based on only monocular input. Experiments on multiple VLN benchmarks show that MonoDream consistently improves monocular navigation performance and significantly narrows the gap with panoramic-based agents.
Shuo Wang 0015, Yongcai Wang, Zhaoxin Fan, Maiyue Chen, Kaihui Wang, Zhizhong Su, Yeying Jin, Deying Li 0001
AAAI6
2026 Real-time integrated 2.04 cm range resolution and 16.14-Gbps bidirectional wireless communication in photonic-assisted millimeter wave band system over 100 m
Wen Zhou 0008, Jingtao Ge, Sicong Xu, Chengzhen Bian, Xiongwei Yang, Kaihui Wang, Jianjun Yu
Sci. China Inf. Sci.12
2026 Quantum Neural Networks for Symbol Recovery in Long-Haul Terahertz Communication Systems
abstract
Terahertz (THz) communication is a key enabling technology for achieving high-capacity, long-distance inter-satellite and satellite-ground communications in the next-generation wireless systems. Photonic-assisted upconversion provides a cost-effective approach to realizing ultra-wideband THz communication systems. To explore the potential of quantum neural network for the photonic-assisted THz communication systems, this work is the first to design the hybrid quantum-classical neural network for quadrature amplitude modulation (QAM) symbol recovery over tens of Gbit/s THz kilometer-level wireless transmission link, which can significantly enhance the receiver performance and reduce the computational complexity. For the proof of concept, we create the 10 Gbaud sub-THz prototype to validate the proposed scheme over 4.6km wireless long-distance. The results demonstrate that the proposed scheme achieves a 0.7dB improvement in receiver sensitivity and a one-hundred-times reduction in real-valued multiplications per symbol (RMps) compared to the classical ones.
Wen Zhou 0008, Lifeng Wang 0002, Sicong Xu, Chengzhen Bian, Xiongwei Yang, Jingtao Ge, Jingwen Lin, Zhihang Ou, Siyue Huang, Kaihui Wang, Jianjun Yu
IEEE Trans. Wirel. Commun.16
2025 Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation
abstract
Vision-Language Navigation is a critical task for developing embodied agents that can follow natural language instructions to navigate in complex real-world environments. Recent advances by finetuning large pretrained models have significantly improved generalization and instruction grounding compared to traditional approaches. However, the role of reasoning strategies in navigation—an action-centric, long-horizon task—remains underexplored, despite Chain-of-Thought reasoning's demonstrated success in static tasks like question answering and visual reasoning. To address this gap, we conduct the first systematic evaluation of reasoning strategies for VLN, including No-Think (direct action prediction), Pre-Think (reason before action), and Post-Think (reason after action). Surprisingly, our findings reveal the Inference-time Reasoning Collaps issue, where inference-time reasoning degrades navigation accuracy, highlighting the challenges of integrating reasoning into VLN. Based on this insight, we propose Aux-Think, a framework that trains models to internalize structured reasoning patterns through CoT supervision during training, while preserving No-Think inference for efficient action prediction. To support this framework, we release R2R-CoT-320k, a large-scale Chain-of-Thought annotated dataset. Empirically, Aux-Think significantly reduces training effort without compromising performance.
Shuo Wang 0015, Yongcai Wang, Maiyue Chen, Kaihui Wang, Zhizhong Su, Deying Li 0001, Zhaoxin Fan
NeurIPS7
2025 Cost-effective 200-Gbps/λ coherent PON enabled by DFB lasers and a pilot-based carrier recovery
Jianjun Yu, Jianyu Long, Bohan Sang, Wen Zhou 0008, Kaihui Wang
Sci. China Inf. Sci.7
2025 Experimental demonstration of 220-GHz terahertz signals wireless transmission over 4.6 km
Yi Wei 0005, Jianjun Yu, Xiongwei Yang, Qiutong Zhang, Jingwen Tan, Wen Zhou 0008, Kaihui Wang, Feng Zhao 0011
Sci. China Inf. Sci.12
2025 Exploiting polarization isolation for high diversity gain in a THz MIMO system
Qiutong Zhang, Jianjun Yu, Jiao Zhang 0005, Junjie Ding, Yi Wei 0005, Kaihui Wang, Wen Zhou 0008
Sci. China Inf. Sci.7
2023 Demonstration of DSM-OFDM-1024QAM transmission over 400 m at 335 GHz
Kaihui Wang, Jianjun Yu, Junjie Ding, Feng Wang 0067, Wen Zhou 0008, Jiao Zhang 0005, Tangyao Xie, Jianguo Yu, Li Zhao 0008, Feng Zhao 0011
Sci. China Inf. Sci.1
2022 A Low-latency Carrier Phase Recovery Hardware for Coherent Optical Communication
abstract
Carrier phase recovery (CPR) determines the accuracy of the receiver in modern coherent optical communication. The accurate estimation and tracking of carriers are particularly vital with the increase of throughput for long-distance transmit. It is a challenge to implement a real-time system because the computational complexity increases with fractional bits. Moreover, the conversion between polar coordinates and Cartesian coordinates introduces a high latency. In this paper, we present an FPGA implementation of low latency Viterbi-Viterbi 4thPower Estimation (VV4E) based CPR, which mainly performs the computation in Cartesian coordinates and implements the trigonometric function with a look-up table (LUT). Evaluations on Xilinx ZCU102 show that at a frequency of 370MHz, it introduces a 22-cycle latency to handle the 29.6 GBd QPSK signals, which is the minimum value to our knowledge.
Liyu Lin, Kaihui Wang, Yun Chen 0001, Jianjun Yu, Xiaoyang Zeng
ISCAS2
2022 Delivery of 335GHz OFDM Terahertz Signal over 400 meters Employing Advanced DSP Algorithms
abstract
We have experimentally demonstrated terahertz (THz) wireless transmission of 27.9 Gbit/s orthogonal frequency division multiplexing (OFDM) signal at 335 GHz over 10 km fiber and 400 m wireless link, employing advanced digital signal processing (DSP) algorithms. As we know, based on the photonics-aided scheme, THz-wireless OFDM signal transmission is achieved for the first time over record-breaking 400 m wireless distance.
Jianjun Yu, Yanyi Wang, Kaihui Wang, Wen Zhou 0008
PIMRC5
2021 Demonstration of High-Speed 4096QAM Millimeter-Wave Signal Wireless Transmission at E and D-bands
abstract
We realized the bi-directional transmission of millimeter-wave (MMW) 4096-ary quadrature amplitude modulation (4096QAM) orthogonal frequency division multiplexing (OFDM) signal with frequencies of 83.5 GHz and 73.3 GHz at E-band over 2 m wireless distance with a net transmission rate of 91.46 Gbit/s. Meanwhile, we also achieved the transmission of the MMW 4096QAM OFDM signal at 117 GHz in D-band with a net transmission rate of 57.21 Gbit/s over 13.42 m wireless distance. With the help of probabilistic shaping (PS) technique and Volterra nonlinearity compensation (VNC) technology, both of the transmission systems can satisfy the 0.8 normalized generalized mutual information (NGMI) threshold with 25% soft-decision forward-error-correction (SD-FEC) overhead. In addition, the performances of different modulation formats are also be compared.
Yuxuan Tan, Kaihui Wang, Li Zhao 0008, Junjie Ding, Jianjun Yu
VTC Fall2
2019 Spark-based real-time proactive image tracking protection model
abstract
With rapid development of the Internet, images are spreading more and more quickly and widely. The phenomenon of image illegal usage emerges frequently, and this has marked impacts on people’s normal life. Therefore, it is of great importance to protect image security and image owner’s rights. At present, most image protection is passive. Most of the time, only when the images had been used illegally and serious adverse consequences had appeared did the image owners discover it. In this paper, a Spark-based real-time proactive image tracking protection model (SRPITP) is proposed to monitor the status of images under protection in real time. Whenever illegal use is found, an alert will be issued to image owners. The model mainly includes image fingerprint extraction module, image crawling module, and image matching module. The experimental results show that in SRPITP, the image matching accuracy rate is above 98.9%, and compared with its stand-alone counterpart, the corresponding time reduction for image extraction and matching are about 58.78% and 61.67%.
Yahong Hu, Xia Sheng, Jiafa Mao, Kaihui Wang, Danhong Zhong
EURASIP J. Inf. Secur.4
2018 GrabCut algorithm for dental X-ray images based on full threshold segmentation
abstract
Teeth are difficult to be destroyed due to their corrosion resistance, high melting point and hardness. Dental biometrics can therefore provide assistance in human forensic identification, especially to the unknown corpses. One of the key issue in dental based human identification is the segmentation of Dental X‐ray images. In this paper, a novel segmentation algorithm has been proposed for this purpose. The proposed algorithm is based on full threshold segmentation. We first obtain the outline image set Iwhole n and crown image set Icrown m of the complete target tooth. Morphological open operation is then applied to the difference images of Iwhole n and Icrown m . Subsequently, the most complete target tooth image and its corresponding crown image are selected. Getting independent target tooth image I contour and its crown image I crown from these two images. Median filtering is applied to the synthetic image of I contour and I crown , and the resulted image will be used as the Mask for GrabCut to obtain the target tooth image. Experimental results show our proposed algorithm can effectively overcome the problems of uneven grayscale distribution and adhesion of adjacent crowns in dental X‐ray images. It can also achieve a high segmentation accuracy and outperform related methods to be compared.
Jiafa Mao, Kaihui Wang, Yahong Hu, Weiguo Sheng 0001, Qixin Feng
IET Image Process.2