Jiawei Yao

dblp:203/3900 · DBLP profile ↗
← Back
22ranked-venue papers
13as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 SpikeGS: Reconstruct 3D Scene Captured by a Fast-Moving Bio-Inspired Camera
abstract
3D Gaussian Splatting (3DGS) has been proven to exhibit exceptional performance in reconstructing 3D scenes. However, the effectiveness of 3DGS heavily relies on sharp images, and fulfilling this requirement presents challenges in real-world scenarios particularly when utilizing fast-moving cameras. This limitation severely constrains the practical application of 3DGS and may compromise the feasibility of real-time reconstruction. To mitigate these challenges, we proposed Spike Gaussian Splatting (SpikeGS), the first framework that integrates the Bayer-pattern spike streams into the 3DGS pipeline to reconstruct 3D scenes captured by a fast-moving high temporal color spike camera in one second. With accumulation rasterization, interval supervision, and a special designed pipeline, SpikeGS realizes continuous spatiotemporal perception while extracts detailed structure and texture from Bayer-pattern spike stream which is unstable and lacks details. Extensive experiments on both synthetic and real-world datasets demonstrate the superiority of SpikeGS compared with existing spike-based and deblur 3D scene reconstruction methods.
Yijia Guo, Liwen Hu 0002, Yuanxi Bai, Jiawei Yao, Lei Ma 0008, Tiejun Huang 0001
AAAI4
2025 Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation
abstract
Multimodal large language models have experienced rapid growth, and numerous different models have emerged. The interpretability of LVLMs remains an under-explored area. Especially when faced with more complex tasks such as chain-of-thought reasoning, its internal mechanisms still resemble a black box that is difficult to decipher. By studying the interaction and information flow between images and text, we noticed that in models such as LLaVA1.5, image tokens that are semantically related to text are more likely to have information flow convergence in the LLM decoding layer, and these image tokens receive higher attention scores. However, those image tokens that are less relevant to the text do not have information flow convergence, and they only get very small attention scores. To efficiently utilize the image information, we propose a new image token reduction method, Simignore, which aims to improve the complex reasoning ability of LVLMs by computing the similarity between image and text embeddings and ignoring image tokens that are irrelevant and unimportant to the text. Through extensive experiments, we demonstrate the effectiveness of our method for complex reasoning tasks.
Fanshuo Zeng, Yihao Quan, Zheng Hui, Jiawei Yao
AAAI5
2025 LP-Gaussians: Learnable Parametric Gaussian Splatting for Efficient Dynamic Reconstruction of Single-View Scenes
abstract
With the popularity of short video platforms, the number of single-view videos has increased significantly. Existing NeRF-based methods can reconstruct dynamic scenes in a single-view setting, but slow rendering speed and low rendering quality limit their practical applications. To address these challenges, we propose a fast single-view scenes reconstruction framework based on 3D Gaussian Splatting. Our method uses point clouds obtained with depth priors as the Gaussian initialization and introduces learnable parametric functions to model the time-dependent deformation of Gaussians. The explicit deformation modeling for Gaussians significantly reduces training and rendering time. Furthermore, to improve the rendering quality of challenging areas, we adopt an adaptive sampling strategy to densify Gaussians. For occlusion problems from single-view videos, we design a smooth loss function to restore the color of the occluded areas. Experimental results demonstrate that our method significantly reduces training time, enhances rendering quality, and accelerates rendering speed. Project page: https://github.com/LPGaussians.
Shaoqi Wu, Weixing Xie, Youhong Peng, Jiawei Yao, Junfeng Yao
ICASSP5
2025 KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems
abstract
As scaling large language models faces prohibitive costs, multi-agent systems emerge as a promising alternative, though challenged by static knowledge assumptions and coordination inefficiencies. We introduce Knowledge-Aware Bayesian Bandits (KABB), a novel framework that enhances multi-agent system coordination through semantic understanding and dynamic adaptation. The framework features three key innovations: a customized knowledge distance model for deep semantic understanding, a dual-adaptation mechanism for continuous expert optimization, and a knowledge-aware Thompson Sampling strategy for efficient expert selection. Extensive evaluation demonstrates KABB achieves an optimal cost-performance balance, maintaining high performance while keeping computational demands relatively low in multi-agent coordination.
Jusheng Zhang, Zimeng Huang, Yijia Fan, Ningyuan Liu, Zhuojie Yang, Jiawei Yao, Jian Wang 0100, Keze Wang
ICML7
2025 DepthSSC: Monocular 3D Semantic Scene Completion via Depth-Spatial Alignment and Voxel Adaptation
abstract
The task of 3D semantic scene completion using monocular cameras is gaining significant attention in the field of autonomous driving. This task aims to predict the occupancy status and semantic labels of each voxel in a 3D scene from partial image inputs. Despite numerous existing methods, many face challenges such as inaccurately predicting object shapes and misclassifying object boundaries. To address these issues, we propose DepthSSC, an advanced method for semantic scene completion using only monocular cameras. DepthSSC integrates the Spatial Transformation Graph Fusion (ST-GF) module with Geometric-Aware Voxelization (GAV), enabling dynamic adjustment of voxel resolution to accommodate the geometric complexity of 3D space. This ensures precise alignment between spatial and depth information, effectively mitigating issues such as object boundary distortion and incorrect depth perception found in previous methods. Evaluations on the SemanticKITTI and SSCBench-KITTI-360 dataset demonstrate that DepthSSC not only captures intricate 3D structural details effectively but also achieves state-of-the- art performance.
Jiawei Yao, Jusheng Zhang, Xiaochao Pan, Canran Xiao
WACV1
2024 Text-Guided Mixup Towards Long-Tailed Image Categorization
Richard Franklin, Jiawei Yao, Deyang Zhong, Qi Qian 0001, Juhua Hu
BMVC2
2024 Improving Depth Gradient Continuity in Transformers: A Comparative Study on Monocular Depth Estimation with CNN
Jiawei Yao
BMVC1
2024 Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering
abstract
Multiple clustering has gained significant attention in recent years due to its potential to reveal multiple hidden structures of data from different perspectives. The advent of deep multiple clustering techniques has notably advanced the performance by uncovering complex patterns and relationships within large datasets. However, a major challenge arises as users often do not need all the clusterings that algorithms generate, and figuring out the one needed requires a substantial understanding of each clustering result. Traditionally, aligning a user's brief keyword of interest with the corresponding vision components was challenging, but the emergence of multi-modal and large language models (LLMs) has begun to bridge this gap. In response, given unlabeled target visual data, we propose Multi-MaP, a novel method employing a multi-modal proxy learning process. It leverages CLIP encoders to extract coherent text and image embeddings, with GPT-4 integrating users' interests to formulate effective textual contexts. Moreover, reference word constraint and concept-level constraint are designed to learn the optimal text proxy according to the user's interest. Multi-MaP not only adeptly captures a user's interest via a keyword but also facilitates identifying relevant clusterings. Our extensive experiments show that Multi-MaP consistently outperforms state-of-the-art methods in all benchmark multi-clustering vision tasks. Our code is available at https://github.com/Alexander-Yao/Multi-MaP.
Jiawei Yao, Qi Qian 0001, Juhua Hu
CVPR1
2024 Building Lane-Level Maps from Aerial Images
abstract
Detecting lane lines from sensors is becoming an increasingly significant part of autonomous driving systems. However, less development has been made on high-definition lane-level mapping based on aerial images, which could automatically build and update offline maps for auto-driving systems. To this end, our work focuses on extracting fine-level detailed lane lines together with their topological structures. This task is challenging since it requires large amounts of data covering different lane types, terrain and regions. In this paper, we introduce for the first time a large-scale aerial image dataset built for lane detection, with high-quality polyline lane annotations on high-resolution images of around 80 kilometers of road. Moreover, we developed a baseline deep learning lane detection method from aerial images, called AerialLaneNet, consisting of two stages. The first stage is to produce coarse-grained results at point level, and the second stage exploits the coarse-grained results and feature to perform the vertex-matching task, producing fine-grained lanes with topology. The experiments show our approach achieves significant improvement compared with the state-of-the-art methods on our new dataset. Our code and new dataset are available at https://github.com/Jiawei-Yao0812/AerialLaneNet.
Jiawei Yao, Xiaochao Pan
ICASSP1
2024 Hierarchical Adaptive Position Encoding-Based Transformer for Point Cloud Analysis
Jiawei Yao, Junfeng Yao, Chunyang Huang
ICONIP (1)1
2024 HarmonicNeRF: Geometry-Informed Synthetic View Augmentation for 3D Scene Reconstruction in Driving Scenarios
Xiaochao Pan, Jiawei Yao, Hongrui Kou, Canran Xiao
ACM Multimedia2
2024 QE-BEV: Query Evolution for Bird's Eye View Object Detection in Varied Contexts
Jiawei Yao, Yingxin Lai, Hongrui Kou, Ruixi Liu
ACM Multimedia1
2024 Customized Multiple Clustering via Multi-Modal Subspace Proxy Learning
abstract
Multiple clustering aims to discover various latent structures of data from different aspects. Deep multiple clustering methods have achieved remarkable performance by exploiting complex patterns and relationships in data. However, existing works struggle to flexibly adapt to diverse user-specific needs in data grouping, which may require manual understanding of each clustering. To address these limitations, we introduce Multi-Sub, a novel end-to-end multiple clustering approach that incorporates a multi-modal subspace proxy learning framework in this work. Utilizing the synergistic capabilities of CLIP and GPT-4, Multi-Sub aligns textual prompts expressing user preferences with their corresponding visual representations. This is achieved by automatically generating proxy words from large language models that act as subspace bases, thus allowing for the customized representation of data in terms specific to the user’s interests. Our method consistently outperforms existing baselines across a broad set of datasets in visual multiple clustering tasks. Our code is available at https://github.com/Alexander-Yao/Multi-Sub.
Jiawei Yao, Qi Qian 0001, Juhua Hu
NeurIPS1
2024 Swift Sampler: Efficient Learning of Sampler by 10 Parameters
abstract
Data selection is essential for training deep learning models. An effective data sampler assigns proper sampling probability for training data and helps the model converge to a good local minimum with high performance. Previous studies in data sampling are mainly based on heuristic rules or learning through a huge amount of time-consuming trials. In this paper, we propose an automatic swift sampler search algorithm, SS, to explore automatically learning effective samplers efficiently. In particular, SS utilizes a novel formulation to map a sampler to a low dimension of hyper-parameters and uses an approximated local minimum to quickly examine the quality of a sampler. Benefiting from its low computational expense, SS can be applied on large-scale data sets with high efficiency. Comprehensive experiments on various tasks demonstrate that SS powered sampling can achieve obvious improvements (e.g., 1.5% on ImageNet) and transfer among different neural networks. Project page: https://github.com/Alexander-Yao/Swift-Sampler.
Jiawei Yao, Chuming Li, Canran Xiao
NeurIPS1
2024 Dual-disentangled Deep Multiple Clustering
abstract
Multiple clustering has gathered significant attention in recent years due to its potential to reveal multiple hidden structures of the data from different perspectives. Most of multiple clustering methods first derive feature representations by controlling the dissimilarity among them, subsequently employing traditional clustering methods (e.g., k-means) to achieve the final multiple clustering outcomes. However, the learned feature representations can exhibit a weak relevance to the ultimate goal of distinct clustering. Moreover, these features are often not explicitly learned for the purpose of clustering. Therefore, in this paper, we propose a novel Dual-Disentangled deep Multiple Clustering method named DDMC by learning disentangled representations. Specifically, DDMC is achieved by a variational Expectation-Maximization (EM) framework. In the E-step, the disentanglement learning module employs coarse-grained and fine-grained disentangled representations to obtain a more diverse set of latent factors from the data. In the M-step, the cluster assignment module utilizes a cluster objective function to augment the effectiveness of the cluster output. Our extensive experiments demonstrate that DDMC consistently outperforms state-of-the-art methods across seven commonly used tasks. Our code is available at https://github.com/Alexander-Yao/DDMC.
Jiawei Yao, Juhua Hu
SDM1
2024 Blind Beamforming for Coverage Enhancement With Intelligent Reflecting Surface
abstract
Conventional policy for configuring an intelligent reflecting surface (IRS) typically requires channel state information (CSI), thus incurring substantial overhead costs and facing incompatibility with the current network protocols. This paper proposes a blind beamforming strategy in the absence of CSI, aiming to boost the minimum signal-to-noise ratio (SNR) among all the receiver positions, namely the coverage enhancement. Although some existing works already consider the IRS-assisted coverage enhancement without CSI, they assume certain position-channel models through which the channels can be recovered from the geographic locations. In contrast, our approach solely relies on the received signal power data, not assuming any position-channel model. We examine the achievability and converse of the proposed blind beamforming method. If the IRS has N reflective elements and there are U receiver positions, then our method guarantees the minimum SNR of$\Omega (N^{2}/U)$—which is fairly close to the upper bound$O(N+N^{2}\sqrt {\ln (NU)}/\sqrt [{4}]{U})$. Aside from the simulation results, we justify the practical use of blind beamforming in a field test at 2.6 GHz. According to the real-world experiment, the proposed blind beamforming method boosts the minimum SNR across seven random positions in a conference room by 18.22 dB, while the position-based method yields a boost of 12.08 dB.
Fan Xu 0001, Jiawei Yao, Wenhai Lai, Kaiming Shen, Xin Li 0112, Xin Chen 0062, Zhi-Quan Luo
IEEE Trans. Wirel. Commun.2
2023 Blind Beamforming for Multiple Intelligent Reflecting Surfaces
abstract
Channel acquisition is a major challenge faced by the conventional beamforming methods when dealing with multiple intelligent reflecting surfaces (IRSs), because the number of unknown channels grows exponentially with the number of IRSs. This work proposes to sidestep channel estimation and to configure the IRSs blindly based on the statistical information which is extracted from a set of random samples of the received signal power. The proposed blind beamforming method has provable performance in terms of the signal-to-noise ratio (SNR) boost. For instance, it yields a quartic SNR boost of$\Theta(N^{4})$for a double-IRS system under certain condition, where$N$is the number of reflected elements of each IRS. We remark that the above$\Theta(N^{4})$result is more sophisticated than the existing ones about the double-IRS system in the literature. Furthermore, we numerically demonstrate the advantage of the proposed blind beamforming method through prototype tests with multiple IRSs.
Jiawei Yao, Fan Xu 0001, Wenhai Lai, Kaiming Shen, Xin Li 0112, Xin Chen 0062, Zhi-Quan Luo
ICC1
2023 NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates Space
abstract
Monocular 3D Semantic Scene Completion (SSC) has garnered significant attention in recent years due to its potential to predict complex semantics and geometry shapes from a single image, requiring no 3D inputs. In this paper, we identify several critical issues in current state-of-the-art methods, including the Feature Ambiguity of projected 2D features in the ray to the 3D space, the Pose Ambiguity of the 3D convolution, and the Computation Imbalance in the 3D convolution across different depth levels. To address these problems, we devise a novel Normalized Device Coordinates scene completion network (NDC-Scene) that directly extends the 2D feature map to a Normalized Device Coordinates (NDC) space, rather than to the world space directly, through progressive restoration of the dimension of depth with deconvolution operations. Experiment results demonstrate that transferring the majority of computation from the target 3D space to the proposed normalized device coordinates space benefits monocular SSC tasks. Additionally, we design a Depth-Adaptive Dual Decoder to simultaneously upsample and fuse the 2D and 3D feature maps, further improving overall performance. Our extensive experiments confirm that the proposed method consistently outperforms state-of-the-art methods on both outdoor SemanticKITTI and indoor NYUv2 datasets. Our code are available at https://github.com/Jiawei-Yao0812/NDCScene.
Jiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai, Hao Li 0069, Wanli Ouyang, Hongsheng Li 0001
ICCV1
2022 Maps Vision: A Computer Vision-based System for Detecting Discrepancies in Map Textual Labels
abstract
We demonstrate MapsVision, a computer vision-based framework capable of identifying discrepancies across different map providers for similar geographical locations. In this study, we primarily focus on three map providers including: (a) Bing Maps, (b) Google Maps, and (c) OpenStreetMap. MapsVision detects textual data discrepancies such as: (1) missing location labels (2) misspelled or different keywords, (3) shifted labels, and (4) level of significance manifested by text or label font-size and color. For a given location, our MapsVision framework compares textual labels based on a ground truth entered manually to those that exist in the three map providers. We then use the results of the textual extraction to determine the accuracy of textual data appearing on map providers. Our framework intelligently identifies the set of techniques for each map providers' that can maximize the overall detection accuracy. MapsVision is composed of three main building blocks including: (a) a capturing module that captures map tiles from map providers, (b) an analysis tool that uses computer vision and text-analytic techniques, and (c) a rich visualization interface for displaying statistical and real-time analytics. The objective of MapsVision is to help map editors improve the textual quality of their maps compared to other map providers.
Adel A. Sabour, Jiawei Yao, Abdulrahman Salama, Cordel Hampshire, Eyhab Al-Masri, Mohamed Ali 0002, Harsh Govind, Vashutosh Agrawal, Egor Maresov, Ravi Prakash 0007
MDM2
2022 A Geospatial Method for Detecting Map-Based Road Segment Discrepancies
abstract
Today, people's lives are enriched by the integration of electronic maps via smartphones. Electronic maps are required for a variety of commercial activities, such as catering, movie viewing, and tourism. Route planning and navigation are particularly intrinsically linked to electronic maps. As a result, it is critical that the roads on the electronic map are complete and accurate. At the present time, there are discrepancies between the map roads of various providers. This paper evaluates the roads on various map providers' maps. Due to the varied terrain depicted on the map, assessing the road properties can be challenging. Additionally, roads of varying thicknesses exist within a tile image, making it difficult to quantify the map's road lengths. This paper proposes a method for extracting road segments using an image binarization technique and employs edge erosion to assist in automatically computing the length of roads within maps. Throughout the paper, we provide comparison and statistical analysis on using our proposed road length detection model across map providers. Results show that our detection model can identify road length accurately and hence provide an overall measure of quality of maps.
Jiawei Yao, Eyhab Al-Masri, Mohamed Ali 0002, Vashutosh Agrawal, Harsh Govind, Adel A. Sabour, Abdulrahman Salama, Reuben Keller, Dino Jazvin, Ravi Prakash 0007, Egor Maresov
MDM1
2022 Context-Aware Feature Learning for Noise Robust Person Search
abstract
Person search aims to localize and identify specific pedestrians from numerous surveillance scene images. In this work, we focus on the noise in person search. We categorize the noise into scene-inherent noise and human-introduced noise. Scene-inherent noise comes from congestion, occlusion, and illumination changes. Human-introduced noise originates from the labeling process. For scene-inherent noise, we propose a novel context contrastive loss to take advantage of the latent contextual information from scene images. Features from context regions are utilized to construct contrastive pairs to constrain the feature discrimination among pedestrians in scene images while maintaining the feature consistency of the same identity. The network can thus learn to distinguish congested and overlapped pedestrians and more robust features can be obtained. For human-introduced noise, we propose a noise-discovery and noise-suppression training process for mislabeling robust person search. After the first training pass, the relation between feature prototypes of different identities is analyzed and the mislabeled pedestrians are discovered. During the second training pass, the label noise is suppressed to reduce the negative influence of mislabeled data. Experiments show that the proposed context-aware noise-robust (CANR) person search can achieve competitive performance. Further ablation studies confirm the effectiveness of CANR.
Cairong Zhao, Shuguang Dou, Zefan Qu, Jiawei Yao, Jun Wu 0006, Duoqian Miao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2012 Product Recommendation Based on Search Keywords
abstract
Recommender systems have been widely deployed on E-commerce websites. The cold start problem of making effective recommendations to new users without any historical data on the website is still challenging. These new users often have some available information, such as search keywords, before visiting the website. It is natural to use the information to predict users' preference, such that an immediate recommendation is possible. In this paper, we propose a new product recommendation approach for new users based on the implicit relationships between search keywords and products. The relationships between keywords and products are represented in a graph and relevance of keywords to products is derived from attributes of the graph. The relevance information will be utilized to predict preferences of new users. A preliminary experiment is conducted and shows that our approach outperforms the traditional approach (Recommending Most Popular Products).
Jiawei Yao, Jiajun Yao, Zhenyu Chen 0001
WISA1