VLDB 2026 Research / reviewers in the wild / expert
Guang Zhou
dblp:91/2785
· DBLP profile ↗
22ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Restoring neural radiance fields performance under adverse weather conditions
Ying He 0006, Gan Chen, F. Richard Yu, Ming Li 0073, Fei Ma 0006, Guang Zhou |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | VS-DSN: Variable-Speed Dual-Stream Network for Continuous Sign Language Recognition
Guang Zhou, Bo Yang 0061, Wan Tang, Jixing Yang |
CGI (2) | 1 |
| 2025 | DEP-SLAM: A Dynamic Environment Perception SLAM System with Large Language ModelsabstractInderscience is a global company, a dynamic leading independent journal publisher disseminates the latest research across the broad fields of science, engineering and technology; management, public and business administration; environment, ecological economics and sustainable development; computing, ICT and internet/web services, and related areas. Ying He 0006, F. Richard Yu, Fei Ma 0006, Ming Li 0073, Guang Zhou |
ICASSP | 5 |
| 2025 | Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object ReconstructionabstractRecent advancements in implicit 3D reconstruction methods, e.g., neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive object with n movable parts, requiring 2^n separate models to represent all discrete states. To overcome this limitation, we propose Inter3D, a new benchmark and approach for novel state synthesis of human-interactive objects. We introduce a self-collected dataset featuring commonly encountered interactive objects and a new evaluation pipeline, where only individual part states are observed during training, while part combination states remain unseen. We also propose a strong baseline approach that leverages Space Discrepancy Tensors to efficiently modelling all states of an object. To alleviate the impractical constraints on camera trajectories across training states, we propose a Mutual State Regularization mechanism to enhance the spatial density consistency of movable parts. In addition, we explore two occupancy grid sampling strategies to facilitate training efficiency. We conduct extensive experiments on the proposed benchmark, showcasing the challenges of the task and the superiority of our approach. The code and data are publicly available at https://github.com/Inter3D-ui/Inter3D. Gan Chen, Ying He 0006, Mulin Yu, F. Richard Yu, Fei Ma 0006, Ming Li 0073, Guang Zhou |
IJCAI | 8 |
| 2025 | TSceneJAL: Joint Active Learning of Traffic Scenes for 3D Object DetectionabstractMost autonomous driving (AD) datasets incur substantial costs for collection and labeling, inevitably yielding a plethora of low-quality and redundant data instances, thereby compromising performance and efficiency. Many applications in AD systems necessitate high-quality training datasets using both existing datasets and newly collected data. In this paper, we propose a traffic scene joint active learning (TSceneJAL) framework that can efficiently sample the balanced, diverse, and complex traffic scenes from both labeled and unlabeled data. The novelty of this framework is threefold: 1) a scene sampling scheme based on a category entropy, to identify scenes containing multiple object classes, thus mitigating class imbalance for the active learner; 2) a similarity sampling scheme, estimated through the directed graph representation and a marginalize kernel algorithm, to pick sparse and diverse scenes; 3) an uncertainty sampling scheme, predicted by a mixture density network, to select instances with the most unclear or complex regression outcomes for the learner. Finally, the integration of these three schemes in a joint selection strategy yields an optimal and valuable subdataset. Experiments on the KITTI, Lyft, nuScenes and SUScape datasets demonstrate that our approach outperforms existing state-of-the-art methods on 3D object detection tasks with up to 12% improvements. Chenyang Lei, Weiyuan Peng, Guang Zhou, Meiying Zhang, Qi Hao 0003, Chunlin Ji, Cheng-Zhong Xu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Temporal Fusion Network for Continuous Sign Language Recognition
Jixing Yang, Bo Yang 0061, Xincheng Hu, Guang Zhou |
CGI (1) | 4 |
| 2024 | Progressive Image Synthesis from Semantics to Details with Denoising Diffusion GANabstractAlthough denoising diffusion probabilistic models (DDPMs) have shown remarkable progress in image generation, they typically face two main challenges: the time-expensive sampling process and the semantically meaningless latent space, which are often addressed separately in previous works. In particular, the latest representative work Denoising Diffusion GAN reduces the sampling steps to as few as two but ignores the semantics of the latent space. To address the two challenges simultaneously, we propose a two-stage framework to make the latent space of Denoising Diffusion GAN more semantically meaningful while enjoying its efficiency. Extensive results on three benchmark datasets demonstrate that our proposed diffusion model achieves competitive results with only two sampling steps in unconditional image generation. More importantly, the latent space of our diffusion model trained for unconditional image generation is shown to be semantically meaningful, which can be exploited on various downstream tasks (e.g., attribute editing) without further training. Guoxing Yang, Haoyu Lu, Chongxuan Li, Guang Zhou, Zhiwu Lu 0001 |
ICASSP | 4 |
| 2024 | OTOcc: Optimal Transport for Occupancy Prediction
Pengteng Li, Ying He 0006, F. Richard Yu, Pinhao Song, Xingchen Zhou, Guang Zhou |
IJCAI | 6 |
| 2024 | SPP-SLAM: Dynamic Visual SLAM with Multiple Constraints based on Semantic Masks and Probabilistic PropagationabstractVisual simultaneous localization and mapping (vS-LAM) has attracted great attentions in mobile robots. Most vSLAM systems assume that the objects are stationary in static environments. However, in the real world, there are many objects that are non-stationary in dynamic environments, which will cause performance degradation of vSLAM systems. In this paper, we propose a novel vSLAM system suitable for dynamic environments, named as SPP-SLAM, which is based on semantic masks and probabilistic propagation. Prior motion probabilities of feature points are obtained using semantic mask constraints and multi-view geometric constraints. Then the dynamic probability of each feature point is obtained via the probability propagation model, and highly dynamic feature points are rejected. In addition, we propose a missed detection compensation module combined with inertial measurement unit (IMU) information to recover the semantic masks of the missed objects. Experimental results on the OpenLORIS-Scen and TUM RGB-D datasets demonstrate that the proposed approach can improve the performance of vSLAM systems in a variety of challenging scenarios. Run Qiu, Ying He 0006, F. Richard Yu, Guang Zhou |
WCNC | 4 |
| 2024 | Chinese Low-Latitude Atmosphere and Ionosphere Radar (CLAIR) of Meridian Project II: System Description and Initial Observation ResultsabstractWith the support of the Meridian Project II of China, Qinzhou mesosphere-stratosphere-troposphere (MST) radar, also named Chinese low-latitude atmosphere ionosphere radar (CLAIR), was built by Wuhan University in Qinzhou, Guangxi, China. The radar has been completed construction in March 2024. Two independent radar systems make up the CLAIR: one is a typical MST radar operated at 50-MHz frequency (CLAIR A) and the other is a dual-frequency radar working at 160- and 200-MHz frequencies (CLAIR B). CLAIR A has a very large quasi-circular antenna array of 155 m diameter composed of 1261 Yagi antennas with ~2-MW peak power. It is a full digital array radar and each antenna connects an independent signal channel. The hardware and software structure of CLAIR B is the same as that of CLAIR A, but the antenna array of CLAIR B is smaller with 32.7 m diameter and composed of 931 log-periodic antennas with ~0.58-MW peak power. This radar has two operating modes for atmosphere observation, including the ST mode for stratosphere and troposphere observation and the M mode for mesosphere observation. The observation results of the stratospheric and tropospheric wind field by the CLAIR with the three operating frequencies and the rawinsonde present good consistency, demonstrating the reliability of the CLAIR observations. The peak power of 2 MW enables CLAIR A to well observe wind field variations at an altitude of 60–110-km altitude. Moreover, the observation result of ionospheric field-aligned irregularities in the E-region is also displayed. Shao-Dong Zhang, Gang Chen 0026, Wanlin Gong, Xiao-Ming Zhou, Jinpeng Tao, Yaxian Li, Guang Zhou |
IEEE Trans. Geosci. Remote. Sens. | 12 |
| 2023 | IGG: Improved Graph Generation for Domain Adaptive Object DetectionabstractDomain Adaptive Object Detection (DAOD) transfers an object detector from a labeled source domain to a novel unlabeled target domain. Recent works bridge the domain gap by aligning cross-domain pixel-pairs in the non-euclidean graphical space and minimizing the domain discrepancy for adapting semantic distribution. Though great successes, these methods model graphs roughly with coarse semantic sampling due to ignoring the non-informative noises and failing to concentrate on precise semantics alignment. Besides, the coarse graph generation inevitably contains abnormal nodes. These challenges result in biased domain adaptation. Therefore, we propose an Improved Graph Generation (IGG) framework which conducts high-quality graph generation for DAOD. Specifically, we design an Intensive Node Refinement (INR) module that reconstructs the noisy sampled nodes with a memory bank, and contrastively regularizes the noisy features. For better semantics alignment, we decouple the domain-specific style and category-invariant content encoded in graph covariance and selectively eliminate only the domain-specific style. Then, a Precision Graph Optimization (PGO) adaptor is proposed which utilizes the variational inference to down-weight abnormal nodes. Comprehensive experiments on three adaptation benchmarks demonstrate that IGG achieves state-of-the-art results in unsupervised domain adaptation. Pengteng Li, Ying He 0006, F. Richard Yu, Pinhao Song, Dongfu Yin, Guang Zhou |
ACM Multimedia | 6 |
| 2023 | A Novel Visual SLAM System for Autonomous Vehicles in Dynamic EnvironmentsabstractWith the development of autonomous vehicles and intelligent robots, visual simultaneous localization and mapping (SLAM) has attracted great attentions. Most existing visual SLAM systems assume that the objects are stationary in static environments. However, in the real world, there are many objects that are non-stationary in dynamic environments, which will cause performance degradation of visual SLAM systems. In this paper, to address this issue, we propose a novel visual SLAM system based on multi-task deep neural networks. Specifically, we apply multi-task deep neural networks to extract oriented keypoints and perceive dynamic semantic regions, which are used to perform outlier rejection in the SLAM system. We evaluate our method on public datasets, and the results show that our method outperforms existing visual SLAM systems. The presentation video url is: https://youtu.be/qGE1OvaJvV0. Xinyu Zeng, Ying He 0006, F. Richard Yu, Guang Zhou |
VTC Fall | 4 |
| 2021 | IC solder joint inspection via generator-adversarial-network based template
Nian Cai, Zhuokun Mo, Guang Zhou, Han Wang 0017 |
Mach. Vis. Appl. | 4 |
| 2020 | Pop Music Generation: From Melody to Multi-style ArrangementabstractMusic plays an important role in our daily life. With the development of deep learning and modern generation techniques, researchers have done plenty of works on automatic music generation. However, due to the special requirements of both melody and arrangement, most of these methods have limitations when applying to multi-track music generation. Some critical factors related to the quality of music are not well addressed, such as chord progression, rhythm pattern, and musical style. In order to tackle the problems and ensure the harmony of multi-track music, in this article, we propose an end-to-end melody and arrangement generation framework to generate a melody track with several accompany tracks played by some different instruments. To be specific, we first develop a novel Chord based Rhythm and Melody Cross-Generation Model to generate melody with a chord progression. Then, we propose a Multi-Instrument Co-Arrangement Model based on multi-task learning for multi-track music arrangement. Furthermore, to control the musical style of arrangement, we design a Multi-Style Multi-Instrument Co-Arrangement Model to learn the musical style with adversarial training. Therefore, we can not only maintain the harmony of the generated music but also control the musical style for better utilization. Extensive experiments on a real-world dataset demonstrate the superiority and effectiveness of our proposed models. Hongyuan Zhu 0001, Qi Liu 0003, Nicholas Jing Yuan, Kun Zhang 0015, Guang Zhou, Enhong Chen |
ACM Trans. Knowl. Discov. Data | 5 |
| 2018 | XiaoIce Band: A Melody and Arrangement Generation Framework for Pop MusicabstractWith the development of knowledge of music composition and the recent increase in demand, an increasing number of companies and research institutes have begun to study the automatic generation of music. However, previous models have limitations when applying to song generation, which requires both the melody and arrangement. Besides, many critical factors related to the quality of a song such as chord progression and rhythm patterns are not well addressed. In particular, the problem of how to ensure the harmony of multi-track music is still underexplored. To this end, we present a focused study on pop music generation, in which we take both chord and rhythm influence of melody generation and the harmony of music arrangement into consideration. We propose an end-to-end melody and arrangement generation framework, called XiaoIce Band, which generates a melody track with several accompany tracks played by several types of instruments. Specifically, we devise a Chord based Rhythm and Melody Cross-Generation Model (CRMCG) to generate melody with chord progressions. Then, we propose a Multi-Instrument Co-Arrangement Model (MICA) using multi-task learning for multi-track music arrangement. Finally, we conduct extensive experiments on a real-world dataset, where the results demonstrate the effectiveness of XiaoIce Band. Hongyuan Zhu 0001, Qi Liu 0003, Nicholas Jing Yuan, Chuan Qin 0002, Kun Zhang 0015, Guang Zhou, Furu Wei, Yuanchun Xu, Enhong Chen |
KDD | 7 |
| 2016 | Robust real-time UAV based power line detection and trackingabstractPower line inspection is an essential but costly task while automated UAV (unmanned aerial vehicle) inspection can greatly reduce such costs. However, navigating along power lines is a challenging task due to the narrow width and limited features of power lines. Existing power line tracking methods have threshold selection problems and cannot work well for complex and changing backgrounds. We make use of power line specific knowledge to build a model to achieve predictive and continuous parameter selection so that the best thresholds are selected for changing scenarios. Experimental studies show that our solution is much better than existing ones. Specifically, we have built the first fully automated UAV that successfully tracks power lines in the real world. Guang Zhou, Jinwei Yuan, I-Ling Yen, Farokh B. Bastani |
ICIP | 1 |
| 2016 | A Self-Stabilizing Algorithm for the Foraging Problem in Swarm Robotic SystemsabstractThe foraging problem, which evolved from ants finding food and delivering them to the nest, is for a swarm of robots to transport objects from a source to a destination. To avoid collision of robots or congestions, object transportation is generally conducted in a pipelining manner. Existing solutions towards the foraging problem either have performance problems or cannot tolerate failures. We present a self-stabilizing solution for the swarm robotic system to form a pipeline structure to transport objects and achieve the foraging task. In this paper, we first define the stable state for the swarm, which includes two requirements: (1) the robots operate in non-overlapping regions, i.e., transport objects in a pipeline structure and (2) the transportation rate of the system is optimal. Then, we introduce our self-stabilizing algorithm for the foraging problem and prove its convergence and the correctness of its convergence properties. Due to the self-stabilization nature, our solution is decentralized and fault tolerant. The swarm can achieve the foraging task as long as there is at least one working robot. Due to the definition of the stable state, our swarm system, when converged, can achieve optimal performance. We conduct simulations to evaluate the effectiveness of our algorithm, and the experimental results show that from any state, our algorithm can converge very quickly to reach the stable state and provide optimal performance for object transportation. Guang Zhou, Farokh B. Bastani, Wei Zhu 0002, I-Ling Yen |
IROS | 1 |
| 2015 | A PT-SOA Model for CPS/IoT ServicesabstractService computing technologies have been widely applied to many application domains to facilitate rapid system integration for desired goals. However, existing service models need to be enhanced for cyber-physical systems (CPS) and internet-of-things (IoT). In this paper, we develop an ontology model for the specification of services in CPS/IoT. First, we discuss the major differences in modeling software services and services in CPS/IoT. Then, we propose a novel PT-SOA (PT stands for physical things) model, which is mainly extended from OWL-S, to enhance existing service models for CPS/IoT systems. Finally, we show a case study system to illustrate how our model can facilitate proper service selection and composition. Wei Zhu 0002, Guang Zhou, I-Ling Yen, Farokh B. Bastani |
ICWS | 2 |
| 2015 | A Smart Physical World Based on Service Technologies, Big Data, and Game-Based Crowd SourcingabstractThe physical world is becoming smarter and smarter due to the advances in smart devices and CPS/IoT technologies. In this paper, we investigate the roles various cutting-edge technologies, such as service computing, big data analytics, crowd sourcing, gaming technologies, etc., can play to significantly enhance the intelligence of our physical environment and subsequently benefit the society. We consider a smart physical world (SPW) consisting of physical entities, cyber entities, and human. Service technologies can be used to model the entities in SPW and specify their capabilities. Service discovery and composition techniques can be used to, based on the modeling, compose these capabilities to solve real-world problems. To enable higher level intelligence, we further discuss how various technologies, such as big data analytics and artificial intelligence, can be used in the smart world to reason from sensor inputs to derive the situation facts, and from the situation facts to derive the reactive actions. From the derived actions (tasks), the service technologies can be used to compose the capabilities of the entities to realize the task. However, current technologies and artificial intelligence may fall short in many situations. We further propose a gaming-based crowd sourcing platform to make use of human intelligence to enable successful completion of some challenging reasoning and control tasks. I-Ling Yen, Guang Zhou, Wei Zhu 0002, Farokh B. Bastani, San-Yih Hwang |
ICWS | 2 |
| 2013 | Toward Ontology and Service Paradigm for Enhanced Carbon Footprint Management and LabelingabstractAs the green house gas emission becomes a serious problem, a lot of researches now focus on how to monitor and manage carbon footprint (CF) of a production process or transportation, especially in the supply chain. Usually, most of the carbon footprint management systems are based on databases. But database is not sufficient in describing the production and transportation processes and the facilities used in these processes. In this paper, we develop an SOA based model for the carbon footprint management and labeling (CFML), using ontology and OWL-S techniques. We use OWL-S to describe the processes and workflows for production and transportation and extend it to specify the methods for deriving CF of them. We use the existing energy conversion formula to derive the CF when the energy data can be collected separately. We also derive an approach to separate the CFs when data of different processes have to be collected together. In a supply chain, a production company may have different choices of suppliers to provide certain components. To balance the tradeoff between carbon dioxide emission and cost of the overall production process, we design a supplier selection algorithm to derive the optimal solution. Wei Zhu 0002, Guang Zhou, I-Ling Yen, San-Yih Hwang |
ICWS | 2 |
| 2012 | Service-Oriented Robotic Swarm Systems: Model and Structuring AlgorithmsabstractIn this paper, we consider the design issues in building robotic swarm systems as a service and develop a multi-level service-oriented robotic swarm (SORS) framework. First we consider how to support easy composition of services into workflows to accomplish user tasks. We identify a set of primitive virtual services for this purpose. Virtual services are not physical robotic swarm actions, but are high-level primitive services geared toward easy specification of the workflows of many tasks. To realize virtual services, SORS incorporates the layers of planning services and micro services to plan and control the swarm of robots. Two example micro services and the detailed algorithms in their implementation are given in the paper. Also, a case study system is developed to show the use of virtual services to specify the workflow for an example task and to illustrate the design of micro services to realize virtual services. Guang Zhou, Yansheng Zhang, Farokh B. Bastani, I-Ling Yen |
ISORC | 1 |
| 2005 | SNR-Based Frame-Level Video Bit Rate AllocationabstractQuality fluctuation has a major negative effect on perceptive video quality. In [1], we derived accurate approximations in close-form for the highly nonlinear rate-distortion (R-D) and distortion-quantization (D-Q) relationships, all at the frame-level. Based on the two close forms, we can allocate the bit rate at the frame-level rather easily as far as a target distortion for each frame could be established. In [1], a target distortion was set up for each frame based on a hypothesis that maintaining constant distortion over frames would boast video quality smoothing and extensive experiments showed the constant-distortion bit allocation (CDBA) scheme significantly outperformed the popular constant bit allocation (CBA) scheme in terms of delivered video quality. Maintaining constant distortion is no different from maintaining constant Peak-signal-to-Noise-Ratio (PSNR). In scene changes, however, the picture energy often dramatically changes, producing significantly different Signal-to-Noise-Ratio (SNR) if constant distortion or constant PSNR is maintained. Although computationally more complex, SNR represents a more objective measure than PSNR in assessing picture/video quality. In the paper, an SNR-based bit allocation scheme is developed for video quality smoothing. The algorithm uses a single pass and attempts to maintain constant SNR at the frame level throughout the video sequence. Experimental results on all testing video sequences show that the proposed CSNRBA scheme provides smooth video quality in terms of natural color and sharp objects and silhouette significantly better than both the CBA and CDBA schemes. Xinhua Zhuang, Xiangui Kang, Li Liu 0009, Junqiang Lan, Guang Zhou, Guangdong Wu |
MMSP | 5 |