VLDB 2026 Research / reviewers in the wild / expert
Qiong Chang
dblp:219/6333
· DBLP profile ↗
34ranked-venue papers
8as first author
31since 2021 · last 2026
0000-0002-4447-0480ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GPU-Accelerated Dependency Graph Construction and Conflict Analysis for Preventing Read-Only Anomalies
Reo Chiyomaru, Takamitsu Shioi, Qiong Chang, Jun Miyazaki |
DaWaK | 3 |
| 2025 | Unified Schema-Driven Graph Polystore: Achieving Transparency in Multi-model Integration and Migration
Fumihiro Yamashita, Qiong Chang, Jun Miyazaki |
DEXA (2) | 2 |
| 2025 | FSAC-IA: A Hierarchical Constructed SAC-IA Algorithm for Point Cloud Alignment AccelerationabstractPoint cloud alignment plays a crucial role in numerous fields and applications, enabling the integration, analysis, and reconstruction of 3D data. SAmple Consensus Initial Alignment (SAC-IA), a coarse alignment algorithm, provides the initial position for precise alignment. While SAC-IA can effectively capture both global and local information, it incurs high computational costs associated with the correspondence matching of high-dimensional features. To address this challenge, we propose FSAC-IA, a faster SAC-IA-based algorithm designed to accelerate LiDAR point cloud alignment. FSAC-IA introduces a light Fast Point Feature Histogram (FPFH) descriptor to optimize the feature presentation, ensuring high matching speed without sacrificing accuracy. We also replace the original k-d tree structure in the SAC-IA, with a Hierarchical Navigable Small World (HNSW) structure for descriptor-matching, enabling efficient large-scale point cloud alignment. Experiments on fine-grained LiDAR point clouds demonstrate that FSAC-IA achieves a 7x speedup over traditional k-d tree-based alignment algorithms, establishing a new benchmark for efficient and accurate point cloud alignment. Ziyang Yu 0004, Qiong Chang, Jun Miyazaki |
ICIP | 2 |
| 2025 | Fast Approximate Aggregation with Error Guarantee Using Encoded Bit-Slice Indexing
Kakeru Ito, Ryogo Maeda, Qiong Chang, Jun Miyazaki |
iiWAS | 3 |
| 2025 | Deep Dual Internal Learning for Hyperspectral Image Super-Resolution
Yongqing Sun, Hong Liu 0009, Qiong Chang, Xianhua Han |
MMM (1) | 3 |
| 2025 | 3D-NLM: Voxel-based non-local means for 3D point cloud noise detection and smoothing
Weimin Wang 0007, Yu Liu 0012, Qiong Chang |
Comput. Graph. | 5 |
| 2025 | Direction-aware convolutional autoencoder based on positional encoding for one-dimensional anomaly detection
Qien Yu, Qiong Chang, Tinghui Ouyang, Takio Kurita 0001, Ran Dong |
Inf. Sci. | 2 |
| 2025 | Accelerating Nearest Neighbor Search in 3D Point Cloud Registration on GPUsabstractThe Iterative Closest Points (ICP) algorithm is the most widely used method for estimating rigid transformation in 3D point cloud registration. However, the ICP relies on repeatedly performing computationally intensive nearest neighbor searches (NNS) within 3D space. This dependency becomes a significant bottleneck when processing large datasets, thereby hindering the practical deployment of point cloud technologies in real-world applications. To address this issue, we propose two approximate nearest neighbor search (ANNS) acceleration strategies for efficient improvement of the processing speed of the NNS. Our strategies first voxelize target cloud points and then fill voxels in the 3D coordinate space around the source point cloud in two different ways, which can convert the global nearest neighbor search to a local search. Both the proposed methods are suited to be parallelized on GPUs with a low computational load. Extensive experiments show that our methods significantly accelerate NNS processing while maintaining high accuracy, outperforming most of the currently known approaches. Qiong Chang, Weimin Wang 0007, Jun Miyazaki |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | 3D GNLM: Efficient 3D Non-Local Means Kernel with Nested Reuse Strategies for Embedded GPUsabstractThe 3D Non-Local Means (NLM) algorithm has become a crucial preprocessing technique for 3D image datasets due to its effectiveness in denoising while preserving fine details. This method has been proven to be highly efficient in high-demand tasks within industrial applications such as medical imaging and remote sensing. The 3D NLM algorithm computes the filtered value for each voxel by calculating the weighted average of all voxels within a 3D search window, where the weights are determined by the similarity between pairs of 3D template windows. Therefore, the computational burden becomes significant, especially in embedded GPUs with limited computational power and memory resources. To address this issue, we propose an efficient GPU parallel kernel to minimize redundant computations and memory accesses. The kernel integrates three nested reuse strategies to handle redundant computations in three dimensions: for columns, we leverage the fast data exchange mechanism to reuse column computation results via on-chip registers; for rows, we use a sliding window strategy, utilizing GPU global memory as an intermediary to store and reuse similarity values between filtered rows; and for channels, we introduce a zigzag scanning strategy that enables simultaneous computation across multiple channels and employs on-chip registers to facilitate channel computation reuse. Experimental results demonstrate that our kernel achieves an average speedup of 7.7x on the embedded Jetson AGX Xavier platform across a range of 3D image datasets compared to existing methods, showcasing exceptional performance. Xiang Li 0110, Qiong Chang, Yun Li 0015, Jun Miyazaki |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | Faster than Fast: Accelerating Oriented FAST Feature Detection on Low-end Embedded GPUsabstractThe visual-based SLAM (Simultaneous Localization and Mapping) is a technology widely used in applications such as robotic navigation and virtual reality, which primarily focuses on detecting feature points from visual images to construct an unknown environmental map and simultaneously determines its own location. It usually imposes stringent requirements on hardware power consumption, processing speed, and accuracy. Currently, the ORB (Oriented FAST and Rotated BRIEF)-based SLAM systems have exhibited superior performance in terms of processing speed and robustness. However, they still fall short of meeting the demands for real-time processing on mobile platforms. This limitation is primarily due to the time-consuming Oriented FAST calculations accounting for approximately half of the entire SLAM system. This article presents two methods to accelerate the Oriented FAST feature detection on low-end embedded GPUs. These methods optimize the most time-consuming steps in Oriented FAST feature detection: FAST feature point detection and Harris corner detection, which is achieved by implementing a binary-level encoding strategy to determine candidate points quickly and a separable Harris detection strategy with efficient low-level GPU hardware-specific instructions. Extensive experiments on a Jetson TX2 embedded GPU demonstrate an average speedup of over 7.3 times compared to widely used OpenCV with GPU support. This significant improvement highlights its effectiveness and potential for real-time applications in mobile and resource-constrained environments. Qiong Chang, Xiang Li 0110, Weimin Wang 0007, Jun Miyazaki |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2024 | $k$-Way In-Place Merge by CPU-GPU Cooperative ProcessingabstractWe propose a high-performance$k$-way in-place merging algorithm by CPU-GPU cooperative processing. Current merging algorithms are either$k$-way in-place with CPUs only or in-place for data smaller than the GPU memory size. To solve this problem, we extend these merging algorithms to be more adaptable to a wider range of applications by improving the segmentation, reordering, and inter-segment merging methods used in the current merging algorithms by using both CPUs and GPUs cooperatively. We conducted experiments to evaluate the performance of our proposed algorithm and analyzed its detailed behavior by changing parameters and data distributions. We also applied the proposed merging algorithm to the merge sort algorithm to verify its performance. Shinya Miura, Qiong Chang, Jun Miyazaki |
ASAP | 2 |
| 2024 | Extension of Parallel Primitives and Their Applications to Large-Scale Data Processing
Masashi Nakano, Qiong Chang, Jun Miyazaki |
DEXA (2) | 2 |
| 2024 | Attribution-Based Scanline Perturbation Attack on 3d Detectors of Lidar Point CloudsabstractLiDAR point cloud data is widely utilized in autonomous driving systems and has significantly improved the 3D detection performance with well-designed deep neural network models. However, due to the complexity of real-world environments and model vulnerability, false detections or malicious attacks may cause severe accidents in unseen situations. In this paper, we propose a novel attack approach, Attribution-based Scanline Perturbation (ASP), an efficient and physically possible adversarial attack method for 3D detectors. ASP first utilizes attribution methods to identify critical points for the detection model and perturbs them along the laser beams by simulating the situation in which particles exist between the LiDAR sensor and objects, which can actually occur in snow or sandstorm weather. Extensive experiments on practical 3D detectors validate the effectiveness of our approach in misleading the model and causing both false and missed detections. Ziyang Yu 0004, Qiong Chang, Yu Liu 0012, Weimin Wang 0007 |
ICASSP | 3 |
| 2024 | A Data Model of a Data Lineage Management System for Database Repair and Simulation
Wei Jun Wong, Kyoko Yasuda, Qiong Chang, Jun Miyazaki |
iiWAS (2) | 3 |
| 2024 | An Optimized GPU Implementation for GIST DescriptorabstractThe GIST descriptor is a classic feature descriptor primarily used for scene categorization and recognition tasks. It drives a bank of Gabor filters, which respond to edges and textures at various scales and orientations to capture the spatial structures in an image. Compared to other scene recognition algorithms that rely on detailed object detection, GIST has lower computational complexity, allowing it to be widely applied. However, its internal multi-scale and multi-orientation Gabor filters also mean that systems based on it cannot be executed fast enough. This article proposes an optimized GPU kernel for the GIST descriptor. It fully takes advantage of the symmetry of Gabor filters and proposes different optimization strategies for both oblique and orthogonal orientations. Extensive experiments demonstrate that the proposed kernel is adaptable to images of various scales and different GPUs. Compared to the cuFFT library, our kernel achieves 12.09× and 3.86× acceleration on an RTX 3080 GPU and a Jetson AGX Xavier GPU, respectively. Xiang Li 0110, Qiong Chang, Aolong Zha, Shijie Chang, Yun Li 0015, Jun Miyazaki |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | TinyStereo: A Tiny Coarse-to-Fine Framework for Vision-Based Depth Estimation on Embedded GPUsabstractStereo vision, a popular depth estimation technology in computing vision, finds wide-ranging applications in embedded systems, including robotics vision and autonomous driving. These applications demand both high accuracy and fast processing speeds. To address hardware limitations, most current embedded systems rely on nonlearning algorithms for fast matching, sacrificing accuracy. Some recent studies have explored using convolutional neural networks (CNNs) to improve matching accuracy, but the computational load of existing learning-based systems hampers real-world applicability. This article presents significant contributions: 1) a novel stereo matching framework that greatly enhances accuracy on real-time embedded platforms and 2) a two-pronged approach combining a nonlearning-based algorithm and a lightweight super-resolution residual neural network (sRRNet). The nonlearning-based algorithm yields a low-resolution disparity map, while the lightweight sRRNet generates a high-resolution disparity map. Experimental results on benchmark data demonstrate that the proposed method achieves a low matching error rate of 5.17% and a real-time processing speed of 51 fps using the embedded Jetson AGX GPU. The proposed method outperforms all existing real-time embedded systems. Qiong Chang, Aolong Zha, Meng Joo Er, Yongqing Sun, Yun Li 0015 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | GPU Acceleration of Multi-Object Tracking with Motion Vector Interpolation and Affine TransformationabstractIn recent studies of object detection and tracking, neural networks have been widely used, and their accuracy has improved. However, its computational complexity is very high and requires the use of high-end GPUs. In order to achieve realtime inference on edge devices, it is necessary to reduce the computational complexity of the network by scaling it down, but this leads to a loss of accuracy. To avoid this loss of accuracy, a method has been proposed in which object detection is performed using a neural network at regular intervals, and in the frames in between, the detected object positions are interpolated using motion prediction. In this research, we propose a method to improve the accuracy of interpolation even when the camera is moving by using an affine transformation used for image stabilization. We also show its realtime computation method on Jetson TX2, one of the lowest power embedded GPUs. The proposed method enables realtime processing of object detection using Yolov5s and tracking of the detected objects at the edge. Yoshiki Kunimoto, Qiong Chang, Yoshiki Yamaguchi, Tsutomu Maruyama |
ASAP | 2 |
| 2023 | How Does the System Perceive Me? - A Transparent and Tunable Recommender System
Mingman Xu, Qiong Chang, Jun Miyazaki |
DEXA (2) | 2 |
| 2023 | VAN-ICP: GPU-Accelerated Approximate Nearest Neighbor Search for ICP Registration via Voxel DilationabstractThe Iterative Closest Points (ICP) algorithm and its variants have been widely applied for 3D point cloud registration which estimates the rigid transformation. As the most computationally intensive step in ICP, nearest neighbor search (NNS) takes up most of the execution time, hindering the practical applications of ICP registration. To overcome the bottleneck, we propose a novel GPU-friendly approximate nearest neighbor search (ANNS) acceleration scheme, named Voxel dilAtioN (VAN), which can efficiently convert the global search to local $\left. {\mathcal{O}(n)} \right)$. Extensive experiments demonstrate that our VAN can drastically boost the NNS processing while keeping high registration accuracy. Specifically, our GPU-based VAN-ICP achieves 2.7x, 7.6x, and 13.4x speedup on three datasets compared with the CPU-based ICP implementation of Point Cloud Library (PCL). Source codes are available at https://github.com/mfxox/VAN-ICP. Weimin Wang 0007, Qiong Chang |
ICASSP | 2 |
| 2023 | StereoVAE: A lightweight stereo-matching system using embedded GPUsabstractWe propose a lightweight system for stereo-matching using embedded graphic processing units (GPUs). The proposed system overcomes the trade-off between accuracy and processing speed in stereo matching, thus further improving the matching accuracy while ensuring real-time processing. The basic idea is to construct a tiny neural network based on a variational autoencoder (VAE) to achieve the upscaling and refinement a small size of coarse disparity map. This map is initially generated using a traditional matching method. The proposed hybrid structure maintains the advantage of low computational complexity found in traditional methods. Additionally, it achieves matching accuracy with the help of a neural network. Extensive experiments on the KITTI 2015 benchmark dataset demonstrate that our tiny system exhibits high robustness in improving the accuracy of coarse disparity maps generated by different algorithms, while running in real-time on embedded GPUs. Qiong Chang, Xiang Li 0110, Xin Liu 0020, Yun Li 0015, Jun Miyazaki |
ICRA | 1 |
| 2023 | Optimizing Local Feature Representations of 3D Point Clouds with Anisotropic Edge Modeling
Haoyi Xiu, Xin Liu 0020, Weimin Wang 0007, Kyoung-Sook Kim 0001, Takayuki Shinohara, Qiong Chang, Masashi Matsuoka |
MMM (1) | 6 |
| 2023 | Enhancing Lidar and Radar Fusion for Vehicle Detection in Adverse Weather via Cross-Modality Semantic Consistency
Qiong Chang, Weimin Wang 0007 |
PRCV (3) | 3 |
| 2023 | An incremental SAT-based approach for solving the real-time taxi-sharing service problemabstractThis paper deals with a combinatorial optimization problem that models real-time taxi-sharing services. Because finding an optimal solution to this problem is NP-hard, most previous studies have developed the corresponding optimization algorithms based on well-known metaheuristics for its simplified problem. In this study, we focus on assigning an appropriate taxi and re-planning its route for each newly arisen demand so that the service can minimize the sum of the planned travel time for transporting all passengers who are allocated to this taxi. We propose a novel algorithm based on incremental Boolean satisfiability solving, to optimize the taxi allocation for demands occurring in real time. In our experiment, we generate instances by simulating urban traffic based on a road network. The experimental result shows that our new approach is overall faster than its equivalent integer programming-based solving method. Aolong Zha, Qiong Chang, Itsuki Noda |
Discret. Appl. Math. | 2 |
| 2023 | Diffusion unit: Interpretable edge enhancement and suppression learning for 3D point cloud segmentationabstract3D point clouds are discrete samples of continuous surfaces which can be used for various applications. However, the lack of true connectivity information, i.e., edge information, makes point cloud recognition challenging. Recent edge-aware methods incorporate edge modeling into network designs to better describe local structures. Although these methods show that incorporating edge information is beneficial, how edge information helps remains unclear, making it difficult for users to analyze its usefulness. To shed light on this issue, in this study, we propose a new algorithm called Diffusion Unit (DU) that handles edge information in a principled and interpretable manner while providing decent improvement. First, we theoretically show that DU learns to perform task-beneficial edge enhancement and suppression. Second, we experimentally observe and verify the edge enhancement and suppression behavior. Third, we empirically demonstrate that this behavior contributes to performance improvement. Extensive experiments and analyses performed on challenging benchmarks verify the effectiveness of DU. Specifically, our method achieves state-of-the-art performance in object part segmentation using ShapeNet part and scene segmentation using S3DIS. Our source code is available at https://github.com/martianxiu/DiffusionUnit. Haoyi Xiu, Xin Liu 0020, Weimin Wang 0007, Kyoung-Sook Kim 0001, Takayuki Shinohara, Qiong Chang, Masashi Matsuoka |
Neurocomputing | 6 |
| 2023 | Multi-directional Sobel operator kernel on GPUs
Qiong Chang, Xiang Li 0110, Yun Li 0015, Jun Miyazaki |
J. Parallel Distributed Comput. | 1 |
| 2023 | A deep learning framework for realistic robot motion generation
Ran Dong, Qiong Chang, Soichiro Ikuno |
Neural Comput. Appl. | 2 |
| 2022 | Acceleration of video stabilization using embedded GPUabstractVideo stabilization is a technique used to eliminate the shakiness in video. It plays an important role to improve the quality of the videos captured by cameras mounted on drones and autonomous robots, and handheld cameras. In these cases, it is generally required to achieve real-time video stabilization using low-power and low-cost devices. In this research, we focus on software video stabilization, and propose a new implementation method using an embedded GPU. Our target device is Nvidia Jetson Nano, which is one of the smallest and least power consumption embedded GPUs. In our implementation, 1) multi-threads on the CPU on Jetson Nano and the GPU run asynchronously and in parallel to achieve high performance, 2) the size of the search area is minimized to reduce the amount of computation assuming that sufficiently fast processing speed is possible using the GPU, and 3) the data transfer from the global memory of the GPU is minimized to hide the high latency. Our implementation achieves 81.4 fps for full HD videos, and its quality is high enough for practical use. Yuzuki Mimura, Qiong Chang, Tsutomu Maruyama |
ASAP | 2 |
| 2022 | Jointly Learning Propagating Features on the Knowledge Graph for Movie Recommendation
Yun Liu 0044, Jun Miyazaki, Qiong Chang |
DEXA (1) | 3 |
| 2022 | Efficient stereo matching on embedded GPUs with zero-means cross correlation
Qiong Chang, Aolong Zha, Weimin Wang 0007, Xin Liu 0020, Masaki Onishi, Meng Joo Er, Tsutomu Maruyama |
J. Syst. Archit. | 1 |
| 2021 | Enhancing Local Feature Learning for 3D Point Cloud Processing using Unary-Pairwise Attention
Haoyi Xiu, Xin Liu 0020, Kyoung-Sook Kim 0001, Takayuki Shinohara, Qiong Chang, Masashi Matsuoka |
BMVC | 6 |
| 2021 | Fast SQL/Row Pattern Recognition Query Processing Using Parallel Primitives on GPUs
Tsubasa Ohara, Qiong Chang, Jun Miyazaki |
DEXA (1) | 2 |
| 2020 | Z2-ZNCC: ZigZag Scanning based Zero-means Normalized Cross Correlation for Fast and Accurate Stereo Matching on Embedded GPUabstractMobile stereo matching systems are becoming more important in many applications such as auto-driving and autonomous robots. However, to maintain its low power consumption, mobile platforms have only limited hardware resources. Accurate stereo matching methods require a high computational complexity, and it is difficult to maintain both acceptable accuracy and processing speed on the mobile platforms. To solve this trade-off, in this paper, we propose a novel acceleration approach for a well-known matching algorithm Zero-means Normalized Cross Correlation (ZNCC), and show its effectiveness on a Jetson TX2 embedded GPU. By combining our new approach, Z2- ZNCC, with the Semi-Global Matching (SGM) algorithm, our system achieves a low error rate of 7.76% while keeping 28 fps for 1242×375 pixels images with the maximum disparity of 128 on the KITTI 2015 dataset. This performance is higher than previous state-of-the-art system on the same hardware platform. Qiong Chang, Aolong Zha, Weimin Wang 0007, Xin Liu 0020, Masaki Onishi, Tsutomu Maruyama |
ICCD | 1 |
| 2020 | CNF Encodings for the Min-Max Multiple Traveling Salesmen ProblemabstractIn this study, we consider the multiple traveling salesmen problem (mTSP) with the min-max objective of minimizing the longest tour length. We begin by reviewing an existing integer programming (IP) formulation of this problem. Then, we present several novel conjunctive normal form (CNF) encodings and an approach based on modifying a maximum satisfiability (MaxSAT) algorithm for the min-max mTSP. The correctness and the space complexity of each encoding are analyzed. In our experiments, we compare the performance of solving the TSP benchmark instances using an existing encoding and our new encodings comparing the results achieved using an implemented group MaxSAT solver to those achieved using the IP method. The results show that for the same problem, the new encodings significantly reduce the number of generated clauses over the existing CNF encoding. Although the proposals are still not competitive compared to the IP method, one of them may be more effective on relatively large-scale problems, and it has an advantage over the IP method in solving an instance with a small ratio of the number of cities to the number of salesmen. Aolong Zha, Rongxuan Gao, Qiong Chang, Miyuki Koshimura, Itsuki Noda |
ICTAI | 3 |
| 2018 | Real-Time High-Quality Stereo Matching System on a GPUabstractIn this paper, we propose a low error rate and realtime stereo vision system on G PU. Many stereo vision systems on G PU have been proposed to date. In those systems, the error rates and the processing speed are in trade-off relationship. We propose a real-time stereo vision system on GPU for the high resolution images. This system also maintains a low error rate compared to other fast systems. In our approach, we have implemented the cost aggregation (CA), cross-checking and median filter on GPU in order to realize the real-time processing. Its processing speed is 40 fps for 1436×992 pixels images when the maximum disparity is 145, and its error rate is the lowest among the GPU systems which are faster than 30 fps. Qiong Chang, Tsutomu Maruyama |
ASAP | 1 |