EDBT 2026 Demo / reviewers in the wild / expert
Qichuan Geng
dblp:217/2419
· DBLP profile ↗
26ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0002-0046-5794ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GeoBayes: Probabilistic Image Geo-Localization Inference via Sequential Bayesian UpdatingabstractImage geo-localization aims to determine the geographic location of a query image. While Multimodal Large Language Models (MLLMs) show potential for this task due to their rich world knowledge and explainable abilities, they often struggle with confirmation bias, i.e., committing to early, potentially incorrect guesses driven by visual clues with varied geographic likelihoods. In this paper, we propose GeoBayes, a novel training-free framework that formulates geolocalization as a Maximum a Posteriori (MAP) estimation task over multiple geographic hypotheses and performs probabilistic thought via sequential Bayesian reasoning. GeoBayes treats each visual object and its associated geographic clues as probabilistic evidence, integrating them iteratively through a Hypothesize–Verify–Update loop. At each step, it evaluates how new evidence supports existing hypotheses and updates their posterior probabilities, gradually converging on the most probable location. This allows GeoBayes to explicitly quantify and fuse the varied geographic probabilities implied by various visual elements, reducing the risk of overcommitting to misleading clues. Furthermore, considering the natural hierarchy of geographic labels (e.g., country, city), GeoBayes introduces a state memory mechanism that stores hypotheses, inference context, and evidence scores across levels. This design enables the framework to propagate prior knowledge across levels of the geographic hierarchy and incorporate geographic structural constraints into the Bayesian update process, achieving a coarse-to-fine geo-localization. Experiments on IM2GPS3k and YFCC4K show that GeoBayes improves MLLM-based geo-localization accuracy without extra training. This demonstrates the effectiveness of probabilistic reasoning for robust and interpretable geo-localization. Kaige Li, Junhao Fang, Qichuan Geng, Zhong Zhou |
AAAI | 6 |
| 2026 | Out-of-Distribution Semantic Segmentation With Disentangled and Calibrated RepresentationabstractOut-of-distribution (OoD) semantic segmentation aims to recognize pixels of classes undefined in the training dataset. Existing methods mostly focus on training the model to fit real OoD data samples to identify OoD pixels, which requires extra data collection and annotation efforts. By contrast, synthesizing OoD data with training data provides a more resource-efficient alternative. However, synthetic data generated from controlled settings lacks diversity, causing the model to suffer from overfitting. To this end, we propose a disentangled representation learning (DRL) method to guide the model to disentangle semantic-related and semantic-unrelated features from synthetic OoD data. DRL encourages the model to utilize the former to identify semantic categories, rather than overfitting to such semantic-unrelated features as synthetic artificiality. Specifically, DRL first incorporates two disentanglers to extract the semantic-related and -unrelated features and then applies a shuffle and reconstruction mechanism to regularize the disentangled features. Furthermore, to facilitate disentangling, we propose a pixel-wise feature similarity calibration (PSC) module, which utilizes more accurate ID-OoD similarity to calibrate inaccurate ID-OoD similarity learned exclusively from ID data. Thus, PSC delivers accurate and stable pixel-wise features for effective disentangling. Extensive experiments illustrate that the proposed method exhibits strong generalization ability. It attains 74.04% AuPRC and 20.82% FPR on Road Anomaly, 69.85% AuPRC and 5.78% FPR on Fishyscapes LostAndFound Validation Set, using SegFormer with the MiT-B5 backbone. Source code is available at https://github.com/WanMotion/DisentangledOoDSeg. Maoxian Wan, Kaige Li, Qichuan Geng, Binyi Su, Zhong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | MaskScene: Hierarchical Conditional Masked Models for Real-Time 3D Indoor Scene SynthesisabstractIndoor scene synthesis is essential for creative industries. Driven by this demand, recent advances in scene synthesis using diffusion and autoregressive models have shown promising results. However, existing models struggle to achieve real-time performance, high visual fidelity, and flexible scene editing simultaneously. To tackle this challenge, we propose MaskScene, a novel hierarchical conditional masked model for real-time 3D indoor scene synthesis and editing. Specifically, MaskScene introduces a hierarchical scene representation that explicitly encodes scene relationships, semantics, and tokens. Based on this representation, we design a hierarchical conditional masked modeling architecture that enables parallel and iterative decoding, conditioned on both semantics and relationships. MaskScene leverages local object masking and hierarchical scene structures to infer occluded regions from partial observations, enabling rapid construction of 3D indoor environments that accurately reflect real-world scenes. Compared to state-of-the-art methods, MaskScene achieves 80× faster generation speed and improves scene quality by 10%, while also supporting zero-shot editing, such as scene completion and rearrangement, without extra fine-tuning. Qichuan Geng, Zhong Zhou, Wenfeng Song |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | Enhancing open-vocabulary scene understanding via push-pull alignment in gaussian splatting
Shengjia Liang, Yuan Xiong, Qichuan Geng, Zhong Zhou |
Vis. Comput. | 5 |
| 2026 | CleanSplat: curriculum structural Gaussian splatting for spot-free novel view synthesisabstract3D Gaussian splatting (3DGS) has gained significant attention for its real-time, photorealistic rendering in novel view synthesis. However, its performance degrades severely when applied to real-world scenes with transients that break cross-view consistency. Existing methods typically attempt to identify and mask out these transients, but inherent masking errors often leave behind incoherent floating Gaussians, resulting in spotted artifacts that degrade the final scene quality. To address these limitations, we propose CleanSplat, a curriculum structural Gaussian splatting method that purifies the scene Gaussians through curriculum 3DGS optimization for transient removal and structural pruning of spotted artifacts. We introduce a curriculum-guided masking paradigm that generates coarse-to-fine transient masks from multi-scale features. The progressive optimization is driven by modulating the masking supervision based on current training state. To clear the spots, we propose a structure-aware handling strategy that employs a superpoint graph (SPG) partitioning of the Gaussians to perform principled identification and hierarchical pruning. This allows for the filtering of both intra-superpoint outliers and entire spurious superpoints based on local 3D coherence instead of only simple photometric consistency. By integrating curriculum 3DGS optimization and structural pruning, our method effectively separates the transients and purifies the static scene Gaussians. Extensive experiments on challenging datasets demonstrate that CleanSplat significantly outperforms state-of-the-art methods, delivering more detailed and cleaner novel view synthesis. Shengjia Liang, Qichuan Geng, Yuan Xiong, Zhong Zhou |
Virtual Real. Intell. Hardw. | 3 |
| 2025 | Spatial-Aware Symmetric Alignment for Text-Guided Medical Image SegmentationabstractText-guided Medical Image Segmentation has shown considerable promise for medical image segmentation, with rich clinical text serving as an effective supplement for scarce data. However, current methods have two key bottlenecks. On one hand, they struggle to process diagnostic and descriptive texts simultaneously, making it difficult to identify lesions and establish associations with image regions. On the other hand, existing approaches focus on lesions description and fail to capture positional constraints, leading to critical deviations. Specifically, with the text “in the left lower lung”, the segmentation results may incorrectly cover both sides of the lung. To address the limitations, we propose the Spatial-aware Symmetric Alignment (SSA) framework to enhance the capacity of referring hybrid medical texts consisting of locational, descriptive, and diagnostic information. Specifically, we propose symmetric optimal transport alignment mechanism to strengthen the associations between image regions and multiple relevant expressions, which establishes bi-directional fine-grained multimodal correspondences. In addition, we devise a composite directional guidance strategy that explicitly introduces spatial constraints in the text by constructing region-level guidance masks. Extensive experiments on public benchmarks demonstrate that SSA achieves state-of-the-art (SOTA) performance, particularly in accurately segmenting lesions characterized by spatial relational constraints. Linglin Liao, Qichuan Geng |
BIBM | 3 |
| 2025 | Incremental Few-Shot Semantic Segmentation Via Multi-Level Switchable Visual Prompts
Maoxian Wan, Kaige Li, Qichuan Geng, Zhong Zhou |
ICCV | 3 |
| 2025 | 3D Lane Detection Based on Projection-Consistent Reference Points and Intra- & Inter-lane Contextabstract3D lane detection aims to identify lane categories and trends in 3D space, which is a vital and challenging task in autonomous driving. Existing methods introduce various priors to guide 3D lane prediction, which generally consist of a series of reference points for context aggregation. However, due to the misalignment between these reference points and the lanes, it is difficult to obtain complete and discriminative context for complex instances. In this paper, we are devoted to introducing 3D priors adaptive to lane appearances, which serve as references to aggregate the lane context. Specifically, we propose a projection-consistent reference generation strategy to keep the projected 3D reference points geometrically consistent with the corresponding lanes in images. In addition, a segmentation-lifting denoising strategy is designed to improve the ability of the model to map the lane segmentation into 3D space. To leverage more lane-related information, we propose a decoupled lane-context aggregation module by considering the perspectives of individual geometries and integrated layout, namely intra-lane and inter-lane context. Extensive experiments on the OpenLane dataset show that our approach outperforms previous methods and achieves the state-of-the-art performance. The code will be made publicly available. Yiqiu Bing, Huilin Niu, Hong Zhang 0009, Zhong Zhou, Qichuan Geng |
ICRA | 6 |
| 2025 | Understanding Matters: Semantic-Structural Determined Visual Relocalization for Large ScenesabstractScene Coordinate Regression (SCR) estimates 3D scene coordinates from 2D images, and has become an important approach in visual relocalization. Existing methods exhibit high localization accuracy in small scenes, but still face substantial challenges in large-scale scenes, which usually have significant variations in depth, scale, and occlusion. Although structure-guided scene partitioning is commonly adopted, the over-partitioned elements and large feature variances within subscenes impede the estimation of the 3D coordinates, introducing misleading information for subsequent processing. To address the above-mentioned issues, we propose the Semantic-Structural Determined Visual Relocalization method for SCR, which leverages semantic-structural partition learning and partition-determined pose refinement to better understand the semantic and structural information on large scenes. Firstly, we partition the scene into small subscenes with label assignments, ensuring semantic consistency and structural continuity within each subscene. A classifier is then trained with sampling-based learning to predict these labels. Secondly, the partition predictions are encoded into embeddings and integrated with local features for intra-class compactness and inter-class separation, producing partition-aware features. To further decrease feature variances, we employ a discriminability metric and suppress ambiguous points, improving subsequent computations. Experimental results on the Cambridge Landmarks dataset demonstrate that the proposed method achieves significant improvements with fewer training costs on large-scale scenes, reducing the median error by 38% compared to the state-of-the-art SCR method DSAC*. Code is available: https://gitee.com/VR_NAVE/ss-dvr. Jingyi Nie, Liangliang Cai, Qichuan Geng, Zhong Zhou |
IJCAI | 3 |
| 2025 | Boosting Few-Shot Open-Set Object Detection via Prompt Learning and Robust Decision BoundaryabstractFew-shot Open-set Object Detection (FOOD) poses a challenge in many open-world scenarios. It aims to train an open-set detector to detect known objects while rejecting unknowns with scarce training samples. Existing FOOD methods are subject to limited visual information, and often exhibit an ambiguous decision boundary between known and unknown classes. To address these limitations, we propose the first prompt-based few-shot open-set object detection framework, which exploits additional textual information and delves into constructing a robust decision boundary for unknown rejection. Specifically, as no available training data for unknown classes, we select pseudo-unknown samples with Attribution-Gradient based Pseudo-unknown Mining (AGPM), which leverages the discrepancy in attribution gradients to quantify uncertainty. Subsequently, we propose Conditional Evidence Decoupling (CED) to decouple and extract distinct knowledge from selected pseudo-unknown samples by eliminating opposing evidence. This optimization process can enhance the discrimination between known and unknown classes. To further regularize the model and form a robust decision boundary for unknown rejection, we introduce Abnormal Distribution Calibration (ADC) to calibrate the output probability distribution of local abnormal features in pseudo-unknown samples. Our method achieves superior performance over previous state-of-the-art approaches, improving the average recall of unknown class by 7.24% across all shots in VOC10-5-5 dataset settings and 1.38% in VOC-COCO dataset settings. Our source code is available at https://gitee.com/VR_NAVE/ced-food. Zhaowei Wu, Binyi Su, Qichuan Geng, Zhong Zhou |
IJCAI | 3 |
| 2025 | DGDiff: Immersive 3D Indoor Scene Synthesis via Dialog-Graph Conditioned DiffusionabstractImmersive 3D indoor scene synthesis is essential for applications such as AR/VR and 3D content creation. However, existing approaches fail to meet the immersive AR/VR requirements for fidelity, user-system interaction, and production speed simultaneously. Traditional scene synthesis methods are overly rigid, limiting user interactivity, whereas large language model (LLM)-based approaches suffer from slow response times and imprecise spatial structuring. To address these issues, we propose DGDiff, a novel dialog-graph conditioned diffusion framework for immersive, controllable, continuous synthesis and editing of 3D indoor scenes. This framework combines a conversational module powered by LLMs with a multimodal diffusion model. The conversational module translates user dialogue into structured semantic graphs, while the diffusion model integrates textual and graph-based conditions to synthesize realistic, editable indoor scenes. Experimental results demonstrate that DGDiff outperforms single-modality baselines, achieving an improvement of over 10 % in FID and a reduction of approximately 30 % in response time for dynamic scene interactive editing, offering an immersive and user-friendly synthesis experience. Project page: https://gitee.com/VR_NAVE/dgdiff.git Qichuan Geng, Zhong Zhou |
ISMAR | 3 |
| 2025 | An Internal Knowledge Maintaining Mechanism in Pre-trained Student Model for Enhancing Generalization Ability of CoT Distillation from LLMs
Zunbo Liu, Qichuan Geng, Chenhan Yang, Xiangjie Pan |
PRCV (12) | 2 |
| 2025 | AU-Guided Feature Aggregation for Micro-Expression RecognitionabstractABSTRACT Micro‐expressions (MEs) are spontaneous and transient facial movements that reflect real internal emotions and have been widely applied in various fields. Recent deep learning‐based methods have been rapidly developing in micro‐expression recognition (MER).Still, it is typical to focus on the one‐sided nature of MEs, covering only representational features or low‐ranking Action Unit (AU) features. The subtle changes in MEs characterize its feature representation weak and inconspicuous, making it tough to analyze MEs only from a single piece or a small amount of information to achieve a considerable recognition effect. In addition, the lower‐order information can only distinguish MEs from a single low‐dimensional perspective and neglects the potential of corresponding MEs and AU combinations to each other. To address these issues, we first explore how the higher‐order relations of different AU combinations correspond with MEs through statistical analysis. Afterward, based on this attribute, we propose an end‐to‐end multi‐stream model that integrates global feature learning and local muscle movement representation guided by AU semantic information. The comparative experiments were performed on benchmark datasets, with better performance than the state‐of‐art methods. Also, the ablation experiments demonstrate the necessity of our model to introduce the information of AU and its relationship to MER. Xiaohui Tan, Jiazheng Wu, Hao Geng, Qichuan Geng |
Comput. Animat. Virtual Worlds | 5 |
| 2025 | MAP: Masked Adversarial Perturbation for Boosting Black-Box Attack TransferabilityabstractThe transferability of adversarial examples is vital for black-box attacks, as it enables the adversary to deceive the target model without knowing its internals. Despite numerous methods focusing on transferability, they still struggle with transferring across models with distinct architectural components (e.g., CNNs and ViTs). In this work, we argue that the limited adversarial perturbation diversity leads to overfitting of the surrogate model, which acts as a key factor in reducing transferability. To this end, we propose a Masked Adversarial Perturbation (MAP) method to boost adversarial transferability across various architectures from a novel perspective of diversifying perturbation. Specifically, MAP randomly masks perturbation patches during iterations and compels the remaining ones to retain the attack effect, which diversifies perturbations to mitigate their overfitting to the surrogate model. Naturally, MAP spreads perturbation over local patches to alleviate their co-adaptation and prevent perturbations from overly relying on specific patterns. Consequently, it can deceive convolution operation and self-attention mechanism indiscriminately by attacking their basic input units, i.e., a single patch, showing superior transferability over previous methods. Extensive experiments illustrate that MAP consistently and significantly boosts diverse black-box attacks to achieve state-of-the-art performance. Kaige Li, Maoxian Wan, Qichuan Geng, Xiaochun Cao, Zhong Zhou |
IEEE Trans. Image Process. | 3 |
| 2024 | A survey on person and vehicle re-identificationabstractAbstract Person/vehicle re‐identification aims to use technologies such as cross‐camera retrieval to associate the same person (same vehicle) in the surveillance videos at different locations, different times, and images captured by different cameras so as to achieve cross‐surveillance image matching, person retrieval and trajectory tracking. It plays an extremely important role in the fields of intelligent security, criminal investigation etc. In recent years, the rapid development of deep learning technology has significantly propelled the advancement of re‐identification (Re‐ID) technology. An increasing number of technical methods have emerged, aiming to enhance Re‐ID performance. This paper summarises four popular research areas in the current field of re‐identification, focusing on the current research hotspots. These areas include the multi‐task learning domain, the generalisation learning domain, the cross‐modality domain, and the optimisation learning domain. Specifically, the paper analyses various challenges faced within these domains and elaborates on different deep learning frameworks and networks that address these challenges. A comparative analysis of re‐identification tasks from various classification perspectives is provided, introducing mainstream research directions and current achievements. Finally, insights into future development trends are presented. Zhaofa Wang, Zhi-Ping Shi 0002, Qichuan Geng |
IET Comput. Vis. | 5 |
| 2024 | Exploring Scale-Aware Features for Real-Time Semantic Segmentation of Street ScenesabstractReal-time semantic segmentation of street scenes is an essential and challenging task for autonomous driving systems, which needs to achieve both high accuracy and efficiency. Moreover, numerous objects and stuff at different scales in street scenes further increase the difficulty of this task. To address this challenge, we develop a lightweight and high-accuracy network termed Scale-Aware Network (SANet), which aims to selectively aggregate multi-scale features while maintaining high efficiency. In SANet, we first design a Selective Context Encoding (SCE) module, which considers the intrinsic differences of various pixels to selectively encode private contexts for each pixel, thus learning more desirable contextual features while reducing redundancy. With the context embedding in hand, we then design a Selective Feature Fusion (SFF) module to recursively fuses them with multiple features at different levels or scales to generate scale-aware features, where each feature map contains scale-specific information. Extensive experiments on challenging street scene datasets, i.e., Cityscapes and CamVid, illustrate that our SANet achieves a leading trade-off between segmentation accuracy and speed. Concretely, our method yields$78.1\%$mIoU at$109.0$FPS on the Cityscapes test set and$77.2\%$mIoU at$250.4$FPS on the CamVid test set. Code will be available at https://github.com/kaigelee/SANet. Kaige Li, Qichuan Geng, Zhong Zhou |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Exploiting 3D Human Recovery for Action Recognition with Spatio-Temporal Bifurcation FusionabstractAction recognition utilizes information in images or videos to analyze and classify human behaviors. The existing methods usually exploit 2D pose to improve classification features. Due to the lack of 3D cues, some approximate behaviors in 2D perspective cannot be recognized. In this paper, we propose a novel action recognition method with 3D human recovery and spatio-temporal bifurcations fusion. It consists of 3D spatial branch and 2D temporal branch. The 3D spatial branch exploits overlapping human models from 3D recovery to learn 3D recognition feature. The 2D temporal branch utilizes channel attention mechanism to enhance dynamic associated feature. Two types of features are fused by adaptive weights module, in order to improve the recognition of approximate behavior in 2D perspective. Extensive experiments show that this proposed method outperforms most of the state-of-the-art methods on the Olympic Sport, Diving48 and Human3.6M datasets. Qichuan Geng, Zhi-Ping Shi 0002 |
ICASSP | 3 |
| 2023 | RDEPD: Re-Exploring Depth Estimation for Pedestrian DetectionabstractPedestrian detection is a fundamental task in computer vision field. It remains challenging due to perspective affine and object occlusion. To alleviate these problems, this paper re-thinks the assistance of depth estimation, and then proposes a novel pedestrian detection method named RDEPD. The method consists of a data augmentation module based on depth (DamD) and a detection framework with learnable attention module based on depth(LAmD) and self-suitable NMS(S2NMS). DamD provides hard examples with mosaics at the depth gradient, which can improve the generalization ability. The detection framework pays more attention on these instances with occlusion or perspective affine by LamD and S2NMS. Where LAmD is responsible for integrating depth and RGB cues to guide localization, and S2NMS exploits every possible predicting box to improve the detecting precision. Extensive experimental results demonstrate that the proposed RDEPD significantly outperforms most state-of-the-art methods on the authoritative MOT17det, CrowdHuman, and Citypersons datasets, especially for occlusion situations. Yifei Pei, Zhi-Ping Shi 0002, Qichuan Geng, Zhaofa Wang, Yongkang Zhang 0001 |
ICIP | 3 |
| 2023 | Context and Spatial Feature Calibration for Real-Time Semantic SegmentationabstractContext modeling or multi-level feature fusion methods have been proved to be effective in improving semantic segmentation performance. However, they are not specialized to deal with the problems of pixel-context mismatch and spatial feature misalignment, and the high computational complexity hinders their widespread application in real-time scenarios. In this work, we propose a lightweight Context and Spatial Feature Calibration Network (CSFCN) to address the above issues with pooling-based and sampling-based attention mechanisms. CSFCN contains two core modules: Context Feature Calibration (CFC) module and Spatial Feature Calibration (SFC) module. CFC adopts a cascaded pyramid pooling module to efficiently capture nested contexts, and then aggregates private contexts for each pixel based on pixel-context similarity to realize context feature calibration. SFC splits features into multiple groups of sub-features along the channel dimension and propagates sub-features therein by the learnable sampling to achieve spatial feature calibration. Extensive experiments on the Cityscapes and CamVid datasets illustrate that our method achieves a state-of-the-art trade-off between speed and accuracy. Concretely, our method achieves 78.7% mIoU with 70.0 FPS and 77.8% mIoU with 179.2 FPS on the Cityscapes and CamVid test sets, respectively. The code is available at https://nave.vr3i.com/ and https://github.com/kaigelee/CSFCN. Kaige Li, Qichuan Geng, Maoxian Wan, Xiaochun Cao, Zhong Zhou |
IEEE Trans. Image Process. | 2 |
| 2022 | Multi-Task Feature Decomposition Based Marginal Distribution for Person SearchabstractPerson search is a composite task, aiming at locating and identifying a query person from uncropped images. It requires jointly solving Pedestrian Detection and Person Re-identification. One major challenge in person search is the contradictory goals of detection and re-identification. The model has to simultaneously model the universality and specificity of persons. In this paper, we propose a novel parameter-free approach called Feature Decomposition Person Search (FDPS) to separate various tasks. FDPS decomposes the ROI feature map to extract sub-features based on the marginal distribution for different tasks. Also, we find that the Online Instance Match loss pays imbalanced attention to positive and negative categories. We present a Balance Online Instance Match (BOIM) loss to enhance the contribution of negative categories during training. Our method achieves the state-of-the-art performance in one-step methods on two prevailing benchmarks, with high efficiency. Yuanzhe Yang, Qichuan Geng, Chengxiang Chu, Zhong Zhou |
ICME | 3 |
| 2022 | Part-Level Car Parsing and Reconstruction in Single Street View ImagesabstractPart information has been proven to be resistant to occlusions and viewpoint changes, which are main difficulties in car parsing and reconstruction. However, in the absence of datasets and approaches incorporating car parts, there are limited works that benefit from it. In this paper, we propose the first part-aware approach for joint part-level car parsing and reconstruction in single street view images. Without labor-intensive part annotations on real images, our approach simultaneously estimates pose, shape, and semantic parts of cars. There are two contributions in this paper. First, our network introduces dense part information to facilitate pose and shape estimation, which is further optimized with a novel 3D loss. To obtain part information in real images, a class-consistent method is introduced to implicitly transfer part knowledge from synthesized images. Second, we construct the first high-quality dataset containing 348 car models with physical dimensions and part annotations. Given these models, 60K synthesized images with randomized configurations are generated. Experimental results demonstrate that part knowledge can be effectively transferred with our class-consistent method, which significantly improves part segmentation performance on real street views. By fusing dense part information, our pose and shape estimation results achieve the state-of-the-art performance on the ApolloCar3D and outperform previous approaches by large margins in terms of both A3DP-Abs and A3DP-Rel. Qichuan Geng, Hong Zhang 0009, Feixiang Lu, Xinyu Huang 0001, Sen Wang 0003, Zhong Zhou, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Frequency Domain Image Translation: More Photo-realistic, Better Identity-preservingabstractImage-to-image translation has been revolutionized with GAN-based methods. However, existing methods lack the ability to preserve the identity of the source domain. As a result, synthesized images can often over-adapt to the reference domain, losing important structural characteristics and suffering from suboptimal visual quality. To solve these challenges, we propose a novel frequency domain image translation (FDIT) framework, exploiting frequency information for enhancing the image generation process. Our key idea is to decompose the image into low-frequency and high-frequency components, where the high-frequency feature captures object structure akin to the identity. Our training objective facilitates the preservation of frequency information in both pixel space and Fourier spectral space. We broadly evaluate FDIT across five large-scale datasets and multiple tasks including image translation and GAN inversion. Extensive experiments and ablations show that FDIT effectively preserves the identity of the source image, and produces photo-realistic images. FDIT establishes state-of-the-art performance, reducing the average FID score by 5.6% compared to the previous best method. Mu Cai, Hong Zhang 0009, Huijuan Huang 0001, Qichuan Geng, Yixuan Li 0001, Gao Huang 0001 |
ICCV | 4 |
| 2021 | Gated Path Selection Network for Semantic SegmentationabstractSemantic segmentation is a challenging task that needs to handle large scale variations, deformations, and different viewpoints. In this paper, we develop a novel network named Gated Path Selection Network (GPSNet), which aims to adaptively select receptive fields while maintaining the dense sampling capability. In GPSNet, we first design a two-dimensional SuperNet, which densely incorporates features from growing receptive fields. And then, a Comparative Feature Aggregation (CFA) module is introduced to dynamically aggregate discriminative semantic context. In contrast to previous works that focus on optimizing sparse sampling locations on regular grids, GPSNet can adaptively harvest free form dense semantic context information. The derived adaptive receptive fields and dense sampling locations are data-dependent and flexible which can model various contexts of objects. On two representative semantic segmentation datasets, i.e., Cityscapes and ADE20K, we show that the proposed approach consistently outperforms previous methods without bells and whistles. Qichuan Geng, Hong Zhang 0009, Xiaojuan Qi 0001, Gao Huang 0001, Ruigang Yang, Zhong Zhou |
IEEE Trans. Image Process. | 1 |
| 2020 | Automatic façade recovery from single nighttime image
Qichuan Geng, Zhong Zhou, Wei Wu 0008 |
Frontiers Comput. Sci. | 2 |
| 2020 | The ApolloScape Open Dataset for Autonomous Driving and Its ApplicationabstractAutonomous driving has attracted tremendous attention especially in the past few years. The key techniques for a self-driving car include solving tasks like 3D map construction, self-localization, parsing the driving road and understanding objects, which enable vehicles to reason and act. However, large scale data set for training and system evaluation is still a bottleneck for developing robust perception models. In this paper, we present the ApolloScape dataset [1] and its applications for autonomous driving. Compared with existing public datasets from real scenes, e.g., KITTI [2] or Cityscapes [3] , ApolloScape contains much large and richer labelling including holistic semantic dense point cloud for each site, stereo, per-pixel semantic labelling, lanemark labelling, instance segmentation, 3D car instance, high accurate location for every frame in various driving videos from multiple sites, cities and daytimes. For each task, it contains at lease 15x larger amount of images than SOTA datasets. To label such a complete dataset, we develop various tools and algorithms specified for each task to accelerate the labelling process, such as joint 3D-2D segment labeling, active labelling in videos etc. Depend on ApolloScape, we are able to develop algorithms jointly consider the learning and inference of multiple tasks. In this paper, we provide a sensor fusion scheme integrating camera videos, consumer-grade motion sensors (GPS/IMU), and a 3D semantic map in order to achieve robust self-localization and semantic segmentation for autonomous driving. We show that practically, sensor fusion and joint learning of multiple tasks are beneficial to achieve a more robust and accurate system. We expect our dataset and proposed relevant algorithms can support and motivate researchers for further development of multi-sensor fusion and multi-task learning in the field of computer vision. Xinyu Huang 0001, Peng Wang 0001, Xinjing Cheng, Dingfu Zhou, Qichuan Geng, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Survey of recent progress in semantic image segmentation with CNNs
Qichuan Geng, Zhong Zhou, Xiaochun Cao |
Sci. China Inf. Sci. | 1 |