VLDB 2026 Research / reviewers in the wild / expert
Qi Chen 0014
dblp:66/6320-14
· DBLP profile ↗
51ranked-venue papers
14as first author
38since 2021 · last 2026
0000-0001-8732-8049ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 12 first-author · 25 since 2021Artificial intelligence and machine learning · 33 · 10 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured VideosabstractMulti-object tracking (MOT) aims to track multiple objects while maintaining consistent identities across frames of a given video. In unmanned aerial vehicle (UAV) recorded videos, frequent viewpoint changes and complex UAV-ground relative motion dynamics pose significant challenges, which often lead to unstable affinity measurement and ambiguous association. Existing methods typically model motion and appearance cues separately, overlooking their spatio-temporal interplay and resulting in suboptimal tracking performance. In this work, we propose AMOT, which jointly exploits appearance and motion cues through two key components: an Appearance-Motion Consistency (AMC) matrix and a Motion-aware Track Continuation (MTC) module. Specifically, the AMC matrix computes bi-directional spatial consistency under the guidance of appearance features, enabling more reliable and context-aware identity association. The MTC module complements AMC by reactivating unmatched tracks through appearance-guided predictions that align with Kalman-based predictions, thereby reducing broken trajectories caused by missed detections. Extensive experiments on three UAV benchmarks, including VisDrone2019, UAVDT, and VT-MOT-UAV, demonstrate that our AMOT outperforms current state-of-the-art methods and generalizes well in a plug-and-play and training-free manner. Jianbo Ma 0001, Hui Luo 0002, Qi Chen 0014, Yuankai Qi, Yumei Sun, Amin Beheshti, Jianlin Zhang 0001, Ming-Hsuan Yang 0001 |
AAAI | 3 |
| 2026 | MMCLIP: Cross-Modal Attention Masked Modelling for Medical Language-Image Pre-TrainingabstractVision-and-language pretraining (VLP) in medicine leverages contrastive learning on image-text pairs, often enhanced with masked modeling.However, existing methods face two challenges: difficulty reconstructing key pathological features due to limited data, and reliance on either paired or image-only datasets without combining both.To address this, we propose MMCLIP (Masked Medical Contrastive Language-Image Pre-training), which introduces two modules: AttMIM, masking image features highly correlated with text to improve reconstruction of fine medical details, and EntMLM, masking key medical entities in text and reconstructing them using visual cues.Furthermore, MMCLIP incorporates unpaired data through disease-kind prompts, achieving state-of-the-art performance in zero-shot and fine-tuning across five benchmarks.Code Biao Wu 0006, Yutong Xie 0001, Zeyu Zhang 0006, Vu Minh Hieu Phan, Qi Chen 0014, Ling Chen 0006, Qi Wu 0001 |
ACL (1) | 5 |
| 2026 | OUGS: Active View Selection via Object-aware Uncertainty Estimation in 3DGSabstractAbstract Recent advances in 3D Gaussian Splatting (3DGS) have achieved state‐of‐the‐art results for novel view synthesis. However, efficiently capturing high‐fidelity reconstructions of specific objects within complex scenes remains a significant challenge. A key limitation of existing active reconstruction methods is their reliance on scene‐level uncertainty metrics, which are often biased by irrelevant background clutter and lead to inefficient view selection for object‐centric tasks. We present OUGS, a novel framework that addresses this challenge with a more principled, physically‐grounded uncertainty formulation for 3DGS. Our core innovation is to derive uncertainty directly from the explicit physical parameters of the 3D Gaussian primitives (e.g., position, scale, rotation). By propagating the covariance of these parameters through the rendering Jacobian, we establish a highly interpretable uncertainty model. This foundation allows us to then seamlessly integrate semantic segmentation masks to produce a targeted, object‐aware uncertainty score that effectively disentangles the object from its environment. This allows for a more effective active view selection strategy that prioritizes views critical to improving object fidelity. Experimental evaluations on public datasets demonstrate that our approach significantly improves the efficiency of the 3DGS reconstruction process and achieves higher quality for targeted objects compared to existing state‐of‐the‐art methods, while also serving as a robust uncertainty estimator for the global scene. Haiyi Li, Qi Chen 0014, Denis Kalkofen, Hsiang-Ting Chen |
Comput. Graph. Forum | 2 |
| 2026 | Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
Zhuo Tao, Liang Li 0003, Qi Chen 0014, Yunbin Tu, Zhengjun Zha, Amin Beheshti, Qingming Huang, Yuankai Qi, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | Multi-contrast low-field MRI acceleration with k-space progressive learning and image-space hybrid attention fusion
Xiaohan Xing, Qi Chen 0014, Lequan Yu, Lingting Zhu, Lei Xing 0001, Lianli Liu |
Medical Image Anal. | 2 |
| 2026 | A comprehensive analysis of Mamba for 3D volumetric medical image segmentation
Chaohan Wang, Yutong Xie 0001, Qi Chen 0014, Yuyin Zhou, Qi Wu 0001 |
Pattern Recognit. | 3 |
| 2026 | Text-Driven Tumor SynthesisabstractTumor synthesis can generate challenging cases that AI often misses or over-detects. Training on these cases improves AI performance. However, most existing synthesis methods are either unconditional- generating images from random variables-or conditioned only on tumor shape. As a result, they lack control over clinically important tumor characteristics, such as texture, heterogeneity, boundary, and pathology. The generated tumors are therefore overly similar or duplicates of existing training cases, failing to effectively address AI's weaknesses. We propose a new text-driven tumor synthesis approach, termed TextoMorph, that provides textual control over tumor characteristics in conjunction with mask control. This approach is particularly beneficial for examples that confuse the AI the most, such as early tumor detection (improving Sensitivity by + 6.5%), tumor segmentation for precise radiotherapy (improving NSD by + 3.1%), and classification between benign and malignant tumors (improving Sensitivity by + 8.2%). By incorporating text mined from radiology reports into the synthesis process, we increase the variability and controllability of the synthetic tumors to target AI's failure cases more precisely. Moreover, TextoMorph uses contrastive learning across different texts and CT scans, significantly reducing dependence on scarce image-report pairs (only 141 pairs used in this study) by leveraging a large corpus of 34,035 radiology reports. Finally, we have developed rigorous tests to evaluate synthetic tumors, showing that our synthetic tumors is realistic and diverse in texture, heterogeneity, boundary, and pathology. Code and models are available at https://github.com/MrGiovanni/TextoMorph. Yi Shuai, Qi Chen 0014, Dong Yang 0005, Can Zhao 0001, Pedro R. A. S. Bassi, Daguang Xu, Kang Wang 0016, Yang Yang 0009, Alan L. Yuille, Zongwei Zhou |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-TuningabstractRecent advancements in Multimodal Large Language Models (MLLMs) have generated significant interest in their ability to autonomously interact with and interpret Graphical User Interfaces (GUIs). A major challenge in these systems is grounding—accurately identifying critical GUI components such as text or icons based on a GUI image and a corresponding text query. Traditionally, this task has relied on fine-tuning MLLMs with specialized training data to predict component locations directly. However, in this paper, we propose a novel Tuning-free Attention-driven Grounding (TAG) method that leverages the inherent attention patterns in pretrained MLLMs to accomplish this task without the need for additional fine-tuning. Our method involves identifying and aggregating attention maps from specific tokens within a carefully constructed query prompt. Applied to MiniCPM-Llama3-V 2.5, a state-of-the-art MLLM, our tuning-free approach achieves performance comparable to tuning-based methods, with notable success in text localization. Additionally, we demonstrate that our attention map-based grounding technique significantly outperforms direct localization predictions from MiniCPM-Llama3-V 2.5, highlighting the potential of using attention maps from pretrained MLLMs and paving the way for future innovations in this domain. Qi Chen 0014, Lei Wang 0001, Lingqiao Liu |
AAAI | 2 |
| 2025 | Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question AnsweringabstractKnowledge-Based visual question answering (KBVQA) separates image interpretation and knowledge retrieval into separate processes, motivated in part by the fact that they are very different tasks. In this paper, we transform the KB-VQA into linguistic question-answering tasks so that we can leverage the rich world knowledge and strong reasoning abilities of Large Language Models (LLMs). The caption-then-question approach to KBVQA has been effective, but relies on the captioning method to describe the detail required to answer every possible question. We propose instead a Question-Aware Captioner (QACap), which uses the question as guidance to extract correlated visual information from the image and generate a question-related caption. To train such a model, we utilize GPT-4 to build a corresponding high-quality question-aware caption dataset on top of existing KBVQA datasets. Extensive experiments demonstrate that our QACap model and dataset significantly improve KBVQA performance. Our method, QA-Cap, achieves 68.2% accuracy on the OKVQA validation set, 73.4% on the direct-answer part of the A-OKVQA validation set, and 74.8% on the multiple-choice part, all setting new SOTA benchmarks. Zhuo Tao, Qi Chen 0014, Liang Li 0003, Yuankai Qi, Anton van den Hengel, Qingming Huang |
CVPR | 3 |
| 2025 | Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
Qi Chen 0014, Xinze Zhou, Hao Chen 0011, Zekun Jiang, Ziyan Huang, Dexin Yu, Junjun He, Yefeng Zheng 0001, Ling Shao 0001, Alan L. Yuille, Zongwei Zhou |
ICCV | 1 |
| 2025 | Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual GroundingabstractVisual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features corresponding to textual descriptions, improving model transparency and trustworthiness for wider adoption of deep learning models in clinical practice. Current models struggle to associate textual descriptions with disease regions due to inefficient attention mechanisms and a lack of fine-grained token representations. In this paper, we empirically demonstrate two key observations. First, current VLMs assign high norms to background tokens, diverting the model's attention from regions of disease. Second, the global tokens used for cross-modal learning are not representative of local disease tokens. This hampers identifying correlations between the text and disease tokens. To address this, we introduce simple, yet effective Disease-Aware Prompting (DAP) process, which uses the explainability map of a VLM to identify the appropriate image features. This simple strategy amplifies disease-relevant regions while suppressing background interference. Without any additional pixel-level annotations, DAP improves visual grounding accuracy by 20.74% compared to state-of-the-art methods across three major chest X-ray datasets. Ta Duc Huy, Duy Anh Huynh, Yutong Xie 0001, Yuankai Qi, Qi Chen 0014, Phi-Le Nguyen, Sen Kim Tran, Son Lam Phung, Anton van den Hengel, Zhibin Liao, Minh-Son To, Johan Verjans, Vu Minh Hieu Phan |
ICCV | 5 |
| 2025 | OVG-HQ: Online Video Grounding with Hybrid-Modal QueriesabstractVideo grounding (VG) task focuses on locating specific moments in a video based on a query, usually in text form. However, traditional VG struggles with some scenarios like streaming video or queries using visual cues. To fill this gap, we present a new task named Online Video Grounding with Hybrid-modal Queries (OVG-HQ), which enables online segment localization using text, images, video segments, and their combinations. This task poses two new challenges: limited context in online settings and modality imbalance during training, where dominant modalities overshadow weaker ones. To address these, we propose OVG-HQ-Unify, a unified framework featuring a Parametric Memory Block (PMB) that retain previously learned knowledge to enhance current decision and a cross-modal distillation strategy that guides the learning of non-dominant modalities. This design enables a single model to effectively handle hybrid-modal queries. Due to the lack of suitable datasets, we construct QVHighlights-Unify, an expanded dataset with multi-modal queries. Besides, since offline metrics overlook prediction timeliness, we adapt them to the online setting, introducing oR@n, IoU=m, and online mean Average Precision (omAP) to evaluate both accuracy and efficiency. Experiments show that our OVG-HQ-Unify outperforms existing models, offering a robust solution for online, hybrid-modal video grounding. Source code and datasets are available at https://github.com/maojiaqi2324/OVG-HQ. Runhao Zeng, Jiaqi Mao, Minghao Lai, Minh Hieu Phan, Yanjie Dong 0003, Wei Wang 0077, Qi Chen 0014, Xiping Hu |
ICCV | 7 |
| 2025 | Localizing Before Answering: A Benchmark for Grounded Medical Visual Question AnsweringabstractMedical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in current medical LMMs: instead of analyzing relevant pathological regions, they often rely on linguistic patterns or attend to irrelevant image areas when responding to disease-related queries. To address this, we introduce HEAL-MedVQA (Hallucination Evaluation via Localization MedVQA), a comprehensive benchmark designed to evaluate LMMs' localization abilities and hallucination robustness. HEAL-MedVQA features (i) two innovative evaluation protocols to assess visual and textual shortcut learning, and (ii) a dataset of 67K VQA pairs, with doctor-annotated anatomical segmentation masks for pathological regions. To improve visual reasoning, we propose the Localize-before-Answer (LobA) framework, which trains LMMs to localize target regions of interest and self-prompt to emphasize segmented pathological areas, generating grounded and reliable answers. Experimental results demonstrate that our approach significantly outperforms state-of-the-art biomedical LMMs on the challenging HEAL-MedVQA benchmark, advancing robustness in medical VQA. Minh Khoi Ho, Ta Duc Huy, Thanh Tam Nguyen, Qi Chen 0014, Kumar Rav, Quy Duong Dang, Satwik Ramchandre, Son Lam Phung, Zhibin Liao, Minh-Son To, Johan Verjans, Phi-Le Nguyen, Vu Minh Hieu Phan |
IJCAI | 5 |
| 2025 | Controllable Image Synthesis Workflow for Enhancing Cervical Cell Detection
Yihuang Hu, Qi Chen 0014, Linbo Liao, Weiping Lin, Huisi Wu, Liansheng Wang 0002 |
MICCAI (13) | 2 |
| 2025 | PedCLIP: A Vision-Language Model for Pediatric X-Rays with Mixture of Body Part Experts
Ta Duc Huy, Abin Shoby, Sen Kim Tran, Yutong Xie 0001, Qi Chen 0014, Phi-Le Nguyen, Akshay Gole, Lingqiao Liu, Antonios Perperidis, Mark Friswell, Rebecca Linke, Andrea Glynn, Minh-Son To, Anton van den Hengel, Johan Verjans, Zhibin Liao, Minh Hieu Phan |
MICCAI (5) | 5 |
| 2025 | CausalMVC: Causal Content-Style Representation Learning for Deep Multi-View Clustering
Shifeng Bao, Zhe Xue, Qi Chen 0014, Shilong Ou, Amin Beheshti, Quan Z. Sheng, Anton van den Hengel, Yuankai Qi |
ACM Multimedia | 3 |
| 2025 | PanTS: The Pancreatic Tumor Segmentation DatasetabstractPanTS is a large-scale, multi-institutional dataset curated to advance research in pancreatic CT analysis. It contains 36,390 CT scans from 145 medical centers, with expert-validated, voxel-wise annotations of over 993,000 anatomical structures, covering pancreatic tumors, pancreas head, body, and tail, and 24 surrounding anatomical structures such as vascular/skeletal structures and abdominal/thoracic organs. Each scan includes metadata such as patient age, sex, diagnosis, contrast phase, in-plane spacing, slice thickness, etc. AI models trained on PanTS achieve significantly better performance in pancreatic tumor detection, localization, and segmentation than those trained on existing public datasets. Our analysis indicates that these gains are directly attributable to the 16× larger-scale tumor annotations and indirectly supported by the 24 additional surrounding anatomical structures. As the largest and most comprehensive resource of its kind, PanTS offers a new benchmark for developing and evaluating AI models in pancreatic CT analysis. Xinze Zhou, Qi Chen 0014, Pedro R. A. S. Bassi, Xiaoxi Chen, Zheren Zhu, Kang Wang 0016, Yang Yang 0009, Yucheng Tang, Daguang Xu, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 3 |
| 2025 | Are Pixel-Wise Metrics Reliable for Computerized Tomography Reconstruction?abstractWidely adopted evaluation metrics for sparse-view CT reconstruction, such as Structural Similarity Index Measure and Peak Signal-to-Noise Ratio, prioritize pixel-wise fidelity but often fail to capture the completeness of critical anatomical structures, particularly small or thin regions that are easily missed. To address this limitation, we propose a suite of novel anatomy-aware evaluation metrics designed to assess structural completeness across anatomical structures, including large organs, small organs, intestines, and vessels. Building on these metrics, we introduce CARE, a Completeness-Aware Reconstruction Enhancement framework that incorporates structural penalties during training to encourage anatomical preservation of significant structures. CARE is model-agnostic and can be seamlessly integrated into analytical, implicit, and generative methods.
When applied to these methods, CARE substantially improves structural completeness in CT reconstructions, achieving up to **32%** improvement for large organs, **22%** for small organs, **40%** for intestines, and **36%** for vessels. Chuntung Zhuang, Qi Chen 0014, Yuanhao Cai, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 4 |
| 2025 | MMR-Mamba: Multi-modal MRI reconstruction with Mamba and spatial-frequency information fusion
Lanqing Liu, Qi Chen 0014, Zhanli Hu, Xiaohan Xing, Harry Qin |
Medical Image Anal. | 3 |
| 2025 | CIT: Rethinking class-incremental semantic segmentation with a Class Independent TransformationabstractClass-incremental semantic segmentation (CSS) requires that a model learn to segment new classes without forgetting how to segment previous ones: this is typically achieved by distilling the current knowledge and incorporating the latest data. However, bypassing iterative distillation by directly transferring outputs of initial classes to the current learning task is not supported in existing class-specific CSS methods. Via Softmax, they enforce dependency between classes and adjust the output distribution at each learning step, resulting in a large probability distribution gap between initial and current tasks. We introduce a simple, yet effective Class Independent Transformation (CIT) that converts the outputs of existing semantic segmentation models into class-independent forms with negligible cost or performance loss. By utilizing class-independent predictions facilitated by CIT, we establish an accumulative distillation framework, ensuring equitable incorporation of all class information. We conduct extensive experiments on various segmentation architectures, including DeepLabV3, Mask2Former, and SegViTv2. Results from these experiments show minimal task forgetting across different datasets, with less than 5% for ADE20K in the most challenging 11 task configurations and less than 1% across all configurations for the PASCAL VOC 2012 dataset. • Softmax interdependency causes incremental forgetting in continual learning. • We introduce a class-independent transformation (CIT) to reduce forgetting. • CIT reformulates segmentation as class-agnostic, enhancing CSS training pipelines. • Our method significantly reduces forgetting on ADE20K compared to CSS baselines. • CIT achieves near-zero forgetting ( ≤ 1%) in Pascal-VOC 2012 settings. Jinchao Ge, Bowen Zhang 0009, Akide Liu, Vu Minh Hieu Phan, Qi Chen 0014, Yangyang Shu, Yang Zhao 0019 |
Pattern Recognit. | 5 |
| 2025 | Improving Video Moment Retrieval by Auxiliary Moment-Query Pairs With Hyper-InteractionabstractMost existing video moment retrieval (VMR) benchmark datasets face a common issue of sparse annotations-only a few moments being annotated. We argue that videos contain a broader range of meaningful moments that, if leveraged, could significantly enhance performance. Existing methods typically follow a generate-then-select paradigm, focusing primarily on generating moment-query pairs while neglecting the crucial aspect of selection. In this paper, we propose a new method, HyperAux, to yield auxiliary moment-query pairs by modeling the multi-modal hyper-interaction between video and language. Specifically, given a set of candidate moment-query pairs from a video, we construct a hypergraph with multiple hyperedges, each corresponding to a moment-query pair. Unlike traditional graphs where each edge connects only two nodes (frames or queries), each hyperedge connects multiple nodes, including all frames within a moment, semantically related frames outside the moment, and an input query. This design allows us to consider the frames within a moment as a whole, rather than modeling individual frame-query relationships separately. More importantly, constructing the relationships among all moment-query pairs within a video into a large hypergraph facilitates selecting higher-quality data from such pairs. On this hypergraph, we employ a hypergraph neural network to aggregate node information, update the hyperedge, and propagate video-language hyper-interactions to each connected node, resulting in context-aware node representations. This enables us to use node relevance to select high-quality moment-query pairs and refine the moments’ boundaries. We also exploit the discrepancy in semantic matching within and outside moments to construct a loss function for training the HGNN without human annotations. Our auxiliary data enhances the performance of twelve VMR models under fully-supervised, weakly-supervised, and zero-shot settings across three widely used VMR datasets: ActivityNet Captions, Charades-STA, and QVHighlights. We will release the source code and models publicly. Runhao Zeng, Yishen Zhuo, Yunjin Yang, Huisi Wu, Qi Chen 0014, Xiping Hu, Victor C. M. Leung |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | WebVLN: Vision-and-Language Navigation on WebsitesabstractVision-and-Language Navigation (VLN) task aims to enable AI agents to accurately understand and follow natural language instructions to navigate through real-world environments, ultimately reaching specific target locations. We recognise a promising opportunity to extend VLN to a comparable navigation task that holds substantial significance in our daily lives, albeit within the virtual realm: navigating websites on the Internet. This paper proposes a new task named Vision-and-Language Navigation on Websites (WebVLN), where we use question-based instructions to train an agent, emulating how users naturally browse websites. Unlike the existing VLN task that only pays attention to vision and instruction (language), the WebVLN agent further considers underlying web-specific content like HTML, which could not be seen on the rendered web pages yet contain rich visual and textual information. Toward this goal, we contribute a dataset, WebVLN-v1, and introduce a novel approach called Website-aware VLN Network (WebVLN-Net), which is built upon the foundation of state-of-the-art VLN techniques. Experimental results show that WebVLN-Net outperforms current VLN and web-related navigation methods. We believe that the introduction of the newWebVLN task and its dataset will establish a new dimension within the VLN domain and contribute to the broader vision-and-language research community. Code is available at: https://github.com/WebVLN/WebVLN. Qi Chen 0014, Dileepa Pitawela, Chongyang Zhao 0003, Gengze Zhou, Hsiang-Ting Chen, Qi Wu 0001 |
AAAI | 1 |
| 2024 | Act Like a Radiologist: Radiology Report Generation Across Anatomical Regions
Qi Chen 0014, Yutong Xie 0001, Biao Wu 0006, Minh-Son To, Xiaojun Chang, Qi Wu 0001 |
ACCV (6) | 1 |
| 2024 | Towards Generalizable Tumor SynthesisabstractTumor synthesis enables the creation of artificial tumors in medical images, facilitating the training of AI models for tumor detection and segmentation. However, success in tumor synthesis hinges on creating visually realistic tumors that are generalizable across multiple organs and, furthermore, the resulting AI models being capable of detecting real tumors in images sourced from different domains (e.g., hospitals). This paper made a progressive stride toward generalizable tumor synthesis by leveraging a critical observation: early-stage tumors (< 2cm) tend to have similar imaging characteristics in computed tomography (CT), whether they originate in the liver, pancreas, or kidneys. We have ascertained that generative AI models, e.g., Diffusion Models, can create realistic tumors generalized to a range of organs even when trained on a limited number of tumor examples from only one organ. Moreover, we have shown that AI models trained on these synthetic tumors can be generalized to detect and segment real tumors from CT volumes, encompassing a broad spectrum of patient demographics, imaging protocols, and healthcare facilities. Qi Chen 0014, Xiaoxi Chen, Haorui Song, Zhiwei Xiong, Alan L. Yuille, Chen Wei 0002, Zongwei Zhou |
CVPR | 1 |
| 2024 | G-NeRF: Geometry-enhanced Novel View Synthesis from Single-View ImagesabstractNovel view synthesis aims to generate new view images of a given view image collection. Recent attempts address this problem relying on 3D geometry priors (e.g., shapes, sizes, and positions) learned from multi-view images. However, such methods encounter the following limitations: 1) they require a set of multi-view images as training data for a specific scene (e.g., face, car or chair), which is often un-available in many real-world scenarios; 2) they fail to ex-tract the geometry priors from single-view images due to the lack of multi-view supervision. In this paper, we pro-pose a Geometry-enhanced NeRF (G-NeRF), which seeks to enhance the geometry priors by a geometry-guided multi-view synthesis approach, followed by a depth-aware training. In the synthesis process, inspired that existing 3D GAN models can unconditionally synthesize high-fidelity multi-view images, we seek to adopt off-the-shelf 3D GAN models, such as EG3D, as a free source to provide geometry priors through synthesizing multi-view data. Simultaneously, to further improve the geometry quality of the synthetic data, we introduce a truncation method to effectively sample la-tent codes within 3D GAN models. To tackle the absence of multi-view supervision for single-view images, we design the depth-aware training approach, incorporating a depth-aware discriminator to guide geometry priors through depth maps. Experiments demonstrate the effectiveness of our method in terms of both qualitative and quantitative results. Zixiong Huang, Qi Chen 0014, Naizhou Wang, Qi Wu 0001, Mingkui Tan |
CVPR | 2 |
| 2024 | PairAug: What Can Augmented Image-Text Pairs Do for Radiology?abstractCurrent vision-language pre-training (VLP) methodologies predominantly depend on paired image-text datasets, a resource that is challenging to acquire in radiology due to privacy considerations and labelling complexities. Data augmentation provides a practical solution to overcome the issue of data scarcity, however, most augmentation methods exhibit a limited focus, prioritising either image or text augmentation exclusively. Acknowledging this limitation, our objective is to devise a framework capable of concurrently augmenting medical image and text data. We design a Pairwise Augmentation (PairAug) approach that contains an Inter-patient Augmentation (InterAug) branch and an Intra-patient Augmentation (IntraAug) branch. Specifically, the InterAug branch of our approach generates radiology images using synthesised yet plausible reports derived from a Large Language Model (LLM). The generated pairs can be considered a collection of new patient cases since they are artificially created and may not exist in the original dataset. In contrast, the IntraAug branch uses newly generated reports to manipulate images. This process allows us to create new paired data for each individual with diverse medical conditions. Our extensive experiments on various downstream tasks covering medical image classification zero-shot and fine-tuning analysis demonstrate that our PairAug, concurrently expanding both image and text data, substantially outperforms image-/text-only expansion baselines and advanced medical VLP baselines. Our code is released at https://github.com/YtongXie/PairAug. Yutong Xie 0001, Qi Chen 0014, Sinuo Wang, Minh-Son To, Iris Lee, Ee Win Khoo, Kerolos Hendy, Daniel Koh, Yong Xia 0001, Qi Wu 0001 |
CVPR | 2 |
| 2024 | Learning Multiscale Consistency for Self-Supervised Electron Microscopy Instance SegmentationabstractElectron microscopy (EM) images are notoriously challenging to segment due to their complex structures and lack of effective annotations. Fortunately, large-scale self-supervised pretraining offers a promising solution by allowing us to acquire prior knowledge of cell and subcellular tissue structures, which can significantly improve EM instance segmentation results. However, most existing pretraining methods fail to capture the crucial local information that is essential for EM images, instead focusing only on high-level semantic information. In this paper, we propose a novel pretraining framework that leverages multiscale visual representations to adapt to the complex structures of EM images. Our framework achieves instance-level alignment by maximizing the consistency between strongly and weakly augmented images, while also incorporating a cross-attention mechanism to match multiscale features and encode more low-level information into high-level semantics. Most importantly, our approach employs multi-task optimization on the feature pyramid, enabling multiscale pixel restoration and feature comparison. We extensively pretrain our method on four large-scale EM datasets and demonstrate significant gains on neuron and mitochondria segmentation tasks. Code is available at https://github.com/ydchen0806/MS-Con-EM-Seg. Yinda Chen, Wei Huang 0036, Xiaoyu Liu 0006, Shiyu Deng, Qi Chen 0014, Zhiwei Xiong |
ICASSP | 5 |
| 2024 | Accelerated Multi-contrast MRI Reconstruction via Frequency and Spatial Mutual Learning
Qi Chen 0014, Xiaohan Xing, Zhen Chen 0013, Zhiwei Xiong |
MICCAI (7) | 1 |
| 2024 | Towards Lightweight Super-Resolution With Dual Regression LearningabstractDeep neural networks have exhibited remarkable performance in image super-resolution (SR) tasks by learning a mapping from low-resolution (LR) images to high-resolution (HR) images. However, the SR problem is typically an ill-posed problem and existing methods would come with several limitations. First, the possible mapping space of SR can be extremely large since there may exist many different HR images that can be super-resolved from the same LR image. As a result, it is hard to directly learn a promising SR mapping from such a large space. Second, it is often inevitable to develop very large models with extremely high computational cost to yield promising SR performance. In practice, one can use model compression techniques to obtain compact models by reducing model redundancy. Nevertheless, it is hard for existing model compression methods to accurately identify the redundant components due to the extremely large SR mapping space. To alleviate the first challenge, we propose a dual regression learning scheme to reduce the space of possible SR mappings. Specifically, in addition to the mapping from LR to HR images, we learn an additional dual regression mapping to estimate the downsampling kernel and reconstruct LR images. In this way, the dual mapping acts as a constraint to reduce the space of possible mappings. To address the second challenge, we propose a dual regression compression (DRC) method to reduce model redundancy in both layer-level and channel-level based on channel pruning. Specifically, we first develop a channel number search method that minimizes the dual regression loss to determine the redundancy of each layer. Given the searched channel numbers, we further exploit the dual regression manner to evaluate the importance of channels and prune the redundant ones. Extensive experiments show the effectiveness of our method in obtaining accurate and efficient SR models. Mingkui Tan, Zeshuai Deng, Jingdong Wang 0001, Qi Chen 0014, Jiezhang Cao, Yanwu Xu 0001, Jian Chen 0011 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | WASPSYN: A Challenge for Domain Adaptive Synapse Detection in Microwasp Brain ConnectomesabstractThe size of image volumes in connectomics studies now reaches terabyte and often petabyte scales with a great diversity of appearance due to different sample preparation procedures. However, manual annotation of neuronal structures (e.g., synapses) in these huge image volumes is time-consuming, leading to limited labeled training data often smaller than 0.001% of the large-scale image volumes in application. Methods that can utilize in-domain labeled data and generalize to out-of-domain unlabeled data are in urgent need. Although many domain adaptation approaches are proposed to address such issues in the natural image domain, few of them have been evaluated on connectomics data due to a lack of domain adaptation benchmarks. Therefore, to enable developments of domain adaptive synapse detection methods for large-scale connectomics applications, we annotated 14 image volumes from a biologically diverse set of Megaphragma viggianii brain regions originating from three different whole-brain datasets and organized the WASPSYN challenge at ISBI 2023. The annotations include coordinates of pre-synapses and post-synapses in the 3D space, together with their one-to-many connectivity information. This paper describes the dataset, the tasks, the proposed baseline, the evaluation method, and the results of the challenge. Limitations of the challenge and the impact on neuroscience research are also discussed. The challenge is and will continue to be available at https://codalab.lisn.upsaclay.fr/competitions/9169. Successful algorithms that emerge from our challenge may potentially revolutionize real-world connectomics research and further the cause that aims to unravel the complexity of brain structure and function. Yicong Li 0002, Wanhua Li 0001, Qi Chen 0014, Wei Huang 0036, Yuda Zou, Kazunori Shinomiya, Pat Gunn, Nishika Gupta, Alexey Polilov, Yongchao Xu, Yueyi Zhang 0001, Zhiwei Xiong, Hanspeter Pfister, Donglai Wei 0001, Jingpeng Wu |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Prompt Switch: Efficient CLIP Adaptation for Text-Video RetrievalabstractIn text-video retrieval, recent works have benefited from the powerful learning capabilities of pre-trained text-image foundation models (e.g., CLIP) by adapting them to the video domain. A critical problem for them is how to effectively capture the rich semantics inside the video using the image encoder of CLIP. To tackle this, state-of-the-art methods adopt complex cross-modal modeling techniques to fuse the text information into video frame representations, which, however, incurs severe efficiency issues in large-scale retrieval systems as the video representations must be recomputed online for every text query. In this paper, we discard this problematic cross-modal fusion process and aim to learn semantically-enhanced representations purely from the video, so that the video representations can be computed offline and reused for different texts. Concretely, we first introduce a spatial-temporal "Prompt Cube" into the CLIP image encoder and iteratively switch it within the encoder layers to efficiently incorporate the global video semantics into frame representations. We then propose to apply an auxiliary video captioning objective to train the frame representations, which facilitates the learning of detailed video semantics by providing fine-grained guidance in the semantic space. With a naive temporal fusion strategy (i.e., mean-pooling) on the enhanced frame representations, we obtain state-of-the-art performances on three benchmark datasets, i.e., MSR-VTT, MSVD, and LSMDC. Chaorui Deng, Qi Chen 0014, Pengda Qin, Da Chen 0003, Qi Wu 0001 |
ICCV | 2 |
| 2023 | Self-Supervised Neuron Segmentation with Multi-Agent Reinforcement LearningabstractThe performance of existing supervised neuron segmentation methods is highly dependent on the number of accurate annotations, especially when applied to large scale electron microscopy (EM) data. By extracting semantic information from unlabeled data, self-supervised methods can improve the performance of downstream tasks, among which the mask image model (MIM) has been widely used due to its simplicity and effectiveness in recovering original information from masked images. However, due to the high degree of structural locality in EM images, as well as the existence of considerable noise, many voxels contain little discriminative information, making MIM pretraining inefficient on the neuron segmentation task. To overcome this challenge, we propose a decision-based MIM that utilizes reinforcement learning (RL) to automatically search for optimal image masking ratio and masking strategy. Due to the vast exploration space, using single-agent RL for voxel prediction is impractical. Therefore, we treat each input patch as an agent with a shared behavior policy, allowing for multi-agent collaboration. Furthermore, this multi-agent model can capture dependencies between voxels, which is beneficial for the downstream segmentation task. Experiments conducted on representative EM datasets demonstrate that our approach has a significant advantage over alternative self-supervised methods on the task of neuron segmentation. Code is available at https://github.com/ydchen0806/dbMiM. Yinda Chen, Wei Huang 0036, Shenglong Zhou 0002, Qi Chen 0014, Zhiwei Xiong |
IJCAI | 4 |
| 2022 | V2C: Visual Voice CloningabstractExisting Voice Cloning (VC) tasks aim to convert a para-graph text to a speech with desired voice specified by a ref-erence audio. This has significantly boosted the development of artificial speech applications. However, there also exist many scenarios that cannot be well reflected by these VC tasks, such as movie dubbing, which requires the speech to be with emotions consistent with the movie plots. To fill this gap, in this work we propose a new task named Vi-sual Voice Cloning (V2C), which seeks to convert a para-graph of text to a speech with both desired voice speci-fied by a reference audio and desired emotion specified by a reference video. To facilitate research in this field, we construct a dataset, V2C-Animation, and propose a strong baseline based on existing state-of-the-art (SoTA) VC techniques. Our dataset contains 10,217 animated movie clips covering a large variety of genres (e.g., Comedy, Fantasy) and emotions (e.g., happy, sad). We further design a set of evaluation metrics, named MCD-DTW-SL, which help eval-uate the similarity between ground-truth speeches and the synthesised ones. Extensive experimental results show that even SoTA VC methods cannot generate satisfying speeches for our V2C task. We hope the proposed new task together with the constructed dataset and evaluation metric will fa-cilitate the research in the field of voice cloning and broader vision-and-language community. Source code and dataset will be released in https://github.com/chenqi008/V2C. Qi Chen 0014, Mingkui Tan, Yuankai Qi, Jiaqiu Zhou, Yuanqing Li 0001, Qi Wu 0001 |
CVPR | 1 |
| 2022 | Mask Rearranging Data Augmentation for 3D Mitochondria Segmentation
Qi Chen 0014, Mingxing Li 0003, Jiacheng Li 0004, Bo Hu 0014, Zhiwei Xiong |
MICCAI (4) | 1 |
| 2022 | Learning Distinct and Representative Modes for Image CaptioningabstractOver the years, state-of-the-art (SoTA) image captioning methods have achieved promising results on some evaluation metrics (e.g., CIDEr). However, recent findings show that the captions generated by these methods tend to be biased toward the "average" caption that only captures the most general mode (a.k.a, language pattern) in the training corpus, i.e., the so-called mode collapse problem. Affected by it, the generated captions are limited in diversity and usually less informative than natural image descriptions made by humans. In this paper, we seek to avoid this problem by proposing a Discrete Mode Learning (DML) paradigm for image captioning. Our innovative idea is to explore the rich modes in the training caption corpus to learn a set of "mode embeddings", and further use them to control the mode of the generated captions for existing image captioning models. Specifically, the proposed DML optimizes a dual architecture that consists of an image-conditioned discrete variational autoencoder (CdVAE) branch and a mode-conditioned image captioning (MIC) branch. The CdVAE branch maps each image caption to one of the mode embeddings stored in a learned codebook, and is trained with a pure non-autoregressive generation objective to make the modes distinct and representative. The MIC branch can be simply modified from an existing image captioning model, where the mode embedding is added to the original word embeddings as the control signal. In the experiments, we apply the proposed DML to two widely used image captioning models, Transformer and AoANet. The results show that the learned mode embedding successfully facilitates these models to generate high-quality image captions with different modes, further leading to better performance for both diversity and quality on the MS COCO dataset. Qi Chen 0014, Chaorui Deng, Qi Wu 0001 |
NeurIPS | 1 |
| 2022 | Towards Accurate and Compact Architectures via Neural Architecture TransformerabstractDesigning effective architectures is one of the key factors behind the success of deep neural networks. Existing deep architectures are either manually designed or automatically searched by some Neural Architecture Search (NAS) methods. However, even a well-designed/searched architecture may still contain many nonsignificant or redundant modules/operations (e.g., some intermediate convolution or pooling layers). Such redundancy may not only incur substantial memory consumption and computational cost but also deteriorate the performance. Thus, it is necessary to optimize the operations inside an architecture to improve the performance without introducing extra computational cost. To this end, we have proposed a Neural Architecture Transformer (NAT) method which casts the optimization problem into a Markov Decision Process (MDP) and seeks to replace the redundant operations with more efficient operations, such as skip or null connection. Note that NAT only considers a small number of possible replacements/transitions and thus comes with a limited search space. As a result, such a small search space may hamper the performance of architecture optimization. To address this issue, we propose a Neural Architecture Transformer++ (NAT++) method which further enlarges the set of candidate transitions to improve the performance of architecture optimization. Specifically, we present a two-level transition rule to obtain valid transitions, i.e., allowing operations to have more efficient types (e.g., convolution → separable convolution) or smaller kernel sizes (e.g., 5×5 → 3×3). Note that different operations may have different valid transitions. We further propose a Binary-Masked Softmax (BMSoftmax) layer to omit the possible invalid transitions. Last, based on the MDP formulation, we apply policy gradient to learn an optimal policy, which will be used to infer the optimized architectures. Extensive experiments show that the transformed architectures significantly outperform both their original counterparts and the architectures optimized by existing methods. Mingkui Tan, Qi Chen 0014, Jian Chen 0011, Peilin Zhao, Junzhou Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Contrastive Neural Architecture Search With Neural Architecture ComparatorsabstractOne of the key steps in Neural Architecture Search (NAS) is to estimate the performance of candidate architectures. Existing methods either directly use the validation performance or learn a predictor to estimate the performance. However, these methods can be either computationally expensive or very inaccurate, which may severely affect the search efficiency and performance. Moreover, as it is very difficult to annotate architectures with accurate performance on specific tasks, learning a promising performance predictor is often non-trivial due to the lack of labeled data. In this paper, we argue that it may not be necessary to estimate the absolute performance for NAS. On the contrary, we may need only to understand whether an architecture is better than a baseline one. However, how to exploit this comparison information as the reward and how to well use the limited labeled data remains two great challenges. In this paper, we propose a novel Contrastive Neural Architecture Search (CTNAS) method which performs architecture search by taking the comparison results between architectures as the reward. Specifically, we design and learn a Neural Architecture Comparator (NAC) to compute the probability of candidate architectures being better than a baseline one. Moreover, we present a baseline updating scheme to improve the baseline iteratively in a curriculum learning manner. More critically, we theoretically show that learning NAC is equivalent to optimizing the ranking over architectures. Extensive experiments in three search spaces demonstrate the superiority of our CTNAS over existing methods. Yaofo Chen, Qi Chen 0014, Minli Li, Wei Zeng 0006, Yaowei Wang 0001, Mingkui Tan |
CVPR | 3 |
| 2021 | R-GAN: Exploring Human-like Way for Reasonable Text-to-Image Synthesis via Generative Adversarial NetworksabstractDespite recent significant progress on generative models, context-rich text-to-image synthesis depicting multiple complex objects is still non-trivial. The main challenges lie in the ambiguous semantic of a complex description and the intricate scene of an image with various objects, different positional relationship and diverse appearances. To address these challenges, we propose R-GAN, which can generate reasonable images according to the given text in a human-like way. Specifically, just like humans will first find and settle the essential elements to create a simple sketch, we first capture a monolithic-structural text representation by building a scene graph to find the essential semantic elements. Then, based on this representation, we design a bounding box generator to estimate the layout with position and size of target objects, and a following shape generator, which draws a fine-detailed shape for each object. Different from previous work only generating coarse shapes blindly, we introduce a coarse-to-fine shape generator based on a shape knowledge base. At last, to finish the final image synthesis, we propose a multi-modal geometry-aware spatially-adaptive generator conditioned on the monolithic-structural text representation and the geometry-aware map of the shapes. Extensive experiments on the real-world dataset MSCOCO show the superiority of our method in terms of both quantitative and qualitative metrics. Yanyuan Qiao, Qi Chen 0014, Chaorui Deng, Yuankai Qi, Mingkui Tan, Xincheng Ren, Qi Wu 0001 |
ACM Multimedia | 2 |
| 2020 | Modular Graph Attention Network for Complex Visual Relational Reasoning
Yihan Zheng, Zhiquan Wen, Mingkui Tan, Runhao Zeng, Qi Chen 0014, Yaowei Wang 0001, Qi Wu 0001 |
ACCV (6) | 5 |
| 2020 | Intelligent Home 3D: Automatic 3D-House Design From Linguistic Descriptions OnlyabstractHome design is a complex task that normally requires architects to finish with their professional skills and tools. It will be fascinating that if one can produce a house plan intuitively without knowing much knowledge about home design and experience of using complex designing tools, for example, via natural language. In this paper, we formulate it as a language conditioned visual content generation problem that is further divided into a floor plan generation and an interior texture (such as floor and wall) synthesis task. The only control signal of the generation process is the linguistic expression given by users that describe the house details. To this end, we propose a House Plan Generative Model (HPGM) that first translates the language input to a structural graph representation and then predicts the layout of rooms with a Graph Conditioned Layout Prediction Network (GC-LPN) and generates the interior texture with a Language Conditioned Texture GAN (LCT-GAN). With some post-processing, the final product of this task is a 3D house model. To train and evaluate our model, we build the first Text-to-3D House Model dataset, which will be released at: https:// hidden-link-for-submission. Qi Chen 0014, Qi Wu 0001, Rui Tang 0015, Yuhan Wang 0001, Mingkui Tan |
CVPR | 1 |
| 2020 | Closed-Loop Matters: Dual Regression Networks for Single Image Super-ResolutionabstractDeep neural networks have exhibited promising performance in image super-resolution (SR) by learning a nonlinear mapping function from low-resolution (LR) images to high-resolution (HR) images. However, there are two underlying limitations to existing SR methods. First, learning the mapping function from LR to HR images is typically an ill-posed problem, because there exist infinite HR images that can be downsampled to the same LR image. As a result, the space of the possible functions can be extremely large, which makes it hard to find a good solution. Second, the paired LR-HR data may be unavailable in real-world applications and the underlying degradation method is often unknown. For such a more general case, existing SR models often incur the adaptation problem and yield poor performance. To address the above issues, we propose a dual regression scheme by introducing an additional constraint on LR data to reduce the space of the possible functions. Specifically, besides the mapping from LR to HR images, we learn an additional dual regression mapping estimates the down-sampling kernel and reconstruct LR images, which forms a closed-loop to provide additional supervision. More critically, since the dual regression process does not depend on HR images, we can directly learn from LR images. In this sense, we can easily adapt SR models to real-world data, e.g., raw video frames from YouTube. Extensive experiments with paired training data and unpaired real-world data demonstrate our superiority over existing methods. Jian Chen 0011, Jingdong Wang 0001, Qi Chen 0014, Jiezhang Cao, Zeshuai Deng, Yanwu Xu 0001, Mingkui Tan |
CVPR | 4 |
| 2020 | Object as Hotspots: An Anchor-Free 3D Object Detection Approach via Firing of Hotspots
Qi Chen 0014, Lin Sun 0004, Zhixin Wang, Kui Jia, Alan L. Yuille |
ECCV (21) | 1 |
| 2020 | Dynamic Extension Nets for Few-shot Semantic SegmentationabstractSemantic segmentation requires a large amount of densely annotated data for training and may generalize poorly to novel categories. In real-world applications, we have an urgent need for few-shot semantic segmentation which aims to empower a model to handle unseen object categories with limited data. This task is non-trivial due to several challenges. First, it is difficult to extract the class-relevant information to handle the novel class as only a few samples are available. Second, since the image content can be very complex, the novel class information may be suppressed by the base categories due to limited data. Third, one may easily learn promising base classifiers based on a large amount of training data, but it is non-trivial to exploit the knowledge to train the novel classifiers. More critically, once a novel classifier is built, the output probability space will change. How to maintain the base classifiers and dynamically include the novel classifiers remains an open question. To address the above issues, we propose a Dynamic Extension Network (DENet) in which we dynamically construct and maintain a classifier for the novel class by leveraging the knowledge from the base classes and the information from novel data. More importantly, to overcome the information suppression issue, we design a Guided Attention Module (GAM), which can be plugged into any framework to help learn class-relevant features. Last, rather than directly train the model with limited data, we propose a dynamic extension training algorithm to predict the weights of novel classifiers, which is able to exploit the knowledge of base classifiers by dynamically extending classes during training. The extensive experiments show that our proposed method achieves state-of-the-art performance on the PASCAL-5i and COCO-20i datasets. The source code is available at https://github.com/lizhaoliu-Lec/DENet. Lizhao Liu, Junyi Cao, Minqian Liu, Qi Chen 0014, Mingkui Tan |
ACM Multimedia | 5 |
| 2020 | Every View Counts: Cross-View Consistency in 3D Object Detection with Hybrid-Cylindrical-Spherical VoxelizationabstractRecent voxel-based 3D object detectors for autonomous vehicles learn point cloud representations either from bird eye view (BEV) or range view (RV, a.k.a. the perspective view). However, each view has its own strengths and weaknesses. In this paper, we present a novel framework to unify and leverage the benefits from both BEV and RV. The widely-used cuboid-shaped voxels in Cartesian coordinate system only benefit learning BEV feature map. Therefore, to enable learning both BEV and RV feature maps, we introduce Hybrid-Cylindrical-Spherical voxelization. Our findings show that simply adding detection on another view as auxiliary supervision will lead to poor performance. We proposed a pair of cross-view transformers to transform the feature maps into the other view and introduce cross-view consistency loss on them. Comprehensive experiments on the challenging NuScenes Dataset validate the effectiveness of our proposed method by virtue of joint optimization and complementary information on both views. Remarkably, our approach achieved mAP of 55.8%, outperforming all published approaches by at least 3% in overall performance and up to 16.5% in safety-crucial categories like cyclist. Qi Chen 0014, Lin Sun 0004, Ernest Cheung, Alan L. Yuille |
NeurIPS | 1 |
| 2020 | Scripted Video Generation With a Bottom-Up Generative Adversarial NetworkabstractGenerating videos given a text description (such as a script) is non-trivial due to the intrinsic complexity of image frames and the structure of videos. Although Generative Adversarial Networks (GANs) have been successfully applied to generate images conditioned on a natural language description, it is still very challenging to generate realistic videos in which the frames are required to follow both spatial and temporal coherence. In this paper, we propose a novel Bottom-up GAN (BoGAN) method for generating videos given a text description. To ensure the coherence of the generated frames and also make the whole video match the language descriptions semantically, we design a bottom-up optimisation mechanism to train BoGAN. Specifically, we devise a region-level loss via attention mechanism to preserve the local semantic alignment and draw details in different sub-regions of video conditioned on words which are most relevant to them. Moreover, to guarantee the matching between text and frame, we introduce a frame-level discriminator, which can also maintain the fidelity of each frame and the coherence across frames. Last, to ensure the global semantic alignment between whole video and given text, we apply a video-level discriminator. We evaluate the effectiveness of the proposed BoGAN on two synthetic datasets (i.e., SBMG and TBMG) and two real-world datasets (i.e., MSVD and KTH). Qi Chen 0014, Qi Wu 0001, Jian Chen 0011, Qingyao Wu, Anton van den Hengel, Mingkui Tan |
IEEE Trans. Image Process. | 1 |
| 2019 | NAT: Neural Architecture Transformer for Accurate and Compact ArchitecturesabstractDesigning effective architectures is one of the key factors behind the success of deep neural networks. Existing deep architectures are either manually designed or automatically searched by some Neural Architecture Search (NAS) methods. However, even a well-searched architecture may still contain many non-significant or redundant modules or operations (e.g., convolution or pooling), which may not only incur substantial memory consumption and computation cost but also deteriorate the performance. Thus, it is necessary to optimize the operations inside an architecture to improve the performance without introducing extra computation cost. Unfortunately, such a constrained optimization problem is NP-hard. To make the problem feasible, we cast the optimization problem into a Markov decision process (MDP) and seek to learn a Neural Architecture Transformer (NAT) to replace the redundant operations with the more computationally efficient ones (e.g., skip connection or directly removing the connection). Based on MDP, we learn NAT by exploiting reinforcement learning to obtain the optimization policies w.r.t. different architectures. To verify the effectiveness of the proposed strategies, we apply NAT on both hand-crafted architectures and NAS based architectures. Extensive experiments on two benchmark datasets, i.e., CIFAR-10 and ImageNet, demonstrate that the transformed architecture by NAT significantly outperforms both its original form and those architectures optimized by existing methods. Mingkui Tan, Qi Chen 0014, Jian Chen 0011, Peilin Zhao, Junzhou Huang |
NeurIPS | 4 |
| 2019 | Auto-Embedding Generative Adversarial Networks For High Resolution Image SynthesisabstractGenerating images via a generative adversarial network (GAN) has attracted much attention recently. However, most of the existing GAN-based methods can only produce lowresolution images of limited quality. Directly generating highresolution images using GANs is nontrivial, and often produces problematic images with incomplete objects. To address this issue, we develop a novel GAN called auto-embedding generative adversarial network, which simultaneously encodes the global structure features and captures the fine-grained details. In our network, we use an autoencoder to learn the intrinsic high-level structure of real images and design a novel denoiser network to provide photo-realistic details for the generated images. In the experiments, we are able to produce 512 × 512 images of promising quality directly from the input noise. The resultant images exhibit better perceptual photo-realism, that is, with sharper structure and richer details, than other baselines on several datasets, including Oxford-102 Flowers, Caltech-UCSD Birds (CUB), High-Quality Large-scale CelebFaces Attributes (CelebAHQ), Large-scale Scene Understanding (LSUN), and ImageNet. Qi Chen 0014, Jian Chen 0011, Qingyao Wu, Qinfeng Shi, Mingkui Tan |
IEEE Trans. Multim. | 2 |
| 2018 | UnrealStereo: Controlling Hazardous Factors to Analyze Stereo VisionabstractA reliable stereo algorithm is critical for many robotics applications. But textureless and specular regions can easily cause failure by making feature matching difficult. Understanding whether an algorithm is robust to these hazardous regions is important. Although many stereo benchmarks have been developed to evaluate performance, it is hard to quantify the effect of hazardous regions in real images because the location and severity of these regions are unknown. In this paper, we develop a synthetic image generation tool enabling to control hazardous factors, such as making objects more specular or transparent, to produce hazardous regions at different degrees. The densely controlled sampling strategy in virtual worlds enables to effectively stress test stereo algorithms by varying the types and degrees of the hazard. We generate a large synthetic image dataset with automatically computed hazardous regions and analyze algorithms on these regions. The observations from synthetic images are further validated by annotating hazardous regions in real-world datasets Middlebury and KITTI (which gives a sparse sampling of the hazards). Our synthetic image generation tool is based on a game engine Unreal Engine 4 and will be open-source along with the virtual scenes in our experiments. Many publicly available realistic game contents can be used by our tool to provide an enormous resource for development and evaluation of algorithms. Yi Zhang 0099, Weichao Qiu, Qi Chen 0014, Xiaolin Hu 0001, Alan L. Yuille |
3DV | 3 |
| 2018 | SampleAhead: Online Classifier-Sampler Communication for Learning from Synthesized Data
Qi Chen 0014, Weichao Qiu, Yi Zhang 0099, Lingxi Xie, Alan L. Yuille |
BMVC | 1 |
| 2018 | A unified framework with a benchmark dataset for surveillance event detection
Zhicheng Zhao 0001, Xuanchong Li, Xingzhong Du, Qi Chen 0014, Yanyun Zhao, Xiaojun Chang, Alex Hauptmann 0001 |
Neurocomputing | 4 |
| 2015 | Part-based deep network for pedestrian detection in surveillance videosabstractAccurate pedestrian detection in highly crowded surveillance videos is a challenging task, since the regions of pedestrians in the videos may be largely occluded by other pedestrians. In this paper, we propose an effective part-based deep network cascade (HsNet) to solve this problem. In this model, the part-based scheme effectively restrains the appearance variations of pedestrians caused by heavy occlusion. The deep network captures discriminative information of visible body parts. In addition, the cascade architecture enables very fast detection. We make experiments on one of the largest surveillance video dataset, namely TRECVid SED Pedestrian Dataset (SED-PD). It is shown that in highly crowded surveillance videos, our proposed method achieves very competitive performance compared with state-of-the-art methods. More importantly, our method is significantly faster. Qi Chen 0014, Wenhui Jiang 0001, Yanyun Zhao, Zhicheng Zhao 0001 |
VCIP | 1 |