Jinghao Zhou

dblp:69/5505 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Representation and self-supervised learning · 35% Robot manipulation · 26% Video understanding and tracking · 22%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder
0.812024
Exploring Target Representations for Masked Autoencoders · ICLR 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised visual representation learning
0.812024
Exploring Target Representations for Masked Autoencoders · ICLR 2024
Visual content generation and editing › style transfer
3d content stylization
0.812024
Scene-Conditional 3D Object Stylization and Composition · ECCV (68) 2024
Visual content generation and editing › 3d content creation
3d object composition
0.812024
Scene-Conditional 3D Object Stylization and Composition · ECCV (68) 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
non-contrastive learning
0.712023
Non-Contrastive Learning Meets Language-Image Pre-Training · CVPR 2023
Robotics › Robot manipulation › soft robotics
soft actuator
0.712023
A Restorable, Variable Stiffness Pneumatic Soft Gripper Based on Jamming of Strings of Beads · IEEE Trans. Robotics 2023
Robotics › Robot manipulation › grasping
soft gripper
0.712023
A Restorable, Variable Stiffness Pneumatic Soft Gripper Based on Jamming of Strings of Beads · IEEE Trans. Robotics 2023
Robotics › Robot manipulation › actuator design › variable stiffness
variable stiffness gripper
0.712023
A Restorable, Variable Stiffness Pneumatic Soft Gripper Based on Jamming of Strings of Beads · IEEE Trans. Robotics 2023
Computer vision › Vision and language
vision-language pretraining
0.712023
Non-Contrastive Learning Meets Language-Image Pre-Training · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling
0.612022
Image BERT Pre-training with Online Tokenizer · ICLR 2022
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.612022
Image BERT Pre-training with Online Tokenizer · ICLR 2022
Computer vision › Video understanding and tracking
object tracking
0.412020
Discriminative and Robust Online Learning for Siamese Visual Tracking · AAAI 2020
Computer vision › Video understanding and tracking › object tracking › learning-based tracking
online learning for tracking
0.412020
Discriminative and Robust Online Learning for Siamese Visual Tracking · AAAI 2020
Computer vision › Video understanding and tracking › object tracking › deep tracking
siamese tracking
0.412020
Discriminative and Robust Online Learning for Siamese Visual Tracking · AAAI 2020
Computer vision › Video understanding and tracking › object tracking › adaptive tracking
template update
0.412020
Discriminative and Robust Online Learning for Siamese Visual Tracking · AAAI 2020
Computer vision › Image recognition and object detection
object detection
0.112007
A boosting regression approach to medical anatomy detection · CVPR 2007

Methods — techniques the papers use, named apart from their topics

neural rendering · 0.8knowledge distillation · 0.8bootstrapped teacher · 0.8pneumatic actuation · 0.7particle jamming · 0.7non-contrastive learning · 0.7multi-task learning · 0.7contrastive learning · 0.7online tokenization · 0.6masked image modeling · 0.6attention mechanism · 0.4
YearPublicationVenuePosition
2024 Scene-Conditional 3D Object Stylization and Composition
Jinghao Zhou, Tomas Jakab, Philip Torr 0001, Christian Rupprecht 0001
ECCV (68)1
2024 Exploring Target Representations for Masked Autoencoders
abstract
Masked autoencoders have become popular training paradigms for self-supervised visual representation learning. These models randomly mask a portion of the input and reconstruct the masked portion according to assigned target representations. In this paper, we show that a careful choice of the target representation is unnecessary for learning good visual representation since different targets tend to derive similarly behaved models. Driven by this observation, we propose a multi-stage masked distillation pipeline and use a randomly initialized model as the teacher, enabling us to effectively train high-capacity models without any effort to carefully design the target representation. On various downstream tasks, the proposed method to perform masked knowledge distillation with bootstrapped teachers (dbot) outperforms previous self-supervised methods by nontrivial margins. We hope our findings, as well as the proposed method, could motivate people to rethink the roles of target representations in pre-training masked autoencoders.
Xingbin Liu, Jinghao Zhou, Tao Kong, Xianming Lin, Rongrong Ji
ICLR2
2023 Non-Contrastive Learning Meets Language-Image Pre-Training
abstract
Contrastive language-image pre-training (CLIP) serves as a de-facto standard to align images and texts. Nonetheless, the loose correlation between images and texts of webcrawled data renders the contrastive objective data inefficient and craving for a large training batch size. In this work, we explore the validity of non-contrastive language-image pre-training (nCLIP), and study whether nice properties exhibited in visual self-supervised models can emerge. We empirically observe that the non-contrastive objective benefits representation learning while sufficiently underperforming under zero-shot recognition. Based on the above study, we further introduce xCLIP, a multi-tasking framework combining CLIP and nCLIP, and show that nCLIP aids CLIP in enhancing feature semantics. The synergy between two objectives lets xCLIP enjoy the best of both worlds: superior performance in both zero-shot transfer and representation learning. Systematic evaluation is conducted spanning a wide variety of downstream tasks including zero-shot classification, out-of-domain classification, retrieval, visual representation learning, and textual representation learning, showcasing a consistent performance gain and validating the effectiveness of xCLIP. The code and pre-trained models will be publicly available at https://github.com/shallowtoil/xclip.
Jinghao Zhou, Li Dong 0004, Zhe Gan, Furu Wei
CVPR1
2023 A Restorable, Variable Stiffness Pneumatic Soft Gripper Based on Jamming of Strings of Beads
abstract
Soft robots based on particle jamming cannot return to the initial position and initial mechanical state due to the accumulation of particles after removing the particle jamming, which means poor restorability, and the compliance of the robots during deformation will be reduced because of the jamming effect. Here, we present the design, fabrication, and tests of a novel soft actuator with good restorability and compliance. To improve the restorability of the actuator, we used cotton threads to connect the spherical acrylic beads into form strings instead of discrete beads. The beads could be pulled to the initial position by the threads, the actuator also returns to the initial state. To avoid the jamming effect during the deformation of the actuator, we used compressed air to drive the actuator and injected the beads into the actuator after the active deformation. To reduce the driving pressure and facilitate the flow of the beads, an initial noncontact, frame-type strain constraint structure was designed for the soft actuator. Experimental data show that the actuator was flexible during bending and the stiffness can increase more than 12-fold to resist the external load. By pulling the threads, the actuator could be restored to the initial state with an error of less than 3% of the actuator length after an operation cycle. The soft gripper based on the actuator can grasp repeatedly or laterally. The gripper can grasp soft objects such as a piece of tofu and a balloon of water, and the maximum weight that can be stably grasped is 2.744 kg.
Fenglin Han, Lei Fei, Run Zou, Jinghao Zhou, Haiming Zhao
IEEE Trans. Robotics5
2022 Image BERT Pre-training with Online Tokenizer
Jinghao Zhou, Chen Wei 0005, Wei Shen 0002, Cihang Xie, Alan L. Yuille, Tao Kong
ICLR1
2022 Pluggable Weakly-Supervised Cross-View Learning for Accurate Vehicle Re-Identification
abstract
Learning cross-view consistent feature representation is the key for accurate vehicle Re-identification (ReID), since the visual appearance of vehicles changes significantly under different viewpoints. To this end, many existing approaches resort to the supervised cross-view learning using extensive extra viewpoints annotations, which however, is difficult to deploy in real applications due to the expensive labelling cost and the continous viewpoint variation that makes it hard to define discrete viewpoint labels. In this study, we present a pluggable Weakly-supervised Cross-View Learning (WCVL) module for vehicle ReID. Through hallucinating the cross-view samples as the hardest positive counterparts with small luminance difference and large local feature variance, we can learn the consistent feature representation via minimizing the cross-view feature distance based on vehicle IDs only without using any viewpoint annotation. More importantly, the proposed method can be seamlessly plugged into most existing vehicle ReID baselines for cross-view learning without re-training the baselines. To demonstrate its efficacy, we plug the proposed method into a bunch of off-the-shelf baselines and obtain significant performance improvement on four public benchmark datasets, i.e., VeRi-776, VehicleID, VRIC and VRAI.
Lu Yang 0016, Hongbang Liu, Lingqiao Liu, Jinghao Zhou, Lei Zhang 0054, Peng Wang 0015, Yanning Zhang 0001
ICMR4
2022 Semi-Supervised Segmentation of Radiation-Induced Pulmonary Fibrosis From Lung CT Scans With Multi-Scale Guided Dense Attention
abstract
Computed Tomography (CT) plays an important role in monitoring radiation-induced Pulmonary Fibrosis (PF), where accurate segmentation of the PF lesions is highly desired for diagnosis and treatment follow-up. However, the task is challenged by ambiguous boundary, irregular shape, various position and size of the lesions, as well as the difficulty in acquiring a large set of annotated volumetric images for training. To overcome these problems, we propose a novel convolutional neural network called PF-Net and incorporate it into a semi-supervised learning framework based on Iterative Confidence-based Refinement And Weighting of pseudo Labels (I-CRAWL). Our PF-Net combines 2D and 3D convolutions to deal with CT volumes with large inter-slice spacing, and uses multi-scale guided dense attention to segment complex PF lesions. For semi-supervised learning, our I-CRAWL employs pixel-level uncertainty-based confidence-aware refinement to improve the accuracy of pseudo labels of unannotated images, and uses image-level uncertainty for confidence-based image weighting to suppress low-quality pseudo labels in an iterative training process. Extensive experiments with CT scans of Rhesus Macaques with radiation-induced PF showed that: 1) PF-Net achieved higher segmentation accuracy than existing 2D, 3D and 2.5D neural networks, and 2) I-CRAWL outperformed state-of-the-art semi-supervised learning methods for the PF lesion segmentation task. Our method has a potential to improve the diagnosis of PF and clinical assessment of side effects of radiotherapy for lung cancers.
Guotai Wang, Shuwei Zhai, Giovanni Lasio, Baoshe Zhang, Byong Yi, Shifeng Chen, Thomas J. Macvittie, Dimitris N. Metaxas, Jinghao Zhou, Shaoting Zhang 0001
IEEE Trans. Medical Imaging9
2020 Discriminative and Robust Online Learning for Siamese Visual Tracking
abstract
The problem of visual object tracking has traditionally been handled by variant tracking paradigms, either learning a model of the object's appearance exclusively online or matching the object with the target in an offline-trained embedding space. Despite the recent success, each method agonizes over its intrinsic constraint. The online-only approaches suffer from a lack of generalization of the model they learn thus are inferior in target regression, while the offline-only approaches (e.g., convolutional siamese trackers) lack the target-specific context information thus are not discriminative enough to handle distractors, and robust enough to deformation. Therefore, we propose an online module with an attention mechanism for offline siamese networks to extract target-specific features under L2 error. We further propose a filter update strategy adaptive to treacherous background noises for discriminative learning, and a template update strategy to handle large target deformations for robust learning. Effectiveness can be validated in the consistent improvement over three siamese baselines: SiamFC, SiamRPN++, and SiamMask. Beyond that, our model based on SiamRPN++ obtains the best results over six popular tracking benchmarks and can operate beyond real-time.
Jinghao Zhou, Peng Wang 0015
AAAI1
2009 3D Meshless Prostate Segmentation and Registration in Image Guided Radiotherapy
Ting Chen 0001, Sung N. Kim, Jinghao Zhou, Dimitris N. Metaxas, Gunaretnam Rajagopal, Ning J. Yue
MICCAI (1)3
2007 A boosting regression approach to medical anatomy detection
abstract
The state-of-the-art object detection algorithm learns a binary classifier to differentiate the foreground object from the background. Since the detection algorithm exhaustively scans the input image for object instances by testing the classifier, its computational complexity linearly depends on the image size and, if say orientation and scale are scanned, the number of configurations in orientation and scale. We argue that exhaustive scanning is unnecessary when detecting medical anatomy because a medical image offers strong contextual information. We then present an approach to effectively leveraging the medical context, leading to a solution that needs only one scan in theory or several sparse scans in practice and only one integral image even when the rotation is considered. The core is to learn a regression function, based on an annotated database, that maps the appearance observed in a scan window to a displacement vector, which measures the difference between the configuration being scanned and that of the target object. To achieve the learning task, we propose an image-based boosting ridge regression algorithm, which exhibits good generalization capability and training efficiency. Coupled with a binary classifier as a confidence scorer, the regression approach becomes an effective tool for detecting left ventricle in echocardiogram, achieving improved accuracy over the state-of-the-art object detection algorithm with significantly less computation.
Shaohua Kevin Zhou, Jinghao Zhou, Dorin Comaniciu
CVPR2
2007 Registration of Lung Tissue Between Fluoroscope and CT Images: Determination of Beam Gating Parameters in Radiotherapy
Sukmoon Chang, Jinghao Zhou, Qingshan Liu 0001, Dimitris N. Metaxas, Bruce G. Haffty, Sung N. Kim, Salma J. Jabbour, Ning J. Yue
MICCAI (1)2
2006 Automatic Detection and Segmentation of Ground Glass Opacity Nodules
Jinghao Zhou, Sukmoon Chang, Dimitris N. Metaxas, Binsheng Zhao, Lawrence H. Schwartz, Michelle S. Ginsberg
MICCAI (1)1