Chong Zhou

dblp:42/5942 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Leveraging Inter-Generational Knowledge Transfer in Large-Scale Global Optimization
abstract
Large-scale global optimization (LSGO) presents significant challenges due to the high dimensionality and complexity of the search space. We propose an IGKT (Inter-Generational Knowledge Transfer) optimization method incorporating a novel knowledge transfer mechanism to address these challenges. The proposed mechanism enables the algorithm to reduce reliance on stochastic exploration and enhance convergence efficiency by transferring information from the best-performing individuals across generations, guiding the population toward promising regions in the search space. Experimental results on the CEC2013 LSGO benchmark suite demonstrate that IGKT outperforms several state-of-the-art algorithms across various tested functions, achieving superior convergence speed and solution quality. Additionally, the IGKT framework handles both separable and non-separable functions and tasks, including those with overlapping and highly coupled variables. In summary, IGKT represents a powerful tool for addressing complex, high-dimensional optimization problems, providing a robust and adaptable solution for LSGO.
Yuefeng Xu, Rui Zhong 0004, Chong Zhou, Chao Zhang 0030, Jun Yu 0012
CEC3
2025 EdgeTAM: On-Device Track Anything Model
abstract
On top of Segment Anything Model (SAM), SAM 2 further extends its capability from image to video inputs through a memory bank mechanism and obtains a remarkable performance compared with previous methods, making it a foundation model for video segmentation task. In this paper, we aim at making SAM 2 much more efficient so that it even runs on mobile devices while maintaining a comparable performance. Despite several works optimizing SAM for better efficiency, we find they are not sufficient for SAM 2 because they all focus on compressing the image encoder, while our benchmark shows that the newly introduced memory attention blocks are also the latency bottleneck. Given this observation, we propose EdgeTAM, which leverages a novel 2D Spatial Perceiver to reduce the computational cost. In particular, the proposed 2D Spatial Perceiver encodes the densely stored frame-level memories with a lightweight Transformer that contains a fixed set of learnable queries. Given that video segmentation is a dense prediction task, we find preserving the spatial structure of the memories is essential so that the queries are split into global-level and patch-level groups. We also propose a distillation pipeline that further improves the performance without inference overhead. As a result, EdgeTAM achieves 87.7, 70.0, 72.3, and 71.7 ${\mathcal{J}}{{\& }}{\mathcal{F}}$ on DAVIS 2017, MOSE, SA-V val, and SA-V test, while running at 16 FPS on iPhone 15 Pro Max. The code and models are available here.
Chong Zhou, Chenchen Zhu, Yunyang Xiong, Saksham Suri, Fanyi Xiao, Lemeng Wu, Raghuraman Krishnamoorthi, Bo Dai 0002, Chen Change Loy, Vikas Chandra, Bilge Soran
CVPR1
2025 Efficient Track Anything
abstract
Segment Anything Model 2 (SAM 2) has emerged as a powerful tool for video object segmentation and tracking anything. Key components of SAM 2 that drive the impressive video object segmentation performance include a large multistage image encoder for frame feature extraction and a memory mechanism that stores memory contexts from past frames to help current frame segmentation. The high computation complexity of multistage image encoder and memory module has limited its applications in real-world tasks, e.g., video object segmentation on mobile devices. To address this limitation, we propose EfficientTAMs, lightweight track anything models that produce high-quality results with low latency and model size. Our idea is based on revisiting the plain, nonhierarchical Vision Transformer (ViT) as an image encoder for video object segmentation, and introducing an efficient memory module, which reduces the complexity for both frame feature extraction and memory computation for current frame segmentation. We take vanilla lightweight ViTs and efficient memory module to build EfficientTAMs, and train the models on SA-1B and SA-V datasets for video object segmentation and track anything tasks. We evaluate on multiple video segmentation benchmarks including semi-supervised VOS and promptable video segmentation, and find that our proposed EfficientTAM with vanilla ViT perform comparably to SAM 2 model (HieraB+SAM 2) with ~2x speedup on A100 and ~2.4x parameter reduction. On segment anything image tasks, our EfficientTAMs also perform favorably over original SAM with ~20x speedup on A100 and ~20x parameter reduction. On mobile devices such as iPhone 15 Pro Max, our EfficientTAMs can run at ~10 FPS for performing video object segmentation with reasonable quality, highlighting the capability of small models for on-device video object segmentation applications.
Yunyang Xiong, Chong Zhou, Xiaoyu Xiang, Lemeng Wu, Chenchen Zhu, Zechun Liu, Saksham Suri, Balakrishnan Varadarajan, Ramya Akula, Forrest N. Iandola, Raghuraman Krishnamoorthi, Bilge Soran, Vikas Chandra
ICCV2
2025 Vision Transformers with Self-Distilled Registers
abstract
Vision Transformers (ViTs) have emerged as the dominant architecture for visual processing tasks, demonstrating excellent scalability with increased training data and model size. However, recent work has identified the emergence of artifact tokens in ViTs that are incongruous with local semantics. These anomalous tokens degrade ViT performance in tasks that require fine-grained localization or structural coherence. An effective mitigation of this issue is the addition of register tokens to ViTs, which implicitly ''absorb'' the artifact term during training. Given the availability of existing large-scale pre-trained ViTs, in this paper we seek to add register tokens to existing models without retraining the models from scratch, which is infeasible considering their size. Specifically, we propose Post Hoc Registers (**PH-Reg**), an efficient self-distillation method that integrates registers into an existing ViT without requiring additional labeled data and full retraining. PH-Reg initializes both teacher and student networks from the same pre-trained ViT. The teacher remains frozen and unmodified, while the student is augmented with randomly initialized register tokens. By applying test-time augmentation to the teacher’s inputs, we generate denoised dense embeddings free of artifacts, which are then used to optimize only a small subset of unlocked student weights. We show that our approach can effectively reduce the number of artifact tokens, improving the segmentation and depth prediction of the student ViT under zero-shot and linear probing.
Zipeng Yan, Yinjie Chen, Chong Zhou, Bo Dai 0002, Andrew Luo 0001
NeurIPS3
2025 A binary particle swarm optimization with dual encoding mechanism for feature selection
Chong Zhou, Rumeng Liang, Sirui Niu
Eng. Appl. Artif. Intell.1
2025 EdgeSAM: Prompt-In-the-Loop Distillation for SAM
Chong Zhou, Xiangtai Li, Chen Change Loy, Bo Dai 0002
Int. J. Comput. Vis.1
2024 Open-Vocabulary SAM: Segment and Recognize Twenty-Thousand Classes Interactively
Haobo Yuan, Xiangtai Li, Chong Zhou, Kai Chen 0026, Chen Change Loy
ECCV (43)3
2023 Rubik's Cube: High-Order Channel Interactions with a Hierarchical Receptive Field
abstract
Image restoration techniques, spanning from the convolution to the transformer paradigm, have demonstrated robust spatial representation capabilities to deliver high-quality performance.Yet, many of these methods, such as convolution and the Feed Forward Network (FFN) structure of transformers, primarily leverage the basic first-order channel interactions and have not maximized the potential benefits of higher-order modeling. To address this limitation, our research dives into understanding relationships within the channel dimension and introduces a simple yet efficient, high-order channel-wise operator tailored for image restoration. Instead of merely mimicking high-order spatial interaction, our approach offers several added benefits: Efficiency: It adheres to the zero-FLOP and zero-parameter principle, using a spatial-shifting mechanism across channel-wise groups. Simplicity: It turns the favorable channel interaction and aggregation capabilities into element-wise multiplications and convolution units with $1 \times 1$ kernel. Our new formulation expands the first-order channel-wise interactions seen in previous works to arbitrary high orders, generating a hierarchical receptive field akin to a Rubik's cube through the combined action of shifting and interactions. Furthermore, our proposed Rubik's cube convolution is a flexible operator that can be incorporated into existing image restoration networks, serving as a drop-in replacement for the standard convolution unit with fewer parameters overhead. We conducted experiments across various low-level vision tasks, including image denoising, low-light image enhancement, guided image super-resolution, and image de-blurring. The results consistently demonstrate that our Rubik's cube operator enhances performance across all tasks. Code is publicly available at https://github.com/zheng980629/RubikCube.
Naishan Zheng, Man Zhou 0003, Chong Zhou, Chen Change Loy
NeurIPS3
2022 PAC-Net: Highlight Your Video via History Preference Modeling
Penghao Zhou, Chong Zhou, Zhao Zhang 0018, Xing Sun 0001
ECCV (34)3
2022 Extract Free Dense Labels from CLIP
Chong Zhou, Chen Change Loy, Bo Dai 0002
ECCV (28)1
2022 YOLACT++ Better Real-Time Instance Segmentation
abstract
We present a simple, fully-convolutional model for real-time ( fps) instance segmentation that achieves competitive results on MS COCO evaluated on a single Titan Xp, which is significantly faster than any previous state-of-the-art approach. Moreover, we obtain this result after training on only one GPU. We accomplish this by breaking instance segmentation into two parallel subtasks: (1) generating a set of prototype masks and (2) predicting per-instance mask coefficients. Then we produce instance masks by linearly combining the prototypes with the mask coefficients. We find that because this process doesn't depend on repooling, this approach produces very high-quality masks and exhibits temporal stability for free. Furthermore, we analyze the emergent behavior of our prototypes and show they learn to localize instances on their own in a translation variant manner, despite being fully-convolutional. We also propose Fast NMS, a drop-in 12 ms faster replacement for standard NMS that only has a marginal performance penalty. Finally, by incorporating deformable convolutions into the backbone network, optimizing the prediction head with better anchor scales and aspect ratios, and adding a novel fast mask re-scoring branch, our YOLACT++ model can achieve 34.1 mAP on MS COCO at 33.5 fps, which is fairly close to the state-of-the-art approaches while still running at real-time.
Daniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae Lee
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 NOH-NMS: Improving Pedestrian Detection by Nearby Objects Hallucination
abstract
Greedy-NMS inherently raises a dilemma, where a lower NMS threshold will potentially lead to a lower recall rate and a higher threshold introduces more false positives. This problem is more severe in pedestrian detection because the instance density varies more intensively. However, previous works on NMS don't consider or vaguely consider the factor of the existent of nearby pedestrians. Thus, we propose \heatmapname (\heatmapnameshort ), which pinpoints the objects nearby each proposal with a Gaussian distribution, together with \nmsname, which dynamically eases the suppression for the space that might contain other objects with a high likelihood. Compared to Greedy-NMS, our method, as the state-of-the-art, improves by $3.9%$ AP, $5.1%$ Recall, and $0.8%$ MR\textsuperscript-2 on CrowdHuman to $89.0%$ AP and $92.9%$ Recall, and $43.9%$ MR\textsuperscript-2 respectively.
Penghao Zhou, Chong Zhou, Junlong Du, Xing Sun 0001, Feiyue Huang
ACM Multimedia2
2020 Generative Adversarial Active Learning for Unsupervised Outlier Detection
abstract
Outlier detection is an important topic in machine learning and has been used in a wide range of applications. In this paper, we approach outlier detection as a binary-classification issue by sampling potential outliers from a uniform reference distribution. However, due to the sparsity of data in high-dimensional space, a limited number of potential outliers may fail to provide sufficient information to assist the classifier in describing a boundary that can separate outliers from normal data effectively. To address this, we propose a novel Single-Objective Generative Adversarial Active Learning (SO-GAAL) method for outlier detection, which can directly generate informative potential outliers based on the mini-max game between a generator and a discriminator. Moreover, to prevent the generator from falling into the mode collapsing problem, the stop node of training should be determined when SO-GAAL is able to provide sufficient information. But without any prior information, it is extremely difficult for SO-GAAL. Therefore, we expand the network structure of SO-GAAL from a single generator to multiple generators with different objectives (MO-GAAL), which can generate a reasonable reference distribution for the whole dataset. We empirically compare the proposed approach with several state-of-the-art outlier detection methods on both synthetic and real-world datasets. The results show that MO-GAAL outperforms its competitors in the majority of cases, especially for datasets with various cluster types or high irrelevant variable ratio. The experiment codes are available at: https://github.com/leibinghe/GAAL-based-outlier-detection.
Ye-Zheng Liu 0001, Zhe Li 0070, Chong Zhou, Yuan-Chun Jiang, Jianshan Sun, Meng Wang 0001, Xiangnan He 0001
IEEE Trans. Knowl. Data Eng.3
2019 YOLACT: Real-Time Instance Segmentation
abstract
We present a simple, fully-convolutional model for real-time instance segmentation that achieves 29.8 mAP on MS COCO at 33.5 fps evaluated on a single Titan Xp, which is significantly faster than any previous competitive approach. Moreover, we obtain this result after training on only one GPU. We accomplish this by breaking instance segmentation into two parallel subtasks: (1) generating a set of prototype masks and (2) predicting per-instance mask coefficients. Then we produce instance masks by linearly combining the prototypes with the mask coefficients. We find that because this process doesn't depend on repooling, this approach produces very high-quality masks and exhibits temporal stability for free. Furthermore, we analyze the emergent behavior of our prototypes and show they learn to localize instances on their own in a translation variant manner, despite being fully-convolutional. Finally, we also propose Fast NMS, a drop-in 12 ms faster replacement for standard NMS that only has a marginal performance penalty.
Daniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae Lee
ICCV2
2019 Trilateral Smooth Filtering for Hyperspectral Image Feature Extraction
abstract
Traditional bilateral filtering (BF) cannot extract hyperspectral image (HSI) features well when the center pixel of the neighborhood pixel set is a noise point in the process of filtering the HSI. In this letter, a trilateral smooth filtering (TRSF) is presented. The proposed algorithm avoids the above-mentioned limitation problem in the BF algorithm. TRSF is successfully applied to the feature extraction of three actual HSIs. To prove the effectiveness of the proposed algorithm, support vector machines are used to classify the extracted features. Experimental results show that the proposed feature extraction method is simple and effective.
Junjun Jiang, Chong Zhou, Xinwei Jiang, Shaoyuan Fu, Zhihua Cai
IEEE Geosci. Remote. Sens. Lett.3
2018 Enhanced θ dominance and density selection based evolutionary algorithm for many-objective optimization problems
Chong Zhou, Guangming Dai, Maocai Wang
Appl. Intell.1
2018 Entropy based evolutionary algorithm with adaptive reference points for many-objective optimization problems
Chong Zhou, Guangming Dai, Cuijun Zhang, Xiangping Li
Inf. Sci.1
2018 Indicator and reference points co-guided evolutionary algorithm for many-objective optimization problems
Guangming Dai, Chong Zhou, Maocai Wang, Xiangping Li
Knowl. Based Syst.2
2017 Anomaly Detection with Robust Deep Autoencoders
abstract
Deep autoencoders, and other deep neural networks, have demonstrated their effectiveness in discovering non-linear features across many problem domains. However, in many real-world problems, large outliers and pervasive noise are commonplace, and one may not have access to clean training data as required by standard deep denoising autoencoders. Herein, we demonstrate novel extensions to deep autoencoders which not only maintain a deep autoencoders' ability to discover high quality, non-linear features but can also eliminate outliers and noise without access to any clean training data. Our model is inspired by Robust Principal Component Analysis, and we split the input data X into two parts, $X = L_{D} + S$, where $L_{D}$ can be effectively reconstructed by a deep autoencoder and $S$ contains the outliers and noise in the original data X. Since such splitting increases the robustness of standard deep autoencoders, we name our model a "Robust Deep Autoencoder (RDA)". Further, we present generalizations of our results to grouped sparsity norms which allow one to distinguish random anomalies from other types of structured corruptions, such as a collection of features being corrupted across many instances or a collection of instances having more corruptions than their fellows. Such "Group Robust Deep Autoencoders (GRDA)" give rise to novel anomaly detection approaches whose superior performance we demonstrate on a selection of benchmark problems.
Chong Zhou, Randy C. Paffenroth
KDD1
2016 Entropy determined hybrid two-stage multi-objective evolutionary algorithm combining locally linear embedding
abstract
For some probabilistic model-based multi-objective evolutionary algorithms (MOEAs), the probability model established may not accurate enough due to the lack of effective distribution information in the early evolutionary stage. To improve this problem, a novel hybrid multi-objective optimization algorithm is proposed in this paper. Specifically, traditional crossover and mutation operation are used in the early evolutionary stage to explore the promising search areas. Moreover, the locally linear embedding (LLE) with low neighbor parameter approach is involved to enhance the exploitation ability of the proposed algorithm. In addition, an entropy-based criterion is introduced to judge whether certain regularity is presented in population's distribution. The probabilistic model-based approach will be used to reproduce new offspring if some certain regularity is presented. The hybrid two-stage multi-objective evolutionary algorithm proposed in this paper is called entropy determined hybrid two-stage multi-objective evolutionary algorithm combining locally linear embedding (EHMOEA_LLE). To verify the performance of EHMOEA_LLE, several test problems used widely are employed to conduct the comparison experiments with two state-of-the-art multi-objective evolutionary algorithms NSGA-II and RM-MEDA. The simulation results show that the entropy-based criterion is effective and the proposed algorithm is better optimization performance.
Chong Zhou, Guangming Dai, Ruixue Hu
CEC2
2016 Robust design optimization based on multi-objective particle swarm optimization
abstract
For real world problems, there are inevitably perturbations in the design parameters or (and) variables. If an optimal solution is sensitive to the small perturbations of design parameters or variables, it may be inappropriate or risk for practical use. Robust design optimization can find solutions which are good in optimality and good in robustness simultaneously. Traditional robust optimization searched for robust solution by converting the original problem into a single-objective optimization problem. But only one solution can be obtained from one run of optimization using these methods. This paper applied a multi-objective optimization approach to get the robust optimal solutions. A novel robustness measurement is proposed and is compared with other methods of estimating robustness. The theoretical and experimental results verify that the proposed method outperforms the conventional methodologies. In solving the converted multi-objective optimization problems, a more efficient multi-objective particle swarm optimization is used. The test functions proved that the proposed method is more efficient than the conventional ones.
Guangming Dai, Chong Zhou, Lei Peng 0001
CEC4
2013 Co-training over domain-independent and domain-dependent features for sentiment analysis of an online cancer support community
abstract
Sentiment analysis has been widely researched in the domain of online review sites with the aim of getting summarized opinions of product users about different aspects of the products. However, there has been little work focusing on identifying the polarity of sentiments expressed by users in online health communities such as cancer support forums, etc. Online health communities act as a medium through which people share their health concerns with fellow members of the community and get social support. Identifying sentiments expressed by members in a health community can be helpful in understanding dynamics of the community such as dominant health issues, emotional impacts of interactions on members, etc. In this work, we perform sentiment classification of user posts in an online cancer support community (Cancer Survivors Network). We use Domain-dependent and Domain-independent sentiment features as the two complementary views of a post and use them for post classification in a semi-supervised setting using the co-training algorithm. Experimental results demonstrate effectiveness of our methods.
Prakhar Biyani, Cornelia Caragea, Prasenjit Mitra 0001, Chong Zhou, John Yen, Greta E. Greer, Kenneth Portier
ASONAM4
2012 HeteRecom: a semantic-based recommendation systemin heterogeneous networks
abstract
Making accurate recommendations for users has become an important function of e-commerce system with the rapid growth of WWW. Conventional recommendation systems usually recommend similar objects, which are of the same type with the query object without exploring the semantics of different similarity measures. In this paper, we organize objects in the recommendation system as a heterogeneous network. Through employing a path-based relevance measure to evaluate the relatedness between any-typed objects and capture the subtle semantic containing in each path, we implement a prototype system (called HeteRecom) for semantic based recommendation. HeteRecom has the following unique properties: (1) It provides the semantic-based recommendation function according to the path specified by users. (2) It recommends the similar objects of the same type as well as related objects of different types. We demonstrate the effectiveness of our system with a real-world movie data set.
Chuan Shi 0001, Chong Zhou, Xiangnan Kong, Philip S. Yu, Gang Liu 0008, Bai Wang 0001
KDD2
2010 Predicate-based indexing for desktop search
Cristian Duda, Donald Kossmann, Chong Zhou
VLDB J.3
2009 AJAX Crawl: Making AJAX Applications Searchable
abstract
Current search engines such as Google and Yahoo! are prevalent for searching the Web. Search on dynamic client-side Web pages is, however, either inexistent or far from perfect, and not addressed by existing work, for example on Deep Web. This is a real impediment since AJAX and Rich Internet Applications are already very common in the Web. AJAX applications are composed of states which can be seen by the user, but not by the search engine, and changed by the user using client-side events. Current search engines either ignore AJAX applications or produce false negatives. The reason is that crawling client-side code is a difficult problem that cannot be solved naively by invoking user events. The challenges are: lack of caching, duplicate states detection, very granular events, reducing the number of AJAX calls and infinite event invocation. This paper sets the stage for this new search challenge and proposes a solution: it shows how an AJAX Web application can be crawled in the granularity of the application states. A model of AJAX Web sites is presented. An AJAX Crawler and optimizations for caching and duplicate elimination are defined, and finally, the gain in search result quality and corresponding performance price are evaluated on YouTube, a real AJAX application.
Cristian Duda, Gianni Frey, Donald Kossmann, Reto Matter, Chong Zhou
ICDE5
2008 AJAXSearch: crawling, indexing and searching web 2.0 applications
abstract
Current search engines such as Google and Yahoo! are prevalent for searching the Web. Search in dynamic pages, however, is either inexistent or far from perfect. AJAX and Rich Internet Application are such applications. They are increasingly frequent on the Web (in YouTube, Amazon, GMail, Yahoo!Mail) or mobile devices and are offering a high degree of interactivity to the user, by seamlessly loading content from the server without the need to refresh the page. Current search engines cannot correctly index AJAX applications. This produces false positives and false negatives, because search engines do not understand the application logic that loads content dynamically. Crawling an AJAX application is a difficult problem. Since the user invokes events on the page, crawling must identify the different application states generated by the client-side logic. This demo sets the stage for this new type of search and shows that a search engine for AJAX can be built. Among others, the challenges, as opposed to traditional search engines, are: automatically identifying states by triggering events, efficiently crawling application states, avoiding the invocation of potentially very numerous events, scalability in the number of events, duplicate elimination of states, result presentation and aggregation, ranking. The demo presents the AJAX search engine: crawler, indexer and query processor, applied on a real application and showcases challenges and solutions.
Cristian Duda, Gianni Frey, Donald Kossmann, Chong Zhou
Proc. VLDB Endow.4
2007 Wavelet Synopsis: Setting Unselected Coefficients to Zero Is Not Optimal
Yansheng Lu, Chong Zhou
DEXA3