EDBT 2026 Demo / reviewers in the wild / expert
Zijun Wei
dblp:157/3589
· DBLP profile ↗
27ranked-venue papers
9as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-level adaptive feature enhancement and attention mask-guided aggregation for visual place recognition
Longhao Wang, Chaozhen Lan, Beibei Wu, Fushan Yao, Zijun Wei, Tian Gao 0006, Hanyang Yu |
Knowl. Based Syst. | 5 |
| 2026 | A Conditional GAN-Based Framework for Sparse sEMG Data Augmentation With Muscle Synergy Prior ConstraintsabstractThe scarcity of high-quality surface electromyography (sEMG) data, caused by ethical constraints, privacy concerns, and noise interference, poses significant challenges for developing robust deep learning models in sEMG analysis. Multi-channel sEMG signals exhibit complex inter-channel correlations reflecting neuromuscular coordination. However, existing generative methods suffer from error accumulation in sequential channel generation, insufficient inter-channel relationship modeling, and lack of physiological constraints, producing data-driven valid but physiologically implausible signals that compromise biological fidelity for clinical applications. To address these fundamental limitations, we propose a Muscle Synergy-Constrained Conditional GAN (MS-cGAN) framework that simultaneously generates multi-channel sEMG signals while preserving bio-mechanical fidelity. Firstly, A novel Graph Convolutional Network (GCN)-based generator architecture specifically tailored for sparse sEMG signals, which enables the generator to capture and model complex inter-channel relationship features through graph-based representation learning, thereby circumventing error accumulation issues by leveraging the inherent inter-channel correlations. Secondly, Integration of Muscle Synergy (MS) prior constraints as dynamic loss functions based on MS theory, which enforces generator optimization within a physiologically plausible parameter space and ensures signals maintain synergistic consistency with underlying physiological mechanisms. Lastly, experiments on IRASS datasets and public datasets (NinaPro DB1 and DB2) demonstrate that MS-cGAN significantly improves signal authenticity and enhances downstream task performance compared to traditional GANs and state-of-the-art diffusion models. The generated data effectively supplement scarce sEMG datasets and improve kinematic prediction precision for deep learning models. Meiju Li, Zijun Wei, Zhiqiang Zhang 0001, Shengquan Xie |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | A Transformer Framework Informed by Muscle Anatomy and Sequence-to-Sequence Translation for Continuous Joint Kinematics Prediction Using sEMGabstractThe key to achieving assist-as-needed (AAN) control in rehabilitation robots lies in accurately predicting patient motion intentions. This study, for the first time, redefines motion intention prediction from the perspective of sequence-to-sequence translation by analogizing sEMG signals and joint angles to the source language and target language, respectively. The proposed 3DCNN-TF model achieves precise translation of neural control signals into kinematic representations. This model comprises three modules: an sEMG “sentence” generation module that compiles multiple sEMG sliding windows into a “sentence,” a 3DCNN module based on muscle anatomy and electrode placement to extract muscle synergy features from each “word” in the “sentence,” and a Transformer (TF) module that autoregressively generates the next joint angle as the translation result. Experimental results indicate that the 3DCNN-TF model achieves superior overall performance compared to eight baseline models and existing studies in continuously predicting wrist and knee flexion/extension angles across varying speeds. Moreover, the 3DCNN-TF achieves an optimal balance between prediction accuracy and computational efficiency while exhibiting exceptional robustness and generalizability. Specifically, the 3DCNN-TF achieves average nRMSE and R2values of (6.2%/95.5%) and (5.5%/96.2%) on wrist and knee datasets, respectively, with an average training time of less than two minutes. Additionally, the 3DCNN-TF can predict joint angles up to 300 ms in advance without compromising accuracy, which is critical for real-time AAN control in rehabilitation robots. Zijun Wei, Zhiqiang Zhang 0001, Shengquan Xie |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Refer to Any Segmentation Mask Group with Vision-Language PromptsabstractRecent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language and vision. This limitation reduces their effectiveness in applications that require user-friendly interactions driven by vision-language prompts. To bridge this gap, we introduce a novel task of omnimodal referring expression segmentation (ORES). In this task, a model produces a group of masks based on arbitrary prompts specified by text only or text plus reference visual entities. To address this new challenge, we propose a novel framework to "Refer to Any Segmentation Mask Group" (RAS), which augments segmentation models with complex multimodal interactions and comprehension via a mask-centric large multimodal model. For training and benchmarking ORES models, we create datasets MaskGroups-2M and MaskGroups-HQ to include diverse mask groups specified by text and reference entities. Through extensive evaluation, we demonstrate superior performance of RAS on our new ORES task, as well as classic referring expression segmentation (RES) and generalized referring expression segmentation (GRES) tasks. Project page: https://Ref2Any.github.io. Shengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu, Lingzhi Zhang, Jiuxiang Gu, Hyunjoon Jung, Liangyan Gui, Yu-Xiong Wang |
ICCV | 2 |
| 2025 | A Depth Semantic Perception Network for Camouflage Object DetectionabstractCamouflage object detection (COD) aims to identify objects that blend seamlessly into complex backgrounds, making it inherently more challenging than conventional object detection. However, most existing COD methods fail to strike a balance between local details and global understanding, leading to false positives or missed detections by the model. To address this issue, we propose a Depth Semantic Perception Network (DSP-Net), which can capture depth semantic information from images and accurately model the spatial relationship between objects and their backgrounds, thereby improving detection accuracy. Specifically, we first propose a Semantic Localization Module (SLM) that integrates multi-scale features from the backbone network and obtains rich semantic information, which is used to modulate the fused features obtained in the Cross-Level Fusion Module (CLFM) to generate higher-level semantic representations. Next, we propose an Adaptive Channel Fusion Module (ACFM), which aggregates multi-scale features through weighted fusion of input features and dynamically learns the weights for each channel. Finally, we propose a Cross-Level Fusion Module (CLFM), which captures depth semantic information through prior guidance and cross-level feature fusion, and balances the information against high spatial resolution information, thereby enhancing the accuracy of COD. Extensive experiments on four benchmark COD datasets show that our DSP-Net outperforms other state-of-the-art models. Zijun Wei, Zhenhong Jia, Haochu Ku |
ICME | 1 |
| 2025 | Continuous Prediction of Wrist Joint Kinematics Using Surface Electromyography From the Perspective of Muscle Anatomy and Muscle Synergy Feature ExtractionabstractPost-stroke upper limb dysfunction severely impacts patients' daily life quality. Utilizing sEMG signals to predict patients' motion intentions enables more effective rehabilitation by precisely adjusting the assistance level of rehabilitation robots. Employing the muscle synergy (MS) features can establish more accurate and robust mappings between sEMG and motion intentions. However, traditional matrix factorization algorithms based on blind source separation still exhibit certain limitations in extracting MS features. This paper proposes four deep learning models to extract MS features from four distinct perspectives: spatiotemporal convolutional kernels, compression and reconstruction of sEMG, graph topological structure, and the anatomy of target muscles. Among these models, the one based on 3DCNN predicts motion intentions from the muscle anatomy perspective for the first time. It reconstructs 1D sEMG samples collected at each time point into 2D sEMG frames based on the anatomical distribution of target muscles and sEMG electrode placement. These 2D frames are then stacked as video segments and input into 3DCNN for MS feature extraction. Experimental results on both our wrist motion dataset and public Ninapro DB2 dataset demonstrate that the proposed 3DCNN model outperforms other models in terms of prediction accuracy, robustness, training efficiency, and MS feature extraction for continuous prediction of wrist flexion/extension angles. Specifically, the average nRMSE and R2values of 3DCNN on these two datasets are (0.14/0.93) and (0.04/0.95), respectively. Furthermore, compared to existing studies, the 3DCNN outperforms musculoskeletal models based on direct collocation optimization, physics-informed GANs, and CNN-LSTM-based deep Kalman filter models when evaluated on our dataset. Zijun Wei, Meiju Li, Zhiqiang Zhang 0001, Shengquan Xie |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Tag-grounded Visual Instruction Tuning with Retrieval AugmentationabstractPlease describe this photo in Daiqing Qi, Handong Zhao, Zijun Wei, Sheng Li 0001 |
EMNLP | 3 |
| 2024 | Uncertainty-aware Fine-tuning of Segmentation Foundation ModelsabstractThe Segment Anything Model (SAM) is a large-scale foundation model that has revolutionized segmentation methodology. Despite its impressive generalization ability, the segmentation accuracy of SAM on images with intricate structures is often unsatisfactory. Recent works have proposed lightweight fine-tuning using high-quality annotated data to improve accuracy on such images. However, here we provide extensive empirical evidence that this strategy leads to forgetting how to "segment anything": these models lose the original generalization abilities of SAM, in the sense that they perform worse for segmentation tasks not represented in the annotated fine-tuning set. To improve performance without forgetting, we introduce a novel framework that combines high-quality annotated data with a large unlabeled dataset. The framework relies on two methodological innovations. First, we quantify the uncertainty in the SAM pseudo labels associated with the unlabeled data and leverage it to perform uncertainty-aware fine-tuning. Second, we encode the type of segmentation task associated with each training example using a $\textit{task prompt}$ to reduce ambiguity. We evaluated the proposed Segmentation with Uncertainty Model (SUM) on a diverse test set consisting of 14 public benchmarks, where it achieves state-of-the-art results. Notably, our method consistently surpasses SAM by 3-6 points in mean IoU and 4-7 in mean boundary IoU across point-prompt interactive segmentation rounds. Code is available at https://github.com/Kangningthu/SUM Kangning Liu, Brian L. Price, Jason Kuen, Zijun Wei, Luis Figueroa, Krzysztof J. Geras, Carlos Fernandez-Granda |
NeurIPS | 5 |
| 2024 | Camouflaged Object Detection Based on Feature Aggregation and Global Semantic Learning
Kuan Wang 0007, Zijun Wei, Lining Wan |
PRCV (12) | 5 |
| 2024 | Classify Traffic Rather Than Flow: Versatile Multi-Flow Encrypted Traffic Classification With Flow ClusteringabstractEncrypted Traffic Classification (ETC) can provide necessary information support for network management and security. The state-of-the-art ETC methods take a single flow as the unit and only use sequence features based on in-flow relationships. In an actual network, one-time access to an application will generate multiple flows. Taking a single flow as the classification unit will produce many repeated and potentially erroneous results, which dramatically reduces the classification efficiency and prevents the results from being used for effective network management and security. In this paper, we propose a multi-flow ETC method. Since multiple flows generated by an application cannot be directly bound in a complex multi-application scenario, we first cluster the encrypted traffic to acquire flow bunches through the proposed Time Sequential Hierarchical Clustering with Sliding Windows (TSHC-SW) algorithm. Then, based on the flow bunches, we propose five different multi-flow classification schemas that can realize multi-flow classification effectively with model-independent. Open-world experiments show that our method is versatile in that it can pursue classification accuracy, speed, or sample covering rate, respectively, according to the actual demand and network environment constraints. In flow clustering, we achieve 95% adjusted Rand Index and 98% purity. In the multi-flow classification, we have over 99% F1-score, 79% prediction time saving, and 5% sample covering rate increasing, which is far superior to the state-of-the-art single-flow methods. Zihan Chen 0003, Guang Cheng 0001, Zijun Wei, Dandan Niu, Nan Fu |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2023 | Automatic High Resolution Wire Segmentation and RemovalabstractWires and powerlines are common visual distractions that often undermine the aesthetics of photographs. The manual process of precisely segmenting and removing them is extremely tedious and may take up hours, especially on high-resolution photos where wires may span the entire space. In this paper, we present an automatic wire clean-up system that eases the process of wire segmentation and removal/inpainting to within a few seconds. We observe several unique challenges: wires are thin, lengthy, and sparse. These are rare properties of subjects that common segmentation tasks cannot handle, especially in high-resolution images. We thus propose a two-stage method that leverages both global and local contexts to accurately segment wires in high-resolution images efficiently, and a tile-based inpainting strategy to remove the wires given our predicted segmentation masks. We also introduce the first wire segmentation benchmark dataset, WireSegHR. Finally, we demonstrate quantitatively and qualitatively that our wire clean-up system enables fully automated wire removal with great generalization to various wire appearances. Mang Tik Chiu, Xuaner Cecilia Zhang, Zijun Wei, Yuqian Zhou, Eli Shechtman, Connelly Barnes, Zhe Lin 0001, Florian Kainz, Sohrab Amirghodsi, Humphrey Shi |
CVPR | 3 |
| 2023 | LightPainter: Interactive Portrait Relighting with Freehand ScribbleabstractRecent portrait relighting methods have achieved realistic results of portrait lighting effects given a desired lighting representation such as an environment map. However, these methods are not intuitive for user interaction and lack precise lighting control. We introduce LightPainter, a scribble-based relighting system that allows users to interactively manipulate portrait lighting effect with ease. This is achieved by two conditional neural networks, a delighting module that recovers geometry and albedo optionally conditioned on skin tone, and a scribble-based module for re-lighting. To train the relighting module, we propose a novel scribble simulation procedure to mimic real user scribbles, which allows our pipeline to be trained without any human annotations. We demonstrate high-quality and flexible portrait lighting editing capability with both quantitative and qualitative experiments. User study comparisons with commercial lighting editing tools also demonstrate consistent user preference for our method. Yiqun Mei, He Zhang 0004, Xuaner Cecilia Zhang, Jianming Zhang 0001, Zhixin Shu, Yilin Wang 0002, Zijun Wei, Hyunjoon Jung, Vishal M. Patel |
CVPR | 7 |
| 2023 | Interactive Portrait Harmonization
Jeya Maria Jose Valanarasu, He Zhang 0004, Jianming Zhang 0001, Yilin Wang 0002, Zhe Lin 0001, Jose Echevarria, Yinglan Ma, Zijun Wei, Kalyan Sunkavalli, Vishal M. Patel |
ICLR | 8 |
| 2023 | Predicting Visual Attention in Graphic Design DocumentsabstractWe present a model for predicting visual attention during the free viewing of graphic design documents. While existing works on this topic have aimed at predicting static saliency of graphic designs, our work is the first attempt to predict both spatial attention and dynamic temporal order in which the document regions are fixated by gaze using a deep learning based model. We propose a two-stage model for predicting dynamic attention on such documents, with webpages being our primary choice of document design for demonstration. In the first stage, we predict the saliency maps for each of the document components (e.g. logos, banners, texts, etc. for webpages) conditioned on the type of document layout. These component saliency maps are then jointly used to predict the overall document saliency. In the second stage, we use these layout-specific component saliency maps as the state representation for an inverse reinforcement learning model of fixation scanpath prediction during document viewing. To test our model, we collected a new dataset consisting of eye movements from 41 people freely viewing 450 webpages (the largest dataset of its kind). Experimental results show that our model outperforms existing models in both saliency and scanpath prediction for webpages, and also generalizes very well to other graphic design documents such as comics, posters, mobile UIs, etc. and natural images. Souradeep Chakraborty, Zijun Wei, Conor Kelton, Seoyoung Ahn, Aruna Balasubramanian, Gregory J. Zelinsky, Dimitris Samaras |
IEEE Trans. Multim. | 2 |
| 2022 | Lite Vision Transformer with Enhanced Self-AttentionabstractDespite the impressive representation capacity of vision transformer models, current light-weight vision transformer models still suffer from inconsistent and incorrect dense predictions at local regions. We suspect that the power of their self-attention mechanism is limited in shallower and thinner networks. We propose Lite Vision Transformer (LVT), a novel light-weight transformer network with two enhanced self-attention mechanisms to improve the model performances for mobile deployment. For the low-level features, we introduce Convolutional Self-Attention (CSA). Unlike previous approaches of merging convolution and self-attention, CSA introduces local self-attention into the convolution within a kernel of size$3\times 3$to enrich low-level features in the first stage of LVT. For the high-level features, we propose Recursive Atrous Self-Attention (RASA), which utilizes the multi-scale context when calculating the similarity map and a recursive mechanism to increase the representation capability with marginal extra parameter cost. The superiority of LVT is demonstrated on ImageNet recognition, ADE20K semantic segmentation, and COCO panoptic segmentation. The code is made publicly available11https://github.com/Chenglin-Yang/LVT. Yilin Wang 0002, Jianming Zhang 0001, He Zhang 0004, Zijun Wei, Zhe Lin 0001, Alan L. Yuille |
CVPR | 5 |
| 2022 | Higher Layers, Better Results: Application Layer Feature Engineering in Encrypted Traffic Classification
Zihan Chen 0003, Guang Cheng 0001, Zijun Wei, Nan Fu |
WASA (2) | 3 |
| 2021 | Sequence-to-Segments Networks for Detecting Segments in VideosabstractDetecting segments of interest from videos is a common problem for many applications. And yet it is a challenging problem as it often requires not only knowledge of individual target segments, but also contextual understanding of the entire video and the relationships between the target segments. To address this problem, we propose the Sequence-to-Segments Network (S2N), a novel and general end-to-end sequential encoder-decoder architecture. S2N first encodes the input video into a sequence of hidden states that capture information progressively, as it appears in the video. It then employs the Segment Detection Unit (SDU), a novel decoding architecture, that sequentially detects segments. At each decoding step, the SDU integrates the decoder state and encoder hidden states to detect a target segment. During training, we address the problem of finding the best assignment of predicted segments to ground truth using the Hungarian Matching Algorithm with Lexicographic Cost. Additionally we propose to use the squared Earth Mover's Distance to optimize the localization errors of the segments. We show the state-of-the-art performance of S2N across numerous tasks, including video highlighting, video summarization, and human action proposal generation. Zijun Wei, Boyu Wang 0001, Minh Hoai, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Radomír Mech, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Learning Visual Emotion Representations From Web DataabstractWe present a scalable approach for learning powerful visual features for emotion recognition. A critical bottleneck in emotion recognition is the lack of large scale datasets that can be used for learning visual emotion features. To this end, we curate a webly derived large scale dataset, StockEmotion, which has more than a million images. StockEmotion uses 690 emotion related tags as labels giving us a fine-grained and diverse set of emotion labels, circumventing the difficulty in manually obtaining emotion annotations. We use this dataset to train a feature extraction network, EmotionNet, which we further regularize using joint text and visual embedding and text distillation. Our experimental results establish that EmotionNet trained on the StockEmotion dataset outperforms SOTA models on four different visual emotion tasks. An aded benefit of our joint embedding training approach is that EmotionNet achieves competitive zero-shot recognition performance against fully supervised baselines on a challenging visual emotion dataset, EMOTIC, which further highlights the generalizability of the learned emotion features. Zijun Wei, Jianming Zhang 0001, Zhe Lin 0001, Joon-Young Lee, Niranjan Balasubramanian, Minh Hoai, Dimitris Samaras |
CVPR | 1 |
| 2020 | Predicting Goal-Directed Human Attention Using Inverse Reinforcement LearningabstractHuman gaze behavior prediction is important for behavioral vision and for computer vision applications. Most models mainly focus on predicting free-viewing behavior using saliency maps, but do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We propose the first inverse reinforcement learning (IRL) model to learn the internal reward function and policy used by humans during visual search. We modeled the viewer's internal belief states as dynamic contextual belief maps of object locations. These maps were learned and then used to predict behavioral scanpaths for multiple target categories. To train and evaluate our IRL model we created COCO-Search18, which is now the largest dataset of high-quality search fixations in existence. COCO-Search18 has 10 participants searching for each of 18 target-object categories in 6202 images, making about 300,000 goal-directed fixations. When trained and evaluated on COCO-Search18, the IRL model outperformed baseline models in predicting search fixation scanpaths, both in terms of similarity to human search behavior and search efficiency. Finally, reward maps recovered by the IRL model reveal distinctive target-dependent patterns of object prioritization, which we interpret as a learned object context. Zhibo Yang 0002, Lihan Huang, Yupei Chen, Zijun Wei, Seoyoung Ahn, Gregory J. Zelinsky, Dimitris Samaras, Minh Hoai |
CVPR | 4 |
| 2019 | SmartEye: Assisting Instant Photo Taking via Integrating User Preference with Deep View Proposal NetworkabstractInstant photo taking and sharing has become one of the most popular forms of social networking. However, taking high-quality photos is difficult as it requires knowledge and skill in photography that most non-expert users lack. In this paper we present SmartEye, a novel mobile system to help users take photos with good compositions in-situ. The back-end of SmartEye integrates the View Proposal Network (VPN), a deep learning based model that outputs composition suggestions in real time, and a novel, interactively updated module (P-Module) that adjusts the VPN outputs to account for personalized composition preferences. We also design a novel interface with functions at the front-end to enable real-time and informative interactions for photo taking. We conduct two user studies to investigate SmartEye qualitatively and quantitatively. Results show that SmartEye effectively models and predicts personalized composition preferences, provides instant high-quality compositions in-situ, and outperforms the non-personalized systems significantly. Shuai Ma 0005, Zijun Wei, Feng Tian 0001, Xiangmin Fan, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Jin Huang 0009, Radomír Mech, Dimitris Samaras, Hongan Wang |
CHI | 2 |
| 2019 | Reading detection in real-timeabstractObservable reading behavior, the act of moving the eyes over lines of text, is highly stereotyped among the users of a language, and this has led to the development of reading detectors-methods that input windows of sequential fixations and output predictions of the fixation behavior during those windows being reading or skimming. The present study introduces a new method for reading detection using Region Ranking SVM (RRSVM). An SVM-based classifier learns the local oculomotor features that are important for real-time reading detection while it is optimizing for the global reading/skimming classification, making it unnecessary to hand-label local fixation windows for model training. This RRSVM reading detector was trained and evaluated using eye movement data collected in a laboratory context, where participants viewed modified web news articles and had to either read them carefully for comprehension or skim them quickly for the selection of keywords (separate groups). Ground truth labels were known at the global level (the instructed reading or skimming task), and obtained at the local level in a separate rating task. The RRSVM reading detector accurately predicted 82.5% of the global (article-level) reading/skimming behavior, with accuracy in predicting local window labels ranging from 72-95%, depending on how tuned the RRSVM was for local and global weights. With this RRSVM reading detector, a method now exists for near real-time reading detection without the need for hand-labeling of local fixation windows. With real-time reading detection capability comes the potential for applications ranging from education and training to intelligent interfaces that learn what a user is likely to know based on previous detection of their reading behavior. Conor Kelton, Zijun Wei, Seoyoung Ahn, Aruna Balasubramanian, Samir Ranjan Das, Dimitris Samaras, Gregory J. Zelinsky |
ETRA | 2 |
| 2019 | Radio-Frequency Power Amplifier Based on CVD Graphene Field-Effect TransistorabstractIn this work, we investigate the performance of radio-frequency (RF) power amplifier based on chemical vapor deposition (CVD) graphene field-effect transistors (GFETs). The GFETs with gate length of 300 nm were fabricated on SiO2/Si substrate, which show an extrinsic current gain cut-off frequency (fT) of 18.6 GHz and an extrinsic maximum oscillation frequency (fmax) of 19.8 GHz. The parameters of compact large-signal model for the GFETs were extracted from the measured direct current (DC) and RF characteristics of GFETs, followed by implementation of the compact model using Verilog-A for circuit simulation. Using electronic design automation (EDA) tools, we designed a GFET-based power amplifier. The power amplifier working at 2.5 GHz shows a gain of ~7 dB, an output power of ~0 dBm (1 mW), a power added efficiency (PAE) of 2.8% and third order intermodulation distortion (IMD) of ~20 dBc, at the 1 dB compression point. Pei Peng 0002, Zidong Wang 0009, Zijun Wei, Zhongzheng Tian, Muchan Li, Liming Ren, Yunyi Fu |
ISCAS | 3 |
| 2018 | Good View Hunting: Learning Photo Composition From Dense View PairsabstractFinding views with good photo composition is a challenging task for machine learning methods. A key difficulty is the lack of well annotated large scale datasets. Most existing datasets only provide a limited number of annotations for good views, while ignoring the comparative nature of view selection. In this work, we present the first large scale Comparative Photo Composition dataset, which contains over one million comparative view pairs annotated using a cost-effective crowdsourcing workflow. We show that these comparative view annotations are essential for training a robust neural network model for composition. In addition, we propose a novel knowledge transfer framework to train a fast view proposal network, which runs at 75+ FPS and achieves state-of-the-art performance in image cropping and thumbnail generation tasks on three benchmark datasets. The superiority of our method is also demonstrated in a user study on a challenging experiment, where our method significantly outperforms the baseline methods in producing diversified well-composed views. Zijun Wei, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Radomír Mech, Minh Hoai, Dimitris Samaras |
CVPR | 1 |
| 2018 | Sequence-to-Segment Networks for Segment DetectionabstractDetecting segments of interest from an input sequence is a challenging problem which often requires not only good knowledge of individual target segments, but also contextual understanding of the entire input sequence and the relationships between the target segments. To address this problem, we propose the Sequence-to-Segment Network (S$^2$N), a novel end-to-end sequential encoder-decoder architecture. S$^2$N first encodes the input into a sequence of hidden states that progressively capture both local and holistic information. It then employs a novel decoding architecture, called Segment Detection Unit (SDU), that integrates the decoder state and encoder hidden states to detect segments sequentially. During training, we formulate the assignment of predicted segments to ground truth as bipartite matching and use the Earth Mover's Distance to calculate the localization errors. We experiment with S$^2$N on temporal action proposal generation and video summarization and show that S$^2$N achieves state-of-the-art performance on both tasks. Zijun Wei, Boyu Wang 0001, Minh Hoai, Jianming Zhang 0001, Zhe Lin 0001, Xiaohui Shen, Radomír Mech, Dimitris Samaras |
NeurIPS | 1 |
| 2016 | Region Ranking SVM for Image ClassificationabstractThe success of an image classification algorithm largely depends on how it incorporates local information in the global decision. Popular approaches such as average-pooling and max-pooling are suboptimal in many situations. In this paper we propose Region Ranking SVM (RRSVM), a novel method for pooling local information from multiple regions. RRSVM exploits the correlation of local regions in an image, and it jointly learns a region evaluation function and a scheme for integrating multiple regions. Experiments on PASCAL VOC 2007, VOC 2012, and ILSVRC2014 datasets show that RRSVM outperforms the methods that use the same feature type and extract features from the same set of local regions. RRSVM achieves similar to or better than the state-of-the-art performance on all datasets. Zijun Wei, Minh Hoai |
CVPR | 1 |
| 2016 | Learned Region Sparsity and Diversity Also Predicts Visual AttentionabstractLearned region sparsity has achieved state-of-the-art performance in classification tasks by exploiting and integrating a sparse set of local information into global decisions. The underlying mechanism resembles how people sample information from an image with their eye movements when making similar decisions. In this paper we incorporate the biologically plausible mechanism of Inhibition of Return into the learned region sparsity model, thereby imposing diversity on the selected regions. We investigate how these mechanisms of sparsity and diversity relate to visual attention by testing our model on three different types of visual search tasks. We report state-of-the-art results in predicting the locations of human gaze fixations, even though our model is trained only on image-level labels without object location annotations. Notably, the classification performance of the extended model remains the same as the original. This work suggests a new computational perspective on visual attention mechanisms and shows how the inclusion of attention-based mechanisms can improve computer vision techniques. Zijun Wei, Hossein Adeli, Minh Hoai, Gregory J. Zelinsky, Dimitris Samaras |
NIPS | 1 |
| 2014 | 3D Pose-by-Detection of Vehicles via Discriminatively Reduced Ensembles of Correlation Filters
Yair Movshovitz-Attias, Yaser Sheikh, Vishnu Naresh Boddeti, Zijun Wei |
BMVC | 4 |