EDBT 2026 Demo / reviewers in the wild / expert
Wei Jiang 0001
dblp:21/3839-1
· DBLP profile ↗
29ranked-venue papers
16as first author
15since 2021 · last 2026
0000-0001-6672-9783ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 16 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCONE: A Practical, Constraint-Aware Plug-in for Latent Encoding in Learned DNA StorageabstractDNA storage has matured from concept to practical stage, yet its integration with neural compression pipelines remains inefficient. Early DNA encoders applied redundancy-heavy constraint layers atop raw binary data-workable but primitive. Recent neural codecs compress data into learned latent representations with rich statistical structure, yet still convert these latents to DNA via naive binary-to-quaternary transcoding, discarding the entropy model's optimization. This mismatch undermines compression efficiency and complicates the encoding stack. A plug-in module that collapses latent compression and DNA encoding into a single step. SCONE performs quaternary arithmetic coding directly on the latent space in DNA bases. Its Constraint-Aware Adaptive Coding module dynamically steers the entropy encoder's learned probability distribution to enforce biochemical constraints-GC balance and homopolymer suppression-deterministically during encoding, eliminating post-hoc correction. The design preserves full reversibility and exploits the hyperprior model's learned priors without modification. Experiments show SCONE achieves near-perfect constraint satisfaction with negligible computational overhead (< 2% latency), establishing a latent-agnostic interface for end-to-end DNA-compatible learned codecs. Cihan Ruan, Lebin Zhou, Rongduo Han, Linyi Han, Bingqing Zhao, Chenchen Zhu, Wei Jiang 0001, Wei Wang 0526, Nam Ling |
ISCAS | 7 |
| 2026 | Salient-Channel-Guided Knowledge Distillation in Learned Image Compression
Jinhao Wang, Cihan Ruan, Nam Ling, Wei Wang 0526, Wei Jiang 0001 |
ISCAS | 5 |
| 2026 | Analysis of Converged 3D Gaussian Splatting Solutions: Density Effects and Prediction LimitsabstractWe investigate what structure emerges in 3D Gaussian Splatting (3DGS) solutions from standard multi-view optimization. We term these Rendering-Optimal References (RORs) and analyze their statistical properties, revealing stable patterns-log-normal scales and bimodal radiance-across diverse scenes. To understand what determines these parameters, we apply learnability probes: training predictors to reconstruct RORs from point clouds without rendering supervision. Our analysis uncovers fundamental density-stratification: dense regions exhibit geometry-correlated parameters amenable to render-free prediction, while sparse regions show systematic failure across architectures. We formalize this through variance decomposition, demonstrating that visibility heterogeneity creates covariance-dominated coupling between geometric and appearance parameters in sparse regions. This reveals RORs' dual character-geometric primitives where point clouds suffice, view synthesis primitives where multi-view constraints are essential. We provide density-aware strategies that improve training robustness and discuss architectural implications for systems that adaptively balance feed-forward prediction and rendering-based refinement. Cihan Ruan, Jingchuan Xiao, Chuqing Shi, Wei Jiang 0001, Nam Ling |
ISCAS | 5 |
| 2026 | $\mathtt {M^3VIR}$ for benchmarking sparse-view novel view synthesis and 3D object removalabstractAbstract The demand for editable high-fidelity 3D content is increasing with the growth of immersive gaming and virtual reality. The $$\mathtt {M^3VIR}$$ dataset was introduced as a large-scale multi-modal benchmark with accurate geometry and semantic annotations. This work extends the use of the dataset to two challenging settings: sparse-view novel-view synthesis (NVS) and 3D object removal. While the original dataset focused on dense observations, practical capture often involves limited viewpoints and scenes that require object editing. Pipelines that rely on Structure-from-Motion (SfM) for initialization frequently become unstable under sparse views and can produce incomplete geometry. We examine whether the annotations provided by $$\mathtt {M^3VIR}$$ , including metric depth, precise camera poses, and instance-level semantic masks, can support reconstruction and editing under these constraints. For sparse-view NVS, we replace SfM initialization with a depth-based TSDF reconstruction that generates a geometric prior for neural rendering methods. This initialization improves stability for both 3D Gaussian Splatting and NeRFacto when only a small number of views are available. For object removal, we evaluate a mask-guided 2D-inpaint-to-3D-lift workflow in which objects are removed in individual views and the edited images are consolidated through multi-view optimization. Experimental results show that geometric and semantic annotations improve reconstruction stability in sparse-view settings and enable consistent scene editing after object removal. The study establishes baseline evaluation protocols for geometry-guided reconstruction and editing on multi-modal datasets. Yuanzhi Li, Lebin Zhou, Nam Ling, Wei Jiang 0001 |
Multim. Tools Appl. | 6 |
| 2025 | HybridFlow-DNA: A Deep Generative Compression Framework for DNA Storage of ImagesabstractDNA storage has emerged as a promising solution to address the exponentially growing demand for storage capacity, offering advantages in density, stability, and long-term preservation potential. Currently, image compression for DNA storage has evolved into learned image compression (LIC), particularly through the application of deep learning methods based on artificial neural networks. The present study proposes a novel image compression framework for DNA Storage, named HybridFlow-DNA. HybridFlow-DNA is established by integration of VQGAN and MLIC with the adaptive dynamic DNA fountain encoding scheme. Experimental results demonstrate that HybridFlow-DNA achieves a high virtual information capacity while effectively maintaining the fidelity of the reconstruction of images. Cihan Ruan, Rongduo Han, Wei Jiang 0001, Wei Wang 0311, Nam Ling |
ISCAS | 5 |
| 2025 | (RichMediaGAI'25) 3rd International Workshop on Rich Media with Generative AIabstractThe goal of this workshop is to showcase the latest advancements in generative AI (GAI) for creating, editing, restoring, and compressing rich media data, including images, videos, and 3D content. GAI models such as VAEs, GANs, and diffusion models have demonstrated remarkable impact in both academic research and industrial applications. For example, GAI enables users to design and generate synthetic yet realistic content without requiring professional artistic or technical expertise, driving significant market growth in gaming and entertainment. Beyond creative applications, GAI also provides crucial simulated data for training embodied AI agents. When applied to media restoration and synthesis, GAI techniques can further alleviate transmission challenges by offloading computation to client devices. To advance this field, the workshop will host four competition tracks using novel industry-level data, solicit high-quality paper submissions, and invite leading speakers from academia and industry to foster collaboration and innovation. In particular, the competition focuses on media generation and transmission with GAI. The first three tracks address reducing computation and transmission costs for efficient media delivery, while the fourth track focuses on controlled novel content creation. To support these challenges, a large-scale multi-modality, multi-view dataset named M3VIR is provided. This dataset comprises a diverse collection of videos simulated using the UE5 Unreal Engine, with carefully matched content serving as ground truth for the competition tasks. Wei Jiang 0001, Dong Xu 0001 |
ACM Multimedia | 1 |
| 2025 | HDCompression: Hybrid-Diffusion Image Compression for Ultra-low Bitrates
Yanzhi Wang 0001, Wei Wang 0526, Wei Jiang 0001 |
PRICAI (5) | 5 |
| 2024 | HybridFlow: Infusing Continuity into Masked Codebook for Extreme Low-Bitrate Image CompressionabstractThis paper investigates the challenging problem of learned image compression (LIC) with extreme low bitrates. Previous LIC methods based on transmitting quantized continuous features often yield blurry and noisy reconstruction due to the severe quantization loss. While previous LIC methods based on learned codebooks that discretize visual space usually give poor-fidelity reconstruction due to the insufficient representation power of limited codewords in capturing faithful details. We propose a novel dual-stream framework, HyrbidFlow, which combines the continuous-feature-based and codebook-based streams to achieve both high perceptual quality and high fidelity under extreme low bitrates. The codebook-based stream benefits from the high-quality learned codebook priors to provide high quality and clarity in reconstructed images. The continuous feature stream targets at maintaining fidelity details. To achieve the ultra low bitrate, a masked token-based transformer is further proposed, where we only transmit a masked portion of codeword indices and recover the missing indices through token generation guided by information from the continuous feature stream. We also develop a bridging correction network to merge the two streams in pixel decoding for final image reconstruction, where the continuous stream features rectify biases of the codebook-based pixel decoder to impose reconstructed fidelity details. Experimental results demonstrate superior performance across several datasets under extremely low bitrates, compared with existing single-stream codebook-based or continuous-feature-based LIC methods. Yanyue Xie, Wei Jiang 0001, Wei Wang 0311, Xue Lin 0001, Yanzhi Wang 0001 |
ACM Multimedia | 3 |
| 2024 | Neural Image Compression Using Masked Sparse Visual RepresentationabstractWe study neural image compression based on the Sparse Visual Representation (SVR), where images are embedded into a discrete latent space spanned by learned visual codebooks. By sharing codebooks with the decoder, the encoder transfers integer codeword indices that are efficient and cross-platform robust, and the decoder retrieves the embedded latent feature using the indices for reconstruction. Previous SVR-based compression lacks effective mechanism for rate-distortion tradeoffs, where one can only pursue either high reconstruction quality or low transmission bitrate. We propose a Masked Adaptive Codebook learning (M-AdaCode) method that applies masks to the latent feature subspace to balance bitrate and reconstruction quality. A set of semantic-class-dependent basis codebooks are learned, which are weighted combined to generate a rich latent feature for high-quality reconstruction. The combining weights are adaptively derived from each input image, providing fidelity information with additional transmission costs. By masking out unimportant weights in the encoder and recovering them in the decoder, we can trade off reconstruction quality for transmission bits, and the masking rate controls the balance between bitrate and distortion. Experiments over the standard JPEG-AI dataset demonstrate the effectiveness of our M-AdaCode approach. Wei Jiang 0001, Wei Wang 0311 |
WACV | 1 |
| 2023 | FVC: An End-to-End Framework Towards Deep Video Compression in Feature SpaceabstractDeep video compression is attracting increasing attention from both deep learning and video processing community. Recent learning-based approaches follow the hybrid coding paradigm to perform pixel space operations for reducing redundancy along both spatial and temporal dimentions, which leads to inaccurate motion estimation or less effective motion compensation. In this work, we propose a feature-space video coding framework (FVC), which performs all major operations (i.e., motion estimation, motion compression, motion compensation and residual compression) in the feature space. Specifically, a new deformable compensation module, which consists of motion estimation, motion compression and motion compensation, is proposed for more effective motion compensation. In our deformable compensation module, we first perform motion estimation in the feature space to produce the motion information (i.e., the offset maps). Then the motion information is compressed by using the auto-encoder style network. After that, we use the deformable convolution operation to generate the predicted feature for motion compensation. Finally, the residual information between the feature from the current frame and the predicted feature from the deformable compensation module is also compressed in the feature space. Motivated by the conventional codecs, in which the blocks with different sizes are used for motion estimation, we additionally propose two new modules called resolution-adaptive motion coding (RaMC) and resolution-adaptive residual coding (RaRC) to automatically cope with different types of motion and residual patterns at different spatial locations. Comprehensive experimental results demonstrate that our proposed framework achieves the state-of-the-art performance on three benchmark datasets including HEVC, UVG and MCL-JCV. Dong Xu 0001, Guo Lu, Wei Jiang 0001, Wei Wang 0311, Shan Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | LSVC: A Learning-based Stereo Video Compression FrameworkabstractIn this work, we propose the first end-to-end optimized framework for compressing automotive stereo videos (i.e., stereo videos from autonomous driving applications) from both left and right views. Specifically, when compressing the current frame from each view, our framework reduces temporal redundancy by performing motion compensation using the reconstructed intra-view adjacent frame and at the same time exploits binocular redundancy by conducting disparity compensation using the latest reconstructed cross-view frame. Moreover, to effectively compress the introduced motion and disparity offsets for better compensation, we further propose two novel schemes called motion residual compression and disparity residual compression to respectively generate the predicted motion offset and disparity offset from the previously compressed motion offset and disparity offset, such that we can more effectively compress residual offset information for better bit-rate saving. Overall, the entire framework is implemented by the fully-differentiable modules and can be optimized in an end-to-end manner. Our comprehensive experiments on three automotive stereo video benchmarks Cityscapes, KITTI 2012 and KITTI 2015 demonstrate that our proposed framework outperforms the learning-based single-view video codec and the traditional hand-crafted multi-view video codec. Guo Lu, Shan Liu 0001, Wei Jiang 0001, Dong Xu 0001 |
CVPR | 5 |
| 2022 | Coarse-To-Fine Deep Video Coding with Hyperprior-Guided Mode PredictionabstractThe previous deep video compression approaches only use the single scale motion compensation strategy and rarely adopt the mode prediction technique from the traditional standards like H.264/H.265 for both motion and residual compression. In this work, we first propose a coarse-to-fine (C2F) deep video compression framework for better motion compensation, in which we perform motion estimation, compression and compensation twice in a coarse to fine manner. Our C2F framework can achieve better motion compensation results without significantly increasing bit costs. Observing hyperprior information (i.e., the mean and variance values) from the hyperprior networks contains discriminant statistical information of different patches, we also propose two efficient hyperprior-guided mode prediction methods. Specifically, using hyper-prior information as the input, we propose two mode prediction networks to respectively predict the optimal block resolutions for better motion coding and decide whether to skip residual information from each block for better residual coding without introducing additional bit cost while bringing negligible extra computation cost. Comprehensive experimental results demonstrate our proposed C2F video compression framework equipped with the new hyperprior-guided mode prediction methods achieves the state-of-the-art performance on HEVC, UVG and MCL-JCV datasets. Guo Lu, Jinyang Guo 0002, Shan Liu 0001, Wei Jiang 0001, Dong Xu 0001 |
CVPR | 5 |
| 2022 | Hardware-Friendly Acceleration for Deep Neural Networks with Micro-Structured CompressionabstractDeep Neural Network (DNN) compression techniques including weight pruning and quantization have made great success in reducing the amount of model parameters and computations for various applications. However, the existing studies hardly consider two critical targets jointly, i.e., enhancing the computation and resource utilization efficiency that is essential for DNN acceleration on hardware, and at the same time maintaining the original model performance, such as the accuracy in classification tasks, or the peak signal-to-noise ratio (PSNR) in super resolution tasks. Approaches like coarse-grained structured (filter, channel, etc.) pruning and low-precision (binary, ternary, fixed-point with 4-bit or less) quantization suffer from non-negligible accuracy loss, and unstructured pruning incurs extra indexing overhead and degradation in computation parallelism. Mengshu Sun, Sheng Lin 0001, Shan Liu 0001, Songnan Li, Yanzhi Wang 0001, Wei Jiang 0001, Wei Wang 0311 |
FCCM | 6 |
| 2022 | Substitutional Neural Image CompressionabstractWe describe Substitutional Neural Image Compression (SNIC), a general approach for enhancing any neural image compression model, that requires no data or additional tuning of the trained model. It boosts compression performance toward a flexible distortion metric and enables bit-rate control using a single model instance. The key idea is to replace the image to be compressed with a substitutional one that outperforms the original one in a desired way. Finding such a substitute is inherently difficult for conventional codecs, yet surprisingly favorable for neural compression models thanks to their fully differentiable structures. With gradients of a particular loss back-propogated to the input, a desired substitute can be efficiently crafted iteratively. We demonstrate the effectiveness of SNIC, when combined with various neural compression models and target metrics, in improving compression quality and performing bit-rate control measured by rate-distortion curves. Xiao Wang 0028, Ding Ding 0004, Wei Jiang 0001, Wei Wang 0311, Xiaozhong Xu, Shan Liu 0001, Brian Kulis, Sang (Peter) Chin |
PCS | 3 |
| 2022 | Overview of the Neural Network Compression and Representation (NNR) StandardabstractNeural Network Coding and Representation (NNR) is the first international standard for efficient compression of neural networks (NNs). The standard is designed as a toolbox of compression methods, which can be used to create coding pipelines. It can be either used as an independent coding framework (with its own bitstream format) or together with external neural network formats and frameworks. For providing the highest degree of flexibility, the network compression methods operate per parameter tensor in order to always ensure proper decoding, even if no structure information is provided. The NNR standard contains compression-efficient quantization and deep context-adaptive binary arithmetic coding (DeepCABAC) as core encoding and decoding technologies, as well as neural network parameter pre-processing methods like sparsification, pruning, low-rank decomposition, unification, local scaling and batch norm folding. NNR achieves a compression efficiency of more than 97% for transparent coding cases, i.e. without degrading classification quality, such as top-1 or top-5 accuracies. This paper provides an overview of the technical features and characteristics of NNR. Heiner Kirchhoffer, Paul Haase, Wojciech Samek, Karsten Müller 0001, Hamed Rezazadegan Tavakoli, Francesco Cricri, Emre Aksu, Miska M. Hannuksela, Wei Jiang 0001, Wei Wang 0311, Shan Liu 0001, Swayambhoo Jain, Shahab Hamidi-Rad, Fabien Racapé, Werner Bailer |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2017 | Face detection and recognition for home service robots with end-to-end deep neural networksabstractThis paper proposes an effective end-to-end face detection and recognition framework based on deep convolutional neural networks for home service robots. We combine the state-of-the-art region proposal based deep detection network with the deep face embedding network into an end-to-end system, so that the detection and recognition networks can share the same deep convolutional layers, enabling significant reduction of computation through sharing convolutional features. The detection network is robust to large occlusion, and scale, pose, and lighting variations. The recognition network does not require explicit face alignment, which enables an effective training strategy to generate a unified network. A practical robot system is also developed based on the proposed framework, where the system automatically asks for a minimum level of human supervision when needed, and no complicated region-level face annotation is required. Experiments are conducted over WIDER and LFW benchmarks, as well as a personalized dataset collected from an office setting, which demonstrate state-of-the-art performance of our system. Wei Jiang 0001, Wei Wang 0311 |
ICASSP | 1 |
| 2012 | Grouplet-Based Distance Metric Learning for Video Concept DetectionabstractWe investigate general concept detection in unconstrained videos. A distance metric learning algorithm is developed to use the information of the group let structure for improved detection. A group let is defined as a set of audio and/or visual code words that are grouped together according to their strong correlations in videos. By using the entire group lets as building elements, concepts can be more robustly detected than using discrete audio or visual code words. Compared with the traditional method of generating aggregated group let-based features for classification, our group let-based distance metric learning approach directly learns distances between data points, which better preserves the group let structure. Specifically, our algorithm uses an iterative quadratic programming formulation where the optimal distance metric can be effectively learned based on the large-margin nearest-neighbor setting. The framework is quite flexible, where various types of distances can be computed using individual group lets, and through the same distance metric learning algorithm the distances computed over individual group lets can be combined for final classification. We extensively evaluate our method over the large-scale Columbia Consumer Video set. Experiments demonstrate that our approach can achieve consistent and significant performance improvements. Wei Jiang 0001, Alexander C. Loui |
ICME | 1 |
| 2011 | Automatic consumer video summarization by audio and visual analysisabstractVideo summarization provides a condensed version of a video stream by analyzing the video content. Automatic summarization of consumer videos is an important tool that facilitates efficient browsing, searching, and album creation in large consumer video collections. This paper studies automatic video summarization in the consumer domain where most previous methods cannot be easily applied due to the challenging issues for content analysis, i.e., consumer videos are captured with uncontrolled conditions such as uneven illumination, clutter, and large camera motion, and with poor-quality soundtrack as a mix of multiple sound sources under severe noise. To pursue reliable summarization, a case study with actual consumer users is conducted, from which a set of consumer-oriented guidelines is obtained. The guidelines reflect the high-level semantic rules, in both visual and audio aspects, which are recognized by consumers as important to produce good video summaries. Following these guidelines, an automatic video summarization algorithm is developed where both visual and audio information are used to generate improved summaries. To the best of our knowledge, this is a first systematic study on automatic summarization of consumer-quality videos. Experimental evaluations from consumer subjects show the effectiveness of our approach. Wei Jiang 0001, Courtenay V. Cotton, Alexander C. Loui |
ICME | 1 |
| 2011 | Audio-visual grouplet: temporal audio-visual interactions for general video concept classificationabstractWe investigate general concept classification in unconstrained videos by joint audio-visual analysis. A novel representation, the Audio-Visual Grouplet (AVG), is extracted by studying the statistical temporal audio-visual interactions. An AVG is defined as a set of audio and visual codewords that are grouped together according to their strong temporal correlations in videos. The AVGs carry unique audio-visual cues to represent the video content, based on which an audio-visual dictionary can be constructed for concept classification. By using the entire AVGs as building elements, the audio-visual dictionary is much more robust than traditional vocabularies that use discrete audio or visual codewords. Specifically, we conduct coarse-level foreground/background separation in both audio and visual channels, and discover four types of AVGs by exploring mixed-and-matched temporal audio-visual correlations among the following factors: visual foreground, visual background, audio foreground, and audio background. All of these types of AVGs provide discriminative audio-visual patterns for classifying various semantic concepts. We extensively evaluate our method over the large-scale Columbia Consumer Video set. Experiments demonstrate that the AVG-based dictionaries can achieve consistent and significant performance improvements compared with other state-of-the-art approaches. Wei Jiang 0001, Alexander C. Loui |
ACM Multimedia | 1 |
| 2010 | Automatic aesthetic value assessment in photographic imagesabstractThe automatic assessment of aesthetic values in consumer photographic images is an important issue for content management, organizing and retrieving images, and building digital image albums. This paper explores automatic aesthetic estimation in two different tasks: (1) to estimate fine-granularity aesthetic scores ranging from 0 to 100, a novel regression method, namely Diff-RankBoost, is proposed based on RankBoost and support vector techniques; and (2) to predict coarse-granularity aesthetic categories (e.g., visually “very pleasing” or “not pleasing”), multi-category classifiers are developed. A set of visual features describing various characteristics related to image quality and aesthetic values are used to generate multidimensional feature spaces for aesthetic estimation. Experiments over a consumer photographic image collection with user ground-truth indicate that the proposed algorithms provide promising results for automatic image aesthetic assessment. Wei Jiang 0001, Alexander C. Loui, Cathleen Daniels Cerosaletti |
ICME | 1 |
| 2010 | Audio-visual atoms for generic video concept classificationabstractWe investigate the challenging issue of joint audio-visual analysis of generic videos targeting at concept detection. We extract a novel local representation, Audio-Visual Atom (AVA), which is defined as a region track associated with regional visual features and audio onset features. We develop a hierarchical algorithm to extract visual atoms from generic videos, and locate energy onsets from the corresponding soundtrack by time-frequency analysis. Audio atoms are extracted around energy onsets. Visual and audio atoms form AVAs, based on which discriminative audio-visual codebooks are constructed for concept detection. Experiments over Kodak's consumer benchmark videos confirm the effectiveness of our approach. Wei Jiang 0001, Courtenay V. Cotton, Shih-Fu Chang, Daniel P. W. Ellis, Alexander C. Loui |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2009 | Visual saliency with side informationabstractWe propose novel algorithms for organizing large image and video datasets using both the visual content and the associated side-information, such as time, location, authorship, and so on. Earlier research have used side-information as pre-filter before visual analysis is performed, and we design a machine learning algorithm to model the join statistics of the content and the side information. Our algorithm, diverse-density contextual clustering (D2C2), starts by finding unique patterns for each sub-collection sharing the same side-info, e.g., scenes from winter. It then finds the common patterns that are shared among all subsets, e.g., persistent scenes across all seasons. These unique and common prototypes are found with multiple instance learning and subsequent clustering steps. We evaluate D2C2on two Web photo collections from Flickr and one news video collection from TRECVID. Results show that not only the visual patterns found by D2C2are intuitively salient across different seasons, locations and events, classifiers constructed from the unique and common patterns also outperform state-of-the-art bag-of-features classifiers. Wei Jiang 0001, Lexing Xie, Shih-Fu Chang |
ICASSP | 1 |
| 2009 | Short-term audio-visual atoms for generic video concept classificationabstractWe investigate the challenging issue of joint audio-visual analysis of generic videos targeting at semantic concept detection. We propose to extract a novel representation, the Short-term Audio-Visual Atom (S-AVA), for improved concept detection. An S-AVA is defined as a short-term region track associated with regional visual features and background audio features. An effective algorithm, named Short-Term Region tracking with joint Point Tracking and Region Segmentation (STR-PTRS), is developed to extract S-AVAs from generic videos under challenging conditions such as uneven lighting, clutter, occlusions, and complicated motions of both objects and camera. Discriminative audio-visual codebooks are constructed on top of S-AVAs using Multiple Instance Learning. Codebook-based features are generated for semantic concept detection. We extensively evaluate our algorithm over Kodak's consumer benchmark video set from real users. Experimental results confirm significant performance improvements - over 120% MAP gain compared to alternative approaches using static region segmentation without temporal tracking. The joint audio-visual features also outperform visual features alone by an average of 8.5% (in terms of AP) over 21 concepts, with many concepts achieving more than 20%. Wei Jiang 0001, Courtenay V. Cotton, Shih-Fu Chang, Daniel P. W. Ellis, Alexander C. Loui |
ACM Multimedia | 1 |
| 2008 | Semantic Concept Classification by Joint Semi-supervised Learning of Feature Subspaces and Support Vector Machines
Wei Jiang 0001, Shih-Fu Chang, Tony Jebara, Alexander C. Loui |
ECCV (4) | 1 |
| 2008 | Cross-domain learning methods for high-level visual concept classificationabstractExploding amounts of multimedia data increasingly require automatic indexing and classification, e.g. training classifiers to produce high-level features, or semantic concepts, chosen to represent image content, like car, person, etc. When changing the applied domain (i.e. from news domain to consumer home videos), the classifiers trained in one domain often perform poorly in the other domain due to changes in feature distributions. Additionally, classifiers trained on the new domain alone may suffer from too few positive training samples. Appropriately adapting data/models from an old domain to help classify data in a new domain is an important issue. In this work, we develop a new cross-domain SVM (CDSVM) algorithm for adapting previously learned support vectors from one domain to help classification in another domain. Better precision is obtained with almost no additional computational cost. Also, we give a comprehensive summary and comparative study of the state-of-the-art SVM-based cross-domain learning methods. Evaluation over the latest large-scale TRECVID benchmark data set shows that our CDSVM method can improve mean average precision over 36 concepts by 7.5%. For further performance gain, we also propose an intuitive selection criterion to determine which cross-domain learning method to use for each concept. Wei Jiang 0001, Eric Zavesky, Shih-Fu Chang, Alexander C. Loui |
ICIP | 1 |
| 2008 | Semantic event detection for consumer photo and video collectionsabstractThe automatic detection of semantic events in userspsila image and video collections is an important technique for content management and retrieval. In this paper we propose a novel semantic event detection approach by considering an event-level bag-of-features (BOF) representation to model typical events. Based on this BOF representation, semantic events are detected in a concept space instead of the original low-level visual feature space. There are two advantages of our approach: we can avoid the sensitivity problem by decreasing the influence of difficult or erroneous images or videos in measuring the event-level similarity; also we can utilize the power of higher-level concept scores in describing semantic events. Experiments over a large real consumer database confirm the effectiveness of our approach. Wei Jiang 0001, Alexander C. Loui |
ICME | 1 |
| 2007 | Kernel Sharing With Joint Boosting For Multi-Class Concept DetectionabstractObject/scene detection by discriminative kernel-based classification has gained great interest due to its promising performance and flexibility. In this paper, unlike traditional approaches that independently build binary classifiers to detect individual concepts, we proposed a new framework for multi-class concept detection based on kernel sharing and joint learning. By sharing "good" kernels among concepts, accuracy of individual weak detectors can be greatly improved; by joint learning of common detectors among classes, the required kernels and the computational complexity for detecting each individual concept can be reduced. We demonstrated our approach by developing an extended JointBoost framework, which was used to choose the optimal kernel and subset of sharing classes in an iterative boosting process. In addition, we constructed multi-resolution visual vocabularies by hierarchical clustering and computed kernels based on spatial matching. We tested our method in detecting 12 concepts (objects, scenes, etc) over 80+ hours of broadcast news videos from the challenging TRECVID 2005 corpus. Significant performance gains were achieved -10% in mean average precision (MAP) and up to 34% average precision (AP) for some concepts like maps, building, and boat-ship. Extensive analysis of the results also revealed interesting and important underlying relations among concepts. Wei Jiang 0001, Shih-Fu Chang, Alexander C. Loui |
CVPR | 1 |
| 2007 | Context-Based Concept Fusion with Boosted Conditional Random FieldsabstractThe contextual relationships among different semantic concepts provide important information for automatic concept detection in images/videos. We propose a new context-based concept fusion (CBCF) method for semantic concept detection. Our work includes two folds. (1) We model the inter-conceptual relationships by a conditional random field (CRF) that improves detection results from independent detectors by taking into account the inter-correlation among concepts. CRF directly models the posterior probability of concept labels and is more accurate for the discriminative concept detection than previous statistical inferencing techniques. The boosted CRF framework is incorporated to further enhance performance by combining the power of boosting with CRF. (2) We develop an effective criterion to predict which concepts may benefit from CBCF. As reported in previous works, CBCF has inconsistent performance gain on different concepts. With accurate prediction, computational and data resources can be allocated to enhance concepts that are promising to gain performance. Evaluation on TRECVID2005 development set demonstrates the effectiveness of our algorithm. Wei Jiang 0001, Shih-Fu Chang, Alexander C. Loui |
ICASSP (1) | 1 |
| 2006 | Active Context-Based Concept Fusionwith Partial User LabelsabstractIn this paper we propose a new framework, called active context-based concept fusion, for effectively improving the accuracy of semantic concept detection in images and videos. Our approach solicits user annotations for a small number of concepts, which are used to refine the detection of the rest of concepts. In contrast with conventional methods, our approach is active, by using information theoretic criteria to automatically determine the optimal concepts for user annotation. Our experiments over TRECVID 2005 development set (about 80 hours) show significant performance gains. In addition, we have developed an effective method to predict concepts that may benefit from context-based fusion. Wei Jiang 0001, Shih-Fu Chang, Alexander C. Loui |
ICIP | 1 |