Ning Bi

dblp:13/3540 · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AortaDiff: A Unified Multitask Diffusion Framework for Contrast-Free AAA Imaging
abstract
While contrast-enhanced CT (CECT) is standard for assessing abdominal aortic aneurysms (AAA) , the required iodinated contrast agents pose significant risks, including patient allergies and environmental harm. To reduce contrast agent use, methods have focused on generating synthetic CECT from non-contrast CT (NCCT) scans. However, most adopt a multi-stage pipeline that first generates images and then performs segmentation, which leads to error accumulation and fails to leverage shared information. To address this, we propose a unified framework that generates synthetic CECT images from NCCT scans while simultaneously segmenting the aortic lumen and thrombus. Our approach integrates conditional diffusion models (CDM) with multi-task learning, enabling end-to-end joint optimization of image synthesis and anatomical segmentation. Unlike previous multitask diffusion models, our approach requires no initial predictions (e.g., a coarse segmentation mask), shares both encoder and decoder parameters across tasks, and employs a semi-supervised training strategy to learn from scans with missing segmentation labels, a common constraint in clinical data. Evaluated on a cohort of 264 patients, our method consistently outperformed state-of-the-art single-task and multi-stage models. For image synthesis, it achieved a PSNR of 25.61 dB, compared to 23.80 dB from a single-task CDM. For segmentation, it improved the lumen Dice score to 0.89 from 0.87 and the challenging thrombus Dice score to 0.53 from 0.48 (nnU-Net). These segmentation enhancements led to more accurate clinical measurements, reducing the lumen diameter MAE to 4.19 mm from 5.78 mm and the thrombus area error to 33.85% from 41.45%. Code is at https://github.com/yuxuanou623/AortaDiff.
Yuxuan Ou, Ning Bi, Jiazheng Pan, Jiancheng Yang, Boliang Yu, Usama Zidan, Regent Lee, Vicente Grau
WACV2
2025 EgoPrivacy: What Your First-Person Camera Says About You?
abstract
While the rapid proliferation of wearable cameras has raised significant concerns about egocentric video privacy, prior work has largely overlooked the unique privacy threats posed to the camera wearer. This work investigates the core question: How much privacy information about the camera wearer can be inferred from their first-person view videos? We introduce EgoPrivacy, the first large-scale benchmark for the comprehensive evaluation of privacy risks in egocentric vision. EgoPrivacy covers three types of privacy (demographic, individual, and situational), defining seven tasks that aim to recover private information ranging from fine-grained (e.g., wearer's identity) to coarse-grained (e.g., age group). To further emphasize the privacy threats inherent to egocentric vision, we propose Retrieval-Augmented Attack, a novel attack strategy that leverages ego-to-exo retrieval from an external pool of exocentric videos to boost the effectiveness of demographic privacy attacks. An extensive comparison of the different attacks possible under all threat models is presented, showing that private information of the wearer is highly susceptible to leakage. For instance, our findings indicate that foundation models can effectively compromise wearer privacy even in zero-shot settings by recovering attributes such as identity, scene, gender, and race with 70–80% accuracy. Our code and data are available at https://github.com/williamium3000/ego-privacy.
Yijiang Li, Genpei Zhang, Yi Li 0051, Xiaojun Shan, Dashan Gao 0001, Jiancheng Lyu, Ning Bi, Nuno Vasconcelos
ICML9
2025 SegMorph: Concurrent Motion Estimation and Segmentation for Cardiac MRI Sequences
abstract
We propose a novel recurrent variational netwo8=k]irk, SegMorph, to perform concurrent segmentation and motion estimation on cardiac cine magnetic resonance image (CMR) sequences. Our model establishes a recurrent latent space that captures spatiotemporal features from cine-MRI sequences for multitask inference and synthesis. The proposed model follows a recurrent variational auto-encoder framework and adopts a learnt prior from the temporal inputs. We utilise a multi-branch decoder to handle bi-ventricular segmentation and motion estimation simultaneously. In addition to the spatiotemporal features from the latent space, motion estimation enriches the supervision of sequential segmentation tasks by providing pseudo-ground truth. On the other hand, the segmentation branch helps with motion estimation by predicting deformation vector fields (DVFs) based on anatomical information. Experimental results demonstrate that the proposed method performs better than state-of-the-art approaches qualitatively and quantitatively for both segmentation and motion estimation tasks. We achieved an 81% average Dice Similarity Coefficient (DSC) and a less than 3.5 mm average Hausdorff distance on segmentation. Meanwhile, we achieved a motion estimation Dice Similarity Coefficient of over 79%, with approximately 0.14% of pixels displaying a negative Jacobian determinant in the estimated DVFs.
Ning Bi, Arezoo Zakeri, Yan Xia 0002, Nina Cheng, Alejandro F. Frangi, Ali Gooya
IEEE Trans. Medical Imaging1
2024 AUEditNet: Dual-Branch Facial Action Unit Intensity Manipulation with Implicit Disentanglement
abstract
Facial action unit (AU) intensity plays a pivotal role in quantifying fine-grained expression behaviors, which is an effective condition for facial expression manipulation. How-ever, publicly available datasets containing intensity annotations for multiple AUs remain severely limited, often featuring a restricted number of subjects. This limitation places challenges to the AU intensity manipulation in images due to disentanglement issues, leading researchers to resort to other large datasets with pretrained AU intensity estimators for pseudo labels. In addressing this constraint and fully leveraging manual annotations of AU intensities for precise manipulation, we introduce AUEditNet. Our proposed model achieves impressive intensity manipulation across 12 AUs, trained effectively with only 18 subjects. Utilizing a dual-branch architecture, our approach achieves comprehensive disentanglement of facial attributes and identity without necessitating additional loss functions or implementing with large batch sizes. This approach offers a potential solution to achieve desired facial attribute editing despite the dataset's limited subject count. Our experiments demonstrate AUEdit-Net's superior accuracy in editing AU intensities, affirming its capability in disentangling facial attributes and identity within a limited subject pool. AUEditNet allows conditioning by either intensity values or target images, eliminating the need for constructing AU combinations for specific facial expression synthesis. Moreover, AU intensity estimation, as a downstream task, validates the consistency between real and edited images, confirming the effectiveness of our proposed AU intensity manipulation method.
Shiwei Jin, Zhen Wang 0009, Lei Wang 0018, Peng Liu 0039, Ning Bi, Truong Q. Nguyen
CVPR5
2023 ReDirTrans: Latent-to-Latent Translation for Gaze and Head Redirection
abstract
Learning-based gaze estimation methods require large amounts of training data with accurate gaze annotations. Facing such demanding requirements of gaze data collection and annotation, several image synthesis methods were proposed, which successfully redirected gaze directions pre-cisely given the assigned conditions. However, these methods focused on changing gaze directions of the images that only include eyes or restricted ranges of faces with low res-olution (less than$128\times 128$) to largely reduce interference from other attributes such as hairs, which limits application scenarios. To cope with this limitation, we proposed a portable network, called ReDirTrans, achieving latent-to-latent translation for redirecting gaze directions and head orientations in an interpretable manner. ReDirTrans projects input latent vectors into aimed-attribute embed-dings only and redirects these embeddings with assigned pitch and yaw values. Then both the initial and edited embeddings are projected back (deprojected) to the initial latent space as residuals to modify the input latent vec-tors by subtraction and addition, representing old status re-moval and new status addition. The projection of aimed at-tributes only and subtraction-addition operations for status replacement essentially mitigate impacts on other attributes and the distribution of latent vectors. Thus, by combining ReDirTrans with a pretrained fixed e4e-StyleGAN pair, we created ReDirTrans-GAN, which enables accurately redi-recting gaze in full-face images with$1024\times 1024$resolution while preserving other attributes such as identity, expres-sion, and hairstyle. Furthermore, we presented improvements for the downstream learning-based gaze estimation task, using redirected samples as dataset augmentation.
Shiwei Jin, Zhen Wang 0009, Lei Wang 0018, Ning Bi, Truong Q. Nguyen
CVPR4
2023 Relative Position Embedding Asymmetric Siamese Network for Offline Handwritten Mathematical Expression recognition
Hurunqi Luo, Xiaqing Rao, Ning Bi, Jun Tan 0001
ICDAR (1)5
2023 GSMorph: Gradient Surgery for Cine-MRI Cardiac Deformable Registration
Haoran Dou, Ning Bi, Luyi Han, Yuhao Huang 0001, Ritse Mann, Xin Yang 0009, Dong Ni 0001, Nishant Ravikumar, Alejandro F. Frangi, Yunzhi Huang
MICCAI (10)2
2023 PVP: Personalized Video Prior for Editable Dynamic Portraits using StyleGAN
abstract
Abstract Portrait synthesis creates realistic digital avatars which enable users to interact with others in a compelling way. Recent advances in StyleGAN and its extensions have shown promising results in synthesizing photorealistic and accurate reconstruction of human faces. However, previous methods often focus on frontal face synthesis and most methods are not able to handle large head rotations due to the training data distribution of StyleGAN. In this work, our goal is to take as input a monocular video of a face, and create an editable dynamic portrait able to handle extreme head poses. The user can create novel viewpoints, edit the appearance, and animate the face. Our method utilizes pivotal tuning inversion (PTI) to learn a personalized video prior from a monocular video sequence. Then we can input pose and expression coefficients to MLPs and manipulate the latent vectors to synthesize different viewpoints and expressions of the subject. We also propose novel loss functions to further disentangle pose and expression in the latent space. Our algorithm shows much better performance over previous approaches on monocular video datasets, and it is also capable of running in real‐time at 54 FPS on an RTX 3080.
Kai-En Lin, Alex Trevithick, Ke-Li Cheng, Michel Sarkis, Mohsen Ghafoorian, Ning Bi, Gerhard Reitmayr, Ravi Ramamoorthi
Comput. Graph. Forum6
2023 DragNet: Learning-based deformable registration for realistic cardiac MR sequence generation from a single frame
abstract
Deformable image registration (DIR) can be used to track cardiac motion. Conventional DIR algorithms aim to establish a dense and non-linear correspondence between independent pairs of images. They are, nevertheless, computationally intensive and do not consider temporal dependencies to regulate the estimated motion in a cardiac cycle. In this paper, leveraging deep learning methods, we formulate a novel hierarchical probabilistic model, termed DragNet, for fast and reliable spatio-temporal registration in cine cardiac magnetic resonance (CMR) images and for generating synthetic heart motion sequences. DragNet is a variational inference framework, which takes an image from the sequence in combination with the hidden states of a recurrent neural network (RNN) as inputs to an inference network per time step. As part of this framework, we condition the prior probability of the latent variables on the hidden states of the RNN utilised to capture temporal dependencies. We further condition the posterior of the motion field on a latent variable from hierarchy and features from the moving image. Subsequently, the RNN updates the hidden state variables based on the feature maps of the fixed image and the latent variables. Different from traditional methods, DragNet performs registration on unseen sequences in a forward pass, which significantly expedites the registration process. Besides, DragNet enables generating a large number of realistic synthetic image sequences given only one frame, where the corresponding deformations are also retrieved. The probabilistic framework allows for computing spatio-temporal uncertainties in the estimated motion fields. Our results show that DragNet performance is comparable with state-of-the-art methods in terms of registration accuracy, with the advantage of offering analytical pixel-wise motion uncertainty estimation across a cardiac cycle and being a motion generator. We will make our code publicly available.
Arezoo Zakeri, Alireza Hokmabadi, Ning Bi, Isuru Wijesinghe, Michael G. Nix, Steffen E. Petersen, Alejandro F. Frangi, Zeike A. Taylor, Ali Gooya
Medical Image Anal.3
2022 CCLSL: Combination of Contrastive Learning and Supervised Learning for Handwritten Mathematical Expression Recognition
Qiqiang Lin, Xiaonan Huang, Ning Bi, Ching Y. Suen, Jun Tan 0001
ACCV (2)3
2022 Face Relighting with Geometrically Consistent Shadows
abstract
Most face relighting methods are able to handle diffuse shadows, but struggle to handle hard shadows, such as those cast by the nose. Methods that propose techniques for handling hard shadows often do not produce geometrically consistent shadows since they do not directly leverage the estimated face geometry while synthesizing them. We propose a novel differentiable algorithm for synthesizing hard shadows based on ray tracing, which we incorporate into training our face relighting model. Our proposed algorithm directly utilizes the estimated face geometry to synthesize geometrically consistent hard shadows. We demonstrate through quantitative and qualitative experiments on Multi-PIE and FFHQ that our method produces more geometrically consistent shadows than previous face relighting methods while also achieving state-of-the-art face relighting performance under directional lighting. In addition, we demonstrate that our differentiable hard shadow modeling improves the quality of the estimated face geometry over diffuse shading models.
Andrew Z. Hou, Michel Sarkis, Ning Bi, Yiying Tong, Xiaoming Liu 0002
CVPR3
2022 From Local to Holistic: Self-supervised Single Image 3D Face Reconstruction Via Multi-level Constraints
abstract
Single image 3D face reconstruction with accurate geometric details is a critical and challenging task due to the similar appearance on the face surface and fine details in organs. In this work, we introduce a self-supervised 3D face reconstruction approach from a single image that can recover detailed textures under different camera settings. The proposed network learns high-quality disparity maps from stereo face images during the training stage, while just a single face image is required to generate the 3D model in real applications. To recover fine details of each organ and facial surface, the framework introduces facial landmark spatial consistency to constrain the face recovering learning process in local point level and segmentation scheme on facial organs to constrain the correspondences at the organ level. The face shape and textures will further be refined by establishing holistic constraints based on the varying light illumination and shading information. The proposed learning framework can recover more accurate 3D facial details both quantitatively and qualitatively compared with state-of-the-art 3DMM and geometry-based reconstruction algorithms based on a single image.
Yawen Lu, Michel Sarkis, Ning Bi, Guoyu Lu 0001
IROS3
2022 Perceptual Consistency in Video Segmentation
abstract
In this paper, we present a novel perceptual consistency perspective on video semantic segmentation, which can capture both temporal consistency and pixel-wise correctness. Given two nearby video frames, perceptual consistency measures how much the segmentation decisions agree with the pixel correspondences obtained via matching general perceptual features. More specifically, for each pixel in one frame, we find the most perceptually correlated pixel in the other frame. Our intuition is that such a pair of pixels are highly likely to belong to the same class. Next, we assess how much the segmentation agrees with such perceptual correspondences, based on which we derive the perceptual consistency of the segmentation maps across these two frames. Utilizing perceptual consistency, we can evaluate the temporal consistency of video segmentation by measuring the perceptual consistency over consecutive pairs of segmentation maps in a video. Furthermore, given a sparsely labeled test video, perceptual consistency can be utilized to aid with predicting the pixel-wise correctness of the segmentation on an unlabeled frame. More specifically, by measuring the perceptual consistency between the predicted segmentation and the available ground truth on a nearby frame and combining it with the segmentation confidence, we can accurately assess the classification correctness on each pixel. Our experiments show that the proposed perceptual consistency can more accurately evaluate the temporal consistency of video segmentation as compared to flow-based measures. Furthermore, it can help more confidently predict segmentation accuracy on unlabeled test frames, as compared to using classification confidence alone. Finally, our proposed measure can be used as a regularizer during the training of segmentation models, which leads to more temporally consistent video segmentation while maintaining accuracy.
Yizhe Zhang 0001, Shubhankar Borse, Ying Wang 0051, Ning Bi, Xiaoyun Jiang, Fatih Porikli
WACV5
2021 Towards High Fidelity Face Relighting With Realistic Shadows
abstract
Existing face relighting methods often struggle with two problems: maintaining the local facial details of the subject and accurately removing and synthesizing shadows in the relit image, especially hard shadows. We propose a novel deep face relighting method that addresses both problems. Our method learns to predict the ratio (quotient) image between a source image and the target image with the desired lighting, allowing us to relight the image while maintaining the local facial details. During training, our model also learns to accurately modify shadows by using estimated shadow masks to emphasize on the high-contrast shadow borders. Furthermore, we introduce a method to use the shadow mask to estimate the ambient light intensity in an image, and are thus able to leverage multiple datasets during training with different global lighting intensities. With quantitative and qualitative evaluations on the Multi-PIE and FFHQ datasets, we demonstrate that our proposed method faithfully maintains the local facial details of the subject and can accurately handle hard shadows while achieving state-of-the-art face relighting performance.
Andrew Z. Hou, Michel Sarkis, Ning Bi, Yiying Tong, Xiaoming Liu 0002
CVPR4
2021 Real-Time Selfie Video Stabilization
abstract
We propose a novel real-time selfie video stabilization method. Our method is completely automatic and runs at 26 fps. We use a 1D linear convolutional network to directly infer the rigid moving least squares warping which implicitly balances between the global rigidity and local flexibility. Our network structure is specifically designed to stabilize the background and foreground at the same time, while providing optional control of stabilization focus (relative importance of foreground vs. background) to the users. To train our network, we collect a selfie video dataset with 1005 videos, which is significantly larger than previous selfie video datasets. We also propose a grid approximation to the rigid moving least squares that enables the real-time frame warping. Our method is fully automatic and produces visually and quantitatively better results than previous real-time general video stabilization methods. Compared to previous offline selfie video methods, our approach produces comparable quality with a speed improvement of orders of magnitude. Our code and selfie video dataset is available at https://github.com/jiy173/selfievideostabilization.
Jiyang Yu, Ravi Ramamoorthi, Ke-Li Cheng, Michel Sarkis, Ning Bi
CVPR5
2020 An improved SIFT algorithm for robust emotion recognition under various face poses and illuminations
Zhao Lv, Ning Bi, Chao Zhang 0047
Neural Comput. Appl.3
2020 A new upper bound of p for lp-minimization in compressed sensing
Kaihao Liang, Ning Bi
Signal Process.2
2020 Visual Tracking With Multiview Trajectory Prediction
abstract
Recent progresses in visual tracking have greatly improved the tracking performance. However, challenges such as occlusion and view change remain obstacles in real world deployment. A natural solution to these challenges is to use multiple cameras with multiview inputs, though existing systems are mostly limited to specific targets (e.g. human), static cameras, and/or require camera calibration. To break through these limitations, we propose a generic multiview tracking (GMT) framework that allows camera movement, while requiring neither specific object model nor camera calibration. A key innovation in our framework is a cross-camera trajectory prediction network (TPN), which implicitly and dynamically encodes camera geometric relations, and hence addresses missing target issues such as occlusion. Moreover, during tracking, we assemble information across different cameras to dynamically update a novel collaborative correlation filter (CCF), which is shared among cameras to achieve robustness against view change. The two components are integrated into a correlation filter tracking framework, where features are trained offline using existing single view tracking datasets. For evaluation, we first contribute a new generic multiview tracking dataset (GMTD) with careful annotations, and then run experiments on the GMTD and CAMPUS datasets. The proposed GMT algorithm shows clear advantages in terms of robustness over state-of-the-art ones.
Minye Wu, Haibin Ling, Ning Bi, Shenghua Gao, Qiang Hu 0003, Hao Sheng 0001, Jingyi Yu 0001
IEEE Trans. Image Process.3
2019 PPGNet: Learning Point-Pair Graph for Line Segment Detection
abstract
In this paper, we present a novel framework to detect line segments in man-made environments. Specifically, we propose to describe junctions, line segments and relationships between them with a simple graph, which is more structured and informative than end-point representation used in existing line segment detection methods. In order to extract a line segment graph from an image, we further introduce the PPGNet, a convolutional neural network that directly infers a graph from an image. We evaluate our method on published benchmarks including York Urban and Wireframe datasets. The results demonstrate that our method achieves satisfactory performance and generalizes well on all the benchmarks. The source code of our work is available at https://github.com/svip-lab/PPGNet.
Ning Bi, Jia Zheng 0002, Kun Huang 0001, Weixin Luo, Yanyu Xu 0001, Shenghua Gao
CVPR3
2019 Residual BiRNN Based Seq2Seq Model with Transition Probability Matrix for Online Handwritten Mathematical Expression Recognition
abstract
In this paper, we present a Seq2Seq model for online handwritten mathematical expression recognition (OHMER), which consists of two major parts: a residual bidirectional RNN (BiRNN) based encoder that takes handwritten traces as the input and a transition probability matrix introduced decoder that generates LaTeX notations. We employ residual connection in the BiRNN layers to improve feature extraction. Markovian transition probability matrix is introduced in decoder and long-term information can be used in each decoding step through joint probability. Furthermore, we analyze the impact of the novel encoder and transition probability matrix through several specific instances. Experimental results on the CROHME 2014 and CROHME 2016 competition tasks show that our model outperforms the previous state-of-the-art single model by only using the official training dataset.
Zelin Hong, Ning You, Jun Tan 0001, Ning Bi
ICDAR4
2019 DAC: Data-Free Automatic Acceleration of Convolutional Networks
abstract
Deploying a deep learning model on mobile/IoT devices is a challenging task. The difficulty lies in the trade-off between computation speed and accuracy. A complex deep learning model with high accuracy runs slowly on resource-limited devices, while a light-weight model that runs much faster loses accuracy. In this paper, we propose a novel decomposition method, namely DAC, that is capable of factorizing an ordinary convolutional layer into two layers with much fewer parameters. DAC computes the corresponding weights for the newly generated layers directly from the weights of the original convolutional layer. Thus, no training (or fine-tuning) or any data is needed. The experimental results show that DAC reduces a large number of floating-point operations (FLOPs) while maintaining high accuracy of a pre-trained model. If 2% accuracy drop is acceptable, DAC saves 53% FLOPs of VGG16 image classification model on ImageNet dataset, 29% FLOPS of SSD300 object detection model on PASCAL VOC2007 dataset, and 46% FLOPS of a multi-person pose estimation model on Microsoft COCO dataset. Compared to other existing decomposition methods, DAC achieves better performance.
Xin Li 0080, Shuai Zhang 0009, Bolan Jiang, Yingyong Qi, Mooi Choo Chuah, Ning Bi
WACV6
2019 The Handwritten Chinese Character Recognition Uses Convolutional Neural Networks with the GoogLeNet
abstract
With the outstanding performance in 2014 at the ImageNet Large-Scale Visual Recognition Challenge 2014 (ILSVRC14), an effective convolutional neural network (CNN) model named GoogLeNet has drawn the attention of the mainstream machine learning field. In this paper we plan to take an insight into the application of the GoogLeNet in the Handwritten Chinese Character Recognition (HCCR) on the database HCL2000 and CASIA-HWDB with several necessary adjustments and also state-of-the-art improvement methods for this end-to-end approach. Through the experiments we have found that the application of the GoogLeNet for the Handwritten Chinese Character Recognition (HCCR) results into significant high accuracy, to be specific more than 99% for the final version, which is encouraging for us to further research.
Ning Bi, Jun Tan 0001
Int. J. Pattern Recognit. Artif. Intell.1
2019 A multi-feature selection approach for gender identification of handwriting based on kernel mutual information
Ning Bi, Ching Y. Suen, Nicola Nobile, Jun Tan 0001
Pattern Recognit. Lett.1
2018 A Modified PSRoI Pooling with Spatial Information
Yiqing Zheng, Xiaolu Hu, Ning Bi, Jun Tan 0001
PRCV (2)3
2018 An Embedded Method for Feature Selection Using Kernel Parameter Descent Support Vector Machine
Haiqing Zhu, Ning Bi, Jun Tan 0001, Dongjie Fan
PRCV (3)2
2018 High-dimensional supervised feature selection via optimized kernel mutual information
Ning Bi, Jun Tan 0001, Jian-Huang Lai, Ching Y. Suen
Expert Syst. Appl.1
2016 Multi-feature Selection of Handwriting for Gender Identification Using Mutual Information
abstract
This paper presents a new flexible approach to predict the gender of the writers from their handwriting samples. Handwriting features can be extracted from different methods. Therefore, the multi-feature sets are irrelevant and redundant. The conflict of the features exists in the sets, which affects the accuracy of classification and the computing cost. This paper proposes a Mutual Information (MI) approach, that focuses on feature selection. The approach can decrease redundancies and conflicts. In addition, it extracts an optimal subset of features from the writing samples produced by male and female writers. The classification is carried out using a Support Vector Machine (SVM) on two databases. The first database comes from the ICDAR 2013 competition on gender prediction, the other database contains the Registration-Document-Form (RDF) database in Chinese. The proposed and compared methods were evaluated on both databases. Results from the methods highlight the importance of feature selection for gender prediction from handwriting.
Jun Tan 0001, Ning Bi, Ching Y. Suen, Nicola Nobile
ICFHR2
2007 Robust Image Watermarking Based on Multiband Wavelets and Empirical Mode Decomposition
abstract
In this paper, we propose a blind image watermarking algorithm based on the multiband wavelet transformation and the empirical mode decomposition. Unlike the watermark algorithms based on the traditional two-band wavelet transform, where the watermark bits are embedded directly on the wavelet coefficients, in the proposed scheme, we embed the watermark bits in the mean trend of some middle-frequency subimages in the wavelet domain. We further select appropriate dilation factor and filters in the multiband wavelet transform to achieve better performance in terms of perceptually invisibility and the robustness of the watermark. The experimental results show that the proposed blind watermarking scheme is robust against JPEG compression, Gaussian noise, salt and pepper noise, median filtering, and ConvFilter attacks. The comparison analysis demonstrate that our scheme has better performance than the watermarking schemes reported recently.
Ning Bi, Qiyu Sun, Daren Huang, Zhihua Yang, Jiwu Huang
IEEE Trans. Image Process.1
2005 The M-band wavelets in image watermarking
abstract
Multi-band (M-band) wavelet domain presents a novelty to host the watermark. In this paper, a new family of M-band wavelets, which is symmetric and parameterized with a variable /spl lambda/, is proposed and applied to image watermarking. The parameter /spl lambda/ also can be used as a key in watermark detection to improve the security of watermark. The multi-resolution analysis (MRA) of M-band wavelet transform, integrating with the CDMA (code division multiple access) encoding techniques is studied and employed to watermarking. The security, imperceptibility, and the robustness against JPEG compression and Gaussian noise, are analyzed for the proposed watermarking scheme. The experiments of watermarking based on M-band wavelet transform provide more encouraging results than those based on 2-band wavelets.
Yanmei Fang, Ning Bi, Daren Huang, Jiwu Huang
ICIP (1)2
2004 M-band wavelets application to palmprint recognition based on texture features
abstract
In this paper, a new texture feature set based on M-band wavelet analysis is proposed for online palmprint identification. First, the central part of a palm was decomposed by M-band wavelets transformation, then, l-norm energy was extracted out as our features, at last, the candidate was found by matching process. The presented experimental results demonstrate the validity of our approach.
Ning Bi, Daren Huang, Dvaid Zhang
ICIP2
2002 A robust speech recognition system embedded in CDMA cellular phone chipsets
abstract
PureVoice™ VR is a speech recognition system characterized by noise robustness performance, small footprint, and limited computing requirement. The system was embedded into Qualcomm's Code Division Multiple Access (CDMA) Mobile Station Modem (MSM) chipsets. The PureVoice™ VR engine enabled many manufactures of CDMA cellular phones to add highly accurate voice-activated dialing features to their handsets and hands-free car kits without significant increase in hardware cost. Users of these cellular phones could make calls by simply speaking names or phone numbers. The current embedded system has the capacity to recognize 100 speaker-dependent (SD) nametags and user defined commands, or 30 speaker-independent (SI) digits and commands in realtime. It also works in a hybrid mode, in which an input utterance can be either SD namtags or SI commands. PureVoice™ VR has successfully supported the voice-activated dialing function for millions of CDMA handsets sold in the world market.
Ning Bi, Harinath Garudadri, Chienchung Chang, Andrew DeJaco, Yingyong Qi, Naren Malayath, William Huang
ICASSP1
1997 Application of speech conversion to alaryngeal speech enhancement
abstract
Two existing speech conversion algorithms were modified and used to enhance alaryngeal speech. The modifications were aimed at reducing the spectral distortion (bandwidth increase) in a vector-quantization (VQ) based system and the spectral discontinuity in a linear multivariate regression (LMR) based system. Spectral distortion was compensated for by formant enhancement using the chirp z-transform and cepstral weighting. Spectral discontinuity was alleviated using overlapping clusters during the construction of the conversion mapping function. The modified VQ and LMR algorithms were used to enhance alaryngeal speech. The results of perceptual evaluation indicated that listeners generally preferred to listen to the alaryngeal speech samples enhanced by the modified conversions over original samples.
Ning Bi, Yingyong Qi
IEEE Trans. Speech Audio Process.1