Search results for "Speech recognition"
showing 10 items of 357 documents
The Pursuit of Happiness in Music: Retrieving Valence with Contextual Music Descriptors
2009
In the study of music emotions, Valence is usually referred to as one of the dimensions of the circumplex model of emotions that describes music appraisal of happiness, whose scale goes from sad to happy. Nevertheless, related literature shows that Valence is known as being particularly difficult to be predicted by a computational model. As Valence is a contextual music feature, it is assumed here that its prediction should also require contextual music descriptors in its predicting model. This work describes the usage of eight contextual (also known as higher-level) descriptors, previously developed by us, to calculate happiness in music. Each of these descriptors was independently tested …
Part of speech tagging with Naïve Bayes methods
2014
Supramodal neural processing of abstract information conveyed by speech and gesture
2013
Abstractness and modality of interpersonal communication have a considerable impact on comprehension. They are relevant for determining thoughts and constituting internal models of the environment. Whereas concrete object-related information can be represented in mind irrespective of language, abstract concepts require a representation in speech. Consequently, modality-independent processing of abstract information can be expected. Here we investigated the neural correlates of abstractness (abstract vs. concrete) and modality (speech vs. gestures), to identify an abstractness-specific supramodal neural network. During fMRI data acquisition 20 participants were presented with videos of an ac…
2014
Due to its millisecond-scale temporal resolution, EEG allows to assess neural correlates with precisely defined temporal relationship relative to a given event. This knowledge is generally lacking in data from functional magnetic resonance imaging (fMRI) which has a temporal resolution on the scale of seconds so that possibilities to combine the two modalities are sought. Previous applications combining event-related potentials (ERPs) with simultaneous fMRI BOLD generally aimed at measuring known ERP components in single trials and correlate the resulting time series with the fMRI BOLD signal. While it is a valuable first step, this procedure cannot guarantee that variability of the chosen …
Heart rate accelerations during four active encoding tasks — pilot results
1997
Implicit Wiener Filtering for Speech Enhancement In Non-Stationary Noise
2021
Speech quality is degraded in the presence of background noise, which reduces the quality of experience (QoE) of the end-user and therefore motivates the usage of speech enhancement algorithms. A large number of approaches have been proposed in this context. However most of them have focused on the case where the noise is stationary, an assumption that seldom holds in practice. For instance, in mobile telephony, noise sources with a marked non-stationary spectral signature include vehicles, machines, and other speakers to name a few. On the other hand, the usage of frequency-domain information in existing algorithms for speech enhancement in non-stationary noise environments can be made mor…
Concatenated trial based Hilbert-Huang transformation on event-related potentials
2010
Time-frequency analysis is critical to study event-related potentials (ERPs) now. ERPs are usually generated through averaging over a number of trials, and such averaging limits the application of a nonlinear time-frequency analysis method—Hilbert-Huang transformation (HHT). This is because HHT usually requires very long recordings to sufficiently decompose the complicated signal into oscillations and the averaged ERP trace tends to possess only hundreds of samples. Thus, this study designs the concatenated trial based HHT to release the limitation on the decomposition. Such a paradigm may reveal better temporal and spectral properties of an ERP than the conventional wavelet transformation …
MUSIC-LS Modal Channel Estimation for an OFDM-OQAM System
2008
Orthogonal frequency division multiplexing based on offset quadrature amplitude modulation (OFDM-OQAM) signaling over frequency selective multipath channels shows inter-symbol interference (ISI) and Inter-Channel Interference (ICI) that degrade its performance. Channel equalization and channel estimation are needed to combat these intrinsic interferences. In this paper a novel modal channel estimator based on the MUltiple signal classification (MUSIC) and least squares (LS) algorithms for a wideband passband OFDM-OQAM signaling over static multipath channels is presented. The effects of the frequency selective channel on the received signal are described considering a wideband OFDM-OQAM sys…
A pre-processing technique based on the wavelet transform for linear autoassociators with applications to face recognition
1997
In order to improve the performance of a linear autoassociator (which is a neural network model), we explore the use of several preprocessing techniques. The gist of our approach is to store, in addition to the original pattern, one or several pre-processed (i.e. filtered) versions of the patterns to be stored in a neural network. First, we compare the performance of several pre-processing techniques (a plain vanilla version of the autoassociator as a control, a Sobel operator, a Canny-Deriche operator, and a multiscale Canny-Deriche operator) on an example of a pattern completion task using a noise degraded version of a face stored in an autoassociator. We found that the multiscale Canny-D…
Analyzing and organizing the sonic space of vocal imitations
2015
The sonic space that can be spanned with the voice is vast and complex and, therefore, it is difficult to organize and explore. In order to devise tools that facilitate sound design by vocal sketching we attempt at organizing a database of short excerpts of vocal imitations. By clustering the sound samples on a space whose dimensionality has been reduced to the two principal components, it is experimentally checked how meaningful the resulting clusters are for humans. Eventually, a representative of each cluster, chosen to be close to its centroid, may serve as a landmark in the exploration of the sound space, and vocal imitations may serve as proxies for synthetic sounds.