Search results for "Speech recognition"

showing 10 items of 357 documents

The Pursuit of Happiness in Music: Retrieving Valence with Contextual Music Descriptors

2009

In the study of music emotions, Valence is usually referred to as one of the dimensions of the circumplex model of emotions that describes music appraisal of happiness, whose scale goes from sad to happy. Nevertheless, related literature shows that Valence is known as being particularly difficult to be predicted by a computational model. As Valence is a contextual music feature, it is assumed here that its prediction should also require contextual music descriptors in its predicting model. This work describes the usage of eight contextual (also known as higher-level) descriptors, previously developed by us, to calculate happiness in music. Each of these descriptors was independently tested …

MusicologyComputational modelMusic psychologyComputer scienceSpeech recognitionmedia_common.quotation_subjectHappinessLinear modelMusic information retrievalValence (psychology)Musical formmedia_common
researchProduct

Part of speech tagging with Naïve Bayes methods

2014

Naive Bayes classifierbusiness.industryPart-of-speech taggingComputer scienceSpeech recognitionArtificial intelligencecomputer.software_genrebusinesscomputerNatural language processing2014 18th International Conference on System Theory, Control and Computing (ICSTCC)
researchProduct

Supramodal neural processing of abstract information conveyed by speech and gesture

2013

Abstractness and modality of interpersonal communication have a considerable impact on comprehension. They are relevant for determining thoughts and constituting internal models of the environment. Whereas concrete object-related information can be represented in mind irrespective of language, abstract concepts require a representation in speech. Consequently, modality-independent processing of abstract information can be expected. Here we investigated the neural correlates of abstractness (abstract vs. concrete) and modality (speech vs. gestures), to identify an abstractness-specific supramodal neural network. During fMRI data acquisition 20 participants were presented with videos of an ac…

Neural correlates of consciousnessModality (human–computer interaction)Cognitive NeuroscienceSpeech recognitionspeechfMRIRepresentation (systemics)Context (language use)Interpersonal communicationemblematic gesturesSemanticslcsh:RC321-571ComprehensionBehavioral NeuroscienceNeuropsychology and Physiological Psychologytool-use gesturesabstract semanticsgestureOriginal Research ArticlePsychologylcsh:Neurosciences. Biological psychiatry. NeuropsychiatryGestureNeuroscienceFrontiers in Behavioral Neuroscience
researchProduct

2014

Due to its millisecond-scale temporal resolution, EEG allows to assess neural correlates with precisely defined temporal relationship relative to a given event. This knowledge is generally lacking in data from functional magnetic resonance imaging (fMRI) which has a temporal resolution on the scale of seconds so that possibilities to combine the two modalities are sought. Previous applications combining event-related potentials (ERPs) with simultaneous fMRI BOLD generally aimed at measuring known ERP components in single trials and correlate the resulting time series with the fMRI BOLD signal. While it is a valuable first step, this procedure cannot guarantee that variability of the chosen …

Neural correlates of consciousnessgenetic structuresmedicine.diagnostic_testGeneral NeuroscienceSpeech recognitionElectroencephalographyEEG-fMRIbehavioral disciplines and activitiesIndependent component analysisTask (project management)nervous systemTemporal resolutionmedicineGeneralizability theoryFunctional magnetic resonance imagingPsychologypsychological phenomena and processesFrontiers in Neuroscience
researchProduct

Heart rate accelerations during four active encoding tasks — pilot results

1997

Neuropsychology and Physiological PsychologyComputer sciencePhysiology (medical)General NeuroscienceSpeech recognitionEncoding (memory)Heart rateInternational Journal of Psychophysiology
researchProduct

Implicit Wiener Filtering for Speech Enhancement In Non-Stationary Noise

2021

Speech quality is degraded in the presence of background noise, which reduces the quality of experience (QoE) of the end-user and therefore motivates the usage of speech enhancement algorithms. A large number of approaches have been proposed in this context. However most of them have focused on the case where the noise is stationary, an assumption that seldom holds in practice. For instance, in mobile telephony, noise sources with a marked non-stationary spectral signature include vehicles, machines, and other speakers to name a few. On the other hand, the usage of frequency-domain information in existing algorithms for speech enhancement in non-stationary noise environments can be made mor…

Noise powerComputer scienceSpeech recognitionWiener filterSpectral densityComputer Science::Computation and Language (Computational Linguistics and Natural Language and Speech Processing)Context (language use)Background noiseSpeech enhancementNoisesymbols.namesakeComputer Science::SoundFrequency domainsymbols2021 11th International Conference on Information Science and Technology (ICIST)
researchProduct

Concatenated trial based Hilbert-Huang transformation on event-related potentials

2010

Time-frequency analysis is critical to study event-related potentials (ERPs) now. ERPs are usually generated through averaging over a number of trials, and such averaging limits the application of a nonlinear time-frequency analysis method—Hilbert-Huang transformation (HHT). This is because HHT usually requires very long recordings to sufficiently decompose the complicated signal into oscillations and the averaged ERP trace tends to possess only hundreds of samples. Thus, this study designs the concatenated trial based HHT to release the limitation on the decomposition. Such a paradigm may reveal better temporal and spectral properties of an ERP than the conventional wavelet transformation …

Nonlinear systemTransformation (function)WaveletEvent-related potentialSpeech recognitionSpectral propertiesSignalMathematicsTime–frequency analysisTRACE (psycholinguistics)The 2010 International Joint Conference on Neural Networks (IJCNN)
researchProduct

MUSIC-LS Modal Channel Estimation for an OFDM-OQAM System

2008

Orthogonal frequency division multiplexing based on offset quadrature amplitude modulation (OFDM-OQAM) signaling over frequency selective multipath channels shows inter-symbol interference (ISI) and Inter-Channel Interference (ICI) that degrade its performance. Channel equalization and channel estimation are needed to combat these intrinsic interferences. In this paper a novel modal channel estimator based on the MUltiple signal classification (MUSIC) and least squares (LS) algorithms for a wideband passband OFDM-OQAM signaling over static multipath channels is presented. The effects of the frequency selective channel on the received signal are described considering a wideband OFDM-OQAM sys…

OFDM-OQAMOrthogonal frequency-division multiplexingComputer scienceSettore ING-INF/03 - TelecomunicazioniSpeech recognitionData_CODINGANDINFORMATIONTHEORYLeast squaresChannel modelingMUSICIntersymbol interferenceInterference (communication)Adjacent-channel interferenceWidebandAlgorithmQuadrature amplitude modulationComputer Science::Information TheoryCommunication channel
researchProduct

A pre-processing technique based on the wavelet transform for linear autoassociators with applications to face recognition

1997

In order to improve the performance of a linear autoassociator (which is a neural network model), we explore the use of several preprocessing techniques. The gist of our approach is to store, in addition to the original pattern, one or several pre-processed (i.e. filtered) versions of the patterns to be stored in a neural network. First, we compare the performance of several pre-processing techniques (a plain vanilla version of the autoassociator as a control, a Sobel operator, a Canny-Deriche operator, and a multiscale Canny-Deriche operator) on an example of a pattern completion task using a noise degraded version of a face stored in an autoassociator. We found that the multiscale Canny-D…

Operator (computer programming)Artificial neural networkComputer scienceSpeech recognitionPattern recognition (psychology)ComputingMethodologies_IMAGEPROCESSINGANDCOMPUTERVISIONWavelet transformSobel operatorNoise (video)Facial recognition systemEdge detection
researchProduct

Analyzing and organizing the sonic space of vocal imitations

2015

The sonic space that can be spanned with the voice is vast and complex and, therefore, it is difficult to organize and explore. In order to devise tools that facilitate sound design by vocal sketching we attempt at organizing a database of short excerpts of vocal imitations. By clustering the sound samples on a space whose dimensionality has been reduced to the two principal components, it is experimentally checked how meaningful the resulting clusters are for humans. Eventually, a representative of each cluster, chosen to be close to its centroid, may serve as a landmark in the exploration of the sound space, and vocal imitations may serve as proxies for synthetic sounds.

PCALandmarkSettore INF/01 - InformaticaComputer scienceSound designSpeech recognitionCentroidSpace (commercial competition)ClusteringLandmarkPrincipal component analysisVocal imitationsCluster analysisCurse of dimensionality
researchProduct