Search results for "speech recognition"
showing 10 items of 357 documents
Probing neural mechanisms of music perception, cognition, and performance using multivariate decoding.
2012
Recent neuroscience research has shown increasing use of multivariate decoding methods and machine learning. These methods, by uncovering the source and nature of informative variance in large data sets, invert the classical direction of inference that attempts to explain brain activity from mental state variables or stimulus features. However, these techniques are not yet commonly used among music researchers. In this position article, we introduce some key features of machine learning methods and review their use in the field of cognitive and behavioral neuroscience of music. We argue for the great potential of these methods in decoding multiple data types, specifically audio waveforms, e…
Non-negative matrix factorization Vs. FastICA on mismatch negativity of children
2009
In this presentation two event-related potentials, mismatch negativity (MMN) and P3a, are extracted from EEG by non-negative matrix factorization (NMF) simultaneously. Typically MMN recordings show a mixture of MMN, P3a, and responses to repeated standard stimuli. NMF may release the source independence assumption and data length limitations required by Fast independent component analysis (FastICA). Thus, in theory NMF could reach better separation of the responses. In the current experiment MMN was elicited by auditory duration deviations in 102 children. NMF was performed on the time-frequency representation of the raw data to estimate sources. Support to Absence Ratio (SAR) of the MMN co…
Visual Cortex Performs a Sort of Non-linear ICA
2010
Here, the standard V1 cortex model optimized to reproduce image distortion psychophysics is shown to have nice statistical properties, e.g. approximate factorization of the PDF of natural images. These results confirm the efficient encoding hypothesis that aims to explain the organization of biological sensors by information theory arguments.
Rapid neural encoding of the contrast between native and nonnative speech in the alpha band
2021
AbstractMore than half of the world’s population is multilingual, yet it is not known how the human brain encodes the perception of native vs. nonnative speech. To find out, we asked German native speakers to detect the onset of native and nonnative (English and Turkish) vowels in a roving standard stimulation. Using EEG, we show that nonnativeness is robustly registered by an increase in phase coherence in the alpha band (8-12 Hz), beginning as early as ∼100 ms after stimulus onset and lasting more than 200 ms. The alpha band effect is speech-specific, successfully predicts the response speed advantage of nonnative speech, and grants ∼90% decoding accuracy in distinguishing native vs. nonn…
Perceptual Interactions in Complex Odor Mixtures
2014
The perception of everyday odors relies on elemental or configural processing of complex mixtures of odorants. Theoretically, the configural processing of a mixture could lead to the perception of a single specific odor for the mixture; however, such a type of perception has hardly been proven in human studies. Here, we report the results of a sorting task demonstrating that a six-component mixture carries an odor clearly distinct from the odors of its components. These results suggest a blending effect of individual components’ odors in mixtures containing more than three odorants.
Interaction of sight and sound in the perception and experience of musical performance
2016
Recently, Vuoskoski, Thompson, Clarke, and Spence (2014) demonstrated that visual kinematic performance cues may be more important than auditory performance cues in terms of observers’ ratings of expressivity perceived in audiovisual excerpts of piano playing, and that visual kinematic performance cues had crossmodal effects on the perception of auditory expressivity. The present study was designed to extend these findings, and to provide additional information about the roles of sight and sound in the perception and experience of musical performance. Experiment 1 investigated the relative contributions of auditory and visual kinematic performance features to participants’ subjective emotio…
Sound Event Envelope Estimation in Polyphonic Mixtures
2019
Sound event detection is the task of identifying automatically the presence and temporal boundaries of sound events within an input audio stream. In the last years, deep learning methods have established themselves as the state-of-the-art approach for the task, using binary indicators during training to denote whether an event is active or inactive. However, such binary activity indicators do not fully describe the events, and estimating the envelope of the sounds could provide more precise modeling of their activity. This paper proposes to estimate the amplitude envelopes of target sound event classes in polyphonic mixtures. For training, we use the amplitude envelopes of the target sounds…
Sketching Sound with Voice and Gesture
2015
Voice and gestures are natural sketching tools that can be exploited to communicate sonic interactions. In product and interaction design, sounds should be included in the early stages of the design process. Scientists of human motion have shown that auditory stimuli are important in the performance of difficult tasks and can elicit anticipatory postural adjustments in athletes. These findings justify the attention given to sound in interaction design for gaming, especially in action and sports games that afford the development of levels of virtuosity. The sonic manifestations of objects can be designed by acting on their mechanical qualities and by augmenting the objects with synthetic and…
On the Dissociation of Word/Nonword Repetition Effects in Lexical Decision: An Evidence Accumulation Account
2016
A number of models of visual-word recognition assume that the repetition of an item in a lexical decision experiment increases that item's familiarity/wordness. This would produce not only a facilitative repetition effect for words, but also an inhibitory effect for nonwords (i.e., more familiarity/wordness makes the negative decision slower). We conducted a two-block lexical decision experiment to examine word/nonword repetition effects in the framework of a leading “familiarity/wordness” model of the lexical decision task, namely, the diffusion model (Ratcliff et al., 2004). Results showed that while repeated words were responded to faster than the unrepeated words, repeated nonwords were…
Glottal Source Features for Automatic Speech-Based Depression Assessment
2017
Depression is one of the most prominent mental disorders, with an increasing rate that makes it the fourth cause of disability worldwide. The field of automated depression assessment has emerged to aid clinicians in the form of a decision support system. Such a system could assist as a pre-screening tool, or even for monitoring high risk populations. Related work most commonly involves multimodal approaches, typically combining audio and visual signals to identify depression presence and/or severity. The current study explores categorical assessment of depression using audio features alone. Specifically, since depression-related vocal characteristics impact the glottal source signal, we exa…