Example: dental hygienist

SOUND RE-SYNTHESIS FROM RHYTHM PATTERN FEATURES …

SOUND RE-SYNTHESIS FROM RHYTHM PATTERN FEATURES audible INSIGHT INTO A MUSIC feature EXTRACTION PROCESST homas LidyGeorg P olzlbauerAndreas RauberVienna University of TechnologyDepartment of Software Technology and Interactive SystemsVienna, Austria{lidy, poelzlbauer, tasks like musical genre identification and similaritysearches in audio databases, audio files have to be de-scribed by suitable feature sets. Since these feature setsusually try to capture diverse discriminative characteris-tics, it is interesting and desirable to create an acousticrepresentation of the feature set to support intuitive eval-uation. In this paper, we present an approach for makinga specific feature set, namely RHYTHM patterns , instantlyhuman comprehensible by re-assembling SOUND from thenumerical descriptors. The re-synthesized audio chunksrepresent clearly perceivable rhythmical characteristics oncritical frequency bands of the original INTRODUCTIONThe Music Information Retrieval research domain gainedincreasing attention in recent years.}

SOUND RE-SYNTHESIS FROM RHYTHM PATTERN FEATURES – AUDIBLE INSIGHT INTO A MUSIC FEATURE EXTRACTION PROCESS Thomas Lidy Georg Polzlbauer Andreas Rauber¨

Tags:

  Form, Feature, Insights, Synthesis, Rhythm, Patterns, Audible, Synthesis from rhythm pattern features, Synthesis from rhythm pattern features audible insight

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of SOUND RE-SYNTHESIS FROM RHYTHM PATTERN FEATURES …

1 SOUND RE-SYNTHESIS FROM RHYTHM PATTERN FEATURES audible INSIGHT INTO A MUSIC feature EXTRACTION PROCESST homas LidyGeorg P olzlbauerAndreas RauberVienna University of TechnologyDepartment of Software Technology and Interactive SystemsVienna, Austria{lidy, poelzlbauer, tasks like musical genre identification and similaritysearches in audio databases, audio files have to be de-scribed by suitable feature sets. Since these feature setsusually try to capture diverse discriminative characteris-tics, it is interesting and desirable to create an acousticrepresentation of the feature set to support intuitive eval-uation. In this paper, we present an approach for makinga specific feature set, namely RHYTHM patterns , instantlyhuman comprehensible by re-assembling SOUND from thenumerical descriptors. The re-synthesized audio chunksrepresent clearly perceivable rhythmical characteristics oncritical frequency bands of the original INTRODUCTIONThe Music Information Retrieval research domain gainedincreasing attention in recent years.}

2 The sheer amount ofmusic titles available in repositories calls for sophisticatedsearch, retrieval and organization techniques. These, inturn, require representative descriptors for pieces of mu-sic in order to measure similarities between music titles. Itis, however, unclear, what constitutes the essential mean-ing of music. Most approaches thus focus on perceptu-ally relevant FEATURES of music. The descriptors, or fea-tures, are derived from the plain audio signal. Althoughsubstantial reduction of data is desired, the descriptorshave to contain sufficient information from the data to rep-resent some kind of semantics of the descriptors build the basis for many in-formation retrieval tasks, such as similarity based searches(query-by-example, query-by-humming, etc.), organiza-tion and clustering tasks, classification tasks, etc.

3 Numer-ous different types of descriptors have been proposed. Allthese descriptors are more or less suitable as a representa-tion of the content of audio, frequently depending on thespecific application. The performance of FEATURES in sim-ilarity retrieval or classification tasks has been evaluatedin numerous experiments, some of which showed that acombination of several feature sets improves the our work we concentrate on RHYTHM patterns as fea-tures, describing the loudness amplitude modulation indifferent frequency bands. The feature set does not merelyrepresent RHYTHM , or beat, it describes fluctuations in nu-merous frequency regions covering the complete audiblefrequency question frequently raised, particularly for non-stan-dard feature sets, is on the cognitive characteristics of theextracted numbers. What is it, that the FEATURES actuallyrepresent?

4 For humans it is sometimes difficult to get anotion of the feature set as a whole, as the feature space isoften high-dimensional. Relations between the attributescan be elusive, thus the chance for insight into the datais diminished. Since this issue of acoustic interpretationof the feature set has been raised several times since thefeature set s inception [10], in this paper we present anacoustical RE-SYNTHESIS of the feature set. We thus seekto make the numerical descriptors instantly comprehensi-ble to humans allowing to verify characteristics present inthe feature set intuitively. Furthermore, the synthesizedsound can serve as a control technique for the feature ex-traction process and provides a notion of the suitabilityof the feature set for content-based description of musi-cal data. With the audible feature set, one can evaluate theeffectiveness of the feature extraction through asking a hu-man for the same task as the computer, working only withthe substantially reduced information from the aggregateddescriptor, : Can you discriminate musical genres pro-vided only with the information from the feature set?

5 In the evaluation of the re-synthesized audible vec-tors we experienced, that the rhythmical structure ( modulation) on all frequency regions is satis-factorily reassembled. This serves as an indication, thatthe RHYTHM patterns we chose for content description provesuitable to represent characteristics of a given piece of au-dio, thus appearing appropriate for classification, organi-zation and retrieval remainder of the paper is structured as follows: InSection 2 we give an overview of related work. Section 3introduces the feature extraction algorithm that forms thebase of the SOUND synthesis approach. The synthesis pro-cess is outlined in Section 4. Section 5 states evaluationresults, followed by conclusions in Section RELATED WORKIn recent years, audio analysis received by far more atten-tion than audio synthesis .

6 As stated in the introduction,this is due to the currently strong interest in music infor-mation retrieval tasks. Approaches for deriving content-based audio descriptors are manifold and include the ex-traction of tempo, beat [2, 5], RHYTHM [3], pitch [7, 14],and melody [4], to just name a the music information retrieval domain, clearly, thereis little work on synthesis of audio. Recently, SOUND syn-thesis is applied in computer music and in conjunctionwith animation and art. A work about additive synthe-sis using the Inverse Fast Fourier Transformation, both ofwhich is used in our approach, dates back to 1992 [12]. In[9] the authors present a method for preventing artefactsin the RE-SYNTHESIS of a signal that was previously anal-ysed using the Short Time Fourier Transform with win-dow recent method from the music information retrievaldomain applies signal synthesis during automatic drumdetection [15].

7 Drum SOUND is iteratively derived and re-synthesised for progressive detection of further percussivesound in the input signal. Signal analysis and subsequentsynthesis is also applied in [1, 8], modelling the timbre ofa musical feature EXTRACTIONThe feature set our work is based on is denominated as RHYTHM patterns . Describing amplitude modulationson various frequency bands covering the complete humanaudible frequency range, it contains far more than whatis commonly considered as RHYTHM . The RHYTHM Pat-terns FEATURES are derived analysing the spectral data ofthe music signal plus incorporating psycho-acoustic phe-nomenons. At the final stage they represent fluctuationsper modulation frequency on 24 frequency bands accord-ing to human perception. The algorithm is described indetail in [11]. In the following, we give a brief outline ofthe extraction process, depicted in Figure algorithm processes audio tracks in standard digi-tal PCM format with kHz sampling frequency as in-put.

8 First, the audio track is segmented into pieces of 6seconds , a short time Fast Fourier Trans- form (STFT) is used to retrieve the energy per frequencybin, the spectrum, every ms, resulting in a spec-trogram of the 6 second segment. To reduce the amount ofdata, the frequency bins of the spectrogram are summedup to 24 so-called critical bands, according to the Barkscale [16]. A further psycho-acoustical phenomenon in-corporated is spectral masking, the occlusion of onesound by another SOUND . This phenomenon is coped witha spreading function [13]. Successively, the data is trans-formed into the logarithmic decibel scale, equal-loudnesscurves are accounted for [16], resulting in a transforma-tion into the unit Phon and afterwards into the unit Sone,reflecting the specific loudness sensation of the humanauditory system.

9 At this point, we computed the spe-cific loudness sensation over time on 24 critical frequency1 The duration of the segment is actually seconds, which has anappropriate number of samples (218) for effective processing throughthe two Fast Fourier Transforms. Nevertheless, we denote the segment 6 second segment throughout the diagram of the feature extraction process. Ar-rows with broken lines do not form part of the feature extrac-tion, but indicate typical post-extraction approaches. Our newapproach is the RE-SYNTHESIS of extracted feature Still, we have a time-dependent signal, althoughreduced to 512 sample values at the time axis due to thewindow size in the order to obtain a time-independent representation ofthe data, another Fourier Transform is applied. The ideais to regard the varying energy on a frequency band of thespectrogram as a modulation of the amplitude over the second Fourier Transform, the spectrum of thismodulation signal is computed.

10 It is a time-invariant sig-nal that denotes the modulation frequency on the abscissa,and the magnitude of modulation on the ordinate. A highamplitude at the modulation frequency of 2 Hz for exam-ple indicates a strong RHYTHM at 120 bpm (beats per minute= modulation frequency * 60). The abscissa ranges Hz to 43 Hz, with 43 Hz corresponding to 2580bpm, which is far beyond what any auditory system isable to perceive as RHYTHM . The notion of RHYTHM alreadyends above 15 Hz where the sensation of roughness startsand goes up to 150 Hz, the limit where only three sepa-rately audible tones are perceivable. For that reason, thedata used for the derived FEATURES is cut after a modula-tion frequency of 10 Hz, which means, that on each of the24 critical bands, only 60 values are preserved. The fi-nal feature vector thus has 24*60 dimensions, containinga time-invariant representation of fluctuation strength be-tween Hz and 10 Hz.


Related search queries