{"title": "Monaural Speech Separation", "book": "Advances in Neural Information Processing Systems", "page_first": 1245, "page_last": 1252, "abstract": null, "full_text": " \n\n \n \n \n \n \n \n\nMonaural Speech Separation \n\nGuoning Hu \n\nBiophysics Program \n\nThe Ohio State University \n\nColumbus, OH 43210 \n\nhu.117@osu.edu \n\nDepartment of Computer and Information \n\nScience & Center of Cognitive Science \n\nThe Ohio State University, Columbus, OH 43210 \n\nDeLiang Wang \n\ndwang@cis.ohio-state.edu \n\nAbstract \n\nthat deals with \n\nMonaural speech separation has been studied in previous systems that \nincorporate auditory scene analysis principles. A major problem for \nthese systems is their inability to deal with speech in the high-\nfrequency range. Psychoacoustic evidence suggests that different \nperceptual mechanisms are involved in handling resolved and \nunresolved harmonics. Motivated by this, we propose a model for \nmonaural separation \nlow-frequency and high-\nfrequency signals differently. For resolved harmonics, our model \ngenerates segments based on temporal continuity and cross-channel \ncorrelation, and groups them according to periodicity. For unresolved \nharmonics, the model generates segments based on amplitude \nmodulation (AM) in addition to temporal continuity and groups them \naccording to AM repetition rates derived from sinusoidal modeling. \nUnderlying the separation process is a pitch contour obtained \naccording to psychoacoustic constraints. Our model is systematically \nevaluated, and it yields substantially better performance than previous \nsystems, especially in the high-frequency range. \n\n1 Introduction \n\nIn a natural environment, speech usually occurs simultaneously with acoustic \ninterference. An effective system for attenuating acoustic interference would greatly \nfacilitate many applications, including automatic speech recognition (ASR) and \nspeaker identification. Blind source separation using independent component analysis \n[10] or sensor arrays for spatial filtering require multiple sensors. In many situations, \nsuch as telecommunication and audio retrieval, a monaural (one microphone) solution \nis required, in which intrinsic properties of speech or interference must be considered. \nVarious algorithms have been proposed for monaural speech enhancement [14]. These \nmethods assume certain properties of interference and have difficulty in dealing with \ngeneral acoustic interference. Monaural separation has also been studied using phase-\nbased decomposition [3] and statistical learning [17], but with only limited evaluation. \nWhile speech enhancement remains a challenge, the auditory system shows a \nremarkable capacity for monaural speech separation. According to Bregman [1], the \nauditory system separates the acoustic signal into streams, corresponding to different \nsources, based on auditory scene analysis (ASA) principles. Research in ASA has \ninspired considerable work to build computational auditory scene analysis (CASA) \n\n\f \n\nsystems for sound separation [19] [4] [7] [18]. Such systems generally approach speech \nseparation in two main stages: segmentation (analysis) and grouping (synthesis). In \nsegmentation, the acoustic input is decomposed into sensory segments, each of which \nis likely to originate from a single source. In grouping, those segments that likely come \nfrom the same source are grouped together, based mostly on periodicity. In a recent \nCASA model by Wang and Brown [18], segments are formed on the basis of similarity \nbetween adjacent filter responses (cross-channel correlation) and temporal continuity, \nwhile grouping among segments is performed according to the global pitch extracted \nwithin each time frame. In most situations, the model is able to remove intrusions and \nrecover low-frequency (below 1 kHz) energy of target speech. However, this model \ncannot handle high-frequency (above 1 kHz) signals well, and it loses much of target \nspeech in the high-frequency range. In fact, the inability to deal with speech in the \nhigh-frequency range is a common problem for CASA systems. \nWe study monaural speech separation with particular emphasis on the high-frequency \nproblem in CASA. For voiced speech, we note that the auditory system can resolve the \nfirst few harmonics in the low-frequency range [16]. It has been suggested that \ndifferent perceptual mechanisms are used to handle resolved and unresolved harmonics \n[2]. Consequently, our model employs different methods to segregate resolved and \nunresolved harmonics of target speech. More specifically, our model generates \nsegments for resolved harmonics based on temporal continuity and cross-channel \ncorrelation, and these segments are grouped according to common periodicity. For \nunresolved harmonics, it is well known that the corresponding filter responses are \nstrongly amplitude-modulated and the response envelopes fluctuate at the fundamental \nfrequency (F0) of target speech [8]. Therefore, our model generates segments for \nunresolved harmonics based on common AM in addition to temporal continuity. The \nsegments are grouped according to AM repetition rates. We calculate AM repetition \nrates via sinusoidal modeling, which is guided by target pitch estimated according to \ncharacteristics of natural speech. \nSection 2 describes the overall system. In section 3, systematic results and a \ncomparison with the Wang-Brown system are given. Section 4 concludes the paper. \n\n2 M odel description \n\nOur model is a multistage system, as shown in Fig. 1. Description for each stage is \ngiven below. \n\n2 . 1 \n\nIni t i a l pr oc e ssi ng \n\nFirst, an acoustic input is analyzed by a standard cochlear filtering model with a bank \nof 128 gammatone filters [15] and subsequent hair cell transduction [12]. This \nperipheral processing is done in time frames of 20 ms long with 10 ms overlap between \nconsecutive frames. As a result, the input signal is decomposed into a group of time-\nfrequency (T-F) units. Each T-F unit contains the response from a certain channel at a \ncertain frame. The envelope of the response is obtained by a lowpass filter with \n\n \n\n \nMixture \n\nPeripheral and \n\nInitial \n\nsegregation \n\nmid-level \nprocessing \n Figure 1. Schematic diagram of the proposed multistage system. \n\nlabeling \n\nsegregation \n\n \n\nPitch \ntracking \n\nUnit \n\nFinal \n\nResynthesis \n\n \n\nSegregated \n\nSpeech \n\n\f \n\npassband [0, 1 kHz] and a Kaiser window of 18.25 ms. \nMid-level processing is performed by computing a correlogram (autocorrelation \nfunction) of the individual responses and their envelopes. These autocorrelation \nfunctions reveal response periodicities as well as AM repetition rates. The global pitch \nis obtained from the summary correlogram. For clean speech, the autocorrelations \ngenerally have peaks consistent with the pitch and their summation shows a dominant \npeak corresponding to the pitch period. With acoustic interference, a global pitch may \nnot be an accurate description of the target pitch, but it is reasonably close. \nBecause a harmonic extends for a period of time and its frequency changes smoothly, \ntarget speech likely activates contiguous T-F units. This is an instance of the temporal \ncontinuity principle. In addition, since the passbands of adjacent channels overlap, a \nresolved harmonic usually activates adjacent channels, which leads to high cross-\nchannel correlations. Hence, in initial segregation, the model first forms segments by \nmerging T-F units based on temporal continuity and cross-channel correlation. Then \nthe segments are grouped into a foreground stream and a background stream by \ncomparing the periodicities of unit responses with global pitch. A similar process is \ndescribed in [18]. \nFig. 2(a) and Fig. 2(b) illustrate the segments and the foreground stream. The input is a \nmixture of a voiced utterance and a cocktail party noise (see Sect. 3). Since the \nintrusion is not strongly structured, most segments correspond to target speech. In \naddition, most segments are in the low-frequency range. The initial foreground stream \nsuccessfully groups most of the major segments. \n\n2 . 2 Pi t c h tr a cki ng \n\nIn the presence of acoustic interference, the global pitch estimated in mid-level \nprocessing is generally not an accurate description of target pitch. To obtain accurate \npitch information, target pitch is first estimated from the foreground stream. At each \nframe, the autocorrelation functions of T-F units in the foreground stream are \nsummated. The pitch period is the lag corresponding to the maximum of the summation \nin the plausible pitch range: [2 ms, 12.5 ms]. Then we employ the following two \nconstraints to check its reliability. First, an accurate pitch period at a frame should be \nconsistent with the periodicity of the T-F units at this frame in the foreground stream. \nAt frame j, let t( j) represent the estimated pitch period, and A(i,j,t) the autocorrelation \nfunction of uij, the unit in channel i. uij agrees with t( j) if \n\n,(\niA\n\nt\n(,\n\nj\n\nj\n\n/))\n\n,(\niA\n\nt\n,\nj\n\n>)\n\nq\n\n \n\nd\n\nm\n\n \n\n \n\n \n\n \n\n (1) \n\n(a)\n\n(b)\n\n5000\n\n)\nz\nH\n\n(\n \ny\nc\nn\ne\nu\nq\ne\nr\nF\n\n2335\n\n1028\n\n 387\n\n 80\n0\n\n5000\n\n2335\n\n1028\n\n 387\n\n0.5\n1\nTime (Sec)\n\n1.5\n\n 80\n0\n\n0.5\n1\nTime (Sec)\n\n1.5\n\n \n\nFigure 2. Results of initial segregation for a speech and cocktail-party mixture. (a) \nSegments formed. Each segment corresponds to a contiguous black region. (b) \nForeground stream. \n\n\f \n\nHere, q d=0.95, the same threshold used in [18], and tm is the lag corresponding to the \nmaximum of A(i,j,t) within [2 ms, 12.5 ms]. t( j) is considered reliable if more than \nhalf of the units in the foreground stream at frame j agree with it. Second, pitch periods \nin natural speech vary smoothly in time [11]. We stipulate the difference between \nreliable pitch periods at consecutive frames be smaller than 20% of the pitch period, \njustified from pitch statistics. Unreliable pitch periods are replaced by new values \nextrapolated from reliable pitch points using temporal continuity. As an example, \nsuppose at two consecutive frames j and j+1 that t( j) is reliable while t( j+1) is not. All \nthe channels corresponding to the T-F units agreeing with t( j) are selected. t( j+1) is \nthen obtained from the summation of the autocorrelations for the units at frame j+1 in \nthose selected channels. Then the re-estimated pitch is further verified with the second \nconstraint. For more details, see [9]. \nFig. 3 illustrates the estimated pitch periods from the speech and cocktail-party \nmixture, which match the pitch periods obtained from clean speech very well. \n\n2 . 3 Uni t l a be li ng \n\nWith estimated pitch periods, (1) provides a criterion to label T-F units according to \nwhether target speech dominates the unit responses or not. This criterion compares an \nestimated pitch period with the periodicity of the unit response. It is referred as the \nperiodicity criterion. It works well for resolved harmonics, and is used to label the units \nof the segments generated in initial segregation. \nHowever, the periodicity criterion is not suitable for units responding to multiple \nharmonics because unit responses are amplitude-modulated. As shown in Fig. 4, for a \nfilter response that is strongly amplitude-modulated (Fig. 4(a)), the target pitch \ncorresponds to a local maximum, indicated by the vertical line, in the autocorrelation \ninstead of the global maximum (Fig. 4(b)). Observe that for a filter responding to \nmultiple harmonics of a harmonic source, the response envelope fluctuates at the rate \nof F0 [8]. Hence, we propose a new criterion for labeling the T-F units corresponding \nto unresolved harmonics by comparing AM repetition rates with estimated pitch. This \ncriterion is referred as the AM criterion. \nTo obtain an AM repetition rate, the entire response of a gammatone filter is half-wave \nrectified and then band-pass filtered to remove the DC component and other possible \n\n)\ns\nm\n\n(\n \n\nd\no\ni\nr\ne\nP\nh\nc\nt\ni\n\n \n\nP\n\n14\n\n12\n\n10\n\n8\n\n6\n\n4\n0\n\n0.5\n\nTime (Sec)\n\n1\n\n \n\nspeech \n\nFigure 3. Estimated target pitch for \nthe \ncocktail-party \nmixture, marked by \u201cx\u201d. The solid \nline \nthe pitch contour \nobtained from clean speech. \n\nindicates \n\nand \n\n(a)\n\n180\n\n185\n\n190\n\n195\n\nTime (ms)\n\n200\n\n205\n\n210\n\n(b)\n\n0\n\n2\n\n8\n\n10\n\n12\n\n4\n\n6\n\nLag (ms)\n\n \nFigure 4. AM effects. (a) Response of a \nfilter with center frequency 2.6 kHz. (b) \nCorresponding autocorrelation. The vertical \nline marks the position corresponding to the \npitch period of target speech. \n\n\f \n\nharmonics except for the F0 component. The rectified and filtered signal is then \nnormalized by its envelope to remove the intensity fluctuations of the original signal, \nwhere the envelope is obtained via the Hilbert Transform. Because the pitch of natural \nspeech does not change noticeably within a single frame, we model the corresponding \nnormalized signal within a T-F unit by a single sinusoid to obtain the AM repetition \nrate. Specifically, \n\nf\n\n,\n\nf\n\nij\n\nij\n\n=\n\narg\nf\n\nmin\nf\n,\n\nM\n\n=\n1\n\nk\n\n,(\u02c6[\nir\n\nkTj\n\n)\n\np\n2\nsin(\n\nfk\n\n/\n\nf\n\nS\n\n+\n\nf\n\n2\n\n)]\n\n, for f\u02db\n\n[80 Hz, 500 Hz], (2) \n\n),(\u02c6\ntir\n\n is the normalized filter response, fS is the \nwhere a square error measure is used. \nsampling frequency, M spans a frame, and T=10 ms is the progressing period from one \nframe to the next. In the above equation, fij gives the AM repetition rate for unit uij. \nNote that in the discrete case, a single sinusoid with a sufficiently high frequency can \nalways match these samples perfectly. However, we are interested in finding a \nfrequency within the plausible pitch range. Hence, the solution does not reduce to a \ndegenerate case. With appropriately chosen initial values, this optimization problem \ncan be solved effectively using iterative gradient descent (see [9]). \nThe AM criterion is used to label T-F units that do not belong to any segments \ngenerated in initial segregation; such segments, as discussed earlier, tend to miss \nunresolved harmonics. Specifically, unit uij is labeled as target speech if the final \nsquare error is less than half of the total energy of the corresponding signal and the AM \nrepetition rate is close to the estimated target pitch: \n\nt\nf\nij\n\n|\n\n|1)(\nj\n\n<\n\nq\n\n. \n\nf\n\n \n\n \n\n \n\n \n\n \n\n (3) \n\nPsychoacoustic evidence suggests that to separate sounds with overlapping spectra \nrequires 6-12% difference in F0 [6]. Accordingly, we choose q f to be 0.12. \n\n2 . 4 Fi na l se gr eg a t i on a nd r e sy nt he si s \n\nFor adjacent channels responding to unresolved harmonics, although their responses \nmay be quite different, they exhibit similar AM patterns and their response envelopes \nare highly correlated. Therefore, for T-F units labeled as target speech, segments are \ngenerated based on cross-channel envelope correlation in addition to temporal \ncontinuity. \nThe spectra of target speech and intrusion often overlap and, as a result, some segments \ngenerated in initial segregation contain both units where target speech dominates and \nthose where intrusion dominates. Given unit labels generated in the last stage, we \nfurther divide the segments in the foreground stream, SF, so that all the units in a \nsegment have the same label. Then the streams are adjusted as follows. First, since \nsegments for speech usually are at least 50 ms long, segments with the target label are \nretained in SF only if they are no shorter than 50 ms. Second, segments with the \nintrusion label are added to the background stream, SB, if they are no shorter than 50 \nms. The remaining segments are removed from SF, becoming undecided. \nFinally, other units are grouped into the two streams by temporal and spectral \ncontinuity. First, SB expands iteratively to include undecided segments in its \nneighborhood. Then, all the remaining undecided segments are added back to SF. For \nindividual units that do not belong to either stream, they are grouped into SF iteratively \nif the units are labeled as target speech as well as in the neighborhood of SF. The \nresulting SF is the final segregated stream of target speech. \nFig. 5(a) shows the new segments generated in this process for the speech and cocktail-\nparty mixture. Fig. 5(b) illustrates the segregated stream from the same mixture. Fig. \n5(c) shows all the units where target speech is stronger than intrusion. The foreground \n\n-\n-\n\n-\n\f \n\nstream generated by our algorithm contains most of the units where target speech is \nstronger. In addition, only a small number of units where intrusion is stronger are \nincorrectly grouped into it. \nA speech waveform is resynthesized from the final foreground stream. Here, the \nforeground stream works as a binary mask. It is used to retain the acoustic energy from \nthe mixture that corresponds to 1\u2019s and reject the mixture energy corresponding to 0\u2019s. \nFor more details, see [19]. \n\n3 Evaluation and comparison \n\nthe \n\nthan \n\nis greater \n\nOur model is evaluated with a corpus of 100 mixtures composed of 10 voiced \nutterances mixed with 10 intrusions collected by Cooke [4]. The intrusions have a \nconsiderable variety. Specifically, they are: N0 - 1 kHz pure tone, N1 - white noise, N2 \n- noise bursts, N3 - \u201ccocktail party\u201d noise, N4 - rock music, N5 - siren, N6 - trill \ntelephone, N7 - female speech, N8 - male speech, and N9 - female speech. \nGiven our decomposition of an input signal into T-F units, we suggest the use of an \nideal binary mask as the ground truth for target speech. The ideal binary mask is \nconstructed as follows: a T-F unit is assigned one if the target energy in the \ncorresponding unit \nintrusion energy and zero otherwise. \nTheoretically speaking, an ideal binary mask gives a performance ceiling for all binary \nmasks. Figure 5(c) illustrates the ideal mask for the speech and cocktail-party mixture. \nIdeal masks also suit well the situations where more than one target need to be \nsegregated or the target changes dynamically. The use of ideal masks is supported by \nthe auditory masking phenomenon: within a critical band, a weaker signal is masked by \na stronger one [13]. In addition, an ideal mask gives excellent resynthesis for a variety \nof sounds and is similar to a prior mask used in a recent ASR study that yields \nexcellent recognition performance [5]. \nThe speech waveform resynthesized from the final foreground stream is used for \nevaluation, and it is denoted by S(t). The speech waveform resynthesized from the ideal \nbinary mask is denoted by I(t). Furthermore, let e1(t) denote the signal present in I(t) \nbut missing from S(t), and e2(t) the signal present in S(t) but missing from I(t). Then, \nthe relative energy loss, REL, and the relative noise residue, RNR, are calculated as \nfollows: \n\n0.5\nTime (Sec)\n\n1\n\n0\n\n0.5\nTime (Sec)\n\n \nFigure 5. Results of final segregation for the speech and cocktail-party mixture. (a) \nNew segments formed in the final segregation. (b) Final foreground stream. (c) \nUnits where target speech is stronger than the intrusion. \n\n1\n\n0\n\n0.5\nTime (Sec)\n\n1\n\n2\n\nI\n\n)(\nt\n\n, \n\n2\n\n)(\ntS\n\n. \n\nt\n\nt\n\n2\n)(\nte\n1\n\n2\n)(\nte\n2\n\nt\n\nt\n\n(a)\n\n \n\n \n\n(b)\n\n \n\n \n\n \n\n \n\n (4a) \n\n (4b) \n\n \n\n \n\n(c)\n\n=\n\u0001=\n\nREL\n\nRNR\n\n5000\n\n)\nz\nH\n\n(\n \ny\nc\nn\ne\nu\nq\ne\nr\nF\n\n2355\n\n1054\n\n 387\n\n 80\n0\n\n\n\u0001\n\f Table 1: REL and RNR \n \n\nIntrusion \n\nProposed \n\nmodel \n\nWang-Brown \n\nmodel \n\nREL (%) RNR (%) REL (%) RNR (%) \n\nN0 \nN1 \nN2 \nN3 \nN4 \nN5 \nN6 \nN7 \nN8 \nN9 \n\nAverage 3.40 \n\n0.02 \n2.12 \n3.55 \n4.66 \n1.30 \n1.38 \n2.72 \n3.83 \n2.27 \n4.00 \n0.10 \n2.83 \n0.30 \n1.61 \n2.18 \n3.21 \n1.48 \n1.82 \n8.57 19.33 \n3.32 \n\n0 \n6.99 \n1.61 \n28.96 \n0.71 \n5.77 \n1.92 \n21.92 \n1.41 \n10.22 \n0 \n7.47 \n0.48 \n5.99 \n4.23 \n8.61 \n0.48 \n7.27 \n15.81 33.03 \n11.91 \n4.39 \n\n \n\n20\n\n15\n\n10\n\n5\n\n0\n\n)\n\nB\nd\n(\n \n\nR\nN\nS\n\n\u22125\n\nN0 N1 N2 N3 N4 N5 N6 N7 N8 N9\n\nIntrusion Type\n\n \nFigure 6. SNR results for segregated \nspeech. White bars show the results \nfrom the proposed model, gray bars \nthose from the Wang-Brown system, \nand black bars those of the mixtures. \n\nThe results from our model are shown in Table 1. Each value represents the average of \none intrusion with 10 voiced utterances. A further average across all intrusions is also \nshown in the table. On average, our system retains 96.60% of target speech energy, and \nthe relative residual noise is kept at 3.32%. As a comparison, Table 1 also shows the \nresults from the Wang-Brown model [18], whose performance is representative of \ncurrent CASA systems. As shown in the table, our model reduces REL significantly. In \naddition, REL and RNR are balanced in our system. \nFinally, to compare waveforms directly we measure a form of signal-to-noise ratio \n(SNR) in decibels using the resynthesized signal from the ideal binary mask as ground \ntruth: \n\nSNR\n\n=\n\n10\n\nlog\n\n[\n\n10\n\n2\n\nI\n\n)(\nt\n\n)((\ntI\n\n(\ntS\n\n))\n\n2\n\n]\n\n. \n\n \n\n \n\n (5) \n\nt\n\nt\n\nThe SNR for each intrusion averaged across 10 target utterances is shown in Fig. 6, \ntogether with the results from the Wang-Brown system and the SNR of the original \nmixtures. Our model achieves an average SNR gain of around 12 dB and 5 dB \nimprovement over the Wang-Brown model. \n\n4 Discussion \n\nThe main feature of our model lies in using different mechanisms to deal with resolved \nand unresolved harmonics. As a result, our model is able to recover target speech and \nreduce noise interference in the high-frequency range where harmonics of target speech \nare unresolved. \nThe proposed system considers the pitch contour of the target source only. However, it \nis possible to track the pitch contour of the intrusion if it has a harmonic structure. With \ntwo pitch contours, one could label a T-F unit more accurately by comparing whether \nits periodicity is more consistent with one or the other. Such a method is expected to \nlead to better performance for the two-speaker situation, e.g. N7 through N9. As \nindicated in Fig. 6, the performance gain of our system for such intrusions is relatively \nlimited. Our model is limited to separation of voiced speech. In our view, unvoiced \nspeech poses the biggest challenge for monaural speech separation. Other grouping \ncues, such as onset, offset, and timbre, have been demonstrated to be effective for \nhuman ASA [1], and may play a role in grouping unvoiced speech. In addition, one \nshould consider the acoustic and phonetic characteristics of individual unvoiced \nconsonants. We plan to investigate these issues in future study. \n\n\n\n-\n\f \n\nAc k nowl e dg me nt s \n\nWe thank G. J. Brown and M. Wu for helpful comments. Preliminary versions of this \nwork were presented in 2001 IEEE WASPAA and 2002 IEEE ICASSP. This research \nwas supported in part by an NSF grant (IIS-0081058) and an AFOSR grant (F49620-\n01-1-0027). \n\nRe f er e nce s \n\n[1] A. S. Bregman, Auditory scene analysis, Cambridge MA: MIT Press, 1990. \n\n[2] R. P. Carlyon and T. M. Shackleton, \u201cComparing the fundamental frequencies of resolved \nand unresolved harmonics: evidence for two pitch mechanisms?\u201d J. Acoust. Soc. Am., Vol. \n95, pp. 3541-3554, 1994. \n\n[3] G. Cauwenberghs, \u201cMonaural separation of independent acoustical components,\u201d In Proc. \n\nof IEEE Symp. Circuit & Systems, 1999. \n\n[4] M. Cooke, Modeling auditory processing and organization, Cambridge U.K.: Cambridge \n\nUniversity Press, 1993. \n\n[5] M. Cooke, P. Green, L. Josifovski, and A. Vizinho, \u201cRobust automatic speech recognition \n\nwith missing and unreliable acoustic data,\u201d Speech Comm., Vol. 34, pp. 267-285, 2001. \n\n[6] C. J. Darwin and R. P. Carlyon, \u201cAuditory grouping,\u201d in Hearing, B. C. J. Moore, Ed., San \n\nDiego CA: Academic Press, 1995. \n\n[7] D. P. W. Ellis, Prediction-driven computational auditory scene analysis, Ph.D. Dissertation, \n\nMIT Department of Electrical Engineering and Computer Science, 1996. \n\n[8] H. Helmholtz, On the sensations of tone, Braunschweig: Vieweg & Son, 1863. (A. J. Ellis, \n\nEnglish Trans., Dover, 1954.) \n\n[9] G. Hu and D. L. Wang, \u201cMonaural speech segregation based on pitch tracking and \namplitude modulation,\u201d Technical Report TR6, Ohio State University Department of \nComputer and Information Science, 2002. (available at www.cis.ohio-state.edu/~hu) \n\n[10] A. Hyv\u00e4rinen, J. Karhunen, and E. Oja, Independent component analysis, New York: \n\nWiley, 2001. \n\n[11] W. J. M. Levelt, Speaking: From intention to articulation, Cambridge MA: MIT Press, \n\n1989. \n\n[12] R. Meddis, \u201cSimulation of auditory-neural transduction: further studies,\u201d J. Acoust. Soc. \n\nAm., Vol. 83, pp. 1056-1063, 1988. \n\n[13] B. C. J. Moore, An Introduction to the psychology of hearing, 4th Ed., San Diego CA: \n\nAcademic Press, 1997. \n\n[14] D. O\u2019Shaughnessy, Speech communications: human and machine, 2nd Ed., New York: \n\nIEEE Press, 2000. \n\n[15] R. D. Patterson, I. Nimmo-Smith, J. Holdsworth, and P. Rice, \u201cAn efficient auditory \nfilterbank based on the gammatone function,\u201d APU Report 2341, MRC, Applied \nPsychology Unit, Cambridge U.K., 1988. \n\n[16] R. Plomp and A. M. Mimpen, \u201cThe ear as a frequency analyzer II,\u201d J. Acoust. Soc. Am., \n\nVol. 43, pp. 764-767, 1968. \n\n[17] S. Roweis, \u201cOne microphone source separation,\u201d In Advances in Neural Information \n\nProcessing Systems 13 (NIPS\u201900), 2001. \n\n[18] D. L. Wang and G. J. Brown, \u201cSeparation of speech from interfering sounds based on \n\noscillatory correlation,\u201d IEEE Trans. Neural Networks, Vol. 10, pp. 684-697, 1999. \n\n[19] M. Weintraub, A theory and computational model of auditory monaural sound separation, \n\nPh.D. Dissertation, Stanford University Department of Electrical Engineering, 1985. \n\n\f", "award": [], "sourceid": 2314, "authors": [{"given_name": "Guoning", "family_name": "Hu", "institution": null}, {"given_name": "Deliang", "family_name": "Wang", "institution": null}]}