{"title": "Noise Suppression Based on Neurophysiologically-motivated SNR Estimation for Robust Speech Recognition", "book": "Advances in Neural Information Processing Systems", "page_first": 821, "page_last": 827, "abstract": null, "full_text": "Noise suppression based on \n\nneurophysiologically-motivated  SNR \n\nestimation for  robust  speech recognition \n\nJ iirgen Tcharz \n\nMedical Physics Group \nOldenburg University \n\n26111  Oldenburg \n\nGermany \n\ntch@medi.physik.uni-oldenburg.de \n\nMichael Kleinschmidt \nMedical Physics Group \nOldenburg University \n\n26111  Oldenburg \n\nGermany \n\nBirger Kallmeier \n\nMedical Physics Group \nOldenburg University \n\n26111  Oldenburg \n\nGermany \n\nAbstract \n\nA  novel  noise  suppression  scheme  for  speech  signals  is  proposed \nwhich  is  based  on  a  neurophysiologically-motivated estimation  of \nthe  local  signal-to-noise  ratio  (SNR)  in  different  frequency  chan(cid:173)\nnels.  For  SNR-estimation,  the  input  signal  is  transformed  into \nso-called  Amplitude Modulation Spectrograms  (AMS),  which rep(cid:173)\nresent both spectral and temporal characteristics of the respective \nanalysis  frame,  and  which  imitate  the  representation  of modula(cid:173)\ntion  frequencies  in  higher  stages  of the  mammalian  auditory sys(cid:173)\ntem.  A neural network is  used to analyse AMS  patterns generated \nfrom  noisy  speech  and  estimates  the  local  SNR.  Noise  suppres(cid:173)\nsion  is  achieved  by  attenuating  frequency  channels  according  to \ntheir SNR. The noise suppression algorithm is evaluated in speaker(cid:173)\nindependent  digit  recognition  experiments  and  compared to noise \nsuppression by Spectral Subtraction. \n\n1 \n\nIntroduction \n\nOne of the major problems in automatic speech recognition  (ASR)  systems is their \nlack of robustness in  noise,  which  severely degrades their usefulness  in  many prac(cid:173)\ntical  applications.  Several proposals have been made to increase the robustness of \nASR systems, e.g.  by model compensation or more noise-robust feature extraction \n[1,  2].  Another method  to increase robustness  of ASR  systems  is  to  suppress  the \nbackground noise before feature extraction.  Classical approaches for  single-channel \nnoise  suppression are Spectral Subtraction  [3]  and  related schemes,  e.g.  [4],  where \n\n\fthe  noise  spectrum  is  usually  measured  in  detected  speech  pauses  and  subtracted \nfrom  the  signal.  In these  approaches,  stationarity of the  noise  has  to be  assumed \nwhile  speech  is  active.  Furthermore,  portions detected  as  speech  pauses  must  not \ncontain  any  speech  in  order to  allow  for  correct  noise  measurement.  At  the same \ntime,  all  actual  speech  pauses  should  be  detected  for  a  fast  update  of the  noise \nmeasurement.  In reality,  however, these partially conflicting requirements are often \nnot met. \nThe  noise  suppression  algorithm  outlined  in  this  work  directly  estimates  the  lo(cid:173)\ncal  SNR  in  a  range  of  frequency  channels  even  if  speech  and  noise  are  present \nat the  same time,  i.e.,  no  explicit  detection  of speech  pauses  and no  assumptions \non  noise  stationarity  during  speech  activity  are  necessary.  For  SNR  estimation, \nthe  input  signal  is  transformed  into  spectro-temporal  input  features,  which  are \nneurophysiologically-motivated:  experiments  on  amplitude  modulation  processing \nin  higher  stages  of the  auditory  system  in  mammals  show  that  modulations  are \nrepresented in  \"periodotopical\" gradients, which  are almost orthogonal to the tono(cid:173)\ntopical  organization  of center  frequencies  [5].  Thus,  both  spectral  and  temporal \ninformation is represented in two-dimensional maps.  These findings were applied to \nsignal  processing  in  a  binaural  noise  suppression  system  [6]  with  the introduction \nof so-called  Amplitude  Modulation  Spectrograms  (AMS) ,  which  contain  informa(cid:173)\ntion on  both  center frequencies  and modulation frequencies.  In the present study, \nthe different  representations of speech  and noise in  AMS  patterns are detected  by \na  neural  network,  which  estimates  the  local  SNR  in  each  frequency  channel.  For \nnoise  suppression,  the frequency  bands  are attenuated  according to the estimated \nlocal  SNR in the different frequency channels. \nThe  proposed  noise  suppression  scheme  is  evaluated  in  isolated-digit  recognition \nexperiments.  As recognizer, a combination of an auditory-based front end  [2]  and a \nlocally-recurrent neural network [7]  is used.  This combination was found to allow for \nmore robust isolated-digit recognition rates, compared to a standard recognizer with \nmel-cepstral features  and HMM  modeling [8,  9].  Thus, the recognition experiments \nin this study were  conducted with this particular combination to evaluate whether \na further increase of robustness can be achieved with  additional noise suppression. \n\n2  The recognition system \n\n2.1  Noise suppression \n\nFigure 1 shows the processing steps which are performed for  noise suppression.  To \ngenerate AMS  patterns which are used for  SNR estimation, the input signal (16 kHz \nsampling rate)  is  short-term level  adjusted,  i.e.,  each 32  ms segment which is  later \ntransformed into an AMS pattern is scaled to the same root-mean-square value.  The \nlevel-adjusted signal is then subdivided into overlapping segments of 4.0 ms duration \nwith a progression of 0.25 ms for each new  segment.  Each segment is multiplied by \na  Hanning  window  and padded with zeros to obtain a  frame  of 128 samples  which \nis  transformed with  a  FFT into a  complex  spectrum,  with a  spectral resolution of \n125 Hz.  The resulting 64  complex samples are considered as a function of time, i.e., \nas a band pass filtered complex time signal.  Their respective envelopes are extracted \nby squaring.  This envelope  signal is  again segmented into overlapping segments of \n128 samples (32ms) with an overlap of 64  samples.  Each segment is  multiplied with \na  Hanning  window  and  padded  with  zeros  to  obtain  a  frame  of 256  samples.  A \nfurther  FFT  is  computed  and  supplies  a  modulation  spectrum  in  each  frequency \nchannel,  with  a  modulation  frequency  resolution  of  15.6  Hz.  By  an  appropriate \nsummation of neighbouring FFT bins the frequency  axis  is  transformed to a  Bark \nscale with 15  channels,  with  center frequencies from  100-7300 Hz.  The modulation \n\n\fscaled \ninput signal \n\ninput signal \n\n-+- ---c=J \n\n~ftl;l \nov-add  LJ \n\nlevel \n\nnormalization \n\nanalysis \n\nbandpass \ntime signals \n\nenvelope \n\nmodulation \nspectrogram \nftl : !  \n\nft~ \nFFI \n~U \n~ LJ ------' \n\nrescale, \nlog(cid:173)\namplitude \n\n-[ \n\n-[ \n\noutput signal \n\n-+-\n\nFigure 1:  Processing stages of AMS-based noise suppression. \n\nfrequency  spectrum  is  scaled  logarithmically  by  appropriate  summation,  which  is \nmotivated  by  psychoacoustical  findings  about  the  shape  of  auditory  modulation \nfilters  [10).  The modulation frequency  spectrum is  restricted to the range between \n50-400  Hz  and  has  a  resolution  of 15  channels.  Thus,  the  fundamental  frequency \nof typical  voiced  speech  is  represented  in  the  modulation  spectrum.  The  AMS \nrepresentation is  restricted to a  15 times 15 pattern to limit the amount of training \ndata which is necessary to train the fully  connected perceptron.  In a last processing \nstep,  the amplitude  range  is  log-compressed.  Examples for  AMS  patterns  can  be \nseen  in  Fig.  2.  The  AMS  pattern  on  the  left  side  was  generated  from  a  voiced \nspeech portion.  The periodicity at the fundamental frequency  (approx.  110 Hz)  is \nrepresented in each center frequency  band.  The AMS  pattern on the right side was \ngenerated from  speech  simulating noise.  The typical spectral tilt  can be seen,  but \nthere is  no  structure across modulation frequencies. \n\nFor  classifying  AMS  patterns and  estimating the  narrow-band  SNR of each  AMS \npattern, a  feed-forward  neural network is  employed.  The net  consists of 225  input \nneurons  (15*15,  the AMS  resolution of center frequencies  and modulation frequen(cid:173)\ncies,  respectively),  a  hidden  layer  with  160  neurons,  and  an  output layer  with  15 \nneurons.  The  activity  of each  output  neuron  indicates the SNR in  one  of the  15 \ncenter frequency channels.  For training, the narrow-band SNRs in 15 channels were \nmeasured  for  each  AMS  analysis  frame  of  the  training  material  prior  to  adding \nspeech  and  noise.  The  neural  network  was  trained with  AMS  patterns generated \nfrom  72  min of noisy speech from  400 talkers and 41  natural noise types,  using the \nmomentum  backpropagation  algorithm.  After  training,  AMS  patterns  generated \nfrom  \"unknown\"  sound material are presented to the network.  The 15  output neu(cid:173)\nron activities that appear for each pattern serve as SNR estimates for the respective \nfrequency  channels.  In  a  detailed  study  on  AMS-based  broad-band  SNR estima(cid:173)\ntion [11)  it was shown that harmonicity which is well  represented in  AMS patterns \nis  the  most  important  cue  for  the  neural  network  to  distinguish  between  speech \nand  noise.  However,  harmonicity  is  not  the  only  cue,  as  the  algorithm  allows  for \nreliable  discrimination  between  unvoiced  speech  and  noise.  The  accuracy  of SNR \n\n\f55 \n\n73  100  135  192  246  333 \nModulation  Frequency [Hz] \n\n55 \n\n73  100  135  192  246  333 \nModulation  Frequency [Hz] \n\nFigure 2:  AMS  patterns generated from  a  voiced  speech  segment  (left),  and from \nspeech  simulating noise  (right).  Each AMS  pattern represents  a  32  ms  portion of \nthe input signal.  Bright and dark areas indicate high and low energies, respectively. \n\nestimation in terms of mean deviation between the actual and the estimated SNR in \neach frame,  for  each frequency  channel, was  determined  with  \"unknown\"  test  data \n(36  min  of noisy speech).  The average deviation  across all frequency  channels was \n5.4  dB,  with a  decrease of accuracy towards  higher  frequency  channels.  Sub-band \nSNR  estimates  are  utilized  for  noise  suppression  by  attenuating  frequency  chan(cid:173)\nnels according to their local SNR.  The gain function  which was  applied is  given by \nUk  =  (SNRk / (SNRk + 1))X , where k denotes the frequency channel, SNR the signal(cid:173)\nto-noise ratio on a linear scale, and x  is  an exponent which controls the strength of \nthe attenuation, and which was set to 1.5 for  the experiments described below. \nNoise suppression based on AMS-derived SNR estimation is performed in the FFT(cid:173)\ndomain.  The  input  signal  is  segmented  into  overlapping  frames  with  a  window \nlength of 32  ms,  and  a  shift  of 16  ms  is  applied,  i.e.,  each  window  corresponds to \none  AMS  analysis frame.  The FFT is  computed in every window.  The magnitude \nin  each  frequency  bin  is  multiplied  by the corresponding gain  computed from  the \nAMS-based  SNR  estimation.  The  gain  in  frequency  bins  which  are  not  covered \nby  the  center  frequencies  from  the  SNR  estimation  is  linearly  interpolated  from \nneighboring estimation frequencies.  The phase of the input signal is unchanged and \napplied to the attenuated magnitude spectrum.  An  inverse FFT is  computed,  and \nthe enhanced speech is  attained by overlapping and adding. \n\n2.2  Auditory-based  ASR feature  extraction \n\nThe  front  end  which  is  used  in  the  recognition  system  is  based  on  a  quantita(cid:173)\ntive model  of the  \"effective\"  peripheral  auditory  processing.  The model simulates \nboth  spectral and temporal properties of sound  processing in the  auditory system \nwhich  were  found  in  psychoacoustical  and  physiological  experiments.  The  model \nwas originally developed for  describing human performance in typical psychoacous(cid:173)\ntical spectral and temporal masking experiments, e.g., predicting the thresholds in \nbackward, simultaneous, and forward-masking experiments [12,  13].  The main pro(cid:173)\ncessing  stages of the  auditory model  are gammatone filtering,  envelope extraction \nin each frequency  channel,  adaptive amplitude compression,  and low  pass filtering \nof the envelope in each band.  The adaptive  compression stage compresses  steady(cid:173)\nstate  portions  of the  input  signal  logarithmically.  Changes  like  onsets  or  offsets, \nin contrast, are transformed linearly.  A  detailed  description  of the auditory-based \nfront end is  given in  [2]. \n\n\f2.3  Neural network recognizer \n\nFor scoring of the input features,  a locally recurrent neural network  (LRNN)  is em(cid:173)\nployed with three layers of neurons (150 input, 289 hidden, and 10 output neurons). \nHidden  layer  neurons  have  recurrent  connections  to  their  24  nearest  neighbours. \nThe  input  matrix  consists  of 5  times  the  auditory  model  feature  vector  with  30 \nelements, glued together in order to allow the network to memorize a time sequence \nof input matrices.  The network was trained using the Backpropagation-trough-time \nalgorithm with  200  iterations (see  [7]  for  a  detailed  description of the recognizer) . \n\n3  Recognition experiments \n\n3.1  Setup \n\nThe  speech  material  for  training of the  word  models  and  scoring  was  taken from \nthe  ZIFKOM  database of Deutsche  Telekom  AG.  Each  German  digit  was  spoken \nonce  by  200  different  speakers  (100  males,  100  females).  The  recording  sessions \ntook place in soundproof booths or quiet offices.  The speech material was  sampled \nat 16  kHz. \nThree  different  types  of  noise  were  added  to  the  speech  material  at  different \nsignal-to-noise ratios before feature extraction:  a)  white Gaussian noise, b)  speech(cid:173)\nsimulating  noise  which  is  characterized by a  long-term  speech  spectrum  and  am(cid:173)\nplitude modulations which  reflect an uncorrelated superposition of 6 speakers,  and \nc)  background noise recorded in  a  printing room which  strongly fluctuates  in  both \namplitude and spectral shape.  The background noises were added to the utterances \nwith signal-to-noise ratios ranging from 20 to -10 dB. The word models were trained \nwith features from  100 undisturbed and unprocessed utterances of each digit.  Fea(cid:173)\ntures for  testing  were  calculated from  another  100  utterances  of each  digit  which \nwere  distorted  by  additive noise  before  preprocessing.  The recognition  rates  were \nmeasured without noise suppression and with noise suppression as described in Sec(cid:173)\ntion 2.1. \nFor  comparison, the recognition rates were  measured with noise suppression based \non Spectral Subtraction including residual noise reduction [3]  before feature extrac(cid:173)\ntion.  Two  methods for  noise estimation  were  applied.  In the first  method,  speech \npauses in the noisy signals were detected using Voice Activity Detection (VAD)  [14]. \nThe noise measure was updated in speech pauses using a low  pass filter  with a time \nconstant  of  40  ms.  In  the  second  method,  the  noise  spectrum  was  measured  in \nspeech pauses which were detected from  the  clean utterances using an energy crite(cid:173)\nrion  (thus,  perfect  speech  pause information is  provided,  which  is  not  available  in \nreal applications). \n\n3.2  Results \n\nThe speaker-independent isolated-digit recognition rates which were obtained in the \nexperiments are plotted in Fig. 3 for  three types of background noise as a  function \nof the  SNR.  In  all  tested  noises,  noise  suppression  with  the  proposed  algorithm \nincreases the  recognition  rate  in  comparison  with  the  unprocessed  data and  with \nSpectral  Subtraction  with  VAD-based  noise  measurement.  Spectral  Subtraction \nwith  perfect  speech  pause  detection  allows  for  higher  recognition  rates  than  the \nAMS-based approach in stationary white noise.  Here, the noise measure for Spectral \nSubtraction  is  very  accurate  during  speech  activity  and  allows  for  effective  noise \nremoval.  AMS-based noise suppression estimates the SNR in every analysis frame, \nand no a priori information on speech-free segments is provided to the algorithm.  In \n\n\fWhite noise \n\nPrinting room  noise \n\n1 00  ~=,..,...--r-.---'----r---'----'----. \n\n. ~>!.:~~:-<~----\" \n\n~ 90 \n80 \n~  70 \n<:  60 \n:8  50 \n0 \nx \n'2: \nCl  40  AMS-based  ____ , ___ .  , ,-.. \n\u00a7  30 \n\". \na:  20 \n. . \n\n*  \" \nnoalgo  --+-- ~. \n\nSS  VAD  . .. ,.... \nSSJ)erf  \u00b7\u00b7\u00b7\n.... \n\n0 \n\n1 0  L...L--H-.....L:::....L-----'-----''---'-----'---'----' \n\nclean  20  15  1 0  5  0 \n\n-5  -10 \n\n20  15  1 0  5  0 \n\n-5  -10 \n\nSNR [dB) \n\nSpeech simulating noise \n\n100  r---~\u00b7 .. ~. F . .. ~ .. ~_~.~'--.-r-, \n\n~ 90 \nCD  80 \n~  70 \n<:  60 \n:8  50 \nnoalgo  --+--\n'2: \nCl  40  AMS-based  ____ , ___ . \n\u00a7  30 \nSS  VAD  ... ,. ... \na:  20 \nSSJ)erfo \n\n1 0  L...L--H-.....L::....L-----'-----''---'-----'---'----' \n\nclean  20  15  1 0  5  0 \n\n-5  -10 \n\nSNR [dB) \n\nFigure  3:  Speaker-independent,  iso(cid:173)\nlated  digit  recognition  rates  for  three \ntypes of noise as a function  of the SNR \nwithout  noise  suppression \n(noalgo ), \nwith  AMS-based  noise  suppression, \nSpectral  Subtraction  with  VAD-based \nnoise  measurement,  and  Spectral  Sub(cid:173)\ntraction  with  perfect  speech  pause  in(cid:173)\nformation. \n\nSNR [dB) \n\nspeech simulation noise, which fluctuates in level but not in spectral shape, Spectral \nSubtraction with  perfect  speech  pause  detection  works  slightly  better than  AMS(cid:173)\nbased  noise  suppression.  In  printing  room  noise,  which  fluctuates  in  both  level \nand  spectrum,  the  AMS-based  approach  yields  the  best  results.  Here,  Spectral \nSubtraction  even  degrades  the  recognition  rates  in  some  SNRs,  compared  to  the \nunprocessed  data.  The  noise  measure  from  VAD-based  or  perfect  speech  pause \ndetection cannot be updated while speech is  active.  Thus, an incorrect spectrum is \nsubtracted  and leads  to  artifacts  and  degraded  recognition  performance.  In  clean \nspeech,  recognition  rates  of  99.5%  for  unprocessed  speech,  99.1%  after  Spectral \nSubtraction, and 98.9% after AMS-based noise suppression were  obtained. \n\n4  Discussion \n\nThe proposed neurophysiologically-motivated noise  suppression scheme was  shown \nto significantly improve  digit  recognition  in  noise  in  comparison  with  unprocessed \ndata  and  with  Spectral  Subtraction  using  VAD-based  noise  measures.  A  perfect \nspeech  pause  detection  (which  is  not  available  yet  in  real  systems)  allows  for  a \nreliable  estimation  of the noise  floor  in  stationary noise.  In  non-stationary noise, \nhowever,  the  AMS  pattern-based signal  classification  and  noise  suppression  is  ad(cid:173)\nvantageous,  as  it  does  not  depend  on  speech  pause  detection  and  no  assumption \nis  necessary about the noise  being  stationary while speech is  active.  Spectral Sub(cid:173)\ntraction  as  described  in  [3]  produces  musical  tones,  i.e.  fast  fluctuating  spectral \npeaks.  The  neurophysiologically-based  noise  suppression  scheme  outlined  in  this \npaper does  not produce  such fast  fluctuating  artifacts.  In  general,  a  good  quality \nof speech is  maintained.  The choice of the attenuation exponent  x  has only  little \nimpact  on the  quality  of speech  in favourable  SNRs.  With  decreasing  SNR,  how(cid:173)\never,  there  is  a  tradeoff between  the amount  of noise  suppression  and  distortions \n\n\fof the  speech.  A  typical  distortion  of speech  in  poor  signal-to-noise  ratios  is  an \nunnatural spectral  \"coloring\",  rather than fast  fluctuating  distortions.  In informal \ntests, most listeners did not have the impression that the algorithm improves speech \nintelligibility,  but  clearly  preferred  the  processed signal over  the unprocessed  one, \nas  the  background  noise  was  significantly  suppressed  without  annoying  artifacts. \nClean  speech  is  almost  perfectly  preserved after  processing.  The performance and \ncharacteristics of the algorithm of course strongly depends on the training data, as \nonly lttle knowledge on the differences between speech and noise is  \"hard wired\". \n\nAcknowledgments \n\nWe  thank  Klaus  Kasper  and  Herbert  Reininger  from  Institut  fUr  Angewandte \nPhysik,  Universitat  Frankfurt/M.  for  supplying  us  with  their  LRNN  implemen(cid:173)\ntation. \n\nReferences \n\n[1]  Hermansky,  H.  and Morgan,  N.  (1994).  RASTA  processing  of speech.  IEEE  Trans. \n\nSpeech Audio Processing  2(4),  pp.  578-589 \n\n[2]  Tchorz,  J.  and Kollmeier,  B.  (1999) .  A  Model  of Auditory  Perception  as  Front  End \n\nfor  Automatic Speech  Recognition.  J.  Acoust.  Soc.  Am.  106,  pp.  2040-2050 \n\n[3]  Boll,  S.  (1979).  Suppression  of acoustic  noise  in  speech  using  spectral  subtraction. \n\nIEEE Trans.  Acoust ., Speech, Signal  Processing  27(2) ,  pp.  113- 120 \n\n[4]  Ephraim,  Y.  and  Malah,  M.  (1984).  Speech  enhancement  using  a  minimum mean(cid:173)\n\nsquare error short-time spectral  amplitude estimator.  IEEE Trans.  Acoust.,  Speech, \nSignal  Processing  32(6),  pp.  1109-1121 \n\n[5]  Langner,  G.,  Sams,  M.,  Heil,  P.,  and Schulze, H.,  (1997) .  Frequency and  periodicity \nare  represented  in  orthogonal  maps  in  the  human  auditory  cortex:  evidence  from \nmagnetoencephalography.  J.  Compo  Physiol.  A  181,  pp.  665- 676 \n\n[6]  Kollmeier,  B.  and Koch,  R.,  (1994) .  Speech enhancement based on physiological  and \npsycho acoustical models of modulation perception and binaural interaction. J. Acoust. \nSoc.  Am. 95,  pp.  1593- 1602 \n\n[7]  Kasper,  K.,  Reininger,  H.,  Wolf,  D.,  and Wiist,  H.  (1995).  A  speech recognizer  with \nlow  complexity based  on  RNN.  In:  Neural  Networks  for  Signal  Processing  V,  Proc. \nof the IEEE workshop,  Cambridge  (MA),  pp.  272- 281 \n\n[8]  Kasper,  K.,  Reininger,  R.,  and Wolf,  D.  (1997).  Exploiting the potential of auditory \npreprocessing for robust speech recognition by locally recurrent neural networks. Proc. \nInt. Conf.  Acoustics, Speech  and Signal  Processing  (ICASSP)  2,  pp.  1223- 1227 \n\n[9]  Kleinschmidt,  M.,  Tchorz,  J .,  and Kollmeier,  B.  (2000).  Combining speech  enhance(cid:173)\nment and auditory feature  extraction for  robust speech recognition.  Speech  Commu(cid:173)\nnication,  Special  issue  on robust  ASR (accepted) \n\n[10]  Ewert,  S.  and  Dau,  T .  (1999) .  Frequency  selectivity  in  amplitude-modulation  pro(cid:173)\n\ncessing.  J.  Acoust. Soc.  Am.  (submitted) \n\n[11]  Tchorz,  J.  and  Kollmeier,  B.  (2000).  Estimation  of  the  signal-to-noise  ratio  with \n\namplitude modulation spectrograms.  Speech Communication  (submitted) \n\n[12]  Dau, T ., Piischel,  D., and Kohlrausch,  A. (1996) . A  quantitative model of the  \"effec(cid:173)\ntive\"  signal  processing in the auditory system:  II.  Simulations and measurements.  J. \nAcoust.  Soc.  Am 99,  pp. 3623- 3631 \n\n[13]  Dau,  T.,  Kollmeier,  B.,  and  Kohlrausch,  A.  (1997).  Modeling  auditory  processing \nof  amplitude  modulation:  I.  Modulation  Detection  and  masking  with  narrowband \ncarriers. J . Acoust.  Soc.  Am 102,  pp.  2892- 2905 \n[14]  Recommendation ITU-T  G.729  Annex B,  1996 \n\n\f", "award": [], "sourceid": 1902, "authors": [{"given_name": "J\u00fcrgen", "family_name": "Tchorz", "institution": null}, {"given_name": "Michael", "family_name": "Kleinschmidt", "institution": null}, {"given_name": "Birger", "family_name": "Kollmeier", "institution": null}]}