{"title": "Glove-TalkII: Mapping Hand Gestures to Speech Using Neural Networks", "book": "Advances in Neural Information Processing Systems", "page_first": 843, "page_last": 850, "abstract": null, "full_text": "Glove-TalkII:  Mapping Hand Gestures to \n\nSpeech Using Neural Networks \n\nS.  Sidney Fels \n\nGeoffrey Hinton \n\nDepartment of Computer Science \n\nDepartment of Computer Science \n\nUniversity of Toronto \nToronto, ON,  M5S  lA4 \nssfels@ai.toronto.edu \n\nUniversity of Toronto \nToronto, ON,  M5S  lA4 \nhinton@ai.toronto.edu \n\nAbstract \n\nGlove-TaikII is  a  system which  translates  hand gestures  to speech \nthrough  an  adaptive interface.  Hand  gestures  are  mapped  contin(cid:173)\nuously  to  10  control  parameters  of a  parallel formant  speech  syn(cid:173)\nthesizer.  The mapping allows the hand to act  as  an  artificial  vocal \ntract  that  produces  speech  in  real  time.  This  gives  an  unlimited \nvocabulary  in addition to direct  control of fundamental  frequency \nand  volume.  Currently,  the  best  version  of Glove-TalkII  uses  sev(cid:173)\neral  input  devices  (including  a  CyberGlove,  a  ContactGlove,  a  3-\nspace  tracker,  and a foot-pedal),  a  parallel formant speech  synthe(cid:173)\nsizer  and  3 neural  networks.  The gesture-to-speech  task  is  divided \ninto  vowel  and  consonant  production  by  using  a  gating  network \nto weight  the outputs of a  vowel  and  a  consonant  neural  network. \nThe  gating  network  and  the  consonant  network  are  trained  with \nexamples  from  the  user.  The  vowel  network  implements  a  fixed, \nuser-defined  relationship  between  hand-position  and  vowel  sound \nand does  not require any training examples from the user.  Volume, \nfundamental  frequency  and  stop  consonants  are  produced  with  a \nfixed  mapping from  the  input devices.  One subject  has  trained  to \nspeak  intelligibly with Glove-TalkII. He  speaks slowly with speech \nquality  similar  to a  text-to-speech  synthesizer  but  with  far  more \nnatural-sounding pitch  variations. \n\n\f844 \n\nS.  Sidney  Fe Is,  Geoffrey  Hinton \n\nIntroduction \n\n1 \nThere  are many different  possible schemes  for  converting  hand gestures  to speech. \nThe  choice  of scheme  depends  on  the  granularity of the  speech  that  you  want  to \nproduce.  Figure 1 identifies a spectrum defined  by possible divisions of speech based \non  the  duration  of the  sound  for  each  granularity.  What  is  interesting  is  that  in \ngeneral,  the  coarser  the  division  of speech,  the  smaller  the  bandwidth  necessary \nfor  the  user.  In  contrast,  where  the granularity of speech  is  on  the order  of artic(cid:173)\nulatory  muscle  movements  (i.e.  the  artificial  vocal  tract  [AVT])  high  bandwidth \ncontrol is  necessary for good speech.  Devices which  implement this model of speech \nproduction  are  like  musical  instruments  which  produce  speech  sounds.  The  user \nmust  control  the  timing  of sounds  to  produce  speech  much  as  a  musician  plays \nnotes  to  produce  music.  The  AVT  allows  unlimited  vocabulary,  control  of pitch \nand  non-verbal  sounds.  Glove-TalkII  is  an  adaptive  interface  that  implements  an \nAVT. \nTranslating gestures  to speech  using an  AVT model  has a long history beginning in \nthe  late  1700's.  Systems  developed  include a  bellows-driven  hand-varied  resonator \ntube with auxiliary controls (1790's [9]),  a  rubber-moulded skull  with actuators for \nmanipulating tongue and jaw position  (1880's  [1])  and  a  keyboard-footpedal  inter(cid:173)\nface  controlling  a  set  of linearly  spaced  bandpass  frequency  generators  called  the \nYoder  (1940 [3]).  The Yoder was demonstrated  at the World's Fair in  1939 by oper(cid:173)\nators who had trained continuously for  one year  to learn to speak  with the system. \nThis suggests that the task of speaking with a gestural interface is very difficult and \nthe  training  times could  be significantly  decreased  with  a  better  interface.  Glove(cid:173)\nTalkII  is  implemented  with  neural  networks  which  allows  the  system  to  learn  the \nuser's  interpretation of an articulatory model of speaking. \n\nThis paper begins  with  an overview of the whole  Glove-TalkII system.  Then,  each \nneural  network  is  described  along  with  its  training  and  test  results.  Finally,  a \nqualitative analysis is provided of the speech  produced  by a single subject  after  100 \nhours of speaking  with Glove-TalkI!. \n\nArtificial \n\nVocal Tract \n(AVT) \n\nI \n10-30 \n\nPhoneme \nGenerator \n\nFinger \nSpelling \n\nSyllable \nGenerator \n\nWord \n\nGenerator \n\n> \n\n100 \n\n130 \n\n200 \n\n600 \n\nApproximate time per gesture (msec) \n\nFigure  1:  Spectrum  of  gesture-to-speech  mappings  based  on  the  granularity  of \nspeech. \n\n2  Overview of Glove-TalkII \nThe Glove-TalkIl system  converts  hand  gestures  to speech ,  based  on  a gesture-to(cid:173)\nformant  model.  The  gesture  vocabulary  is  based  on  a  vocal-articulator  model  of \nthe hand.  By  dividing the mapping tasks into independent  subtasks, a  substantial \nreduction  in network size  and training time  is  possible (see  [4]). \n\nFigure 2 illustrates the whole Glove-TalkIl system.  Important features  include the \n\n\f845 \n\n)))S~ \n\nGlove- Talkll \n\nRlabt Hand data \n\nll.y.:r: \n\nroll.l!itch. yaw \nevery 1/60 IIOCUld \n\n~~ \n\n~, \n\nk-.--'-\n\n10 flell analea \n\n4  abduction anclea \nthumb and pinkie rotation \nwrist pitch and yaw \nevery 1/100 aecond \n\nFllled Pitch \nMappina \n\nVIC Doc:ision \n\nNetwork \n\nVowel \nNetwork \n\nCcnsonant \nNetwork \n\nFllled SlOp \nMappina \n\nI---~ Combinina \nFunction \n\nFigure  2:  Block  diagram of Glove-TalkII:  input  from  the  user  is  measured  by  the \nCyberglove,  polhemus,  keyboard  and  foot  pedal,  then  mapped  using  neural  net(cid:173)\nworks  and fixed  functions  to formant  parameters which  drive  the  parallel formant \nsynthesizer  [8]. \n\nthree  neural  networks  labeled  vowel/consonant  decision  (V /C),  vowel,  and  conso(cid:173)\nnant.  The V /C network is trained on data collected from the user to decide whether \nhe wants to produce a vowel or a consonant sound .  Likewise,  the consonant network \nis trained to produce consonant sounds based on user-generated  examples based on \nan  initial  gesture  vocabulary.  In  contrast,  the  vowel  network  implements  a  fixed \nmapping  between  hand-positions  and  vowel  phonemes  defined  by  the  user.  Nine \ncontact  points  measured  on  the  user's  left  hand  by  a  ContactGlove  designate  the \nnine stop consonants  (B,  D,  G, J, P,  T, K,  CH,  NG), because  the dynamics of such \nsounds  proved  too  fast  to  be  controlled  by  the  user.  The  foot  pedal  provides  a \nvolume  control  by  adjusting the speech  amplitude and  this  mapping is  fixed.  The \nfundamental frequency,  which  is  related to the pitch of the speech,  is  determined by \na  fixed  mapping from  the  user's  hand height.  The output  of the  system  drives  10 \ncontrol  parameters of a  parallel formant speech  synthesizer  every  10  msec.  The  10 \ncontrol parameters are:  nasal amplitude (ALF), first,  second  and third formant fre(cid:173)\nquency  and amplitude (F1,  A1,  F2, A2,  F3,  A3), high  frequency  amplitude (AHF), \ndegree  of voicing (V)  and fundamental frequency  (FO).  Each of the control  param(cid:173)\neters  is  quantized  to 6  bits. \n\nOnce  trained,  Glove-Talk II  can  be  used  as  follows: \nto  initiate  speech,  the  user \nforms  the  hand shape of the first  sound she  intends  to  produce.  She  depresses  the \nfoot  pedal  and  the  sound  comes  out  of the  synthesizer.  Vowels  and  consonants \nof various  qualities  are  produced  in  a  continuous  fashion  through  the  appropriate \nco-ordination of hand  and foot  motions.  Words  are  formed  by  making  the  correct \nmotions;  for  example,  to say  \"hello\"  the  user  forms  the  \"h\"  sound,  depresses  the \nfoot pedal and quickly moves her hand to produce the  \"e\"  sound, then the  \"I\" sound \nand finally  the  \"0\"  sound.  The user  has  complete control of the  timing and quality \nof the  individual sounds.  The  articulatory  mapping  between  gestures  and  speech \n\n\f846 \n\nS.  Sidney  Fe/s,  Geoffrey  Hinton \n\nFigure 3:  Hand-position  to  Vowel  Sound \nMapping.  The  coordinates  are  specified \nrelative to the origin at the sound A.  The \nX  and  Y  coordinates  form  a  horizontal \nplane parallel to the floor when the user is \nsitting.  The  11  cardinal  phoneme targets \nare  determined  with  the  text-to-speech \nsynthesizer. \n\n... .-),~O\"\"':--+--\"'--7--\n... -~-~c- - - - : - -\n\n... --!----\"~ -Y(cm) \n\nu \n\nI .. \n\nis  decided  a  priori.  The  mapping  is  based  on  a  simplistic  articulatory  phonetic \ndescription  of speech  (5].  The  X,Y  coordinates  (measured  by  the  polhemus)  are \nmapped  to something like tongue position and  height l  producing vowels when  the \nuser's  hand  is  in  an  open  configuration  (see  figure  2  for  the  correspondence  and \ntable  1  for  a  typical  vowel  configuration).  Manner  and  place  of articulation  for \nnon-stop  consonants  are  determined  by  opposition  of the  thumb  with  the  index \nand  middle fingers  as  described  in  table  1.  The  ring  finger  controls  voicing.  Only \nstatic articulatory configurations are used as training points for the neural networks, \nand  the  interpolation  between  them  is  a  result  of the learning  but is  not explicitly \ntrained.  Ideally,  the  transitions  should  also  be  learned,  but  in  the  text-to-speech \nformant  data we  use for  training [6]  these  transitions  are  poor,  and  it is  very  hard \nto extract formant  trajectories from  real  speech  accurately. \n\n2.1  The Vowel/Consonant  (VIC)  Network \nThe VIC  network  decides,  on  the  basis  of the  current  configuration  of the  user's \nhand,  to emit a  vowel  or a  consonant sound.  For  the quantitative results  reported \nhere,  we  used  a  10-5-1  feed-forward  network  with  sigmoid  activations  [7].  The  10 \ninputs are  ten  scaled  hand  parameters  measured  with  a  Cyberglove:  8 flex  angles \n(knuckle  and  middle joints of the  thumb,  index,  middle and  ring  fingers),  thumb \nabduction  angle  and  thumb  rotation  angle.  The output  is  a  single  number  repre(cid:173)\nsenting  the probability that  the hand  configuration  indicates a  vowel.  The output \nof the  VIC  network  is  used  to  gate  the  outputs of the  vowel  and  consonant  net(cid:173)\nworks,  which  then  produce a  mixture of vowel  and  consonant formant  parameters. \nThe training data available includes only user-produced  vowel or consonant sounds. \nThe network interpolates between hand configurations to create a smooth but fairly \nrapid  transition between  vowels and consonants. \n\nFor  quantitative  analysis,  typical  training  data  consists  of 2600  examples  of con(cid:173)\nsonant  configurations  (350  approximants,  1510  fricatives  [and  aspirant],  and  740 \nnasals)  and  700  examples  of vowel  configurations.  The  consonant  examples  were \nobtained from training data collected  for  the consonant  network  by  an expert user. \nThe vowel examples were collected from the user  by  requiring him to move his hand \nin  vowel  configurations  for  a  specified  amount  of time.  This  procedure  was  per(cid:173)\nformed  in several sessions.  The test set consists of 1614 examples (1380 consonants \nand  234 vowels).  After  training,2  the mean squared  error on  the training and  test \n\nlIn reality,  the  XY coordinates  map  more closely  to changes in  the first  two  formants, \nFI  and  F2 of vowels.  From  the user's perspective though,  the link to tongue  movement is \nuseful. \n2The  V Ie network,  the  vowel  network  and  the  consonant  network  are  trained  using \n\n\fGlove-Talkll \n\n847 \n\nF \n\nDH \n\n'.:~:~:::. \n\n'.~ \n.. \n;J(il. \n~ \u2022 \u2022\u2022 :$ \u2022\u2022\u2022 \n\n, . . .;-\n~ ~  ~ ~ \n7\u00b7: .... :\u00b7 , I.!:.  ~ .:::~::.  ~ ~:;\u00a7::. \nto:.\u00b7.\u00b7.  -\":\"~\"'\" \n\" ~':\" ':: ..  -' ... ~: \n~ ~ ~ ~ \n\n. (, \n. ~ \n\":'s \n~~. \n\n.. ' \n....... \n'.>' \n\nTH \n\n., \n. ... ). ... \n\nSH \n\nM \n\n''':;;1.. \n\nL \n\nR \n\n:<: \u2022 \n\nH \n\nS \n\nN \n\n, \n\n' . \n\u2022  0:' :~: \n\n~.,,> \u2022  ~ \n\n\u2022 \u2022\u2022\u2022 '<1. \u2022\u2022 \n\nV \n\nW \n\nZ \n\nZH \n\nvowel \n\nTable 1:  Static Gesture-to-Consonant Mapping for all phonemes.  Note, each gesture \ncorresponds to a static non-stop consonant phoneme generated by the text-to-speech \nsynthesizer. \n\nset  was  less  than  10-4 . \nDuring  normal  speaking  neither  network  made  perceptual  errors.  The  decision \nboundary  feels  quite sharp,  and  provides  very  predictable,  quick  transitions  from \nvowels to consonants and back.  Also,  vowel sounds are produced when  the user  hy(cid:173)\nperextends his hand.  Any unusual configurations that would intuitively be expected \nto produce  consonant sounds do  indeed  produce  consonant sounds. \n2.2  The Vowel  Network \nThe  vowel  network  is  a  2-11-8  feed  forward  network.  The  11  hidden  units  are \nnormalized  radial  basis functions  (RBFs)  [2]  which  are  centered  to respond  to one \nof 11  cardinal  vowels.  The  outputs  are  sigmoid  units  representing  8  synthesizer \ncontrol parameters (ALF,  F1,  AI,  F2,  A2,  F3,  A3,  AHF). The radial basis function \nused  is: \n\nL(Wji-O.)~ \n\n(1) \nwhere  OJ  is the (un-normalized) output of the RBF unit, Wji  is the weight from unit \ni  to unit j, 0i  is  the output of input unit i,  and (1/  is  the variance of the RBF. The \nnormalization used  is: \n\noj=e-\n\n<l'j2 \n\nnj  =  L \n\nO\u00b7 \nJ \n\nmEpom \n\n(2) \n\nwhere  nj is  the normalized output of unit j  and the summation is over  all the units \nin  the  group  of normalized  RBF  units.  The  centres  of the  RBF  units  are  fixed \n\nconjugate gradient descent  and  a  line search. \n\n\f848 \n\nS.  Sidney  Fels,  Geoffrey  Hinton \n\naccording to the X and Y values of each of the 11  vowels  in the predefined mapping \n(see  figure  2).  The variances of the  11  RBF's are set  to 0.025. \n\nThe weights  from  the  RBF  units  to the output units are  trained.  For  the  training \ndata,  100  identical examples  of each  vowel  are generated  from  their  corresponding \nX  and Y  positions  in  the user-defined  mapping,  providing 1100 examples.  Noise  is \nthen  added  to the  scaled X  and  Y  coordinates for  each  example.  The added  noise \nis  uniformly distributed  in  the  range -0.025  to 0.025.  In  terms  of unscaled  ranges, \nthese  correspond  to an X  range of approximately \u00b1  0.5 cm and a Y  range of \u00b1 0.26 \ncm. \n\nThree different  test sets  were  created.  Each test  set  had 50  examples of each  vowel \nfor  a  total  of 550  examples.  The  first  test  set  used  additive  uniform  noise  in  the \ninterval \u00b1  0.025.  The second  and third test sets  used  additive uniform noise  in  the \ninterval \u00b1  0.05 and \u00b1 0.1  respectively. \nThe mean squared  error  on the training set  was  0.0016.  The  MSE on  the  additive \nnoise  test  sets  (noise  =  \u00b1  0.025,  0.05  and  0.01)  was  0.0018,  0.0038,  0.0120  which \ncorresponds  to expected  errors  of 1.1 %,  3.1 % and  5.5%  in  the formant parameters, \nrespectively.  This network  performs  well  perceptually.  The  key  feature  is  the nor(cid:173)\nmalization of the RBF units.  Often, when speaking, the user will overshoot cardinal \nvowel  positions (especially  when she  is  producing dipthongs)  and all  the RBF units \nwill be quite suppressed.  However,  the normalization magnifies any slight difference \nbetween  the  activities of the  units  and  the  sound  produced  will  be  dominated  by \nthe cardinal  vowel  corresponding  to the one  whose  centre  is  closest  in  hand space. \n\n2.3  The Consonant  Network \nThe  consonant  network  is  a  10-14-9  feed-forward  network.  The  14  hidden  units \nare  normalized  RBF  units.  Each  RBF  is  centred  at  a  hand  configuration  deter(cid:173)\nmined from  training data collected  from  the user  corresponding  to one  of 14  static \nconsonant  phonemes.  The target consonants are created with a  text-to-speech  syn(cid:173)\nthesizer.  Figure  1 defines  the  initial  mapping for  each  of the  14  consonants.  The \n9  sigmoid  output  units  represent  9  control  parameters  of the  formant  synthesizer \n(ALF,  F1,  AI,  F2,  A2,  F3,  A3,  AHF,  V).  The  voicing parameter  is  required  since \nconsonant sounds  have different  degrees  of voicing.  The inputs are  the same as  for \nthe manager  V Ie network. \nTraining and  test  data for  the  consonant  network  is  obtained  from  the  user .  Tar(cid:173)\nget  data  is  created  for  each  of the  14  consonant  sounds  using  the  text-to-speech \nsynthesizer .  The scheme  to collect  data for  a single consonant  is: \n\n1.  The target consonant is  played for  100 msec through the speech synthesizer; \n2.  the user  forms  a  hand configuration corresponding to the consonant; \n3.  the  user  depresses  the foot  pedal to begin  recording; the start of recording \n\nis  indicated  by  the appearance of a green  square; \n\n4.  10-15 time steps of hand data are collected and stored with the correspond(cid:173)\n\ning  formant  parameter  targets  and  phoneme  identifier;  the  end  of  data \ncollection  is  indicated  by  turning the green  square  red; \n\n5.  the user  chooses  whether to save the data to a file,  and whether  to redo the \n\ncurrent  target or  move  to the next one. \n\n\fGlove- Talkll \n\n849 \n\nUsing this procedure 350 approximants, 1510 fricatives and 700 nasals were collected \nand scaled  for  the training data.  The hand data were  averaged for  each  consonant \nsound to form  the  RBF centres.  For the test  data, 255  approximants, 960 fricatives \nand  165  nasals  were  collected  and  scaled .  The RBF  variances  were  set  to 0.05. \n\nThe  mean  square  error  on  the  training set  was  0.005  and  on  the  testing  set  was \n0.01  corresponding to expected  errors of 3.3% and 4.7% in the formant parameters, \nrespectively.  Listening  to  the  output  of the  network  reveals  that  each  sound  is \nproduced reasonably well  when the user's hand is  held  in  a fixed  position.  The only \ndifficulty is  that the Rand L sounds are very sensitive to motion of the index finger. \n\n3  Qualitative Performance of Glove-TalkII \nOne subject,  who is  an accomplished  pianist, has been  trained extensively to speak \nwith  Glove-TalkII.  We  expected  that  his  pianistic skill  in  forming  finger  patterns \nand  his  musical  training  would  help  him  learn  to  speak  with  Glove-TalkII. After \n100  hours  of  training,  his  speech  with  Glove-TalklI  is  intelligible  and  somewhat \nnatural-sounding.  He still finds  it difficult to speak  quickly,  pronounce  polysyllabic \nwords,  and speak spontaneously. \n\nDuring his  training, Glove-TalkII also adapted to suit changes  required  by  the sub(cid:173)\nject.  Initially, good performance of the VIC  network  is  critical for  the user  to learn \nto speak.  If the V Ie network performs poorly the user hears a mixture of vowel and \nconsonant sounds making it difficult to adjust his hand configurations to say differ(cid:173)\nent  utterances.  For  this  reason,  it  is  important  to  have  the  user  comfortable  with \nthe  initial  mapping so  that  the  training data collected  leads  to  the  VIC  network \nperforming well.  In the  100  hours of practice,  Glove-Talk II was  retrained  about 10 \ntimes.  Four  significant  changes  were  made from  the  original system  analysed  here \nfor  the new  subject.  First, the NG  sound was added  to the non-stop  consonant list \nby adding an additional hand shape, namely the user touches his pinkie to his thumb \non his right hand.  To accomodate this change, the consonant and VIC network  had \ntwo inputs added to represent  the two flex  angles of the pinkie.  Also,  the consonant \nnetwork  has  an  extra hidden  unit  for  the  NG  sound.  Second,  the  consonant  net(cid:173)\nwork  was  trained  to allow  the  RBF  centres  to change.  After  the  hidden-to-output \nweights were trained until little improvement was seen,  the input-to-hidden weights \n(i.e.  the  RBF  centres)  were  also  allowed  to  adapt .  This noticeably  improved  per(cid:173)\nformance  for  the  user.  Third,  the  vowel  mapping  was  altered  so  that  the  I  was \nmoved  closer  to  the  EE sound  and  the  entire  mapping  was  reduced  to  75%  of its \nsize.  Fourth,  for  this subject,  the VIC  network  needed  was  a  10-10-1  feed-forward \nsigmoid unit network.  Understanding the interaction between the user's adaptation \nand  Glove-TalkII's adaptation remains  an interesting  research  pursuit. \n\n4  Summary \nThe initial  mapping is  loosely  based  on  an  articulatory  model  of speech.  An  open \nconfiguration  of  the  hand  corresponds  to  an  unobstructed  vocal  tract,  which  in \nturn  generates  vowel  sounds.  Different  vowel  sounds  are  produced  by  movements \nof the hand  in  a  horizontal  X-Y  plane  that  corresponds  to  movements  of the  first \ntwo formants  which  are  roughly related  to tongue position .  Consonants other than \nstops are produced by closing the index,  middle, or ring fingers or flexing the thumb, \nrepresenting  constrictions  in  the  vocal  tract .  Stop  consonants  are  produced  by \n\n\f850 \n\nS.  Sidney  Fels,  Geoffrey  Hinton \n\ncontact switches  worn on  the user's  left  hand.  FO  is  controlled  by  hand height  and \nspeaking intensity  by foot  pedal depression. \nGlove-TaikII learns the user's  interpretation of this initial mapping.  The VIC net(cid:173)\nwork  and  the  consonant  network  learn  the  mapping  from  examples  generated  by \nthe user  during phases of training.  The vowel  network  is  trained on examples  com(cid:173)\nputed  from  the  user-defined  mapping  between  hand-position  and  vowels.  The  FO \nand volume mappings are  non-adaptive. \n\nOne subject was  trained  to use  Glove-TalkII. After  100  hours of practice  he  is  able \nto speak  intelligibly.  His  speech  is  fairly  slow  (1.5  to  3  times  slower  than  normal \nspeech)  and somewhat robotic.  It sounds similar to speech  produced with a text-to(cid:173)\nspeech synthesizer but has a more natural intonation contour which greatly improves \nthe intelligibility and naturalness of the speech.  Reading novel  passages  intelligibly \nusually  requires  several  attempts,  especially  with  polysyllabic  words.  Intelligible \nspontaneous speech  is  possible but difficult. \n\nAcknowledgements \nWe  thank  Peter  Dayan,  Sageev  Oore  and  Mike  Revow  for  their  contributions. \nThis  research  was  funded  by  the  Institute  for  Robotics  and  Intelligent  Systems \nand  NSERC.  Geoffrey  Hinton  is  the  Noranda fellow  of the  Canadian  Institute  for \nAdvanced  Research. \n\nReferences \n\n[1]  A.  G.  Bell.  Making a  talking-machine. In  Beinn Bhreagh  Recorder, pages 61-72, \n\nNovember  1909~ \n\n[2]  D. Broomhead and D. Lowe.  Multivariable functional interpolation and adaptive \n\nnetworks.  Complex Systems,  2:321-355,  1988. \n\n[3]  Homer Dudley,  R.  R.  Riesz,  and S.  S.  A.  Watkins.  A synthetic speaker.  Journal \n\nof the  Franklin  Institute,  227(6):739-764, June  1939. \n\n[4]  S.  S.  Fels.  Building adaptive interfaces  using  neural  networks:  The Glove-Talk \n\npilot study.  Technical  Report CRG-TR-90-1, University of Toronto,  1990. \n\n[5]  P.  Ladefoged.  A  course  in  Phonetics  (2  ed.).  Harcourt  Brace  Javanovich,  New \n\nYork,  1982. \n\n[6]  E.  Lewis.  A  'C' implementation of the JSRU  text-to-speech  system.  Technical \n\nreport,  Computer Science  Dept.,  University of Bristol,  1989. \n\n[7]  D.  E.  Rumelhart,  G.  E.  Hinton,  and  R.  J.  Williams.  Learning  internal  repre(cid:173)\n\nsentations by  back-propagating errors.  Nature,  323:533-536, 1986. \n\n[8]  J. M.  Rye and J. N.  Holmes.  A versatile software parallel-formant speech synthe(cid:173)\n\nsizer.  Technical  Report  JSRU-RR-1016, Joint  Speech  Research  Unit,  Malvern, \nUK,  1982. \n\n[9]  Wolfgang Ritter von  Kempelen.  Mechanismus  der menschlichen  Sprache  nebst \nBeschreibungeiner  sprechenden  Maschine.  Mit  einer  Einleitung  vonHerbert  E. \nBrekle  und  Wolfgang  Wild.  Stuttgart-Bad Cannstatt  F.  Frommann, Stuttgart, \n1970. \n\n\f", "award": [], "sourceid": 892, "authors": [{"given_name": "Sidney", "family_name": "Fels", "institution": null}, {"given_name": "Geoffrey", "family_name": "Hinton", "institution": null}]}