{"title": "Learning Temporal Dependencies in Connectionist Speech Recognition", "book": "Advances in Neural Information Processing Systems", "page_first": 1051, "page_last": 1058, "abstract": null, "full_text": "Learning Temporal Dependencies in \nConnectionist Speech Recognition \n\nSteve Renals \n\nMike Hocbberg \n\nTony Robinson \n\nCambridge University Engineering Department \n\nCambridge CB2 IPZ, UK \n\n{sjr,mmh,ajr}@eng.cam.ac.uk \n\nAbstract \n\nHybrid connectionistfHMM systems model time both using a Markov \nchain and through properties of a connectionist network. In this paper, \nwe discuss the nature of the time dependence currently employed in our \nsystems using recurrent networks (RNs) and feed-forward multi-layer \nperceptrons (MLPs). In particular, we introduce local recurrences into a \nMLP to produce an enhanced input representation. This is in the form \nof an adaptive gamma filter and incorporates an automatic approach for \nlearning temporal dependencies. We have experimented on a speaker(cid:173)\nindependent phone recognition task using the TIMIT database. Results \nusing the gamma filtered input representation have shown improvement \nover the baseline MLP system. Improvements have also been obtained \nthrough merging the baseline and gamma filter models. \n\n1 \n\nINTRODUCTION \n\nThe most common approach to large-vocabulary, talker-independent speech recognition \nhas been statistical modelling with hidden Markov models (HMMs). The HMM has an \nexplicit model for time specified by the Markov chain parameters. This temporal model \nis governed by the grammar and phonology of the language being modelled. The acoustic \nsignal is modelled as a random process of the Markov chain and adjoining local temporal \ninformation is assumed to be independent. This assumption is certainly not the case and a \ngreat deal of research has addressed the problem of modelling acoustic context. \n\nStandard HMM techniques for handling the context dependencies of the signal have ex-\n\n1051 \n\n\f1052 \n\nRenals, Hochberg, and Robinson \n\nplicitly modelled all the n-tuples of acoustic segments (e.g., context-dependent triphone \nmodels). Typically, these systems employ a great number of parameters and, subsequently, \nrequire massive amounts of training data and/or care in smoothing of the parameters. Where \nthe context of the model is greater than two segments, an additional problem is that it is \nvery likely that contexts found in testing data are never observed in the training data. \n\nRecently, we have developed state-of-the-art continuous speech recognition systems using \nhybrid connectionistlHMM methods (Robinson, 1994; Renals et aI., 1994). These hybrid \nconnectionistlHMM systems model context at two levels (although these levels are not \nnecessarily at distinct scales). As in the traditional HMM, a Markov process is used to \nspecify the duration and lexical constraints on the model. The connectionist framework \nprovides a conditional likelihood estimate of the local (in time) acoustic waveform given \nthe Markov process. Acoustic context is handled by either expanding the network input to \ninclude multiple, adjacent input frames, or using recurrent connections in the network to \nprovide some memory of the previous acoustic inputs. \n\n2 DEPTH AND RESOLUTION \n\nFollowing Principe et al. (1993), we may characterise the time dependence displayed by a \nparticular model in terms of depth and resolution. Loosely speaking, the depth tells us how \nfar back in time a model is able to look l , and the resolution tells us how accurately the past \nto a given depth may be reconstructed. The baseline models that we currently use are very \ndifferent in terms of these characteristics. \n\nMulti-layer Perceptron \n\nThe feed-forward multi-layer perceptron (MLP) does not naturally model time, but simply \nmaps an input to an output. Crude temporal dependence may be imparted into the system by \nusing a delay-lined input (figure 1 a); an extension of this approach is the time-delay neural \nnetwork (TDNN). The MLP may be interpreted as acting as a FIR filter. A delay-lined \ninput representation may be characterised as having low depth (limited by the delay line \nlength) and high resolution (no smoothing). \n\nRecurrent Network \n\nThe recurrent network (RN) models time dependencies of the acoustic signal via a fully(cid:173)\nconnected, recurrent hidden layer (figure 1 b). The RN has a potentially infinite depth \n(although in practice this is limited by available training algorithms) and low resolution, \nand may be regarded as analogous to an IIR filter. A small amount of future context is \navailable to the RN, through a four frame target delay. \n\nExperiments \n\nExperiments on the DARPA Resource Management (RM) database have indicated that \nthe tradeoff between depth and resolution is important. In Robinson et al. (1993), we \ncompared different acoustic front ends using a MLP and a RN. Both networks used 68 \n\n1 In the language of section 3, the depth may be expressed as the mean duration, relative to the \n\ntarget, of the last kernel in a filter that is convolved with the input. \n\n\fLearning Temporal Dependencies in Connectionist Speech Recognition \n\n1053 \n\np(q\" I X~~), Vk = I , .... K \n\nHidden Layer \n\n512 - 1,024 hidden units \n\ny(t-4) \n\nu(t) \n\nx(t) \n\nXn_c \n\n.,. \n\nxn_1 \n\nxn+1 \n\n... \n\nxn+c \n\n(a) Multi-layer Perceptron \n\n(b) Recurrent Network \n\nFigure 1: Connectionist architectures used for speech recognition. \n\noutputs (corresponding to phones); the MLP used 1000 hidden units and the RN used \n256 hidden units. Both architectures were trained using a training set containing 3990 \nsentences spoken by 109 speakers. Two different resolutions were used in the front-end \ncomputation of mel-frequency cepstral coefficients (MFCCs): one with a 20ms Hamming \nwindow and a lOms frame step (referred to as 20110), the other with a 32ms Hamming \nwindow and a 16ms frame step (referred to as 32116). A priori, we expected the higher \nresolution frame rate (20/10) to produce a higher performance recogniser because rapid \nspeech events would be more accurately modelled. While this was the case for the MLP, \nthe RN showed better results using the lower resolution front end (32/16) (see table 1). For \nthe higher resolution front-end, both models require a greater depth (in frames) for the same \ncontext (in milliseconds). In these experiments the network architectures were constant so \nincreasing the resolution of the front end results in a loss of depth. \n\nNet \nRN \nRN \nMLP \nMLP \n\nFront End \n\n20/10 \n32116 \n20/10 \n32/16 \n\nWord Error Rate % \n\nfeb89 \n6.1 \n5.9 \n5.7 \n6.6 \n\noct89 \n7.6 \n6.3 \n7.1 \n7.8 \n\nfeb91 \n7.4 \n6.1 \n7.6 \n8.5 \n\nsep92 \n12.1 \n11.5 \n12.0 \n15.0 \n\nTable 1: Comparison of acoustic front ends using a RN and a MLP for continuous speech \nrecognition on the RM task, using a wordpair grammar of perplexity 60. The four test sets \n(feb89, oct89, feb91 and sep92, labelled according to their date of release by DARPA) each \ncontain 300 sentences spoken by 10 new speakers. \n\nIn the case of the MLP we were able to explicitly set the memory depth. Previous experi(cid:173)\nments had determined that a memory depth of 6 frames (together with a target delayed by 3 \nframes) was adequate for problems relating to this database. In the case of the RN, memory \n\n\f1054 \n\nRenals, Hochberg, and Robinson \n\nP(qlx) \n\nP(qlx) \n\nOutput Layer \n\nHidden Layer (1000 hidden units) \n\nHidden Layer (1000 hidden units) \n\nx(l) \n\nx(I+2) \n\n(a) Gamma Filtered Input \n\n(b) Gamma Filter + Future Context \n\nFigure 2: Gamma memory applied to the network input. The simple gamma memory in (a) \ndoes not incorporate any information about the future, unless the target is delayed. In (b) \nthere is an explicit delay line to incorporate some future context. \n\ndepth is not determined directly, but results from the interaction between the network archi(cid:173)\ntecture (i.e., number of state units) and the training process (in this case, back-propagation \nthrough time). We hypothesise that the RN failed to make use of the higher resolution front \nend because it did not adapt to the required depth. \n\n3 GAMMA MEMORY STRUCTURE \n\nThe tradeoff between depth and resolution has led us to investigate other network archi(cid:173)\ntectures. The gamma filter, introduced by de Vries and Principe (1992) and Principe et al. \n(1993), is a memory structure designed to automatically determine the appropriate depth \nand resolution (figure 2). This locally recurrent architecture enables lowpass and bandpass \nfilters to be learned from data (using back-propagation through time or real-time recurrent \nlearning) with only a few additional parameters. \n\nWe may regard the gamma memory as a generalisation of a delay line (Mozer, 1993) in \nwhich the kth tap at time t is obtained by convolving the input time series with a kernel \nfunction, g~(t), and where 11 parametrises the Kth order gamma filter, \n\ngg(t) = 8(t) \n\nl<k<K. \n\nThis family of kernels is attractive, since it may be computed incrementally by \n\ndXk(t) \n---;tt = -l1xk(t) + I1Xk-l (t) . \n\nThis is in contrast to some other kernels that have been proposed (e.g., Gaussian kernels \nproposed by Bodenhausen and Waibel (1991) in which the convolutions must be performed \n\n\fLearning Temporal Dependencies in Connectionist Speech Recognition \n\n1055 \n\nexplicitly). In the discrete time case the filter becomes: \n\nXk(t) = (l -\n\nIl)Xk(t - 1) + IlXk-l(t - 1) \n\nThis recursive filter is guaranteed to be stable when 0 < J1 < 2. \n\nIn the experiments reported below we have replaced the input delay line of a MLP with a \ngamma memory structure, using one gamma filter for each input feature. This structure is \nreferred to as a \"focused gamma net\" by de Vries and Principe (1992). \n\nOwing to the effects of anticipatory coarticulation, information about the future is as \nimportant as past context in speech recognition. A simple gamma filtered input (figure 2a) \ndoes not include any future context. There are various ways in which this may be remedied; \n\n\u2022 Use the same architecture, but delay the target (similar to figure Ib); \n\u2022 Explicitly specify future context by adding a delay line from the future (figure 2b); \n\u2022 Use two gamma filters per feature: one forward, one backward in time. \n\nA drawback of the first approach is that the central frame corresponding to the delayed target \nwill have been smoothed by the action of the gamma filter. The third approach necessitates \ntwo passes when either training or running the network. \n\n4 SPEECH RECOGNITION EXPERIMENTS \n\nWe have performed experiments using the standard TIMIT speech database. This database \nis divided into 462 training speakers and 168 test speakers. Each speaker utters eight \nsentences that are used in these experiments, giving a training set of 3696 sentences and a \ntest set of 1344 sentences. We have used this database for a continuous phone recognition \ntask: labelling each sentence using a sequence of symbols, drawn from the standard 61 \nelement phone set. \n\nThe acoustic data was preprocessed using a 12th order perceptual linear prediction (PLP) \nanalysis to produce an energy coefficient plus 12 PLP cepstral coefficients for each frame \nof data. A 20ms Hamming window was used with a lOms frame step. The temporal \nderivatives of each of these features was also estimated (using a linear regression over \u00b1 3 \nadjacent frames) giving a total of 26 features per frame. \n\nThe networks we employed (table 2) were MLPs, with 1000 hidden units, 61 output \nunits (one per phone) and a variety of input representations. The Markov process used \nsingle state phone models, a bigram phone grammar, and a Viterbi decoder was used for \nrecognition. The feed-forward weights in each network were initialised with identical sets \nof small random values. The gamma filter coefficients were initialised to 1.0 (equivalent \nto a delay line). The feed-forward weights were trained using back-propagation and the \ngamma filter coefficients were trained in a forward in time back-propagation procedure \nequivalent to real-time recurrent learning. An important detail is that the gradient step size \nwas substantially lower (by a factor of 10) for the gamma filter parameters compared with \nthe feed-forward weights. This was necessary to prevent the gamma filter parameters from \nbecoming unstable. \nThe baseline system using a delay line (Base) corresponds to figure 1 a, with \u00b1 3 frames of \ncontext. The basic four-tap gamma filter G4 is illustrated in figure 2a (but using 1 fewer \n\n\f1056 \n\nRenals, Hochberg, and Robinson \n\nSystem ID Description \n\nBase Baseline delay line, \u00b1 3 frames of context \n\nG4 Gamma filter, 4 taps \nG7 Gamma filter, 7 taps, delayed target \nG7i G7 initialised using weights from Base \n\nG4F3 Gamma filter, 4 taps, 3 frames future context \nG4F3i G4F3 initialised using weights from Base \n\nTable 2: Input representations used in the experiments. Note that G7i and G4F3i were \ninitialised using a partially trained weight matrix (after six epochs) from Base. \n\ntap than the picture) and G7 is a 7 frame gamma filter with the target delayed for 3 frames, \nthus providing some future context (but at the expense of smoothing the \"centre\" frame). \nFuture context is explicitly incorporated in G4F3, in which the three adjacent future frames \nare included (similar to figure 2b). Systems G7i and G4F3i were both initialised using \na partially trained weight matrix for the delay line system, Base. This was equivalent to \nfixing the value of the gamma filter coefficients to a constant (1.0) during the first six epochs \nof training and only adapting the feed-forward weights, before allowing the gamma filter \ncoefficients to adapt. \n\nThe results of using these systems on the TIM IT phone recognition task are given in table \n3. Table 4 contains the results of some model merging experiments, in which the output \nprobability estimates of 2 or more networks were averaged to produce a merged estimate. \n\nSystem ID Depth Correct% \n\nInsert. % Subst.% Delet.% Error % \n\nBase \nG4 \nG7 \nG7i \nG4F3 \nG4F3i \n\n4.0 \n8.5 \n11.7 \n5.8 \n9.6 \n4.9 \n\n67.6 \n65.8 \n65.5 \n67.3 \n67.8 \n68.0 \n\n4.1 \n4.1 \n4.1 \n3.8 \n3.8 \n3.9 \n\n24.7 \n25.9 \n26.0 \n24.5 \n24.2 \n24.2 \n\n7.7 \n8.3 \n8.5 \n8.2 \n8.0 \n7.8 \n\n36.5 \n38.2 \n38.6 \n36.5 \n36.0 \n35.9 \n\nTable 3: TIMIT phone recognition results for the systems defined in table 2. The Depth \nvalue is estimated as the ratio of filter order to average filter parameter KIJ.!. Future context \nis ignored in the estimate of depth, and the estimates for G7 and G7i are adjusted to account \nfor the delayed target. \n\nSystem ID Correct% \n\nG4F3+Base \nG4F3 +G4F3i \nG7 + Base \nG7+G7i \n\n68.1 \n68.2 \n67.0 \n67.4 \n\nInsert.% Subst.% Delet.% Error% \n\n3.2 \n3.2 \n3.2 \n3.6 \n\n23.7 \n23.5 \n24.4 \n24.4 \n\n8.2 \n8.3 \n8.6 \n8.2 \n\n35.1 \n35.0 \n36.2 \n36.2 \n\nTable 4: Model merging on the TIMIT phone recognition task. \n\n\fLearning Temporal Dependencies in Connectionist Speech Recognition \n\n1057 \n\n- PLP Coefficients \n-\n\nDerivatives \n\nI \n\nC2 C3 C4 C5 C6 C7 CB \n\nC9 Cl0 Cll C12 \n\nI \n\nFeature \n\nO.B \n\n0.6 \n\n0.4 \n\n0.2 \n\nE \n\nCl \n\nFigure 3: Gamma filter coefficients for G4F3. The coefficients correspond to energy (E) \nand 12 PLP cepstral coefficients (C1-C12) and their temporal derivatives. \n\n5 DISCUSSION \n\nSeveral comments may be made about the results in section 4. As can be seen in table 3, \nreplacing a delay line with an adaptive gamma filter can lead to an improvement in per(cid:173)\nformance. Knowledge of future context is important. This is shown by G4, which had no \nfuture context or delayed target information, and had poorer performance than the baseline. \nHowever, incorporating future context using a delay line (G4F3) gives better performance \nthan a pure gamma filter representation with a delayed target (G7). Training the locally \nrecurrent gamma filter coefficients is not trivial. Fixing the gamma filter coefficients to \n1.0 (delay line) whilst adapting the feed-forward weights during the first part of training \nis beneficial. This is demonstrated by comparing the performance of G7 with G7i and \nG4F3 with G4F3i. Finally, table 4 shows that model merging generally leads to improved \nrecognition performance relative to the component models. This also indicates that the \ndelay line and gamma filter input representations are somewhat complementary. \n\nFigure 3 displays the trained gamma filter coefficients for G4F3. There are several points \nto make about the learned temporal dependencies. \n\n\u2022 The derivative parameters are smaller compared with the static PLP parameters. \nThis indicates the derivative filters have greater depth and lower resolution com(cid:173)\npared with the static PLP filters. \n\n\u2022 If a gamma filter is regarded as a lowpass IIR filter, then lower filter coefficients \nindicate a greater degree of smoothing. Better estimated coefficients (e.g., static \nPLP coefficients Cl and C2) give rise to gamma filters with less smoothing. \n\n\u2022 The training schedule has a significant effect on filter coefficients. The depth \nestimates of G4F3 and G4F3i in table 3 demonstrate that very different sets of \nfilters were arrived at for the same architecture with identical initial parameters, \nbut with different training schedules. \n\n\f1058 \n\nRenals, Hochberg, and Robinson \n\nWe are investigating the possibility of using gamma filters to model speaker characteristics. \nPreliminary experiments in which the gamma filters of speaker independent networks were \nadapted to a new speaker have indicated that the gamma filter coefficients are speaker \ndependent. This is an attractive approach to speaker adaptation, since very few parameters \n(26 in our case) need be adapted to a new speaker. \n\nGamma filtering is a simple, well-motivated approach to modelling temporal dependencies \nIt adds minimal complexity to the system \nfor speech recognition and other problems. \n(in our case a parameter increase of 0.01 %), and these initial experiments have shown an \nimprovement in phone recognition performance on the TIM IT database. A further increase \nin performance resulted from a model merging process. We note that gamma filtering and \nmodel merging may be regarded as two sides of the same coin: gamma filtering smooths \nthe input acoustic features, while model merging smooths the output probability estimates. \n\nAcknowledgement \n\nThis work was supported by ESPRIT BRA 6487, WERNICKE. SR was supported by a SERC \npostdoctoral fellowship and a travel grant from the NIPS foundation. TR was supported by \na SERC advanced fellowship. \n\nReferences \n\nBodenhausen, D., & Waibel, A. (1991). The Tempo 2 algorithm: Adjusting time delays \nby supervised learning. In Lippmann, R. P., Moody, J. E., & Touretzky, D. S. (Eds.), \nAdvances in Neural Information Processing Systems, Vol. 3, pp. 155-161. Morgan \nKaufmann, San Mateo CA. \n\nde Vries, B., & Principe, J. C. (1992). The gamma model-a new neural model for temporal \n\nprocessing. Neural Networks, 5,565-576. \n\nMozer, M. C. (1993). Neural net architectures for temporal sequence processing. \n\nIn \nWeigend, A. S., & Gershenfeld, N. (Eds.), Predicting the future and understanding \nthe past. Addison-Wesley, Redwood City CA. \n\nPrincipe, J. C., de Vries, B., & de Oliveira, P. G. (1993). The gamma filter-a new class of \nadaptive IIR filters with restricted feedback. IEEE Transactions on Signal Processing, \n41, 649-656. \n\nRenals, S., Morgan, N., Bourlard, H., Cohen, M., & Franco, H. (1994). Connectionist \nprobability estimators in HMM speech recognition. IEEE Transactions on Speech \nand Audio Processing. In press. \n\nRobinson, A. J., Almeida, L., Boite, J.-M., Bourlard, H., Fallside, F., Hochberg, M., Ker(cid:173)\nshaw, D., Kohn, P., Konig, Y., Morgan, N., Neto, J. P., Renals, S., Saerens, M., & \nWooters, C. (1993). A neural network based, speaker independent, large vocabu(cid:173)\nlary, continuous speech recognition system: the WERNICKE project. In Proceedings \nEuropean Conference on Speech Communication and Technology, pp. 1941-1944 \nBerlin. \n\nRobinson, T. (1994). The application of recurrent nets to phone probability estimation. \n\nIEEE Transactions on Neural Networks. In press. \n\n\f", "award": [], "sourceid": 851, "authors": [{"given_name": "Steve", "family_name": "Renals", "institution": null}, {"given_name": "Mike", "family_name": "Hochberg", "institution": null}, {"given_name": "Tony", "family_name": "Robinson", "institution": null}]}