{"title": "Retrieved context and the discovery of semantic structure", "book": "Advances in Neural Information Processing Systems", "page_first": 1193, "page_last": 1200, "abstract": null, "full_text": "Retrieved context and the discovery of semantic\n\nstructure\n\nVinayak A. Rao, Marc W. Howard\u2217\n\nSyracuse University\n\nDepartment of Psychology\n\n430 Huntington Hall\nSyracuse, NY 13244\n\nvrao@gatsby.ucl.ac.uk, marc@memory.syr.edu\n\nAbstract\n\nSemantic memory refers to our knowledge of facts and relationships between con-\ncepts. A successful semantic memory depends on inferring relationships between\nitems that are not explicitly taught. Recent mathematical modeling of episodic\nmemory argues that episodic recall relies on retrieval of a gradually-changing rep-\nresentation of temporal context. We show that retrieved context enables the de-\nvelopment of a global memory space that re\ufb02ects relationships between all items\nthat have been previously learned. When newly-learned information is integrated\ninto this structure, it is placed in some relationship to all other items, even if that\nrelationship has not been explicitly learned. We demonstrate this effect for global\nsemantic structures shaped topologically as a ring, and as a two-dimensional sheet.\nWe also examined the utility of this learning algorithm for learning a more realistic\nsemantic space by training it on a large pool of synonym pairs. Retrieved context\nenabled the model to \u201cinfer\u201d relationships between synonym pairs that had not yet\nbeen presented.\n\n1 Introduction\n\nSemantic memory refers to our ability to learn and retrieve facts and relationships about concepts\nwithout reference to a speci\ufb01c learning episode. For example, when answering a question such as\n\u201cwhat is the capital of France?\u201d it is not necessary to remember details about the event when this fact\nwas \ufb01rst learned in order to correctly retrieve this information. An appropriate semantic memory for\na set of stimuli as complex as, say, words in the English language, requires learning the relationships\nbetween tens of thousands of stimuli. Moreover, the relationships between these items may describe\na network of non-trivial topology [16]. Given that we can only simultaneously perceive a very small\nnumber of these stimuli, in order to be able to place all stimuli in the proper relation to each other\nthe combinatorics of the problem require us to be able to generalize beyond explicit instruction. Put\nanother way, semantic memory needs to not only be able to retrieve information in the absence of\na memory for the details of the learning event, but also retrieve information for which there is no\nlearning event at all.\nComputational models for automatic extraction of semantic content from naturally-occurring text,\nsuch as latent semantic analysis [12], and probabilistic topic models [1, 7], exploit the temporal\nco-occurrence structure of naturally-occurring text to estimate a semantic representation of words.\nTheir success relies to some degree on their ability to not only learn relationships between words\nthat occur in the same context, but also to infer relationships between words that occur in similar\n\u2217Vinayak Rao is now at the Gatsby Computational Neuroscience Unit, University College London.\n\nhttp://memory.syr.edu.\n\n1\n\n\fcontexts. However, these models operate on an entire corpus of text, such that they do not describe\nthe process of learning per se.\nHere we show that the temporal context model (TCM), developed as a quantitative model of human\nperformance in episodic memory tasks, can provide an on-line learning algorithm that learns appro-\npriate semantic relationships from incomplete information. The capacity for this model of episodic\nmemory to also construct semantic knowledge spaces of multiple distinct topologies, suggests a\nrelatively subtle relationship between episodic and semantic memory.\n\n2 The temporal context model\n\nEpisodic memory is de\ufb01ned as the vivid conscious recollection of information from a speci\ufb01c in-\nstance from one\u2019s life [18]. Many authors describe episodic memory as the result of the recovery\nof some type of a contextual representation that is distinct from the items themselves. If a cue item\ncan recover this \u201cpointer\u201d to an episode, this enables recovery of other items that were bound to the\ncontextual representation without committing to lasting interitem connections between items whose\noccurrence may not be reliably correlated [17].\nLaboratory episodic memory tasks can provide an important clue to the nature of the contextual\nrepresentation that could underlie episodic memory. For instance, in the free recall task, subjects\nare presented with a series of words to be remembered and then instructed to recall all the words\nthey can remember in any order they come to mind. If episodic recall of an item is a consequence\nof recovering a state of context, then the transitions between recalls may tell us something about\nthe ability of a particular state of context to cue recall of other items. Episodic memory tasks show\na contiguity effect\u2014a tendency to make transitions to items presented close together in time, but\nnot simultaneously, with the just-recalled word. The contiguity effect shows an apparently universal\nform across multiple episodic recall tasks, with a characteristic asymmetry favoring forward recall\ntransitions [11] (see Figure 1a).\nThe temporal contiguity effect observed in episodic recall can be simply reconciled with the hypoth-\nesis that episodic recall is the result of recovery of a contextual representation if one assumes that\nthe contextual representation changes gradually over time. The temporal context model (TCM) de-\nscribes a set of rules for a gradually-changing representation of temporal context and how items can\nbe bound to and recover states of temporal context. TCM has been applied to a number of problems\nin episodic recall [9]. Here we describe the model, incorporating several changes that enable TCM\nto describe the learning of stable semantic relationships (detailed in Section 3).1\nTCM builds on distributed memory models which have been developed to provide detailed descrip-\ntions of performance in human memory tasks [14]. In TCM, a gradually-changing state of temporal\ncontext mediates associations between items and is responsible for recency effects and contiguity\neffects. The state of the temporal context vector at time step i is denoted as ti and changes from\nmoment-to-moment according to\n\n(1)\nwhere \u03b2 is a free parameter, tIN\nis the input caused by the item presented at time step i, assumed to\nbe of unit length, and \u03c1i is chosen to ensure that ti is of unit length. Items, represented as unchanging\northonormal vectors f, are encoded in their study contexts by means of a simple outer-product matrix\nconnecting the t layer to the f layer, MT F , which is updated according to:\n\nti = \u03c1iti\u22121 + \u03b2tIN\n\n,\n\ni\n\ni\n\ni = fit0\n\ni\u22121,\n\n\u2206MT F\n\n(2)\nwhere the prime denotes the transpose and the subscripts here re\ufb02ect time steps. Items are probed\nfor recall by multiplying MT F from the right with the current state of t as a cue. This means that\nwhen tj is presented as a cue, each item is activated to the extent that the probe context overlaps\nwith its encoding contexts.\nThe space over which t evolves is obviously determined by the tIN s. We will decompose tIN into\ncIN , a component that does not change over the course of study of this paper, and hIN , a component\n\n1Previous published treatments of TCM have focused on episodic tasks in which items were presented only\nonce. Although the model described here differs from previously published versions in notation and its behavior\nover multiple item repetitions, it is identical to previously-published results described for single presentations\nof items.\n\n2\n\n\fa\n\nb\n\nc\n\nFigure 1: Temporal recovery in episodic memory. a. Temporal contiguity effect in episodic recall.\nGiven that an item from a series has just been recalled, the y-axis gives the probability that the next\nitem recalled came from each serial position relative the just-recalled item. This \ufb01gure is averaged\nacross a dozen separate studies [11]. b. Visualization of the model. Temporal context vectors\nti are hypothesized to reside in extra-hippocampal MTL regions. When an item fi is presented,\nit evokes two inputs to t\u2014a slowly-changing direct cortical input cIN\nand a more rapidly varying\nhippocampal input hIN\n. When an item is repeated, the hippocampal component retrieves the context\nin which the item was presented. c. While the cortical component serves as a temporally-asymmetric\ncue when an item is repeated, the hippocampal component provides a symmetric cue. Combining\nthese in the right proportion enables TCM to describe temporal contiguity effects.\n\ni\n\ni\n\nthat changes rapidly to retrieve the contexts in which an item was presented. Denoting the time steps\nat which a particular item A was presented as Ai, we have\n\nAi+1 \u221d \u03b3 \u02c6hIN\ntIN\n\nAi+1 + (1 \u2212 \u03b3) cIN\nA .\n\n(3)\nwhere the proportionality re\ufb02ects the fact that tIN is always normalized before being used to update\nti as in Eq. 1 and the hat on the hIN term refers to the normalization of hIN . We assume that the\ncIN s corresponding to the items presented in any particular experiment start and remain orthonormal\nto each other. In contrast, hIN starts as zero for each item and then changes according to:\n\n+ tAi\u22121.\n\nAi+1 = hIN\nhIN\nAi\n\n(4)\nIt has been hypothesized that ti re\ufb02ects the pattern of activity at extra-hippocampal medial temporal\nlobe (MTL) regions, in particular the entorhinal cortex [8]. The notation cIN and hIN re\ufb02ects\nthe hypothesis that the consistent and rapidly-changing parts of tIN re\ufb02ect inputs to the entorhinal\ncortex from cortical and hippocampal sources, respectively (Figure 1b).\nAccording to TCM, associations between items are not formed directly, but rather are mediated by\nthe effect that items have on the state of context which is then used to probe for recall of other items.\nWhen an item is repeated as a probe, this induces a correlation between the tIN of the probe context\nand the study context of items that were neighbors of the probe item when it was initially presented.\nThe consistent part of tIN is an effective cue for items that followed the initial presentation of\nthe probe item (open symbols, Figure 1c). In contrast, recovery of the state of context that was\npresent before the probe item was initially presented is a symmetric cue (\ufb01lled symbols, Figure 1).\nCombining these two components in the proper proportions provides an excellent description of\ncontiguity effects in episodic memory [8].\n\n3 Constructing global semantic information from local events\n\nIn each of the following simulations, we specify a to-be-learned semantic structure by imagining\nitems as the nodes of a graph with some topology. We generated training sequences by randomly\nsampling edges from the graph.2 Each edge only contains a limited amount of information about\n2The pairs are chosen randomly, so that any across-pair learning would be uninformative with respect to\nthe overall structure of the graph. To further ensure that learning across pairs from simple contiguity could not\ncontribute to our results, we set \u03b2 in Eq. 1 to one when the \ufb01rst member of each pair was presented. This means\nthat the temporal context when the second item is presented is effectively isolated from the previous pair.\n\n3\n\ncfthiih-5-4-3-2-1012345Lag00.20.40.60.81Cue strengthHippocampalCortical\fa\n\nb\n\nc\n\nd\n\nFigure 2: Learning of a one-dimensional structure using contextual retrieval. a. The graph used to\ngenerate the training pairs. b-c. Associative strength between items after training (higher strength\ncorresponds to darker cells). b. The model without contextual retrieval (\u03b3 = 0). c. The model with\ncontextual retrieval (\u03b3 > 0). d. Two dimensional MDS solution for the log of the data in c. Lines\nconnect points corresponding to nodes connected by an edge.\n\nB includes the tIN\n\nthe global structure. For the model is to learn the global structure of the graph, it must somehow\nintegrate the learning events into a coherent whole.\nAfter training we evaluated the ability of the model to capture the topology of the graph by ex-\namining the cue strength between each item. The cue strength from item A to B is de\ufb01ned as\nf0\nA . This re\ufb02ects the overlap between the cIN and hIN components of A and the contexts\nBMT F tIN\nin which B was presented.3\nBecause tIN\nis caused by presentation of item i, we can think of the tIN s as a representation of the\nset of items. Learning can be thought of as a mixing of the tIN s according to the temporal structure\nof experience. Because the cIN s are \ufb01xed, changes in the representation are solely due to changes\nin the hIN s. Suppose that two items, A and B are presented in sequence. If context is retrieved,\nthen after presentation of the pair A-B hIN\nA that obtained when A was presented.\nThis includes the current state of hIN\nA . If at some later time B is now\npresented as part of the sequence B-C , then because tIN\nA , item C is learned in a\ncontext that resembles tIN\nA , despite the fact that A and C were not actually presented close together\nin time. After learning A-B and B-C , tIN\nC will resemble each other. This ability to\nrate as similar items that were not presented together in the same context, but that were presented in\nsimilar contexts, is a key property of latent models of semantic learning [12].\nTo isolate the importance of retrieved context for the ability to extract global structure, we will\ncompare a version of the model with \u03b3 = 0 to one with \u03b3 > 0.4 With \u03b3 = 0, the model functions\nas a simple co-occurrence detector in that the cue strength between A and B is non-zero only if cIN\nA\nwas part of the study contexts of B. In the absence of contextual retrieval, this requires that B was\npreceded by A during study.\nUltimately, the tis and hIN\ns can be expressed as a combination of the cIN vectors. We therefore\ntreated these as orthonormal basis vectors in the simulations that follow. MT F and the hIN s were\ninitialized as a matrix and vectors of zeros, respectively. The parameter \u03b2 for the second member of\na pair was \ufb01xed at 0.6.\n\nA as well as the \ufb01xed state cIN\n\nB is similar to tIN\n\nA and tIN\n\ni\n\ni\n\n3.1 1-D: Rings\n\nFor this simulation we sampled edges chosen from a ring of ten items (Fig. 2a). We treated the ring\nas an undirected graph, in that we sampled an edge A-B equally often as B-A . We presented the\nmodel with 300 pairs chosen randomly from the ring. For example, the training pairs might include\nthe sub-sequence C-D , A-B , F-E , B-C .\n\n3In this implementation of TCM, hIN\n\nAMT F . This need not be the case in general, as one\ncould alter the learning rate, or even the structure of Eqs. 2 and/or 4 without changing the basic idea of the\nmodel.\n\n4In the simulations reported below, this value is set to 0.6. The precise value does not affect the qualitative\n\nA is identical to f0\n\nresults we report as long as it is not too close to one.\n\n4\n\nABCDGHEFIJABCDEFGHIJABCDEFGHIJABCDEFGHIJABCDEFGHIJ\u203a2\u203a1.5\u203a1\u203a0.500.511.52Dimension 1\u203a1.4\u203a1.2\u203a1\u203a0.8\u203a0.6\u203a0.4\u203a0.200.20.40.60.811.21.4Dimension 2\fa\n\nb\n\nc\n\nFigure 3: Reconstruction of a 2-dimensional spatial representation. a. The graph used to construct\nsequences. b. 2-dimensional MDS solution constructed from the temporal co-occurrence version\nof TCM \u03b3 = 0 using the log of the associative strength as the metric. Lines connect stimuli from\nadjacent edges. c. Same as b, but for TCM with retrieved context. The model accurately places the\nitems in the correct topology.\n\nFigure 2b shows the cue strength between each pair of items as a grey-scale image after training the\nmodel without contextual retrieval (\u03b3 = 0). The diagonal is shaded re\ufb02ecting the fact that an item\u2019s\ncue strength to itself is high. In addition, one row on either side of the diagonal is shaded. This\nre\ufb02ects the non-zero cue strength between items that were presented as part of the same training\npair. That is, the model without contextual retrieval has correctly learned the relationships described\nby the edges of the graph. However, without contextual retrieval the model has learned nothing\nabout the relationships between the items that were not presented as part of the same pair (e.g.\nthe cue strength between A and C is zero). Figure 2c shows the cue strength between each pair\nof items for the model with contextual retrieval \u03b3 > 0. The effect of contextual retrieval is that\npairs that were not presented together have non-zero cue strength and this cue strength falls off with\nthe number of edges separating the items in the graph. This happens because contextual retrieval\nenables similarity to \u201cspread\u201d across the edges of the graph, reaching an equilibrium that re\ufb02ects\nthe global structure. Figure 2d shows a two-dimensional MDS (multi-dimensional scaling) solution\nconducted on the log of the cue strengths of the model with contextual retrieval. The model appears\nto have successfully captured the topology of the graph that generated the pairs. More precisely,\nwith contextual retrieval, TCM can place the items in a space that captures the topology of the graph\nused to generate the training pairs.\nOn the one hand, the relationships that result from contextual retrieval in this simulation seem in-\ntuitive and satisfying. Viewed from another perspective, however, this could be seen as undesirable\nbehavior. Suppose that the training pairs accurately sample the entire set of relationships that are\nactually relevant. Moreover, suppose that one\u2019s task were simply to remember the pairs, or alterna-\ntively, to predict the next item that would be presented after presenting the \ufb01rst member of a pair.\nUnder these circumstances, the co-occurrence model performs better than the model equipped with\ncontextual retrieval.\nIt should be noted that people form associations across pairs (e.g. A-C ) after learning lists of paired\nassociates with a linked temporal structure like the rings shown in Figure 2a [15]. In addition, rats\ncan also generalize across pairs, but this ability depends on an intact hippocampus [2]. These \ufb01nding\nsuggest that the mechanism of contextual retrieval capture an important property of how we learn in\nsimilar circumstance.\n\n3.2 2-D: Spatial navigation\n\nThe ring illustrated in Figure 2 demonstrates the basic idea behind contextual retrieval\u2019s ability to\nextract semantic spaces, but it is hard to imagine an application where such a simple space would\nneed to be extracted. In this simulation will illustrate the ability of retrieved context to discover\nrelationships between stimuli arranged in a two-dimensional sheet. The use of a two-dimensional\nsheet has an analog in spatial navigation.\nIt has long been argued that the medial temporal lobe has a special role in our ability to store and\nretrieve information from a spatial map. Eichenbaum [5] has argued that the MTL\u2019s role in spatial\n\n5\n\nABFCGDHLEJKIMNOPQRSTUVWXY\u203a6\u203a5\u203a4\u203a3\u203a2\u203a10123456Dimension 1\u203a6\u203a5\u203a4\u203a3\u203a2\u203a10123456Dimension 2\u203a5\u203a4\u203a3\u203a2\u203a1012345Dimension 1\u203a5\u203a4\u203a3\u203a2\u203a1012345Dimension 2\fnavigation is merely a special case of more general role in organizing disjointed experiences into\nintegrated representations. The present model can be seen as a computational mechanism that could\nimplement this idea.\nIn our typical experience, spatial information is highly correlated with temporal information. Be-\ncause of our tendency to move in continuous paths through our environment, locations that are close\ntogether in space also tend to be experienced close together in time. However, insofar as we travel\nin more-or-less straight paths, the combinatorics of the problem place a premium on the ability to\nintegrate landmarks experienced on different paths into a coherent whole. At the outset we should\nemphasize that our extremely simple simulation here does not capture many of the aspects of actual\nspatial navigation\u2014the model is not provided with metric spatial information, nor gradually chang-\ning item inputs, nor do we discuss how the model could select an appropriate trajectory to reach a\ngoal [3].\nWe constructed a graph arranged as a 5\u00d75 grid with horizontal and vertical edges (Figure 3a). We\npresented the model with 600 edges from the graph in a randomly-selected order. One may think\nof the items as landmarks in a city with a rectangular street plan. The \u201ctraveler\u201d takes trips of one\nblock at a time (perhaps teleporting out of the city between journeys).5 The problem here is not\nonly to integrate pairs into rows and columns as in the 1-dimensional case, but to place the rows and\ncolumns into the correct relationship to each other.\nFigure 3b shows the two-dimensional MDS solution calculated on the log of the cue strengths for the\nco-occurrence model. Without contextual retrieval the model places the items in a high-dimensional\nstructure that re\ufb02ects their co-occurrence. Figure 3c shows the same calculation for TCM with\ncontextual retrieval. Contextual retrieval enables the model to place the items on a two-dimensional\nsheet that preserves the topology of the graph used to generate the pairs. It is not a map\u2014there is\nno sense of North nor an accurate metric between the points\u2014but it is a semantic representation\nthat captures something intuitive about the organization that generated the pairs. This illustrates the\nability of contextual retrieval to organize isolated experiences, or episodes, into a coherent whole\nbased on the temporal structure of experience.\n\n3.3 More realistic example: Synonyms\n\nThe preceding simulations showed that retrieved context enables learning of simple topologies with\na few items. It is possible that the utility of the model in discovering semantic relationships is limited\nto these toy examples. Perhaps it does not scale up well to spaces with large numbers of stimuli, or\nperhaps it will be fooled by more realistic and complex topologies.\nIn this subsection we demonstrate that retrieved context can provide bene\ufb01ts in learning relationships\namong a large number of items with a more realistic semantic structure. We assembled a large list of\nEnglish words (all unique strings in the TASA corpus) and used these as probes to generate a list of\nnearly 114,000 synonym pairs using WordNet. We selected 200 of these synonym pairs at random\nas a test list. The word pairs organize into a large number of connected graphs of varying sizes. The\nlargest of these contained slightly more than 26,000 words; there were approximately 3,500 clusters\nwith only two words. About 2/3 of the pairs re\ufb02ect edges within the \ufb01ve largest clusters of words.\nWe tested performance by comparing the cue strength of the cue word with its synonym to the\nassociative strength to three lures that were synonyms of other cue words\u2014if the correct answer had\nthe highest cue strength, it was counted as correct.6 We averaged performance over ten shuf\ufb02es of\nthe training pairs. We preserved the order of the synonym pairs, so that this, unlike the previous two\nsimulations, described a directed graph.\nFigure 4a shows performance on the training list as a function of learning. The lower curve shows\n\u201cco-occurrence\u201d TCM without contextual retrieval, \u03b3 = 0. The upper curve shows TCM with con-\ntextual retrieval, \u03b3 > 0. In the absence of contextual retrieval, the model learns linearly, performing\nperfectly on pairs that have been explicitly presented. However, contextual retrieval enables faster\nlearning of the pairs, presumably due to the fact that it can \u201cinfer\u201d relationships between words\n\n5We also observed the same results when we presented the model with complete rows and columns of the\n\n6In instances where the cue strength was zero for all the choices, as at the beginning of training, this was\n\nsheet as a training set rather than simply pairs.\n\ncounted as 1/4 of a correct answer.\n\n6\n\n\fa\n\nb\n\nFigure 4: Retrieved context aids in learning synonyms that have not been presented. a. Performance\non the synonym test. The curve labeled \u201cTCM\u201d denotes the performance of TCM with contextual\nretrieval. The curve labeled \u201cCo-occurrence\u201d is the performance of TCM without contextual re-\ntrieval. b. Same as a, except that the training pairs were shuf\ufb02ed to omit any of the test pairs from\nthe middle region of the training sequence.\n\nthat were never presented together. To con\ufb01rm that this property holds, we constructed shuf\ufb02es of\nthe training pairs such that the test synonyms were not presented for an extended period (see Fig-\nure 4b). During this period, the model without contextual retrieval does not improve its performance\non the test pairs because they are not presented. In contrast, TCM with contextual retrieval shows\nconsiderable improvement during that interval.7\n\n4 Discussion\n\nWe showed that retrieval of temporal context, an on-line learning method developed for quantita-\ntively describing episodic recall data, can also integrate distinct learning events into a coherent and\nintuitive semantic representation. It would be incorrect to describe this representation as a semantic\nspace\u2014the cue strength between items is in general asymmetric (Figure 1c). The model thus has\nthe potential to capture some effects of word order and asymmetry. However, one can also think of\nthe set of tIN s corresponding to the items as a semantic representation that is also a proper space.\nExisting models of semantic memory, such as LSA and LDA, differ from TCM in that they are off-\nline learning algorithms. More speci\ufb01cally, these algorithms form semantic associations between\nwords by batch-processing large collections of natural text (e.g., the TASA corpus). While it would\nbe interesting to compare results generated by running TCM on such a corpus with these models,\nconstraints of syntax and style complicate this task. Unlike the simple examples employed here, tem-\nporal proximity is not a perfect indicator of local similarity in real world text. The BEAGLE model\n[10] describes the semantic representation of a word as a superposition of the words that occurred\nwith it in the same sentence. This enables BEAGLE to describe semantic relations beyond simple\ncooccurrence, but precludes the development of a representation that captures continuously-varying\nrepresentations (e.g., Fig. 3). It may be possible to overcome this limitation of a straightforward\napplication of TCM to naturally-occurring text by generating a predictive representation, as in the\nsyntagmatic-paradigmatic model [4].\nThe present results suggest that retrieved temporal context\u2014previously hypothesized to be essen-\ntial for episodic memory\u2014could also be important in developing coherent semantic representations.\nThis could re\ufb02ect similar computational mechanisms contribute to separate systems, or it could in-\ndicate a deep connection between episodic and semantic memory. A key \ufb01nding is that adult-onset\namnesics with impaired episodic memory retain the ability to express previously-learned semantic\nknowledge but are impaired at learning new semantic knowledge [19]. Previous connectionist mod-\nels have argued that the hippocampus contributes to classical conditioning by learning compressed\nrepresentations of stimuli, and that these representations are eventually transferred to entorhinal cor-\n\n7To ensure that this property wasn\u2019t simply a consequence of backward associations for the model with\nretrieved context, we re-ran the simulations presenting the pairs simultaneously rather than in sequence (so that\nthe co-occurrence model would also learn backward associations) and obtained the same results.\n\n7\n\n020406080100Number of pairs presented (1k)00.20.40.60.81P(correct)TCMCo-occurrence020406080100Number of pairs presented (1k)00.20.40.60.81P(correct)TCMCo-occurrence\ftex [6]. This could be implemented in the context of the current model by allowing slow plasticity\nto change the cIN s over long time scales [13].\n\nAcknowledgments\n\nSupported by NIH award MH069938-01. Thanks to Mark Steyvers, Tom Landauer, Simon Dennis,\nand Shimon Edelman for constructive criticism of the ideas described here at various stages of\ndevelopment. Thanks to Hongliang Gai and Aditya Datey for software development and Jennifer\nProvyn for reading an earlier version of this paper.\n\nReferences\n[1] D. Blei, A. Ng, and M. Jordan. Latent Dirichlet allocation. Journal of Machine Learning\n\nResearch, 3:993\u20131022, 2003.\n\n[2] M. Bunsey and H. B. Eichenbaum. Conservation of hippocampal memory function in rats and\n\nhumans. Nature, 379(6562):255\u2013257, 1996.\n\n[3] P. Byrne, S. Becker, and N. Burgess. Remembering the past and imagining the future: a neural\n\nmodel of spatial memory and imagery. Psychological Review, 114(2):340\u201375, 2007.\n\n[4] S. Dennis. A memory-based theory of verbal cognition. Cognitive Science, 29:145\u2013193, 2005.\n[5] H. Eichenbaum. The hippocampus and declarative memory: cognitive mechanisms and neural\n\ncodes. Behavioural Brain Research, 127(1-2):199\u2013207, 2001.\n\n[6] M. A. Gluck, C. E. Myers, and M. Meeter. Cortico-hippocampal interaction and adaptive\nstimulus representation: A neurocomputational theory of associative learning and memory.\nNeural Networks, 18:1265\u20131279, 2005.\n\n[7] T. L. Grif\ufb01ths, M. Steyvers, and J. B. Tenenbaum. Topics in semantic representation. Psycho-\n\nlogical Review, 114(2):211\u201344, 2007.\n\n[8] M. W. Howard, M. S. Fotedar, A. V. Datey, and M. E. Hasselmo. The temporal context model\nin spatial navigation and relational learning: Toward a common explanation of medial temporal\nlobe function across domains. Psychological Review, 112(1):75\u2013116, 2005.\n\n[9] M. W. Howard and M. J. Kahana. A distributed representation of temporal context. Journal of\n\nMathematical Psychology, 46(3):269\u2013299, 2002.\n\n[10] M. N. Jones and D. J. K. Mewhort. Representing word meaning and order information com-\n\nposite holographic lexicon. Psychological Review, 114:1\u201332, 2007.\n\n[11] M. J. Kahana, M.W. Howard, and S.M. Polyn. Associative processes in episodic memory. In\nH. L. Roediger, editor, Learning and Memory - A Comprehensive Reference. Elsevier, in press.\n[12] T. K. Landauer and S. T. Dumais. Solution to Plato\u2019s problem : The latent semantic analy-\nsis theory of acquisition, induction, and representation of knowledge. Psychological Review,\n104:211\u2013240, 1997.\n\n[13] J. L. McClelland, B. L. McNaughton, and R. C. O\u2019Reilly. Why there are complementary\nlearning systems in the hippocampus and neocortex: insights from the successes and failures\nof connectionist models of learning and memory. Psychological Review, 102(3):419\u201357, 1995.\n[14] B. B. Murdock. Context and mediators in a theory of distributed associative memory (TO-\n\nDAM2). Psychological Review, 1997:839\u2013862, 1997.\n\n[15] N. J. Slamecka. An analysis of double-function lists. Memory & Cognition, 4:581\u2013585, 1976.\n[16] M. Steyvers and J. Tenenbaum. The large scale structure of semantic networks: statistical\n\nanalyses and a model of semantic growth. Cognitive Science, 29:41\u201378, 2005.\n\n[17] T. J. Teyler and P. DiScenna. The hippocampal memory indexing theory. Behavioral Neuro-\n\nscience, 100(2):147\u201354, 1986.\n\n[18] E. Tulving. Elements of Episodic Memory. Oxford, New York, 1983.\n[19] R. Westmacott and M. Moscovitch. Names and words without meaning: incidental postmorbid\nsemantic learning in a person with extensive bilateral medial temporal damage. Neuropsychol-\nogy, 15(4):586\u201396, 2001.\n\n8\n\n\f", "award": [], "sourceid": 344, "authors": [{"given_name": "Vinayak", "family_name": "Rao", "institution": null}, {"given_name": "Marc", "family_name": "Howard", "institution": null}]}