{"title": "Using body-anchored priors for identifying actions in single images", "book": "Advances in Neural Information Processing Systems", "page_first": 1072, "page_last": 1080, "abstract": "This paper presents an approach to the visual recognition of human actions using only single images as input. The task is easy for humans but difficult for current approaches to object recognition, because action instances may be similar in terms of body pose, and often require detailed examination of relations between participating objects and body parts in order to be recognized. The proposed approach applies a two-stage interpretation procedure to each training and test image. The first stage produces accurate detection of the relevant body parts of the actor, forming a prior for the local evidence needed to be considered for identifying the action. The second stage extracts features that are \u2018anchored\u2019 to the detected body parts, and uses these features and their feature-to-part relations in order to recognize the action. The body anchored priors we propose apply to a large range of human actions. These priors allow focusing on the relevant regions and relations, thereby significantly simplifying the learning process and increasing recognition performance.", "full_text": "Usingbody-anchoredpriorsforidentifyingactionsin\n\nsingleimages\n\nLeonidKarlinsky\n\nShimonUllman\n\nDepartmentofComputerScience\nWeizmannInstituteofScience\n\nMichaelDinerstein\nRehovot76100,Israel\n\n{leonid.karlinsky, michael.dinerstein, shimon.ullman} @weizmann.ac.il\n\nAbstract\n\nThispaperpresentsanapproachtothevisualrecognitionofhumanactionsusing\nonly single images as input. The task is easy for humans but difficult for current\napproaches to object recognition, because instances of different actions may be\nsimilarintermsofbodypose,andoftenrequiredetailedexaminationofrelations\nbetween participating objects and body parts in order to be recognized. The pro-\nposed approach applies a two-stage interpretation procedure to each training and\ntest image. The first stage produces accurate detection of the relevant body parts\nof the actor, forming a prior for the local evidence needed to be considered for\nidentifyingtheaction. Thesecondstageextractsfeaturesthatareanchoredtothe\ndetected body parts, and uses these features and their feature-to-part relations in\norder to recognize the action. The body anchored priors we propose apply to a\nlargerangeofhumanactions. Thesepriorsallowfocusingontherelevantregions\nandrelations,therebysignificantlysimplifyingthelearningprocessandincreasing\nrecognitionperformance.\n\nIntroduction\n\n1\nThis paper deals with the problem of recognizing transitive actions in single images. A transitive\naction is often described by a transitive verb and involves a number of components, or thematic\nroles [1], including an actor, a tool, and in some cases a recipient of the action. Simple examples\naredrinkingfromaglass,talkingonthephone,eatingfromaplatewithaspoon,orbrushingteeth.\nTransitive actions are characterized visually by the posture of the actor, the tool she/he is holding,\nthe type of grasping, and the presence of the action recipient. In many cases, such actions can be\nreadily identified by human observers from only a single image (see figure 1a). We will consider\nbelow the problem of static action recognition (SAR for short) from a single image, without using\nmotioninformationthatisexploitedbyapproachesdealingwithdynamicactionrecognitioninvideo\nsequences, such as [2]. The problem is of interest first, because in a short observation interval, the\nuse of motion information for identifying an action (e.g.\ntalking on the phone) may be limited.\nSecond, as a natural human capacity, it is of interest for both cognitive and brain studies. Several\nstudies [3, 4, 5, 6] have shown evidence for the presence of SAR related mechanisms in both the\nventral and dorsal areas of the visual cortex, and computational modeling of SAR may shed new\nlight on these mechanisms. Unlike the more common task of detecting individual objects such as\nfaces and cars, SAR depends on detecting object configurations. Different actions may involve\nthe same type of objects (eg. person, phone) but appearing in different configurations (answering,\ndialing), sometimes differing in subtle details, making their identification difficult compared with\nindividualobjectrecognition.\nOnly a few approaches to date have dealt with the SAR problem. [7] studied the recognition of\nsports actions using the pose of the actor.\n[8] used scene interpretation in terms of objects and\n\n1\n\n\f(a)\n\n(b)\n\ndrinking\ndrinking\ndrinking\ndrinking no cup\ndrinking no cup\ndrinking no cup\neating with spoon\neating with spoon\neating with spoon\nphone talking\nphone talking\nphone talking\nphone with bottle\nphone with bottle\nphone with bottle\nscratching\nscratching\nscratching\nsinging with mike\nsinging with mike\nsinging with mike\nsmoking\nsmoking\nsmoking\nteeth brushing\nteeth brushing\nteeth brushing\ntoasting\ntoasting\ntoasting\nwaving\nwaving\nwaving\nwearing glasses\nwearing glasses\nwearing glasses\n\nFigure 1: (a) Examples of similar transitive actions identifiable by humans from single images\n(brushing teeth, talking on a cell phone and wearing glasses). (b) Illustration of a single run of the\nproposedtwo-stageapproach. Inthefirststagethepartsaredetectedintheface\u2192hand\u2192elbowor-\nder. Inthesecondstageweapplybothactionlearningandactionrecognitionusingtheconfiguration\nofthedetectedpartsandthefeaturesanchoredtothehandregion; thebargraphontherightshows\nrelativelog-posteriorestimatesforthedifferentactions.\ntheir relative configuration to distinguish between different sporting events such as badminton and\nsailing. [9] recognized static intransitive actions such as walking and jumping based on a human\nbodyposerepresentedbyavariantoftheHOGdescriptor. [10]discriminatedbetweenplayingand\nnot playing musical instruments using a star-like model. The most detailed static schemes to date\n[11, 12] recognized static transitive sports actions, such as the tennis forehand and the volleyball\nsmash. [11]usedafullbodymask,bagoffeaturesfordescribingscenecontext,andthedetectionof\nthe objects relevant for the action, such as bats, balls, etc., while [12] learned joint models of body\nposeandobjectsspecifictoeachaction. [11]usedGrabCut[13]toextractthebodymask,andboth\n[11]and[12]performedfullysupervisedtrainingfortheaprioriknownrelevantobjectsandscenes.\nInthispaperweconsiderthetaskofdifferentiatingbetweensimilartypesoftransitiveactions,such\nas smoking a cigarette, drinking from a cup, eating from a cup with a spoon, talking on the phone,\netc., given only a single image as input. The similarity between the body poses in such actions\ncreates a difficulty for approaches that rely on pose analysis [7, 9, 11]. The relevant differences\nbetween similar actions in terms of the actor body configuration can be at a fine level of detail.\nTherefore,onecannotrelyonafixedpre-determinednumberofconfigurationtypes[7,9,11];rather,\noneneedstobeabletomakeasfinediscriminationsasrequiredbythetask. Objectsparticipatingin\ndifferent actions may be very small, occupying only a few pixels in a low resolution image (brush,\nphone,Fig. 1a). Inaddition,theseobjectsmaybeunknownapriori,suchasinthenaturalcasewhen\nthe learning is weakly supervised, i.e. we know only the action label of the training images, while\ntheparticipatingobjectsarenotannotatedandcannotbeindependentlylearnedasin[8,11]. Finally,\nthe background scene, used by [8, 11] to recognize sports actions and events, is uninformative for\nmanytransitiveactionsofinterest,andcannotbedirectlyutilized.\nSince SAR is a version of an object recognition problem, a natural question to ask is whether it\ncan be solved by directly applying state-of-the-art techniques of object recognition. As shown in\nthe results section 3, the problem is significantly more difficult for current methods compared with\nmorestandardobjectrecognitionapplications. Theproposedmethodidentifiesmoreaccuratelythe\nfeaturesandgeometricrelationshipsthatcontributetocorrectrecognitioninthisdomain,leadingto\nbetter recognition. It is further shown that integrating standard object recognition approaches into\ntheproposedframeworksignificantlyimprovestheirresultsintheSARdomain.\nThe main contribution of this paper is an approach, employing the so-called body anchored strat-\negyexplainedbelow,forrecognizinganddistinguishingbetweensimilartransitiveactionsinsingle\nimages. Inboththelearningandthetestsettings,theapproachappliesatwo-stageinterpretationto\neach (training or test) image. The first stage produces accurate detection and localization of body\nparts, and the second then extracts and uses features from locations anchored to body parts. In the\nimplementationofthefirststage,thefaceisdetectedfirst,anditsdetectionisextendedtoaccurately\nlocalize the elbow and the hand of the actor. In the second stage, the relative part locations and the\nhand region are analyzed for action related learning and recognition. During training, this allows\ntheautomaticdiscoveryandconstructionofimplicitnon-parametricmodelsfordifferentimportant\naspects of the actions, such as accurate relative part locations, relevant objects, types of grasping,\nand part-object configurations. During testing, this allows the approach to focus on image regions,\n\n2\n\n\f(a)\n\nO\nFH\n\nOo\n,\nfhn\nHE\n\nF\n\nm\n\n,1\n\n,\n\nnk\n\nA\n\no\nhen\nmB\nm\n\nf\n\nmn\n\n(b)\n\nFigure2: (a)Examplesofthecomputedbinarymasks(cyan)forsearchingforelbowlocationgiven\nthe detected hand and face marked by a red-green star and magenta rectangle respectively. The\nyellow square marks the detected elbow; (b) Graphical representation (in plate notation) of the\nproposedprobabilisticmodelforactionrecognition(seesection2.2fordetails).\nfeatures and relations that contain all of the relevant information for recognizing the action. As a\nresult,weeliminatetheneedtohaveapriorimodelsfortheobjectsrelevantfortheactionthatwere\nusedin[11,12]. Focusinginabody-anchoredmannerontherelevantinformationnotonlyincreases\nefficiency,butalsoconsiderablyenhancesrecognitionresults. Theapproachisillustratedinfig. 1b.\nThe rest of the paper is organized as follows. Section 2 describes the proposed approach and its\nimplementation details. Section 3 describes the experimental validation. Summary and discussion\nareprovidedinsection4.\n2 Method\nAsoutlinedabove,theapproachproceedsintwomainstages. Thefirststageisbodyinterpretation,\nwhichisbyitselfasequentialprocess. First, thepersonisdetectedbydetectingher/hisface. Next,\nthe face detection is extended to detect the hands and elbows of the person. This is achieved in a\nnon-parametric manner by following chains of features connecting the face to the part of interest\n(hand,elbow),byanextensionof[14]. Inthesecondstage,featuresgatheredfromthehandregion\nand the relative locations of the hand, face and elbow, are used to model and recognize the static\nactionofinterest. Thefirststageoftheprocess,dealingwiththeface,handandelbowdetection,is\ndescribedinsection2.1. Thestaticactionmodelingandrecognitionisdescribedinsection2.2and\nadditionalimplementationdetailsareprovidedinsection2.3.\n2.1 Bodypartsdetection\nBody parts detection in static images is a challenging problem, which has recently been addressed\nby several studies [14, 15, 16, 17, 18]. The most difficult parts to detect are the most flexible parts\nof the body - the lower arms and the hands. This is due to large pose and appearance variability\nand the small size typical to these parts. In our approach, we have adopted an extension of the\nnon-parametric method for the detection of parts of deformable objects recently proposed by [14].\nThis method can operate in two modes. The first mode is used for the independent detection of\nsufficiently large and rigid objects and object parts, such as the face. The second mode allows\npropagating from some of the parts, which are independently detected, to additional parts, which\nare more difficult to detect independently, such as hands and elbows. The method extends the so-\ncalled star model by allowing features to vote for the detection target either directly, or indirectly,\nvia features participating in feature-chains going towards the target. In the independent detection\nmode, these feature chains may start anywhere in the image, whereas in the propagation mode\nthese chains must originate from already detected parts. The method employs a non-parametric\ngenerative probabilistic model, which can efficiently learn to detect any part from a collection of\ntraining sequences with marked target (e.g., hand) and source (e.g., face) parts (or only the target\nparts in the independent detection mode). The details of this model are described in [14]. In our\napproach,thefaceisdetectedintheindependentdetectionmodeof[14],andthehandandtheelbow\naredetectedbychains-propagationfromthefacedetection(treatedasthesourcepart). Themethod\nistrainedusingacollectionofshortvideosequences,eachhavingtheface,thehandandtheelbow\nmarkedbythreepoints. Thecodeforthemethodof[14]wasextendedtoallowrestricteddetection\nofdependentparts,suchashandandelbow. Insomecases,theelbowismoredifficulttodetectthan\nthe hand, as it has less structure. For each (training or test) image In, we therefore constrain the\nelbow detection by a binary mask of possible elbow locations gathered from training images with\n\n3\n\n\fn \u2212 xf\n\nn , 1\n\nH = of h\n\nn , ohe\n\nE = ohe\n\nn \u2261 xe\n\nn \u2212 xh\n\nn,andtheelbowby xe\n\nn,andthepatchfeatures {F m = f m\n\nthe sufficiently similar hand-face offset (within 0.25 face width) to the one detected on In. Figure\n2a shows some examples of the detected faces, hands and elbows together with the elbow masks\nderivedfromthedetectedface-handoffset.\n2.2 Modelingandrecognitionofstaticactions\nGivenanimage In (trainingortest),wefirstintroducethefollowingnotation(lowerindexrefersto\nthe image, upper indices to parts). Denote the instance of the action contained in In by an (known\nfortrainingandunknownfortestimages). Denotethedetectedlocationsofthefaceby xf\nn,thehand\nby xh\nn. Alsodenotethewidthofthedetectedfaceby sn. Throughoutthepaper,\nwewillexpressallsizeanddistanceparametersin sn units,inordertoeliminatethedependenceon\nthescaleofthepersonperformingtheaction. Formanytransitiveactionsmostofthediscriminating\ninformation about the action resides in regions around specific body parts [19]. Here we focus on\nhand regions for hand-related actions, but for other actions their respective parts and part regions\ncanbelearnedandused. Werepresenttheinformationfromthehandregionbyasetofrectangular\npatchfeaturesextractedfromthisregion. Allfeaturesaretakenfromacircularregionwitharadius\n0.75 \u2219 sn aroundthehandlocation xh\nn. Fromthisregionweextract sn \u00d7 sn pixelrectangularpatch\nfeaturescenteredatallCannyedgepointssub-sampledwitha 0.2 \u2219 sn pixelgrid. Denotethesetof\npatch features extracted from image In bynf m\nn is\nn(cid:1)io, where SIF T m\nn =hSIF T m\nsn(cid:0)xm\nn \u2212 xf\nthe SIFT descriptor [20] of the m-th feature, xm\nn is its image location, 1\nn(cid:1) is the offset\nsn(cid:0)xm\n(in sn units) between the feature and the face, and square brackets denote a row vector. The index\nm enumerates the features in arbitrary order for each image. Denote by kn the number of patch\nfeaturesextractedfromimage In.\nThe probabilistic generative model explaining all the gathered data is defined as follows. The ob-\nserved variables of the model are: the face-hand offset OF\nn, the hand-elbow\noffset OH\nn }. Theunobservedvariablesofthe\nmodelaretheactionlabelvariable A, andthesetofbinaryvariables {Bm}, oneforeachextracted\npatch feature. The meaning of Bm = 1 is that the m-th patch feature was generated by the action\nA, while the meaning of Bm = 0 is that the m-th patch feature was generated independently of\nA. Throughout the paper we will use a shorthand form of variable assignments, e.g., P(cid:0)of h\nn (cid:1)\ninstead of P(cid:0)OF\nn (cid:1). We define the joint distribution of the model that generates\nthedataforimage In as:\nP (Bm) \u2219 P(cid:0)f m\nn , Bm(cid:1)(1)\nP(cid:0)A,{Bm} , of h\nHere P (A)isaprioractiondistribution,whichwetaketobeuniform,and:\nn (cid:1)\n(2)\nn (cid:1)\nThe P (Bm) = \u03b1, is the prior probability for the m-th feature to be generated from the action,\nand we assume it maintains the following relation: P (Bm = 1) = \u03b1 (cid:28) (1 \u2212 \u03b1) = P (Bm = 0)\nreflectingthefactthatmostpatchfeaturesarenotrelatedtotheaction. Figure2bshowsthegraphical\nrepresentationoftheproposedmodel.\nAs shown in the Appendix A, in order to find the action label assignment to A that maximizes the\nposterioroftheproposedprobabilisticgenerativemodel,itissufficienttocompute:\nn (cid:1)\nn (cid:17)\n\n(3)\nAs can be seen from eq. 3, and as shown in the Appendix A, the inference is independent of the\nexact value of \u03b1 (as long as \u03b1 (cid:28) (1 \u2212 \u03b1)). In section 2.3 we explain how to empirically estimate\ntheprobabilities P(cid:0)f m\n\nknYm=1\nn }(cid:1) = P (A) \u2219 P(cid:0)of h\nn (cid:1) \u2219\nn (cid:12)(cid:12)A, of h\nn , Bm(cid:1) =(cid:26) P(cid:0)f m\nn (cid:12)(cid:12)of h\nP(cid:0)f m\n\nn (cid:12)(cid:12)A, of h\nP(cid:0)f m\n\nP(cid:0)f m\nP(cid:16)f m\n\nn , ohe\nn , ohe\n\nn , A, of h\nn , of h\n\nknXm=1\nn (cid:1)thatarenecessarytocompute3.\n\nn (cid:1)and P(cid:0)f m\n\nn }(cid:1) = arg max\n\nlog P(cid:0)A(cid:12)(cid:12)of h\n\nn (cid:12)(cid:12)A, of h\n\nn \u2261 xh\n\nn \u2212 xf\n\nH = of h\n\nn , OH\n\nE = ohe\n\nif Bm = 1\notherwise\n\nn , ohe\n\nn ,{f m\n\nn , A, of h\n\nn , ohe\n\nn , of h\n\nn , ohe\n\nn , ohe\nn , ohe\n\narg max\n\nA\n\nn , ohe\n\nn ,{f m\n\nA\n\nn , ohe\n\nn , ohe\n\nn , ohe\n\n4\n\n\f1: drinking\n1: drinking\n\n2: drinking no cup\n2: drinking no cup\n\n1\n11\n\n3: eating with spoon\n3: eating with spoon\n\n5: phone talking with bottle\n5: phone talking with bottle\n\n9: teeth brushing\n9: teeth brushing\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n5\n55\n\n9\n99\n\n6: scratching head\n6: scratching head\n\n10: toasting\n10: toasting\n\n2\n22\n\n6\n66\n\n10\n1010\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n7: singing with mike\n7: singing with mike\n\n11: waving\n11: waving\n\n4: phone talking\n4: phone talking\n\n8: smoking\n8: smoking\n\n12: wearing glasses\n12: wearing glasses\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n1\n1\n2\n2\n3\n3\n4\n4\n5\n5\n6\n6\n7\n7\n8\n8\n9\n9\n10\n10\n11\n11\n12\n12\n\n3\n33\n\n7\n77\n\n11\n1111\n\n4\n44\n\n8\n88\n\n12\n1212\n\nn , A = a, of h\n\nn , ohe\n\nFigure 3: Examples of similar static transitive action recognition on our 12-actions / 10-people\n(\u201812/10\u2019) dataset. On all examples, the detected face, hand and elbow are shown by cyan circle,\nred-green star and yellow square, respectively. At the right hand side of each image, the bar graph\nshows the estimated log-posterior of the action variable A. Each example shows a zoomed-in ROI\noftheaction. Additionalexamplesareprovidedinsupplementarymaterial.\n2.3 Modelprobabilities\nThemodelprobabilitiesareestimatedfromthetrainingdatausingKernelDensityEstimation(KDE)\n[21]. Assumewearegivenasetofsamples {Y1, . . . , YR}fromsomedistributionofinterest. Givena\nnewsample Y fromthesamedistribution,asymmetricGaussianKDEestimate P (Y )fortheproba-\nbilityof Y canbeapproximatedas: P (Y ) \u2248 1\nR \u2219PYr\u2208N N (Y ) exp(cid:16)\u22120.5 \u2219 kY \u2212 Yrk2.\u03c32(cid:17)where\nN N (Y ) is the set of nearest neighbors of Y within the given set of samples. When the number of\nsamples Rislarge,brute-forcesearchforthe N N (Y )setbecomesinfeasible. Therefore,weuseAp-\nproximateNearestNeighbor(ANN)search(usingtheimplementationof[22])tocomputetheKDE.\nTo compute P(cid:0)f m\nn(cid:1)i\nn =hSIF T m\nsn(cid:0)xm\nn \u2212 xh\nin test image In, we search for the nearest neighbors of the row vectorhf m\nn i in a\nohe\nt in training images It, s.t. at = aousinganANN\nsetofrowvectors: nhf r\nquery. Recallthat sn wasdefinedasthewidthofthedetectedfaceinimage In,andhence 1\nsn isthe\nscalefactorthatweusefortheoffsetsinthequery. Thequeryreturnsasetof K nearestneighbors,\nand the Gaussian KDE with \u03c3 = 0.2, is applied to this set to compute the estimated probability\nn (cid:1). In our experiments we found that it is sufficient to use K = 25. The\nP(cid:0)f m\nP(cid:0)f m\n3 Results\nTotestourapproach,wehaveappliedittotwostatictransitiveactionrecognitiondatasets. Thefirst\ndataset, denoted \u201812/10\u2019 dataset, was created by us and contained 12 similar transitive actions per-\nformedby10differentpeople,appearingagainstdifferentnaturalbackgrounds. Theseconddataset\nwas compiled by [11] for dynamic transitive action recognition. It contains 9 different people per-\nforming 6 general transitive actions. Although originally designed and used by Gupta et al in [11]\nfordynamicactionrecognition,wetransformeditintoastaticactionrecognitiondatasetbyassign-\ningactionlabelstoframesactuallycontainingtheactionsandtreatingeachsuchframeasaseparate\nstatic instance of the action. Since successive frames are not independent, the experiments con-\nducted on both datasets were all performed in a person-leave-one-out manner, meaning that during\nthe training we completely excluded all the frames of the tested person. Section 3.1 provides more\ndetails on the relevant parts (face, hand, and elbow) detection in our experiments complementing\nsection 2.1. Sections 3.2 and 3.3 describe the \u201812/10\u2019 and the Gupta et al datasets in more detail\n\nn (cid:1) for the m-th patch feature f m\nt i(cid:12)(cid:12)(cid:12) all f r\nn (cid:1)iscomputedas: P(cid:0)f m\n\nn (cid:1) =Pa P(cid:0)f m\n\nn , A = a, of h\nn , of h\n\nn , ohe\n\nn , 1\nn , 1\nof h\nsn\n\nn , A = a, of h\n\nn , ohe\n\nn (cid:1).\n\nn , of h\n\nn , ohe\n\nn , ohe\n\nn , 1\nsn\n\nt , 1\nst\n\nof h\nt\n\n, 1\nst\n\nohe\n\n5\n\n\f(a)\n(a)\n\n \n \n\n(b)\n(b)\n\n1\n1\n1\n11\n\n4\n4\n\n5\n5\n\n6\n6\n\n7\n7\n\n8\n8\n\n2\n2\n2\n22\n\n4\n4\n44\n4\n\n12\n12\n12\n1212\n\n2\n2\n\n3\n3\n\n1\n1\n1\n1\n2\n2\n2\n2\n3\n3\n3\n3\n4\n4\n4\n4\n5\n5\n5\n5\n6\n6\n6\n6\n7\n7\n7\n7\n8\n8\n8\n8\n9\n9\n9\n9\n10\n10\n10\n10\n11\n11\n11\n11\n12\n12\n12\n12\n\n1\n1\n1\n1\n2\n2\n2\n2\n3\n3\n3\n3\n4\n4\n4\n4\n5\n5\n5\n5\n6\n6\n6\n6\n7\n7\n7\n7\n8\n8\n8\n8\n9\n9\n9\n9\n10\n10\n10\n10\n11\n11\n11\n11\n12\n12\n12\n12\n\n \n \n\n1\n1\n\n1\n1\n1\n1\n2\n2\n2\n2\n3\n3\n3\n3\n4\n4\n4\n4\n5\n5\n5\n5\n6\n6\n6\n6\n7\n7\n7\n7\n8\n8\n8\n8\n9\n9\n9\n9\n10\n10\n10\n10\n11\n11\n11\n11\n12\n12\n12\n12\n\n1\n1\n1\n1\n2\n2\n2\n2\n3\n3\n3\n3\n4\n4\n4\n4\n5\n5\n5\n5\n6\n6\n6\n6\n7\n7\n7\n7\n8\n8\n8\n8\n9\n9\n9\n9\n10\n10\n10\n10\n11\n11\n11\n11\n12\n12\n12\n12\n\n9 10 11 12 13\n9 10 11 12 13\n\n1: drinking\n1: drinking\n2: drinking no cup\n2: drinking no cup\n3: eating with spoon\n3: eating with spoon\n4: phone\n4: phone\n5: phone with bottle\n5: phone with bottle\n6: scratching\n6: scratching\n7: singing with mike\n7: singing with mike\n8: smoking\n8: smoking\n9: teeth brushing\n9: teeth brushing\n10: toasting\n10: toasting\n11: waving\n11: waving\n12: wearing glasses\n12: wearing glasses\n13: no action\n13: no action\n\n0.7\n0.7\n0.6\n0.6\n0.5\n0.5\n0.4\n0.4\n0.3\n0.3\n0.2\n0.2\n0.1\n0.1\nFigure 4: (a) Average static action confusion matrix obtained by leave-one-out cross validation of\nthe proposed method on the \u201812/10\u2019 dataset; (b) Some interesting failures (red boxes), on the right\nof each failure there is a successfully recognized instance of an action with which the method has\nconfused. Themeaningofthebar-graphisasinfigure3. Additionalfailureexamplesareprovided\ninthesupplementarymaterial.\ntogether with the respective static action recognition experiments performed on them. All exper-\niments were performed on grayscale versions of the images. Figures 3 and 6a illustrate the two\ntesteddatasetsalsoshowingexamplesofsuccessfullyrecognizedstatictransitiveactions,andfigure\n4bshowssomeinterestingfailures.\n3.1 Partdetectiondetails\nOurapproachisbasedonpriorpartdetectionanditsperformanceisboundedfromabovebythepart\ndetectionperformance. Thedetectionratesofstate-of-the-artmethodsforlocalizingbodypartsina\ngeneralsettingarecurrentlyasignificantlimitingfactor. Forexample,[14]thatweusehere,obtains\nan average of 66% correct hand detection (comparing favorably to other state-of-the-art methods)\nin the general setting experiments, when both the person and the background are unseen during\npart detector training. However, as shown in [14], average 85% part detection performance can be\nachievedinmorerestrictedsettings. Onesuchsetting(denotedself-trained)iswhenanindependent\nshort part detection training period of several seconds is allowed for each test person, as for e.g.\nin the human-computer interaction applications. Another setting (denoted environment-trained) is\nwhentheenvironmentinwhichpeopleperformtheactionisfixed,e.g. inapplicationswherewecan\ntrainpartdetectorsonsomepeople,andthenapplythemtonewunseenpeople,butappearinginthe\nsameenvironment. Asdemonstratedinthemethodscomparisonexperimentinsection3.2,itappears\nthatpartdetectionisanessentialcomponentofsolvingSAR.Currentperformanceinautomaticbody\nparts detection is well below human performance, but the area is now a focus of active research\nwhich is likely to reduce this current performance gap. In our experiments we adopted the more\nconstrained (but still useful) part detection settings described above, the self-trained for the 12-10\ndataset (having each person in different environment) and the environment-trained for the Gupta et\nal. dataset(havingallthepeopleinthesamegeneralenvironment).\nIn the 12-10 dataset experiments, the part detection models for the face, hand and elbow described\nin section 2.1, were trained using 10 additional short movies, one for each person, in which the\nactors randomly moved their hands. On these 10 movies, face, hand and elbow locations were\nmanuallymarked. Thelearnedmodelswerethenappliedtodetecttheirrespectivepartsonthe120\nmovie sequences of our dataset. The hand detection performance was 94.1% (manually evaluated\non a subset of frames). Qualitative examples are provided in the supplementary material. The part\ndetection for the Gupta et al. dataset was performed in a person-leave-one-out manner. For each\nperson the parts (face, hand) were detected using models trained on other people. The mean hand\ndetection performance was 88% (manually evaluated). Since most people in the dataset wear very\ndark clothing, in many cases the elbow is invisible and therefore it was not used in this experiment\n(itisstraightforwardtoremoveitfromthemodelbyassigningafixedvaluetothehand-elbowoffset\ninbothtrainingandtest).\n3.2 The\u201812/10\u2019datasetexperiments\nThe \u201812/10\u2019 dataset consists of 120 videos of 10 people performing 12 similar transitive actions,\nnamelydrinking,talkingonthephone,scratching,toasting,waving,brushingteeth,smoking,wear-\ning glasses, eating with a spoon, singing to a microphone, and also drinking without a cup and\n\n6\n\n\f1\n\n0.5\n\n0.5\n\ndrinking\n\nscratching\n\n(a)\n1eating with spoon\n1drinking without cup\n1\n0.5\n0.5\n0.5\n0\n0\n0\n0\n1\n1\n0\n1singing with mike\n1phone talking with bottle\n1\n0.5\n0.5\n0.5\n0\n0\n0\n0\n1 teeth brushing\n1\n1\n0.5\n0.5\n0.5\n0\n0\n0\n\n0.5\ntoasting\n\n0.5\nwaving\n\n0.5\n\n0.5\n\n1\n\n0\n\n0\n\n0\n\n0\n\n0.5\n\n0\n\n0\n\n0.5\n\n0.5\n\n1\n\n1\n\n1\n\n1\n\n1\n\n0\n\n0.5\nsmoking\n\n1 phone talking\n0.5\n0\n1\n0.5\n0\n0\n1 wearing glasses\n0.5\n0\n\n0.5\n\n0.5\n\n0\n\n1\n\n1\n\n1\n\n(b)\nFull person \nbounding box (no \nanchoring)\n24.6 \u00b112.5%\n8.8 \u00b11%\n16.6 \u00b16.6%\n\nHand \nanchored\nregion\n58.5 \u00b19.1%\n37.7 \u00b12.7%\n45.7 \u00b19.1%\n\nMethod / \nexperiment\nSAR method \n(section 2.2)\n\n[2 ]\n\nBoWSVM\n\nFigure5: (a)ROCbasedcomparisonwiththestate-of-the-artmethodofobjectdetection[23]applied\ntorecognizestaticactions. Foreachaction,thebluelineistheaverageROCof[23],andthemagenta\nline is the average ROC of the proposed method. (b) Comparing state-of-the-art object recognition\nmethodsontheSARtaskwithandwithout\u2018bodyanchoring\u2019.\nmakingaphonecallwithabottle. Allpeopleexceptonewerefilmedagainstdifferentbackgrounds.\nAll backgrounds were natural indoor / outdoor scenes containing both clutter and people passing\nby. The drinking and toasting actions were performed with 2-4 different tools, and phone talking\nwas performed with mobile and regular phones. Overall, the dataset contains 44,522 frames. Not\nall frames contain actions of interest (e.g. in drinking there are frames where the person reaches to\n/ puts down a cup). The ground-truth action labels were manually assigned to the relevant frames.\nEachoftheresulting23,277relevantactionframeswasconsideredaseparateinstanceofanaction.\nTheremainingframeswerelabeled\u2018no-action\u2019. Theaveragerecognitionaccuracywas 59.3\u00b18.6for\nthe13actions(includingno-action)and 58.5 \u00b1 9.1%forthe12mainactions(excludingno-action).\nFigure4ashowstheconfusionmatrixfor13actionsaveragedoverthe10testpeople.\nAsmentionedintheintroduction,oneoftheimportantquestionswewishtoansweristheneedfor\nthe detection of the fine details of the person, such as the accurate hand, and elbow locations, in\norder to recognize the action. To test this issue, we have compared the results of three approaches:\ndeformable parts model [23], Bag-of-Words (BoW) SVM ([24]), and our approach described in\nsection 2.2, in two settings. In the first setting the methods were trained to distinguish between the\nactions based on a bounding box of the entire person (i.e. without focusing on the fine details such\nas provided by the hand and elbow detection). In the second, body anchored setting, the methods\nwere applied to the hand anchored regions (small regions around the detected hand as described in\nsection2.2). Themethodof[23]isoneofthestate-of-the-artobjectrecognitionschemes,achieving\ntop scores on recent PASCAL-VOC competitions [25], and BoW SVM is a popular method in\nthe literature also obtaining state-of-the art results for some datasets. Figure 5b shows the results\nobtainedbythethreemethodsinthetwosettings. Figure5aprovidesROC-basedcomparisonofthe\nresultsofourfullapproachwiththeonesobtainedby[23]. Theobtainedresultsstronglysuggestthat\nbodyanchoringisapowerfulpriorforthetaskofdistinguishingbetweensimilartransitiveactions.\n3.3 Guptaetaldatasetexperiments\nThisdatasetwascompiledby[11]. Itconsistsof46moviesequencesof9differentpeopleperform-\ning 6 distinct transitive actions: drinking, spraying, answering the phone, making a call, pouring\nfrom a pitcher and lighting a flashlight. In each movie, we manually assigned action labels to all\nthe frames actually containing the action, labeling the remainder of the frames \u2018no-action\u2019. Since\nthe distinction between \u2018making a call\u2019 and \u2018answering phone\u2019 was in the presence or absence of\nthe \u2018dialing\u2019 action in the respective video, we re-labeled the frames of these actions into \u2018phone\ntalking\u2019 and \u2018dialing\u2019. The action recognition performance was measured using the person-leave-\none-out cross-validation, in the same manner as for our dataset. The average accuracy over the 7\nstaticactions(includingno-action)was 82 \u00b1 11.5%,andwas 86 \u00b1 14.4%excludingno-action. The\naverage 7-action confusion matrix is shown in figure 6b. The presented results are for the static\naction recognition, and hence are not directly comparable with the results obtained on this dataset\nfor the dynamic action recognition by [11], who obtained 93.34% recognition (out of the 46 video\n\n7\n\n\f1\n1\n1\n2\n2\n2\n3\n3\n3\n4\n4\n4\n5\n5\n5\n6\n6\n6\n\n1\n1\n1\n2\n2\n2\n3\n3\n3\n4\n4\n4\n5\n5\n5\n6\n6\n6\n\n1\n1\n1\n2\n2\n2\n3\n3\n3\n4\n4\n4\n5\n5\n5\n6\n6\n6\n\n1\n1\n1\n2\n2\n2\n3\n3\n3\n4\n4\n4\n5\n5\n5\n6\n6\n6\n\n \n \n\n1\n1\n\n(a)\n(a)\n\n(b)\n(b)\n\n \n \n\n5\n5\n55\n\n6\n6\n66\n\n00.10.20.30.40.50.60.70.80.9\n00.10.20.30.40.50.60.70.80.9\n\n1\n1\n1\n2\n2\n2\n3\n3\n3\n4\n4\n4\n5\n5\n5\n6\n6\n6\n\n1\n1\n1\n2\n2\n2\n3\n3\n3\n4\n4\n4\n5\n5\n5\n6\n6\n6\n\n2\n2\n22\n\n3\n3\n33\n\n1\n1\n11\n\n4\n4\n44\n\n4\n4\n\n5\n5\n\n6\n6\n\n7\n7\n\n2\n2\n\n3\n3\n\n1: dialing\n1: dialing\n2: drinking\n2: drinking\n3: flashlight\n3: flashlight\n4: phone talking\n4: phone talking\n5: pouring\n5: pouring\n6: spraying\n6: spraying\n7: no action\n7: no action\nFigure6: (a)somesuccessfullyidentifiedactionexamplesfromthedatasetof[11]; (b)meanstatic\nactionconfusionmatrixforleave-one-outcrossvalidationexperimentsontheGuptaetal. dataset.\nsequencesandnotframes)usingboththetemporalinformation(parttracks,etc.) andapriorimodels\nfortheparticipatingobjects(cup,pitcher,flashlight,spraybottleandphone).\n4 Discussion\nWe have presented a method for recognizing transitive actions from single images. This task is\nperformed naturally and efficiently by humans, but performance by current recognition methods is\nseverely limited. The proposed method can successfully handle both similar transitive actions (the\n\u201812/10\u2019 dataset), and general transitive actions (the Gupta et al dataset). The method uses priors\nthatfocusonbodypartanchoredfeaturesandrelations. Ithasbeenshownthatmostcommonverbs\nare associated with specific body parts [19]; the actions considered here were all hand-related in\nthis sense. The detection of hands and elbows therefore provided useful priors in terms of regions\nand properties likely to contribute to the SAR task in this setting. The proposed approach can be\ngeneralizedtodealwithotheractionsbydetectingallthebodypartsassociatedwithcommonverbs,\nautomaticallydetectingtherelevantpartsforeachspecificactionduringtraining,andfinallyapply-\ning the body anchored SAR model described in section 2.2. The comparisons show that without\nusing the body anchored priors there is a highly significant drop in SAR performance even when\nemploying state-of-the-art methods for object recognition. The main reasons for this drop are the\nfine details and local nature of the relevant evidence for distinguishing between actions, the huge\nnumberofpossiblelocations,anddetailedfeaturesthatneedtobesearchedifbody-anchoredpriors\narenotused. Directionsforfuturestudiesthereforeincludeamorecompleteandaccuratebodyparts\ndetectionandtheiruseinprovidingusefulpriorsforstaticactionrecognitionandinterpretation.\nA Log-posteriorderivation\nHere we derive the equivalent form of log-posterior (eq. 3) of the proposed probabilistic action\nrecognition model defined in eq. 1. In 4, the symbol \u223c means equivalent in terms of maximizing\noverthevaluesoftheactionvariable A.\nlog P(cid:0)A(cid:12)(cid:12)of h\nn }(cid:1) \u223c logP{Bm}\nP(cid:0)A,{Bm} , of h\nn ,{f m\nloghP (A) \u2219 P(cid:0)of h\nm=1hP1\nn (cid:12)(cid:12)A, of h\nBm=0 P (Bm) \u2219 P(cid:0)f m\nn (cid:1) \u2219Qkn\nn , ohe\nn , Bm(cid:1)i =\nm=1 loghP1\nn (cid:12)(cid:12)A, of h\nBm=0 P (Bm) \u2219 P(cid:0)f m\nPkn\nn (cid:12)(cid:12)A, of h\nn (cid:12)(cid:12)of h\nm=1 log(cid:2)\u03b1 \u2219 P(cid:0)f m\nn (cid:1)(cid:3) =\nn (cid:1) + (1 \u2212 \u03b1) \u2219 P(cid:0)f m\nPkn\nn )(cid:21) +Pkn\nm=1 log(cid:20)1 +\nn (cid:12)(cid:12)of h\nm=1 log(cid:2)(1 \u2212 \u03b1) \u2219 P(cid:0)f m\nPkn\nn )(cid:21) \u2248(\u2217)Pkn\nm=1 log(cid:20)1 +\nn ) \u223cPkn\nPkn\nPkn\n(4)\nIn eq. 4, \u03b3 = \u03b1/ (1 \u2212 \u03b1), the termPkn\nn (cid:1)(cid:3) is independent of the\nn (cid:12)(cid:12)of h\nm=1 log(cid:2)(1 \u2212 \u03b1) \u2219 P(cid:0)f m\naction(constantforagivenimage In)andthuscanbedropped,and (\u2217)followsfrom log (1 + \u03b5) \u2248 \u03b5\nfor \u03b5 (cid:28) 1andfrom \u03b3 beinglargeduetoourassumptionthat \u03b1 (cid:28) (1 \u2212 \u03b1).\n\nn }(cid:1) =\nn , Bm(cid:1)ii \u223c\nn (cid:1)(cid:3) \u223c\n\nn |A,of h\nP (f m\nn |of h\n\u03b3\u2219P (f m\nn |A,of h\nP (f m\nn |of h\n\u03b3\u2219P (f m\nn ,ohe\nn )\nn )\nn ,ohe\n\nP (f m\n\u03b3\u2219P (f m\n\nn |A,of h\nn |of h\n\nn ,ohe\nn ,ohe\nn ,ohe\nn ,ohe\n\nn ,ohe\nn ,ohe\n\nn )\nn ) \u223c\n\nP (f m\nP (f m\n\nn |A,of h\nn |of h\n\nP (f m\nP (f m\n\nn ,A,of h\nn ,of h\n\nn , ohe\n\nn ,{f m\n\nn )\n\nn ,ohe\nn ,ohe\n\nn , ohe\n\nn , ohe\n\nn , ohe\n\nn , ohe\n\nn )\n\nm=1\n\nm=1\n\nn , ohe\n\nn )\n\nm=1\n\nn , ohe\n\nn , ohe\n\n8\n\n\fmovies. In: CVPR.(2008)1\u20138\nAnnalsofNeurology(2007)\n\nReferences\n[1] Jackendoff,R.: Semanticinterpretationingenerativegrammar. TheMITPress(1972)\n[2] Laptev, I., Marszalek, M., Schmid, C., Rozenfeld, B.: Learning realistic human actions from\n[3] Iacoboni,M.,Mazziotta,J.C.: Mirrorneuronsystem: basicfindingsandclinicalapplications.\n[4] Kim,J.,Biederman,I.: Wheredoobjectsbecomescenes? JournalofVision(2009)\n[5] Helbig,H.,Graf,M.,Kiefer,M.: Theroleofactionrepresentationsinvisualobjectrecognition.\nExperimentalBrainResearch(2006)\n[6] Sakata, H., Taira, M., Kusunoki, M., Murata, A., Tanaka, Y., Tsutsui, K.: Neural coding of\n3dfeaturesofobjectsforhandactionintheparietalcortexofthemonkey. PhilosTransRSoc\nLondBBiolSci.(1998)\n[7] Wang, Y., Jiang, H., Drew, M.S., nian Li, Z., Mori, G.: Unsupervised discovery of action\nclasses. In: CVPR.(2006) 5\n[8] Li,L.,Fei-Fei,L.: What,whereandwho? classifyingeventsbysceneandobjectrecognition.\nIn: ICCV.(2007)1\u20138\n[9] Thurau, C., Hlavac, V.: Pose primitive based human action recognition in videos or still\nimages. In: CVPR.(2008)1\u20138\n[10] Yao, B., Fei-Fei, L.: Grouplet: A structured image representation for recognizing human and\nobjectinteractions. CVPR(2010)\n[11] Gupta, A., Kembhavi, A., Davis, L.: Observing human-object interactions: Using spatial and\nfunctionalcompatibilityforrecognition. PAMI(2009)\n[12] Yao, B., Fei-Fei, L.: Modeling mutual context of object and human pose in human-object\ninteractionactivities. CVPR(2010)\n[13] Blake,A.,Rother,C.,Brown,M.,Perez,P.,Torr,P.: Interactiveimagesegmentationusingan\nadaptivegmmrfmodel. ECCV(2004)\n[14] Karlinsky,L.,Dinerstein,M.,Harari,D.,Ullman,S.: Thechainsmodelfordetectingpartsby\ntheircontext. CVPR(2010)\n[15] Ferrari, V., Marin, M., Zisserman, A.: Progressive search space reduction for human pose\nestimation. CVPR(2008)\n[16] Andriluka, M., Roth, S., Schiele, B.: Pictorial structures revisited: People detection and\narticulatedposeestimation. CVPR(2009)\n[17] Felzenszwalb,P.,Huttenlocher,D.: Pictorialstructuresforobjectrecognition. IJCV61(2005)\n55\u201379\n[18] Ramanan, D., Forsyth, D.A., Barnard, K.: Building models of animals from video. PAMI\n(2006)\n[19] Maouene, J., Hidaka, S., Smith, L.B.: Bodypartsandearly-learnedverbs. CognitiveScience\n(2008)\n[20] Lowe,D.: Distinctiveimagefeaturesfromscale-invariantkeypoints. IJCV(2004)\n[21] Duda,R.,Hart,P.: Patternclassificationandsceneanalysis. Wiley(1973)\n[22] Mount, D., Arya, S.: Ann: A library for approximate nearest neighbor searching. CGC 2nd\nAnnualWorkshoponComp.Geometry(1997)\n[23] Felzenszwalb, P., McAllester, D., Ramanan, D.: A discriminatively trained, multiscale, de-\nformablepartmodel. CVPR(2008)1\u20138\n[24] Zhang,J.,Marszalek,M.,Lazebnik,S.,Schmid,C.: Localfeaturesandkernelsforclassifica-\ntionoftextureandobjectcategories: Acomprehensivestudy. IJCV(2007)\n[25] Everingham, M., Van Gool, L., Williams, C., Winn, J., Zisserman, A.: The pascal visual ob-\nject classes challenge 2007 results. http://pascallin.ecs.soton.ac.uk/challenges/VOC/voc2007\n(2007)\n\n9\n\n\f", "award": [], "sourceid": 441, "authors": [{"given_name": "Leonid", "family_name": "Karlinsky", "institution": null}, {"given_name": "Michael", "family_name": "Dinerstein", "institution": null}, {"given_name": "Shimon", "family_name": "Ullman", "institution": null}]}