Proposal for Shape-sequence Descriptor for motion-description

  • 김회율

초록

Shape is a very important feature for human visual perception. Shape descriptors defined in the current MPEG-7 XM/WD are to describe still regions with the functionality of image matching. Shape variation in a video segment often plays an important role in characterizing the content of the segment. For example, finding out a similar swing form of a golfer or baseball player or such swing segments in sports video, or 3-D motion of 3-D objects, to name a few. For video retrieval tasks, motion information becomes crucial. For that reason, in the current set of MPEG-7 visual descriptors, motion descriptors such as Camera Motion, Motion Trajectory, Parametric Motion, Motion Activity are developed for that purpose. Among them, Motion trajectory and Parametric Motion are to describe the global motion of the whole objects, but they do not capture the associated shape information of the objects. As a result, for example, the sequences of walking animal and human cannot be differentiated. They may appear to be the same motion, but obviously are in different motion. Of course, shape, color, texture or other visual features can be incorporated along with the motion. When the current descriptor is to be used for video image retrieval in terms of shape and motion, each frame of segmented video image needs to be described in terms of existing descriptors. However, since those visual features can only be defined on a static image, it may be often the case that the amount of these visual features may vary drastically due to the three-dimensional characteristics of the object, for example, walking human or moving car in a parking lot. The existing spatio-temporal locator can also be used along with a shape descriptor. However, due to the 3-D nature of a 3-D object, for example, in 3-D object animation, the amount of information that needs to be extracted for the retrieval would be prohibitive not to mention the matching complexity. Therefore, there exists a need for a utility that allows to retrieve candidate video clips with a relatively inexpensive amount of features. In this document, we propose a new shape-sequence image descriptor that can describe shape variation in successive frames similar to the one suggested in M6425. The proposed descriptor can describe shapes caused by non-rigid body motion. The functionality of the descriptor is as described in the document, to retrieve a similar video sequence that may be represented by a key frame. In order to describe the motion information of such sequence, each frame of the video sequence is binarized to segment the object in motion and accumulated to a separate image plane called as shape-sequence map. After a proper normalization, a set of Zernike moments is computed from the map to describe the shape-sequence of an object. Experimental results show that the proposed shape-sequence descriptor is able to retrieve shape variation of the object. This proposal is organized as follows: Section 2 describes the procedure for extracting the shape-sequence. In section 3, we show the experimental results of the proposed descriptor and conclude in section 4.

제목
Proposal for Shape-sequence Descriptor for motion-description
저자
김회율
발행일
2001-01-15
학회명
ISO/IEC JTC1/SC29/WG11 (MPEG)
개최지
Italy/Pisa