archiving
Pixels and samples, the building blocks of audiovisual data, cannot be 'read' directly, as they lack explicit meaning. An extra step is needed to bridge the semantic gap between form and content, allowing us to describe audiovisual data in words, archive it, and make it searchable. Traditionally, manual documentation by archivists provided this link. Recently, structured data captured at various stages of the production pipeline is also being used. On top of that, we are gradually incorporating automated techniques as helpful tools.
Both manual and automated descriptions are essentially 'representations' of data. They represent a specific perspective: the view of a person with a particular background (such as an editor or documentalist), or an automated algorithm focused on a specific modality (speech, image) or elements, like sounds, words, speakers, or emotions in audio, and faces, objects, and actions in video. These descriptions are structured within metadata or stored separately as distinct descriptions.
Read more about metadata.