Overview
ISO/IEC 15938-4:2002/Amd 1:2004 - “Information technology - Multimedia content description interface - Part 4: Audio - Amendment 1: Audio extensions” is an amendment to the MPEG‑7 audio description part. It extends the audio metadata model to better support multi‑channel (surround) audio, enriched audio‑quality descriptors, and enhanced speech/lexicon elements. The amendment updates XML schema types (e.g., WordLexiconType, SpokenContentLinkType) and introduces the AudioSignalQualityType and handling rules for channel numbering and channel attributes.
Key topics and technical requirements
-
Multichannel handling and channel numbering
- Recommends a convention for mapping common surround tags (L, C, R, LS, RS, LFE) to channel numbers.
- Introduces the channels attribute (defined in ISO/IEC 15938-5/Amd.1 MDS) to specify which channels a descriptor applies to and to enable per‑channel operations (e.g., mean computation).
- Advises counting order (optional center, left→right, top→bottom, front→back) and ignoring channel indices that exceed the media file’s declared channel count.
-
Audio quality metadata (AudioSignalQualityType)
- Defines descriptors for signal quality and events, including BackgroundNoiseLevel, RelativeDelay, Balance, DcOffset, CrossChannelCorrelation, Bandwidth, TransmissionTechnology, and an ErrorEventList for discrete errors.
- Attributes such as IsOriginalMono and BroadcastReady support practical quality assessments.
-
Speech and lexicon extensions
- WordLexiconType extended with attributes like linguisticUnit (word, syllable, morpheme, phrase, nonspeech, etc.) and representation (orthographic/nonorthographic) to better index ASR outputs and lexicons.
- SpokenContentLinkType adds acousticScore plus existing probability and nodeOffset, enabling richer lattice/link scoring for speech recognition output.
Applications and who uses it
- Digital asset managers, broadcast engineers, and archivists who need standardized audio metadata for:
- Searching and filtering large audio/video collections by quality or channel content.
- Selecting appropriate files for download, streaming, or broadcast based on signal quality metadata.
- Developers of audio analysis, speech recognition (ASR) and spoken‑document retrieval systems:
- Use lexicon and spoken‑link enhancements to store and weight different linguistic units and acoustic scores.
- Multimedia metadata designers and integrators working with MPEG‑7 and ISO/IEC 15938 profiles for interoperable audio description and delivery.
Related standards
- ISO/IEC 15938 (MPEG‑7) - multimedia content description framework
- ISO/IEC 15938-5/Amd.1 (MDS) - AudioD and AudioDS types and the channels attribute referenced by this amendment
- MPEG‑7 XML schema namespaces (mpeg7) used throughout the amendment
Keywords: ISO/IEC 15938-4, MPEG‑7, audio metadata, multichannel audio, audio quality, WordLexiconType, SpokenContentLinkType, AudioSignalQualityType, surround 5.1, channels attribute, speech recognition metadata.