Overview
ISO 24611:2012 - Morpho-syntactic annotation framework (MAF) defines a standardized framework for representing morpho-syntactic annotations of word-forms in texts. It provides a meta-model that links tokens, word-forms and their morpho‑syntactic properties to persistent semantic descriptors in the ISOCat data category registry (ISO 12620). The standard also specifies an XML serialization for MAF and explains equivalences with TEI (Text Encoding Initiative) guidelines, enabling interoperable annotated language resources.
Key topics and technical requirements
- MAF meta-model: conceptual model describing relationships among tokens, word-forms, lexical references and annotation objects.
- Token segmentation: formal representation of tokens (including embedded, stand‑off and TEI‑based variants) and rules for joining or overlapping tokens.
- Word-form representation: how word-forms attach to tokens (one-to-one, many-to-one, discontinuous spans, zero-token cases) and how they reference lexical entries.
- Morpho‑syntactic content: use of feature structures, compact tagsets, and FSR libraries to encode grammatical properties (gender, number, POS, etc.) in a machine-readable way.
- Tagset design and element: mechanisms for defining and reusing morpho-syntactic tagsets within MAF.
- Ambiguity handling: explicit constructs for representing lexical, word-form and structural ambiguities (e.g., , local lattices, mixed linear/lattice representations).
- XML Serialization & Elements: normative elements documented (for example, , , , ) and attribute classes for standardized encoding.
- Linking to ISOCat: requirement to reference standardized data categories (persistent identifiers) so annotations are semantically interoperable.
Practical applications and users
ISO 24611:2012 is intended for anyone creating or consuming annotated linguistic corpora or tools that require reliable, interoperable morpho-syntactic data:
- NLP engineers building POS taggers, morphological analyzers, parsers and training datasets.
- Corpus linguists and language resource developers who need standard tokenization and annotation formats.
- Lexicographers and terminologists linking word-forms to lexical resources.
- Tool providers integrating TEI documents or providing import/export pipelines for annotated resources.
- Research groups sharing annotated corpora and aiming for reproducible, comparable annotations.
Benefits include improved interoperability, clearer semantics via ISOCat references, and consistent XML encodings that align with TEI.
Related standards
- ISO 12620 - Data category registry (ISOCat)
- TEI Guidelines - for text encoding equivalences and TEI-compliant document handling
- ISO/TC 37 family of language resource standards (context for language resource management)
Keywords: ISO 24611:2012, MAF, morpho-syntactic annotation, ISOCat, data category registry, XML serialization, TEI, tokens, word-forms, language resource management.