Overview
SIST ISO 24611:2013 (ISO 24611:2012) - Morpho-syntactic annotation framework (MAF) defines a standardized framework for morpho-syntactic annotation of word-forms in texts. It provides a meta-model that links tokens, word-forms, lexical references and morpho‑syntactic properties to an authoritative registry of semantic descriptors (the ISOCat / ISO 12620 data category registry). The standard also specifies an XML serialization for MAF and describes equivalences with TEI (Text Encoding Initiative) guidelines to promote interoperability.
Key technical topics and requirements
- MAF meta-model: conceptual separation of text segmentation (tokens), linguistic units (word-forms), and their morpho‑syntactic descriptions.
- Tokenization strategies: inline, stand-off and TEI‑aligned notations; rules for joining, overlapping and spanning tokens.
- Word-form modeling: attachment patterns (one-to-one, one-to-many, discontinuous or zero-token forms), compound word-forms and lexical pointers.
- Morpho-syntactic content: representation using feature structures, compact tagsets, and FSR libraries; guidance for designing reusable tagsets.
- Handling ambiguities: mechanisms for encoding alternative analyses (e.g., , local lattices, simplified linear/mixed representations).
- XML elements and serialization: normative element names and structures (e.g., , , , , , , ) are specified in the normative annex.
- Data category linkage: mandatory referencing of data categories in ISOCat/ISO 12620 for stable semantic interoperability.
- Normative and informative annexes: includes encoded examples and a full MAF specification to support implementation.
Practical applications and users
MAF is intended to enable consistent, interoperable morpho‑syntactic annotation across tools and corpora. Typical applications:
- Corpus annotation and standards-compliant corpora for linguistics research
- POS tagging and morphological analysis for NLP pipelines
- Creation and exchange of annotated training data for machine learning
- Interchange between TEI-encoded texts and NLP tools
- Lexicography and language resource management
Primary users:
- Corpus linguists, computational linguists, NLP engineers
- Language resource managers and standards implementers
- Digital humanities researchers and TEI practitioners
- Tool developers creating annotation editors, converters, or corpus platforms
Related standards
- ISO 12620 - Data category registry (ISOCat): provides the semantic descriptors MAF references.
- TEI guidelines - MAF specifies equivalences to TEI encoding for easier integration with TEI‑based resources.
By adopting ISO 24611 (MAF), organizations ensure transparent morpho-syntactic encoding, stronger interoperability, and clearer semantic grounding for annotated language resources.