Overview - ISO 24614-1:2010 (Word segmentation, basic concepts)
ISO 24614-1:2010 is an international standard in language resource management that defines the basic concepts and general principles for word segmentation of written texts. It provides language‑independent guidelines to segment text into reproducible units called word segmentation units (WSU). The standard is intended to make segmentation consistent across languages and applications so that processes like counting, lookup and automated processing are reliable and comparable.
Key topics and technical coverage
- Scope and objectives: Establishes a universal framework for dividing written text into WSUs for a wide range of languages and scripts.
- Terms and definitions: Precise definitions for core linguistic concepts used in segmentation - e.g. morpheme, lexeme, lemma, stem, word form, compound, multiword expression (MWE), WSU.
- Basic framework: Conceptual relationships among morphemes, lexemes and WSUs; treatment of compounding, agglutination, and lexicalization.
- General principles: Language‑independent rules and considerations (spaces, punctuation, abbreviations, numerics, multiword units, idioms) to achieve reliable and reproducible segmentation.
- Annex A (informative): Guidance on representing word segmentation in XML for interoperability and tooling.
- Standardization context: Prepared by ISO/TC 37 (Terminology and language resources); Part 2 addresses CJK specifics and Part 3 is planned for other languages.
Practical applications
ISO 24614-1 is applicable wherever accurate word boundaries matter:
- Translation & localization: consistent word counts, translation memory and CAT tool segmentation.
- Natural Language Processing (NLP): preprocessing for morphosyntactic analyzers, parsers, tokenizers, spellcheckers, text classification and corpus annotation.
- Speech technologies: lexicon lookup, TTS prosody and speech synthesis that require consistent lexical units.
- Content management & search: indexing and search require clear word boundaries for matching and retrieval.
- Lexicography & terminology management: consistent corpus counts and lexicon construction.
Who should use this standard
- NLP engineers and data scientists building tokenizers and text pipelines
- CAT tool, TMS and CMS developers and integrators
- Speech technology vendors (TTS, ASR)
- Corpus linguists, lexicographers and terminology managers
- Project managers needing reproducible word counts for costing and QA
Related standards
- ISO 24614-2 - Word segmentation for Chinese, Japanese and Korean (CJK)
- Future ISO 24614‑3 - planned for additional language‑specific rules
Using ISO 24614-1 helps ensure consistent, interoperable word segmentation across tools and languages, improving accuracy in NLP workflows, translation workflows, search and lexicon management.