ISO 24619:2011 (PISA) - Overview
ISO 24619:2011, titled Language resource management - Persistent identification and sustainable access (PISA), defines requirements for using persistent identifiers (PIDs) to reference, cite and access language resources reliably over time. The standard targets digital language resources - for example digital dictionaries, terminological resources, machine‑translation lexica, annotated corpora and multimedia/multimodal language data - and specifies minimum requirements for PID frameworks and PID usage to ensure persistence, resolvability and reproducible citation.
Key topics and technical requirements
- PID framework requirements: mandates that persistent references be implemented through an established PID framework and imposes minimum functional and operational requirements on such frameworks (resolver services, metadata association, persistence guarantees).
- PID usage and citation: describes how PIDs can be used as references and citations both in documents and embedded inside language resources; supports inclusion of citation metadata tied to identifiers.
- Granularity and parts identification: addresses how to identify and access resource parts, snapshots and versions; advises on granularity decisions for complex, federated or multi‑layered language resources.
- Collections and virtual incarnations: covers published collections and virtual assemblies (resource collection incarnations) that are referenced as a whole while enabling access to constituent parts.
- Complementary and practical constraints: discusses issues of persistence, sustainability (avoiding fragile dependencies), and best practices to enable machine and human resolvability.
- Terminology and definitions: provides clear definitions for resources, collections, archives, repositories, fragments, snapshots, versions and other core concepts relevant to language data management.
Practical applications and users
ISO 24619:2011 is intended for:
- Computational and applied linguists who create or cite corpora, lexica and annotated datasets.
- Information specialists, digital librarians and repository managers responsible for long‑term access and citation of language resources.
- Research projects and archives that need stable linking between distributed resources, parts, and multimedia elements.
- Developers of language technology platforms that must resolve identifiers reliably for automated processing (e.g., NLP pipelines, e‑Science workflows).
Typical applications:
- Assigning PIDs to corpora, datasets or lexical entries to enable reproducible research and machine‑actionable citations.
- Using resolver services to keep links valid when resources move locations or are mirrored.
- Referencing specific parts or snapshots of complex or federated resources in publications and metadata records.
Related standards (examples cited in ISO 24619)
- Citation formats: ISO 690, APA, MLA
- Fragment/part identification: ISO/IEC 21000‑17 (MPEG‑21), XPointer, IETF RFC 5147
- PID systems and implementations: DOI, Handle System, ARK, PURL
- Web and metadata technologies referenced for PID use and metadata attachment
ISO 24619:2011 helps organizations and researchers ensure persistent identification, sustainable access, and reliable citation of language resources - critical for reproducibility, long‑term preservation and machine‑readable scholarly infrastructures.