Overview
ISO/IEC 23092-6:2023 - "Information technology - Genomic information representation - Part 6: Coding of genomic annotations" defines a normative representation and coding model for genomic annotations. The standard covers structured encoding for multiple annotation types including variants with genotyping information, functional annotations, genome browser tracks, expression matrices, and contact matrices (from Hi‑C or similar experiments). It specifies data structures, descriptors, attribute semantics, compression/decompression transforms and decoding/output formats to enable interoperable storage, transmission and processing of genomic annotation data.
Key topics and technical requirements
- Scope and terminology: formal terms, conventions and abbreviated terms used across the standard.
- Data structures: annotation parameter sets, tile configuration, descriptor configuration, attribute parameter sets, annotation access units and block layout for efficient representation and access.
- Descriptors & attributes: semantic definitions for genomic intervals, variants, functional annotations, contact matrices and typed attributes used to describe annotation content.
- Compression/transforms: specification of inverse transforms and decompression building blocks (e.g., Lempel‑Ziv‑Welch, binarization, sparse transform, delta transform, run‑length encoding, serialization).
- Decompression algorithms: guidance on decoding using algorithms listed in the standard (for example, Context‑Adaptive Binary Arithmetic Coding, Lempel‑Ziv‑Markov Chain Algorithm, Zstandard, JBIG, block sorting coder).
- Decoding process: step‑by‑step procedures for decoding access units and descriptors (variant access units, functional annotation access units, expression and contact matrix access units).
- Output formats: standardized output record definitions (variant site records, genotype records) including semantics and initialization procedures.
- Semantics and data types: clear rules for typed data, order of operations, logical/arithmetic operators and precise semantics for consistent interpretation.
Applications and who should use it
- Bioinformatics tool developers building variant callers, annotation pipelines, genome browsers and visualization tools.
- Sequencing centers, clinical genomics labs and data repositories needing standardized storage, compression, and exchange of large annotation datasets.
- Software vendors and standards bodies implementing interoperable genomic data formats for research, clinical and population genomics.
- Data engineers optimizing long‑term archival and high‑performance retrieval of genomic annotations, expression matrices and Hi‑C contact maps.
Related standards
- Other parts of the ISO/IEC 23092 series (genomic information representation) for complementary specifications on container formats and metadata.
- Industry and community formats (e.g., VCF, BAM/CRAM, BED) - ISO/IEC 23092‑6 focuses on a normative binary coding model useful for high‑efficiency exchange and storage.
Keywords: ISO/IEC 23092-6:2023, genomic annotations, genomic information representation, coding of genomic annotations, variants, expression matrices, contact matrices, Hi‑C, compression, decompression codecs, descriptors, access units.