If you have difficulty in submitting comments on draft standards you can use a commenting template and email it to admin.start@bsigroup.com. The commenting template can be found here.
This document specifies the Semantic Similarity Markup Language (SSML) for the uniform and interoperable representation of semantic similarity information across languages, granularities, corpora and methods. SSML provides a formal specification for representing three core of a semantic similarity assertion. With respect to annotation granularity, this document applies to semantic similarity annotation at various levels: lexical (words, phrases), sentential (clauses, sentences), discoursal (paragraphs, sections, and documents), and non-textual (vectors). This document does not prescribe methods for computing semantic similarity, normalisation strategies, aggregation algorithms, or evaluation procedures. Also, it does not prescribe the content of semantic similarity annotation guidelines, nor does it define particular similarity measurement tasks or evaluation benchmark formats.
Semantic similarity is a foundational concept in natural language processing and computational linguistics, used to measure the degree of semantic relatedness between two semantic objects — including words, phrases, sentences, paragraphs, and long texts. Semantic similarity underpins the core logic of a wide range of language technology applications, including information retrieval, question answering, machine translation, paraphrase recognition, textual entailment inference, semantic search, text summarisation, and cross-lingual information extraction. With the widespread adoption of largescale pre-trained language models, semantic similarity measurement has also become a central part in language model alignment and evaluation. In the era of big data, the demand for interoperability of semantic similarity data across systems, languages, and tasks is growing rapidly in both industry and academia. However, no unified representation schema for semantic similarity information currently exists. Individual research datasets, evaluation benchmarks, and application systems each define their own data formats. For instance, STS (Semantic Textual Similarity) evaluation datasets typically store similarity scores in tab-separated plain text; lexical similarity datasets (e.g., SimLex-999, WordSim-353) record word pairs and scores in custom XML or CSV formats; and cross-lingual similarity tasks (e.g., multilingual STS-B) introduce additional language alignment fields. This fragmentation hinders the comparison and interoperability of similarity results across sources, makes it difficult to trace and reproduce similarity judgements produced by different methods within a unified framework, and prevents the integration of semantic similarity data across systems due to format incompatibility. This proposal develops SSML within the ISO 24617 SemAF series, aligning with established principles (abstract/concrete syntax separation, standoff annotation, language neutrality). It provides comprehensive, extensible, interoperable, context-aware and replicable similarity annotation, supporting accessibility use cases such as lexical simplification for users with dyslexia or poor literacy.
Required form fields are indicated by an asterisk (*) character.
You are now following this standard. Weekly digest emails will be sent to update you on the following activities:
You can manage your follow preferences from your Account. Please check your mailbox junk folder if you don't receive the weekly email.
You have successfully unsubscribed from weekly updates for this standard.
Comment by: