Fingerprints¶
Ponyglot computes a fingerprint of every segment’s source text. When the fingerprint changes, the segment’s translations become stale. Ponyglot’s value is authoritative.
Connectors may compute the same fingerprint for their own change detection and send it as
fingerprint. If it differs from Ponyglot’s value, the push still succeeds, Ponyglot uses its
own value, and the segment key is listed in fingerprint_mismatches. That is a hint that the
connector’s implementation differs.
Rules¶
The rules are frozen for API v1. Changing them would make every translation stale.
Normalize to Unicode NFC.
Replace CRLF and CR line endings with LF.
For
plainandmarkdown: remove trailing whitespace from every line, collapse runs of spaces and tabs inside a line to one space, and remove leading and trailing blank lines. Line breaks are kept.For
htmlandrich_text: collapse every run of whitespace, including line breaks, to one space, then trim. Markup is not re-serialized, so attribute order and quoting count.The fingerprint is the lowercase hex SHA-256 of the UTF-8 encoded result.
Examples¶
|
|
Normalized |
|---|---|---|
|
|
|
|
|
|
|
|
|
Reference implementation¶
import hashlib
import re
import unicodedata
_INLINE_SPACE = re.compile(r"[ \t]+")
_ANY_SPACE = re.compile(r"\s+")
def normalize(text, format="plain"):
text = unicodedata.normalize("NFC", text).replace("\r\n", "\n").replace("\r", "\n")
if format in {"html", "rich_text"}:
return _ANY_SPACE.sub(" ", text).strip()
lines = [_INLINE_SPACE.sub(" ", line).rstrip() for line in text.split("\n")]
return "\n".join(lines).strip("\n").strip()
def fingerprint(text, format="plain"):
return hashlib.sha256(normalize(text, format).encode()).hexdigest()
Python’s \s and str.strip() treat all Unicode whitespace as whitespace, including
non-breaking spaces. Match that behaviour if you implement the rules in another language.