The content model¶
Ponyglot translates django CMS pages, Wagtail pages and model instances from django-parler and django-modeltranslation. They store content in very different ways: plugin trees, StreamField blocks, translation tables, one column per language. Ponyglot maps all of them to one simple model, so the same translation memory, glossary and QA work for everything.
Units and segments¶
A unit is one translatable object: a page, a blog post, a product. A segment is one translatable piece of text inside it: a title, a paragraph, an alt text, a meta description.
unit djangocms:page:42 "Pricing" /pricing/
├─ segment title "Pricing"
├─ segment meta_description "Simple pricing for multilingual Django sites: …"
├─ segment plugin:101:body "<h1>Simple, transparent pricing</h1> …"
└─ segment plugin:103:body "<p>€79 per site and month …</p>" (parent plugin:102)
Units and segments have stable keys chosen by the connector. The keys are how Ponyglot
recognises “the same text” across edits, so they must not change when an editor adds a
paragraph above or reorders blocks. That’s why keys name the plugin or block (plugin:103:body),
not its position.
Structure (position, parent_key) is stored separately. A connector needs it to rebuild a
plugin tree or nested blocks in the target language, but moving a block doesn’t change what
it says, so it doesn’t make translations stale.
Snapshots instead of change events¶
A connector always sends the complete current state of a unit, never a list of changes. Ponyglot compares the snapshot with what it knows and works out what was added, changed or removed.
This makes the protocol forgiving. A connector doesn’t need to remember what it sent before. If a request fails, it sends the same snapshot again; pushing an unchanged snapshot changes nothing. If Ponyglot missed an edit, the next push corrects it.
Fingerprints and delta sync¶
For every segment Ponyglot computes a fingerprint: a hash of the normalized source text. Normalization ignores differences that don’t change the meaning, such as trailing spaces or Windows line endings. Every translation remembers the fingerprint of the source it was made from.
When an editor changes a sentence, its fingerprint changes. Now the translations of that sentence were made from a different source: they are stale. The sentences around it keep their fingerprints and stay up to date. A delta job translates exactly the stale and missing segments, which is why only what changed gets retranslated.
Ponyglot computes fingerprints itself and trusts only its own value. A connector with a slightly different implementation therefore can’t corrupt the stale tracking. The rules are public and frozen for API v1 (see Fingerprints), because changing them would make every translation stale at once.
Stale is computed, not stored¶
Ponyglot doesn’t flip a “stale” flag when the source changes. It compares two fingerprints whenever it needs the state. This has useful consequences:
If an editor undoes a change, the old translations are current again automatically.
Nothing needs to be updated in bulk when a large page changes.
The state is always consistent with the current source.
Likewise, a missing translation is simply the absence of one. Adding a target language doesn’t create any placeholder rows; all segments are missing in that language until a job translates them.
Removed and deleted content¶
A segment that disappears from a snapshot is removed, but not forgotten: if it comes back, for example after an editor undoes a deletion, its translations come back with it. The same applies to units: a deleted unit is restored when the connector pushes it again.
Languages¶
Target languages are set per site in the dashboard. Whether Ponyglot may write a language
also depends on the website: django-modeltranslation, for example, needs a database column per
language. The connector’s handshake reports which languages the website has, and a unit can
narrow them further with writable_languages. Ponyglot translates only languages the website
can store.