diff --git a/.claude/skills/fix-issue/findings/mdl-other.jsonl b/.claude/skills/fix-issue/findings/mdl-other.jsonl index 7dc32ef07d..afdf7116b0 100644 --- a/.claude/skills/fix-issue/findings/mdl-other.jsonl +++ b/.claude/skills/fix-issue/findings/mdl-other.jsonl @@ -50,3 +50,5 @@ {"area": "mdl/offlinepaths", "date": "2026-08-27", "symptom": "Adding an offline navigation profile makes a build fail with **CE6206** (\"Attribute paths with multiple steps cannot be used on pages that are accessible through an offline-based navigation\") in pages the statement never mentioned", "cause": "An offline profile restricts every page it can REACH to at most one association hop. The pages were valid before; creating the profile is what invalidated them, and mxcli writes multi-step paths happily and said nothing", "file": "`mdl/offlinepaths/scan.go` (the stored-document scan), `mdl/executor/offline_profile_warning.go` (the report), `mdl/types/navigation_profile_kind.go` (`IsOfflineProfileKind`)", "insight": "**The threshold is TWO hops, not \"any indirect reference\"** — one hop is explicitly allowed, pinned by a control inside one page where mxbuild flags column 2 of 2 and accepts column 1. Scan for the stored `DomainModels$AttributeRef` whose `EntityRef.Steps` holds ≥2 steps, keying on `$Type` and not on a table of widget property names — a DataGrid2 column's binding is nested several levels inside a `CustomWidgets$WidgetValue`, so a property-name scan reports a clean project for the widget people actually use. Do NOT match every `IndirectEntityRef`: a data source that navigates an association is a different element and CE6206 does not reject it. Report as a **warning**, never a refusal — reachability needs the whole page graph (home pages, menu items, and every page those open), which mxcli does not walk, and the end-to-end control shows mxbuild at 0 errors while the flagged page is unreachable and CE6206 the moment the profile's home page points at it. ako/mxcli-maintenance", "ce": ["CE6206"]} {"area": "mdl/types", "date": "2026-08-28", "symptom": "`rename entity` / `rename module` reports success but prints no \"Updated N reference(s)\" line, and the next `mx check` fails **CE0174** \"Cannot resolve object name 'MyFirstModule.Period'\" on a view entity", "cause": "A view's OQL is the ONE place a qualified name is stored **embedded in a sentence** rather than as a property of its own. Both engines' rename walkers match a string that *equals* the old name or *begins with* it — correct for a BY_NAME property, and blind to `from MyFirstModule.Period as p` mid-query", "file": "`mdl/types/oql_rename.go` (`RewriteOQLQualifiedName`), wired at the `Oql` key in `mdl/backend/modelsdk/infrastructure_write.go` (`replaceQNInDocCounted`) and `sdk/mpr/writer_rename.go` (`replaceStringsInDoc`)", "insight": "**Ask where each reference is STORED, not just which documents reference it** — a scan built for whole-string properties silently covers 0 of the embedded ones, and reports 0, which reads like \"nothing referenced it\". Scope the rewrite to the `Oql` key: a blanket substring replace across every string reaches documentation and expressions, where a similar-looking name is not a reference. Three name-shaped ways it goes wrong, all covered by tests: a **longer name starting with the old one** (`M.PeriodDetail`) must not move, a **quoted** reference must stay quoted (bare is CE0174 — quoting is how a view names an entity called after an OQL reserved word), and a **module rename arrives as a prefix pair** (`\"Old.\" -> \"New.\"`) with no entity half to split, so it needs its own path or module rename stays broken while entity rename looks fixed. Where a query-local alias is spelled exactly like the module, the rewrite **declines** rather than guessing — 0 references reported is a visible non-event, a corrupted query is not. Both engines, one shared rewrite. ako/mxcli-captrack", "ce": ["CE0174"]} {"area": "mdl/translations", "date": "2026-08-30", "raw": "| `create or modify translations in for ` reports success (\"Set 212 nl_NL translation(s) across 20 document(s)\") and the app's pages switch language while the **menu does not** | `mdl/translations/outofscope.go` (new), `mdl/executor/cmd_translations.go`, `cmd/mxcli/syntax/features_misc.go` | The **navigation is a project-level document**, not a module one, so `in ` never reaches it. Measured on the reporting project: 151 strings scoped against 546 unscoped, and re-running the same file unscoped landed 65 more across 22 further documents. Nothing warned — the document count was the only tell, and only if you knew what number to expect. A scoped run now names **the file's own entries** it did not reach (`translations.OutOfScope`), not \"the project has other strings\", which is true of every scoped run and would warn forever — the per-module workflow is exactly what the scoping exists to support. Second, load-bearing half: those entries were previously swept into the **drift** warning, whose premise (\"no text has this as its source\") is *false* about them — they matched, out of scope. They are subtracted from it, so \"the text may have been deleted\" is only said where it is true. Controls: an unscoped run of the same file reports nothing new and lands the strings; a key matching nothing anywhere is still reported as drift. Reported as ledger #137 |", "refs": ["#137"]} +{"area": "mdl/catalog", "date": "2026-09-03", "symptom": "`SHOW LANGUAGES` omits a language the project really has (ar_DZ absent from a list of 8 where the project has 9), and `search ''` returns \"No matches found\" for a string `DESCRIBE TRANSLATIONS` lists. Nothing errors and the catalog builds clean.", "cause": "CATALOG.strings was filled by hand-written per-type extractors reaching five sites (page title, enum caption, three microflow message templates), so a text anywhere else — every widget caption, tooltip, validation message, client template — was never indexed.", "file": "mdl/catalog/builder_strings.go", "insight": "A language present only on an unindexed site is INVISIBLE, not undercounted, so it vanishes from SHOW LANGUAGES entirely and from lint rule QUAL005, which discovers its language set from the same table. The fix is not a sixth case — that is how five was ever the number. Index from the type-agnostic walk DESCRIBE TRANSLATIONS already uses (translations.SitesInUnit over ListRawUnitsByType(\"\")), leaving only non-Texts$Text strings in the typed path (URLs, log nodes, REST paths, documentation, and Microflows$StringTemplate, which holds a plain Text and cannot carry a translation). Derive ObjectType from the unit $Type mechanically rather than via a table. Measured before: 69 of 3265 texts, 8 of 9 languages, 66 en_US of 1045. After: 1496 rows, 9 languages, counts identical to an independent BSON walk. Atlas design templates are ~70% of the corpus and are indexed rather than excluded, because CREATE TRANSLATIONS writes them and a SHOW LANGUAGES that excluded them would reopen the same split. CONTROL: stub the walk and the run reports `strings: 3` with SHOW LANGUAGES reporting nothing at all.", "refs": ["#250"]} +{"area": "mdl/linter", "date": "2026-09-03", "symptom": "Lint rule QUAL005 reports no missing translation for an enumeration where only one value is translated (11 real gaps unreported), and likewise for a page's sibling action buttons.", "cause": "The rule grouped by (QualifiedName, StringContext) while ElementId sat unused in the strings table, so every sibling element of one type collapsed into one group and a single translated value made the set look complete.", "file": "mdl/linter/rules/missing_translations.go", "insight": "Add ElementId to the SELECT, the ORDER BY and the elementKey struct. No test caught it because the harness synthesized ElementId from QualifiedName+StringContext, giving every sibling the same value and reproducing the defect inside the fixture — a fixture that encodes the bug cannot detect it. CONTROL: with every sibling translated the run must stay at 0 violations, or the new violation is an artifact of splitting the group rather than the missing translation.", "refs": ["#250"]} diff --git a/CHANGELOG.md b/CHANGELOG.md index 3ce8ccf96a..087a6c3dbd 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,12 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). ## [Unreleased] +- **`CATALOG.strings` indexes every translatable string, not five hand-picked kinds** — `SHOW LANGUAGES` listed 8 of a project's 9 languages and `search` could not find a widget caption that `DESCRIBE TRANSLATIONS` had just listed. The index was filled by per-type extractors reaching five sites (page title, enum caption, three microflow message templates), so a text anywhere else was never indexed: measured on a stock 11.13 app, **69 of 3265 texts and 8 of 9 languages**. A language present only on an unindexed site is *invisible* rather than undercounted, which also blinded lint rule QUAL005 — it discovers its language set from the same table. + + The rows now come from the type-agnostic `Texts$Text` walk that `DESCRIBE TRANSLATIONS` already uses, so the two subsystems cannot disagree about what the project contains; the typed path keeps only the strings that are *not* translatable (URLs, log node names, REST paths, documentation, and the `Microflows$StringTemplate` a workflow name is stored in). `StringContext` now names the site — `Forms$ActionButton.Caption` rather than `page_title` — and `ObjectType` is derived from the unit `$Type` mechanically, so a document type Mendix adds later is named correctly with nobody maintaining a list. Same project after: 1496 rows, 9 languages, counts identical to an independent BSON walk. Atlas design templates are ~70% of the corpus and are indexed rather than dropped, because `CREATE TRANSLATIONS` writes them and a `SHOW LANGUAGES` that excluded them would reopen the same split; `ObjectType` is how a consumer filters them. + +- **QUAL005 reports the sibling elements it used to fold together** — the rule grouped by `(QualifiedName, StringContext)` while `ElementId` sat unused in the table, so an enumeration's twelve values became one group and translating any single value made the whole set look complete. Grouping now includes `ElementId`. The existing test harness synthesized that column from `QualifiedName+StringContext`, which is why no test caught it. + - **`call external action` now types its return value and its parameters** (mendixlabs/mxcli#1020) — a call against a consumed OData service was written with neither the result variable's type nor its parameters' types, so Mendix reported `CE7269` ("the return type for remote action … has changed") and `CE7252` ("the parameters … have changed"), and re-running `CREATE OR MODIFY EXTERNAL ENTITIES` never cleared them. It never could. Both codes are defined on `CallExternalAction.cs` — they are raised by the microflow **activity**, not by the entity — which is what made the reported remedy the wrong lever. Two omissions of the same shape, each a `DataTypes$` sub-document that was never written: @@ -43,6 +49,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). - **`describe navigation` no longer drops a profile's login page and not-found page** — the two clauses were written correctly and sat on disk, but the modelsdk reader type-asserted only the `$Type`s `modelsdk/gen` declares for those slots, and neither is what the documents carry: `LoginPageSettings` is stored as `Forms$FormSettings` (page under `Form`) and `NotFoundHomepage` as `Navigation$HomePage`. A failed type assertion leaves the field empty, so `DESCRIBE NAVIGATION` printed neither clause and pasting its output back — the copy workflow the docs recommend — deleted both from the profile. + The legacy engine read the same bytes correctly the whole time, which is both the diagnosis and the control: the two engines now print identical output for the same document. The two slots were wrong in opposite directions. For the login page a blank app's own navigation document and `generated/metamodel` agree with the writers, so **gen** is wrong. For the not-found page — Studio Pro's **"Fallback page"** — gen and `generated/metamodel` agree with each other and mxcli's three **writers** were the odd one out, storing `Navigation$HomePage` where Studio Pro stores `Navigation$NotFoundHomePage`; the writers are corrected and now reproduce a Studio Pro-authored fallback page exactly. mxbuild accepts either spelling, so only a reference document could separate them. The reader keeps accepting both, because every not-found page mxcli wrote before this carries the old spelling and has to keep round-tripping. + The legacy engine read the same bytes correctly the whole time, which is both the diagnosis and the control: the two engines now print identical output for the same document. The two slots were wrong in opposite directions. For the login page a blank app's own navigation document and `generated/metamodel` agree with the writers, so **gen** is wrong. For the not-found page — Studio Pro's **"Fallback page"** — gen and `generated/metamodel` agree with each other and mxcli's three **writers** were the odd one out, storing `Navigation$HomePage` where Studio Pro stores `Navigation$NotFoundHomePage`; the writers are corrected and now reproduce a Studio Pro-authored fallback page exactly. **Correction to an earlier claim in this entry:** it previously said mxbuild accepts either spelling, and that only a reference document could separate them. That is wrong, and it understated the bug. Measured on 11.13 against a build emitting the old spelling, both `mx check` and `mxbuild --target=deploy` exit 1 with `Object of type 'Mendix.Modeler.WebUI.Navigation.HomePage' cannot be converted to type 'Mendix.Modeler.WebUI.Navigation.NotFoundHomePage'` — the project cannot be **loaded**, so every check downstream is lost with it. Nothing caught it because nothing had ever *built* a project with a fallback page set: the automated `mx check` coverage runs `doctype-tests/` only, and no script there sets one — the first that does was added by this fix. The reader keeps accepting both spellings for a different reason than stated: a project written before this does not build at all, and mxcli parses the BSON directly, so reading the old spelling is what lets it open that project and repair it. @@ -491,7 +499,6 @@ Headline: **A statement mxcli accepts is now a statement mxcli honours.** This r - An additive chain keeps its operators in the order they were written. - A building-block datasource override is rebound by widget type, not by one that happens to be present already. - ## [0.18.0] - 2026-08-14 Headline: **mxcli can now maintain a project it did not author.** Marketplace modules install and update headlessly — carrying the GUIDs the database keys on, the role grants that live outside the module, and the MPR v2 format `mx module-import` silently collapses — and `marketplace diff` reports which elements were edited locally before an update replaces them. Alongside that, a write that changes nothing no longer touches the file, five more document types become authorable (task queues, scheduled events, regular expressions, validation rules, menus), and a long tail of activities that could be written but not read back stop disappearing from the describe → edit → re-exec loop. Separately, the Windows and macOS binaries stop shipping the embedded tunnel — 13.5 MB smaller, and no longer carrying the tunnelling stack that had Defender and enterprise EDR blocking mxcli on managed corporate endpoints. diff --git a/docs-site/src/internals/catalog-schema.md b/docs-site/src/internals/catalog-schema.md index 49f48079e3..01e3177e26 100644 --- a/docs-site/src/internals/catalog-schema.md +++ b/docs-site/src/internals/catalog-schema.md @@ -186,12 +186,31 @@ CREATE TABLE PERMISSIONS ( ```sql CREATE VIRTUAL TABLE STRINGS USING fts5( - name, -- Document qualified name - kind, -- Document type - strings, -- All text content concatenated - tokenize='porter unicode61' -); -``` + QualifiedName, -- Document qualified name, e.g. MyModule.Home + ObjectType, -- Document type, derived from the unit $Type: PAGE, + -- PAGE_TEMPLATE, BUILDING_BLOCK, MICROFLOW, ENUMERATION, ... + StringValue, -- The string itself + StringContext, -- Where it lives. For translatable text this is + -- ., e.g. Forms$ActionButton.Caption. + -- Non-translatable strings keep a plain label: page_url, + -- log_node, documentation, rest_path, task_name, ... + Language, -- Language code, empty for non-translatable strings + ElementId, -- The owning element's $ID — what distinguishes an + -- enumeration's twelve values from each other + ModuleName +); +``` + +Every `Texts$Text` in the project is indexed, found by a type-agnostic walk +rather than per-document-type extraction, so a caption in a document type mxcli +cannot otherwise read is still searchable. That includes Atlas's design +templates (`PAGE_TEMPLATE`, `BUILDING_BLOCK`), which are roughly 70% of a stock +project's text and never render in a running app — filter them out with +`ObjectType` when you want only the app's own strings. + +An empty translation — a text that exists but is not translated yet — is **not** +a row, so a language's presence in this table means it is actually translated +somewhere. ### SOURCE (FTS5) diff --git a/docs-site/src/language/translations.md b/docs-site/src/language/translations.md index 0d4751ae57..7b031b7749 100644 --- a/docs-site/src/language/translations.md +++ b/docs-site/src/language/translations.md @@ -61,7 +61,7 @@ modules ship translations in **nine**, so "other languages already have translations here" is true and misleading. > `SHOW LANGUAGES` lists languages that **have translations**, which is a -> different list — a stock app reports eight while one is enabled. The enabled +> different list — a stock app reports nine while one is enabled. The enabled > list is in `DESCRIBE SETTINGS`. ## Drift: a source string that was edited diff --git a/docs/11-proposals/PROPOSAL_translations.md b/docs/11-proposals/PROPOSAL_translations.md index 9d331e40fa..065b925a53 100644 --- a/docs/11-proposals/PROPOSAL_translations.md +++ b/docs/11-proposals/PROPOSAL_translations.md @@ -1,6 +1,6 @@ --- title: Translations — preserve, describe, author, and auto-translate -status: draft +status: partial date: 2026-08-23 related: - PROPOSAL_catalog_integration.md @@ -9,7 +9,7 @@ related: # Proposal: Translations — preserve, describe, author, and auto-translate -**Status:** Draft +**Status:** Partial — all four slices shipped; see Implementation Status. **Date:** 2026-08-23 A Mendix app ships its user-visible strings in every language it supports. @@ -307,6 +307,38 @@ makes deduplication and hand-editing work. This is the translation-memory problem, and TM tools answer it the same way: keep the dictionary source-keyed, flag "source changed", do not guess. +## Implementation Status + +All four slices below are shipped. What remains is one convenience flag and the +open questions at the end of this document; the feature itself is usable. + +| Slice | State | Landed in | +|-------|-------|-----------| +| 1 — Preservation | done | #245 | +| 2 — The `Texts$Text` overlay | done | #250 | +| 3 — `DESCRIBE TRANSLATIONS` | done | #250 | +| 4 — `CREATE TRANSLATIONS` + drift report | done, except `--untranslated` | #250 | + +Shipped beyond the original plan, because running the feature exposed the need: + +- **Enabled-language statements** — `ALTER SETTINGS ADD/REMOVE/ADD OR MODIFY + LANGUAGE`, which is what Open Question 2 turned into (#257 and follow-ups). + `create translations` now *warns* when the language is not enabled rather than + refusing or enabling it silently. +- **An out-of-scope report** (`mdl/translations/outofscope.go`) — a scoped run + names the entries its own scope kept it from reaching. Ledger #137: `in Ledger` + never reaches the project-level NAVIGATION, so the pages went Dutch and the + sidebar did not, under a message that read as success. +- **A removal form** — an `OR REPLACE` naming nothing takes a language's + translations back out, which is otherwise unexpressible. +- **A lint rule** (`mdl/linter/rules/missing_translations.go`), a skill + (`.claude/skills/mendix/translations/SKILL.md`), and a manual page + (`docs-site/src/language/translations.md`). + +Still unbuilt: **`--untranslated`** on `DESCRIBE`, to emit only the empty targets. +It is a cost optimisation for the LLM loop on an already-translated project, not +a correctness gap — the full describe is what the loop uses today. + ## Implementation Plan Four slices, each shippable alone. **Slice 1 is the bug and should go first.** @@ -419,23 +451,60 @@ observed uniformly (3299 of 3299) on 11.13.0 and is the same value ## Open Questions +Two of the five are settled by shipped work; the resolutions are recorded here +rather than deleted, since each was decided by a measurement. + +### Settled + +2. **The enabled-language list.** ~~Does a translation for an unenabled language + do anything?~~ **Measured**: a stock 11.13 app enables exactly one language + (`en_US`) while its documents carry translations in nine, all from marketplace + modules — so Studio Pro stores and keeps them, and `mx check` passes. But the + app does not *serve* an unenabled language, which makes translating 411 + strings into a language nobody can select the quiet failure here. Resolved by + doing neither of the two things this question offered: `CREATE TRANSLATIONS` + does not enable the language (that is a settings change, and not this + statement's business) and does not refuse (the translations are legitimately + stored) — it **warns**, and `ALTER SETTINGS ADD LANGUAGE` is the fix it names. + Note the trap the skill now documents: `show languages` lists languages that + have *translations*, not enabled ones, so it reports 8 where 1 is enabled. + + +3. **Catalog coverage.** ~~`CATALOG.strings` misses widget captions; worth + widening so `SHOW LANGUAGES` and `search` reflect reality.~~ **Done**, and + wider than the question assumed. Measured on a stock project the index held + **69 of 3265 texts and 8 of 9 languages** — a language present only on an + unindexed site was *invisible*, not undercounted, so `SHOW LANGUAGES` named 8 + and `search` returned nothing for a caption `DESCRIBE TRANSLATIONS` had just + listed. It also blinded the QUAL005 lint rule that shipped with slice 4, which + discovers its language set from the same table. + + The resolution was not to add cases. The five extractors were hand-written per + type, so a sixth site cost a sixth case; the index is now built from **this + proposal's own walk** (`translations.SitesInUnit`), which is what makes the two + subsystems structurally unable to disagree about what the project contains. + `StringContext` names the site (`Forms$ActionButton.Caption`) and `ObjectType` + is derived from the unit `$Type`, so a document type Mendix adds later is + handled with no list to maintain. Atlas design templates are ~70% of the corpus + and are indexed rather than excluded, because `CREATE TRANSLATIONS` writes them + — `ObjectType` is how a consumer filters them out. + +5. **Ordering inside `Items`.** ~~Does Studio Pro care, as it did for widget + `PropertyTypes`?~~ **No** — the patch preserves existing order and appends, + and projects patched this way open in Studio Pro and build clean. What *did* + bite, and was not this, is the **`$ID` form**: a new `Texts$Translation` must + carry a 16-byte `$ID` as its **first** property or the build fails with + `Expected '$ID' as the first property of a storage object`. See + `mdl-examples/bug-tests/translation-id-form.mdl`. + +### Still open + 1. **Homographs.** One source string needing different translations in different - contexts. Mendix's Excel export has the same limitation, so matching it is - defensible — and `in ` scoping covers most real cases. Worth deciding - explicitly rather than discovering. -2. **The enabled-language list.** `Settings$LanguageSettings.Languages` exists in - gen (`modelsdk/gen/settings/types.go:768`) and is surfaced nowhere — - `describe settings` shows only `DefaultLanguageCode`. Does adding a translation - for a language that is not enabled on the project do anything useful? Needs one - measurement in Studio Pro. If not, `CREATE TRANSLATIONS` should enable it or - refuse. -3. **Catalog coverage.** `CATALOG.strings` indexes 21 contexts and misses widget - captions — 39 texts in a page, 2 indexed. Worth widening so `SHOW LANGUAGES` - and `search` reflect reality, but it is a separate change and this proposal - does not depend on it. -4. **The 2 texts with translations but no default language.** Harmless to skip on - export, but the import should not silently create a default-language entry for - them. Confirm what Studio Pro does with such a text. -5. **Ordering inside `Items`.** The patch preserves existing order and appends. - Whether Studio Pro cares (it did for widget `PropertyTypes` — CE0463) is - unverified for texts; a Studio Pro open after an import settles it. + contexts. Unchanged, and now with usage behind it: `DESCRIBE` flags a + conflicting source (3 on a stock app), and `in ` resolves the common + case. Mendix's own Excel export has the same limitation. Worth deciding + explicitly rather than discovering, but nothing has forced the decision yet. + +4. **The 2 texts with translations but no default language.** Skipped on export, + as planned. What Studio Pro does with such a text is still unconfirmed, and + the import still does not invent a default-language entry for them. diff --git a/docs/11-proposals/README.md b/docs/11-proposals/README.md index a6b78f60ca..5ad06d2367 100644 --- a/docs/11-proposals/README.md +++ b/docs/11-proposals/README.md @@ -135,7 +135,7 @@ for display in this README): | [Replace Generated Playwright Tests with playwright-cli](proposal-playwright-cli.md) | Draft | The current approach (documented in proposal-playwright-testing.md) has Claude Code generate TypeScript test files (.spec.ts), then run them | | [Self-Describing Syntax Feature Registry](syntax-feature-registry.md) | Draft | Branch: research/recursive-help-discovery | | [Structured description of irreducible microflow graphs](PROPOSAL_structured_microflow_description.md) | Draft | DESCRIBE MICROFLOW renders a microflow's control flow as nested if/then/else. | -| [Translations — preserve, describe, author, and auto-translate](PROPOSAL_translations.md) | Draft | A Mendix app ships its user-visible strings in every language it supports. | +| [Translations — preserve, describe, author, and auto-translate](PROPOSAL_translations.md) | Partial | A Mendix app ships its user-visible strings in every language it supports. | | [Version-Aware Agent Support](PROPOSAL_version_aware_agent_support.md) | Draft | Three use cases require mxcli to be version-aware at the MDL level: | | [warm dev loop — Docker-free run and iPad split-screen preview](PROPOSAL_mxcli_dev_warm_loop.md) | Draft | Relates to: PROPOSAL_check_mxbuild_gap_heuristics.md (the static-check gate that | | [Workflow / Microflow Syntax Alignment](PROPOSAL_workflow_microflow_syntax_alignment.md) | Draft | MDL spells the same concept differently depending on which document type you are | diff --git a/mdl-examples/bug-tests/catalog-strings-coverage.mdl b/mdl-examples/bug-tests/catalog-strings-coverage.mdl new file mode 100644 index 0000000000..3920c868e7 --- /dev/null +++ b/mdl-examples/bug-tests/catalog-strings-coverage.mdl @@ -0,0 +1,58 @@ +-- CATALOG.strings reached five hand-written sites, so most of a project's +-- translatable text was never indexed — and a language present only on an +-- unindexed site was INVISIBLE rather than undercounted. +-- +-- Measured on testdata/expr-checker/minimal.mpr (MPR v2), before the fix: +-- +-- indexed actually in the project +-- translatable texts ~69 3265 +-- languages 8 9 <- ar_DZ missing +-- en_US translations 66 1045 +-- extraction sites 5 17 +-- +-- The five were page titles, enum captions and three microflow message +-- templates, each hand-written against a typed reader. Everything else — every +-- widget caption, tooltip, validation message and client template — was +-- unreachable, and a sixth site cost another hand-written case. +-- +-- Two commands showed the split, on the same project: +-- +-- describe translations for nl_NL -> 'Save' as 'Opslaan' +-- search 'Opslaan' -> No matches found. +-- +-- The fix indexes from the type-agnostic Texts$Text walk that DESCRIBE +-- TRANSLATIONS already uses (translations.SitesInUnit), so the two subsystems +-- cannot disagree about what the project contains. +-- +-- After the fix, on the same project: 1496 rows, 9 languages, en_US 1045, +-- nl_NL 333, ar_DZ 4 — identical to an independent BSON walk of the units. +-- +-- CONTROL: stub buildTranslatableStrings and re-run. `strings: 3` (only the +-- non-translatable rows survive), and `show languages` reports nothing at all. +-- Without that line, "9 languages" is equally consistent with a build that +-- never had the fix. + +refresh catalog full; + +-- Every language in the project, not only the ones on a page title. +show languages; + +-- A widget caption is translatable text; before the fix this found nothing. +search 'Opslaan'; + +-- The context now names the site a text lives at, e.g. +-- Forms$ActionButton.Caption, rather than a hand-picked label like page_title. +select StringContext, count(*) as n +from CATALOG.strings +where Language != '' +group by StringContext +order by n desc; + +-- Atlas design templates are ~70% of the corpus and never render in the app. +-- They ARE indexed (DESCRIBE TRANSLATIONS reaches them, so SHOW LANGUAGES must +-- agree), and ObjectType is what lets a consumer filter them out. +select ObjectType, count(*) as n +from CATALOG.strings +where Language != '' +group by ObjectType +order by n desc; diff --git a/mdl/catalog/builder_strings.go b/mdl/catalog/builder_strings.go index 3267792bd4..5f20eaf9f6 100644 --- a/mdl/catalog/builder_strings.go +++ b/mdl/catalog/builder_strings.go @@ -3,6 +3,13 @@ package catalog import ( + "sort" + "strings" + "unicode" + + "go.mongodb.org/mongo-driver/v2/bson" + + "github.com/mendixlabs/mxcli/mdl/translations" "github.com/mendixlabs/mxcli/sdk/microflows" "github.com/mendixlabs/mxcli/sdk/workflows" ) @@ -32,27 +39,22 @@ func (b *Builder) buildStrings() error { count++ } - // Extract from pages (title, URL) — using cached list + // Every TRANSLATABLE string comes from the walk below, not from here. What + // remains in the typed extractions is the strings that are not Texts$Text + // and so are invisible to it: URLs, log node names, REST paths, + // documentation, and the workflow templates (Microflows$StringTemplate, + // which holds a plain Text and cannot carry a translation). + + // Page URL (no language) pageList, err := b.cachedPages() if err == nil { for _, pg := range pageList { + if pg.URL == "" { + continue + } moduleID := b.hierarchy.findModuleID(pg.ContainerID) moduleName := b.hierarchy.getModuleName(moduleID) - qn := moduleName + "." + pg.Name - - pageID := string(pg.ID) - - // Page title translations (with language code) - if pg.Title != nil && pg.Title.Translations != nil { - for lang, t := range pg.Title.Translations { - insert(qn, "PAGE", t, "page_title", lang, pageID, moduleName) - } - } - - // Page URL (no language) - if pg.URL != "" { - insert(qn, "PAGE", pg.URL, "page_url", "", pageID, moduleName) - } + insert(moduleName+"."+pg.Name, "PAGE", pg.URL, "page_url", "", string(pg.ID), moduleName) } } @@ -66,36 +68,12 @@ func (b *Builder) buildStrings() error { mfID := string(mf.ID) - // Documentation (no language) + // Documentation (no language). The activities' message templates + // are Texts$Text and come from the walk. if mf.Documentation != "" { insert(qn, "MICROFLOW", mf.Documentation, "documentation", "", mfID, moduleName) } - - // Extract strings from activities - extractActivityStrings(mf.ObjectCollection, qn, "MICROFLOW", moduleName, insert) - } - } - - // Extract from enumerations (value captions) — using cached list - enums, err := b.cachedEnumerations() - if err == nil { - for _, enum := range enums { - moduleID := b.hierarchy.findModuleID(enum.ContainerID) - moduleName := b.hierarchy.getModuleName(moduleID) - qn := moduleName + "." + enum.Name - - enumID := string(enum.ID) - for _, val := range enum.Values { - if val.Caption != nil && val.Caption.Translations != nil { - valID := string(val.ID) - if valID == "" { - valID = enumID - } - for lang, t := range val.Caption.Translations { - insert(qn, "ENUMERATION", t, "enum_caption", lang, valID, moduleName) - } - } - } + extractLogNodeNames(mf.ObjectCollection, qn, "MICROFLOW", moduleName, insert) } } @@ -151,10 +129,131 @@ func (b *Builder) buildStrings() error { } } + b.buildTranslatableStrings(insert) + b.report("strings", count) return nil } +// buildTranslatableStrings indexes every Texts$Text in the project. +// +// It walks the RAW units rather than the typed readers, deliberately. The typed +// path reached five sites because each was hand-written, and a sixth cost +// another case; this reaches all of them — 17 distinct sites in a stock 11.13 +// app — with no per-type code, and covers document types mxcli cannot otherwise +// round-trip. Measured on that app, the typed path indexed ~69 of 3265 texts and +// saw 8 of 9 languages, so `ar_DZ` was invisible to SHOW LANGUAGES and to +// QUAL005 rather than merely undercounted. +// +// Atlas design templates (Forms$PageTemplate, Forms$BuildingBlock) are ~70% of +// the corpus and their captions never render in a running app. They are indexed +// anyway, with ObjectType naming the document type so a consumer can filter: +// DESCRIBE TRANSLATIONS reaches them and CREATE TRANSLATIONS writes them, so a +// SHOW LANGUAGES that excluded them would disagree with the statement that +// changes them — the same split this is closing. +// The caller's insert closure counts the rows it writes, so nothing is counted +// here — doing both reported twice the rows the table actually holds. +func (b *Builder) buildTranslatableStrings(insert func(string, string, string, string, string, string, string)) { + units, err := b.reader.ListRawUnitsByType("") + if err != nil { + return + } + + for _, u := range units { + if len(u.Contents) == 0 { + continue + } + var named struct { + Name string `bson:"Name"` + } + _ = bson.Unmarshal(u.Contents, &named) + + moduleName := b.hierarchy.getModuleName(b.hierarchy.findModuleID(u.ContainerID)) + qn := named.Name + if moduleName != "" && qn != "" { + qn = moduleName + "." + qn + } + + for _, r := range translatableRows(u.Type, qn, moduleName, u.Contents) { + insert(r.QualifiedName, r.ObjectType, r.StringValue, r.StringContext, r.Language, r.ElementID, r.ModuleName) + } + } +} + +// stringRow is one row of the strings index. +type stringRow struct { + QualifiedName string + ObjectType string + StringValue string + StringContext string + Language string + ElementID string + ModuleName string +} + +// translatableRows turns one unit's stored bytes into index rows, one per +// (text, language). A language present with an empty string is a text that is +// not translated yet and is skipped: indexing it would make the language look +// present everywhere it is not, which is the opposite of what QUAL005 asks. +func translatableRows(unitType, qualifiedName, moduleName string, raw []byte) []stringRow { + sites, err := translations.SitesInUnit(raw) + if err != nil { + return nil + } + objType := catalogObjectType(unitType) + + var out []stringRow + for _, site := range sites { + ctx := site.OwnerType + "." + site.Property + for _, lang := range sortedLangs(site.Targets) { + if site.Targets[lang] == "" { + continue + } + out = append(out, stringRow{ + QualifiedName: qualifiedName, + ObjectType: objType, + StringValue: site.Targets[lang], + StringContext: ctx, + Language: lang, + ElementID: site.ElementID, + ModuleName: moduleName, + }) + } + } + return out +} + +// sortedLangs keeps row order deterministic — a map iteration here would make +// the catalog's bytes differ between two builds of an unchanged project. +func sortedLangs(m map[string]string) []string { + out := make([]string, 0, len(m)) + for k := range m { + out = append(out, k) + } + sort.Strings(out) + return out +} + +// catalogObjectType turns a unit's stored $Type into the catalog's object-type +// vocabulary — "Forms$PageTemplate" to "PAGE_TEMPLATE". Derived rather than +// looked up in a table, so a document type Mendix adds later is named correctly +// without anyone maintaining a list. It agrees with the hand-written values on +// every type they both cover (PAGE, MICROFLOW, ENUMERATION). +func catalogObjectType(unitType string) string { + name := unitType + if i := strings.LastIndex(name, "$"); i >= 0 { + name = name[i+1:] + } + var b strings.Builder + for i, r := range name { + if unicode.IsUpper(r) && i > 0 { + b.WriteByte('_') + } + b.WriteRune(unicode.ToUpper(r)) + } + return b.String() +} + // extractWorkflowFlowStrings extracts strings from workflow activities recursively. func extractWorkflowFlowStrings(flow *workflows.Flow, qn, moduleName string, insert func(string, string, string, string, string, string, string)) { for _, act := range flow.Activities { @@ -207,42 +306,21 @@ func extractWorkflowFlowStrings(flow *workflows.Flow, qn, moduleName string, ins } } -// extractActivityStrings extracts string literals from microflow/nanoflow activities. -func extractActivityStrings(oc *microflows.MicroflowObjectCollection, qn, objType, moduleName string, insert func(string, string, string, string, string, string, string)) { +// extractLogNodeNames indexes the one microflow-activity string that is NOT a +// Texts$Text. The message templates that used to be extracted here — log, show +// message, validation feedback — are Texts$Text and come from the walk in +// buildTranslatableStrings, which also reaches the ones this never listed. +func extractLogNodeNames(oc *microflows.MicroflowObjectCollection, qn, objType, moduleName string, insert func(string, string, string, string, string, string, string)) { if oc == nil { return } - for _, obj := range oc.Objects { act, ok := obj.(*microflows.ActionActivity) if !ok || act.Action == nil { continue } - - actID := string(act.ID) - - switch a := act.Action.(type) { - case *microflows.LogMessageAction: - if a.MessageTemplate != nil && a.MessageTemplate.Translations != nil { - for lang, t := range a.MessageTemplate.Translations { - insert(qn, objType, t, "log_message", lang, actID, moduleName) - } - } - if a.LogNodeName != "" { - insert(qn, objType, a.LogNodeName, "log_node", "", actID, moduleName) - } - case *microflows.ShowMessageAction: - if a.Template != nil && a.Template.Translations != nil { - for lang, t := range a.Template.Translations { - insert(qn, objType, t, "show_message", lang, actID, moduleName) - } - } - case *microflows.ValidationFeedbackAction: - if a.Template != nil && a.Template.Translations != nil { - for lang, t := range a.Template.Translations { - insert(qn, objType, t, "validation_message", lang, actID, moduleName) - } - } + if a, ok := act.Action.(*microflows.LogMessageAction); ok && a.LogNodeName != "" { + insert(qn, objType, a.LogNodeName, "log_node", "", string(act.ID), moduleName) } } } diff --git a/mdl/catalog/builder_strings_test.go b/mdl/catalog/builder_strings_test.go new file mode 100644 index 0000000000..305f74632c --- /dev/null +++ b/mdl/catalog/builder_strings_test.go @@ -0,0 +1,152 @@ +// SPDX-License-Identifier: Apache-2.0 + +package catalog + +import ( + "testing" + + "go.mongodb.org/mongo-driver/v2/bson" + + "github.com/mendixlabs/mxcli/modelsdk/mpr" +) + +func txt(pairs ...string) bson.D { + items := bson.A{int32(3)} + for i := 0; i+1 < len(pairs); i += 2 { + items = append(items, bson.D{ + {Key: "$Type", Value: "Texts$Translation"}, + {Key: "LanguageCode", Value: pairs[i]}, + {Key: "Text", Value: pairs[i+1]}, + }) + } + return bson.D{{Key: "$Type", Value: "Texts$Text"}, {Key: "Items", Value: items}} +} + +func marshal(t *testing.T, d bson.D) []byte { + t.Helper() + raw, err := bson.Marshal(d) + if err != nil { + t.Fatal(err) + } + return raw +} + +// The gap this replaces: a widget caption is translatable text the hand-written +// extractor never reached, because it only ever read a page's Title. Measured on +// a stock project, the index held ~69 of 3265 texts. +func TestTranslatableRows_ReachesAWidgetCaptionNotJustThePageTitle(t *testing.T) { + raw := marshal(t, bson.D{ + {Key: "$ID", Value: mpr.IDToBsonBinary("11111111-1111-1111-1111-111111111111")}, + {Key: "$Type", Value: "Forms$Page"}, + {Key: "Name", Value: "Home"}, + {Key: "Title", Value: txt("en_US", "Home")}, + {Key: "Widgets", Value: bson.A{int32(3), + bson.D{ + {Key: "$ID", Value: mpr.IDToBsonBinary("22222222-2222-2222-2222-222222222222")}, + {Key: "$Type", Value: "Forms$ActionButton"}, + {Key: "Caption", Value: txt("en_US", "Save", "nl_NL", "Opslaan")}, + }, + }}, + }) + + rows := translatableRows("Forms$Page", "MyModule.Home", "MyModule", raw) + + var caption *stringRow + for i := range rows { + if rows[i].StringValue == "Opslaan" { + caption = &rows[i] + } + } + if caption == nil { + t.Fatalf("the button caption never reached the index; rows = %+v", rows) + } + if caption.Language != "nl_NL" { + t.Errorf("Language = %q, want nl_NL", caption.Language) + } + if caption.StringContext != "Forms$ActionButton.Caption" { + t.Errorf("StringContext = %q, want Forms$ActionButton.Caption", caption.StringContext) + } + if caption.ElementID != "22222222-2222-2222-2222-222222222222" { + t.Errorf("ElementID = %q, want the button's own", caption.ElementID) + } + if caption.ObjectType != "PAGE" { + t.Errorf("ObjectType = %q, want PAGE", caption.ObjectType) + } + + // The title must still be there — the walk replaces the typed extraction, + // it does not trade one site for another. + var sawTitle bool + for _, r := range rows { + if r.StringValue == "Home" && r.StringContext == "Forms$Page.Title" { + sawTitle = true + } + } + if !sawTitle { + t.Errorf("page title lost; rows = %+v", rows) + } +} + +// A language reaching the index at all is what decides whether SHOW LANGUAGES +// can list it and whether QUAL005 can reason about it. Measured: the old +// extractor saw 8 of the project's 9 languages, and ar_DZ was invisible rather +// than undercounted. +func TestTranslatableRows_ALanguageOnlyOnAWidgetStillReachesTheIndex(t *testing.T) { + raw := marshal(t, bson.D{ + {Key: "$Type", Value: "Forms$Page"}, + {Key: "Title", Value: txt("en_US", "Home")}, + {Key: "Widgets", Value: bson.A{int32(3), + bson.D{ + {Key: "$Type", Value: "Forms$Label"}, + {Key: "Caption", Value: txt("en_US", "Hello", "ar_DZ", "مرحبا")}, + }, + }}, + }) + + rows := translatableRows("Forms$Page", "MyModule.Home", "MyModule", raw) + + langs := map[string]bool{} + for _, r := range rows { + langs[r.Language] = true + } + if !langs["ar_DZ"] { + t.Fatalf("ar_DZ never reached the index, so SHOW LANGUAGES cannot list it; languages = %v", langs) + } +} + +// An empty translation is a text that exists but is not translated yet. Writing +// it as a row would make the language look present everywhere it is not, which +// is the opposite of what QUAL005 is for. +func TestTranslatableRows_AnEmptyTranslationIsNotARow(t *testing.T) { + raw := marshal(t, bson.D{ + {Key: "$Type", Value: "Forms$Page"}, + {Key: "Title", Value: txt("en_US", "Home", "nl_NL", "")}, + }) + + for _, r := range translatableRows("Forms$Page", "MyModule.Home", "MyModule", raw) { + if r.Language == "nl_NL" { + t.Fatalf("an untranslated nl_NL was indexed as though it were translated: %+v", r) + } + } +} + +// Atlas design templates are ~70% of a project's texts and their captions never +// render in the app, so a consumer has to be able to filter them out. They are +// indexed rather than dropped because DESCRIBE TRANSLATIONS reaches them, and a +// SHOW LANGUAGES that disagreed with it would be the same split being fixed here. +func TestCatalogObjectType_DerivesFromTheUnitTypeWithoutATable(t *testing.T) { + for unitType, want := range map[string]string{ + "Forms$Page": "PAGE", + "Forms$PageTemplate": "PAGE_TEMPLATE", + "Forms$BuildingBlock": "BUILDING_BLOCK", + "Microflows$Microflow": "MICROFLOW", + "Microflows$Nanoflow": "NANOFLOW", + "Enumerations$Enumeration": "ENUMERATION", + "Texts$SystemTextCollection": "SYSTEM_TEXT_COLLECTION", + "Navigation$NavigationDocument": "NAVIGATION_DOCUMENT", + "Forms$Layout": "LAYOUT", + } { + if got := catalogObjectType(unitType); got != want { + t.Errorf("catalogObjectType(%q) = %q, want %q", unitType, got, want) + } + } +} diff --git a/mdl/linter/rules/missing_translations.go b/mdl/linter/rules/missing_translations.go index 3c2e0a8ec6..e9ea544ffb 100644 --- a/mdl/linter/rules/missing_translations.go +++ b/mdl/linter/rules/missing_translations.go @@ -57,12 +57,18 @@ func (r *MissingTranslationsRule) Check(ctx *linter.LintContext) []linter.Violat } // Step 2: Find elements that have translations in some languages but not all. - // Group by (QualifiedName, StringContext) — each group should have all languages. + // Group by (QualifiedName, StringContext, ElementId) — each group should have + // all languages. + // + // ElementId is load-bearing: sibling elements of one type share a + // QualifiedName and a StringContext (an enumeration's twelve values, a page's + // action buttons), so grouping without it folds them into one group and a + // single translated value makes the whole set look complete. rows, err := db.Query(` - SELECT QualifiedName, ObjectType, StringContext, Language, StringValue + SELECT QualifiedName, ObjectType, StringContext, Language, StringValue, ElementId FROM strings WHERE Language != '' - ORDER BY QualifiedName, StringContext, Language + ORDER BY QualifiedName, StringContext, ElementId, Language `) if err != nil { return nil @@ -73,6 +79,7 @@ func (r *MissingTranslationsRule) Check(ctx *linter.LintContext) []linter.Violat type elementKey struct { QualifiedName string StringContext string + ElementID string } type elementInfo struct { ObjectType string @@ -82,11 +89,11 @@ func (r *MissingTranslationsRule) Check(ctx *linter.LintContext) []linter.Violat elements := make(map[elementKey]*elementInfo) for rows.Next() { - var qn, objType, sctx, lang, value string - if err := rows.Scan(&qn, &objType, &sctx, &lang, &value); err != nil { + var qn, objType, sctx, lang, value, elemID string + if err := rows.Scan(&qn, &objType, &sctx, &lang, &value, &elemID); err != nil { continue } - key := elementKey{qn, sctx} + key := elementKey{qn, sctx, elemID} info, ok := elements[key] if !ok { info = &elementInfo{ObjectType: objType, Languages: make(map[string]bool)} diff --git a/mdl/linter/rules/missing_translations_test.go b/mdl/linter/rules/missing_translations_test.go index 381a1e5eee..1c99afdafa 100644 --- a/mdl/linter/rules/missing_translations_test.go +++ b/mdl/linter/rules/missing_translations_test.go @@ -4,6 +4,7 @@ package rules import ( "database/sql" + "strings" "testing" "github.com/mendixlabs/mxcli/mdl/catalog" @@ -160,3 +161,76 @@ func TestMissingTranslationsRuleMetadata(t *testing.T) { t.Errorf("expected severity Warning, got %v", rule.DefaultSeverity()) } } + +// setupTranslationsDBWithIDs is setupTranslationsDB with the ElementId given +// rather than synthesized, which is what lets a test put two sibling elements in +// one document. +// Rows are [QualifiedName, ObjectType, StringValue, StringContext, Language, ElementId, ModuleName]. +func setupTranslationsDBWithIDs(t *testing.T, rows [][]string) catalog.CatalogDB { + t.Helper() + db, err := sql.Open("sqlite", ":memory:") + if err != nil { + t.Fatalf("failed to open in-memory db: %v", err) + } + if _, err := db.Exec(`CREATE VIRTUAL TABLE strings USING fts5( + QualifiedName, ObjectType, StringValue, StringContext, Language, ElementId, ModuleName + )`); err != nil { + t.Fatalf("failed to create strings table: %v", err) + } + stmt, err := db.Prepare(`INSERT INTO strings (QualifiedName, ObjectType, StringValue, StringContext, Language, ElementId, ModuleName) + VALUES (?, ?, ?, ?, ?, ?, ?)`) + if err != nil { + t.Fatalf("failed to prepare insert: %v", err) + } + defer stmt.Close() + for _, row := range rows { + if len(row) != 7 { + t.Fatalf("expected 7 columns, got %d", len(row)) + } + if _, err := stmt.Exec(row[0], row[1], row[2], row[3], row[4], row[5], row[6]); err != nil { + t.Fatalf("failed to insert row: %v", err) + } + } + return catalog.WrapSqlDB(db) +} + +// Sibling elements of one type share a QualifiedName and a StringContext — an +// enumeration's twelve values, a page's action buttons. Grouping without the +// ElementId folds them into one, so translating a single value makes the whole +// set look complete and eleven missing translations go unreported. +func TestMissingTranslations_SiblingElementsAreNotOneGroup(t *testing.T) { + db := setupTranslationsDBWithIDs(t, [][]string{ + // "Low" is translated. + {"MyModule.Priority", "ENUMERATION", "Low", "Enumerations$EnumerationValue.Caption", "en_US", "id-low", "MyModule"}, + {"MyModule.Priority", "ENUMERATION", "Laag", "Enumerations$EnumerationValue.Caption", "nl_NL", "id-low", "MyModule"}, + // "High" is not. + {"MyModule.Priority", "ENUMERATION", "High", "Enumerations$EnumerationValue.Caption", "en_US", "id-high", "MyModule"}, + }) + defer db.Close() + + violations := NewMissingTranslationsRule().Check(linter.NewLintContextFromDB(db)) + + if len(violations) != 1 { + t.Fatalf("want 1 violation for the untranslated value, got %d: %v", len(violations), violations) + } + if !strings.Contains(violations[0].Message, "High") { + t.Errorf("the violation names the wrong value: %q", violations[0].Message) + } +} + +// The control for the test above: with every sibling translated there is nothing +// to report, so the violation there is the missing translation and not an +// artifact of splitting the group. +func TestMissingTranslations_SiblingElementsAllTranslatedIsClean(t *testing.T) { + db := setupTranslationsDBWithIDs(t, [][]string{ + {"MyModule.Priority", "ENUMERATION", "Low", "Enumerations$EnumerationValue.Caption", "en_US", "id-low", "MyModule"}, + {"MyModule.Priority", "ENUMERATION", "Laag", "Enumerations$EnumerationValue.Caption", "nl_NL", "id-low", "MyModule"}, + {"MyModule.Priority", "ENUMERATION", "High", "Enumerations$EnumerationValue.Caption", "en_US", "id-high", "MyModule"}, + {"MyModule.Priority", "ENUMERATION", "Hoog", "Enumerations$EnumerationValue.Caption", "nl_NL", "id-high", "MyModule"}, + }) + defer db.Close() + + if v := NewMissingTranslationsRule().Check(linter.NewLintContextFromDB(db)); len(v) != 0 { + t.Fatalf("want 0 violations, got %d: %v", len(v), v) + } +} diff --git a/mdl/translations/sites.go b/mdl/translations/sites.go new file mode 100644 index 0000000000..50ce37b5e7 --- /dev/null +++ b/mdl/translations/sites.go @@ -0,0 +1,100 @@ +// SPDX-License-Identifier: Apache-2.0 + +package translations + +import ( + "go.mongodb.org/mongo-driver/v2/bson" + + "github.com/mendixlabs/mxcli/modelsdk/codec" +) + +// Site is one translatable text together with where it sits: the nearest +// enclosing storage object and the property the text hangs off. +// +// It exists so a consumer can name a text's location without knowing the +// document type. The catalog's string index used to reach five sites because +// each one was hand-written against a typed reader (page titles, enum captions, +// three microflow activities); the walk below reaches every site in the project +// with no per-type code — 17 distinct ones in a stock 11.13 app, and whatever a +// future Mendix version adds, for free. +type Site struct { + // OwnerType is the $Type of the nearest enclosing storage object, e.g. + // "Forms$ActionButton". It is the unit's own $Type for a text on the root. + OwnerType string + // Property is the key the Texts$Text hangs off, e.g. "Caption". + Property string + // ElementID is the owner's $ID as a UUID string, empty when it has none. + // + // Load-bearing for grouping: several sibling elements of one type carry the + // same OwnerType and Property, so a consumer that groups without this folds + // an enumeration's twelve values into one and calls the set complete as soon + // as any one value is translated. + ElementID string + // Targets is language code → text, exactly as stored. A language present + // with an empty string is a text that exists but is not translated yet. + Targets map[string]string +} + +// SitesIn returns every translatable text in a decoded document, in document +// order. Order is stable so a caller writing rows gets a deterministic result. +func SitesIn(doc bson.D) []Site { + var out []Site + var walk func(v any, ownerType, ownerID, prop string) + walk = func(v any, ownerType, ownerID, prop string) { + switch n := v.(type) { + case bson.D: + if ty, _ := lookup(n, "$Type").(string); ty == "Texts$Text" { + out = append(out, Site{ + OwnerType: ownerType, + Property: prop, + ElementID: ownerID, + Targets: translationsOf(n), + }) + return + } + // A node with its own $Type becomes the owner for everything below + // it. A node without one (an anonymous sub-document) leaves the + // owner as it was, so the text is still attributed to a real + // element rather than to nothing. + nt, nid := ownerType, ownerID + if ty, _ := lookup(n, "$Type").(string); ty != "" { + nt = ty + nid = elementIDOf(n) + } + for _, e := range n { + walk(e.Value, nt, nid, e.Key) + } + case bson.A: + // An array element inherits the property its array hangs off, so a + // text inside `Widgets` is not reported as living at "Widgets". + for _, e := range n { + walk(e, ownerType, ownerID, prop) + } + } + } + walk(doc, "", "", "") + return out +} + +// SitesInUnit is SitesIn over a unit's raw stored bytes. +func SitesInUnit(raw []byte) ([]Site, error) { + var doc bson.D + if err := bson.Unmarshal(raw, &doc); err != nil { + return nil, err + } + return SitesIn(doc), nil +} + +// elementIDOf reads a storage object's $ID as a UUID string. Real documents +// store it as a 16-byte binary in .NET field order; tests and a few synthetic +// documents use a plain string, which is passed through unchanged. +func elementIDOf(d bson.D) string { + switch v := lookup(d, "$ID").(type) { + case bson.Binary: + return codec.BinaryToUUID(v.Data) + case string: + return v + default: + return "" + } +} diff --git a/mdl/translations/sites_test.go b/mdl/translations/sites_test.go new file mode 100644 index 0000000000..ad0841df4e --- /dev/null +++ b/mdl/translations/sites_test.go @@ -0,0 +1,129 @@ +// SPDX-License-Identifier: Apache-2.0 + +package translations + +import ( + "testing" + + "go.mongodb.org/mongo-driver/v2/bson" + + "github.com/mendixlabs/mxcli/modelsdk/mpr" +) + +// The point of the walk: it names the site a text lives at without knowing +// anything about the document type. A widget caption and a page title are found +// by the same code, which is what the hand-written catalog extractor could not do. +func TestSitesIn_NamesTheOwnerAndPropertyOfEveryText(t *testing.T) { + pageID := "11111111-1111-1111-1111-111111111111" + btnID := "22222222-2222-2222-2222-222222222222" + + doc := bson.D{ + {Key: "$ID", Value: mpr.IDToBsonBinary(pageID)}, + {Key: "$Type", Value: "Forms$Page"}, + {Key: "Title", Value: text(tr("en_US", "Home"))}, + {Key: "Widgets", Value: bson.A{ + bson.D{ + {Key: "$ID", Value: mpr.IDToBsonBinary(btnID)}, + {Key: "$Type", Value: "Forms$ActionButton"}, + {Key: "Caption", Value: text(tr("en_US", "Save"), tr("nl_NL", "Opslaan"))}, + {Key: "Tooltip", Value: text(tr("en_US", "Save this record"))}, + }, + }}, + } + + got := SitesIn(doc) + if len(got) != 3 { + t.Fatalf("want 3 sites, got %d: %+v", len(got), got) + } + + byProp := map[string]Site{} + for _, s := range got { + byProp[s.OwnerType+"."+s.Property] = s + } + + title, ok := byProp["Forms$Page.Title"] + if !ok { + t.Fatalf("page title site missing; got %v", siteKeys(byProp)) + } + if title.Targets["en_US"] != "Home" { + t.Errorf("title en_US = %q, want %q", title.Targets["en_US"], "Home") + } + if title.ElementID != pageID { + t.Errorf("title ElementID = %q, want the page's %q", title.ElementID, pageID) + } + + // The caption is the one the hand-written builder never reached. + capt, ok := byProp["Forms$ActionButton.Caption"] + if !ok { + t.Fatalf("action button caption site missing; got %v", siteKeys(byProp)) + } + if capt.Targets["nl_NL"] != "Opslaan" { + t.Errorf("caption nl_NL = %q, want %q", capt.Targets["nl_NL"], "Opslaan") + } + if capt.ElementID != btnID { + t.Errorf("caption ElementID = %q, want the button's %q, not the page's", capt.ElementID, btnID) + } + + if _, ok := byProp["Forms$ActionButton.Tooltip"]; !ok { + t.Errorf("tooltip site missing; got %v", siteKeys(byProp)) + } +} + +// Two texts on sibling elements of the same type must stay distinguishable, or a +// consumer grouping by (document, property) folds them into one — which is the +// QUAL005 defect the ElementID exists to let a caller avoid. +func TestSitesIn_SiblingElementsOfOneTypeKeepSeparateIDs(t *testing.T) { + lowID := "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa" + highID := "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb" + + doc := bson.D{ + {Key: "$Type", Value: "Enumerations$Enumeration"}, + {Key: "Values", Value: bson.A{ + bson.D{ + {Key: "$ID", Value: mpr.IDToBsonBinary(lowID)}, + {Key: "$Type", Value: "Enumerations$EnumerationValue"}, + {Key: "Caption", Value: text(tr("en_US", "Low"), tr("nl_NL", "Laag"))}, + }, + bson.D{ + {Key: "$ID", Value: mpr.IDToBsonBinary(highID)}, + {Key: "$Type", Value: "Enumerations$EnumerationValue"}, + {Key: "Caption", Value: text(tr("en_US", "High"))}, // untranslated + }, + }}, + } + + got := SitesIn(doc) + if len(got) != 2 { + t.Fatalf("want 2 sites, got %d", len(got)) + } + if got[0].ElementID == got[1].ElementID { + t.Fatalf("sibling values share ElementID %q — a consumer cannot tell them apart", got[0].ElementID) + } + if got[0].ElementID != lowID || got[1].ElementID != highID { + t.Errorf("ElementIDs = %q, %q; want %q, %q", got[0].ElementID, got[1].ElementID, lowID, highID) + } +} + +// A text directly on the unit root has no enclosing element but is still a real +// site — dropping it would silently lose the document's own title. +func TestSitesIn_TextOnTheUnitRootIsStillASite(t *testing.T) { + doc := bson.D{ + {Key: "$Type", Value: "Forms$Page"}, + {Key: "Title", Value: text(tr("en_US", "Home"))}, + } + got := SitesIn(doc) + if len(got) != 1 { + t.Fatalf("want 1 site, got %d", len(got)) + } + if got[0].OwnerType != "Forms$Page" || got[0].Property != "Title" { + t.Errorf("site = %s.%s, want Forms$Page.Title", got[0].OwnerType, got[0].Property) + } +} + +func siteKeys(m map[string]Site) []string { + var out []string + for k := range m { + out = append(out, k) + } + return out +}