A hand-written, dependency-minimal codec for the OpenDocument Format (ODF — OASIS/ISO 26300):
.odt/.ods/.odp/.odg/.odf/.odb/.odmand their template variants, built on Zod 4 codecs.
odf.js is the ODF sibling of ooxml.js, mirroring its architecture: a lossless ZIP-of-XML core that round-trips any package byte-for-content-faithful, with typed readers layered on top. Two ODF-specific differences shape the design: ODF has no relationship mechanism (inter-part references are direct paths, with an exhaustive META-INF/manifest.xml), and ODF has no inline/direct formatting — every formatting difference must be a named "automatic style," so odf.js owns a style-interning subsystem (src/styles/) with no OOXML equivalent.
This package does not depend on ooxml.js. ooxml.js's branding and SBOM are scoped to ECMA-376/OOXML; depending on it would be the wrong signal for an OASIS-standard codec and would force a breaking ooxml.js release for every ODF-only fix. The small generic ZIP/XML/Package layer is duplicated, kept structurally identical so TypeScript's structural typing makes both packages' values interchangeable for a shared consumer like documents.js. Both depend on document-schema.js for the genuinely identical ContentDocument/LayoutDocument content model.
graph TD
schema("document-schema.js")
ooxml("ooxml.js")
odf("odf.js")
pdfcodec("pdf-codec")
mdcodec("markdown-codec")
bytecodec("byte-codec")
documents("documents.js")
mcp("document-mcp")
cli("document-cli")
schema --> ooxml
schema --> odf
schema --> pdfcodec
schema --> mdcodec
schema --> documents
ooxml --> documents
odf --> documents
pdfcodec --> documents
mdcodec --> documents
bytecodec --> pdfcodec
bytecodec --> documents
documents --> mcp
pdfcodec --> mcp
documents --> cli
odf --> cli
pdfcodec --> cli
click schema "https://github.com/ExaDev/document-schema.js" "document-schema.js"
click ooxml "https://github.com/ExaDev/ooxml.js" "ooxml.js"
click odf "https://github.com/ExaDev/odf.js" "odf.js"
click pdfcodec "https://github.com/ExaDev/pdf-codec" "pdf-codec"
click mdcodec "https://github.com/ExaDev/markdown-codec" "markdown-codec"
click bytecodec "https://github.com/ExaDev/byte-codec" "byte-codec"
click documents "https://github.com/ExaDev/documents.js" "documents.js"
click mcp "https://github.com/ExaDev/document-mcp" "document-mcp"
click cli "https://github.com/ExaDev/document-cli" "document-cli"
style odf fill:#f9a825,stroke:#333,stroke-width:3px
Under active development. Built and shipped:
- Lossless core — ZIP-of-XML primitives (
Package/XmlNode/XmlElement, XML parse/build, zip/unzip, base64, thepackageCodec/xmlCodecz.codec()pairs). - Namespaces, media types, mimetype, manifest (
src/ns.ts,src/media-type.ts,src/mimetype.ts,src/manifest.ts) — full read and write, includingMETA-INF/manifest.xmland the mimetype part's mandatory first-entry/stored/uncompressed layout. - Style interning (
src/styles/) —StyleRegistryadopts existing automatic styles, finds-or-mints onintern(), fingerprints on canonical serialized properties (neverJSON.stringify), collision-checked across all four style containers. - Shared typed primitives (
src/typed/shared/) — unit parsing, A1 cell-reference computation with repeat-count cursor advancement, colour/geometry/master-page parsing intodocument-schema.jstypes, whitespace-run decoding, the read-side style cascade, sharedreadOdfParagraph/readOdfTable, thedraw:transform/draw:ggroup-flattening geometry resolver, ansvg:d/draw:pointspath parser, andmeta.xmlreading. - Typed readers —
readOdt(wordprocessing),readOdp(presentation),readOdg(drawing: vector primitives indraw:z-index-aware paint order), andreadOds(spreadsheet: everyoffice:value-type, cell/page-anchored images, and embedded sub-documents) each resolve aPackageintodocument-schema.js's ownContentSection/ContentSlide/ContentDrawPage/ContentSheetshapes. readOdfFormula— resolves a standalone/embedded.odfformula's bare-MathMLcontent.xmlinto raw MathML nodes plus a StarMath annotation.readOdfFormulaDocumentwraps that into a real'formula'-kindContentDocument.readOdm— resolves a.odmmaster document into an ordered list of chapter references ({ name, href, filterName? }); chapters are genuinely external.odtfiles by ODF design, never cached.readOdbInventory— resolves a.odbinto connection info, table names, query definitions ({ name, command, escapeProcessing? }with real SQL text), and form/report{ name, href }pairs. A sub-document directory is named after an opaque persistent name (forms/Obj11), not the user-visible name.readOdbForm/readOdbReport— extract one sub-document's static structure, executing nothing: a form's control tree and data bindings, or a report's band stack, recursive group tree, bound fields, and computed expressions.
Not yet built: live-view editors and the .odb database-table-export subsystem. A general-purpose SQL query engine for rendering a Report against its data is deliberately not attempted — building even a bounded SQL engine means reimplementing HSQLDB's/Firebird's query semantics, a materially different undertaking from decoding their file formats, with unreviewed licensing questions. Gated on the requesting engineer's explicit sign-off.
Requires Node.js >=20 and pnpm 11.6.0 (pinned via packageManager in package.json).
pnpm installInstall as a dependency in another project:
pnpm add odf.js
# or
npm install odf.jspnpm build # turbo run _build -> tsdown (dist/: ESM + CJS + .d.ts)
pnpm typecheck # turbo run _typecheck -> tsc -p tsconfig.json && tsc -p tsconfig.node.json
pnpm lint # turbo run _lint -> eslint . --fix --cache --max-warnings 0
pnpm test # turbo run _test -> vitest run --project unit
pnpm test:workers # turbo run _test:workers -> vitest run --config vitest.workers.config.ts (the package parsing and ODF content readers run inside a real Cloudflare Workers isolate, proving they carry zero Node-only API usage)To run a single test file: pnpm vitest run src/path/to/file.test.ts.
The lossless core — the only public surface stable enough to document with real examples right now:
import { decodePackage, encodePackage } from 'odf.js';
// .odt / .ods / .odp bytes -> faithful JSON Package
const pkg = decodePackage(new Uint8Array(await file.arrayBuffer()));
// ...inspect pkg.parts...
// Package -> bytes (content-identical, mimetype-first/stored, manifest untouched)
const bytes = encodePackage(pkg);Manifest and mimetype, ODF's own package-identity mechanism (no relationships, unlike OOXML):
import { readManifest, syncManifest, setDocumentMediaType, readMimetype } from 'odf.js';
const manifest = readManifest(pkg); // { entries: [{ fullPath, mediaType }, ...] }
setDocumentMediaType(pkg, 'application/vnd.oasis.opendocument.text'); // updates mimetype + manifest root entry atomically
syncManifest(pkg); // rebuilds manifest.xml to exactly match pkg's current parts
readMimetype(pkg); // 'application/vnd.oasis.opendocument.text'Every module is also importable directly by its own subpath, without going through the barrel:
import { parseOdfLength } from 'odf.js/typed/shared/units';
parseOdfLength('2.5cm'); // 70.86614173228347Any src/**/*.ts module (excluding tests and test-support/ fixtures) resolves at its path relative to src/ — src/manifest.ts as odf.js/manifest, src/typed/odt/read.ts as odf.js/typed/odt/read, and so on.
Layered from a lossless core outward, mirroring ooxml.js:
src/model/—Package/XmlNode/XmlElement: a duplicate-by-design copy ofooxml.js's equivalent.src/xml/— XML parse/build (fast-xml-parser), production element/text-node construction, entity encoding, and tree-query helpers.src/image/—sniffImageFormat: a PNG/JPEG magic-byte sniffer consumed bysrc/manifest.tsandsrc/typed/draw/shapes.ts.src/zip.ts— takes ordered[path, entry]tuples, not aRecord, so the mimetype-first/stored/uncompressed requirement doesn't depend on insertion order surviving a Zod round trip.src/package-io/—write.tshoistsmimetypefirst (stored) andMETA-INF/manifest.xmlsecond if present; never fabricates either as a side effect.src/manifest.ts— full manifest read/write; the manifest is ODF's one mandatory part, unlikeooxml.js's read-only OPC-relationship stance.src/styles/—properties.ts/serialize.ts(canonical property-bag ↔ XML attributes),registry.ts(StyleRegistry, the mandatory style-interning layer),span.ts(character-rangetext:spanwrapping).src/typed/shared/— ODF-specific typed primitives every reader builds on (units, A1 cursors, colour/geometry, whitespace runs, style cascade, shared paragraph/table readers, transform/path parsing, metadata).src/typed/odt/,odp/,odg/,ods/— thereadOdt/readOdp/readOdg/readOdsreaders.src/typed/draw/— the shareddraw:frame/draw:g/vector shape vocabulary (shapes.ts), plusembedded.ts(readDrawObjectReference,readDrawImageBlock).src/typed/formula/,odm/—readOdfFormula/readOdfFormulaDocumentandreadOdm.src/typed/odb/—readOdbInventory,readOdbForm/readOdbReport,resolveOdbComponent,subDocumentPackage.
- Zod-first schema/type/guard, matching
ooxml.js/document-schema.js: every model type is inferred from its Zod schema, never hand-written. - Recursive types use a hand-written structural guard, not
z.lazy(collapses tounknownin the pinned Zod version). - No type assertions anywhere —
assertionStyle: 'never',noInlineConfig: true. - Ground truth over memory for every ODF spec fact — namespace URIs, media types, and attribute names are verified against the OASIS spec or real LibreOffice output, never assumed from an OOXML analogue (see Gotchas).
- Several ODF namespace URIs are not what you'd guess from the prefix.
draw:is...drawing:1.0,number:is...datastyle:1.0,fo:/svg:/smil:are*-compatible:1.0. Seesrc/ns.ts. .odb's media type isapplication/vnd.oasis.opendocument.base, not...database.dc:creatorrecords whoever most recently saved the document, not the author — the original author ismeta:initial-creator.meta:keywordappears once per keyword, unlike OOXML's single comma-separatedcp:keywords.table:number-columns-repeated/-rows-repeatedmust be cursor-advanced, never materialized — real sheets have trailing repeat counts over a million.- ODF cells carry no explicit cell-reference attribute (unlike xlsx's
r="B7") —typed/shared/a1.tscomputes references from a running cursor. - A rotated
draw:rect/ellipse/path/custom-shapereads its ownrotationDegvia the sameresolveOdfShapeGeometrymachinerydraw:frameuses, composing any enclosingdraw:grotation. - Every
ContentShape/ContentVectorcarries a resolvedpaintOrderso true relative paint order survives across the independently-orderedshapes/vectorsarrays. svg:fill-ruleanddraw:strokemap ontoContentVector.fillRule/ContentStroke.style. A dotted pattern and"double"stroke have no ODF vector-stroke counterpart and remain unread.readOds/readTableCellresolve cellbackground/borders/alignment/verticalAlignmentfrom the real style cascade. An explicitfo:border-*of"none"/"hidden"clears an inherited edge.readOdsreads sheet-anchored drawings — cell-anchoreddraw:frames (coordinates relative to the cell) and page-anchored ones (intable:shapes). A sheet cannot carry a floating text box, bare vector, or embedded chart; each is skipped.readDrawObjectReferenceresolves a frame's embedded sub-document kind from its owncontent.xml, not the manifest. Adraw:objectmust be checked before the frame's preview image, since an embedded-object frame also carries a previewdraw:image.- An embedded Math object in a spreadsheet cell reads as
objectKind: 'formula'— itscontent.xmlroot is the MathML root, soreadDrawObjectReferencefalls back tofindMathRootand dispatches toreadOdfFormulaDocument. - A
draw:frame's alternative text (svg:title, falling back tosvg:desc) reads intoContentImageBlock.altText. readOdbInventory'squeriescarry realdb:commandSQL text, not just names — a breaking rename fromstring[]toOdbQueryInfo[]..odbForm/Report structure extraction is real (readOdbForm/readOdbReport), grounded in a genuine fixture. A SQL/rpt:rendering engine to execute a query or evaluate report totals is deliberately not attempted — see the Status section. Even a fully bounded SQL engine would not suffice to render a report: grouping breaks (rpt:HASCHANGED), prefix functions (rpt:LEFT), and running totals (rpt:SUM) are evaluated by Report Builder's ownrpt:formula language, not by SQL.
.github/workflows/ci.yml runs commitlint, lint, typecheck, unit, and smoke tests on every push/PR. On a push to main where all pass, release.config.ts drives semantic-release: version bump from commit history, CHANGELOG.md/package.json committed back, GitHub Release cut, and npm publish via OIDC trusted publishing (no NPM_TOKEN). Once a release publishes (detected by diffing package.json's version): a sibling-released event dispatches to documents.js/document-cli, the build republishes under @exadev/odf.js to GitHub Packages, and an SPDX SBOM plus build-provenance attestation are signed against the tarball.
Conventional Commits (feat:, fix:, test:, chore:, …), enforced by commitlint via a husky commit-msg hook and a CI job. A pre-commit hook runs lint-staged (eslint --fix on staged *.ts); pre-push runs the test suite. Single main branch, no open PR workflow.
- ooxml.js — the OOXML sibling; architecturally mirrored, deliberately not depended on.
- document-schema.js — the canonical
ContentDocument/LayoutDocumentschema both packages depend on. - documents.js — the downstream consumer; its
readOdtContent/readOdpContent/readOdsContent/readOdgContentare thin adapters over this package's readers.
MIT