LibRed: ACE-parity SQL semantics, standard SQL extensions and file-format write fixes; Jet Math translation fixes (11.0.0-alpha.3) - #301
Merged
Conversation
DropTable freed the index roots, the TDEF and whatever the table's data-page map named -- two pages for a 60-row memo table. A Memo/OLE column keeps its LVAL pages in a per-column usage map, pointed at from the TDEF by column id, which that path never looked at. So for a memo-heavy table the drop returned almost nothing, and the content of the table stayed allocated forever. Measured against ACE on the same file: dropping it through ACE returned 123 pages, through LibRed 2. Refilling afterwards then grew the file from 167 pages to 286, because ACE reuses exactly what the global free map offers it and nothing else. Now matched, in three parts. The per-column owned maps are walked and their pages freed, clearing each page's bit on the way out as releasing a single value does. The map records are retired from their holder page with the existing ReclaimRow, and the holder itself goes back once no live row is left on it -- the exclusivity check matters, since a holder can carry records for several columns or tables and releasing a shared one would hand away a live page. And the released TDEF is marked 0x08, which is what the unexplained 0x08 pages in real files turn out to be: exactly one byte of the 4,096 changes across an ACE drop. Both engines now free the same 123 pages, ACE reuses the space rather than growing the file, and 163 of the 167 pages are left byte-identical. Of the four that differ, one is the opening user's commit slot, which moves for any write at all; the other three are catalog index roots holding identical entries in identical order, where LibRed compacts a leaf harder than ACE does. That is index maintenance rather than anything to do with dropping, and is recorded in the spec rather than changed -- matching it would mean deliberately compacting less. The TDEF is read through the chain reader: a wide table's definition spans continuation pages and parsing only the first throws on the declared length. The ACE suites caught that, ALTER COLUMN reaching it through RewriteColumn; the cross-platform ones did not, since the failing path needs a table wide enough. Also documents page type 0x09 -- a released, emptied page that no operation reachable through the SQL surface produces. The eliminations are the useful part and are written down; it is left unnamed in PageType rather than guessed at. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… cap ACE pads an all-fixed row's fixed region out to two bytes, so the shortest record it writes is five. It is a floor and not an alignment -- a three-BYTE table keeps its odd 3-byte region -- and a row carrying a variable trailer is exempt, keeping a region of 0 or 1. LibRed sized the region from the table definition alone, one or two bytes short of ACE for the same table. That is not cosmetic. An all-Boolean table of eight columns or fewer encoded to a 3-byte record, and ACE reads every Boolean in such a row as False. Measured both ways round -- ACE's own table filled by LibRed read False, LibRed's table filled by ACE read True -- which places the fault in the record rather than the definition. The TDEF keeps the true unpadded length either way, so the rounding belongs at row-write time and nowhere else. The rows-per-page ceiling drops from 256 to 255 for a related reason. The row pointer can name 256 slots, but ACE writes at most 255 and, reading, takes the full 16-bit count and then caps at 256 slots per page regardless -- so a 256-row page is a shape Access never produces and cannot fully read. Its cap is not about space: a page ACE had filled to 255 still held 2,297 bytes free, and it drops such a page from the table's free-pages map as if it were full, which LibRed already did by superseding the tail page. With both, a 900-row fill of a one-BYTE table lands on the page layout, the row bytes and the free-space figures ACE's own fill produces. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The four bytes at data-page header 0x08 -- Jet 4 only, which page-01 recorded as "purpose unknown, zero on every page observed" and which mdbtools describes the same way -- are a version tag on a long-value chain. They are zero everywhere except the FIRST page of a multi-page chain, where they repeat the pointing descriptor's own bytes at 0x08, until now documented as reserved. ACE enforces that the two agree. Measured by patching an ACE-written file behind its back: setting both copies to DEADBEEF reads back fine, setting both to zero reads back fine, and changing either one alone -- to any value, zero included -- makes ACE refuse to materialise the value. Later chunk pages are not bound; DEADBEEF in all of them with the descriptor and first page untouched changes nothing ACE notices. So the value is arbitrary and only the agreement is checked, and only at the chain's entry point, every chunk after it being reached from a page already validated. ACE stamps GetTickCount() and mints a fresh one per write of the chain: rewriting the value restamps both copies, updating another column of the same row leaves them alone, and shrinking to the single-page form drops both to zero. The check fires when the long value is materialised, not when the row is read -- ACE will update another column of a row whose stamp disagrees and refuse the moment anything reads the memo. It reports a mismatch as "you and another user are attempting to change the same data at the same time", which is the point of it: a chain's pages can be freed and reused, and the descriptor is the only way in that a stale pointer can arrive by. LibRed now writes the same tag in both places and refuses a chain whose entry page disagrees. Like the database creation date the value does not reproduce between runs, which is accepted. LibRed is single-writer and cannot produce the interleaving ACE guards against, but it can be handed a file another engine wrote -- and this is the check multi-user concurrency will need. Files written by earlier LibRed carry zero in both places, which agrees, and keep reading. Also corrects the long-value form table, whose single-page range read "66 ... 3816": a memo steps two bytes per character, so 65 had never been asked. Swept through an OLE column, which takes any byte length: 64 inline / 65 on a page, 3816 one page / 3817 chained, LibRed choosing the same form as ACE at all 337 sizes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…hrough a rebuild
Two faults in the ALTER COLUMN paths, both reached only by shapes nothing covered.
RewriteColumn rebuilt each untouched column's spec without its CalculatedExpression/CalculatedResultType. Those live in the LvProp blob rather than the 25-byte descriptor, so the RawDescriptor passthrough that carries SystemFlags, compression and the rest cannot reach them, and the rebuilt column came back plain. The rebuild also re-inserted each row's cached calculated value, which an insert refuses outright ("field not updateable"), so with the expression restored the rebuild threw on every row -- the value is a cache and is recomputed from its expression instead. Reached by any full rebuild, a Memo retype for instance, on a table holding a calculated column.
RecycleOwnedMapRow wrote step (2) of ACE's owned-map recycle as a MOVE of the appended record into the old row's freed slot. ACE re-lays the page instead: the old record is reclaimed, every later row keeps its number while its record slides up, and the fresh map takes the position freed at the end of the live region. The two produce identical bytes whenever the recycled row is the LAST on its holder page -- the only kind an ACE-built schema gives, since a long-value column declared in CREATE TABLE takes its map rows before the index's -- and diverge as soon as a Memo or OLE column is added AFTER an index. The move then points a slot back up the page, which no reader can walk (a row's extent runs to where the previous slot begins), and AlterColumnTypeInPlace scans the table through that map immediately afterwards, so the ALTER failed outright.
Step (1) is unchanged and still verbatim: the appended row is where the new root's bit is set, and the bytes it abandons stay in free space. A whole-file diff against ACE turns on a single byte of them.
Both halves hide from a different measurement, which is how the description came to be half-right. The abandoned copy lies below the lowest live record, inside the region free space already covers, so slot offsets and free-space arithmetic read the page as though it were not there. The re-lay is invisible to offsets alone, a moved record and a slid record occupying the same places; only identifying records by content separates them. OwnedMapRecycleAccessTests restores the whole-file byte comparison for this area -- ACE builds the schema, ACE and LibRed each alter a copy, every byte of every page must match bar page 0 and the MSysObjects timestamp -- across four shapes including the memo-after-index one that nothing covered. It is the only instrument that sees both halves.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Page type 0x09 is an emptied packed long-value page. Values in the single-page form (<= 3816 bytes) are packed several to an LVAL page; each delete retires that value's row to a 0-length deleted+overflow tombstone and re-lays the page, the survivors packing from the page end in slot order, and once nothing live is left the page is stamped 0x09, cleared from the column's owned and free maps, and freed. A chained value never produces one -- those own their pages outright and go back at 0x01. That is why the type went unexplained: every sweep that ruled out DELETE, UPDATE, DROP COLUMN and DROP TABLE used memos large enough to chain, and the remaining hypothesis on record was that only the Access UI could produce one. It needs no UI at all, and reproduces identically through OLE DB SQL, DAO SQL and a DAO recordset. Measured against ACE, 12 rows of 400-character memos across three pages, deleting 4 and then all 12: start 0x01 n=5 free=72 [3296,2496,1696,896,96] 4 gone 0x01 n=5 free=3272 [4096DO,4096DO,4096DO,4096DO,3296] all 0x09 n=5 free=4072 [4096DO x5] Finding it exposed a larger gap: FreeLongValue returned early for the single-page form, so LibRed never reclaimed a packed long value at all. Every deleted short memo leaked its row, and its page with it -- invisible, because every read stayed correct and the file merely grew. Both are fixed together: ReleasePackedValue reproduces ACE's page exactly, and a surviving page goes back into the column's free map, having room again. PackedLongValueReleaseAccessTests compares every byte of every page against ACE's own delete across the partial, full, multi-page and chained cases. The only pages it skips are index pages, whose post-removal compaction difference is recorded and accepted in page-03-04 10.4a. Docs: 0x08 and 0x09 now have their own files, matching the one-per-page-type convention, and both are listed in the README page-type table, the file list and the appendix -- neither had been. The material that had accumulated in page-05 moves there rather than being duplicated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An outbound differential sweep - LibRed writes, ACE judges - found four places where LibRed produced something the real engine would not take. ACE 15 (version byte 0x04) can no longer be created. ACE refuses to open a file carrying it, an empty one included, and restamping 0x14 to 0x03 opens the same bytes. Reading one still works, so the guard sits on CreateEmpty alone. The spec had called the byte "reserved" by inference from absence; it is now measured as actively rejected. A DECIMAL wider than its column's declared precision is refused, as ACE does on INSERT, UPDATE, INSERT...SELECT and an ALTER that narrows the declaration. The 17-byte payload cannot enforce the declaration itself, and LibRed reading its own file back agreed with itself, so only a second engine could see it. Excess decimal scale truncates toward zero instead of rounding half-to-even, matching ACE. IndexKeyEncoder.EncodeFixedPoint changes with JetTypeCodec.EncodeNumeric: the key is the same unscaled integer the row stores, so quantising one differently files a value under a number its row lacks. Two conversions that leaked raw InvalidCastException now refuse with a typed exception naming the column - JetTypeCodec's byte[] fallback and TableCreator.ConvertValue on an ALTER COLUMN rebuild. Precision and scale are validated on the declaration too, and an unspecified precision resolves to ACE's own default of 18 rather than reaching disk as 0, which is a column its OLE DB reader cannot materialise at all. The sweep and its Core-API arm land as probes over a shared AceValidityLadder: open, enumerate schema, read every row, ACE writes, LibRed re-reads. A failed typed read is re-asked of the engine as text before the file is called bad, because ACE's provider breaks on values its own engine handles - DATETIME2 outright, and decimals wider than their declaration. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Page 0's minor byte at 0x15 is 0x01 only on a database created in the 2010 format. ACE writes 0x00 there whenever it raises the version, whatever the target: BIGINT and DATETIME2 raising a created 2010 file clear it, and a calculated column raising a 2007 file onto 0x03 leaves it 0x00 rather than stamping the 0x01 a created 2010 file carries. RaiseFormatVersion moved only 0x14, so LibRed left a 2010 file at (0x05, 0x01), a pair ACE never writes. It now clears 0x15 in the same page-0 write, so the reset still joins the statement's transaction. The raise tests start ACE and LibRed from copies of one base file and diff page 0 outside the commit-byte table, replacing the test that pinned the old behaviour. The spec also records 0x6A as the creating engine's build number rather than a fixed constant: ACE stamps 4518 whatever its version, while files created by Jet 4 carry their msjet40.dll build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The spec had accumulated the record of how each fact was found alongside the fact itself: test and probe names, fixture file names, corpus counts, dated sweeps, and "an earlier revision said" corrections. That material described the research, not the bytes, so it is gone. Verified/inferred status stays, and a correction survives only where the wrong reading is one a reader would reach on their own, restated as a rule. The pass also resolved contradictions it surfaced: - A column's 0x0B-0x0E bytes carry the database's LCID, not a constant 0x0409, and 0x0D is the sort id rather than half of a version word. - A long value is chained from 3817 bytes, not from one 4076-byte row. - ACE writes at most 255 rows to a data page. - Ligatures are encoded by decomposition, and the locale tailorings include digraphs and contractions, not only single characters. - The primary-key usage-map conversion point is marked not reconciled with ACE rather than verified. - The 510-byte entry limit moves out of the middle of the key-encoding list to its own section, 10.4b, leaving 10.5 for insertion and splitting. - Page-1 allocator notes stranded under the released long-value heading return to the global free-pages map section. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nothing constructs or calls any of these: - JetDateTimeExpressionVisitor, an obsolete compile-time visitor neither provider registers. JetDateTimeRangeConverter's comment named it. - RowInserter.SlotBytes and PageCache.EvictFrom. - AlterColumnTypeInPlaceTdef on TableCreator and JetDatabase, a seam for a byte-diff probe deleted long ago. On its own it burns a column id and moves the fixed slot without re-laying rows or rebuilding indexes, so any caller with data corrupts the table. AlterColumnTypeInPlace never used it. AlterColumnTypeInPlace's summary also described the all-fixed, unindexed subset it started as; it now states the shapes it handles. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI installs the ACE 2016 redistributable, which refuses any 0x05 or 0x06 file as needing a newer Access. The write-validity sweep assumed every format LibRed creates would open, so its empty-database theory failed at both versions and every generated workload drawn at them failed on its first statement, saying nothing about LibRed. The theory now skips a version newer than the engine opens, and the sweep clamps a drawn version down to it. Clamping rather than dropping keeps the random stream, and so each seed's statements, the same on every engine. The newest version comes from the same BIGINT/DATETIME2 probe the rest of the suite gates on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
LibRed.Engine.AccessTests puts every class that drives ACE in one xunit collection, because ACE faults when two of them run at once. The ACE half of Core relied on xunit.runner.json turning collection parallelism off instead. Core now does what Engine does: the same AceCollection, carried by every class in the project, and no runner configuration. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Page 0's 0x18 and 0x1C are [row][page] pointers to the two global usage maps: the free-pages map and a released-pages map, page 1 rows 0 and 1 in every file ACE writes. ACE follows both wherever they point and never allocates a page set in the released map. LibRed now finds both maps through the pointers, skips released pages when allocating, and refuses a writable open whose pointers name the same record, a page outside the file or anything but a usage map. A read-only open still reads such a file, as ACE does. ACE holds most freed pages until the connection closes: a deleted row's long values, a dropped index (DROP INDEX, DROP CONSTRAINT, an ALTER COLUMN rebuild) and a dropped table. Only the long value an UPDATE replaces is free at once. PageAllocator.Release keeps such a page on the handle, staged with the open transaction, kept on commit and dropped on rollback or a savepoint rollback. Closing a writable JetDatabase returns those pages, and any already in the released map, to the free map and clears the released map. Before that it sizes the released map the way ACE's close does: lengthen the inline record from its start page, or move its window to the lowest released page, or convert it to reference form with a bitmap page for each range holding a released page. The inline record's leftover bytes stay on the page as ACE leaves them. DROP TABLE now frees everything the table owns. Index pages beyond the root and each index's map record were missed, which also kept the map holder page live. A multi-page definition's continuation pages and a reference-form map's bitmap pages were missed too. The map records are retired in ACE's order: long-value maps, then index maps, then the table's own maps. Each retirement slides the records below it, so the order shows in the bytes left behind. Whole-file diffs against ACE drops of memo, indexed, 255-column and reference-map tables, across sessions, are byte-identical outside page 0 and the catalog's own pages. The case with released pages in the first and third bitmap ranges builds a 272 MB file and is marked Explicit. The spec records the pointers, the release-at-close rules, the released map's sizing and the drop's retire order. The README's DROP TABLE leak gap is closed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
LibRed's extended mode shared JetSqlTranslatingExpressionVisitor with the Jet provider, so Math.Max, Math.Min, EF.Functions.Greatest/Least and a Max() or Min() over an inline collection all became nested CASE comparisons, the only form ACE can run. Extended mode now has its own copy of the visitor, which generates GREATEST and LEAST instead. LibRedSqlTranslatingExpressionVisitorFactory picks it by SQL mode, as the query SQL generator factory already does, so compatible mode keeps the shared visitor and its SQL. The engine gains both functions. Like COALESCE they are standard SQL that ACE lacks. NULL arguments are ignored and the result is NULL only when every argument is, which is how SQL Server and PostgreSQL define them and what EF Core's translations assume. Arguments compare as the comparison operators compare, and the result declares the type the arguments unify to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ACE accepts CLUSTERED or NONCLUSTERED straight after PRIMARY KEY or UNIQUE in any constraint clause: a CREATE TABLE column or table constraint, and ALTER TABLE's ADD CONSTRAINT, ADD COLUMN and ALTER COLUMN. It stores nothing for either word. The file is byte-identical without it, and DAO reports Clustered = False even for an index created with Clustered = True. ACE rejects the word everywhere else: after FOREIGN KEY, between PRIMARY and KEY, on a bare column and in CREATE INDEX. LibRed's grammar now takes the optional word in the same two places and drops it. Both words are reserved, as they are in ACE, so an unbracketed table, column or alias of either name is a syntax error. The parser is regenerated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…IDENTITY REFERENCES T with no column list, in a column constraint, a FOREIGN KEY clause or ADD COLUMN, references T's primary key, pairing the child columns with the key's columns in order. As in ACE, a parent without a primary key refuses the relationship (a unique index does not count), a column count that differs is refused, and a table referencing itself in CREATE TABLE needs its primary key declared earlier in the statement. The resolution lives in Core's TableCreator beside the parent lookup, so CREATE TABLE and ALTER TABLE share one path and one "cannot find table" message, and the Engine builds both relationship specs through one method. Every relationship now checks that each child column has its parent column's storage type. Lengths are ignored, so TEXT and CHAR pair at any length, DECIMAL at any precision and scale, and BINARY with VARBINARY. An AutoNumber counts as a Long on either side. IDENTITY [(seed [, increment])] is a column attribute, as ACE parses it: after the type, NULL/NOT NULL or another IDENTITY, and before DEFAULT or a constraint. It makes a Long column an AutoNumber with its own seed and increment, 1 when omitted, and is ignored on any other type. It still works as a type on its own. The multi-word INT IDENTITY type names are gone. A NOT NULL AutoNumber now writes the Required property, as ACE does. ADD COLUMN applies its PRIMARY KEY and UNIQUE constraints to the new column. A second primary key is refused, on every path that adds one. A primary key or DISALLOW NULL index over existing rows with a NULL key is refused, in the same scan that checks unique keys. An AutoNumber added to a table with rows numbers them 1 to n in table order. The default 1/1 counter carries on after them, and any other seed or increment restarts at its seed. The spec records the Required rule, the numbering of an added AutoNumber, the single primary key, required indexes over existing rows, the absent clustered flag, relationship type matching and REFERENCES without columns. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each operator was probed against ACE and corrected, keeping LibRed's own result types: - + and &: + concatenates only two texts and otherwise reads text as a number; & is Null only when both sides are; values are written as CStr writes them. - - * / \ MOD ^: text is read as a number, overflow and division by zero are errors, \ and MOD round half to even, and a Decimal result keeps the places of its operands, cut as it goes into the result column. Double to Decimal uses JetDecimalConverter, now shared with EFCore.Jet.Data. - Comparisons: text against a number reads as a number, dates compare by serial, and = or <> against a True/False literal tests truth. - NOT AND OR, plus new XOR EQV IMP: a value is False when it reads as 0. - LIKE: a new matcher with ANSI-92 wildcards, bracket lists, ss/ae expansion and lazily reported invalid patterns; NULL LIKE '%' is False. - BETWEEN is its own node, with bounds in either order and a literal-bound index seek; its lower bound no longer swallows a following AND. - IN skips Null items; unary + is added, unary minus reads text and dates; BAND BOR BXOR BNOT work on the operands' bits, as ACE does for 16-bit values. - Precedence follows VBA: \ and MOD have their own levels, & sits below + and -, BNOT sits with NOT, each bitwise operator with its logical one, and a minus written against a number is part of it. '--' followed by a number is negation rather than a comment. The EF translators no longer emit an ESCAPE clause for a single wildcard character, which neither ACE nor LibRed accepts; they bracket it instead. Expression walkers now share ExpressionTree.Operands/MapOperands, so a new node type is added once, and every walker now descends into CASE. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each function was probed against ACE and corrected, keeping LibRed's own result types. Arguments are read as the conversion functions read them (text as a number or date, a date as its serial, True as -1); a Null argument gives Null where ACE raises an error; out-of-range arguments are an invalid procedure call and values past a type an overflow. - Conversion and numeric: CInt/CLng/CByte/CCur/CSng/CDbl/CBool/CDate, Fix, Int, Abs, Sgn, Round on the value's decimal form, and the trigonometric functions with ACE's domain errors. - Text: Len, Mid, InStr, InStrRev, Replace, StrComp and friends compare in the database sort order (JetTextComparer); Asc/Chr work in the ANSI code page; StrConv handles every mode ACE accepts, with an LCID. - Dates: text is read as OLE Automation reads it (VbaDateText); DateAdd, DateDiff, DatePart, Weekday, WeekdayName and DateSerial take the first day of the week and first week of the year, carry out-of-range parts, and refuse dates outside 100-9999. A #time# literal is on 1899-12-30. - Str, Val, Hex and Oct write and read numbers as VBA does. - IIf rounds its condition; Choose truncates its index; Choose, Switch and IsError evaluate every argument; IsNumeric reads text as + does; TypeName/VarType tell Currency from Decimal and report Int64 as LongLong; Partition, RGB and QBColor validate their arguments. - Format is a new engine for named, number, date and text formats; FormatNumber, FormatCurrency, FormatPercent and FormatDateTime follow the regional patterns. - Financial functions follow the VBA runtime's algorithms, to the bit. - Aggregates read text, dates and Booleans as numbers; Var/StDev use ACE's formula; Min/Max treat empty text as ACE does. - A Single divided by Singles, Integers or Booleans stays a Single. - x IN (..., NULL) is Null on a miss, as the standard has it. JetDecimalConverter is now compiled into LibRed.Core and gains ToDecimal, which every double-to-decimal conversion in LibRed uses, so writes and index keys no longer pick up .NET 11's exact-binary expansion. Calculated columns accept single-quoted text and read a memo's text rather than its long-value descriptor. The operator and function test classes share one database per class. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
With JetConfiguration.UseConnectionPooling on, Open() looks the pool up by the connection string it rebuilt from the user's, while Close() put the inner connection back under the user's own string. Whenever the two differed the lookup never hit, so every Open() created a new native connection and the pool held every one of them open. Jet 4.0 allows 64 open sessions per process, so the 65th query failed with "Unspecified error" from OleDbConnectionInternal's constructor. 10.0.1 made this near-certain: its rebuild lower-cases keys such as "Jet OLEDB:Database Locking Mode", so a string that earlier versions rebuilt unchanged no longer matches. Close() now returns the connection under the string Open() used. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Newer ICU data, which .NET uses on Linux and macOS, puts a narrow no-break space (U+202F) before AM/PM in the time pattern. Windows' regional settings, and so ACE, use a plain space; LibRed now writes that everywhere. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ity tests A bare True/False could not tell a real difference from a statement one engine cannot run at all. A failure now names the statement and the error for each side. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Task.Run waits for the thread pool, which the parallel suite keeps busy; on the ARM runner that wait passed the 2 s timeout. The tests now start dedicated threads and wait longer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The test built its expectation from the culture's time pattern, which newer ICU data gives a narrow no-break space. LibRed writes a plain space on every platform, as ACE does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI's ACE 2016 redistributable refuses the type as a syntax error, which says nothing about IDENTITY. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…sults as ACE does A CASE, IIF, COALESCE, GREATEST or LEAST column and a set operation's column now declare one type and convert each value to it, so the type a reader reports is the type it returns. - CASE, IIF, COALESCE, GREATEST and LEAST widen numbers on one ladder: the wider whole number, a Single with a Byte or an Integer, a Double for a Single with anything wider or with a Double, a Decimal with a whole number, and a Double for Currency with a Large Number. IIF no longer declares nothing when its branches differ. - UNION, INTERSECT and EXCEPT type each column from both queries, as ACE does: the same numeric ladder with a Boolean as the Integer -1 or 0; a bare NULL takes the other query's type; a GUID or binary value with anything else makes a binary column of each value's bytes; any other mix is text. Values are converted before rows are compared. - BNOT of a Boolean or an Integer, and BAND, BOR and BXOR of two of them, give an Integer; anything else gives a Long, or an Int64. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each form was probed against ACE and matched:
- Every SET reads the joined row as it was before any of them, so
SET x = y, y = x swaps; a column set twice is refused ("Duplicate
output destination"). An unqualified SET names the one table with
that column, and an ambiguous one is refused.
- UPDATE and DELETE take a comma list of tables, as FROM does.
- A SET on the side of an outer join with no matching row adds a row
there, with the SET values and the table's defaults, one per joined
row. A DELETE skips such rows but counts them.
- A join onto a bracketed group is laid out as ACE accepts it: an inner
join onto a group whose ON reads its leading inner tables, and a LEFT
join onto a table followed by LEFT joins, null-extended as a whole.
- A derived table over one table or a join, filtered, ordered and cut
by TOP or OFFSET, is written through to its tables' rows, under the
names its projection gives the columns, anywhere a table can be. A
DELETE through a join is refused, as in ACE.
INSERT's row checks move into InsertNewRow, which UPDATE now shares.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A whole number written as an ORDER BY item names the output column at that position, through a star, a set operation or a grouping; any other constant still sorts nothing, and a position that names no column is an error. The planner replaces the position with the projected expression, or, where the sorted rows are already the output, with a new OutputColumnPosition read from the row. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A number written with a decimal point takes its written value, and text is read straight into a Decimal, instead of both going through a Double and keeping 15 significant digits. CCur('12345678901234.5678') now keeps every place, as ACE does, and CDec keeps all of a long literal or text.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The aggregates take OVER (…) over the standard's default frame: with an ORDER BY, the partition's rows up to the current row and its peers; without one, the whole partition. Their values and types are the grouped aggregates', because both now run through RunningAggregate, which replaces the separate SUM, statistic and MIN/MAX helpers. A windowed DISTINCT is refused, and only COUNT takes *. NTILE, PERCENT_RANK, CUME_DIST, LAG and LEAD follow the standard. LAG and LEAD take each row's own offset and default, and share a type between the value and the default as CASE does. A window function's type may now depend on any of its arguments, and every value it returns is converted to that type. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…egates FIRST_VALUE, LAST_VALUE and NTH_VALUE read a row of the frame; NTH_VALUE counts from either end with FROM FIRST or FROM LAST, and they, LAG and LEAD take RESPECT NULLS or IGNORE NULLS. Access's First and Last take OVER as FIRST_VALUE and LAST_VALUE. A window takes the standard's frame clause: ROWS, RANGE or GROUPS, with any bounds and an EXCLUDE. A RANGE offset measures a number or date key, and a Null key has only its peers. Offsets are evaluated per row. The default frame is the same code path, and a frame that only grows at one end is aggregated row by row. The words these clauses add are not reserved, so a column named Range or a table named Last still works. PERCENTILE_CONT, PERCENTILE_DISC and LISTAGG take WITHIN GROUP, grouped or over a window. A windowed aggregate may be DISTINCT, and every aggregate takes FILTER (WHERE ...). The standard's STDDEV_SAMP, STDDEV_POP, VAR_SAMP and VAR_POP are the Access statistics, and CORR, COVAR_POP, COVAR_SAMP and the REGR_ functions read (y, x) pairs, their sums kept by the Youngs-Cramer update. RunningAggregate is now the one list of the aggregates it computes, and each is a window function too. A window over a grouped query runs over the groups HAVING keeps, so RANK() OVER (ORDER BY SUM(x)) and SUM(SUM(x)) OVER () rank and total groups; a window whose argument holds an aggregate makes the query grouped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Trim, LTrim and RTrim strip U+3000 as well as the space, in any mixture, and nothing else: not a tab, CR, LF, no-break space, the other Unicode spaces or a zero-width one (verified vs ACE). A trailing U+3000 still counts in a comparison, as it does in ACE. DATALENGTH, SQL Server's, gives the bytes a value takes as Access stores it: text two per character, binary its length (a fixed column's padded width), and each number, date and GUID its stored size — Decimal 17, Currency 8. A Boolean, a bit of the null bitmap on disk, counts 1 as a SQL Server bit does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e2 parameters Dates compared by their OLE Automation serial alone, which holds whole milliseconds, so DATETIME2 values within one millisecond compared equal and sorted as ties. Two dates with the same serial are now settled by their ticks, later first below the epoch as the serial runs. A parameter typed DbType.DateTime2 is no longer truncated to the millisecond, as SqlClient passes one through; a DateTime value alone still infers DbType.DateTime and is truncated, so a Date/Time column's WHERE d = @p keeps matching. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The DAO probes release nothing they create, and DAO is apartment-threaded, so each Workspace, Database, TableDef and Field was torn down inside ACE on a COM-created thread whenever a GC happened to run - usually in the middle of a later test that was itself inside ACE. ACE faults with two threads in it: on CI that showed as RPC_E_SERVERFAULT from a DAO call and then a hang in the next ACE test, in a different test each run. AceTestDatabase.ReleaseAbandonedComObjects runs the finalizers at points where nothing is in ACE: after every test in LibRed.Core.AccessTests, on opening an ACE connection, and before creating a DAO engine, which every probe now does through AceTestDatabase.CreateDaoEngine. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Remove the DAO probes whose findings now live in DatabaseCreator and the format docs: the complex-system-table and page-layout dumps, the database-creation probe, and AceDdlOnLibRedDatabaseProbeTest's six isolating probes. Its regression guard stays. CollationSurveyProbeTests, the generator behind part of Collation.cs, and DaoLocaleCollationProbeTest, the evidence Collation.cs cites, run only with LIBRED_ACE_SURVEYS=1; an empty survey batch now skips before it touches DAO. AceDateTime2UpgradeTests starts from a Northwind copy instead of a DAO-created file, and DateTime2LocaleAccessTests asks whether ACE takes DATETIME2 before compacting through DAO. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The millisecond interval counts in Int64, but the column was declared a Long Integer like every other interval. A written 'ms' interval now declares Int64; one read from a parameter or column still declares the Long Integer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…past a Long A Currency result is a Currency at every step, as in ACE: +, - and * round it to four places, half to even, and overflow past the Currency range, as CCur does. A whole-number literal too big for a Long counts as a Decimal of no places when working out a result's type, as ACE reads it, so Currency * 864000000000 is not a Currency and 864000000000 / 7 has no places; its value stays an Int64. MOD and \ work in Int64 and need only the result to fit its type, Int32 unless an operand is an Int64. A remainder always fits, so a Double, Decimal or Currency past a Long MOD a Long now answers where ACE overflows; a quotient past a Long is still an overflow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CLngLng reads its argument as CLng does - text as a number, a date as its serial, True as -1, half to even - into an Int64, and past an Int64 it is an overflow. Text is read exactly rather than through a Double, so the ends of an Int64 survive. ACE has no such function, so it is a LibRed extension. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…eter A parameter is typed by the value bound to it, as a literal is, so an expression over one declares its result instead of leaving the reader to guess from the first row, which a Null there made Object. A date less a TimeSpan or TimeOnly parameter came back as a day count: the ADO layer had turned the span into a date, and a date less a date is a Double. The span now reaches the engine as itself. ParameterBag still resolves it to the time on the 1899-12-30 epoch wherever it is read - saved, compared, passed to a function - but remembers it was a span, and a date plus or less a span expression moves by it to the tick, declared a date. A span is a span parameter, negated, added to or less another, times or divided by a number, or chosen by IIF, COALESCE or CASE from spans and Nulls. A negative span now moves the date back instead of landing on the epoch's far side. A time written into the SQL is still a date, as in Access. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…cimal in a choice Round, Abs, Int and Fix keep their operand's type but declared none, so in IIF, CASE or COALESCE a whole number in another arm declared the column while the Round arm returned a Decimal. They now declare what they return - a Decimal, Double, Single or Int64 as it is, a narrower integer or Boolean as a Long Integer, text as a Double, a date as a date for Int and Fix - and Sgn declares an Integer. A number written with a decimal point now counts as a Decimal when the alternatives of IIF, CASE, COALESCE, GREATEST, LEAST and LAG/LEAD settle on one type, as ACE reads it. EF writes a decimal zero as 0.0, so its Sum-with-a-default IIF(SUM(x) IS NULL, 0.0, SUM(x)) made every money sum a Double; it stays an exact Decimal now. A Double still wins where one is present. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Math.Floor was translated to FIX, which rounds towards zero - that is Math.Truncate - so Floor(-1.5) came back as -1. Access's INT rounds towards negative infinity, which is Floor. Math.Ceiling's IIF(FIX(x) = x, FIX(x), FIX(x) + 1) was one too high for every negative non-integer, and +1 for anything between -1 and 0; -INT(-x) is right. Both keep the argument's type. Verified against ACE and LibRed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Atan2 was Atn(y / x), which puts x < 0 in the wrong quadrant and divides by zero at x = 0. It now adds or takes away Pi for x < 0 and gives Sgn(y) * Pi / 2 for x = 0; ACE's IIF only evaluates the branch it takes. Asin and Acos divided by zero at x = 1 and x = -1; they now use the half-angle form 2 * Atn(x / (Sqr(1 - x * x) + 1)), whose divisor is never below 1. Verified against ACE and LibRed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Floor, Ceiling/Ceil, Sign, Sqrt, Ln, Log10, Log(base, x), Power, Asin, Acos, Atan, Atan2, Sinh, Cosh, Tanh, Degrees, Radians and Pi, none of which ACE has. Where Access has the function under another name it is that function - Floor is Int, Sqrt is Sqr, Ln is Log, Atan is Atn, Sign is Sgn and Power is the ^ operator - so it reads its arguments and fails as that one does. Floor and Ceiling keep their operand's type. Log with one argument is still Access's natural log; with two it takes the base first, as the standard does. An argument outside a function's domain is an invalid procedure call. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
functions.md now has the shared rules first, then the Access / VBA functions ACE has, then LibRed's extensions, each grouped by category. The extensions that sat among the Access ones - CLngLng, CDec, the standard statistic names and set functions - moved to their own section, which also gains Coalesce, NullIf, Greatest and Least and the millisecond date intervals. Extended functions are for queries only: anything stored in the file - a view, a procedure, a CHECK constraint, a DEFAULT or a calculated column - is read by ACE and Access too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
LibRed gets its own method-call translator provider, which uses the shared JetMathTranslator in compatible mode, whose SQL ACE has to run, and LibRed's copy of it in extended mode. The copy translates to the engine's standard functions: Floor, Ceiling, Sqrt, Ln, Log10, Power, Sign, Atan, Asin, Acos, Atan2, Radians and Degrees where ACE needs Int, Atn, Sqr and arithmetic, and Sinh, Cosh and Tanh, which ACE cannot write at all. Math.Log(x, n) becomes LOG(n, x), as the engine's LOG takes the base first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A definition's chain holds the definition and then its 8-byte trailing reserve, which spills onto a page of its own when it does not fit, so a continuation page can hold reserve bytes and no definition. The reader worked out the page count from the length alone, so it rejected every ACE table whose definition ends within 8 bytes of a page boundary. The writer crashed on those lengths at the first boundary and, at later ones, put definition bytes on the extra page where ACE puts only reserve. Rewriting an existing definition now matches ACE too, growing or shrinking: the first page is rewritten in place, and alone only the reserve past the new end is zeroed, older bytes beyond it left as they were. Continuation data always goes to freshly allocated pages, the last allocated first from the lowest free page, and the old continuation pages are released untouched instead of reused. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
DROP COLUMN refused a memo or OLE column because its long values were not handled. ACE removes the column's long-value map entry from the definition, retires its owned and free map records from their holder as DROP TABLE retires a long-value column's, and returns every page it owned to the global free map at close, leaving the pages themselves as they were. LibRed now does the same, through the steps DROP TABLE already had, now shared by both. Retiring those records, ACE clears the owned pages' bits except the page still in the column's free map, whose bit stays in both records; DROP TABLE cleared that one too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adding a relationship puts an incoming block in the parent table's definition, and that edited the first page in place: it left the 8 bytes past the new end as they were, and refused a parent whose definition spans or would spill onto a continuation page. ACE zeroes those 8 bytes when a definition grows, as when it shrinks, so over the tail an earlier drop leaves behind LibRed kept bytes ACE clears. The block now goes in through the same rewrite as every other definition change, which also lets a relationship reach a wide parent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ACE records every relationship as a type-8 MSysObjects object in the Relationships container, besides its MSysRelationships rows: named after it, flags 0, taking the next id from the sequence queries draw on, with two MSysACEs rows. LibRed wrote only the MSysRelationships rows. It now creates the object through the allocation views already use, and DROP CONSTRAINT - and so DROP TABLE of the referencing table - removes it with its permission rows. A relationship name another relationship already has is now refused, as ACE refuses it, before anything is written; it may still match a table's or a query's. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An index's statistics block holds a total entry count and a unique entry count. ACE advances the unique count on INSERT only, by one when the row brings a key the index does not hold at that moment - so a non-unique index counts a new key, collation-equal keys and repeated Nulls once, a key whose last row was deleted again - and once per real index. Building an index over existing rows, whether CREATE INDEX, a foreign key's backing index or an ALTER COLUMN's rebuild, sets both counts from the rows present: the entries it holds and its distinct keys. Nothing else changes them. LibRed advanced only unique indexes, counted twice where a relationship shares a primary key, wrote 0/0 for a built index and 1 as a rebuilt one's total, and its whole-table rebuild for a Memo/OLE retype recounted every index, the referencing tables' included; each now matches ACE. An OLE column has no index key, and ACE refuses it on every route into an index - CREATE INDEX, a primary key or unique constraint, a foreign key, an ALTER COLUMN to OLE of an indexed column - before anything is written. LibRed accepted all of these on an empty table, after which every insert failed; it now refuses them with ACE's message. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
LOG takes the standard's base as a second argument in LibRed, beyond ACE's one-argument Log, so the arity test now expects one or two. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An explicit commit and every close forced the OS cache to disk with FlushFileBuffers. ACE never does: under its default settings and with Implicit Commit Sync or User Commit Sync flipped, it issues no flush on commit or on close and opens the file without write-through. A synchronous commit only means its pages reach the OS before the commit returns. LibRed now does the same, handing the writes to the OS without forcing them to disk. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Outside a transaction PageCount measured the stream, a file-information syscall, and every read path range-checks its pointers against it: each row fetched by an index seek, each index, definition, long-value and usage-map page followed. Reads never open a transaction, so a query paid one syscall per pointer - nearly free on Linux, but on Windows each goes through the filter stack. A CREATE TABLE made 196 of them and the first INSERT after it 133. Only LibRed's own writes change the file's length, and all of them go through PageChannel, so the shared per-file cache now holds the committed length: measured once when the first channel opens the file, updated when a publish grows it or a rollback truncates it. PageCount and the commit paths read that instead, and the same statements now make none. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The committed file length was read under the page cache's lock, which every cache hit already takes and every connection on the file shares. PageCount reads it on each range check - once per row an index seek fetches - so it added a lock acquisition to the hottest read path. It is a single value, so it is now read and written atomically without the lock. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ChrisJollyAU
force-pushed
the
nextfeature1
branch
from
September 20, 2026 06:55
fca037e to
091e9fa
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This branch brings LibRed's SQL engine and file writer much closer to ACE, adds standard-SQL extensions for extended mode, and fixes several Math translations in the shared Jet dialect. Nearly every behaviour change was measured against ACE first, and the file-format changes are pinned by whole-file byte comparisons against ACE's own output. Version bumped to 11.0.0-alpha.3.
EF Core providers (Jet and shared
EFCore.Jet.Common)Math.Floorwas translated toFIX, which truncates toward zero; it is nowINT.Math.Ceilingis now-INT(-x); before, it was off by one for negative non-integers.Math.Atan2now gets the quadrant right for x < 0 and no longer divides by zero at x = 0.Math.AsinandMath.Acosuse the half-angle form, so they no longer divide by zero at ±1.Close()now returns a pooled connection under the keyOpen()looks it up by. Before, every open created a new native connection, and the 65th query failed.ESCAPEclause, which neither ACE nor LibRed accepts.LibRed EF Core provider (extended mode)
Math.MaxandMath.Min(andEF.Functions.Greatest/Least, andMax()/Min()over an inline collection) translate toGREATEST/LEAST.Mathtranslates to the engine's standard functions:FLOOR,CEILING,SQRT,LN,LOG10,LOG(base, x),POWER,SIGN,ATAN,ASIN,ACOS,ATAN2,SINH/COSH/TANH,RADIANS,DEGREES. Compatible mode keeps the shared Jet translator.LibRed SQL engine
LIKE,BETWEEN,INand bitwise operators, with VBA precedence.Format, financial and aggregate functions, with ACE's domain and overflow errors.MODand\compute in Int64.CCur/CDecread their arguments exactly.Trimalso strips the ideographic space.CASE,IIF,COALESCE,GREATEST/LEASTand set operations settle on one type, as ACE does.DATEDIFF('ms')is Int64.DATETIME2keeps its sub-millisecond ticks.UPDATE/DELETEbehave as in ACE.ORDER BY nnames the nth output column.REFERENCESwithout a column list resolves to the parent's primary key.IDENTITYis accepted.CLUSTERED/NONCLUSTEREDare accepted where ACE accepts them.EXCLUDE,IGNORE/RESPECT NULLSandNTH_VALUE … FROM LAST.DISTINCT,FILTER (WHERE …)and windows over grouped queries.PERCENTILE_CONT/PERCENTILE_DISC … WITHIN GROUP,LISTAGG.CORR/COVAR_*/REGR_*.DATALENGTH,CLngLng.docs/functions.mdis reorganised into Access/VBA functions and LibRed extensions. Extensions are for queries only, never for anything stored in the file.LibRed file format — writes that now match ACE
DROP TABLEnow frees everything the table owns: long-value pages, all index pages and map records, continuation pages.0x09.ALTER COLUMN: fixes for indexed columns, and calculated columns carried through a rebuild.DROP COLUMNof Memo/OLE columns is now supported, retiring their long-value maps and pages as ACE does.MSysObjectsobject with theirMSysACEsrows, as ACE records them.Tests and CI
DROP COLUMN, the trailing-reserve spill, index statistics, and the OLE index refusal.Docs
src/LibRed/docs/format): tidied down to what the format is, and updated with everything above that was measured.CLAUDE.md/AGENTS.md: the version line is updated.