From 55b1748cddcb9a262a1cf48900f876fa007a1ae0 Mon Sep 17 00:00:00 2001 From: Sergii Demianchuk Date: Sat, 3 Oct 2026 09:45:49 -0400 Subject: [PATCH] docs: RELEASE-NOTES.md for v1.3.0 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Replaces the v1.2.0 notes, as each release does. Every figure is measured rather than recalled: 48 suites and 854 cases, counted the way scripts/draft-release-notes.mjs counts them (v1.2.0 that way is 40 and 599, which is what its notes say); the per-suite counts from each suite's own last line; the context windows by running both releases' catalog over the model ids; the token lifetimes from the backend's EditorToken; the plan each model starts at from the live pricing endpoint. The draft script itself could not be used: it refuses to run once the tag exists, and v1.3.0 was tagged first. So these notes land after the tag, as v1.2.0's did. Described relative to v1.2.0, not to this cycle's own review rounds — a state that only ever existed between two commits here is not something a user saw. One claim is flagged in the text as unproven on screen: the agent's browser preview. The chain is tested up to the call that opens the tab. --- RELEASE-NOTES.md | 147 ++++++++++++++++++++--------------------------- 1 file changed, 63 insertions(+), 84 deletions(-) diff --git a/RELEASE-NOTES.md b/RELEASE-NOTES.md index 3aa7b2b..350b98c 100644 --- a/RELEASE-NOTES.md +++ b/RELEASE-NOTES.md @@ -1,128 +1,107 @@ -# LevelCode v1.2.0 +# LevelCode v1.3.0 -Take a screenshot, `⌘V` into the composer, ask why the layout is wrong. That is the whole feature, and most of this release is the work underneath it — a token meter that tells the truth about pixels, a store that keeps your screenshots on your own disk, and a gate that stops an image reaching a model that cannot see it. The transcript also stops narrating itself, and Sessions is one click from the chat tab. +When a LevelCode Cloud session ended, v1.2.0 said so with `API 401: Signature has expired` in red — beside an account popover that still claimed you were signed in. This release replaces that with one card and a **Sign in** button, shown before you send anything. Mostly you will not see even the card: the session is renewed ahead of time, a lapsed token is renewed wherever a request is sent rather than only in the chat, and thirty days of sign-in now run from the last time you used the editor. GPT-6 Astra and Claude Fable 5 / 5.1 are sized as the models they are, and the agent's browser preview can open for the first time. ## Highlights -### Paste a screenshot, ask about it +### An ended session is a Sign in button, not a red error -Three ways in, in the order you will actually use them: +The card is headed **Your session has expired**, greets you by name, and says: *Sign in again to keep using LevelCode Cloud — your chat and files here are untouched.* It has two buttons: **Sign in**, and **Use my own key instead**, which opens the provider-mode setting. -- **Paste.** `⌘V` a screenshot straight into the composer. No dialog, no upload step, no file to manage. -- **The image button** in the composer, or **AI: Attach Image to Chat** from the palette. It offers the image tabs you already have open before it offers a file dialog — the picture you want is usually one you were just looking at. -- **Drag a file from Finder** onto the chat. This one needed a core patch; see below. +- **It is found before your first message.** The session is checked when the chat opens and every time the window comes back to the front. The check reads the token's own expiry, so a session with hours left costs no request. +- **The popover and the footer stop saying you are signed in** at the same moment. That was the worst part of the old failure: two parts of the same window disagreeing about whether you were signed in. +- **The card stays with the conversation.** New Chat, a resumed session and a checkpoint restore each rewrite the transcript, and each puts the card back while the expiry is unanswered. Signing in, or choosing your own key, takes it away. +- **Inline edit and Agent Sketch say it too.** An ended session reads *Your LevelCode Cloud session has expired. Sign in again to continue.* there — with a **Sign in** button on the inline-edit toast — where v1.2.0 showed the gateway's 401. -`png`, `jpeg`, `gif` and `webp` are accepted and nothing else — that is the list the vision APIs take, so a wider net would only fail later and further away. Each attachment becomes a chip under the composer with a thumbnail, its cost in tokens, and a remove button. Up to **five images per message** by default, and when there is more than one they are introduced to the model as `Image 1:`, `Image 2:` so that "the second screenshot" has something to refer to on this turn and every turn after it. +**Only the server saying no ends a session.** Being offline, a 5xx or a malformed reply means nothing is known yet: you stay signed in, and the request shows its own error. A renewal that hangs is given ten seconds and then treated the same way. Clearing credentials on a network blip would sign you out for closing your laptop on a train. -### Your screenshots never leave your machine +If you use LevelCode on your own key with no account, none of this reaches you — no card, no prompt. Gateway mode with no token is what a fresh install is, and it still falls back to your key. -Bytes are written beside the session that used them, content-addressed by SHA-256: +### Thirty days now run from the last time you used it -``` -~/.levelcode/sessions//media/.png -``` +A cloud sign-in is two credentials: an access token good for 8 hours, and a refresh token good for 30 days that buys the next access token. Through v1.2.0 the editor kept the refresh token it was handed at sign-in for that token's whole life, so the thirty days ran from the day you signed in. Someone who used LevelCode every day was signed out on day thirty all the same. -The conversation, the session log and the token meter all carry a **reference**, never the bytes. That is not a size optimisation: listing your History re-parses every session file in a project whenever the index is missing or on an older schema, and inlined base64 would make drawing a list of session titles parse every screenshot in every session you have ever taken. +The cloud now issues a new refresh token with every renewal, and the editor stores it. The thirty days slide: keep using the editor and you do not reach the wall; leave it for a month and you sign in again. -Nothing is uploaded. There is no bucket, no signed URL and no retention policy to read, because there is nothing on our side to retain — which is also the only shape that works for BYOK, where the editor talks to your provider directly and a detour through our infrastructure would contradict the promise that we are not in the middle. - -Because the same screenshot pasted twice hashes to the same file, a re-paste after a failed send costs a hash and a stat rather than a second copy. - -**Media is swept, not orphaned.** Sessions are append-only and deleting one writes a lifecycle event rather than removing the transcript, so "the images go away with the session" was never going to be true. A sweep runs when a session is sealed and removes media nothing refers to any more — with a **seven-day age floor**, because a plain unreferenced-means-delete rule would delete the images of the conversation you have open right now. - -### The context meter stopped lying about images - -The estimator measured `JSON.stringify(messages).length / 4`, which is sound for text and catastrophic for an image. Base64 books about a third of its byte count as tokens, so a 1 MB screenshot read as roughly **333,000 tokens** — larger than most context windows — for something that really costs about 4,800. Storing refs instead of bytes then swung it the other way and reported a ~1,800-token image as about 18. - -Images are now counted by what they actually cost. Claude sees an image as 28×28 patches, so the price is `⌈w/28⌉ × ⌈h/28⌉` visual tokens, capped per model tier: - -| Model tier | Long edge | Token cap | -| --- | --- | --- | -| Claude 4.7 and later, including the 5 line | 2576 px | 4784 | -| Everything else | 1568 px | 1568 | - -An unknown model falls to the standard tier and an unknown size assumes the cap, so the meter fails toward over-counting rather than under. - -Worth being plain about the scope: **today this only misreported the meter you look at.** Compaction cuts on message count and goal boundaries and never reads a token number, so nothing was being silently evicted. It becomes a correctness bug the day anything automatic keys off that figure, which is why it is fixed now rather than later. - -### Screenshots are resized before they are sent +| What | Value | +| --- | --- | +| Access token | 8 hours | +| Sign-in | 30 days, counted from the last renewal | +| Renewed ahead of expiry | when under 5 minutes remain — checked at chat open, and on window focus at most every 10 minutes | +| A renewal that hangs | given up on after 10 seconds — you stay signed in | -The long edge is capped at **2000 px** in the webview before anything is stored or sent. An image already under the cap is **passed through untouched, in its original format** — re-encoding a screenshot of text only stacks compression artifacts on the thing that most needs to stay legible. One that is over gets a single resize and a single WebP pass at quality 0.92. +### A lapsed token is renewed wherever a request is sent -The cap is 2000 rather than the more obvious 1568 because the server caps the *cost* at 4784 tokens either way, so the extra pixels buy legibility on small editor text for tokens that were already being spent. Measured on a dense 4K editor screenshot — small text edge to edge, the worst case for re-encoding — **770 KB → 189 KB on the wire**, and 4784 → 2952 tokens. Shots with more flat UI in them compress harder than that. +The chat and the agent already renewed a lapsed access token and sent the request again. Nothing else did. Once the eight hours were up, everything around the chat failed until a chat message happened to renew the token — for a session that was perfectly renewable. -| Source | Sent as | Visual tokens | +| Where | v1.2.0 | Now | | --- | --- | --- | -| 4K screenshot 3840×2160 | 2000×1125 | 2952 | -| macOS retina window 3024×1964 | 2000×1299 | 3384 | -| 1080p screenshot 1920×1080 | unchanged | 2691 | -| Half-screen 1280×1440 | unchanged | 2392 | - -Budget roughly **2,400–3,900 tokens per screenshot**, and remember it rides along on every subsequent turn in that conversation. +| Inline edit | *LevelCode AI edit failed: … API 401: Signature has expired* | renewed, and the edit asked once more | +| Agent Sketch — Run, the board command, Generate flow | the gateway's 401 on the nodes that ran | renewed; nodes that fail together share one renewal | +| Inline completion | silently nothing, on every pause in typing | renewed quietly | +| Compact and the session-memory summary | failed | renewed | -### Dropping a file on the chat needed a core patch +What it deliberately will not do: -VS Code's workbench claims OS file drops before a webview iframe ever sees them, so a drag out of Finder arrived with an empty `dataTransfer.files` and the editor helpfully opened your screenshot in an image tab instead. The chat now handles the drop in `editorDropTarget.ts`, reads the paths, and hands them to the extension. +- **Send anything twice.** One retry, and none once part of an answer has arrived — a second send would repeat it. +- **Outlive a Cancel.** Stop, Cancel or the next keystroke ends the wait at once, and a request nobody is waiting for starts no renewal. +- **Cross accounts or hosts.** A request stays with the account and the cloud host it was sent on. Sign in as someone else while it is out — in this window or another — or point the editor at a different host, and it is not sent again on the new credentials. +- **Turn typing into refreshes.** Ghost text is sent on every pause in typing. It may start one renewal a minute, and waits on one that is already out. -**If you build from source, this is a core patch, not an extension change** — a fresh `vscode/` clone needs `patches/levelcode-core.patch` applied by `bootstrap.sh` before drag-and-drop works. Two things cost real time here and are written down in `docs/IMAGES.md` so they cost nobody else any: extension webview view types are rewritten with a `mainThreadWebview-` prefix before they reach the drop target, so matching the bare id makes the patch silently inert; and holding Shift takes a different code path entirely, which means "drag with Shift works" was never evidence that the patch worked. +### One session, every window -### An image only goes to a model that can actually see +The session is stored once and every window shares it, but a window only heard about the changes it made itself. One left open in the background kept showing "signed in" after another window had signed out. -The provider **and** the model must both declare vision. Reading only the model id meant a custom OpenAI-compatible endpoint returned true for any model whose *name* looked like a vision model, and images went to an endpoint nobody had said could read them. +Each window now compares what is stored with what it is showing whenever it comes to the front, and catches up: the card and a signed-out popover if the session ended elsewhere, the popover alone if you signed out there, and the card taken down if you signed back in there. -Four providers declare it: **Anthropic**, **OpenAI**, **OpenRouter** and **xAI**. Ollama and custom endpoints deliberately do not — a custom endpoint that does serve a vision model needs `vision: true` on its registry entry, because the honest place to declare a provider's capabilities is the provider registry, not a per-user override. In gateway mode the check runs against the gateway's own model, so LevelCode Cloud gets images wherever the model supports them. +### GPT-6 Astra and Claude Fable 5 / 5.1, sized as they are -The composer refuses an attachment *before* you type anything, and re-checks at send — you can switch models between attaching a screenshot and pressing enter, and that used to produce a provider error instead of a sentence. +v1.2.0 had no entry for any of them and fell back to a guess: a 200k-token window for all three, and no images for Astra. In gateway mode that meant the model picker offered Astra while the composer refused your screenshot, and the context meter was sized for a fifth of the real window. -### The activity group reads as text, not as a widget +| Model | Context window | Images | +| --- | --- | --- | +| GPT-6 Astra | 1,050,000 tokens | yes | +| Claude Fable 5 and 5.1 | 1,000,000 tokens | yes | -The collapsed header said "3 steps". It now says what actually happened — which files were read, which command ran — because a count is the one thing you can already see. Context is announced **once** when it enters the conversation rather than re-stated every turn, the chevron trails the thing it discloses instead of leading it, and the group rows are inset inside a single container rather than nested in two with a rail down the side. +The rows hold under every id the models arrive as: `openai/gpt-6-astra` and `anthropic/claude-fable-5.1` from the gateway and OpenRouter, and Anthropic's own `claude-fable-5-1`. -### Sessions and Project Memory are one click from the chat +On LevelCode Cloud, Fable 5 and 5.1 are included from **Pro** and GPT-6 Astra from **Pro+**. They are expensive models, and the pricing page says so in turns rather than leaving you to find out: Pro's 2,000 credits are about 260 turns on Kimi K2.7 and about 20 on Fable 5.1. -Both have a button on the chat tab, and **AI: Project Memory** is a command now. The row actions inside the Sessions panel — Rename, Done, Delete, Pin — were rendered but invisible, showing an empty grey box on hover; they render, and they are visible. +### The agent's browser preview opens -## New settings +When the agent starts a web server in the background and it prints a local address, LevelCode opens it in the built-in browser beside the chat — without taking focus, and once per address, so closing the tab is final. That is what `levelcode.ai.preview.autoOpen` has promised since v0.9.2. -| Setting | Default | Description | -| --- | --- | --- | -| `levelcode.ai.chat.maxImagesPerMessage` | `5` | How many images may be attached to one message. Clamped to 1–20 at the boundary — 20 is the API's own ceiling | +It never happened. A logging call on that path referred to a name that was out of scope, and threw before the browser was asked for. Every release from v0.9.2 to v1.2.0 carries it; this is the first that can open the tab. -## New commands +**What is proven, and what is not.** A test now runs the real chain — the agent, the tool, the command's output — up to the call that opens the browser. The tab itself appearing in a packaged build is the one step no test covers. -| Command | Does | -| --- | --- | -| `AI: Attach Image to Chat` | Offers open image tabs first, then a file dialog | -| `AI: Project Memory` | Opens the project's memory from anywhere | +Only local addresses are ever opened: `localhost`, `127.0.0.1` and `[::1]`. ## Also fixed -- **The context and review bars sat flush left** instead of in the transcript column, because a `margin` shorthand overwrote the `margin-inline: auto` that centred them. The guard that now prevents it reads *every* declaration rather than the first — the original check used `.exec()` without the global flag, so it inspected one rule and reported the file clean. -- **Every image thumbnail was broken.** The chat's Content-Security-Policy declared `default-src 'none'` with no `img-src`, so the composer chip rendered as a broken-image glyph. -- **An image with no words returned a 400.** Sending a screenshot with an empty composer produced an empty text block alongside it, which Anthropic rejects. Text-only turns also stay a plain string rather than becoming a single-element array, so cached prefixes do not churn. -- **Agent mode dropped every pasted image.** The send path stored the bytes and then called the agent with the text alone. -- **Opening without a folder refused images outright**, on the same guard that used to refuse everything else. -- **The × on an image chip could not remove it**, and the target was too small to hit reliably. Both fixed, and the cap moved into one setting rather than being written in two places. -- **A send can no longer outrun its own attachment.** Normalisation is async, so pressing enter mid-decode posted an attachment with no data — refused at the host, and the image vanished from a message you had watched it attach to. The send path now waits on in-flight work and refuses a placeholder outright. -- **The provider boundary fails loudly.** Translating a conversation for an OpenAI-compatible provider silently dropped content blocks it did not recognise. It now throws, because a request that quietly discards its own subject is the exact failure this feature exists to avoid. Extended-thinking blocks remain an explicit, deliberate drop. +- **Three icons were printed as words.** Opening the chat in an editor tab showed a banner reading *layout MOVED TO THE EDITOR*; the preview chip and the recall and project-memory chips showed `globe` and `history` as text on ungrouped timeline rows. The glyphs are in, and a test now runs every icon name written down on either side through the real renderer. +- **A sign-in replaces the whole session.** One that arrived without a refresh token kept the previous session's, and the next renewal was made with it: refused, it ended the session you had just started; still valid, it renewed the *previous* account under the new one's name. Latent — the cloud sends a refresh token with every sign-in — and closed. +- **A background command that failed to start** raised an unhandled rejection instead of being logged. Same out-of-scope name as the preview. ## Not in this release -**The Sessions panel still does not search.** No filter, no fuzzy switcher, no keyboard jump. It was the stated gap in v1.0.5 and in v1.1.0, and it is the stated gap again — the chat surface took another cycle. +**The Sessions panel still does not search.** It was the stated gap in v1.0.5, in v1.1.0 and in v1.2.0, and it is the stated gap again. This cycle went to sign-in. -**Images are re-sent on every turn.** Base64 rides along with each subsequent request in a conversation, so a screenshot you attached ten turns ago is still being uploaded. The Files API fixes this properly by uploading once and referencing thereafter, but it is Anthropic-direct only, so it cannot be the primary path in a multi-provider client. Deferred deliberately, and the cost is bounded by the resize. +**A session with no refresh token cannot be renewed, and is not ended either.** When its access token lapses, the request still shows the gateway's 401. The cloud issues a refresh token with every sign-in, so reaching this takes a sign-in that arrives without one — but it is the one path left where that error can appear. -**Custom endpoints cannot opt into vision** without editing the provider registry. That is the correct default and the wrong end state; a per-endpoint capability declaration is the missing piece. +**Nothing renews the token in a window that simply stays in front.** The check runs when the chat opens and when the window regains focus. A window you never leave finds a lapsed token on its next request, which is renewed and sent again — a moment's delay rather than an error, but not the same as never lapsing. -**The sweep's seven-day floor is hard-coded.** It is the right default and it should probably be a setting. +**"Use my own key instead" is only on the chat's card.** The inline-edit toast offers **Sign in** alone, and Agent Sketch shows the sentence with no button at all. ## Test coverage -- **40 suites**, **599 cases** across the bundled extensions — all green. v1.1.0 measured the same way was 36 suites and 527 cases. -- `test/imageAttach.test.js` (30 cases) — the wire shape end to end: images lead and text follows, refs materialise into base64 only when a request is built, the vision gate at attach and again at send, the per-message cap, and multi-image labelling. -- `test/imageCost.test.js` (9 cases) — the arithmetic, pinned against every worked example in the vision documentation: 1092² → 1521, 1920×1080 → 2691, 3840×2160 → 2576×1449 at 4784, and that nothing is ever scaled *up*. -- `test/imageStore.test.js` (11 cases) — content addressing, the 5 MB ceiling, path-traversal refusal on a ref, a missing file throwing rather than sending a request without its subject, and the sweep's age floor. -- `test/contextAnnounce.test.js` (5 cases) — context enters the conversation once and is not re-announced. -- `test/translate.test.js` gained 92 lines covering the boundary that now throws instead of dropping blocks. +- **48 suites**, **854 cases** across the bundled extensions — all green. v1.2.0 measured the same way was 40 suites and 599 cases. +- `test/sessionExpiredHost.test.js` (70 cases) — the host's own functions, sliced out of `extension.js` and run against stores that answer a turn late: the expiry found at open and on focus, late answers that must not touch a newer session, a sign-out in the middle of a request, and what other windows do to the store. +- `test/authRetryCallers.test.js` (76 cases) — the real inline edit, Agent Sketch and completion modules, driven as the editor drives them against a gateway that answers 401 to a lapsed token. +- `test/authRetry.test.js` (45 cases) — the renew-and-retry rules on their own: one retry, one renewal at a time, and a request that stays with its account and host. +- `test/session.test.js` (19 cases), `test/sessionExpiredUi.test.js` (17) and `test/sessionExpiredCallers.test.js` (16) — the token arithmetic, the card's routing in the webview, and what inline edit and Agent Sketch say. +- `test/chatIcons.test.js` (6 cases) and `test/agentRunCommand.test.js` (4) — the icon guard, and the preview chain. + +**The suites now run on every pull request.** Until this release they ran in CI at one moment only — when a tag was pushed — so a new suite's first run on Linux was the release gate, after the merge. The first v1.3.0 tag failed its gate that way: a new test counted turns of the event loop across a real timer, which holds on a Mac and not on the runner. The test waits by the clock now, and the gate is one script, `scripts/test-extensions.sh`, shared by the release workflow and a pull-request check. -**Full changelog:** https://github.com/levelcodeai/levelcode/compare/v1.1.0...v1.2.0 +**Full changelog:** https://github.com/levelcodeai/levelcode/compare/v1.2.0...v1.3.0