Report Grok usage with recorded and estimated spend - #3135
Conversation
|
🦞👀 Pull request received. I will update this pull request when review starts. ClawSweeper review completeClawSweeper finished reviewing this revision. The review result is being finalized. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 09cf7edb0f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Codex review: blocked before merge. Reviewed September 7, 2026, 9:32 PM ET / September 8, 2026, 01:32 UTC. ClawSweeper reviewWhat this changesThe PR counts completed Grok CLI turns, reports recorded spend with labeled price estimates where needed, and imports qualifying OpenCodex OAuth usage into Usage & Spend. Merge readiness⛔ Blocked before merge - 4 items remain Keep open: the accounting improvement remains absent from main and v0.56.8, and no blocking code finding remains. The unresolved owner decision about displaying Grok dollars to existing users still prevents landing. Priority: P2 Review scores
Verification
How this fits togetherCodexBar combines local usage logs with provider snapshots to populate its menu and spend dashboard. This change updates Grok accounting and filters optional OpenCodex imports using each recorded attempt’s credential provenance. flowchart TD
A[Grok CLI session logs] --> B[Bounded completed-turn scan]
B --> C[Recorded spend or price estimate]
D[Optional OpenCodex logs] --> E[Reported Grok OAuth attempts]
E --> F[Windowed spend dashboard]
C --> F
C --> G[Grok menu]
Decision needed
Why: The independent accounting evidence supports the implementation, but it cannot settle the requested reconsideration of the earlier display-policy ruling. Before merge
Agent review detailsSecurityNone. Review metrics
Root-cause clusterRelationship: Members:
Proposal only: this assessment does not dispatch repair, suppress jobs, mutate sibling items, close, or merge anything. Merge-risk optionsMaintainer options:
Technical reviewBest possible solution: Keep completed-turn accounting and clear cost-source disclosures, with the owner explicitly approving the dollar-display behavior for existing users. Do we have a high-confidence way to reproduce the issue? Yes for the token-accounting defect: current main reads context occupancy instead of completed-turn consumption, corroborated by two independent corpora. This read-only review did not execute a reproduction. Is this the best way to solve the issue? Yes for accounting: bounded completed-turn scans, recorded-cost precedence, and attempt-specific import are coherent repairs. Whether dollars should appear under existing settings remains a product decision. AGENTS.md: found and applied where relevant. Codex review notes: model internal, reasoning medium; reviewed against 9810f24b0609. LabelsLabel justifications:
EvidenceWhat I checked:
Likely related people:
Rating scale
Overall follows the weaker of proof and patch quality. Workflow
HistoryReview history (54 earlier review cycles; latest 8 shown)
|
08360b5 to
e3cd3b9
Compare
|
Both automated findings are addressed, plus the review's other checklist items. The inline comments were left against P1 — Preserve the Grok fallback on repeated probe failuresFixed in if provider == .grok {
if self.tokenSnapshotPublicationForCurrentProviderConfig(for: provider) == nil {
Task { @MainActor [weak self] in
await self?.scanAndPublishGrokLocalTokenSnapshot(...)
}
}
} else if Self.tokenCostRequiresProviderSnapshot(provider) {
self.clearTokenSnapshot(for: provider)
}Regression coverage is in P2 — Refresh pricing before scanning Grok sessionsCorrect, and thank you — this was a genuine gap and not one the local tests would have surfaced. Fixed in Note the inline comment still points at Coverage: Real-session evidence
The same corpus on Those figures were cross-checked against an independent reimplementation of the pricing formula over the same logs; the two agree to the cent. Merge risk / branch stateRebased onto current One thing deliberately left undone: no |
|
Addressed both current findings in
Validation on the exact pushed head:
@clawsweeper re-review |
|
🦞🧹 I asked ClawSweeper to review this item again. |
ClawSweeper found two window-projection defects that the 365-day Grok snapshot newly exposes, since a maximum-window snapshot is now narrowed per consumer. `grokLocalTokenSnapshot` recomputed tokens and requests from the retained days but copied the published cost total, so a 30-day menu view could render the 365-day dollar amount beside a 30-day token count. It now sums the retained rows. Both that projection and `CostUsageTokenSnapshot.narrowed(toHistoryDays:)` carried the snapshot-wide provenance into the derived window, so a window that excluded every recorded row still claimed recorded spend — newly reachable now that a Grok window can be mixed. The Grok projection derives the disclosure from the coverage counts of the days it kept, because its scanner counts a recorded turn as priced and a card fallback as estimated. The generic narrowing applies the existing `CostProvenance.forWindow` rule instead, which stops a costless window from claiming a provenance without guessing what the surviving rows of another provider mean. Each fix carries a regression confirmed to fail without it.
|
Rebased onto current main The existing opt-in OpenCodex import now accepts only reported physical Grok OAuth attempts with request-time provenance from OpenCodex #3642. It excludes API-key/historic/unknown records, avoids combo-parent and duplicate counting, preserves provenance through the derived cache, and keeps OpenCodex list-price estimates distinct from native CLI-recorded spend. The menu remains backed by native CLI logs. Validation on this head:
The PR description contains the exact commands and evidence. The default Grok dollar surface remains an owner/product decision before merge. @clawsweeper re-review |
|
🦞🧹 I asked ClawSweeper to review this item again. |
|
Addressed both custom-pricing findings on
The new regression suite reproduced both findings before the repair. Validation on the final tree: The PR body now contains the final-head validation and the explicit standalone-price contract. The Grok dollar-display decision remains with the maintainer. Please review this final description and head. @clawsweeper re-review |
|
🦞🧹 I asked ClawSweeper to review this item again. |
|
Synced #3135 with main The generated parser-hash conflict is resolved by regeneration from the merged source ( Main's Linux cache fixes, stacked-chart rendering, provider/widget changes, and release metadata are preserved. The two Grok changelog entries now live under 0.56.8 Unreleased; published release sections match main. The evidence report also records that the producer contract landed through the attributed OpenCodex #3762 carry; its original captured bytes remain pinned and are not represented as a new capture of that carry. Validation on this final head:
The PR description now contains the final source, commands, and evidence. The default Grok dollar display remains an owner decision. Please review the final head and description. @clawsweeper re-review |
|
🦞🧹 I asked ClawSweeper to review this item again. |
|
Updated to Current-head validation: The default Grok dollar display remains an owner decision, and the prior maintainer change-request review still needs reconsideration. @clawsweeper re-review |
|
🦞🧹 I asked ClawSweeper to review this item again. |
|
@steipete The latest ClawSweeper review of |
|
Updated to
@clawsweeper re-review |
|
🦞🧹 I asked ClawSweeper to review this item again. |
|
Final validation for
The PR body is updated with the current source and validation. The default Grok dollar-display policy and the earlier maintainer changes-requested review still require the owner's decision. |
|
Synced with main The parser hash is regenerated as
@clawsweeper re-review |
|
🦞🧹 I asked ClawSweeper to review this item again. Re-review progress:
|
|
I retrieved the landed OpenCodex carry directly through the GitHub API to clarify the dependency-inspection limitation in the latest review. #3762 is merged, with immutable merge commit
For content identity, the inspected schema/persistence blob is This verifies the landed source contract. The executable capture remains pinned to @clawsweeper re-review |
|
🦞🧹 I asked ClawSweeper to review this item again. Re-review progress:
|
|
Final validation for
The PR body now has the final validation results and the immutable landed-producer source audit. The owner display-policy decision and older changes-requested review remain unresolved. The review environment's inability to retrieve the later producer carry is also recorded; the supplied source audit does not expand the pinned executable capture into a claim about a released producer version. |


Problem and resulting behavior
Grok's local fallback reads ending context occupancy from
signals.json, which is not completed-turn usage, and publishes no local cost. This PR reads completed turns from bounded native CLI logs and carries recorded-versus-estimated cost through the menu and Usage & Spend.session/updateand_x.ai/session/updaterecords inupdates.jsonl, with daily, model, and request breakdowns.usage.costUsdTicks / 1e10as the authoritative turn total. Count it once. Show nested model dollars only when all nested ticks exist and reconcile to that total; otherwise retain model tokens with unknown model dollars. Missing recorded cost falls back to disclosed public xAI list prices.Grok CLI-recorded spend, list price where unrecorded · not a bill.OpenCodex integration
With the existing Include OpenCodex usage logs switch enabled (off by default), Usage & Spend also includes reported physical Grok OAuth attempts. The producer contract originally proposed in OpenCodex #3642 has now landed in
devthrough the attributed carry #3762, merged asf00f2bcaea251ebe7ad4de9e38337b4be0ccee47. It writesattempts[].credentialSourcefrom the resolved upstream transport. Compatibility is verified against the pinned producer commit in the evidence below; availability in a released OpenCodex version is not claimed.The landed carry's immutable source contract has additionally been inspected directly, including resolved-adapter stamping, persistence normalization, and OAuth/API-key regression assertions. The executable capture remains pinned to its documented producer revision; this source audit does not claim a new capture or released-version coverage.
Only
provider: "xai",credentialSource: "grok-oauth"attempts with an upstream send and reported token usage qualify. API-key traffic, historic rows, locally answered requests, unknown sources, and estimated/unreported token counts stay excluded. Current configuration and top-level credential metadata cannot retroactively classify usage.Combo requests contribute each qualifying attempt's own token counts, never the parent aggregate. Duplicate request IDs are resolved before grouping; duplicate attempt ordinals are rejected. SQLite cache schema 3 retains attempt metadata and rebuilds older derived caches from the log. OpenCodex dollars use list prices and remain estimates, including when combined with native CLI-recorded spend. Missing token classes and unknown prices retain tokens without inventing a dollar value. The Grok menu continues to use native CLI logs.
Producer-to-dashboard evidence and review fixes
The committed evidence report and raw ledger fixture contain output from OpenCodex's unmodified production handlers and durable usage writer at
146ed679c9633e5d68726217fcadc8e0b107339b. Two localhost HTTP requests exercised OAuth 401 replay through Responses and native Chat API-key dispatch. Upstream responses and credentials used isolated fixtures; this is production-path capture evidence, not live vendor authentication or billing evidence. The capture helper rejects unexpected external fetches and is reproducible against that pinned checkout.CodexBar imports the exact captured bytes through the production disk loader, with no injected entries or loader closure, and reopens the persisted cache:
Standalone reports retain explicit xAI custom-price estimates from either the caller overlay or the application overlay, including known zero. Raw xAI records without an explicit price remain token-only, and the subscription fan-out still excludes API-key and historic records.
opencodeandopencode-freeretain their existing catalog and custom prices.Application overlays first match the original model name before the provider-qualified catalog name. Bare keys retain precedence when both keys exist; incomplete rates stay unknown, cached-input accounting is preserved, and caller-supplied custom pricing still takes precedence over the application overlay. New regressions cover these cases through the application-overlay parameter and the standalone disk/cache loader. Coverage includes known-zero overrides, incomplete rates, cache accounting, and both snapshot and application overlays.
Both async native Grok scan entry points propagate the executor's cancellation callback through discovery, JSONL reads, and aggregation. An in-flight regression cancels after parsing begins, observes
CancellationError, and confirms the next queued scan runs within one second. Cancelled partial parses are uncacheable and cannot establish complete history. The final serial proof stopped at 4,874/40,000 decoded records and released the queue after 0.001864583 seconds. The evidence report also retains the initial measurement.Bounds and compatibility
Native scans run on the dedicated executor with limits of 64 MiB / 20,000 turns per file, 1 MiB per record, 256 sessions / 256 MiB / 100,000 turns per scan, and 4,096 discovery entries. The process cache retains at most 64 files or 50,000 turns. Cancellation and truncated history cannot publish complete coverage.
Merged main
912eac22372d7c3f1a1afff77acb97dc502117f4, retaining its shared report accumulation, bounded pricing resolver, provider fixes, and 0.56.8 release. Model-target mapping is colocated with the existing Codex resolver extension without changing its API or mapping; the separate xAI pricing fingerprint remains intact. The architecture gate retains the same provider-reference fingerprints at the updated source locations.Regenerated the native parser hash from the merged source:
c3a879df4eff7187. Current-main hash9ca89383b9957b07, prior-PR hash1a4afd74939160fd, and intermediate upstream hashba2eca901de4c53dremain compatible because these report and lookup changes preserve native parsed rows and checkpoints. Existing predecessors are retained. Parameterized SQLite adoption tests cover these hashes without rebuilding. Grok release notes are under0.56.9 — Unreleased; the published 0.56.8 and earlier sections match main.Validation
Head:
c35734ad5a747a83c46b3844c8ce33fd46dbefe1.make checkpassed: zero SwiftLint violations in 2,143 files.make testpassed: 1,036 selections, 87/87 groups successful on the first attempt, zero failures, retries, or timeouts (947.9 seconds execution; 964.5 seconds total).git diff --checkpassed.c3919a224: 136 cache/pricing/architecture/chart tests and 206 native Grok/producer-import/pricing/cache/cancellation tests passed on that revision. The current integration is validated separately by the regression checks above and the full suite.The supplemental proof and local-corpus measurements below were captured at
c3919a224. They remain historical, source-linked evidence; current-head validation is recorded separately above. The supplemental proof uses the repository’s serial execution mode because those suites share scanner-cache counters.The September 6 local-corpus proof retained
vendorMeteredfor populated windows: 7 days = 95,891,889 tokens / $12.57875572; 30 days = 98,631,812 tokens / $12.94957366. The empty 1-day window retainedunknown. Controlled JSONL fixtures also proved recorded-only and estimate-only windows after narrowing mixed source history.The OpenCodex regressions cover mixed OAuth/API-key/provider attempts, historical and malformed records, duplicate suppression, cache reopen/incremental append/schema upgrade, missing-price behavior, recorded-plus-estimated date filtering, and the dashboard's opt-in switch. Existing native regressions cover bounded parsing, outer/nested cost reconciliation, local fallback freshness, account isolation, and menu/dashboard provenance.
All ordinary tests suppress Keychain access and isolate provider files. The optional native proof reads local Grok logs and the cached pricing catalog without authentication, browser-cookie import, or live provider requests. Presentation evidence uses production menu/dashboard models; no app-bundle screenshot is claimed.
Maintainer decision
Please revisit the 2026-08-21 Grok cost ruling before merge. It preferred public-card pricing based partly on my incorrect
1e9divisor; #3345 established1e10, and the corrected measurement explains the difference. This branch uses recorded spend with public-card fallback.Whether existing Grok users should receive this disclosed dollar surface by default remains an owner decision. The accounting corrections and source labeling do not override that decision. Maintainer approval is still required.