Code Mower Cloud sharing is optional. The local OSS package remains useful without a CodeMower.com account or token.
This document defines the public metadata-only data boundary for Code Mower Cloud uploads.
- OSS user: installs Code Mower, runs local checks/reports, optionally gets a CodeMower.com team token, inspects a dry run, then uploads metadata.
- CodeMower.com operator: runs Supabase, Vercel, OAuth providers, DNS, service-role/admin secrets, retention, abuse handling, and hosted reporting.
OSS users should not need Supabase, Vercel, OAuth-app, DNS, database, service-role, or hosted-secret access.
Default cloud bundles exclude:
- source code;
- raw diffs;
- raw model transcripts;
- raw stdout/stderr;
- auth probe output; and
- secrets.
Report text is not uploaded by default. Uploading report contents requires an
explicit --include-reports flag, and the hosted service may still discard
report text depending on operator retention settings.
code-mower cloud export writes an inspectable local directory:
.code-mower/cloud-benchmark-bundle/
code-mower-cloud-bundle.json
README.md
reports/
The manifest uses schema code_mower.cloudUpload.v1. It contains metadata such
as privacy mode, upload mode, install id, optional team id, optional repository
slug, report count and report kinds, structured event count and event types,
excluded-content declaration, and copied report file metadata.
code-mower lanes status --repo OWNER/REPO and
code-mower board serve --repo OWNER/REPO use local schemas
code_mower.laneStatus.v1 and code_mower.board.v1. The explicit local
history commands use code_mower.boardEvent.v1 and
code_mower.boardEventStore.v1 under .code-mower/board/events.jsonl, and
code_mower.boardRecord.v1 for explicit write acknowledgements. Optional local
agent card adapters use code_mower.boardAgentAdapters.v1 from
.code-mower/board/agents/*.json. Board admin commands use
code_mower.boardDoctor.v1 for local diagnostics and
code_mower.boardReset.v1 for explicit local-history reset acknowledgements.
code-mower productivity report --repo OWNER/REPO emits local-only
code_mower.productivityReport.v1 by reading Board history, reviewer-spend
rows, and optional aggregate productivity_summary event files.
Those payloads are operator visibility data, not cloud upload data.
Current board/status JSON and local board event-store data are local-only and
are not uploaded by default. The explicit
code-mower cloud board-snapshot --repo-slug OWNER/REPO --json command exports
one summarized board_snapshot event with zero report text; adding --yes
uploads that metadata-only summary. Any future dashboard mirror expansion must
land as a paired OSS and dashboard change, update both this contract and
Board Data Contract, keep the hosted service
backward-compatible with v0.6/v0.7 uploads, and preserve the metadata-only
boundary: no source, raw diffs, transcripts, issue body text, raw stdout/stderr,
auth output, browser history, local secret values, or secrets.
Metadata events use schema code_mower.benchmarkEvent.v1. They are intended to
capture reviewer and workflow facts without raw code artifacts.
Supported event types include:
adoption_rundogfood_uploadbuilder_runreviewer_runcalibration_runvalue_report_snapshotlane_policy_snapshotprovider_catalog_snapshotwork_orderworkflow_runboard_snapshotcontroller_decisionmerge_decisionqueue_state_snapshotowner_interventionpr_outcomeproductivity_summaryreviewer_finding_outcome
Events may include provider/lens names, timing, cost, verdict, useful finding counts, false-positive counts, repository slug, install id, and coarse runtime metadata. They must not include source code, raw diffs, raw transcripts, stdout/stderr, auth output, or secrets.
The OSS uploader normalizes and validates every structured event before it is
written into a bundle. Required fields use simple JSON object/string shapes,
metrics, dimensions, and tool remain objects, event types come from the
supported list above, and additive fields are allowed as long as they pass the
same metadata-only privacy scan.
provider_catalog_snapshot events are special: they describe configured
provider lanes and safe tool/model/version coverage. They are useful for setup
and benchmark trust diagnostics, but they are not reviewer accuracy evidence and
must not be counted as useful findings, false positives, or lane-promotion
support. Cloud bundle provenance summaries therefore keep both raw upload counts
and benchmark-evidence counts: catalog snapshots still appear in raw inventory,
but they are excluded from benchmark_* provenance coverage fields.
work_order events are also operational metadata, not reviewer accuracy
evidence. They connect a GitHub issue planning flow to later builder/reviewer
runs by recording issue/work-order provenance, role lenses, review lanes, and
optional delivery metadata: PR URL/number/state, reviewer-check names/statuses,
merge SHA, and merged-at time. They must not include issue bodies, source code,
raw diffs, transcripts, stdout/stderr, auth output, or secrets.
builder_run events are authoring-side provenance, not reviewer approval. They
record who/what produced a branch or PR from an issue/work-order contract:
builder provider, executor surface, issue/PR identifiers, branch name, optional
safe run URL, and optional coarse metrics such as elapsed time, cost, or human
intervention count. They exist so CodeMower.com can connect
issue -> plan -> work order -> builder run -> PR -> reviewer checks -> merge
without receiving source, issue bodies, diffs, prompts, transcripts,
stdout/stderr, auth output, or secrets.
board_snapshot events are explicit CodeMower.com Board mirror metadata. They
use dimensions.snapshot_schema=code_mower.cloudBoardSnapshot.v1, include zero
reports, and summarize the local Board/status surface into whitelisted
metadata: repository, generated time, next action, remote availability, gate
status, PR number/branch/author/draft/merge/check/label summaries, workflow run
summaries, owner-queue reason summaries, opt-in agent card summaries, and
verdict/spend group counts. In v1.0, the same event may also include
dimensions.supervised_pilot, a compact controller-backed summary with the
supervised-pilot schema, controller mode, cycle state, decision state,
stop condition, next action/detail, lane/check references, reviewer outcome
states, queue metrics, active lane counts, active PR metadata, and active ready
issue metadata. Matching top-level metrics may include
supervised_open_pr_count, supervised_ready_issue_count, and
supervised_owner_action_count. The uploader intentionally does not send the
full local Board payload, PR titles, owner note titles, issue titles, local cwd
paths, PIDs, full head SHAs, gate rerun commands, source, raw diffs,
transcripts, issue body text, raw stdout/stderr, auth output, browser history,
local secret values, or secrets. The event type and supervised-pilot fields are
additive and optional, so CodeMower.com must continue accepting v0.6/v0.7/v0.9
uploads that omit them.
Supervised-pilot events are additive v1.0 metadata for controller and merge
visibility. controller_decision, merge_decision, queue_state_snapshot,
and owner_intervention events use
dimensions.supervised_pilot_schema=code_mower.supervisedPilot.v1 and record
state, next action, lane/check references, owner-action reasons, and coarse
counts without raw work content. See
Supervised Pilot Contract for the product
boundary, stop conditions, reviewer outcome references, example fixtures, and
privacy rules. CodeMower.com must keep accepting v0.9.x uploads that omit these
event types.
Controller-produced controller_decision, merge_decision,
queue_state_snapshot, and owner_intervention events may include optional
dimensions.orchestrator_provider. It is the host provider explicitly supplied
through controller run --orchestrator PROVIDER, or lower-precedence
CODE_MOWER_HOST, and mirrors the local report's orchestrator_provider.
Known values are codex, claude, cursor, devin, grok-bot, antigravity,
and muse. The explicit extension form is custom:<slug> with 1–64 lowercase
letters, digits, or hyphens, starting with a letter. Validation rejects empty,
unknown, reviewer-only, malformed, and secret-like values. Extensions identify
providers only; they must not carry user or machine identity, local paths,
session/lease identifiers, prompts, work content, auth output, or secrets.
When identity is not supplied, the field is omitted; no process-state inference
is performed. This dimension is additive and optional within the existing v1
schemas. Consumers must keep accepting older events without it. Top-level
provider and tool still describe Code Mower, and upload privacy remains
metadata-only.
pr_outcome is the additive atomic event for Dashboard 2.0. It uses the normal
code_mower.benchmarkEvent.v1 envelope with
dimensions.pr_outcome_schema=code_mower.prOutcome.v1. One event describes one
PR observation; repo_slug plus dimensions.pr_number is the PR identity.
Producers should make retries idempotent with the same event_id. When a newer
lifecycle observation uses another event id, consumers select the latest
created_at observation per PR before aggregating.
Required dimensions are pr_outcome_schema, positive integer-string
pr_number, ISO 8601 opened_at, outcome (open, merged,
closed_unmerged, or reverted), and cost_coverage (complete, partial,
or unknown). Merged and reverted outcomes require merged_at, closed-unmerged
requires closed_at, and reverted requires reverted_at. Timestamps include a
UTC offset and cannot precede opened_at; reverted_at cannot precede
merged_at. The reverted outcome is reserved for producers that hold
rollback evidence; GitHub's PR-list state field alone cannot infer it.
Metrics are atomic values, never precomputed dashboard rates:
pr_countis always1; optionalfix_round_count,reviewer_catch_count, andblocking_bug_countare non-negative integers.reported_cost_usdis the observed builder-plus-reviewer spend for the PR.cost_reported_run_countandcost_expected_run_countexpose source coverage.cost_covered_pr_countis1only for complete cost coverage and0otherwise. Partial coverage requires0 < reported runs < expected runs. Unknown coverage omitsreported_cost_usd; missing cost is never zero.
Dashboard Total Spend may sum reported_cost_usd across complete and partial
observations while displaying their coverage mix. Dashboard Cost per PR uses
only merged observations with cost_coverage=complete and computes
sum(reported_cost_usd) / sum(cost_covered_pr_count). No cost_per_pr value is
uploaded. Events that omit pr_outcome, including v0.6 through v1.0 uploads,
remain valid.
The code-mower cloud pr-outcomes command joins local builder_run events
(from .code-mower/builder-runs/*.cloud-event.json), reviewer_run events
derived from reviewer-spend rows, and the live GitHub PR list to produce one
pr_outcome event per PR. Attempts are deduplicated by their source identity
(event_id, or dimensions.spend_run_id for converted reviewer-spend rows)
for cost totals, but attempts with missing or duplicate source identities are
still counted as expected attempts with unknown cost so they cannot inflate
complete coverage. Outcome event identifiers are stable for unchanged
observation content, making repeated uploads idempotent, while a deterministic
digest of the observed run evidence is reported as
dimensions.pr_outcome_observation_version and included in the event_id.
The observation-version dimension is optional: historical valid
code_mower.prOutcome.v1 events predate it, so consumers must accept its
absence, though a present value must be a non-empty string.
The command records a local metadata-only observation state
(.code-mower/pr-outcome-observations.json, fingerprints and timestamps only)
so an unchanged retry reproduces the same created_at and event_id, and
corrected or late-arriving local evidence produces a new event_id whose
created_at never regresses below the previously emitted observation — even
when GitHub updatedAt and the run timestamps did not advance. Cost coverage is
complete when every observed builder/reviewer attempt reports cost_usd,
partial when at least one but not all attempts report cost, and unknown
when no attempt reports cost. The optional
dimensions.missing_cost_sources list names observed lane/provider
identifiers that did not report cost; it is metadata-only and must never
contain commands, paths, auth output, or secrets. Missing cost is omitted,
never serialized as zero.
Builder evidence fails closed: a *.cloud-event.json file that cannot be
read, parsed, or recognized as a builder_run event is never silently
omitted. A failure attributable to a PR via its filename is recorded on that
PR as an expected attempt with unknown cost under the fixed
unreadable-evidence source label and surfaced as a bounded per-PR error;
a failure that cannot be attributed suppresses complete coverage for every
emitted outcome. Diagnostics carry PR numbers and fixed labels only — never
paths or file contents. If the observation-state file cannot be persisted,
the command aborts before export/upload so a later correction cannot tie on
created_at.
The event contains identifiers, timestamps, categorical outcomes, and numeric counts/cost only. Its dimension and metric names are closed in v1, so undeclared fields are rejected rather than becoming accidental prose channels. It must not contain PR or issue prose, source, diffs, prompts, transcripts, issue body text, raw stdout/stderr, auth output, local paths, or secrets.
reviewer_finding_outcome is the additive blocker-scoped disposition event for
reviewer evidence and confirmed-catch claims. It uses the normal
code_mower.benchmarkEvent.v1 envelope with
dimensions.reviewer_finding_outcome_schema=code_mower.reviewerFindingOutcome.v1.
One event describes one finding outcome observation; the stable opaque
finding_id plus repo_slug, pr_number, head_sha, and lane_id form the
finding identity. Producers should make retries idempotent with the same
event_id. When a newer observation uses another event id, consumers select the
latest observed_at observation per finding before aggregating.
Required dimensions are reviewer_finding_outcome_schema, opaque finding_id
(at most 160 characters), repo_slug (in OWNER/REPO format), positive
integer-string pr_number, head_sha (7-64 characters), lane_id (at most 80
characters), severity (blocker, major, minor, or info), disposition
(accepted_fixed, false_positive, accepted_risk, owner_decision,
duplicate, infrastructure, or insufficient_context), ISO 8601 observed_at
with a UTC offset, and source (automated or manual). Optional dimensions
are resolved_at (which cannot precede observed_at), fix_commit_sha (7-64
characters, required for accepted_fixed disposition), and decision_id (at
most 160 characters, required for owner_decision disposition).
Metrics are atomic: finding_outcome_count is always 1. No precomputed
dashboard rates are uploaded.
The event contains identifiers, timestamps, severity, disposition, and optional
linkage only. Its dimension and metric names are closed in v1, so undeclared
fields are rejected rather than becoming accidental prose channels. It must not
contain finding prose, titles, details, source, diffs, prompts, transcripts,
issue body text, raw stdout/stderr, auth output, file paths, local paths, or
secrets. Events that omit reviewer_finding_outcome, including all uploads
before this contract, remain valid.
Work-type dimensions are additive, versioned metadata that may be attached
only to builder_run, reviewer_run, work_order, and productivity_summary
events so Dashboard 2.0 can compare builders and reviewers by development
work type. They use dimensions.work_type_schema=code_mower.workType.v1.
Events that omit these dimensions, including all uploads before this
contract, remain valid; validation only runs when work_type_schema is
present. work_type_schema on any other event type (for example
workflow_run or dogfood_upload) is rejected outright.
dimensions.work_type is one of web, backend, ios, macos, android,
infrastructure, documentation, or unknown. dimensions.work_type_source
records classification precedence: explicit_user metadata wins first,
deterministic repository_metadata (for example a GitHub primary language)
or coarse file_category_metadata (a bucket label such as web-frontend,
never a filename) is checked second, and unknown otherwise. work_type
unknown requires source unknown or explicit_user.
dimensions.work_type_role is event-shape-pinned, not free text:
builder_run events must record work_type_role=builder and
reviewer_run events must record work_type_role=reviewer; any other value
on those event types is rejected rather than silently reattributed.
work_order and productivity_summary events leave work_type_role
optional and it is never guessed — when absent, work_type_attribution must
also be absent.
When work_type_role is present it stays distinct from
dimensions.work_type_attribution, which is builder_credit,
reviewer_credit, or excluded_self_review. Builder role requires
builder_credit; reviewer role must use reviewer_credit or
excluded_self_review. When work_type_lane_id equals
work_type_builder_lane_id, an author lane cannot count as independent
review: the reviewer role must record excluded_self_review rather than
reviewer_credit.
Optional work_type_provider and work_type_model mirror the same
provider/model identity already carried on tool. When both the work-type
identity and tool.provider/tool.model are present, they must agree after
the same whitespace-collapsing normalization tool provenance already uses;
disagreement is rejected. Either side may be omitted, which keeps
provider/model-free work-type events backward-compatible.
Work-type metadata must not include filenames, source, diffs, prompts, or issue text; repository and file-category inputs are already-coarse category labels, never paths.
adoption_run is the additive atomic event for release-qualification
campaigns. It uses the normal code_mower.benchmarkEvent.v1 envelope with
dimensions.adoption_run_schema=code_mower.adoptionRun.v1. One event describes
one code_mower.adoptionResult.v1 observation produced by
code-mower release qualify; the release identity is release_tag plus
normalized_version, qualification_context, and provider. The converter
derives a deterministic event_id from the source result content, so retrying
an export/upload of the same result file reuses the same event id and stays
idempotent. Newer observations of the same campaign use another event id, and
consumers select the latest created_at observation per release before
aggregating.
Two local routes produce these events, and both go through one converter:
code-mower cloud dogfood --event adoption_run=path/to/result.json (and cloud export) converts one result file at a time, while code-mower release campaign upload converts every terminal (complete or blocked) provider result a
campaign holds -- passing and failing evidence alike. When a campaign uses
compatibility-default posture (older campaigns and campaigns created without
--required-providers), both routes omit provider_posture so the same result
yields the identical event shape and event id across routes. When a campaign
explicitly configures posture via --required-providers, the campaign route
records that posture in the optional dimension, intentionally yielding a
posture-specific stable event id. The campaign route previews by default and
posts only with --yes, using the identical event set both times; providers
without terminal evidence are counted as skipped, and a terminal provider
whose stored result no longer validates stops the upload with a bounded error
instead of publishing a partial set. A terminal state that contradicts its
bound result outcome (complete with a failing/incomplete outcome, or
blocked with a passing one) is likewise rejected with the bounded
adoption_result_state_mismatch reason, never normalized into evidence.
Required dimensions are adoption_run_schema, release_tag
(v<major>.<minor>.<patch>[-<stage>.<num>]), package_identity
(code-mower), normalized_version, qualification_context
(cold_install, upgrade, or unknown), provider, executor,
host_class (local, ci, github_actions, or unknown),
runtime_class (python_<major>.<minor> or unknown), execution_state
(planned or executed), outcome (pass, pass_with_warnings, fail,
or incomplete), result_timestamp (ISO 8601 with a UTC offset), and
provenance_coverage (complete, partial, or unknown). Optional
dimensions are starting_version and ending_version (which must be empty or
normalized versions), and provider_posture (whose only accepted values are
required and informational). Standalone/file-export conversion, older
campaigns, and campaigns created without --required-providers omit
provider_posture and preserve the exact existing event shape and event id.
Campaigns with explicitly configured posture supply provider_posture,
incorporating it into deterministic event identity so the same result used under
different postures cannot collide. Tag and spec versions must
agree: the tag-derived normalized version must equal normalized_version.
Upgrade context requires a starting_version lower than the target; other
contexts must leave it empty. Executed runs must not report incomplete, and
planned runs must report incomplete or fail. Complete provenance coverage
requires a known provider, executor, host class, and runtime class.
Metrics are atomic values, never precomputed dashboard rates:
adoption_run_countis always1;step_countis the observed step total, andstep_pass_count,step_warn_count,step_fail_count,step_unavailable_count, andstep_planned_countare non-negative integers that must sum tostep_count.elapsed_secondsis the observed qualification wall time and must be finite and non-negative. Step timings must sum to within 1.0 second of this total (rounding/overhead tolerance).warning_countandowner_action_countare non-negative integer summaries across steps. Apassoutcome requires zero owner actions, matching the local adoption-result contract.
Conversion applies the same semantic validation as local campaign paths:
executed-result timestamp bounds (not older than 2020-01-01T00:00:00Z, not
more than 300 seconds in the future; planned previews exempt), the built-in
step taxonomy (board, doctor, lanes_status, overhead,
package_install) or an
explicit <namespace>__<name> provider extension, and the timing and
owner-action rules above. Rejections use bounded errors that never echo result
content, paths, auth output, or raw provider output.
Missing model, token, cost, and optional measurements stay unavailable and
omitted, never zero-filled: the closed metric set contains no cost, token, or
model metrics, and the reporter tool provenance uses
model_source=not_applicable because no AI model generated the operational
event. No cost_per_run or dashboard rate is uploaded.
The event contains identifiers, coarse environment classes, categorical
outcomes, and numeric counts/timings only. Its dimension and metric names are
closed in v1, so undeclared fields are rejected rather than becoming accidental
prose, path, or output channels. Dimensions must also stay single-line and
path-free. It must not contain report text (never uploaded by default),
command output, source, diffs, prompts, transcripts, issue body text, raw
stdout/stderr, auth output, local paths, or secrets. Events that omit
adoption_run, including v0.6 through v1.0.4 uploads, remain valid, and
CodeMower.com must keep accepting uploads that omit this event type.
productivity_summary is the v1.0.1 aggregate metric event for local reports
and CodeMower.com dashboards. It uses the normal
code_mower.benchmarkEvent.v1 envelope with
dimensions.productivity_schema=code_mower.productivityMetrics.v1. It is an
aggregate snapshot, not a raw trace: dashboards may group it by repository,
lane, provider, builder, reviewer, issue, PR, or release without receiving
source, diffs, prompts, transcripts, issue body text, raw stdout/stderr, auth
output, local paths, or secrets.
Required dimensions:
productivity_schema:code_mower.productivityMetrics.v1;repo_slug:OWNER/REPO;window_startandwindow_end: UTC timestamps for the measured window;window_granularity:cycle,day,week,release, orcustom; andaggregation_subject:repo,lane,provider,builder,reviewer,issue,pr, orrelease.
Optional dimensions are intentionally small and metadata-only:
aggregation_key, lane_id, provider, builder_provider,
reviewer_provider, role, issue_number, pr_number, branch, release,
pilot_posture, policy_state, merge_state, verdict, manual_outcome,
automated_vs_manual, owner_action_reason, and event_source.
Metric names and units are stable:
- Time metrics use seconds:
cycle_time_seconds,active_time_seconds,wait_time_seconds,queue_wait_seconds,time_to_first_review_seconds,time_to_green_seconds,time_to_merge_seconds, andowner_wait_seconds. - Count metrics use integer counts:
builder_run_count,reviewer_run_count,audit_pass_count,audit_blocked_count,reviewer_catch_count,blocking_bug_count,blocked_finding_count,false_blocker_count,missed_blocker_count,fix_round_count,owner_intervention_count,manual_override_count,automerge_eligible_count,automerge_requested_count,automerge_completed_count,merged_pr_count,abandoned_pr_count,reverted_pr_count,checks_failed_count,checks_passed_count, andpost_merge_defect_count. - Cost metrics use US dollars:
cost_usd. - Token metrics use integer token counts:
input_tokens,output_tokens,total_tokens,cached_input_tokens, andreasoning_tokens.
Missing metrics mean unknown, not zero. Derived percentiles, rates, and
rankings are report/dashboard calculations and should reference the source
metric names used to compute them. The OSS validator rejects unknown
productivity_summary metric names, negative or non-finite values, and
fractional count/token values so future producers do not drift from the
contract accidentally. CodeMower.com must continue accepting v0.9.x and v1.0
uploads that omit productivity_summary.
Consumers may total count, token, and cost metrics across multiple
productivity_summary events only within one headline aggregation subject. The
recommended headline subject priority is repo, then release, issue, then
pr. Time metrics describe a measured window and must stay latest-window or
unknown unless a producer emits an explicit aggregate window event. Provider-,
builder-, reviewer-, and lane-scoped productivity_summary events are
scorecard inputs; consumers should not add them into headline repo/release
totals for the same window. Scorecard promotion recommendations remain advisory
until reviewed against docs/lane-promotion-policy.md.
code_mower.productivityWindow.v1 is the additive local observation shape for
deterministic normalized repository/release productivity windows (issue #738).
Operators derive one observation file per window from GitHub lifecycle metadata
(PR opened/merged/closed and check timestamps, revert references) plus local
controller/audit timing when available, using numeric aggregates only:
code-mower cloud dogfood --event productivity_summary=window.json --json
code-mower cloud export --event productivity_summary=window.json --repo-slug OWNER/REPO --json
code-mower cloud repo-sync --repo OWNER/REPO=/path/to/repo \
--event productivity_summary=window.json --jsonThe OSS uploader converts each observation into a productivity_summary event
with dimensions.productivity_window_schema=code_mower.productivityWindow.v1.
An observation that omits repo_slug is filled from --repo-slug (export)
or the detected repo (dogfood, and repo-sync per synced repo); an explicit
observation slug that disagrees with the repo-sync target is rejected.
The converter separates elapsed time (cycle_time_seconds, always the
window_end minus window_start span), observed active agent time
(active_time_seconds), queue/wait (queue_wait_seconds, with an explicit
wait_time_seconds aggregate only when the operator supplies one), review
(time_to_first_review_seconds), time-to-green (time_to_green_seconds),
merge (time_to_merge_seconds), and owner-wait (owner_wait_seconds).
PR counts, fix rounds, interventions, reverts, and post-merge defect linkage
are included only when explicitly observed in the input.
Missing values stay unavailable, never zero: unobserved timings and counts are
omitted, and explicit active_time_coverage/defect_coverage dimensions
(observed or unavailable) record what was measured so incomplete
active-time data cannot be presented as complete. comparison_basis
(code_mower_window, pre_code_mower, operator_selected, or unknown)
supports pre-Code-Mower or operator-selected comparison windows, and every
windowed event carries causal_claim=none: before/after deltas are
correlation context, never causal proof. timing_provenance
(github_lifecycle, github_lifecycle_and_local_timing,
operator_supplied, or unknown) records where the numbers came from.
Windowed events use a closed dimension vocabulary and repo/release subjects
only, so undeclared fields are rejected rather than becoming accidental prose
channels. The event id is a deterministic UUIDv5 over the canonical window
content and created_at is the window end, so repeated syncs over the same
observation re-emit the same event id and bytes; only changed observation
content yields a new event. Repo-sync forwards --event entries into each
repo's dogfood step and summarizes window coverage under
productivity_baseline in data_class_summary, alongside the existing
current-dogfood, imported-history, and reviewer-evidence classes.
Observations and events must not contain source, diffs, prompts, transcripts,
issue bodies, raw output, local paths, auth output, or secrets. Events that
omit the window stamp, including all uploads before this producer, remain
valid, and CodeMower.com must keep accepting uploads that omit these
dimensions.
Local code_mower.authoringRun.v1 artifacts from builder-experiment run may
also be passed as --event builder_run=PATH. The OSS uploader converts them to
the normalized builder_run event shape and uploads only metadata such as
provider, model, branch/PR, elapsed time, command hash, and
command_output_capture: disabled; local privacy markers and executor details
that use raw-output vocabulary are not uploaded as event keys.
Calibration and value-report uploads may add optional automated-vs-manual
metadata to reviewer summaries and reviewer_run-shaped rows:
manual_outcome (pass, blocked, or unknown), automated_vs_manual
(match, missed_blocker, false_blocker, or unknown), and aggregate
profile counters such as auto_manual_match_runs,
auto_manual_missed_blocker_runs, and auto_manual_false_blocker_runs. These
fields compare automated reviewer status with manual/adjudicated calibration
truth. They are additive and CodeMower.com must continue accepting beta.40
uploads that omit them. They must not include plan text, issue body text,
source code, raw diffs, prompts, transcripts, stdout/stderr, auth output, or
secrets.
Plan-context audit prompts only read manifest-listed documents/previews that
resolve inside the repository root. Default manifests are read from the trusted
base ref rather than mutable working-tree files; explicit manifest paths are
operator-pinned. The Codex wrapper runs codex exec review --base without a
supplemental stdin prompt; trusted plan and decision context are passed to the
structured verdict conversion prompt.
Auto-inferred builder_run events may add metadata-only dimensions such as
auto_inferred, builder_inference_confidence, builder_inference_signals,
and pr_author. The inference signals are marker names only, for example a bot
author, branch prefix, or detected hosted-agent URL marker; the PR body text and
footer text used for inference are not stored.
Cursor inference only accepts Cursor agent/background-agent URLs or explicit
Cursor-agent footer markers; generic cursor.com links are ignored.
When metadata signals disagree, the highest-priority provider signal wins and
provider-specific run URLs are emitted only for that winning provider.
Audit CLIs may also append local spend rows to reviewer-spend.json using
schema code_mower.reviewerSpend.v1. The file remains backward-compatible with
beta.40 aggregate files that only contain profiles; new clients add an
append-only runs list. Each run may include run_id, created_at, lane,
repo, pr_number, head_sha, model, wall_seconds, verdict,
cost_usd, and token counters such as input_tokens, output_tokens,
total_tokens, cached-input counters, or reasoning_tokens. These are
metadata-only fields. The ledger must not contain source, diffs, prompts,
transcripts, stdout/stderr, issue bodies, auth output, or secrets.
code-mower cloud export --spend reviewer-spend.json, cloud dogfood, and
cloud reviewer-runs convert spend runs into reviewer_run events.
reviewer-runs reads .code-mower/reviewer-spend.json automatically when the
ledger is present and merges spend metrics into matching verdict events before
upload so dashboards do not double-count reviewer attempts. When --verdicts
narrows the exported artifacts, unmatched spend rows are ignored by default; use
--include-unmatched-spend only for deliberate reviewer-spend backfill. The
derived event places latency/cost/token numbers under metrics, PR/SHA/lane
identifiers under dimensions, and model/tool identity under tool.
CodeMower.com should accept uploads without these fields from beta.40 clients
and treat missing spend rows as unknown, not zero measured spend.
Generated self-hosted local audit workflows may automatically call
cloud reviewer-runs and cloud dogfood after audit attempts when a team
configures CODE_MOWER_CLOUD_TOKEN; runner-temp spend rows travel with the
reviewer-run upload. This is an upload-path change, not an
event-shape change: verdict artifacts still become reviewer_run events,
spend rows still become reviewer_run events, work-order sidecars remain
work_order events, and beta.40 through beta.46 uploads that omit these
automated events remain valid. When a verdict artifact records audit runtime,
the exported reviewer-run metrics include the legacy duration_seconds_total
and dashboard-compatible duration_seconds and wall_seconds aliases. The
workflow uses trusted default-branch support files and must not upload source,
diffs, prompts, transcripts, stdout/stderr, issue body text, or secrets.
Fixture-shaped or quarantined audit verdict artifacts are excluded from
reviewer_run export and upload so local wrapper tests cannot become dashboard
or calibration evidence.
Before conversion, provider verdict artifacts are checked for the local
code_mower.auditVerdictArtifact.v1 shape: repository, PR number, verdict, and
comment body are required, known verdict values are enforced, and timing fields
must be finite non-negative metadata when present.
reviewer_run events may include per-lane audit comment attribution in
dimensions.audit_comment_lane_id, dimensions.audit_comment_identity_source,
and dimensions.audit_comment_trailer_prefix. When present,
audit_comment_identity_source=trailer means the hidden audit-state trailer
was the authoritative lane signal for that result. CodeMower.com must continue
to treat beta.40 uploads that only include dimensions.lane_id as valid
metadata-only reviewer evidence.
Each event may also include a tool object using schema
code_mower.toolProvenance.v1. This object is the benchmark-grade provenance
surface for AI tool/version/model data:
role:builder,reviewer,workflow, or another explicit lane role;tool_nameandtool_version: the local CLI, GitHub App, hosted reviewer, or agent surface that produced the event;provider,model, andmodel_version_raw: the AI provider/model identity when known;model_source: where the normalized model identity came from, such asenv,profile:<name>,default,vendor_hidden,not_applicable, ormissing;version_source: where the tool/package version came from, such ascli_version_probe,package_version,vendor_hidden,not_applicable,not_probed, ormissing;integrationandruntime_environment: for examplecli,github_app,hosted,local, orgithub_actions; andlensandprompt_pack_version: the review lens/prompt bundle that shaped the run.
Model identity can come from explicit environment configuration, the selected
Code Mower provider profile, a safe default, safe provider metadata, or
structured provider summary stats. For example, Google-compatible CLI summaries
may report multiple internal models; Code Mower records the main review model
when it can identify one, and leaves the model blank when it cannot do so
safely. CodeMower.com should display model_source and version_source
alongside tool/model rows so benchmark readers can tell the difference between
exact configured provenance, profile-derived provenance, defaults, and missing
metadata.
For local CLI lanes, prefer explicit model environment variables such as
CODE_MOWER_CODEX_MODEL, CODE_MOWER_GEMINI_MODEL, or the provider's native
model variable when the CLI does not expose model identity through safe
metadata. Model identifiers are benchmark metadata, not secrets.
Hosted/manual reviewer lanes may report model_source=vendor_hidden and
version_source=vendor_hidden when the review service does not expose the
underlying model or app version. That is known provenance about the provider
surface, not a configured model id. Code Mower's own reporter events use
model_source=not_applicable because no AI model generated the operational
event. Local CLI or API lanes that omit a model remain missing until the user
configures the relevant model environment variable, provider profile, or
default.
Code Mower treats missing tool/model provenance as acceptable for operational dogfood, but incomplete for benchmark claims. CodeMower.com therefore displays provenance coverage separately from upload volume.
The OSS client fails closed for structured events that contain unsafe field names such as raw output, transcripts, tokens, secrets, auth previews, or secret-like values. Fix the event producer instead of relying on cloud upload to silently scrub sensitive data.
CodeMower.com uses team ingest tokens for upload authorization. Users create or receive a token, then store it locally with:
code-mower cloud setup \
--token-stdin \
--team-id "your-team-slug" \
--install-id "your-install-id" \
--out ~/.config/code-mower/tokens/your-install-id.envcloud setup writes a sourceable 0600 env file and records it as the current
local cloud profile. Upload commands resolve tokens from the live env first,
then explicit --token-file, then --install-id, then the current profile, then
one unambiguous stored profile. If multiple stored token files exist without a
current selection, Code Mower refuses to guess and reports filenames only.
The hosted service stores token hashes and short prefixes, not full token values. A token can be revoked without rotating every team credential.
The recommended flow is dry-run first:
code-mower cloud export \
--report value-report=.code-mower/reviewer-value-report.md \
--output-dir .code-mower/cloud-benchmark-bundle \
--anonymous \
--json
code-mower cloud upload .code-mower/cloud-benchmark-bundle --dry-run --jsonNothing uploads unless --yes is supplied:
code-mower cloud upload .code-mower/cloud-benchmark-bundle --yes --jsonThe current hosted service stores upload ids and timestamps, token/team linkage, repository slug when supplied, report summaries and counts, structured metadata events, cost/latency/usefulness fields when supplied, and recommendation inputs derived from metadata.
It should not store source, raw diffs, raw transcripts, stdout/stderr, auth output, or secrets by default.
Current controls:
- uploads are opt-in and dry-run-first;
- team ingest tokens can be revoked;
- full token values are not stored after creation; and
- report text is excluded unless explicitly included by the uploader;
- signed-in team members can export team metadata; and
- team owners/admins can delete uploaded metadata and related summaries/events.
Known gap:
- automated retention jobs and user-configurable retention windows are not implemented yet.
For early adopter pilots, deletion/export basics are live, but broad cloud-data collection should wait until a published retention policy and automated retention jobs are available.
Before broad public adoption, Code Mower Cloud should add retention settings, clearer anonymization/cohort rules, schema migration notes, and public examples of useful aggregate benchmark outputs.