Skip to content

RLCR feedback: pre-review audit needs a correctness-family checklist + verify declared CI gates locally (4 suggestions) #227

Description

@ZziTaiLeo

RLCR Methodology Analysis

Scope: a single-file automation deliverable targeting a constrained legacy runtime, with an
offline behavior-test suite and a static-analysis CI gate. Analyzed purely as process; all
domain, code, and identity details are deliberately omitted.

Session Shape (sanitized)

  • 1 build round implementing the entire plan in one pass, containing an in-plan
    adversarial self-audit sub-step that surfaced 6 must-fix items and folded them in before
    any external review. The build also self-found and fixed 2 runtime-trap bugs and captured
    them once as a reusable lesson.
  • 3 implementation-review rounds (build round + 2 follow-ups), each closing exactly 1
    semantic gap, converging to full acceptance-criteria coverage.
  • 4 code-review-phase rounds with strictly decreasing severity: a production-credential
    exposure → rerun/idempotence semantics → non-actionable guidance in error messages →
    passability of the static-analysis gate.
  • Finalize: minimal cleanup (one dead parameter removed), tests reconfirmed.

Test count grew monotonically (~25 → ~37), no regressions, no reopened items, stall count 0.

Verdict: the methodology worked well

Iteration efficiency, feedback quality, and communication were strong. The findings below are
a small number of real gaps, not a systemic critique.

What worked:

  • Front-loaded adversarial audit paid off. The in-plan self-audit converted a prior-session
    lesson (an evidence standard discovered serially, round by round) into a proactive step. It
    removed 6 issues pre-review; the implementation-review phase then needed only 2 short rounds
    to reach full coverage. This measurably reduced round count versus the serial-discovery baseline.
  • Near-zero false positives. Across 7 review rounds every finding was a genuine defect, each
    verified by actually running commands and citing observed output — not speculative. Signal-to-
    noise was excellent.
  • Tight scope control. Every round declared a single objective plus an explicit "do not do"
    / queued list, so deferred items were tracked, never silently dropped, and never caused drift.
  • Anti-drift tracking. An immutable goal/criteria section plus a mutable evolution log kept a
    persistent anchor across all rounds; each change mapped back to specific criteria.
  • Clear, templated communication. Summaries and reviews followed stable structures
    (implemented / files / validation / remaining), making round-to-round progress legible.

Gaps + concrete RLCR improvements

1. The pre-review audit missed a whole correctness family, costing 3 serial rounds.
Observed: the self-audit covered the leak/injection/permission threat family well, but three
later code-review rounds shared one root cause — "a default value is not the same as an unused
value" for stateful/persistent resources (already-initialized state, load-bearing defaults,
credentials on externally reachable surfaces). These were discovered one per round rather than
batched.
Improvement: give the in-plan adversarial audit an explicit correctness-family checklist to
enumerate against — at minimum "persistent resource already initialized," "default value that is
nonetheless load-bearing," and "documented remediation must be executable." Add a root-cause
clustering pass so issues sharing a cause are fixed in one round, not serialized.

2. Security/usage-correctness review was gated behind coverage review.
Observed: two review lenses ran in sequence — acceptance-coverage first, real-usage correctness/
security second — and the second lens found the highest-severity issue (a credential exposure).
A P1 was therefore discovered only after a whole phase of coverage review.
Improvement: run a lightweight usage-correctness/security pass concurrently with the first
review rounds. Do not let a coverage-complete verdict precede any security scrutiny; the highest-
severity class should be probed earliest, not last.

3. A declared verification gate went unrun for the entire loop.
Observed: the deliverable included a static-analysis CI gate, but the tool "wasn't installed
locally," so passability was deferred to CI — and surfaced as a P1 only in the final review round.
The gate was assertable all along via a one-shot/ephemeral tool runner.
Improvement: treat any acceptance criterion that names a CI/gate as blocking until executed
locally in the round that introduces it
. "CI will check it later" must count as an unmet
verification, i.e. a review finding — never as evidence.

4. The true end-to-end path was proven by proxy, not exercised.
Observed: the full startup path could not be stood up (environment/port collision with the
working environment), so it was honestly recorded as a limitation and covered by command-
construction assertions instead. Good transparency, but "does it actually start once" stayed
unproven.
Improvement: for provisioning/bootstrap-type deliverables, budget an isolated ephemeral
environment (throwaway namespace/ports) so the real end-to-end path is exercised at least once,
rather than substituted entirely by indirect assertions.

Bottom line

Efficient loop, high-quality and honest feedback, no stagnation, clean plan-to-execution
alignment. The one structural lesson: the pre-review audit and the review lenses were strong on
the threat families they targeted but blind to a stateful-correctness family and to an unrun
declared gate — both fixable by (a) an explicit correctness-family checklist with root-cause
clustering, and (b) a rule that declared gates and highest-severity classes are probed earliest
and only counted as done when actually executed.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions