Skip to content

Add Tasksmith PR-to-Harbor generation with LangGraph, Pi and OpenCode - #89

Draft
adithya-s-k wants to merge 74 commits into
mainfrom
codex/dynamic-harbor-curation
Draft

adithya-s-k wants to merge 74 commits into
mainfrom
codex/dynamic-harbor-curation

Conversation

@adithya-s-k

@adithya-s-k adithya-s-k commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • Add Tasksmith, an experimental PR-to-Harbor workflow using LangGraph checkpoints, Pi or OpenCode coding authors, and Modal or Daytona workspaces. It freezes PR pins, explores the change remotely, builds and caches dependencies, establishes reference readiness, and selects a useful task/verifier design.
  • Review the public developer request before private verifier construction. A separate editor can correct it using public evidence only; the approved text is then frozen. Emit offline Harbor tasks with protected deterministic grading, verified solver-to-grader source transfer, and explicit task, episode, verifier, oracle and materialization contracts. Baseline, repeated oracle, incorrect/valid alternatives, tamper and independent challenges precede Sonnet, Opus and adversarial rollouts and an eight-dimension quality review.
  • Bootstrap repair distinguishes missing dependencies from invalid pytest selectors. Pinned source and collection reports use complete size/hash-verified file transfer. A faulty generated smoke permits one exact command correction only after independent evidence review, followed by all readiness checks on the same image/reference.
  • Retain source identity, deadlines, cumulative budgets, artifacts and failed attempts across recovery. Verified partial drafts can be repaired without repeating discovery. Author script failures with concrete execution and cleanup receipts return for construction repair; infrastructure failures earn no behavioral reward. A failed learner observation remains diagnostic while complete protected tests determine its reward. A new revision must pass every admission gate.

Test plan

  • Full local test suite: 2,449 passed, nine optional/live checks skipped. Ruff and diff checks pass. Tests mock inference and providers; target/generated task code executes only remotely.
  • Focused tests for public-only editing, full paged evidence reads, frozen instructions, script-failure classification, protected grading, native author adapters and durable recovery.
  • Minimal-install emitter/runner checks: 73 passed, one optional Harbor integration skipped; the execution-extra CI job exercises the integrations.
  • Remote CPU provider fixtures on both Modal and Daytona: execution, file transfer, blocked external networking and confirmed cleanup.
  • Fixed five-PR POC with complete protected controls, Sonnet/Opus/adversarial trajectories and final review. 0/5 accepted. Two PRs emitted packages. Both completed ordinary control sequences and both solvers earned reward1, but additional independently executed mutations exposed verifier gaps: Accelerate3969 accepts corrupted fixed-size padding; Accelerate3850 accepts a str/bool-only implementation that fails its explicit integer-list example. Both remain held. A further dim1 tensor-order gap is static evidence only. The final Accelerate3850 adversary contains test-context heuristics, but a production failure of its final collected version is unproven; an earlier failing direct call belonged to an earlier draft.
  • Final review: the first dossier exceeded the old bound; the second paged review hit its cumulative reservation limit. A new review-only API authenticated all retained trials and delivered a reversible complete dossier in one bounded request, but the model returned wrong rubric keys. The schema now exposes the exact eight keys; the invalid response remains retained and cannot establish acceptance. Publication holds remain binding independently of model scores.
  • Remaining inputs: PEFT3350 stopped after a second missing test dependency; the repair now requires complete selected-module collection before rebuilding. TRL6206 passed an independently approved pickle-to-dill smoke correction, then hit a CPU/GPU test-configuration mismatch. Futile dependency repair was stopped and provider cleanup verified. A CPU fixture adaptation is integrated: independent review must authorize omitted-use_cpu selection, then every original readiness command and changed-test coverage must pass on the same image. Controller tests and CI pass; live CPU adaptation remains unproven. TRL6116 completed discovery and built its image on Modal, then hit the same bf16/CPU configuration issue. Its CPU assessment spent$0.8039835 before the next paged request exceeded the original$3 allowance. The captured request also assumed Transformers4.56.2 while the image actually used4.57.6; no valid CPU approval or task package was produced. Both remote resources have independently confirmed termination. CPU assessment policy 2 now supplies complete evidence in one structured request and requires an observed installed Transformers version before inference. All 2,449 local tests and current-commit CI pass; no live CPU adaptation has succeeded. Earlier review costs and readable evidence remain required for explicit recovery; the original bootstrap window has since expired. Its stopped construction is reconciled as failed with confirmed cleanup and no retry authorization. The stopped TRL6206 bootstrap was reconciled with confirmed cleanup; its original workspace deadline expired, so it is not being replayed. Original source pins, failed attempts, limits and all spending remain retained.

Out of scope

The initial execution profile is CPU Python source tasks. GPU, native rebuild, multi-service and performance profiles require further conformance work. Provider fixtures do not establish complete Harbor task validity. Earlier curation runs and all charges/reservations remain retained; the original 30-task target has not been achieved. This PR remains a draft until the POC evidence is complete.

Review-only recovery preserves trial receipts and earlier judge costs; normal validation now uses the complete-evidence inline reviewer and commits reversible projection and exact delivery receipts, including invalid responses. Recovery accepts old paged records and verifies the new metadata when present. Both Accelerate lineages and PEFT have reached their original absolute deadlines; this PR does not reset them or count the held tasks.

See Tasksmith documentation and RFC0012 for usage, contracts, recovery and limitations.

@adithya-s-k adithya-s-k changed the title Add dynamic PR curation with LangGraph and Harbor Add cloud PR-to-Harbor curation with Pi, OpenCode and LangGraph Sep 5, 2026
@adithya-s-k adithya-s-k changed the title Add cloud PR-to-Harbor curation with Pi, OpenCode and LangGraph Add Tasksmith PR-to-Harbor generation with LangGraph, Pi and OpenCode Sep 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant