What happens today
core.check_certainty(...) hands you back 0.83. core.requirement_check(...) hands you back a float.
Requirement.validate() hands you back a ValidationResult. In each case the value looks the same
whether it came from an aLoRA adapter, a LoRA adapter, a hand-written prompt run against the base model,
or LLM-as-a-judge.
There is no way to ask, for a given result, which of those produced it. ValidationResult carries
result, reason, score, thunk, context and error. The thunk field is populated on every
path, so it cannot be used to tell them apart either, and its docstring still describes it as the output
"produced during LLM-as-a-Judge validation", which stopped being true when requirement rerouting landed.
Why that is a problem
The mechanism behind a score used to be fixed for a given setup. Three changes make it variable:
Mellea's own sampling strategies consume these scores to decide whether to accept a candidate, repair it,
or try again. So a rejection-sampling loop can start making different decisions because the evidence
behind its verdicts changed, and neither the user nor the loop can see that happen. Anyone using these
scores as a gate in a pipeline has the same problem, with the added wrinkle that a prompt-backed score is
weaker than one from trained weights.
How it should behave
A caller who needs to know should be able to ask, per result, which reality produced it. Sampling
strategies should be able to record it. At minimum, the existing route should be documented.
What already exists
backend.resolve_adapter(name).weights.binding_type returns local_file, embedded, server_mediated
or prompt. It is public and works today. It answers "what is registered right now" rather than "what
produced this result", and it is not mentioned anywhere in the docs.
Telemetry has the right information already: the adapter_function_invocation_complete hook payload
carries binding_type and revision. That needs the hooks extra plus a registered plugin, so it serves
operators rather than application code.
Detail: possible shapes
- Document the
resolve_adapter(...).weights.binding_type route. Cheapest first step, and it could land
alongside the prompt-fallback documentation.
- Record the binding type, and ideally the resolved revision, on the result metadata. There is an
established home for this: mot.generation already carries model, provider and usage.
- Give
ValidationResult an explicit provenance field so sampling strategies can log or branch on it,
and correct the thunk docstring while we are there.
Options 2 and 3 touch surfaces shared across every backend, which is why this is filed for discussion
rather than bolted onto the PR that exposed it.
Related
Found while reviewing #1654, which introduces the weightless prompt reality. A worked example of the
mechanism behind a score changing without any signal is #1679.
What happens today
core.check_certainty(...)hands you back0.83.core.requirement_check(...)hands you back a float.Requirement.validate()hands you back aValidationResult. In each case the value looks the samewhether it came from an aLoRA adapter, a LoRA adapter, a hand-written prompt run against the base model,
or LLM-as-a-judge.
There is no way to ask, for a given result, which of those produced it.
ValidationResultcarriesresult,reason,score,thunk,contextanderror. Thethunkfield is populated on everypath, so it cannot be used to tell them apart either, and its docstring still describes it as the output
"produced during LLM-as-a-Judge validation", which stopped being true when requirement rerouting landed.
Why that is a problem
The mechanism behind a score used to be fixed for a given setup. Three changes make it variable:
chosen automatically when no trained adapter exists for the target model (feat(adapters)!: prompt fallback for adapter functions without a trained adapter #1654),
as a side effect of any adapter-function call.
Mellea's own sampling strategies consume these scores to decide whether to accept a candidate, repair it,
or try again. So a rejection-sampling loop can start making different decisions because the evidence
behind its verdicts changed, and neither the user nor the loop can see that happen. Anyone using these
scores as a gate in a pipeline has the same problem, with the added wrinkle that a prompt-backed score is
weaker than one from trained weights.
How it should behave
A caller who needs to know should be able to ask, per result, which reality produced it. Sampling
strategies should be able to record it. At minimum, the existing route should be documented.
What already exists
backend.resolve_adapter(name).weights.binding_typereturnslocal_file,embedded,server_mediatedor
prompt. It is public and works today. It answers "what is registered right now" rather than "whatproduced this result", and it is not mentioned anywhere in the docs.
Telemetry has the right information already: the
adapter_function_invocation_completehook payloadcarries
binding_typeandrevision. That needs the hooks extra plus a registered plugin, so it servesoperators rather than application code.
Detail: possible shapes
resolve_adapter(...).weights.binding_typeroute. Cheapest first step, and it could landalongside the prompt-fallback documentation.
established home for this:
mot.generationalready carriesmodel,providerandusage.ValidationResultan explicit provenance field so sampling strategies can log or branch on it,and correct the
thunkdocstring while we are there.Options 2 and 3 touch surfaces shared across every backend, which is why this is filed for discussion
rather than bolted onto the PR that exposed it.
Related
Found while reviewing #1654, which introduces the weightless prompt reality. A worked example of the
mechanism behind a score changing without any signal is #1679.