Skip to content

Host Service: DeepSeek + TTS → atomic Programme #45

Description

@zeevenn

This was generated by AI during triage.

Parent

#39 — PRD: AI Radio — Host-driven Channel experience with chat timeline

Goal

Implement the V1 Host generation pipeline. Programme Manager supplies a CandidateSet and bounded Session Context; the Host makes one structured model-planning call, validates the plan in deterministic application code, synthesises retained Interludes, materialises playable local URIs, and atomically returns a complete Programme.

The Host is the domain abstraction. “Agent” describes its implementation style; V1 has no tool loop and does not expose Player controls.

V1 provider decisions

  • Use AI SDK Core with the direct DeepSeek provider.
  • Model: deepseek-v4-flash with Thinking disabled.
  • Use schema-validated structured output; do not parse free-form JSON with a regular expression.
  • V1 has one model provider. DeepSeek and TTS use user-owned BYOK credentials stored outside the app bundle.
  • Inject TTSProvider and audio-cache boundaries. TTS returns audio bytes plus media type; the cache materialises a native file URI or Web Blob URL.
  • The final TTS provider and Channel voice IDs remain gated by the Mandarin blind-listening spike.

Contracts

  • ProgrammeHost.generateProgramme(request) -> Promise<Programme> receives Generation ID, generation reason, Channel, CandidateSet, bounded Session Context, optional Intervention, and AbortSignal.
  • HostPlanner returns a structured HostPlan containing Track IDs and Interlude scripts.
  • Versioned application-owned Host Policy defines duration, ordering, cadence, and validation rules; Channel or user text cannot override it.
  • TTSProvider accepts an Interlude script, Host Voice ID, and AbortSignal, and returns bytes plus media type.
  • InterludeAudioCache materialises and removes generated audio artifacts.

Programme rules

  • Target 30–45 minutes instead of a fixed Track count; a small CandidateSet may produce a shorter Programme.
  • Select only CandidateSet Tracks and never repeat a Track within one Programme.
  • Programme Manager deterministically withholds recently used/current Programme Tracks and reintroduces the least recent Tracks only when needed for sufficient duration.
  • Channel entry and Intervention may start with at most one opening/acknowledgement Interlude. Background refill must not have an opening Interlude.
  • Transition Interludes follow every 2–3 Tracks according to Host Policy, and a Programme cannot end with an Interlude.
  • Track Request takes effect after the current Segment and places the requested Track first, after at most one acknowledgement Interlude.

Failure and lifecycle rules

  • Invalid or empty model output receives bounded retry, then degrades to a deterministic music-only Programme.
  • Each TTS failure receives bounded retry, then only that failed Interlude is omitted.
  • Only the latest Generation ID may commit. AbortSignal propagates to model, TTS, and cache operations; cancellation never triggers fallback.
  • Stale/cancelled generation artifacts are removed.
  • Player receives only a complete Programme whose retained Interludes already have playable URIs.
  • Provider diagnostics may record provider, model, and Host Policy version, but never API keys.

Acceptance criteria

  • Anthropic SDK and hard-coded Claude usage are removed from Host.
  • DeepSeek uses AI SDK Core, deepseek-v4-flash, Thinking disabled, and schema-validated output.
  • No regular-expression JSON extraction remains.
  • Domain validation rejects unknown Track IDs, duplicate Tracks, invalid Segment ordering, excessive Interlude cadence, and plans outside available duration constraints.
  • Track Request, generation reason, Session Context, deterministic recency, cancellation, and stale-generation races have contract tests.
  • TTS and audio-cache boundaries are injected; tests use fakes and no real credentials.
  • A failed TTS Interlude is omitted without making the Programme unplayable.
  • Model failure after bounded retry produces a deterministic 30–45 minute music-only fallback when enough music is available.
  • BYOK credentials use secure device storage and both credentials gate Radio before provider-backed generation is enabled.
  • Store tests and Store/Mobile TypeScript checks pass.

Out of scope

  • Speech-to-text or microphone input
  • Host control of volume, pause, or Track skipping
  • Long-term memory or model-owned conversation history
  • Multiple model providers or user-selectable models
  • Final TTS provider selection and final voice IDs
  • Flow-hosted proxy/API, server-side key custody, billing, or quotas
  • Radio UI redesign and persistent chat history

References

  • Radio and Store Context documents
  • ADRs 0007–0018

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestready-for-agentReady for an AI agent to implement

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions