Skip to content
Hunter-124Public

About

Local-first Windows AI home assistant in C++/Qt Quick, with voice, memory, camera vision, and optional external-agent delegation.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Polymath — a local-first AI home assistant

Polymath is a Windows desktop AI assistant built in C++ and Qt Quick/QML. It combines local language and speech inference with persistent goals, semantic memory, camera vision, and a voice-first interface.

Local core, optional network features. Core inference runs on your own machine via llama.cpp, whisper.cpp, and ONNX Runtime. Optional external-agent CLI sessions, web search/fetch, and embedded browsing can contact external services and send data outside the machine; model downloads also require network access. Review tool allow-lists and external-provider settings before using those integrations.

The application and codebase use Polymath; Hearth is a legacy workspace alias for the same project, not a separate application (naming context).

UI preview

Polymath chat UI capture with a persona selector and stub conversation

Committed chat UI capture from the GUI audit, rendered offscreen with stub/demo data—not a live inference run or a real camera result. This earlier UI capture illustrates the chat layout; see the capture report for provenance.

What it does

  • Voice loop — VAD-gated wake word → speech-to-text (whisper.cpp, on-demand GPU) → local LLM (llama.cpp) → text-to-speech (Piper, streaming sentences). Push-to-talk or hands-free; barge-in supported.
  • Agent harness v2 — multi-step goals that plan → execute → reflect, persist across turns/restarts, and deliver results to chat + notifications (+ optional TTS). Not just a chat box with leaf tools.
  • Skills — declarative, hot-reloaded workflow bundles (data/skills/*/skill.json) the model can run or author (run_skill / save_skill). Ships with starters such as morning_brief, research_brief, and slop_mode.
  • External agent sessions — spawn/monitor Claude Code (and Codex / generic PTY adapters) as live cards; voice/toast when a session needs input; reply from chat or voice.
  • Surfaces — the AI can compose on-screen content (ui_control + SurfaceHost: placeholder, image, web/video). Real embedded browser via QtWebEngine with adblock + YouTube clean-mode.
  • Holographic UI — frameless shell, aurora glass theme, per-section hues, command palette (Ctrl+K), settings page, notification center + toasts, dashboard HUD.
  • Personalities — drop-in historical personas (persona.json bundles); ships with Marcus Aurelius and Ada Lovelace.
  • Vision — per-camera pipeline over ESP32-CAM / MJPEG: motion gating, person detection (YOLOv8n), face recognition (SCRFD + ArcFace), and “where did I last see …” via a VLM.
  • Memory — long-term semantic memory (vector recall over EmbeddingGemma), daily summarizer, per-category retention — wired into the harness prompts and tools.
  • Toolset — ~25 tools with risk classes: web search/fetch, browser automation, docs & lab reports, print, shopping, reminders/tasks, cameras/who's-home, memory, skills, agent sessions, UI control.
  • Tiered inference — resident Fast model for live voice (4k ctx, q8 KV on 8 GB cards), on-demand Vision/Embedding, honest VRAM budgeter. Heavy local 27B is optional; deep work prefers agent-session delegation on Max-Q GPUs.

Privacy & security. Local inference does not make every feature offline. External CLI providers and web tools have their own network/data policies. Microphone, ambient transcription, cameras, face recognition, and screen capture default ON, with per-feature toggles and a master kill-switch; retention is configurable. SQLCipher builds encrypt the database using a per-install key protected by Windows DPAPI, but encryption can be inactive if the codec or key is unavailable. See docs/PRIVACY.md for retention, enrollment, and the actual encryption boundary.

Hardware note (this project’s target machine)

Budgeted for Intel i7-9750H, 32 GB RAM, RTX 2070 Max-Q 8 GB (sm_75). Older docs that mentioned a 3080 Ti are wrong. Resource table: docs/overhaul/04_VOICE_RESOURCES.md.

Install (end users)

Grab the latest installer from the repo's Releases and run it:

  • Polymath-<version>-win64-cuda-Setup.exe — NVIDIA GPU build (CUDA, much faster).
  • Polymath-<version>-win64-cpu-Setup.exe — CPU-only build (works anywhere, slower).

On first launch with no models, Polymath guides you through a model fetch + a GPU/driver check rather than dropping you into a dead app. Models are not bundled (they're ~GBs); the first-run wizard downloads them. See docs/PACKAGING.md.

Build from source (developers)

Windows 10/11, MSVC 2022, CMake ≥ 3.25. Native engines (llama.cpp, whisper.cpp, SQLCipher, …) are vendored and built from source; Qt 6.6, OpenCV, ONNX Runtime and the small vcpkg libs come from build/deps + vcpkg. Full detail in docs/BUILD.md.

# CPU build (no GPU needed) — configures, builds, runs ctest, deploys runtime DLLs
pwsh scripts/build-cpu.ps1

# Fetch the default local models into build/cpu/bin/Release/data/models
pwsh scripts/fetch-models.ps1            # add -Minimal to skip the big optional ones

# GPU / CUDA build (NVIDIA, sm_75 on this machine). Assumes CPU prereqs exist.
pwsh scripts/build-gpu.ps1               # -> build/cuda/bin/Polymath.exe

# Run
build/cuda/bin/Polymath.exe                # or build/cpu/bin/Release/Polymath.exe

UI-only visual loop (no CUDA, minutes not tens of minutes):

cmake --build build/cpu --config Release --target capture_views
$env:QT_QPA_PLATFORM='offscreen'
build/cpu/bin/Release/capture_views.exe <outdir> [--empty]   # 13 views → PNG

Run the test suite: ctest --test-dir build/cpu -C Release (14 suites: core, tools, audio, agent, vision, inference, memory, privacy, integration, ui, j_phase2, harness, skills, sessions). CI entry point: scripts/ci.ps1.

Package a distributable + build the installer:

pwsh scripts/package.ps1 -Flavor cuda    # stages dist/ + a portable zip
& "$env:LOCALAPPDATA\Programs\Inno Setup 6\ISCC.exe" /DAppVersion=0.3.0 /DFlavor=cuda scripts/installer/polymath.iss

See docs/SHIP.md for the full release checklist. Live status and verify gates: docs/STATUS.md.

Repository layout

src/core/        shared contracts: EventBus, DB/schema (+ SQLCipher), config, privacy, retention
src/inference/   llama.cpp backend, tiered model manager, VRAM budget, GBNF grammar
src/scheduler/   deep-work task queue, idle detector, proactive engine
src/audio/       capture, VAD-gated wake, whisper ASR, Piper TTS (async workers)
src/vision/      camera workers, motion, YOLO, face recognition, visual memory / finder
src/memory/      SQLite store, vector index, daily summarizer
src/agent/       AgentLoop v2, tools (~25), skills registry, personas
src/sessions/    external agent session service + CLI providers
src/personality/ hot-loadable persona bundle manager
src/app/         AppController facade + main.cpp
src/ui/          Qt Quick (QML) holographic shell, views, models, capture_views
firmware/esp32cam/  ESP32-CAM MJPEG streaming firmware + flashing guide
data/skills/     starter skill.json bundles
docs/            architecture, build, models, privacy, packaging, ship, status
docs/overhaul/   2026-07 overhaul plan + PROGRESS ledger (source of truth for that work)

Architecture deep-dive: docs/ARCHITECTURE.md.
Overhaul plan (if contributing to in-flight DAG work): docs/overhaul/00_MASTER_PLAN.md.

About

Local-first Windows AI home assistant in C++/Qt Quick, with voice, memory, camera vision, and optional external-agent delegation.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages