add multi-platform inference framework - #123
Merged
Merged
Conversation
Prepare the repo for the four-platform story (Python / macOS / iOS / Android) ahead of the unified inference framework: - move the Python framework (pyproject, src, tests, scripts, examples) into python/; add python/README.md (pyproject readme target); docs stay at the repo root - reserve macos/, ios/, android/ platform directories; their READMEs describe the released engines (source drops 2026-09-30) - CI: run both jobs with working-directory python/ - hygiene suite: scan from the repo root so it covers all platform dirs; exclude .pdf; MLX-boundary check follows python/src - fix(edge0 convert-adapters): resolve scripts/ via parents[2] — the command pointed at a non-existent src/scripts since before the move - .gitignore: python/scripts/alignment_results.json Verified: pytest -m 'not slow' → 79 passed, 1 skipped (fresh venv, editable install from python/). Co-Authored-By: Claude Code <noreply@anthropic.com>
Reorganize both READMEs as News -> About -> Getting Started -> Roadmap -> Contributing -> Citation -> Contact Us: - News: 2026-09-30 three-platform engine release (iOS / macOS / Android, source open-sourced in-repo), --prefill-ondemand, the arXiv report, and the initial release - About: core mechanisms, a platform matrix (Python available now, three engines open-sourced, Windows on the roadmap), models, design (backend isolation scoped to the Python framework), quality and benchmark tables carried over verbatim - Getting Started: full Python walkthrough on the new python/ paths; macOS / iOS / Android point at each directory's README - Roadmap: the unified inference framework (one access layer, runtime auto-adapting to iOS / macOS / Android / Windows / Python) by the end of October 2026, plus CUDA backend and Windows - Citation: BibTeX for arXiv 2609.18063 - Contact Us: placeholder channels (email / Discord / WeChat) with a TODO for the team to fill in Co-Authored-By: Claude Code <noreply@anthropic.com>
…and demo.py
End-to-end testing of the README quick-start surfaced three issues:
- edge0 convert-adapters was broken. Its subparser defined no options and
cmd_convert re-exec'd the legacy script via runpy without resetting
sys.argv, so the subcommand name leaked into the script's argparse
("unrecognized arguments: convert-adapters"), while --npz-dir / --force
were rejected by the top-level parser. Define --force / --npz-dir on the
subparser and delegate to the legacy main(argv) with a clean argv.
- examples/demo.py --model <tier> raised FileNotFoundError because tier
env-var resolution lived only in the CLI. Resolve $EDGE0_<TIER>_MODEL the
same way; an explicit --model-dir still wins, and a missing env var now
fails with a clear message.
- The edge0-35b / edge0-8b profiles overrode the serving port (8085 / 8083)
while the CLI defaulted to 8000. Drop the overrides so every tier
inherits the unified port 8000; update the profile assertions and the
docs/models/*.md serve/curl examples to match.
Also ignore downloaded model checkpoint dirs (/models/, /python/models/) so
multi-GB weights are never committed; the pattern is anchored to avoid
matching the source package python/src/edge0/models/.
Verified: convert-adapters end to end on synthetic npz (2 sources -> 4
safetensors, key normalization, owners, 35b language_model. prefix,
idempotent skip, --force rewrite); pytest -m 'not slow' -> 79 passed.
Co-Authored-By: Claude Code <noreply@anthropic.com>
- Drop the section-heading emoji (News / About / Getting Started / Roadmap / Contributing / Citation / Contact Us) and repair the in-repo anchors that the removal invalidated (#-roadmap -> #roadmap) in both root READMEs and the macos / ios / android placeholder READMEs. - Add the required "model" field to the /v1/chat/completions curl example; the server rejects requests without it (surfaced during testing). - Fill in the Contact Us email; sync the Chinese README with all of the above. Co-Authored-By: Claude Code <noreply@anthropic.com>
Split the Roadmap into "Platforms & systems" (the existing unified framework / Windows / CUDA items) and "Models & algorithms — Q4 2026", summarizing the algorithm team's Q4 plan for a public audience: - Next-gen architecture support (Qwen3.8-Flash class): hybrid linear attention (GDN + QSA), gated multi-branch residual, and N-gram embedding — designs that suit SSD streaming offload (O(1)-state attention, lookup-only N-gram tables). - Latent thinking + batched expert pre-prediction: move expert routing from per-position to once per block so load volume decouples from the reasoning loop count, with cross-block prefetch; measured as end-to-end thinking-phase time at matched accuracy. Internal research detail (training configs, supervision ablations, quantified speedup targets, weight-availability contingencies) is intentionally omitted. EN and ZH READMEs updated in sync. Co-Authored-By: Claude Code <noreply@anthropic.com>
… un-anchored) The bare `models/` rule in windows/.gitignore matched at any depth, so the frontend source module windows/app/src/models/cards.tsx was silently dropped at `git add` time on fresh clones -> tsc TS2307 on models.tsx/welcome.tsx, `npm run build` (and thus tauri build) impossible from a clone. Anchor to /models/ (the local model download dir) and add the source file. Verified in a git-archive sandbox: typecheck passes with the file present.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
add inference framework for ios,mac,android and windows