Skip to content

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

ArduinoLLM-7B

An open-source, LoRA-tuned Qwen2.5-Coder-7B specialist for embedded systems wiring and code generation — Arduino C++, Raspberry Pi Python, MicroPython, and CircuitPython, across 9 boards (ESP32, ESP32-S3, ESP32-C3, Raspberry Pi, Raspberry Pi Pico, Arduino Uno, ESP8266, Teensy 4.0, M5Stack Core2).

Given a component and a board, it produces a wiring table, an ASCII wiring diagram, working code, and a short explanation — in one consistent format.

v2: Corpus-Enhanced Training

v1 was trained entirely on synthetic examples generated by LLMs (Claude, Gemini, and a locally-hosted Qwen2.5-Coder-32B). That approach is fast and controllable, but it has a real ceiling: a model trained only on another model's description of correct wiring can't be more accurate than the model that generated its training data.

v2 adds a genuine second training stage: continued pretraining directly on real library source code — 240,634 real files (~2.59 billion tokens) from the actual Arduino core libraries, micropython-lib, and the Adafruit CircuitPython Bundle. This release trained on the first ~500 million tokens (~19% of the corpus) of that real code, then re-ran instruction fine-tuning on top to restore clean output formatting.

Why this matters

Manual testing of v1 found three concrete, reproducible failure modes:

  • A wiring diagram that internally contradicted its own wiring table
  • A fabricated, non-functional I2C address-conflict "fix" using invented register names
  • A weather-station answer that invented a CO2 reading from a sensor (BME280) that cannot measure CO2

All three were re-tested after v2 training and no longer reproduce.

Measured results (12-question held-out benchmark)

Metric v1 (synthetic only) v2 (corpus + synthetic)
Format compliance 83% 100%
Runtime-contamination-free 100% 100%
Correct runtime API present 92% 92%
Code syntax valid ~100% 100%
Generation speed (RTX 5090) 73.2 tok/s 37.5 tok/s

Known tradeoff: v2 is currently slower than v1 due to being loaded from a locally re-merged checkpoint rather than a pre-optimized quantized repo. A faster-quantized export may follow.

How it was built

  1. Dataset generation — synthetic wiring/code examples generated via the Anthropic API, Google's Gemini API, and a locally-hosted Qwen2.5-Coder-32B (via vLLM), validated for format compliance and runtime contamination. 1,607 unique examples in full_dataset.jsonl.
  2. Continued pretraining — prepare_pretraining_corpus.py chunks real library source into 2048-token blocks; train_continued_pretrain.py runs LoRA rank-32 training with embed_tokens/lm_head unlocked (needed to absorb new vocabulary/patterns from raw code, not just new behavior).
  3. Merge — merge_pretrain_checkpoint.py combines the pretraining adapter into a standalone base model.
  4. Instruction fine-tuning — train_lora.py re-runs standard LoRA fine-tuning (rank 16, attention/MLP only) on top of the corpus-enhanced base, using the same synthetic dataset as v1.
  5. Final merge — merge_final_release.py combines both stages into the single published model.
  6. Evaluation — evaluate_model.py runs the 12-question automated benchmark; manual testing checks specific documented failure modes via ask_model.py.

Licensing note

Built on Qwen2.5-Coder-7B-Instruct (Apache 2.0). Synthetic training data was generated using Anthropic's and Google's APIs, and a locally-run open-weight model — OpenAI and DeepSeek were deliberately not used for data generation, as both explicitly prohibit using their outputs to train competing models in their terms of service. The real-code corpus (Arduino core libraries, micropython-lib, CircuitPython Bundle) is used under each project's own open-source license.

Limitations

  • Reliable on single-component, well-represented combinations; less reliable on genuinely novel multi-sensor combinations or troubleshooting scenarios not well-represented in the training data.
  • v2 trained on only ~19% of the available real-code corpus; the remaining ~81% has not yet been used.
  • Not a substitute for checking your own wiring against a component's actual datasheet, especially for voltage-sensitive components.

Usage

from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="EzioDevio/ArduinoLLM-7B",
    max_seq_length=2048,
    load_in_4bit=True,
)
FastLanguageModel.for_inference(model)

See ask_model.py for a full interactive example.

Releases

Packages

Used by

Contributors

Languages