diff --git a/docs/cli.md b/docs/cli.md index 7dce5523..8e54bda1 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -109,6 +109,12 @@ Available component flags include: Component loading requires a compatible set of encoders, VAE, diffusion model, and optional vision components for the selected model family. +The `--diffusion-model` path can be a **pre-quantized GGUF produced by +`ed-convert`** (see [Pre-quantized GGUF](#pre-quantized-gguf-with-ed-convert)), +so you can quantize just the transformer offline and keep the encoders/VAE in +their original precision — a common way to shrink the largest component while +loading everything else from the standard component files. + ## Text-to-Image @@ -440,33 +446,76 @@ Notes: - Most useful for large models, where on-load quantization can take tens of seconds to minutes while a pre-quantized GGUF loads in seconds. -#### Distilled models - -Both distilled layouts convert and load fine; how the family is recorded differs: - -- **Full-directory distills** (FLUX.1-schnell, SD3.5-Medium-Turbo) convert exactly - like their base model — point `--model` at the directory. The family lands in - the GGUF metadata, so the output loads standalone (`ed-cli --model schnell-q8.gguf`). -- **Transformer-only distills** (Qwen-Image-Lightning, Wan2.1-Distill, - Kontext-Lightning — shipped as a bare DiT `.safetensors` or shard index) - convert the transformer weights alone. `ed-convert` prints - `model version is unknown` because a bare DiT stack carries no config to - identify the family, so the family is **not** written to the GGUF. That is - expected: these distills already require the base model's text-encoder / VAE / - scheduler, so load the GGUF as the `--diffusion-model` on top of the base - directory, which supplies the family: - - ```bash - # convert the distilled transformer weights once - ./build-cuda/bin/ed-convert \ - --model path/to/models/qwen-image-lightning/transformer/diffusion_pytorch_model.safetensors.index.json \ - --type q8_0 --output qwen-lightning-q8.gguf - - # load it on top of the base model (base provides text encoder / VAE / scheduler) - ./build-cuda/bin/ed-cli --backend cuda \ - --model path/to/models/qwen-image --diffusion-model qwen-lightning-q8.gguf \ - --prompt "..." --steps 4 --cfg-scale 1.0 --output out.png - ``` +#### Per-component quantization (quantize the transformer, load by components) + +The transformer is by far the largest component, so a common workflow is to +**pre-quantize only the transformer** and load it together with the original +encoders / VAE via component flags — no need to convert the whole model. +`ed-convert` reads the transformer's `config.json` to detect the family and +writes it into the GGUF, so the quantized transformer loads standalone as a +`--diffusion-model`. + +```bash +# 1) quantize just the transformer to a portable q8_0 GGUF (q4_k also works) +./build-cuda/bin/ed-convert \ + --model /path/to/stable-diffusion-3-medium/transformer/diffusion_pytorch_model.safetensors \ + --type q8_0 --output sd3-dit-q8.gguf + +# 2) load it by components: quantized DiT + original VAE + CLIP encoders +./build-cuda/bin/ed-cli --backend cuda \ + --diffusion-model sd3-dit-q8.gguf \ + --vae /path/to/stable-diffusion-3-medium/vae/diffusion_pytorch_model.safetensors \ + --clip_l /path/to/stable-diffusion-3-medium/text_encoder/model.safetensors \ + --clip_g /path/to/stable-diffusion-3-medium/text_encoder_2/model.safetensors \ + --no-t5 --prompt "a glass teapot on a wooden table" \ + -W 1024 -H 1024 --cfg-scale 5.0 --flow-shift 3.0 --output sd3.png +``` + +Qwen-Image works the same way — its transformer ships as a shard index, which +`ed-convert` accepts directly. Qwen loads its text encoder through `--llm`: + +```bash +# 1) quantize the Qwen-Image transformer (shard index) to q8_0 +./build-cuda/bin/ed-convert \ + --model /path/to/Qwen-Image/transformer/diffusion_pytorch_model.safetensors.index.json \ + --type q8_0 --output qwen-dit-q8.gguf + +# 2) quantized DiT + original VAE + LLM text encoder +./build-cuda/bin/ed-cli --backend cuda \ + --diffusion-model qwen-dit-q8.gguf \ + --vae /path/to/Qwen-Image/vae/diffusion_pytorch_model.safetensors \ + --llm /path/to/Qwen-Image/text_encoder/model.safetensors.index.json \ + --prompt "a glass teapot on a wooden table" \ + -W 1024 -H 1024 --cfg-scale 4.0 --output qwen.png +``` + +Alternatively, `ed-convert` can **merge** the transformer with external +encoders / VAE into one standalone GGUF — pass the components at convert time +and load the single output file with `--model`: + +```bash +./build-cuda/bin/ed-convert \ + --model /path/to/stable-diffusion-3-medium/transformer/diffusion_pytorch_model.safetensors \ + --vae /path/to/stable-diffusion-3-medium/vae/diffusion_pytorch_model.safetensors \ + --clip_l /path/to/stable-diffusion-3-medium/text_encoder/model.safetensors \ + --clip_g /path/to/stable-diffusion-3-medium/text_encoder_2/model.safetensors \ + --no-t5 --type q8_0 --output sd3-merged-q8.gguf + +./build-cuda/bin/ed-cli --backend cuda --model sd3-merged-q8.gguf --no-t5 \ + --prompt "a glass teapot on a wooden table" \ + -W 1024 -H 1024 --cfg-scale 5.0 --flow-shift 3.0 --output sd3.png +``` + +`ed-convert` component flags: `--vae`, `--clip_l`, `--clip_g`, `--t5xxl`, +`--llm`, and `--no-t5` (skip merging a T5). + +This also covers **transformer-only distilled checkpoints** (e.g. +Qwen-Image-Lightning, Wan2.1-Distill, Kontext-Lightning), which ship as a bare +DiT `.safetensors` or shard index: `ed-convert` recovers the family from the +transformer's `config.json` and records it in the GGUF, so the quantized DiT +loads as a `--diffusion-model` on top of the base model's encoders / VAE. A full +diffusers-directory checkpoint (base or distilled) can also be converted whole +with `--model ` and then loaded standalone. ### Activation-calibrated imatrix quantization diff --git a/examples/cli/convert_main.cpp b/examples/cli/convert_main.cpp index 3ef44ccb..eefa66d7 100644 --- a/examples/cli/convert_main.cpp +++ b/examples/cli/convert_main.cpp @@ -35,6 +35,12 @@ void print_usage(const char* prog) { " -m, --model Input model: diffusers dir, .safetensors, shard index, or .gguf (required)\n" " -o, --output Output path; .safetensors -> safetensors, else GGUF (required)\n" " --vae Optional external VAE weights, merged under the vae. prefix\n" + " --clip_l Optional external CLIP-L text encoder, merged for a standalone GGUF\n" + " --clip_g Optional external CLIP-G text encoder, merged for a standalone GGUF\n" + " --t5xxl Optional external T5-XXL text encoder, merged for a standalone GGUF\n" + " --no-t5 Do not merge a T5 encoder (offline equivalent of ed-cli --no-t5)\n" + " Merging DiT + VAE + CLIP(+T5) lets a bare transformer convert into a\n" + " single GGUF that reloads standalone (the loader recovers the version).\n" " --type Target weight type: f32, f16, bf16, q4_0, q4_1, q5_0, q5_1,\n" " q8_0, q2_k, q3_k, q4_k, q5_k, q6_k. Default: q8_0\n" " --tensor-type-rules \n" @@ -84,6 +90,9 @@ int main(int argc, char** argv) { const char* model_path = nullptr; const char* output_path = nullptr; const char* vae_path = nullptr; + const char* clip_l_path = nullptr; + const char* clip_g_path = nullptr; + const char* t5xxl_path = nullptr; const char* tensor_type_rules = nullptr; const char* imatrix_path = nullptr; ed_dtype_t output_type = ED_DTYPE_Q8_0; @@ -113,6 +122,19 @@ int main(int argc, char** argv) { } else if (std::strcmp(key, "--vae") == 0) { vae_path = require_value(key); if (!vae_path) return 1; + } else if (std::strcmp(key, "--clip_l") == 0) { + clip_l_path = require_value(key); + if (!clip_l_path) return 1; + } else if (std::strcmp(key, "--clip_g") == 0) { + clip_g_path = require_value(key); + if (!clip_g_path) return 1; + } else if (std::strcmp(key, "--t5xxl") == 0) { + t5xxl_path = require_value(key); + if (!t5xxl_path) return 1; + } else if (std::strcmp(key, "--no-t5") == 0) { + // Offline equivalent of ed-cli --no-t5: simply do not merge a T5. + // Accepted as a no-op flag for symmetry / self-documenting commands. + t5xxl_path = nullptr; } else if (std::strcmp(key, "--type") == 0 || std::strcmp(key, "--weight-type") == 0) { const char* v = require_value(key); if (!v) return 1; @@ -148,7 +170,8 @@ int main(int argc, char** argv) { std::fprintf(stderr, "converting '%s' -> '%s' (type=%d)\n", model_path, output_path, (int)output_type); - if (!convert(model_path, vae_path, output_path, output_type, tensor_type_rules, convert_name, imatrix_path)) { + if (!convert(model_path, vae_path, clip_l_path, clip_g_path, t5xxl_path, output_path, output_type, + tensor_type_rules, convert_name, imatrix_path)) { std::fprintf(stderr, "conversion failed\n"); return 2; } diff --git a/scripts/build_cpu.sh b/scripts/build_cpu.sh index ffb65d69..5f4085f3 100755 --- a/scripts/build_cpu.sh +++ b/scripts/build_cpu.sh @@ -36,7 +36,7 @@ fi "${CMAKE_BIN}" -S . -B "${BUILD_DIR}" \ "-DCMAKE_BUILD_TYPE=${BUILD_TYPE}" \ -DED_BUILD_EXAMPLES=ON \ - "${ONEDNN_ARGS[@]}" + ${ONEDNN_ARGS[@]+"${ONEDNN_ARGS[@]}"} "${CMAKE_BIN}" --build "${BUILD_DIR}" -j diff --git a/src/core/runtime/model_loader.cpp b/src/core/runtime/model_loader.cpp index cf63a0c6..5569ddf5 100644 --- a/src/core/runtime/model_loader.cpp +++ b/src/core/runtime/model_loader.cpp @@ -1021,6 +1021,44 @@ void ModelLoader::add_tensor_storage(const TensorStorage& tensor_storage) { tensor_storage_map_[tensor_storage.name] = tensor_storage; } +std::string ModelLoader::resolve_bare_transformer_prefix(const std::string& resolved_path, + const std::string& prefix) { + // A caller-supplied prefix wins (e.g. the "--diffusion-model" component path + // passes "model.diffusion_model." directly). Only a body load with an empty + // prefix needs us to recover the version and pick a component prefix. + if (!prefix.empty()) { + return prefix; + } + + std::string lower_name = fs::path(resolved_path).filename().string(); + std::transform(lower_name.begin(), lower_name.end(), lower_name.begin(), [](unsigned char c) { + return static_cast(std::tolower(c)); + }); + if (contains(lower_name, "flux")) { + version_ = contains(lower_name, "kontext") ? VERSION_FLUX_KONTEXT : VERSION_FLUX; + return "transformer."; + } + + // A bare diffusers transformer file/shard-index (…/transformer/…) carries no + // top-level config, so get_ld_version() -- which keys on canonical names -- + // cannot recover the family before convert_tensors_name() runs, and that name + // mapping itself needs the version (chicken-and-egg). Seed version_ from the + // sibling transformer/config.json (_class_name), then return the "transformer." + // component prefix so convert_tensor_name rewrites the DiT to + // "model.diffusion_model.*" (matching how init_from_diffusers_directory loads + // the transformer/ subdir). This lets offline convert of a standalone + // transformer record the right family in the GGUF metadata and canonicalize + // names, so the result loads standalone or via --diffusion-model. + if (version_ == VERSION_COUNT) { + const SDVersion transformer_file_version = infer_transformer_file_version(resolved_path); + if (transformer_file_version != VERSION_COUNT) { + version_ = transformer_file_version; + return "transformer."; + } + } + return prefix; +} + bool ModelLoader::init_from_file(const std::string& file_path, const std::string& prefix) { last_error_.clear(); const std::string resolved_path = resolve_model_path(file_path); @@ -1055,20 +1093,13 @@ bool ModelLoader::init_from_file(const std::string& file_path, const std::string } if (is_safetensors_file(resolved_path)) { LOG_INFO("load %s using safetensors format", resolved_path.c_str()); - std::string effective_prefix = prefix; - std::string lower_name = fs::path(resolved_path).filename().string(); - std::transform(lower_name.begin(), lower_name.end(), lower_name.begin(), [](unsigned char c) { - return static_cast(std::tolower(c)); - }); - if (effective_prefix.empty() && contains(lower_name, "flux")) { - version_ = contains(lower_name, "kontext") ? VERSION_FLUX_KONTEXT : VERSION_FLUX; - effective_prefix = "transformer."; - } + const std::string effective_prefix = resolve_bare_transformer_prefix(resolved_path, prefix); return init_from_safetensors_file(resolved_path, effective_prefix); } if (is_safetensors_index_file(resolved_path)) { LOG_INFO("load %s using safetensors shard index format", resolved_path.c_str()); - return init_from_safetensors_index_file(resolved_path, prefix); + const std::string effective_prefix = resolve_bare_transformer_prefix(resolved_path, prefix); + return init_from_safetensors_index_file(resolved_path, effective_prefix); } set_error(file_exists(resolved_path) ? "unsupported model format: " + resolved_path : "model path not found: " + resolved_path); diff --git a/src/core/runtime/model_loader.h b/src/core/runtime/model_loader.h index e78cdf60..66eee7fd 100644 --- a/src/core/runtime/model_loader.h +++ b/src/core/runtime/model_loader.h @@ -227,6 +227,16 @@ class ModelLoader final { bool init_from_diffusers_directory(const std::string& dir_path, const std::string& prefix = ""); + // For a bare diffusers transformer file/shard-index loaded as the model body + // (empty prefix): recover version_ from the file name ("flux") or the sibling + // transformer/config.json, and return the component prefix ("transformer.") + // that convert_tensor_name needs to canonicalize the DiT to + // "model.diffusion_model.*". When a caller already passes a prefix (e.g. the + // "--diffusion-model" path passes "model.diffusion_model."), it is returned + // unchanged. Shared by the safetensors single-file and shard-index branches. + std::string resolve_bare_transformer_prefix(const std::string& resolved_path, + const std::string& prefix); + // Weight binding: the original load_tensors is likewise no longer exposed to the Engine. bool load_tensors(on_new_tensor_cb_t on_new_tensor_cb, int n_threads = 0, diff --git a/src/utils/model_io/convert.cpp b/src/utils/model_io/convert.cpp index 7efc4d3e..3537799d 100644 --- a/src/utils/model_io/convert.cpp +++ b/src/utils/model_io/convert.cpp @@ -152,6 +152,9 @@ static std::map> load_and_remap_imatrix(const ch bool convert(const char* input_path, const char* vae_path, + const char* clip_l_path, + const char* clip_g_path, + const char* t5xxl_path, const char* output_path, ed_dtype_t output_type, const char* tensor_type_rules, @@ -170,6 +173,28 @@ bool convert(const char* input_path, return false; } } + // Optional external text encoders, merged under the same canonical prefixes + // the runtime uses (model_loader.cpp component loading). This lets a bare + // transformer be combined with its encoders/VAE into one standalone GGUF. + if (clip_l_path != nullptr && strlen(clip_l_path) > 0) { + if (!model_loader.init_from_file(clip_l_path, "text_encoders.clip_l.transformer.")) { + LOG_ERROR("init model loader from file failed: '%s'", clip_l_path); + return false; + } + } + if (clip_g_path != nullptr && strlen(clip_g_path) > 0) { + if (!model_loader.init_from_file(clip_g_path, "text_encoders.clip_g.transformer.")) { + LOG_ERROR("init model loader from file failed: '%s'", clip_g_path); + return false; + } + } + // Skipping t5xxl_path is the offline equivalent of --no-t5. + if (t5xxl_path != nullptr && strlen(t5xxl_path) > 0) { + if (!model_loader.init_from_file(t5xxl_path, "text_encoders.t5xxl.transformer.")) { + LOG_ERROR("init model loader from file failed: '%s'", t5xxl_path); + return false; + } + } if (convert_name) { model_loader.convert_tensors_name(); } @@ -222,10 +247,16 @@ bool convert(const char* input_path, if (output_is_safetensors) { success = write_safetensors_file(output_path, tensors, &error); } else { - // Persist the true model version (known from the diffusers config.json - // that ModelLoader read on init) into the GGUF metadata, so the loader - // no longer has to guess FLUX-Kontext / Qwen-Image-Edit from the file name. - const SDVersion version = model_loader.version(); + // Persist the true model version into the GGUF metadata, so the + // loader no longer has to guess FLUX-Kontext / Qwen-Image-Edit from + // the file name. version() reads it from a diffusers config.json; + // when the source is a bare transformer (no config), fall back to + // get_ld_version() -- names are already canonical here, so it can + // recover the family from signature tensors (e.g. SD3 "joint_blocks."). + SDVersion version = model_loader.version(); + if (version == VERSION_COUNT) { + version = model_loader.get_ld_version(); + } const std::string model_ver = version != VERSION_COUNT ? ed_version_name(version) : ""; success = write_gguf_file(output_path, tensors, model_ver, &error); } diff --git a/src/utils/model_io/convert.h b/src/utils/model_io/convert.h index 10b9682f..f2fee4b2 100644 --- a/src/utils/model_io/convert.h +++ b/src/utils/model_io/convert.h @@ -6,8 +6,12 @@ // Offline model converter / quantizer. // // Loads a model from `input_path` (diffusers directory, .safetensors, shard -// index, or .gguf), optionally merges an external VAE from `vae_path`, applies -// the target weight type `output_type` plus optional per-tensor +// index, or .gguf), optionally merges external components -- VAE (`vae_path`), +// CLIP-L (`clip_l_path`), CLIP-G (`clip_g_path`), and T5-XXL (`t5xxl_path`) -- +// so a bare transformer file can be combined with its text encoders / VAE into +// one self-contained GGUF that reloads standalone. Pass nullptr / "" for any +// component to skip it; skipping `t5xxl_path` is the offline equivalent of +// `--no-t5`. Applies the target weight type `output_type` plus optional per-tensor // `tensor_type_rules` (same "name-regex=ggml-type,..." syntax as the CLI // --tensor-type-rules), and writes a single file to `output_path`. The output // format is chosen by extension: ".safetensors" -> safetensors, anything else @@ -37,6 +41,9 @@ // plain quantization (identical to before this feature). bool convert(const char* input_path, const char* vae_path, + const char* clip_l_path, + const char* clip_g_path, + const char* t5xxl_path, const char* output_path, ed_dtype_t output_type, const char* tensor_type_rules = nullptr, diff --git a/src/utils/model_io/gguf_io.cpp b/src/utils/model_io/gguf_io.cpp index 38d3195a..addcad65 100644 --- a/src/utils/model_io/gguf_io.cpp +++ b/src/utils/model_io/gguf_io.cpp @@ -94,6 +94,16 @@ bool read_gguf_file(const std::string& file_path, int64_t version_key_id = gguf_find_key(ctx_gguf_, "edgedit.model_version"); if (version_key_id >= 0 && gguf_get_kv_type(ctx_gguf_, version_key_id) == GGUF_TYPE_STRING) { *model_version = gguf_get_val_str(ctx_gguf_, version_key_id); + } else { + // Fall back to the standard "general.architecture" key so a third-party + // or official diffusers-structured GGUF (which carries no edge-dit key, + // e.g. "general.architecture" = "sd3" / "flux") is still version-tagged. + // ed_version_from_name() ignores values it doesn't recognize, so this is + // a safe best-effort hint that never regresses edge-dit's own GGUFs. + int64_t arch_key_id = gguf_find_key(ctx_gguf_, "general.architecture"); + if (arch_key_id >= 0 && gguf_get_kv_type(ctx_gguf_, arch_key_id) == GGUF_TYPE_STRING) { + *model_version = gguf_get_val_str(ctx_gguf_, arch_key_id); + } } } diff --git a/third_party/ggml b/third_party/ggml index b3899351..e621959f 160000 --- a/third_party/ggml +++ b/third_party/ggml @@ -1 +1 @@ -Subproject commit b38993515cc6129cbb7ea00041ff44484bac3baf +Subproject commit e621959f6baea1670d407f801c62a8cde253c9a7