diff --git a/.agents/docs/2026-09-07-general-build-infrastructure-gaps-design.md b/.agents/docs/2026-09-07-general-build-infrastructure-gaps-design.md new file mode 100644 index 00000000..1998de3d --- /dev/null +++ b/.agents/docs/2026-09-07-general-build-infrastructure-gaps-design.md @@ -0,0 +1,895 @@ +# 通用构建基础设施:缺口、归属与验证 + +2026-09-07。本文处理**通用构建基础设施**:其他构建系统都有对应物、其存在理由不引用 +任何领域概念的那些能力。异构方向特有的东西不在缺口清单里,但 §7 给出归属规则, +§10 给出用一个陌生厂商证伪整条架构主张的实验。 + +本文经过一次修订(§14 变更记录)。修订推翻了初稿的两处判断,均已就地更正。 + +## 0. 范围 + +### 0.1 判据:什么算"通用" + +一项能力进入本文,当且仅当 CMake / Meson / Autotools / Cargo 中至少两个有对应物, +且它的存在理由不引用任何领域概念(设备、加速器、内核)。 + +| 项 | 对应物 | 结论 | +|---|---|---| +| 导出面 | CMake `CXX_VISIBILITY_PRESET`、`WINDOWS_EXPORT_ALL_SYMBOLS`;Meson `vs_module_defs` | 通用 | +| 通用链接标志 | Cargo `cargo:rustc-link-arg` | 通用 | +| 包内布局 | `install(FILES ...)` / Meson `install_data` | 通用 | +| 探测库 | `check_include_file` / `check_function_exists` / `check_type_size` | 通用 | +| 配置头 | `configure_file` | 通用,但**不是引擎改动**(§5) | +| 开放词表的"空"取值 | CMake 无对应物;Meson/Cargo 亦无 —— 但它们的谓词词表是封闭的,**因而不需要**;开放词表是本生态自选的性质,这条规则随之而来 | 通用(§6.1) | +| 生成输入的依赖粒度 | CMake `add_custom_command(OUTPUT ...)` 的逐文件依赖;Ninja 的 `order_only` | 通用(§6.2) | +| 多语言(Fortran) | 四者皆有 | 通用(§6.3) | + +### 0.2 RDC 为什么不在本文(初稿判断已更正) + +relocatable device code 的存在理由是"设备编译器把 `__device__` 函数的跨 TU 调用 +推迟到一次设备链接"。这句话无法脱离设备概念陈述,因此按 §0.1 的准入线它不是通用 +基础设施。 + +**更要紧的是:它也不是引擎缺口。** 初稿把 RDC 列为"真缺"并暗示需要引擎支持,按 §7 +的归属规则复核后不成立。它需要的原语今天全部存在:规则包用 `action` 以 `-rdc=true` +编每个设备 TU(各自以 `object` 归宿进入链接),再用一个 `action` 跑设备链接步骤, +其产物同样以 `object` 归宿加入普通链接;设备运行时库走 `link_lib`。**没有新的边种类, +没有新的产物性质,引擎零改动。** + +结论:RDC 属于 `rules-cuda` 的工作量,与本文各项**并行**,不构成依赖。 + +## 1. 现状,实测 + +四条,均在 2026-09-07 于 `main` 上核实: + +1. 引擎认识 25 条 `mcpp:` 指令。链接相关只有三条:`link-lib`、`link-search`、 + `link-script`(即 `-T`)。**不存在通用的链接标志出口。** +2. 共享库的导出面在两个平台上都是"全导出":ELF 走默认 visibility;PE 由 + `src/build/coff_exports.cppm` 自动生成 `.def`,语义对齐 CMake 的 + `WINDOWS_EXPORT_ALL_SYMBOLS`。**不存在收窄导出面的声明。** +3. `[runtime].artifacts` 已经是"包内相对路径 + role + provenance"的声明,role 是 + 封闭白名单。**包内布局不需要新 section。** +4. 配置头生成的四块拼图都在:`toolchain_dir()` / `sysroot_dir()` 给出正确的编译器 + 与 sysroot,build.mcpp 是真正的程序因而能写文件,`include-dir` 让本包 TU 看见, + `rerun-if-changed` 保证增量。**这一项不是引擎缺口。** + +## 2. 缺口一:导出面 + +### 2.1 问题 + +一个 `.so` / `.dylib` / `.dll` 目前只有一种导出策略:全部导出。这在两类项目上不成立。 + +**ICD / 插件。** Vulkan loader 只按名字取 `vk_icdGetInstanceProcAddr` 与 +`vk_icdNegotiateLoaderICDInterfaceVersion`。一个把内部符号一并导出的 ICD,会与 +loader 以及同进程内的另一个 ICD 撞名。 + +**一个镜像里两个 C++ 运行时。** 实测:SYCL 示例构建时 mcpp 自己的重复符号检查报告 + + warning: sycl-saxpy: 68 symbols in this image are also provided by a library + it loads. _Unwind_DeleteException() _Unwind_GetGR() ... + +`libsycl.so` 对着 libstdc++ 编译,mcpp 产物链 libc++,双方都导出 unwinder 符号。 +收窄导出面是这类问题的标准解法。 + +绕过办法今天存在:`cflag` / `cxxflag` 塞 `-fvisibility=hidden`。它能用,但它是标志 +不是声明,而且 PE 上没有对应物 —— 那里只有"全导出"或作者自己手写 `.def`。 + +### 2.2 设计:一个中立声明,三种平台渲染 + +```toml +[targets.mydriver] +kind = "shared" +soname = "libmydriver.so.1" +exports = "abi/mydriver.exports" # 或内联:exports = ["vk_icd*"] +``` + +文件内容是符号模式,一行一条,`#` 起注释: + +``` +vk_icdGetInstanceProcAddr +vk_icdNegotiateLoaderICDInterfaceVersion +``` + +引擎按平台渲染同一份声明: + +| 平台 | 渲染为 | +|---|---| +| ELF | version script,经 `-Wl,--version-script=` | +| Mach-O | `-exported_symbols_list` | +| PE | `.def`,**取代**自动生成的全导出版本 | + +**这与 `[runtime]` 的既有先例同构。** `[runtime]` 存在的理由正是"一句中立的话,由 +引擎按方言渲染",而不是让作者写三份平台专用文件。导出面是同一形状的第二个实例, +因此它不是一个新概念,是一条既有原则的应用。 + +### 2.3 声明 `exports` 不隐含编译期 hidden(实现时更正) + +初稿写的是"声明 `exports` 时引擎同时把编译期默认置为隐藏"。**实现时更正为不隐含。** + +三种格式上的收窄都是**链接期**属性:version script 限制的是动态符号表, +`-exported_symbols_list` 与 `.def` 同理。因此隐含一个编译期效果会让一个键有两个效果, +而第二个效果还改变**本库各翻译单元之间**如何看见彼此 —— 那是一个有独立理由的独立决定。 + +`-fvisibility=hidden` 仍可经 `[build] cxxflags` 使用以取得代码生成收益。相应地,初稿的 +判据 C3(声明后夹具应链接失败)**作废**,因为它断言的正是这个被取消的耦合。 + +
+初稿原文(已作废) + +仅有 version script 会收窄动态符号表,但对象里的符号仍是默认可见性,链接期优化拿 +不到收益,而且 Mach-O 与 PE 的渲染需要编译期配合。因此:**声明 `exports` 时,引擎 +同时把编译期默认置为隐藏**(ELF/Mach-O 的 `-fvisibility=hidden`)。 + +这是一处行为变化,必须写进文档:一个此前依赖默认可见性做跨 DSO 内部调用的项目, +在声明 `exports` 之后会链接失败。这正是作者声明 `exports` 时所要求的语义,失败点 +也在链接期而非运行期,因此是可接受的。 + +
+ +### 2.4 不做什么 + +- **不做符号版本的完整语法。** `foo@@LIB_1.0` 与 `foo@LIB_0.9` 并存是 ELF 独有的 + 能力,无法中立表达。需要它的项目把 map 文件签入仓库,经 `[build] ldflags` 使用; + 需要**生成** map 的项目走 §3 的出口。 +- **不做 per-symbol 的属性宏。** `__declspec(dllexport)` 那一套是源码的事。 + +## 3. 缺口二:通用链接标志 + +### 3.1 问题 + +`link-lib`、`link-search`、`link-script` 之外没有出口,因此**构建程序算出来的**链接 +标志无法送达。三个具体场景: + +| 标志 | 谁需要 | +|---|---| +| `-Wl,--version-script=<生成的 map>` | 导出面随 feature 组合变化的库(§2.4) | +| `-Wl,--wrap=malloc` | 接管 C 库符号的运行时:内存池、tracing、sanitizer | +| `-Wl,--exclude-libs,ALL` | 静态吞入的第三方库不得再导出,否则其符号成为本包 ABI 的一部分 | + +第三条与 §2 是同一问题的两半:一半管自己的符号,一半管吞进来的符号。 + +### 3.2 设计 + +``` +mcpp:link-flag= mcpp::link_flag(s) +``` + +按发出顺序追加,位置在 manifest 的 `[build] ldflags` 之后。 + +### 3.3 传播性:到达消费者(实现时更正) + +初稿判它私有,与 `include-dir` 同规。**实现时更正:这个类比是假的,而且代码就是证据。** + +`linkUsage.ldflags` 是 `buildConfig.ldflags` 的一份拷贝,`propagateLinkFlags` 把依赖的 +每一条 ldflag 推到消费者 —— **引擎今天没有"私有链接标志"这个策略可表达**。 + +更要紧的是类比本身错在哪:`include-dir` 私有,是因为编译接口有一个声明式的公开对应物 +(`[build] include_dirs`),构建期程序若能加宽它就是绕过 manifest。链接标志没有这个 +分裂 —— `[build] ldflags` 本来就传播。让"算出来"的形态与它自己的声明式孪生行为不同, +才是不一致,而不是防护。 + +**后果写明而不藏起来**:依赖发出的 `--version-script` 也会落到消费者链接行上,而这通常 +不是它的本意。这个隐患**不是新的** —— 依赖在 `[build] ldflags` 里写同一条标志一直如此 +—— 所以这条指令加宽的是**谁能算出这个值**,不是**这个值能到达哪里**。 + +相应地,C4 的反向断言("不出现在消费者的链接行上")作废。 + +## 4. 缺口三:包内布局 + +### 4.1 问题不是 `install()` + +mcpp 世界里没有系统前缀:消费者解析包,不扫路径。`[resources]` 是把资产**嵌入产物** +(图标、版本元数据),不是安装。因此 `install(FILES ... DESTINATION /usr/share)` 这个 +形状在这里是错的。 + +真实需求窄得多,且**只有被第三方按路径扫描的文件才有**: + +| 机制 | 谁扫 | +|---|---| +| `/usr/share/vulkan/icd.d/*.json`,或 `VK_DRIVER_FILES` 指向的文件 | Vulkan loader | +| `OCL_ICD_FILENAMES`(追加语义)/ `OCL_ICD_VENDORS`(替换语义) | OpenCL ICD loader | +| 任意 dlopen 插件目录 | 宿主程序 | + +关键点:**那个 JSON 不是给 mcpp 消费者读的,是给 loader 读的。** 它必须是包内某个 +确定相对路径上的真实文件。 + +### 4.2 设计:扩 role,不开新 section + +`[runtime].artifacts` 已经是"包内相对路径 + role + provenance"的声明。按 docs/05 +附录 A 第二条(一个键若重复了别处已给出的答案则不予准入),这里**不得**新开 section。 + +新增一个 role: + +```toml +[runtime] +artifacts = [ + { role = "library", path = "lib/libmydriver.so.1", provenance = "built" }, + { role = "manifest", path = "share/vulkan/icd.d/mydriver.json", provenance = "built" }, +] +``` + +`role = "manifest"` 的含义:**一个被本包之外的加载器按路径读取的数据文件**。它与 +`library` 的区别不是格式,是读者 —— 这是 role 白名单里唯一缺的那一类。 + +**已知约束**:打包之后 `runtime.artifacts` 是封闭白名单,而已发布的描述符会跳过它 +不认识的键。因此新增 role **必须**先落地引擎、发布,再由包使用;顺序反了会让老 +引擎静默丢掉这条 artifact。这与 SPEC-004 §4.3 的规则同源。 + +### 4.3 生成文件与路径回填 + +ICD JSON 的内容里要写 `.so` 的位置,而那是构建期才知道的。两条约束: + +1. JSON 里写的**必须**是相对于包根的路径,不得是构建目录的绝对路径。否则包一经 + 移动或分发即失效 —— 这是本仓库已经付过学费的形态(载荷内嵌绝对路径)。 +2. 因此生成它的是 build.mcpp,而声明它的是 `[runtime].artifacts`。二者的接缝需要 + 一条指令让构建程序贡献一个 artifact 条目: + +``` +mcpp:artifact== mcpp::artifact(role, relpath) +``` + +**开放问题已查清**(2026-09-07 读 `src/pack/manifest_emit.cppm`): + +`[[runtime.artifacts]]` 的发出**完全由 packer 自己决定**,它从 `doc.legs`(构建出来的 +库)加一条 `interface` 条目生成,**不携带作者在源 manifest 里写的 `[runtime].artifacts`**。 +`role` 在 manifest 解析侧是自由字符串,没有白名单;白名单效应来自读者:`prebuilt.cppm` +认 `static-library` / `shared-library` / `interface`,`prepare.cppm` 认前两个。 + +因此本项的工作量比初估大,且落在 packer 而非 manifest 解析: + +1. `mcpp:artifact==` 指令(与 §3 的 `link-flag` 同形,一行表项); +2. packer 要把该文件**拷进产物**并把条目**写进描述符** —— 这是新行为,今天的 packer + 只发它自己产出的东西; +3. 该文件的内容通常由构建程序生成(ICD JSON 里要写 `.so` 的位置),所以第 2 步接收的 + 是构建目录里的一个路径,而落点是包内相对路径。 + +结论:**这一项不是"加一个 role",是给 packer 增加一条"携带被声明的文件"的通路。** +分期不变(三期),但依据从"小改动"改为"边界清楚、工作量中等,且不阻塞其他各项"。 + +## 5. 缺口四:探测库(是包,不是引擎) + +### 5.1 现状 + +配置头生成今天就能写(§1 第 4 条)。缺的不是能力,是**公共实现**:每个移植过来的 C 项目 +都要自己写一遍"这个头在不在""这个函数能不能链上""这个类型多宽"。不做的后果是 +CMake 模块生态碎片化的重演 —— 每个项目一份略有差异的 `check_function_exists`。 + +### 5.2 一条硬约束:探测不得读宿主 + +这是本文唯一一条会被写错的设计。autotools 的探测按构造读宿主,而本生态的不变量是 +相反的:**探测必须用生态解析出的编译器与 sysroot 进行**。 + +因此探测库的每个入口都经 `toolchain_dir()` / `toolchain_sysroot()` / +`toolchain_binutils_dir()` 组装命令行,任何一条走 `/usr/bin/cc` 的实现都是错的。 +这与 rule 包驱动第二编译器时的规则是同一条(v2 设计 §1 第 1、2 条推论)。 + +### 5.3 形状 + +一个包 `mcpplibs:probe`,供 build.mcpp 导入: + +```cpp +import mcpp; +import mcpp.probe; + +int main() { + mcpp::probe::Ctx cx; // 从 toolchain_* 组装,不读宿主 + bool mman = cx.has_header("sys/mman.h"); + bool slcpy = cx.links("strlcpy", "#include "); + int lw = cx.sizeof_type("long"); + mcpp::probe::configure_file(cx, "config.h.in", out / "config.h"); + mcpp::include_dir(out); +} +``` + +探测结果必须**按工具链指纹缓存**,否则每次构建重探。缓存键取 +`toolchain_fingerprint` 已有的值,不新造。 + +## 6. 新增的三处缺口(修订加入) + +### 6.1 开放词表不能靠枚举取反 + +**问题。** CPU 回退今天只能这样写: + +```toml +[target.'cfg(not(any(accelerator = "cuda", accelerator = "vulkan")))'.build] +sources = ["src/cpu/*.cpp"] +``` + +`accelerator` 的取值是**开放的** —— docs/20 明确说"第五个后端是一个包,不是引擎改动"。 +因此这条谓词的含义会随生态增长**静默改变**:新增一个后端之后,每个工程的回退谓词都 +必须被编辑。漏一个的后果是 CPU 实现与设备实现同时进入编译集,或者该编时没编。 + +这不是措辞问题,是一条**一般规则**: + +> 一个开放词表的键,`not(any(<枚举>))` 永远不等价于"该词表为空"。因此每个开放词表的 +> 层键**必须**提供一个不依赖枚举的"空"取值。 + +**设计。** 取值 `none`: + +```toml +[target.'cfg(accelerator = "none")'.build] +sources = ["src/cpu/*.cpp"] +``` + +`accelerator = "none"` 为真当且仅当本次构建的加速器集合为空。`not(accelerator = "none")` +自然表达"有任意设备后端"。 + +**为什么是 `none` 而不是 `cpu`。** 三条: + +1. **本仓库已有这个拼法。** `os = "none"` 就是裸机(docs/05 §2.7.2)。同一份 manifest + 里,同一个词,同一个意思。 +2. `cpu` 引入歧义。`accel = "cuda"` 时 `cfg(accelerator = "cpu")` 是真是假?为真则 CPU + 源码永远参与编译,破坏互斥语义;为假则必须写 `accel = "cuda, cpu"`,改动既有 manifest。 +3. `accel` 这条轴回答的是"哪个设备编译器、哪个架构"。CPU 不需要设备编译器,它不在 + 这条轴上。把它塞进去会让这条轴同时承载两种问题。 + +**"CPU 后端与设备后端并存"是另一个需求,今天已经可写。** 那种形态(如 ggml 的 CPU +backend 与 CUDA backend 同时编入一个产物)不需要本项:CPU 源码放无条件的 +`[build] sources`,设备源码放 `cfg(accelerator = "x")`。`not(...)` 只在**互斥接缝** +(要么这个实现,要么那个)时才需要,而互斥接缝正是本项要修的场景。 + +**实现规模。** cfg 求值器里 `accelerator` 走的是集合成员判定,`none` 需要一条特判: +集合为空时为真。不是免费的,但是一条条件。 + +**推广。** 同一条规则适用于每个开放词表的层键。`compiler`、`c-abi`、`compiler-runtime` +是否也需要 `none`,应在实现本项时一并裁定,而不是逐个再议 —— 否则这条规则会以每次 +一个键的方式被重新发现。 + +### 6.2 生成输入的依赖粒度 + +**问题。** `mcpp:generated=` / `source` 归宿的语义是(docs/07 原文)"**本包每一条编译边 +都等它**"。一个生成的头因此构成**包级栅栏**。 + +134 个 shader 无所谓 —— 它们是叶子,没有别的 TU 等它们。但一个**被少数 TU 包含的生成 +头**会让全包的编译边排在它后面。在驱动、编译器这一档的规模上,这是"并行构建"与 +"分阶段构建"的差别。 + +**设计方向。** 需要"某条编译边依赖某个具体生成文件",而不是"全包等全部生成物"。 +两个候选形状: + +- 让 `action` 的输出可被具体源码 glob 引用(声明式的边) +- 让生成物携带一个标签,源码侧按标签声明依赖 + +**本文不选型。** 这一项与 ninja 图的构造方式耦合较深,选型前需要读 `src/build/plan` +与 `ninja_backend`,确认哪种形状不会与既有的 dyndep/BMI 调度冲突。本文的职责是**指出 +它是通用缺口并给出触发条件**:当一个包同时具备(a)生成的头文件,且(b)编译边数量 +达到千级时,包级栅栏成为主要瓶颈。 + +### 6.3 多语言:Fortran + +**问题。** mcpp 覆盖 C / C++ / 汇编。异构与 HPC 栈里 Fortran 不是边缘 —— 参考 LAPACK、 +大量求解器,以及 Fortran + OpenMP target 这个在科学计算代码里非常常见的组合。没有 +Fortran,数值栈的一大块进不来。 + +**归属。** 按 §7 三测试:不点名厂商(测试 1 不触发),不改变产物性质(测试 2 不触发), +但**新增一种编译边的种类**(测试 3 触发)⇒ **引擎侧**。 + +**规模诚实说。** 这一项比本文其余各项都大:它要求工具链模型承认第三种编译器、模块/ +接口文件(`.mod`)有自己的依赖图(与 BMI 类似但不同)、以及 Fortran/C 的名字修饰与 +调用约定。**本文把它记为已识别的缺口,不给设计** —— 给它一个半成品设计比不给更坏。 + +## 7. 归属:引擎侧还是插件侧 + +不列清单,给三条测试。任一为是即归该侧;测试按顺序应用。 + +| # | 测试 | 归属 | +|---|---|---| +| 1 | 它是否点名某个厂商或工具? | **插件**。仓库已在强制这条:`tests/unit/test_core_vendor_probes.cpp` 以文件数为分母,断言 `src/` 去注释后不含厂商工具名 | +| 2 | 它是否改变**产物是什么**(符号面、包内布局、身份)? | **引擎**。packer、索引、消费者三方必须就此达成一致,而这种一致无法住在插件里 | +| 3 | 它是否**新增一种边或节点**? | **引擎**。插件声明边,不发明边的种类 | + +三条测试判的是**引擎侧还是插件侧**,不判**通用还是领域**。后者由 §0.1 的准入线判。 +两把尺子会交叉,`kind = "device"` 就落在交叉格里:它是引擎侧(测试 2),但它谈的是 +`accel`,按 §0.1 的准入线属**领域**。因此它**不在本文的缺口清单、判据与分期里**, +归 docs/20;此处列出只是为了让归属表完整。 + +应用到已识别的各项: + +| 项 | 归属 | 触发的测试 | +|---|---|---| +| `exports` | 引擎 | 2 | +| `link-flag` | 引擎 | 3(边的属性) | +| `role = "manifest"` | 引擎 | 2 | +| `accelerator = "none"` | 引擎 | 3(谓词词表) | +| 生成输入粒度 | 引擎 | 3 | +| Fortran | 引擎 | 3 | +| `kind = "device"` | 引擎,但**属领域侧**(见下) | 2 —— 它让 `mcpp pack` 能**测量**出 `accel` 而不是抄声明 | +| 探测库 | 插件/包 | 1 | +| **RDC** | **插件** | 均不触发(见 §0.2) | +| `.omp` 岛、`.stdpar` 岛 | 插件 | 1 | + +## 8. 编程模型的三分法,与岛化的边界 + +docs/20 目前是二分:SYCL 分得开,OpenMP offload 与 stdpar 分不开。**这个二分不成立, +应改为三分。** + +| 模型 | 可岛化? | 依据 | +|---|---|---| +| SYCL | **天然** | kernel 是 `submit` 里的闭包,本来就隔离 | +| OpenMP offload | **重构后可以** | 把 `target` 区域提成函数放进自己的 TU,该 TU 用 offload 标志编;**调用方不需要任何 offload 标志** | +| stdpar | **重构后可以** | 见下 | + +### 8.1 stdpar 可以岛化(初稿判断已更正) + +初稿称 stdpar 与 mcpp 的收益"互斥"。这个说法过头了,应更正。 + +`-stdpar` 拆开是两件事,**两件都可表达**: + +| 层面 | 是什么 | 表达为 | +|---|---|---| +| 编译侧 | 含并行算法调用点的 TU 由该编译器生成 kernel | 岛,与 `.sycl` 同形 | +| 链接侧 | 整个进程的分配器换成托管内存 | **链接行属性**,`[runtime]` 的 link intent | + +关键在于分配器替换是**链接期**的事。一个普通 TU 里分配的内存,只要最终链接进了托管 +分配器,就是设备可见的;不需要每个 TU 都由该编译器编译。 + +而且 mcpp 对"依赖强加给消费者的全镜像属性"已有先例:`cxx_runtime`。C++ 运行时的选择 +正是这个形状 —— 一个包的选择决定整个镜像并沿依赖传播。托管分配器是同一类东西。 + +**真正丢掉的是卖点,不是能力。** stdpar 的价值主张是"源码一行不改",而岛化要求一次 +重构。这是产品张力,不是架构矛盾。规范应当这样陈述,而不是宣称做不到。 + +### 8.2 唯一真正的对立,以及它会自己过期 + +只有一处:**整目标编译 + C++20 模块**。没有 offloading 编译器接受具名模块,所以一个 +整目标 target 就是一个没有模块的 target。 + +这是**编译器能力**造成的,不是设计取舍,并且是**时限性的** —— 等 offloading 编译器 +支持模块,这条对立自行消失。一条会过期的约束与一条设计取舍应分开记录,处理方式不同。 + +### 8.3 OpenMP 岛的已知边界 + +当 `declare target` 的全局数据跨 TU 时,岛会漏:设备镜像要求那个全局也在设备侧发出, +而它定义在宿主 TU 里。`.omp` 岛成立的条件是"offload 区域对设备全局自包含",这条限制 +必须写进规则包的文档,不能留给用户在链接错误里发现。 + +## 9. 术语:"island" 这个词 + +概念是真的:**存在一组 TU,由另一个编译器二进制处理,其产物加入普通链接。** 这是一条 +构建系统的轴,业界没有为它命名 —— CMake 直接把 CUDA 当一门"语言"绕过去了。 + +但这个词有两个问题: + +1. **它与业界既有的轴交叉。** 业界的二分是 single-source / separate-source。按那条轴 + **CUDA 是单源**(一个 `.cu` 里既有 `__global__` 又有主机代码,由 nvcc 内部拆分), + 而 docs/20 把 CUDA 称作岛。熟悉 CUDA 的读者会在这里卡住,因为两条轴用了同一批例子 + 给出相反的归类。 +2. **`island` 在链接器词汇里已被占用。** ARM / Mach-O 的 *branch island*(veneer)是 + 长跳转桩。mcpp 是一个谈链接的构建系统,这个碰撞是实际的。 + +**建议:保留词,补一句对齐**,写进 docs/20: + +> 本文的"岛"是**构建系统**的轴:哪些 TU 交给另一个编译器二进制。它与编程模型的 +> single-source / separate-source 轴正交。CUDA 在后者是单源,在前者是岛。 + +不建议改词:"岛 / 接缝"这对比喻自洽,且已进入多份文档;改名的代价大于这句对齐。 + +## 10. 验证:用一个陌生厂商证伪 + +本文各项都是"补齐"。而整条架构最强的主张是另一句:**引擎里没有厂商知识,加一个厂商 +等于加一个规则包。** 它有测试在守(§7 测试 1),但**从未被一个引擎没见过的厂商检验 +过** —— 现有四条 lane 全在 NVIDIA / Khronos 谱系里。 + +### 10.1 靶子 + +昇腾 CANN 栈。三条理由: + +1. 完全不同的 ISA、编译器与运行时 API,与既有四条 lane 无谱系重叠。 +2. Ascend C 是岛形态 —— docs/20 已把它列为岛的例子,但那是**推断**,这里可变成实测。 +3. 垂直完整,从框架适配到驱动。 + +### 10.2 不做全栈,做一个跨层切片 + +全栈不可行,五条理由:驱动是内核态(按不变量本就排除);毕昇编译器是 LLVM 量级; +Python 层(`pyasc`、`pypto`、框架适配)不在覆盖内;规模是数人年;**没有真机或模拟器 +则退化为编译验证** —— 正是本仓库给现有 lane 打的差评,代价放大百倍。 + +还有一条更微妙的:CANN 是普通 C++ 而非模块化 C++,所以这个实验测的是 mcpp 作为通用 +构建系统的能力,**不测它的差异化能力**。这不是反对理由,但必须清楚测的是哪一半。 + +切片: + +| 做 | 不做 | +|---|---| +| `rules-ascendc` 规则包 | 从源码构建毕昇编译器 | +| 毕昇**作为载荷**(厂商 URL,与 dpcpp 同待遇) | `ge` 图引擎 | +| 运行时 host API 作为载荷 | Python 层 | +| 一个算子库的**少量算子**原生构建 | `driver` | +| 算子注册元数据(若为 JSON,需核实)→ `role = "manifest"` 的真实用例 | `ops-nn` 全量 | +| 一个最小消费者,在真机或模拟器上跑出正确结果 | | + +这条切片贯穿应用 → 库(导出设备代码)→ 运行时 → 设备编译器 → 设备执行,并恰好压在 +`exports`、`role = "manifest"`、探测库、以及 RDC 的昇腾对应物上。**四项够用则通过; +不够则暴露第五项 —— 那正是想要的产出。** + +### 10.3 调研结果(2026-09-07 实测) + +三处不确定已调研,**全部证实,无阻碍项**(§10.3.3 的初次否定结论已被推翻)。仓库全部可匿名 clone +(`gitcode.com/cann/*`,分支 `8.5.0`),许可为 CANN Open Software License Agreement 2.0。 + +#### 10.3.1 无硬件执行:存在,且是厂商自己的测试路径(证实) + +Ascend C 有**三种**运行模式,经 `-DCMAKE_ASC_RUN_MODE=` 选择: + +| 模式 | 需要硬件 | 走不走岛 | +|---|---|---| +| `npu`(默认) | 是 | 是 | +| `sim`(NPU 仿真) | 否 | **是**(见下) | +| `cpu`(CPU 调试) | 否 | **否** | + +仿真器**随工具链发布**,按 SoC 分库:`${ASCEND_DIR}/*/simulator//lib`, +`Findpvmodel.cmake` 里的目标是 `pvmodel_ascend910` / `pvmodel_ascend310p` / +`pvmodel_ascend610` 与 `pem_davinci_ascend910B1` / `pem_davinci_ascend310B` / +`pem_davinci_ascend610Lite`。**asc-devkit 自己的单元测试就链接它们** +(`tests/unit/basic_api/ut/CMakeLists.txt`),所以这是厂商既有的无卡测试路径, +不是我们发明的用法。 + +**关键区分,而且它决定判据能不能用:** + +`cpu` 模式链接 `tikicpulib::${SOC_VERSION}` 并使用 `compiler/tikcpp/` 的头文件 +(实测于 `cmake/asc/legacy_modules/function.cmake`)。也就是说**同一份 kernel 源码由 +宿主编译器编译**,构建图里根本没有岛。用 `cpu` 模式跑绿,证明的是 kernel 数值对, +**不是岛的机制对** —— 它没走那条路径。这正是本仓库反复付学费的形态:判据施加在 +错误的对象上。 + +`sim` 模式保留岛。**这一条已由源码证实,不再是推断**(`asc-devkit/tools/ascc/cmake/` +与 `cmake/asc/legacy_modules/`): + +| 证据 | 内容 | +|---|---| +| `CMakeDetermineASCCompiler.cmake:47` | `find_program(CMAKE_ASC_COMPILER NAMES "bisheng" PATHS ".../ccec_compiler/bin/")` | +| `CMakeASCInformation.cmake:49` | `CMAKE_ASC_COMPILE_OBJECT = " … -c -x asc "` | +| `host_config.cmake:69` | `CCEC_LINKER = /ccec_compiler/bin/ld.lld` —— 设备侧链接 | +| **全仓 `RUN_MODE` 判断** | **每一处都是 `STREQUAL "cpu"`,不存在 `sim` 分支** | + +最后一行是决定性的:构建期只区分 `cpu` 与**非** `cpu`,因此 **`sim` 与 `npu` 走完全 +相同的编译路径**,bisheng 被调用,岛成立。两者的差别在**运行期**(加载哪套运行时 / +仿真库),不在构建图。 + +分支点上唯一的额外动作是:非 `cpu` 分支多跑一个 `update_host_stub.py` —— 生成宿主侧 +的启动桩,而那正是"存在一个真实设备二进制需要被拉起"的标志。 + +**为什么这个疑问值得问。** 反过来的设计是存在的,而且有正当理由:指令级仿真比跑宿主 +代码慢几个数量级;宿主编译器能给 gdb / ASAN / printf。CANN 的选择是把这两种诉求拆开 +—— `cpu` 模式**就是**那个设计,所以 `sim` 若不执行真实设备指令便与 `cpu` 重复,没有 +存在的理由。三种模式而非两种,本身就是答案。 + +#### 10.3.2 算子注册元数据:JSON,且按 SoC 分(证实) + +实测 `ops-math`(72 MB,`math/` 下 1451 个 `.cpp`、927 个 `.h`): + +``` +math//op_kernel/ 设备侧 +math//op_host/ 宿主侧 +math//op_host/config//_binary.json +math//op_host/config//_simplified_key.ini +``` + +`math/` 一棵树里 166 个 JSON、143 个 INI。JSON 的内容是算子签名到**设备二进制文件名** +的映射: + +```json +{ "op_type": "Abs", + "op_list": [ { "bin_filename": "Abs_1c4543fdfe...", + "inputs": [ { "dtype": "bfloat16", "format": "ND", ... } ] } ] } +``` + +**这是 `role = "manifest"` 的教科书用例**:一个数据文件,由本包之外的运行时按路径读取, +以文件名指向设备产物。 + +而且 SoC 目录有六个 —— `ascend310p`、`ascend910`、`ascend910_93`、`ascend910_95`、 +`ascend910b`、`kirinx90`。**注册文件是按目标条件化的**,因此这个切片同时压在 +`role = "manifest"` 与 SPEC-004 的目标轴上。 + +另外两项确认:**源码布局本身已经是岛** —— 每个算子的 `op_kernel/` 与 `op_host/` 是分开 +的目录;设备架构标志是 `CMAKE_ASC_ARCHITECTURES=dav-2201`,与 `sm_89` 同类。 + +#### 10.3.3 工具包的获取:官方镜像,可匿名拉取(实测,已推翻初次结论) + +**初次调研把这一条判成"未能证实、构成阻碍"。复查后推翻。** + +首先纠正一个框架错误:初次把"毕昇"与"模拟器"当成两个获取问题。**它们在同一个包里** —— + +| 组件 | 路径 | +|---|---| +| 设备编译器 | `${ASCEND_DIR}/compiler/ccec_compiler/bin/bisheng` | +| 模拟器 | `${ASCEND_DIR}/*/simulator//lib` | + +两者都在 CANN 工具包内部,所以这是**一个**获取问题,不是两个。 + +**官方获取方式是 Docker 镜像**,而不是 `.run` 安装包(实测于 `ops-math/QUICKSTART.md`): + +``` +swr.cn-south-1.myhuaweicloud.com/ascendhub/cann:8.5.0-910b-ubuntu22.04-py3.10-ops +``` + +**它可以匿名拉取。** 实测过程与结论: + +| 步骤 | 结果 | +|---|---| +| `GET /v2/` | `401` | +| 匿名 token:`GET /swr/auth/v2/registry/auth?service=dockyard&scope=repository:ascendhub/cann:pull` | **签发**(2107 字节 JWT) | +| 持 token 取 manifest | **200**,manifest list,含 `arm64` 与 `amd64` | + +**那个 401 差点让我下错结论。** 裸的 401 与"需要凭据"读数相同,但它同样是匿名 token +握手的第一步(Docker Hub 就是这样)。只看第一步会把"可匿名获取"误判成"不可得" —— +这正是本仓库记过的形态:判据的"否"与"没测成"同读数。**判据必须走完握手。** + +**合规性:正好落在不变量允许的那一档。** 本生态的规则是"专有厂商用户态要么在它已经 +在的地方链接,要么从**厂商自己发布的 URL** 取 —— 绝不拷进 xlings-res 的发布物"。 +从华为自己的 registry 拉官方镜像**就是**这一档,不需要再分发权。 + +结论:**这一条不再是阻碍。** 剩下的是工程问题(镜像里取哪些目录、怎么做成载荷), +不是许可问题。 + +#### 10.3.3b AscendNPU IR 已开放,存在第二个切入点 + +毕昇开放了 **AscendNPU IR**(昇腾自有的 MLIR dialect),并给出 Triton 的完整路径: +Triton IR → Linalg IR → AscendNPU IR → 算子二进制,配套 `triton-ascend`。 + +这不改变"毕昇二进制在工具包里"这个事实,但它意味着**规则包有两个可选的切入高度**: + +| 切入点 | 输入 | 代价 | +|---|---|---| +| Ascend C(§10.2 的切片) | `op_kernel/*.cpp` | 直接对应现有算子库 | +| Triton / AscendNPU IR | Triton kernel | 与 CUDA 侧的 Triton 生态同构,但离现有 CANN 算子库更远 | + +**本文不选型**,只记录第二条存在。对"验证 mcpp 完备性"这个目的而言 Ascend C 更直接, +因为它才是 `ops-math` 里那 1451 个 `.cpp` 的实际形态。 + +#### 10.3.4 一个附带发现:CMake 把 Ascend C 当作一门"语言" + +`cmake/asc/ASCConfig.cmake` 注释写着 `plugin support ASC language`,并从 +`$ENV{ASCEND_HOME_PATH}/compiler/tikcpp/ascendc_kernel_cmake/ASC_CMake/FindASC.cmake` +引入。开源仓库里只有 `include()`,**真正的设备编译机制在工具包内部**。 + +两个推论: + +1. 移植意味着**读工具包自带的 cmake 来学会它发什么标志**,与 `rules-sycl` 驱动 dpcpp + 的做法同形,不是新问题。 +2. 这是 §9 那条轴的第三方佐证:CMake 用 `enable_language(ASC)` 把它放进**引擎**, + mcpp 用规则包把它放进**包**。**同一条轴,不同的归属** —— 而归属正是 §7 要裁决的事。 + +### 10.4 可行性闸(两次修正后) + +| 目的 | 需要 | 结论 | +|---|---|---| +| **架构验证**(本实验) | 无 | **可开工**。裁决判据 `git diff src/` 为空完全在构建期 | +| **设备执行判据** | `sim` 模式 + 工具包 | **可得**。工具包镜像匿名可拉(§10.3.3),`sim` 无需硬件(§10.3.1) | +| 真机验证 | 昇腾硬件 | 可选,不阻塞以上任何一项 | + +**没有阻碍项。** 初稿列的三条前置里,两条证实、一条被推翻;剩下的都是工程量。 + +**没有待确认的技术问题。** 初稿留的最后一条(`sim` 是否仍调用 `bisheng`)已由源码 +证实为"是"(§10.3.1)。 + +### 10.5 裁决判据 + +阶段一(`rules-ascendc` + 一个 hello kernel)两条: + +1. **设备执行。** 在 `sim` 模式下跑出正确结果,**断言设备名/执行路径而非结果数值** + (数值在 `cpu` 模式与静默回退时同样正确)。**不得用 `cpu` 模式充当这条** —— + 理由见 §10.3.1。`sim` 调用 `bisheng` 已证实,因此这条落在正确的对象上。 +2. **引擎无厂商知识。** 见下,这条初稿写错了。 + +#### 10.5.1 第二条判据的更正:`git diff src/` 为空分不开两种失败 + +初稿把它写成"`git diff src/` 为空"。**这个判据的"否"有两个成因,而它分不开:** + +| 若 `src/` 有改动 | 含义 | 是否推翻主张 | +|---|---|---| +| 改动里出现昇腾专有标识 | 引擎吸收了厂商知识 | **推翻** | +| 改动是 `exports` / `link-flag` 这类**通用**能力 | 引擎缺一项通用基础设施 | **不推翻** —— 那正是 §2–§6 在补的东西 | + +两者读数相同,所以这条判据在通用缺口落地**之前**跑,必然把第二种误报成第一种。 + +**更正后的两级判据:** + +| 级 | 判据 | 何时可跑 | +|---|---|---| +| **主** | 移植完成后,`tests/unit/test_core_vendor_probes.cpp` **仍然绿** —— 它以文件数为分母,断言 `src/` 去注释后不含厂商工具名 | **任何时候**。它按性质判定,不按有没有改动判定 | +| 严 | `git diff src/` 为空 | 仅在 §13 的一、二、三期落地**之后**才有意义 | + +主判据才是这条架构主张的直接检验,而且它不需要前置。严判据是附加的更强陈述。 + +主判据为否 —— 即为了让昇腾跑起来,不得不把厂商标识写进 `src/` —— 那就是"这条架构 +主张在第一个陌生厂商面前没有成立"。这个结论比移植成功更有价值,也更该早点知道。 + +### 10.6 范围警告 + +交付物是**验证物,不是承诺长期维护的 fork**。算子库切片钉在 `8.5.0`,不承诺跟随。 + +## 11. 准入自检(docs/05 附录 A) + +| 项 | 是否重复了别处已给出的答案 | +|---|---| +| `exports` | 否。没有任何 section 回答"这个产物发布哪些符号" | +| `link-flag` | 否。`ldflags` 是声明式的,本项是构建程序算出来的 | +| `role = "manifest"` | 否,且**刻意复用** `[runtime].artifacts` 而非新开 section | +| `accelerator = "none"` | 否,且**刻意复用** `os = "none"` 的既有拼法,不新造词 | +| 生成输入粒度 | 否。现有语义是包级,本项是同一概念的细化,不是第二个概念 | +| Fortran | 否 | +| 探测库 | 不是键 | + +四项均为封闭语法、开放词表:`exports` 的内容由作者定,引擎只负责渲染; +`link-flag` 的内容引擎不解释;role 的白名单加一项而语义由读者定义。 + +## 12. 判据 + +每条都要求两侧可测 —— 拿掉实现会红,而不是"没测成"与"通过"同读数。 + +| # | 判据 | +|---|---| +| C1 | 声明 `exports` 的共享库,`nm -D --defined-only` 只列出声明的符号;不声明时列出全部。两侧都断言,否则"少了几个"与"根本没链上"读数相同 | +| C2 | 同一份 `exports` 在 ELF 与 PE 上各渲染一次,两边导出集合**相同**。跨平台是这条设计的全部理由,单平台绿零信息量 | +| ~~C3~~ | **作废**(§2.3):它断言的隐含 hidden 在实现时被取消,一个键一个效果 | +| C4 | 构建程序**算出来的** `link-flag` 到达链接器。判据是**链接器的行为**而非命令行文本:e2e 620 让程序算出 `-Wl,--defsym=…=42`,产物打印那个符号的地址。grep build.ninja 会对"写下了但没交给链接器"同样成立 | +| C5 | `role = "manifest"` 的文件在打包后位于声明的相对路径上,内容里的路径为包内相对路径。判据读**打包后的产物**,不读构建目录 | +| C6 | 探测库在一台**没有宿主编译器**的机器上仍能完成探测。这是 §5.2 唯一能证伪的判据 | +| C7 | 声明 `cfg(accelerator = "none")` 的回退源码,在 `accel` 为空时编译、在任意后端被命名时不编译;**且新增一个后端后该谓词的行为不变** —— 这条才是本项的理由,单后端下绿零信息量 | +| C8 | 一个包同时具备生成头与千级编译边时,生成头**不**阻塞与它无关的编译边。对照组是今天的包级栅栏 | +| **C9** | **验证项(§10)。** 昇腾切片移植完成后:(a) `test_core_vendor_probes.cpp` 仍绿 —— 主判据,按性质判定,任何时候可跑;(b) `sim` 模式下跑出正确结果且**断言执行路径而非数值**;(c) 四项通用缺口够用,否则暴露的第五项本身就是产出。**不得用 `cpu` 模式充当 (b)** —— 它不走岛(§10.3.1) | + +**Fortran(§6.3)没有判据,因为本文没有给它设计。** 这是有意的:一个没有设计的条目配上 +一条判据,会让它看起来比实际成熟。它在 §13 里也不占期次。 + +C6 值得单独说明:它是这批里唯一无法在开发机上验证的判据 —— 开发机总有 +`/usr/bin/cc`,一个错误读宿主的实现在那里永远绿。它必须跑在 hermetic 容器里, +本仓库已有 `hermetic e2e (no host toolchain, container)` 这个 job。 + +## 13. 分期 + +| 期 | 内容 | 依据 | +|---|---|---| +| 一(**已实现** 2026.9.6.5) | `accelerator = "none"` | 最小,且它修的是一条**会随生态增长而静默失效**的写法 —— 越晚落地,要改的既有 manifest 越多 | +| 一(**已实现** 2026.9.6.5) | `link-flag` | 同样最小:一条指令 + 一处透传。且它是 §2.4 的逃生口,应先于 `exports` 落地 | +| 二(**已实现** 2026.9.6.5,不含隐含 hidden) | `exports` | 打开"可发布稳定 ABI 的 `.so`"这一档,同时惠及运行时与驱动 | +| 三 | `role = "manifest"` + `mcpp:artifact` | 只有驱动这一档需要;且受 §4.2 的发布顺序约束,越早落地引擎越好 | +| 四 | `mcpplibs:probe` | 与引擎正交,任何时候可做;但 C6 要求它一开始就跑在 hermetic job 里 | +| 五 | 生成输入粒度 | 触发条件明确(§6.2),未达到该规模前不做 | +| — | Fortran | 已识别,本文不给设计 | +| 并行 | RDC(`rules-cuda`)、`.omp` 岛 | 插件侧,不依赖以上任何一项 | +| **验证** | **昇腾切片(§10)** | 见下 | + +一、二、三期合计的引擎改动量小于 RDC 一项,且互不阻塞。RDC 归插件侧之后,**关键路径 +上不再有大件**。 + +### 13.1 验证项的位置 + +**主判据(C9a)不依赖任何一期,随时可跑。** 它按性质判定 —— 移植后 `src/` 里有没有 +出现厂商标识 —— 而不是按有没有改动判定。 + +**严判据(`git diff src/` 为空)必须在一、二、三期之后。** 在那之前跑,一次"缺通用 +能力"会被误读成"引擎吸收了厂商知识"(§10.5.1)。 + +因此推荐的顺序是:**一、二、三期落地 → 昇腾切片**。但如果想更早拿到信息,只跑主判据 +也是有效的,而且它可能提前暴露第五项通用缺口 —— 那种情况下,缺口清单本身就被验证 +补全了一次,这比等到三期做完再发现要便宜。 + +**这一项不是引擎工作量,是判断整套设计对不对的实验。** 它的产出是一份带判据的报告 +加一个规则包(§10.6),不是一个要长期维护的 fork。 + +## 14b. 昇腾工具包的取得路径(修订八:比修订三所写的更简单) + +修订三的结论是"官方镜像可匿名拉取",这仍然成立,但它不是最简的那条路,而当时没有再往下 +找一层。实测(2026-09-07,HEAD 请求,无任何凭据): + +| 版本 | `Ascend-cann-toolkit__linux-x86_64.run` | 大小 | +|---|---|---| +| 8.0.RC1 / 8.0.RC2 / 8.0.0 | 200 | 1.95 / 2.00 / 2.07 GB | +| 8.1.RC1 | 200 | 2.16 GB | +| 8.2.RC1 / 8.3.RC1 | 200 | 2.30 / 2.44 GB | +| 8.5.0 | 200 | 1.12 GB | + +URL 形如 +`https://ascend-repo.obs.cn-east-2.myhuaweicloud.com/CANN/CANN%20/Ascend-cann-toolkit__linux-x86_64.run`。 + +这条路径把 `xim:cann-toolkit` 从"要在配方里走 docker registry 的 token 握手、拉 10 个层、 +再从层里刨目录"降成"下载一个 `.run`、非交互安装、导出目录" —— 而索引里已有 `.run` 安装 +器的先例(`qemu-user-aarch64`),docker 拉取则一个先例都没有(`openclaw` 里那条是别的 +东西)。**没有先例的机制不是不可以引入,但在有先例的机制够用时引入它是净损失。** + +合规上这落在用户既定不变量的第一档:**闭源、不可再分发的东西从厂商自有 URL 直取,不做 +CN 镜像**。厂商 URL 本身就在国内,所以镜像这一层在这里也没有意义。 + +一个记法上的更正:修订三写的是"毕昇与模拟器在同一个包里,是一个获取问题不是两个"。这仍然 +对,但"那个包"现在指的是这个 `.run`,不是那个镜像。 + +## 15. 第二个后端才暴露的五处缺口(修订七) + +前十四节的缺口清单来自**读代码**与**一个后端的实践**。写第二个后端的时候,五处新缺口 +在同一个下午暴露出来,而它们没有一处能靠继续读代码找到 —— 每一处的读数都是"构建成功" +或"一条指向别处的错误消息"。 + +这本身是对方法的一次更正:**接口是否完整,第二个实例才回答**(§7 的归属规则同理)。 + +### 15.1 设备源可以什么都没编到,而构建成功 + +设备类源是引擎唯一没有编译规则的源。它经 `MCPP_DEVICE_SOURCES` 交给构建程序,要么作为 +action 回来,要么根本不被编译 —— **而没有任何东西检查它回来了**。 + +读数:链接期 `undefined reference to opkit_cuda_saxpy`。那条消息点的是符号,从不是那个 +本该定义它的文件;`kind = "lib"` 的目标连这条都没有,因为静态库不做符号解析,产物只是 +少了一个成员。 + +判据取 **action 的输入**,不取"构建程序跑过了" —— 跑了却什么都没认领是常见情形(见 +15.2)。它同时是 action 本就该满足的条件:编译某文件却不把它声明为输入的 action,在该 +文件变化时不会重跑。落地:e2e 623(三条腿,含 `--no-accel` 的反向腿)。 + +### 15.2 一条规则拿走了全部设备源 + +`mcpp::device_sources()` 是本包设备源的**全集**,同一个构建程序里每条规则读到同一个值。 +四条官方规则都拿走全集 —— 这在"一次构建只有一条规则"的前提下永远正确,而那个前提从来 +没有被写下来过。 + +实测(两条规则同时导入,未修版本): + + cuda:scale clang -x cuda shaders/scale.comp -o scale.cu.o + mcpp.rules.spirv: src/kernels/saxpy.cu has no shader stage. + +两个失败,**安静的那个更糟**:CUDA 规则没有拒绝那个着色器,它把着色器**当 CUDA 编译了** +并产出了一个 `.o`。 + +归属按 §7 的三条判据:这没有发明新的边或节点种类,也没有改变产物是什么 —— 它是"一条 +规则如何挑自己的输入",落在**插件侧**。引擎侧只补 15.1 的那条判据,因为"没有任何规则 +认领它"只有引擎看得见。落地:mcpp-plugins 0.2.2 + `tests/multi-rule-consumer`。 + +### 15.3 `accelerator` 被当成了"被解析出来的层" + +五个真正的层(`c-abi`、`compiler` …)由依赖解析回答,所以以它们为谓词的 `[xlings]` 表 +被拒绝是对的:载荷要在解析之前装好。`accelerator` 不是它们中的一个 —— 它是 `--accel`, +或 `[build] accel`,在查找第一个包之前就已读入。 + +代价每次构建都在付,而且付在最不该付的那次上:厂商工具包只能无条件声明或者干脆不声明, +于是**不带加速器的那次构建**(最便宜的、也是 CI 跑的那次)会为一个它没在编译的设备下载 +数 GB。这一条也是"把载荷放进规则包的 feature"这个提议(§15.5)成立与否的前提。 + +修法是把层键按**日程**而非按主题分成两组,并且只把"晚"的那组交给第二趟合并。 +落地:e2e 625(两个方向),以及一处必然的连带 —— `[target.'cfg(accelerator = ...)' +.dependencies]` 现在生效,因为加速器不是图给出的答案,不存在循环。 + +### 15.4 命名后端子集被当成不匹配 + +受约束 glob 的拒绝原本是对的:为 sm_89 写的文件在 targeting sm_80 的构建里不是一个变体。 +但那条判据被施加在了**跨后端**上:一个同时有 CUDA 岛与 Vulkan 岛的包,用 +`--accel "vulkan1.2"` 构建时被整体拒绝。 + +后果是结构性的:**一个包可以有多个设备后端,当且仅当每次构建都全要**。这正好是可加式 +后端(库的形态,§15 的整个语境)做不到的事。 + +修法:后端**没被命名**的 glob 像 `--no-accel` 一样被留在外面;后端被命名而架构不覆盖的 +才是不匹配。让这条安全的是同时新增的一条:glob 的后端必须在 `[package] accelerators` 里 +—— 否则 `accel = "cude12.9"` 会从"被拒绝"变成"永远不编译且无人提起"。落地:e2e 626 四条腿。 + +### 15.5 规则声明的载荷,消费者的构建程序够不到 + +规则的代码跑在**消费者的**构建程序里,所以 `mcpp::xpkg_dir` 是在那边被问的;而载荷是规则 +在自己的 `[feature-xlings]` 里声明的。依赖图早就会安装它 —— 缺的只是回答: +`fillXpkgDirs` 只读一份 manifest,于是地址被下载、解包,然后对唯一想用它的那段代码不可见。 + +读数是"工具包没装",而它就在盘上。落地:e2e 622,并在已发布的 2026.9.6.4 上跑过对照。 + +**这一条打开了一个设计选项,但没有替我们做选择。** 把厂商载荷声明进规则包的 feature, +让"启用插件即得到工具"成立,现在机制上是通的。是否要这么做取决于 15.3:feature 由依赖边 +激活,与加速器无关,所以在 15.3 之前那样做等于让每个启用 `rules-cuda` 的工程无条件下载 +整个工具包。15.3 之后仍需逐规则判断 —— 规则**自己的编译器**(glslang、dpcpp)与**厂商 +工具包**不是一回事,后者更适合留在工程侧并由工程按加速器开关。 + +### 15.6 对清单的影响 + +§6 的三处缺口是读代码找到的,§15 的五处是**用起来**找到的,而后者更多。这不说明读代码 +没用 —— §2 到 §5 的四处至今没有被推翻 —— 它说明**接口的完整性只能由第二个实例回答**, +所以"再写一个后端"应当是每一轮设计的收尾动作,而不是下一轮的开头。 + +## 14. 变更记录 + +| 版本 | 变更 | +|---|---| +| 初稿 | 四处通用缺口:`exports`、`link-flag`、包内布局、探测库。配置头与包内布局在核实 main 后各自缩小 | +| 修订一 | 两处判断被推翻并就地更正:**RDC 归插件侧,引擎零改动**(§0.2);**stdpar 可岛化,初稿的"互斥"说法过头**(§8.1)。新增三处缺口(§6)、归属规则(§7)、编程模型三分法(§8)、术语对齐建议(§9)、以及用陌生厂商证伪的实验(§10) | +| 修订二(第三处结论已被修订三推翻) | §10 的三处不确定实地调研(§10.3),两处证实一处未证实。第三处更正:**"有硬件或模拟器"不是开工硬闸**,架构验证完全在构建期(§10.4);真实前置只有毕昇的合规取得。新增两条实测结论:`cpu` 模式**不走岛**因而不能充当设备判据(§10.3.1),以及 CMake 把 Ascend C 当作一门语言 —— 同一条轴、不同归属的第三方佐证(§10.3.4) | +| 修订三 | §10.3.3 的结论被**推翻**:工具包官方镜像 `swr.cn-south-1.myhuaweicloud.com/ascendhub/cann` **可匿名拉取**(走完 token 握手实测),且从厂商自有 registry 取正落在不变量允许的一档。同时纠正一个框架错误:毕昇与模拟器**在同一个包里**,是一个获取问题不是两个。**至此该实验没有阻碍项。** 另记 AscendNPU IR 已开放,规则包存在第二个切入高度(§10.3.3b) | +| 修订六 | 一、二期实现落地(mcpp 2026.9.6.5),实现过程中三处判断被更正:**`link-flag` 到达消费者**(§3.3,`linkUsage.ldflags` 是 `buildConfig.ldflags` 的拷贝,私有形态引擎无从表达);**`exports` 不隐含 hidden**(§2.3,一个键一个效果);**C3 作废、C4 改写**。另查清 §4.3 的开放问题:packer 不携带作者声明的 artifacts,所以三期是给 packer 增加一条通路而非加一个 role | +| 修订五 | 综合复核。两处补齐:验证项此前**既无判据也无期次**,现补为 C9 与 §13.1;`kind = "device"` 此前是孤儿条目,现按"两把尺子"说明它落在引擎侧但属领域,归 docs/20。一处更正:C9 的严判据 `git diff src/` 为空**分不开两种失败**(吸收了厂商知识 vs 缺一项通用能力),改为两级判据,主判据用既有的 vendor-probe 测试(§10.5.1) | +| 修订七 | 写第二个后端时暴露的五处新缺口,全部实现并带判据(§15):设备源可以什么都没编到而构建成功;一条规则拿走全部设备源;`accelerator` 被误分类为被解析的层,导致载荷无法按设备开关;命名后端子集被当成不匹配,使可加式多后端在结构上不可能;规则声明的载荷消费者够不到。同时更正一条方法结论 —— 读代码找到四处、用起来找到五处,**接口完整性只能由第二个实例回答**(§15.6) | +| 修订四 | 最后一条待确认项闭合:`sim` **确实调用 bisheng**,由 `asc-devkit` 内 vendored 的 `ASC_CMake` 证实 —— ASC 是一门 CMake 语言,其编译器就是 `bisheng`,且**全仓 `RUN_MODE` 判断只区分 `cpu` 与非 `cpu`,不存在 `sim` 分支**。至此 §10 无待确认项 | diff --git a/.github/tools/build_examples.sh b/.github/tools/build_examples.sh index 5f4e46d9..97876ab1 100755 --- a/.github/tools/build_examples.sh +++ b/.github/tools/build_examples.sh @@ -24,6 +24,19 @@ BUILD=( examples/03-pack-static examples/04-workspace examples/08-build-rules/app + # The CPU-only path of the multi-backend example: no payloads, and it is + # where `cfg(accelerator = "none")` is exercised. The device paths are + # opt-in via --accel and are covered by the rule packages' own CI. + # Built here for one reason worth the cost: it is the only example whose + # CPU-only configuration exercises `cfg(accelerator = "none")` and two rule + # packages in one build program, and both of those are engine paths that a + # description cannot cover. Its `[toolchain] default = "llvm@22.1.8"` means + # this job installs an LLVM payload it otherwise would not -- the CUDA leg + # takes the clang route, because the nvcc route on the 12.9 line is refused + # by nvcc's own front end and the 13.x line raises the driver floor to r580. + # The device payloads are NOT installed: they are gated on the accelerator, + # and this builds without one. + examples/09-heterogeneous/multi-backend ) # `key|reason`. @@ -38,6 +51,7 @@ SKIP=( "examples/09-heterogeneous/hip/app|same, for the HIP payloads" "examples/09-heterogeneous/sycl/app|needs the dpcpp payload (over a gigabyte) and a device its runtime accepts" "examples/09-heterogeneous/vulkan/app|built AND RUN by the next step of this job, on the lavapipe payload, which needs no GPU" + "examples/09-heterogeneous/cann/app|its device leg needs the Ascend DRIVER, which a runner does not have: the kernel compiles and the object links, and then `libascend_hal.so` is missing, which is correct on a machine with no NPU. Its CPU leg does build -- and is not built here only because the plugins pin would make this job resolve a fifth rule package for one example. Covered by the measurements in its README" ) # Every ROOT manifest in the tree: a directory with an `mcpp.toml` that has no diff --git a/CHANGELOG.md b/CHANGELOG.md index bdbca7ac..f868fa8f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,128 @@ ## [Unreleased] +### 写第二个后端的时候,才发现接口只对第一个成立 + +多后端示例本身没建起来,而它暴露的四件事都不是示例的问题。 + +**设备源可以什么都没编到,而构建成功。** 设备类源是引擎唯一没有编译规则的源:它交给 +本包的构建程序,要么作为 action 回来,要么根本不被编译 —— 而**没有任何东西在检查它 +回来了**。示例没有 `build.mcpp`,于是两个设备 glob 被静默丢掉,读数是链接期 +`undefined reference to opkit_cuda_saxpy`:那条消息点的是符号,从不是那个本该定义它的 +文件;`kind = "lib"` 的目标连这条都没有,因为静态库不做解析。现在拒绝,并点名文件, +且区分「根本没有构建程序」与「程序跑了但没有 action 认领它们」——两者的修法不同。 +判据取 action 的**输入**,这同时是 action 本就该满足的条件:编译某文件却不声明它为 +输入的 action,在该文件变化时不会重跑。 + +**规则声明的载荷,消费者的构建程序够不到。** 规则的代码跑在**消费者的**构建程序里, +所以 `mcpp::xpkg_dir` 是在那边被问的,而载荷是规则在自己的 `[feature-xlings]` 里声明 +的。依赖图早就会安装它;缺的只是回答 —— `fillXpkgDirs` 只读一份 manifest,于是地址被 +下载、解包,然后对唯一想用它的那段代码不可见,读数是「工具包没装」而它就在盘上。 + +**`accelerator` 被当成了「解析出来的层」。** 五个真正的层(`c-abi`、`compiler` …)由 +依赖解析回答,所以以它们为谓词的 `[xlings]` 表被拒绝是对的。`accelerator` 不是:它是 +`--accel`,或 `[build] accel`,在查找第一个包之前就已读入。混在一起的代价每次构建都在 +付 —— 厂商工具包只能无条件声明或者不声明,于是**不带加速器的那次构建**(最便宜的、 +也是 CI 跑的那次)会为一个它没在编译的设备下载数 GB。现在 +`[target.'cfg(accelerator = "cuda")'.xlings.workspace]` 与同谓词下的 +`[dependencies]` 都生效。 + +**忘了 `host-module = true`,消息里没有这个词。** 实测读数是 GCC 的 +`failed to read compiled module` 加一句「imports 必须先被构建」—— 都对,而既没点出会 +提供这个模块的包,也没点出那个让它可导入的键。构建程序能编译的名字是一个闭集 +(`std`、`std.compat`、内置 `mcpp`、可导入的 host 模块),所以集合之外的名字在编译器 +被调用**之前**就被拒绝,并列出声明了却没写 `host-module = true` 的依赖。 + +配套判据:e2e 622(规则载荷跨到消费者)、623(设备源三条腿,含 `--no-accel` 的反向腿)、 +624(点名拒绝,且它给的修法真的能修好)、625(按加速器开关载荷,两个方向)。622 与 625 +都在已发布的 2026.9.6.4 上跑过对照并如期失败。 + +### 两个新示例,以及一处所有既有示例都写错了的地方 + +**`examples/09-heterogeneous/multi-backend`** —— 多个后端进**同一个产物**,运行期选择。 +既有四个示例每个都是**一道接缝**:设备文件与 CPU 文件定义同一个符号、永不同时进入一次 +链接,所以恰好存在一个,选择在构建期做完。那是程序的正确形态。库做不了这个选择:它 +只编译一次,而消费者的机器各不相同,所以它的后端是**叠加**的。这个示例是那种形态。 + +它同时是 `cfg(accelerator = "none")` 的第一个真实用例:分发器要在 cuda、在 vulkan、在 +两者同时、以及在一个还不存在的后端下都被构建,而 `not(any(cuda, vulkan))` 每次生态新增 +后端都要改一遍 —— 忘掉的那次是静默的。默认 `accel` 为空,所以 `mcpp build` 不需要任何 +载荷,CI 因此能真正构建它,`none` 那条路径也就真的被覆盖。 + +**`examples/09-heterogeneous/cann`** —— 昇腾。**目前构建不了**,README 点明缺的两块 +(`mcpp.rules.ascendc` 规则包,以及承载毕昇与仿真器的 `xim:cann-toolkit`),并记下已经 +查实的三件事:两者在同一个工具包里;工具包镜像可匿名拉取;`sim` 模式无需硬件且保留岛, +而 `cpu` 模式**不保留**因而不能充当设备判据。manifest 写出来而不是描述出来,是为了让 +形状具体。CANN 自己的算子库本来就是 `op_kernel/` 与 `op_host/` 分开的,岛不是 mcpp 强加 +给昇腾的形状。 + +### 规则包应当声明在 `[build-dependencies]` + +既有四个示例都把 `mcpp:plugins` 写在 `[dependencies]` 里,而 docs/05 §2.6.1 自己立的规则 +恰恰点名这种情形:**库绝不该到达目标,而它的规则仍然被需要**。两条轴是分开的 —— +`host-module = true` 说要哪种构建期产物,section 说这个包是否到达目标,规则包在第二条轴 +上答"否"。四个示例与两份文档已改正。 + +写在 `[dependencies]` 里同样能工作,这正是这条区分必须被**陈述**而不能指望由一次失败来 +教会的原因。 + + +### 共享库能说出自己发布哪些符号:`exports` + +两个平台的默认都是"全导出":ELF 给符号默认可见性,PE 由引擎自动生成列出全部符号的 +`.def`。**缺的是反方向** —— 声明式地只发布一组。 + +两类工程需要它。有稳定 ABI 的运行时只发布一份经评审的集合;与同类并存的插件不能撞名 +—— 一个把内部符号也导出的 Vulkan ICD 会与 loader 以及同进程内另一个 ICD 相撞。本仓库 +自己就有现成的例子:SYCL 示例构建时重复符号检查报的那 68 个 `_Unwind_*`,是一个镜像里 +两个 C++ 运行时都在导出 unwinder 符号。 + +`exports` 接受一个符号模式文件或一个内联数组,由引擎按平台渲染成 version script / +`-exported_symbols_list` / `.def` —— 一句中立的话三种渲染,与 `[runtime]` 已确立的形状 +相同,而不是让作者写三份平台专用文件。 + +**它不隐含编译期 hidden。** 三种格式上收窄都是链接期属性,所以一个键只有一个效果; +`-fvisibility=hidden` 仍可经 `[build] cxxflags` 取得代码生成收益,而那是单独的决定, +因为它同时改变本库各 TU 之间如何看见彼此。符号**版本化**(`foo@@LIB_1.0`)不在此列, +它是 ELF 独有、无法中立表达的能力。 + +判据 e2e 621 把同一份源码构建两次并要求两次读数**不同**:只断言公开符号在,会对"导出 +全部"同样成立(那正是本特性之前的状态);只断言内部符号不在,分不开"正确地隐藏了"与 +"根本没链上"。 + + +### 构建程序能发出它算出来的链接标志:`mcpp:link-flag` + +`link-lib`、`link-search`、`link-script` 各自命名一类东西,于是一条**算出来的**标志无处 +可去:内容依赖 feature 组合的版本脚本、运行时接管 C 库符号用的 `-Wl,--wrap=malloc`、 +以及 `-Wl,--exclude-libs,ALL`(静态吞入的第三方不得成为本包 ABI 的一部分)。签入仓库 +的标志一直可以走 `[build] ldflags`,生成的不行。 + +新增 `mcpp:link-flag=` 与 `mcpp::link_flag()`(协议 v8)。原样传递 —— 引擎不解析链接器 +词汇。**它到达消费者**,与 `[build] ldflags` 一致:编译接口有声明式公开对应物因而 +`include-dir` 必须私有,链接标志没有这个分裂,让"算出来"的形态与它自己的声明式孪生 +行为不同才是不一致。后果写明:依赖发出的 `--version-script` 也会落到消费者链接行上, +而这个隐患不是新的。 + +判据是 e2e 620,它断言**链接器的行为**而不是命令行文本:程序算出 +`-Wl,--defsym=mcpp_e2e_620=42`,产物打印那个符号的地址。值只可能来自链接器真的收到了 +这条标志。 + +### `cfg(accelerator = "none")` —— 开放词表不能靠枚举取反 + +CPU 回退此前只能写成 `not(any(accelerator = "cuda", accelerator = "vulkan"))`。 +`accelerator` 的取值是**开放的**(docs/20:第五个后端是一个包,不是引擎改动),所以这条 +谓词的含义会随生态增长**静默改变** —— 新增一个后端之后,每个已写好的回退谓词都开始把 +"命名了新后端的构建"当成"没有加速器",于是 CPU 实现与设备实现一起编进去。 + +`accelerator = "none"` 为真当且仅当加速器集合为空。拼法沿用本仓库已有的 +`os = "none"`(裸机),不新造词。不用 `cpu`:那会让这条轴同时承载两个问题,并且 +`cfg(accelerator = "cpu")` 在 `accel = "cuda"` 下的真假无法自洽地定下来。 + +判据 `test_cfg_accelerator_none.cpp` 直接**模拟第五个后端到来**:枚举写法当场开始说谎, +`none` 不变。这是这项改动的全部理由,单后端下跑绿零信息量。 + + ### 工具也有两条解析轴:`[target..xlings…]` 一条工具条目回答的是两个不同问题中的一个:它是在构建机上执行的(宿主),还是产物编译 diff --git a/docs/01-examples.md b/docs/01-examples.md index 9bddab1a..8cb133b5 100644 --- a/docs/01-examples.md +++ b/docs/01-examples.md @@ -31,11 +31,13 @@ examples. | 06 | [`examples/06-openkal-cross`](../examples/06-openkal-cross/) | One program asking each machine what it is, built for four targets from any host | `--target`, openkal, cross-compilation without editing the source | | 07 | [`examples/07-project-subos`](../examples/07-project-subos/) | A build program that finds its tools in the environment the project declared | `[xlings] subos`, `[xlings.workspace]`, a build program whose `PATH` is the environment the project named | | 08 | [`examples/08-build-rules`](../examples/08-build-rules/) | Two rule packages and a project that uses both | `host-module = true`, `[build-dependencies]`, `mcpp::action` with `role = "check"` | -| 09 | [`examples/09-heterogeneous`](../examples/09-heterogeneous/) | One computation on a device, in four programming models, with a CPU fallback in each | `accel`, constrained source globs, the seam module, rule packages from `mcpp:plugins`, `cfg(accelerator = …)` | +| 09 | [`examples/09-heterogeneous`](../examples/09-heterogeneous/) | One computation on a device, in several programming models, with a CPU fallback in each; plus one artifact carrying several backends at once | `accel`, constrained source globs, the seam module, rule packages from `mcpp:plugins`, `cfg(accelerator = …)` | | 09a | [`…/cuda`](../examples/09-heterogeneous/cuda/) | A CUDA kernel behind a seam module | `mcpp.rules.cuda`, `mcpp::action` with `role = "object"`, the driver stated as a fact and a floor | | 09b | [`…/vulkan`](../examples/09-heterogeneous/vulkan/) | The same computation as a Vulkan compute shader, on a GPU or on the CPU | `mcpp.rules.spirv`, `mcpp::action` with `role = "source"`, generated headers, a software driver as a payload | | 09c | [`…/sycl`](../examples/09-heterogeneous/sycl/) | The same computation as a SYCL kernel, compiled by a second compiler | `mcpp.rules.sycl`, the `.sycl` device extension, a chained `mcpp::action` for the device link, `compat:sycl-runtime` | | 09d | [`…/hip`](../examples/09-heterogeneous/hip/) | The same computation in HIP, reaching an NVIDIA device | `mcpp.rules.hip`, HIP as a header layer over the CUDA runtime, a two-chunk `accel` | +| 09e | [`…/multi-backend`](../examples/09-heterogeneous/multi-backend/) | Several backends in ONE artifact, chosen at run time — the library shape, not the program shape | `accel` as a set, `cfg(accelerator = "none")` and its negation, a dispatch chain, a module seam over a C island boundary | +| 09f | [`…/cann`](../examples/09-heterogeneous/cann/) | An Ascend C kernel behind the same seam. **Does not build yet** — its README names the two missing pieces | the `.asc` device extension, `op_kernel`/`op_host` as an island CANN already has, `accelerator = "none"` for the fallback | ## Suggested Reading Order diff --git a/docs/05-mcpp-toml.md b/docs/05-mcpp-toml.md index a05fea2c..9800c529 100644 --- a/docs/05-mcpp-toml.md +++ b/docs/05-mcpp-toml.md @@ -145,6 +145,48 @@ the loader opens and the import library the linker consumes, with the export list generated from the objects on the MSVC ABI (which exports nothing without `__declspec(dllexport)` or a `.def`). See `tests/e2e/08`, `257` and `259`. +#### `exports` — which symbols the artifact publishes (mcpp 2026.9.6.5+) + +```toml +[targets.mydriver] +kind = "shared" +soname = "libmydriver.so.1" +exports = "abi/mydriver.exports" # or inline: exports = ["vk_icd*"] +``` + +**Omitting the key publishes everything, which is what both platforms already +do** — ELF gives symbols default visibility, and PE gets an auto-generated +`.def` listing every symbol. `exports` narrows that. + +Two projects need the narrowing. A **runtime with a stable ABI** publishes a +reviewed set and nothing else, so that what is not in the set stays free to +change. A **plugin loaded beside its rivals** must not collide: a Vulkan ICD is +found by name for `vk_icdGetInstanceProcAddr`, and one that also exports its +internals collides with the loader and with the other ICDs in the process. + +The file lists one symbol pattern per line, `#` starts a comment, and `*` is the +only wildcard. The inline array says the same thing and is for the two or three +entry points where a separate file would be ceremony. + +One statement, three renderings: + +| Platform | Rendered as | +|---|---| +| ELF | a version script, `-Wl,--version-script=` | +| Mach-O | `-Wl,-exported_symbols_list` (the leading underscore is supplied by the engine) | +| PE | the `.def`, replacing the auto-generated all-exports one | + +**It does not change compile-time visibility, and that is deliberate.** The +narrowing is a link-time property on all three formats, so one key has one +effect. `-fvisibility=hidden` remains available through `[build] cxxflags` for +the code-generation benefit it brings, and it is a separate decision because it +also changes how this library's own translation units see each other. + +**Symbol versioning is not this key.** `foo@@LIB_1.0` alongside `foo@LIB_0.9` +is an ELF-only capability that cannot be stated neutrally; a package that needs +it writes the version script itself and passes it through `[build] ldflags`, or +computes it and emits `mcpp:link-flag=` (docs/07). + A `soname` is meaningful on `kind = "lib"` too — see [`dependency_linkage`](#dependency_linkage--static-or-shared-is-the-consumers-decision) below, where the form a library takes becomes the consumer's decision. @@ -1160,15 +1202,22 @@ for arch/env conditions and combinators. - **Predicate keys**: `os`, `arch`, `family`, `env` — the triple's coordinates — and, from mcpp 2026.9.1.1, the five target-side layer names `compiler`, `compiler-runtime`, `kernel-abi`, `c-abi`, `c++-abi` - ([14 — The Target Side](14-target-side.md)). Barewords `linux` / `macos` / + ([14 — The Target Side](14-target-side.md)). `accelerator` is a key here too + and is answered from this build's own `accel` — the backend names in + `--accel` or `[build] accel` — so it is a membership test over a set, and + `accelerator = "none"` is how a section says "this build named no backend" + without enumerating the ones it is not. Barewords `linux` / `macos` / `windows` / `unix` are sugar for the matching `os` / `family` test. A key outside this set is reported as a schema warning and the section does not apply — it used to answer false in silence, which is indistinguishable from a section that correctly did not match. -- **A layer predicate cannot select dependencies.** A layer is resolved *from* - the dependency graph, so a dependency chosen by one would decide the answer - it is asking for. `[target.'cfg(c-abi = "musl")'.dependencies]` is reported - and ignored; the `build` inputs under the same predicate do apply. +- **A resolved-layer predicate cannot select dependencies.** A layer is + resolved *from* the dependency graph, so a dependency chosen by one would + decide the answer it is asking for. `[target.'cfg(c-abi = "musl")'.dependencies]` + is reported and ignored; the `build` inputs under the same predicate do apply. + `accelerator` is not one of these (mcpp 2026.9.6.5): it is an input to the + build rather than an answer from the graph, so + `[target.'cfg(accelerator = "cuda")'.dependencies]` applies. - **Precedence**: an exact-triple table wins over a `cfg`/alias table; multiple matching predicate tables have their flags concatenated. Conditional entries are appended **after** the unconditional `[build]` ones, so under GNU @@ -2084,14 +2133,35 @@ says which targets, the feature says whether at all. "xim:shaderc" = "2026.3" ``` -**A selector here must not name a target-side layer.** `accelerator`, `c-abi`, -`c++-abi`, `compiler`, `compiler-runtime` and `kernel-abi` are answered by -dependency resolution, which happens after tools are installed and after build -programs run. A tool conditioned on one would be declared and never installed — -a build that succeeds with the tool simply absent — so such a manifest is -refused, naming both the tool and the predicate. Condition it on the target, or -gate it on a feature: `[feature-xlings.]` is known before anything is -provisioned, which is why it is the form that answers this case. +**A selector here must not name a RESOLVED layer.** `c-abi`, `c++-abi`, +`compiler`, `compiler-runtime` and `kernel-abi` are answered by dependency +resolution, which happens after tools are installed and after build programs +run. A tool conditioned on one would be declared and never installed — a build +that succeeds with the tool simply absent — so such a manifest is refused, +naming both the tool and the predicate. Condition it on the target, or gate it +on a feature: `[feature-xlings.]` is known before anything is +provisioned. + +**`accelerator` is the exception, and is admitted** (mcpp 2026.9.6.5). It is +not resolved from anything: it is `--accel`, or `[build] accel`, read before the +first package is looked up. A payload predicated on it is merged in the same +pass as a triple predicate and installed like any other. + +```toml +[target.'cfg(accelerator = "cuda")'.xlings.workspace] +"xim:cuda-nvcc" = "12.9.86" +"xim:cuda-cudart" = "12.9.79" +``` + +This is the form a project with a device island should use. Without it the +vendor toolkit is declared unconditionally or not at all, so `mcpp build` with +no accelerator — the cheapest build, and the one CI usually runs — downloaded +gigabytes for a device it was not compiling for. + +The same rule governs dependencies: `[target.'cfg(accelerator = "cuda")'.dependencies]` +is honoured, while a dependency conditioned on a resolved layer is not, because +that one would decide the answer it is asking for. Nothing about the +accelerator is circular. The selector is the only place the condition is written. A value under a selector that also carries platform keys states one fact twice, and is refused @@ -2337,9 +2407,18 @@ rules-spirv = { sources = ["rules/spirv.cppm"] } # export module mcpp.rules.spir ```toml # a consumer -[dependencies.mcpp] -plugins = { version = "0.1.1", features = ["rules-spirv"], host-module = true } -``` +[build-dependencies.mcpp] +plugins = { version = "0.2.1", features = ["rules-spirv"], host-module = true } +``` + +**`[build-dependencies]`, not `[dependencies]`** — a rule package is the case +§2.6.1 describes exactly: its library must never reach the target while its +rule is still wanted. The two axes are separate, so `host-module = true` says +*which build-time product* is wanted and the section says *whether the package +reaches the target*; a rule package answers "no" on the second axis, and the +section is where that is said. Written in `[dependencies]` it still works, and +that is precisely why the distinction has to be stated rather than enforced by +a failure. The module set is the feature set: a unit whose feature is not active is not compiled, and importing it fails as an unknown module. `mcpp:plugins` is the diff --git a/docs/07-build-mcpp.md b/docs/07-build-mcpp.md index a5d3b117..335fc2f0 100644 --- a/docs/07-build-mcpp.md +++ b/docs/07-build-mcpp.md @@ -53,6 +53,7 @@ is ignored, so diagnostics may be logged freely. | `mcpp:include-dir=` *(0.0.100+)* | add a **private** include directory (`-I`) for this package's own TUs (absolute, or relative to the package root; normalized). Replaces the `cxxflag=-I` + `cflag=-I` double emission | | `mcpp:include-dir-after=` *(0.0.100+)* | like `include-dir`, but searched **after** the system directories (`-idirafter`) — for payload trees that shadow system headers | | `mcpp:runner=` *(2026.8.19.2+)* | one argv token of the command that EXECUTES this build's artifact, when the host cannot. Emitted once per token, in order; the artifact path is appended (or substituted for `{}`). Reaches the **consumer**. Emit the executable as an ABSOLUTE path, and only **one** dependency may supply it | +| `mcpp:link-flag=` *(2026.9.6.5+)* | add a **linker flag** this program computed, verbatim. The outlet `link-lib` / `link-search` / `link-script` leave open: a generated version script (`-Wl,--version-script=`), `-Wl,--wrap=malloc` for a runtime that takes over a C-library symbol, `-Wl,--exclude-libs,ALL` so a statically absorbed third party does not become part of this package's ABI. Appended after `[build] ldflags`, in emission order. **Reaches the consumer**, exactly as `[build] ldflags` does — see below | | `mcpp:link-script=` *(2026.8.19+)* | link with this **linker script** (`-T`; relative resolves against the package root, and the emitted path is absolute because the link runs in the build directory). Reaches the **consumer**, unlike `include-dir` — a board's memory layout is the one thing a consumer cannot write for itself | | `mcpp:warning=` *(2026.8.21.2+)* | say something to the user and **keep going**. The one directive that changes no compile line, no link line and no source set. Survives the build cache — see below | | `mcpp:fact==` *(2026.9.5.2+)* | state something the program **established about the machine** (`cuda.driver=12.4`). Compared against floors before anything is compiled; see below | @@ -65,6 +66,19 @@ registry dependency — the dependency graph stays declarative in `mcpp.toml` (including platform-conditional `[target.windows.dependencies]`). `build.mcpp` is for *leaf* decisions: flags, codegen, link requirements. +`link-flag` is deliberately **not** private, and the reason is worth stating +because the opposite looks safer. A compile interface has a declarative public +counterpart (`[build] include_dirs`), so a build-time program widening it would +go behind the manifest's back — hence `include-dir`'s privateness. Link flags +have no such split: `[build] ldflags` already propagates to consumers, so a +private computed form would behave differently from its own declarative twin. + +The consequence is stated rather than hidden. A dependency emitting +`-Wl,--version-script=` puts it on the consumer's link line too, which is +usually not what that dependency meant. That hazard is not new — a dependency +writing the same flag in `[build] ldflags` has always done this — so this +directive widens *who can compute the value*, not *what the value can reach*. + `include-dir`/`include-dir-after` are deliberately **private** (Cargo discipline): they color only this package's own TUs and are never propagated to consumers. An include directory consumers must see is part of the public @@ -104,6 +118,7 @@ int main() { | `mcpp::rerun_if_changed(p)` / `mcpp::rerun_if_env_changed(v)` | the matching `rerun-*` directives | | `mcpp::rerun_if_changed_glob(pat)` *(2026.8.6.2+)* | `mcpp:rerun-if-changed-glob=` — re-run when the **set** of files matching `pat` changes (see below) | | `mcpp::dep_bin(pkg, tool)` *(2026.8.5.1+)* | reads `MCPP_DEP__BIN_` — the absolute path of a **host tool** built by a dependency (see below) | +| `mcpp::link_flag(s)` *(2026.9.6.5+)* | `mcpp:link-flag=` | | `mcpp::link_script(p)` *(2026.8.19+)* | `mcpp:link-script=` | | `mcpp::runner(tok)` *(2026.8.19.2+)* | `mcpp:runner=` — see below | | `mcpp::xpkg_dir(ns, name)` / `mcpp::xpkg_dir(name)` *(2026.8.19+)* | the payload directory of a package this manifest declared in `[xlings.workspace]`; `""` when it was not declared or is not installed (see below) | @@ -668,6 +683,45 @@ The guidance below generalises from `mcpplibs.grpcgen`, the first such package, with each of its traits judged individually. It is guidance and not a rule because none of it admits a criterion the engine could check. +**Every device source must reach an action** *(2026.9.5.2+ contract, enforced +from 2026.9.6.5)*. A device-kind file is the one source the engine has no +compile rule for: it is handed to the package's build program through +`MCPP_DEVICE_SOURCES` and comes back as an action, or it is not compiled at +all. mcpp refuses a build in which one did not, naming the files: + + error: `opkit`: device sources that no action compiles: + src/backends/cuda/saxpy.cu + src/backends/vulkan/saxpy.comp + +The criterion is the action **inputs**, not that a build program ran: a program +that ran and claimed nothing is the common case, because a rule takes the +extensions it knows and leaves the rest. It is also the condition an action +needs anyway — one that compiles a file it does not declare as an input does +not rerun when that file changes — so a rule that satisfies it is a rule that +rebuilds correctly. What it replaces is an undefined reference at the link +naming a symbol and never the file, and for a `kind = "lib"` target not even +that, because an archive is not resolved. + +**A rule takes the extensions it claims.** `mcpp::device_sources()` is the +package's whole device set, and every rule in one build program reads the same +value. A project with two backends puts a `.cu` and a `.comp` in that one list, +so a rule that consumes all of it hands its compiler a file the compiler does +not accept. A rule selects by extension, and returns without complaint when +this build names no backend it serves — a build program with several rules +calls them all. + +**An import nothing provides is refused by name.** A build program may import +`std`, `std.compat`, the bundled `mcpp`, and the host modules its dependency +edges asked for. Anything else is refused before the compiler is reached, with +the key that would have made it importable: + + error: build.mcpp imports 'mcpp.rules.spirv', and no dependency provides it + as a host module. + ... + [build-dependencies.] + = { version = "...", host-module = true } + declared without `host-module = true`: mcpp.plugins (in [build-dependencies]) + **The module name is declared by the rule's source, and `mcpp.*` is reserved.** A host module is registered under the name its interface unit declares, not under the package name, so `export module mcpp.rules.spirv;` is what a consumer diff --git a/docs/20-heterogeneous-builds.md b/docs/20-heterogeneous-builds.md index d371bb5b..db4061aa 100644 --- a/docs/20-heterogeneous-builds.md +++ b/docs/20-heterogeneous-builds.md @@ -87,6 +87,7 @@ does not accept C++20 modules. |---|---| | CUDA, HIP | `.cu`, `.hip` | | SYCL | `.sycl` (2026.9.6.1+) | +| Ascend C | `.asc`, `.cce` (2026.9.6.5+) | | GLSL, by stage | `.comp`, `.vert`, `.frag`, `.geom`, `.tesc`, `.tese`, `.mesh`, `.task`, `.rgen`, `.rint`, `.rahit`, `.rchit`, `.rmiss`, `.rcall` | | GLSL, stage-less | `.glsl` | | HLSL | `.hlsl` | diff --git a/docs/zh/01-examples.md b/docs/zh/01-examples.md index aabf1955..a7670518 100644 --- a/docs/zh/01-examples.md +++ b/docs/zh/01-examples.md @@ -28,11 +28,13 @@ mcpp build && mcpp run | 06 | [`examples/06-openkal-cross`](../../examples/06-openkal-cross/) | 同一个程序问每台机器它是什么,从任意宿主构建到四个目标 | `--target`、openkal、不改源码的交叉编译 | | 07 | [`examples/07-project-subos`](../../examples/07-project-subos/) | 构建程序在工程声明的环境里找工具,而不是问机器上恰好有什么 | `[xlings] subos`、`[xlings.workspace]`、构建程序的 `PATH` 来自工程声明的那个环境 | | 08 | [`examples/08-build-rules`](../../examples/08-build-rules/) | 两个规则包,以及同时用到它们的工程 | `host-module = true`、`[build-dependencies]`、`role = "check"` 的 `mcpp::action` | -| 09 | [`examples/09-heterogeneous`](../../examples/09-heterogeneous/) | 同一个计算在设备上跑,写成四种编程模型,每种都带 CPU 回退 | `accel`、带约束的 source glob、接缝模块、来自 `mcpp:plugins` 的规则包、`cfg(accelerator = …)` | +| 09 | [`examples/09-heterogeneous`](../../examples/09-heterogeneous/) | 同一个计算在设备上跑,写成多种编程模型,每种都带 CPU 回退;外加一个同时携带多个后端的产物 | `accel`、带约束的 source glob、接缝模块、来自 `mcpp:plugins` 的规则包、`cfg(accelerator = …)` | | 09a | [`…/cuda`](../../examples/09-heterogeneous/cuda/) | 接缝模块背后的 CUDA kernel | `mcpp.rules.cuda`、`role = "object"` 的 `mcpp::action`、把驱动陈述为 fact 与 floor | | 09b | [`…/vulkan`](../../examples/09-heterogeneous/vulkan/) | 同一个计算写成 Vulkan compute shader,在 GPU 上或在 CPU 上 | `mcpp.rules.spirv`、`role = "source"` 的 `mcpp::action`、生成的头文件、作为载荷的软件驱动 | | 09c | [`…/sycl`](../../examples/09-heterogeneous/sycl/) | 同一个计算写成 SYCL kernel,由第二个编译器编译 | `mcpp.rules.sycl`、`.sycl` 设备扩展名、为 device link 串起来的 `mcpp::action`、`compat:sycl-runtime` | | 09d | [`…/hip`](../../examples/09-heterogeneous/hip/) | 同一个计算写成 HIP,够到一台 NVIDIA 设备 | `mcpp.rules.hip`、HIP 作为 CUDA 运行时之上的一层头文件、两段式的 `accel` | +| 09e | [`…/multi-backend`](../../examples/09-heterogeneous/multi-backend/) | 多个后端进**同一个产物**,运行期选择 —— 这是库的形态,不是程序的形态 | `accel` 作为集合、`cfg(accelerator = "none")` 及其否定、分发链、C 岛边界之上的模块接缝 | +| 09f | [`…/cann`](../../examples/09-heterogeneous/cann/) | 同一道接缝背后的 Ascend C kernel。**目前还构建不了** —— README 里点明了缺的两块 | `.asc` 设备扩展名、CANN 本来就有的 `op_kernel`/`op_host` 岛、回退用 `accelerator = "none"` | ## 推荐阅读顺序 diff --git a/docs/zh/05-mcpp-toml.md b/docs/zh/05-mcpp-toml.md index a792c7dd..4e603498 100644 --- a/docs/zh/05-mcpp-toml.md +++ b/docs/zh/05-mcpp-toml.md @@ -138,6 +138,42 @@ soname = "libmylib.so.1" # 可选: Linux/ELF ABI 名称,运行时会生成同 MSVC ABI 上从对象生成导出表(该 ABI 没有 `__declspec(dllexport)` 或 `.def` 时 不导出任何符号)。参见 `tests/e2e/08`、`257`、`259`。 +#### `exports` —— 产物发布哪些符号(mcpp 2026.9.6.5+) + +```toml +[targets.mydriver] +kind = "shared" +soname = "libmydriver.so.1" +exports = "abi/mydriver.exports" # 或内联:exports = ["vk_icd*"] +``` + +**不写这个键就发布全部,而那正是两个平台今天的默认**——ELF 给符号默认可见性,PE 会 +自动生成列出全部符号的 `.def`。`exports` 把它收窄。 + +两类工程需要收窄。**有稳定 ABI 的运行时**只发布一份经过评审的集合,不在集合里的东西 +才保持可改。**与同类并存的插件**不能撞名:Vulkan loader 按名字找 +`vk_icdGetInstanceProcAddr`,一个把内部符号也导出的 ICD 会与 loader 以及同进程内另一个 +ICD 相撞。 + +文件一行一条符号模式,`#` 起注释,`*` 是唯一的通配符。内联数组说的是同一件事,用于 +只有两三个入口、单开一个文件反而是仪式的场合。 + +一句话,三种渲染: + +| 平台 | 渲染为 | +|---|---| +| ELF | version script,`-Wl,--version-script=` | +| Mach-O | `-Wl,-exported_symbols_list`(前导下划线由引擎补) | +| PE | `.def`,取代自动生成的全导出版本 | + +**它不改变编译期可见性,这是有意的。** 三种格式上收窄都是链接期属性,所以一个键只有 +一个效果。`-fvisibility=hidden` 仍可经 `[build] cxxflags` 使用以取得代码生成上的收益, +而它是一个**单独**的决定,因为它同时改变本库各翻译单元之间如何看见彼此。 + +**符号版本化不是这个键。** `foo@@LIB_1.0` 与 `foo@LIB_0.9` 并存是 ELF 独有的能力, +无法中立表达;需要它的包自己写 version script 经 `[build] ldflags` 传入,或者算出来后 +用 `mcpp:link-flag=` 发出(docs/07)。 + `soname` 对 `kind = "lib"` 同样有意义 —— 见下文的 `dependency_linkage`, 库以何种形态出现是**消费者**的决定。 @@ -1013,13 +1049,18 @@ cxxflags = ["-march=x86-64-v2"] Linux 构建**根本不会下载** `[target.windows]` 依赖。 - **谓词的键**:`os`、`arch`、`family`、`env` —— 三元组的坐标 —— 以及自 mcpp 2026.9.1.1 起的五个目标侧层名 `compiler`、`compiler-runtime`、`kernel-abi`、 - `c-abi`、`c++-abi`(见[14 —— 目标侧](14-target-side.md))。裸词 + `c-abi`、`c++-abi`(见[14 —— 目标侧](14-target-side.md))。`accelerator` 同样 + 是这里的键,由本次构建自己的 `accel`(`--accel` 或 `[build] accel` 里的后端名) + 回答,因此它是对一个集合的成员判定;`accelerator = "none"` 则是一段用来说 + 「本次构建没有命名任何后端」的写法,而不必枚举它不是的那些后端。裸词 `linux` / `macos` / `windows` / `unix` 是对应 `os` / `family` 判定的糖。 集合之外的键会被报成一条 schema 警告,且该段不生效 —— 它过去静默地求值为假, 而那与「这一段本就不该匹配」读数完全相同。 -- **层谓词不能选择依赖。** 层是**从**依赖图解析出来的,因此由它选出的依赖会决定它 - 正在询问的那个答案。`[target.'cfg(c-abi = "musl")'.dependencies]` 会被报出并忽略; - 同一谓词下的 `build` 输入照常生效。 +- **被解析的层的谓词不能选择依赖。** 层是**从**依赖图解析出来的,因此由它选出的 + 依赖会决定它正在询问的那个答案。`[target.'cfg(c-abi = "musl")'.dependencies]` + 会被报出并忽略;同一谓词下的 `build` 输入照常生效。`accelerator` 不在此列 + (mcpp 2026.9.6.5):它是构建的输入而不是图给出的答案,所以 + `[target.'cfg(accelerator = "cuda")'.dependencies]` 生效。 - **优先级**:精确三元组表胜过 `cfg`/别名表;多个命中的谓词表,其 flag 按序拼接。 条件项追加在无条件 `[build]` 项**之后**,因此在 GNU「最后一个 flag 生效」的 规则下,条件规则会覆盖更宽的无条件规则。这正是让按 OS **移除**成为可表达的原因: @@ -1776,12 +1817,29 @@ feature 说的是要不要。 "xim:shaderc" = "2026.3" ``` -**这里的 selector 禁止命名目标侧层。** `accelerator`、`c-abi`、`c++-abi`、 -`compiler`、`compiler-runtime`、`kernel-abi` 由依赖解析回答,而依赖解析发生在工具安装 -之后、构建程序运行之后。按层条件化的工具会被声明却永远装不上——构建照常成功,工具 +**这里的 selector 禁止命名被解析的层。** `c-abi`、`c++-abi`、`compiler`、 +`compiler-runtime`、`kernel-abi` 由依赖解析回答,而依赖解析发生在工具安装之后、 +构建程序运行之后。按这些层条件化的工具会被声明却永远装不上——构建照常成功,工具 就是不在——所以这样的 manifest 会被拒绝,并把工具与谓词都点出来。改成按目标条件化, -或者用 feature 做门:`[feature-xlings.]` 在任何东西被供给之前就已知,这正是 -它能回答这个场景的原因。 +或者用 feature 做门:`[feature-xlings.]` 在任何东西被供给之前就已知。 + +**`accelerator` 是例外,它被接受**(mcpp 2026.9.6.5)。它不由任何东西解析而来: +它是 `--accel`,或 `[build] accel`,在查找第一个包之前就已读入。以它为谓词的载荷 +与三元组谓词在同一趟合并,并像其它载荷一样被安装。 + +```toml +[target.'cfg(accelerator = "cuda")'.xlings.workspace] +"xim:cuda-nvcc" = "12.9.86" +"xim:cuda-cudart" = "12.9.79" +``` + +带设备孤岛的工程应当用这种写法。没有它,厂商工具包只能无条件声明或者干脆不声明, +于是不带加速器的 `mcpp build`——最便宜的那次构建,也是 CI 通常跑的那次——会为一个 +它根本没在编译的设备下载数 GB。 + +依赖同理:`[target.'cfg(accelerator = "cuda")'.dependencies]` 生效,而以被解析的层 +为条件的依赖不生效,因为后者会决定它正在询问的那个答案。加速器这条路径上没有任何 +循环。 条件只写在 selector 一处。selector 之下的值如果又带平台键,就是同一件事说了两遍, 会被拒绝,并把两半都指出来: @@ -1991,10 +2049,17 @@ rules-spirv = { sources = ["rules/spirv.cppm"] } # export module mcpp.rules.spir ```toml # 消费者 -[dependencies.mcpp] -plugins = { version = "0.1.1", features = ["rules-spirv"], host-module = true } +[build-dependencies.mcpp] +plugins = { version = "0.2.1", features = ["rules-spirv"], host-module = true } ``` +**用 `[build-dependencies]` 而不是 `[dependencies]`** —— 规则包正是 §2.6.1 描述的那种 +情形:它的库绝不该到达目标,而它的规则仍然被需要。两条轴是分开的: +`host-module = true` 说的是**要哪一种构建期产物**,而 section 说的是**这个包是否到达 +目标**;规则包在第二条轴上的答案是"否",而 section 就是说这件事的地方。写在 +`[dependencies]` 里同样能工作 —— 这恰恰是为什么这条区分必须被**陈述**,而不能指望由 +一次失败来教会。 + 模块集合就是 feature 集合:feature 未激活的单元不编译,import 它会以未知模块失败。 `mcpp:plugins` 是 mcpp 项目维护的集合(仓库 `mcpp-community/mcpp-plugins`);其成员 命名为 `mcpp.rules.`(规则包)与 `mcpp.tools.`(构建期工具)。 diff --git a/docs/zh/07-build-mcpp.md b/docs/zh/07-build-mcpp.md index 69ea8bbf..ef2aaf4c 100644 --- a/docs/zh/07-build-mcpp.md +++ b/docs/zh/07-build-mcpp.md @@ -50,6 +50,7 @@ mcpp build # 编译 + 运行 build.mcpp,然后构建工程 | `mcpp:include-dir=` *(0.0.100+)* | 为本包自身 TU 增加一个**私有** include 目录(`-I`;绝对路径或相对包根,自动规范化)。取代过去 `cxxflag=-I` + `cflag=-I` 的双重裸发 | | `mcpp:include-dir-after=` *(0.0.100+)* | 同 `include-dir`,但排在系统目录**之后**搜索(`-idirafter`)——用于会遮蔽系统头的 payload 源树 | | `mcpp:runner=` *(2026.8.19.2+)* | 执行本次构建产物的命令的**一个 argv token**(宿主跑不了它时)。一个 token 一次调用、按顺序;产物路径会被追加(或替换 `{}`)。**到达消费者**。可执行文件要发**绝对路径**,且**只能有一个**依赖提供它 | +| `mcpp:link-flag=` *(2026.9.6.5+)* | 加一条本程序**算出来的**链接标志,原样传递。这是 `link-lib` / `link-search` / `link-script` 各自命名一类东西之后留下的出口:生成的版本脚本(`-Wl,--version-script=`)、运行时接管 C 库符号用的 `-Wl,--wrap=malloc`、以及 `-Wl,--exclude-libs,ALL`(静态吞入的第三方不得成为本包 ABI 的一部分)。按发出顺序追加在 `[build] ldflags` 之后。**到达消费者**,与 `[build] ldflags` 一致 —— 理由见下 | | `mcpp:link-script=` *(2026.8.19+)* | 用这个**链接脚本**链接(`-T`;相对路径按包根解析,发出的是绝对路径,因为链接是在构建目录里跑的)。与 `include-dir` 不同,它**到达消费者** —— 板子的内存布局恰恰是消费者写不出来的那一项 | | `mcpp:warning=` *(2026.8.21.2+)* | 对用户说一句话并**继续**。唯一一条不改变编译行、链接行与源码集的指令。它**穿过构建缓存** —— 见下 | | `mcpp:fact==` *(2026.9.5.2+)* | 陈述程序**测得的机器事实**(`cuda.driver=12.4`)。在编译任何东西之前与 floor 比较;见下 | @@ -61,6 +62,15 @@ mcpp build # 编译 + 运行 build.mcpp,然后构建工程 `mcpp.toml` 里声明式管理(包括平台条件依赖 `[target.windows.dependencies]`)。 `build.mcpp` 用于*叶子*决策:开关、代码生成、链接需求。 +`link-flag` 刻意**不**私有,而这一点值得说明,因为相反的选择看上去更安全。编译接口有 +一个声明式的公开对应物(`[build] include_dirs`),所以构建期程序若能加宽它就是绕过了 +manifest —— 这正是 `include-dir` 私有的理由。链接标志没有这个分裂:`[build] ldflags` +本来就传播给消费者,因此一个私有的"算出来"形态会与它自己的声明式孪生行为不一致。 + +后果写明而不藏起来:一个依赖发出 `-Wl,--version-script=`,该标志也会落到消费者的链接 +行上,而那通常不是它的本意。这个隐患不是新的 —— 依赖在 `[build] ldflags` 里写同一条 +标志一直如此 —— 所以这条指令加宽的是**谁能算出这个值**,不是**这个值能到达哪里**。 + `include-dir`/`include-dir-after` 刻意保持**私有**(Cargo 纪律):只染色本包自身的 TU,绝不向消费者传播。需要消费者可见的 include 目录属于公共接口,应写在声明式 manifest/描述符里(`[build] include_dirs`),而不是构建期程序里。 @@ -97,6 +107,7 @@ int main() { | `mcpp::rerun_if_changed(p)` / `mcpp::rerun_if_env_changed(v)` | 对应的 `rerun-*` 指令 | | `mcpp::rerun_if_changed_glob(pat)` *(2026.8.6.2+)* | `mcpp:rerun-if-changed-glob=` —— 匹配 `pat` 的文件**集合**发生变化时重跑(见下) | | `mcpp::dep_bin(pkg, tool)` *(2026.8.5.1+)* | 读 `MCPP_DEP__BIN_` —— 依赖构建出的 **host 工具**的绝对路径(见下) | +| `mcpp::link_flag(s)` *(2026.9.6.5+)* | `mcpp:link-flag=` | | `mcpp::link_script(p)` *(2026.8.19+)* | `mcpp:link-script=` | | `mcpp::runner(tok)` *(2026.8.19.2+)* | `mcpp:runner=` —— 见下 | | `mcpp::xpkg_dir(ns, name)` / `mcpp::xpkg_dir(name)` *(2026.8.19+)* | 本 manifest 在 `[xlings.workspace]` 里声明的包的载荷目录;没声明或没安装时返回 `""`(见下) | @@ -575,6 +586,37 @@ shim,而可用的那份就在项目自己的环境里,根本不在 `PATH` 上。 下面这些从第一个规则包 `mcpplibs.grpcgen` 归纳而来,每一条特征都单独判过是必然还是偶然。 它们是指引而非规则,因为其中没有一条能给出引擎可以检查的判据。 +**每一个设备源都必须到达某个 action**(2026.9.5.2+ 的契约,自 2026.9.6.5 起强制)。 +设备类源是引擎唯一没有编译规则的源:它经 `MCPP_DEVICE_SOURCES` 交给本包的构建程序, +要么作为 action 回来,要么根本不会被编译。若有源没有回来,mcpp 拒绝这次构建并点名文件: + + error: `opkit`: device sources that no action compiles: + src/backends/cuda/saxpy.cu + src/backends/vulkan/saxpy.comp + +判据是 action 的**输入**,而不是「构建程序跑过了」:跑了却什么都没认领恰恰是常见情形, +因为一条规则只取它认识的扩展名、把其余留给别人。这同时也是 action 本就需要满足的条件 +—— 编译某个文件却不把它声明为输入的 action,在那个文件变化时不会重跑 —— 所以满足这条 +判据的规则也就是能正确增量的规则。它取代的读数是链接期的 undefined reference:那条消息 +点的是符号而从不是那个文件;而 `kind = "lib"` 的目标连这条都没有,因为静态库不做解析。 + +**一条规则只取它认领的扩展名。** `mcpp::device_sources()` 是本包设备源的**全集**, +同一个构建程序里的每条规则读到的是同一个值。带两个后端的工程会把一个 `.cu` 和一个 +`.comp` 放进这一份清单,于是把全集拿走的规则会把编译器不接受的文件递给它。规则按扩展名 +挑选,并在本次构建没有命名它所服务的后端时安静返回 —— 带多条规则的构建程序会把它们 +全部调用一遍。 + +**没有任何东西提供的 import 会被点名拒绝。** 一个构建程序可以 import 的是:`std`、 +`std.compat`、内置的 `mcpp`,以及依赖边要来的 host 模块。此外的名字在编译器被调用之前 +就被拒绝,并给出那个本该让它可导入的键: + + error: build.mcpp imports 'mcpp.rules.spirv', and no dependency provides it + as a host module. + ... + [build-dependencies.] + = { version = "...", host-module = true } + declared without `host-module = true`: mcpp.plugins (in [build-dependencies]) + **模块名由规则的源码声明,`mcpp.*` 是保留前缀。** host 模块以其接口单元声明的名字注册, 而不是以包名注册,所以 `export module mcpp.rules.spirv;` 就是消费者 import 的那个名字。 官方插件集中在一个包里,`mcpp:plugins`(仓库 `mcpp-community/mcpp-plugins`):规则包命名为 diff --git a/docs/zh/20-heterogeneous-builds.md b/docs/zh/20-heterogeneous-builds.md index 1f8cf975..0c9b358a 100644 --- a/docs/zh/20-heterogeneous-builds.md +++ b/docs/zh/20-heterogeneous-builds.md @@ -70,6 +70,7 @@ C++20 modules 的编译器。 |---|---| | CUDA、HIP | `.cu`、`.hip` | | SYCL | `.sycl`(2026.9.6.1+) | +| Ascend C | `.asc`、`.cce`(2026.9.6.5+) | | GLSL(按 stage) | `.comp`、`.vert`、`.frag`、`.geom`、`.tesc`、`.tese`、`.mesh`、`.task`、`.rgen`、`.rint`、`.rahit`、`.rchit`、`.rmiss`、`.rcall` | | GLSL(无 stage) | `.glsl` | | HLSL | `.hlsl` | diff --git a/examples/05-lib-distribution/consumer/mcpp.toml b/examples/05-lib-distribution/consumer/mcpp.toml index f144f3d5..53cffc31 100644 --- a/examples/05-lib-distribution/consumer/mcpp.toml +++ b/examples/05-lib-distribution/consumer/mcpp.toml @@ -7,6 +7,25 @@ license = "Apache-2.0" # A packed library is an ordinary package: a path dependency, a git dependency, # a downloaded archive and an index entry all reach it the same way. Point this # at whatever `mcpp pack mathkit` left under ../producer/target/dist/. + +# THE TOOLCHAIN IS PINNED BECAUSE THE CONSUMER NAMES THE TAG. +# +# `mcpp pack` writes a directory whose name carries the ABI tag -- +# `mathkit-0.1.0-x86_64-linux-gnu-gcc16-libstdcxx16-c++23` -- and the consumer +# beside this one names that directory literally, which is what the README +# tells a reader to do. The tag therefore depends on which toolchain built it, +# and without this pin the pair only agrees on a machine whose DEFAULT +# toolchain happens to be gcc 16. Measured: a CI runner whose sandbox had been +# left with an LLVM default produced +# `mathkit-0.1.0-x86_64-linux-gnu-clang22-libcxx22-c++23`, and the consumer +# named a directory that did not exist. +# +# An example that hard-codes an ABI tag has to name the toolchain that +# produces it. That is not a workaround: a packed artifact IS specific to the +# ABI it was built against, which is the whole subject of this example. +[toolchain] +default = "gcc@16.1.0" + [dependencies] mathkit = { path = "../producer/target/dist/mathkit-0.1.0-x86_64-linux-gnu-gcc16-libstdcxx16-c++23" } diff --git a/examples/05-lib-distribution/producer/mcpp.toml b/examples/05-lib-distribution/producer/mcpp.toml index c3563b41..230d3bff 100644 --- a/examples/05-lib-distribution/producer/mcpp.toml +++ b/examples/05-lib-distribution/producer/mcpp.toml @@ -4,6 +4,25 @@ version = "0.1.0" description = "Demo: shipping a prebuilt library with both a header and a module interface" license = "Apache-2.0" + +# THE TOOLCHAIN IS PINNED BECAUSE THE CONSUMER NAMES THE TAG. +# +# `mcpp pack` writes a directory whose name carries the ABI tag -- +# `mathkit-0.1.0-x86_64-linux-gnu-gcc16-libstdcxx16-c++23` -- and the consumer +# beside this one names that directory literally, which is what the README +# tells a reader to do. The tag therefore depends on which toolchain built it, +# and without this pin the pair only agrees on a machine whose DEFAULT +# toolchain happens to be gcc 16. Measured: a CI runner whose sandbox had been +# left with an LLVM default produced +# `mathkit-0.1.0-x86_64-linux-gnu-clang22-libcxx22-c++23`, and the consumer +# named a directory that did not exist. +# +# An example that hard-codes an ABI tag has to name the toolchain that +# produces it. That is not a workaround: a packed artifact IS specific to the +# ABI it was built against, which is the whole subject of this example. +[toolchain] +default = "gcc@16.1.0" + [build] sources = ["src/*.cppm", "src/*.cpp", "src/*.c"] # The public headers. Published whole — see include/mathkit_c.h. diff --git a/examples/09-heterogeneous/cann/app/README.md b/examples/09-heterogeneous/cann/app/README.md new file mode 100644 index 00000000..0447fd15 --- /dev/null +++ b/examples/09-heterogeneous/cann/app/README.md @@ -0,0 +1,81 @@ +# Ascend, through the same island + +`op_kernel/` beside `op_host/` is how CANN's own operator libraries are already +laid out — every operator in `ops-math` splits that way on disk. The island is +therefore not a shape mcpp imposes on Ascend; it is the shape Ascend already +has, and this example writes it in mcpp's vocabulary. + +`.asc` is that seam made checkable, exactly as `.sycl` is in the SYCL example: +the file's content is C++ and nothing in it would tell a reader otherwise. What +makes it a device translation unit is that it goes to BiSheng — a compiler with +a device back end, and one that does not accept C++20 modules. + +## What builds, and what does not + +```bash +mcpp run --no-accel # builds and runs +mcpp build --accel "ascend8.5+{dav-c220}" # compiles the kernel, links, then + # stops on the missing driver +``` + +Measured on an x86_64 machine with **no Ascend hardware and no Ascend driver**: + +| step | result | +|---|---| +| `xim:cann-toolkit` provisioned | 2.9 GB, no root, no driver | +| `build.mcpp` compiles and runs | `mcpp.rules.ascendc` imported from `mcpp:plugins` | +| the kernel compiles | `bisheng -x asc --cce-aicore-arch=dav-c220` | +| the object joins the ordinary link | mixed mode: an x86-64 object carrying the device binary | +| the host half links | ACL, plus the six-library closure the rule names | +| the artifact starts | **no** -- `libascend_hal.so` is missing | +| `--no-accel` | builds and runs: `12 24 36 48`, `device: cpu` | + +`libascend_hal.so` belongs to the **driver**, not the toolkit, and is the role +`libcuda.so.1` plays for CUDA: in ABI lockstep with the kernel module, not +redistributable, and absent on a machine with no NPU. A device build of this +example therefore completes everywhere and *runs* only on an Ascend machine -- +which is the same statement `examples/09-heterogeneous/cuda` makes about a +machine with no NVIDIA driver, and the reason both are skipped by CI. + +The payloads are gated on the accelerator, so `mcpp run --no-accel` installs +nothing: the CPU leg costs a C++ compile and no download at all. + +## What the toolkit turned out to be + +The design this follows expected the toolkit to be hard to obtain. It is not. +Every CANN toolkit from 8.0.RC1 to 8.5.0 is a plain `.run` on Huawei's own OBS, +answering 200 to an anonymous HEAD request, and it installs unattended: + +```bash +./Ascend-cann-toolkit_8.5.0_linux-x86_64.run --install \ + --install-path= --quiet +``` + +Both halves the lane needs are inside it: + +| | | +|---|---| +| `-linux/ccec_compiler/bin/{ccec,bisheng}` | the device compiler, clang 15.0.5 | +| `-linux/simulator//lib/libpem_davinci.so` | **38 SoCs**, no hardware required | + +The second is why the lane is verifiable without an NPU at all, and it is the +next thing this example should use: the kernel compiles today, and running it +under `libpem_davinci` is a separate piece of work with its own contract. + +## The seam is a C function, and that is measured rather than stylistic + +BiSheng's own launcher for a `__global__` function is **C++-mangled** even when +the kernel is declared `extern "C"`. Calling it from the host half would make +this program depend on BiSheng and the project's C++ compiler agreeing about +mangling -- two different compilers, one of them clang 15 and the other +whatever the project chose. The `.asc` file therefore exports an `extern "C"` +wrapper, and the `<<<...>>>` launch spelling never leaves the translation unit +the device compiler owns. + +## Why `accelerator = "none"` for the fallback + +`not(accelerator = "ascend")` would work today and rot tomorrow: `accelerator` +is an open vocabulary, so a fallback written by enumerating what it is not +changes meaning every time the ecosystem gains a backend. Ascend is itself an +instance of that growth, which is the neatest possible argument for the +spelling. diff --git a/examples/09-heterogeneous/cann/app/build.mcpp b/examples/09-heterogeneous/cann/app/build.mcpp new file mode 100644 index 00000000..df15ada6 --- /dev/null +++ b/examples/09-heterogeneous/cann/app/build.mcpp @@ -0,0 +1,14 @@ +import std; +import mcpp; +import mcpp.rules.ascendc; + +// The kernel is compiled by BiSheng, which mcpp does not drive. Everything the +// rule needs is already in the manifest -- the architecture in `[build] accel`, +// the kernel in the constrained glob, the toolkit under `[xlings.workspace]` -- +// so this program names the project's own include directory and says go. +int main() { + mcpp::rerun_if_changed_glob("src/kernels/**/*.asc"); + mcpp::rules::ascendc::options opt; + opt.includes = { "include" }; + return mcpp::rules::ascendc::compile(opt) ? 0 : 1; +} diff --git a/examples/09-heterogeneous/cann/app/include/saxpy/saxpy.h b/examples/09-heterogeneous/cann/app/include/saxpy/saxpy.h new file mode 100644 index 00000000..feb12726 --- /dev/null +++ b/examples/09-heterogeneous/cann/app/include/saxpy/saxpy.h @@ -0,0 +1,24 @@ +// The Ascend island's interface. +// +// `extern "C"` and free of standard-library types, for the reason every island +// in this directory gives: the device half is compiled by BiSheng, a compiler +// mcpp did not choose, so the two sides share no C++ ABI. +#ifndef MCPP_EXAMPLE_ASCEND_SAXPY_H +#define MCPP_EXAMPLE_ASCEND_SAXPY_H + +#ifdef __cplusplus +extern "C" { +#endif + +// out[i] = a * x[i] + y[i], computed on the NPU. Returns 0 on success. +int saxpy_device(float a, const float* x, const float* y, float* out, unsigned n); + +// Which device the last successful call ran on, or "" if none has. Both +// implementations of this seam produce the same numbers, so the numbers alone +// do not separate a device run from the fallback. +const char* saxpy_device_name(void); + +#ifdef __cplusplus +} +#endif +#endif diff --git a/examples/09-heterogeneous/cann/app/mcpp.toml b/examples/09-heterogeneous/cann/app/mcpp.toml new file mode 100644 index 00000000..63abed31 --- /dev/null +++ b/examples/09-heterogeneous/cann/app/mcpp.toml @@ -0,0 +1,61 @@ +[package] +name = "ascend-saxpy" +namespace = "example" +version = "0.1.0" +description = "An Ascend C kernel behind a seam module, with a CPU fallback" +accelerators = ["ascend"] + +[language] +standard = "c++23" +modules = true +import_std = true + +# The rule that compiles the kernel, selected by its feature; `build.mcpp` +# imports it as `mcpp.rules.ascendc`. `[build-dependencies]`, because a rule +# package's library must never reach the target while its rule is wanted. +[build-dependencies.mcpp] +plugins = { version = "0.2.3", features = ["rules-ascendc"], host-module = true } + +# The toolkit carries BOTH halves this lane needs, so it is one payload rather +# than two: +# +# /cann/-linux/ccec_compiler/bin/bisheng the device compiler +# /cann/-linux/simulator//lib 38 SoCs, no hardware +# +# Gated on the accelerator: it is 2.9 GB installed, and `mcpp build` with no +# accelerator has no use for it. That gating needs mcpp 2026.9.6.5; before it +# the only spellings were "unconditionally" and "not at all". +[target.'cfg(accelerator = "ascend")'.xlings.workspace] +"xim:cann-toolkit" = "8.5.0" + +[build] +# `dav-2201` is the device architecture, the role `sm_89` plays for CUDA. The +# rule package derives BiSheng's own flag from it. +accel = "ascend8.5+{dav-c220}" +sources = [ + "src/*.cppm", + "src/*.cpp", + # The kernel carries the accel it is for. It is never offered to the C++ + # compiler; the constrained glob routes it to the build program instead. + { glob = "src/kernels/*.asc", accel = "ascend8.5+{dav-c220}" }, +] +include_dirs = ["include"] + +# The CPU variant of the same seam. `accelerator = "none"` rather than +# `not(accelerator = "ascend")`: `accelerator` is an open vocabulary, so a +# fallback written by enumeration changes meaning as the ecosystem grows. +[target.'cfg(accelerator = "none")'.build] +sources = ["src/cpu/*.cpp"] + +# The HOST half of the island: it launches the kernel through ACL, and declines +# when no NPU is present -- which is every machine that is not an Ascend one, +# including the one this example was written on. Declining is the contract the +# seam is built around, so a device build on a hostless machine still links and +# still runs; it simply reports that it found no device. +[target.'cfg(accelerator = "ascend")'.build] +sources = ["src/ascend/*.cpp"] +ldflags = ["-lascendcl"] + +[targets.ascend-saxpy] +kind = "bin" +main = "src/main.cpp" diff --git a/examples/09-heterogeneous/cann/app/src/app.cppm b/examples/09-heterogeneous/cann/app/src/app.cppm new file mode 100644 index 00000000..cbaf8857 --- /dev/null +++ b/examples/09-heterogeneous/cann/app/src/app.cppm @@ -0,0 +1,25 @@ +// The seam, as a module. +// +// Its reason for existing is not that BiSheng rejects modules. It is that this +// is the one place a backend can be exchanged: the island underneath can become +// CUDA or a CPU fallback without a single importer changing. +module; +#include "saxpy/saxpy.h" +export module app.saxpy; +import std; + +export namespace app { + +std::optional> +saxpy(float a, std::span x, std::span y) { + if (x.size() != y.size()) return std::nullopt; + std::vector out(x.size()); + if (saxpy_device(a, x.data(), y.data(), out.data(), + static_cast(x.size())) != 0) + return std::nullopt; + return out; +} + +std::string_view device_name() { return saxpy_device_name(); } + +} // namespace app diff --git a/examples/09-heterogeneous/cann/app/src/ascend/saxpy.cpp b/examples/09-heterogeneous/cann/app/src/ascend/saxpy.cpp new file mode 100644 index 00000000..e4692ee0 --- /dev/null +++ b/examples/09-heterogeneous/cann/app/src/ascend/saxpy.cpp @@ -0,0 +1,93 @@ +// The Ascend island's HOST half: it launches the kernel through ACL. +// +// IT DECLINES WHEN NO NPU IS PRESENT, and that is the contract rather than a +// shortcut. `aclInit` and `aclrtSetDevice` fail on a machine with no Ascend +// device -- which is every machine this example was developed on -- so the +// function returns non-zero and the caller falls back. A device build +// therefore still links and still runs on a machine with no device; it simply +// says so, which is the property that lets one artifact serve both. +// +// The kernel itself was compiled by BiSheng into a Da Vinci object and linked +// into this binary by the ordinary link. Nothing here compiles device code. +#include "saxpy/saxpy.h" + +#include + +#include +#include +#include + +// The launcher the `.asc` file exports. `extern "C"` for the reason the seam +// header gives, and for a sharper one measured here: BiSheng's own launcher +// symbol is C++-mangled, so a C++ declaration would make this program depend +// on BiSheng and g++ agreeing about mangling. The wrapper is one function in +// the device translation unit and removes that dependency. +extern "C" void saxpy_launch(std::uint32_t blockDim, void* stream, + std::uint8_t* x, std::uint8_t* y, std::uint8_t* out, + float a, std::uint32_t n); + +namespace { +char g_ran_on[128] = ""; + +// One place to leave ACL in the state it was found in, whichever step failed. +struct acl_session { + bool inited = false, device = false; + aclrtStream stream = nullptr; + ~acl_session() { + if (stream) aclrtDestroyStream(stream); + if (device) aclrtResetDevice(0); + if (inited) aclFinalize(); + } +}; +} // namespace + +extern "C" const char* saxpy_device_name(void) { return g_ran_on; } + +extern "C" int saxpy_device(float a, const float* x, const float* y, + float* out, unsigned n) { + acl_session s; + if (aclInit(nullptr) != ACL_SUCCESS) return 1; + s.inited = true; + + std::uint32_t count = 0; + if (aclrtGetDeviceCount(&count) != ACL_SUCCESS || count == 0) return 1; + if (aclrtSetDevice(0) != ACL_SUCCESS) return 1; + s.device = true; + if (aclrtCreateStream(&s.stream) != ACL_SUCCESS) return 1; + + const std::size_t bytes = static_cast(n) * sizeof(float); + void *dx = nullptr, *dy = nullptr, *dout = nullptr; + auto release = [&] { + if (dx) aclrtFree(dx); + if (dy) aclrtFree(dy); + if (dout) aclrtFree(dout); + }; + if (aclrtMalloc(&dx, bytes, ACL_MEM_MALLOC_HUGE_FIRST) != ACL_SUCCESS + || aclrtMalloc(&dy, bytes, ACL_MEM_MALLOC_HUGE_FIRST) != ACL_SUCCESS + || aclrtMalloc(&dout, bytes, ACL_MEM_MALLOC_HUGE_FIRST) != ACL_SUCCESS) { + release(); + return 1; + } + if (aclrtMemcpy(dx, bytes, x, bytes, ACL_MEMCPY_HOST_TO_DEVICE) != ACL_SUCCESS + || aclrtMemcpy(dy, bytes, y, bytes, ACL_MEMCPY_HOST_TO_DEVICE) != ACL_SUCCESS) { + release(); + return 1; + } + + // One block: this example is about the build, and a tiling strategy would + // be the subject of a different one. + saxpy_launch(1, s.stream, + static_cast(dx), static_cast(dy), + static_cast(dout), a, n); + if (aclrtSynchronizeStream(s.stream) != ACL_SUCCESS) { release(); return 1; } + + if (aclrtMemcpy(out, bytes, dout, bytes, ACL_MEMCPY_DEVICE_TO_HOST) != ACL_SUCCESS) { + release(); + return 1; + } + release(); + + const char* name = aclrtGetSocName(); + std::snprintf(g_ran_on, sizeof g_ran_on, "%s", name ? name : "ascend"); + return 0; +} diff --git a/examples/09-heterogeneous/cann/app/src/cpu/saxpy.cpp b/examples/09-heterogeneous/cann/app/src/cpu/saxpy.cpp new file mode 100644 index 00000000..ac70711c --- /dev/null +++ b/examples/09-heterogeneous/cann/app/src/cpu/saxpy.cpp @@ -0,0 +1,13 @@ +#include "saxpy/saxpy.h" + +// The CPU variant of the same seam, selected by `cfg(accelerator = "none")`. +// This file and the Ascend host half define the same symbols and are never in +// one link. +extern "C" int saxpy_device(float a, const float* x, const float* y, + float* out, unsigned n) { + for (unsigned i = 0; i < n; ++i) out[i] = a * x[i] + y[i]; + return 0; +} +extern "C" const char* saxpy_device_name(void) { + return "cpu (this build names no accelerator)"; +} diff --git a/examples/09-heterogeneous/cann/app/src/kernels/saxpy.asc b/examples/09-heterogeneous/cann/app/src/kernels/saxpy.asc new file mode 100644 index 00000000..3c5b4334 --- /dev/null +++ b/examples/09-heterogeneous/cann/app/src/kernels/saxpy.asc @@ -0,0 +1,39 @@ +// The Ascend C kernel. +// +// `.asc` is the seam made checkable, the way `.sycl` is in the SYCL example: +// its content is C++, and nothing in the file would tell a reader otherwise. +// What makes it a device translation unit is that it goes to BiSheng -- a +// compiler with a device back end, and one that does not accept C++20 modules. +// +// CANN's own operator libraries already have this shape on disk: every operator +// in `ops-math` splits into `op_kernel/` and `op_host/`, so the island is not +// something mcpp imposes here. +#include "kernel_operator.h" + +extern "C" __global__ __aicore__ void saxpy_kernel( + GM_ADDR x, GM_ADDR y, GM_ADDR out, float a, uint32_t n) { + AscendC::GlobalTensor gx, gy, go; + gx.SetGlobalBuffer((__gm__ float*)x, n); + gy.SetGlobalBuffer((__gm__ float*)y, n); + go.SetGlobalBuffer((__gm__ float*)out, n); + for (uint32_t i = AscendC::GetBlockIdx(); i < n; i += AscendC::GetBlockNum()) + go.SetValue(i, a * gx.GetValue(i) + gy.GetValue(i)); +} + +// THE SEAM, compiled by BiSheng along with the kernel above. +// +// BiSheng's own launcher for a `__global__` function is C++-mangled even when +// the kernel is declared `extern "C"`, so calling it from the host half would +// make the program depend on two different compilers agreeing about name +// mangling. An `extern "C"` wrapper here costs one function and removes that +// dependency entirely: the host half sees a C symbol, which is what the seam +// header already promises. +// +// `<<>>` is BiSheng syntax and exists only in this +// translation unit. That is the island: the launch spelling never leaves the +// file the device compiler owns. +extern "C" void saxpy_launch(uint32_t blockDim, void* stream, + GM_ADDR x, GM_ADDR y, GM_ADDR out, + float a, uint32_t n) { + saxpy_kernel<<>>(x, y, out, a, n); +} diff --git a/examples/09-heterogeneous/cann/app/src/main.cpp b/examples/09-heterogeneous/cann/app/src/main.cpp new file mode 100644 index 00000000..c298e781 --- /dev/null +++ b/examples/09-heterogeneous/cann/app/src/main.cpp @@ -0,0 +1,14 @@ +import std; +import app.saxpy; + +int main() { + const std::vector x{1, 2, 3, 4}, y{10, 20, 30, 40}; + auto out = app::saxpy(2.0f, x, y); + if (!out) { std::println("device unavailable"); return 1; } + for (auto v : *out) std::print("{} ", v); + std::println(""); + // After the call, never before: the name records a run that happened. + std::println("device: {}", app::device_name()); + const std::vector want{12, 24, 36, 48}; + return *out == want ? 0 : 1; +} diff --git a/examples/09-heterogeneous/cuda/app/mcpp.toml b/examples/09-heterogeneous/cuda/app/mcpp.toml index 64dabe2b..ba18eb22 100644 --- a/examples/09-heterogeneous/cuda/app/mcpp.toml +++ b/examples/09-heterogeneous/cuda/app/mcpp.toml @@ -19,7 +19,7 @@ default = "llvm@22.1.8" # The rule that compiles the island lives in the official plugin collection, # selected by its feature; `build.mcpp` imports it as `mcpp.rules.cuda`. -[dependencies.mcpp] +[build-dependencies.mcpp] plugins = { version = "0.2.1", features = ["rules-cuda"], host-module = true } # The driver's userspace library, reached through an index package that owns diff --git a/examples/09-heterogeneous/hip/app/mcpp.toml b/examples/09-heterogeneous/hip/app/mcpp.toml index 24eead2e..015ac711 100644 --- a/examples/09-heterogeneous/hip/app/mcpp.toml +++ b/examples/09-heterogeneous/hip/app/mcpp.toml @@ -16,7 +16,7 @@ import_std = true [toolchain] default = "llvm@22.1.8" -[dependencies.mcpp] +[build-dependencies.mcpp] plugins = { version = "0.2.1", features = ["rules-hip"], host-module = true } # The driver's userspace library. HIP reaches the device through the CUDA diff --git a/examples/09-heterogeneous/multi-backend/README.md b/examples/09-heterogeneous/multi-backend/README.md new file mode 100644 index 00000000..5a312e2a --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/README.md @@ -0,0 +1,104 @@ +# Several backends in one artifact, chosen at run time + +The four examples beside this one are each **one seam**: a device file and a CPU +file define the same symbol and are never in one link, so exactly one exists and +the choice is made at build time. That is the right shape for a program. + +A library cannot make that choice. It is compiled once and consumed by people +whose machines differ, so its backends are **additive** — several land in one +artifact and the choice moves to run time. This is that shape, in the smallest +form that still shows it. + +## What it demonstrates + +| | | +|---|---| +| `accel` is a **set** | `--accel "cuda12.9+{sm_89}, vulkan1.2"` compiles both islands into one artifact | +| A constrained glob gates itself | `{ glob = "…/*.cu", accel = "…" }` reaches the build program only when this build's `accel` accepts it, so device sources need no `cfg` block | +| `cfg(accelerator = "none")` | the CPU-only variant, selected without enumerating the backends it is not | +| `cfg(not(accelerator = "none"))` | the dispatcher, built whenever **some** backend is named | +| A module seam | `main.cpp` does `import opkit;`; the C boundary exists only where an island requires it | + +## Building it + +```bash +mcpp run # CPU only, no payloads +mcpp run --accel "vulkan1.2" # + the Vulkan island +mcpp run --accel "cuda12.9+{sm_89}" # + the CUDA island +mcpp run --accel "cuda12.9+{sm_89}, vulkan1.2" # both, one artifact +``` + +Measured, on a machine with an RTX 4080 and a 12.4 driver: + +| build | prints | +|---|---| +| `mcpp run` | `backend: cpu (only backend in this build)` | +| `--accel "vulkan1.2"` | `backend: vulkan (NVIDIA GeForce RTX 4080)` | +| `--accel "cuda12.9+{sm_89}"` | `backend: cuda` | +| both | `backend: cuda` — the chain's first entry answers | + +All four print `12 24 36 48`, which is exactly why the backend is printed +first: every backend returns the same four numbers, so the numbers alone cannot +separate a device run from the reference one. + +**Naming a subset is not a mismatch.** `--accel "vulkan1.2"` leaves the `.cu` +glob out the way `--no-accel` leaves both out, and the `cfg(accelerator = +"cuda")` section carrying that backend's host half does not activate either, so +the two halves stay together. Only an accelerator this build *does* name whose +architecture it does not cover is refused (mcpp 2026.9.6.5). + +**Nothing is installed for a device this build did not name.** The payloads sit +under `[target.'cfg(accelerator = ...)'.xlings.workspace]`, so `mcpp run` +fetches neither the CUDA toolkit nor the shader compiler. That gating needs +mcpp 2026.9.6.5; before it, the only spellings available were "unconditionally" +and "not at all", and the cheapest build paid for the most expensive one. + +## The CUDA leg takes the clang route + +`[toolchain] default = "llvm@22.1.8"`, and the reason is measured rather than +stylistic. On the 12.9 line the nvcc route is refused by nvcc's own front end: +the toolkit headers redeclare the C23 `cospi`, `sinpi` and `rsqrt` for the host +without `noexcept` while the C library declares them with it. Driving an older +`xim:gcc` payload does not help — the declarations come from the C library, not +from the host compiler, and this was tried. The 13.x line fixes it and raises +the driver floor to r580, which is a requirement on the machine rather than a +decision the project gets to make. The clang route never includes that header +and runs on any driver from r525 onward. + +## Why the dispatcher's predicate matters + +The registry has to be built for cuda, for vulkan, for both, and for a backend +that does not exist yet. Written as + +```toml +[target.'cfg(not(any(accelerator = "cuda", accelerator = "vulkan")))'.build] +``` + +it would have to be edited every time the ecosystem gains a backend — and +`accelerator` is an **open** vocabulary, so a third backend is a package rather +than an engine change. The edit that is forgotten is silent: the CPU-only +dispatcher and the registry would both compile, or neither would. + +`accelerator = "none"` says "this build named no backend" directly, and keeps +saying it after the vocabulary grows. + +## Where the rule package is declared + +```toml +[build-dependencies.mcpp] +plugins = { version = "0.2.2", features = ["rules-cuda", "rules-spirv"], host-module = true } +``` + +Two rules, in one build program, which is what an additive-backend package +needs and what no other example here has. `mcpp::device_sources()` is the +package's whole device set, so in a build naming both backends that one list +holds a `.cu` and a `.comp`: each rule takes the extensions it claims and +leaves the rest, which is what 0.2.2 fixed. A device source no rule claims is +not silently dropped either — mcpp refuses a device source that reached no +action, naming the file. + +`[build-dependencies]`, not `[dependencies]`: a rule package's library must +never reach the target while its rule is still wanted, which is the case +docs/05 section 2.6.1 exists for. `host-module = true` says which build-time +product is wanted; the section says whether the package reaches the target. +Two axes, and a rule package answers "no" on the second. diff --git a/examples/09-heterogeneous/multi-backend/build.mcpp b/examples/09-heterogeneous/multi-backend/build.mcpp new file mode 100644 index 00000000..f8490140 --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/build.mcpp @@ -0,0 +1,30 @@ +import std; +import mcpp; +import mcpp.rules.cuda; +import mcpp.rules.spirv; + +// TWO RULES IN ONE BUILD PROGRAM, which is what a project with several +// backends needs and what the four examples beside this one never exercised. +// +// Both are called unconditionally. Neither is told which backends this build +// named -- each parses `[build] accel` itself and returns immediately when its +// own is absent, so `mcpp build` with no accel runs both and compiles nothing. +// That is also why the order here carries no meaning. +// +// Each rule takes the device sources whose EXTENSION it claims and leaves the +// rest: `mcpp::device_sources()` is the package's whole device set, and in a +// build that names both backends it holds a `.cu` and a `.comp`. A device +// source no rule claims is not silently dropped -- the engine refuses a device +// source that reached no action. +int main() { + mcpp::rerun_if_changed_glob("src/backends/**/*.cu"); + mcpp::rerun_if_changed_glob("src/backends/**/*.comp"); + + mcpp::rules::cuda::options cu; + cu.includes = { "include" }; + if (!mcpp::rules::cuda::compile(cu)) return 1; + + mcpp::rules::spirv::options sp; + sp.includes = { "src/backends/vulkan" }; + return mcpp::rules::spirv::compile(sp) ? 0 : 1; +} diff --git a/examples/09-heterogeneous/multi-backend/include/opkit/opkit.h b/examples/09-heterogeneous/multi-backend/include/opkit/opkit.h new file mode 100644 index 00000000..9e53573d --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/include/opkit/opkit.h @@ -0,0 +1,44 @@ +// opkit -- one operator, several device backends, one artifact. +// +// `extern "C"` and free of standard-library types, for the reason the other +// examples in this directory give: a device island is compiled by a compiler +// mcpp did not choose, so the two sides share no C++ ABI. +#ifndef MCPP_EXAMPLE_OPKIT_H +#define MCPP_EXAMPLE_OPKIT_H + +#ifdef __cplusplus +extern "C" { +#endif + +// out[i] = a * x[i] + y[i]. Returns 0 on success. +// +// WHICH backend answers is decided at RUN time among those this build +// compiled in, which is what makes this an operator library rather than four +// separate programs. +int opkit_saxpy(float a, const float* x, const float* y, float* out, unsigned n); + +// The backend that served the last successful call, or "" before one. +// +// An operator library that computes and does not say where cannot be checked: +// every backend returns the same numbers, so the numbers alone do not +// distinguish a device run from the reference one. +const char* opkit_backend(void); + +// Each backend supplies these two. A backend that is not compiled in is not +// declared, so the dispatcher's list is decided at compile time. +int opkit_cpu_saxpy(float, const float*, const float*, float*, unsigned); +#ifdef OPKIT_HAVE_CUDA +int opkit_cuda_saxpy(float, const float*, const float*, float*, unsigned); +#endif +#ifdef OPKIT_HAVE_VULKAN +int opkit_vulkan_saxpy(float, const float*, const float*, float*, unsigned); +// The device the Vulkan backend last ran on. A Vulkan build may find a +// discrete GPU, an integrated one or a CPU rasteriser, and which of those +// answered is not derivable from the numbers -- they are the same numbers. +const char* opkit_vulkan_device_name(void); +#endif + +#ifdef __cplusplus +} +#endif +#endif diff --git a/examples/09-heterogeneous/multi-backend/mcpp.toml b/examples/09-heterogeneous/multi-backend/mcpp.toml new file mode 100644 index 00000000..46ee3633 --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/mcpp.toml @@ -0,0 +1,158 @@ +[package] +name = "opkit-multi-backend" +namespace = "example" +version = "0.1.0" +description = "One operator, several device backends in one artifact, chosen at run time" +accelerators = ["cuda", "vulkan"] + +[language] +standard = "c++23" +modules = true +import_std = true + +# WHAT THIS EXAMPLE SHOWS THAT THE FOUR BESIDE IT DO NOT. +# +# Each of cuda/, hip/, sycl/ and vulkan/ is ONE seam: a device file and a CPU +# file define the same symbol and are never in one link, so exactly one exists +# and the choice is made at BUILD time. That is the right shape for a program. +# +# A library cannot make that choice. It is compiled once and consumed by people +# whose machines differ, so its backends are ADDITIVE -- several land in one +# artifact and the choice moves to RUN time. This is that shape. +# +# `accel` is deliberately absent, so a plain `mcpp build` produces the CPU-only +# variant and needs no payloads at all. The device backends are opt-in: +# +# mcpp build --accel "vulkan1.2" +# mcpp build --accel "cuda12.9+{sm_89}" +# mcpp build --accel "cuda12.9+{sm_89}, vulkan1.2" # both, one artifact +# clang, because the CUDA rule follows the project's toolchain to pick its +# route and the clang route is the one this line supports. It costs the +# CPU-only build an LLVM payload it would not otherwise need, which is the +# honest price of having the device leg work on the driver a developer already +# has. +[toolchain] +default = "llvm@22.1.8" + +# BOTH rules, in one build program. `host-module = true` compiles their module +# interfaces for the build program to import; `[build-dependencies]` keeps the +# package out of the target, which is the case docs/05 section 2.6.1 exists for. +# The rules are declared unconditionally because `build.mcpp` imports them +# unconditionally -- each returns immediately when its own backend is absent. +[build-dependencies.mcpp] +plugins = { version = "0.2.2", features = ["rules-cuda", "rules-spirv"], host-module = true } + +# ── the payloads, gated on the device they are for ────────────────────────── +# +# `cfg(accelerator = ...)` in an `[xlings]` table is what keeps `mcpp build` +# free: a CPU-only build of this project installs neither the CUDA toolkit nor +# the shader compiler, because neither predicate holds. Unconditional pins -- +# the only spelling available before mcpp 2026.9.6.5 -- would have made the +# cheapest build the most expensive one, and that is the build CI runs. +# +# `accelerator` is admitted here and the five resolved layer keys (`c-abi`, +# `compiler`, ...) are not, because it is an INPUT to the build rather than an +# answer from the dependency graph: `--accel` is read before the first package +# is resolved, while a C library is chosen by the resolution a payload would +# have to precede. +# The 12.9 line and the CLANG route, which is the combination this repository +# verifies everywhere: examples/09-heterogeneous/cuda and every mcpp-plugins +# fixture use it, and it runs on any driver from r525 onward. +# +# The nvcc route on the same line does not work: the 12.9 headers redeclare the +# C23 `cospi`/`sinpi`/`rsqrt` for the host without `noexcept` while the C +# library declares them with it, and nvcc's front end refuses the pair. Moving +# to 13.x fixes that and raises the driver floor to r580, which is a machine +# requirement rather than a project decision. The clang route never includes +# that header and imposes no such floor. +[target.'cfg(accelerator = "cuda")'.xlings.workspace] +"xim:cuda-nvcc" = "12.9.86" +"xim:cuda-cudart" = "12.9.79" +# clang's CUDA wrapper includes a cuRAND header for every device unit, and that +# header includes from CCCL. Neither is called by this kernel; on a +# developer machine the host's /usr/include used to supply them silently. +"xim:libcurand" = "10.3.10.19" +"xim:cuda-cccl" = "12.9.27" + +# The host-link stub is Linux-only, and the OS goes in the PREDICATE rather +# than in the value: a value carrying platform keys under an already-predicated +# table would state the condition twice and let the two disagree. +[target.'cfg(all(accelerator = "cuda", linux))'.xlings.workspace] +"xim:libcuda-host-link" = "0.0.1" + +# The driver's userspace library, reached through an index package rather than +# the host: mcpp's private loader does not consult /usr/lib, so a statically +# linked CUDA runtime cannot otherwise dlopen the driver. It is the one CUDA +# component that cannot be an ordinary payload -- the licence forbids +# redistributing it and it is in ABI lockstep with the kernel module. +[target.'cfg(accelerator = "cuda")'.dependencies.compat] +cuda-driver = "2026.09.05" + +[target.'cfg(accelerator = "vulkan")'.xlings.workspace] +# The shader compiler, and a Vulkan driver that is always present because it is +# the CPU. The second is what makes the Vulkan leg runnable on a machine with +# no GPU, which is what every CI runner in this ecosystem is. +"xim:glslang" = "15.1.0" +"xim:mesa-lavapipe" = "26.2.1" + +# The Khronos loader, built by the index rather than taken from the host, and +# the adapter that makes the host's own ICDs reachable from a binary running +# under mcpp's private loader. Neither is a driver: a driver has to match the +# kernel module on the machine it runs on, which is why the software one above +# is a payload and the hardware ones are the host's. +[target.'cfg(accelerator = "vulkan")'.dependencies.compat] +vulkan = "1.4.357.0" +vulkan-runtime = "2026.09.07" + +[build] +# The device sources carry the accel they are for. A CONSTRAINED glob gates +# itself -- it is offered to the build program only when this build's `accel` +# accepts it -- so these need no `cfg` block, and the engine never offers a +# `.cu` or a `.comp` to the C++ compiler. +sources = [ + "src/*.cppm", + "src/main.cpp", + "src/cpu/*.cpp", + { glob = "src/backends/cuda/*.cu", accel = "cuda12.9+{sm_89}" }, + { glob = "src/backends/vulkan/*.comp", accel = "vulkan1.2" }, +] +include_dirs = ["include"] + +# ── the backends, additive ────────────────────────────────────────────── +# +# Each block activates when its backend is named, and several may activate at +# once -- `accelerator` is a SET, not a choice. `defines` is what tells the +# dispatcher which backends exist, so the registry's list and the sources +# actually compiled cannot drift apart. +# What each block carries is the HOST half of a backend plus the define that +# admits it to the dispatcher. The device half is gated by its constrained glob +# above, and the two must agree -- which is why the define lives here, beside +# the host file that implements the entry point the dispatcher will call. +[target.'cfg(accelerator = "cuda")'.build] +defines = ["OPKIT_HAVE_CUDA=1"] +# Linked statically, and only when a device build asks for it. NO ABSOLUTE +# PATHS: the rule package puts the payload's library directory on the link line +# from `mcpp::xpkg_dir`, so this names libraries only. +ldflags = ["-lcudart_static", "-lrt", "-lpthread", "-ldl"] + +[target.'cfg(accelerator = "vulkan")'.build] +sources = ["src/backends/vulkan/*.cpp"] +defines = ["OPKIT_HAVE_VULKAN=1"] + +# ── the dispatcher, and why neither predicate enumerates ───────────────── +# +# `not(accelerator = "none")` means "this build named at least one backend". +# Spelled `not(any(accelerator = "cuda", accelerator = "vulkan"))` it would +# have to be edited every time the ecosystem gains a backend -- and +# `accelerator` is open by design, so a third one is a package rather than an +# engine change. The edit that is forgotten is silent: the CPU-only dispatcher +# and the registry would both compile, or neither would. +[target.'cfg(not(accelerator = "none"))'.build] +sources = ["src/dispatch/registry.cpp"] + +[target.'cfg(accelerator = "none")'.build] +sources = ["src/dispatch/cpu_only.cpp"] + +[targets.opkit-multi-backend] +kind = "bin" +main = "src/main.cpp" diff --git a/examples/09-heterogeneous/multi-backend/src/backends/cuda/saxpy.cu b/examples/09-heterogeneous/multi-backend/src/backends/cuda/saxpy.cu new file mode 100644 index 00000000..3bd93b95 --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/src/backends/cuda/saxpy.cu @@ -0,0 +1,33 @@ +// The CUDA island. Compiled only when the build names `cuda`, by the compiler +// mcpp.rules.cuda resolves -- never by mcpp's own. +#include + +__global__ void opkit_saxpy_kernel(float a, const float* x, const float* y, + float* out, unsigned n) { + unsigned i = blockIdx.x * blockDim.x + threadIdx.x; + if (i < n) out[i] = a * x[i] + y[i]; +} + +extern "C" int opkit_cuda_saxpy(float a, const float* x, const float* y, + float* out, unsigned n) { + // A backend that is compiled in may still find no device. Returning + // non-zero is how it declines, and the dispatcher moves to the next one -- + // the reason this is a chain rather than a build-time choice. + int count = 0; + if (cudaGetDeviceCount(&count) != cudaSuccess || count == 0) return 1; + + float *dx = nullptr, *dy = nullptr, *dout = nullptr; + const size_t bytes = size_t(n) * sizeof(float); + if (cudaMalloc(&dx, bytes) != cudaSuccess) return 1; + if (cudaMalloc(&dy, bytes) != cudaSuccess) { cudaFree(dx); return 1; } + if (cudaMalloc(&dout, bytes) != cudaSuccess) { cudaFree(dx); cudaFree(dy); return 1; } + + cudaMemcpy(dx, x, bytes, cudaMemcpyHostToDevice); + cudaMemcpy(dy, y, bytes, cudaMemcpyHostToDevice); + opkit_saxpy_kernel<<<(n + 255) / 256, 256>>>(a, dx, dy, dout, n); + const bool ok = cudaDeviceSynchronize() == cudaSuccess; + if (ok) cudaMemcpy(out, dout, bytes, cudaMemcpyDeviceToHost); + + cudaFree(dx); cudaFree(dy); cudaFree(dout); + return ok ? 0 : 1; +} diff --git a/examples/09-heterogeneous/multi-backend/src/backends/vulkan/host.cpp b/examples/09-heterogeneous/multi-backend/src/backends/vulkan/host.cpp new file mode 100644 index 00000000..54931a5c --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/src/backends/vulkan/host.cpp @@ -0,0 +1,296 @@ +// The Vulkan backend's host half -- the real one, ported from +// examples/09-heterogeneous/vulkan. +// +// A stub that declined would have been simpler and would have taught the wrong +// thing: an example about run-time dispatch whose second backend never answers +// prints `backend: cpu` for a Vulkan build, which reads as "Vulkan failed on +// this machine" rather than "this file was never written". +// +// It declines by RETURNING NON-ZERO when no usable device is present, which is +// the contract every backend in the chain follows and what makes the +// dispatcher a fallback chain rather than a build-time choice. +#include "opkit/opkit.h" + +#include + +#include +#include +#include +#include + +// Generated by mcpp.rules.spirv into the build directory, which the rule puts on +// the include path. The symbol name is derived from the file name: +// `src/backends/vulkan/saxpy.comp` -> `saxpy_comp_spv`. +#include "saxpy_comp.h" + +namespace { + +struct push_constants { float a; std::uint32_t n; }; + +// A single allocation holds x, y and out back to back, so the shader needs one +// binding and this file one memory mapping. +constexpr std::uint32_t kVectorsInBuffer = 3; + +// Any memory that is host-visible and coherent will do. A discrete GPU offers +// a non-device-local heap with those properties and a software rasteriser has +// nothing else, so this is the one requirement both satisfy; a real workload +// would stage through device-local memory and this one would gain nothing from +// it. +int find_memory_type(VkPhysicalDevice phys, std::uint32_t bits, + VkMemoryPropertyFlags want) { + VkPhysicalDeviceMemoryProperties props{}; + vkGetPhysicalDeviceMemoryProperties(phys, &props); + for (std::uint32_t i = 0; i < props.memoryTypeCount; ++i) + if ((bits & (1u << i)) && (props.memoryTypes[i].propertyFlags & want) == want) + return static_cast(i); + return -1; +} + +// The first device with a compute queue. Not "the fastest": what this program +// demonstrates is that it runs wherever a driver exists, and a machine whose +// only driver is lavapipe is the case the example is written for. +bool pick_device(VkInstance inst, VkPhysicalDevice& out, std::uint32_t& family) { + std::uint32_t n = 0; + vkEnumeratePhysicalDevices(inst, &n, nullptr); + if (n == 0) return false; + std::vector devices(n); + vkEnumeratePhysicalDevices(inst, &n, devices.data()); + for (auto d : devices) { + std::uint32_t qn = 0; + vkGetPhysicalDeviceQueueFamilyProperties(d, &qn, nullptr); + std::vector qs(qn); + vkGetPhysicalDeviceQueueFamilyProperties(d, &qn, qs.data()); + for (std::uint32_t i = 0; i < qn; ++i) + if (qs[i].queueFlags & VK_QUEUE_COMPUTE_BIT) { out = d; family = i; return true; } + } + return false; +} + +// THE NAME OF THE DEVICE, RECORDED BY THE RUN. +// +// Set only after a successful call, and read through the seam. The CPU +// fallback produces the same four numbers as this island does, so without a +// name printed by the program nothing distinguishes a device run from a silent +// fallback -- which is the one confusion an example about heterogeneous +// compute must not leave in place. +char g_ran_on[256] = ""; + +} // namespace + + +extern "C" const char* opkit_vulkan_device_name(void) { return g_ran_on; } + +extern "C" int opkit_vulkan_saxpy(float a, const float* x, const float* y, + float* out, unsigned n) { + if (n == 0) return 0; + + VkApplicationInfo app{}; + app.sType = VK_STRUCTURE_TYPE_APPLICATION_INFO; + app.apiVersion = VK_API_VERSION_1_1; + VkInstanceCreateInfo ici{}; + ici.sType = VK_STRUCTURE_TYPE_INSTANCE_CREATE_INFO; + ici.pApplicationInfo = &app; + + VkInstance inst{}; + if (vkCreateInstance(&ici, nullptr, &inst) != VK_SUCCESS) return 1; + + VkPhysicalDevice phys{}; + std::uint32_t family = 0; + if (!pick_device(inst, phys, family)) { vkDestroyInstance(inst, nullptr); return 1; } + + const float priority = 1.0f; + VkDeviceQueueCreateInfo qci{}; + qci.sType = VK_STRUCTURE_TYPE_DEVICE_QUEUE_CREATE_INFO; + qci.queueFamilyIndex = family; + qci.queueCount = 1; + qci.pQueuePriorities = &priority; + VkDeviceCreateInfo dci{}; + dci.sType = VK_STRUCTURE_TYPE_DEVICE_CREATE_INFO; + dci.queueCreateInfoCount = 1; + dci.pQueueCreateInfos = &qci; + + VkDevice dev{}; + if (vkCreateDevice(phys, &dci, nullptr, &dev) != VK_SUCCESS) { + vkDestroyInstance(inst, nullptr); + return 1; + } + + int rc = 1; + VkQueue queue{}; + vkGetDeviceQueue(dev, family, 0, &queue); + + const VkDeviceSize bytes = VkDeviceSize(n) * kVectorsInBuffer * sizeof(float); + VkBufferCreateInfo bci{}; + bci.sType = VK_STRUCTURE_TYPE_BUFFER_CREATE_INFO; + bci.size = bytes; + bci.usage = VK_BUFFER_USAGE_STORAGE_BUFFER_BIT; + bci.sharingMode = VK_SHARING_MODE_EXCLUSIVE; + + VkBuffer buffer{}; + VkDeviceMemory memory{}; + VkShaderModule shader{}; + VkDescriptorSetLayout setLayout{}; + VkPipelineLayout pipeLayout{}; + VkPipeline pipeline{}; + VkDescriptorPool pool{}; + VkCommandPool cmdPool{}; + VkFence fence{}; + + if (vkCreateBuffer(dev, &bci, nullptr, &buffer) != VK_SUCCESS) goto done; + + { + VkMemoryRequirements req{}; + vkGetBufferMemoryRequirements(dev, buffer, &req); + const int type = find_memory_type(phys, req.memoryTypeBits, + VK_MEMORY_PROPERTY_HOST_VISIBLE_BIT + | VK_MEMORY_PROPERTY_HOST_COHERENT_BIT); + if (type < 0) goto done; + VkMemoryAllocateInfo mai{}; + mai.sType = VK_STRUCTURE_TYPE_MEMORY_ALLOCATE_INFO; + mai.allocationSize = req.size; + mai.memoryTypeIndex = static_cast(type); + if (vkAllocateMemory(dev, &mai, nullptr, &memory) != VK_SUCCESS) goto done; + if (vkBindBufferMemory(dev, buffer, memory, 0) != VK_SUCCESS) goto done; + + void* mapped = nullptr; + if (vkMapMemory(dev, memory, 0, bytes, 0, &mapped) != VK_SUCCESS) goto done; + auto* v = static_cast(mapped); + std::memcpy(v, x, n * sizeof(float)); + std::memcpy(v + n, y, n * sizeof(float)); + std::memset(v + 2 * n, 0, n * sizeof(float)); + vkUnmapMemory(dev, memory); + } + + { + VkShaderModuleCreateInfo smci{}; + smci.sType = VK_STRUCTURE_TYPE_SHADER_MODULE_CREATE_INFO; + smci.codeSize = sizeof saxpy_comp_spv; + smci.pCode = saxpy_comp_spv; + if (vkCreateShaderModule(dev, &smci, nullptr, &shader) != VK_SUCCESS) goto done; + + VkDescriptorSetLayoutBinding binding{}; + binding.binding = 0; + binding.descriptorType = VK_DESCRIPTOR_TYPE_STORAGE_BUFFER; + binding.descriptorCount = 1; + binding.stageFlags = VK_SHADER_STAGE_COMPUTE_BIT; + VkDescriptorSetLayoutCreateInfo dslci{}; + dslci.sType = VK_STRUCTURE_TYPE_DESCRIPTOR_SET_LAYOUT_CREATE_INFO; + dslci.bindingCount = 1; + dslci.pBindings = &binding; + if (vkCreateDescriptorSetLayout(dev, &dslci, nullptr, &setLayout) != VK_SUCCESS) + goto done; + + VkPushConstantRange range{}; + range.stageFlags = VK_SHADER_STAGE_COMPUTE_BIT; + range.size = sizeof(push_constants); + VkPipelineLayoutCreateInfo plci{}; + plci.sType = VK_STRUCTURE_TYPE_PIPELINE_LAYOUT_CREATE_INFO; + plci.setLayoutCount = 1; + plci.pSetLayouts = &setLayout; + plci.pushConstantRangeCount = 1; + plci.pPushConstantRanges = ⦥ + if (vkCreatePipelineLayout(dev, &plci, nullptr, &pipeLayout) != VK_SUCCESS) goto done; + + VkComputePipelineCreateInfo cpci{}; + cpci.sType = VK_STRUCTURE_TYPE_COMPUTE_PIPELINE_CREATE_INFO; + cpci.stage.sType = VK_STRUCTURE_TYPE_PIPELINE_SHADER_STAGE_CREATE_INFO; + cpci.stage.stage = VK_SHADER_STAGE_COMPUTE_BIT; + cpci.stage.module = shader; + cpci.stage.pName = "main"; + cpci.layout = pipeLayout; + if (vkCreateComputePipelines(dev, VK_NULL_HANDLE, 1, &cpci, nullptr, &pipeline) + != VK_SUCCESS) + goto done; + } + + { + VkDescriptorPoolSize size{VK_DESCRIPTOR_TYPE_STORAGE_BUFFER, 1}; + VkDescriptorPoolCreateInfo dpci{}; + dpci.sType = VK_STRUCTURE_TYPE_DESCRIPTOR_POOL_CREATE_INFO; + dpci.maxSets = 1; + dpci.poolSizeCount = 1; + dpci.pPoolSizes = &size; + if (vkCreateDescriptorPool(dev, &dpci, nullptr, &pool) != VK_SUCCESS) goto done; + + VkDescriptorSetAllocateInfo dsai{}; + dsai.sType = VK_STRUCTURE_TYPE_DESCRIPTOR_SET_ALLOCATE_INFO; + dsai.descriptorPool = pool; + dsai.descriptorSetCount = 1; + dsai.pSetLayouts = &setLayout; + VkDescriptorSet set{}; + if (vkAllocateDescriptorSets(dev, &dsai, &set) != VK_SUCCESS) goto done; + + VkDescriptorBufferInfo info{buffer, 0, bytes}; + VkWriteDescriptorSet write{}; + write.sType = VK_STRUCTURE_TYPE_WRITE_DESCRIPTOR_SET; + write.dstSet = set; + write.descriptorCount = 1; + write.descriptorType = VK_DESCRIPTOR_TYPE_STORAGE_BUFFER; + write.pBufferInfo = &info; + vkUpdateDescriptorSets(dev, 1, &write, 0, nullptr); + + VkCommandPoolCreateInfo cpi{}; + cpi.sType = VK_STRUCTURE_TYPE_COMMAND_POOL_CREATE_INFO; + cpi.queueFamilyIndex = family; + if (vkCreateCommandPool(dev, &cpi, nullptr, &cmdPool) != VK_SUCCESS) goto done; + + VkCommandBufferAllocateInfo cbai{}; + cbai.sType = VK_STRUCTURE_TYPE_COMMAND_BUFFER_ALLOCATE_INFO; + cbai.commandPool = cmdPool; + cbai.level = VK_COMMAND_BUFFER_LEVEL_PRIMARY; + cbai.commandBufferCount = 1; + VkCommandBuffer cmd{}; + if (vkAllocateCommandBuffers(dev, &cbai, &cmd) != VK_SUCCESS) goto done; + + VkCommandBufferBeginInfo begin{}; + begin.sType = VK_STRUCTURE_TYPE_COMMAND_BUFFER_BEGIN_INFO; + begin.flags = VK_COMMAND_BUFFER_USAGE_ONE_TIME_SUBMIT_BIT; + vkBeginCommandBuffer(cmd, &begin); + vkCmdBindPipeline(cmd, VK_PIPELINE_BIND_POINT_COMPUTE, pipeline); + vkCmdBindDescriptorSets(cmd, VK_PIPELINE_BIND_POINT_COMPUTE, pipeLayout, + 0, 1, &set, 0, nullptr); + const push_constants pc{a, n}; + vkCmdPushConstants(cmd, pipeLayout, VK_SHADER_STAGE_COMPUTE_BIT, + 0, sizeof pc, &pc); + // The workgroup size is in the shader (`local_size_x = 64`); the count + // here has to agree with it, which is why the shader also bounds-checks. + vkCmdDispatch(cmd, (n + 63) / 64, 1, 1); + vkEndCommandBuffer(cmd); + + VkFenceCreateInfo fci{}; + fci.sType = VK_STRUCTURE_TYPE_FENCE_CREATE_INFO; + if (vkCreateFence(dev, &fci, nullptr, &fence) != VK_SUCCESS) goto done; + + VkSubmitInfo submit{}; + submit.sType = VK_STRUCTURE_TYPE_SUBMIT_INFO; + submit.commandBufferCount = 1; + submit.pCommandBuffers = &cmd; + if (vkQueueSubmit(queue, 1, &submit, fence) != VK_SUCCESS) goto done; + if (vkWaitForFences(dev, 1, &fence, VK_TRUE, ~0ull) != VK_SUCCESS) goto done; + + void* mapped = nullptr; + if (vkMapMemory(dev, memory, 0, bytes, 0, &mapped) != VK_SUCCESS) goto done; + std::memcpy(out, static_cast(mapped) + 2 * n, n * sizeof(float)); + vkUnmapMemory(dev, memory); + rc = 0; + } + +done: + if (fence) vkDestroyFence(dev, fence, nullptr); + if (cmdPool) vkDestroyCommandPool(dev, cmdPool, nullptr); + if (pool) vkDestroyDescriptorPool(dev, pool, nullptr); + if (pipeline) vkDestroyPipeline(dev, pipeline, nullptr); + if (pipeLayout) vkDestroyPipelineLayout(dev, pipeLayout, nullptr); + if (setLayout) vkDestroyDescriptorSetLayout(dev, setLayout, nullptr); + if (shader) vkDestroyShaderModule(dev, shader, nullptr); + if (memory) vkFreeMemory(dev, memory, nullptr); + if (buffer) vkDestroyBuffer(dev, buffer, nullptr); + if (rc == 0) { + VkPhysicalDeviceProperties properties{}; + vkGetPhysicalDeviceProperties(phys, &properties); + std::snprintf(g_ran_on, sizeof g_ran_on, "%s", properties.deviceName); + } + vkDestroyDevice(dev, nullptr); + vkDestroyInstance(inst, nullptr); + return rc; +} diff --git a/examples/09-heterogeneous/multi-backend/src/backends/vulkan/saxpy.comp b/examples/09-heterogeneous/multi-backend/src/backends/vulkan/saxpy.comp new file mode 100644 index 00000000..742c2715 --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/src/backends/vulkan/saxpy.comp @@ -0,0 +1,15 @@ +#version 450 + +// The device side of the same computation the `cuda` example runs on +// CUDA: out = a*x + y. One storage buffer holds all three vectors so the host +// side needs one allocation and one descriptor. +layout(local_size_x = 64) in; + +layout(std430, binding = 0) buffer Data { float v[]; }; +layout(push_constant) uniform Push { float a; uint n; } push; + +void main() { + const uint i = gl_GlobalInvocationID.x; + if (i >= push.n) return; + v[2u * push.n + i] = push.a * v[i] + v[push.n + i]; +} diff --git a/examples/09-heterogeneous/multi-backend/src/cpu/saxpy.cpp b/examples/09-heterogeneous/multi-backend/src/cpu/saxpy.cpp new file mode 100644 index 00000000..7dd2dade --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/src/cpu/saxpy.cpp @@ -0,0 +1,14 @@ +#include "opkit/opkit.h" + +// The reference implementation, ALWAYS built. +// +// This is the difference between an operator library and the single-seam +// examples beside it: there a CPU file and a device file define the SAME +// symbol and are never in one link, so exactly one exists. Here the reference +// is a backend like any other and every build has it, which is what makes a +// runtime fallback possible on a machine whose device turns out to be absent. +extern "C" int opkit_cpu_saxpy(float a, const float* x, const float* y, + float* out, unsigned n) { + for (unsigned i = 0; i < n; ++i) out[i] = a * x[i] + y[i]; + return 0; +} diff --git a/examples/09-heterogeneous/multi-backend/src/dispatch/cpu_only.cpp b/examples/09-heterogeneous/multi-backend/src/dispatch/cpu_only.cpp new file mode 100644 index 00000000..446dce57 --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/src/dispatch/cpu_only.cpp @@ -0,0 +1,23 @@ +#include "opkit/opkit.h" + +// The CPU-ONLY build, `cfg(accelerator = "none")`. +// +// `mcpp build` with no accelerator named produces this: no dispatch chain, no +// registry, a direct call. That is a real difference rather than a cosmetic +// one -- a consumer who wants the operator and no device machinery gets an +// artifact that contains none of it. +// +// The predicate is the reason this file can exist at all. `accelerator` is an +// OPEN vocabulary, so "this build names no backend" cannot be said by listing +// the backends it is not; `none` says it directly and keeps saying it after the +// ecosystem gains a fifth. +static const char* g_backend = ""; + +extern "C" const char* opkit_backend(void) { return g_backend; } + +extern "C" int opkit_saxpy(float a, const float* x, const float* y, + float* out, unsigned n) { + int rc = opkit_cpu_saxpy(a, x, y, out, n); + if (rc == 0) g_backend = "cpu (only backend in this build)"; + return rc; +} diff --git a/examples/09-heterogeneous/multi-backend/src/dispatch/registry.cpp b/examples/09-heterogeneous/multi-backend/src/dispatch/registry.cpp new file mode 100644 index 00000000..2e41282f --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/src/dispatch/registry.cpp @@ -0,0 +1,38 @@ +#include "opkit/opkit.h" + +// THE DISPATCHER, built only when this build names at least one device +// backend -- `cfg(not(accelerator = "none"))`. +// +// The predicate needs no enumeration of the backends, which is the whole point: +// this file must be built for cuda, for vulkan, for both, and for a backend +// that does not exist yet. Written as `not(any(accelerator = "cuda", +// accelerator = "vulkan"))` it would have to be edited every time the ecosystem +// gains a backend, and the edit that is forgotten is silent. +#include + +static const char* g_backend = ""; +#ifdef OPKIT_HAVE_VULKAN +static char g_vulkan[288]; +#endif + +extern "C" const char* opkit_backend(void) { return g_backend; } + +extern "C" int opkit_saxpy(float a, const float* x, const float* y, + float* out, unsigned n) { + // Device backends first, in declaration order; the reference last. A + // device that is compiled in may still be absent at run time, which is why + // this is a fallback chain rather than a single choice made at build time. +#ifdef OPKIT_HAVE_CUDA + if (opkit_cuda_saxpy(a, x, y, out, n) == 0) { g_backend = "cuda"; return 0; } +#endif +#ifdef OPKIT_HAVE_VULKAN + if (opkit_vulkan_saxpy(a, x, y, out, n) == 0) { + std::snprintf(g_vulkan, sizeof g_vulkan, "vulkan (%s)", + opkit_vulkan_device_name()); + g_backend = g_vulkan; + return 0; + } +#endif + if (opkit_cpu_saxpy(a, x, y, out, n) == 0) { g_backend = "cpu (fallback)"; return 0; } + return 1; +} diff --git a/examples/09-heterogeneous/multi-backend/src/main.cpp b/examples/09-heterogeneous/multi-backend/src/main.cpp new file mode 100644 index 00000000..2e2a1e31 --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/src/main.cpp @@ -0,0 +1,18 @@ +import std; +import opkit; + +int main() { + const std::vector x{1, 2, 3, 4}, y{10, 20, 30, 40}; + + auto out = opkit::saxpy(2.0f, x, y); + if (!out) { std::println("no backend served the call"); return 1; } + + // Printed after the call, never before: the name records a run that + // happened rather than predicting one. + std::println("backend: {}", opkit::backend()); + for (auto v : *out) std::print("{} ", v); + std::println(""); + + const std::vector want{12, 24, 36, 48}; + return *out == want ? 0 : 1; +} diff --git a/examples/09-heterogeneous/multi-backend/src/opkit.cppm b/examples/09-heterogeneous/multi-backend/src/opkit.cppm new file mode 100644 index 00000000..32d71e37 --- /dev/null +++ b/examples/09-heterogeneous/multi-backend/src/opkit.cppm @@ -0,0 +1,32 @@ +// The seam, as a module. +// +// The backends underneath are C: an island is compiled by a compiler mcpp did +// not choose, so the boundary carries raw pointers and a count and nothing that +// depends on a C++ ABI. This module is where that becomes C++ again, and it is +// the only file a consumer imports -- which is what lets a backend be added, +// removed or reordered without any importer changing. +module; +#include "opkit/opkit.h" +export module opkit; +import std; + +export namespace opkit { + +// The operator. Which backend serves it is decided inside, at run time, among +// those this build compiled in. +std::optional> +saxpy(float a, std::span x, std::span y) { + if (x.size() != y.size()) return std::nullopt; + std::vector out(x.size()); + if (opkit_saxpy(a, x.data(), y.data(), out.data(), + static_cast(x.size())) != 0) + return std::nullopt; + return out; +} + +// WHICH backend answered. An operator library that computes and does not say +// where cannot be checked: every backend returns the same numbers, so the +// numbers alone do not separate a device run from the reference one. +std::string_view backend() { return opkit_backend(); } + +} // namespace opkit diff --git a/examples/09-heterogeneous/sycl/app/mcpp.toml b/examples/09-heterogeneous/sycl/app/mcpp.toml index a85ed154..e9af3ad9 100644 --- a/examples/09-heterogeneous/sycl/app/mcpp.toml +++ b/examples/09-heterogeneous/sycl/app/mcpp.toml @@ -16,7 +16,7 @@ import_std = true [toolchain] default = "llvm@22.1.8" -[dependencies.mcpp] +[build-dependencies.mcpp] plugins = { version = "0.2.1", features = ["rules-sycl"], host-module = true } # The SYCL runtime, on the artifact's runtime search path. mcpp's private diff --git a/examples/09-heterogeneous/vulkan/app/mcpp.toml b/examples/09-heterogeneous/vulkan/app/mcpp.toml index f601281f..fcf45c2a 100644 --- a/examples/09-heterogeneous/vulkan/app/mcpp.toml +++ b/examples/09-heterogeneous/vulkan/app/mcpp.toml @@ -12,7 +12,7 @@ import_std = true # The rule that compiles the shaders lives in the official plugin collection, # selected by its feature; `build.mcpp` imports it as `mcpp.rules.spirv`. -[dependencies.mcpp] +[build-dependencies.mcpp] plugins = { version = "0.2.1", features = ["rules-spirv"], host-module = true } # The Khronos loader, built by the index rather than taken from the host, and diff --git a/mcpp.toml b/mcpp.toml index f70aa05b..d2d091b8 100644 --- a/mcpp.toml +++ b/mcpp.toml @@ -1,6 +1,6 @@ [package] name = "mcpp" -version = "2026.9.6.4" +version = "2026.9.6.5" description = "Modern C++ build & package management tool" license = "Apache-2.0" authors = ["mcpp-community"] diff --git a/modules/buildmcpp/src/directives.cppm b/modules/buildmcpp/src/directives.cppm index b6cc87dd..1ad395c3 100644 --- a/modules/buildmcpp/src/directives.cppm +++ b/modules/buildmcpp/src/directives.cppm @@ -217,7 +217,7 @@ struct Def { int sinceProtocol; }; -inline constexpr std::array kTable{{ +inline constexpr std::array kTable{{ // wire tag slot scope transform must missingPrefix missingSuffix since {"cxxflag", "cxxflag", Slot::CxxFlags, Scope::PackagePrivate, Transform::Verbatim, false, "", "", 1}, {"cflag", "cflag", Slot::CFlags, Scope::PackagePrivate, Transform::Verbatim, false, "", "", 1}, @@ -267,6 +267,35 @@ inline constexpr std::array kTable{{ {"runner-longlived", "runner-longlived", Slot::RunnerLongLived, Scope::RunGlobal, Transform::Verbatim, false, "", "", 6}, {"run-exclusive", "run-exclusive", Slot::RunExclusive, Scope::RunGlobal, Transform::Verbatim, false, "", "", 6}, {"link-script", "ldflag", Slot::LdFlags, Scope::LinkGlobal, Transform::LinkerScript, false, "", "", 3}, + // THE OUTLET THE LINK FAMILY WAS MISSING (v8). + // + // `link-lib`, `link-search` and `link-script` each name one KIND of thing. + // A flag the program COMPUTED belongs to none of them: a generated version + // script (`-Wl,--version-script=`), `-Wl,--wrap=malloc` for a runtime that + // takes over a C-library symbol, `-Wl,--exclude-libs,ALL` so a statically + // absorbed third party does not become part of this package's ABI. + // + // Scope::LinkGlobal, AND THAT IS THE CORRECTION OF AN EARLIER DESIGN. + // The design doc first ruled it PackagePrivate by analogy with + // `include-dir`. The analogy is false. `include-dir` is private because a + // compile interface has a declarative public counterpart + // (`[build] include_dirs`) and a build-time program must not widen it + // behind the manifest's back. Link flags have no such split: the + // declarative `[build] ldflags` ALREADY propagates to consumers, and + // `linkUsage.ldflags` is a copy of `buildConfig.ldflags`. A private link + // flag is not a policy this engine can express today, and making the + // computed form behave differently from its declarative twin would be the + // inconsistency, not the safeguard. + // + // The consequence is stated rather than hidden: a dependency emitting + // `-Wl,--version-script=` puts it on the consumer's link too. That hazard + // is not new -- a dependency writing the same flag in `[build] ldflags` + // has always done this -- so this row widens who can compute the value, + // not what the value can reach. + // + // Verbatim: the engine does not parse linker flags. `-Wl,` forms, `-z` + // pairs and vendor spellings are the linker's vocabulary, not this table's. + {"link-flag", "ldflag", Slot::LdFlags, Scope::LinkGlobal, Transform::Verbatim, false, "", "", 8}, {"include-dir", "include-dir", Slot::IncludeDirs, Scope::PackagePrivate, Transform::AbsPath, false, "", "", 1}, {"include-dir-after", "include-dir-after", Slot::IncludeDirsAfter, Scope::PackagePrivate, Transform::AbsPath, false, "", "", 1}, {"rerun-if-changed", "", Slot::RerunFiles, Scope::RerunKey, Transform::Verbatim, false, "", "", 1}, diff --git a/modules/buildmcpp/src/program_protocol.cppm b/modules/buildmcpp/src/program_protocol.cppm index 88b92656..fab31c4d 100644 --- a/modules/buildmcpp/src/program_protocol.cppm +++ b/modules/buildmcpp/src/program_protocol.cppm @@ -60,7 +60,13 @@ export namespace mcpp::build::program_protocol { // fact about the machine and the floor it needs of it; the engine compares // them before compiling. Same cost as v5's: a package calling `mcpp::fact()` // fails on an older engine at the build.mcpp COMPILE, not through a refusal. -inline constexpr int kProtocolVersion = 7; +// v8: adds `link-flag` -- the generic linker-flag outlet. `link-lib`, +// `link-search` and `link-script` cover a library, a search path and a layout; +// a flag a program COMPUTED (a generated version script, `--wrap`, +// `--exclude-libs`) had no way out. Same cost as v5's: a package calling +// `mcpp::link_flag()` fails on an older engine at the build.mcpp COMPILE, +// because that engine's bundled module has no such function. +inline constexpr int kProtocolVersion = 8; // ── Cache-format epoch ───────────────────────────────────────────────────── // diff --git a/modules/manifest/src/mangle.cppm b/modules/manifest/src/mangle.cppm index c1df0e76..496dd51e 100644 --- a/modules/manifest/src/mangle.cppm +++ b/modules/manifest/src/mangle.cppm @@ -39,6 +39,19 @@ std::string mangle_name(std::string_view base, std::string_view version); // bare partition declarations are ignored. std::vector declared_module_roots(std::string_view source); +// The module names a source IMPORTS, in order of appearance, deduplicated. +// `import N;` and `export import N;` yield `N`; `import N:P;` yields `N`; +// header units (`import ;`, `import "x.h";`) and bare partition +// imports (`import :P;`) yield nothing, because none of them names a module +// another package could provide. +// +// Same line-based matcher as `declared_module_roots`, and the same limits: +// the keyword must open a logical line, so a declaration quoted inside a +// comment whose line starts with `import` is read as one. That is acceptable +// for its one caller -- a diagnostic that suggests a package, and suggests +// nothing when the name matches none. +std::vector imported_module_names(std::string_view source); + // Rewrite a single .cppm file's module / import declarations: // * `(export )?module N;` → `(export )?module rename[N];` // * `(export )?module N:P;` → `(export )?module rename[N]:P;` @@ -141,6 +154,34 @@ std::vector declared_module_roots(std::string_view source) { return roots; } +std::vector imported_module_names(std::string_view source) { + std::vector names; + std::size_t lineStart = 0; + while (lineStart < source.size()) { + auto eol = source.find('\n', lineStart); + if (eol == std::string_view::npos) eol = source.size(); + auto line = source.substr(lineStart, eol - lineStart); + + auto cur = skip_ws(line, 0); + if (auto p = consume_keyword(line, cur, "export"); + p != std::string::npos) { + cur = skip_ws(line, p); + } + if (auto afterImport = consume_keyword(line, cur, "import"); + afterImport != std::string::npos) { + cur = skip_ws(line, afterImport); + auto [nameEnd, name] = read_name(line, cur); + if (nameEnd != std::string::npos + && std::ranges::find(names, name) == names.end()) { + names.emplace_back(name); + } + } + if (eol == source.size()) break; + lineStart = eol + 1; + } + return names; +} + std::string rewrite_module_decls( std::string_view source, const std::map& rename) diff --git a/modules/manifest/src/toml.cppm b/modules/manifest/src/toml.cppm index e97f318c..a292d06a 100644 --- a/modules/manifest/src/toml.cppm +++ b/modules/manifest/src/toml.cppm @@ -1042,6 +1042,65 @@ std::expected parse_string(std::string_view content, if (auto msg = validate_target_soname(t, std::format("targets.{}.", tname))) { return std::unexpected(error(origin, *msg)); } + // `exports` -- a file of symbol patterns, or the patterns inline. + // + // BOTH FORMS, because the two are used at different scales and the ABI + // contract of a real library is reviewed as a unit. A file is what a + // library with a stable ABI wants (it is the contract, and it belongs + // in review beside the headers); the inline array is for the two or + // three entry points a plugin publishes, where a separate file would + // be ceremony. + if (auto eit = tt.find("exports"); eit != tt.end()) { + if (eit->second.is_string()) { + t.exportsFile = eit->second.as_string(); + // READ IT HERE, so every later stage sees one representation. + // The alternative -- carrying the path and reading it at plan + // time -- would need the package root at three call sites (the + // root, a path dependency, a store dependency) and is the + // "same decision derived in N places" shape this repository + // has paid for. `origin` is the manifest's own path, so the + // root is its parent, and that is true for all three. + auto file = origin.parent_path() / t.exportsFile; + std::ifstream in(file); + if (!in) + return std::unexpected(error(origin, std::format( + "targets.{}.exports names '{}', which does not exist " + "(looked at '{}')", tname, t.exportsFile, + file.generic_string()))); + for (std::string line; std::getline(in, line); ) { + if (auto h = line.find('#'); h != std::string::npos) + line.erase(h); + auto b = line.find_first_not_of(" \t\r"); + if (b == std::string::npos) continue; + auto e = line.find_last_not_of(" \t\r"); + t.exportPatterns.push_back(line.substr(b, e - b + 1)); + } + if (t.exportPatterns.empty()) + return std::unexpected(error(origin, std::format( + "targets.{}.exports: '{}' names no symbol. An empty " + "export list is not how a library says 'publish " + "everything' -- omitting the key is.", + tname, t.exportsFile))); + } else if (eit->second.is_array()) { + for (auto& v : eit->second.as_array()) { + if (!v.is_string()) + return std::unexpected(error(origin, std::format( + "targets.{}.exports: every entry must be a symbol " + "pattern (a string)", tname))); + t.exportPatterns.push_back(v.as_string()); + } + if (t.exportPatterns.empty()) + return std::unexpected(error(origin, std::format( + "targets.{}.exports is an empty list. A library that " + "publishes nothing is not what an empty list means " + "here -- omit the key to publish everything, which is " + "the default, or name the symbols.", tname))); + } else { + return std::unexpected(error(origin, std::format( + "targets.{}.exports must be a path to a file of symbol " + "patterns, or an inline list of them", tname))); + } + } // Per-target flags (entry-scoped) + required-features gate. auto read_list = [&](const char* key, std::vector& out) { @@ -1069,7 +1128,7 @@ std::expected parse_string(std::string_view content, // must reach SHARED code is intentionally not a target key; point users // at the right axis (workspace / features / profile). static constexpr std::string_view kKnownTargetKeys[] = { - "kind", "main", "soname", + "kind", "main", "soname", "exports", "cflags", "cxxflags", "defines", "required_features", }; for (auto& [key, _] : tt) { diff --git a/modules/manifest/src/types.cppm b/modules/manifest/src/types.cppm index 2dbed459..3ceb9e96 100644 --- a/modules/manifest/src/types.cppm +++ b/modules/manifest/src/types.cppm @@ -117,6 +117,22 @@ struct Target { enum Kind { Library, Binary, SharedLibrary, TestBinary } kind; std::string main; // for binary / test std::string soname; // ABI name for shared libraries, e.g. libfoo.so.1 + // WHICH SYMBOLS THIS ARTIFACT PUBLISHES. Empty = every symbol, which is + // what both platforms do today (ELF default visibility; PE gets an + // auto-generated .def listing everything, mcpp.build.coff_exports). + // + // ONE NEUTRAL STATEMENT, THREE RENDERINGS -- the shape `[runtime]` already + // established, and the reason this is a manifest key rather than three + // platform-specific flag lists. ELF gets a version script, Mach-O an + // `-exported_symbols_list`, PE a `.def` that REPLACES the all-exports one. + // + // The entries are symbol patterns, one per line in the named file or one + // per element inline. `*` is the only wildcard, because it is the only one + // all three renderings share. + std::vector exportPatterns; + // Where they came from, for diagnostics and for the re-read that a changed + // file must trigger. Empty when `exports` was written inline. + std::string exportsFile; // Per-target compile flags. SCOPE: applied ONLY to this target's exclusive // entry source (its `main`) — never to shared module/impl objects, which are // compiled once and linked into every target (the build's compile-once model; diff --git a/modules/source-kind/src/source_kind.cppm b/modules/source-kind/src/source_kind.cppm index db59dc75..0cbed781 100644 --- a/modules/source-kind/src/source_kind.cppm +++ b/modules/source-kind/src/source_kind.cppm @@ -264,8 +264,15 @@ constexpr std::string_view kHeaderExtensions[] = { ".h", ".hpp", ".hh", ".hxx" } // extension, so a rule package refuses a stage-less name — which is the right // place for that message, and the reason this table does not need to know // which of these extensions name a stage. +// +// `.asc` is Ascend C, compiled by `ccec` (BiSheng) from the CANN toolkit. It is +// here for the reason `.sycl` is: the language is ordinary C++ and the +// criterion is the compiler. CANN's own operator libraries already split +// `op_kernel/` from `op_host/`, and CMake registers ASC as a LANGUAGE of its +// own -- so the island is the shape Ascend already has, not one mcpp imposes. +// `.cce` is the older spelling of the same thing and is accepted beside it. constexpr std::string_view kDeviceExtensions[] = { - ".cu", ".hip", ".sycl", + ".cu", ".hip", ".sycl", ".asc", ".cce", ".comp", ".vert", ".frag", ".geom", ".tesc", ".tese", ".mesh", ".task", ".rgen", ".rint", ".rahit", ".rchit", ".rmiss", ".rcall", ".glsl", ".hlsl", ".cl", ".metal", diff --git a/modules/versioning/src/version.cppm b/modules/versioning/src/version.cppm index 21c09bdb..fc835a05 100644 --- a/modules/versioning/src/version.cppm +++ b/modules/versioning/src/version.cppm @@ -31,6 +31,6 @@ import std; export namespace mcpp { -inline constexpr std::string_view MCPP_VERSION = "2026.9.6.4"; +inline constexpr std::string_view MCPP_VERSION = "2026.9.6.5"; } // namespace mcpp diff --git a/src/build/build_program.cppm b/src/build/build_program.cppm index 14dc4c25..f0706959 100644 --- a/src/build/build_program.cppm +++ b/src/build/build_program.cppm @@ -15,6 +15,7 @@ export module mcpp.build.build_program; import std; import mcpp.manifest; import mcpp.platform; +import mcpp.pm.mangle; // imported_module_names -- what build.mcpp asks for import mcpp.platform.process; import mcpp.toolchain.cppfly; // std_flag (dialect- and c++fly-aware -std= spelling) import mcpp.toolchain.dialect; // CommandDialect — gnu vs cl.exe spellings @@ -812,6 +813,88 @@ std::expected run_build_program( hm.logical, hm.logical)); } + // …and the complementary case: an import NOTHING provides. + // + // Left to the compiler this is a raw `failed to read compiled module` with + // a note that imports must be built before being imported -- true, and it + // names neither the package that would provide the module nor the key that + // would make it importable. Measured on a project that declared its rule + // package correctly except for `host-module`: + // + // mcpp.rules.spirv: error: failed to read compiled module: No such + // file or directory + // mcpp.rules.spirv: note: imports must be built before being imported + // + // The set of names that CAN compile here is closed -- `std`, `std.compat`, + // the bundled `mcpp` module, and the importable host modules -- so a name + // outside it cannot become valid later and is refused rather than warned + // about. Header units (`import ;`, `import "x.h";`) name no module + // and `imported_module_names` does not report them. + // + // ONLY build.mcpp's own imports. A rule interface compiled alongside it may + // import a module that is present as a prerequisite and not importable + // here, which is the case the check above states. + { + std::set available{"std", "std.compat", "mcpp"}; + for (auto const& hm : env.hostModules) + if (hm.importable) available.insert(hm.logical); + for (auto const& want : mcpp::pm::imported_module_names(srcText)) { + if (available.contains(want)) continue; + // A dependency that could plausibly be meant: one declared without + // `host-module = true`. Named as a hint, not as a claim -- the + // engine has not resolved this package's lib root here and does not + // know which module it would declare. + // + // Ranked by the one relation available without resolving anything: + // a rule package's modules are expected to open with its namespace, + // which is what `reserved_prefix_warning` is about. A dependency + // whose key shares the imported name's first segment is listed + // alone; when none does, all of them are listed rather than none, + // because the convention is a convention and a package may declare + // a module outside its own prefix. + const auto dot = want.find('.'); + const auto head = dot == std::string::npos ? want : want.substr(0, dot); + std::string near, all; + auto scan = [&](const std::map& deps, + std::string_view section) { + for (auto const& [k, s] : deps) { + if (s.hostModule) continue; + const auto one = std::format("{} (in [{}])", k, section); + auto& dst = k.starts_with(head + ".") || k == head ? near : all; + if (!dst.empty()) dst += ", "; + dst += one; + } + }; + scan(m.buildDependencies, "build-dependencies"); + scan(m.dependencies, "dependencies"); + const std::string candidates = near.empty() ? all : near; + std::string importable; + for (auto const& hm : env.hostModules) + if (hm.importable) { + if (!importable.empty()) importable += ", "; + importable += hm.logical; + } + return std::unexpected(std::format( + "build.mcpp imports '{}', and no dependency provides it as a " + "host module.\n" + " A package's module is compiled for the build program " + "only when the dependency\n" + " edge asks for it, which is a separate question from " + "whether the package reaches\n" + " the target:\n" + " [build-dependencies.]\n" + " = {{ version = \"...\", host-module = true }}\n" + " importable here: {}\n" + "{}", + want, + importable.empty() ? "(none)" : importable, + candidates.empty() + ? " This package declares no dependency that could provide it." + : std::format(" declared without `host-module = true`: {}", + candidates))); + } + } + // Fast path: declared inputs + contract unchanged → reapply cached // directives, no run. CacheRecord cache = read_cache(bdir); diff --git a/src/build/hostprogram.cppm b/src/build/hostprogram.cppm index 9d525d6c..50efaa59 100644 --- a/src/build/hostprogram.cppm +++ b/src/build/hostprogram.cppm @@ -128,6 +128,11 @@ inline void floor(const char* spec) { std::printf("mcpp:floor=%s\n // (like link_lib/link_search, unlike include_dir), because the package that // knows a board's layout is not the package being built. inline void link_script(const char* path) { std::printf("mcpp:link-script=%s\n", path); } +// A linker flag this program COMPUTED. The outlet `link_lib` / `link_search` / +// `link_script` leave open: a generated version script, `--wrap`, +// `--exclude-libs`. Reaches the consumer's link line, as `[build] ldflags` +// already does -- see the table row for why a private form is not offered. +inline void link_flag(const char* flag) { std::printf("mcpp:link-flag=%s\n", flag); } // ── Build-graph nodes (mcpp 2026.8.5.1+) ──────────────────────────────── // Declare WORK instead of doing it. A build program is a good place to decide // what the build looks like and a bad place to perform it: work done here is diff --git a/src/build/ninja_backend.cppm b/src/build/ninja_backend.cppm index fe1402b7..941e58eb 100644 --- a/src/build/ninja_backend.cppm +++ b/src/build/ninja_backend.cppm @@ -272,6 +272,55 @@ std::string shared_soname_flag(const LinkUnit& lu, const BuildPlan& plan) { return lu.soname.empty() ? "" : "-Wl,-soname," + lu.soname; } +// WHICH SYMBOLS A SHARED LIBRARY PUBLISHES, rendered per platform from one +// neutral list. +// +// The default on both platforms is "everything": ELF gives symbols default +// visibility, and PE gets an auto-generated .def listing every symbol +// (mcpp.build.coff_exports, CMake's WINDOWS_EXPORT_ALL_SYMBOLS semantics). +// Declaring `exports` narrows that, which is what a runtime with a stable ABI +// and a plugin loaded beside its rivals both need -- an ICD that exports its +// internals collides with the loader and with the other ICDs in the process. +// +// THREE RENDERINGS OF ONE STATEMENT, and that is the reason this is a manifest +// key rather than three flag lists in three `[target.]` blocks. It is the +// same shape `[runtime]` already established for link intent. +// +// The file is written into the build directory rather than read from the +// package, because the inline form has no file to read and because a version +// script's syntax is not what the author wrote. +std::string exports_file_contents(const LinkUnit& lu, std::string_view os) { + std::string out; + if (os == "macos") { + // One symbol per line. Mach-O symbols carry a leading underscore that + // the C++ source never writes, so it is added here -- the author names + // the symbol, not the object format's spelling of it. + for (auto const& p : lu.exportPatterns) out += "_" + p + "\n"; + return out; + } + // ELF version script. One anonymous version node: naming versions is a + // separate capability (symbol VERSIONING, `foo@@LIB_1.0`) that cannot be + // stated neutrally, and a package needing it writes the map itself and + // passes it with `[build] ldflags`. + out = "{\n global:\n"; + for (auto const& p : lu.exportPatterns) out += " " + p + ";\n"; + out += " local:\n *;\n};\n"; + return out; +} + +// The flag that names the file. PE is absent on purpose: there the export set +// is the `.def`, which lu.defFile already declares and the def-generating step +// already writes, so narrowing it is that step's business rather than a second +// flag on the link line. +std::string exports_flag(const LinkUnit& lu, std::string_view os, + const std::filesystem::path& file) { + if (lu.kind != LinkUnit::SharedLibrary || lu.exportPatterns.empty()) return ""; + if (os == "windows") return ""; + if (os == "macos") + return "-Wl,-exported_symbols_list," + file.generic_string(); + return "-Wl,--version-script=" + file.generic_string(); +} + // Write only when the bytes would actually change. // // One of these files is a BUILD INPUT: `obj/mcpp_ios_init.c`, the generated @@ -2081,6 +2130,31 @@ std::string emit_ninja_string(const BuildPlan& plan) { implicit.empty() ? std::string{} : " |" + implicit); if (auto flag = shared_soname_flag(lu, plan); !flag.empty()) out_line += " soname_flag = " + flag + "\n"; + // The export set, written beside the artifact and named on the link. + // + // Folded into `soname_flag` rather than given a rule variable of its + // own: both are "a property of THIS shared library's link", the rule + // already interpolates that variable in the right place, and a second + // one would have to be added to every link rule that mentions the + // first -- the "same decision in N places" shape this repository keeps + // paying for. + if (lu.kind == LinkUnit::SharedLibrary && !lu.exportPatterns.empty()) { + const auto tr = mcpp::toolchain::triple::parse(plan.toolchain.targetTriple); + const std::string os = tr ? tr->os + : (mcpp::platform::is_macos ? "macos" + : mcpp::platform::is_windows ? "windows" : "linux"); + const auto file = plan.outputDir / "obj" + / (lu.targetName + ".exports.gen"); + std::error_code mkec; + std::filesystem::create_directories(file.parent_path(), mkec); + write_file(file, exports_file_contents(lu, os)); + if (auto ef = exports_flag(lu, os, file); !ef.empty()) { + if (out_line.find(" soname_flag = ") == std::string::npos) + out_line += " soname_flag = " + ef + "\n"; + else + out_line.insert(out_line.rfind('\n'), " " + ef); + } + } // Where the linker is told to write it. A rule-level `$out.lib` would // spell the msvc case `foo.dll.lib`, i.e. a name nothing else in mcpp // agrees with — the name belongs to plan.cppm's import_library_for, and diff --git a/src/build/plan.cppm b/src/build/plan.cppm index 7462b0ab..27198faa 100644 --- a/src/build/plan.cppm +++ b/src/build/plan.cppm @@ -94,6 +94,11 @@ struct LinkUnit { // wins, so the emitter puts this after every other linker argument. // Deciding it stays here; placing it is the emitter's business. std::string loaderTagFlag; + // The symbol patterns this unit publishes, empty when it publishes all. + // Carried as the neutral list rather than a rendered flag: the file to + // write and the flag that names it are both platform-shaped, and the + // backend is where the target OS is known. + std::vector exportPatterns; std::filesystem::path output; // relative to plan.outputDir // The import library a PE shared library also produces — empty on ELF and // Mach-O, and empty for every non-shared unit. It is a SECOND output of the @@ -1675,6 +1680,7 @@ make_plan(const mcpp::manifest::Manifest& manifest, if (msvcTarget && !lu.importLibrary.empty()) lu.defFile = std::filesystem::path("bin") / (dep.target.name + ".def"); lu.soname = dep.target.soname; + lu.exportPatterns = dep.target.exportPatterns; lu.runtimeAliases = runtime_aliases_for_target(dep.target, naming); lu.loaderTagFlag = loader_tag_flag(lu.kind); append_package_objects(lu, dep.packageName); @@ -1707,6 +1713,7 @@ make_plan(const mcpp::manifest::Manifest& manifest, if (msvcTarget && !lu.importLibrary.empty()) lu.defFile = std::filesystem::path("bin") / (t.name + ".def"); lu.soname = t.soname; + lu.exportPatterns = t.exportPatterns; lu.runtimeAliases = runtime_aliases_for_target(t, naming); } else if (t.kind == mcpp::manifest::Target::TestBinary) { lu.kind = LinkUnit::TestBinary; diff --git a/src/build/prepare.cppm b/src/build/prepare.cppm index 12fafaa4..db1097fe 100644 --- a/src/build/prepare.cppm +++ b/src/build/prepare.cppm @@ -301,15 +301,23 @@ export void merge_conditional_xlings(mcpp::manifest::Manifest& m, m.xlings.featurePins.insert_or_assign(addr, pin); } -// A `[target..xlings…]` selector MUST NOT name a target-side layer. +// A `[target..xlings…]` selector MUST NOT name a RESOLVED layer. // -// Not a style rule, a schedule one. The five layer keys (`accelerator`, -// `c-abi`, `compiler`, …) are answered by dependency RESOLUTION, so a -// predicate naming one is held back to the second merge pass further down — -// which runs after every package's build.mcpp has already run and after the -// root's tool provisioning. An entry admitted there would be declared and -// never installed, and the failure it produces is the worst-shaped one there -// is: the build succeeds and the tool is simply absent. +// Not a style rule, a schedule one. The five resolved layer keys (`c-abi`, +// `compiler`, …) are answered by dependency RESOLUTION, so a predicate naming +// one is held back to the second merge pass further down — which runs after +// every package's build.mcpp has already run and after the root's tool +// provisioning. An entry admitted there would be declared and never installed, +// and the failure it produces is the worst-shaped one there is: the build +// succeeds and the tool is simply absent. +// +// `accelerator` IS ADMITTED, and used to be refused here with the rest. It is +// not resolved from anything: it is `--accel`, or `[build] accel`, read before +// the first package is looked up, so a payload predicated on it is merged in +// the FIRST pass and installed like any other. Refusing it had a cost paid on +// every build of every project with a device island — the vendor toolkit is +// declared unconditionally or not at all, so a CPU-only build downloaded +// gigabytes for a device it was not compiling for. // // Refused rather than deferred, and refused at the earliest point that can // see the predicate. The gate plane answers the case this shape is reached @@ -337,9 +345,10 @@ layer_predicated_xlings_refusal(const mcpp::manifest::Manifest& m) "which happens after tools are installed and after build programs " "run, so a tool conditioned on one would be declared and never " "installed. Condition it on the target instead " - "(`[target.'cfg(os = \"linux\")'.xlings.workspace]`), or gate it " - "on a feature (`[feature-xlings.]`), which is known " - "before anything is provisioned. See docs/05 section 2.13.", + "(`[target.'cfg(os = \"linux\")'.xlings.workspace]`), on the " + "accelerator (`[target.'cfg(accelerator = \"cuda\")'.xlings" + ".workspace]`, which IS answered before provisioning), or on a " + "feature (`[feature-xlings.]`). See docs/05 section 2.13.", cc.predicate, named); } return std::nullopt; @@ -1797,11 +1806,17 @@ prepare_build(bool print_fingerprint, "not know, so the section never applies (ignored). {}", cc.predicate, names, cfgpred::vocabulary_sentence())); } - // A layer is resolved AFTER dependency resolution, so a dependency - // selected by one would form a cycle with the resolution that produces - // the answer — docs/14 states this. The section's build inputs are - // honoured by the second pass; its dependencies cannot be, and saying - // so is the difference between a documented limit and a silent drop. + // A RESOLVED layer is answered AFTER dependency resolution, so a + // dependency selected by one would form a cycle with the resolution + // that produces the answer — docs/14 states this. The section's build + // inputs are honoured by the second pass; its dependencies cannot be, + // and saying so is the difference between a documented limit and a + // silent drop. + // + // `accelerator` is not one of these (see kCfgEarlyLayerKeys), so + // `[target.'cfg(accelerator = "cuda")'.dependencies]` is honoured and + // never reaches this warning: nothing about it is circular, because the + // accel is an input to the build rather than an answer from the graph. if (cfgpred::uses_layer(cc.predicate) && !(cc.dependencies.empty() && cc.devDependencies.empty() && cc.buildDependencies.empty() && cc.featureDeps.empty())) { @@ -4773,6 +4788,24 @@ prepare_build(bool print_fingerprint, std::map> hostModulesByConsumer; + // The same providers by INDEX, and the reason they are needed twice. + // + // A rule's code runs inside its CONSUMER's build program, so + // `mcpp::xpkg_dir("cuda-nvcc")` is asked there -- while the payload that + // answers it was declared by the RULE, under `[feature-xlings.]`, which + // is where it belongs: which packages a device compiler needs is the + // rule's knowledge and no project should have to rediscover it. + // + // The graph pass already INSTALLS what a dependency declares. Only the + // answer was missing: `fillXpkgDirs` read one manifest, so the address was + // fetched, unpacked, and then unreachable from the only code that wanted + // it -- a failure that reads as "the toolkit is not installed" while it + // sits on disk. + // + // The set is the host-module providers rather than every dependency: the + // code that can call `xpkg_dir` in this build program is the consumer's + // own `build.mcpp` plus exactly the rule modules compiled into it. + std::map> hostModuleProvidersByConsumer; // #359: who can see which build-time provision. Computed once by the // provisioning pass below (a fixpoint over `dependencyEdges`, the same // shape as computeUsageRequirements) and read by every consumer of the @@ -4887,7 +4920,8 @@ prepare_build(bool print_fingerprint, std::map namedRunnerProvider; auto fillXpkgDirs = [&](mcpp::build::BuildProgramEnv& e, - const mcpp::manifest::Manifest& owner) { + const mcpp::manifest::Manifest& owner, + std::size_t consumer) { // `[feature-xlings.]` is provisioned when `` is active, so it has // to be answerable here too. Before this, a tool a feature declared was // downloaded and installed and then `mcpp::xpkg_dir` returned "" for it @@ -4905,6 +4939,28 @@ prepare_build(bool print_fingerprint, for (auto const& address : it->second) if (std::ranges::find(declared, address) == declared.end()) declared.push_back(address); + // …and what the rule packages compiled INTO this build program + // declared. Their own active features, not the consumer's: the + // consumer asked for `features = ["rules-cuda"]` on the edge, and that + // is what decides which of the rule's `[feature-xlings]` tables apply. + if (auto pit = hostModuleProvidersByConsumer.find(consumer); + pit != hostModuleProvidersByConsumer.end()) { + for (auto q : pit->second) { + if (q >= packages.size()) continue; + auto const& pm = packages[q].manifest; + auto want = [&](const std::string& address) { + if (std::ranges::find(declared, address) == declared.end()) + declared.push_back(address); + }; + for (auto const& address : pm.xlings.deps) want(address); + const auto& pf = q < activeFeaturesByPackage.size() + ? activeFeaturesByPackage[q] : std::vector{}; + for (auto const& f : pf) + if (auto it = pm.xlings.featureDeps.find(f); + it != pm.xlings.featureDeps.end()) + for (auto const& address : it->second) want(address); + } + } if (declared.empty()) return; auto cfg = get_cfg(); if (!cfg) return; @@ -7129,8 +7185,66 @@ prepare_build(bool print_fingerprint, " and say so only at the link, or never.", pkg.manifest.package.name, sc.glob, sc.accel)); } + // A backend the package never declared. Checked BEFORE + // the build's own accel is consulted, because it is a + // property of the manifest alone and because the exclusion + // below would otherwise turn `accel = "cude12.9"` into a + // glob that is quietly never built. Only when the package + // states its backends -- `[package] accelerators` is + // optional, and a package that omits it has said nothing to + // contradict. + if (!pkg.manifest.package.accelerators.empty()) { + for (auto const& w : mcpp::pack::parse_accel(sc.accel)) { + if (std::ranges::find(pkg.manifest.package.accelerators, + w.backend) + != pkg.manifest.package.accelerators.end()) continue; + std::string declared; + for (auto const& a : pkg.manifest.package.accelerators) + declared += (declared.empty() ? "" : ", ") + a; + return std::unexpected(std::format( + "`{}`: [build] sources entry '{}' names accelerator " + "backend \"{}\", which this package does not declare.\n" + " [package] accelerators = [{}]\n" + " A constrained glob is left out of builds that do " + "not name its\n" + " backend, so a backend spelled wrong here is a file " + "that is never\n" + " compiled and never mentioned.\n" + " fix: correct the spelling, or add the backend to " + "`[package] accelerators`.", + pkg.manifest.package.name, sc.glob, w.backend, declared)); + } + } if (buildAccel.empty()) { excludedGlobs.insert(sc.glob); continue; } const auto want = mcpp::pack::parse_accel(sc.accel); + + // A GLOB WHOSE BACKEND THIS BUILD NEVER NAMED IS NOT A + // MISMATCH, IT IS ABSENT. + // + // The refusal below is about a real disagreement: a file + // written for sm_89 in a build that targets sm_80 is not a + // variant. Across DIFFERENT backends there is no such + // disagreement. A package with a CUDA island and a Vulkan + // one, built with `--accel vulkan1.2`, is asking for the + // Vulkan half; refusing it made a build that names a SUBSET + // of a package's backends impossible, so a package could + // have several device backends only if every build took all + // of them. + // + // The glob is dropped exactly as `--no-accel` drops it, and + // the `cfg(accelerator = ...)` section carrying that + // backend's host half does not activate either, so the two + // halves stay together. + // + // What keeps a TYPO from becoming a silent exclusion is the + // check below, against `[package] accelerators`: a backend + // the package never declared is refused before this point. + bool backendNamed = false; + for (auto const& w : want) + for (auto const& b : buildAccel) + if (b.backend == w.backend) backendNamed = true; + if (!backendNamed) { excludedGlobs.insert(sc.glob); continue; } + if (!mcpp::pack::accel_accepts(buildAccel, want)) { refusal::record(refusal::Code::AccelMismatch); return std::unexpected(std::format( @@ -7402,6 +7516,11 @@ prepare_build(bool print_fingerprint, if (auto r = visit(visit, p); !r) return std::unexpected(r.error()); + // Every provider on this consumer's rule closure, transitive + // ones included -- `done` is exactly that set, and a rule + // imported by another rule declares payloads just as directly. + hostModuleProvidersByConsumer[c].assign(done.begin(), done.end()); + if (auto clash = prov::host_module_collision(ordered)) return std::unexpected(*clash); for (auto const& hm : ordered) { @@ -7844,7 +7963,7 @@ prepare_build(bool print_fingerprint, // …and the xlings packages this package itself declared. Its own // manifest, not the root's: a dependency's `[xlings] deps` is what // its build.mcpp asks about. - fillXpkgDirs(bpEnv, packages[i].manifest); + fillXpkgDirs(bpEnv, packages[i].manifest, i); // #355: the host tools THIS package requested (resolved above). if (auto tit = toolEnvByConsumer.find(i); tit != toolEnvByConsumer.end()) bpEnv.toolPaths = tit->second; @@ -8763,7 +8882,7 @@ prepare_build(bool print_fingerprint, bpEnv.features = feature_closure(*m, parse_feature_request(overrides.features)); // mcpp#241 (root): consumer index 0, same owner as the dep loop. fillDepDirs(bpEnv, 0); - fillXpkgDirs(bpEnv, *m); + fillXpkgDirs(bpEnv, *m, 0); // #355: the host tools the ROOT package requested (consumer index 0). if (auto tit = toolEnvByConsumer.find(0u); tit != toolEnvByConsumer.end()) bpEnv.toolPaths = tit->second; @@ -8831,6 +8950,70 @@ prepare_build(bool print_fingerprint, bcRoot.ldflags.begin() + rldN, bcRoot.ldflags.end()); } + // ── Every device source must reach some action ───────────────────────── + // + // A device-kind file is the one source the engine has no compile rule for. + // It is handed to the package's build program (MCPP_DEVICE_SOURCES) and + // comes back as an action, or it is not compiled at all. Nothing checked + // that it came back. Two ways it does not, both silent until now: + // + // - the package has no `build.mcpp`. The engine computed the list and + // dropped it. Both run sites above are guarded on that file existing, + // so there was not even a program to ignore it. + // - a program runs but no imported rule claims the extension. A project + // with a `.cu` and a `.comp` that imports only `mcpp.rules.spirv` is + // this case, and it is the ordinary case for a project with two + // backends: a rule takes the extensions it knows and leaves the rest. + // + // What they produce today is an undefined reference at the link, naming a + // symbol and never the file that would have defined it -- and for a + // `kind = "lib"` target not even that, because an archive is not resolved. + // A device source that compiles nothing is never what was meant, so it is + // refused here, where both halves of the fact are still in hand. + // + // THE CRITERION IS THE ACTION INPUTS, not "a build program ran": a program + // that ran and consumed nothing is exactly the second case. It is also the + // condition an action needs anyway -- one that compiles a file it does not + // declare as an input does not rerun when that file changes -- so a rule + // that satisfies it is a rule that rebuilds correctly. + for (std::size_t i = 0; i < packages.size(); ++i) { + auto const& pkg = packages[i]; + auto dit = deviceSourcesByPackage.find(pkg.root.string()); + if (dit == deviceSourcesByPackage.end() || dit->second.empty()) continue; + auto const& mm = (i == 0) ? *m : pkg.manifest; + std::set consumed; + for (auto const& a : mm.buildConfig.actions) + for (auto const& in : a.inputs) { + std::filesystem::path ip(in); + consumed.insert((ip.is_absolute() ? ip : pkg.root / ip).lexically_normal()); + } + std::string orphans; + for (auto const& rel : dit->second) + if (!consumed.contains((pkg.root / rel).lexically_normal())) + orphans += " " + rel + "\n"; + if (orphans.empty()) continue; + std::error_code hasEc; + const bool hasProgram = std::filesystem::exists(pkg.root / "build.mcpp", hasEc); + return std::unexpected(std::format( + "`{}`: device sources that no action compiles:\n{}" + " A device-kind source is compiled by this package's build program\n" + " and by nothing else -- the engine has no rule for these extensions\n" + " and never will.\n" + "{}", + mm.package.name, orphans, + hasProgram + ? " `build.mcpp` ran but declared no action taking them as inputs.\n" + " fix: import the rule package that claims these extensions and\n" + " call it, or drop them from `[build] sources`. A rule that\n" + " compiles a file must also declare it as an action input, or the\n" + " action will not rerun when the file changes." + : " This package has no `build.mcpp`, so nothing was ever offered\n" + " them.\n" + " fix: add a `build.mcpp` importing the rule for these files (e.g.\n" + " `mcpp.rules.cuda` for `.cu`, `mcpp.rules.spirv` for shaders), or\n" + " drop them from `[build] sources`.")); + } + // [targets.*] required_features gate: a target is emitted only when ALL its // required features are active in this build; otherwise it is silently // skipped. A pure build-selection knob — it runs before the modgraph/plan diff --git a/src/build/prepare_inputs.cppm b/src/build/prepare_inputs.cppm index 69208cbf..da61e85d 100644 --- a/src/build/prepare_inputs.cppm +++ b/src/build/prepare_inputs.cppm @@ -92,9 +92,37 @@ struct Ctx { // backends are enabled". Membership everywhere keeps `any`/`all`/`not` // pure boolean combinators, and a single-backend build still answers // `accelerator = "cuda"` true and `accelerator = "rocm"` false. + // `none` IS THE EMPTY SET, AND AN OPEN VOCABULARY CANNOT SAY THAT BY + // ENUMERATION. + // + // A CPU fallback used to be written `not(any(accelerator = "cuda", + // accelerator = "vulkan"))`. `accelerator`'s vocabulary is OPEN by design + // -- docs/20 states that a fifth backend is a package rather than an + // engine change -- so that predicate's meaning changes the day a fifth one + // exists: every fallback already written silently starts matching a build + // that named the new backend. The failure is that the CPU implementation + // and the device implementation compile together, or that neither does. + // + // The spelling is the one this manifest already uses for the same idea: + // `os = "none"` is bare metal (docs/05 section 2.7.2). One word, one + // meaning, no new vocabulary. + // + // NOT `cpu`. That would put a second question on this axis -- the axis + // answers "which device compiler, which architecture", and the CPU needs + // neither -- and it would leave `cfg(accelerator = "cpu")` undecided under + // `accel = "cuda"`: true makes the fallback compile alongside the device + // path and destroys the mutual exclusion the seam exists for; false forces + // every existing manifest to write `accel = "cuda, cpu"`. + // + // A build where BOTH a CPU path and a device path are wanted needs none of + // this: the CPU sources go in the unconditional `[build] sources` and the + // device sources under `cfg(accelerator = "x")`. `not(...)` was only ever + // needed for a mutually exclusive seam, which is the case this repairs. bool layer_matches(std::string_view k, std::string_view v) const { - if (k == "accelerator") + if (k == "accelerator") { + if (v == "none") return accelerators.empty(); return std::ranges::find(accelerators, v) != accelerators.end(); + } return layer_value(k) == v; } }; @@ -146,17 +174,49 @@ inline Ctx context_for(std::string_view targetTriple) { inline constexpr std::string_view kCfgTripleKeys[] = { "arch", "env", "family", "os", }; +// THE LAYER KEYS SPLIT BY SCHEDULE, not by subject matter. +// +// The five in `kCfgLayerKeys` are answered BY dependency resolution: which C +// library, which compiler, which compiler runtime the graph settled on. A +// predicate naming one cannot be evaluated before the graph exists, which is +// why the second merge pass owns them and why a dependency conditioned on one +// is refused -- it would decide the answer it is asking for. +// +// `accelerator` is not like them. It is an INPUT: `--accel`, or `[build] +// accel`, read near the top of prepare() and known before the first package is +// resolved. Grouping it with the five made three things wrong at once. A +// payload could not be gated on the device it is for, so a CPU-only build of a +// project that also has a CUDA island downloaded the whole vendor toolkit. A +// dependency under `cfg(accelerator = ...)` was warned about and dropped, +// though nothing about it is circular. And the section was carried to the late +// pass for no reason at all. +// +// Both sets are the cfg VOCABULARY, so `is_cfg_layer_key` still answers for +// either; only the schedule question (`uses_layer`) distinguishes them. +inline constexpr std::string_view kCfgEarlyLayerKeys[] = { + "accelerator", +}; inline constexpr std::string_view kCfgLayerKeys[] = { - "accelerator", "c++-abi", "c-abi", "compiler", "compiler-runtime", + "c++-abi", "c-abi", "compiler", "compiler-runtime", "kernel-abi", }; inline constexpr std::string_view kCfgBarewords[] = { "linux", "macos", "unix", "windows", }; -inline bool is_cfg_layer_key(std::string_view k) { +// Answerable before resolution. Its value comes from the build's own accel. +inline bool is_cfg_early_layer_key(std::string_view k) { + return std::ranges::find(kCfgEarlyLayerKeys, k) != std::end(kCfgEarlyLayerKeys); +} +// Answerable only after resolution -- what the second merge pass owns. +inline bool is_cfg_late_layer_key(std::string_view k) { return std::ranges::find(kCfgLayerKeys, k) != std::end(kCfgLayerKeys); } +// The vocabulary question: is this a layer key at all. Both sets, because an +// unknown token must stay unknown and `accelerator` is not one. +inline bool is_cfg_layer_key(std::string_view k) { + return is_cfg_early_layer_key(k) || is_cfg_late_layer_key(k); +} // Recursive-descent evaluator over the inside of `cfg(...)`: // expr := all(list) | any(list) | not(expr) | key="value" | bareword @@ -206,11 +266,15 @@ struct Parser { if (k == "arch") return c.arch == v; if (k == "family") return c.family == v; if (k == "env") return c.env == v; - // A layer key is not answerable until the target side is resolved. In - // the first (triple-only) pass this returns false and the section is - // skipped — which is correct, because the second pass owns it and would - // otherwise append the same inputs twice through `append()`. - if (is_cfg_layer_key(k)) + // `accelerator` is answerable whenever the context carries the build's + // accel, which is from the first pass onward -- see kCfgEarlyLayerKeys. + if (is_cfg_early_layer_key(k)) return c.layer_matches(k, v); + // The other layer keys are not answerable until the target side is + // resolved. In the first (triple-only) pass this returns false and the + // section is skipped — which is correct, because the second pass owns + // it and would otherwise append the same inputs twice through + // `append()`. + if (is_cfg_late_layer_key(k)) return c.layersKnown && c.layer_matches(k, v); return false; } @@ -302,8 +366,12 @@ inline PredicateScan scan_predicate(const std::string& predicate) { // both would contribute its inputs twice. inline bool uses_layer(const std::string& predicate) { auto scan = scan_predicate(predicate); + // The LATE keys only. A predicate naming `accelerator` is answered in the + // first pass, so claiming it here would move it to a pass that adds + // nothing and takes away the ability to gate a payload or a dependency on + // the device it is for. return std::ranges::any_of(scan.keys, - [](auto const& k) { return is_cfg_layer_key(k); }); + [](auto const& k) { return is_cfg_late_layer_key(k); }); } // Tokens outside the vocabulary. A predicate naming one of these used to diff --git a/tests/e2e/606_constrained_source_globs_narrow_to_the_build.sh b/tests/e2e/606_constrained_source_globs_narrow_to_the_build.sh index 0aa9e382..b689ee84 100755 --- a/tests/e2e/606_constrained_source_globs_narrow_to_the_build.sh +++ b/tests/e2e/606_constrained_source_globs_narrow_to_the_build.sh @@ -49,9 +49,36 @@ kind = "bin" main = "src/main.cpp" EOF cat > build.mcpp <<'EOF' +import std; import mcpp; + +// A `check` action per device source, and the reason it is here rather than a +// bare read of the variable: mcpp refuses a device source that reached no +// action (2026.9.6.5). A build program that only LOOKS at +// `mcpp::device_sources()` models a project whose device files compile to +// nothing, which is the defect that refusal exists to catch -- so this fixture +// declares the edge a real rule package would declare, and asserts on the +// variable as before. int main() { const char* d = mcpp::device_sources(); + std::string flat(d); + std::size_t n = 0, start = 0; + while (start <= flat.size()) { + auto nl = flat.find('\n', start); + auto one = flat.substr(start, nl == std::string::npos ? flat.size() - start : nl - start); + start = nl == std::string::npos ? flat.size() + 1 : nl + 1; + if (one.empty()) continue; + auto stamp = std::string(mcpp::out_dir()) + "/dev-" + std::to_string(n++) + ".stamp"; + mcpp::action a; + a.id = "seen"; + a.role = "check"; + a.description = "account for a device source"; + auto abs = std::string(mcpp::manifest_dir()) + "/" + one; + a.arg("cp").arg(abs.c_str()).arg(stamp.c_str()); + a.input(abs.c_str()); + a.output(stamp.c_str()); + a.submit(); + } mcpp::warning((*d ? d : "(no device sources)")); return 0; } diff --git a/tests/e2e/609_shader_sources_are_device_sources.sh b/tests/e2e/609_shader_sources_are_device_sources.sh index ffa04b85..a3bf22d7 100755 --- a/tests/e2e/609_shader_sources_are_device_sources.sh +++ b/tests/e2e/609_shader_sources_are_device_sources.sh @@ -39,8 +39,33 @@ mkdir -p shaders cat > build.mcpp <<'EOF' import std; import mcpp; + +// A `check` action per device source, and the reason it is here rather than a +// bare read of the variable: mcpp refuses a device source that reached no +// action (2026.9.6.5). A build program that only LOOKS at +// `mcpp::device_sources()` models a project whose device files compile to +// nothing, which is the defect that refusal exists to catch -- so this fixture +// declares the edge a real rule package would declare, and asserts on the +// variable as before. int main() { std::string flat(mcpp::device_sources()); + std::size_t n = 0, start = 0; + while (start <= flat.size()) { + auto nl = flat.find('\n', start); + auto one = flat.substr(start, nl == std::string::npos ? flat.size() - start : nl - start); + start = nl == std::string::npos ? flat.size() + 1 : nl + 1; + if (one.empty()) continue; + auto stamp = std::string(mcpp::out_dir()) + "/dev-" + std::to_string(n++) + ".stamp"; + mcpp::action a; + a.id = "seen"; + a.role = "check"; + a.description = "account for a device source"; + auto abs = std::string(mcpp::manifest_dir()) + "/" + one; + a.arg("cp").arg(abs.c_str()).arg(stamp.c_str()); + a.input(abs.c_str()); + a.output(stamp.c_str()); + a.submit(); + } for (auto& c : flat) if (c == '\n') c = ' '; mcpp::warning(("device=[" + flat + "]").c_str()); return 0; diff --git a/tests/e2e/613_sycl_sources_are_device_sources.sh b/tests/e2e/613_sycl_sources_are_device_sources.sh index cedf934d..0e65f965 100755 --- a/tests/e2e/613_sycl_sources_are_device_sources.sh +++ b/tests/e2e/613_sycl_sources_are_device_sources.sh @@ -34,8 +34,33 @@ mkdir -p src/kernels cat > build.mcpp <<'EOF2' import std; import mcpp; + +// A `check` action per device source, and the reason it is here rather than a +// bare read of the variable: mcpp refuses a device source that reached no +// action (2026.9.6.5). A build program that only LOOKS at +// `mcpp::device_sources()` models a project whose device files compile to +// nothing, which is the defect that refusal exists to catch -- so this fixture +// declares the edge a real rule package would declare, and asserts on the +// variable as before. int main() { std::string flat(mcpp::device_sources()); + std::size_t n = 0, start = 0; + while (start <= flat.size()) { + auto nl = flat.find('\n', start); + auto one = flat.substr(start, nl == std::string::npos ? flat.size() - start : nl - start); + start = nl == std::string::npos ? flat.size() + 1 : nl + 1; + if (one.empty()) continue; + auto abs = std::string(mcpp::manifest_dir()) + "/" + one; + auto stamp = std::string(mcpp::out_dir()) + "/dev-" + std::to_string(n++) + ".stamp"; + mcpp::action a; + a.id = "seen"; + a.role = "check"; + a.description = "account for a device source"; + a.arg("cp").arg(abs.c_str()).arg(stamp.c_str()); + a.input(abs.c_str()); + a.output(stamp.c_str()); + a.submit(); + } for (auto& c : flat) if (c == '\n') c = ' '; mcpp::warning(("device=[" + flat + "]").c_str()); return 0; diff --git a/tests/e2e/619_a_tool_may_not_be_conditioned_on_a_layer.sh b/tests/e2e/619_a_tool_may_not_be_conditioned_on_a_layer.sh index c313be81..ad8da2b3 100755 --- a/tests/e2e/619_a_tool_may_not_be_conditioned_on_a_layer.sh +++ b/tests/e2e/619_a_tool_may_not_be_conditioned_on_a_layer.sh @@ -3,8 +3,8 @@ # `[target..xlings…]` accepts a target predicate. It must REFUSE one # that names a target-side layer. # -# The reason is schedule, not style. The five layer keys (`accelerator`, -# `c-abi`, `compiler`, ...) are answered by dependency RESOLUTION, so a +# The reason is schedule, not style. The five RESOLVED layer keys (`c-abi`, +# `compiler`, ...) are answered by dependency RESOLUTION, so a # predicate naming one is held back to the second merge pass -- which runs after # tools are provisioned and after every build.mcpp. An entry admitted there is # declared and never installed, and the build that results is the worst-shaped diff --git a/tests/e2e/620_link_flag_reaches_the_link_line.sh b/tests/e2e/620_link_flag_reaches_the_link_line.sh new file mode 100755 index 00000000..867b8f85 --- /dev/null +++ b/tests/e2e/620_link_flag_reaches_the_link_line.sh @@ -0,0 +1,71 @@ +#!/usr/bin/env bash +# requires: gcc +# A linker flag a build program COMPUTED reaches the link line. +# +# `link-lib`, `link-search` and `link-script` each name one KIND of thing, so a +# flag the program worked out for itself -- a generated version script, +# `-Wl,--wrap=`, `-Wl,--exclude-libs` -- had no way out of build.mcpp. This is +# that outlet. +# +# THE CRITERION IS THE LINKER'S BEHAVIOUR, NOT THE COMMAND LINE. Asserting that +# a string appears in build.ninja would pass for a flag that was written down +# and never handed to the linker, and it would break the day the flag is +# rendered with different spacing. `-Wl,--defsym==` DEFINES a +# symbol at link time, so the program's own output is the evidence: the value +# can only be there if the linker saw the flag. +set -e + +TMP=$(mktemp -d) +trap "rm -rf $TMP" EXIT +cd "$TMP" + +mkdir -p src +cat > mcpp.toml <<'TOML' +[package] +name = "linkflag" +version = "0.1.0" + +[targets.linkflag] +kind = "bin" +main = "src/main.cpp" +TOML + +# The program computes the value rather than hard-coding it, because a value +# the manifest could have written is a case `[build] ldflags` already served. +cat > build.mcpp <<'CPP' +#include +#include +import mcpp; +int main() { + int computed = 40 + 2; + std::string flag = "-Wl,--defsym=mcpp_e2e_620=" + std::to_string(computed); + mcpp::link_flag(flag.c_str()); + return 0; +} +CPP + +cat > src/main.cpp <<'CPP' +#include +extern "C" char mcpp_e2e_620; +int main() { + // The linker put the value in the SYMBOL'S ADDRESS, which is how --defsym + // works; reading the object would read memory that was never written. + std::printf("defsym=%lld\n", + (long long)(unsigned long long)(void*)&mcpp_e2e_620); + return 0; +} +CPP + +if ! out=$("$MCPP" run 2>&1); then + echo "$out" + echo "FAIL: the project did not build or run" + exit 1 +fi +echo "$out" + +case "$out" in + *"defsym=42"*) ;; + *) echo "FAIL: the computed link flag did not reach the linker"; exit 1 ;; +esac + +echo "PASS: a link flag computed by build.mcpp reaches the link line" diff --git a/tests/e2e/621_exports_narrows_the_symbol_surface.sh b/tests/e2e/621_exports_narrows_the_symbol_surface.sh new file mode 100755 index 00000000..3add8d0b --- /dev/null +++ b/tests/e2e/621_exports_narrows_the_symbol_surface.sh @@ -0,0 +1,98 @@ +#!/usr/bin/env bash +# requires: elf gcc +# A shared library publishes what `exports` names, and everything otherwise. +# +# The default on both platforms is "everything": ELF gives symbols default +# visibility, and PE gets an auto-generated .def listing every symbol. Narrowing +# it is what a runtime with a stable ABI needs, and what a plugin loaded beside +# its rivals needs -- an ICD exporting its internals collides with the loader +# and with the other ICDs in the same process. +# +# THE CRITERION IS TWO-SIDED, AND BOTH SIDES ARE ASSERTED. Checking only that +# the public symbol is present would pass for a library exporting everything, +# which is the state before this feature. Checking only that the internal one is +# absent cannot tell "correctly hidden" from "never linked at all" -- so the +# same source is built twice, once with the key and once without, and the two +# readings must differ. +set -e + +TMP=$(mktemp -d) +trap 'rm -rf "$TMP"' EXIT +cd "$TMP" + +mkdir -p src abi +cat > src/lib.cpp <<'CPP' +extern "C" int mcpp_e2e_621_public(int x) { return x + 1; } +extern "C" int mcpp_e2e_621_internal(int x) { return x + 2; } +CPP + +cat > abi/lib.exports <<'EXPORTS' +# One symbol pattern per line; `#` starts a comment. +mcpp_e2e_621_public +EXPORTS + +emit_manifest() { +cat > mcpp.toml </dev/null | awk '{print $NF}' | grep '^mcpp_e2e_621' | sort +} + +# leg 1: no `exports` -- both symbols are published +emit_manifest "" +rm -rf target +"$MCPP" build >/dev/null 2>&1 || { echo "FAIL: the default build failed"; exit 1; } +before=$(dynsyms) +echo "default: $(echo "$before" | tr '\n' ' ')" +case "$before" in + *mcpp_e2e_621_public*) ;; + *) echo "FAIL: the default build did not publish the public symbol"; exit 1 ;; +esac +case "$before" in + *mcpp_e2e_621_internal*) ;; + *) echo "FAIL: the default is supposed to publish everything and did not." + echo " Without this leg the second one proves nothing."; exit 1 ;; +esac + +# leg 2: with `exports` -- only the named symbol +emit_manifest 'exports = "abi/lib.exports"' +rm -rf target +"$MCPP" build >/dev/null 2>&1 || { echo "FAIL: the build with exports failed"; exit 1; } +after=$(dynsyms) +echo "exports: $(echo "$after" | tr '\n' ' ')" +case "$after" in + *mcpp_e2e_621_public*) ;; + *) echo "FAIL: the declared symbol is not published"; exit 1 ;; +esac +case "$after" in + *mcpp_e2e_621_internal*) echo "FAIL: an undeclared symbol is still published"; exit 1 ;; +esac +[ "$before" != "$after" ] || { echo "FAIL: the two legs read identically"; exit 1; } + +# the inline form is the same statement +emit_manifest 'exports = ["mcpp_e2e_621_public"]' +rm -rf target +"$MCPP" build >/dev/null 2>&1 || { echo "FAIL: the inline form failed to build"; exit 1; } +inline=$(dynsyms) +[ "$inline" = "$after" ] || { + echo "FAIL: the inline list and the file disagree" + echo " file: $after" + echo " inline: $inline"; exit 1; } + +echo "PASS: exports narrows the published symbol set, and the two forms agree" diff --git a/tests/e2e/622_rule_declared_payload_reaches_the_consumers_build_program.sh b/tests/e2e/622_rule_declared_payload_reaches_the_consumers_build_program.sh new file mode 100755 index 00000000..1c512b01 --- /dev/null +++ b/tests/e2e/622_rule_declared_payload_reaches_the_consumers_build_program.sh @@ -0,0 +1,127 @@ +#!/usr/bin/env bash +# requires: elf gcc +# A RULE PACKAGE'S `[feature-xlings.]` must be findable from the CONSUMER's +# build program, because that is where the rule's code runs. +# +# The sibling case (618) is a dependency reading its OWN declaration from its +# OWN build.mcpp, and that works. This one is different in the only way that +# matters: a rule is compiled INTO its consumer's build program, so +# `mcpp::xpkg_dir` is asked in the consumer's environment while the payload was +# declared in the rule's manifest. `fillXpkgDirs` read one manifest, so the +# address was fetched, unpacked, and then unreachable from the only code that +# wanted it -- an answer of "" that reads as "the toolkit is not installed" +# while it sits on disk. +# +# THE CRITERION IS THE PATH THE RULE ANSWERS WITH, not that the build +# succeeded: a build whose rule silently found nothing succeeds too. +set -e + +TMP=$(mktemp -d) +trap "rm -rf $TMP" EXIT +cd "$TMP" + +# Same reasoning as 618: the payload has to be one mcpp does not install for +# its own reasons, or the criterion selects an object that is present anyway. +TOOL=shaderc +TOOL_VERSION="2026.3" + +mkdir -p rule/src +cat > rule/src/rule.cppm < rule/mcpp.toml < app/src/main.cpp <<'EOF' +int main() { return 0; } +EOF +cat > app/build.mcpp <<'EOF' +#include +#include +import mcpp; +import rule; +int main() { + // TO A FILE, NOT stdout: mcpp prints a build program's output only when it + // FAILS, so an assertion grepping the build log would be unreachable on + // exactly the run that is supposed to produce it. + std::string out = std::string(mcpp::manifest_dir()) + "/rule-saw.txt"; + std::FILE* f = std::fopen(out.c_str(), "w"); + if (f == nullptr) return 3; + std::fprintf(f, "%s\n", testrule::tool_dir().c_str()); + std::fclose(f); + return 0; +} +EOF +cat > app/mcpp.toml <<'EOF' +[package] +name = "consumer" +version = "0.1.0" +[language] +standard = "c++23" +modules = true +import_std = true +# `[build-dependencies]`, because the rule must never reach the target, and +# `host-module = true`, because its module must be compiled for the build +# program. Two axes, and this package answers them separately. +[build-dependencies] +rule = { path = "../rule", features = ["usestool"], host-module = true } +[targets.consumer] +kind = "bin" +main = "src/main.cpp" +EOF + +# An isolated home, for 618's reason: the ambient registry very likely holds +# the payload already, and then this passes on a broken engine. +export MCPP_HOME="$TMP/home" +mkdir -p "$MCPP_HOME" + +cd app +if ! "$MCPP" build >build.log 2>&1; then + echo "FAIL: the consumer did not build" + grep -iE 'error|not provisioned' build.log | head -5 + exit 1 +fi +seen="$TMP/app/rule-saw.txt" +[ -s "$seen" ] || { + echo "FAIL: the rule left no record" + tail -10 build.log + exit 1 +} +answer=$(cat "$seen") +echo "the rule, inside the consumer's build program, answered: '${answer}'" +case "$answer" in + *xim-x-$TOOL*) ;; + "") echo "FAIL: the rule got an empty answer -- its payload was installed and unreachable"; exit 1 ;; + *) echo "FAIL: the answer does not name $TOOL"; exit 1 ;; +esac +echo "PASS: a rule package's [feature-xlings] payload is reachable from the consumer's build program" diff --git a/tests/e2e/623_device_sources_must_reach_an_action.sh b/tests/e2e/623_device_sources_must_reach_an_action.sh new file mode 100755 index 00000000..95d61231 --- /dev/null +++ b/tests/e2e/623_device_sources_must_reach_an_action.sh @@ -0,0 +1,83 @@ +#!/usr/bin/env bash +# requires: elf gcc +# A device-kind source that no action compiles is refused, and the refusal +# names the file. +# +# The engine has no compile rule for `.cu` and never will: a device source is +# handed to the package's build program and comes back as an action, or it is +# not compiled at all. Nothing checked that it came back, so a project whose +# build program does not claim the extension -- or that has no build program -- +# got an undefined reference at the link naming a SYMBOL and never the file +# that would have defined it. For a `kind = "lib"` target not even that, +# because an archive is not resolved. +# +# TWO LEGS, because the two situations have different fixes and the message +# has to distinguish them. +set -e + +TMP=$(mktemp -d) +trap "rm -rf $TMP" EXIT +cd "$TMP" + +mk() { + rm -rf p; mkdir -p p/src/kernels + cat > p/src/main.cpp <<'EOF' +extern "C" void k(); +int main() { return 0; } +EOF + cat > p/src/kernels/k.cu <<'EOF' +extern "C" __global__ void k() {} +EOF + cat > p/mcpp.toml <<'EOF' +[package] +name = "orphan" +version = "0.1.0" +accelerators = ["cuda"] +[build] +accel = "cuda12.9+{sm_89}" +sources = [ + "src/*.cpp", + { glob = "src/kernels/*.cu", accel = "cuda12.9+{sm_89}" }, +] +[targets.orphan] +kind = "bin" +main = "src/main.cpp" +EOF +} + +# ── leg 1: no build program at all ────────────────────────────────────────── +mk +cd p +out=$("$MCPP" build 2>&1) && { echo "FAIL: a package with an uncompilable device source built"; exit 1; } +echo "$out" | grep -q 'src/kernels/k.cu' || { echo "FAIL: the refusal does not name the file"; echo "$out" | tail -5; exit 1; } +echo "$out" | grep -q 'no `build.mcpp`' || { echo "FAIL: the refusal does not say the build program is missing"; echo "$out" | tail -5; exit 1; } +echo "ok: no build program -- named the file and the missing program" +cd .. + +# ── leg 2: a build program that claims nothing ────────────────────────────── +# +# The ordinary shape for a project with two backends: a rule takes the +# extensions it knows and leaves the rest, so a file no imported rule claims is +# left over. Distinguished from leg 1 because the fix is different. +mk +cat > p/build.mcpp <<'EOF' +import std; +import mcpp; +int main() { return 0; } +EOF +cd p +out=$("$MCPP" build 2>&1) && { echo "FAIL: an unclaimed device source built"; exit 1; } +echo "$out" | grep -q 'src/kernels/k.cu' || { echo "FAIL: the refusal does not name the file"; echo "$out" | tail -5; exit 1; } +echo "$out" | grep -q 'ran but declared no action' || { echo "FAIL: the refusal does not distinguish a program that claimed nothing"; echo "$out" | tail -5; exit 1; } +echo "ok: build program present -- named the file and said no action took it" +cd .. + +# ── the negative leg: --no-accel leaves the glob out, so there is nothing to +# refuse. Without this the two legs above would pass on an engine that refused +# EVERY device source, which is a different and wrong behaviour. +mk +cd p +"$MCPP" build --no-accel >/dev/null 2>&1 || { echo "FAIL: --no-accel must not be refused; the constrained glob is left out"; exit 1; } +echo "ok: --no-accel builds -- no device source, nothing to account for" + +echo "PASS: a device source that reaches no action is refused, and only then" diff --git a/tests/e2e/624_a_missing_host_module_names_the_key.sh b/tests/e2e/624_a_missing_host_module_names_the_key.sh new file mode 100755 index 00000000..90be3147 --- /dev/null +++ b/tests/e2e/624_a_missing_host_module_names_the_key.sh @@ -0,0 +1,88 @@ +#!/usr/bin/env bash +# requires: elf gcc +# An import no dependency provides is refused BY NAME, before the compiler is +# reached. +# +# Left to the compiler the message is +# +# mcpp.rules.spirv: error: failed to read compiled module: No such file or +# directory +# mcpp.rules.spirv: note: imports must be built before being imported +# +# which is true and names neither the package that would provide the module nor +# the key that would make it importable. `host-module = true` and the section a +# dependency is written in are separate axes -- the section says whether the +# package reaches the target, `host-module` whether its module is compiled for +# the build program -- and forgetting the first while getting the second right +# is the ordinary mistake. +# +# THE SET OF NAMES THAT CAN COMPILE HERE IS CLOSED (`std`, `std.compat`, the +# bundled `mcpp`, and the importable host modules), so a name outside it cannot +# become valid later and is refused rather than warned about. +set -e + +TMP=$(mktemp -d) +trap "rm -rf $TMP" EXIT +cd "$TMP" + +mkdir -p rule/src +cat > rule/src/rule.cppm <<'EOF' +export module rule; +export namespace testrule { inline int answer() { return 42; } } +EOF +cat > rule/mcpp.toml <<'EOF' +[package] +name = "rule" +version = "0.1.0" +[language] +standard = "c++23" +modules = true +import_std = true +[build] +sources = ["src/rule.cppm"] +[targets.rule] +kind = "lib" +EOF + +mkdir -p app/src +cat > app/src/main.cpp <<'EOF' +int main() { return 0; } +EOF +cat > app/build.mcpp <<'EOF' +import std; +import mcpp; +import rule; +int main() { return testrule::answer() == 42 ? 0 : 1; } +EOF +# Declared, and declared in the right section -- only `host-module` is absent. +# That is the point: the manifest looks correct. +cat > app/mcpp.toml <<'EOF' +[package] +name = "consumer" +version = "0.1.0" +[language] +standard = "c++23" +modules = true +import_std = true +[build-dependencies] +rule = { path = "../rule" } +[targets.consumer] +kind = "bin" +main = "src/main.cpp" +EOF + +cd app +out=$("$MCPP" build 2>&1) && { echo "FAIL: the build succeeded with an unprovided import"; exit 1; } +echo "$out" | grep -q "imports 'rule'" || { echo "FAIL: the refusal does not name the module"; echo "$out" | tail -6; exit 1; } +echo "$out" | grep -q 'host-module = true' || { echo "FAIL: the refusal does not name the key"; echo "$out" | tail -6; exit 1; } +echo "$out" | grep -q 'rule (in \[build-dependencies\])' || { echo "FAIL: the refusal does not name the candidate dependency"; echo "$out" | tail -6; exit 1; } +echo "$out" | grep -qi 'failed to read compiled module' && { echo "FAIL: the compiler was reached; the check must run first"; exit 1; } +echo "ok: refused by name, before the compiler" + +# The same project with the key added must build -- otherwise the check could +# be refusing something that was always going to work. +sed -i 's|rule = { path = "../rule" }|rule = { path = "../rule", host-module = true }|' mcpp.toml +"$MCPP" build >/dev/null 2>&1 || { echo "FAIL: adding host-module = true did not make it build"; exit 1; } +echo "ok: the fix the message names is the fix that works" + +echo "PASS: an import no dependency provides is refused by name, and the named fix works" diff --git a/tests/e2e/625_a_payload_can_be_gated_on_the_accelerator.sh b/tests/e2e/625_a_payload_can_be_gated_on_the_accelerator.sh new file mode 100755 index 00000000..b6dfa86f --- /dev/null +++ b/tests/e2e/625_a_payload_can_be_gated_on_the_accelerator.sh @@ -0,0 +1,71 @@ +#!/usr/bin/env bash +# requires: elf gcc +# A payload predicated on the ACCELERATOR is installed for a device build and +# not for a CPU-only one. +# +# `accelerator` was grouped with the five resolved layer keys (`c-abi`, +# `compiler`, ...) and refused in `[target.'cfg(...)'.xlings]` for their reason: +# a layer is answered by dependency resolution, which runs after provisioning. +# That reason does not hold for this one key. The accel is an INPUT -- `--accel` +# or `[build] accel` -- read before the first package is resolved. +# +# The cost of the old grouping was paid on every build of every project with a +# device island: the vendor toolkit is declared unconditionally or not at all, +# so a CPU-only build downloaded gigabytes for a device it was not compiling +# for. There was no third spelling. +# +# TWO LEGS AND THE SECOND IS THE POINT. A test that only checked the device leg +# would pass on an engine that installs the payload always. +set -e + +TMP=$(mktemp -d) +trap "rm -rf $TMP" EXIT +cd "$TMP" + +# Same selection rule as 618: a payload mcpp does not install for its own +# reasons, or the criterion measures something that is present anyway. +TOOL=shaderc +TOOL_VERSION="2026.3" + +mkdir -p p/src +cat > p/src/main.cpp <<'EOF' +int main() { return 0; } +EOF +cat > p/mcpp.toml <cpu.log 2>&1 || { echo "FAIL: the CPU-only build was refused"; tail -20 cpu.log; exit 1; } +if [ -d "$store" ]; then + echo "FAIL: a CPU-only build installed the device payload at $store" + exit 1 +fi +echo "ok: no accel -- the gated payload was not installed" + +# ── leg 2: the accelerator named -- it must be ───────────────────────────── +"$MCPP" build --accel "vulkan1.2" >dev.log 2>&1 || { echo "FAIL: the device build was refused"; tail -20 dev.log; exit 1; } +[ -d "$store" ] || { + echo "FAIL: the device build did not install the payload its predicate names" + tail -20 dev.log + exit 1 +} +echo "ok: --accel vulkan1.2 -- the gated payload was installed" + +echo "PASS: a payload can be gated on the accelerator, in both directions" diff --git a/tests/e2e/626_naming_a_subset_of_backends_is_not_a_mismatch.sh b/tests/e2e/626_naming_a_subset_of_backends_is_not_a_mismatch.sh new file mode 100755 index 00000000..862b86c4 --- /dev/null +++ b/tests/e2e/626_naming_a_subset_of_backends_is_not_a_mismatch.sh @@ -0,0 +1,93 @@ +#!/usr/bin/env bash +# requires: elf gcc +# A constrained glob whose BACKEND this build never named is left out, not +# refused. A glob whose backend IS named but whose architecture is not covered +# still is. +# +# The refusal is about a real disagreement -- a file written for sm_89 in a +# build targeting sm_80 is not a variant. Across DIFFERENT backends there is no +# disagreement, and refusing there made a build that names a SUBSET of a +# package's backends impossible: a package could have several device backends +# only if every build took all of them. That is the opposite of what an +# additive-backend library needs. +# +# FOUR LEGS, because each states something the others cannot. Leg 2 is the +# change; leg 3 is what must NOT have changed with it; leg 4 is what keeps +# leg 2 from turning a typo into a file that is never compiled. +set -e + +TMP=$(mktemp -d) +trap "rm -rf $TMP" EXIT +cd "$TMP" + +mk() { # $1 = the glob's accel, $2 = [package] accelerators list + rm -rf p; mkdir -p p/src/kernels + cat > p/src/main.cpp <<'EOF' +int main() { return 0; } +EOF + cat > p/src/kernels/k.cu <<'EOF' +extern "C" __global__ void k() {} +EOF + cat > p/mcpp.toml < p/build.mcpp <<'EOF' +import std; +import mcpp; +int main() { return 0; } +EOF +} + +# ── leg 1: the backend is named -- the glob comes through ─────────────────── +mk 'cuda12.9+{sm_89}' '"cuda"' +cd p +out=$("$MCPP" build --accel 'cuda12.9+{sm_89}' 2>&1) && { echo "FAIL: the glob did not reach the build"; exit 1; } +echo "$out" | grep -q 'no action compiles' || { echo "FAIL: expected the device-source audit, got:"; echo "$out" | tail -4; exit 1; } +echo "ok: backend named -- the glob is in the build (the audit sees it)" +cd .. + +# ── leg 2: a DIFFERENT backend -- the glob is left out ────────────────────── +mk 'cuda12.9+{sm_89}' '"cuda"' +cd p +"$MCPP" build --accel 'vulkan1.2' >leg2.log 2>&1 || { + echo "FAIL: naming a different backend was refused" + tail -8 leg2.log + exit 1 +} +echo "ok: a different backend -- the glob is left out, as --no-accel leaves it" +cd .. + +# ── leg 3: the SAME backend, an architecture it does not cover -- refused ─── +mk 'cuda12.9+{sm_89}' '"cuda"' +cd p +out=$("$MCPP" build --accel 'cuda12.9+{sm_80}' 2>&1) && { echo "FAIL: an uncovered architecture built"; exit 1; } +echo "$out" | grep -q 'does not cover' || { echo "FAIL: expected the accel mismatch refusal, got:"; echo "$out" | tail -4; exit 1; } +echo "ok: same backend, uncovered architecture -- still refused" +cd .. + +# ── leg 4: a backend the package never declared -- refused ────────────────── +# +# Without this, leg 2 would turn `accel = "cude12.9"` into a glob that is +# quietly never built and never mentioned. +mk 'cude12.9+{sm_89}' '"cuda"' +cd p +out=$("$MCPP" build --accel 'cuda12.9+{sm_89}' 2>&1) && { echo "FAIL: an undeclared backend built"; exit 1; } +echo "$out" | grep -q 'does not declare' || { echo "FAIL: expected the undeclared-backend refusal, got:"; echo "$out" | tail -4; exit 1; } +echo "$out" | grep -q 'cude' || { echo "FAIL: the refusal does not name the misspelling"; exit 1; } +echo "ok: a backend [package] accelerators does not list -- refused, naming it" + +echo "PASS: a subset of backends builds; a mismatch and a misspelling do not" diff --git a/tests/unit/test_cfg_accelerator.cpp b/tests/unit/test_cfg_accelerator.cpp index f6f3a1af..953ebda4 100644 --- a/tests/unit/test_cfg_accelerator.cpp +++ b/tests/unit/test_cfg_accelerator.cpp @@ -61,18 +61,47 @@ TEST(CfgAccelerator, ComposesWithTripleKeys) { EXPECT_FALSE(m(R"(cfg(all(windows, accelerator = "cuda")))", cuda)); } -TEST(CfgAccelerator, UnresolvedTargetSideDoesNotMatch) { - // The first pass runs before dependency resolution and cannot answer a - // layer key. Returning false there is correct: the second pass owns it and - // would otherwise contribute the same inputs twice. +TEST(CfgAccelerator, AnsweredBeforeResolutionUnlikeTheOtherLayers) { + // `accelerator` is answerable in the FIRST merge pass, and the difference + // from the other five layer keys is the schedule, not the subject. Its + // value is an input to the build -- `--accel`, or `[build] accel` -- read + // before the first package is resolved, while `c-abi` is an answer the + // dependency graph produces. + // + // The consequence is what this states: a payload, or a dependency, can be + // gated on the device it is for. Grouped with the resolved keys it could + // not be, and a CPU-only build of a project with a device island + // downloaded the whole vendor toolkit. auto early = cfgpred::context_for("x86_64-linux-gnu"); ASSERT_FALSE(early.layersKnown); - EXPECT_FALSE(m(R"(cfg(accelerator = "cuda"))", early)); + early.accelerators = {"cuda"}; + EXPECT_TRUE (m(R"(cfg(accelerator = "cuda"))", early)); + EXPECT_FALSE(m(R"(cfg(accelerator = "vulkan"))", early)); + + // A resolved layer key in the same position still answers false, and must: + // the second pass owns it and would otherwise contribute the same inputs + // twice through `append()`. + EXPECT_FALSE(m(R"(cfg(c-abi = "glibc"))", early)); + + // No accel named: false for any backend, and `none` is how a section says + // so without enumerating the backends it is not. + auto nothing = cfgpred::context_for("x86_64-linux-gnu"); + EXPECT_FALSE(m(R"(cfg(accelerator = "cuda"))", nothing)); + EXPECT_TRUE (m(R"(cfg(accelerator = "none"))", nothing)); } -TEST(CfgAccelerator, IsALayerKeyNotATripleKey) { - EXPECT_TRUE (cfgpred::uses_layer(R"(cfg(accelerator = "cuda"))")); +TEST(CfgAccelerator, IsInTheVocabularyButNotAResolvedLayer) { + // Three separate questions, and only the middle one changed. It is a known + // key (so a misspelling is still reported); it is not a triple key; and it + // does NOT send its section to the pass that runs after resolution. + EXPECT_TRUE (cfgpred::unknown_tokens(R"(cfg(accelerator = "cuda"))").empty()); + EXPECT_FALSE(cfgpred::uses_layer(R"(cfg(accelerator = "cuda"))")); EXPECT_FALSE(cfgpred::uses_layer(R"(cfg(arch = "x86_64"))")); + EXPECT_TRUE (cfgpred::uses_layer(R"(cfg(c-abi = "musl"))")); + // A predicate mixing the two belongs to the late pass, which can answer + // both. Ownership is by membership, not by which leg matched. + EXPECT_TRUE (cfgpred::uses_layer( + R"(cfg(all(accelerator = "cuda", c-abi = "musl")))")); } TEST(CfgAccelerator, MisspellingIsReportedNotSilentlyFalse) { diff --git a/tests/unit/test_cfg_accelerator_none.cpp b/tests/unit/test_cfg_accelerator_none.cpp new file mode 100644 index 00000000..5c1f6e28 --- /dev/null +++ b/tests/unit/test_cfg_accelerator_none.cpp @@ -0,0 +1,69 @@ +#include + +import std; +import mcpp.build.prepare_inputs; + +namespace cfgpred = mcpp::build::cfgpred; + +namespace { + +cfgpred::Ctx with(std::vector accelerators) { + auto c = cfgpred::context_for("x86_64-linux-gnu"); + c.layersKnown = true; + c.accelerators = std::move(accelerators); + return c; +} + +bool m(std::string_view predicate, const cfgpred::Ctx& c) { + return cfgpred::matches(std::string(predicate), c); +} + +} // namespace + +// `accelerator = "none"` is the empty set, and the reason it has to exist is +// that `accelerator`'s vocabulary is OPEN: docs/20 states a fifth backend is a +// package rather than an engine change. A fallback written by enumeration +// therefore changes meaning the day a fifth backend exists, silently. + +TEST(CfgAcceleratorNone, NoneIsTheEmptySet) { + EXPECT_TRUE (m(R"(cfg(accelerator = "none"))", with({}))); + EXPECT_FALSE(m(R"(cfg(accelerator = "none"))", with({"cuda"}))); + EXPECT_FALSE(m(R"(cfg(accelerator = "none"))", with({"cuda", "vulkan"}))); + + // The negation is the other half a seam needs: "some device backend". + EXPECT_FALSE(m(R"(cfg(not(accelerator = "none")))", with({}))); + EXPECT_TRUE (m(R"(cfg(not(accelerator = "none")))", with({"vulkan"}))); +} + +TEST(CfgAcceleratorNone, NoneDoesNotDisturbMembership) { + // `none` must not become a member of the set it describes the emptiness + // of, or `not(accelerator = "cuda")` would start answering for it. + EXPECT_TRUE (m(R"(cfg(accelerator = "cuda"))", with({"cuda"}))); + EXPECT_FALSE(m(R"(cfg(accelerator = "vulkan"))", with({"cuda"}))); + // A build that named a backend literally spelled "none" is not a case this + // engine has to serve; what matters is that the empty set stays the only + // thing `none` reports, which the first test pins. +} + +// THE REASON THIS IS NOT A COSMETIC CHANGE. +// +// The two spellings agree today and part company on the day a backend is +// added. That is the whole point, so the test simulates the addition rather +// than describing it: `enumerated` is what a project wrote against the +// two-backend world, `none` is what it should have written. +TEST(CfgAcceleratorNone, EnumerationRotsAndNoneDoesNot) { + constexpr auto enumerated = + R"(cfg(not(any(accelerator = "cuda", accelerator = "vulkan"))))"; + constexpr auto stable = R"(cfg(accelerator = "none"))"; + + // The world the fallback was written in: the two agree. + EXPECT_EQ(m(enumerated, with({})), m(stable, with({}))); + EXPECT_EQ(m(enumerated, with({"cuda"})), m(stable, with({"cuda"}))); + + // A fifth backend arrives. The enumeration now claims "no accelerator" for + // a build that named one -- the CPU fallback would compile beside the + // device implementation -- while `none` is unchanged. + auto fifth = with({"ascend"}); + EXPECT_TRUE (m(enumerated, fifth)) << "the enumeration is expected to rot"; + EXPECT_FALSE(m(stable, fifth)) << "`none` must not rot"; +}