Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
# Cumulative and Averaging Fission of Beliefs — Fusion Must Know Whether Evidence Is Independent

**Paper:** *Cumulative and Averaging Fission of Beliefs*, Audun Jøsang (UniK, University of Oslo)
**arXiv:** 0712.1182v1
**Source:** https://arxiv.org/abs/0712.1182
**Read grade:** READ-IN-FULL (2026-10-08, coresearch "CE64 as BF16 for thinking")

## Punchline

> **There are two fusions, not one. Cumulative fusion adds independent evidence; averaging fusion combines two readings of the same evidence. Using the cumulative rule on dependent evidence counts it twice. And cumulative fusion has an inverse — fission — that removes a known contribution.**

## What the paper establishes

- Binomial opinion `(b, d, u, a)`, `b + d + u = 1`; `b + d = 0` is vacuous (§2.1).
- **Cumulative fusion** (independent observers, disjoint periods): Case I `b = (bA·uB + bB·uA)/(uA + uB − uA·uB)`, `u = uA·uB/(uA + uB − uA·uB)`. It is commutative, associative and **non-idempotent** (Theorem 1). It equals a-posteriori updating of Dirichlet evidence: evidence counts add.
- **Dogmatic case** (`uA = uB = 0`, Case II): `b = γ·bA + (1 − γ)·bB`, `u = 0`, with `γ = lim uB/(uA + uB)`. Associativity in Case II requires carrying γ as extra state (Theorem 1 and text after it).
- **Averaging fusion** (dependent observers, same evidence): `b = (bA·uB + bB·uA)/(uA + uB)`, `u = 2uA·uB/(uA + uB)`. It is commutative and **idempotent**, not associative (Theorem 2).
- **Fission** is the inverse: given a fused opinion and one contributor, recover the other (Theorems 3, 4). Cumulative fission is non-commutative, non-associative.
- Dempster's rule is not used in subjective logic (§1).

## Mapping to lance-graph (Q-A: the CE64 revision operator)

NARS revision is cumulative fusion in evidence space: with `w = c/(1−c)` (horizon k = 1 in `isa::truth::revision`), pooled weights add (`ws = w1 + w2`). That makes it non-idempotent by construction — so **revising an edge with itself raises c**, which is the measured self-revision defect (`entries/2026-10-08-coresearch-ce64-moore-masking-wiring.md`, "no stamp check"). Jøsang names the missing distinction: the operator must know whether its two inputs are independent. NARS answers this with evidential-base (stamp) overlap; CE64's revision has no stamp, so it cannot tell cumulative from averaging situations.

| paper concept | CE64 / lance-graph | status |
|---|---|---|
| cumulative fusion = evidence counts add | `revision`: `ws = w1 + w2` | HAVE |
| averaging fusion (idempotent) for dependent inputs | none; self-revision double counts | MISSING — a stamp/evidential-base gate decides which rule applies |
| dogmatic Case II with relative weight γ | c=255 + c=255 computes NaN -> 0 | the paper's math confirms the case needs a declared answer; adopting SL as a second calculus was REJECTED earlier (`moore-masking-wiring.md:119-120`) |
| fission (remove a contribution) | in w-space, `w_A = w_C − w_B` | NEW reading: evidence-space revision is invertible, which gives detached/retractable revision (cf. `.arxiv/2608.16333`) |

## Harvest

1. Revision is cumulative fusion; it is only correct on independent evidence. The stamp gate (STAMP-GATE, open) is not an optimisation, it selects the fusion rule.
2. Working in evidence space makes revision a commutative monoid that is cancellative: a known contribution can be subtracted exactly (fission), which suits detached revision and replay. This holds only away from the dogmatic edge, where w is infinite.
3. Subjective logic stays rejected as a second calculus; only its case analysis (dogmatic limit, cumulative vs averaging) is harvested.

## Kill condition

The fission reading is dropped if, on the u8 grid, `revision(revision(a, b), fission-inverse of b)` does not return `a` within one code for c < 254 (quantisation destroys invertibility at 8 bits).
39 changes: 39 additions & 0 deletions .arxiv/1905.12322_narrow_storage_wide_accumulate_round_once.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
# BFLOAT16 — Narrow Storage, Wide Accumulate, Round Once

**Paper:** *A Study of BFLOAT16 for Deep Learning Training*, Kalamkar, Mudigere, Mellempudi, Das, et al. (Intel Labs, Facebook)
**arXiv:** 1905.12322v3
**Source:** https://arxiv.org/abs/1905.12322
**Read grade:** READ-IN-FULL (2026-10-08, coresearch "CE64 as BF16 for thinking")

## Punchline

> **A narrow format wins when it keeps the range, gives up precision, and is only ever an operand: every sum lives in a wider accumulator and is narrowed once, at the store.**

## What the paper establishes

- BF16 is FP32 with the low 16 mantissa bits dropped: same 8-bit exponent, 7-bit mantissa (Table 1). The claim to fame is the range, not the precision: "no hyper-parameter tuning is required", whereas FP16 needs loss scaling and INT16 needs block scaling (§1, §3).
- All three half-precision methods share one schema: 16-bit operands, 32-bit accumulators (§1). GEMMs accept BF16 inputs and accumulate into FP32 outputs (Figure 1).
- Conversion is round-to-nearest-even on the dropped bits (§1, Quantlib).
- Truncation instead of RNE costs a small but real amount: Deep & Cross log-loss 0.44372 (RNE) vs 0.44393 (truncation); the authors call 0.001 log-loss "unacceptable in practice" (§4.4, Table 5).
- Bias tensors stay FP32; the weight update uses an FP32 master copy (§3).

## Mapping to lance-graph (Q-A only: CE64 F/C, bits 24..39)

CE64 is a 64-bit tagged instruction word, not a number format (premise audit, entries/2026-10-08-coresearch-ce64-bf16-for-thinking.md). The BF16 analogy applies only to the two u8 truth scalars F and C and to the truth ALU in `causal_edge::isa::truth`.

| BF16 property | CE64 F/C counterpart | status |
|---|---|---|
| narrow operand, wide accumulate | `revision` decodes u8 -> f32, pools in f32, re-encodes once per call (`isa.rs:216-261`) | HAVE per call |
| round once, at the store | chains (`replay_step`) and `NarsTables` re-quantise per hop / into 16 c-bins | MISSING across chains |
| RNE at the narrowing | `(x*255).round()` is half-away-from-zero; exact .5 ties exist in revision's integer form (e.g. c1=c2=45 -> 76.5) | rounding mode undeclared |
| master copy in wide precision | none: the u8 code is the only state | WRONG-TO-COPY (the record is the canonical copy; a fold-side wide summary is the lawful analogue) |

## Harvest

1. Evidence is accumulated in the wide domain (`w = c/(1-c)`, or the integer form `c/(255-c)`), and narrowed to u8 C exactly once, at the store. Multi-step reasoning that re-encodes every hop is the "per-layer truncation" BF16 avoided.
2. For populations (D-RPF-4), the wide accumulator is a fold-side summary (`Σw`, `Σw·f`), never a second stored copy in the row.
3. Rounding mode is a declared property of the narrowing step, not an accident of `f32::round`.

## Kill condition

If a chain of N weak same-sign revisions re-encoded per hop never differs by more than one code from the wide-accumulate-round-once result, the "round once" rule is cosmetic for CE64 (probe P-X1 in the coresearch entry).
41 changes: 41 additions & 0 deletions .arxiv/2209.05433_special_values_are_declared_not_computed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
# FP8 Formats — Special Values Are Declared, Not Computed

**Paper:** *FP8 Formats for Deep Learning*, Micikevicius, Stosic, Judd, Kamalu, Oberman, Shoeybi, Siu, Wu (NVIDIA); Burgess, Ha, Grisenthwaite (Arm); Mellempudi, Cornea, Heinecke, Dubey (Intel)
**arXiv:** 2209.05433v2
**Source:** https://arxiv.org/abs/2209.05433
**Read grade:** READ-IN-FULL (2026-10-08, coresearch "CE64 as BF16 for thinking")

## Punchline

> **At 8 bits, every code is precious and the edge codes are policy. E4M3 drops infinity, keeps a single NaN pattern, and saturates overflow to the largest finite value — a format decision, written down, not an arithmetic accident.**

## What the paper establishes

- Two encodings: E4M3 (weights, activations) and E5M2 (gradients). E5M2 keeps IEEE special values; E4M3 represents no infinities and only one mantissa pattern for NaN (S.1111.111), gaining one binade (max 448 instead of 240) (§3.1, Table 1).
- Operations on FP8 inputs produce higher-precision outputs, optionally narrowed to FP8 before the store (§2).
- Overflowing values "are then saturated to the maximum representable value"; skipping updates on overflow is a poor fit at 8 bits because overflows are frequent (§2).
- Converting a wide Inf or NaN to E4M3 yields NaN; a non-saturating conversion mode can be offered for strict overflow handling (§2).
- "Rounding mode (round to nearest even, stochastic, etc.) choice is orthogonal to the interchange format and is left up to the implementation" (§2).
- Scaling lives in software, per tensor, not in a programmable exponent bias (§3.2).

## Mapping to lance-graph (Q-A: CE64 C, bits 32..39)

The CE64 confidence byte has its own edge-code problem. Code 255 decodes to exactly 1.0, and `evidence_weight(1.0)` is `f32::MAX`; revising two 255 operands computes `inf/inf = NaN`, which the saturating `as u8` cast stores as **0** — "certain plus certain becomes no evidence" (`isa.rs:216-261`, pinned non-normative in `tests/ce64_isa_golden.rs:253-273`).

The integer form of revision (with `w = c/(255-c)`) is `c_out·255 = 255·N/D`, `N = c1(255-c2) + c2(255-c1)`, `D = 255² - c1c2`. At `c1 = c2 = 255`, N = D = 0: the dogmatic case is **0/0 even in exact arithmetic**. So it is not a floating-point bug that better arithmetic would fix; like E4M3's top code, it needs a declared answer.

| FP8 move | CE64 C equivalent | fit |
|---|---|---|
| saturate overflow to max finite | arithmetic never writes 255 (saturate at 254), decode of 255 unchanged | fits, gated per I-LEGACY-API-FEATURE-GATED |
| one reserved NaN code | give 255 a new stored meaning ("dogmatic") | CONFLICTS-ANCHOR: re-reads frozen v2 bits |
| rounding mode left to the implementation | `isa::contracts` should declare it | fits (additive) |
| per-tensor software scale | truth has a fixed decode; no scale | WRONG-TO-COPY |

## Harvest

1. The c=255 cell is a policy cell. The smallest fit is the prior-art fix: cap c before `evidence_weight` (as the f64 reference already does), behind a version gate, with the golden flipped from "pins the defect" to "pins the policy".
2. NaN -> 0 through a saturating cast erases information silently. FP8's own documentation of special-value conversion is the precedent for declaring the mapping in the op's contract.

## Kill condition

If capping c at 254 before `evidence_weight` yields 0 NaN, 0 monotonicity violations, and agreement with the exact integer form on all 65,535 non-(255,255) (c1,c2) pairs, no special code is needed and option "reserved dogmatic code" is closed (probe P-X2).
36 changes: 36 additions & 0 deletions .arxiv/2310.10537_shared_scale_lives_outside_the_element.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
# Microscaling (MX) — The Shared Scale Lives Outside the Element, and the Axis Matters

**Paper:** *Microscaling Data Formats for Deep Learning*, Bita Darvish Rouhani, Ritchie Zhao, et al. (Microsoft, AMD, Intel, Meta, NVIDIA, Qualcomm)
**arXiv:** 2310.10537v3
**Source:** https://arxiv.org/abs/2310.10537
**Read grade:** READ-IN-FULL (2026-10-08, coresearch "CE64 as BF16 for thinking")

## Punchline

> **Give a block of k narrow elements one shared power-of-two scale, kept beside them rather than inside them. It works — but the scale is a lossy quantisation, and quantising along one axis does not commute with transposing to another.**

## What the paper establishes

- An MX block is k scalar elements plus one shared scale X; value `v_i = X·P_i` (§2, Figure 1). "The layout of an MX block is not prescribed — an implementation may store X contiguously with or separately from the elements."
- All concrete MX formats use k = 32 and an E8M0 scale (Table 1). The scale is `2^(floor(log2 max|V_i|) − emax_elem)`; elements that overflow are clamped to the element maximum; Inf and NaN are not clamped (Algorithm 1). A NaN scale makes the whole block NaN; the scale never encodes Inf (§2.1).
- Dot products produce outputs in a scalar float format (BF16/FP32); vector ops stay in scalar float; a master FP32 weight copy is kept (§4.1).
- "Conversion to MX format and transposing are not commutative operations", so the quantised weights and their transpose are stored as two tensors (§3, §4.1).
- Rounding: RNE for inference conversions, round-half-away-from-zero in the training runs (§4.3, §4.5).

## Mapping to lance-graph (Q-A: folds over F/C; D-RPF-4)

CE64 truth has a fixed decode (`u8/255`) with no dynamic range to scale, so a per-record scale has nothing to do — WRONG-TO-COPY inside the row, which the frozen v2 layout forbids anyway.

The one tempting place is a population fold (D-RPF-4: `Σw`, `Σw·f` over a masked population). Two findings from the paper argue against using MX there:

1. A block-scaled summary is lossy, and the lane-fold capstone rule is that a lossy decomposition is never a fold. An exact fold exists: the integer evidence weights `c/(255−c)` summed in i128 (or u32 for capped c — `254·256·65536 < 2³²`) need no scale.
2. The axis lesson transfers directly: a summary built along rows (a fold over a population) and one built along fields (a bit-sliced plane) are different quantisations. They cannot be swapped without recomputation, which is why a bit-sliced copy is a derived lane, never an equivalent representation.

## Harvest

- MX's "scale beside the element, not inside it" is the same move as the D-RPF rule "summaries are fold-side, the row stays as stored".
- For truth folds, prefer the exact integer evidence fold; a block-scaled summary may at most be reported as a documented approximation beside it (the status D-RPF-7 gives an SVD truncation).

## Kill condition

The block-scale idea is closed if a fixed global Q8 scale over capped evidence weights gives at most one code of error on every test population at 64K rows (probe P-X8). If it does not, the exact fold is still preferred; MX would only be a reporting format.
39 changes: 39 additions & 0 deletions .arxiv/2603.24161_stagnation_is_a_rounding_property.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
# Limited-Precision Stochastic Rounding — Stagnation Is a Rounding Property

**Paper:** *Probabilistic Error Analysis of Limited-Precision Stochastic Rounding: Horner's Algorithm and Pairwise Summation*, El-Mehdi El Arar, Massimiliano Fasi, Silviu-Ioan Filip, Mantas Mikaitis
**arXiv:** 2603.24161v1
**Source:** https://arxiv.org/abs/2603.24161
**Read grade:** READ-IN-FULL (2026-10-08, coresearch "CE64 as BF16 for thinking")

## Punchline

> **Round-to-nearest stagnates when small same-sign increments meet a large running total; stochastic rounding fixes it with a handful of random bits. But it buys accuracy in expectation, not per run — a different contract from replayable state.**

## What the paper establishes

- Stochastic rounding (SR) rounds up with probability proportional to the distance to the lower neighbour, so `E[SR(x)] = x` (Def. 1).
- Deterministic rounding has worst-case error bounds growing as O(n·u); SR gives probabilistic bounds O(√n·u) (§1).
- Limited-precision SR with r random bits is biased: `E[SR_{p,r}(x)] = fl_{p+r}(x)` (Def. 2, Eq. 4). Its bounds are ∝ √n·u_p + n·u_{p+r}; the rule of thumb is r ≈ ⌈log2(k)/2⌉ for error chains of length k (Remarks 2, 4).
- Stagnation under RN appears with same-sign data (coefficients in [0,1]); with mixed signs RN errors partly cancel and SR's advantage shrinks (§5.1). Pairwise summation largely avoids stagnation (§5.2).
- IEEE P3109 already specifies three limited-precision SR variants (Remark 1).

## Mapping to lance-graph (Q-A: CE64 C under repeated revision)

Repeated revision of an edge with small same-sign evidence is the same shape as summing small positive terms into a large total in a coarse format: a step whose Δw moves c by less than half a code rounds back to the same u8 C. Whether this happens on the real path is **unmeasured**.

The measured failure points the other way: chain confidence 200 -> 224 -> 237 saturates because `replay_step` revises on every hop and nothing stops double counting (`entries/2026-10-08-coresearch-ce64-moore-masking-wiring.md`). SR does nothing for over-accumulation.

| SR property | CE64 fit |
|---|---|
| unbiased in expectation | conflicts with replayable, deterministic state unless the random bits are hashed from (stamp, step) |
| changes `revision` outputs | I-LEGACY-API-FEATURE-GATED: feature gate or new entry point |
| error bound assumes mean independence | I-NOISE-FLOOR-JIRAK: CE64 bits are weakly dependent; bounds must cite weak-dependence rates |

## Harvest

1. Measure stagnation before choosing a remedy. The cheaper remedy is BF16's (accumulate wide, round once), which is deterministic.
2. If SR is ever adopted, the random bits must be a deterministic hash, the bit count follows ⌈log2(chain length)/2⌉, and it lives behind a feature gate.

## Kill condition

SR is dropped if a grid of same-sign weak-revision chains (start codes {0..254}, increment codes {1,4,16,64}, N up to 1000) shows zero steps where the per-hop RN code stalls while the wide accumulator advances (probe P-X1/X6).
Loading
Loading